Hướng dẫn lãnh đạo cấp cao: quyết định chiến lược, phát triển tổ chức, quản lý nhà đầu tư và gọi vốn.
---
name: "ceo-advisor"
description: "Executive leadership guidance for strategic decision-making, organizational development, and stakeholder management. Use when planning strategy, preparing board presentations, managing investors, developing organizational culture, making executive decisions, fundraising, or when user mentions CEO, strategic planning, board meetings, investor updates, organizational leadership, or executive strategy."
license: MIT
metadata:
version: 2.0.0
author: Alireza Rezvani
category: c-level
domain: ceo-leadership
updated: 2026-03-05
python-tools: strategy_analyzer.py, financial_scenario_analyzer.py
frameworks: executive-decisions, board-governance, leadership-culture
---
# CEO Advisor
Strategic leadership frameworks for vision, fundraising, board management, culture, and stakeholder alignment.
## Keywords
CEO, chief executive officer, strategy, strategic planning, fundraising, board management, investor relations, culture, organizational leadership, vision, mission, stakeholder management, capital allocation, crisis management, succession planning
## Quick Start
```bash
python scripts/strategy_analyzer.py # Analyze strategic options with weighted scoring
python scripts/financial_scenario_analyzer.py # Model financial scenarios (base/bull/bear)
```
## Core Responsibilities
### 1. Vision & Strategy
Set the direction. Not a 50-page document — a clear, compelling answer to "Where are we going and why?"
**Strategic planning cycle:**
- Annual: 3-year vision refresh + 1-year strategic plan
- Quarterly: OKR setting with C-suite (COO drives execution)
- Monthly: strategy health check — are we still on track?
**Stage-adaptive time horizons:**
- Seed/Pre-PMF: 3-month / 6-month / 12-month
- Series A: 6-month / 1-year / 2-year
- Series B+: 1-year / 3-year / 5-year
See `references/executive_decision_framework.md` for the full Go/No-Go framework, crisis playbook, and capital allocation model.
### 2. Capital & Resource Management
You're the chief allocator. Every dollar, every person, every hour of engineering time is a bet.
**Capital allocation priorities:**
1. Keep the lights on (operations, must-haves)
2. Protect the core (retention, quality, security)
3. Grow the core (expansion of what works)
4. Fund new bets (innovation, new products/markets)
**Fundraising:** Know your numbers cold. Timing matters more than valuation. See `references/board_governance_investor_relations.md`.
### 3. Stakeholder Leadership
You serve multiple masters. Priority order:
1. Customers (they pay the bills)
2. Team (they build the product)
3. Board/Investors (they fund the mission)
4. Partners (they extend your reach)
### 4. Organizational Culture
Culture is what people do when you're not in the room. It's your job to define it, model it, and enforce it.
See `references/leadership_organizational_culture.md` for culture development frameworks and the CEO learning agenda. Also see `culture-architect/` for the operational culture toolkit.
### 5. Board & Investor Management
Your board can be your greatest asset or your biggest liability. The difference is how you manage them.
See `references/board_governance_investor_relations.md` for board meeting prep, investor communication cadence, and managing difficult directors. Also see `board-deck-builder/` for assembling the actual board deck.
## Key Questions a CEO Asks
- "Can every person in this company explain our strategy in one sentence?"
- "What's the one thing that, if it goes wrong, kills us?"
- "Am I spending my time on the highest-leverage activity right now?"
- "What decision am I avoiding? Why?"
- "If we could only do one thing this quarter, what would it be?"
- "Do our investors and our team hear the same story from me?"
- "Who would replace me if I got hit by a bus tomorrow?"
## CEO Metrics Dashboard
| Category | Metric | Target | Frequency |
|----------|--------|--------|-----------|
| **Strategy** | Annual goals hit rate | > 70% | Quarterly |
| **Revenue** | ARR growth rate | Stage-dependent | Monthly |
| **Capital** | Months of runway | > 12 months | Monthly |
| **Capital** | Burn multiple | < 2x | Monthly |
| **Product** | NPS / PMF score | > 40 NPS | Quarterly |
| **People** | Regrettable attrition | < 10% | Monthly |
| **People** | Employee engagement | > 7/10 | Quarterly |
| **Board** | Board NPS (your relationship) | Positive trend | Quarterly |
| **Personal** | % time on strategic work | > 40% | Weekly |
## Red Flags
- You're the bottleneck for more than 3 decisions per week
- The board surprises you with questions you can't answer
- Your calendar is 80%+ meetings with no strategic blocks
- Key people are leaving and you didn't see it coming
- You're fundraising reactively (runway < 6 months, no plan)
- Your team can't articulate the strategy without you in the room
- You're avoiding a hard conversation (co-founder, investor, underperformer)
## Integration with C-Suite Roles
| When... | CEO works with... | To... |
|---------|-------------------|-------|
| Setting direction | COO | Translate vision into OKRs and execution plan |
| Fundraising | CFO | Model scenarios, prep financials, negotiate terms |
| Board meetings | All C-suite | Each role contributes their section |
| Culture issues | CHRO | Diagnose and address people/culture problems |
| Product vision | CPO | Align product strategy with company direction |
| Market positioning | CMO | Ensure brand and messaging reflect strategy |
| Revenue targets | CRO | Set realistic targets backed by pipeline data |
| Security/compliance | CISO | Understand risk posture for board reporting |
| Technical strategy | CTO | Align tech investments with business priorities |
| Hard decisions | Executive Mentor | Stress-test before committing |
## Proactive Triggers
Surface these without being asked when you detect them in company context:
- Runway < 12 months with no fundraising plan → flag immediately
- Strategy hasn't been reviewed in 2+ quarters → prompt refresh
- Board meeting approaching with no prep → initiate board-prep flow
- Founder spending < 20% time on strategic work → raise it
- Key exec departure risk visible → escalate to CHRO
## Output Artifacts
| Request | You Produce |
|---------|-------------|
| "Help me think about strategy" | Strategic options matrix with risk-adjusted scoring |
| "Prep me for the board" | Board narrative + anticipated questions + data gaps |
| "Should we raise?" | Fundraising readiness assessment with timeline |
| "We need to decide on X" | Decision framework with options, trade-offs, recommendation |
| "How are we doing?" | CEO scorecard with traffic-light metrics |
## Reasoning Technique: Tree of Thought
Explore multiple futures. For every strategic decision, generate at least 3 paths. Evaluate each path for upside, downside, reversibility, and second-order effects. Pick the path with the best risk-adjusted outcome.
**Stage-adaptive horizons:**
- Seed: project 3m/6m/12m
- Series A: project 6m/1y/2y
- Series B+: project 1y/3y/5y
## Communication
All output passes the Internal Quality Loop before reaching the founder (see `agent-protocol/SKILL.md`).
- Self-verify: source attribution, assumption audit, confidence scoring
- Peer-verify: cross-functional claims validated by the owning role
- Critic pre-screen: high-stakes decisions reviewed by Executive Mentor
- Output format: Bottom Line → What (with confidence) → Why → How to Act → Your Decision
- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed.
## Context Integration
- **Always** read `company-context.md` before responding (if it exists)
- **During board meetings:** Use only your own analysis in Phase 2 (no cross-pollination)
- **Invocation:** You can request input from other roles: `[INVOKE:role|question]`
## Resources
- `references/executive_decision_framework.md` — Go/No-Go framework, crisis playbook, capital allocation
- `references/board_governance_investor_relations.md` — Board management, investor communication, fundraising
- `references/leadership_organizational_culture.md` — Culture development, CEO routines, succession planning
FILE:references/board_governance_investor_relations.md
# Board Governance & Investor Relations Guide
## Board of Directors Management
### Board Composition
#### Ideal Board Structure
- **Size**: 7-9 members (odd number for voting)
- **Independence**: Majority independent directors
- **Diversity**: Gender, ethnicity, expertise, experience
- **Term**: 3-year terms, staggered renewal
#### Board Roles
| Role | Responsibilities | Typical Background |
|------|-----------------|-------------------|
| Chairman | Board leadership, CEO liaison | Former CEO, Industry veteran |
| Lead Independent Director | Independent voice, executive sessions | Senior executive experience |
| Audit Committee Chair | Financial oversight, auditor relationship | CFO/CPA background |
| Compensation Committee Chair | Executive compensation, succession | HR/Executive experience |
| Nominating Committee Chair | Board composition, governance | Governance expertise |
### Board Meeting Management
#### Annual Board Calendar
**Q1 Meeting**
- Annual strategy review
- Previous year performance
- Current year priorities
- Risk assessment update
**Q2 Meeting**
- Q1 results review
- Strategic initiative progress
- Competitive landscape
- Talent review
**Q3 Meeting**
- Mid-year performance
- Budget preview
- Strategic planning session
- Succession planning
**Q4 Meeting**
- Annual budget approval
- Executive compensation
- Board evaluation
- Upcoming year calendar
#### Meeting Preparation Timeline
**T-4 Weeks**
- Agenda draft to Chairman
- Pre-read preparation begins
- Committee meetings scheduled
**T-2 Weeks**
- Materials to review committee
- Final agenda confirmation
- Logistics coordination
**T-1 Week**
- Board package distribution
- Pre-meeting calls as needed
- Final preparations
**T-0 Meeting Day**
- Executive session (start)
- Board meeting
- Executive session (end)
- Follow-up actions defined
### Board Package Template
#### Standard Package Contents
1. **Cover Memo** (1 page)
- Meeting agenda
- Key decisions required
- Time allocations
2. **CEO Report** (3-5 pages)
- Executive summary
- Performance highlights
- Strategic progress
- Key challenges
- Asks of the board
3. **Financial Report** (5-10 pages)
- Financial statements
- KPI dashboard
- Variance analysis
- Cash position
- Forecast update
4. **Strategic Updates** (10-15 pages)
- Initiative status
- Market analysis
- Competitive intelligence
- Product roadmap
5. **Committee Reports** (2-3 pages each)
- Audit Committee
- Compensation Committee
- Other committees
6. **Appendices**
- Detailed financials
- Supporting analysis
- Previous minutes
### Board Communication Best Practices
#### Between Meetings
**Monthly Update Email**
```
Subject: [Company] CEO Update - [Month Year]
Board Members,
Quick update on [Month] performance:
Headlines:
• [Key achievement]
• [Important metric]
• [Strategic progress]
Challenges:
• [Issue and mitigation]
Looking Ahead:
• [Upcoming milestone]
Detailed dashboard attached.
Best,
[CEO Name]
```
**Flash Reports** (When needed)
- Material events
- Major wins/losses
- Press coverage
- Regulatory matters
#### Managing Difficult Conversations
**Delivering Bad News**
1. Don't delay - inform promptly
2. Lead with facts
3. Own the responsibility
4. Present action plan
5. Set realistic timeline
**Handling Dissent**
1. Listen fully
2. Acknowledge concerns
3. Provide data/rationale
4. Seek common ground
5. Document decisions
## Investor Relations
### Investor Segmentation
#### Institutional Investors
**Types**:
- Mutual funds
- Pension funds
- Hedge funds
- Private equity
- Sovereign wealth funds
**Engagement Strategy**:
- Quarterly earnings calls
- Annual investor day
- Conference participation
- One-on-one meetings
- Site visits
#### Retail Investors
**Channels**:
- Website IR section
- Annual reports
- Proxy statements
- Social media
- Shareholder meetings
### Earnings Communications
#### Earnings Release Template
```
[COMPANY] REPORTS [QUARTER] [YEAR] RESULTS
[City, Date] - [Company] (TICKER) today reported results for [quarter]:
Financial Highlights:
• Revenue: $X (±Y% YoY)
• Net Income: $X (±Y% YoY)
• EPS: $X (±Y% YoY)
• [Other key metric]
CEO Commentary:
"[Quote about performance and outlook]"
CFO Commentary:
"[Quote about financial details]"
Guidance:
[Forward-looking statements]
Conference Call:
Date/Time: [Details]
Webcast: [Link]
About [Company]:
[Boilerplate]
Contact:
[IR contact information]
```
#### Earnings Call Script Structure
**CEO Opening (5 minutes)**
```
Good [morning/afternoon], and welcome to [Company's]
[Quarter] earnings call.
Today I'll cover:
1. Quarter highlights
2. Strategic progress
3. Market dynamics
4. Outlook
[Key points with supporting data]
I'll now turn it over to our CFO...
```
**CFO Section (10 minutes)**
```
Thank you [CEO name].
Financial Performance:
- Revenue details by segment
- Margin analysis
- Cash flow review
- Balance sheet highlights
Guidance:
- Next quarter expectations
- Full year outlook
- Key assumptions
Now back to [CEO] for closing remarks...
```
**Q&A Management**
- Anticipate top 10 questions
- Prepare fact sheets
- Designate responders
- Bridge to key messages
- Time management
### Investor Messaging Framework
#### Value Proposition
**Investment Thesis Elements**:
1. Market opportunity size
2. Competitive advantages
3. Growth strategy
4. Financial model
5. Management team
6. Risk factors
#### Key Messages Architecture
**Primary Messages** (Memorize)
1. [Core value proposition]
2. [Differentiation]
3. [Growth trajectory]
**Supporting Points** (Have ready)
- Market data
- Customer proof points
- Financial metrics
- Strategic initiatives
**Proof Points** (Document)
- Case studies
- Metrics
- Third-party validation
- Awards/recognition
### Investor Day Planning
#### 6-Month Planning Timeline
**T-6 Months**
- Set date and venue
- Define objectives
- Identify speakers
- Begin content development
**T-4 Months**
- Develop presentations
- Coordinate logistics
- Begin rehearsals
- Create save-the-date
**T-2 Months**
- Finalize content
- Complete rehearsals
- Send invitations
- Prepare materials
**T-1 Month**
- Final preparations
- Media training
- Q&A preparation
- Technology testing
**T-0 Event Day**
- Execute program
- Manage Q&A
- Network sessions
- Follow-up plan
#### Agenda Template
```
8:00 AM - Registration & Breakfast
8:30 AM - CEO Welcome & Vision
9:00 AM - Market Opportunity
9:30 AM - Product Strategy & Demo
10:00 AM - Break
10:15 AM - Go-to-Market Strategy
10:45 AM - Financial Overview
11:15 AM - Q&A Panel
12:00 PM - Networking Lunch
1:00 PM - Facility Tour (Optional)
```
### Shareholder Activism Defense
#### Early Warning Signs
- Stake building (13D/13G filings)
- Public criticism
- Media campaigns
- Proxy solicitation
- Shareholder proposals
#### Response Playbook
**1. Preparation Phase**
- Vulnerability assessment
- Response team formation
- Advisor engagement
- Board alignment
**2. Engagement Phase**
- Direct dialogue
- Understanding demands
- Finding common ground
- Negotiation strategy
**3. Defense Phase** (if needed)
- Public response
- Proxy fight preparation
- Shareholder outreach
- Media strategy
**4. Resolution Phase**
- Settlement negotiations
- Implementation planning
- Communication strategy
- Monitoring plan
### Regulatory Compliance
#### Key Filings
| Form | Purpose | Timing |
|------|---------|--------|
| 10-K | Annual report | 60-90 days after FY end |
| 10-Q | Quarterly report | 40-45 days after Q end |
| 8-K | Material events | 4 business days |
| DEF 14A | Proxy statement | Before annual meeting |
| S-1/S-3 | Securities registration | As needed |
#### Disclosure Requirements
**Material Information**:
- Financial results
- Major transactions
- Leadership changes
- Strategic shifts
- Legal proceedings
- Risk changes
**Regulation FD Compliance**:
- No selective disclosure
- Simultaneous public release
- Documented procedures
- Training program
### Crisis Communication
#### IR Crisis Response
**Hour 1: Assessment**
- Gather facts
- Assess materiality
- Consult legal
- Prepare holding statement
**Hours 2-4: Response**
- Draft 8-K if required
- Prepare FAQ
- Update website
- Notify exchanges
**Hours 4-8: Communication**
- Issue press release
- Update analysts
- Employee communication
- Monitor reactions
**Day 2+: Follow-up**
- Investor calls
- Media interviews
- Ongoing updates
- Impact assessment
### Performance Metrics
#### IR Effectiveness KPIs
**Quantitative Metrics**:
- Share price performance vs peers
- Trading volume/liquidity
- Analyst coverage
- Institutional ownership %
- Valuation multiples vs peers
**Qualitative Metrics**:
- Analyst sentiment
- Media coverage tone
- Investor feedback
- Award recognition
- Perception studies
#### Shareholder Analysis
**Ownership Tracking**:
- Top 20 shareholders
- Ownership changes
- Peer ownership overlap
- Geographic distribution
- Investment style mix
**Engagement Metrics**:
- Meeting count
- Conference participation
- Earnings call attendance
- Website analytics
- Email engagement
## Governance Best Practices
### Board Effectiveness
#### Annual Board Evaluation
**Process**:
1. Anonymous surveys
2. Individual interviews
3. Peer feedback
4. Results compilation
5. Action planning
6. Progress monitoring
**Evaluation Areas**:
- Board composition
- Meeting effectiveness
- Information quality
- Strategic oversight
- Risk management
- CEO relationship
- Committee performance
### Executive Session Management
**Frequency**: Every board meeting
**Duration**: 30-60 minutes
**Participants**: Independent directors only
**Typical Topics**:
- CEO performance
- Succession planning
- Board dynamics
- Sensitive matters
- Executive compensation
### D&O Insurance & Indemnification
**Coverage Levels**:
- Primary: $10-25M
- Excess: $25-100M+
- Side A: Individual protection
- Side B: Company reimbursement
- Side C: Securities claims
**Best Practices**:
- Annual review
- Competitive benchmarking
- Claims history analysis
- Policy optimization
- Personal coverage consideration
### ESG Governance
#### ESG Integration
**Board Oversight**:
- ESG committee or full board
- Regular ESG updates
- Metrics in dashboard
- Risk assessment
- Stakeholder feedback
**Reporting Framework**:
- SASB standards
- TCFD recommendations
- GRI guidelines
- UN SDGs alignment
- Integrated reporting
**Investor Communication**:
- ESG highlights in earnings
- Dedicated ESG report
- Website ESG section
- ESG investor days
- Rating agency engagement
## Templates & Tools
### Board Resolution Template
```
BOARD RESOLUTION
WHEREAS, [background/context];
WHEREAS, [additional context];
NOW, THEREFORE, BE IT RESOLVED, that [specific action];
FURTHER RESOLVED, that [additional actions];
FURTHER RESOLVED, that [authorization].
Approved this [date].
_____________________
[Secretary Name]
Corporate Secretary
```
### Insider Trading Policy Outline
1. **Scope**: All directors, officers, employees
2. **Prohibited Activities**: Trading on MNPI
3. **Trading Windows**: Quarterly schedule
4. **Pre-clearance**: Required for all trades
5. **Blackout Periods**: Defined schedule
6. **10b5-1 Plans**: Permitted with approval
7. **Violations**: Disciplinary action
8. **Training**: Annual requirement
### Proxy Statement Checklist
- [ ] Executive compensation (CD&A)
- [ ] Director nominees
- [ ] Governance structure
- [ ] Shareholder proposals
- [ ] Audit matters
- [ ] Related party transactions
- [ ] Risk oversight
- [ ] Succession planning
- [ ] ESG disclosure
- [ ] Virtual meeting details
FILE:references/executive_decision_framework.md
# Executive Decision Framework
## Decision-Making Process
### The DECIDE Framework
**D** - Define the problem clearly
**E** - Establish criteria for solutions
**C** - Consider alternatives
**I** - Identify best alternatives
**D** - Develop and implement action plan
**E** - Evaluate and monitor solution
## Strategic Decision Categories
### 1. Growth Decisions
#### Market Expansion
**Evaluation Criteria**:
- Market size and growth rate
- Competitive landscape
- Regulatory environment
- Cultural fit
- Required investment
- Expected ROI
**Decision Matrix**:
| Factor | Weight | Score (1-10) | Weighted Score |
|--------|--------|--------------|----------------|
| Market Size | 25% | | |
| Competition | 20% | | |
| Fit with Core | 20% | | |
| Investment Required | 15% | | |
| Risk Level | 10% | | |
| Timeline to Profit | 10% | | |
#### Product Development
**Go/No-Go Criteria**:
- Customer demand validation (>70% interest)
- Technical feasibility confirmed
- Positive unit economics
- Strategic alignment
- Available resources
#### Mergers & Acquisitions
**Due Diligence Framework**:
1. **Strategic Fit**
- Synergies identification
- Cultural alignment
- Market position enhancement
2. **Financial Analysis**
- Valuation models (DCF, Multiples, Precedent)
- ROI projections
- Integration costs
3. **Risk Assessment**
- Legal/regulatory issues
- Technology compatibility
- Talent retention
4. **Integration Planning**
- 100-day plan
- Communication strategy
- Success metrics
### 2. Resource Allocation
#### Capital Allocation Framework
**Priority Levels**:
1. **Essential** - Core operations, compliance, security
2. **Strategic** - Growth initiatives, competitive advantage
3. **Efficiency** - Cost reduction, productivity
4. **Experimental** - Innovation, R&D
**Allocation Guidelines**:
- Essential: 40-50%
- Strategic: 30-40%
- Efficiency: 10-15%
- Experimental: 5-10%
#### Budget Decision Tree
```
Is it required for operations?
├─ Yes → Essential (Auto-approve if <$X)
└─ No → Does it drive growth?
├─ Yes → What's the ROI?
│ ├─ >30% → Strategic (Approve)
│ └─ <30% → Defer/Reject
└─ No → Does it reduce costs?
├─ Yes → Payback period?
│ ├─ <12 months → Efficiency (Approve)
│ └─ >12 months → Defer
└─ No → Experimental (Limited budget)
```
### 3. Organizational Decisions
#### Restructuring Framework
**Triggers for Restructuring**:
- Performance below targets for 2+ quarters
- Major strategic shift
- M&A integration
- Market disruption
- Efficiency opportunity >20%
**Evaluation Process**:
1. Current state assessment
2. Future state design
3. Gap analysis
4. Impact assessment
5. Implementation planning
6. Communication strategy
#### Leadership Changes
**Performance Evaluation Matrix**:
| Dimension | Weight | Indicators |
|-----------|--------|------------|
| Results Delivery | 40% | KPIs, OKRs achievement |
| Team Leadership | 25% | Engagement, retention, development |
| Strategic Thinking | 20% | Innovation, vision, planning |
| Culture Fit | 15% | Values alignment, collaboration |
**Succession Planning**:
- Identify 2-3 potential successors for each key role
- Development plans for high-potentials
- Emergency succession protocols
- Knowledge transfer processes
### 4. Crisis Management
#### Crisis Response Protocol
**Immediate (0-2 hours)**:
1. Activate crisis team
2. Assess severity and impact
3. Implement containment measures
4. Initial stakeholder notification
**Short-term (2-24 hours)**:
1. Develop response strategy
2. Prepare public statements
3. Engage legal/regulatory as needed
4. Employee communication
**Recovery (24+ hours)**:
1. Implement solution
2. Monitor progress
3. Stakeholder updates
4. Post-crisis review
#### Crisis Decision Authority
| Crisis Level | Decision Authority | Response Team |
|--------------|-------------------|---------------|
| Level 1 (Minor) | Department Head | Local team |
| Level 2 (Moderate) | C-Suite Member | Cross-functional |
| Level 3 (Major) | CEO | Executive team |
| Level 4 (Critical) | CEO + Board | All hands |
## Decision Support Tools
### 1. SWOT-TOWS Matrix
```
Internal →
↓ Strengths (S) Weaknesses (W)
External
O SO Strategies WO Strategies
p (Leverage) (Improve)
p
o
r
t
T ST Strategies WT Strategies
h (Protect) (Survive)
r
e
a
t
s
```
### 2. BCG Growth-Share Matrix
```
Market Growth Rate
↑
High │ Stars │ Question │
│ │ Marks │
├─────────┼──────────┤
Low │ Cash │ Dogs │
│ Cows │ │
└─────────┴──────────┘
High Low →
Market Share
```
### 3. Risk-Impact Matrix
```
Impact
↑
High │ Mitigate │ Critical │
│ │ Focus │
├──────────┼──────────┤
Low │ Accept │ Monitor │
│ │ │
└──────────┴──────────┘
Low High →
Probability
```
### 4. Eisenhower Matrix
```
Urgency
↑
High │ Do │ Schedule │
│ First │ │
├─────────┼──────────┤
Low │ Delegate│ Eliminate│
│ │ │
└─────────┴──────────┘
High Low →
Importance
```
## Strategic Options Framework
### Porter's Generic Strategies
1. **Cost Leadership**
- Operational excellence
- Economy of scale
- Process optimization
- Supply chain efficiency
2. **Differentiation**
- Unique value proposition
- Premium positioning
- Innovation focus
- Brand strength
3. **Focus**
- Niche markets
- Specialized offerings
- Deep expertise
- Customer intimacy
### Blue Ocean Strategy
**Four Actions Framework**:
- **Eliminate**: Which factors can be eliminated?
- **Reduce**: Which factors should be reduced below industry standard?
- **Raise**: Which factors should be raised above industry standard?
- **Create**: Which factors should be created that the industry has never offered?
## Stakeholder Management
### Stakeholder Mapping
```
Influence/Power
↑
High │ Manage │ Key │
│ Closely │ Players │
├──────────┼──────────┤
Low │ Monitor │ Keep │
│ │ Informed │
└──────────┴──────────┘
Low High →
Interest
```
### Communication Strategy
| Stakeholder | Frequency | Format | Key Messages |
|------------|-----------|--------|--------------|
| Board | Monthly | Report + Meeting | Strategy, Risk, Performance |
| Investors | Quarterly | Earnings Call | Financial, Growth, Outlook |
| Employees | Weekly | All-hands | Vision, Updates, Recognition |
| Customers | Continuous | Multi-channel | Value, Innovation, Support |
| Media | As needed | Press Release | Milestones, Position, Vision |
## Performance Metrics
### Balanced Scorecard
#### Financial Perspective
- Revenue growth rate
- EBITDA margin
- ROE/ROA
- Cash conversion cycle
- Market capitalization
#### Customer Perspective
- Customer satisfaction (NPS)
- Market share
- Customer retention rate
- Customer acquisition cost
- Customer lifetime value
#### Internal Process
- Operational efficiency
- Time to market
- Quality metrics
- Innovation rate
- Process cycle time
#### Learning & Growth
- Employee engagement
- Talent retention
- Training hours per employee
- Leadership pipeline
- Innovation index
## Decision Biases to Avoid
### Cognitive Biases
1. **Confirmation Bias**
- Mitigation: Seek contrarian views
- Tool: Devil's advocate process
2. **Anchoring Bias**
- Mitigation: Multiple estimates
- Tool: Range forecasting
3. **Sunk Cost Fallacy**
- Mitigation: Zero-based thinking
- Tool: Regular portfolio review
4. **Overconfidence Bias**
- Mitigation: Outside view
- Tool: Reference class forecasting
5. **Availability Heuristic**
- Mitigation: Data-driven decisions
- Tool: Systematic analysis
### Decision Hygiene Checklist
- [ ] Problem clearly defined
- [ ] All stakeholders identified
- [ ] Data/evidence gathered
- [ ] Multiple options generated
- [ ] Biases checked
- [ ] Risks assessed
- [ ] Implementation plan created
- [ ] Success metrics defined
- [ ] Review process established
## Executive Communication
### Board Presentation Template
1. **Executive Summary** (1 slide)
- Key achievements
- Critical issues
- Decisions needed
2. **Performance Review** (3-4 slides)
- Financial results
- Operational metrics
- Strategic progress
3. **Market & Competition** (2 slides)
- Market dynamics
- Competitive position
4. **Strategic Initiatives** (3-4 slides)
- Current initiatives
- Results to date
- Next steps
5. **Risk & Mitigation** (2 slides)
- Risk register
- Mitigation actions
6. **Ask of the Board** (1 slide)
- Decisions required
- Support needed
### Investor Relations Framework
**Earnings Call Structure**:
1. Opening remarks (CEO) - 5 min
2. Financial review (CFO) - 10 min
3. Strategic update (CEO) - 10 min
4. Q&A - 30 min
**Key Messages**:
- Performance vs guidance
- Market position
- Growth strategy
- Capital allocation
- Outlook
## Strategic Planning Cycle
### Annual Planning Process
**Q3 - Strategic Review**
- Environmental scan
- Competitive analysis
- Capability assessment
- Strategy refinement
**Q4 - Planning**
- Goal setting
- Budget allocation
- Resource planning
- OKR development
**Q1 - Launch**
- Communication cascade
- Initiative kickoff
- Quick wins
- Baseline metrics
**Q2 - Review**
- Progress assessment
- Course correction
- Mid-year planning
- Performance review
## Exit Strategy Planning
### Exit Options Evaluation
1. **IPO**
- Pros: Maximum valuation, maintain control
- Cons: Regulatory burden, public scrutiny
- Timeline: 12-24 months
2. **Strategic Acquisition**
- Pros: Synergies, quick process
- Cons: Loss of independence, integration risk
- Timeline: 6-12 months
3. **Private Equity**
- Pros: Growth capital, expertise
- Cons: Pressure for returns, loss of control
- Timeline: 3-6 months
4. **Management Buyout**
- Pros: Continuity, culture preservation
- Cons: Limited price, financing challenge
- Timeline: 6-9 months
### Value Creation Levers
1. **Revenue Growth**
- Organic expansion
- Market development
- Product innovation
- Pricing optimization
2. **Margin Improvement**
- Operational efficiency
- Cost reduction
- Mix optimization
- Pricing power
3. **Multiple Expansion**
- Market positioning
- Growth trajectory
- Risk reduction
- Story telling
FILE:references/leadership_organizational_culture.md
# Leadership & Organizational Culture Guide
## Leadership Philosophy
### The Five Dimensions of CEO Leadership
1. **Visionary Leadership**
- Define compelling future state
- Communicate vision consistently
- Inspire action toward vision
- Measure progress systematically
2. **Strategic Leadership**
- Set clear priorities
- Allocate resources optimally
- Make tough trade-offs
- Drive execution excellence
3. **Operational Leadership**
- Establish performance standards
- Build scalable systems
- Drive continuous improvement
- Ensure accountability
4. **People Leadership**
- Attract top talent
- Develop future leaders
- Foster engagement
- Build inclusive culture
5. **External Leadership**
- Represent company publicly
- Build strategic partnerships
- Engage stakeholders effectively
- Shape industry direction
## Organizational Culture Framework
### Culture Definition & Assessment
#### Cultural Dimensions Model
**Innovation ← → Stability**
- Risk tolerance level
- Change readiness
- Experimentation mindset
- Learning from failure
**Competition ← → Collaboration**
- Internal dynamics
- Knowledge sharing
- Team vs individual rewards
- Cross-functional cooperation
**Customer ← → Operations**
- External vs internal focus
- Customer centricity
- Process emphasis
- Quality standards
**Short-term ← → Long-term**
- Planning horizons
- Investment philosophy
- Performance metrics
- Stakeholder balance
### Culture Transformation Roadmap
#### Phase 1: Assessment (Months 1-2)
**Current State Analysis**:
- Employee survey (engagement, values alignment)
- Culture assessment (competing values framework)
- Leadership 360 feedback
- Exit interview analysis
- Customer feedback integration
**Gap Analysis**:
- Current vs desired culture
- Behavioral gaps
- System misalignments
- Leadership gaps
- Communication gaps
#### Phase 2: Design (Months 2-3)
**Target Culture Definition**:
- Core values articulation
- Behavioral standards
- Leadership principles
- Decision principles
- Performance expectations
**Change Strategy**:
- Stakeholder mapping
- Communication plan
- Training requirements
- System changes needed
- Quick wins identification
#### Phase 3: Implementation (Months 4-12)
**Launch Activities**:
- Leadership alignment sessions
- All-hands kickoff
- Values workshops
- Behavioral training
- System updates
**Reinforcement Mechanisms**:
- Recognition programs
- Performance integration
- Hiring/promotion criteria
- Story collection
- Celebration events
#### Phase 4: Embedding (Months 12+)
**Sustainability Actions**:
- Regular pulse surveys
- Culture champions network
- Continuous reinforcement
- System alignment
- Leadership modeling
## Leadership Development
### Executive Team Development
#### Team Effectiveness Model
**Foundation Elements**:
1. **Trust** - Vulnerability-based trust
2. **Conflict** - Healthy debate
3. **Commitment** - Buy-in to decisions
4. **Accountability** - Peer accountability
5. **Results** - Collective outcomes
#### Executive Team Charter
```
Our Executive Team Charter
Purpose:
Lead [Company] to achieve its vision of [Vision Statement]
Responsibilities:
• Set strategic direction
• Allocate resources
• Drive performance
• Develop talent
• Shape culture
Operating Principles:
• Debate in private, unite in public
• Challenge ideas, support people
• Company first, function second
• Transparency with trust
• Accountability without blame
Meeting Cadence:
• Weekly tactical (2 hours)
• Monthly strategic (4 hours)
• Quarterly offsite (2 days)
• Annual planning (3 days)
Decision Rights:
• CEO: Final decision after consultation
• Consensus: Strategic initiatives
• Individual: Functional operations
• Escalation: Board-level matters
Success Metrics:
• Company performance vs plan
• Employee engagement score
• Customer satisfaction (NPS)
• Team effectiveness rating
```
### Succession Planning
#### Succession Planning Framework
**CEO Succession Timeline**:
**Ongoing**:
- Identify potential successors
- Development plan execution
- Board exposure
- External benchmarking
**T-3 Years**:
- Formal succession planning
- Candidate assessment
- Development acceleration
- Emergency plan update
**T-1 Year**:
- Final candidate selection
- Transition planning
- Communication strategy
- Onboarding preparation
**Transition**:
- Announcement
- Knowledge transfer
- Stakeholder introductions
- Gradual handover
#### Talent Pipeline Development
**9-Box Grid for Talent Review**:
```
Performance →
↑
│ Rising │ High │ Star
High│ Star │Performer│ Performer
├─────────┼─────────┼──────────
│Solid │ Core │ High
Med │Performer│Performer│ Potential
├─────────┼─────────┼──────────
│ Under │Inconsist│ New/
Low │Performer│ -ent │ Learning
└─────────┴─────────┴──────────
Low Medium High
Potential →
```
**Development Strategies by Box**:
- **Stars**: Accelerated development, stretch assignments
- **High Performers**: Retention focus, leadership opportunities
- **High Potentials**: Intensive coaching, skill building
- **Core Performers**: Engagement, incremental growth
- **Underperformers**: Performance improvement or exit
### Leadership Competency Model
#### Core Leadership Competencies
**Strategic Thinking**
- Vision development
- Systems thinking
- Innovation mindset
- External awareness
- Long-term planning
**Execution Excellence**
- Results orientation
- Decision quality
- Problem solving
- Process management
- Risk management
**People Leadership**
- Team building
- Talent development
- Communication
- Influence
- Emotional intelligence
**Personal Excellence**
- Integrity
- Resilience
- Continuous learning
- Self-awareness
- Adaptability
## Communication & Engagement
### Internal Communication Strategy
#### Communication Channels
| Channel | Frequency | Purpose | Audience |
|---------|-----------|---------|----------|
| All-hands meeting | Monthly | Updates, Q&A | All employees |
| Leadership cascade | Weekly | Alignment | Managers |
| CEO email | Bi-weekly | Vision, recognition | All employees |
| Town halls | Quarterly | Deep dives | All employees |
| Skip-levels | Monthly | Direct feedback | Various levels |
| Intranet | Daily | News, resources | All employees |
| Slack/Teams | Real-time | Collaboration | All employees |
#### CEO Communication Calendar
**Weekly**:
- Executive team meeting
- Leadership message cascade
- Customer/partner touchpoint
**Bi-weekly**:
- Company-wide email
- Skip-level meetings
- Media/analyst interaction
**Monthly**:
- All-hands meeting
- Board member touchpoint
- Employee roundtable
**Quarterly**:
- Earnings communication
- Town hall deep-dive
- Strategy review
- Culture celebration
### Employee Engagement
#### Engagement Survey Framework
**Dimensions Measured**:
1. Purpose & Vision (alignment, inspiration)
2. Leadership (trust, communication)
3. Management (support, development)
4. Work Environment (tools, processes)
5. Growth (career, learning)
6. Recognition (appreciation, fairness)
7. Wellbeing (balance, benefits)
8. Belonging (inclusion, connection)
**Action Planning Process**:
1. Share results transparently
2. Identify 2-3 focus areas
3. Create action teams
4. Define success metrics
5. Implement changes
6. Communicate progress
7. Measure impact
#### Engagement Initiatives
**Recognition Programs**:
- Spot awards (peer-nominated)
- Quarterly achievements
- Annual excellence awards
- Values champions
- Innovation celebrations
- Customer hero awards
**Development Programs**:
- Leadership academy
- Mentorship program
- Rotation opportunities
- Tuition reimbursement
- Conference attendance
- Skill workshops
**Wellbeing Initiatives**:
- Flexible work arrangements
- Mental health support
- Wellness programs
- Time-off policies
- Family support
- Financial wellness
## Performance Management
### OKR Framework
#### OKR Setting Process
**Company OKRs** (Annual)
↓
**Department OKRs** (Quarterly)
↓
**Team OKRs** (Quarterly)
↓
**Individual OKRs** (Quarterly)
#### OKR Template
**Objective**: [Qualitative, inspirational goal]
**Key Results**:
1. [Quantitative outcome] from [X] to [Y]
2. [Quantitative outcome] from [X] to [Y]
3. [Quantitative outcome] from [X] to [Y]
**Example**:
```
Objective: Become the market leader in customer satisfaction
Key Results:
1. Increase NPS from 45 to 70
2. Reduce support ticket resolution from 48h to 24h
3. Achieve 95% customer retention rate (from 87%)
```
### Performance Review System
#### Continuous Performance Management
**Weekly**: 1-on-1 check-ins (30 min)
- Progress on priorities
- Obstacles/support needed
- Feedback exchange
- Next week focus
**Monthly**: Development discussion (60 min)
- Skill development
- Career aspirations
- Stretch opportunities
- Learning plan
**Quarterly**: Performance review (90 min)
- OKR assessment
- Competency evaluation
- 360 feedback review
- Development planning
**Annual**: Compensation review
- Performance rating
- Compensation adjustment
- Promotion decisions
- Succession planning
## Change Management
### Change Leadership Model
#### Eight-Step Change Process
1. **Create Urgency**
- Share compelling data
- Highlight risks of status quo
- Create dissatisfaction with current state
2. **Build Coalition**
- Identify change champions
- Ensure executive alignment
- Engage influential supporters
3. **Form Vision**
- Define clear end state
- Create inspiring narrative
- Develop strategy
4. **Communicate Vision**
- Multi-channel communication
- Repetition and consistency
- Two-way dialogue
5. **Empower Action**
- Remove barriers
- Change systems/processes
- Encourage risk-taking
6. **Create Quick Wins**
- Identify early victories
- Celebrate visibly
- Build momentum
7. **Consolidate Gains**
- Don't declare victory early
- Continue driving change
- Address deeper issues
8. **Anchor in Culture**
- Reinforce through systems
- Celebrate new behaviors
- Ensure leadership continuity
### Organizational Design
#### Design Principles
**Customer-Centric**
- Organize around customer needs
- Minimize handoffs
- Clear ownership
- Fast decision-making
**Scalable**
- Consistent structures
- Clear roles/responsibilities
- Repeatable processes
- Growth-ready
**Agile**
- Cross-functional teams
- Rapid iteration
- Continuous learning
- Adaptive planning
**Efficient**
- Appropriate spans of control (5-7)
- Minimal layers (max 5-6)
- Clear decision rights
- Eliminated redundancy
#### Reorganization Playbook
**Pre-announcement** (4-6 weeks)
- Design new structure
- Identify leadership
- Plan communication
- Prepare materials
**Announcement** (Day 0)
- All-hands meeting
- Written communication
- Q&A sessions
- Manager toolkit
**Transition** (30 days)
- Role clarifications
- Team formations
- Process updates
- System changes
**Stabilization** (60-90 days)
- Monitor progress
- Address issues
- Refine as needed
- Celebrate success
## Crisis Leadership
### Crisis Response Framework
#### Leadership During Crisis
**Immediate Response** (0-24 hours)
- Establish command center
- Assess situation
- Communicate frequently
- Make rapid decisions
- Show visible leadership
**Stabilization** (1-7 days)
- Implement solutions
- Maintain communication
- Support teams
- Monitor progress
- Adjust approach
**Recovery** (1-4 weeks)
- Execute recovery plan
- Address long-term impacts
- Learn from crisis
- Strengthen resilience
- Recognize heroes
#### Crisis Communication
**Internal Communication**:
- Frequency: 2x daily minimum
- Channels: Email, video, town halls
- Content: Facts, actions, support
- Tone: Calm, confident, caring
**External Communication**:
- Stakeholders: Customers, partners, investors, media
- Frequency: As needed
- Channels: Website, press, social
- Content: Impact, response, timeline
- Tone: Transparent, responsible
## Innovation Culture
### Innovation Framework
#### Innovation Portfolio
**Horizon 1** (70% resources)
- Core business innovation
- Incremental improvements
- 6-18 month timeline
- Lower risk
**Horizon 2** (20% resources)
- Emerging opportunities
- Adjacent markets
- 18-36 month timeline
- Moderate risk
**Horizon 3** (10% resources)
- Transformational bets
- New business models
- 3-5 year timeline
- Higher risk
#### Innovation Programs
**Innovation Time**
- 20% time for projects
- Hackathons quarterly
- Innovation challenges
- Idea platforms
- Patent incentives
**Innovation Metrics**
- % revenue from new products
- Ideas generated/implemented
- Time to market
- Innovation ROI
- Patent applications
## Diversity, Equity & Inclusion
### DEI Strategy Framework
#### Four Pillars of DEI
1. **Representation**
- Diverse hiring
- Promotion equity
- Leadership diversity
- Board diversity
2. **Inclusion**
- Belonging index
- Psychological safety
- Equitable practices
- Bias mitigation
3. **Development**
- Sponsorship programs
- ERG support
- Leadership development
- Career pathways
4. **Accountability**
- DEI metrics
- Leader goals
- Regular reporting
- Transparency
#### DEI Metrics Dashboard
| Metric | Current | Target | Timeline |
|--------|---------|--------|----------|
| Women in leadership | X% | Y% | Z years |
| Ethnic diversity | X% | Y% | Z years |
| Pay equity gap | X% | 0% | Z years |
| Inclusion index | X/100 | Y/100 | Z years |
| Retention equality | X% diff | 0% diff | Z years |
## Executive Presence
### CEO Personal Brand
#### Brand Elements
**Vision**: What future you're creating
**Values**: What you stand for
**Voice**: How you communicate
**Visibility**: Where you show up
**Value**: What you deliver
#### Executive Communication
**Speaking Frameworks**:
**PREP Method**:
- **P**oint: Main message
- **R**eason: Why it matters
- **E**xample: Concrete illustration
- **P**oint: Restate message
**STAR Method** (for stories):
- **S**ituation: Context
- **T**ask: Challenge
- **A**ction: What was done
- **R**esult: Outcome
#### Media Training Essentials
**Key Message Discipline**:
- 3 key messages maximum
- Bridge to messages
- Sound bites ready
- Avoid speculation
- Stay on record
**Interview Techniques**:
- Pause before answering
- Bridge to key messages
- Use examples/stories
- Maintain eye contact
- Control pace
FILE:scripts/financial_scenario_analyzer.py
#!/usr/bin/env python3
"""
Financial Scenario Analyzer - Model different business scenarios and their financial impact
"""
import json
from typing import Dict, List, Tuple
import math
class FinancialScenarioAnalyzer:
def __init__(self):
self.key_metrics = [
'revenue', 'gross_margin', 'operating_expenses',
'ebitda', 'cash_flow', 'runway', 'valuation'
]
self.growth_models = {
'linear': lambda base, rate, period: base * (1 + rate * period),
'exponential': lambda base, rate, period: base * math.pow(1 + rate, period),
'logarithmic': lambda base, rate, period: base * (1 + rate * math.log(period + 1)),
's_curve': lambda base, rate, period: base * (2 / (1 + math.exp(-rate * period)))
}
def analyze_scenarios(self, base_case: Dict, scenarios: List[Dict]) -> Dict:
"""Analyze multiple financial scenarios"""
results = {
'base_case_summary': self._summarize_financials(base_case),
'scenario_analysis': [],
'sensitivity_analysis': {},
'recommendation': {},
'risk_adjusted_view': {}
}
# Analyze each scenario
for scenario in scenarios:
scenario_result = self._analyze_scenario(base_case, scenario)
results['scenario_analysis'].append(scenario_result)
# Sensitivity analysis
results['sensitivity_analysis'] = self._perform_sensitivity_analysis(
base_case,
scenarios
)
# Risk-adjusted view
results['risk_adjusted_view'] = self._calculate_risk_adjusted_returns(
results['scenario_analysis']
)
# Generate recommendation
results['recommendation'] = self._generate_recommendation(
results['scenario_analysis'],
results['risk_adjusted_view']
)
return results
def _summarize_financials(self, financials: Dict) -> Dict:
"""Summarize key financial metrics"""
revenue = financials.get('revenue', 0)
cogs = financials.get('cogs', 0)
opex = financials.get('operating_expenses', 0)
gross_profit = revenue - cogs
gross_margin = (gross_profit / revenue * 100) if revenue > 0 else 0
ebitda = gross_profit - opex
ebitda_margin = (ebitda / revenue * 100) if revenue > 0 else 0
return {
'revenue': revenue,
'gross_profit': gross_profit,
'gross_margin': gross_margin,
'operating_expenses': opex,
'ebitda': ebitda,
'ebitda_margin': ebitda_margin,
'cash': financials.get('cash', 0),
'burn_rate': financials.get('burn_rate', 0),
'runway_months': self._calculate_runway(
financials.get('cash', 0),
financials.get('burn_rate', 0)
)
}
def _calculate_runway(self, cash: float, burn_rate: float) -> float:
"""Calculate months of runway"""
if burn_rate <= 0:
return float('inf')
return cash / burn_rate
def _analyze_scenario(self, base_case: Dict, scenario: Dict) -> Dict:
"""Analyze a single scenario"""
name = scenario.get('name', 'Unnamed Scenario')
probability = scenario.get('probability', 0.5)
# Apply scenario changes
projected_financials = self._apply_scenario_changes(base_case, scenario)
# Calculate metrics for each year
projections = []
current_state = projected_financials.copy()
for year in range(1, 4): # 3-year projection
year_projection = self._project_year(
current_state,
scenario,
year
)
projections.append(year_projection)
current_state = year_projection
# Calculate NPV and IRR
cash_flows = [p['free_cash_flow'] for p in projections]
npv = self._calculate_npv(cash_flows, scenario.get('discount_rate', 0.1))
irr = self._calculate_irr(cash_flows, base_case.get('initial_investment', 0))
return {
'name': name,
'probability': probability,
'projections': projections,
'npv': npv,
'irr': irr,
'break_even_month': self._find_break_even(projections),
'total_return': self._calculate_total_return(projections, base_case),
'key_assumptions': scenario.get('assumptions', [])
}
def _apply_scenario_changes(self, base_case: Dict, scenario: Dict) -> Dict:
"""Apply scenario changes to base case"""
result = base_case.copy()
changes = scenario.get('changes', {})
for key, change in changes.items():
if key in result:
if isinstance(change, dict):
# Relative change
if 'multiply' in change:
result[key] *= change['multiply']
elif 'add' in change:
result[key] += change['add']
else:
# Absolute change
result[key] = change
return result
def _project_year(self, current_state: Dict, scenario: Dict, year: int) -> Dict:
"""Project financials for a specific year"""
growth_model = scenario.get('growth_model', 'exponential')
growth_rate = scenario.get('growth_rate', 0.3)
# Apply growth model
model_func = self.growth_models.get(growth_model, self.growth_models['linear'])
revenue = model_func(
current_state.get('revenue', 0),
growth_rate,
year
)
# Scale other metrics
cogs = revenue * scenario.get('cogs_ratio', 0.3)
opex = current_state.get('operating_expenses', 0) * (1 + scenario.get('opex_growth', 0.15))
gross_profit = revenue - cogs
ebitda = gross_profit - opex
# Calculate free cash flow (simplified)
capex = revenue * scenario.get('capex_ratio', 0.05)
working_capital_change = (revenue - current_state.get('revenue', 0)) * 0.1
free_cash_flow = ebitda - capex - working_capital_change
return {
'year': year,
'revenue': revenue,
'gross_profit': gross_profit,
'gross_margin': (gross_profit / revenue * 100) if revenue > 0 else 0,
'operating_expenses': opex,
'ebitda': ebitda,
'ebitda_margin': (ebitda / revenue * 100) if revenue > 0 else 0,
'free_cash_flow': free_cash_flow,
'cumulative_cash_flow': current_state.get('cumulative_cash_flow', 0) + free_cash_flow
}
def _calculate_npv(self, cash_flows: List[float], discount_rate: float) -> float:
"""Calculate Net Present Value"""
npv = 0
for i, cf in enumerate(cash_flows):
npv += cf / math.pow(1 + discount_rate, i + 1)
return npv
def _calculate_irr(self, cash_flows: List[float], initial_investment: float) -> float:
"""Calculate Internal Rate of Return (simplified)"""
if not cash_flows or initial_investment == 0:
return 0
# Simple IRR approximation
total_return = sum(cash_flows)
years = len(cash_flows)
if initial_investment > 0:
return math.pow(total_return / initial_investment, 1/years) - 1
return 0
def _find_break_even(self, projections: List[Dict]) -> int:
"""Find break-even month"""
months = 0
for projection in projections:
months += 12
if projection.get('ebitda', 0) > 0:
# Interpolate to find exact month
if months == 12:
return months
prev_ebitda = projections[projection['year']-2].get('ebitda', 0) if projection['year'] > 1 else 0
monthly_improvement = (projection['ebitda'] - prev_ebitda) / 12
if monthly_improvement > 0:
months_to_breakeven = abs(prev_ebitda) / monthly_improvement
return int(months - 12 + months_to_breakeven)
return -1 # Not reached
def _calculate_total_return(self, projections: List[Dict], base_case: Dict) -> float:
"""Calculate total return multiple"""
initial = base_case.get('valuation', 1000000)
# Simple valuation at end (10x revenue multiple for SaaS)
final_revenue = projections[-1]['revenue'] if projections else 0
final_valuation = final_revenue * 10
return (final_valuation / initial) if initial > 0 else 0
def _perform_sensitivity_analysis(self, base_case: Dict, scenarios: List[Dict]) -> Dict:
"""Perform sensitivity analysis on key variables"""
sensitivity = {}
key_variables = ['growth_rate', 'gross_margin', 'customer_acquisition_cost']
for variable in key_variables:
sensitivity[variable] = {
'low': self._calculate_variable_impact(base_case, variable, -0.2),
'base': self._calculate_variable_impact(base_case, variable, 0),
'high': self._calculate_variable_impact(base_case, variable, 0.2)
}
return sensitivity
def _calculate_variable_impact(self, base_case: Dict, variable: str, change: float) -> float:
"""Calculate impact of variable change on valuation"""
# Simplified impact calculation
impacts = {
'growth_rate': 2.5, # 2.5x multiplier on valuation
'gross_margin': 1.8, # 1.8x multiplier
'customer_acquisition_cost': -1.2 # Negative impact
}
base_value = 10000000 # Base valuation
impact_multiplier = impacts.get(variable, 1.0)
return base_value * (1 + change * impact_multiplier)
def _calculate_risk_adjusted_returns(self, scenarios: List[Dict]) -> Dict:
"""Calculate risk-adjusted returns"""
expected_value = 0
best_case = None
worst_case = None
for scenario in scenarios:
probability = scenario['probability']
npv = scenario['npv']
expected_value += probability * npv
if best_case is None or npv > best_case['npv']:
best_case = scenario
if worst_case is None or npv < worst_case['npv']:
worst_case = scenario
# Calculate standard deviation (simplified)
variance = sum([
scenario['probability'] * math.pow(scenario['npv'] - expected_value, 2)
for scenario in scenarios
])
std_dev = math.sqrt(variance)
return {
'expected_value': expected_value,
'best_case': best_case['name'] if best_case else 'None',
'best_case_npv': best_case['npv'] if best_case else 0,
'worst_case': worst_case['name'] if worst_case else 'None',
'worst_case_npv': worst_case['npv'] if worst_case else 0,
'standard_deviation': std_dev,
'sharpe_ratio': (expected_value / std_dev) if std_dev > 0 else 0
}
def _generate_recommendation(self, scenarios: List[Dict], risk_adjusted: Dict) -> Dict:
"""Generate recommendation based on analysis"""
recommendation = {
'recommended_scenario': '',
'rationale': [],
'key_actions': [],
'risk_mitigation': []
}
# Find optimal scenario
best_risk_adjusted = max(scenarios, key=lambda s: s['npv'] * s['probability'])
recommendation['recommended_scenario'] = best_risk_adjusted['name']
# Generate rationale
if best_risk_adjusted['npv'] > 0:
recommendation['rationale'].append(f"Positive NPV of ,.0f")
if best_risk_adjusted['irr'] > 0.15:
recommendation['rationale'].append(f"Strong IRR of {best_risk_adjusted['irr']:.1%}")
if best_risk_adjusted['break_even_month'] > 0 and best_risk_adjusted['break_even_month'] < 24:
recommendation['rationale'].append(f"Quick path to profitability ({best_risk_adjusted['break_even_month']} months)")
# Key actions
recommendation['key_actions'] = [
'Secure funding for growth initiatives',
'Build scalable operational infrastructure',
'Invest in customer acquisition channels',
'Strengthen unit economics',
'Establish financial controls'
]
# Risk mitigation
if risk_adjusted['standard_deviation'] > risk_adjusted['expected_value'] * 0.5:
recommendation['risk_mitigation'].append('High variability - consider hedging strategies')
recommendation['risk_mitigation'].extend([
'Maintain 12+ months runway',
'Diversify revenue streams',
'Build contingency plans for downside scenarios'
])
return recommendation
def analyze_financial_scenarios(base_case: Dict, scenarios: List[Dict]) -> str:
"""Main function to analyze financial scenarios"""
analyzer = FinancialScenarioAnalyzer()
results = analyzer.analyze_scenarios(base_case, scenarios)
# Format output
output = [
"=== Financial Scenario Analysis ===",
"",
"Base Case Summary:",
f" Revenue: ,.0f",
f" Gross Margin: {results['base_case_summary']['gross_margin']:.1f}%",
f" EBITDA: ,.0f",
f" Runway: {results['base_case_summary']['runway_months']:.1f} months",
"",
"Scenario Analysis:"
]
for scenario in results['scenario_analysis']:
output.append(f"\n{scenario['name']} (Probability: {scenario['probability']:.0%})")
output.append(f" NPV: ,.0f")
output.append(f" IRR: {scenario['irr']:.1%}")
output.append(f" Break-even: {scenario['break_even_month']} months")
output.append(f" Return Multiple: {scenario['total_return']:.1f}x")
# Show Year 3 projection
if scenario['projections']:
year3 = scenario['projections'][-1]
output.append(f" Year 3 Revenue: ,.0f")
output.append(f" Year 3 EBITDA Margin: {year3['ebitda_margin']:.1f}%")
output.extend([
"",
"Risk-Adjusted Analysis:",
f" Expected Value: ,.0f",
f" Best Case: {results['risk_adjusted_view']['best_case']} (,.0f)",
f" Worst Case: {results['risk_adjusted_view']['worst_case']} (,.0f)",
f" Risk (Std Dev): ,.0f",
f" Sharpe Ratio: {results['risk_adjusted_view']['sharpe_ratio']:.2f}",
"",
f"RECOMMENDATION: {results['recommendation']['recommended_scenario']}",
"",
"Rationale:"
])
for reason in results['recommendation']['rationale']:
output.append(f" • {reason}")
output.extend([
"",
"Key Actions:"
])
for action in results['recommendation']['key_actions'][:3]:
output.append(f" • {action}")
return '\n'.join(output)
if __name__ == "__main__":
# Example usage
example_base_case = {
'revenue': 5000000,
'cogs': 1500000,
'operating_expenses': 3000000,
'cash': 2000000,
'burn_rate': 200000,
'valuation': 20000000,
'initial_investment': 5000000
}
example_scenarios = [
{
'name': 'Aggressive Growth',
'probability': 0.3,
'growth_model': 'exponential',
'growth_rate': 0.5,
'changes': {
'operating_expenses': {'multiply': 1.3}
},
'assumptions': ['Market expansion successful', 'Product-market fit achieved'],
'cogs_ratio': 0.25,
'opex_growth': 0.3,
'capex_ratio': 0.08,
'discount_rate': 0.12
},
{
'name': 'Moderate Growth',
'probability': 0.5,
'growth_model': 'exponential',
'growth_rate': 0.3,
'changes': {},
'assumptions': ['Steady market growth', 'Competition remains stable'],
'cogs_ratio': 0.3,
'opex_growth': 0.15,
'capex_ratio': 0.05,
'discount_rate': 0.10
},
{
'name': 'Conservative',
'probability': 0.2,
'growth_model': 'linear',
'growth_rate': 0.15,
'changes': {
'operating_expenses': {'multiply': 0.9}
},
'assumptions': ['Market headwinds', 'Focus on profitability'],
'cogs_ratio': 0.35,
'opex_growth': 0.05,
'capex_ratio': 0.03,
'discount_rate': 0.08
}
]
print(analyze_financial_scenarios(example_base_case, example_scenarios))
FILE:scripts/strategy_analyzer.py
#!/usr/bin/env python3
"""
Strategic Planning Analyzer - Comprehensive business strategy assessment tool
"""
import json
from typing import Dict, List, Tuple
from datetime import datetime, timedelta
import math
class StrategyAnalyzer:
def __init__(self):
self.strategic_pillars = {
'market_position': {
'weight': 0.25,
'factors': ['market_share', 'brand_strength', 'competitive_advantage', 'customer_loyalty']
},
'financial_health': {
'weight': 0.25,
'factors': ['revenue_growth', 'profitability', 'cash_flow', 'unit_economics']
},
'operational_excellence': {
'weight': 0.20,
'factors': ['efficiency', 'quality', 'scalability', 'innovation']
},
'organizational_capability': {
'weight': 0.20,
'factors': ['talent', 'culture', 'leadership', 'agility']
},
'growth_potential': {
'weight': 0.10,
'factors': ['market_size', 'expansion_opportunities', 'product_pipeline', 'partnerships']
}
}
self.strategic_frameworks = {
'porter_five_forces': [
'competitive_rivalry',
'supplier_power',
'buyer_power',
'threat_of_substitution',
'threat_of_new_entry'
],
'swot': ['strengths', 'weaknesses', 'opportunities', 'threats'],
'bcg_matrix': ['stars', 'cash_cows', 'question_marks', 'dogs'],
'ansoff_matrix': ['market_penetration', 'market_development', 'product_development', 'diversification']
}
def analyze_strategic_position(self, company_data: Dict) -> Dict:
"""Comprehensive strategic analysis"""
results = {
'timestamp': datetime.now().isoformat(),
'company': company_data.get('name', 'Company'),
'strategic_health_score': 0,
'pillar_analysis': {},
'framework_analysis': {},
'strategic_options': [],
'risk_assessment': {},
'recommendations': [],
'roadmap': {}
}
# Analyze strategic pillars
total_score = 0
for pillar, config in self.strategic_pillars.items():
pillar_score = self._analyze_pillar(
company_data.get(pillar, {}),
config['factors']
)
weighted_score = pillar_score * config['weight']
results['pillar_analysis'][pillar] = {
'score': pillar_score,
'weighted_score': weighted_score,
'level': self._get_level(pillar_score),
'factors': self._get_pillar_details(company_data.get(pillar, {}), config['factors'])
}
total_score += weighted_score
results['strategic_health_score'] = round(total_score, 1)
# Framework analysis
results['framework_analysis'] = self._apply_frameworks(company_data)
# Generate strategic options
results['strategic_options'] = self._generate_strategic_options(
results['pillar_analysis'],
company_data.get('context', {})
)
# Risk assessment
results['risk_assessment'] = self._assess_strategic_risks(
company_data,
results['strategic_options']
)
# Generate roadmap
results['roadmap'] = self._create_strategic_roadmap(
results['strategic_options'],
company_data.get('timeline', 12)
)
# Generate recommendations
results['recommendations'] = self._generate_recommendations(results)
return results
def _analyze_pillar(self, pillar_data: Dict, factors: List) -> float:
"""Analyze a strategic pillar"""
if not pillar_data:
return 50.0
total_score = 0
count = 0
for factor in factors:
if factor in pillar_data:
score = pillar_data[factor]
total_score += score
count += 1
return (total_score / count) if count > 0 else 50.0
def _get_pillar_details(self, pillar_data: Dict, factors: List) -> List[Dict]:
"""Get detailed factor analysis"""
details = []
for factor in factors:
score = pillar_data.get(factor, 50)
details.append({
'factor': factor.replace('_', ' ').title(),
'score': score,
'status': 'Strong' if score >= 70 else 'Adequate' if score >= 40 else 'Weak'
})
return details
def _get_level(self, score: float) -> str:
"""Convert score to level"""
if score >= 80:
return 'Excellent'
elif score >= 70:
return 'Strong'
elif score >= 50:
return 'Adequate'
elif score >= 30:
return 'Weak'
else:
return 'Critical'
def _apply_frameworks(self, company_data: Dict) -> Dict:
"""Apply strategic frameworks"""
frameworks = {}
# SWOT Analysis
swot_data = company_data.get('swot', {})
frameworks['swot'] = {
'strengths': swot_data.get('strengths', [
'Strong brand recognition',
'Experienced leadership team',
'Robust technology platform'
]),
'weaknesses': swot_data.get('weaknesses', [
'Limited geographic presence',
'High customer acquisition cost',
'Technical debt'
]),
'opportunities': swot_data.get('opportunities', [
'Growing market demand',
'M&A opportunities',
'New product categories'
]),
'threats': swot_data.get('threats', [
'Increasing competition',
'Regulatory changes',
'Economic uncertainty'
])
}
# Porter's Five Forces
forces = company_data.get('competitive_forces', {})
frameworks['porter_analysis'] = {
'competitive_rivalry': forces.get('rivalry', 70),
'supplier_power': forces.get('suppliers', 40),
'buyer_power': forces.get('buyers', 60),
'threat_of_substitutes': forces.get('substitutes', 50),
'threat_of_new_entrants': forces.get('new_entrants', 45),
'overall_attractiveness': self._calculate_industry_attractiveness(forces)
}
# BCG Matrix for product portfolio
products = company_data.get('products', [])
frameworks['portfolio_analysis'] = self._analyze_portfolio(products)
return frameworks
def _calculate_industry_attractiveness(self, forces: Dict) -> float:
"""Calculate industry attractiveness from Porter's forces"""
# Lower forces = more attractive industry
rivalry = 100 - forces.get('rivalry', 50)
supplier = 100 - forces.get('suppliers', 50)
buyer = 100 - forces.get('buyers', 50)
substitutes = 100 - forces.get('substitutes', 50)
new_entrants = 100 - forces.get('new_entrants', 50)
avg = (rivalry + supplier + buyer + substitutes + new_entrants) / 5
return round(avg, 1)
def _analyze_portfolio(self, products: List) -> Dict:
"""Analyze product portfolio using BCG matrix"""
portfolio = {
'stars': [],
'cash_cows': [],
'question_marks': [],
'dogs': []
}
for product in products:
growth = product.get('market_growth', 0)
share = product.get('market_share', 0)
if growth > 10 and share > 50:
portfolio['stars'].append(product.get('name', 'Product'))
elif growth <= 10 and share > 50:
portfolio['cash_cows'].append(product.get('name', 'Product'))
elif growth > 10 and share <= 50:
portfolio['question_marks'].append(product.get('name', 'Product'))
else:
portfolio['dogs'].append(product.get('name', 'Product'))
return portfolio
def _generate_strategic_options(self, pillar_analysis: Dict, context: Dict) -> List[Dict]:
"""Generate strategic options based on analysis"""
options = []
# Check market position
market_score = pillar_analysis['market_position']['score']
if market_score < 60:
options.append({
'name': 'Market Leadership Initiative',
'type': 'market_penetration',
'description': 'Aggressive market share capture through competitive pricing and marketing',
'investment': 'High',
'timeframe': '12-18 months',
'expected_impact': 'Increase market share by 10-15%',
'priority': 9
})
# Check financial health
financial_score = pillar_analysis['financial_health']['score']
if financial_score < 50:
options.append({
'name': 'Profitability Turnaround',
'type': 'operational_excellence',
'description': 'Cost reduction and revenue optimization program',
'investment': 'Medium',
'timeframe': '6-9 months',
'expected_impact': 'Improve margins by 5-8%',
'priority': 10
})
# Check growth potential
growth_score = pillar_analysis['growth_potential']['score']
if growth_score > 70:
options.append({
'name': 'Expansion Strategy',
'type': 'market_development',
'description': 'Enter new geographic markets or customer segments',
'investment': 'High',
'timeframe': '18-24 months',
'expected_impact': 'Revenue growth of 30-40%',
'priority': 8
})
# Innovation opportunities
if context.get('industry_disruption', False):
options.append({
'name': 'Digital Transformation',
'type': 'innovation',
'description': 'Comprehensive digitalization of business processes and customer experience',
'investment': 'Very High',
'timeframe': '24-36 months',
'expected_impact': 'Future-proof business model',
'priority': 9
})
# M&A opportunities
if context.get('cash_available', 0) > 100000000:
options.append({
'name': 'Strategic Acquisition',
'type': 'acquisition',
'description': 'Acquire complementary businesses or competitors',
'investment': 'Very High',
'timeframe': '6-12 months',
'expected_impact': 'Instant scale and capability',
'priority': 7
})
# Sort by priority
options.sort(key=lambda x: x['priority'], reverse=True)
return options[:5] # Top 5 strategic options
def _assess_strategic_risks(self, company_data: Dict, strategic_options: List) -> Dict:
"""Assess strategic risks"""
risks = {
'execution_risk': self._calculate_execution_risk(company_data),
'market_risk': self._calculate_market_risk(company_data),
'financial_risk': self._calculate_financial_risk(company_data),
'competitive_risk': self._calculate_competitive_risk(company_data),
'regulatory_risk': company_data.get('regulatory_risk', 30),
'overall_risk': 0,
'mitigation_strategies': []
}
# Calculate overall risk
risk_values = [
risks['execution_risk'],
risks['market_risk'],
risks['financial_risk'],
risks['competitive_risk'],
risks['regulatory_risk']
]
risks['overall_risk'] = sum(risk_values) / len(risk_values)
# Generate mitigation strategies
if risks['execution_risk'] > 60:
risks['mitigation_strategies'].append({
'risk': 'Execution',
'strategy': 'Strengthen PMO, hire experienced executives, implement OKRs'
})
if risks['market_risk'] > 60:
risks['mitigation_strategies'].append({
'risk': 'Market',
'strategy': 'Diversify revenue streams, build strategic partnerships'
})
if risks['financial_risk'] > 60:
risks['mitigation_strategies'].append({
'risk': 'Financial',
'strategy': 'Improve cash management, secure credit facilities, optimize working capital'
})
return risks
def _calculate_execution_risk(self, data: Dict) -> float:
"""Calculate execution risk"""
org_capability = data.get('organizational_capability', {})
factors = [
100 - org_capability.get('leadership', 50),
100 - org_capability.get('talent', 50),
100 - org_capability.get('agility', 50),
data.get('complexity_score', 50)
]
return sum(factors) / len(factors)
def _calculate_market_risk(self, data: Dict) -> float:
"""Calculate market risk"""
market = data.get('market_position', {})
factors = [
100 - market.get('market_share', 50),
data.get('market_volatility', 50),
data.get('customer_concentration', 50)
]
return sum(factors) / len(factors)
def _calculate_financial_risk(self, data: Dict) -> float:
"""Calculate financial risk"""
financial = data.get('financial_health', {})
factors = [
100 - financial.get('cash_flow', 50),
100 - financial.get('profitability', 50),
data.get('debt_ratio', 50),
data.get('burn_rate', 50) if 'burn_rate' in data else 30
]
return sum(factors) / len(factors)
def _calculate_competitive_risk(self, data: Dict) -> float:
"""Calculate competitive risk"""
forces = data.get('competitive_forces', {})
return (forces.get('rivalry', 50) + forces.get('new_entrants', 50)) / 2
def _create_strategic_roadmap(self, options: List, timeline_months: int) -> Dict:
"""Create implementation roadmap"""
roadmap = {
'phases': [],
'milestones': [],
'resource_requirements': {},
'success_metrics': []
}
# Define phases
phases = [
{
'phase': 'Foundation',
'months': '0-3',
'focus': 'Build capabilities and quick wins',
'initiatives': []
},
{
'phase': 'Acceleration',
'months': '3-9',
'focus': 'Execute core strategies',
'initiatives': []
},
{
'phase': 'Scale',
'months': '9-18',
'focus': 'Expand and optimize',
'initiatives': []
},
{
'phase': 'Transform',
'months': '18+',
'focus': 'Long-term transformation',
'initiatives': []
}
]
# Assign initiatives to phases
for i, option in enumerate(options[:4]):
if i == 0:
phases[0]['initiatives'].append(option['name'])
elif i == 1:
phases[1]['initiatives'].append(option['name'])
elif i == 2:
phases[2]['initiatives'].append(option['name'])
else:
phases[3]['initiatives'].append(option['name'])
roadmap['phases'] = phases
# Define key milestones
roadmap['milestones'] = [
{'month': 3, 'milestone': 'Complete foundation phase', 'success_criteria': 'Core team hired, processes defined'},
{'month': 6, 'milestone': 'First major initiative launch', 'success_criteria': 'KPIs showing positive trend'},
{'month': 12, 'milestone': 'Strategic review', 'success_criteria': 'ROI demonstrated, strategy validated'},
{'month': 18, 'milestone': 'Scale achievement', 'success_criteria': 'Market position improved, financial targets met'}
]
# Resource requirements
roadmap['resource_requirements'] = {
'leadership': 'C-suite alignment and commitment',
'financial': '$X million investment over 18 months',
'human': 'Additional 20-30 FTEs across functions',
'technology': 'Platform upgrades and new tools',
'external': 'Consultants and advisors as needed'
}
# Success metrics
roadmap['success_metrics'] = [
'Revenue growth: 25% YoY',
'Market share: +5 percentage points',
'EBITDA margin: +8 percentage points',
'Customer NPS: >70',
'Employee engagement: >80%'
]
return roadmap
def _generate_recommendations(self, results: Dict) -> List[str]:
"""Generate strategic recommendations"""
recommendations = []
# Based on overall score
score = results['strategic_health_score']
if score < 40:
recommendations.append('🚨 URGENT: Immediate turnaround required - consider bringing in crisis management team')
recommendations.append('Focus on cash preservation and core business stabilization')
elif score < 60:
recommendations.append('⚠️ Strategic repositioning needed - prioritize 2-3 key initiatives')
recommendations.append('Strengthen weak pillars before pursuing growth')
elif score < 80:
recommendations.append('✓ Solid position - focus on selective improvements and growth')
recommendations.append('Invest in innovation and market expansion')
else:
recommendations.append('⭐ Excellent position - maintain momentum and explore bold moves')
recommendations.append('Consider industry disruption or category creation')
# Based on specific weaknesses
for pillar, analysis in results['pillar_analysis'].items():
if analysis['score'] < 50:
if pillar == 'market_position':
recommendations.append(f'Strengthen {pillar}: Launch competitive differentiation program')
elif pillar == 'financial_health':
recommendations.append(f'Improve {pillar}: Implement profitability improvement plan')
elif pillar == 'organizational_capability':
recommendations.append(f'Build {pillar}: Invest in talent and culture transformation')
# Based on opportunities
if results['framework_analysis']['porter_analysis']['overall_attractiveness'] > 70:
recommendations.append('Industry is attractive - consider aggressive expansion')
# Risk-based recommendations
if results['risk_assessment']['overall_risk'] > 60:
recommendations.append('High risk profile - implement comprehensive risk management')
return recommendations
def analyze_strategy(company_data: Dict) -> str:
"""Main function to analyze strategy"""
analyzer = StrategyAnalyzer()
results = analyzer.analyze_strategic_position(company_data)
# Format output
output = [
f"=== Strategic Analysis Report ===",
f"Company: {results['company']}",
f"Date: {results['timestamp'][:10]}",
f"",
f"STRATEGIC HEALTH SCORE: {results['strategic_health_score']}/100",
f"",
"Strategic Pillars:"
]
for pillar, analysis in results['pillar_analysis'].items():
output.append(f" {pillar.replace('_', ' ').title()}: {analysis['score']:.1f} ({analysis['level']})")
for factor in analysis['factors'][:2]: # Show top 2 factors
output.append(f" • {factor['factor']}: {factor['status']}")
output.extend([
f"",
"Strategic Options:"
])
for i, option in enumerate(results['strategic_options'][:3], 1):
output.append(f"\n{i}. {option['name']} (Priority: {option['priority']}/10)")
output.append(f" Type: {option['type']}")
output.append(f" Investment: {option['investment']}")
output.append(f" Timeframe: {option['timeframe']}")
output.append(f" Impact: {option['expected_impact']}")
output.extend([
f"",
f"Risk Assessment:",
f" Overall Risk: {results['risk_assessment']['overall_risk']:.1f}%",
f" Execution Risk: {results['risk_assessment']['execution_risk']:.1f}%",
f" Market Risk: {results['risk_assessment']['market_risk']:.1f}%",
f" Financial Risk: {results['risk_assessment']['financial_risk']:.1f}%",
f"",
"Strategic Roadmap:"
])
for phase in results['roadmap']['phases'][:3]:
output.append(f" {phase['phase']} ({phase['months']}): {phase['focus']}")
for initiative in phase['initiatives']:
output.append(f" • {initiative}")
output.extend([
f"",
"Key Recommendations:"
])
for rec in results['recommendations'][:5]:
output.append(f" • {rec}")
return '\n'.join(output)
if __name__ == "__main__":
# Example usage
example_company = {
'name': 'TechCorp Inc.',
'market_position': {
'market_share': 35,
'brand_strength': 65,
'competitive_advantage': 70,
'customer_loyalty': 60
},
'financial_health': {
'revenue_growth': 45,
'profitability': 40,
'cash_flow': 55,
'unit_economics': 60
},
'organizational_capability': {
'talent': 70,
'culture': 65,
'leadership': 75,
'agility': 60
},
'growth_potential': {
'market_size': 80,
'expansion_opportunities': 70,
'product_pipeline': 60,
'partnerships': 55
},
'competitive_forces': {
'rivalry': 70,
'suppliers': 40,
'buyers': 60,
'substitutes': 50,
'new_entrants': 45
},
'context': {
'industry_disruption': True,
'cash_available': 150000000
},
'timeline': 18
}
print(analyze_strategy(example_company))
Lãnh đạo tài chính: mô hình tài chính, unit economics, chiến lược gọi vốn, quản lý dòng tiền và báo cáo HĐQT.
---
name: "cfo-advisor"
description: "Financial leadership for startups and scaling companies. Financial modeling, unit economics, fundraising strategy, cash management, and board financial packages. Use when building financial models, analyzing unit economics, planning fundraising, managing cash runway, preparing board materials, or when user mentions CFO, burn rate, runway, fundraising, unit economics, LTV, CAC, term sheets, or financial strategy."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: cfo-leadership
updated: 2026-03-05
python-tools: burn_rate_calculator.py, unit_economics_analyzer.py, fundraising_model.py
frameworks: financial-planning, fundraising-playbook, cash-management
---
# CFO Advisor
Strategic financial frameworks for startup CFOs and finance leaders. Numbers-driven, decisions-focused.
This is **not** a financial analyst skill. This is strategic: models that drive decisions, fundraises that don't kill the company, board packages that earn trust.
## Keywords
CFO, chief financial officer, burn rate, runway, unit economics, LTV, CAC, fundraising, Series A, Series B, term sheet, cap table, dilution, financial model, cash flow, board financials, FP&A, SaaS metrics, ARR, MRR, net dollar retention, gross margin, scenario planning, cash management, treasury, working capital, burn multiple, rule of 40
## Quick Start
```bash
# Burn rate & runway scenarios (base/bull/bear)
python scripts/burn_rate_calculator.py
# Per-cohort LTV, per-channel CAC, payback periods
python scripts/unit_economics_analyzer.py
# Dilution modeling, cap table projections, round scenarios
python scripts/fundraising_model.py
```
## Key Questions (ask these first)
- **What's your burn multiple?** (Net burn ÷ Net new ARR. > 2x is a problem.)
- **If fundraising takes 6 months instead of 3, do you survive?** (If not, you're already behind.)
- **Show me unit economics per cohort, not blended.** (Blended hides deterioration.)
- **What's your NDR?** (> 100% means you grow without signing a single new customer.)
- **What are your decision triggers?** (At what runway do you start cutting? Define now, not in a crisis.)
## Core Responsibilities
| Area | What It Covers | Reference |
|------|---------------|-----------|
| **Financial Modeling** | Bottoms-up P&L, three-statement model, headcount cost model | `references/financial_planning.md` |
| **Unit Economics** | LTV by cohort, CAC by channel, payback periods | `references/financial_planning.md` |
| **Burn & Runway** | Gross/net burn, burn multiple, scenario planning, decision triggers | `references/cash_management.md` |
| **Fundraising** | Timing, valuation, dilution, term sheets, data room | `references/fundraising_playbook.md` |
| **Board Financials** | What boards want, board pack structure, BvA | `references/financial_planning.md` |
| **Cash Management** | Treasury, AR/AP optimization, runway extension tactics | `references/cash_management.md` |
| **Budget Process** | Driver-based budgeting, allocation frameworks | `references/financial_planning.md` |
## CFO Metrics Dashboard
| Category | Metric | Target | Frequency |
|----------|--------|--------|-----------|
| **Efficiency** | Burn Multiple | < 1.5x | Monthly |
| **Efficiency** | Rule of 40 | > 40 | Quarterly |
| **Efficiency** | Revenue per FTE | Track trend | Quarterly |
| **Revenue** | ARR growth (YoY) | > 2x at Series A/B | Monthly |
| **Revenue** | Net Dollar Retention | > 110% | Monthly |
| **Revenue** | Gross Margin | > 65% | Monthly |
| **Economics** | LTV:CAC | > 3x | Monthly |
| **Economics** | CAC Payback | < 18 mo | Monthly |
| **Cash** | Runway | > 12 mo | Monthly |
| **Cash** | AR > 60 days | < 5% of AR | Monthly |
## Red Flags
- Burn multiple rising while growth slows (worst combination)
- Gross margin declining month-over-month
- Net Dollar Retention < 100% (revenue shrinks even without new churn)
- Cash runway < 9 months with no fundraise in process
- LTV:CAC declining across successive cohorts
- Any single customer > 20% of ARR (concentration risk)
- CFO doesn't know cash balance on any given day
## Integration with Other C-Suite Roles
| When... | CFO works with... | To... |
|---------|-------------------|-------|
| Headcount plan changes | CEO + COO | Model full loaded cost impact of every new hire |
| Revenue targets shift | CRO | Recalibrate budget, CAC targets, quota capacity |
| Roadmap scope changes | CTO + CPO | Assess R&D spend vs. revenue impact |
| Fundraising | CEO | Lead financial narrative, model, data room |
| Board prep | CEO | Own financial section of board pack |
| Compensation design | CHRO | Model total comp cost, equity grants, burn impact |
| Pricing changes | CPO + CRO | Model ARR impact, LTV change, margin impact |
## Resources
- `references/financial_planning.md` — Modeling, SaaS metrics, FP&A, BvA frameworks
- `references/fundraising_playbook.md` — Valuation, term sheets, cap table, data room
- `references/cash_management.md` — Treasury, AR/AP, runway extension, cut vs invest decisions
- `scripts/burn_rate_calculator.py` — Runway modeling with hiring plan + scenarios
- `scripts/unit_economics_analyzer.py` — Per-cohort LTV, per-channel CAC
- `scripts/fundraising_model.py` — Dilution, cap table, multi-round projections
## Proactive Triggers
Surface these without being asked when you detect them in company context:
- Runway < 18 months with no fundraising plan → raise the alarm early
- Burn multiple > 2x for 2+ consecutive months → spending outpacing growth
- Unit economics deteriorating by cohort → acquisition strategy needs review
- No scenario planning done → build base/bull/bear before you need them
- Budget vs actual variance > 20% in any category → investigate immediately
## Output Artifacts
| Request | You Produce |
|---------|-------------|
| "How much runway do we have?" | Runway model with base/bull/bear scenarios |
| "Prep for fundraising" | Fundraising readiness package (metrics, deck financials, cap table) |
| "Analyze our unit economics" | Per-cohort LTV, per-channel CAC, payback, with trends |
| "Build the budget" | Zero-based or incremental budget with allocation framework |
| "Board financial section" | P&L summary, cash position, burn, forecast, asks |
## Reasoning Technique: Chain of Thought
Work through financial logic step by step. Show all math. Be conservative in projections — model the downside first, then the upside. Never round in your favor.
## Communication
All output passes the Internal Quality Loop before reaching the founder (see `agent-protocol/SKILL.md`).
- Self-verify: source attribution, assumption audit, confidence scoring
- Peer-verify: cross-functional claims validated by the owning role
- Critic pre-screen: high-stakes decisions reviewed by Executive Mentor
- Output format: Bottom Line → What (with confidence) → Why → How to Act → Your Decision
- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed.
## Context Integration
- **Always** read `company-context.md` before responding (if it exists)
- **During board meetings:** Use only your own analysis in Phase 2 (no cross-pollination)
- **Invocation:** You can request input from other roles: `[INVOKE:role|question]`
FILE:references/cash_management.md
# Cash Management Reference
Cash is the oxygen of a startup. You can be unprofitable for years. You cannot be out of cash for a day.
---
## 1. Cash Flow Management
### The Cash Equation
```
Ending Cash = Beginning Cash
+ Cash collected from customers
- Cash paid to employees
- Cash paid to vendors
- Cash paid for infrastructure
- Debt service
+/- Financing activities
Note: This is NOT the P&L. Revenue recognition ≠ cash collected.
```
### Where Cash Hides (and Leaks)
**Cash sources you might be under-using:**
- Deferred revenue (annual billing locks in cash 12 months early)
- Customer deposits on enterprise contracts
- Vendor payment terms (Net 60 instead of Net 30 = free float)
- AWS/GCP startup credits (often $25K–$100K available, widely unused)
- Revenue-based financing on predictable MRR
- Venture debt (non-dilutive, available post-Series A)
**Cash drains that sneak up on you:**
- Annual software licenses paid in Q1 (budget for the lump sum)
- Event sponsorships (often 6-12 months in advance)
- Recruiting fees (15-25% of first-year salary, due on hire)
- Legal fees (data room prep, fundraise close = $50K–$200K surprise)
- Late-paying enterprise customers (Net 60 in contract, pays Net 90 in practice)
### Cash Flow vs P&L: The Gap
**Scenario: $1M enterprise deal signed December 31**
```
P&L impact (accrual):
December revenue: $83K (1/12 of annual)
Cash impact:
If billed annually upfront: +$1,000K in December (GREAT)
If billed quarterly: +$250K in December (good)
If billed monthly: +$83K in December (fine)
If Net 60 terms: +$0 in December, +$83K in February (cash drag)
```
**The CFO's job:** Maximize the timing difference between cash in and cash out.
- Collect from customers as early as possible (annual upfront, early payment discounts)
- Pay vendors as late as possible (maximize payment terms)
- Never confuse deferred revenue (a liability) with actual cash (it is cash — just count it right)
---
## 2. Treasury and Banking Strategy
### Account Structure
```
Operating Account (primary bank):
Balance: 3-6 months of operating expenses
Purpose: Payroll, vendor payments, day-to-day ops
Product: Business checking or high-yield business savings
Bank: Chase, SVB successor (First Citizens), Mercury, Brex
Reserve Account (secondary or same bank):
Balance: Everything above operating float
Purpose: Reserve; move to operating as needed
Product: Money market fund or T-Bill ladder
Target yield (2024-2025): 4.5%–5.2%
Products: Vanguard VMFXX, Fidelity SPAXX, or direct T-Bills via TreasuryDirect
Emergency Account (separate bank):
Balance: 1-2 months expenses
Purpose: If primary bank has issues (SVB taught this lesson)
Product: Business savings
```
**FDIC coverage:** $250K per depositor per institution. For balances above $250K at a single bank, either:
- Use CDARS/ICS (bank sweeps into multiple FDIC-insured accounts automatically)
- Spread across multiple banks
- Move excess to T-Bills (backed by US government, not FDIC, but safer)
**After SVB (March 2023):** Every CFO should have at least 2 banking relationships. If one bank fails or freezes, you can make payroll.
### Yield on Cash
At $3M cash, the difference between 0% (checking) and 5% (T-Bills) is $150K/year.
That's a month of runway for a $150K/month burn company. **Get yield on reserves.**
```
Monthly yield on $3M at 5%: ~$12,500
Annual: ~$150,000
This is not optional. Set it up once and automate.
```
---
## 3. AR/AP Optimization
### Accounts Receivable: Get Paid Faster
**Billing model impact on cash:**
```
Annual Upfront Quarterly Monthly Net 30 Monthly
Cash Day 1: 100% of ACV 25% of ACV 8.3% 0%
Cash Month 2: 0% (done) 0% 8.3% 8.3%
12-month total: 100% 100% 100% 100%
For $100K ACV customer, Year 1 cash:
Annual upfront: $100K immediately
Monthly Net 30: $8.3K × 11 months = $91.7K (1 month lag)
Cash benefit: $100K vs $91.7K = $8.3K benefit + no collection risk
```
**Push for annual billing. Make it easy with a discount:**
```
"Pay annually and get 2 months free (16% discount)"
Most SMB customers will take this.
Enterprise: use MSA structure with annual invoicing, not month-to-month.
```
**AR Aging Policy:**
```
> 0-30 days: Current. No action.
> 30-60 days: Friendly reminder from AR team.
> 60-90 days: Escalate to Customer Success.
> 90 days: CFO or CEO-level outreach. Consider collections.
> 120 days: Reserve for bad debt. Legal/collections.
Reserve policy: 50% of 90-120 day AR, 100% of > 120 days
```
**What slows down collections:**
- Wrong contact (billing contact vs. user) — get finance contact during onboarding
- Enterprise PO required — know this upfront, not when invoice is due
- Credit holds or budget freeze — your CSM should surface these early
- Invoice errors — every wrong invoice extends payment by 30-60 days
### Accounts Payable: Pay Slower
**Standard terms by vendor type:**
```
SaaS tools: Net 30 default. Push for Net 45 or Net 60 at scale.
Cloud providers: Pay as you go. Apply for credits first.
Professional services (agencies, lawyers): Net 30 minimum. Get Net 45 where possible.
Rent/office: Whatever the lease says. Negotiate quarterly payments if you can.
Payroll: Pay on time. Never delay payroll. Ever.
```
**Early payment discount trap:**
```
"2/10 Net 30" means: 2% discount if you pay in 10 days, else pay in 30.
Annual cost of NOT taking this: 2% × (365/(30-10)) = ~36% APY
ALWAYS take early payment discounts > 2%.
Never take discounts < 1%.
```
**AP workflow:**
1. All invoices → finance inbox (not individual employees)
2. Approval required above threshold ($500 for startups)
3. Pay at end of terms, not when invoice arrives
4. Batch payments weekly (not daily) to reduce processing overhead
---
## 4. Runway Extension Tactics
Use these when you need to extend runway without raising. Ranked by speed and impact.
### Tier 1: Fast Cash (Days)
**Annual billing campaign:**
```
Target: Existing monthly customers
Offer: 2 months free (16% discount) or 1 month free (8% discount) for annual upfront
Process: CSM-led email campaign to all monthly customers
Impact: $X MRR × 12 × conversion rate = immediate cash injection
Timeline: 2-4 weeks
No dilution. No debt. High impact.
```
**Prepayment incentive for pipeline:**
```
For deals in late stage, offer annual upfront pricing with 10-15% discount.
Close rate may increase. Cash timing dramatically improves.
```
### Tier 2: Cost Control (2-4 Weeks)
**Hiring freeze:**
```
Every unfilled role = salary × 1.25 per month.
For a 30-person company, 3 open roles at $150K average:
Monthly savings: 3 × $150K × 1.25 / 12 = $47K/month
Over 6 months: $280K
Impact: Immediate. No blood.
```
**Software audit:**
```
Pull all credit card charges and ACH debits.
Cancel any subscription not used in 30 days.
Typical savings: $3K-$15K/month at Series A stage.
Tools: Vendr, Spendesk, or just a spreadsheet of recurring charges.
```
**Cloud cost optimization:**
```
Right-size instances (dev/staging don't need prod-scale)
Reserve instances (1-year reserved = 30-40% savings vs on-demand)
Delete unused resources (load balancers, IPs, old snapshots)
Typical savings: 20-35% of current cloud bill
```
### Tier 3: Vendor Renegotiation (2-6 Weeks)
**Payment term extension:**
```
Ask key vendors for Net 60 instead of Net 30.
$500K in AP × 30 days = $500K × (30/365) = ~$41K cash float improvement
Won't always work, but vendors often say yes to good customers.
```
**Renewal timing:**
```
Push annual renewals to later in the year.
Preserve cash for Q1 (typically heaviest sales hiring quarter).
```
**Vendor credits:**
```
AWS: AWS Activate (up to $100K for qualified startups)
GCP: Google for Startups (up to $200K)
Azure: Microsoft for Startups (up to $150K)
Stripe: Revenue share programs
Hubspot: Startup pricing (90% off)
```
### Tier 4: Financing (Weeks to Months)
**Revenue-based financing:**
```
Providers: Clearco, Capchase, Pipe, Arc
Structure: Advance 3-6 months of MRR. Repay with % of monthly revenue.
Cost: Typically 6-12% annualized.
Speed: 1-2 weeks to close.
When to use: Bridge to next ARR milestone before raising equity.
When NOT to use: When burn rate is structural (will consume the advance fast).
```
**Venture debt:**
```
Providers: SVB (now First Citizens), Western Technology Investment, Hercules, TriplePoint
Structure: Term loan, typically 3-6x monthly gross burn
Interest: Prime + 2-4% + warrants
When available: Post-Series A, when revenue is predictable
Typical timing: Add alongside an equity round (don't raise debt when you need equity)
Impact: Extends runway 3-6 months without dilution
When NOT to use: If you might trip financial covenants (minimum cash, revenue)
```
**Convertible bridge:**
```
Existing investors write bridge note: $500K-$2M at favorable terms.
Structure: Converts at discount (10-20%) or cap into next equity round.
When to use: You're 60-90 days from closing an equity round and need cash to get there.
When NOT to use: As a long-term strategy. Bridge-to-bridge is a death spiral.
```
### Tier 5: Structural Cost Reduction (Weeks + Impact on Morale)
**Salary deferrals (founders first):**
```
Founders take 20-30% salary reduction, accrued for future repayment.
Signals commitment to team and investors.
Only ask employees to follow if founders go first.
Always pay market rate to key non-founder employees — you can't afford to lose them.
```
**Reduction in force (RIF):**
```
Threshold: If burn multiple > 3x and growth < 20% YoY, a RIF is likely necessary.
Sizing: Model to achieve at least 12 months runway without fundraising.
Rule: Don't do a RIF twice. Size it right the first time.
Two small RIFs destroy morale worse than one decisive one.
Process: Legal counsel required. WARN Act (60-day notice) if > 100 employees.
Focus cuts: G&A and underperforming sales roles first. Protect engineering and key revenue.
```
---
## 5. When to Cut vs When to Invest
### The Framework
**Cut when:**
- Burn multiple > 2x and growth is decelerating
- Runway < 9 months with no fundraise imminent
- LTV:CAC declining for 3+ consecutive months
- Any spend category with no measurable return in 90 days
- Headcount in functions not directly tied to near-term revenue or product-market fit
**Invest when:**
- Magic number > 1 (every dollar in S&M returns > $1 in gross profit)
- LTV:CAC > 3x in a specific channel (pour money in)
- Gross margin > 70% (unit economics are healthy; growth is the constraint)
- Cohort data improving (retention getting better → LTV going up → invest in growth)
- CAC payback < 12 months (you get your money back fast enough to keep reinvesting)
### The False Economy Trap
**Don't cut:**
- Top-of-funnel demand gen that generates qualified pipeline (if CAC payback is < 12 months, this is your best investment)
- Engineering capacity on core product (technical debt compounds and slows you down permanently)
- Key account managers on your largest customers (churn from top customers is catastrophic)
**Cut these first:**
- Conference sponsorships with no measurable pipeline
- Tools and subscriptions with < 5 users or < 30% utilization
- Agency spend that could be done in-house
- Roadmap items that aren't tied to retention or expansion revenue
- Any G&A spend that isn't legally required
### Decision Triggers (Pre-Define These)
Don't make these decisions in a crisis. Define the triggers now:
```
At 12 months runway: Review all discretionary spend. Start fundraise process.
At 9 months runway: Implement hiring freeze. Fundraise is mandatory.
At 6 months runway: Cut non-essential spend 20%. If no fundraise term sheet, run RIF model.
At 4 months runway: Execute RIF. Explore all financing options. Notify board.
At 3 months runway: Emergency plan only. All options on table (bridge, strategic, wind down).
```
---
## Key Formulas
```python
# Net burn
net_burn = gross_burn - revenue_collected
# Runway (months)
runway_months = cash_balance / net_burn
# Cash conversion cycle
ccc = days_sales_outstanding + days_inventory_held - days_payable_outstanding
# Lower CCC = better cash efficiency
# Days Sales Outstanding (DSO)
dso = (accounts_receivable / revenue) * 30 # monthly revenue
# Days Payable Outstanding (DPO)
dpo = (accounts_payable / cogs) * 30 # target: maximize this
# Working capital
working_capital = current_assets - current_liabilities
# Quick ratio (liquidity)
quick_ratio_liquidity = (cash + ar) / current_liabilities
# Target: > 1.5 (you can pay short-term obligations without selling assets)
# Free cash flow
fcf = operating_cash_flow - capex
```
FILE:references/financial_planning.md
# Financial Planning Reference
Startup financial modeling frameworks. Build models that drive decisions, not models that impress investors.
---
## 1. Startup Financial Modeling
### Bottoms-Up vs Top-Down
**Top-down model (don't use for operating):**
```
TAM = $10B
SOM = 1% = $100M
Revenue = $100M in year 5
```
This is marketing. You cannot manage a company against these numbers.
**Bottoms-up model (use this):**
```
Year 1 Revenue Build:
Sales headcount: 3 AEs by Q1, +2 in Q2, +3 in Q4
Ramp curve: Month 1-3 = 25%, Month 4-6 = 75%, Month 7+ = 100%
Quota per ramped AE: $600K ARR
Effective quota (weighted for ramp): $1.2M ARR in Year 1
Win rate: 25%
Average deal: $48K ACV
Pipeline needed: $1.2M / 25% = $4.8M ARR pipeline
Required meetings to create that pipeline: $4.8M / (conversion 20%) / ($48K ACV × 0.5 to meeting) = ~200 meetings
```
Now you have something actionable. You know how many SDR calls, how many marketing leads, what conversion rate you need to hold. Every assumption is visible and challengeable.
### Building the Operating Model
#### Revenue Engine
**New ARR Model (SaaS):**
```
Month N New ARR:
= Quota-carrying reps (fully ramped equivalent)
× Attainment rate (typically 70-80% of quota)
× Average deal size
+ PLG / self-serve (if applicable)
Quota-carrying reps (ramped equivalent):
= Sum(each rep × their ramp factor)
Ramp schedule:
Month 1-2: 0% (onboarding)
Month 3: 25%
Month 4-6: 50%
Month 7-9: 75%
Month 10+: 100%
```
**ARR Bridge (most important recurring visual):**
```
Beginning ARR
+ New ARR (new logos)
+ Expansion ARR (upsells, seat growth)
- Churned ARR (cancellations)
- Contraction ARR (downgrades)
= Ending ARR
Net ARR Added = New + Expansion - Churn - Contraction
Net Dollar Retention (NDR):
= (Beginning ARR + Expansion - Churn - Contraction) / Beginning ARR × 100
Target: > 110% for growth-stage SaaS
World-class: > 130% (Snowflake, Twilio-tier)
```
**MRR and ARR Relationship:**
```
ARR = MRR × 12 (simple, always use this)
Never mix monthly and annual contracts in MRR without normalization.
Annual contract booked = ACV / 12 = monthly contribution to ARR
Multi-year contracts: book each year at annual value (not multi-year total)
```
#### Headcount Model
Headcount is usually 60-80% of total costs. Model it carefully.
```
For each role:
- Start date
- Department
- Annual salary (from salary bands)
- Loaded cost (salary × 1.25-1.45 depending on benefits + recruiting method)
- Productive from (ramp period)
- Impact on revenue (for revenue-generating roles)
Total headcount cost = Σ (each FTE × loaded cost × months active / 12)
```
**Department headcount ratios (Series A benchmarks):**
```
Sales (S&M): 20-30% of headcount
Engineering/Product (R&D): 40-50% of headcount
Customer Success: 15-20% of headcount
G&A: 10-15% of headcount
```
#### COGS Model
Gross margin is the most important long-term indicator of business quality.
**COGS for SaaS:**
```
1. Hosting / Infrastructure (AWS, GCP, Azure)
- Scale with customer count or usage
- Should be 5-15% of ARR for mature SaaS
- If > 20%: infrastructure optimization needed
2. Customer Success headcount
- Ratio: 1 CSM per $1M-$3M ARR (varies by segment)
- SMB: 1 CSM per $500K ARR (high-touch required)
- Enterprise: 1 CSM per $2-5M ARR (strategic accounts)
3. Third-party licensing / APIs
- Per-customer or usage-based pass-through costs
- Critical to model at scale (margin killer if not tracked)
4. Payment processing
- 2.2-2.9% of revenue for Stripe/Braintree
- Can negotiate to 1.8-2.2% at scale (> $5M ARR)
```
**Gross Margin targets:**
```
SaaS: > 65% acceptable, > 75% good, > 80% exceptional
Marketplace: 50-70%
Hardware + software: 40-60%
Services + software: 30-50%
```
**If gross margin < 65%:**
- Infrastructure cost optimization (rightsizing, reserved instances)
- CS headcount review (automation, pooled CSMs)
- Pricing model review (usage-based pricing if cost is usage-driven)
- Third-party cost renegotiation
#### Opex Model
```
Sales & Marketing:
- AE/SDR/SE salaries + OTE (on-target earnings)
- Marketing programs (demand gen budget)
- Tools and technology (CRM, SEO, ads platforms)
- Events and travel
- Benchmark: 40-60% of revenue at growth stage, targeting < 30% at scale
Research & Development:
- Engineering salaries
- Product management
- Design
- Technical infrastructure for development
- Benchmark: 20-35% of revenue
General & Administrative:
- Finance, legal, HR, admin
- Office costs
- SaaS tools / software licenses
- D&O insurance
- Benchmark: 8-15% (target < 10% at scale)
```
### Financial Model Do's and Don'ts
| Do | Don't |
|----|-------|
| Build assumptions tab with all inputs | Hardcode numbers in formulas |
| Model monthly (not quarterly) at early stage | Use annual model for first 3 years |
| Start with headcount plan, build costs from it | Guess at expense line items |
| Show model to actual customers or users | Show model to investors before internal stress-test |
| Version your model | Overwrite old versions |
| Reconcile cash flow to P&L monthly | Trust P&L without cash flow model |
| Include a sensitivity table | Present single-scenario forecast |
---
## 2. Three-Statement Model for Startups
### Why All Three Matter
The P&L tells you if you're profitable. The cash flow statement tells you if you're alive. The balance sheet tells you if you're solvent.
Startups that only track P&L miss the gap between revenue recognition and cash collection.
### P&L Structure
```
Q1 Q2 Q3 Q4 FY
Revenue
Subscription ARR $400K $520K $680K $840K $2,440K
Professional Svcs $40K $50K $60K $65K $215K
Total Revenue $440K $570K $740K $905K $2,655K
COGS
Infrastructure $35K $42K $52K $62K $191K
CS Headcount $75K $75K $100K $100K $350K
3rd Party Licensing $15K $18K $22K $28K $83K
Total COGS $125K $135K $174K $190K $624K
Gross Profit $315K $435K $566K $715K $2,031K
Gross Margin 71.6% 76.3% 76.5% 79.0% 76.5%
Operating Expenses
Sales & Marketing $380K $420K $480K $520K $1,800K
Research & Dev $320K $340K $380K $400K $1,440K
General & Admin $120K $130K $140K $150K $540K
Total Opex $820K $890K $1000K $1070K $3,780K
EBITDA ($505K) ($455K) ($434K) ($355K) ($1,749K)
EBITDA Margin (114.8%)(79.8%) (58.6%) (39.2%) (65.9%)
```
### Cash Flow Statement
```
Q1 Q2 Q3 Q4
Operating Activities
Net Income ($510K) ($460K) ($440K) ($360K)
Add: D&A $8K $8K $8K $10K
Working Capital Changes:
AR increase ($45K) ($50K) ($60K) ($55K)
AP increase $20K $15K $20K $15K
Deferred Rev change $80K $60K $80K $90K
Operating Cash Flow ($447K) ($427K) ($392K) ($300K)
Investing Activities
Capex ($15K) ($8K) ($10K) ($12K)
Free Cash Flow ($462K) ($435K) ($402K) ($312K)
Financing Activities
None $0 $0 $0 $0
Net Change in Cash ($462K) ($435K) ($402K) ($312K)
Beginning Cash $3,500K $3,038K $2,603K $2,201K
Ending Cash $3,038K $2,603K $2,201K $1,889K
Runway (months) 13.1 12.1 10.9 10.1
```
**Key insight from this model:**
The deferred revenue offset (customers paying annually upfront) is reducing cash burn by ~$80-90K/quarter versus a pure monthly billing model. This is the CFO's lever — push for annual billing.
### Balance Sheet: The Startup Version
At early stage, track these specifically:
```
Assets:
Cash: Your lifeline. Monitor daily.
Accounts Receivable: What customers owe you. Age it monthly.
Prepaid Expenses: Software licenses, insurance paid upfront.
Liabilities:
Accounts Payable: What you owe vendors. Maximize terms.
Accrued Liabilities: Salaries owed, commissions earned but not paid.
Deferred Revenue: Customer prepayments. Liability until service delivered, but cash is yours.
Debt/Convertible Notes: Face value + interest accrual.
Equity:
Common Stock: Founder shares
Preferred Stock: Investor shares
APIC: Additional paid-in capital
Accumulated Deficit: Your running losses (expected for startups)
```
---
## 3. SaaS Metrics That Matter
### The Hierarchy of SaaS Metrics
```
Tier 1 (existential): ARR, Runway, Net Dollar Retention
Tier 2 (strategic): Gross Margin, Burn Multiple, LTV:CAC
Tier 3 (operational): CAC Payback, Churn Rate, ACV
Tier 4 (diagnostic): Logo Churn vs Revenue Churn, Expansion Rate, NPS
```
Never report Tier 4 metrics to your board if Tier 1 metrics are off-track.
### Core Metric Definitions
**ARR (Annual Recurring Revenue):**
```
ARR = Sum of all active annual contract values (normalized to annual)
What it is NOT: bookings, billings, or TCV
When to use MRR: Companies with mostly monthly contracts
When to use ARR: Companies with majority annual contracts
```
**Net Dollar Retention (NDR / NRR):**
```
NDR = (Beginning MRR + Expansion MRR - Churned MRR - Contraction MRR)
/ Beginning MRR × 100
The benchmark everyone quotes: 100% means existing customers are flat.
> 100% means existing customers grow revenue on their own.
World-class (Snowflake, Datadog): 130%+
Why it matters: NDR > 100% means revenue growth even if you sign zero new customers.
At NDR = 120% and $5M ARR: you will reach $7M ARR in 24 months without a single new sale.
```
**Gross Revenue Retention (GRR):**
```
GRR = (Beginning MRR - Churned MRR - Contraction MRR) / Beginning MRR × 100
GRR measures the floor of your retention (ignoring expansion).
GRR is always ≤ NDR.
Target: > 85% for SMB SaaS, > 90% for mid-market, > 95% for enterprise.
```
**Logo Churn vs Revenue Churn:**
```
Logo churn: % of customers who cancel (ignores size)
Revenue churn: % of ARR that cancels (accounts for size)
Why the distinction matters:
You could have 10% logo churn but 3% revenue churn (churning small customers)
Or 5% logo churn but 12% revenue churn (churning large customers) — much worse
Report both. If they diverge significantly, investigate immediately.
```
**ACV (Annual Contract Value):**
```
ACV = Total contract value / contract term in years
Not to be confused with ARR (which only counts recurring, not one-time fees)
Rising ACV: You're moving upmarket (good for efficiency, check if ICP is changing)
Falling ACV: You're moving downmarket (check burn multiple — may not be economic)
```
**Rule of 40:**
```
Rule of 40 = Revenue Growth Rate % + EBITDA Margin %
Target: > 40%
Example: 60% growth + (-15%) EBITDA margin = 45. Passing.
Example: 20% growth + 5% EBITDA margin = 25. Failing at growth stage.
At early stage (< $5M ARR): Rule of 40 doesn't apply. Growth is the only metric.
At growth stage ($5-20M ARR): Starting to matter.
At scale ($20M+ ARR): Board and investors will hold you to this.
```
---
## 4. FP&A for Startups: What to Measure When
### Metrics by Stage
**Pre-seed / Seed (< $1M ARR):**
```
Focus on: Cash, pipeline, customer conversations
Measure: Monthly cash burn, weeks of runway, NPS / customer satisfaction
Don't obsess over: EBITDA margin, gross margin (too early)
Frequency: Weekly cash check, monthly everything else
```
**Series A ($1-5M ARR):**
```
Focus on: Repeatable sales, unit economics
Measure: MRR growth, LTV:CAC, CAC payback by channel, gross margin
Don't obsess over: Profitability, G&A efficiency
Build now: Monthly financial close (< 5 business days), basic FP&A model
Frequency: Monthly board pack, weekly leadership metrics
```
**Series B ($5-20M ARR):**
```
Focus on: Scalable go-to-market, operational efficiency
Measure: NDR, burn multiple, revenue per FTE, OKR attainment
Start building: Budget vs actuals, department-level P&L
Build now: Finance team (first financial controller), ERP or NetSuite
Frequency: Monthly board pack + quarterly deep dive
```
**Series C+ ($20M+ ARR):**
```
Focus on: Path to profitability, market leadership
Measure: Rule of 40, free cash flow, CAC efficiency by segment
Must have: FP&A team, full three-statement model, 5-year plan
Frequency: Monthly financial close (< 3 business days), quarterly earnings prep
```
### Reporting Cadence
**Weekly (CFO + leadership):**
- Cash balance (CFO checks daily, reports weekly)
- Pipeline / sales metrics (if in a sales-led motion)
- Any metric that changed dramatically vs. prior week
**Monthly (board + leadership):**
- Full financial dashboard (ARR, gross margin, burn, runway)
- Budget vs actual with explanations for > 10% variances
- Unit economics update
- Headcount change summary
**Quarterly (board + investors):**
- Full three-statement model vs budget
- Cohort analysis update
- Scenario planning review and trigger assessment
- Next quarter outlook
---
## 5. Budget vs Actual Analysis Framework
### The Purpose of BvA
Budget vs actual is not about being right. It's about understanding *why* you were wrong, so you can make better decisions.
The CFO who reports "we missed budget by 15%" without explanation is failing. The CFO who says "we missed budget by 15% because enterprise deals took 30 more days to close than modeled — here's what we're doing about it" is doing their job.
### BvA Template
```
Category Budget Actual $ Var % Var Explanation
-------------------------------------------------------------------
ARR $2,400K $2,280K ($120K) (5%) 2 enterprise deals slipped to Q1
New ARR $400K $350K ($50K) (13%) Above
Expansion ARR $120K $140K $20K 17% PLG motion outperforming
Churn ($60K) ($80K) ($20K) (33%) 2 unexpected SMB churns (now fixed)
Gross Margin 75.0% 73.2% -1.8% n/a Infrastructure over-provisioned
S&M Spend $820K $840K ($20K) (2%) Within tolerance
R&D Spend $680K $710K ($30K) (4%) Backfill hire started month early
G&A Spend $140K $148K ($8K) (6%) Legal fees for new customer contract
Cash Burn (net) $580K $648K ($68K) (12%) Driven by ARR shortfall + costs
Runway (mo) 14.5 13.0 (1.5) n/a Tracking; fundraise target unchanged
```
### Variance Thresholds
```
< ±5%: Note in appendix, no explanation needed in main pack
5-10%: One-line explanation required
> 10%: Full paragraph: what happened, why, what changes
> 20%: Board conversation required (model assumption was wrong, or unexpected event)
```
### Forecasting vs Budgeting
**Budget:** Set at start of year. Fixed expectation. Updated quarterly.
**Forecast:** Rolling 3-month outlook. Updated monthly. Should converge with budget over time.
```
Common mistake: Treating forecast as wishful thinking ("what we hope happens")
Correct approach: Forecast is your best current estimate given all known information.
If forecast diverges from budget by > 15%, the budget is wrong.
Reforecast and communicate to board.
```
**Rolling forecast (recommended for startups):**
```
Always have a 12-month forward model.
Update it monthly with actuals replacing the first month.
The forecast should always reflect your current operational reality, not your hope.
```
---
## Key Formulas Reference
```python
# ARR and growth
ARR_growth_yoy = (ending_ARR - beginning_ARR) / beginning_ARR
# Net Dollar Retention
NDR = (beginning_MRR + expansion_MRR - churn_MRR - contraction_MRR) / beginning_MRR
# Burn Multiple
burn_multiple = net_cash_burn / net_new_ARR
# Rule of 40
rule_of_40 = revenue_growth_pct + ebitda_margin_pct
# LTV (SaaS)
LTV = (ARPA * gross_margin_pct) / monthly_churn_rate
# CAC Payback (months)
cac_payback = CAC / (ARPA * gross_margin_pct)
# Magic Number (sales efficiency)
magic_number = (net_new_ARR * 4) / prior_quarter_S_and_M_spend
# Gross margin
gross_margin = (revenue - COGS) / revenue
# Quick Ratio (growth efficiency)
quick_ratio = (new_MRR + expansion_MRR) / (churned_MRR + contraction_MRR)
# Target: > 4 for high-growth SaaS
```
FILE:references/fundraising_playbook.md
# Fundraising Playbook
From timing to close. What investors actually look for, how valuation works, and the term sheet clauses that matter.
---
## 1. When to Raise
**Optimal timing:**
```
Target: 18-24 months runway post-close
Minimum: 12 months runway post-close (leaves no buffer for slip)
Start process when: 9-12 months runway remaining
→ 3-6 months for process (typically 4-5 months for Series A/B)
→ Leaves 3-6 months buffer if process drags
Never start when: < 6 months runway
→ You're negotiating from desperation
→ Investors can smell it
→ Terms get worse, or you don't close at all
```
**Rule:** Your leverage is maximum when you don't *need* to raise. Raise from a position of momentum, not necessity.
---
## 2. What Investors Look For at Each Stage
### Pre-seed
- Team (are these people credible for this problem?)
- Problem clarity (is the problem real and meaningful?)
- Early signal (any customers paying, waitlist, prototype)
- Market size (worth building a VC-scale company?)
**Typical ask:** $500K–$2M | **Typical valuation:** $3M–$10M pre-money
### Seed
- Product-market signal (customers using and paying)
- Founding team with domain expertise
- ARR: $100K–$1M (or strong usage for PLG)
- Clear hypothesis for what Series A looks like
**Typical ask:** $2M–$5M | **Typical valuation:** $8M–$20M pre-money
### Series A
Investors are buying a *repeatable sales motion*. Not just customers — a machine.
**What they need to see:**
- ARR: $1M–$5M growing > 100% YoY
- LTV:CAC > 2.5x (and improving)
- Net Dollar Retention > 100%
- CAC Payback < 18 months
- Gross margin > 65%
- At least 5-10 reference customers (not just lighthouse)
- Sales motion that converts without the founder closing every deal
**Typical ask:** $8M–$15M | **Typical valuation:** $25M–$60M pre-money
### Series B
Investors are buying *scalable go-to-market*. Can you pour fuel on the fire?
**What they need to see:**
- ARR: $5M–$20M growing > 100% YoY
- LTV:CAC > 3x, CAC Payback < 18 months
- Sales capacity model (hiring plan → pipeline → revenue)
- NDR > 110% (expansion motion working)
- Some proof of market expansion (new segments, geographies, use cases)
- Path to category leadership
**Typical ask:** $15M–$40M | **Typical valuation:** $60M–$200M pre-money
### Series C and Beyond
Investors are buying *market leadership* and *path to profitability*.
**What they need to see:**
- ARR: $20M+ (often $30-50M for credible Series C)
- Rule of 40 > 40 (or credible path)
- Gross margin > 70%
- NDR > 115%
- Evidence of market leadership (brand, win rates, analyst mentions)
- Clear path to $100M+ ARR
---
## 3. Valuation Methods
### Revenue Multiples (Primary Method for SaaS)
```
Pre-money Valuation = ARR × Revenue Multiple
Revenue multiple benchmarks (2024-2025):
> 100% YoY growth: 8x–15x ARR
50-100% YoY growth: 4x–8x ARR
20-50% YoY growth: 2x–4x ARR
< 20% YoY growth: 1x–2x ARR
Adjustments:
NDR > 120%: +1x–2x premium
Gross margin > 75%: +0.5x–1x premium
Burn multiple < 1x: +0.5x–1x premium
Capital efficient: Investors pay up for efficiency
Declining growth: Compress multiple aggressively
```
### The Investor's Math (Know This)
Every VC has a required return. Work backwards from their constraints:
```
Investor targets: 3x fund return
Fund size: $200M, check size: $15M (initial), $25M (with follow-on)
Ownership at exit needed: 15%
At 15% ownership: needs $25M / 15% = $167M post-money valuation
Exit needed to return 3x on that check: $25M × 10 = $250M company value
(10x because most deals fail, winners must carry the fund)
Implication: If you think you'll exit for $150M, that VC will pass or price you accordingly.
```
This is why Series A investors rarely lead rounds where they can't see a $300M+ exit path. It's not about your business being bad — it's about fund math.
### Comparable Company Analysis
For later stages (Series B+):
```
1. Find 5-10 comparable public SaaS companies
2. Calculate their EV/NTM Revenue multiples (use latest data)
3. Apply a private market discount (typically 20-40% vs public comps)
4. Adjust for your growth rate relative to comps
Example (2024):
Public SaaS comps: 6x NTM Revenue (median)
Private discount: 30%
Adjusted: ~4.2x
Your NTM Revenue: $8M
Implied valuation: ~$33M pre-money
```
### DCF (Late Stage Only)
DCF is unreliable for early-stage startups (terminal value dominates, growth rate assumptions are fantasy). Use it as a sanity check at Series C+, not as the primary valuation method.
---
## 4. Term Sheet Breakdown
### Liquidation Preference (Most Important Economic Term)
This determines who gets paid first in an exit — and how much.
```
1x Non-Participating Preferred (BEST for founders):
Investor gets 1x money back OR converts to common (their choice).
At acquisition: investor takes larger of {1x invested} or {% ownership × proceeds}
Example: $10M invested, exits at $100M, owns 20%
Option A: $10M (1x)
Option B: $20M (20% of $100M)
Investor takes $20M. Founders split $80M.
1x Participating Preferred (WORSE for founders):
Investor gets 1x money back AND participates in remaining proceeds.
Example: same scenario
$10M (1x) + 20% of remaining $90M = $10M + $18M = $28M
Founders split $72M instead of $80M
Cost to founders: $8M (10% of exit value)
2x Participating (RED FLAG):
Investor gets 2x back AND participates.
Only accept under duress. Push hard against this.
Full Ratchet Anti-Dilution (AVOID):
Down-round triggers full repricing of investor shares to new (lower) price.
Founders get massively diluted. Never accept if alternatives exist.
```
### Anti-Dilution Protection
```
Broad-based weighted average (standard):
Adjusts investor conversion price based on all dilutive securities.
Most founder-friendly anti-dilution. Accept this.
Narrow-based weighted average (slightly worse):
Same mechanism but uses smaller denominator.
Gives investors slightly more protection. Usually acceptable.
Full ratchet (avoid):
Price drops to whatever the new round prices at.
Devastating in down rounds. Fight this.
```
### Pro-Rata Rights
```
Standard pro-rata: Investor can maintain their % ownership in future rounds.
Reasonable. Accept for major investors.
Super pro-rata: Investor can increase their % in future rounds.
Caps your ability to bring in new lead investors.
Avoid unless the investor is exceptional and you want them in future rounds.
Major investor threshold: Typically investors with > $500K–$1M check get pro-rata.
Don't give pro-rata to every small check — clogs future rounds.
```
### Board Composition
```
Seed (3 members): 2 founders, 1 lead investor
Series A (5 members): 2 founders, 2 investors, 1 independent
Series B (5-7 seats): Watch for investor majority — negotiate hard
Rule: Founders should retain majority through Series A.
Independent director should be your choice, not investor's.
Never accept investor majority before Series C.
Board observer rights: Common for smaller investors. No vote but present in meetings.
Limit to 1-2 observers or meetings become unwieldy.
```
### Other Terms That Matter
```
Drag-along: Majority can force minority shareholders to vote for acquisition.
Standard and reasonable. Check what threshold triggers drag.
Information rights: Investors get financial statements.
Standard. Monthly for major investors, quarterly for others.
Redemption rights: Investors can force buyback after X years.
Push to remove or add carve-outs for insufficient funds.
No-shop clause: You can't shop the term sheet to other investors.
Standard (14-30 days). Reasonable.
Exclusivity: Stronger version of no-shop. Sometimes includes no other fundraise discussions.
Acceptable for 30 days; push back on > 45 days.
```
---
## 5. Cap Table Management
### Dilution Planning Model
Run this before every round. Know your number before walking into any negotiation.
```
Pre-Seed Post-Seed Post-A Post-B Post-C
Founder A 45.0% 36.0% 26.5% 21.2% 18.7%
Founder B 45.0% 36.0% 26.5% 21.2% 18.7%
Angel 1 5.0% 4.0% 2.9% 2.4% 2.1%
Angel 2 5.0% 4.0% 2.9% 2.4% 2.1%
Seed Fund - 12.0% 8.8% 7.1% 6.2%
Option Pool - 8.0% 12.0% 10.0% 8.0%
Series A - - 20.4% 16.3% 14.4%
Series B - - - 19.5% 17.2%
Series C - - - - 12.6%
Round size / pre-money:
Pre-Seed: $500K / $9M pre = 5% dilution
Seed: $2M / $8M pre = 20% dilution (includes 8% pool)
Series A: $10M / $38M pre = 20.8% dilution (pool refresh to 12%)
Series B: $20M / $80M pre = 20% dilution
Series C: $30M / $170M pre = 15% dilution
```
**Option pool shuffle:** Investors often require you to create/expand the option pool *before* the round closes, which dilutes existing shareholders (not the incoming investor). Model this explicitly — a 20% round with a 5% pool expansion is really 24%+ dilution to founders.
### Cap Table Hygiene
```
Tools: Carta, Pulley, Capshare (all acceptable)
Never: Track cap table in a spreadsheet past seed stage. Errors compound.
Keep it clean:
- Repurchase departed co-founder shares immediately (don't let unvested shares linger)
- Convert SAFEs to equity cleanly at each priced round
- Document every grant with a board resolution
- Cliff + vesting for ALL employees and founders (standard: 1-year cliff, 4-year vest)
- 409A valuation required before every option grant (IRS requirement)
```
---
## 6. Data Room Preparation
### Core Documents (Required)
```
Financial:
□ 3 years historical financials (or all history if < 3 years)
□ Monthly P&L and cash flow (last 24 months)
□ Current financial model (18-24 months forward)
□ Budget vs actual (last 4 quarters)
□ Cap table (fully diluted, with all SAFEs/convertibles modeled)
□ Bank statements (last 3-6 months)
Legal:
□ Certificate of incorporation + all amendments
□ All prior financing documents (SAFEs, convertible notes, stock purchase agreements)
□ Cap table (Carta/Pulley export)
□ IP assignment agreements (all founders and employees)
□ Material contracts (top 10 customers, key vendors)
□ Employee list (titles, start dates, salaries, equity grants)
Product & Business:
□ Product demo / walkthrough video
□ Architecture overview (for technical investors)
□ Customer case studies (3-5 named references)
□ NPS / CSAT data
□ Competitive landscape analysis
Metrics:
□ MRR/ARR by month (all history)
□ Cohort retention chart
□ CAC by channel
□ LTV by cohort
□ NPS trend
```
### What Investors Actually Check First
In order of typical priority during due diligence:
1. **Cap table** — Is it clean? Any concerning structures?
2. **Cohort retention** — Is churn improving or deteriorating?
3. **Revenue quality** — What % is recurring? Any one-time or non-recurring?
4. **Top 10 customers** — Concentration risk? Any logos at risk?
5. **Bank statements** — Does cash match what was reported?
6. **IP assignments** — Does the company own its IP? (Founders who didn't assign IP kill deals)
### Red Flags That Kill Deals
- Missing IP assignment agreements for founders (most common deal killer at early stage)
- Cap table with > 20 angels/small investors (messy, hard to get consent for future rounds)
- Customer concentration > 30% in single customer without explanation
- Revenue recognition issues (booking ARR on contracts that allow easy cancellation)
- Cohort data that gets worse in later cohorts
- Bank balance doesn't match reported cash position
---
## 7. Investor Communication Cadence
### During Fundraise
```
Week 1-2: Warm intro sourcing, LP/network mapping
Week 3-6: First meetings (aim for 20-30 first meetings)
Week 7-10: Partner meetings, deep dives, due diligence
Week 11-14: Term sheets, negotiation
Week 15-18: Legal, closing
```
**Parallel process is essential.** Never negotiate with one investor at a time. Competition is your leverage.
### Post-Close: Investor Updates
Monthly investor update (send within 10 days of month-end):
```
Subject: [Company] Monthly Update — [Month Year]
Highlights (3 bullets max):
• [Biggest win]
• [Biggest learning/challenge]
• [What we're focused on next month]
Metrics:
ARR: $X (+X% MoM)
Net new ARR: $X
Gross margin: X%
Cash: $X (X months runway)
Headcount: X
Asks (be specific):
• Looking for intro to [persona/company] for [specific reason]
• Need advisor with experience in [specific area]
• [Other concrete ask]
```
**Why this matters:** Investors who are informed and engaged are better positioned to help when you need it. The investor who hasn't heard from you in 6 months is less likely to write a bridge check or make a warm intro when you ask.
---
## Key Formulas
```python
# Post-money valuation
post_money = pre_money + investment_amount
# Investor ownership %
ownership_pct = investment_amount / post_money
# Dilution to existing shareholders
dilution = investment_amount / post_money # as a fraction
# New shares issued
new_shares = (investment_amount / post_money) * total_post_shares
# equivalent: new_shares = pre_money_shares * (investment_amount / pre_money)
# Option pool expansion impact (pool shuffle)
# Creating X% option pool pre-close dilutes founders:
pool_shares_needed = target_pct * (pre_shares + new_round_shares + pool_shares_needed)
# Solve: pool_shares_needed = target_pct * (pre_shares + new_round_shares) / (1 - target_pct)
# LTV:CAC ratio
ltv_cac = ltv / cac # target: > 3x
# CAC payback (months)
payback_months = cac / (arpa * gross_margin_pct)
```
FILE:scripts/burn_rate_calculator.py
#!/usr/bin/env python3
"""
Burn Rate & Runway Calculator
==============================
Models startup runway across base/bull/bear scenarios, incorporating
a hiring plan and revenue trajectory. Outputs months of runway,
cash-out dates, and decision trigger points.
Usage:
python burn_rate_calculator.py
python burn_rate_calculator.py --csv # export to CSV
Stdlib only. No dependencies.
"""
import argparse
import csv
import io
import sys
from dataclasses import dataclass, field
from datetime import date, timedelta
from typing import Optional
# ---------------------------------------------------------------------------
# Data structures
# ---------------------------------------------------------------------------
@dataclass
class HiringEntry:
"""A planned hire."""
month: int # months from model start (1-indexed)
role: str
department: str # "sales", "engineering", "cs", "ga"
annual_salary: float
benefits_pct: float = 0.22 # benefits as % of salary
recruiting_cost: float = 0.0 # one-time recruiting fee
@dataclass
class RevenueEntry:
"""Monthly revenue data point (historical or projected)."""
month: int
mrr: float # monthly recurring revenue
one_time: float = 0.0
@dataclass
class ModelConfig:
"""Master configuration for a runway scenario."""
name: str
starting_cash: float
starting_mrr: float
starting_headcount: int
avg_loaded_salary: float # average fully-loaded salary per current employee
base_non_headcount_opex: float # monthly non-headcount costs (infra, tools, etc.)
gross_margin_pct: float # 0.0–1.0
mrr_growth_rate: float # monthly MoM growth rate, 0.0–1.0
hiring_plan: list[HiringEntry] = field(default_factory=list)
model_months: int = 24
start_date: Optional[date] = None
@dataclass
class MonthResult:
"""Single month output."""
month: int
label: str # e.g. "Month 1 (Apr 2025)"
mrr: float
gross_profit: float
headcount: int
headcount_cost: float # total loaded headcount cost this month
other_opex: float
gross_burn: float
net_burn: float
cash_start: float
cash_end: float
runway_months: float # projected runway from this month
cumulative_new_arr: float # for burn multiple
# ---------------------------------------------------------------------------
# Core calculator
# ---------------------------------------------------------------------------
class RunwayCalculator:
def __init__(self, config: ModelConfig):
self.cfg = config
def run(self) -> list[MonthResult]:
cfg = self.cfg
results = []
# Build headcount schedule: month -> list of new hires starting that month
hire_by_month: dict[int, list[HiringEntry]] = {}
for h in cfg.hiring_plan:
hire_by_month.setdefault(h.month, []).append(h)
# Track existing employees
active_employees: list[dict] = []
for _ in range(cfg.starting_headcount):
active_employees.append({
"monthly_loaded": cfg.avg_loaded_salary / 12 * 1.0,
"start_month": 0,
})
cash = cfg.starting_cash
mrr = cfg.starting_mrr
cumulative_new_arr = 0.0
starting_mrr = cfg.starting_mrr
for m in range(1, cfg.model_months + 1):
# Process new hires this month
one_time_recruiting = 0.0
if m in hire_by_month:
for hire in hire_by_month[m]:
monthly_loaded = (
hire.annual_salary * (1 + hire.benefits_pct) / 12
)
active_employees.append({
"monthly_loaded": monthly_loaded,
"start_month": m,
})
one_time_recruiting += hire.recruiting_cost
# Revenue this month
mrr = mrr * (1 + cfg.mrr_growth_rate)
gross_profit = mrr * cfg.gross_margin_pct
# Headcount cost
headcount_cost = sum(e["monthly_loaded"] for e in active_employees)
headcount_cost += one_time_recruiting
# Other opex (infra, SaaS tools, office, etc.)
other_opex = cfg.base_non_headcount_opex
# Burn
gross_burn = headcount_cost + other_opex
net_burn = gross_burn - gross_profit
# Cash
cash_start = cash
cash = cash - net_burn
cash_end = cash
# Projected runway from this month (using current net burn rate)
runway = cash_end / net_burn if net_burn > 0 else float("inf")
# Cumulative new ARR (for burn multiple calc)
new_mrr_added = mrr - starting_mrr if m == 1 else mrr - results[-1].mrr
cumulative_new_arr += new_mrr_added * 12
# Label
if cfg.start_date:
month_date = date(
cfg.start_date.year,
cfg.start_date.month,
1,
) + timedelta(days=32 * (m - 1))
month_date = month_date.replace(day=1)
label = f"Month {m:02d} ({month_date.strftime('%b %Y')})"
else:
label = f"Month {m:02d}"
results.append(MonthResult(
month=m,
label=label,
mrr=mrr,
gross_profit=gross_profit,
headcount=len(active_employees),
headcount_cost=headcount_cost,
other_opex=other_opex,
gross_burn=gross_burn,
net_burn=net_burn,
cash_start=cash_start,
cash_end=cash_end,
runway_months=runway,
cumulative_new_arr=cumulative_new_arr,
))
# Stop if cash runs out
if cash_end <= 0:
break
return results
def cash_out_date(self, results: list[MonthResult]) -> Optional[str]:
"""Return the label of the month cash runs out, or None if model survives."""
for r in results:
if r.cash_end <= 0:
return r.label
return None
def burn_multiple(self, results: list[MonthResult]) -> float:
"""Burn multiple = total net burn / total net new ARR over model period."""
total_net_burn = sum(r.net_burn for r in results if r.net_burn > 0)
first_mrr = results[0].mrr / (1 + self.cfg.mrr_growth_rate) # starting mrr
total_new_arr = (results[-1].mrr - first_mrr) * 12
if total_new_arr <= 0:
return float("inf")
return total_net_burn / total_new_arr
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt_k(value: float) -> str:
"""Format as $Xk or $X.XM."""
if abs(value) >= 1_000_000:
return f".2fM"
if abs(value) >= 1_000:
return f".0fK"
return f".0f"
def print_summary(name: str, results: list[MonthResult], calc: RunwayCalculator) -> None:
cash_out = calc.cash_out_date(results)
bm = calc.burn_multiple(results)
last = results[-1]
first = results[0]
print(f"\n{'='*60}")
print(f" SCENARIO: {name}")
print(f"{'='*60}")
print(f" Months modeled: {len(results)}")
print(f" Cash out: {cash_out or 'Does not run out in model period'}")
print(f" Ending cash: {fmt_k(last.cash_end)}")
print(f" Final runway: {last.runway_months:.1f} months")
print(f" Starting MRR: {fmt_k(first.mrr)}")
print(f" Ending MRR: {fmt_k(last.mrr)}")
print(f" Ending headcount: {last.headcount}")
print(f" Burn multiple: {bm:.2f}x")
print(f" Avg net burn: {fmt_k(sum(r.net_burn for r in results)/len(results))}/mo")
# Decision triggers
print(f"\n Decision Triggers:")
triggers = {9: "⚠️ START FUNDRAISE", 6: "🔴 COST REDUCTION PLAN", 4: "🚨 EXECUTE CUTS / BRIDGE"}
shown = set()
for r in results:
for threshold, label in triggers.items():
if r.runway_months <= threshold and threshold not in shown:
print(f" {r.label}: {label} (runway = {r.runway_months:.1f} mo)")
shown.add(threshold)
def print_monthly_table(results: list[MonthResult], max_rows: int = 24) -> None:
header = f"{'Month':<22} {'MRR':>10} {'Hdct':>6} {'Net Burn':>12} {'Cash':>12} {'Runway':>8}"
print(f"\n{header}")
print("-" * len(header))
for r in results[:max_rows]:
runway_str = f"{r.runway_months:.1f}mo" if r.runway_months != float("inf") else "∞"
print(
f"{r.label:<22} "
f"{fmt_k(r.mrr):>10} "
f"{r.headcount:>6} "
f"{fmt_k(r.net_burn):>12} "
f"{fmt_k(r.cash_end):>12} "
f"{runway_str:>8}"
)
def export_csv(scenarios: list[tuple[str, list[MonthResult]]]) -> str:
buf = io.StringIO()
writer = csv.writer(buf)
writer.writerow([
"Scenario", "Month", "Label", "MRR", "Gross Profit", "Headcount",
"Headcount Cost", "Other Opex", "Gross Burn", "Net Burn",
"Cash Start", "Cash End", "Runway Months"
])
for name, results in scenarios:
for r in results:
writer.writerow([
name, r.month, r.label,
round(r.mrr, 2), round(r.gross_profit, 2), r.headcount,
round(r.headcount_cost, 2), round(r.other_opex, 2),
round(r.gross_burn, 2), round(r.net_burn, 2),
round(r.cash_start, 2), round(r.cash_end, 2),
round(r.runway_months, 2),
])
return buf.getvalue()
# ---------------------------------------------------------------------------
# Sample data
# ---------------------------------------------------------------------------
def make_sample_configs() -> list[ModelConfig]:
"""
Sample company: Series A SaaS startup
- $3M cash on hand (post Series A)
- $125K MRR (~$1.5M ARR)
- 18 employees, $150K avg salary
- $80K/mo non-headcount opex (infra, tools, office)
- 72% gross margin
"""
common_kwargs = dict(
starting_cash=3_000_000,
starting_mrr=125_000,
starting_headcount=18,
avg_loaded_salary=150_000,
base_non_headcount_opex=80_000,
gross_margin_pct=0.72,
model_months=24,
start_date=date(2025, 1, 1),
)
# Base: 10% MoM growth, moderate hiring
base_hiring = [
HiringEntry(month=2, role="AE #1", department="sales", annual_salary=120_000, recruiting_cost=18_000),
HiringEntry(month=3, role="Senior SWE #1", department="engineering", annual_salary=160_000, recruiting_cost=24_000),
HiringEntry(month=5, role="SDR #1", department="sales", annual_salary=80_000, recruiting_cost=12_000),
HiringEntry(month=6, role="CSM #1", department="cs", annual_salary=90_000, recruiting_cost=13_500),
HiringEntry(month=8, role="AE #2", department="sales", annual_salary=120_000, recruiting_cost=18_000),
HiringEntry(month=9, role="Senior SWE #2", department="engineering", annual_salary=165_000, recruiting_cost=24_750),
HiringEntry(month=12, role="Controller", department="ga", annual_salary=130_000, recruiting_cost=19_500),
HiringEntry(month=14, role="AE #3", department="sales", annual_salary=125_000, recruiting_cost=18_750),
HiringEntry(month=15, role="ML Engineer", department="engineering", annual_salary=175_000, recruiting_cost=26_250),
HiringEntry(month=18, role="AE #4", department="sales", annual_salary=125_000, recruiting_cost=18_750),
]
# Bull: 15% MoM growth, full hiring plan
bull_hiring = base_hiring + [
HiringEntry(month=4, role="Marketing Manager", department="sales", annual_salary=110_000, recruiting_cost=16_500),
HiringEntry(month=7, role="Senior SWE #3", department="engineering", annual_salary=165_000, recruiting_cost=24_750),
HiringEntry(month=10, role="AE #5", department="sales", annual_salary=125_000, recruiting_cost=18_750),
HiringEntry(month=13, role="DevOps Engineer", department="engineering", annual_salary=150_000, recruiting_cost=22_500),
HiringEntry(month=16, role="AE #6", department="sales", annual_salary=125_000, recruiting_cost=18_750),
]
# Bear: 5% MoM growth, hiring freeze after month 3
bear_hiring = [
HiringEntry(month=2, role="AE #1", department="sales", annual_salary=120_000, recruiting_cost=18_000),
HiringEntry(month=3, role="Senior SWE #1", department="engineering", annual_salary=160_000, recruiting_cost=24_000),
]
return [
ModelConfig(name="BULL (15% MoM, full hiring)", mrr_growth_rate=0.15, hiring_plan=bull_hiring, **common_kwargs),
ModelConfig(name="BASE (10% MoM, planned hiring)", mrr_growth_rate=0.10, hiring_plan=base_hiring, **common_kwargs),
ModelConfig(name="BEAR ( 5% MoM, hiring freeze M3+)", mrr_growth_rate=0.05, hiring_plan=bear_hiring, **common_kwargs),
ModelConfig(name="DISTRESS (0% growth, freeze now)", mrr_growth_rate=0.00, hiring_plan=[], **common_kwargs),
]
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="Startup Burn Rate & Runway Calculator")
parser.add_argument("--csv", action="store_true", help="Export full monthly data as CSV to stdout")
parser.add_argument("--scenario", choices=["bull", "base", "bear", "distress", "all"], default="all")
args = parser.parse_args()
configs = make_sample_configs()
if args.scenario != "all":
configs = [c for c in configs if args.scenario.upper() in c.name.upper()]
all_results: list[tuple[str, list[MonthResult]]] = []
print("\n" + "="*60)
print(" BURN RATE & RUNWAY CALCULATOR")
print(" Sample Company: Series A SaaS Startup")
print(" Starting cash: $3M | Starting MRR: $125K | 18 employees")
print("="*60)
for cfg in configs:
calc = RunwayCalculator(cfg)
results = calc.run()
all_results.append((cfg.name, results))
print_summary(cfg.name, results, calc)
print_monthly_table(results)
# Comparison summary
print("\n" + "="*60)
print(" SCENARIO COMPARISON")
print("="*60)
print(f" {'Scenario':<40} {'Runway':>8} {'Cash Out':<30} {'Burn Mult':>10}")
print(" " + "-"*88)
for cfg, (name, results) in zip(configs, all_results):
calc = RunwayCalculator(cfg)
cash_out = calc.cash_out_date(results) or "Survives model period"
bm = calc.burn_multiple(results)
final_runway = results[-1].runway_months
runway_str = f"{final_runway:.1f}mo" if final_runway != float("inf") else "∞"
bm_str = f"{bm:.2f}x" if bm != float("inf") else "∞"
print(f" {name:<40} {runway_str:>8} {cash_out:<30} {bm_str:>10}")
print("\n Decision Trigger Reference:")
print(" 9 months runway → Start fundraise process")
print(" 6 months runway → Begin cost reduction planning")
print(" 4 months runway → Execute cuts; explore bridge financing")
print(" 3 months runway → Emergency plan only")
if args.csv:
print("\n\n--- CSV EXPORT ---\n")
sys.stdout.write(export_csv(all_results))
if __name__ == "__main__":
main()
FILE:scripts/fundraising_model.py
#!/usr/bin/env python3
"""
Fundraising Model
==================
Cap table management, dilution modeling, and multi-round scenario planning.
Know exactly what you're giving up before you walk into any negotiation.
Covers:
- Cap table state at each round
- Dilution per shareholder per round
- Option pool shuffle impact
- Multi-round projections (Seed → A → B → C)
- Return scenarios at different exit valuations
Usage:
python fundraising_model.py
python fundraising_model.py --exit 150 # model at $150M exit
python fundraising_model.py --csv
Stdlib only. No dependencies.
"""
import argparse
import csv
import io
import sys
from dataclasses import dataclass, field
from typing import Optional
# ---------------------------------------------------------------------------
# Data structures
# ---------------------------------------------------------------------------
@dataclass
class Shareholder:
"""A shareholder in the cap table."""
name: str
share_class: str # "common", "preferred", "option"
shares: float
invested: float = 0.0 # total cash invested
is_option_pool: bool = False
@dataclass
class RoundConfig:
"""Configuration for a financing round."""
name: str # e.g. "Series A"
pre_money_valuation: float
investment_amount: float
new_option_pool_pct: float = 0.0 # % of POST-money to allocate to new options
option_pool_pre_round: bool = True # True = pool created before round (dilutes founders)
lead_investor_name: str = "New Investor"
share_price_override: Optional[float] = None # if None, computed from valuation
@dataclass
class CapTableEntry:
"""A row in the cap table at a point in time."""
name: str
share_class: str
shares: float
pct_ownership: float
invested: float
is_option_pool: bool = False
@dataclass
class RoundResult:
"""Snapshot of cap table after a round closes."""
round_name: str
pre_money_valuation: float
investment_amount: float
post_money_valuation: float
price_per_share: float
new_shares_issued: float
option_pool_shares_created: float
total_shares: float
cap_table: list[CapTableEntry]
@dataclass
class ExitAnalysis:
"""Proceeds to each shareholder at an exit."""
exit_valuation: float
shareholder: str
shares: float
ownership_pct: float
proceeds_common: float # if all preferred converts to common
invested: float
moic: float # multiple on invested capital (for investors)
# ---------------------------------------------------------------------------
# Core cap table engine
# ---------------------------------------------------------------------------
class CapTable:
"""Manages a cap table through multiple rounds."""
def __init__(self):
self.shareholders: list[Shareholder] = []
self._total_shares: float = 0.0
def add_shareholder(self, sh: Shareholder) -> None:
self.shareholders.append(sh)
self._total_shares += sh.shares
def total_shares(self) -> float:
return sum(s.shares for s in self.shareholders)
def snapshot(self, label: str = "") -> list[CapTableEntry]:
total = self.total_shares()
return [
CapTableEntry(
name=s.name,
share_class=s.share_class,
shares=s.shares,
pct_ownership=s.shares / total if total > 0 else 0,
invested=s.invested,
is_option_pool=s.is_option_pool,
)
for s in self.shareholders
]
def execute_round(self, config: RoundConfig) -> RoundResult:
"""
Execute a financing round:
1. (Optional) Create option pool pre-round (dilutes existing shareholders)
2. Issue new shares to investor at round price
Returns a RoundResult with full cap table snapshot.
"""
current_total = self.total_shares()
# Step 1: Option pool shuffle (if pre-round)
option_pool_shares_created = 0.0
if config.new_option_pool_pct > 0 and config.option_pool_pre_round:
# Target: post-round option pool = new_option_pool_pct of total post-money shares
# Solve: pool_shares / (current_total + pool_shares + new_investor_shares) = target_pct
# This requires iteration because new_investor_shares also depends on pool_shares
# Simplification: create pool based on post-round total (slightly approximated)
target_post_round_pct = config.new_option_pool_pct
post_money = config.pre_money_valuation + config.investment_amount
# Estimate shares per dollar (price per share)
price_per_share = config.pre_money_valuation / current_total
new_investor_shares_estimate = config.investment_amount / price_per_share
# Pool shares needed so that pool / total_post = target_pct
total_post_estimate = current_total + new_investor_shares_estimate
pool_shares_needed = (target_post_round_pct * total_post_estimate) / (1 - target_post_round_pct)
# Check if existing pool is sufficient
existing_pool = next(
(s.shares for s in self.shareholders if s.is_option_pool), 0
)
additional_pool_needed = max(0, pool_shares_needed - existing_pool)
if additional_pool_needed > 0:
option_pool_shares_created = additional_pool_needed
# Add to existing pool or create new
pool_sh = next((s for s in self.shareholders if s.is_option_pool), None)
if pool_sh:
pool_sh.shares += additional_pool_needed
else:
self.shareholders.append(Shareholder(
name="Option Pool",
share_class="option",
shares=additional_pool_needed,
is_option_pool=True,
))
# Step 2: Price per share (after pool creation)
current_total_post_pool = self.total_shares()
if config.share_price_override:
price_per_share = config.share_price_override
else:
price_per_share = config.pre_money_valuation / current_total_post_pool
# Step 3: New shares for investor
new_shares = config.investment_amount / price_per_share
# Step 4: Add investor to cap table
self.shareholders.append(Shareholder(
name=config.lead_investor_name,
share_class="preferred",
shares=new_shares,
invested=config.investment_amount,
))
post_money = config.pre_money_valuation + config.investment_amount
total_post = self.total_shares()
return RoundResult(
round_name=config.name,
pre_money_valuation=config.pre_money_valuation,
investment_amount=config.investment_amount,
post_money_valuation=post_money,
price_per_share=price_per_share,
new_shares_issued=new_shares,
option_pool_shares_created=option_pool_shares_created,
total_shares=total_post,
cap_table=self.snapshot(),
)
def analyze_exit(self, exit_valuation: float) -> list[ExitAnalysis]:
"""
Simple exit analysis: all preferred converts to common, proceeds split pro-rata.
(Does not model liquidation preferences — see fundraising_playbook.md for that.)
"""
total = self.total_shares()
price_per_share = exit_valuation / total
results = []
for s in self.shareholders:
if s.is_option_pool:
continue # unissued options don't receive proceeds
proceeds = s.shares * price_per_share
moic = proceeds / s.invested if s.invested > 0 else 0.0
results.append(ExitAnalysis(
exit_valuation=exit_valuation,
shareholder=s.name,
shares=s.shares,
ownership_pct=s.shares / total,
proceeds_common=proceeds,
invested=s.invested,
moic=moic,
))
return sorted(results, key=lambda x: x.proceeds_common, reverse=True)
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt(value: float, prefix: str = "$") -> str:
if value == float("inf"):
return "∞"
if abs(value) >= 1_000_000:
return f"{prefix}{value/1_000_000:.2f}M"
if abs(value) >= 1_000:
return f"{prefix}{value/1_000:.0f}K"
return f"{prefix}{value:.2f}"
def print_round_result(result: RoundResult, prev_cap_table: Optional[list[CapTableEntry]] = None) -> None:
print(f"\n{'='*70}")
print(f" {result.round_name.upper()}")
print(f"{'='*70}")
print(f" Pre-money valuation: {fmt(result.pre_money_valuation)}")
print(f" Investment: {fmt(result.investment_amount)}")
print(f" Post-money valuation: {fmt(result.post_money_valuation)}")
print(f" Price per share: {fmt(result.price_per_share, '$')}")
print(f" New shares issued: {result.new_shares_issued:,.0f}")
if result.option_pool_shares_created > 0:
print(f" Option pool created: {result.option_pool_shares_created:,.0f} shares")
print(f" ⚠️ Pool created pre-round: dilutes existing shareholders, not new investor")
print(f" Total shares post: {result.total_shares:,.0f}")
print(f"\n {'Shareholder':<22} {'Shares':>12} {'Ownership':>10} {'Invested':>10} {'Δ Ownership':>12}")
print(" " + "-"*68)
prev_map = {e.name: e.pct_ownership for e in prev_cap_table} if prev_cap_table else {}
for entry in result.cap_table:
delta = ""
if entry.name in prev_map:
change = (entry.pct_ownership - prev_map[entry.name]) * 100
delta = f"{change:+.1f}pp"
elif not entry.is_option_pool:
delta = "new"
invested_str = fmt(entry.invested) if entry.invested > 0 else "-"
print(
f" {entry.name:<22} {entry.shares:>12,.0f} "
f"{entry.pct_ownership*100:>9.2f}% {invested_str:>10} {delta:>12}"
)
def print_exit_analysis(results: list[ExitAnalysis], exit_valuation: float) -> None:
print(f"\n{'='*70}")
print(f" EXIT ANALYSIS @ {fmt(exit_valuation)} (all preferred converts to common)")
print(f"{'='*70}")
print(f"\n {'Shareholder':<22} {'Ownership':>10} {'Proceeds':>12} {'Invested':>10} {'MOIC':>8}")
print(" " + "-"*65)
for r in results:
moic_str = f"{r.moic:.1f}x" if r.moic > 0 else "n/a"
invested_str = fmt(r.invested) if r.invested > 0 else "-"
print(
f" {r.shareholder:<22} {r.ownership_pct*100:>9.2f}% "
f"{fmt(r.proceeds_common):>12} {invested_str:>10} {moic_str:>8}"
)
print(f"\n Note: Does not model liquidation preferences.")
print(f" Participating preferred reduces founder proceeds in most real exits.")
print(f" See references/fundraising_playbook.md for full liquidation waterfall.")
def print_dilution_summary(rounds: list[RoundResult]) -> None:
print(f"\n{'='*70}")
print(f" DILUTION SUMMARY — FOUNDER PERSPECTIVE")
print(f"{'='*70}")
# Find all founders (common shareholders who aren't investors or option pool)
founder_names = []
for entry in rounds[0].cap_table:
if entry.share_class == "common" and not entry.is_option_pool:
founder_names.append(entry.name)
if not founder_names:
print(" No common shareholders found in initial cap table.")
return
header = f" {'Round':<16}" + "".join(f" {n:<16}" for n in founder_names) + f" {'Total Inv':>12}"
print(header)
print(" " + "-" * (16 + 18 * len(founder_names) + 14))
for result in rounds:
cap_map = {e.name: e for e in result.cap_table}
total_invested = sum(e.invested for e in result.cap_table if not e.is_option_pool)
row = f" {result.round_name:<16}"
for name in founder_names:
pct = cap_map[name].pct_ownership * 100 if name in cap_map else 0
row += f" {pct:>6.2f}% "
row += f" {fmt(total_invested):>12}"
print(row)
def export_csv_rounds(rounds: list[RoundResult]) -> str:
buf = io.StringIO()
writer = csv.writer(buf)
writer.writerow(["Round", "Shareholder", "Share Class", "Shares", "Ownership Pct",
"Invested", "Pre Money", "Post Money", "Price Per Share"])
for r in rounds:
for entry in r.cap_table:
writer.writerow([
r.round_name, entry.name, entry.share_class,
round(entry.shares, 0), round(entry.pct_ownership * 100, 4),
round(entry.invested, 2), round(r.pre_money_valuation, 0),
round(r.post_money_valuation, 0), round(r.price_per_share, 4),
])
return buf.getvalue()
# ---------------------------------------------------------------------------
# Sample data: typical two-founder Series A/B/C startup
# ---------------------------------------------------------------------------
def build_sample_model() -> tuple[CapTable, list[RoundResult]]:
"""
Sample company:
- 2 founders, started with 10M shares each
- 1M shares for early advisor
- Raises Pre-seed → Seed → Series A → Series B → Series C
"""
cap = CapTable()
SHARES_PER_FOUNDER = 4_000_000
SHARES_ADVISOR = 200_000
# Founding state
cap.add_shareholder(Shareholder("Founder A (CEO)", "common", SHARES_PER_FOUNDER))
cap.add_shareholder(Shareholder("Founder B (CTO)", "common", SHARES_PER_FOUNDER))
cap.add_shareholder(Shareholder("Advisor", "common", SHARES_ADVISOR))
rounds: list[RoundResult] = []
prev_cap = cap.snapshot()
# Round 1: Pre-seed — $500K at $4.5M pre, 10% option pool created
r1 = cap.execute_round(RoundConfig(
name="Pre-seed",
pre_money_valuation=4_500_000,
investment_amount=500_000,
new_option_pool_pct=0.10,
option_pool_pre_round=True,
lead_investor_name="Angel Syndicate",
))
rounds.append(r1)
prev_r1 = r1.cap_table[:]
# Round 2: Seed — $2M at $9M pre, expand option pool to 12%
r2 = cap.execute_round(RoundConfig(
name="Seed",
pre_money_valuation=9_000_000,
investment_amount=2_000_000,
new_option_pool_pct=0.12,
option_pool_pre_round=True,
lead_investor_name="Seed Fund",
))
rounds.append(r2)
# Round 3: Series A — $12M at $38M pre, refresh option pool to 15%
r3 = cap.execute_round(RoundConfig(
name="Series A",
pre_money_valuation=38_000_000,
investment_amount=12_000_000,
new_option_pool_pct=0.15,
option_pool_pre_round=True,
lead_investor_name="Series A Fund",
))
rounds.append(r3)
# Round 4: Series B — $25M at $95M pre, refresh pool to 12%
r4 = cap.execute_round(RoundConfig(
name="Series B",
pre_money_valuation=95_000_000,
investment_amount=25_000_000,
new_option_pool_pct=0.12,
option_pool_pre_round=True,
lead_investor_name="Series B Fund",
))
rounds.append(r4)
# Round 5: Series C — $40M at $185M pre, refresh pool to 10%
r5 = cap.execute_round(RoundConfig(
name="Series C",
pre_money_valuation=185_000_000,
investment_amount=40_000_000,
new_option_pool_pct=0.10,
option_pool_pre_round=True,
lead_investor_name="Series C Fund",
))
rounds.append(r5)
return cap, rounds
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="Fundraising Model — Cap Table & Dilution")
parser.add_argument("--exit", type=float, default=250.0,
help="Exit valuation in $M for return analysis (default: 250)")
parser.add_argument("--csv", action="store_true", help="Export round data as CSV to stdout")
args = parser.parse_args()
exit_valuation = args.exit * 1_000_000
print("\n" + "="*70)
print(" FUNDRAISING MODEL — CAP TABLE & DILUTION ANALYSIS")
print(" Sample Company: Two-founder SaaS startup")
print(" Pre-seed → Seed → Series A → Series B → Series C")
print("="*70)
cap, rounds = build_sample_model()
# Print each round
prev = None
for r in rounds:
print_round_result(r, prev)
prev = r.cap_table
# Dilution summary table
print_dilution_summary(rounds)
# Exit analysis at specified valuation
exit_results = cap.analyze_exit(exit_valuation)
print_exit_analysis(exit_results, exit_valuation)
# Also print at 2x and 5x for sensitivity
print("\n Exit Sensitivity — Founder A Proceeds:")
print(f" {'Exit Valuation':<20} {'Founder A %':>12} {'Founder A $':>14} {'MOIC':>8}")
print(" " + "-"*56)
for mult in [0.5, 1.0, 1.5, 2.0, 3.0, 5.0]:
val = rounds[-1].post_money_valuation * mult
ex = cap.analyze_exit(val)
founder_a = next((r for r in ex if r.shareholder == "Founder A (CEO)"), None)
if founder_a:
print(f" {fmt(val):<20} {founder_a.ownership_pct*100:>11.2f}% "
f"{fmt(founder_a.proceeds_common):>14} {'n/a':>8}")
print("\n Key Takeaways:")
final = rounds[-1].cap_table
total = sum(e.shares for e in final)
founder_a_final = next((e for e in final if e.name == "Founder A (CEO)"), None)
if founder_a_final:
print(f" Founder A final ownership: {founder_a_final.pct_ownership*100:.2f}%")
total_raised = sum(e.invested for e in final)
print(f" Total capital raised: {fmt(total_raised)}")
print(f" Total shares outstanding: {total:,.0f}")
print(f" Final post-money: {fmt(rounds[-1].post_money_valuation)}")
print("\n Run with --exit <$M> to model proceeds at different exit valuations.")
print(" Example: python fundraising_model.py --exit 500")
if args.csv:
print("\n\n--- CSV EXPORT ---\n")
sys.stdout.write(export_csv_rounds(rounds))
if __name__ == "__main__":
main()
FILE:scripts/unit_economics_analyzer.py
#!/usr/bin/env python3
"""
Unit Economics Analyzer
========================
Per-cohort LTV, per-channel CAC, payback periods, and LTV:CAC ratios.
Never blended averages — those hide what's actually happening.
Usage:
python unit_economics_analyzer.py
python unit_economics_analyzer.py --csv
Stdlib only. No dependencies.
"""
import argparse
import csv
import io
import sys
from dataclasses import dataclass, field
from typing import Optional
# ---------------------------------------------------------------------------
# Data structures
# ---------------------------------------------------------------------------
@dataclass
class CohortData:
"""
Revenue data for a group of customers acquired in the same period.
Revenue is tracked monthly: revenue[0] = month 1, revenue[1] = month 2, etc.
"""
label: str # e.g. "Q1 2024"
acquisition_period: str # human-readable label
customers_acquired: int
total_cac_spend: float # total S&M spend to acquire this cohort
monthly_revenue: list[float] # revenue per month from this cohort
gross_margin_pct: float = 0.70 # blended gross margin for this cohort
@dataclass
class ChannelData:
"""Acquisition cost and customer data for a single channel."""
channel: str
spend: float
customers_acquired: int
avg_arpa: float # average revenue per account (monthly)
gross_margin_pct: float = 0.70
avg_monthly_churn: float = 0.02 # monthly churn rate for customers from this channel
@dataclass
class UnitEconomicsResult:
"""Computed unit economics for a cohort or channel."""
label: str
customers: int
cac: float
arpa: float # average revenue per account per month
gross_margin_pct: float
monthly_churn: float
ltv: float
ltv_cac_ratio: float
payback_months: float
# Cohort-specific
m1_revenue: Optional[float] = None
m6_revenue: Optional[float] = None
m12_revenue: Optional[float] = None
m24_revenue: Optional[float] = None
m12_ltv: Optional[float] = None # realized LTV through month 12
retention_m6: Optional[float] = None # % of M1 revenue retained at M6
retention_m12: Optional[float] = None
# ---------------------------------------------------------------------------
# Calculators
# ---------------------------------------------------------------------------
def calc_ltv(arpa: float, gross_margin_pct: float, monthly_churn: float) -> float:
"""
LTV = (ARPA × Gross Margin) / Monthly Churn Rate
Assumes constant churn (simplified; cohort method is more accurate).
"""
if monthly_churn <= 0:
return float("inf")
return (arpa * gross_margin_pct) / monthly_churn
def calc_payback(cac: float, arpa: float, gross_margin_pct: float) -> float:
"""
CAC Payback (months) = CAC / (ARPA × Gross Margin)
"""
denominator = arpa * gross_margin_pct
if denominator <= 0:
return float("inf")
return cac / denominator
def analyze_cohort(cohort: CohortData) -> UnitEconomicsResult:
"""Compute full unit economics for a cohort."""
n = cohort.customers_acquired
if n == 0:
raise ValueError(f"Cohort {cohort.label}: customers_acquired cannot be 0")
cac = cohort.total_cac_spend / n
# ARPA from month 1 revenue
m1_rev = cohort.monthly_revenue[0] if cohort.monthly_revenue else 0
arpa = m1_rev / n if n > 0 else 0
# Observed monthly churn from cohort data
# Use revenue decline from M1 to M12 to estimate churn
months_available = len(cohort.monthly_revenue)
if months_available >= 12:
m12_rev = cohort.monthly_revenue[11]
# Revenue retention over 12 months: (M12/M1)^(1/11) per month on average
# Implied monthly retention rate
if m1_rev > 0 and m12_rev > 0:
monthly_retention = (m12_rev / m1_rev) ** (1 / 11)
monthly_churn = 1 - monthly_retention
else:
monthly_churn = 0.02 # default
elif months_available >= 6:
m6_rev = cohort.monthly_revenue[5]
if m1_rev > 0 and m6_rev > 0:
monthly_retention = (m6_rev / m1_rev) ** (1 / 5)
monthly_churn = 1 - monthly_retention
else:
monthly_churn = 0.02
else:
monthly_churn = 0.02 # default if < 6 months data
# Clamp to reasonable range
monthly_churn = max(0.001, min(monthly_churn, 0.30))
ltv = calc_ltv(arpa, cohort.gross_margin_pct, monthly_churn)
payback = calc_payback(cac, arpa, cohort.gross_margin_pct)
ltv_cac = ltv / cac if cac > 0 else float("inf")
# Snapshot revenues
def rev_at(month_idx: int) -> Optional[float]:
if months_available > month_idx:
return cohort.monthly_revenue[month_idx]
return None
m6 = rev_at(5)
m12 = rev_at(11)
m24 = rev_at(23)
# Realized LTV through observed months (actual gross profit)
m12_ltv = sum(cohort.monthly_revenue[:12]) * cohort.gross_margin_pct if months_available >= 12 else None
# Retention rates
ret_m6 = (m6 / m1_rev) if (m6 is not None and m1_rev > 0) else None
ret_m12 = (m12 / m1_rev) if (m12 is not None and m1_rev > 0) else None
return UnitEconomicsResult(
label=cohort.label,
customers=n,
cac=cac,
arpa=arpa,
gross_margin_pct=cohort.gross_margin_pct,
monthly_churn=monthly_churn,
ltv=ltv,
ltv_cac_ratio=ltv_cac,
payback_months=payback,
m1_revenue=m1_rev,
m6_revenue=m6,
m12_revenue=m12,
m24_revenue=m24,
m12_ltv=m12_ltv,
retention_m6=ret_m6,
retention_m12=ret_m12,
)
def analyze_channel(ch: ChannelData) -> UnitEconomicsResult:
"""Compute unit economics for an acquisition channel."""
if ch.customers_acquired == 0:
raise ValueError(f"Channel {ch.channel}: customers_acquired cannot be 0")
cac = ch.spend / ch.customers_acquired
ltv = calc_ltv(ch.avg_arpa, ch.gross_margin_pct, ch.avg_monthly_churn)
payback = calc_payback(cac, ch.avg_arpa, ch.gross_margin_pct)
ltv_cac = ltv / cac if cac > 0 else float("inf")
return UnitEconomicsResult(
label=ch.channel,
customers=ch.customers_acquired,
cac=cac,
arpa=ch.avg_arpa,
gross_margin_pct=ch.gross_margin_pct,
monthly_churn=ch.avg_monthly_churn,
ltv=ltv,
ltv_cac_ratio=ltv_cac,
payback_months=payback,
)
# ---------------------------------------------------------------------------
# Blended metrics (for comparison)
# ---------------------------------------------------------------------------
def blended_cac(channels: list[ChannelData]) -> float:
total_spend = sum(c.spend for c in channels)
total_customers = sum(c.customers_acquired for c in channels)
return total_spend / total_customers if total_customers > 0 else 0
def blended_ltv(channels: list[ChannelData]) -> float:
"""Weighted average LTV by customers acquired."""
total_customers = sum(c.customers_acquired for c in channels)
if total_customers == 0:
return 0
weighted = sum(
calc_ltv(c.avg_arpa, c.gross_margin_pct, c.avg_monthly_churn) * c.customers_acquired
for c in channels
)
return weighted / total_customers
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt(value: float, prefix: str = "$", decimals: int = 0) -> str:
if value == float("inf"):
return "∞"
if abs(value) >= 1_000_000:
return f"{prefix}{value/1_000_000:.2f}M"
if abs(value) >= 1_000:
return f"{prefix}{value/1_000:.1f}K"
return f"{prefix}{value:.{decimals}f}"
def pct(value: Optional[float]) -> str:
if value is None:
return "n/a"
return f"{value*100:.1f}%"
def rating(ltv_cac: float, payback: float) -> str:
if ltv_cac == float("inf"):
return "∞"
if ltv_cac >= 5 and payback <= 12:
return "🟢 Excellent"
if ltv_cac >= 3 and payback <= 18:
return "🟡 Good"
if ltv_cac >= 2 and payback <= 24:
return "🟠 Marginal"
return "🔴 Poor"
def print_cohort_analysis(results: list[UnitEconomicsResult]) -> None:
print("\n" + "="*80)
print(" COHORT ANALYSIS")
print("="*80)
print(f" {'Cohort':<12} {'Cust':>5} {'CAC':>8} {'ARPA/mo':>9} {'Churn/mo':>10} "
f"{'LTV':>10} {'LTV:CAC':>8} {'Payback':>9} {'Ret@M12':>8}")
print(" " + "-"*88)
for r in results:
payback_str = f"{r.payback_months:.1f}mo" if r.payback_months != float("inf") else "∞"
ltv_str = fmt(r.ltv) if r.ltv != float("inf") else "∞"
ltv_cac_str = f"{r.ltv_cac_ratio:.1f}x" if r.ltv_cac_ratio != float("inf") else "∞"
print(
f" {r.label:<12} {r.customers:>5} {fmt(r.cac):>8} {fmt(r.arpa):>9} "
f"{pct(r.monthly_churn):>10} {ltv_str:>10} {ltv_cac_str:>8} "
f"{payback_str:>9} {pct(r.retention_m12):>8}"
)
# Trend analysis
print("\n Cohort Trend (is the business getting better or worse?):")
if len(results) >= 3:
ltv_cac_values = [r.ltv_cac_ratio for r in results if r.ltv_cac_ratio != float("inf")]
cac_values = [r.cac for r in results]
churn_values = [r.monthly_churn for r in results]
if len(ltv_cac_values) >= 2:
ltv_cac_trend = "↑ Improving" if ltv_cac_values[-1] > ltv_cac_values[0] else "↓ Deteriorating"
else:
ltv_cac_trend = "n/a"
cac_trend = "↓ Decreasing (good)" if cac_values[-1] < cac_values[0] else "↑ Increasing"
churn_trend = "↓ Improving" if churn_values[-1] < churn_values[0] else "↑ Worsening"
print(f" LTV:CAC: {ltv_cac_trend}")
print(f" CAC: {cac_trend}")
print(f" Churn rate: {churn_trend}")
def print_channel_analysis(results: list[UnitEconomicsResult], channels: list[ChannelData]) -> None:
print("\n" + "="*80)
print(" CHANNEL ANALYSIS (Per-Channel vs Blended)")
print("="*80)
print(f" {'Channel':<22} {'Spend':>9} {'Cust':>5} {'CAC':>8} {'LTV':>10} {'LTV:CAC':>8} {'Payback':>9} {'Rating'}")
print(" " + "-"*90)
for r, ch in zip(results, channels):
payback_str = f"{r.payback_months:.1f}mo" if r.payback_months != float("inf") else "∞"
ltv_str = fmt(r.ltv) if r.ltv != float("inf") else "∞"
ltv_cac_str = f"{r.ltv_cac_ratio:.1f}x" if r.ltv_cac_ratio != float("inf") else "∞"
print(
f" {r.label:<22} {fmt(ch.spend):>9} {r.customers:>5} {fmt(r.cac):>8} "
f"{ltv_str:>10} {ltv_cac_str:>8} {payback_str:>9} {rating(r.ltv_cac_ratio, r.payback_months)}"
)
# Blended comparison
b_cac = blended_cac(channels)
b_ltv = blended_ltv(channels)
b_ltv_cac = b_ltv / b_cac if b_cac > 0 else 0
total_spend = sum(c.spend for c in channels)
total_customers = sum(c.customers_acquired for c in channels)
avg_payback = sum(
calc_payback(b_cac, c.avg_arpa, c.gross_margin_pct) * c.customers_acquired
for c in channels
) / total_customers
print(" " + "-"*90)
print(
f" {'BLENDED (dangerous)':<22} {fmt(total_spend):>9} {total_customers:>5} "
f"{fmt(b_cac):>8} {fmt(b_ltv):>10} {b_ltv_cac:.1f}x{'':<7} "
f"{avg_payback:.1f}mo{'':<4} {rating(b_ltv_cac, avg_payback)}"
)
print("\n ⚠️ Blended numbers hide channel-level problems. Manage channels individually.")
# Budget reallocation
print("\n Recommended Budget Reallocation:")
sorted_results = sorted(zip(results, channels), key=lambda x: x[0].ltv_cac_ratio, reverse=True)
for r, ch in sorted_results:
if r.ltv_cac_ratio >= 3:
action = "✅ Scale"
elif r.ltv_cac_ratio >= 2:
action = "🔄 Optimize"
else:
action = "❌ Cut / pause"
print(f" {ch.channel:<22} LTV:CAC = {r.ltv_cac_ratio:.1f}x → {action}")
def export_csv_results(cohort_results: list[UnitEconomicsResult], channel_results: list[UnitEconomicsResult]) -> str:
buf = io.StringIO()
writer = csv.writer(buf)
writer.writerow(["Type", "Label", "Customers", "CAC", "ARPA_Monthly", "Gross_Margin_Pct",
"Monthly_Churn", "LTV", "LTV_CAC_Ratio", "Payback_Months",
"Retention_M6", "Retention_M12"])
for r in cohort_results:
writer.writerow(["cohort", r.label, r.customers, round(r.cac, 2), round(r.arpa, 2),
r.gross_margin_pct, round(r.monthly_churn, 4),
round(r.ltv, 2) if r.ltv != float("inf") else "inf",
round(r.ltv_cac_ratio, 2) if r.ltv_cac_ratio != float("inf") else "inf",
round(r.payback_months, 2) if r.payback_months != float("inf") else "inf",
round(r.retention_m6, 3) if r.retention_m6 else "",
round(r.retention_m12, 3) if r.retention_m12 else ""])
for r in channel_results:
writer.writerow(["channel", r.label, r.customers, round(r.cac, 2), round(r.arpa, 2),
r.gross_margin_pct, round(r.monthly_churn, 4),
round(r.ltv, 2) if r.ltv != float("inf") else "inf",
round(r.ltv_cac_ratio, 2) if r.ltv_cac_ratio != float("inf") else "inf",
round(r.payback_months, 2) if r.payback_months != float("inf") else "inf",
"", ""])
return buf.getvalue()
# ---------------------------------------------------------------------------
# Sample data
# ---------------------------------------------------------------------------
def make_sample_cohorts() -> list[CohortData]:
"""
Series A SaaS company, 8 quarters of cohort data.
Shows a business improving on all dimensions over time.
"""
return [
CohortData(
label="Q1 2023", acquisition_period="Jan-Mar 2023",
customers_acquired=12, total_cac_spend=54_000,
gross_margin_pct=0.68,
monthly_revenue=[
10_200, 9_600, 9_100, 8_700, 8_300, 8_000, # M1-M6
7_800, 7_600, 7_400, 7_200, 7_000, 6_800, # M7-M12
6_700, 6_600, 6_500, 6_400, 6_300, 6_200, # M13-M18
6_100, 6_000, 5_900, 5_800, 5_700, 5_600, # M19-M24
],
),
CohortData(
label="Q2 2023", acquisition_period="Apr-Jun 2023",
customers_acquired=15, total_cac_spend=60_000,
gross_margin_pct=0.69,
monthly_revenue=[
13_500, 12_900, 12_500, 12_100, 11_800, 11_500,
11_300, 11_100, 10_900, 10_700, 10_500, 10_300,
10_200, 10_100, 10_000, 9_900, 9_800, 9_700,
],
),
CohortData(
label="Q3 2023", acquisition_period="Jul-Sep 2023",
customers_acquired=18, total_cac_spend=63_000,
gross_margin_pct=0.70,
monthly_revenue=[
16_200, 15_800, 15_400, 15_100, 14_800, 14_600,
14_400, 14_200, 14_000, 13_900, 13_800, 13_700,
13_600, 13_500, 13_400, 13_300,
],
),
CohortData(
label="Q4 2023", acquisition_period="Oct-Dec 2023",
customers_acquired=22, total_cac_spend=70_400,
gross_margin_pct=0.71,
monthly_revenue=[
20_900, 20_500, 20_200, 19_900, 19_700, 19_500,
19_300, 19_100, 19_000, 18_900, 18_800, 18_700,
],
),
CohortData(
label="Q1 2024", acquisition_period="Jan-Mar 2024",
customers_acquired=28, total_cac_spend=81_200,
gross_margin_pct=0.72,
monthly_revenue=[
27_200, 26_900, 26_600, 26_400, 26_200, 26_000,
25_800, 25_700, 25_600, 25_500,
],
),
CohortData(
label="Q2 2024", acquisition_period="Apr-Jun 2024",
customers_acquired=34, total_cac_spend=91_800,
gross_margin_pct=0.72,
monthly_revenue=[
33_300, 33_000, 32_800, 32_600, 32_400, 32_200,
],
),
CohortData(
label="Q3 2024", acquisition_period="Jul-Sep 2024",
customers_acquired=40, total_cac_spend=100_000,
gross_margin_pct=0.73,
monthly_revenue=[
39_600, 39_400, 39_200,
],
),
CohortData(
label="Q4 2024", acquisition_period="Oct-Dec 2024",
customers_acquired=47, total_cac_spend=112_800,
gross_margin_pct=0.73,
monthly_revenue=[
47_000,
],
),
]
def make_sample_channels() -> list[ChannelData]:
"""
Q4 2024 channel breakdown. Blended looks fine; per-channel reveals problems.
"""
return [
ChannelData("Organic / SEO", spend=9_500, customers_acquired=14, avg_arpa=950, gross_margin_pct=0.73, avg_monthly_churn=0.015),
ChannelData("Paid Search (SEM)", spend=48_000, customers_acquired=18, avg_arpa=980, gross_margin_pct=0.73, avg_monthly_churn=0.020),
ChannelData("Paid Social", spend=32_000, customers_acquired=8, avg_arpa=900, gross_margin_pct=0.72, avg_monthly_churn=0.025),
ChannelData("Content / Inbound", spend=11_000, customers_acquired=6, avg_arpa=1100, gross_margin_pct=0.74, avg_monthly_churn=0.012),
ChannelData("Outbound SDR", spend=22_000, customers_acquired=4, avg_arpa=1200, gross_margin_pct=0.73, avg_monthly_churn=0.022),
ChannelData("Events / Webinars", spend=18_500, customers_acquired=3, avg_arpa=1050, gross_margin_pct=0.72, avg_monthly_churn=0.028),
ChannelData("Partner / Referral", spend=7_800, customers_acquired=7, avg_arpa=1000, gross_margin_pct=0.73, avg_monthly_churn=0.013),
]
# ---------------------------------------------------------------------------
# Entry point
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="Unit Economics Analyzer")
parser.add_argument("--csv", action="store_true", help="Export results as CSV to stdout")
args = parser.parse_args()
cohorts = make_sample_cohorts()
channels = make_sample_channels()
print("\n" + "="*80)
print(" UNIT ECONOMICS ANALYZER")
print(" Sample Company: Series A SaaS | Q4 2024 Snapshot")
print(" Gross Margin: ~72% | Monthly Churn: derived from cohort data")
print("="*80)
cohort_results = [analyze_cohort(c) for c in cohorts]
channel_results = [analyze_channel(c) for c in channels]
print_cohort_analysis(cohort_results)
print_channel_analysis(channel_results, channels)
# Health summary
print("\n" + "="*80)
print(" HEALTH SUMMARY")
print("="*80)
latest = cohort_results[-1]
prev = cohort_results[-4] if len(cohort_results) >= 4 else cohort_results[0]
print(f"\n Latest Cohort ({latest.label}):")
print(f" CAC: {fmt(latest.cac)}")
ltv_str = fmt(latest.ltv) if latest.ltv != float("inf") else "∞"
ltv_cac_str = f"{latest.ltv_cac_ratio:.1f}x" if latest.ltv_cac_ratio != float("inf") else "∞"
payback_str = f"{latest.payback_months:.1f} months" if latest.payback_months != float("inf") else "∞"
print(f" LTV: {ltv_str}")
print(f" LTV:CAC: {ltv_cac_str} (target: > 3x)")
print(f" CAC Payback: {payback_str} (target: < 18mo)")
print(f" Rating: {rating(latest.ltv_cac_ratio, latest.payback_months)}")
# Trend vs 4 quarters ago
print(f"\n Trend vs {prev.label}:")
cac_delta = (latest.cac - prev.cac) / prev.cac * 100
ltv_delta_str = "n/a"
if latest.ltv != float("inf") and prev.ltv != float("inf"):
ltv_delta = (latest.ltv - prev.ltv) / prev.ltv * 100
ltv_delta_str = f"{ltv_delta:+.1f}%"
cac_str = "↓ Better" if cac_delta < 0 else "↑ Worse"
print(f" CAC: {cac_delta:+.1f}% ({cac_str})")
print(f" LTV: {ltv_delta_str}")
print("\n Benchmark Reference:")
print(" LTV:CAC > 5x → Scale aggressively")
print(" LTV:CAC 3-5x → Healthy; grow at current pace")
print(" LTV:CAC 2-3x → Marginal; optimize before scaling")
print(" LTV:CAC < 2x → Acquiring unprofitably; stop and fix")
print(" Payback < 12mo → Outstanding capital efficiency")
print(" Payback 12-18mo → Good for B2B SaaS")
print(" Payback > 24mo → Requires long-dated capital to scale")
if args.csv:
print("\n\n--- CSV EXPORT ---\n")
sys.stdout.write(export_csv_results(cohort_results, channel_results))
if __name__ == "__main__":
main()
Lập chiến lược nội dung, quyết định nội dung cần tạo và chủ đề cần phủ, gồm cụm chủ đề và lịch biên tập.
---
name: content-strategy
description: When the user wants to plan a content strategy, decide what content to create, or figure out what topics to cover. Also use when the user mentions "content strategy," "what should I write about," "content ideas," "blog strategy," "topic clusters," "content planning," "editorial calendar," "content marketing," "content roadmap," "what content should I create," "blog topics," "content pillars," or "I don't know what to write." Use this whenever someone needs help deciding what content to produce, not just writing it. For writing individual pieces, see copywriting. For SEO-specific audits, see seo-audit. For social media content specifically, see social.
metadata:
version: 2.1.1
---
# Content Strategy
You are a content strategist. Your goal is to help plan content that drives traffic, builds authority, and generates leads by being either searchable, shareable, or both.
## Before Planning
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Gather this context (ask if not provided):
### 1. Business Context
- What does the company do?
- Who is the ideal customer?
- What's the primary goal for content? (traffic, leads, brand awareness, thought leadership)
- What problems does your product solve?
### 2. Customer Research
- What questions do customers ask before buying?
- What objections come up in sales calls?
- What topics appear repeatedly in support tickets?
- What language do customers use to describe their problems?
### 3. Current State
- Do you have existing content? What's working?
- What resources do you have? (writers, budget, time)
- What content formats can you produce? (written, video, audio)
### 4. Competitive Landscape
- Who are your main competitors?
- What content gaps exist in your market?
---
## Treat Content Like a Product
Every piece is its own launch. Content isn't overhead—it's **brand surface area**: each published piece is a new entry point where a stranger can discover you, and hundreds of pieces compound into hundreds of doorways working 24/7. Plan, ship, and promote each piece with the same intent you'd bring to a product release. A post that's written and forgotten has almost no surface area; a post that's distributed (see **Create Once, Distribute Twice** below) multiplies it.
This section covers the searchable/shareable lens, then the execution and prioritization layer: which pieces to make (scoring), how the calendar splits, and per-format discipline.
## Searchable vs Shareable
Every piece of content must be searchable, shareable, or both. Prioritize in that order—search traffic is the foundation.
**Searchable content** captures existing demand. Optimized for people actively looking for answers.
**Shareable content** creates demand. Spreads ideas and gets people talking.
### When Writing Searchable Content
- Target a specific keyword or question
- Match search intent exactly—answer what the searcher wants
- Use clear titles that match search queries
- Structure with headings that mirror search patterns
- Place keywords in title, headings, first paragraph, URL
- Provide comprehensive coverage (don't leave questions unanswered)
- Include data, examples, and links to authoritative sources
- Optimize for AI/LLM discovery: clear positioning, structured content, brand consistency across the web
### When Writing Shareable Content
- Lead with a novel insight, original data, or counterintuitive take
- Challenge conventional wisdom with well-reasoned arguments
- Tell stories that make people feel something
- Create content people want to share to look smart or help others
- Connect to current trends or emerging problems
- Share vulnerable, honest experiences others can learn from
---
## Content Types
### Searchable Content Types
**Use-Case Content**
Formula: [persona] + [use-case]. Targets long-tail keywords.
- "Project management for designers"
- "Task tracking for developers"
- "Client collaboration for freelancers"
**Hub and Spoke**
Hub = comprehensive overview. Spokes = related subtopics.
```
/topic (hub)
├── /topic/subtopic-1 (spoke)
├── /topic/subtopic-2 (spoke)
└── /topic/subtopic-3 (spoke)
```
Create hub first, then build spokes. Interlink strategically.
**Note:** Most content works fine under `/blog`. Only use dedicated hub/spoke URL structures for major topics with layered depth (e.g., Atlassian's `/agile` guide). For typical blog posts, `/blog/post-title` is sufficient.
**Template Libraries**
High-intent keywords + product adoption.
- Target searches like "marketing plan template"
- Provide immediate standalone value
- Show how product enhances the template
### Shareable Content Types
**Thought Leadership**
- Articulate concepts everyone feels but hasn't named
- Challenge conventional wisdom with evidence
- Share vulnerable, honest experiences
**Data-Driven Content**
- Product data analysis (anonymized insights)
- Public data analysis (uncover patterns)
- Original research (run experiments, share results)
**Expert Roundups**
15-30 experts answering one specific question. Built-in distribution.
**Case Studies**
Structure: Challenge → Solution → Results → Key learnings
**Meta Content**
Behind-the-scenes transparency. "How We Got Our First $5k MRR," "Why We Chose Debt Over VC."
### Link-Earning Formats
When the goal of a piece is backlinks specifically, format choice matters more than production effort. Foundation Inc.'s B2B Backlink Intelligence Report (March 2026 — a single vendor study of B2B SaaS sites, so treat as directional) measured each format's share of backlinks relative to its share of pages:
| Format | Backlinks vs. page share |
|---|---|
| Statistics / data roundups | **4.25x** |
| Glossary / definition pages | 1.47x |
| Interactive tools / calculators (see **free-tools**) | 1.38x |
| How-to / tutorials | 1.36x |
| Original research / reports | 0.80x |
| Ultimate guides | 0.77x |
| Thought leadership | 0.74x |
| Templates / frameworks | 0.68x |
The counterintuitive read: **curating statistics earns ~5x the links of producing original research.** Writers link to whatever makes citation easiest — a maintained stat-roundup page is citation infrastructure, while original research often gets cited *via* the roundups that aggregate it. Implications: (1) publish a stats page for your category and keep it fresh — it's cheap and compounds, and citable one-line stats are also what LLMs lift, making it an AI-visibility play (see **ai-seo**); (2) when you do run original research, pair it with your own stat-roundup page that presents the findings as citable one-liners, so you capture the links your data generates. The formats at the bottom aren't dead — guides, templates, and thought leadership earn their keep on rankings, conversions, and brand. Judge each piece by the job it's for, and don't expect links from formats that don't earn them.
For programmatic content at scale, see **programmatic-seo** skill.
---
## Content Pillars and Topic Clusters
Content pillars are the 3-5 core topics your brand will own. Each pillar spawns a cluster of related content.
Most of the time, all content can live under `/blog` with good internal linking between related posts. Dedicated pillar pages with custom URL structures (like `/guides/topic`) are only needed when you're building comprehensive resources with multiple layers of depth.
### How to Identify Pillars
1. **Product-led**: What problems does your product solve?
2. **Audience-led**: What does your ICP need to learn?
3. **Search-led**: What topics have volume in your space?
4. **Competitor-led**: What are competitors ranking for?
### Pillar Structure
```
Pillar Topic (Hub)
├── Subtopic Cluster 1
│ ├── Article A
│ ├── Article B
│ └── Article C
├── Subtopic Cluster 2
│ ├── Article D
│ ├── Article E
│ └── Article F
└── Subtopic Cluster 3
├── Article G
├── Article H
└── Article I
```
### Pillar Criteria
Good pillars should:
- Align with your product/service
- Match what your audience cares about
- Have search volume and/or social interest
- Be broad enough for many subtopics
---
## Keyword Research by Buyer Stage
Map topics to the buyer's journey using proven keyword modifiers:
### Awareness Stage
Modifiers: "what is," "how to," "guide to," "introduction to"
Example: If customers ask about project management basics:
- "What is Agile Project Management"
- "Guide to Sprint Planning"
- "How to Run a Standup Meeting"
### Consideration Stage
Modifiers: "best," "top," "vs," "alternatives," "comparison"
Example: If customers evaluate multiple tools:
- "Best Project Management Tools for Remote Teams"
- "Asana vs Trello vs Monday"
- "Basecamp Alternatives"
### Decision Stage
Modifiers: "pricing," "reviews," "demo," "trial," "buy"
Example: If pricing comes up in sales calls:
- "Project Management Tool Pricing Comparison"
- "How to Choose the Right Plan"
- "[Product] Reviews"
### Implementation Stage
Modifiers: "templates," "examples," "tutorial," "how to use," "setup"
Example: If support tickets show implementation struggles:
- "Project Template Library"
- "Step-by-Step Setup Tutorial"
- "How to Use [Feature]"
---
## Content Ideation Sources
### 1. Keyword Data
If user provides keyword exports (Ahrefs, SEMrush, GSC), analyze for:
- Topic clusters (group related keywords)
- Buyer stage (awareness/consideration/decision/implementation)
- Search intent (informational, commercial, transactional)
- Quick wins (low competition + decent volume + high relevance)
- Content gaps (keywords competitors rank for that you don't)
Output as prioritized table:
| Keyword | Volume | Difficulty | Buyer Stage | Content Type | Priority |
### 2. Call Transcripts
If user provides sales or customer call transcripts, extract:
- Questions asked → FAQ content or blog posts
- Pain points → problems in their own words
- Objections → content to address proactively
- Language patterns → exact phrases to use (voice of customer)
- Competitor mentions → what they compared you to
Output content ideas with supporting quotes.
### 3. Survey Responses
If user provides survey data, mine for:
- Open-ended responses (topics and language)
- Common themes (30%+ mention = high priority)
- Resource requests (what they wish existed)
- Content preferences (formats they want)
### 4. Forum Research
Use web search to find content ideas:
**Reddit:** `site:reddit.com [topic]`
- Top posts in relevant subreddits
- Questions and frustrations in comments
- Upvoted answers (validates what resonates)
**Quora:** `site:quora.com [topic]`
- Most-followed questions
- Highly upvoted answers
**Other:** Indie Hackers, Hacker News, Product Hunt, industry Slack/Discord
Extract: FAQs, misconceptions, debates, problems being solved, terminology used.
### 5. Competitor Analysis
Use web search to analyze competitor content:
**Find their content:** `site:competitor.com/blog`
**Analyze:**
- Top-performing posts (comments, shares)
- Topics covered repeatedly
- Gaps they haven't covered
- Case studies (customer problems, use cases, results)
- Content structure (pillars, categories, formats)
**Identify opportunities:**
- Topics you can cover better
- Angles they're missing
- Outdated content to improve on
### 6. Sales and Support Input
Extract from customer-facing teams:
- Common objections
- Repeated questions
- Support ticket patterns
- Success stories
- Feature requests and underlying problems
---
## Prioritizing Content Ideas
Score each idea on four factors:
### 1. Customer Impact (40%)
- How frequently did this topic come up in research?
- What percentage of customers face this challenge?
- How emotionally charged was this pain point?
- What's the potential LTV of customers with this need?
### 2. Content-Market Fit (30%)
- Does this align with problems your product solves?
- Can you offer unique insights from customer research?
- Do you have customer stories to support this?
- Will this naturally lead to product interest?
### 3. Search Potential (20%)
- What's the monthly search volume?
- How competitive is this topic?
- Are there related long-tail opportunities?
- Is search interest growing or declining?
### 4. Resource Requirements (10%)
- Do you have expertise to create authoritative content?
- What additional research is needed?
- What assets (graphics, data, examples) will you need?
### Scoring Template
| Idea | Customer Impact (40%) | Content-Market Fit (30%) | Search Potential (20%) | Resources (10%) | Total |
|------|----------------------|-------------------------|----------------------|-----------------|-------|
| Topic A | 8 | 9 | 7 | 6 | 8.0 |
| Topic B | 6 | 7 | 9 | 8 | 7.1 |
Score 1-10 per factor, multiply by the weight, sum for the total. Rank the list; make the top-scoring pieces first.
---
## Calendar Split: 60/30/10
Balance the editorial calendar so search compounds while shareable pieces keep you visible:
- **60% searchable** — the foundation. Demand you can capture predictably (use-case content, hub/spoke, how-tos).
- **30% shareable** — thought leadership, original data, opinion. Creates demand and earns links/mentions.
- **10% experimental** — new formats, channels, or bets. Cheap insurance against a stale mix.
This is a starting ratio, not a rule. A brand-new blog may over-index on searchable to build a base; an established brand chasing category leadership may push shareable higher.
---
## Per-Format Execution Discipline
Treating content like a product means each format has a production standard, not just a topic:
- **Blog post** — write **10 title options** before drafting (the title does most of the work; pick the strongest). Plan **~5 editing passes** (structure, clarity, evidence, line edit, headline/SEO). For the writing itself, see **copywriting**.
- **Long-form guide** — the flagship of a pillar. Comprehensive enough to be *the* resource; structured with a table of contents and internal links to spokes. Build the hub before the spokes.
- **Video** — script the hook first; front-load the payoff. Repurpose into short-form clips at creation time (see **social**).
- **Podcast** — one interview yields a transcript, quote graphics, short clips, and a written recap. Design the episode knowing it will be atomized.
- **Email** — one idea per send; the subject line is the title—write several and pick. For sequences and lifecycle, see **emails**.
---
## Create Once, Distribute Twice
Creating content is half the job—distribution is the other half, and most teams skip it. The philosophy: **one exceptional piece, reformatted and repurposed across every channel, not a fresh piece per platform.** Pouring effort into a single flagship and then distributing it everywhere beats spreading thin effort across many mediocre platform-native posts.
Build **distribution hooks into the piece at creation time**, not after: write subheads that stand alone as social posts, structure sections to be lifted out modularly, and pull quotes/stats you already know you'll graphic-ify. A well-designed guide is a distribution kit in disguise.
**The ORB Framework as a funnel** — route attention from borrowed → rented → owned, which maps to discovery → engagement → conversion:
- **Borrowed** (other people's audiences: podcasts, guest posts, partnerships) — discovery / breakthrough reach.
- **Rented** (social platforms, ad networks) — engagement, but you don't own the audience or the algorithm.
- **Owned** (email list, blog, community) — conversion and the only durable asset. Everything upstream should funnel here.
ORB mechanics live in the **launch** skill (channel-type playbook) and content atomization/repurposing lives in **social**; the value here is consolidating the *distribute* half of content strategy so it has a home.
**Failure modes to avoid:**
- **Spray-and-pray** — posting everywhere with no flagship and no repurposing plan. Effort scatters, nothing compounds.
- **Platform dependency** — building on rented land. Facebook organic reach fell from ~20% to under 2%; any rented channel can throttle you overnight.
- **The ownership paradox** — teams spend ~90% of effort on channels they don't control (rented/borrowed) and neglect the owned assets that actually convert and can't be taken away.
For the full distribution spine—the Content Distribution Flywheel, platform half-lives, and the atomization checklist—see the reference below.
---
## Output Format
When creating a content strategy, provide:
### 1. Content Pillars
- 3-5 pillars with rationale
- Subtopic clusters for each pillar
- How pillars connect to product
### 2. Priority Topics
For each recommended piece:
- Topic/title
- Searchable, shareable, or both
- Content type (use-case, hub/spoke, thought leadership, etc.)
- Target keyword and buyer stage
- Why this topic (customer research backing)
### 3. Topic Cluster Map
Visual or structured representation of how content interconnects.
---
## Task-Specific Questions
1. What patterns emerge from your last 10 customer conversations?
2. What questions keep coming up in sales calls?
3. Where are competitors' content efforts falling short?
4. What unique insights from customer research aren't being shared elsewhere?
5. Which existing content drives the most conversions, and why?
---
## References
- **[Content Distribution Spine](references/content-distribution.md)**: Create Once Distribute Twice, ORB as a funnel, the ownership paradox, platform half-lives, the Content Distribution Flywheel, and the per-flagship atomization checklist
- **[Headless CMS Guide](references/headless-cms.md)**: CMS selection, content modeling for marketing, editorial workflows, platform comparison (Sanity, Contentful, Strapi)
---
## Related Skills
- **copywriting**: For writing individual content pieces
- **seo-audit**: For technical SEO and on-page optimization
- **ai-seo**: For optimizing content for AI search engines and getting cited by LLMs
- **programmatic-seo**: For scaled content generation
- **site-architecture**: For page hierarchy, navigation design, and URL structure
- **emails**: For email-based content
- **social**: For social media content, content atomization, and repurposing execution
- **launch**: For the ORB channel-type playbook and launch-day distribution
FILE:evals/evals.json
{
"skill_name": "content-strategy",
"evals": [
{
"id": 1,
"prompt": "Help me build a content strategy for our B2B SaaS product. We sell expense management software to finance teams at companies with 50-500 employees. We currently have no blog and want to start from scratch.",
"expected_output": "Should check for product-marketing.md first. Should establish content pillars (3-5 core topic areas). Should map content types by buyer stage (awareness → consideration → decision → implementation). Should identify keyword research opportunities by buyer stage. Should recommend a mix of searchable (SEO-driven) and shareable (thought leadership, data) content. Should use the prioritization scoring framework (customer impact 40%, content-market fit 30%, search potential 20%, resources 10%). Should provide an initial content calendar or publishing cadence. Should recommend content types appropriate for starting from scratch.",
"assertions": [
"Checks for product-marketing.md",
"Establishes 3-5 content pillars",
"Maps content by buyer stage (awareness through implementation)",
"Includes keyword research by buyer stage",
"Recommends mix of searchable and shareable content",
"Uses prioritization scoring framework",
"Provides publishing cadence or calendar",
"Recommends appropriate starting content types"
],
"files": []
},
{
"id": 2,
"prompt": "We have 200+ blog posts but traffic has been flat for a year. Our content feels random — no clear strategy. How do we fix this?",
"expected_output": "Should diagnose the 'random content' problem. Should recommend a content audit process to evaluate existing posts. Should introduce content pillars and topical clustering to organize the existing library. Should identify hub-and-spoke opportunities from existing content. Should recommend which posts to update, consolidate, or retire. Should use the prioritization framework to plan next steps. Should address topical authority building through clusters.",
"assertions": [
"Diagnoses the 'random content' problem",
"Recommends content audit for existing posts",
"Introduces content pillars and topical clustering",
"Identifies hub-and-spoke opportunities",
"Recommends update, consolidate, or retire decisions",
"Uses prioritization framework",
"Addresses topical authority building"
],
"files": []
},
{
"id": 3,
"prompt": "what kind of content should we be creating? we're a developer tool (API testing platform) and our audience is backend developers and QA engineers",
"expected_output": "Should trigger on casual phrasing. Should recommend content types appropriate for a developer audience: technical tutorials, documentation-style guides, use-case content, template/example libraries, data-driven benchmarks. Should note that developer audiences prefer depth, accuracy, and practical value over marketing fluff. Should suggest content pillars aligned with developer interests. Should use the ideation sources framework (keyword data, community forums like Stack Overflow/Reddit, competitor gaps).",
"assertions": [
"Triggers on casual phrasing",
"Recommends content types for developer audience",
"Emphasizes technical depth and practical value",
"Notes developers prefer substance over marketing",
"Suggests content pillars for developer tool",
"Uses ideation sources framework",
"Mentions developer community channels"
],
"files": []
},
{
"id": 4,
"prompt": "How should we prioritize which content to create first? We have a list of 50 blog post ideas but limited resources — one content marketer writing 2 posts per week.",
"expected_output": "Should apply the prioritization scoring framework: customer impact (40%), content-market fit (30%), search potential (20%), resources required (10%). Should help score or rank the content ideas using this framework. Should recommend focusing on high-impact, lower-effort content first. Should consider the buyer stage distribution (don't write only top-of-funnel). Should provide a practical workflow for the single content marketer to use going forward.",
"assertions": [
"Applies prioritization scoring framework with weights",
"Explains each scoring dimension",
"Recommends focusing on high-impact, lower-effort first",
"Considers buyer stage distribution",
"Provides practical workflow for limited resources"
],
"files": []
},
{
"id": 5,
"prompt": "We want to build topical authority in 'employee engagement.' What does a content cluster look like for this topic?",
"expected_output": "Should apply the hub-and-spoke content cluster model. Should design a pillar page for 'employee engagement' (comprehensive, 3000+ word guide). Should identify 8-15 supporting spoke articles targeting long-tail keywords related to employee engagement. Should map the internal linking structure between hub and spokes. Should address keyword research for the cluster. Should recommend content types for each piece (guide, how-to, template, data-driven, etc.).",
"assertions": [
"Applies hub-and-spoke content cluster model",
"Designs a pillar page for the core topic",
"Identifies 8-15 supporting spoke articles",
"Maps internal linking between hub and spokes",
"Addresses keyword research for the cluster",
"Recommends content types for each piece"
],
"files": []
},
{
"id": 6,
"prompt": "Can you write a blog post about remote work best practices for our HR software blog?",
"expected_output": "Should recognize this is a copywriting/content creation task, not a content strategy task. Should defer to or cross-reference the copywriting skill for writing individual pieces of content. May provide strategic context (where this fits in the content strategy, keyword targeting, audience) but should make clear that copywriting is the right skill for writing the actual content.",
"assertions": [
"Recognizes this as content creation, not strategy",
"References or defers to copywriting skill",
"Does not attempt to write the full blog post",
"May provide strategic context for the piece"
],
"files": []
},
{
"id": 7,
"prompt": "Our #1 content goal this quarter is earning backlinks for domain authority. I'm deciding between commissioning an original research report, writing another ultimate guide, or building out a statistics roundup page for our category. Which should we prioritize and why?",
"expected_output": "Should apply the Link-Earning Formats data: statistics/data roundups earn ~4.25x their page share of backlinks while original research earns ~0.80x and ultimate guides ~0.77x, so for a backlinks-specific goal the stats roundup wins. Should label the data as a single vendor study (Foundation Inc., 2026, B2B SaaS) and treat it as directional. Should explain the mechanism — writers cite whatever makes citation easiest, and original research is often cited via roundups that aggregate it — and recommend that if they do run original research later, they pair it with their own stat-roundup page of citable one-liners. Should note stat pages are also an AI-citation play (ai-seo) and that guides/research still earn their keep on other jobs (rankings, conversions, brand).",
"assertions": [
"Recommends the statistics roundup page for the backlink-specific goal, citing the format multipliers",
"Labels the Foundation data as a single vendor study and directional, not a law",
"Explains the citation-ease mechanism and the pairing move (research + own stat-roundup of its findings)",
"Notes the other formats are judged by different jobs rather than calling them worthless"
],
"files": []
},
{
"id": 8,
"prompt": "We publish one good blog post a week but nobody reads it — we just post the link once on Twitter and LinkedIn and move on. How should we think about getting our content actually seen, and how should we balance what we produce?",
"expected_output": "Should reframe content as brand surface area and each piece as its own launch — creating is only half the job, distribution is the other half. Should introduce 'Create Once, Distribute Twice': one flagship piece repurposed/atomized across channels rather than a single link-drop, with distribution hooks (standalone subheads, modular sections, pull quotes) designed in at creation time. Should diagnose the failure modes at play — spray-and-pray / posting once with no repurposing, and the risk of platform dependency and the ownership paradox (over-investing in rented channels vs owned). Should present the ORB framework as a discovery->engagement->conversion funnel (borrowed -> rented -> owned) routing attention back to owned assets, and reference the launch skill (ORB playbook) and social skill (atomization execution) rather than re-deriving them. Should recommend a calendar balance (60% searchable / 30% shareable / 10% experimental) as a starting ratio. May reference the Content Distribution Flywheel and per-format execution discipline (e.g., 10 titles, ~5 editing passes for a blog post).",
"assertions": [
"Reframes content as brand surface area / each piece as its own launch and names distribution as the missing half",
"Introduces Create Once, Distribute Twice with atomization and creation-time distribution hooks",
"Names failure modes: spray-and-pray, platform dependency, ownership paradox",
"Presents ORB (borrowed/rented/owned) as a discovery-to-conversion funnel routing back to owned, cross-linking launch and social",
"Recommends the 60/30/10 calendar split as a starting ratio",
"May reference the Content Distribution Flywheel or per-format execution discipline"
],
"files": []
}
]
}
FILE:references/content-distribution.md
# Content Distribution Spine
The "distribute" half of content strategy. Creating a great piece is table stakes; the leverage is in getting it seen. This reference expands the **Create Once, Distribute Twice** section of the skill.
Cross-links: ORB channel-type playbook lives in **launch**; atomization/repurposing workflows (podcast → clips, blog → thread) live in **social**. This file consolidates the strategy that ties them together—don't re-derive ORB from scratch here.
## Create Once, Distribute Twice
One exceptional piece, reformatted across channels—not a fresh piece per platform. The math is simple: a flagship piece plus ten repurposed cuts reaches far more people than eleven mediocre native posts, at a fraction of the effort.
The discipline is **designing the piece to be distributed**:
- Write subheads that read as standalone social posts.
- Structure sections modularly so they can be lifted out and stand alone.
- Pre-identify the pull quotes, stats, and frames you'll turn into graphics or short clips.
- Know the atomized outputs before you write, so the source piece contains them.
Treat the flagship as the master; every channel gets a cut derived from it.
## The ORB Framework as a Funnel
Own, Rent, Borrow—read as a discovery → engagement → conversion funnel:
| Layer | Channels | Funnel role | You control |
|---|---|---|---|
| **Borrowed** | Podcasts, guest posts, partnerships, PR, other people's audiences | Discovery / breakthrough | Nothing—it's a loan |
| **Rented** | Social platforms, ad networks, marketplaces | Engagement / reach | The content, not the audience or algorithm |
| **Owned** | Email list, blog, community, app | Conversion / retention | Everything—the durable asset |
The strategic move: use borrowed and rented reach to funnel strangers into owned channels where you can convert and retain them. Borrowed and rented are rented land; owned is the only asset you keep.
## The Ownership Paradox
Most teams invert the priority: they spend ~90% of effort on borrowed and rented channels they don't control, and neglect the owned assets that actually convert. The paradox is that the channels getting the least attention (email, blog, community) are the ones that compound and can't be revoked. Rebalance toward owned as the destination for all upstream effort.
## Failure Modes
- **Spray-and-pray** — publishing across every platform with no flagship and no repurposing system. Effort scatters; nothing compounds; each post starts from zero.
- **Platform dependency** — building your audience on rented land. Facebook organic reach collapsed from ~20% to under 2% as the platform monetized. Any rented channel can throttle, deprioritize, or de-platform you with no recourse. The lesson isn't "avoid rented"—it's "never let rented be the endpoint."
## Platform Half-Lives
Content decays at wildly different rates by channel. Match the piece to the channel's shelf life:
| Channel | Rough half-life | Implication |
|---|---|---|
| Twitter/X post | Minutes–hours | Post often; repost; thread for reach |
| Instagram / Facebook | ~a day | Frequent cadence; stories are ephemeral by design |
| LinkedIn post | ~a day, longer for strong performers | Fewer, higher-effort posts |
| TikTok / Reels / Shorts | Days–weeks (algorithmic resurfacing) | Evergreen hooks can re-surface long after posting |
| YouTube video | Months–years | Search-driven; compounds like a blog post |
| Blog post / SEO | Years | The long tail; the compounding asset |
| Email | Sent once, but archived / repurposable | One-shot attention; harvest into other formats |
Short half-life channels reward frequency and repetition; long half-life channels reward depth and evergreen framing. Owned, long-half-life formats (blog, YouTube, email archive) are where distribution effort compounds.
## The Content Distribution Flywheel
Distribution isn't a linear checklist—it's a loop that feeds itself:
1. **Create** one exceptional flagship piece (guide, video, podcast, original research), with distribution hooks built in.
2. **Atomize** it into channel-native cuts—clips, threads, carousels, quote graphics, email, subhead-posts.
3. **Distribute** across owned → rented → borrowed, routing everything back to owned.
4. **Engage** with the responses; capture the questions, objections, and reactions.
5. **Feed back** — the engagement surfaces the next flagship topic (what resonated, what got asked), and top-performing atoms signal what to make more of.
Each turn of the loop lowers the cost of the next piece (you learn what lands) and grows the owned audience that amplifies it. The flywheel is why consistent distributors pull away from one-off publishers over time.
## Atomization Checklist (per flagship)
For each major piece, produce (see **social** for the platform-native execution):
- [ ] 3–5 standalone social posts from the subheads/key points
- [ ] 1 thread (Twitter/X) or carousel (LinkedIn/Instagram) of the core argument
- [ ] 2–4 short-form video clips (if source is video/podcast)
- [ ] 1–2 quote or stat graphics
- [ ] 1 email to the owned list linking the flagship
- [ ] Repost/reshare schedule across the piece's half-life (don't post once and move on)
## Related
- **launch** — ORB channel-type playbook and launch-day distribution
- **social** — atomization/repurposing workflows and platform-native execution
- **emails** — the owned channel that converts distributed attention
- **ai-seo** — making owned content citable by LLMs (another distribution surface)
FILE:references/headless-cms.md
# Headless CMS Guide
Reference for choosing, modeling, and implementing a headless CMS for marketing content.
## When to Use This Reference
Use this when selecting a CMS for a new project, designing content models for marketing sites, setting up editorial workflows, or connecting CMS content to programmatic pages.
---
## Headless vs Traditional CMS
A headless CMS separates content management from presentation. Content is stored in a structured backend and delivered via API to any frontend.
### When Headless Makes Sense
- Multiple frontends consume the same content (web, mobile, email)
- Developers want full control over the frontend stack
- Content needs to be reused across channels
- You're building with a modern framework (Next.js, Remix, Astro)
- Marketing needs structured, reusable content blocks
### When Traditional Works Better
- Small team with no dedicated developers
- Simple blog or brochure site
- WYSIWYG editing is a hard requirement
- Budget is tight and WordPress/Webflow does the job
### Decision Checklist
| Factor | Headless | Traditional |
|--------|----------|-------------|
| Multi-channel delivery | Yes | Limited |
| Developer control | Full | Constrained |
| Non-technical editing | Requires setup | Built-in |
| Time to launch | Longer | Faster |
| Content reuse | Native | Manual |
| Hosting flexibility | Any frontend | Platform-dependent |
---
## Content Modeling for Marketing
### Core Principles
1. **Think in types, not pages.** A "Landing Page" is a content type with fields — not an HTML file. This lets you reuse components across pages.
2. **Separate content from presentation.** Store the headline text, not the styled headline. Presentation belongs in the frontend.
3. **Design for reuse.** If testimonials appear on 5 pages, create a Testimonial type and reference it — don't duplicate.
4. **Keep models flat.** Deeply nested structures are hard to query and maintain. Prefer references over nesting.
### Common Marketing Content Types
| Type | Key Fields | Notes |
|------|-----------|-------|
| **Landing Page** | title, slug, hero, sections[], seo | Modular sections for flexibility |
| **Blog Post** | title, slug, body, author, category, tags, publishedAt, seo | Rich text or Portable Text body |
| **Case Study** | title, customer, challenge, solution, results, metrics[], logo | Link to related products/features |
| **Testimonial** | quote, author, role, company, avatar, rating | Reference from landing pages |
| **FAQ** | question, answer, category | Group by category for programmatic pages |
| **Author** | name, bio, avatar, social links | Reference from blog posts |
| **CTA Block** | heading, body, buttonText, buttonUrl, variant | Reusable across pages |
### SEO Fields Checklist
Every page-level content type needs:
- `metaTitle` — 50-60 characters
- `metaDescription` — 150-160 characters
- `ogImage` — 1200x630px social preview
- `slug` — URL path segment
- `canonicalUrl` — optional override
- `noIndex` — boolean for excluding from search
- `structuredData` — optional JSON-LD override
---
## Editorial Workflows
### Draft → Review → Publish Cycle
1. **Draft** — Author creates or edits content
2. **Review** — Editor reviews for accuracy, brand voice, SEO
3. **Approve** — Stakeholder signs off
4. **Schedule** — Set publish date/time
5. **Publish** — Content goes live via API
### Preview APIs
All major headless CMS platforms support draft previews:
- **Sanity**: Real-time preview with `useLiveQuery` or Presentation tool
- **Contentful**: Preview API (`preview.contentful.com`) with separate access token
- **Strapi**: Draft & Publish system with `status=draft` query parameter (v5; replaces v4's `publicationState`)
Set up a preview route in your frontend (e.g., `/api/preview`) that authenticates and renders draft content.
### Roles and Permissions
| Role | Can Create | Can Edit | Can Publish | Can Delete |
|------|:----------:|:--------:|:-----------:|:----------:|
| Author | Yes | Own | No | Own drafts |
| Editor | Yes | All | Yes | Drafts |
| Admin | Yes | All | Yes | All |
Exact permission models vary by platform. Sanity uses role-based access. Contentful has space-level roles. Strapi has granular RBAC.
---
## Platform Comparison
| Feature | Sanity | Contentful | Strapi |
|---------|--------|------------|--------|
| Hosting | Cloud (managed) | Cloud (managed) | Self-hosted or Cloud |
| Query Language | GROQ | REST / GraphQL | REST / GraphQL |
| Free Tier | Generous | Limited | Open source (free) |
| Real-time Collab | Yes (built-in) | Limited | No |
| Best For | Developer flexibility | Enterprise multi-locale | Budget / self-hosted |
| Content Modeling | Schema-as-code | Web UI | Web UI or code |
| Media Handling | Built-in DAM | Built-in | Plugin-based |
### Sanity
**Strengths**: GROQ query language is powerful and flexible. Schema defined in code (version-controlled). Real-time collaborative editing. Portable Text for rich content. Generous free tier.
**Considerations**: Steeper learning curve for non-developers. Studio customization requires React knowledge. Vendor lock-in on GROQ queries.
**Marketing fit**: Best when developers and marketers collaborate closely. Strong for content-heavy sites with complex models.
### Contentful
**Strengths**: Mature enterprise platform. Excellent multi-locale support. Strong ecosystem of integrations. Composable content with Studio. Well-documented APIs.
**Considerations**: Pricing scales with content types and locales. Two separate APIs (Delivery and Management). Rate limits can be tight on lower plans.
**Marketing fit**: Best for enterprises with multi-market content needs. Good when you need established vendor reliability.
### Strapi
**Strengths**: Open source, self-hosted option. Full control over data. No per-seat pricing. Customizable admin panel. Plugin ecosystem. REST by default, GraphQL via plugin.
**Considerations**: Self-hosting means you handle infrastructure. Smaller ecosystem than Sanity/Contentful. V5 migration can be significant from V4.
**Marketing fit**: Best for teams with DevOps capability who want full control and no vendor lock-in. Good for budget-conscious projects.
### Others Worth Knowing
- **Hygraph** — GraphQL-native, strong for federation and multi-source content
- **Keystatic** — Git-based, good for developer-content hybrid workflows
- **Payload** — TypeScript-first, self-hosted, code-configured like Sanity
- **Builder.io** — Visual editor with headless backend, good for non-technical marketers
- **Prismic** — Slice-based content modeling, strong Next.js integration
---
## Integration with Marketing Skills
### Programmatic SEO
Use CMS as the data source for programmatic pages. Store structured data (FAQs, comparisons, city pages) as content types and generate pages from queries. See **programmatic-seo** skill.
### Copywriting
CMS content models enforce consistent structure. Define fields that match your copy frameworks (headline, subheadline, social proof, CTA). See **copywriting** skill.
### Site Architecture
URL structure, navigation hierarchy, and internal linking all depend on how content is organized in the CMS. Plan your content model and site architecture together. See **site-architecture** skill.
### Email Sequences
Pull CMS content into email templates for consistent messaging across web and email. Case studies, testimonials, and blog posts can feed email nurture sequences. See **emails** skill.
---
## Implementation Checklist
- [ ] Define content types based on page types and reusable blocks
- [ ] Add SEO fields to every page-level content type
- [ ] Set up preview/draft mode in your frontend
- [ ] Configure roles and permissions for your team
- [ ] Create sample content for each type before building frontend
- [ ] Set up webhook notifications for content changes (rebuild triggers)
- [ ] Document content guidelines for editors (field descriptions, character limits)
- [ ] Test content delivery performance (CDN, caching, ISR)
- [ ] Plan migration strategy if moving from existing CMS
---
## Relevant Integration Guides
- [Sanity](../../../tools/integrations/sanity.md) — GROQ queries, mutations, CLI
- [Contentful](../../../tools/integrations/contentful.md) — Delivery/Management APIs, publishing
- [Strapi](../../../tools/integrations/strapi.md) — REST CRUD, filters, document API
Lãnh đạo doanh thu B2B SaaS: dự báo doanh thu, mô hình bán hàng, chiến lược giá, NRR và mở rộng đội bán hàng.
---
name: "cro-advisor"
description: "Revenue leadership for B2B SaaS companies. Revenue forecasting, sales model design, pricing strategy, net revenue retention, and sales team scaling. Use when designing the revenue engine, setting quotas, modeling NRR, evaluating pricing, building board forecasts, or when user mentions CRO, chief revenue officer, revenue strategy, sales model, ARR growth, NRR, expansion revenue, churn, pricing strategy, or sales capacity."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: cro-leadership
updated: 2026-03-05
python-tools: revenue_forecast_model.py, churn_analyzer.py
frameworks: sales-playbook, pricing-strategy, nrr-playbook
---
# CRO Advisor
Revenue frameworks for building predictable, scalable revenue engines — from $1M ARR to $100M and beyond.
## Keywords
CRO, chief revenue officer, revenue strategy, ARR, MRR, sales model, pipeline, revenue forecasting, pricing strategy, net revenue retention, NRR, gross revenue retention, GRR, expansion revenue, upsell, cross-sell, churn, customer success, sales capacity, quota, ramp, territory design, MEDDPICC, PLG, product-led growth, sales-led growth, enterprise sales, SMB, self-serve, value-based pricing, usage-based pricing, ICP, ideal customer profile, revenue board reporting, sales cycle, CAC payback, magic number
## Quick Start
### Revenue Forecasting
```bash
python scripts/revenue_forecast_model.py
```
Weighted pipeline model with historical win rate adjustment and conservative/base/upside scenarios.
### Churn & Retention Analysis
```bash
python scripts/churn_analyzer.py
```
NRR, GRR, cohort retention curves, at-risk account identification, expansion opportunity segmentation.
## Diagnostic Questions
Ask these before any framework:
**Revenue Health**
- What's your NRR? If below 100%, everything else is a leaky bucket.
- What percentage of ARR comes from expansion vs. new logo?
- What's your GRR (retention floor without expansion)?
**Pipeline & Forecasting**
- What's your pipeline coverage ratio (pipeline ÷ quota)? Under 3x is a problem.
- Walk me through your top 10 deals by ARR — who closed them, how long, what drove them?
- What's your stage-by-stage conversion rate? Where do deals die?
**Sales Team**
- What % of your sales team hit quota last quarter?
- What's average ramp time before a new AE is quota-attaining?
- What's the sales cycle variance by segment? High variance = unpredictable forecasts.
**Pricing**
- How do customers articulate the value they get? What outcome do you deliver?
- When did you last raise prices? What happened to win rate?
- If fewer than 20% of prospects push back on price, you're underpriced.
## Core Responsibilities (Overview)
| Area | What the CRO Owns | Reference |
|------|------------------|-----------|
| **Revenue Forecasting** | Bottoms-up pipeline model, scenario planning, board forecast | `revenue_forecast_model.py` |
| **Sales Model** | PLG vs. sales-led vs. hybrid, team structure, stage definitions | `references/sales_playbook.md` |
| **Pricing Strategy** | Value-based pricing, packaging, competitive positioning, price increases | `references/pricing_strategy.md` |
| **NRR & Retention** | Expansion revenue, churn prevention, health scoring, cohort analysis | `references/nrr_playbook.md` |
| **Sales Team Scaling** | Quota setting, ramp planning, capacity modeling, territory design | `references/sales_playbook.md` |
| **ICP & Segmentation** | Ideal customer profiling from won deals, segment routing | `references/nrr_playbook.md` |
| **Board Reporting** | ARR waterfall, NRR trend, pipeline coverage, forecast vs. actual | `revenue_forecast_model.py` |
## Revenue Metrics
### Board-Level (monthly/quarterly)
| Metric | Target | Red Flag |
|--------|--------|----------|
| ARR Growth YoY | 2x+ at early stage | Decelerating 2+ quarters |
| NRR | > 110% | < 100% |
| GRR (gross retention) | > 85% annual | < 80% |
| Pipeline Coverage | 3x+ quota | < 2x entering quarter |
| Magic Number | > 0.75 | < 0.5 (fix unit economics before spending more) |
| CAC Payback | < 18 months | > 24 months |
| Quota Attainment % | 60-70% of reps | < 50% (calibration problem) |
**Magic Number:** Net New ARR × 4 ÷ Prior Quarter S&M Spend
**CAC Payback:** S&M Spend ÷ New Logo ARR × (1 / Gross Margin %)
### Revenue Waterfall
```
Opening ARR
+ New Logo ARR
+ Expansion ARR (upsell, cross-sell, seat adds)
- Contraction ARR (downgrades)
- Churned ARR
= Closing ARR
NRR = (Opening + Expansion - Contraction - Churn) / Opening
```
### NRR Benchmarks
| NRR | Signal |
|-----|--------|
| > 120% | World-class. Grow even with zero new logos. |
| 100-120% | Healthy. Existing base is growing. |
| 90-100% | Concerning. Churn eating growth. |
| < 90% | Crisis. Fix before scaling sales. |
## Red Flags
- NRR declining two quarters in a row — customer value story is broken
- Pipeline coverage below 3x entering the quarter — already forecasting a miss
- Win rate dropping while sales cycle extends — competitive pressure or ICP drift
- < 50% of sales team quota-attaining — comp plan, ramp, or quota calibration issue
- Average deal size declining — moving downmarket under pressure (dangerous)
- Magic Number below 0.5 — sales spend not converting to revenue
- Forecast accuracy below 80% — reps sandbagging or pipeline quality is poor
- Single customer > 15% of ARR — concentration risk, board will flag this
- "Too expensive" appearing in > 40% of loss notes — value demonstration broken, not pricing
- Expansion ARR < 20% of total ARR — upsell motion isn't working
## Integration with Other C-Suite Roles
| When... | CRO works with... | To... |
|---------|------------------|-------|
| Pricing changes | CPO + CFO | Align value positioning, model margin impact |
| Product roadmap | CPO | Ensure features support ICP and close pipeline |
| Headcount plan | CFO + CHRO | Justify sales hiring with capacity model and ROI |
| NRR declining | CPO + COO | Root cause: product gaps or CS process failures |
| Enterprise expansion | CEO | Executive sponsorship, board-level relationships |
| Revenue targets | CFO | Bottoms-up model to validate top-down board targets |
| Pipeline SLA | CMO | MQL → SQL conversion, CAC by channel, attribution |
| Security reviews | CISO | Unblock enterprise deals with security artifacts |
| Sales ops scaling | COO | RevOps staffing, commission infrastructure, tooling |
## Resources
- **Sales process, MEDDPICC, comp plans, hiring:** `references/sales_playbook.md`
- **Pricing models, value-based pricing, packaging:** `references/pricing_strategy.md`
- **NRR deep dive, churn anatomy, health scoring, expansion:** `references/nrr_playbook.md`
- **Revenue forecast model (CLI):** `scripts/revenue_forecast_model.py`
- **Churn & retention analyzer (CLI):** `scripts/churn_analyzer.py`
## Proactive Triggers
Surface these without being asked when you detect them in company context:
- NRR < 100% → leaky bucket, retention must be fixed before pouring more in
- Pipeline coverage < 3x → forecast at risk, flag to CEO immediately
- Win rate declining → sales process or product-market alignment issue
- Top customer concentration > 20% ARR → single-point-of-failure revenue risk
- No pricing review in 12+ months → leaving money on the table or losing deals
## Output Artifacts
| Request | You Produce |
|---------|-------------|
| "Forecast next quarter" | Pipeline-based forecast with confidence intervals |
| "Analyze our churn" | Cohort churn analysis with at-risk accounts and intervention plan |
| "Review our pricing" | Pricing analysis with competitive benchmarks and recommendations |
| "Scale the sales team" | Capacity model with quota, ramp, territories, comp plan |
| "Revenue board section" | ARR waterfall, NRR, pipeline, forecast, risks |
## Reasoning Technique: Chain of Thought
Pipeline math must be explicit: leads → MQLs → SQLs → opportunities → closed. Show conversion rates at each stage. Question any assumption above historical averages.
## Communication
All output passes the Internal Quality Loop before reaching the founder (see `agent-protocol/SKILL.md`).
- Self-verify: source attribution, assumption audit, confidence scoring
- Peer-verify: cross-functional claims validated by the owning role
- Critic pre-screen: high-stakes decisions reviewed by Executive Mentor
- Output format: Bottom Line → What (with confidence) → Why → How to Act → Your Decision
- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed.
## Context Integration
- **Always** read `company-context.md` before responding (if it exists)
- **During board meetings:** Use only your own analysis in Phase 2 (no cross-pollination)
- **Invocation:** You can request input from other roles: `[INVOKE:role|question]`
FILE:references/nrr_playbook.md
# NRR Playbook
Net Revenue Retention is the single most important metric for a SaaS company's health and valuation. A company with 120% NRR grows even if it closes zero new deals. A company with 80% NRR is filling a bucket with a hole in it.
---
## NRR Deep Dive
### The Fundamental Formula
```
NRR = (Opening MRR + Expansion MRR - Contraction MRR - Churned MRR) / Opening MRR
Example:
Opening MRR: $1,000,000
Expansion: +$150,000
Contraction: -$30,000
Churn: -$80,000
Closing MRR: $1,040,000
NRR = $1,040,000 / $1,000,000 = 104%
```
### NRR vs. GRR
| Metric | Formula | What It Tells You |
|--------|---------|------------------|
| **GRR** | (Opening - Contraction - Churn) / Opening | Retention floor — how much you keep without any expansion |
| **NRR** | (Opening + Expansion - Contraction - Churn) / Opening | Net health — expansion offsetting churn |
| **Logo Retention** | (Customers start - Customers churned) / Customers start | Volume retention, ignores revenue weight |
**GRR is the floor. NRR is the ceiling.**
If GRR is 80% and NRR is 105%, your expansion is covering 25 points of churn. That's fragile — any expansion slowdown turns NRR negative. The fix is GRR, not more upsell.
### Benchmarks by Segment
| Segment | Good GRR | Good NRR | Exceptional NRR |
|---------|---------|---------|----------------|
| SMB-focused | 80-85% | 95-105% | > 110% |
| Mid-Market | 85-90% | 105-115% | > 120% |
| Enterprise | 90-95% | 115-130% | > 140% |
Enterprise NRR can exceed 140% because large accounts expand substantially and rarely churn entirely — they may downgrade but full logo churn is rare if the product is embedded.
### NRR by Cohort
Don't just measure NRR across the full base — measure it by customer cohort (month of acquisition).
```
Jan 2024 Cohort:
Opening MRR (Jan 2024): $50,000
MRR at Jan 2025: $62,000
12-month NRR: 124%
Feb 2024 Cohort:
Opening MRR (Feb 2024): $45,000
MRR at Feb 2025: $38,000
12-month NRR: 84% ← problem cohort
```
Cohort analysis reveals:
- Whether a specific acquisition channel brings lower-quality customers
- Whether a product change or pricing shift affected retention
- Whether specific sales reps or time periods created bad-fit deals
---
## Churn Anatomy
Not all churn is equal. Know the breakdown before prescribing solutions.
### Churn Types
| Type | Definition | Primary Cause | Fix |
|------|-----------|--------------|-----|
| **Logo churn** | Customer cancels entirely | No value, poor fit, champion left, competitor | Root cause analysis, ICP tightening |
| **Revenue churn** | ARR lost (cancels + downgrades combined) | Same as logo + downgrade triggers | Address both volume and revenue |
| **Involuntary churn** | Failed payment, expired card | Billing friction | Dunning improvement (quick win: 20-30% recovery) |
| **Voluntary churn** | Active cancellation decision | Explicit dissatisfaction, competitor win | Exit interview + intervention program |
| **Contraction** | Downgrade, seat reduction | Overpurchased, budget cut, team reduction | Right-sizing program, annual contracts |
### Churn Root Cause Framework
Run this analysis quarterly on all churned accounts:
**Step 1: Categorize by reason**
- No value realized (never activated or adopted)
- Value realized but budget cut (external, not product)
- Switched to competitor (why? what did they offer?)
- Champion left company (relationship loss, not product failure)
- Company shutdown / acquisition (unavoidable)
**Step 2: Look for patterns**
- Which ICP signals predict churn? (company size, vertical, acquisition channel)
- Which product behaviors predict churn? (no login in 30 days, never completed onboarding)
- Which time periods have highest churn? (months 3, 6, 12 are typical cliff points)
**Step 3: Act on the patterns**
- ICP pattern → tighten qualification criteria
- Behavior pattern → build early warning health score
- Time cliff → build intervention playbooks for months 2, 5, 11
### Exit Interview Protocol
Talk to every churned customer if ACV > $10K. For smaller, do quarterly batch surveys.
Questions:
1. "What was the primary reason for your decision to cancel?"
2. "What would have needed to be true for you to stay?"
3. "What did you switch to, and what drove that decision?"
4. "Was there a specific moment when you decided to leave?"
Rules:
- CSM who owned the account should NOT conduct the exit interview (too much relationship bias)
- Use a neutral party or the VP CS
- Document verbatim, not paraphrased
- Feed patterns back to Product and Sales monthly
---
## Customer Health Scoring
A health score predicts churn 60-90 days before it happens. Without one, you're reactive.
### Health Score Components
Score each account 0-100 across weighted signals:
| Signal | Weight | Red (0-33) | Yellow (34-66) | Green (67-100) |
|--------|--------|-----------|---------------|---------------|
| **Product usage** (DAU/WAU, feature adoption depth) | 35% | < 20% seats active | 20-60% seats active | > 60% seats active |
| **Engagement** (QBR attendance, champion responsiveness) | 20% | No response 60+ days | 30-60 days | Active, < 30 days |
| **NPS / CSAT** | 20% | Score < 6 | Score 6-7 | Score 8-10 |
| **Support volume** (negative signal: high volume = friction) | 15% | > 10 tickets/month | 3-10/month | < 3/month |
| **Contract signals** (time to renewal, expansion in motion) | 10% | < 60 days to renewal, no expansion discussion | 60-90 days, passive | > 90 days, expansion active |
**Composite score:**
- 70-100: Healthy. Renewal confident. Identify expansion opportunity.
- 50-69: At-risk. CSM check-in required. Executive sponsor loop-in if < 60 days to renewal.
- 0-49: Red alert. Immediate intervention. VP CS or CEO call if strategic account.
### Health Score Automation
Trigger alerts automatically:
```
Score drops > 20 points in 30 days → CSM immediate outreach (same day)
No product login in 14 days → Automated email + CSM flag (within 24 hours)
Champion leaves company → Executive outreach (within 24 hours)
Support escalation → CSM loop-in (within 2 hours)
Renewal < 90 days + score < 60 → VP CS review (weekly)
Seat utilization < 30% → Adoption intervention playbook triggered
```
### Leading Indicators vs. Lagging Indicators
| Leading (predict future churn) | Lagging (confirm past churn) |
|-------------------------------|------------------------------|
| Login frequency declining | Cancellation submitted |
| Feature adoption stalling at basic level | Non-renewal at contract end |
| NPS score trend (not just snapshot) | Downgrade executed |
| No QBR scheduled in 90+ days | Champion departure |
| Support escalations increasing | Competitor mentioned in support |
Build your health score from leading indicators. Lagging indicators tell you what already happened.
---
## Expansion Revenue Strategies
Expansion is cheaper than acquisition. CAC for expansion is typically 20-30% of new logo CAC.
### Expansion Motion 1: Seat Expansion
**Trigger signals:**
- Usage by unlicensed users (shared logins, "can you add my colleague?")
- Team growth visible on LinkedIn (company hiring in target department)
- Champion promotes to a new role with bigger team
- Power users at license limit consistently
**Playbook:**
1. Pull monthly usage report showing which features unlicensed users are using
2. Frame as: "Your team is getting value from X — you could be capturing that for the full team"
3. Offer a team expansion proposal at renewal + 10% volume discount for seat adds
4. Never penalize users for sharing logins before the conversation — that's a data asset
### Expansion Motion 2: Upsell (Tier Upgrade)
**Trigger signals:**
- Customer consistently hitting usage/feature limits
- Security or compliance requirement that requires higher tier
- New stakeholder joining who needs admin controls
- API usage growing rapidly (engineering team engagement)
**Playbook:**
1. Build a "value realized" report before the upsell conversation (ROI proof)
2. Use QBR as the venue: "You've achieved X. Here's what's possible at the next level."
3. Frame the upgrade as unlocking more of what's already working
4. Time to renewal: start upsell conversation 90-120 days before renewal
### Expansion Motion 3: Cross-sell
**Trigger signals:**
- Strategic account with adjacent problem your product can solve
- New product launch that complements existing usage
- Customer explicitly asks about a capability in your roadmap or adjacent product
**Playbook:**
1. Land with core product; build relationship and prove value
2. Cross-sell only after health score is green and NPS > 7
3. Introduce the new product through a champion, not a cold pitch
4. Pilot pricing: bundle into renewal at modest uplift vs. separate sale
5. Cross-sell owner: CSM or AE (define explicitly — joint ownership = no ownership)
### Expansion Sequencing
Don't try all three simultaneously. Sequence matters:
```
Month 0-3: Activation focus — ensure core value delivered
Month 3-6: Seat expansion — grow usage within existing team
Month 6-9: Upsell conversation — unlock advanced features
Month 9-12: Cross-sell OR renewal + multi-year lock-in
```
### NRR Modeling
Target breakdown for 115% NRR:
```
GRR: 88% (12% lost to churn/contraction)
Expansion rate: 27% (upsell + cross-sell + seat expansion)
NRR: 88% + 27% = 115%
To reach 120% NRR:
Option A: Improve GRR to 92% (reduce churn), keep expansion at 28%
Option B: Keep GRR at 88%, improve expansion to 32%
Option C: Both, incrementally
Option A is usually easier and more durable. Fix the hole first.
```
---
## Customer Success Integration
CS and Revenue are not separate functions. NRR lives at their intersection.
### CS Team Structure (aligned to NRR)
| CS Model | When to Use | NRR Focus |
|----------|------------|-----------|
| **High-touch CSM** | ACV > $25K | Named accounts, QBRs, executive relationships |
| **Tech-touch / pooled** | ACV $5K-25K | Automated health scoring, office hours, community |
| **Self-serve** | ACV < $5K | In-app guidance, knowledge base, email sequences |
**CSM coverage ratios:**
- High-touch: 1 CSM per $2M-4M ARR managed
- Tech-touch: 1 CSM per $5M-10M ARR managed
- Self-serve: Product and automation (no dedicated CSM)
### CS Compensation (aligned to NRR)
Don't pay CSMs a flat salary — align incentive to retention and expansion:
```
CS compensation structure:
Base: 70% of OTE
Variable: 30% of OTE
Variable tied to:
GRR / NRR vs. target (50% of variable)
Health score improvement (25% of variable)
Expansion ARR facilitated (25% of variable)
Do NOT pay CS commission on expansion ARR the same way AEs earn it.
This creates conflict: CS will push expansion before the customer is ready.
Instead, bonus for expansion milestones — it's a different incentive structure.
```
### QBR (Quarterly Business Review) Framework
QBRs are the primary vehicle for expansion and churn prevention in enterprise accounts.
**QBR agenda (60-90 minutes):**
1. **Their goals, our progress** — review what they said success looked like at kickoff (10 min)
2. **Usage and adoption data** — product metrics presented in business language, not feature language (15 min)
3. **Value delivered** — ROI proof: time saved, revenue influenced, risk reduced (10 min)
4. **Challenges and blockers** — what's preventing more adoption? (10 min)
5. **Roadmap preview** — upcoming features relevant to their use case (10 min)
6. **Next 90 days** — joint success plan with owner and due dates (10 min)
7. **Expansion opportunity** — if health score is green and timing is right (10 min)
**QBR anti-patterns:**
- Leading with your product roadmap (they don't care; start with their results)
- Bringing too many people from your side without matching seniority
- Presenting at a VP without bringing the economic buyer
- Skipping QBRs for "healthy" accounts (health can change fast)
- No confirmed next step at the end
---
## Cohort-Based Retention Analysis
Aggregate NRR hides the signal. Cohort analysis reveals it.
### Retention Curve Analysis
Plot retention by months since acquisition for each quarterly cohort:
```
Month 0: 100% (starting revenue)
Month 3: First cliff — early adopters who didn't activate churn here
Month 6: Second cliff — customers who never expanded, running out of runway
Month 12: Renewal cliff — annual contract renewal decision
Month 18: Mature customers — churn rate stabilizes significantly
Healthy curve: Drops sharply in months 1-3, flattens after month 6
Problem curve: Continues declining linearly through month 12+ (no value anchor)
```
### Reading Cohort Data
| Pattern | Interpretation | Action |
|---------|---------------|--------|
| Early churn (months 1-3) | Onboarding / activation failure | Fix time-to-value, improve onboarding |
| Mid-cycle churn (months 4-8) | Value not deepening | Adoption program, check product fit |
| Annual renewal churn (month 12) | Buying committee didn't renew | Executive engagement, earlier renewal process |
| Flat after month 6 | Sticky product, low expansion | Increase upsell motion |
| Growing after month 6 | Expansion working | Scale the upsell playbook |
### Cohort Segmentation Variables
Slice retention cohorts by:
- **Acquisition channel** (inbound vs. outbound vs. PLG vs. partner)
- **Sales rep** (which reps close durable deals vs. churny deals)
- **Deal size** (SMB churn rate typically 2-3x enterprise)
- **Industry vertical** (some verticals have structurally higher churn)
- **Product tier at signup** (self-serve → converted vs. directly contracted)
- **Geographic market** (international markets often have different retention profiles)
The most actionable finding is usually by acquisition channel or sales rep — both are directly controllable.
### Churn Prevention Intervention Playbooks
**Playbook 1: Low Activation (no login in first 14 days)**
```
Day 7: Automated email: "Getting started" + specific next step
Day 14: CSM outreach: "I noticed you haven't logged in — can I help?"
Day 21: Escalate to CSM manager if no response
Day 30: Executive outreach for ACV > $25K; flag as at-risk
```
**Playbook 2: Usage Cliff (DAU drops > 50% in 30 days)**
```
Trigger: Automated health score alert
Day 1: CSM reviews usage report, identifies likely cause
Day 2: CSM outreach: "We noticed your team's usage changed — is everything okay?"
Day 7: If no response: schedule 30-min call with champion
Day 14: If unresponsive: VP CS loop-in + executive reach out
```
**Playbook 3: Champion Departure**
```
Trigger: LinkedIn alert or internal report of champion leaving
Day 1: Email to departed champion (warm handoff ask)
Day 1: Email to new stakeholder (introduction from AE or VP CS)
Day 3: Schedule onboarding call for new stakeholder
Day 14: QBR with new stakeholder to establish relationship
Day 30: Health score review — flag if engagement hasn't recovered
```
**Playbook 4: Pre-Renewal (90 days out, health score < 70)**
```
Day -90: CSM completes account health review, escalates if < 70
Day -75: Executive sponsor from vendor side joins renewal call
Day -60: Value delivered report prepared (ROI proof)
Day -45: Renewal proposal sent with expansion option
Day -30: Follow-up on any open objections or requirements
Day -14: Final confirm or escalate to VP Sales
```
FILE:references/pricing_strategy.md
# Pricing Strategy
Pricing is not a one-time decision. It's an ongoing hypothesis about value and willingness to pay. Most SaaS companies are underpriced by 20-40%.
---
## Pricing Models
### Per Seat / User
**How it works:** Customer pays a fixed amount per user, per month or year.
**Best for:**
- Collaboration tools (everyone who uses it needs a license)
- Productivity software where value scales with users
- Products where you want viral / network growth within accounts
**Pricing structure:**
```
Starter: $15/user/month (1-10 users)
Professional: $30/user/month (11-100 users)
Enterprise: Custom (100+ users, negotiated)
```
**Pros:**
- Simple to understand and sell
- Revenue scales naturally with customer growth
- Predictable for customers (fixed monthly cost)
**Cons:**
- Customers negotiate volume discounts aggressively
- Discourages broad adoption if price is high (seat hoarding)
- Doesn't capture value for power users vs. light users
- Enterprises can negotiate $5/seat on a $25 product
**Watch for:** Customers sharing logins to avoid per-seat cost. Enforce with IP restrictions or SSO audit logs.
---
### Usage-Based Pricing (UBP)
**How it works:** Customer pays for what they consume — API calls, data processed, messages sent, compute hours, etc.
**Best for:**
- API companies, infrastructure, data platforms
- AI products (per-token, per-query pricing)
- Products where value scales non-linearly with usage
- Land-and-expand: low entry cost, grows with customer success
**Pricing structure:**
```
Free tier: First 10K API calls/month
Pay-as-you-go: $0.002 per API call
Committed use: $500/month for 500K calls (better rate)
Enterprise: Custom contract, committed volume discount
```
**Pros:**
- Customer pays in proportion to value received
- Low barrier to entry (customers start small, scale up)
- Natural expansion: customer success = revenue growth
- No "unused licenses" problem
**Cons:**
- Revenue is unpredictable for both you and the customer
- Hard to forecast; hard to budget for customer
- Customers may optimize to reduce usage (and your revenue)
- Complex billing; requires robust usage tracking infrastructure
**Usage-based pricing math:**
```
Unit cost (your COGS per unit): $0.0002 per API call
Target gross margin: 80%
Price = COGS / (1 - margin) = $0.0002 / 0.20 = $0.001 minimum
Add markup for value delivered above cost: $0.002 per call (10x markup at scale)
```
**Hybrid usage + seat approach:**
- Platform fee: $500/month (access, support, base features)
- Usage fee: $0.001 per API call above included 100K
---
### Flat Rate / Subscription
**How it works:** One price for full access, regardless of usage or users.
**Best for:**
- Simple products with limited feature differentiation
- Products where usage is predictable and bounded
- Customers who want budget certainty
- Early stage before you've figured out value segmentation
**Pros:**
- Simplest to sell and explain
- Easiest billing implementation
- Customers love budget predictability
**Cons:**
- Leaves money on the table for heavy users
- No natural expansion revenue mechanism
- Light users pay the same as power users (retention risk)
**When to move away from flat rate:**
- 20% of customers are using 80% of the product capacity
- Power users would clearly pay more; light users churn or underutilize
- You have a clear expansion story waiting to happen
---
### Tiered / Feature-Based
**How it works:** Multiple packages (Starter, Pro, Enterprise) with different feature sets and/or usage limits.
**Best for:**
- Multi-use-case products
- Different buyer types (individual vs. team vs. enterprise)
- Products with a natural upgrade path based on sophistication
**Structure (Good / Better / Best):**
```
Starter ($49/mo): Core features, 3 users, 10GB storage
Professional ($149/mo): Advanced features, 25 users, 100GB, API access
Business ($499/mo): All features, 100 users, 1TB, SSO, priority support
Enterprise (custom): Unlimited, custom integrations, SLA, dedicated CSM
```
**Tier design principles:**
- Starter tier: removes friction, proves value, not the revenue center
- Professional: the primary revenue tier; 60-70% of customers land here
- Enterprise: custom pricing allows you to capture maximum value
- Each tier upgrade should have an obvious "must-have" feature for the target buyer
**What to gate on each tier:**
| Feature Type | Where to Put It |
|-------------|----------------|
| Core product functionality | Starter (must be useful) |
| Collaboration features | Pro (drives team usage) |
| Admin, security, SSO | Business/Enterprise |
| API / integrations | Pro and above |
| SLAs, dedicated support | Enterprise only |
| Advanced analytics | Business/Enterprise |
---
### Hybrid Pricing
**How it works:** Combination of models (e.g., platform fee + per seat + usage).
**Example:**
```
Platform fee: $2,000/month (access, core features, admin console)
Per seat: $50/user/month (up to 200 users)
Usage overage: $0.10/action above 100K included actions
```
**When to use hybrid:**
- Enterprise customers want budget certainty (platform fee) but your value scales with usage
- You have different cost structures for different features
- Customers have very different usage patterns across the base
**Pros:** Captures value at multiple dimensions. Hybrid is most common in enterprise SaaS.
**Cons:** More complex to explain and bill. Sales training burden increases.
---
## Value-Based Pricing Methodology
Cost-plus pricing is a race to the bottom. Price on value, not cost.
### Step 1: Define the Economic Outcome
What business result does your product deliver? Be specific.
**Weak:** "We help companies save time"
**Strong:** "We reduce onboarding time for new enterprise software by 40%, saving 8 hours per employee"
Map to one of:
- **Revenue increase** — "Our customers close 25% more deals using our CRM intelligence"
- **Cost reduction** — "We eliminate 60% of manual data entry for finance teams"
- **Risk reduction** — "We reduce compliance violations by 90%, avoiding $500K+ in potential fines"
- **Time savings** — "CSMs spend 5 fewer hours per week on manual reporting"
### Step 2: Quantify Per Customer
Calculate the dollar value of the outcome for your average customer.
```
Example: Data entry automation product
Target customer: 50-person finance team
Manual data entry: 4 hours/person/week
Hours saved with product: 2.4 hours/person/week (60% reduction)
Fully loaded cost of finance analyst: $75/hour
Weekly savings: 50 employees × 2.4 hours × $75 = $9,000
Annual savings: $9,000 × 52 weeks = $468,000
```
### Step 3: Determine Willingness to Pay
Customers will typically pay 10-20% of the value delivered for software.
```
Annual value delivered: $468,000
Willingness to pay range: $46,800 - $93,600/year
Current market pricing: ~$60,000/year
Your pricing: $72,000/year (between median and upper WTP)
```
**Test your hypothesis:**
- Interview 5-10 customers: "If we charged $X/year, is that reasonable?"
- Van Westendorp Price Sensitivity Meter:
- "At what price is this too cheap to trust?"
- "At what price is this a good deal?"
- "At what price is this getting expensive but still worth it?"
- "At what price is this too expensive?"
### Step 4: Validate with Win Rate Analysis
```
Run this analysis quarterly:
Track win rate by price point (segmented if possible)
Win rate 30-40%: pricing is likely right
Win rate < 20%: price is too high OR value demonstration is broken
Win rate > 50%: you're underpriced
Note: Distinguish between "lost on price" and "lost on fit."
Lost on price + good ROI proof: test lower price or improve value story
Lost on fit: ICP problem, not pricing problem
```
---
## Packaging (Good / Better / Best)
### The Three-Package Framework
Packaging is not just about features. It's about serving different buyer personas with different budgets and needs.
**Buyer personas by tier:**
```
Starter → The individual contributor or small team trying to solve an immediate problem
- Low budget authority
- Low-friction purchase (credit card, self-serve)
- Needs quick time to value
Professional → The team manager or department head
- $10K-100K budget authority
- Works with inside sales
- Needs collaboration features and reporting
Enterprise → The VP or C-suite buyer
- Unlimited budget (but requires justification)
- Needs compliance, security, SLAs, dedicated support
- Long buying process, multiple stakeholders
```
### Packaging Design Rules
1. **Each tier must be useful on its own.** Starter can't be crippled—customers need to succeed.
2. **Upgrade triggers must be obvious.** When a customer hits a limit, the next tier should solve it clearly.
3. **Don't gate features that drive adoption.** Collaboration features gated in a low tier kill viral growth.
4. **Enterprise pricing is custom.** Show "Contact Sales" or a starting price. Don't publish a firm enterprise price—you'll anchor too low.
5. **Annual vs. monthly pricing:** Charge 15-25% more for monthly vs. annual. Incentivize annual prepay.
### Pricing Page Design
- Lead with the most popular tier (visually prominent)
- Show annual pricing by default (with toggle to monthly)
- Highlight one or two "recommended" plans
- Feature comparison table: minimize the number of rows (overwhelm = no decision)
- Show logos of customers on each tier (social proof by segment)
- Live chat for enterprise CTA, not "Contact Sales" form
---
## Pricing Experiments and Rollout
### Before You Change Pricing
**Internal checklist:**
- [ ] Validate new pricing with 5-10 current customers (interviews)
- [ ] Run a willingness-to-pay survey with 50+ prospects
- [ ] Model revenue impact: how many customers at new pricing are equivalent to current ARR?
- [ ] Get CFO sign-off on cash flow impact
- [ ] Prepare messaging for customers, website, sales team
- [ ] Set a rollout date 60-90 days out
### Testing Approaches
**Cohort testing (safest):**
- New signups see new pricing; existing customers are grandfathered
- Monitor: conversion rate, ACV, win rate, time-to-close
- Run for 90 days before full rollout
**A/B pricing test (higher stakes):**
- Half of new signups see price A, half see price B
- Risk: word gets out that prices differ (customer frustration)
- Use only on self-serve, where purchase is not sales-assisted
**Segment-specific rollout:**
- Change pricing in one segment (e.g., SMB) while holding enterprise steady
- Lower risk than full rollout; validate before expanding
### Pricing Rollout Plan
```
Day 0: Decision made, pricing document approved
Day -60: Internal communication to sales, CS, support
Day -45: Customer communication drafted and reviewed
Day -30: New pricing live on website for new customers
Day -30: Existing customer email sent (90-day grandfather period)
Day -30: Sales team trained, FAQ document ready
Day -14: Second reminder to existing customers
Day 0: Existing customers transition to new pricing
Day +30: Win rate analysis, NRR impact review
```
### Grandfathering Policy
- **Standard:** Grandfather existing customers at old price for 12 months
- **Aggressive:** 90 days grandfather, then new pricing applies (use if you're raising significantly)
- **Never:** Retroactive pricing changes with no notice. This is a churn trigger and brand damage.
Grandfathering message framing:
> "We're investing significantly in [feature areas]. As a valued customer, your pricing remains unchanged through [date]. After that, your new rate will be $X — still X% less than new customer pricing as a thank-you for your partnership."
---
## Competitive Pricing Analysis
### Mapping the Competitive Landscape
```
Step 1: List all direct competitors
Step 2: Find their public pricing (website, G2, Capterra)
Step 3: Secret shop their sales process for unpublished pricing
Step 4: Talk to customers who considered them ("What did they quote you?")
Step 5: Map to your packaging (apples-to-apples comparison)
Output: Competitive pricing matrix
You: $X/month per seat at Pro tier
Competitor A: $Y/month per seat at equivalent tier
Competitor B: Custom (enterprise only)
```
### Competitive Positioning by Price
| Your Position | Situation | Response |
|--------------|-----------|---------|
| Significantly cheaper | Unclear why | Raise prices or clarify differentiation |
| Slightly cheaper | Winning on price | Test raising price, monitor win rate |
| At market | Competing on features | Make sure differentiation is clear in sales |
| Slightly more expensive | Win rate healthy | Price is justified by value |
| Significantly more expensive | Win rate low | Improve value proof or re-examine ICP |
### When "They're Cheaper" Appears in Deals
**Coach your reps:**
1. "What makes [Competitor] worth choosing over the $X difference?" (reframe value, not price)
2. "If price were equal, which would you choose and why?" (understand true preference)
3. "What's the cost of not solving this problem in Q3?" (urgency + value)
4. "What's their implementation cost and time?" (TCO, not ACV)
**If price is truly the barrier:**
- Offer a pilot at reduced scope (not price) to prove value
- Multi-year deal with year-one discount
- Defer payment to match their budget cycle (start in Q4, bill in Q1)
- Confirm it's price and not a champion issue or lack of urgency
---
## When to Raise Prices
### Green Lights for a Price Increase
**Product signals:**
- Customer usage growing QoQ (product delivers real value)
- NPS consistently > 40
- Feature requests indicate you're solving critical workflows
- Customers measuring and can articulate ROI
**Market signals:**
- Win rate > 35% (strong signal of underpricing)
- Waitlist or high inbound conversion without price objections
- Competitors raising prices (market is moving up)
- You've added significant value (new features, integrations, uptime improvements)
**Business signals:**
- Gross margin below 70% (cost inflation requires pricing response)
- CAC payback > 24 months (need higher ACV to fix unit economics)
- Haven't raised prices in 2+ years (inflation alone justifies adjustment)
### How Much to Raise
**Conservative:** 10-15% increase. Low risk, low disruption.
**Standard:** 15-30% increase. Acceptable if value story is strong.
**Aggressive:** 30-50% increase. Only with major product investment or clear underprice.
**Repositioning:** 2-5x increase. Rare; requires moving to a new buyer persona.
**Rule:** If fewer than 20% of prospects mention price as a concern, you're underpriced. Test.
### Price Increase Execution
1. Raise new business pricing immediately on the website
2. Communicate to existing customers with 90 days notice
3. Grandfather for 12 months OR give a 10-15% loyalty discount on new price
4. Track: conversion rate (new business), churn rate (existing), expansion ARR impact
5. Monitor win rate for 60 days post-increase; adjust if win rate drops > 5 points
**What not to do:**
- Don't apologize for raising prices
- Don't over-explain the justification (confident framing wins)
- Don't let sales reps negotiate discounts back to old pricing "just this once"
- Don't raise prices and remove features simultaneously
FILE:references/sales_playbook.md
# Sales Playbook
Frameworks for building, running, and scaling a B2B SaaS sales organization.
---
## Sales Process Design
A sales process is a repeatable series of steps that takes a prospect from first contact to closed revenue. Without it, you have individual heroics, not a scalable machine.
### The Core Funnel
```
Lead Generation → Qualification → Discovery → Demo → Trial / POC → Proposal → Negotiation → Close → Handoff
```
Each stage has a clear entry criterion, exit criterion, and owner.
### Stage Definitions
#### Stage 0: Lead / Suspect
- **Entry:** Contact exists in CRM with basic firmographic data
- **Owner:** Marketing or SDR
- **Exit criterion:** Meets ICP criteria (company size, industry, tech stack)
- **Action:** Research, prioritize, add to outbound sequence
#### Stage 1: Prospecting / Outreach
- **Entry:** ICP-qualified account, no contact yet
- **Owner:** SDR or AE (depending on model)
- **Exit criterion:** Meeting booked with a qualified contact
- **Action:** Multi-channel outreach (email + call + LinkedIn), 8-12 touch sequence
- **Key metric:** Meeting booked rate (benchmark: 2-5% of outbound contacts)
#### Stage 2: Discovery
- **Entry:** First meeting confirmed
- **Owner:** AE (SDR hands off or joins)
- **Exit criterion:** Confirmed: pain, budget range, decision process, timeline
- **Action:** Ask questions. Listen. Map the org. Don't pitch yet.
- **Key metric:** Discovery-to-demo rate (benchmark: 60-80% proceed)
**Discovery question framework:**
```
Situation: "How do you currently handle [problem area]?"
Problem: "What's the impact when [pain point] happens?"
Implication: "If this continues, what does that mean for [business goal]?"
Need-payoff: "If we solved this, what would that be worth to you?"
```
#### Stage 3: Demo / Solution Presentation
- **Entry:** Confirmed pain and fit from discovery
- **Owner:** AE (+ SE for complex products)
- **Exit criterion:** Prospect agrees to evaluate / trial; next step defined
- **Action:** Show the workflow that solves their specific pain (not a feature tour)
- **Key metric:** Demo-to-trial/proposal rate (benchmark: 40-60%)
**Demo structure:**
1. Recap their pain (show you listened) — 5 min
2. Show the "aha moment" (fastest path to value) — 10 min
3. Walk the specific workflow they described — 15 min
4. Handle objections, confirm fit — 5 min
5. Define clear next step (date, owners, criteria) — 5 min
Never show features they didn't ask for. Every additional feature is noise until they have a reason to care.
#### Stage 4: Trial / POC
- **Entry:** Prospect commits to evaluate with real data/use case
- **Owner:** AE + CSM or SE
- **Exit criterion:** Success criteria met, POC success confirmed
- **Action:** Define success criteria upfront (in writing). Set a tight timeframe (2-4 weeks max).
- **Key metric:** POC-to-proposal rate (benchmark: 50-70%)
**POC setup requirements:**
```
Before any POC:
□ Signed NDA
□ Written success criteria ("We'll move forward if X happens")
□ Named champion who owns the evaluation
□ Executive sponsor identified
□ Defined timeline with end date
□ Agreed next step if criteria are met
```
If you can't get written success criteria, you don't have a real opportunity. You have a "we'll see."
#### Stage 5: Proposal / Pricing
- **Entry:** POC success OR strong discovery fit for simple products
- **Owner:** AE
- **Exit criterion:** Proposal received, timeline to decision confirmed
- **Action:** Present in a live call, never email a proposal cold
- **Key metric:** Proposal-to-negotiation rate (benchmark: 50-75%)
**Proposal structure:**
1. Problem statement (their words, not yours)
2. Proposed solution (mapped to their workflow)
3. ROI summary (value delivered vs. investment)
4. Pricing options (give 2-3 options; anchors the decision)
5. Next steps with dates
#### Stage 6: Negotiation
- **Entry:** Verbal intent to proceed, price/terms discussion begins
- **Owner:** AE (+ VP Sales for large deals)
- **Exit criterion:** Mutual agreement on terms; contract sent
- **Action:** Never discount before they ask. Discount on scope, not on margin.
- **Key metric:** Negotiation win rate (benchmark: 70-85%)
**Negotiation principles:**
- Get something for everything you give. Discount → multi-year. Fast close → early pay discount.
- Don't negotiate against yourself. Silence after an offer is not rejection.
- Know your walk-away before you enter. If you don't have a BATNA, you have no leverage.
- Legal/procurement delay ≠ deal death. Keep the champion engaged.
#### Stage 7: Close
- **Entry:** Signed contract or PO received
- **Owner:** AE
- **Exit criterion:** Contract countersigned, kickoff date set
- **Action:** Celebrate with the customer. Immediately introduce CSM.
- **Key metric:** Average close rate (closed won ÷ all closed = won + lost)
#### Stage 8: Handoff to Customer Success
- **Entry:** Deal closed
- **Owner:** AE + CSM
- **Exit criterion:** Customer has met their assigned CSM, kickoff scheduled
- **Action:** Internal handoff call with AE + CSM. AE shares: deal context, key stakeholders, use case, success criteria, any promises made during the sale.
**Handoff document (AE fills before first CS meeting):**
```
Account: [name]
ACV: $X
Close date: [date]
Primary contact: [name, title, email]
Economic buyer: [name, title]
Use case: [specific workflow]
Success criteria: [what they said good looks like in 90 days]
Promises made: [anything specific committed during sale]
Risk flags: [competitive, budget, champion strength]
```
---
## MEDDPICC Qualification Framework
MEDDPICC is the enterprise qualification standard. If you can't answer every letter, you don't have a qualified opportunity — you have a conversation.
### M — Metrics
What is the quantified business impact? What does winning look like in numbers?
- "What's the current cost of [the problem]?"
- "How do you measure success in this area today?"
- "If we achieve X outcome, what does that save or earn you?"
**Red flag:** No metrics = no business case = hard to get budget.
### E — Economic Buyer
Who has final authority to approve the budget?
- "Who else will be involved in the final decision?"
- "Have you purchased solutions in this range before? Who approved that?"
- "When we get to final terms, who needs to sign?"
**Red flag:** You only know the user buyer. Economic buyer hasn't engaged.
### D — Decision Criteria
What factors will they use to evaluate and select a solution?
- "What's most important in your evaluation?"
- "How will you compare options?"
- "What does the ideal solution look like to you?"
**Why it matters:** If you don't know their criteria, you're guessing what to prove. Define the criteria before you compete on them.
### D — Decision Process
What are the steps from evaluation to signed contract?
- "Walk me through your process from here to signed agreement."
- "Does procurement get involved? Legal? InfoSec?"
- "Have you purchased software at this price before? How long did that take?"
**Red flag:** No defined process = unlimited sales cycle.
### P — Paper Process
What's the contract and legal process?
- "Who manages vendor contracts on your side?"
- "What's your standard MSA, or do you use ours?"
- "How long does legal review typically take?"
**Why it matters:** Legal and procurement have killed many "done" deals. Start early. Route to your legal team simultaneously.
### I — Identify Pain
What is the specific, felt pain driving this evaluation?
- "What triggered this initiative now vs. six months ago?"
- "What happens if you don't solve this in Q3?"
- "On a scale of 1-10, how urgent is this for your team?"
**Red flag:** Pain isn't felt by the economic buyer. User pain ≠ budget authority.
### C — Champion
Who will actively sell your solution internally when you're not in the room?
- "Who else have you brought into this evaluation?"
- "Can you help us get access to [economic buyer / IT / security]?"
- "If the decision went the wrong way, who would be disappointed?"
**Red flag:** Your champion is enthusiastic but has no internal influence.
### C — Competition
Who else are they evaluating? What's your position?
- "Are you looking at alternatives?"
- "What made you start with us?"
- "Have you used [Competitor X] before?"
**Why it matters:** Knowing the competitive field tells you what you need to prove and what to neutralize.
### MEDDPICC Scorecard
| Letter | Score 1 | Score 2 | Score 3 |
|--------|---------|---------|---------|
| Metrics | No numbers | Approximate value | Specific ROI model |
| Economic Buyer | Unknown | Named, not engaged | Engaged directly |
| Decision Criteria | Vague | Partially defined | Written, weighted |
| Decision Process | Unknown | Verbal description | Steps confirmed, timeline known |
| Paper Process | Unknown | Basic awareness | Legal contacts, standard process known |
| Identify Pain | No urgency | User-level pain | Executive-level pain with consequences |
| Champion | No advocate | Friendly contact | Actively selling internally |
| Competition | Unknown | Identified | Position mapped, differentiation clear |
**Score each 1-3. Total 16+/24 = qualified opportunity. Under 12 = unqualified, do not forecast.**
---
## Sales Compensation Plans
Comp drives behavior. Design it precisely.
### Base / Variable Split
| Role | Base % | Variable % | Rationale |
|------|--------|-----------|-----------|
| SDR | 60-70% | 30-40% | Activity-based, not purely revenue |
| AE (Inside Sales) | 50% | 50% | Balanced risk/reward |
| AE (Enterprise) | 55-60% | 40-45% | Longer cycle, higher base for stability |
| VP Sales | 50% | 50% | Accountable for team results |
| CSM (retention focus) | 70% | 30% | Less variable, stable relationship role |
| CSM (expansion focus) | 60% | 40% | Expansion quota adds variable |
### Commission Structure
**Standard AE plan:**
```
Base: $80K
Variable: $80K (at 100% quota attainment)
OTE: $160K
Commission rate: OTE variable ÷ Quota
If quota = $800K ARR: commission = $80K ÷ $800K = 10% of ARR closed
Accelerators (performance above quota):
101-125% quota: 1.25x commission rate (12.5% of ARR)
126-150% quota: 1.5x commission rate (15% of ARR)
> 150% quota: 2.0x commission rate (20% of ARR)
```
**Why accelerators matter:**
- They keep top performers motivated past quota
- They make it possible for top reps to earn $200K+ (attracting talent)
- They create the "make it rain" culture
### SDR Compensation
SDRs are measured on output (meetings booked, pipeline created), not closed revenue.
```
Quota: 20 qualified meetings booked per month (or $X pipeline created)
Commission: $150-300 per qualified meeting held
Accelerators:
If a meeting converts to closed won: Bonus $250-500
If monthly meetings > 125% of quota: 1.5x rate on upside meetings
```
### Clawbacks
A clawback recovers commission paid on deals that churn or are fraudulently closed.
**Common clawback rules:**
- Full clawback if customer cancels within 90 days of close
- 50% clawback if customer cancels within 91-180 days
- No clawback after 180 days (AE shouldn't be penalized for future CS failures)
- Clawbacks vest: pay commission immediately but apply against next quarter's payout if triggered
**Why clawbacks matter:**
- Without them, reps are incentivized to close any deal, regardless of fit
- With them, reps self-qualify more carefully
### SPIFFs (Sales Performance Incentive Funds)
Short-term tactical incentives for specific behaviors:
- $5K bonus for closing a new vertical deal this quarter
- 1.5x commission on annual prepay deals in Q4
- $1K for closing a deal in a new geographic territory
Use SPIFFs sparingly. Overuse trains reps to wait for the SPIFF before engaging.
### Multi-Year and Prepay Incentives
Align rep behavior with company cash flow:
- Multi-year deals: Credit full TCV against quota, pay commission upfront on TCV
- Annual prepay: 10-20% uplift on commission rate
- Monthly billing: Standard commission rate
---
## Enterprise vs. SMB vs. Self-Serve Models
### Self-Serve / PLG
**Characteristics:**
- Product is the primary acquisition channel
- Credit card required (no invoicing)
- No human touch in the initial purchase
- Sales engages only at enterprise signals (high usage, team expansion, compliance needs)
**Funnel:**
```
Website → Free trial / Freemium → Activation → PQL → Expansion → Enterprise
```
**Key metrics:**
- Free-to-paid conversion rate (benchmark: 2-5% of signups)
- Time to activation (first core action)
- PQL → expansion conversion rate
- NRR from self-serve base
**Sales involvement triggers (PQL signals):**
- Team size > 10 seats
- Usage spikes (power user patterns)
- Feature limit hits on core features
- Job title change (new economic buyer appears in account)
### SMB Inside Sales
**Characteristics:**
- ACV $5K-25K
- 30-60 day sales cycle
- Inbound-heavy or light outbound
- SDR → AE → CS model
- Phone + email + video; no in-person
**Funnel:**
```
Inbound/MQL → SDR qualifies → AE discovery → Demo → Proposal → Close
```
**Key metrics:**
- MQL-to-SQL rate (benchmark: 15-25%)
- SQL-to-close rate (benchmark: 20-30%)
- Average sales cycle (30-60 days)
- AE productivity: $600K-$1M quota per rep
**Team ratios:**
- 1 SDR supports 3-4 AEs
- 1 CSM manages $1M-2M ARR
### Enterprise Sales
**Characteristics:**
- ACV $50K+
- 90-365 day sales cycle
- Outbound prospecting + inbound from brand
- AE + SE + executive sponsor model
- Multi-stakeholder: champion, economic buyer, IT, legal, procurement
**Funnel:**
```
Account targeting → Executive outreach → Discovery → POC → Security review → Legal → Procurement → Close
```
**Key metrics:**
- Deals in pipeline (volume matters less, quality more)
- POC win rate (benchmark: 60-75%)
- Average sales cycle (3-12 months)
- AE productivity: $1.5M-$3M quota per rep
**Team ratios:**
- 1 SE supports 3-4 AEs
- 1 CSM manages $2M-5M ARR (named accounts, high-touch)
---
## Sales Hiring and Ramp
### What "Good" Looks Like by Role
**SDR (entry level):**
- 1-2 years of outbound experience OR strong track record in customer-facing role
- Resilient: rejection is the job
- Coachable: SDR is a proving ground, not a final destination
- Can write clear, concise prospecting emails without templates
**AE (inside sales):**
- 2-4 years sales experience, preferably SaaS
- Can articulate their process for a discovery call
- Knows their numbers: quota, attainment, average deal size, sales cycle
- Shows how they build pipeline (AEs who only work inbound are a risk)
**AE (enterprise):**
- 4-8 years B2B sales, at least 2 in enterprise
- Has closed deals > $100K ACV
- Can name the stakeholders in a complex deal they navigated
- Understands procurement, security review, multi-year contracts
**VP Sales:**
- Has scaled a team from where you are to 2x your size
- Can build a comp plan from scratch
- Has hiring and firing experience
- Revenue from a repeatable process, not personal relationships
### Interview Process
**3-stage process:**
1. **Recruiter screen** (30 min): Motivation, experience, logistics
2. **Manager interview** (60 min): Structured questions on process, examples, numbers
3. **Panel / role play** (90 min): Mock discovery call + debrief; team fit
**Role play rubric:**
- Did they prepare (knew your product, your ICP)?
- Did they ask before pitching?
- Did they handle pushback without capitulating immediately?
- Did they confirm a next step with a date?
### Onboarding Structure (6-Week Ramp)
| Week | Focus | Activities |
|------|-------|-----------|
| 1 | Company, product, ICP | Onboarding sessions, product sandbox, shadow AE calls |
| 2 | Sales process, tools, messaging | CRM training, call review, write first prospecting emails |
| 3 | First outreach | Send first sequences, book first meetings, shadow closes |
| 4 | Independent discovery | Lead own discovery calls with manager reviewing |
| 5 | Full cycle | Handle pipeline independently, weekly coaching |
| 6 | Quota-bearing | 25% of quota expectation; full accountability begins |
### Performance Management
**Clear standards, no surprises:**
```
Month 3: 25% of quota expected. Miss by > 50% → performance conversation.
Month 4: 50% of quota expected. Miss by > 40% → PIP warning.
Month 5: 75% of quota. Miss by > 30% → formal PIP.
Month 6+: 100% of quota. Consistent miss → exit.
```
**PIP (Performance Improvement Plan) — not for show:**
- Should include specific, measurable targets (not "improve attitude")
- 30-60 day timeline
- Weekly check-ins with manager
- If targets aren't met: exit, no extensions
- A PIP that doesn't lead to improvement or exit is a management failure
**Rule:** Low performers who stay cost you your top performers. They watch what you tolerate.
FILE:scripts/churn_analyzer.py
#!/usr/bin/env python3
"""
Churn & Retention Analyzer
===========================
Customer-level churn and Net Revenue Retention (NRR) analysis for B2B SaaS.
Calculates:
- Gross Revenue Retention (GRR) and Net Revenue Retention (NRR)
- Monthly and annual churn rates (logo + revenue)
- Cohort-based retention curves
- At-risk account identification
- Expansion revenue segmentation
- ARR waterfall (new / expansion / contraction / churn)
Usage:
python churn_analyzer.py
python churn_analyzer.py --csv customers.csv
python churn_analyzer.py --period 2026-Q1 --output summary
Input format (CSV):
customer_id, name, segment, arr, start_date, [churn_date], [expansion_arr], [contraction_arr]
Stdlib only. No dependencies.
"""
import csv
import sys
import json
import argparse
import statistics
from datetime import date, datetime, timedelta
from collections import defaultdict
from io import StringIO
from itertools import groupby
# ---------------------------------------------------------------------------
# Data model
# ---------------------------------------------------------------------------
class Customer:
def __init__(self, customer_id, name, segment, arr, start_date,
churn_date=None, expansion_arr=0.0, contraction_arr=0.0,
health_score=None):
self.customer_id = customer_id
self.name = name
self.segment = segment
self.arr = float(arr)
self.start_date = self._parse_date(start_date)
self.churn_date = self._parse_date(churn_date) if churn_date else None
self.expansion_arr = float(expansion_arr or 0)
self.contraction_arr = float(contraction_arr or 0)
self.health_score = float(health_score) if health_score else None
@staticmethod
def _parse_date(value):
if not value or str(value).strip() in ("", "None", "null"):
return None
for fmt in ("%Y-%m-%d", "%m/%d/%Y", "%d/%m/%Y", "%Y/%m/%d"):
try:
return datetime.strptime(str(value).strip(), fmt).date()
except ValueError:
continue
raise ValueError(f"Cannot parse date: {value!r}")
def is_churned(self):
return self.churn_date is not None
def is_active(self, as_of=None):
as_of = as_of or date.today()
if self.churn_date and self.churn_date <= as_of:
return False
return self.start_date <= as_of
def tenure_days(self, as_of=None):
as_of = as_of or date.today()
end = self.churn_date if self.churn_date else as_of
return (end - self.start_date).days
def tenure_months(self, as_of=None):
return self.tenure_days(as_of) / 30.44
def cohort_month(self):
"""Acquisition cohort: YYYY-MM of start_date."""
return self.start_date.strftime("%Y-%m")
def cohort_quarter(self):
q = (self.start_date.month - 1) // 3 + 1
return f"Q{q} {self.start_date.year}"
def net_arr(self):
"""Current ARR + expansion - contraction."""
return self.arr + self.expansion_arr - self.contraction_arr
def days_since_acquisition(self, as_of=None):
as_of = as_of or date.today()
return (as_of - self.start_date).days
# ---------------------------------------------------------------------------
# Core metrics
# ---------------------------------------------------------------------------
class RetentionAnalyzer:
def __init__(self, customers, as_of=None):
self.customers = customers
self.as_of = as_of or date.today()
def active_customers(self, as_of=None):
as_of = as_of or self.as_of
return [c for c in self.customers if c.is_active(as_of)]
def churned_customers(self, start=None, end=None):
"""Customers who churned in [start, end]."""
result = []
for c in self.customers:
if not c.churn_date:
continue
if start and c.churn_date < start:
continue
if end and c.churn_date > end:
continue
result.append(c)
return result
def arr_waterfall(self, period_start, period_end):
"""
Calculate ARR waterfall for a given period.
Returns dict with opening_arr, new_arr, expansion_arr, contraction_arr,
churned_arr, closing_arr, nrr, grr.
"""
# Opening: active at period start
opening_customers = [c for c in self.customers if c.is_active(period_start)]
opening_arr = sum(c.arr for c in opening_customers)
opening_ids = {c.customer_id for c in opening_customers}
# New: started during the period
new_customers = [
c for c in self.customers
if period_start < c.start_date <= period_end
]
new_arr = sum(c.arr for c in new_customers)
# Churned: were active at start, churn_date within period
churned = [
c for c in opening_customers
if c.churn_date and period_start < c.churn_date <= period_end
]
churned_arr = sum(c.arr for c in churned)
# Expansion and contraction: from customers active at opening
expansion = sum(
c.expansion_arr for c in opening_customers
if not c.is_churned() or (c.churn_date and c.churn_date > period_end)
)
contraction = sum(
c.contraction_arr for c in opening_customers
if not c.is_churned() or (c.churn_date and c.churn_date > period_end)
)
closing_arr = opening_arr + new_arr + expansion - contraction - churned_arr
grr = (opening_arr - contraction - churned_arr) / opening_arr if opening_arr else 0
nrr = (opening_arr + expansion - contraction - churned_arr) / opening_arr if opening_arr else 0
return {
"period_start": period_start.isoformat(),
"period_end": period_end.isoformat(),
"opening_arr": opening_arr,
"new_arr": new_arr,
"expansion_arr": expansion,
"contraction_arr": contraction,
"churned_arr": churned_arr,
"closing_arr": closing_arr,
"net_new_arr": new_arr + expansion - contraction - churned_arr,
"grr": max(0.0, grr),
"nrr": max(0.0, nrr),
}
def logo_churn_rate(self, period_start, period_end):
"""Logo churn rate for a period."""
opening = [c for c in self.customers if c.is_active(period_start)]
churned = [
c for c in opening
if c.churn_date and period_start < c.churn_date <= period_end
]
return len(churned) / len(opening) if opening else 0.0
def revenue_churn_rate(self, period_start, period_end):
"""Gross revenue churn rate for a period."""
opening = [c for c in self.customers if c.is_active(period_start)]
opening_arr = sum(c.arr for c in opening)
churned_arr = sum(
c.arr for c in opening
if c.churn_date and period_start < c.churn_date <= period_end
)
contraction = sum(c.contraction_arr for c in opening)
return (churned_arr + contraction) / opening_arr if opening_arr else 0.0
# ---------------------------------------------------------------------------
# Cohort analysis
# ---------------------------------------------------------------------------
class CohortAnalyzer:
def __init__(self, customers):
self.customers = customers
def build_cohorts(self):
"""Group customers by acquisition cohort (month)."""
cohorts = defaultdict(list)
for c in self.customers:
cohorts[c.cohort_month()].append(c)
return dict(sorted(cohorts.items()))
def retention_at_month(self, cohort_customers, months_after):
"""
What fraction of cohort ARR remains `months_after` months after acquisition?
"""
if not cohort_customers:
return None
opening_arr = sum(c.arr for c in cohort_customers)
if opening_arr == 0:
return None
earliest_start = min(c.start_date for c in cohort_customers)
check_date = earliest_start + timedelta(days=int(months_after * 30.44))
if check_date > date.today():
return None # Future — no data
retained_arr = sum(
c.arr for c in cohort_customers
if c.is_active(check_date)
)
return retained_arr / opening_arr
def retention_curve(self, cohort_customers, max_months=24):
"""Return retention at months 0, 3, 6, 9, 12, 18, 24."""
checkpoints = [0, 3, 6, 9, 12, 18, 24]
checkpoints = [m for m in checkpoints if m <= max_months]
curve = {}
for m in checkpoints:
rate = self.retention_at_month(cohort_customers, m)
if rate is not None:
curve[m] = rate
return curve
def cohort_report(self):
"""Returns dict: cohort → {size, opening_arr, retention_curve}."""
cohorts = self.build_cohorts()
report = {}
for cohort_month, customers in cohorts.items():
curve = self.retention_curve(customers)
report[cohort_month] = {
"customer_count": len(customers),
"opening_arr": sum(c.arr for c in customers),
"churned_count": sum(1 for c in customers if c.is_churned()),
"current_retention": curve.get(12, curve.get(max(curve.keys()) if curve else 0)),
"retention_curve": curve,
}
return report
def identify_at_risk(self, tenure_months_max=6, health_threshold=60):
"""
Identify at-risk customers based on:
- Low health score (if available)
- Short tenure (haven't proved long-term value)
- High contraction signals
"""
at_risk = []
for c in self.customers:
if c.is_churned():
continue
reasons = []
score = 0
# Health score signal
if c.health_score is not None and c.health_score < health_threshold:
reasons.append(f"Health score {c.health_score:.0f} < {health_threshold}")
score += 40
# Early tenure risk
tenure = c.tenure_months()
if tenure < tenure_months_max:
reasons.append(f"Tenure {tenure:.1f} months (< {tenure_months_max})")
score += 20
# Contraction signal
if c.contraction_arr > 0:
contraction_pct = c.contraction_arr / c.arr
reasons.append(f"Contraction {contraction_pct:.0%} of ARR")
score += 30
# No expansion in mature account
if tenure > 12 and c.expansion_arr == 0:
reasons.append("No expansion after 12+ months (stagnant)")
score += 10
if score > 0:
at_risk.append({
"customer_id": c.customer_id,
"name": c.name,
"segment": c.segment,
"arr": c.arr,
"tenure_months": round(tenure, 1),
"health_score": c.health_score,
"risk_score": score,
"risk_reasons": reasons,
})
return sorted(at_risk, key=lambda x: -x["risk_score"])
# ---------------------------------------------------------------------------
# Expansion analysis
# ---------------------------------------------------------------------------
class ExpansionAnalyzer:
def __init__(self, customers):
self.customers = customers
def expansion_summary(self):
active = [c for c in self.customers if not c.is_churned()]
expanding = [c for c in active if c.expansion_arr > 0]
contracting = [c for c in active if c.contraction_arr > 0]
total_arr = sum(c.arr for c in active)
total_expansion = sum(c.expansion_arr for c in active)
total_contraction = sum(c.contraction_arr for c in active)
return {
"active_customers": len(active),
"total_arr": total_arr,
"expanding_count": len(expanding),
"contracting_count": len(contracting),
"expansion_arr": total_expansion,
"contraction_arr": total_contraction,
"expansion_rate": total_expansion / total_arr if total_arr else 0,
"contraction_rate": total_contraction / total_arr if total_arr else 0,
"net_expansion_rate": (total_expansion - total_contraction) / total_arr if total_arr else 0,
}
def expansion_by_segment(self):
active = [c for c in self.customers if not c.is_churned()]
by_segment = defaultdict(lambda: {"arr": 0.0, "expansion": 0.0,
"contraction": 0.0, "count": 0})
for c in active:
seg = c.segment or "Unspecified"
by_segment[seg]["arr"] += c.arr
by_segment[seg]["expansion"] += c.expansion_arr
by_segment[seg]["contraction"] += c.contraction_arr
by_segment[seg]["count"] += 1
result = {}
for seg, data in by_segment.items():
arr = data["arr"]
result[seg] = {
"customer_count": data["count"],
"arr": arr,
"expansion_arr": data["expansion"],
"contraction_arr": data["contraction"],
"expansion_rate": data["expansion"] / arr if arr else 0,
"net_nrr_contribution": (arr + data["expansion"] - data["contraction"]) / arr if arr else 0,
}
return result
def top_expansion_candidates(self, min_tenure_months=6, min_arr=5000):
"""
Customers who are active, healthy tenure, but have zero expansion.
These are upsell/expansion targets.
"""
active = [c for c in self.customers if not c.is_churned()]
candidates = []
for c in active:
tenure = c.tenure_months()
if (tenure >= min_tenure_months
and c.arr >= min_arr
and c.expansion_arr == 0
and (c.health_score is None or c.health_score >= 60)):
candidates.append({
"customer_id": c.customer_id,
"name": c.name,
"segment": c.segment,
"arr": c.arr,
"tenure_months": round(tenure, 1),
"health_score": c.health_score,
})
return sorted(candidates, key=lambda x: -x["arr"])
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt_currency(value):
if value >= 1_000_000:
return f".2fM"
if value >= 1_000:
return f".1fK"
return f".0f"
def fmt_pct(value):
return f"{value * 100:.1f}%"
def nrr_status(nrr):
if nrr >= 1.20:
return "✅ World-class"
if nrr >= 1.10:
return "✅ Healthy"
if nrr >= 1.00:
return "⚠️ Acceptable"
if nrr >= 0.90:
return "🔴 Concerning"
return "🔴 Crisis"
def grr_status(grr):
if grr >= 0.90:
return "✅ Strong"
if grr >= 0.85:
return "⚠️ Acceptable"
return "🔴 Below threshold"
def print_header(title):
width = 70
print()
print("=" * width)
print(f" {title}")
print("=" * width)
def print_section(title):
print(f"\n--- {title} ---")
def print_full_report(customers, period_start, period_end):
analyzer = RetentionAnalyzer(customers, as_of=period_end)
cohort_analyzer = CohortAnalyzer(customers)
expansion_analyzer = ExpansionAnalyzer(customers)
print_header("CHURN & RETENTION ANALYZER")
print(f" Analysis period: {period_start.isoformat()} → {period_end.isoformat()}")
print(f" Total customers in dataset: {len(customers)}")
active = analyzer.active_customers(period_end)
churned_in_period = analyzer.churned_customers(period_start, period_end)
print(f" Active at period end: {len(active)}")
print(f" Churned in period: {len(churned_in_period)}")
# ── ARR Waterfall
print_section("ARR WATERFALL")
wf = analyzer.arr_waterfall(period_start, period_end)
print(f" Opening ARR: {fmt_currency(wf['opening_arr'])}")
print(f" + New Logo ARR: +{fmt_currency(wf['new_arr'])}")
print(f" + Expansion ARR: +{fmt_currency(wf['expansion_arr'])}")
print(f" - Contraction ARR: -{fmt_currency(wf['contraction_arr'])}")
print(f" - Churned ARR: -{fmt_currency(wf['churned_arr'])}")
print(f" {'─'*42}")
print(f" Closing ARR: {fmt_currency(wf['closing_arr'])}")
print(f" Net New ARR: {'+' if wf['net_new_arr'] >= 0 else ''}{fmt_currency(wf['net_new_arr'])}")
# ── NRR / GRR
print_section("RETENTION METRICS")
nrr = wf["nrr"]
grr = wf["grr"]
logo_churn = analyzer.logo_churn_rate(period_start, period_end)
rev_churn = analyzer.revenue_churn_rate(period_start, period_end)
print(f" NRR (Net Revenue Retention): {fmt_pct(nrr)} {nrr_status(nrr)}")
print(f" GRR (Gross Revenue Retention): {fmt_pct(grr)} {grr_status(grr)}")
print(f" Logo Churn Rate (period): {fmt_pct(logo_churn)}")
print(f" Revenue Churn Rate (period): {fmt_pct(rev_churn)}")
if wf["opening_arr"] > 0:
expansion_rate = wf["expansion_arr"] / wf["opening_arr"]
print(f" Expansion Rate (period): {fmt_pct(expansion_rate)}")
print()
print(f" NRR Benchmark: >120% world-class | 100-120% healthy | <100% fix immediately")
# ── Expansion summary
print_section("EXPANSION REVENUE")
exp = expansion_analyzer.expansion_summary()
print(f" Expanding customers: {exp['expanding_count']} / {exp['active_customers']} ({fmt_pct(exp['expanding_count']/exp['active_customers']) if exp['active_customers'] else '—'})")
print(f" Contracting: {exp['contracting_count']} / {exp['active_customers']}")
print(f" Expansion ARR: {fmt_currency(exp['expansion_arr'])} ({fmt_pct(exp['expansion_rate'])} of base)")
print(f" Contraction ARR: {fmt_currency(exp['contraction_arr'])}")
print(f" Net Expansion Rate: {fmt_pct(exp['net_expansion_rate'])}")
# ── Segment breakdown
print_section("SEGMENT BREAKDOWN (NRR Components)")
seg_data = expansion_analyzer.expansion_by_segment()
col_w = [18, 8, 12, 10, 10, 10]
h = (f" {'Segment':<{col_w[0]}} {'Custs':>{col_w[1]}} {'ARR':>{col_w[2]}} "
f"{'Expansion':>{col_w[3]}} {'Contraction':>{col_w[4]}} {'NRR':>{col_w[5]}}")
print(h)
print(" " + "-" * (sum(col_w) + 5))
for seg, data in sorted(seg_data.items(), key=lambda x: -x[1]["arr"]):
print(f" {seg:<{col_w[0]}} {data['customer_count']:>{col_w[1]}} "
f"{fmt_currency(data['arr']):>{col_w[2]}} "
f"{fmt_currency(data['expansion_arr']):>{col_w[3]}} "
f"{fmt_currency(data['contraction_arr']):>{col_w[4]}} "
f"{fmt_pct(data['net_nrr_contribution']):>{col_w[5]}}")
# ── Cohort retention
print_section("COHORT RETENTION CURVES")
cohort_report = cohort_analyzer.cohort_report()
print(f" {'Cohort':<10} {'Custs':>6} {'Opening ARR':>13} {'Mo.3':>8} {'Mo.6':>8} {'Mo.12':>8}")
print(" " + "-" * 57)
for cohort, data in cohort_report.items():
curve = data["retention_curve"]
m3 = fmt_pct(curve[3]) if 3 in curve else " —"
m6 = fmt_pct(curve[6]) if 6 in curve else " —"
m12 = fmt_pct(curve[12]) if 12 in curve else " —"
print(f" {cohort:<10} {data['customer_count']:>6} "
f"{fmt_currency(data['opening_arr']):>13} "
f"{m3:>8} {m6:>8} {m12:>8}")
# ── At-risk accounts
print_section("AT-RISK ACCOUNTS")
at_risk = cohort_analyzer.identify_at_risk()
if at_risk:
print(f" {'Customer':<22} {'Segment':<14} {'ARR':>10} {'Tenure':>8} {'Risk':>6} Reason")
print(" " + "-" * 80)
for acct in at_risk[:10]: # Top 10
reason_short = acct["risk_reasons"][0] if acct["risk_reasons"] else ""
tenure_str = f"{acct['tenure_months']}mo"
print(f" {acct['name']:<22} {acct['segment']:<14} "
f"{fmt_currency(acct['arr']):>10} {tenure_str:>8} "
f"{acct['risk_score']:>5} {reason_short}")
if len(at_risk) > 10:
print(f" ... and {len(at_risk) - 10} more at-risk accounts")
else:
print(" ✅ No at-risk accounts identified")
# ── Expansion candidates
print_section("EXPANSION CANDIDATES (no expansion yet, healthy tenure)")
candidates = expansion_analyzer.top_expansion_candidates()
if candidates:
print(f" {'Customer':<22} {'Segment':<14} {'ARR':>10} {'Tenure':>8} Action")
print(" " + "-" * 70)
for c in candidates[:8]:
action = "Upsell review" if c["arr"] > 20000 else "Seat expansion call"
tenure_str = f"{c['tenure_months']}mo"
print(f" {c['name']:<22} {c['segment']:<14} "
f"{fmt_currency(c['arr']):>10} {tenure_str:>8} {action}")
else:
print(" ✅ All eligible accounts have expansion in motion")
# ── Red flags
print_section("HEALTH FLAGS")
flags = []
if nrr < 1.0:
flags.append("🔴 NRR below 100% — revenue base is shrinking. Fix before scaling sales.")
if grr < 0.85:
flags.append(f"🔴 GRR {fmt_pct(grr)} — gross retention below 85% threshold. Churn is a product/CS problem.")
if logo_churn > 0.05:
flags.append(f"⚠️ Logo churn {fmt_pct(logo_churn)} this period — run cohort analysis to find the pattern.")
if exp["expansion_rate"] < 0.10 and exp["active_customers"] > 10:
flags.append("⚠️ Expansion rate below 10% — upsell motion is weak or non-existent.")
churned_arr_pct = wf["churned_arr"] / wf["opening_arr"] if wf["opening_arr"] else 0
if churned_arr_pct > 0.10:
flags.append(f"🔴 Revenue churn at {fmt_pct(churned_arr_pct)} of opening ARR this period — high urgency.")
if len(at_risk) > len(active) * 0.20:
flags.append(f"⚠️ {len(at_risk)} of {len(active)} active accounts flagged at-risk ({fmt_pct(len(at_risk)/len(active) if active else 0)})")
if flags:
for f in flags:
print(f" {f}")
else:
print(" ✅ No critical health flags")
print()
# ---------------------------------------------------------------------------
# Sample data
# ---------------------------------------------------------------------------
SAMPLE_CSV = """customer_id,name,segment,arr,start_date,churn_date,expansion_arr,contraction_arr,health_score
C001,Acme Manufacturing,Enterprise,120000,2023-01-15,,45000,0,82
C002,TechStart Inc,Mid-Market,28000,2023-02-01,,8000,0,74
C003,Global Retail Co,Enterprise,250000,2023-01-05,,0,25000,45
C004,MedTech Solutions,Mid-Market,45000,2023-03-10,,15000,0,88
C005,FinServ Holdings,Enterprise,185000,2023-01-20,2023-09-15,0,0,
C006,StartupHub Network,SMB,12000,2023-04-01,,0,3000,55
C007,EduPlatform Inc,Mid-Market,32000,2023-02-15,,10000,0,91
C008,BioLab Analytics,Enterprise,95000,2023-01-10,,20000,0,78
C009,RegionalBank Corp,Enterprise,310000,2023-03-01,,75000,0,85
C010,CloudOps Systems,Mid-Market,38000,2023-05-01,2024-01-10,0,0,
C011,InsurTech Platform,Mid-Market,55000,2023-06-15,,0,0,62
C012,LegalAI Corp,SMB,18000,2023-07-01,,5000,0,79
C013,RetailChain Ltd,Enterprise,140000,2023-04-20,,0,20000,41
C014,DataPipeline Co,Mid-Market,42000,2023-08-01,,12000,0,83
C015,NanoTech Startup,SMB,9500,2023-09-15,2024-02-28,0,0,
C016,MedDevice Corp,Enterprise,220000,2023-02-28,,60000,0,92
C017,ConsultingFirm XYZ,SMB,15000,2023-10-01,,0,5000,38
C018,GovTech Solutions,Enterprise,175000,2023-11-15,,0,0,71
C019,AgriData Systems,Mid-Market,31000,2024-01-10,,8000,0,77
C020,HealthcarePlus,Mid-Market,62000,2024-02-01,,0,0,65
"""
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def load_customers_from_csv(csv_text):
reader = csv.DictReader(StringIO(csv_text))
customers = []
errors = []
for i, row in enumerate(reader, start=2):
try:
c = Customer(
customer_id=row.get("customer_id", f"row_{i}"),
name=row.get("name", f"Customer {i}"),
segment=row.get("segment", ""),
arr=row.get("arr", 0),
start_date=row.get("start_date", ""),
churn_date=row.get("churn_date", None) or None,
expansion_arr=row.get("expansion_arr", 0) or 0,
contraction_arr=row.get("contraction_arr", 0) or 0,
health_score=row.get("health_score", None) or None,
)
customers.append(c)
except (ValueError, KeyError) as e:
errors.append(f" Row {i}: {e}")
if errors:
print("⚠️ Skipped rows with errors:")
for err in errors:
print(err)
return customers
def parse_period(period_str):
"""Parse 'YYYY-QN' or 'YYYY-MM' into (start_date, end_date)."""
if not period_str:
today = date.today()
q = (today.month - 1) // 3
start = date(today.year, q * 3 + 1, 1)
# End of current quarter
end_month = start.month + 2
end_year = start.year + (end_month - 1) // 12
end_month = ((end_month - 1) % 12) + 1
import calendar
end_day = calendar.monthrange(end_year, end_month)[1]
return start, date(end_year, end_month, end_day)
import calendar
if "-Q" in period_str:
year, qpart = period_str.split("-Q")
year = int(year)
q = int(qpart)
start_month = (q - 1) * 3 + 1
end_month = start_month + 2
start = date(year, start_month, 1)
end = date(year, end_month, calendar.monthrange(year, end_month)[1])
return start, end
# YYYY-MM
year, month = period_str.split("-")
year, month = int(year), int(month)
start = date(year, month, 1)
end = date(year, month, calendar.monthrange(year, month)[1])
return start, end
def main():
parser = argparse.ArgumentParser(
description="Churn & Retention Analyzer — NRR, cohort analysis, at-risk detection"
)
parser.add_argument(
"--csv", metavar="FILE",
help="CSV file with customer data (uses sample data if not provided)"
)
parser.add_argument(
"--period", metavar="PERIOD",
help='Analysis period: "2026-Q1" or "2026-03" (defaults to current quarter)'
)
parser.add_argument(
"--output", choices=["summary", "full", "json"],
default="full",
help="Output format (default: full)"
)
args = parser.parse_args()
# Load data
if args.csv:
try:
with open(args.csv, "r", encoding="utf-8") as f:
csv_text = f.read()
except FileNotFoundError:
print(f"Error: File not found: {args.csv}", file=sys.stderr)
sys.exit(1)
else:
print("No --csv provided. Using sample customer data.\n")
csv_text = SAMPLE_CSV
customers = load_customers_from_csv(csv_text)
if not customers:
print("No customers loaded. Exiting.", file=sys.stderr)
sys.exit(1)
period_start, period_end = parse_period(args.period)
if args.output == "json":
analyzer = RetentionAnalyzer(customers, as_of=period_end)
cohort_analyzer = CohortAnalyzer(customers)
expansion_analyzer = ExpansionAnalyzer(customers)
wf = analyzer.arr_waterfall(period_start, period_end)
output = {
"period": {"start": period_start.isoformat(), "end": period_end.isoformat()},
"arr_waterfall": wf,
"logo_churn_rate": analyzer.logo_churn_rate(period_start, period_end),
"revenue_churn_rate": analyzer.revenue_churn_rate(period_start, period_end),
"cohort_report": {k: {**v, "retention_curve": {str(m): r for m, r in v["retention_curve"].items()}}
for k, v in cohort_analyzer.cohort_report().items()},
"at_risk_accounts": cohort_analyzer.identify_at_risk(),
"expansion_summary": expansion_analyzer.expansion_summary(),
"expansion_by_segment": expansion_analyzer.expansion_by_segment(),
"expansion_candidates": expansion_analyzer.top_expansion_candidates(),
}
print(json.dumps(output, indent=2))
elif args.output == "summary":
analyzer = RetentionAnalyzer(customers, as_of=period_end)
wf = analyzer.arr_waterfall(period_start, period_end)
print_header("NRR SUMMARY")
print(f" Period: {period_start.isoformat()} → {period_end.isoformat()}")
print(f" NRR: {fmt_pct(wf['nrr'])} {nrr_status(wf['nrr'])}")
print(f" GRR: {fmt_pct(wf['grr'])} {grr_status(wf['grr'])}")
print(f" Opening: {fmt_currency(wf['opening_arr'])}")
print(f" Closing: {fmt_currency(wf['closing_arr'])}")
print(f" Net New: {fmt_currency(wf['net_new_arr'])}")
print()
else:
print_full_report(customers, period_start, period_end)
if __name__ == "__main__":
main()
FILE:scripts/revenue_forecast_model.py
#!/usr/bin/env python3
"""
Revenue Forecast Model
======================
Pipeline-based revenue forecasting for B2B SaaS.
Models:
- Weighted pipeline (stage probability × deal value)
- Historical win rate adjustment (calibrate to actuals)
- Scenario analysis (conservative / base / upside)
- Monthly and quarterly projection with confidence ranges
Usage:
python revenue_forecast_model.py
python revenue_forecast_model.py --csv pipeline.csv
python revenue_forecast_model.py --scenario conservative
Input format (CSV):
deal_id, name, stage, arr_value, close_date, rep, segment
Stdlib only. No dependencies.
"""
import csv
import sys
import json
import argparse
import statistics
from datetime import date, datetime, timedelta
from collections import defaultdict
from io import StringIO
# ---------------------------------------------------------------------------
# Stage configuration
# ---------------------------------------------------------------------------
DEFAULT_STAGE_PROBABILITIES = {
"discovery": 0.10,
"qualification": 0.25,
"demo": 0.40,
"proposal": 0.55,
"poc": 0.65,
"negotiation": 0.80,
"verbal_commit": 0.92,
"closed_won": 1.00,
"closed_lost": 0.00,
}
SCENARIO_MULTIPLIERS = {
"conservative": 0.85, # Win rate 15% below historical
"base": 1.00, # Historical win rate
"upside": 1.15, # Win rate 15% above historical
}
# ---------------------------------------------------------------------------
# Data model
# ---------------------------------------------------------------------------
class Deal:
def __init__(self, deal_id, name, stage, arr_value, close_date, rep="", segment=""):
self.deal_id = deal_id
self.name = name
self.stage = stage.lower().replace(" ", "_").replace("/", "_")
self.arr_value = float(arr_value)
self.close_date = self._parse_date(close_date)
self.rep = rep
self.segment = segment
@staticmethod
def _parse_date(value):
for fmt in ("%Y-%m-%d", "%m/%d/%Y", "%d/%m/%Y", "%Y/%m/%d"):
try:
return datetime.strptime(str(value), fmt).date()
except ValueError:
continue
raise ValueError(f"Cannot parse date: {value!r}")
@property
def quarter(self):
q = (self.close_date.month - 1) // 3 + 1
return f"Q{q} {self.close_date.year}"
@property
def month_key(self):
return self.close_date.strftime("%Y-%m")
def weighted_value(self, stage_probs, scenario="base"):
prob = stage_probs.get(self.stage, 0.0)
multiplier = SCENARIO_MULTIPLIERS.get(scenario, 1.0)
# Clamp probability to [0, 1]
adjusted = min(1.0, max(0.0, prob * multiplier))
return self.arr_value * adjusted
def is_open(self):
return self.stage not in ("closed_won", "closed_lost")
def is_closed_won(self):
return self.stage == "closed_won"
# ---------------------------------------------------------------------------
# Win rate calibration
# ---------------------------------------------------------------------------
def calculate_historical_win_rates(deals):
"""
Calculate actual win rates per stage from closed deals.
Returns a dict: stage → win_rate (float).
Requires deals that were at each stage and are now closed won/lost.
"""
# In a real implementation, you'd have historical stage-at-point-in-time data.
# Here we approximate: among closed deals, what fraction were won?
closed = [d for d in deals if not d.is_open()]
if not closed:
return {}
won = [d for d in closed if d.is_closed_won()]
overall_rate = len(won) / len(closed) if closed else 0.0
# Stage-level calibration: adjust default probs by actual overall rate
# (In production: use CRM historical stage-level conversion data)
calibrated = {}
for stage, default_prob in DEFAULT_STAGE_PROBABILITIES.items():
if overall_rate > 0:
calibrated[stage] = min(1.0, default_prob * (overall_rate / 0.25))
else:
calibrated[stage] = default_prob
return calibrated
# ---------------------------------------------------------------------------
# Forecast engine
# ---------------------------------------------------------------------------
class ForecastEngine:
def __init__(self, deals, stage_probs=None):
self.deals = deals
self.stage_probs = stage_probs or DEFAULT_STAGE_PROBABILITIES
def open_deals(self):
return [d for d in self.deals if d.is_open()]
def closed_won_deals(self):
return [d for d in self.deals if d.is_closed_won()]
def pipeline_by_month(self, scenario="base"):
"""Returns dict: month_key → weighted ARR."""
result = defaultdict(float)
for deal in self.open_deals():
result[deal.month_key] += deal.weighted_value(self.stage_probs, scenario)
return dict(sorted(result.items()))
def pipeline_by_quarter(self, scenario="base"):
"""Returns dict: quarter → weighted ARR."""
result = defaultdict(float)
for deal in self.open_deals():
result[deal.quarter] += deal.weighted_value(self.stage_probs, scenario)
return dict(sorted(result.items()))
def coverage_ratio(self, quota, period_filter=None):
"""
Pipeline coverage = total pipeline ÷ quota.
period_filter: if set, only include deals with close_date in that period.
"""
pipeline = sum(
d.arr_value for d in self.open_deals()
if period_filter is None or d.quarter == period_filter
)
return pipeline / quota if quota else 0.0
def scenario_summary(self, periods=None):
"""
Returns dict: period → {conservative, base, upside, open_pipeline}.
periods: list of month_keys to include; if None, all months.
"""
summaries = {}
all_months = sorted(set(d.month_key for d in self.open_deals()))
target_months = periods or all_months
for month in target_months:
deals_in_month = [d for d in self.open_deals() if d.month_key == month]
if not deals_in_month:
continue
summaries[month] = {
"deal_count": len(deals_in_month),
"open_pipeline": sum(d.arr_value for d in deals_in_month),
"conservative": sum(d.weighted_value(self.stage_probs, "conservative") for d in deals_in_month),
"base": sum(d.weighted_value(self.stage_probs, "base") for d in deals_in_month),
"upside": sum(d.weighted_value(self.stage_probs, "upside") for d in deals_in_month),
}
return summaries
def rep_performance(self):
"""Returns dict: rep → {pipeline, weighted_base, deal_count, avg_deal_size}."""
rep_data = defaultdict(lambda: {"pipeline": 0.0, "weighted_base": 0.0,
"deal_count": 0, "deals": []})
for deal in self.open_deals():
rep_data[deal.rep]["pipeline"] += deal.arr_value
rep_data[deal.rep]["weighted_base"] += deal.weighted_value(self.stage_probs, "base")
rep_data[deal.rep]["deal_count"] += 1
rep_data[deal.rep]["deals"].append(deal.arr_value)
result = {}
for rep, data in rep_data.items():
deals = data["deals"]
result[rep] = {
"pipeline": data["pipeline"],
"weighted_base": data["weighted_base"],
"deal_count": data["deal_count"],
"avg_deal_size": statistics.mean(deals) if deals else 0.0,
}
return result
def segment_breakdown(self, scenario="base"):
"""Returns dict: segment → weighted ARR."""
result = defaultdict(float)
for deal in self.open_deals():
result[deal.segment or "unspecified"] += deal.weighted_value(self.stage_probs, scenario)
return dict(result)
def stage_distribution(self):
"""Returns dict: stage → {count, total_arr, avg_arr}."""
result = defaultdict(lambda: {"count": 0, "total_arr": 0.0})
for deal in self.open_deals():
result[deal.stage]["count"] += 1
result[deal.stage]["total_arr"] += deal.arr_value
out = {}
for stage, data in result.items():
out[stage] = {
"count": data["count"],
"total_arr": data["total_arr"],
"avg_arr": data["total_arr"] / data["count"] if data["count"] else 0,
"probability": self.stage_probs.get(stage, 0.0),
}
return out
def confidence_interval(self, scenario="base", iterations=1000):
"""
Monte Carlo simulation to generate confidence interval around base forecast.
Each deal wins/loses based on its probability; runs iterations times.
Returns (p10, p50, p90) of total expected ARR.
"""
import random
random.seed(42)
totals = []
for _ in range(iterations):
total = 0.0
for deal in self.open_deals():
prob = min(1.0, self.stage_probs.get(deal.stage, 0.0) * SCENARIO_MULTIPLIERS[scenario])
if random.random() < prob:
total += deal.arr_value
totals.append(total)
totals.sort()
n = len(totals)
return (
totals[int(n * 0.10)], # P10 (conservative)
totals[int(n * 0.50)], # P50 (median)
totals[int(n * 0.90)], # P90 (upside)
)
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt_currency(value):
if value >= 1_000_000:
return f".2fM"
if value >= 1_000:
return f".1fK"
return f".0f"
def fmt_pct(value):
return f"{value * 100:.1f}%"
def print_header(title):
width = 70
print()
print("=" * width)
print(f" {title}")
print("=" * width)
def print_section(title):
print(f"\n--- {title} ---")
def print_report(engine, quota=None, current_quarter=None):
open_deals = engine.open_deals()
won_deals = engine.closed_won_deals()
print_header("REVENUE FORECAST MODEL")
print(f" Generated: {date.today().isoformat()}")
print(f" Open deals: {len(open_deals)}")
print(f" Closed Won (in dataset): {len(won_deals)}")
total_pipeline = sum(d.arr_value for d in open_deals)
total_won = sum(d.arr_value for d in won_deals)
print(f" Total open pipeline: {fmt_currency(total_pipeline)}")
print(f" Total closed won: {fmt_currency(total_won)}")
# ── Coverage ratio
if quota:
print_section("PIPELINE COVERAGE")
q = current_quarter or "this quarter"
ratio = engine.coverage_ratio(quota, period_filter=current_quarter)
status = "✅ Healthy" if ratio >= 3.0 else ("⚠️ Thin" if ratio >= 2.0 else "🔴 Critical")
print(f" Quota target: {fmt_currency(quota)}")
print(f" Coverage ratio: {ratio:.1f}x {status}")
print(f" (Minimum healthy = 3x; < 2x = pipeline emergency)")
# ── Stage distribution
print_section("STAGE DISTRIBUTION")
stage_dist = engine.stage_distribution()
col_w = [28, 8, 14, 12, 10]
header = f" {'Stage':<{col_w[0]}} {'Deals':>{col_w[1]}} {'Pipeline':>{col_w[2]}} {'Avg Size':>{col_w[3]}} {'Win Prob':>{col_w[4]}}"
print(header)
print(" " + "-" * (sum(col_w) + 4))
for stage, data in sorted(stage_dist.items(), key=lambda x: -x[1]["total_arr"]):
print(f" {stage:<{col_w[0]}} {data['count']:>{col_w[1]}} "
f"{fmt_currency(data['total_arr']):>{col_w[2]}} "
f"{fmt_currency(data['avg_arr']):>{col_w[3]}} "
f"{fmt_pct(data['probability']):>{col_w[4]}}")
# ── Scenario forecast by month
print_section("MONTHLY FORECAST — ALL SCENARIOS")
summaries = engine.scenario_summary()
col_w2 = [10, 8, 14, 14, 14, 14]
h2 = (f" {'Month':<{col_w2[0]}} {'Deals':>{col_w2[1]}} "
f"{'Pipeline':>{col_w2[2]}} {'Conservative':>{col_w2[3]}} "
f"{'Base':>{col_w2[4]}} {'Upside':>{col_w2[5]}}")
print(h2)
print(" " + "-" * (sum(col_w2) + 5))
for month, data in summaries.items():
print(f" {month:<{col_w2[0]}} {data['deal_count']:>{col_w2[1]}} "
f"{fmt_currency(data['open_pipeline']):>{col_w2[2]}} "
f"{fmt_currency(data['conservative']):>{col_w2[3]}} "
f"{fmt_currency(data['base']):>{col_w2[4]}} "
f"{fmt_currency(data['upside']):>{col_w2[5]}}")
# ── Quarterly rollup
print_section("QUARTERLY FORECAST ROLLUP")
q_conservative = defaultdict(float)
q_base = defaultdict(float)
q_upside = defaultdict(float)
q_pipeline = defaultdict(float)
q_count = defaultdict(int)
for deal in open_deals:
q_conservative[deal.quarter] += deal.weighted_value(engine.stage_probs, "conservative")
q_base[deal.quarter] += deal.weighted_value(engine.stage_probs, "base")
q_upside[deal.quarter] += deal.weighted_value(engine.stage_probs, "upside")
q_pipeline[deal.quarter] += deal.arr_value
q_count[deal.quarter] += 1
quarters = sorted(q_base.keys())
col_w3 = [10, 8, 14, 14, 14, 14]
h3 = (f" {'Quarter':<{col_w3[0]}} {'Deals':>{col_w3[1]}} "
f"{'Pipeline':>{col_w3[2]}} {'Conservative':>{col_w3[3]}} "
f"{'Base':>{col_w3[4]}} {'Upside':>{col_w3[5]}}")
print(h3)
print(" " + "-" * (sum(col_w3) + 5))
for q in quarters:
print(f" {q:<{col_w3[0]}} {q_count[q]:>{col_w3[1]}} "
f"{fmt_currency(q_pipeline[q]):>{col_w3[2]}} "
f"{fmt_currency(q_conservative[q]):>{col_w3[3]}} "
f"{fmt_currency(q_base[q]):>{col_w3[4]}} "
f"{fmt_currency(q_upside[q]):>{col_w3[5]}}")
# ── Monte Carlo confidence interval
print_section("CONFIDENCE INTERVAL (Monte Carlo, 1,000 simulations)")
p10, p50, p90 = engine.confidence_interval("base")
print(f" P10 (conservative floor): {fmt_currency(p10)}")
print(f" P50 (median expected): {fmt_currency(p50)}")
print(f" P90 (upside ceiling): {fmt_currency(p90)}")
print(f" Range spread: {fmt_currency(p90 - p10)}")
# ── Rep performance
print_section("REP PIPELINE PERFORMANCE")
rep_perf = engine.rep_performance()
if rep_perf:
col_w4 = [20, 8, 14, 14, 12]
h4 = (f" {'Rep':<{col_w4[0]}} {'Deals':>{col_w4[1]}} "
f"{'Pipeline':>{col_w4[2]}} {'Weighted':>{col_w4[3]}} {'Avg Size':>{col_w4[4]}}")
print(h4)
print(" " + "-" * (sum(col_w4) + 4))
for rep, data in sorted(rep_perf.items(), key=lambda x: -x[1]["pipeline"]):
print(f" {rep:<{col_w4[0]}} {data['deal_count']:>{col_w4[1]}} "
f"{fmt_currency(data['pipeline']):>{col_w4[2]}} "
f"{fmt_currency(data['weighted_base']):>{col_w4[3]}} "
f"{fmt_currency(data['avg_deal_size']):>{col_w4[4]}}")
# ── Segment breakdown
print_section("SEGMENT BREAKDOWN (Base Forecast)")
seg = engine.segment_breakdown("base")
for segment, value in sorted(seg.items(), key=lambda x: -x[1]):
bar_len = int((value / total_pipeline) * 30) if total_pipeline else 0
bar = "█" * bar_len
print(f" {segment:<20} {fmt_currency(value):>12} {bar}")
# ── Red flags
print_section("FORECAST HEALTH FLAGS")
flags = []
if total_pipeline > 0:
coverage = total_pipeline / quota if quota else None
if coverage and coverage < 2.0:
flags.append("🔴 Pipeline coverage below 2x — serious shortfall risk this quarter")
elif coverage and coverage < 3.0:
flags.append("⚠️ Pipeline coverage below 3x — limited buffer for slippage")
# Stage concentration risk
early_stage_pct = sum(
d.arr_value for d in open_deals
if engine.stage_probs.get(d.stage, 0) < 0.30
) / total_pipeline
if early_stage_pct > 0.60:
flags.append(f"⚠️ {fmt_pct(early_stage_pct)} of pipeline in early stages (< 30% probability)")
# Deal concentration
deal_values = sorted([d.arr_value for d in open_deals], reverse=True)
if deal_values and deal_values[0] / total_pipeline > 0.25:
flags.append(f"⚠️ Top deal is {fmt_pct(deal_values[0]/total_pipeline)} of pipeline — concentration risk")
# Spread between scenarios
total_conservative = sum(d.weighted_value(engine.stage_probs, "conservative") for d in open_deals)
total_upside = sum(d.weighted_value(engine.stage_probs, "upside") for d in open_deals)
spread = (total_upside - total_conservative) / total_conservative if total_conservative else 0
if spread > 0.40:
flags.append(f"⚠️ High scenario spread ({fmt_pct(spread)}) — forecast confidence is low")
if flags:
for f in flags:
print(f" {f}")
else:
print(" ✅ No critical flags detected")
print()
# ---------------------------------------------------------------------------
# Sample data
# ---------------------------------------------------------------------------
SAMPLE_CSV = """deal_id,name,stage,arr_value,close_date,rep,segment
D001,Acme Corp ERP Integration,negotiation,85000,2026-03-15,Sarah Chen,Enterprise
D002,TechStart PLG Expansion,proposal,28000,2026-03-28,Marcus Webb,Mid-Market
D003,Global Retail Co,verbal_commit,220000,2026-03-10,Sarah Chen,Enterprise
D004,BioLab Analytics,poc,62000,2026-04-05,Jamie Park,Mid-Market
D005,FinServ Holdings,demo,150000,2026-04-20,Sarah Chen,Enterprise
D006,MidWest Logistics,qualification,35000,2026-04-30,Marcus Webb,Mid-Market
D007,Edu Platform Inc,negotiation,42000,2026-03-25,Jamie Park,SMB
D008,Healthcare Connect,proposal,95000,2026-05-15,Sarah Chen,Enterprise
D009,Startup Hub Network,demo,18000,2026-04-10,Marcus Webb,SMB
D010,CloudOps Systems,poc,75000,2026-05-01,Jamie Park,Mid-Market
D011,National Bank Corp,verbal_commit,310000,2026-03-31,Sarah Chen,Enterprise
D012,RetailTech Co,qualification,22000,2026-05-20,Marcus Webb,SMB
D013,InsurTech Platform,negotiation,88000,2026-04-15,Jamie Park,Mid-Market
D014,GovTech Solutions,proposal,175000,2026-06-01,Sarah Chen,Enterprise
D015,AgriData Systems,demo,31000,2026-05-10,Marcus Webb,Mid-Market
D016,Legal AI Corp,poc,55000,2026-04-25,Jamie Park,Mid-Market
D017,Closed Won Deal,closed_won,120000,2026-02-15,Sarah Chen,Enterprise
D018,Lost Deal,closed_lost,45000,2026-02-20,Marcus Webb,Mid-Market
"""
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def load_deals_from_csv(csv_text):
reader = csv.DictReader(StringIO(csv_text))
deals = []
errors = []
for i, row in enumerate(reader, start=2):
try:
deal = Deal(
deal_id=row.get("deal_id", f"row_{i}"),
name=row.get("name", ""),
stage=row.get("stage", ""),
arr_value=row.get("arr_value", 0),
close_date=row.get("close_date", ""),
rep=row.get("rep", ""),
segment=row.get("segment", ""),
)
deals.append(deal)
except (ValueError, KeyError) as e:
errors.append(f" Row {i}: {e}")
if errors:
print("⚠️ Skipped rows with errors:")
for err in errors:
print(err)
return deals
def main():
parser = argparse.ArgumentParser(
description="Revenue Forecast Model — pipeline-based ARR forecasting"
)
parser.add_argument(
"--csv", metavar="FILE",
help="CSV file with pipeline data (uses sample data if not provided)"
)
parser.add_argument(
"--quota", type=float, default=1_000_000,
help="Quarterly quota target in ARR (default: $1,000,000)"
)
parser.add_argument(
"--quarter", metavar="QUARTER",
help='Current quarter filter e.g. "Q2 2026" (optional)'
)
parser.add_argument(
"--scenario", choices=["conservative", "base", "upside"],
default="base",
help="Primary scenario to report (default: base)"
)
parser.add_argument(
"--json", action="store_true",
help="Output forecast as JSON instead of formatted report"
)
args = parser.parse_args()
# Load data
if args.csv:
try:
with open(args.csv, "r", encoding="utf-8") as f:
csv_text = f.read()
except FileNotFoundError:
print(f"Error: File not found: {args.csv}", file=sys.stderr)
sys.exit(1)
else:
print("No --csv provided. Using sample pipeline data.\n")
csv_text = SAMPLE_CSV
deals = load_deals_from_csv(csv_text)
if not deals:
print("No deals loaded. Exiting.", file=sys.stderr)
sys.exit(1)
# Calibrate win rates from closed deals
historical_probs = calculate_historical_win_rates(deals)
stage_probs = historical_probs if historical_probs else DEFAULT_STAGE_PROBABILITIES
engine = ForecastEngine(deals, stage_probs=stage_probs)
if args.json:
output = {
"generated": date.today().isoformat(),
"quota": args.quota,
"open_pipeline": sum(d.arr_value for d in engine.open_deals()),
"coverage_ratio": engine.coverage_ratio(args.quota, args.quarter),
"monthly_forecast": engine.scenario_summary(),
"quarterly_base": engine.pipeline_by_quarter("base"),
"confidence_interval": dict(zip(
["p10", "p50", "p90"],
engine.confidence_interval("base")
)),
"rep_performance": engine.rep_performance(),
"segment_breakdown": engine.segment_breakdown("base"),
}
print(json.dumps(output, indent=2))
else:
print_report(engine, quota=args.quota, current_quarter=args.quarter)
if __name__ == "__main__":
main()
Xây dựng, đo lường và phát triển văn hóa công ty: sứ mệnh, giá trị, hành vi, bộ quy tắc văn hóa, đánh giá sức khỏe văn hóa và nghi thức.
---
name: "culture-architect"
description: "Build, measure, and evolve company culture as operational behavior — not wall posters. Covers mission/vision/values workshops, values-to-behaviors translation, culture code creation, culture health assessment, and cultural rituals by stage. Use when building company values, assessing culture health, designing cultural rituals, creating culture codes, handling culture clashes, or when user mentions culture, values, culture debt, founder culture, or culture code."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: culture-leadership
updated: 2026-03-05
frameworks: culture-playbook, culture-code-template
---
# Culture Architect
Culture is what you DO, not what you SAY. This skill builds culture as an operational system — observable behaviors, measurable health, and rituals that scale.
## Keywords
culture, company culture, values, mission, vision, culture code, cultural rituals, culture health, values-to-behaviors, founder culture, culture debt, value-washing, culture assessment, culture survey, Netflix culture deck, HubSpot culture code, psychological safety, culture scaling
## Core Principle
**Culture = (What you reward) + (What you tolerate) + (What you celebrate)**
If your values say "transparency" but you punish bearers of bad news — your real value is "optics." Culture is not aspirational. It's descriptive. The work is closing the gap between stated and actual.
## Frameworks
### 1. Mission / Vision / Values Workshop
Run this conversationally, not as a corporate offsite. Three questions:
**Mission** — Why do we exist (beyond making money)?
- "What would be lost if we disappeared tomorrow?"
- Mission is present-tense. "We reduce preventable falls in elderly care." Not "to be the leading..."
**Vision** — What does winning look like in 5–10 years?
- Specific enough to be wrong. "Every care home in Europe uses our system" beats "be the market leader."
**Values** — What behaviors do we actually model?
- Start with what you observe, not what sounds good. "What did our last great hire do that nobody asked them to?"
- Keep to 3–5. More than 5 and none of them mean anything.
### 2. Values → Behaviors Translation
This is the work. Every value needs behavioral anchors or it's decoration.
| Value | Bad version | Behavioral anchor |
|-------|------------|-------------------|
| Transparency | "We're open and honest" | "We share bad news within 24 hours, including to our manager" |
| Ownership | "We take responsibility" | "We don't hand off problems — we own them until resolved, even across team boundaries" |
| Speed | "We move fast" | "Decisions under €5K happen at team level, same day, no approval needed" |
| Quality | "We don't cut corners" | "We stop the line before shipping something we're not proud of" |
| Customer-first | "Customers are our priority" | "Any team member can escalate a customer issue to leadership, bypassing normal channels" |
**Workshop exercise:** Write your value. Then ask "How would a new hire know we actually live this on day 30?" If you can't answer concretely, it's not a value — it's an aspiration.
### 3. Culture Code Creation
A culture code is a public document that describes how you operate. It should scare off the wrong people and attract the right ones.
**Structure:**
1. Who we are (mission + context)
2. Who thrives here (specific behaviors, not adjectives)
3. Who doesn't thrive here (honest — this is the useful part)
4. How we make decisions
5. How we communicate
6. How we grow people
7. What we expect of leaders
See `templates/culture-code-template.md` for a complete template.
**Anti-patterns to avoid:**
- "We're a family" — families don't fire each other for performance
- Listing only positive traits — the "who doesn't thrive here" section is what makes it credible
- Making it aspirational instead of descriptive
### 4. Culture Health Assessment
Run quarterly. 8–12 questions. Anonymous. See `references/culture-playbook.md` for survey design.
**Core areas to measure:**
1. Psychological safety — "Can I raise a concern without fear?"
2. Clarity — "Do I know how my work connects to company goals?"
3. Fairness — "Are decisions made consistently and transparently?"
4. Growth — "Am I learning and being challenged here?"
5. Trust in leadership — "Do I believe what leadership tells me?"
**Score interpretation:**
| Score | Signal | Action |
|-------|--------|--------|
| 80–100% | Healthy | Maintain, celebrate, document |
| 65–79% | Warning | Identify specific friction — don't over-react |
| 50–64% | Damaged | Urgent leadership attention + specific fixes |
| < 50% | Crisis | Culture emergency — all-hands intervention |
### 5. Cultural Rituals by Stage
Rituals are the delivery mechanism for culture. What works at 10 people breaks at 100.
**Seed stage (< 15 people)**
- Weekly all-hands (30 min): company update + one win + one learning
- Monthly retrospective: what's working, what's not — no hierarchy
- "Default to transparency": share everything unless there's a specific reason not to
**Early growth (15–50 people)**
- Quarterly culture survey: first formal check-in
- Recognition ritual: explicit, public, tied to values (not just results)
- Onboarding buddy program: cultural transmission now requires intentional effort
- Leadership office hours: founders stay accessible as layers appear
**Scaling (50–200 people)**
- Culture committee (peer-driven, not HR): 4–6 people rotating quarterly
- Values-based performance review: culture fit is measured, not assumed
- Manager training: culture now lives or dies in team leads
- Department all-hands + company all-hands separate
**Large (200+ people)**
- Culture as strategy: explicit annual culture plan with owner and KPIs
- Internal NPS for culture ("Would you recommend this company to a friend?")
- Subculture management: engineering culture ≠ sales culture — both must align to company core
### 6. Culture Anti-Patterns
**Value-washing:** Listing values you don't practice. Symptom: employees roll their eyes during values discussions.
- Fix: Run a values audit. Ask "What did the last person who got promoted demonstrate?" If it doesn't match your values, your real values are different.
**Culture debt:** Accumulating cultural compromises over time. "We'll address the toxic star performer later." Later compounds.
- Fix: Act on culture violations faster than you think necessary. One tolerated bad behavior destroys what ten good behaviors build.
**Founder culture trap:** Culture stays frozen at founding team's personality. New hires assimilate or leave.
- Fix: Explicitly evolve values as you scale. What worked at 10 people (move fast, ask forgiveness) may be destructive at 100 (we need process).
**Culture by osmosis:** Assuming culture transmits naturally. It did at 10 people. It doesn't at 50.
- Fix: Make culture intentional. Document it. Teach it. Measure it. Reward it explicitly.
## Culture Integration with C-Suite
| When... | Culture Architect works with... | To... |
|---------|---------------------------------|-------|
| Hiring surge | CHRO | Ensure culture fit is measured, not guessed |
| Org reorg | COO + CEO | Manage culture disruption from structure change |
| M&A or partnership | CEO + COO | Detect and resolve culture clashes early |
| Performance issues | CHRO | Separate culture fit from skill deficit |
| Strategy pivot | CEO | Update values/behaviors that the pivot makes obsolete |
| Rapid growth | All | Scale rituals before culture dilutes |
## Key Questions a Culture Architect Asks
- "Can you name the last person we fired for culture reasons? What did they do?"
- "What behavior got your last promoted employee promoted? Is that in your values?"
- "What would a new hire observe on day 1 that tells them what's really valued here?"
- "What do we tolerate that we shouldn't? Who knows and does nothing?"
- "How does a team lead in Berlin know what the culture is in Madrid?"
## Red Flags
- Values posted on the wall, never referenced in reviews or decisions
- Star performers protected from cultural standards
- Leaders who "don't have time" for culture rituals
- New hires feeling the culture is "different than advertised"
- No mechanism to raise cultural concerns safely
- Culture survey results never shared with the team
## Detailed References
- `references/culture-playbook.md` — Netflix analysis, survey design, ritual examples, M&A playbook
- `templates/culture-code-template.md` — Culture code document template
FILE:references/culture-playbook.md
# Culture Playbook
Reference frameworks for building, measuring, and evolving company culture.
---
## 1. Netflix Culture Deck — What Works, What Doesn't
Reed Hastings published this in 2009. 125 slides. 20M+ views. It changed how tech companies think about culture.
### What works
**"Adequate performance gets a generous severance"** — This is the sentence that made HR professionals uncomfortable. It's also why Netflix has high performers. If you keep B-players, A-players leave.
**Context, not control** — Instead of rules and approvals, Netflix provides context (strategy, goals, constraints) and expects people to make good decisions. This only works if you actually hire people who can.
**"Freedom and responsibility" as a pair** — You can't have one without the other. Freedom without responsibility is chaos. Responsibility without freedom is bureaucracy.
**Publicly stated values actually describe behavior** — The deck is descriptive, not aspirational. It says "here's what we actually do." That's rare and valuable.
### What doesn't work (or doesn't transfer)
**"We are not a family"** — Works at Netflix, lands badly in many cultures (especially European). The principle underneath it is valid: performance matters. The framing is optional.
**"Keeper test"** — "Would I fight to keep this person?" Powerful tool, but managers need coaching to use it well. Without context, it becomes paranoia-inducing.
**No vacation policy** — Works when managers model healthy vacation use. Doesn't work when culture implicitly punishes taking time off. The policy is neutral; the culture around it determines the outcome.
**Radical transparency on compensation** — Netflix publishes pay bands. This works in high-trust, high-fairness environments. In environments with existing pay inequities, it creates problems before it fixes them.
### Key lesson
The Netflix culture deck works because it's honest about tradeoffs. Your culture code should be equally honest. "We move fast, which sometimes means decisions get revisited" is more credible than "we move fast AND we get it right the first time."
---
## 2. Values-to-Behaviors Mapping Framework
Values without behavioral anchors are intentions. Behavioral anchors make values operational.
### The mapping process
**Step 1: List your stated values**
Don't curate. Write down everything on the values list, however it's currently stated.
**Step 2: For each value, find three real examples**
"Describe a time in the last 6 months when someone exemplified [value]."
If you can't find three examples, the value isn't real.
**Step 3: Extract the observable behavior**
From the examples, identify the specific action. Not the feeling, not the intention — the action.
**Step 4: Write the behavioral anchor**
Format: "[Subject] does [specific action] in [specific context]."
**Step 5: Find the counter-example**
For each value, identify a behavior that violates it. This is what you don't tolerate.
Format: "[Subject] does NOT [specific opposite action] even when [temptation/pressure]."
### Example mapping: "Customer Obsession"
| Component | Content |
|-----------|---------|
| Value | Customer Obsession |
| Example 1 | PM delayed a sprint to fix a bug a customer reported on a call, even though it wasn't on the roadmap |
| Example 2 | Support rep escalated a technical issue directly to engineering at 9pm, resolved within 2 hours |
| Example 3 | Sales declined a deal that would have required features that would hurt existing customers |
| Behavioral anchor | "We resolve customer-reported critical issues within 24 hours, regardless of roadmap priority" |
| Counter-example | "We do not close a customer issue as 'resolved' until the customer confirms it's resolved" |
### Common mapping mistakes
**Too vague:** "We put customers first" — this doesn't change behavior.
**Too broad:** "We care about quality in everything we do" — can't be measured or violated.
**Too personal:** "We're passionate" — describes emotion, not action.
**Too aspirational:** "We strive to deliver world-class..." — "strive" lets you off the hook.
---
## 3. Culture Survey Design — 8-12 Questions That Reveal Truth
Most culture surveys are useless because they measure satisfaction, not health. Satisfaction can be high in a dysfunctional culture ("I like my team, my boss, my pay" ≠ healthy culture).
### Survey design principles
1. **Anonymous, always.** If it's not anonymous, people answer what they think you want to hear.
2. **Short enough to complete honestly.** 8–12 questions max. 15 minutes max.
3. **Likert + open text.** "On a scale of 1–5" captures signal. "Why did you give that score?" captures insight.
4. **Action-linked.** Never run a survey unless you're prepared to share results and act on them.
5. **Consistent questions over time.** You want trend data, not one-off snapshots.
### The 10-question core survey
| # | Question | Area measured |
|---|----------|---------------|
| 1 | I can raise concerns or disagreements with my manager without fear of negative consequences. | Psychological safety |
| 2 | I know how my work connects to the company's most important goals. | Clarity/alignment |
| 3 | When I make a mistake, I can be honest about it without hiding it. | Psychological safety |
| 4 | Decisions here are made based on merit and data, not politics or relationships. | Fairness |
| 5 | I trust that leadership tells us the truth, even when it's bad news. | Trust in leadership |
| 6 | I am growing and being challenged in my current role. | Growth |
| 7 | When someone underperforms and nothing happens, I feel that's handled appropriately. | Accountability |
| 8 | I feel comfortable being myself at work. | Inclusion |
| 9 | My manager recognizes my contributions in ways that feel meaningful. | Recognition |
| 10 | I would recommend this company as a great place to work to someone I respect. | Overall health (eNPS) |
### Follow-up open text questions (pick 2–3)
- "What's the one thing leadership could do differently that would most improve the culture?"
- "What do we tolerate that we shouldn't?"
- "What should we protect as we grow that we're at risk of losing?"
- "What's the gap between what we say we value and what we actually do?"
### Analyzing results
**eNPS (question 10):** Score = % Promoters (9–10) minus % Detractors (1–6). Healthy: > 20. Great: > 40.
**Questions 1 and 3 (psychological safety):** If below 70%, you have a leadership problem, not a culture problem. Fix the manager first.
**Question 7 (accountability):** This is the most honest question. Cultures that fail to hold underperformers accountable destroy high-performer retention.
**Biggest drop between surveys:** This is your fire. Don't average it away.
---
## 4. Cultural Ritual Examples by Company Stage
### Seed (< 15 people)
**Weekly "Wins and Learnings" (15 min, Fridays)**
- Each person shares one win (however small) and one learning (failure, insight, mistake)
- No slides. No prep. Just talking.
- Purpose: normalizes imperfection, builds psychological safety early
**"Open book" financials**
- Share revenue, burn, runway with the whole team monthly
- Builds owners, not employees
- Requires trust that people won't misuse the data
**"Postmortem as celebration"**
- When something goes wrong, celebrate the post-mortem publicly
- "We learned X, here's how we'll do it differently"
- Prevents a blame culture from forming early
### Early growth (15–50 people)
**Monthly "Founder's Letter"**
- CEO writes an unfiltered update: what we're winning, what's hard, what's changed
- Not polished. Not PR. Real.
- Distributed internally before it goes external
**Values spotlight in team meetings**
- One agenda item: "Who exemplified [value] this week? What did they do?"
- Takes 3 minutes. Trains the muscle for values-linked recognition.
**New hire "30-day truth sessions"**
- At day 30, every new hire meets with a senior leader (not their manager) and answers: "What surprised you? What's different from what you expected? What would you fix?"
- Captures culture signal while the new hire's eyes are still fresh
### Scaling (50–200 people)
**Quarterly culture review**
- Culture committee reviews survey results, names top issues, proposes 2–3 concrete actions
- Results shared with all-hands within 2 weeks of survey close
- 30-day action accountability check-in
**Manager calibration on culture fit**
- Quarterly: managers share one team member who exemplifies culture, one who struggles
- Group discussion on patterns, not individuals
- Identifies culture outliers early before they become retention or performance crises
**"Culture at the edges" audit**
- Review last 10 performance issues, 10 terminations, 10 promotions
- Ask: "Is the pattern consistent with our stated values?"
- This is the reality check. The data doesn't lie.
### Large (200+ people)
**Subculture alignment mapping**
- Each department articulates its micro-culture
- Cross-reference with company core values
- Identify deviations: healthy variation vs. value violation
**Culture ambassador program**
- Peer-nominated, rotating, not HR
- Run culture rituals, surface issues, connect remote/distributed teams
- Budget: small (recognition, team events), influence: large
---
## 5. How to Evolve Culture Without Losing Identity
Culture must evolve as you scale. The mistake is either: (a) refusing to evolve, preserving founder culture that doesn't scale, or (b) evolving so fast that original identity is lost.
### The evolution framework
**Preserve:** Core values that define who you are. These should be stable across stages. If "move fast" is core, it doesn't go away — but its expression changes.
**Adapt:** Behaviors that worked at one stage but need updating. "Move fast" at 10 people = decide same day. At 200 people = decide within 1 week with the right people in the room.
**Add:** New behaviors required at the new scale. "Documentation culture" wasn't needed at 10. It's essential at 100.
**Retire:** Behaviors that actively hurt at scale. "Ask forgiveness, not permission" works at seed. Creates coordination chaos at Series B.
### The evolution process
1. Annual values review (not a rewrite — an audit)
2. Ask: "Which of our current behaviors are we proud of? Which embarrass us?"
3. Identify behaviors to add/adapt/retire
4. Communicate the evolution explicitly: "Here's what's changing and why"
5. Update the culture code, onboarding, and performance criteria
### Communication of culture change
Never let culture evolution look like hypocrisy. Proactively name it:
"We used to make all decisions quickly at the team level. As we've grown, that's created coordination problems. Here's how we're updating that: [new behavior]. The underlying value — speed — hasn't changed. How we deliver it has."
---
## 6. Handling Culture Clashes in M&A or Rapid Hiring
### M&A culture integration
**Before signing:**
- Culture due diligence is as important as financial DD
- Questions to answer: How do they make decisions? What gets people fired? What gets them promoted? What do they celebrate?
- Red flag: "We have a great culture" with no supporting evidence
**First 90 days:**
- Don't impose culture; conduct a bilateral audit
- Identify: what do they do that we should adopt? What do we do that they should adopt? What conflicts must be resolved?
- Assign an integration lead on each side. Give them actual authority.
**Failure mode:** Assuming acquisition = cultural absorption. The target's culture doesn't disappear. It goes underground and resurfaces as dysfunction.
### Rapid hiring culture dilution
When a company doubles in headcount in 12 months, culture dilution is near-certain. Prevention:
1. **Codify before you scale.** Document the culture before the surge, not after.
2. **Onboarding is cultural transmission.** Not just process, not just paperwork — immersion in how decisions get made, what's celebrated, what's not tolerated.
3. **Hire for culture adds, not fits.** "Fit" means homogeneity. "Add" means the person brings a perspective or behavior that strengthens the culture without violating core values.
4. **Manager density matters.** If you're adding 10 ICs and 0 managers, the new people have nobody to transmit culture to them. Hire managers ahead of the curve.
5. **Culture buddy system.** Pair new hires with culture exemplars for the first 60 days.
FILE:templates/culture-code-template.md
# [Company Name] Culture Code
> This document describes how we work, what we value, and what it's like to be here. It's meant to be honest — which means it will attract some people and repel others. Both outcomes are correct.
---
## Who We Are
[2–3 sentences: what you do, who you serve, what would be lost if you disappeared.]
**Our mission:** [One sentence. Present tense. Specific enough to be wrong.]
**Our vision:** [Where we'll be in 5–10 years. Specific enough to debate.]
---
## What We Value
*Values are behaviors, not adjectives. Each one has a "this is what it looks like" and a "this is what it doesn't look like."*
### [Value 1]
**What this means:** [Behavioral anchor — what someone does when they live this value]
**What this doesn't mean:** [The misconception or violation to guard against]
**Example:** [A real story of this value in action at your company]
---
### [Value 2]
**What this means:** [Behavioral anchor]
**What this doesn't mean:** [The misconception or violation]
**Example:** [Real story]
---
### [Value 3]
**What this means:** [Behavioral anchor]
**What this doesn't mean:** [The misconception or violation]
**Example:** [Real story]
---
*(Repeat for each value. 3–5 total. Never more than 5.)*
---
## Who Thrives Here
*These are specific, observable behaviors — not personality traits or adjectives.*
- You raise problems early, not after they've grown. You don't complain privately and stay silent publicly.
- You own decisions even when the outcome isn't what you expected.
- You say "I don't know" instead of bluffing. Then you find out.
- You give direct feedback to the person who needs to hear it, not to everyone else.
- You make things better, not just done. You notice what's broken and fix it even when it's not your job.
- [Add 2–3 specific to your company]
---
## Who Doesn't Thrive Here
*This is the most useful section. Read it carefully.*
- People who need clear instructions before taking action. We provide context; you figure out the path.
- People who optimize for credit over outcomes. We care what got done, not who gets the headline.
- People who treat bad news as a liability. Here, hiding problems is the problem.
- People who need consensus before every decision. We move faster than that.
- [Add 2–3 specific to your company — be honest]
---
## How We Make Decisions
**Decision types:**
- **Reversible, small scope:** Make it yourself. Don't ask. Tell us what you decided.
- **Reversible, larger scope:** Tell relevant people, move forward unless you hear an objection within 24 hours.
- **Irreversible or high-stakes:** Bring the right people into the room. Write it down. Decide together.
**Default:** Bias toward action. A good decision made fast beats a perfect decision made slow.
**Who decides:** The person closest to the problem, with the most context. Not the most senior person in the room.
---
## How We Communicate
**Default to async.** Most things don't need a meeting. If it can be written, write it.
**Meetings that happen:** [List your recurring meetings and what they're for]
**Meetings that don't happen:** Status updates (use tools), information sharing (write a doc), decisions that one person could make.
**How we give feedback:** Direct, specific, timely. "That report was late and incomplete" not "you should think about your time management." We give feedback to help, not to vent.
**How we share bad news:** Within 24 hours of knowing. To the person who needs to know. Not softened to the point of unclear.
---
## How We Grow People
**We invest in people who invest in themselves.** We provide [budget, learning days, access — be specific]. We don't require you to use them.
**Promotions:** Based on impact already demonstrated, not time served. You're promoted when you're already doing the job you want.
**Performance feedback:** [How often, what format, who delivers it]
**When things aren't working:** We have direct conversations early. We don't let problems simmer for quarterly reviews.
---
## What We Expect of Leaders
Leaders here are multipliers, not heroes. Your job is to make your team better.
- You share context, not just instructions. Your team should be able to make decisions you'd make when you're not there.
- You give credit visibly and take accountability privately.
- You have hard conversations before they become unavoidable.
- You model the culture. If you don't live the values, neither will your team.
- You develop people, including ones who will outgrow their role here.
---
## The Fine Print
This document is descriptive, not aspirational. It describes how we operate today, with the intent to keep improving.
We update this annually. When the update happens, we'll tell you what changed and why.
*Last updated: [Date] | Version: [X.X]*
Tạo hoặc tối ưu chuỗi email, chiến dịch drip, email nuôi dưỡng, chào mừng, kích hoạt lại và chương trình email theo vòng đời.
---
name: "email-sequence"
description: When the user wants to create or optimize an email sequence, drip campaign, automated email flow, or lifecycle email program. Also use when the user mentions "email sequence," "drip campaign," "nurture sequence," "onboarding emails," "welcome sequence," "re-engagement emails," "email automation," or "lifecycle emails." For in-app onboarding, see onboarding-cro.
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Email Sequence Design
You are an expert in email marketing and automation. Your goal is to create email sequences that nurture relationships, drive action, and move people toward conversion.
## Initial Assessment
**Check for product marketing context first:**
If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before creating a sequence, understand:
1. **Sequence Type**
- Welcome/onboarding sequence
- Lead nurture sequence
- Re-engagement sequence
- Post-purchase sequence
- Event-based sequence
- Educational sequence
- Sales sequence
2. **Audience Context**
- Who are they?
- What triggered them into this sequence?
- What do they already know/believe?
- What's their current relationship with you?
3. **Goals**
- Primary conversion goal
- Relationship-building goals
- Segmentation goals
- What defines success?
---
## Core Principles
→ See references/email-sequence-playbook.md for details
## Output Format
### Sequence Overview
```
Sequence Name: [Name]
Trigger: [What starts the sequence]
Goal: [Primary conversion goal]
Length: [Number of emails]
Timing: [Delay between emails]
Exit Conditions: [When they leave the sequence]
```
### For Each Email
```
Email [#]: [Name/Purpose]
Send: [Timing]
Subject: [Subject line]
Preview: [Preview text]
Body: [Full copy]
CTA: [Button text] → [Link destination]
Segment/Conditions: [If applicable]
```
### Metrics Plan
What to measure and benchmarks
---
## Task-Specific Questions
1. What triggers entry to this sequence?
2. What's the primary goal/conversion action?
3. What do they already know about you?
4. What other emails are they receiving?
5. What's your current email performance?
---
## Tool Integrations
For implementation, see the [tools registry](../../tools/REGISTRY.md). Key email tools:
| Tool | Best For | MCP | Guide |
|------|----------|:---:|-------|
| **Customer.io** | Behavior-based automation | - | [customer-io.md](../../tools/integrations/customer-io.md) |
| **Mailchimp** | SMB email marketing | ✓ | [mailchimp.md](../../tools/integrations/mailchimp.md) |
| **Resend** | Developer-friendly transactional | ✓ | [resend.md](../../tools/integrations/resend.md) |
| **SendGrid** | Transactional email at scale | - | [sendgrid.md](../../tools/integrations/sendgrid.md) |
| **Kit** | Creator/newsletter focused | - | [kit.md](../../tools/integrations/kit.md) |
---
## Related Skills
- **cold-email** — WHEN the sequence targets people who have NOT opted in (outbound prospecting). NOT for warm leads or subscribers who have expressed interest.
- **copywriting** — WHEN landing pages linked from emails need copy optimization that matches the email's message and audience. NOT for the email copy itself.
- **launch-strategy** — WHEN coordinating email sequences around a specific product launch, announcement, or release window. NOT for evergreen nurture or onboarding sequences.
- **analytics-tracking** — WHEN setting up email click tracking, UTM parameters, and attribution to connect email engagement to downstream conversions. NOT for writing or designing the sequence.
- **onboarding-cro** — WHEN email sequences are supporting a parallel in-app onboarding flow and need to be coordinated to avoid duplication. NOT as a replacement for in-app onboarding experience.
---
## Communication
Deliver email sequences as complete, ready-to-send drafts — include subject line, preview text, full body, and CTA for every email in the sequence. Always specify the trigger condition and send timing. When the sequence is long (5+ emails), lead with a sequence overview table before individual emails. Flag if any email could conflict with other sequences the audience receives. Load `marketing-context` for brand voice, ICP, and product context before writing.
---
## Proactive Triggers
- User mentions low trial-to-paid conversion → ask if there's a trial expiration email sequence before recommending in-app or pricing changes.
- User reports high open rates but low clicks → diagnose email body copy and CTA specificity before blaming subject lines.
- User wants to "do email marketing" → clarify sequence type (welcome, nurture, re-engagement, etc.) before writing anything.
- User has a product launch coming → recommend coordinating launch email sequence with in-app messaging and landing page copy for consistent messaging.
- User mentions list is going cold → suggest re-engagement sequence with progressive offers before recommending acquisition spend.
---
## Output Artifacts
| Artifact | Description |
|----------|-------------|
| Sequence Architecture Doc | Trigger, goal, length, timing, exit conditions, and branching logic for the full sequence |
| Complete Email Drafts | Subject line, preview text, full body, and CTA for every email in the sequence |
| Metrics Benchmarks | Open rate, click rate, and conversion rate targets per email type and sequence goal |
| Segmentation Rules | Audience entry/exit conditions, behavioral branching, and suppression lists |
| Subject Line Variations | 3 subject line alternatives per email for A/B testing |
FILE:references/email-sequence-playbook.md
# email-sequence reference
## Core Principles
### 1. One Email, One Job
- Each email has one primary purpose
- One main CTA per email
- Don't try to do everything
### 2. Value Before Ask
- Lead with usefulness
- Build trust through content
- Earn the right to sell
### 3. Relevance Over Volume
- Fewer, better emails win
- Segment for relevance
- Quality > frequency
### 4. Clear Path Forward
- Every email moves them somewhere
- Links should do something useful
- Make next steps obvious
---
## Email Sequence Strategy
### Sequence Length
- Welcome: 3-7 emails
- Lead nurture: 5-10 emails
- Onboarding: 5-10 emails
- Re-engagement: 3-5 emails
Depends on:
- Sales cycle length
- Product complexity
- Relationship stage
### Timing/Delays
- Welcome email: Immediately
- Early sequence: 1-2 days apart
- Nurture: 2-4 days apart
- Long-term: Weekly or bi-weekly
Consider:
- B2B: Avoid weekends
- B2C: Test weekends
- Time zones: Send at local time
### Subject Line Strategy
- Clear > Clever
- Specific > Vague
- Benefit or curiosity-driven
- 40-60 characters ideal
- Test emoji (they're polarizing)
**Patterns that work:**
- Question: "Still struggling with X?"
- How-to: "How to [achieve outcome] in [timeframe]"
- Number: "3 ways to [benefit]"
- Direct: "[First name], your [thing] is ready"
- Story tease: "The mistake I made with [topic]"
### Preview Text
- Extends the subject line
- ~90-140 characters
- Don't repeat subject line
- Complete the thought or add intrigue
---
## Sequence Types Overview
### Welcome Sequence (Post-Signup)
**Length**: 5-7 emails over 12-14 days
**Goal**: Activate, build trust, convert
Key emails:
1. Welcome + deliver promised value (immediate)
2. Quick win (day 1-2)
3. Story/Why (day 3-4)
4. Social proof (day 5-6)
5. Overcome objection (day 7-8)
6. Core feature highlight (day 9-11)
7. Conversion (day 12-14)
### Lead Nurture Sequence (Pre-Sale)
**Length**: 6-8 emails over 2-3 weeks
**Goal**: Build trust, demonstrate expertise, convert
Key emails:
1. Deliver lead magnet + intro (immediate)
2. Expand on topic (day 2-3)
3. Problem deep-dive (day 4-5)
4. Solution framework (day 6-8)
5. Case study (day 9-11)
6. Differentiation (day 12-14)
7. Objection handler (day 15-18)
8. Direct offer (day 19-21)
### Re-Engagement Sequence
**Length**: 3-4 emails over 2 weeks
**Trigger**: 30-60 days of inactivity
**Goal**: Win back or clean list
Key emails:
1. Check-in (genuine concern)
2. Value reminder (what's new)
3. Incentive (special offer)
4. Last chance (stay or unsubscribe)
### Onboarding Sequence (Product Users)
**Length**: 5-7 emails over 14 days
**Goal**: Activate, drive to aha moment, upgrade
**Note**: Coordinate with in-app onboarding—email supports, doesn't duplicate
Key emails:
1. Welcome + first step (immediate)
2. Getting started help (day 1)
3. Feature highlight (day 2-3)
4. Success story (day 4-5)
5. Check-in (day 7)
6. Advanced tip (day 10-12)
7. Upgrade/expand (day 14+)
**For detailed templates**: See [references/sequence-templates.md](references/sequence-templates.md)
---
## Email Types by Category
### Onboarding Emails
- New users series
- New customers series
- Key onboarding step reminders
- New user invites
### Retention Emails
- Upgrade to paid
- Upgrade to higher plan
- Ask for review
- Proactive support offers
- Product usage reports
- NPS survey
- Referral program
### Billing Emails
- Switch to annual
- Failed payment recovery
- Cancellation survey
- Upcoming renewal reminders
### Usage Emails
- Daily/weekly/monthly summaries
- Key event notifications
- Milestone celebrations
### Win-Back Emails
- Expired trials
- Cancelled customers
### Campaign Emails
- Monthly roundup / newsletter
- Seasonal promotions
- Product updates
- Industry news roundup
- Pricing updates
**For detailed email type reference**: See [references/email-types.md](references/email-types.md)
---
## Email Copy Guidelines
### Structure
1. **Hook**: First line grabs attention
2. **Context**: Why this matters to them
3. **Value**: The useful content
4. **CTA**: What to do next
5. **Sign-off**: Human, warm close
### Formatting
- Short paragraphs (1-3 sentences)
- White space between sections
- Bullet points for scanability
- Bold for emphasis (sparingly)
- Mobile-first (most read on phone)
### Tone
- Conversational, not formal
- First-person (I/we) and second-person (you)
- Active voice
- Read it out loud—does it sound human?
### Length
- 50-125 words for transactional
- 150-300 words for educational
- 300-500 words for story-driven
### CTA Guidelines
- Buttons for primary actions
- Links for secondary actions
- One clear primary CTA per email
- Button text: Action + outcome
**For detailed copy, personalization, and testing guidelines**: See [references/copy-guidelines.md](references/copy-guidelines.md)
---
FILE:scripts/sequence_analyzer.py
#!/usr/bin/env python3
"""
sequence_analyzer.py — Email sequence quality analyzer
Usage:
python3 sequence_analyzer.py --file sequence.json
python3 sequence_analyzer.py --json
python3 sequence_analyzer.py # demo mode
Input JSON format:
[
{"subject": "...", "body": "...", "delay_days": 0},
{"subject": "...", "body": "...", "delay_days": 2},
...
]
"""
import argparse
import json
import re
import sys
# ---------------------------------------------------------------------------
# Word/pattern lists
# ---------------------------------------------------------------------------
SPAM_TRIGGER_WORDS = [
"free", "guarantee", "guaranteed", "winner", "won", "prize",
"congratulations", "cash", "earn money", "make money", "extra income",
"100% free", "no cost", "risk free", "act now", "limited time",
"click here", "buy now", "order now", "get it now",
"as seen on", "dear friend", "you have been selected",
"this isn't spam", "not spam", "no credit card required",
"special promotion", "special offer", "amazing offer",
"!!!", "!!!", "$$$", "£££",
"increase your", "increase sales", "double your",
"lose weight", "weight loss", "diet", "viagra", "casino",
]
CTA_PATTERNS = re.compile(
r"\b(click|tap|reply|download|sign up|register|buy|purchase|get started|"
r"learn more|read more|visit|go to|check out|schedule|book|claim|try|"
r"subscribe|join|start|access|watch|see|grab|discover)\b",
re.IGNORECASE,
)
PERSONALIZATION_TOKENS = re.compile(
r"\{\{?\s*\w+\s*\}?\}|%\w+%|\[FIRST_NAME\]|\[NAME\]|\[COMPANY\]|\[FIRSTNAME\]",
re.IGNORECASE,
)
# ---------------------------------------------------------------------------
# Per-email analysis
# ---------------------------------------------------------------------------
def analyze_email(email: dict, index: int) -> dict:
subject = email.get("subject", "")
body = email.get("body", "")
delay = email.get("delay_days", 0)
# Subject analysis
subject_len = len(subject)
subject_word_count = len(subject.split())
subject_ok = 30 <= subject_len <= 60
subject_has_number = bool(re.search(r"\d", subject))
subject_question = subject.strip().endswith("?")
subject_all_caps = subject == subject.upper() and len(subject) > 3
# Body analysis
body_words = re.findall(r"\b\w+\b", body)
body_word_count = len(body_words)
# CTA detection
cta_matches = CTA_PATTERNS.findall(body)
has_cta = len(cta_matches) > 0
# Personalization tokens
tokens_in_subject = PERSONALIZATION_TOKENS.findall(subject)
tokens_in_body = PERSONALIZATION_TOKENS.findall(body)
total_tokens = len(tokens_in_subject) + len(tokens_in_body)
# Spam triggers
combined = (subject + " " + body).lower()
spam_found = [w for w in SPAM_TRIGGER_WORDS if w.lower() in combined]
# Spam score (0-100, higher = more spammy)
spam_score = min(100, len(spam_found) * 10)
return {
"email_index": index + 1,
"delay_days": delay,
"subject": {
"text": subject,
"length": subject_len,
"word_count": subject_word_count,
"length_ok": subject_ok,
"has_number": subject_has_number,
"is_question": subject_question,
"all_caps_warning": subject_all_caps,
"personalized": len(tokens_in_subject) > 0,
},
"body": {
"word_count": body_word_count,
"length_verdict": _body_length_verdict(body_word_count),
"has_cta": has_cta,
"cta_phrases": list(set(cta_matches))[:5],
"personalization_tokens": total_tokens,
},
"spam": {
"trigger_words_found": spam_found[:8],
"trigger_count": len(spam_found),
"spam_risk_score": spam_score,
"risk_level": "High" if spam_score >= 40 else "Medium" if spam_score >= 20 else "Low",
},
}
def _body_length_verdict(word_count: int) -> str:
if word_count < 50:
return "Too short (<50 words)"
if word_count <= 150:
return "Short/punchy — good for re-engagement"
if word_count <= 300:
return "Optimal (150-300 words)"
if word_count <= 500:
return "Long — ensure high value throughout"
return "Very long (500+ words) — consider trimming"
# ---------------------------------------------------------------------------
# Sequence-level analysis
# ---------------------------------------------------------------------------
def analyze_pacing(emails: list) -> dict:
if len(emails) <= 1:
return {"note": "Single email — no pacing to analyze"}
delays = [e.get("delay_days", 0) for e in emails]
gaps = [delays[i] - delays[i - 1] for i in range(1, len(delays))]
issues = []
for i, gap in enumerate(gaps):
if gap <= 0:
issues.append(f"Email {i+2}: same-day or before previous — check delay_days")
elif gap == 1:
issues.append(f"Email {i+2}: only 1-day gap — may feel aggressive")
elif gap > 14:
issues.append(f"Email {i+2}: {gap}-day gap — momentum may drop")
# Assess overall cadence
avg_gap = sum(gaps) / len(gaps) if gaps else 0
if avg_gap <= 2:
cadence = "Aggressive (avg <2 days)"
elif avg_gap <= 5:
cadence = "High-frequency (avg 2-5 days)"
elif avg_gap <= 10:
cadence = "Standard (avg 5-10 days)"
else:
cadence = "Low-frequency (avg 10+ days)"
return {
"email_count": len(emails),
"total_duration_days": max(delays) - min(delays),
"avg_gap_days": round(avg_gap, 1),
"cadence_type": cadence,
"gaps": gaps,
"issues": issues,
}
# ---------------------------------------------------------------------------
# Scoring
# ---------------------------------------------------------------------------
def compute_sequence_score(email_analyses: list, pacing: dict) -> dict:
if not email_analyses:
return {"overall": 0}
# Subject score: avg subject length compliance
subject_ok_count = sum(1 for e in email_analyses if e["subject"]["length_ok"])
subject_score = round(subject_ok_count / len(email_analyses) * 100)
# CTA score: % of emails with CTA
cta_count = sum(1 for e in email_analyses if e["body"]["has_cta"])
cta_score = round(cta_count / len(email_analyses) * 100)
# Personalization score
personalized_count = sum(1 for e in email_analyses if e["body"]["personalization_tokens"] > 0)
personalization_score = round(personalized_count / len(email_analyses) * 100)
# Spam score (inverted — low spam = high score)
avg_spam = sum(e["spam"]["spam_risk_score"] for e in email_analyses) / len(email_analyses)
spam_score = max(0, 100 - int(avg_spam))
# Pacing score
pacing_issues = len(pacing.get("issues", []))
pacing_score = max(0, 100 - pacing_issues * 20)
# Body length score
length_ok_count = sum(
1 for e in email_analyses
if "Optimal" in e["body"]["length_verdict"] or "punchy" in e["body"]["length_verdict"]
)
length_score = round(length_ok_count / len(email_analyses) * 100)
weights = {
"subject_quality": 0.20,
"cta_presence": 0.20,
"spam_safety": 0.25,
"personalization": 0.15,
"pacing": 0.10,
"body_length": 0.10,
}
scores = {
"subject_quality": subject_score,
"cta_presence": cta_score,
"spam_safety": spam_score,
"personalization": personalization_score,
"pacing": pacing_score,
"body_length": length_score,
}
overall = round(sum(scores[k] * weights[k] for k in weights))
grade = "A" if overall >= 85 else "B" if overall >= 70 else "C" if overall >= 55 else "D" if overall >= 40 else "F"
return {
"overall": overall,
"grade": grade,
"breakdown": {k: {"score": v, "weight": f"{int(weights[k]*100)}%"} for k, v in scores.items()},
}
# ---------------------------------------------------------------------------
# Demo data
# ---------------------------------------------------------------------------
DEMO_SEQUENCE = [
{
"subject": "{{first_name}}, your free marketing audit is ready",
"body": "Hi {{first_name}},\n\nWe analyzed 500 campaigns like yours and found three quick wins that could double your ROAS in 30 days.\n\nI've put together a custom audit for {{company}}. It's free and takes 10 minutes to review.\n\n→ Click here to see your results: [LINK]\n\nBest,\nSarah",
"delay_days": 0,
},
{
"subject": "Did you see this, {{first_name}}?",
"body": "Quick follow-up.\n\nMost marketers we talk to are sitting on 2-3 easy optimizations that could add 20-40% more revenue from the same ad spend.\n\nHere's the #1 thing we see: landing pages that don't match the ad promise.\n\nWorth 5 minutes? → [Review your audit]\n\nSarah",
"delay_days": 3,
},
{
"subject": "The $50,000 mistake (and how to avoid it)",
"body": "True story.\n\nOne of our clients was spending $8,500/month on Google Ads with a 1.8x ROAS. Technically above break-even, but barely.\n\nWe found that 60% of their budget was going to one keyword that had zero purchase intent.\n\nAfter fixing it: same spend, 4.2x ROAS.\n\nThat's the kind of thing our audit catches. Have you looked at yours yet?\n\n→ [Open your free audit]\n\nSarah\n\nP.S. This offer expires Friday.",
"delay_days": 5,
},
{
"subject": "Last call — your audit expires tonight",
"body": "{{first_name}}, this is the last reminder.\n\nYour personalized audit expires at midnight tonight.\n\nIf growing your ROAS is a priority this quarter, take 10 minutes now.\n\n→ [Claim your audit before it expires]\n\nSarah",
"delay_days": 7,
},
{
"subject": "New case study: {{company}}-style win",
"body": "Since you didn't grab the audit, I wanted to send you something valuable anyway.\n\nHere's a 3-minute case study showing how we helped a B2B SaaS company go from 1.9x to 5.4x ROAS in 45 days.\n\nNo audit required — just solid tactics you can steal.\n\n→ [Read the case study]\n\nHope it helps,\nSarah",
"delay_days": 14,
},
]
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
parser = argparse.ArgumentParser(
description="Email sequence analyzer — scores sequence quality 0-100."
)
parser.add_argument("--file", help="JSON file with email sequence array")
parser.add_argument("--json", action="store_true", help="Output as JSON")
args = parser.parse_args()
if args.file:
with open(args.file, "r", encoding="utf-8") as f:
emails = json.load(f)
else:
emails = DEMO_SEQUENCE
if not args.json:
print("No input provided — running in demo mode (5-email nurture sequence).\n")
email_analyses = [analyze_email(e, i) for i, e in enumerate(emails)]
pacing = analyze_pacing(emails)
scoring = compute_sequence_score(email_analyses, pacing)
if args.json:
output = {
"sequence_score": scoring,
"pacing": pacing,
"emails": email_analyses,
}
print(json.dumps(output, indent=2))
return
# Human-readable
overall = scoring["overall"]
grade = scoring["grade"]
print("=" * 64)
print(f" EMAIL SEQUENCE ANALYSIS Score: {overall}/100 Grade: {grade}")
print("=" * 64)
# Pacing summary
print(f"\n 📅 SEQUENCE PACING")
print(f" Emails: {pacing['email_count']}")
print(f" Duration: {pacing.get('total_duration_days', 0)} days")
print(f" Avg gap: {pacing.get('avg_gap_days', 0)} days")
print(f" Cadence: {pacing.get('cadence_type', 'N/A')}")
if pacing.get("issues"):
for issue in pacing["issues"]:
print(f" ⚠️ {issue}")
print(f"\n 📧 PER-EMAIL BREAKDOWN")
print(f" {'#':<3} {'Subject':<40} {'Words':<6} {'CTA':<4} {'Tokens':<7} {'Spam'}")
print(" " + "─" * 60)
for e in email_analyses:
subj = e["subject"]["text"][:38]
if not e["subject"]["length_ok"]:
subj += "⚠️"
words = e["body"]["word_count"]
cta = "✅" if e["body"]["has_cta"] else "❌"
tokens = e["body"]["personalization_tokens"]
spam_lvl = e["spam"]["risk_level"]
spam_icon = "✅" if spam_lvl == "Low" else ("⚠️ " if spam_lvl == "Medium" else "❌")
spam_str = f"{spam_icon}{spam_lvl}"
print(f" {e['email_index']:<3} {subj:<40} {words:<6} {cta:<4} {tokens:<7} {spam_str}")
if any(e["spam"]["trigger_words_found"] for e in email_analyses):
print(f"\n ⚠️ SPAM TRIGGER WORDS DETECTED")
for e in email_analyses:
if e["spam"]["trigger_words_found"]:
triggers = ", ".join(e["spam"]["trigger_words_found"])
print(f" Email {e['email_index']}: {triggers}")
print(f"\n SCORE BREAKDOWN")
for k, v in scoring["breakdown"].items():
label = k.replace("_", " ").title()
bar_len = round(v["score"] / 10)
bar = "█" * bar_len + "░" * (10 - bar_len)
print(f" {label:<22} [{bar}] {v['score']:>3}/100 (weight {v['weight']})")
print()
print("=" * 64)
print(f" Overall: {overall}/100 Grade: {grade}")
print("=" * 64)
if __name__ == "__main__":
main()
Áp dụng 4 nguyên tắc code của Karpathy khi viết, review hoặc commit: nêu rõ giả định, giữ đơn giản, sửa tối thiểu và đặt mục tiêu kiểm chứng được.
---
name: karpathy-coder
description: Use when writing, reviewing, or committing code to enforce Karpathy's 4 coding principles — surface assumptions before coding, keep it simple, make surgical changes, define verifiable goals. Triggers on "review my diff", "check complexity", "am I overcomplicating this", "karpathy check", "before I commit", or any code quality concern where the LLM might be overcoding.
context: fork
version: 2.9.0
author: claude-code-skills
license: MIT
tags: [code-quality, discipline, karpathy, simplicity, surgical-changes, anti-patterns, review]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# Karpathy Coder — Active Coding Discipline
Derived from [Andrej Karpathy's observations](https://x.com/karpathy/status/2015883857489522876) on LLM coding pitfalls. This is **not just guidelines** — it ships Python tools that detect violations, a review agent, a slash command, and a pre-commit hook.
> "The models make wrong assumptions on your behalf and just run along with them without checking. They don't manage their confusion, don't seek clarifications, don't surface inconsistencies, don't present tradeoffs, don't push back when they should."
>
> "They really like to overcomplicate code and APIs, bloat abstractions, don't clean up dead code... implement a bloated construction over 1000 lines when 100 would do."
>
> "LLMs are exceptionally good at looping until they meet specific goals... Don't tell it what to do, give it success criteria and watch it go."
>
> — Andrej Karpathy
## The four principles
### 1. Think Before Coding
**Don't assume. Don't hide confusion. Surface tradeoffs.**
- State assumptions explicitly. If uncertain, ask.
- If multiple interpretations exist, present them — don't pick silently.
- If a simpler approach exists, say so. Push back when warranted.
- If something is unclear, stop. Name what's confusing. Ask.
### 2. Simplicity First
**Minimum code that solves the problem. Nothing speculative.**
- No features beyond what was asked.
- No abstractions for single-use code.
- No "flexibility" or "configurability" that wasn't requested.
- No error handling for impossible scenarios.
- If you write 200 lines and it could be 50, rewrite it.
**The test:** Would a senior engineer say this is overcomplicated? If yes, simplify.
### 3. Surgical Changes
**Touch only what you must. Clean up only your own mess.**
- Don't "improve" adjacent code, comments, or formatting.
- Don't refactor things that aren't broken.
- Match existing style, even if you'd do it differently.
- If you notice unrelated dead code, mention it — don't delete it.
- Remove imports/variables/functions that YOUR changes made unused.
- Don't remove pre-existing dead code unless asked.
**The test:** Every changed line should trace directly to the user's request.
### 4. Goal-Driven Execution
**Define success criteria. Loop until verified.**
| Instead of... | Transform to... |
|---|---|
| "Add validation" | "Write tests for invalid inputs, then make them pass" |
| "Fix the bug" | "Write a test that reproduces it, then make it pass" |
| "Refactor X" | "Ensure tests pass before and after" |
For multi-step tasks, state a brief plan:
```
1. [Step] → verify: [check]
2. [Step] → verify: [check]
3. [Step] → verify: [check]
```
## Slash command
`/karpathy-check` — Run the full 4-principle review on your staged changes.
## Python tools (`scripts/`)
All tools are stdlib-only. Run with `--help`.
| Script | What it detects |
|---|---|
| `complexity_checker.py` | Over-engineering: too many classes, deep nesting, high cyclomatic complexity, unused params, premature abstractions |
| `diff_surgeon.py` | Diff noise: lines that don't trace to the stated goal — comment changes, style drift, drive-by refactors |
| `assumption_linter.py` | Hidden assumptions in a plan: unasked features, missing clarifications, silent interpretation choices |
| `goal_verifier.py` | Weak success criteria: vague plans without verifiable checks, missing test assertions |
## Sub-agent
`karpathy-reviewer` — Runs all 4 principles against a diff. Dispatched by `/karpathy-check` or manually before committing.
## Pre-commit hook
`hooks/karpathy-gate.sh` — runs `complexity_checker.py` and `diff_surgeon.py` on staged files. Warns (non-blocking) when violations are found. Wire it via `.claude/settings.json` or Husky.
## References
- `references/karpathy-principles.md` — the source quotes, deeper context, when to relax each principle
- `references/anti-patterns.md` — 10+ before/after examples across Python, TypeScript, and shell
- `references/enforcement-patterns.md` — how to wire hooks, CI integration, team adoption
## When to relax
These principles bias toward **caution over speed**. For trivial tasks (typo fixes, obvious one-liners), use judgment. The principles matter most on:
- Non-trivial implementations (>20 lines changed)
- Code you don't fully understand
- Multi-step tasks with unclear requirements
- Anything that will be reviewed by humans
## Cross-tool compatibility
Installs via plugin for Claude Code. For other tools, copy the principles into your schema file:
| Tool | Schema file |
|---|---|
| Claude Code | `CLAUDE.md` (auto-loaded by plugin) |
| Codex CLI | `AGENTS.md` |
| Cursor | `AGENTS.md` or `.cursorrules` |
| Antigravity / OpenCode / Gemini CLI | `AGENTS.md` |
## Related skills (chains via `context: fork`)
- **`self-eval`** — honest quality scoring after completing work
- **`code-reviewer`** — broader code review; karpathy-coder focuses on the 4 LLM-specific pitfalls
- **`llm-wiki`** — compound knowledge; karpathy-coder ensures you don't overcomplicate while building it
FILE:expected_outputs/assumption_linter.json
{
"status": "ok",
"source": "stdin",
"total_findings": 3,
"by_category": {
"assumption-just": 1,
"assumption-obvious": 1,
"missing-format": 1
},
"verdict": "REVIEW",
"findings": [
{
"line": 1,
"category": "assumption-just",
"matched": "just",
"message": "'just' often hides complexity. What's being skipped?",
"context": "I'll just add a function to export all user data"
},
{
"line": 1,
"category": "assumption-obvious",
"matched": "Obviously",
"message": "Signals an unstated assumption. Is it really obvious?",
"context": "Obviously we need caching too"
},
{
"line": 1,
"category": "missing-format",
"matched": "export all user data",
"message": "Export/save/fetch mentioned but format not specified (JSON? CSV? API?)",
"context": "I'll just add a function to export all user data"
}
]
}
FILE:expected_outputs/complexity_checker.json
{
"status": "ok",
"threshold": "medium",
"files_analyzed": 1,
"total_findings": 1,
"average_score": 85.0,
"verdict": "WARN",
"results": [
{
"file": "scripts/complexity_checker.py",
"language": "python",
"lines": 210,
"functions": 7,
"classes": 0,
"imports": 6,
"max_nesting": 5,
"avg_cyclomatic": 3.4,
"score": 85,
"findings": [
{
"rule": "nesting-depth",
"severity": "warn",
"message": "Max nesting depth 5 (max 4). Extract or flatten."
}
]
}
]
}
FILE:expected_outputs/diff_surgeon.json
{
"status": "ok",
"files_in_diff": 3,
"total_change_lines": 45,
"noise_lines": 4,
"noise_ratio": 0.09,
"verdict": "CLEAN",
"file_results": []
}
FILE:expected_outputs/goal_verifier.json
{
"status": "ok",
"source": "plan.md",
"steps_found": 4,
"score": 4,
"max_score": 12,
"percentage": 33.3,
"has_final_verification": false,
"verdict": "MISSING",
"step_results": [
{"title": "1. Add validation to the form", "score": 0, "level": "none", "has_verify_label": false},
{"title": "2. Fix the bug with empty inputs", "score": 0, "level": "none", "has_verify_label": false},
{"title": "3. Write test: submit empty form → expect error → verify: test passes", "score": 2, "level": "reasonable", "has_verify_label": true},
{"title": "4. Run full test suite → verify: all green", "score": 2, "level": "reasonable", "has_verify_label": true}
],
"recommendations": [
"2 step(s) have no verification. Add 'verify: [check]' to each.",
"No final/end-to-end verification step. Add one at the end."
]
}
FILE:references/anti-patterns.md
# Anti-Patterns Gallery
Before/after examples for each of the 4 Karpathy principles. Use these to calibrate your reviews.
## Principle #1 — Think Before Coding
### Silent assumption: scope
**Request:** "Add a feature to export user data"
**Bad:** Immediately writes a function that exports ALL users as JSON to disk, picks column names, and adds CSV support.
**Good:**
```
Before implementing, I need to clarify:
1. Which users — all? current tenant? filtered?
2. Which fields — some may be sensitive (PII)
3. Format — JSON, CSV, or API endpoint?
4. Destination — file download, S3, or API response?
Simplest: add an API endpoint returning paginated JSON for authenticated admins.
Want me to go with that, or do you need something different?
```
### Silent assumption: interpretation
**Request:** "Make the search faster"
**Bad:** Adds Redis caching, async queries, AND database indexes in one commit.
**Good:**
```
"Faster" could mean:
- Lower latency per query (add indexes, ~2h)
- Higher throughput under load (add async, ~4h)
- Faster perceived speed (progressive loading, ~3h)
Current: ~500ms p50. Which dimension matters most?
```
## Principle #2 — Simplicity First
### Over-abstraction
**Request:** "Add a function to calculate discount"
**Bad (150 lines):**
```python
class DiscountStrategy(ABC):
@abstractmethod
def calculate(self, amount: float) -> float: ...
class PercentageDiscount(DiscountStrategy): ...
class FixedDiscount(DiscountStrategy): ...
@dataclass
class DiscountConfig:
strategy: DiscountStrategy
min_purchase: float = 0.0
max_discount: float = float('inf')
class DiscountCalculator:
def __init__(self, config: DiscountConfig): ...
def apply_discount(self, amount: float) -> float: ...
```
**Good (3 lines):**
```python
def calculate_discount(amount: float, percent: float) -> float:
return amount * (percent / 100)
```
Add the Strategy pattern when you actually have multiple discount types. Not before.
### Speculative features
**Request:** "Save user preferences to database"
**Bad:** Adds caching, validation, merge mode, notification system, and a `PreferenceManager` class.
**Good:**
```python
def save_preferences(db, user_id: int, preferences: dict):
db.execute("UPDATE users SET preferences = ? WHERE id = ?",
(json.dumps(preferences), user_id))
```
## Principle #3 — Surgical Changes
### Drive-by refactoring
**Request:** "Fix the bug where empty emails crash the validator"
**Bad diff (touches 15 lines, only 2 fix the bug):**
```diff
def validate_user(user_data):
- # Check email format
+ """Validate user data.""" # ← docstring added (not asked)
+ email = user_data.get('email', '').strip()
...
+ if len(username) < 3: # ← username validation (not asked)
+ raise ValueError("Username too short")
```
**Good diff (touches 3 lines, all fix the bug):**
```diff
def validate_user(user_data):
- if not user_data.get('email'):
+ email = user_data.get('email', '')
+ if not email or not email.strip():
raise ValueError("Email required")
- if '@' not in user_data['email']:
+ if '@' not in email:
```
### Style drift
**Request:** "Add logging to the upload function"
**Bad:** Changes quote style, adds type hints, adds docstring, reformats boolean logic.
**Good:** Adds `import logging`, `logger = logging.getLogger(__name__)`, and 3 `logger.info/error` calls. Matches existing single-quote style. Doesn't touch anything else.
## Principle #4 — Goal-Driven Execution
### Vague vs concrete
**Request:** "Fix the authentication system"
**Bad plan:**
```
1. Review the code
2. Identify issues
3. Make improvements
4. Test the changes
```
**Good plan:**
```
Specific issue: users stay logged in after password change.
1. Write test: change password → old session should be invalid
verify: test fails (reproduces bug)
2. Invalidate all sessions on password change
verify: test passes
3. Check edge: multiple active sessions, concurrent changes
verify: additional tests pass
4. Run full auth test suite
verify: all green, no regressions
```
### Missing final verification
**Bad:** "I've added rate limiting. It should work."
**Good:** "Rate limiting added. Verified: sent 11 requests → first 10 got 200, 11th got 429. Existing tests still pass."
## Quick-reference decision table
| Situation | Principle | Action |
|---|---|---|
| Ambiguous requirement | #1 Think | List interpretations, ask |
| "I need a class for this" | #2 Simplicity | Can it be a function? |
| "While I'm here, I'll fix this too" | #3 Surgical | Mention it, don't fix it |
| "This should work" | #4 Goals | What test proves it? |
| User explicitly asked for abstraction | #2 relaxed | Build the abstraction |
| User said "refactor this file" | #3 relaxed | Broader changes are OK |
| One-liner fix, obvious correctness | all relaxed | Use judgment |
FILE:references/enforcement-patterns.md
# Enforcement Patterns
How to wire the Karpathy principles into your workflow so they're enforced, not just documented.
## Level 1 — Passive (read-only)
Install the plugin. The SKILL.md loads into every Claude Code session as context. The LLM reads it and (usually) follows it.
```
/plugin install karpathy-coder@claude-code-skills
```
**Effectiveness:** ~60%. The LLM sometimes forgets under pressure or for long tasks.
## Level 2 — Active review (on demand)
Run `/karpathy-check` before committing. The review agent catches what the LLM missed.
```
# In Claude Code
/karpathy-check
# Or directly from shell
python scripts/complexity_checker.py src/ --threshold strict
python scripts/diff_surgeon.py
```
**Effectiveness:** ~85%. Catches most violations. Requires the user to remember to run it.
## Level 3 — Automated gate (hook)
Wire `hooks/karpathy-gate.sh` as a pre-commit hook. Non-blocking (warns, doesn't reject) but visible.
### Via Husky (Node.js projects)
```bash
npx husky add .husky/pre-commit "bash path/to/karpathy-gate.sh"
```
### Via Claude Code settings
```json
// .claude/settings.json
{
"hooks": {
"PostToolUse": [{
"matcher": "Bash",
"hooks": [{
"type": "command",
"command": "CLAUDE_PLUGIN_ROOT/hooks/karpathy-gate.sh"
}]
}]
}
}
```
### Via pre-commit framework
```yaml
# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: karpathy-complexity
name: Karpathy complexity check
entry: python engineering/karpathy-coder/scripts/complexity_checker.py
language: python
types: [python]
args: [--threshold, medium]
- id: karpathy-diff
name: Karpathy diff surgeon
entry: python engineering/karpathy-coder/scripts/diff_surgeon.py
language: python
always_run: true
```
**Effectiveness:** ~95%. Violations get flagged before they enter the codebase.
## Level 4 — CI integration
Add the tools to your CI pipeline so PRs get Karpathy-reviewed automatically.
### GitHub Actions
```yaml
# .github/workflows/karpathy-review.yml
name: Karpathy Review
on: [pull_request]
jobs:
karpathy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Complexity check
run: |
python engineering/karpathy-coder/scripts/complexity_checker.py \
$(git diff --name-only origin/main...HEAD | grep -E '\.(py|ts|tsx)$' | tr '\n' ' ') \
--threshold medium --json > complexity.json
- name: Diff noise check
run: |
python engineering/karpathy-coder/scripts/diff_surgeon.py \
--diff origin/main...HEAD --json > noise.json
- name: Report
run: |
echo "## Karpathy Review" >> $GITHUB_STEP_SUMMARY
python -c "
import json
c = json.load(open('complexity.json'))
n = json.load(open('noise.json'))
print(f'Complexity: {c[\"average_score\"]}/100 ({c[\"total_findings\"]} findings)')
print(f'Diff noise: {n[\"noise_ratio\"]*100:.0f}% ({n[\"verdict\"]})')
" >> $GITHUB_STEP_SUMMARY
```
## Team adoption
1. **Start with Level 1** for a week. Let the team see the principles in action.
2. **Add Level 2** when reviewing PRs. Run `/karpathy-check` on every PR.
3. **Add Level 3** when the team agrees the principles are useful. Gate commits.
4. **Add Level 4** for repos with multiple contributors or LLM-heavy workflows.
**Anti-pattern:** Going straight to Level 4 without team buy-in. The principles are opinionated — teams should experience them before enforcing them.
FILE:references/karpathy-principles.md
# Karpathy Principles — Full Context
Source: [Andrej Karpathy on X](https://x.com/karpathy/status/2015883857489522876), January 2026.
## The original observations
Karpathy identified four categories of LLM coding failure:
### 1. Assumption management
> "The models make wrong assumptions on your behalf and just run along with them without checking. They don't manage their confusion, don't seek clarifications, don't surface inconsistencies, don't present tradeoffs, don't push back when they should."
**What this means in practice:**
- User says "export user data" → LLM picks JSON, writes to disk, includes all fields, doesn't ask which users
- User says "make it faster" → LLM adds caching, async, and connection pooling without asking what "faster" means
- User says "fix the bug" → LLM guesses which bug based on context, never confirms
**The fix:** Before writing ANY code, list assumptions explicitly. If there are 2+ valid interpretations, present them and ask. If something is unclear, stop and name the confusion.
### 2. Overcomplexity
> "They really like to overcomplicate code and APIs, bloat abstractions, don't clean up dead code... implement a bloated construction over 1000 lines when 100 would do."
**Why LLMs do this:**
- Training data contains enterprise patterns (Strategy, Factory, Observer) applied at inappropriate scale
- "More thorough" feels safe — the LLM can't be wrong for handling edge cases, even if they're impossible
- No cost pressure — generating 1000 lines takes the same effort as generating 100
**The fix:** Ask "would a senior engineer say this is overcomplicated?" after writing. If a function has one caller, it shouldn't be a class. If an abstraction serves one use case, inline it.
### 3. Orthogonal edits
> "They still sometimes change/remove comments and code they don't sufficiently understand as side effects, even if orthogonal to the task."
**Common manifestations:**
- Reformats quote style while fixing a bug
- Adds type annotations to unchanged functions
- "Improves" a comment near the bug fix
- Renames variables in untouched code
- Adds docstrings to functions that weren't changed
**The fix:** Every changed line must trace to the user's request. If you notice something unrelated that could be improved, mention it — don't change it.
### 4. Weak verification loops
> "LLMs are exceptionally good at looping until they meet specific goals... Don't tell it what to do, give it success criteria and watch it go."
**The insight:** LLMs perform dramatically better with declarative goals ("all tests pass") than imperative instructions ("add a try/except block"). The best workflow:
1. Define success criteria as concrete, verifiable checks
2. Let the LLM loop until all checks pass
3. Each step has its own "verify:" annotation
## When to relax each principle
| Principle | Relax when... |
|---|---|
| Think Before Coding | The request is unambiguous and self-contained (e.g., "add a return statement on line 42") |
| Simplicity First | The user explicitly asked for an abstraction, configuration, or extensibility |
| Surgical Changes | The user said "refactor this file" or "clean up this module" |
| Goal-Driven Execution | The task is a one-liner with obvious correctness (e.g., rename a variable) |
## The 80/20 of enforcement
If you adopt only ONE principle, adopt **Surgical Changes** (#3). It's the most measurable (diff analysis), the most commonly violated (LLMs love to "improve" things), and the easiest to check (does the diff contain lines unrelated to the task?).
If you adopt TWO, add **Simplicity First** (#2). Overcomplexity is the second-most-common failure and the most expensive to fix (you ship abstraction debt, then maintain it forever).
FILE:scripts/assumption_linter.py
#!/usr/bin/env python3
"""
assumption_linter.py — Detect hidden assumptions in a plan or proposal.
Karpathy Principle #1 (Think Before Coding): "State your assumptions
explicitly. If uncertain, ask. If multiple interpretations exist, present
them — don't pick silently."
Reads a markdown plan (or stdin) and flags:
- Phrases that indicate silent choices ("I'll just...", "Obviously...", "Simply...")
- Missing scope boundaries ("export" without specifying what/who/how)
- Format/location assumptions without explicit mention
- Single-interpretation language for ambiguous requirements
- Missing error/edge-case consideration
Usage:
python assumption_linter.py plan.md
echo "I'll add a function to export user data" | python assumption_linter.py -
python assumption_linter.py plan.md --json
This is a heuristic tool, not a proof engine. False positives are expected;
the point is to trigger a conversation about assumptions.
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
# --- Pattern library ---
ASSUMPTION_SIGNALS = [
(re.compile(r"\b(?:I'll just|let me just|we can just|just)\b", re.I),
"assumption-just", "'just' often hides complexity. What's being skipped?"),
(re.compile(r"\b(?:obviously|clearly|of course|naturally)\b", re.I),
"assumption-obvious", "Signals an unstated assumption. Is it really obvious?"),
(re.compile(r"\b(?:simply|straightforward|trivial|easy)\b", re.I),
"assumption-simple", "Minimizing language. Could be hiding real complexity."),
(re.compile(r"\b(?:should be fine|should work|shouldn't be a problem)\b", re.I),
"assumption-hopeful", "Hopeful rather than verified. How will you confirm?"),
(re.compile(r"\b(?:I assume|assuming|I'm guessing|probably)\b", re.I),
"assumption-explicit", "At least it's explicit — but have you verified?"),
(re.compile(r"\b(?:all users|every|everything|always|never)\b", re.I),
"scope-absolute", "Absolute scope. Is that really the case?"),
]
MISSING_CLARIFICATION = [
(re.compile(r"\b(?:export|import|save|load|fetch|send)\b.*\b(?:data|file|users)\b", re.I),
"missing-format", "Export/save/fetch mentioned but format not specified (JSON? CSV? API?)"),
(re.compile(r"\b(?:fix|improve|optimize|refactor|update)\b", re.I),
"vague-action", "Vague action verb. What specifically changes? What's the measurable improvement?"),
(re.compile(r"\b(?:handle|deal with|take care of)\b.*\b(?:error|edge|case)\b", re.I),
"vague-error-handling", "Error handling mentioned vaguely. Which errors? What behavior?"),
(re.compile(r"\b(?:the user|users)\b(?!.*\b(?:who|which|specific|certain|admin|role)\b)", re.I),
"unscoped-user", "Which user(s)? All? Specific role? Authenticated only?"),
]
NO_VERIFICATION = [
(re.compile(r"^(?:(?!(?:test|verify|check|assert|confirm|ensure|validate)).)*$", re.I),
"no-verification", "No verification step found in this block. How will you know it works?"),
]
def lint_text(text, source_name="stdin"):
"""Lint a plan text. Return list of findings."""
findings = []
lines = text.splitlines()
for i, line in enumerate(lines, 1):
stripped = line.strip()
if not stripped or stripped.startswith("#"):
continue
for pattern, category, message in ASSUMPTION_SIGNALS:
for m in pattern.finditer(stripped):
findings.append({
"line": i,
"category": category,
"matched": m.group(0),
"message": message,
"context": stripped[:120],
})
for pattern, category, message in MISSING_CLARIFICATION:
if pattern.search(stripped):
findings.append({
"line": i,
"category": category,
"matched": pattern.search(stripped).group(0),
"message": message,
"context": stripped[:120],
})
# Check if any "plan" or numbered-list block lacks verification
plan_blocks = re.findall(r"(?:^|\n)((?:\d+\.\s+.+\n?)+)", text)
for block in plan_blocks:
has_verify = bool(re.search(r"\b(?:test|verify|check|assert|confirm|ensure|validate)\b", block, re.I))
if not has_verify:
findings.append({
"line": 0,
"category": "missing-verification",
"matched": block[:80].replace("\n", " "),
"message": "Plan block has no verification step. Add 'verify:' checks.",
"context": block[:120].replace("\n", " "),
})
return findings
def main():
p = argparse.ArgumentParser(
description="Detect hidden assumptions in a plan or proposal (Karpathy Principle #1).",
epilog="Reads a markdown file or stdin. Flags silent choices, vague actions, and missing verification.",
)
p.add_argument("input", nargs="?", default="-", help="Markdown file to lint, or - for stdin")
p.add_argument("--json", action="store_true", help="JSON output")
args = p.parse_args()
if args.input == "-":
text = sys.stdin.read()
source = "stdin"
else:
path = Path(args.input)
if not path.exists():
print(f"[error] {path} not found", file=sys.stderr)
sys.exit(1)
text = path.read_text(encoding="utf-8", errors="replace")
source = str(path)
findings = lint_text(text, source)
categories = {}
for f in findings:
categories.setdefault(f["category"], []).append(f)
result = {
"status": "ok",
"source": source,
"total_findings": len(findings),
"by_category": {k: len(v) for k, v in categories.items()},
"verdict": "CLEAN" if len(findings) == 0 else ("REVIEW" if len(findings) < 5 else "CLARIFY"),
"findings": findings,
}
if args.json:
print(json.dumps(result, indent=2))
return
print(f"Assumption Linter — {source}")
print(f"Findings: {len(findings)} Verdict: {result['verdict']}")
if findings:
print()
for cat, items in categories.items():
print(f" [{cat}] ({len(items)})")
for item in items[:5]:
line_ref = f"L{item['line']}: " if item["line"] else ""
print(f" {line_ref}{item['message']}")
print(f" → \"{item['matched']}\" in: {item['context'][:80]}")
if len(items) > 5:
print(f" ... and {len(items) - 5} more")
print()
else:
print("\n Plan looks explicit. Assumptions are surfaced.")
print(f"\nVerdict: {result['verdict']}")
if __name__ == "__main__":
main()
FILE:scripts/complexity_checker.py
#!/usr/bin/env python3
"""
complexity_checker.py — Detect over-engineering in Python/TypeScript files.
Karpathy Principle #2 (Simplicity First): "No abstractions for single-use code.
If you write 200 lines and it could be 50, rewrite it."
Checks:
- Cyclomatic complexity (branches per function)
- Class count relative to file size (too many classes = premature abstraction)
- Nesting depth (deep nesting = hard to read)
- Function length (long functions = doing too much)
- Import count (many imports = over-coupled)
- Abstract base classes / protocols for small files (premature patterns)
Usage:
python complexity_checker.py path/to/file.py
python complexity_checker.py src/ --threshold medium
python complexity_checker.py . --ext py,ts --json
Thresholds:
strict — flags aggressively (good for new code)
medium — balanced (default)
relaxed — flags only egregious cases (good for legacy code)
"""
from __future__ import annotations
import argparse
import json
import os
import re
import sys
from pathlib import Path
# --- Thresholds ---
THRESHOLDS = {
"strict": {
"max_cyclomatic": 5,
"max_nesting": 3,
"max_function_lines": 30,
"max_imports": 10,
"max_classes_per_100_lines": 2,
"max_file_lines": 300,
},
"medium": {
"max_cyclomatic": 8,
"max_nesting": 4,
"max_function_lines": 50,
"max_imports": 15,
"max_classes_per_100_lines": 3,
"max_file_lines": 500,
},
"relaxed": {
"max_cyclomatic": 12,
"max_nesting": 5,
"max_function_lines": 80,
"max_imports": 25,
"max_classes_per_100_lines": 5,
"max_file_lines": 1000,
},
}
# --- Analysis functions ---
BRANCH_KEYWORDS_PY = re.compile(
r"^\s*(if |elif |for |while |except |with |and |or |case )", re.MULTILINE
)
BRANCH_KEYWORDS_TS = re.compile(
r"^\s*(if\s*\(|else if|for\s*\(|while\s*\(|catch\s*\(|case |switch\s*\(|\?\?|&&|\|\|)",
re.MULTILINE,
)
FUNC_DEF_PY = re.compile(r"^\s*(?:async\s+)?def\s+(\w+)", re.MULTILINE)
FUNC_DEF_TS = re.compile(
r"^\s*(?:export\s+)?(?:async\s+)?(?:function\s+(\w+)|(?:const|let)\s+(\w+)\s*=\s*(?:async\s+)?\()",
re.MULTILINE,
)
CLASS_DEF_PY = re.compile(r"^\s*class\s+\w+", re.MULTILINE)
CLASS_DEF_TS = re.compile(r"^\s*(?:export\s+)?(?:abstract\s+)?class\s+\w+", re.MULTILINE)
IMPORT_PY = re.compile(r"^(?:import |from \S+ import )", re.MULTILINE)
IMPORT_TS = re.compile(r"^import\s+", re.MULTILINE)
ABC_PATTERN = re.compile(r"ABC|abstractmethod|Protocol|@abstract|Abstract\w+Base", re.MULTILINE)
INDENT_RE = re.compile(r"^( *)\S", re.MULTILINE)
def detect_lang(path):
ext = path.suffix.lower()
if ext in {".py"}:
return "python"
if ext in {".ts", ".tsx", ".js", ".jsx"}:
return "typescript"
return None
def count_branches(text, lang):
pat = BRANCH_KEYWORDS_PY if lang == "python" else BRANCH_KEYWORDS_TS
return len(pat.findall(text))
def extract_functions(text, lang):
"""Return list of (name, start_line, line_count)."""
pat = FUNC_DEF_PY if lang == "python" else FUNC_DEF_TS
lines = text.splitlines()
funcs = []
for m in pat.finditer(text):
name = m.group(1) or (m.group(2) if m.lastindex and m.lastindex >= 2 else "anonymous")
start = text[:m.start()].count("\n")
# Estimate function length: count indented lines until next same-level def or end
indent = len(m.group(0)) - len(m.group(0).lstrip())
end = start + 1
for i in range(start + 1, len(lines)):
stripped = lines[i].rstrip()
if not stripped:
continue
line_indent = len(stripped) - len(stripped.lstrip())
if line_indent <= indent and stripped.lstrip() and not stripped.lstrip().startswith(("#", "//", "/*", "*")):
if lang == "python" and (stripped.lstrip().startswith("def ") or stripped.lstrip().startswith("class ") or stripped.lstrip().startswith("async def ")):
break
if lang == "typescript" and pat.match(stripped):
break
end = i + 1
funcs.append({"name": name, "start_line": start + 1, "lines": end - start})
return funcs
def max_nesting(text, lang):
"""Return the maximum indentation depth in the file."""
if lang == "python":
unit = 4
else:
unit = 2
depths = []
for m in INDENT_RE.finditer(text):
spaces = len(m.group(1))
depths.append(spaces // unit if unit else 0)
return max(depths) if depths else 0
def analyze_file(path, thresholds):
"""Analyze a single file. Return dict with findings."""
text = path.read_text(encoding="utf-8", errors="replace")
lang = detect_lang(path)
if not lang:
return None
lines = text.splitlines()
line_count = len(lines)
findings = []
# File length
if line_count > thresholds["max_file_lines"]:
findings.append({
"rule": "file-length",
"severity": "warn",
"message": f"File is {line_count} lines (max {thresholds['max_file_lines']}). Consider splitting.",
})
# Import count
imp_pat = IMPORT_PY if lang == "python" else IMPORT_TS
import_count = len(imp_pat.findall(text))
if import_count > thresholds["max_imports"]:
findings.append({
"rule": "import-count",
"severity": "warn",
"message": f"{import_count} imports (max {thresholds['max_imports']}). High coupling?",
})
# Class density
cls_pat = CLASS_DEF_PY if lang == "python" else CLASS_DEF_TS
class_count = len(cls_pat.findall(text))
if line_count > 0:
density = class_count / (line_count / 100)
if density > thresholds["max_classes_per_100_lines"]:
findings.append({
"rule": "class-density",
"severity": "warn",
"message": f"{class_count} classes in {line_count} lines ({density:.1f} per 100). Premature abstraction?",
})
# Premature ABC/Protocol in small files
if class_count > 0 and line_count < 200 and ABC_PATTERN.search(text):
findings.append({
"rule": "premature-abstraction",
"severity": "warn",
"message": "Abstract base class / Protocol in a file under 200 lines. Is this needed yet?",
})
# Nesting depth
depth = max_nesting(text, lang)
if depth > thresholds["max_nesting"]:
findings.append({
"rule": "nesting-depth",
"severity": "warn",
"message": f"Max nesting depth {depth} (max {thresholds['max_nesting']}). Extract or flatten.",
})
# Cyclomatic complexity (file-level)
branches = count_branches(text, lang)
funcs = extract_functions(text, lang)
func_count = max(len(funcs), 1)
avg_cyclomatic = branches / func_count
if avg_cyclomatic > thresholds["max_cyclomatic"]:
findings.append({
"rule": "cyclomatic-complexity",
"severity": "warn",
"message": f"Average cyclomatic complexity {avg_cyclomatic:.1f} (max {thresholds['max_cyclomatic']}). Simplify branching.",
})
# Function length
for f in funcs:
if f["lines"] > thresholds["max_function_lines"]:
findings.append({
"rule": "function-length",
"severity": "warn",
"message": f"Function '{f['name']}' is {f['lines']} lines (max {thresholds['max_function_lines']}). Split it.",
"line": f["start_line"],
})
score = max(0, 100 - len(findings) * 15)
return {
"file": str(path),
"language": lang,
"lines": line_count,
"functions": len(funcs),
"classes": class_count,
"imports": import_count,
"max_nesting": depth,
"avg_cyclomatic": round(avg_cyclomatic, 1),
"score": score,
"findings": findings,
}
def collect_files(target, extensions):
target = Path(target)
if target.is_file():
return [target]
files = []
for ext in extensions:
files.extend(target.rglob(f"*.{ext}"))
# Exclude common non-source dirs
skip = {"node_modules", ".git", "__pycache__", ".venv", "venv", "dist", "build"}
return [f for f in files if not any(p in skip for p in f.parts)]
def main():
p = argparse.ArgumentParser(
description="Detect over-engineering in Python/TypeScript files (Karpathy Principle #2).",
epilog="Thresholds: strict (new code), medium (default), relaxed (legacy).",
)
p.add_argument("target", help="File or directory to analyze")
p.add_argument(
"--threshold",
choices=sorted(THRESHOLDS.keys()),
default="medium",
help="Strictness level (default: medium)",
)
p.add_argument(
"--ext",
default="py,ts,tsx,js,jsx",
help="Comma-separated file extensions to scan (default: py,ts,tsx,js,jsx)",
)
p.add_argument("--json", action="store_true", help="JSON output")
args = p.parse_args()
thresholds = THRESHOLDS[args.threshold]
extensions = [e.strip().lstrip(".") for e in args.ext.split(",")]
files = collect_files(args.target, extensions)
if not files:
msg = f"No files found matching extensions: {extensions}"
if args.json:
print(json.dumps({"status": "error", "message": msg}))
else:
print(f"[error] {msg}", file=sys.stderr)
sys.exit(1)
results = []
for f in sorted(files):
r = analyze_file(f, thresholds)
if r:
results.append(r)
total_findings = sum(len(r["findings"]) for r in results)
avg_score = sum(r["score"] for r in results) / len(results) if results else 100
summary = {
"status": "ok",
"threshold": args.threshold,
"files_analyzed": len(results),
"total_findings": total_findings,
"average_score": round(avg_score, 1),
"verdict": "PASS" if total_findings == 0 else ("WARN" if avg_score >= 50 else "FAIL"),
"results": results,
}
if args.json:
print(json.dumps(summary, indent=2))
return
print(f"Karpathy Simplicity Check — {len(results)} files, threshold: {args.threshold}")
print(f"Average score: {avg_score:.0f}/100 Findings: {total_findings}")
print()
for r in results:
if not r["findings"]:
continue
print(f" {r['file']} (score {r['score']}/100)")
for f in r["findings"]:
line = f" line {f['line']}" if "line" in f else ""
print(f" [{f['severity'].upper()}] {f['rule']}{line}: {f['message']}")
print()
if total_findings == 0:
print(" No findings. Code looks appropriately simple.")
print(f"\nVerdict: {summary['verdict']}")
if __name__ == "__main__":
main()
FILE:scripts/diff_surgeon.py
#!/usr/bin/env python3
"""
diff_surgeon.py — Detect diff noise: changes that don't trace to the stated goal.
Karpathy Principle #3 (Surgical Changes): "Every changed line should trace
directly to the user's request."
Analyzes a git diff and flags:
- Comment-only changes (unrelated to the task)
- Whitespace / formatting changes
- Import additions not used by the new code
- Style changes (quote style, trailing commas, semicolons)
- Docstring additions to unchanged functions
- Variable renames in untouched code
- Type annotation additions to unchanged signatures
Usage:
python diff_surgeon.py # analyze staged diff
python diff_surgeon.py --diff HEAD~1..HEAD # analyze last commit
python diff_surgeon.py --file changes.diff # analyze a diff file
python diff_surgeon.py --json
Exit codes:
0 clean — all changes look intentional
1 noise detected — review before committing
"""
from __future__ import annotations
import argparse
import json
import re
import subprocess
import sys
from pathlib import Path
# --- Noise detectors ---
COMMENT_ONLY = re.compile(r"^[+-]\s*(?:#|//|/\*|\*|<!--)")
WHITESPACE_ONLY = re.compile(r"^[+-]\s*$")
QUOTE_CHANGE = re.compile(r'^[+-]\s*.*["\'].*["\']')
DOCSTRING_ADD = re.compile(r'^[+]\s*"""')
IMPORT_LINE = re.compile(r"^[+]\s*(?:import |from \S+ import |const .* = require)")
TYPE_ANNOTATION = re.compile(r"^[+-].*:\s*(?:str|int|float|bool|list|dict|Optional|Union|Any|string|number|boolean)\b")
SEMICOLON_CHANGE = re.compile(r"^[+-].*;\s*$")
TRAILING_COMMA = re.compile(r"^[+-].*,\s*$")
def get_diff(args):
"""Get diff text from args."""
if args.file:
return Path(args.file).read_text(encoding="utf-8", errors="replace")
diff_range = args.diff or "--staged"
cmd = ["git", "diff", diff_range] if diff_range != "--staged" else ["git", "diff", "--staged"]
try:
result = subprocess.run(cmd, capture_output=True, text=True, timeout=30)
return result.stdout
except (subprocess.TimeoutExpired, FileNotFoundError) as e:
print(f"[error] git diff failed: {e}", file=sys.stderr)
sys.exit(1)
def parse_hunks(diff_text):
"""Parse a unified diff into per-file hunks."""
files = []
current_file = None
current_lines = []
for line in diff_text.splitlines():
if line.startswith("diff --git"):
if current_file:
files.append({"file": current_file, "lines": current_lines})
# Extract filename: diff --git a/path b/path
parts = line.split(" b/")
current_file = parts[-1] if len(parts) > 1 else "unknown"
current_lines = []
elif line.startswith("+++ ") or line.startswith("--- "):
continue
elif line.startswith("@@"):
current_lines.append({"type": "hunk_header", "text": line})
elif line.startswith("+") or line.startswith("-"):
current_lines.append({"type": "change", "text": line})
if current_file:
files.append({"file": current_file, "lines": current_lines})
return files
def classify_line(line_text):
"""Classify a changed line. Returns a noise category or None if intentional."""
if WHITESPACE_ONLY.match(line_text):
return "whitespace"
if COMMENT_ONLY.match(line_text):
return "comment-only"
if DOCSTRING_ADD.match(line_text):
return "docstring-addition"
if SEMICOLON_CHANGE.match(line_text):
# Check if ONLY change is semicolon
stripped = line_text[1:].rstrip(";").rstrip()
if not stripped.strip():
return None
return "semicolon-style"
return None
def analyze_file_diff(file_data):
"""Analyze a single file's diff for noise."""
findings = []
change_lines = [l for l in file_data["lines"] if l["type"] == "change"]
total_changes = len(change_lines)
if total_changes == 0:
return findings
# Detect paired +/- that are only whitespace/style changes
additions = [l["text"] for l in change_lines if l["text"].startswith("+")]
deletions = [l["text"] for l in change_lines if l["text"].startswith("-")]
noise_count = 0
for line_data in change_lines:
category = classify_line(line_data["text"])
if category:
noise_count += 1
findings.append({
"category": category,
"line": line_data["text"][:120],
})
# Detect quote-style swaps (paired changes where only quotes differ)
for a, d in zip(sorted(additions), sorted(deletions)):
a_norm = a[1:].replace('"', "'").strip()
d_norm = d[1:].replace('"', "'").strip()
if a_norm == d_norm and a[1:].strip() != d[1:].strip():
findings.append({
"category": "quote-style-swap",
"line": f"{d[:60]} → {a[:60]}",
})
noise_ratio = noise_count / total_changes if total_changes > 0 else 0
return findings
def main():
p = argparse.ArgumentParser(
description="Detect diff noise — changes that don't trace to the stated goal (Karpathy Principle #3).",
epilog="Run before committing to catch drive-by refactors and style drift.",
)
p.add_argument("--diff", default=None, help="Git diff range (e.g. HEAD~1..HEAD). Default: staged changes.")
p.add_argument("--file", default=None, help="Read diff from a file instead of git")
p.add_argument("--json", action="store_true", help="JSON output")
args = p.parse_args()
diff_text = get_diff(args)
if not diff_text.strip():
result = {"status": "ok", "message": "No diff to analyze", "files": 0, "noise_lines": 0, "verdict": "CLEAN"}
if args.json:
print(json.dumps(result, indent=2))
else:
print("No diff to analyze. Stage changes first (git add) or specify --diff range.")
return
file_diffs = parse_hunks(diff_text)
all_findings = []
file_results = []
for fd in file_diffs:
findings = analyze_file_diff(fd)
if findings:
file_results.append({"file": fd["file"], "findings": findings})
all_findings.extend(findings)
total_noise = len(all_findings)
total_changes = sum(
len([l for l in fd["lines"] if l["type"] == "change"]) for fd in file_diffs
)
noise_ratio = total_noise / total_changes if total_changes > 0 else 0
verdict = "CLEAN" if noise_ratio < 0.1 else ("NOISY" if noise_ratio < 0.3 else "VERY_NOISY")
result = {
"status": "ok",
"files_in_diff": len(file_diffs),
"total_change_lines": total_changes,
"noise_lines": total_noise,
"noise_ratio": round(noise_ratio, 2),
"verdict": verdict,
"file_results": file_results,
}
if args.json:
print(json.dumps(result, indent=2))
return
print(f"Diff Surgeon — {len(file_diffs)} files, {total_changes} changed lines")
print(f"Noise ratio: {noise_ratio:.0%} ({total_noise} noise lines)")
print(f"Verdict: {verdict}")
if file_results:
print()
for fr in file_results:
print(f" {fr['file']}:")
categories = {}
for f in fr["findings"]:
categories.setdefault(f["category"], []).append(f["line"])
for cat, lines in categories.items():
print(f" [{cat}] {len(lines)} instance(s)")
for l in lines[:3]:
print(f" {l}")
if len(lines) > 3:
print(f" ... and {len(lines) - 3} more")
print()
print("Recommendation: review flagged lines. Remove changes that don't trace to your task.")
else:
print("\n All changes look intentional. Clean diff.")
sys.exit(1 if verdict != "CLEAN" else 0)
if __name__ == "__main__":
main()
FILE:scripts/goal_verifier.py
#!/usr/bin/env python3
"""
goal_verifier.py — Check if a plan has verifiable success criteria.
Karpathy Principle #4 (Goal-Driven Execution): "Define success criteria.
Loop until verified. Don't tell it what to do — give it success criteria
and watch it go."
Reads a markdown plan and scores:
- Does each step have a verification check?
- Are success criteria concrete (test, assertion, measurement)?
- Are there vague criteria ("make it work", "looks good")?
- Is there a final verification step?
Usage:
python goal_verifier.py plan.md
python goal_verifier.py plan.md --json
Scoring:
Each plan step gets 0-3 points:
3 = concrete verification (test assertion, metric, command)
2 = reasonable verification (manual check, visual)
1 = vague verification ("should work", "looks right")
0 = no verification mentioned
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
CONCRETE_VERIFY = re.compile(
r"\b(?:test\s+pass|assert|assertEqual|expect\(|\.toBe|\.toEqual|"
r"exit\s+code\s*[=:]\s*0|status\s*[=:]\s*200|curl\s|"
r"grep\s|diff\s|python.*test|npm\s+test|pytest|jest|"
r"measure|benchmark|metric|latency\s*<|throughput\s*>)\b",
re.I,
)
REASONABLE_VERIFY = re.compile(
r"\b(?:verify|check|confirm|inspect|review|compare|validate|"
r"run\s+and\s+see|manually|open\s+in\s+browser|visual|screenshot)\b",
re.I,
)
VAGUE_VERIFY = re.compile(
r"\b(?:should\s+work|looks?\s+(?:good|right|fine|ok)|"
r"seems?\s+(?:correct|fine)|hopefully|probably\s+works?)\b",
re.I,
)
STEP_PATTERN = re.compile(r"^(?:\d+[\.\)]\s+|[-*]\s+\[.\]\s+|[-*]\s+(?:Step\s+\d+))", re.M)
VERIFY_LABEL = re.compile(r"(?:verify|check|success\s+criteria|done\s+when|acceptance)\s*:", re.I)
def extract_steps(text):
"""Extract plan steps from markdown."""
lines = text.splitlines()
steps = []
current_step = None
current_body = []
for line in lines:
if STEP_PATTERN.match(line.strip()):
if current_step:
steps.append({"title": current_step, "body": "\n".join(current_body)})
current_step = line.strip()
current_body = []
elif current_step:
current_body.append(line)
if current_step:
steps.append({"title": current_step, "body": "\n".join(current_body)})
return steps
def score_step(step):
"""Score a step's verification quality (0-3)."""
full_text = step["title"] + "\n" + step["body"]
if CONCRETE_VERIFY.search(full_text):
return 3, "concrete"
if VERIFY_LABEL.search(full_text) and REASONABLE_VERIFY.search(full_text):
return 2, "reasonable"
if REASONABLE_VERIFY.search(full_text):
return 2, "reasonable"
if VAGUE_VERIFY.search(full_text):
return 1, "vague"
return 0, "none"
def analyze_plan(text, source):
"""Analyze a plan for verification quality."""
steps = extract_steps(text)
if not steps:
return {
"status": "ok",
"source": source,
"steps_found": 0,
"message": "No numbered/bulleted plan steps found. Is this a plan?",
"verdict": "NO_PLAN",
"score": 0,
"max_score": 0,
"step_results": [],
}
step_results = []
total_score = 0
max_score = len(steps) * 3
for step in steps:
pts, level = score_step(step)
total_score += pts
step_results.append({
"title": step["title"][:120],
"score": pts,
"level": level,
"has_verify_label": bool(VERIFY_LABEL.search(step["body"])),
})
# Check for final verification
has_final = False
if steps:
last_full = steps[-1]["title"] + steps[-1]["body"]
if re.search(r"\b(?:final|end-to-end|full.*test|regression|all.*pass)\b", last_full, re.I):
has_final = True
pct = (total_score / max_score * 100) if max_score > 0 else 0
if pct >= 70:
verdict = "STRONG"
elif pct >= 40:
verdict = "WEAK"
else:
verdict = "MISSING"
return {
"status": "ok",
"source": source,
"steps_found": len(steps),
"score": total_score,
"max_score": max_score,
"percentage": round(pct, 1),
"has_final_verification": has_final,
"verdict": verdict,
"step_results": step_results,
"recommendations": _recommendations(step_results, has_final),
}
def _recommendations(step_results, has_final):
recs = []
none_steps = [s for s in step_results if s["level"] == "none"]
vague_steps = [s for s in step_results if s["level"] == "vague"]
if none_steps:
recs.append(f"{len(none_steps)} step(s) have no verification. Add 'verify: [check]' to each.")
if vague_steps:
recs.append(f"{len(vague_steps)} step(s) have vague criteria. Replace 'should work' with a concrete check.")
if not has_final:
recs.append("No final/end-to-end verification step. Add one at the end.")
if not recs:
recs.append("Plan has strong verification coverage. Good to go.")
return recs
def main():
p = argparse.ArgumentParser(
description="Check if a plan has verifiable success criteria (Karpathy Principle #4).",
epilog="Scores each step 0-3 based on verification quality.",
)
p.add_argument("input", nargs="?", default="-", help="Markdown plan file, or - for stdin")
p.add_argument("--json", action="store_true", help="JSON output")
args = p.parse_args()
if args.input == "-":
text = sys.stdin.read()
source = "stdin"
else:
path = Path(args.input)
if not path.exists():
print(f"[error] {path} not found", file=sys.stderr)
sys.exit(1)
text = path.read_text(encoding="utf-8", errors="replace")
source = str(path)
result = analyze_plan(text, source)
if args.json:
print(json.dumps(result, indent=2))
return
print(f"Goal Verifier — {source}")
print(f"Steps: {result['steps_found']} Score: {result['score']}/{result['max_score']} ({result['percentage']}%)")
print(f"Verdict: {result['verdict']}")
print()
for sr in result["step_results"]:
icon = {"concrete": "+", "reasonable": "~", "vague": "?", "none": "!"}[sr["level"]]
print(f" [{icon}] {sr['title'][:100]} ({sr['level']}, {sr['score']}/3)")
print()
for rec in result["recommendations"]:
print(f" -> {rec}")
if __name__ == "__main__":
main()
Lập kế hoạch ra mắt sản phẩm, công bố tính năng hoặc chiến lược phát hành, gồm Product Hunt, beta, early access, waitlist và checklist go-to-market.
---
name: launch
description: "When the user wants to plan a product launch, feature announcement, or release strategy. Also use when the user mentions 'launch,' 'Product Hunt,' 'feature release,' 'announcement,' 'go-to-market,' 'beta launch,' 'early access,' 'waitlist,' 'product update,' 'how do I launch this,' 'launch checklist,' 'GTM plan,' or 'we're about to ship.' Use this whenever someone is preparing to release something publicly. For ongoing marketing after launch, see marketing-ideas. For the offer being launched (bonuses, guarantees, scarcity, naming), see offers."
metadata:
version: 2.0.2
---
# Launch Strategy
You are an expert in SaaS product launches and feature announcements. Your goal is to help users plan launches that build momentum, capture attention, and convert interest into users.
## Before Starting
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
---
## Core Philosophy
The best companies don't just launch once—they launch again and again. Every new feature, improvement, and update is an opportunity to capture attention and engage your audience.
A strong launch isn't about a single moment. It's about:
- Getting your product into users' hands early
- Learning from real feedback
- Making a splash at every stage
- Building momentum that compounds over time
---
## The ORB Framework
Structure your launch marketing across three channel types. Everything should ultimately lead back to owned channels.
### Owned Channels
You own the channel (though not the audience). Direct access without algorithms or platform rules.
**Examples:**
- Email list
- Blog
- Podcast
- Branded community (Slack, Discord)
- Website/product
**Why they matter:**
- Get more effective over time
- No algorithm changes or pay-to-play
- Direct relationship with audience
- Compound value from content
**Start with 1-2 based on audience:**
- Industry lacks quality content → Start a blog
- People want direct updates → Focus on email
- Engagement matters → Build a community
**Example - Superhuman:**
Built demand through an invite-only waitlist and one-on-one onboarding sessions. Every new user got a 30-minute live demo. This created exclusivity, FOMO, and word-of-mouth—all through owned relationships. Years later, their original onboarding materials still drive engagement.
### Rented Channels
Platforms that provide visibility but you don't control. Algorithms shift, rules change, pay-to-play increases.
**Examples:**
- Social media (Twitter/X, LinkedIn, Instagram)
- App stores and marketplaces
- YouTube
- Reddit
**How to use correctly:**
- Pick 1-2 platforms where your audience is active
- Use them to drive traffic to owned channels
- Don't rely on them as your only strategy
**Example - Notion:**
Hacked virality through Twitter, YouTube, and Reddit where productivity enthusiasts were active. Encouraged community to share templates and workflows. But they funneled all visibility into owned assets—every viral post led to signups, then targeted email onboarding.
**Platform-specific tactics:**
- Twitter/X: Threads that spark conversation → link to newsletter
- LinkedIn: High-value posts → lead to gated content or email signup
- Marketplaces (Shopify, Slack): Optimize listing → drive to site for more
Rented channels give speed, not stability. Capture momentum by bringing users into your owned ecosystem.
### Borrowed Channels
Tap into someone else's audience to shortcut the hardest part—getting noticed.
**Examples:**
- Guest content (blog posts, podcast interviews, newsletter features)
- Collaborations (webinars, co-marketing, social takeovers)
- Speaking engagements (conferences, panels, virtual summits)
- Influencer partnerships
**Be proactive, not passive:**
1. List industry leaders your audience follows
2. Pitch win-win collaborations
3. Use tools like SparkToro or Listen Notes to find audience overlap
4. Set up affiliate/referral incentives (for channel partner launches, use [Introw](../../tools/integrations/introw.md) to manage deal registration and commissions)
**Example - TRMNL:**
Sent a free e-ink display to YouTuber Snazzy Labs—not a paid sponsorship, just hoping he'd like it. He created an in-depth review that racked up 500K+ views and drove $500K+ in sales. They also set up an affiliate program for ongoing promotion.
Borrowed channels give instant credibility, but only work if you convert borrowed attention into owned relationships.
---
## Readiness Gate: Are You Ready to Launch?
Run this **before** the phased mechanics. Products don't market themselves—but a product that isn't ready won't market either. The launch mechanics only pay off if what you're launching is worth launching.
Two failure modes kill launches from opposite ends:
- **Stealth Mode** — launching too late. "Procrastination in a fancy suit." You keep polishing in private, waiting for the product to be perfect. It never ships, and nobody learns you exist.
- **"Just One More Feature"** — never launching. Every proposed launch date gets pushed for one more thing. The scope creeps forever; the launch never comes.
The middle path is **SLC — Simple, Lovable, Complete** (Jason Cohen), the antidote to shipping a bare MVP that's minimal but unlovable. Don't launch a stub nobody wants; don't wait for a bloated everything-app. A launchable v1 is:
- **Simple** — it does *one* thing. Not many things poorly. One clear job, done well.
- **Lovable** — people *want* to use it, not just tolerate it. An MVP asks users to suffer through a stripped-down experience "to give feedback." SLC gives them something they'd choose. If nobody would be sad to lose it, it isn't lovable yet.
- **Complete** — it's a *whole* experience for that one thing, not a stub with obvious holes. Complete at its chosen scope, not a teaser of a bigger promise.
**The gate:** If it's not yet Simple, Lovable, and Complete, you're in "Just One More Feature" territory only when adding scope is what's missing—otherwise you're in Stealth Mode and should ship. Cut scope until one thing is lovable and complete, then launch that. SLC gives you a real launch now instead of a perfect launch never.
**Quick check before running the phases:**
- [ ] Does it do one clearly-defined thing? (Simple)
- [ ] Would a target user *choose* to use it, not just endure it? (Lovable)
- [ ] Is that one thing a whole experience, with no glaring stubs? (Complete)
- [ ] Are you polishing past this bar? → Stop. You're in Stealth Mode. Ship.
- [ ] Are you still adding new things to the scope? → Stop. You're in "Just One More Feature." Cut back to SLC.
Pass the gate, then run the phases below.
---
## Five-Phase Launch Approach
Launching isn't a one-day event. It's a phased process that builds momentum.
### Phase 1: Internal Launch
Gather initial feedback and iron out major issues before going public.
**Actions:**
- Recruit early users one-on-one to test for free
- Collect feedback on usability gaps and missing features
- Ensure prototype is functional enough to demo (doesn't need to be production-ready)
**Goal:** Validate core functionality with friendly users.
### Phase 2: Alpha Launch
Put the product in front of external users in a controlled way.
**Actions:**
- Create landing page with early access signup form
- Announce the product exists
- Invite users individually to start testing
- MVP should be working in production (even if still evolving)
**Goal:** First external validation and initial waitlist building.
### Phase 3: Beta Launch
Scale up early access while generating external buzz.
**Actions:**
- Work through early access list (some free, some paid)
- Start marketing with teasers about problems you solve
- Recruit friends, investors, and influencers to test and share
**Consider adding:**
- Coming soon landing page or waitlist
- "Beta" sticker in dashboard navigation
- Email invites to early access list
- Early access toggle in settings for experimental features
**Goal:** Build buzz and refine product with broader feedback.
### Phase 4: Early Access Launch
Shift from small-scale testing to controlled expansion.
**Actions:**
- Leak product details: screenshots, feature GIFs, demos
- Gather quantitative usage data and qualitative feedback
- Run user research with engaged users (incentivize with credits)
- Optionally run product/market fit survey to refine messaging
**Expansion options:**
- Option A: Throttle invites in batches (5-10% at a time)
- Option B: Invite all users at once under "early access" framing
**Goal:** Validate at scale and prepare for full launch.
### Phase 5: Full Launch
Open the floodgates.
**Actions:**
- Open self-serve signups
- Start charging (if not already)
- Announce general availability across all channels
**Launch touchpoints:**
- Customer emails
- In-app popups and product tours
- Website banner linking to launch assets
- "New" sticker in dashboard navigation
- Blog post announcement
- Social posts across platforms
- Product Hunt, BetaList, Hacker News, etc.
**Goal:** Maximum visibility and conversion to paying users.
---
## Product Hunt Launch Strategy
Product Hunt can be powerful for reaching early adopters, but it's not magic—it requires preparation.
### Pros
- Exposure to tech-savvy early adopter audience
- Credibility bump (especially if Product of the Day)
- Potential PR coverage and backlinks
### Cons
- Very competitive to rank well
- Short-lived traffic spikes
- Requires significant pre-launch planning
### How to Launch Successfully
**Before launch day:**
1. Build relationships with influential supporters, content hubs, and communities
2. Optimize your listing: compelling tagline, polished visuals, short demo video
3. Study successful launches to identify what worked
4. Engage in relevant communities—provide value before pitching
5. Prepare your team for all-day engagement
**On launch day:**
1. Treat it as an all-day event
2. Respond to every comment in real-time
3. Answer questions and spark discussions
4. Encourage your existing audience to engage
5. Direct traffic back to your site to capture signups
**After launch day:**
1. Follow up with everyone who engaged
2. Convert Product Hunt traffic into owned relationships (email signups)
3. Continue momentum with post-launch content
### Case Studies
**SavvyCal** (Scheduling tool):
- Optimized landing page and onboarding before launch
- Built relationships with productivity/SaaS influencers in advance
- Responded to every comment on launch day
- Result: #2 Product of the Month
**Reform** (Form builder):
- Studied successful launches and applied insights
- Crafted clear tagline, polished visuals, demo video
- Engaged in communities before launch (provided value first)
- Treated launch as all-day engagement event
- Directed traffic to capture signups
- Result: #1 Product of the Day
---
## Post-Launch Product Marketing
Your launch isn't over when the announcement goes live. Now comes adoption and retention work.
### Immediate Post-Launch Actions
**Educate new users:**
Set up automated onboarding email sequence introducing key features and use cases.
**Reinforce the launch:**
Include announcement in your weekly/biweekly/monthly roundup email to catch people who missed it.
**Differentiate against competitors:**
Publish comparison pages highlighting why you're the obvious choice.
**Update web pages:**
Add dedicated sections about the new feature/product across your site.
**Offer hands-on preview:**
Create no-code interactive demo (using tools like Navattic) so visitors can explore before signing up.
### Keep Momentum Going
It's easier to build on existing momentum than start from scratch. Every touchpoint reinforces the launch.
---
## Ongoing Launch Strategy
Don't rely on a single launch event. Regular updates and feature rollouts sustain engagement.
### How to Prioritize What to Announce
Use this matrix to decide how much marketing each update deserves:
**Major updates** (new features, product overhauls):
- Full campaign across multiple channels
- Blog post, email campaign, in-app messages, social media
- Maximize exposure
**Medium updates** (new integrations, UI enhancements):
- Targeted announcement
- Email to relevant segments, in-app banner
- Don't need full fanfare
**Minor updates** (bug fixes, small tweaks):
- Changelog and release notes
- Signal that product is improving
- Don't dominate marketing
### Announcement Tactics
**Space out releases:**
Instead of shipping everything at once, stagger announcements to maintain momentum.
**Reuse high-performing tactics:**
If a previous announcement resonated, apply those insights to future updates.
**Keep engaging:**
Continue using email, social, and in-app messaging to highlight improvements.
**Signal active development:**
Even small changelog updates remind customers your product is evolving. This builds retention and word-of-mouth—customers feel confident you'll be around.
---
## Launch Checklist
### Pre-Launch
- [ ] Landing page with clear value proposition
- [ ] Email capture / waitlist signup
- [ ] Early access list built
- [ ] Owned channels established (email, blog, community)
- [ ] Rented channel presence (social profiles optimized)
- [ ] Borrowed channel opportunities identified (podcasts, influencers)
- [ ] Product Hunt listing prepared (if using)
- [ ] Launch assets created (screenshots, demo video, GIFs)
- [ ] Onboarding flow ready
- [ ] Analytics/tracking in place
### Launch Day
- [ ] Announcement email to list
- [ ] Blog post published
- [ ] Social posts scheduled and posted
- [ ] Product Hunt listing live (if using)
- [ ] In-app announcement for existing users
- [ ] Website banner/notification active
- [ ] Team ready to engage and respond
- [ ] Monitor for issues and feedback
### Post-Launch
- [ ] Onboarding email sequence active
- [ ] Follow-up with engaged prospects
- [ ] Roundup email includes announcement
- [ ] Comparison pages published
- [ ] Interactive demo created
- [ ] Gather and act on feedback
- [ ] Plan next launch moment
---
## Task-Specific Questions
1. What are you launching? (New product, major feature, minor update)
2. What's your current audience size and engagement?
3. What owned channels do you have? (Email list size, blog traffic, community)
4. What's your timeline for launch?
5. Have you launched before? What worked/didn't work?
6. Are you considering Product Hunt? What's your preparation status?
---
## Related Skills
- **marketing-ideas**: For additional launch tactics (#22 Product Hunt, #23 Early Access Referrals)
- **emails**: For launch and onboarding email sequences
- **cro**: For optimizing launch landing pages
- **marketing-psychology**: For psychology behind waitlists and exclusivity
- **programmatic-seo**: For comparison pages mentioned in post-launch
- **sales-enablement**: For launch sales collateral and enablement materials
FILE:evals/evals.json
{
"skill_name": "launch",
"evals": [
{
"id": 1,
"prompt": "We're launching a new B2B SaaS product for design teams in 6 weeks. It's a design review tool. We have a small audience (500 email subscribers, 2k Twitter followers). Help us plan the launch.",
"expected_output": "Should check for product-marketing.md first. Should apply the ORB Framework (Owned, Rented, Borrowed channels) with the user's specific resources. Owned: email list (500 subscribers), website. Rented: Twitter (2k followers). Borrowed: partnerships, communities, Product Hunt. Should recommend the five-phase launch approach with a timeline mapped to the 6-week window: Internal prep, Alpha (existing network), Beta (expanded), Early Access, Full Launch. Should provide specific tactics for each phase. Should recommend building up the audience before launch day. Should include a launch day checklist.",
"assertions": [
"Checks for product-marketing.md",
"Applies ORB Framework (Owned, Rented, Borrowed)",
"Maps to user's specific channels and audience sizes",
"Recommends five-phase launch approach",
"Provides timeline mapped to 6-week window",
"Provides specific tactics for each phase",
"Recommends audience building before launch",
"Includes launch day checklist"
],
"files": []
},
{
"id": 2,
"prompt": "We want to launch on Product Hunt. Any tips? We've never done it before.",
"expected_output": "Should apply the Product Hunt strategy section. Should cover: choosing the right day and time, preparing assets (logo, gallery images, maker video), crafting the tagline and description, building a hunter network, activating supporters on launch day, engaging with comments, and post-launch follow-up. Should recommend preparation timeline (start 2-4 weeks before). Should mention common mistakes to avoid. Should set realistic expectations about outcomes.",
"assertions": [
"Applies Product Hunt strategy section",
"Covers timing (day and time selection)",
"Covers asset preparation",
"Addresses hunter network and supporter activation",
"Recommends preparation timeline",
"Mentions common mistakes to avoid",
"Sets realistic expectations"
],
"files": []
},
{
"id": 3,
"prompt": "we just shipped a major feature update. how should we announce it? it's not a full product launch, just a big new feature.",
"expected_output": "Should trigger on casual phrasing. Should apply the ongoing launch strategy section, specifically the major/medium/minor update matrix. Should identify this as a major feature update. Should recommend appropriate channels and tactics for a feature launch (less than a full product launch but more than a changelog entry). Should include: announcement email, blog post, social media push, in-app notification, and possibly a mini Product Hunt launch. Should provide a feature announcement framework.",
"assertions": [
"Triggers on casual phrasing",
"Applies ongoing launch strategy / update matrix",
"Identifies as major feature update",
"Scales tactics appropriately (not full launch)",
"Recommends announcement channels",
"Includes email, blog, social, and in-app notification",
"Provides feature announcement framework"
],
"files": []
},
{
"id": 4,
"prompt": "Our launch flopped. We launched 3 weeks ago and only got 50 signups. We expected at least 500. What went wrong and what can we do now?",
"expected_output": "Should apply the post-launch product marketing section. Should diagnose potential failure causes: insufficient audience building pre-launch, wrong channels, weak value proposition messaging, poor launch execution, targeting the wrong audience. Should recommend post-launch recovery tactics: iterate on messaging, identify which channels produced the 50 signups and double down, try new distribution channels, leverage early users for testimonials. Should provide a specific 30-day recovery plan.",
"assertions": [
"Applies post-launch product marketing guidance",
"Diagnoses potential failure causes",
"Addresses pre-launch audience building gap",
"Recommends post-launch recovery tactics",
"Suggests analyzing which channels produced signups",
"Provides specific recovery plan"
],
"files": []
},
{
"id": 5,
"prompt": "How do we leverage partnerships and borrowed audiences for our launch? We don't have a big audience of our own.",
"expected_output": "Should focus on the Borrowed channel from the ORB Framework. Should provide specific borrowed audience tactics: podcast guest appearances, co-marketing with complementary tools, influencer partnerships, community engagement (relevant Slack groups, Discord servers, Reddit), guest posts, cross-promotions. Should recommend how to identify and approach potential partners. Should note that borrowed audience strategies take time to build and should start well before launch day.",
"assertions": [
"Focuses on Borrowed channel from ORB Framework",
"Provides specific borrowed audience tactics",
"Mentions partnerships, communities, guest content",
"Recommends how to identify and approach partners",
"Notes borrowed strategies take time to build",
"Suggests starting well before launch day"
],
"files": []
},
{
"id": 6,
"prompt": "Give me some creative marketing ideas to promote our product. We're bootstrapped and don't have a big budget.",
"expected_output": "Should recognize this is a broader marketing ideas request, not specifically a launch strategy task. Should defer to or cross-reference the marketing-ideas skill, which provides 139 marketing ideas organized by category and filtered by budget. May provide some launch-related tactical ideas but should make clear that marketing-ideas is the right skill for a broader brainstorming session.",
"assertions": [
"Recognizes this as broader marketing ideas request",
"References or defers to marketing-ideas skill",
"Does not attempt full marketing brainstorm using launch strategy patterns",
"May provide some launch-related tactical ideas"
],
"files": []
},
{
"id": 7,
"prompt": "We've been building our product in private for 8 months and keep pushing the launch date because there's always one more feature we want to add first. How do we know when we're actually ready to launch?",
"expected_output": "Should apply the Readiness Gate before jumping to launch mechanics. Should name the two failure modes: Stealth Mode (launching too late, 'procrastination in a fancy suit') and 'Just One More Feature' (never launching), and identify that this user is exhibiting both. Should introduce SLC (Simple, Lovable, Complete) by Jason Cohen as the middle path and the alternative to a bare MVP. Should explain each: Simple (does one thing), Lovable (users want to use it, not just tolerate it), Complete (a whole experience, not a stub). Should advise cutting scope to reach SLC on one thing rather than adding more features, and to ship once the gate is passed. Should not just dump the five-phase mechanics without addressing readiness first.",
"assertions": [
"Applies the readiness gate before launch mechanics",
"Names Stealth Mode and 'Just One More Feature' failure modes",
"Diagnoses the user as stuck in these failure modes",
"Introduces SLC (Simple, Lovable, Complete) as the middle path vs a bare MVP",
"Explains Simple, Lovable, and Complete distinctly",
"Advises cutting scope to reach SLC rather than adding features",
"Recommends shipping once the gate is passed"
],
"files": []
}
]
}
Giảm chi phí API LLM: tối ưu token, chọn mô hình phù hợp, triển khai prompt caching, nên dùng khi chi phí AI tăng cao hoặc sắp ra mắt tính năng AI.
--- name: llm-cost-optimizer description: "Use proactively whenever LLM API costs come up -- or should. Triggers include: 'my AI costs are too high', 'optimize token usage', 'which model should I use', 'LLM spend is out of control', 'implement prompt caching', 'we're about to launch an AI feature', 'build me an AI endpoint'. Don't wait for an explicit cost complaint -- if someone is building an AI feature, designing an LLM endpoint, or choosing between models, cost architecture belongs in the conversation. Apply immediately when any of these are true: a system prompt appears that exceeds a few hundred tokens, all requests are hitting the same model, max_tokens is not set, or no per-feature cost logging exists. NOT for RAG pipeline design (use rag-architect). NOT for improving prompt quality or effectiveness (use senior-prompt-engineer)." --- # LLM Cost Optimizer You are an expert in LLM cost engineering with deep experience reducing AI API spend at scale. Your goal is to cut LLM costs by 40–80% without degrading user-facing quality -- using model routing, caching, prompt compression, and observability to make every token count. AI API costs are engineering costs. Treat them like database query costs: measure first, optimize second, monitor always. --- ## Step 0: Classify Before You Ask Before gathering context, classify which mode applies based on what the user has already said. Pull answers from the conversation first -- don't ask for what you already have. | Mode | When to use | |---|---| | **Cost Audit** | Spend exists but no clear picture of where it goes | | **Optimize Existing System** | Cost drivers are known; apply targeted fixes | | **Design Cost-Efficient Architecture** | Building new AI features; wire in cost controls before launch | If the mode is ambiguous, ask in one shot using the context questions below. Only ask what you don't already know. --- ## Context You Need **Current State** - Which LLM providers and models are in use? - Monthly spend? Which features/endpoints drive it? - Token usage logging in place? Cost-per-request visibility? **Goals** - Target cost reduction? (e.g., "cut 50%", "stay under $X/month") - Latency constraints? (affects caching and routing tradeoffs) - Quality floor? (what degradation is acceptable?) **Workload Profile** - Request volume and distribution (p50, p95, p99 token counts)? - Repeated or similar prompts? (caching potential) - Mix of task types? (classification vs. generation vs. reasoning) --- ## Mode 1: Cost Audit Use when spend exists but the breakdown is unknown. Instrument first; optimize second. **Step 1 -- Instrument Every Request** Log per-request: model, input tokens, output tokens, latency, endpoint/feature, user segment, cost (calculated). **Step 2 -- Find the 20% Causing 80% of Spend** Sort by: feature × model × token count. Usually 2–3 endpoints drive the majority of cost. Target those first. **Step 3 -- Classify Requests by Complexity** | Complexity | Characteristics | Right Model Tier | |---|---|---| | Simple | Classification, extraction, yes/no, short output | Small (Haiku, GPT-4o-mini, Gemini Flash) | | Medium | Summarization, structured output, moderate reasoning | Mid (Sonnet, GPT-4o) | | Complex | Multi-step reasoning, code gen, long context | Large (Opus, o3) | **If token logging doesn't exist yet:** That's the first deliverable -- not prompt compression, not routing. You cannot optimize what you cannot see. Provide a logging schema and move to optimization only once baseline data exists. --- ## Mode 2: Optimize Existing System Apply techniques in ROI order. Don't skip ahead -- measure impact at each step before moving to the next. ### 1. Model Routing (60–80% cost reduction on routed traffic) Route by task complexity, not by default. Use a lightweight classifier or rule engine. - **Small models**: classification, extraction, simple Q&A, formatting, short summaries - **Mid models**: structured output, moderate summarization, code completion - **Large models**: complex reasoning, long-context analysis, agentic tasks, code generation Even routing 20% of traffic to a cheaper model produces meaningful savings. Start there. ### 2. Prompt Caching (40–90% reduction on cacheable traffic) Supported by Anthropic (`cache_control`), OpenAI (automatic on some models), Google (context caching). Cache-eligible content: system prompts, static context, document chunks, few-shot examples. Target hit rates: >60% for document Q&A, >40% for chatbots with static system prompts. **Flag immediately** if a system prompt exceeds ~2,000 tokens and is sent on every request -- this is a high-value caching target. ### 3. Output Length Control (20–40% reduction) LLMs over-generate by default. Force conciseness: - Explicit length instructions: "Respond in 3 sentences or fewer." - Schema-constrained output: JSON with defined fields beats free-text - `max_tokens` hard caps: set per endpoint, not globally - Stop sequences: define terminators for list and structured outputs **Flag immediately** if `max_tokens` is not set per endpoint -- every uncapped endpoint is a cost leak. ### 4. Prompt Compression (15–30% input token reduction) Remove filler without losing meaning. Audit each prompt for token efficiency. | Before | After | |---|---| | "Please carefully analyze the following text and provide..." | "Analyze:" | | "It is important that you remember to always..." | "Always:" | | Context already in system prompt, repeated in user message | Remove | | HTML or markdown when plain text works | Strip tags | **Caution:** Over-compression causes hallucination and low-quality outputs, triggering retries that erase the savings. Compress filler; preserve task-critical instructions. ### 5. Semantic Caching (30–60% hit rate on repeated queries) Cache LLM responses keyed by embedding similarity, not exact match. Serve cached responses for semantically equivalent questions. Tools: GPTCache, LangChain cache, custom Redis + embedding lookup. Threshold guidance: cosine similarity >0.95 = safe to serve cached response. ### 6. Request Batching (10–25% reduction via amortized overhead) Batch non-latency-sensitive requests. Process async queues off-peak. --- ## Mode 3: Design Cost-Efficient Architecture Wire these controls in before launch -- retrofitting is more expensive. **Budget Envelopes** -- per feature, per user tier, per day. Set hard limits and soft alerts at 80% of limit. **Routing Layer** -- classify → route → call. Never call the large model by default. **Tier Your Model Access** -- free users do not need the most expensive model. Assign model tiers by user tier at design time. **Cost Observability Dashboard** -- spend by feature, spend by model, cost per active user, week-over-week trend, anomaly alerts. This is not optional; it is the monitoring foundation. **Graceful Degradation** -- when budget is exceeded: switch to smaller model → serve cached response → queue for async processing. --- ## Proactive Flags Surface these without being asked, regardless of which mode is active: | Signal | Action | |---|---| | No per-feature cost breakdown | Instrument logging before any other change | | All requests hitting one model | Model monoculture = #1 overspend pattern; initiate routing design | | System prompt >2,000 tokens, sent every request | Flag as high-value caching target | | `max_tokens` not set per endpoint | Flag as active cost leak | | No cost alerts configured | Spend spikes go undetected for days; set p95 cost-per-request alerts | | Free tier users consuming same model as paid | Tier model access by user tier | --- ## Failure Modes and Recovery | Situation | Response | |---|---| | No token logs exist | Stop. Logging schema is deliverable #1. Return once baseline data is available. | | User can't identify which feature drives spend | Provide an instrumentation plan; schedule a cost review after 2 weeks of data. | | Routing classifier adds latency that exceeds constraint | Fall back to rule-based routing (token count thresholds, endpoint tags) instead of ML classifier. | | Cache hit rate is below 20% | Diagnose: are prompts highly variable? Is context dynamic? Recommend semantic caching or rethink what's being cached. | | Prompt compression degrades quality | Restore compressed section. Flag the specific instruction as compression-resistant. | --- ## Handoff Triggers If the conversation shifts to one of these, pause and invoke the relevant skill rather than continuing inline: - **Prompt quality or effectiveness deteriorates** → invoke `senior-prompt-engineer` - **Retrieval pipeline design comes up** → invoke `rag-architect` - **Broader monitoring stack beyond cost metrics** → invoke `observability-designer` - **Latency profiling becomes the primary concern** → invoke `performance-profiler` --- ## Output Artifacts | Request | Deliverable | |---|---| | Cost audit | Per-feature spend breakdown, top 3 optimization targets, projected savings | | Model routing design | Routing decision tree with model recommendations per task type and estimated cost delta | | Caching strategy | What to cache, cache key design, expected hit rate, implementation pattern | | Prompt optimization | Token-by-token audit with compression suggestions and before/after token counts | | Architecture review | Cost-efficiency scorecard (0–100) with prioritized fixes and projected monthly savings | --- ## Communication Standard - **Bottom line first** -- cost impact before explanation - **What + Why + How** -- every finding includes all three - **Actions have owners and deadlines** -- no vague "consider optimizing..." - **Confidence tagging** -- verified / medium / assumed --- ## Anti-Patterns | Anti-Pattern | Why It Fails | Better Approach | |---|---|---| | Using the largest model for every request | 80%+ of requests are simple tasks a smaller model handles equally well, wasting 5–10x on cost | Implement a routing layer that classifies complexity and selects the cheapest adequate model | | Optimizing prompts without measuring first | You cannot know what to optimize without per-feature spend visibility | Instrument token logging and cost-per-request before any changes | | Caching by exact string match only | Minor phrasing differences cause cache misses on semantically identical queries | Use embedding-based semantic caching with a cosine similarity threshold | | Setting a single global max_tokens | Some endpoints need 2,000 tokens, others need 50 -- a global cap either wastes or truncates | Set max_tokens per endpoint based on measured p95 output length | | Ignoring system prompt size | A 3,000-token system prompt sent on every request is a hidden cost multiplier | Use prompt caching for static system prompts; strip unnecessary instructions | | Treating cost optimization as a one-time project | Model pricing changes, traffic patterns shift, new features launch -- costs drift | Set up continuous cost monitoring with weekly spend reports and anomaly alerts | | Compressing prompts to the point of ambiguity | Over-compressed prompts cause hallucination or low-quality output, requiring retries | Compress filler and redundant context; preserve all task-critical instructions |
Chuyển file markdown thành HTML một file có tương tác nhẹ: tài liệu dài, review code kèm diff và gắn mức độ nghiêm trọng, hoặc bộ slide.
---
name: markdown-html-orchestrator
description: Use when a user wants to convert any markdown file in their Claude project into a single-file, lightly-interactive HTML — long-form documents (specs, plans, RFCs, reports, explainers), code reviews with diffs and severity-tagged annotations, or slide decks. Triggers on "convert this markdown to HTML", "make this an HTML file", "turn this into an interactive document", "render this report as HTML", "PR writeup as HTML", "slides from this markdown". Forks context to route to one of three converter sub-skills (md-document, md-review, md-slides) based on a deterministic doctype classifier, after the user has run the design-system onboarding once. Refuses if input is under 100 lines (per Shihipar — markdown still wins below the threshold) or design-system isn't onboarded. Distinct from Anthropic's official Playground plugin (which is interactive prompt-tuning controls with sliders/knobs/prompt-copy-back) and from marketing/landing/ (which is a landing-page generator).
context: fork
version: 2.10.0
author: Alireza Rezvani
license: MIT
tags: [markdown, html, converter, orchestrator, documentation, code-review, slides, design-system]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# Markdown → HTML — Domain Orchestrator
Thariq Shihipar's argument (Claude Code HTML output essay, Medium 2026): **markdown collapses past 100 lines for agent-generated artifacts.** Long specs, code reviews, and architecture explainers lose density, hierarchy, and lightweight interaction the moment they exceed a screen of text. HTML restores all three — single-file, browser-native, shareable.
This orchestrator forks context, classifies the input markdown deterministically, routes to the right converter sub-skill, and returns a digest with the output path. Heavy intake (full markdown bodies, diffs, slide decks) stays in the forked context.
**Foundation status (v2.10.0):** orchestrator + `design-system` (onboarding + shared brand tokens) are live. Converter sub-skills (`md-document`, `md-review`, `md-slides`) land in v2.10.1 follow-up PRs. Until they land, this skill still runs the classifier and the design-system gate, and surfaces the routing recommendation — it just hands the rendering work back to Claude with the structured brief.
## When to invoke
| Symptom | Sub-skill |
|---|---|
| "Convert this RFC / spec / report / explainer to HTML" — long-form doc | `md-document` |
| "Turn this PR writeup / code review into HTML" — markdown with diff blocks | `md-review` |
| "Make a slide deck from this markdown" — `---` boundaries or H1 cadence | `md-slides` |
## Pre-flight gates (hard refusals)
1. **Below the 100-line threshold.** Markdown wins below 100 lines (Shihipar). The classifier prints `below_min_lines: true` and `route_explainer.py` refuses. Tell the user to keep their input as markdown.
2. **Design-system not onboarded.** If `~/.config/markdown-html/design-system.json` doesn't exist (or its `setup_completed_at` is null), refuse. Point the user at `python3 markdown-html/skills/design-system/scripts/onboard.py` (or `--defaults` for a zero-touch run).
3. **Unwritable save location.** `output_path_resolver.py` refuses if the configured `default_output_dir` (or `--out` override) isn't writable.
## Routing logic (deterministic)
Two-signal threshold pattern lifted from `research-ops/skills/research-ops-skills/SKILL.md`. Filename hint = 2 points; each content signal = 1 point. Silent-route allowed when winner ≥ 3 AND (runner-up = 0 OR winner ≥ 2× runner-up). Below threshold → one clarifying question with a recommended answer.
### Signal table
| Signal class | Filename hints | Content signals | Sub-skill |
|---|---|---|---|
| DOCUMENT | `report.md`, `*-doc.md`, `spec.md`, `rfc-*.md`, `*-analysis.md`, `*-explainer.md` | `## Table of Contents` (2), `^# `, `^## `, markdown table rows, `> [!NOTE]/[!TIP]/[!IMPORTANT]` callouts | `md-document` |
| REVIEW | `review.md`, `*-pr-*.md`, `*.diff.md`, `code-review*.md` | ` ```diff ` (2), `^[-+]{3} ` (2), `^@@` (2), `> [!BLOCKER]/[!MAJOR]/[!MINOR]/[!NIT]` (2), `LGTM`/`nit:`/`blocker:` | `md-review` |
| SLIDES | `deck.md`, `slides.md`, `*-talk.md`, `presentation*.md` | `^---$` ≥ 3 (2 + per-boundary), `<!-- notes:` (2), H1 count ≥ 5 with median gap ≤ 12 lines (2) | `md-slides` |
The pipeline:
```bash
python3 skills/markdown-html-orchestrator/scripts/doctype_classifier.py \
--input <path>.md --output json \
| python3 skills/markdown-html-orchestrator/scripts/route_explainer.py
```
`route_explainer.py` checks the design-system status, applies the < 100-line refusal, and prints one of: `ROUTE_SILENTLY -> md-<type>`, `ASK_USER one question: ...`, or `REFUSE — fix the issues above`.
## Workflow
### Step 1 — Confirm onboarding
If the user has never run onboarding, surface the one-time setup:
```bash
python3 markdown-html/skills/design-system/scripts/onboard.py
```
Ten questions, 1-2 minutes. Captures brand primary + accent + heading/body Google Fonts + design style (editorial/technical/minimal/playful) + default output dir + syntax theme + TOC behavior + optional logo/company. Stored at `~/.config/markdown-html/design-system.json`. Re-runnable with `--scope project` for per-repo overrides.
### Step 2 — Classify the input
Run `doctype_classifier.py` on the markdown. Inspect the verdict.
### Step 3 — Route or ask
Pipe the classification into `route_explainer.py`. If it says `ROUTE_SILENTLY`, forward the original markdown + the design-system config into the named sub-skill's renderer in the forked context. If it says `ASK_USER`, ask ONE question with the recommended answer.
### Step 4 — Resolve the output path
```bash
python3 skills/markdown-html-orchestrator/scripts/output_path_resolver.py \
--input <path>.md --doctype <document|review|slides>
```
Collision handling defaults to `-2 / -3 / ...` suffix; `--on-collision timestamp` for stamped names.
### Step 5 — Hand off to the sub-skill (when shipped)
In v2.10.1+, the converter sub-skill's renderer takes the input markdown, the design-system config, and the resolved output path, and writes a single self-contained HTML file. The orchestrator returns a ≤ 100-word digest: input lines, output path, design style applied, top 3 features used (TOC, search, code-copy, etc.), and one forcing question for the user.
Until v2.10.1, the orchestrator's job stops at step 4 — it returns the classification + routing brief and lets Claude do the rendering inline with the design-system tokens.
## Forcing-question library (Matt Pocock grill-with-docs pattern)
Walk these one at a time, with a recommended answer per question, citing the canon. Lift this list into `/cs:grill-markdown-html` for plan-stage interrogation.
1. **What decision does this HTML drive — is the reader skimming, deciding, or presenting?**
Recommended: name it first; density follows from purpose. Canon: Shihipar — "match output format to consumption context"; Tufte — *Visual Display of Quantitative Information*, ch. 1.
2. **Is the input markdown ≥ 100 lines?**
Recommended: yes — below that, keep it as markdown. Canon: Shihipar — markdown still wins under 100 lines.
3. **Is the design-system onboarded?**
Recommended: yes, globally (`~/.config/markdown-html/design-system.json`). Canon: research-ops onboarding pattern (`research-ops/CLAUDE.md` §8); WCAG 2.2 §1.4.3 (text contrast 4.5:1).
4. **Where does the output save, and will it overwrite anything?**
Recommended: the configured `default_output_dir` with `--on-collision suffix`. Canon: Matt Pocock `handoff` skill — never silently overwrite a working artifact.
5. **Document type confidence — silent-route or one question?**
Recommended: silent-route only when the classifier's verdict is one of `document/review/slides` AND `silent_route_allowed: true`. Otherwise ask. Canon: research-ops two-signal threshold (`research-ops/skills/research-ops-skills/SKILL.md` §"Routing logic").
Never run a sub-skill before the lane is locked.
## Assumptions
1. User has a markdown file ≥ 100 lines they want to convert.
2. User has run onboarding once (`~/.config/markdown-html/design-system.json` exists with `setup_completed_at` populated).
3. Single-file HTML output is acceptable (no multi-file site, no embedded server, no build step).
4. Externals limited to Google Fonts CSS + Prism.js CDN (jsdelivr / cdnjs).
## Non-goals
- Not a landing-page generator (use `marketing/landing/`).
- Not an interactive prompt-tuning playground (use Anthropic's official `playground` plugin).
- Not a static-site generator (no multi-file output, no site index).
- Not a PDF generator (slides use `@media print`; user prints from browser).
- Not a watch / live-reload pipeline (conversion is one-shot).
## Distinct from
- **Anthropic Playground plugin** (`/playground`) — builds interactive controls (sliders, knobs, drag-drop) for prompt tuning, with a copy-prompt-back loop. This plugin converts existing markdown documents to HTML. Different tools for different jobs.
- **`marketing/landing/`** — generates landing pages from scratch (Phase-0 intake → 3 sections → branded TSX/HTML). This plugin converts an existing markdown file you already have.
- **`engineering/handoff/` + `productivity/handoff/`** — preserve session continuity between Claude conversations. Different artifact type (handoff brief vs. document conversion).
## Output artifacts
| Sub-skill | Artifact | Status |
|---|---|---|
| `md-document` | `doc-<slug>.html` (single file, sticky TOC, collapsibles, search, code-copy, scrollspy) | v2.10.1 |
| `md-review` | `review-<slug>.html` (2-col diff + severity margin notes + jump-nav) | v2.10.1 |
| `md-slides` | `deck-<slug>.html` (arrow-key nav + presenter mode + print-to-PDF) | v2.10.1 |
## Anti-patterns (do not)
- ❌ Convert markdown < 100 lines — markdown still wins. Refuse and tell the user.
- ❌ Run the orchestrator before the design-system is onboarded. The output looks broken without tokens.
- ❌ Silently chain two sub-skills (e.g., "convert doc AND make slides from it"). Pick one, finish, ask before chaining.
- ❌ Use external JS frameworks (React/Vue/Svelte). Vanilla JS + IntersectionObserver only. Prism.js CDN is the single exception.
- ❌ Multi-file output (extracted CSS, asset directories). Single file or nothing — that's the whole point.
- ❌ Overwrite an existing output file by default. The path resolver suffixes `-2`, `-3`, …; `--on-collision overwrite` is opt-in only.
## References
- Spec: Thariq Shihipar — "Claude Code HTML output" (Medium, 2026)
- Forking pattern: `research-ops/skills/research-ops-skills/SKILL.md` (`context: fork`, two-signal routing)
- Customization pattern: `research-ops/skills/clinical-research/scripts/` (`onboard.py`, `config_loader.py`)
- Brand palette math: `marketing/landing/skills/landing/scripts/brand_palette_validator.py` (WCAG + HSL derive)
- Information-density canon: Tufte; Shihipar's `thariqs.github.io/html-effectiveness/` gallery; Wattenberger interactive essays; Maggie Appleton digital gardens
FILE:references/information_density_canon.md
# Information Density Canon
**Why this exists:** Thariq Shihipar's central claim is empirical: markdown collapses past ~100 lines because it lacks the visual machinery to manage density. This document anchors that claim in a longer tradition — from Edward Tufte's *Visual Display* to Maggie Appleton's digital gardens — so the orchestrator can defend the 100-line threshold against pushback ("why not 50?", "why not 200?") with cited evidence rather than vibes.
## Core claim
A reader skimming linear markdown loses orientation after roughly 5-7 screens. HTML restores orientation through:
1. **Hierarchy made visible** — typography scale, color, weight, indent, surface
2. **Lateral navigation** — TOC, scrollspy, anchored sections
3. **Lateral structure** — tables, grids, side-by-side comparisons
4. **Lateral interaction** — collapsibles, search, code-copy, hover state
Markdown collapses each of these into the same channel: indented text. HTML opens each into its own channel.
## Sources
### 1. Thariq Shihipar — "Claude Code HTML output" (Medium, 2026)
The spec for this plugin. Key claims used here:
- Threshold ≈ 100 lines: "I stopped reading markdown files past 100 lines. My threshold was about the same. Yours probably is too."
- Three forces converged: agent outputs got longer, editing relationship changed (LLM edits, not human), information became spatial.
- Five advantages: density, clarity, shareability, two-way interaction, context ingestion.
- Examples gallery: `thariqs.github.io/html-effectiveness/` (20 self-contained HTML files across 9 categories).
### 2. Edward Tufte — *The Visual Display of Quantitative Information* (Graphics Press, 1983/2001)
Foundational text on data-ink ratio and small multiples. Specifically:
- Ch. 1, "Graphical Excellence" — graphics should reveal the data; markdown's linear structure conceals comparison.
- Ch. 4, "Data-Ink and Graphical Redesign" — every visual element should earn its place. HTML's collapsibles and tabs are data-ink positive (they reveal more per pixel than the same content laid out linearly).
### 3. Bret Victor — "Up and Down the Ladder of Abstraction" (2011, worrydream.com)
Argues that interactive controls let a reader move fluidly between concrete examples and abstract rules. The "lightweight interactivity" tier of this plugin (search, collapsibles, hover tooltips) is the documents-equivalent: it lets a reader move between TOC abstraction and section detail without losing place.
### 4. Maggie Appleton — *A Brief History & Ethos of the Digital Garden* (2020, maggieappleton.com)
Establishes the "garden" pattern: persistent, interlinked, editable knowledge artifacts rendered as HTML. Reinforces single-file HTML as the right artifact shape for long-form thinking (vs. blog posts as linear sequences). Many of her gardens use the exact patterns this plugin generates: sticky TOC, collapsibles, callouts.
### 5. Amelia Wattenberger — "Why React isn't great for actually building websites" + interactive essay archive (wattenberger.com)
Demonstrates lightweight interactivity in essays without frameworks — IntersectionObserver, vanilla scroll handling, inline SVG. The exact technical patterns md-document will use.
### 6. Bartosz Ciechanowski — *Internal Combustion Engine* and other essays (ciechanow.ski)
The high-water mark of single-page interactive explainers. Each essay is a single HTML file with inline SVG animation and controls. Validates the single-file-HTML-as-artifact thesis at the upper bound.
### 7. GitHub READMEs-as-landing-pages (2021-present)
Empirically, READMEs that exceed ~200 lines either (a) get split into a `docs/` folder or (b) get an HTML-rendered version (e.g., GitBook, Docusaurus, mdBook). The market has already voted on the 100-200-line threshold.
## Practical takeaway for the orchestrator
When `doctype_classifier.below_min_lines` is true, refuse the conversion and quote Shihipar. The threshold is empirically defended and stylistically consistent with the wider canon of information-design discipline.
FILE:references/orchestrator_routing_patterns.md
# Orchestrator Routing Patterns
**Why this exists:** The two-signal routing discipline (silent-route only above a confidence threshold; otherwise ask one question with a recommended answer) is not original to this plugin. It's been established in the research-ops, commercial, and business-operations domains. This document records the canon so the orchestrator never silently chains or guesses below threshold.
## The pattern
Three discrete behaviors based on the classifier's score:
1. **Silent route** — winner ≥ 3 points AND (runner-up = 0 OR winner ≥ 2× runner-up). Hand off to the sub-skill without asking.
2. **Clarify** — winner ≥ 2 points but ratio against runner-up is too close. Ask ONE question, recommend the winner, take the user's confirmation or override.
3. **Ambiguous** — no signals matched. Ask which lane, default to md-document if the user shrugs.
Filename hint counts double (2 points each) because filename is high-signal user intent — a file named `pr-review.md` is almost certainly a code review.
## Sources
### 1. research-ops/skills/research-ops-skills/SKILL.md §"Routing logic (deterministic)"
The two-signal threshold pattern formalized: "Same two-signal threshold pattern as `commercial-skills`. Single-signal → clarifying question. Mixed signals → highest-confidence first, chain second in a follow-up turn. Never silently chain."
### 2. commercial/skills/commercial-skills/SKILL.md
First domain to ship the explicit "never silently chain" rule, with named signal classes (PRICING / DEAL / PARTNERSHIPS / RFP / FORECAST). The discipline is independent of subject matter — same shape for research, for commercial deals, for markdown docs.
### 3. business-operations/skills/business-operations-skills/SKILL.md
The "explore the workspace first" pattern: filenames like `vendor-list.csv` or `sla-tracker.xlsx` resolve the lane without asking. Filename hint = 2 points is calibrated here.
### 4. Matt Pocock — *grill-with-docs* (engineering/grill-with-docs/SKILL.md, MIT)
Five rules formalized:
1. One question per turn — never bundle.
2. Always recommend an answer with citation-backed rationale.
3. Explore before asking.
4. Walk the decision tree depth-first.
5. Track dependencies (don't ask Q3 before Q1's answer determines whether Q3 applies).
### 5. Anthropic — `context: fork` (SKILL.md frontmatter)
The mechanism that makes orchestrator routing efficient: forked sub-skills run in isolated context, so the parent thread doesn't bloat with the full markdown body, the diff hunks, or the slide bodies. Documented in research-ops, commercial, and business-operations orchestrators.
### 6. The "never silently chain" hard rule
Originates from research-ops Sprint 1 design (`documentation/implementation/research-ops-expansion-plan.md`). The rule prevents the worst orchestrator failure mode: routing to two sub-skills in sequence without explicit user acknowledgment of the chain. Markdown-html applies it: "convert this markdown to HTML and also make slides from it" is two operations, asked explicitly.
### 7. NN/g — *Defaults Are the Best Friend of UX* (Jakob Nielsen, 2007)
Recommended answers in clarifying questions reduce decision fatigue. The orchestrator never asks an open question — every clarification ships with "Recommended: <answer>, because <rationale>" so the user can just say "yes."
## Applied to markdown-html
The classifier produces a `total_scores` dict. The orchestrator's decision tree:
```
if below_min_lines: → REFUSE (Shihipar 100-line rule)
elif not setup_completed_at: → REFUSE (point at onboarding)
elif winner_score == 0: → ASK_USER (lane + recommend md-document)
elif silent_route_allowed: → ROUTE_SILENTLY to md-<winner>
elif winner_score >= 2: → ASK_USER (recommend md-<winner>)
else: → ASK_USER (treat as md-document by default)
```
Never two routes in one turn. Never "I'll just do both."
FILE:references/single_file_html_discipline.md
# Single-File HTML Discipline
**Why this exists:** Multi-file HTML output (separate CSS, JS, images, asset folders) breaks the central value proposition: shareability. The recipient can't drop the file into Slack, attach it to an email, or upload it to a static host with one drag. This document codifies the single-file constraint and names the few permitted exceptions.
## The constraint
Every converter (md-document, md-review, md-slides) MUST produce one `.html` file containing all CSS and all JavaScript inline. The only externals permitted are:
1. **Google Fonts CSS** — pulled from `fonts.googleapis.com` via `<link rel="stylesheet">`. Falls back to system stack if blocked.
2. **Prism.js** — pulled from `cdn.jsdelivr.net` or `cdnjs.cloudflare.com` for syntax highlighting. Falls back to plain `<pre>` if blocked.
No other CDN. No build step. No bundler. No framework runtime.
## Why
### Shareability
A single .html file uploads to S3, Vercel, Netlify, or any static host in one operation. It also opens in a recipient's browser without a server, which means it works in:
- Slack DM previews
- Email attachments (Gmail / Outlook web)
- Local `file://` URLs
- GitHub `raw.githubusercontent.com` links
- USB sticks given to a non-technical reviewer
Multi-file output breaks every one of those flows. The marketing/landing/ skill made the same choice for the same reason.
### Portability
Single-file HTML survives copying, archiving, and email-attachment workflows. It's the closest thing to PDF that the web has, with the advantage of being editable and searchable.
### No build-step regret
The moment you require a build step, you require: a Node version, a package.json, a node_modules folder, a transpiler, a watcher, a runtime, and a deployment story. None of that survives "send this to a teammate."
## Permitted CDN externals — discipline
```html
<!-- Google Fonts (CSS only — woff files lazy-loaded by browser) -->
<link rel="preconnect" href="https://fonts.googleapis.com" crossorigin>
<link rel="stylesheet"
href="https://fonts.googleapis.com/css2?family=Inter:wght@400;600&display=swap">
<!-- Prism.js core + theme + autoloader (gracefully degrades without it) -->
<link rel="stylesheet"
href="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/themes/prism-tomorrow.min.css">
<script defer
src="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/components/prism-core.min.js"></script>
<script defer
src="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/plugins/autoloader/prism-autoloader.min.js"></script>
```
Both fall back gracefully: if the CDN is blocked, fonts default to the system stack and code blocks render as plain `<pre>`. The page is still readable, still searchable, still copy-pasteable.
## Anti-patterns
- ❌ External CSS file (`<link rel="stylesheet" href="./style.css">`) — recipient gets a broken page.
- ❌ External JS file (`<script src="./app.js">`) — same problem.
- ❌ External image references for hero/logo (`<img src="./logo.png">`) — base64-embed instead.
- ❌ React/Vue/Svelte/Alpine runtime — vanilla JS only.
- ❌ Tailwind via CDN (`cdn.tailwindcss.com`) — 200 KB of unused CSS; just inline what you use.
- ❌ Web Components requiring a custom-element registry from CDN — same problem.
## Sources
### 1. Thariq Shihipar — "Claude Code HTML output" (Medium, 2026)
"Every playground is a single HTML file with all CSS and JavaScript inlined. No external dependencies. No build step. Open it in any browser." (Playground plugin section — same discipline applies to converted documents.)
### 2. marketing/landing/skills/landing/SKILL.md §"Single-File HTML Discipline"
Established the rule for this repo. Marketing landing pages were the first artifact type to require single-file output; this plugin inherits the discipline directly.
### 3. Tom MacWright — "Big" (github.com/tmcw/big, MIT)
A presentation tool that compiles to a single HTML file. Demonstrates the upper bound of what's possible with the constraint (full slide deck, presenter mode, navigation, in one file).
### 4. Mozilla MDN — *Performance: Reducing HTTP Requests*
Single-file output minimizes round trips. Even on fast networks, a single 200 KB HTML file beats one HTML + three CSS + five JS + four image requests.
### 5. The Web We Lost — Anil Dash (2012, dashes.com)
Argues for portable, host-anywhere web artifacts as a counter to platform lock-in. Single-file HTML is the most portable web artifact possible — no platform, no JS framework, no server.
### 6. Prism.js documentation (prismjs.com)
Lightweight syntax highlighter (~2 KB core + per-language plugins on demand) designed for CDN delivery. The right tradeoff for "single-file with one allowed external."
### 7. Google Fonts API documentation (developers.google.com/fonts/docs/css2)
The `display=swap` parameter ensures system-font fallback while web fonts load, preventing FOUT/FOIT on slow connections. Required parameter for every Google Fonts link the converters emit.
## Applied to markdown-html
`md-document/scripts/html_renderer.py`, `md-review/scripts/review_html_renderer.py`, and `md-slides/scripts/deck_html_renderer.py` all emit single-file output with exactly the two permitted externals. Anything else is a regression.
FILE:scripts/doctype_classifier.py
#!/usr/bin/env python3
"""doctype_classifier.py - Deterministic document-type classifier for markdown-html.
Stdlib-only. Reads a markdown file (or stdin), scans for filename + content signals,
and returns a routing recommendation: document / review / slides / ambiguous.
Routing discipline mirrors research-ops/skills/research-ops-skills/SKILL.md:
- Two-signal threshold: silent-route when score >= 3 OR (winner >= 2 AND
winner >= 2x runner-up). Below threshold => ambiguous, ask the user.
- Filename hint = 2 points; each content signal = 1 point.
- Never silently chain. The orchestrator (Claude) decides; this script
just produces a structured recommendation it can act on.
Hard rule from the article: documents below MIN_LINES are NOT candidates for
HTML conversion — markdown still wins. The classifier surfaces a line_count
field and a below_min_lines boolean so the orchestrator can refuse before
routing.
NO LLM CALLS. Pure regex + counting.
Usage:
python doctype_classifier.py --input report.md --output json
python doctype_classifier.py --input - --output human # stdin
python doctype_classifier.py --sample
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any
MIN_LINES = 100 # Shihipar's threshold — markdown wins below this
FILENAME_HINTS: dict[str, list[str]] = {
"document": [
r"\breport\b", r"-doc\b", r"\bspec\b", r"^rfc-", r"-analysis\b",
r"\bexplainer\b", r"\bguide\b", r"\bplan\b",
],
"review": [
r"\breview\b", r"-pr-", r"\.diff(?:\.md)?$", r"code-review",
r"\bpr-writeup\b",
],
"slides": [
r"\bdeck\b", r"\bslides\b", r"-talk\b", r"\bpresentation\b",
r"\bkeynote\b",
],
}
CONTENT_SIGNALS: dict[str, list[tuple[str, str, int]]] = {
"document": [
# (regex, description, weight)
(r"^## Table of Contents", "TOC heading", 2),
(r"^# .{3,}$", "H1 with title", 1),
(r"^## .{3,}$", "H2 with title", 1),
(r"^\| .+\| .+\|$", "markdown table row", 1),
(r"^> \[!NOTE\]|^> \[!TIP\]|^> \[!IMPORTANT\]", "GFM callout", 1),
],
"review": [
(r"^```diff\b", "diff fence", 2),
(r"^[-+]{3} ", "unified-diff file header", 2),
(r"^@@ .* @@", "unified-diff hunk header", 2),
(r"^> \[!BLOCKER\]|^> \[!MAJOR\]|^> \[!MINOR\]|^> \[!NIT\]", "severity callout", 2),
(r"\bLGTM\b|\bnit:|\bblocker:|\bmajor:", "review-vocab inline", 1),
],
"slides": [
(r"^---\s*$", "HR slide boundary", 1),
(r"<!--\s*notes:", "presenter notes", 2),
(r"^# .{3,}$", "H1 (slide title candidate)", 1),
],
}
def _score_filename(path: Path) -> dict[str, int]:
name = path.name.lower()
out: dict[str, int] = {"document": 0, "review": 0, "slides": 0}
for cls, patterns in FILENAME_HINTS.items():
for p in patterns:
if re.search(p, name):
out[cls] += 2
break
return out
def _score_content(text: str) -> tuple[dict[str, int], dict[str, list[str]]]:
scores: dict[str, int] = {"document": 0, "review": 0, "slides": 0}
evidence: dict[str, list[str]] = {"document": [], "review": [], "slides": []}
lines = text.splitlines()
for cls, sigs in CONTENT_SIGNALS.items():
for pattern, label, weight in sigs:
compiled = re.compile(pattern, re.MULTILINE)
matches = compiled.findall(text)
if matches:
hit_count = len(matches)
scores[cls] += weight * min(hit_count, 5) # cap each signal at 5 hits to avoid runaway
evidence[cls].append(f"{label} x{hit_count}")
# slides special-case: HR slide boundary count >= 3 is a stronger signal
hr_count = len(re.findall(r"^---\s*$", text, re.MULTILINE))
if hr_count >= 3:
scores["slides"] += 2
evidence["slides"].append(f"hr boundaries >= 3 (count={hr_count})")
# slides special-case: many H1s with mostly-empty bodies between
h1_indices = [i for i, ln in enumerate(lines) if re.match(r"^# .{3,}$", ln)]
if len(h1_indices) >= 5:
gaps = [h1_indices[i + 1] - h1_indices[i] for i in range(len(h1_indices) - 1)]
if gaps and sum(g <= 12 for g in gaps) / len(gaps) >= 0.6:
scores["slides"] += 2
evidence["slides"].append(
f"H1 cadence: {len(h1_indices)} H1s, median gap ~{sorted(gaps)[len(gaps)//2]} lines"
)
return scores, evidence
def classify(input_path: Path | None, text: str | None) -> dict[str, Any]:
if text is None:
if input_path is None:
raise ValueError("Need either input_path or text")
text = input_path.read_text(encoding="utf-8")
fn_scores = _score_filename(input_path) if input_path else {"document": 0, "review": 0, "slides": 0}
content_scores, evidence = _score_content(text)
total = {k: fn_scores[k] + content_scores[k] for k in fn_scores}
line_count = len(text.splitlines())
below_min = line_count < MIN_LINES
# Ranking
ranked = sorted(total.items(), key=lambda kv: kv[1], reverse=True)
winner_cls, winner_score = ranked[0]
runner_cls, runner_score = ranked[1]
silent_route = (
winner_score >= 3
and (runner_score == 0 or winner_score >= 2 * runner_score)
)
if winner_score == 0:
verdict = "ambiguous"
recommendation = "Ask the user which document type — no signals matched."
elif silent_route:
verdict = winner_cls
recommendation = f"Route to md-{winner_cls} (score {winner_score} vs runner-up {runner_score})."
elif winner_score >= 2:
verdict = "needs-clarification"
recommendation = (
f"Top candidate is md-{winner_cls} (score {winner_score}) "
f"but md-{runner_cls} also scored {runner_score} — ask user to confirm."
)
else:
verdict = "ambiguous"
recommendation = (
f"Weak signal ({winner_cls}={winner_score}). "
f"Ask user, or treat as md-document by default."
)
return {
"verdict": verdict,
"winner": winner_cls,
"winner_score": winner_score,
"runner_up": runner_cls,
"runner_up_score": runner_score,
"filename_scores": fn_scores,
"content_scores": content_scores,
"total_scores": total,
"evidence": evidence,
"line_count": line_count,
"below_min_lines": below_min,
"min_lines_threshold": MIN_LINES,
"recommendation": recommendation,
"silent_route_allowed": silent_route,
}
def render_human(result: dict[str, Any]) -> str:
lines = []
lines.append(f"Doctype classification: {result['verdict']}")
lines.append(f" recommendation: {result['recommendation']}")
lines.append(f" line count: {result['line_count']} (threshold {result['min_lines_threshold']})")
if result["below_min_lines"]:
lines.append(
f" ! below threshold — markdown still wins under "
f"{result['min_lines_threshold']} lines (Shihipar). Recommend keeping as markdown."
)
lines.append("")
lines.append("Scores:")
for cls in ["document", "review", "slides"]:
fn = result["filename_scores"][cls]
ct = result["content_scores"][cls]
total = result["total_scores"][cls]
lines.append(f" md-{cls:<10s} total={total:<3d} (filename={fn}, content={ct})")
lines.append("")
lines.append("Evidence:")
for cls, sigs in result["evidence"].items():
if sigs:
lines.append(f" md-{cls}: {', '.join(sigs)}")
return "\n".join(lines)
def main(argv: list[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--input", help="Path to markdown file, or '-' for stdin")
parser.add_argument("--output", choices=["human", "json"], default="human")
parser.add_argument("--sample", action="store_true",
help="Classify a built-in sample (a Shihipar-style 200-line spec)")
args = parser.parse_args(argv)
if args.sample:
sample_text = SAMPLE_MARKDOWN
result = classify(None, sample_text)
elif args.input:
if args.input == "-":
text = sys.stdin.read()
result = classify(None, text)
else:
path = Path(args.input)
if not path.exists():
print(f"error: input not found: {path}", file=sys.stderr)
return 2
result = classify(path, None)
else:
parser.print_help()
return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
SAMPLE_MARKDOWN = """# Implementation Plan: Payment Gateway Integration
## Table of Contents
- Goals
- Architecture
- Risks
## Goals
We will integrate Stripe Connect with the existing checkout flow.
| Phase | Timeline | Owner |
|---|---|---|
| Design | Week 1 | jane |
| Build | Week 2-3 | dev team |
| Ship | Week 4 | jane |
## Architecture
The integration will use webhooks for async events.
> [!NOTE]
> All webhook handlers must be idempotent.
## Risks
1. Webhook delivery delays
2. Tax calculation edge cases
3. Refund cascading
""" + "\n" * 120 # pad to > MIN_LINES
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/output_path_resolver.py
#!/usr/bin/env python3
"""output_path_resolver.py - Resolve the final output path for a conversion.
Stdlib-only. Given:
- the input markdown filename
- an optional --out user override
- the design-system config's default_output_dir
- a --doctype hint (document/review/slides) for naming convention
Returns the final absolute path the converter should write to. Handles
collisions by suffixing -2, -3, ... or by inserting an ISO-8601 stamp,
depending on --on-collision mode. Refuses if the chosen parent isn't
writable (matches onboard.py's hard rule).
Pattern (kebab slug + collision detection + timestamp fallback) lifted from
marketing/landing/skills/landing/scripts/kebab_slug_generator.py and
adapted: doctype prefix in the filename, --out override, design-system
default_output_dir as the fallback root.
NO LLM CALLS. Pure path math.
Usage:
python output_path_resolver.py --input report.md
python output_path_resolver.py --input report.md --out ./docs/ --doctype document
python output_path_resolver.py --input PR-123.md --doctype review --on-collision timestamp
"""
from __future__ import annotations
import argparse
import datetime as _dt
import json
import os
import re
import sys
from pathlib import Path
from typing import Any
# Bridge to the design-system config
_DESIGN_SYSTEM_SCRIPTS = (
Path(__file__).resolve().parent.parent.parent / "design-system" / "scripts"
)
sys.path.insert(0, str(_DESIGN_SYSTEM_SCRIPTS))
try:
import config_loader as cfg
except ImportError:
cfg = None
DOCTYPE_PREFIXES = {
"document": "doc",
"review": "review",
"slides": "deck",
}
def kebab_slug(name: str) -> str:
"""Convert a filename or title to a clean kebab-case slug.
'My Report v2!.md' -> 'my-report-v2'
' Spaces And Stuff ' -> 'spaces-and-stuff'
"""
base = name.rsplit(".", 1)[0] if "." in name else name
# Strip non-alphanumerics, collapse to hyphen
slug = re.sub(r"[^a-zA-Z0-9]+", "-", base).strip("-").lower()
return slug or "untitled"
def _writable(path: Path) -> bool:
p = path.expanduser()
parent = p.parent if p.suffix else p
while not parent.exists():
if parent.parent == parent:
return False
parent = parent.parent
return os.access(parent, os.W_OK)
def _resolve_base_dir(out_override: str | None) -> Path:
if out_override:
return Path(out_override).expanduser()
if cfg is not None and not os.environ.get("MARKDOWN_HTML_NO_CONFIG") == "1":
config = cfg.load_config()
default_dir = config.get("default_output_dir") or "./markdown-html-out/"
else:
default_dir = "./markdown-html-out/"
return Path(default_dir).expanduser()
def resolve(
input_path: str,
out_override: str | None = None,
doctype: str | None = None,
on_collision: str = "suffix",
) -> dict[str, Any]:
"""Resolve the final output path. Returns a structured dict with the path
and any collision-handling that happened.
"""
in_p = Path(input_path)
slug = kebab_slug(in_p.name)
prefix = DOCTYPE_PREFIXES.get(doctype or "", "")
filename_base = f"{prefix}-{slug}" if prefix else slug
base_dir = _resolve_base_dir(out_override)
base_dir.mkdir(parents=True, exist_ok=True)
target = base_dir / f"{filename_base}.html"
collision_info: dict[str, Any] = {"existed": False, "strategy": None}
if target.exists():
collision_info["existed"] = True
if on_collision == "timestamp":
stamp = _dt.datetime.now().strftime("%Y%m%dT%H%M%S")
target = base_dir / f"{filename_base}-{stamp}.html"
collision_info["strategy"] = "timestamp"
elif on_collision == "overwrite":
collision_info["strategy"] = "overwrite"
# target unchanged
else: # suffix
for n in range(2, 1000):
candidate = base_dir / f"{filename_base}-{n}.html"
if not candidate.exists():
target = candidate
collision_info["strategy"] = f"suffix-{n}"
break
else:
stamp = _dt.datetime.now().strftime("%Y%m%dT%H%M%S")
target = base_dir / f"{filename_base}-{stamp}.html"
collision_info["strategy"] = "timestamp-after-suffix-exhausted"
return {
"input": str(in_p),
"slug": slug,
"doctype": doctype,
"prefix": prefix,
"base_dir": str(base_dir),
"output_path": str(target),
"writable": _writable(target),
"collision": collision_info,
}
def main(argv: list[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--input", help="Input markdown filename or path")
parser.add_argument("--out", help="Override the output directory (else uses config default)")
parser.add_argument("--doctype", choices=["document", "review", "slides"],
help="Doc type — controls filename prefix (doc-, review-, deck-)")
parser.add_argument("--on-collision", choices=["suffix", "timestamp", "overwrite"],
default="suffix")
parser.add_argument("--output", choices=["human", "json"], default="human",
dest="output_format")
parser.add_argument("--sample", action="store_true",
help="Show a resolved-path example without touching disk semantics")
args = parser.parse_args(argv)
if args.sample:
result = resolve("example-report.md", None, "document", "suffix")
elif args.input:
result = resolve(args.input, args.out, args.doctype, args.on_collision)
else:
parser.print_help()
return 0
if not result["writable"]:
print(
f"refusing: target parent '{result['base_dir']}' is not writable. "
f"Re-run onboarding or pass --out to a writable dir.",
file=sys.stderr,
)
return 3
if args.output_format == "json":
print(json.dumps(result, indent=2))
else:
print(f"output -> {result['output_path']}")
if result["collision"]["existed"]:
print(f" (collision handled: {result['collision']['strategy']})")
print(f" base_dir: {result['base_dir']}")
print(f" slug: {result['slug']}, prefix: {result['prefix']}")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/route_explainer.py
#!/usr/bin/env python3
"""route_explainer.py - Print the routing decision in a form the LLM can act on.
Stdlib-only. Takes the JSON output of doctype_classifier.py (or runs the
classifier itself), and prints a short routing brief: which sub-skill to
invoke, what evidence supports the decision, and what to ask the user if
the verdict is ambiguous.
This is the "never silently chain" enforcer — it prints the recommendation
in a structured form that makes it obvious whether the orchestrator should
route silently, ask one clarifying question, or refuse outright (because
the input is below the 100-line threshold or design-system isn't onboarded).
NO LLM CALLS. Pure formatting + decision-tree branching.
Usage:
python doctype_classifier.py --input X.md --output json | python route_explainer.py
python route_explainer.py --classification-file classification.json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from pathlib import Path
from typing import Any
# Bridge to the design-system config so we can refuse if not onboarded
_DESIGN_SYSTEM_SCRIPTS = (
Path(__file__).resolve().parent.parent.parent / "design-system" / "scripts"
)
sys.path.insert(0, str(_DESIGN_SYSTEM_SCRIPTS))
try:
import config_loader as cfg
except ImportError:
cfg = None
def _design_system_status() -> dict[str, Any]:
if cfg is None:
return {"onboarded": False, "reason": "config_loader not importable"}
if os.environ.get("MARKDOWN_HTML_NO_CONFIG") == "1":
return {"onboarded": True, "reason": "bypass env set", "bypass": True}
if cfg.setup_completed():
c = cfg.load_config()
return {
"onboarded": True,
"default_output_dir": c.get("default_output_dir"),
"design_style": c.get("design_style"),
"brand_primary": (c.get("brand") or {}).get("primary"),
"completed_at": c.get("setup_completed_at"),
}
return {"onboarded": False, "reason": "no setup_completed_at in config"}
def explain(classification: dict[str, Any]) -> dict[str, Any]:
verdict = classification["verdict"]
line_count = classification["line_count"]
below_min = classification["below_min_lines"]
ds = _design_system_status()
refusals: list[str] = []
if below_min:
refusals.append(
f"Input is {line_count} lines (< {classification['min_lines_threshold']}). "
f"Per Shihipar's threshold, markdown wins below 100 lines. "
f"Recommend keeping this as markdown and re-running only on longer documents."
)
if not ds.get("onboarded"):
refusals.append(
"Design-system has not been onboarded. Run "
"`python3 markdown-html/skills/design-system/scripts/onboard.py` "
"(or `--defaults`) before conversion, so the converters have brand tokens to apply."
)
next_action = ""
sub_skill = None
if refusals:
next_action = "REFUSE — fix the issues above before routing."
elif verdict in ("document", "review", "slides"):
sub_skill = f"md-{verdict}"
next_action = (
f"ROUTE_SILENTLY -> {sub_skill}. "
f"Evidence: {classification['winner']} won with score "
f"{classification['winner_score']} (runner-up {classification['runner_up']}="
f"{classification['runner_up_score']})."
)
elif verdict == "needs-clarification":
winner = classification["winner"]
runner = classification["runner_up"]
next_action = (
f"ASK_USER one question: 'I see signals for both md-{winner} (score "
f"{classification['winner_score']}) and md-{runner} (score "
f"{classification['runner_up_score']}). Recommended: md-{winner}. "
f"Confirm or override?'"
)
else: # ambiguous
next_action = (
"ASK_USER one question: 'Which document type is this — long-form "
"document, code review with diff, or slide deck? "
"Recommended: md-document (safe default).'"
)
return {
"decision": "REFUSE" if refusals else next_action.split(" ", 1)[0],
"sub_skill": sub_skill,
"next_action": next_action,
"refusals": refusals,
"classification_verdict": verdict,
"line_count": line_count,
"design_system": ds,
}
def render_human(explanation: dict[str, Any]) -> str:
out = []
out.append(f"Routing decision: {explanation['decision']}")
if explanation["sub_skill"]:
out.append(f" sub-skill: {explanation['sub_skill']}")
out.append(f" next action: {explanation['next_action']}")
if explanation["refusals"]:
out.append("")
out.append("Refusals:")
for r in explanation["refusals"]:
out.append(f" - {r}")
out.append("")
out.append("Design-system:")
ds = explanation["design_system"]
out.append(f" onboarded: {ds.get('onboarded')}")
if ds.get("onboarded"):
out.append(f" default_output_dir: {ds.get('default_output_dir')}")
out.append(f" design_style: {ds.get('design_style')}")
out.append(f" brand_primary: {ds.get('brand_primary')}")
else:
out.append(f" reason: {ds.get('reason')}")
return "\n".join(out)
def main(argv: list[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--classification-file",
help="Path to a doctype_classifier JSON output. Default: read stdin.")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.classification_file:
with open(args.classification_file, encoding="utf-8") as f:
classification = json.load(f)
else:
if sys.stdin.isatty():
parser.print_help()
return 0
classification = json.load(sys.stdin)
explanation = explain(classification)
if args.output == "json":
print(json.dumps(explanation, indent=2))
else:
print(render_human(explanation))
return 0 if explanation["decision"] != "REFUSE" else 3
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Gợi ý ý tưởng, chiến lược và chiến thuật marketing, tăng trưởng cho sản phẩm SaaS hoặc phần mềm.
---
name: marketing-ideas
description: "When the user needs marketing ideas, inspiration, or strategies for their SaaS or software product. Also use when the user asks for 'marketing ideas,' 'growth ideas,' 'how to market,' 'marketing strategies,' 'marketing tactics,' 'ways to promote,' 'ideas to grow,' 'what else can I try,' 'I don't know how to market this,' 'brainstorm marketing,' or 'what marketing should I do.' Use this as a starting point whenever someone is stuck or looking for inspiration on how to grow. For specific channel execution, see the relevant skill (ads, social, emails, etc.)."
metadata:
version: 2.0.1
---
# Marketing Ideas for SaaS
You are a marketing strategist with a library of 139 proven marketing ideas. Your goal is to help users find the right marketing strategies for their specific situation, stage, and resources.
## How to Use This Skill
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
When asked for marketing ideas:
1. Ask about their product, audience, and current stage if not clear
2. Suggest 3-5 most relevant ideas based on their context
3. Provide details on implementation for chosen ideas
4. Consider their resources (time, budget, team size)
---
## Ideas by Category (Quick Reference)
| Category | Ideas | Examples |
|----------|-------|----------|
| Content & SEO | 1-10 | Programmatic SEO, Glossary marketing, Content repurposing |
| Competitor | 11-13 | Comparison pages, Marketing jiu-jitsu |
| Free Tools | 14-22 | Calculators, Generators, Chrome extensions |
| Paid Ads | 23-34 | LinkedIn, Google, Retargeting, Podcast ads |
| Social & Community | 35-44 | LinkedIn audience, Reddit marketing, Short-form video |
| Email | 45-53 | Founder emails, Onboarding sequences, Win-back |
| Partnerships | 54-64 | Affiliate programs, Integration marketing, Newsletter swaps |
| Events | 65-72 | Webinars, Conference speaking, Virtual summits |
| PR & Media | 73-76 | Press coverage, Documentaries |
| Launches | 77-86 | Product Hunt, Lifetime deals, Giveaways |
| Product-Led | 87-96 | Viral loops, Powered-by marketing, Free migrations |
| Content Formats | 97-109 | Podcasts, Courses, Annual reports, Year wraps |
| Unconventional | 110-122 | Awards, Challenges, [Guerrilla marketing](references/guerrilla-marketing.md) |
| Platforms | 123-130 | App marketplaces, Review sites, YouTube |
| International | 131-132 | Expansion, Price localization |
| Developer | 133-136 | DevRel, Certifications |
| Audience-Specific | 137-139 | Referrals, Podcast tours, Customer language |
**For the complete list with descriptions**: See [references/ideas-by-category.md](references/ideas-by-category.md)
**Deep dives** (full framework + case library for a single idea):
- **Guerrilla marketing (#121)**: [references/guerrilla-marketing.md](references/guerrilla-marketing.md) — the direct-mail 3-rule framework (relevance / relationship-building / precision targeting), ROI discipline, "think in stories, not campaigns," "test small before going big," and a named case library (WePay, Xero, Red Bull, Antimetal, Arrows, ProfitWell, Buzzsprout, Wistia).
---
## Implementation Tips
### By Stage
**Pre-launch:**
- Waitlist referrals (#79)
- Early access pricing (#81)
- Product Hunt prep (#78)
**Early stage:**
- Content & SEO (#1-10)
- Community (#35)
- Founder-led sales (#47)
**Growth stage:**
- Paid acquisition (#23-34)
- Partnerships (#54-64)
- Events (#65-72)
**Scale:**
- Brand campaigns
- International (#131-132)
- Media acquisitions (#73)
### By Budget
**Free:**
- Content & SEO
- Community building
- Social media
- Comment marketing
**Low budget:**
- Targeted ads
- Sponsorships
- Free tools
**Medium budget:**
- Events
- Partnerships
- PR
**High budget:**
- Acquisitions
- Conferences
- Brand campaigns
### By Timeline
**Quick wins:**
- Ads, email, social posts
**Medium-term:**
- Content, SEO, community
**Long-term:**
- Brand, thought leadership, platform effects
---
## Top Ideas by Use Case
### Need Leads Fast
- Google Ads (#31) - High-intent search
- LinkedIn Ads (#28) - B2B targeting
- Engineering as Marketing (#15) - Free tool lead gen
### Building Authority
- Conference Speaking (#70)
- Book Marketing (#104)
- Podcasts (#107)
### Low Budget Growth
- Easy Keyword Ranking (#1)
- Reddit Marketing (#38)
- Comment Marketing (#44)
### Product-Led Growth
- Viral Loops (#93)
- Powered By Marketing (#87)
- In-App Upsells (#91)
### Enterprise Sales
- Investor Marketing (#133)
- Expert Networks (#57)
- Conference Sponsorship (#72)
---
## Output Format
When recommending ideas, provide for each:
- **Idea name**: One-line description
- **Why it fits**: Connection to their situation
- **How to start**: First 2-3 implementation steps
- **Expected outcome**: What success looks like
- **Resources needed**: Time, budget, skills required
---
## Task-Specific Questions
1. What's your current stage and main growth goal?
2. What's your marketing budget and team size?
3. What have you already tried that worked or didn't?
4. What competitor tactics do you admire?
---
## Related Skills
- **marketing-plan**: When the user wants a comprehensive plan instead of standalone ideas. Section 12 of the plan cross-references all 139 ideas here against AARRR stages and client-specific status.
- **programmatic-seo**: For scaling SEO content (#4)
- **competitors**: For comparison pages (#11)
- **emails**: For email marketing tactics
- **free-tools**: For engineering as marketing (#15)
- **referrals**: For viral growth (#93)
FILE:evals/evals.json
{
"skill_name": "marketing-ideas",
"evals": [
{
"id": 1,
"prompt": "I need marketing ideas for my SaaS product. We're a bootstrapped team of 3, sell a $49/month analytics tool for e-commerce, and have about 200 customers. Budget is tight — maybe $500/month for marketing.",
"expected_output": "Should check for product-marketing.md first. Should filter ideas by low budget and early-stage constraints. Should pull relevant ideas from the 139 marketing ideas organized by category. Should provide ideas appropriate for bootstrapped SaaS: content marketing, community building, SEO, partnerships, referral programs, social media, Product Hunt, and others that don't require large budgets. Output should follow the format: idea name, why it fits, how to start, expected outcome, resources needed. Should prioritize by likely impact given their stage.",
"assertions": [
"Checks for product-marketing.md",
"Filters ideas by low budget constraint",
"Provides ideas from the 139 marketing ideas catalog",
"Ideas are appropriate for bootstrapped SaaS stage",
"Output follows structured format per idea",
"Includes why it fits, how to start, expected outcome",
"Prioritizes by likely impact",
"Includes resources needed per idea"
],
"files": []
},
{
"id": 2,
"prompt": "What's the fastest way to get more leads? We sell enterprise security software and have a $10k/month marketing budget.",
"expected_output": "Should apply the 'top ideas by use case' — specifically 'leads fast' recommendations. Should recommend paid channels (Google Ads, LinkedIn Ads for enterprise), outbound (cold email, LinkedIn outreach), and content-based lead magnets. Should filter for enterprise-appropriate tactics. Should provide the structured output with why each idea fits, how to start, expected timeline, and resources needed. Should note that 'fast leads' typically means paid or outbound channels.",
"assertions": [
"Applies 'leads fast' use case filter",
"Recommends paid channels appropriate for enterprise",
"Recommends outbound tactics",
"Filters for enterprise-appropriate tactics",
"Provides structured output per idea",
"Notes that fast leads means paid or outbound",
"Includes timeline expectations"
],
"files": []
},
{
"id": 3,
"prompt": "how do we grow without spending money on ads? we're a PLG (product-led growth) company with a freemium model",
"expected_output": "Should trigger on casual phrasing. Should apply the 'PLG' use case filter from top ideas. Should recommend PLG-specific tactics: product virality features, referral programs, community building, content marketing, SEO, free tools/calculators, open-source contributions, social proof loops. Should avoid ad-dependent ideas given the constraint. Should provide structured output with implementation guidance.",
"assertions": [
"Triggers on casual phrasing",
"Applies PLG use case filter",
"Recommends PLG-specific tactics",
"Avoids ad-dependent ideas",
"Includes virality and referral tactics",
"Provides structured output with implementation guidance"
],
"files": []
},
{
"id": 4,
"prompt": "We want to build authority and thought leadership in the HR tech space. We're a newer company and nobody knows who we are yet.",
"expected_output": "Should apply the 'authority building' use case filter. Should recommend thought leadership tactics: original research/surveys, guest posting, podcast appearances, speaking engagements, LinkedIn content, industry report publishing, expert roundups. Should note that authority building is a longer-term play. Should provide structured output with how to start each idea, expected outcomes, and timeline.",
"assertions": [
"Applies 'authority building' use case filter",
"Recommends thought leadership tactics",
"Includes original research and content",
"Includes community and media appearances",
"Notes authority building is longer-term",
"Provides structured output with timelines"
],
"files": []
},
{
"id": 5,
"prompt": "Give me 20 marketing ideas. We sell project management software.",
"expected_output": "Should provide a curated list of ~20 ideas from the catalog. Should organize them by category or by effort/impact. Should provide brief implementation context for each. Should vary the ideas across categories (content, community, partnerships, product, paid, etc.) for a well-rounded set. Output should follow the structured format with at least idea name and brief description for each.",
"assertions": [
"Provides approximately 20 ideas",
"Ideas span multiple categories",
"Organizes by category or effort/impact",
"Provides brief implementation context per idea",
"Output follows structured format",
"Ideas are relevant to project management software"
],
"files": []
},
{
"id": 6,
"prompt": "We want to set up a referral program. How should we structure it?",
"expected_output": "Should recognize this is specifically a referral program design request. Should defer to or cross-reference the referrals skill, which provides detailed guidance on referral loop design, incentive structures, implementation, and optimization. May briefly mention referral programs as a marketing idea but should make clear that referrals is the right skill for detailed program design.",
"assertions": [
"Recognizes this as a referral program design request",
"References or defers to referrals skill",
"Does not attempt detailed referral program design",
"May briefly mention as a marketing idea"
],
"files": []
},
{
"id": 7,
"prompt": "We sell a $30k/year data platform and can't get marketing VPs at target accounts to reply to email or take calls. We have budget for something creative. What guerrilla tactics could break through?",
"expected_output": "Should route to the guerrilla-marketing deep-dive reference. Should surface the direct-mail 3-rule framework (relevance, relationship-building, precision targeting) and the ROI discipline that a $50-200 package to an unqualified lead is malpractice — qualify the named accounts first, test small before scaling. Should note the deal size ($30k/year) justifies high-cost mailers. Should draw on the case library for inspiration (e.g. Corey's Unicorn Floatie Campaign for hard-to-reach marketing leaders, Antimetal's $15k pizza campaign converting ~75 customers and ~$1M revenue). Should apply the operating principles: think in stories not campaigns, and test small before going big.",
"assertions": [
"Routes to the guerrilla-marketing reference",
"Surfaces the direct-mail 3-rule framework (relevance, relationship-building, precision targeting)",
"Applies the ROI discipline (qualify before spending on high-cost packages)",
"Recommends testing small before scaling",
"Notes deal size justifies high-cost direct mail",
"Pulls named cases as inspiration (e.g. Unicorn Floatie, Antimetal pizza)",
"Applies 'think in stories, not campaigns' principle"
],
"files": []
}
]
}
FILE:references/guerrilla-marketing.md
# Guerrilla Marketing (#121)
Unconventional stunts that earn attention in unexpected places — then get turned into repeatable strategy. The point isn't the one-off spectacle. It's finding a memorable, low-cost way to break through, proving it works small, and running it again.
Two operating principles govern everything below:
- **Think in stories, not campaigns.** People don't share your funnel. They share a story worth retelling. Design the moment so the person who receives it wants to tell someone else.
- **Test small before going big.** A guerrilla idea is a hypothesis. Prove it on a handful of targets before you spend real money scaling it.
---
## Direct Mail: The 3-Rule Framework
Physical mail is the most under-used guerrilla channel in SaaS because most people do it lazily — blasting cheap swag at a bought list. Done right, it cuts through inbox noise precisely because nobody else shows up in the mailbox. Every direct-mail play must satisfy all three rules:
1. **Relevance** — The item connects to your product, your message, or something the recipient specifically cares about. A random branded mug is noise. A gift that lands a joke about their exact problem is a story.
2. **Relationship-building** — The mailer opens or deepens a relationship. It's a gift, not a pitch. If the recipient feels sold to, you've spent money to annoy someone.
3. **Precision targeting** — You know exactly who is receiving it and why. Direct mail is expensive per unit, so it only works when aimed at a short, hand-qualified list.
### The ROI discipline
A $50–200 package sent to an unqualified lead is malpractice. The math only works when the target is qualified enough that a single closed deal pays for dozens of packages.
- **Qualify before you spend.** Reserve high-cost mailers for named accounts and hard-to-reach decision-makers where the deal size justifies the cost.
- **Test small before scaling.** Send to a handful first. Measure reply and meeting rates. Only scale the version that actually earns responses.
### Corey's Unicorn Floatie Campaign
To reach marketing leaders who ignore cold email and screen their calls, Corey mailed **giant inflatable unicorn pool floaties** to a short list of hard-to-reach targets. The floatie is absurd, oversized, and impossible to ignore — it lands as a story, not a pitch (relevance + relationship), aimed at a hand-picked list where each closed deal dwarfs the package cost (precision + ROI discipline). Recipients remembered it, mentioned it, and replied.
---
## Case Library (inspiration fuel)
Steal the *pattern*, not the prop. Each of these turned a small, unconventional act into outsized attention or revenue.
- **WePay's 600-lb ice block** — At PayPal's developer conference, WePay dropped a 600-pound block of ice with money frozen inside and a sign reading **"PayPal Freezes Your Accounts."** A pointed, physical jab at a competitor's real weakness, staged exactly where the audience was.
- **Xero skywriting** — Xero paid for **skywriting over TechCrunch Disrupt**, hijacking a competitor-heavy event's attention from above without buying a booth.
- **Red Bull's Felix Baumgartner space jump** — A supersonic freefall from the edge of space drew **8 million concurrent live viewers** — a story so big it dwarfed any ad buy.
- **Antimetal's $15K pizza campaign** — Antimetal spent roughly **$15,000 sending ~1,000 pizzas** to target accounts, converting **~75 customers** and driving **~$1M in revenue**. Precision-targeted, story-worthy, and measured — the direct-mail framework at scale.
- **Arrows personalized memos** — Personalized, hand-crafted memos sent to specific prospects, treating each as an audience of one.
- **ProfitWell trading cards** — Custom trading cards that turned the team and community into collectible characters, making the brand fun to share.
- **Buzzsprout GIPHY library** — A branded library of GIFs on GIPHY, so the brand rode along inside other people's conversations for free.
- **Wistia "Gear Squad vs Dr. Boring"** — Wistia produced an original, over-the-top branded film — entertainment first, marketing second — because a story people actually want to watch travels further than an ad.
---
## How to Use This With a Client
1. **Pick the story.** What's the one memorable thing a target would retell? Start from the story, then reverse-engineer the tactic.
2. **Qualify the list.** For any high-cost play (direct mail especially), name the exact accounts and confirm the deal size justifies the spend.
3. **Run the 3-rule check** on physical mail: relevance, relationship-building, precision targeting. If it fails any one, redesign it.
4. **Test small.** Ship to a handful, measure replies/meetings/coverage, and only scale what earns a response.
5. **Turn the stunt into a system.** If it works, make it repeatable — a recurring mailer program, a series of films, an ongoing library — rather than a one-time spike.
*Source: Corey Haines, Founding Marketing, ch. 13.*
FILE:references/ideas-by-category.md
# The 139 Marketing Ideas
Complete list of proven marketing approaches organized by category.
## Contents
- Content & SEO (1-10)
- Competitor & Comparison (11-13)
- Free Tools & Engineering (14-22)
- Paid Advertising (23-34)
- Social Media & Community (35-44)
- Email Marketing (45-53)
- Partnerships & Programs (54-64)
- Events & Speaking (65-72)
- PR & Media (73-76)
- Launches & Promotions (77-86)
- Product-Led Growth (87-96)
- Content Formats (97-109)
- Unconventional & Creative (110-122)
- Platforms & Marketplaces (123-130)
- International & Localization (131-132)
- Developer & Technical (133-136)
- Audience-Specific (137-139)
## Content & SEO (1-10)
1. **Easy Keyword Ranking** - Target low-competition keywords where you can rank quickly. Find terms competitors overlook—niche variations, long-tail queries, emerging topics.
2. **SEO Audit** - Conduct comprehensive technical SEO audits of your own site and share findings publicly. Document fixes and improvements to build authority.
3. **Glossary Marketing** - Create comprehensive glossaries defining industry terms. Each term becomes an SEO-optimized page targeting "what is X" searches.
4. **Programmatic SEO** - Build template-driven pages at scale targeting keyword patterns. Location pages, comparison pages, integration pages—any pattern with search volume.
5. **Content Repurposing** - Transform one piece of content into multiple formats. Blog post becomes Twitter thread, YouTube video, podcast episode, infographic.
6. **Proprietary Data Content** - Leverage unique data from your product to create original research and reports. Data competitors can't replicate creates linkable assets.
7. **Internal Linking** - Strategic internal linking distributes authority and improves crawlability. Build topical clusters connecting related content.
8. **Content Refreshing** - Regularly update existing content with fresh data, examples, and insights. Refreshed content often outperforms new content.
9. **Knowledge Base SEO** - Optimize help documentation for search. Support articles targeting problem-solution queries capture users actively seeking solutions.
10. **Parasite SEO** - Publish content on high-authority platforms (Medium, LinkedIn, Substack) that rank faster than your own domain.
---
## Competitor & Comparison (11-13)
11. **Competitor Comparison Pages** - Create detailed comparison pages positioning your product against competitors. "[Your Product] vs [Competitor]" pages capture high-intent searchers.
12. **Marketing Jiu-Jitsu** - Turn competitor weaknesses into your strengths. When competitors raise prices, launch affordability campaigns.
13. **Competitive Ad Research** - Study competitor advertising through tools like SpyFu or Facebook Ad Library. Learn what messaging resonates.
---
## Free Tools & Engineering (14-22)
14. **Side Projects as Marketing** - Build small, useful tools related to your main product. Side projects attract users who may later convert.
15. **Engineering as Marketing** - Build free tools that solve real problems. Calculators, analyzers, generators—useful utilities that naturally lead to your paid product.
16. **Importers as Marketing** - Build import tools for competitor data. "Import from [Competitor]" reduces switching friction.
17. **Quiz Marketing** - Create interactive quizzes that engage users while qualifying leads. Personality quizzes, assessments, and diagnostic tools generate shares.
18. **Calculator Marketing** - Build calculators solving real problems—ROI calculators, pricing estimators, savings tools. Calculators attract links and rank well.
19. **Chrome Extensions** - Create browser extensions providing standalone value. Chrome Web Store becomes another distribution channel.
20. **Microsites** - Build focused microsites for specific campaigns, products, or audiences. Dedicated domains can rank faster.
21. **Scanners** - Build free scanning tools that audit or analyze something. Website scanners, security checkers, performance analyzers.
22. **Public APIs** - Open APIs enable developers to build on your platform, creating an ecosystem.
---
## Paid Advertising (23-34)
23. **Podcast Advertising** - Sponsor relevant podcasts to reach engaged audiences. Host-read ads perform especially well.
24. **Pre-targeting Ads** - Show awareness ads before launching direct response campaigns. Warm audiences convert better.
25. **Facebook Ads** - Meta's detailed targeting reaches specific audiences. Test creative variations and leverage retargeting.
26. **Instagram Ads** - Visual-first advertising for products with strong imagery. Stories and Reels ads capture attention.
27. **Twitter Ads** - Reach engaged professionals discussing industry topics. Promoted tweets and follower campaigns.
28. **LinkedIn Ads** - Target by job title, company size, and industry. Premium CPMs justified by B2B purchase intent.
29. **Reddit Ads** - Reach passionate communities with authentic messaging. Transparency wins on Reddit.
30. **Quora Ads** - Target users actively asking questions your product answers. Intent-rich environment.
31. **Google Ads** - Capture high-intent search queries. Brand terms, competitor terms, and category terms.
32. **YouTube Ads** - Video ads with detailed targeting. Pre-roll and discovery ads reach users consuming related content.
33. **Cross-Platform Retargeting** - Follow users across platforms with consistent messaging.
34. **Click-to-Messenger Ads** - Ads that open direct conversations rather than landing pages.
---
## Social Media & Community (35-44)
35. **Community Marketing** - Build and nurture communities around your product. Slack groups, Discord servers, Facebook groups.
36. **Quora Marketing** - Answer relevant questions with genuine expertise. Include product mentions where naturally appropriate.
37. **Reddit Keyword Research** - Mine Reddit for real language your audience uses. Discover pain points and desires.
38. **Reddit Marketing** - Participate authentically in relevant subreddits. Provide value first.
39. **LinkedIn Audience** - Build personal brands on LinkedIn for B2B reach. Thought leadership builds authority.
40. **Instagram Audience** - Visual storytelling for products with strong aesthetics. Behind-the-scenes and user stories.
41. **X Audience** - Build presence on X/Twitter through consistent value. Threads and insights grow followings.
42. **Short Form Video** - TikTok, Reels, and Shorts reach new audiences with snackable content.
43. **Engagement Pods** - Coordinate with peers to boost each other's content engagement.
44. **Comment Marketing** - Thoughtful comments on relevant content build visibility.
---
## Email Marketing (45-53)
45. **Mistake Email Marketing** - Send "oops" emails when something genuinely goes wrong. Authenticity generates engagement.
46. **Reactivation Emails** - Win back churned or inactive users with targeted campaigns.
47. **Founder Welcome Email** - Personal welcome emails from founders create connection.
48. **Dynamic Email Capture** - Smart email capture that adapts to user behavior. Exit intent, scroll depth triggers.
49. **Monthly Newsletters** - Consistent newsletters keep your brand top-of-mind.
50. **Inbox Placement** - Technical email optimization for deliverability. Authentication and list hygiene.
51. **Onboarding Emails** - Guide new users to activation with targeted sequences.
52. **Win-back Emails** - Re-engage churned users with compelling reasons to return.
53. **Trial Reactivation** - Expired trials aren't lost causes. Targeted campaigns can recover them.
---
## Partnerships & Programs (54-64)
54. **Affiliate Discovery Through Backlinks** - Find potential affiliates by analyzing who links to competitors.
55. **Influencer Whitelisting** - Run ads through influencer accounts for authentic reach.
56. **Reseller Programs** - Enable agencies to resell your product. White-label options create distribution partners.
57. **Expert Networks** - Build networks of certified experts who implement your product.
58. **Newsletter Swaps** - Exchange promotional mentions with complementary newsletters.
59. **Article Quotes** - Contribute expert quotes to journalists. HARO connects experts with writers.
60. **Pixel Sharing** - Partner with complementary companies to share remarketing audiences.
61. **Shared Slack Channels** - Create shared channels with partners and customers.
62. **Affiliate Program** - Structured commission programs for referrers.
63. **Integration Marketing** - Joint marketing with integration partners.
64. **Community Sponsorship** - Sponsor relevant communities, newsletters, or publications.
---
## Events & Speaking (65-72)
65. **Live Webinars** - Educational webinars demonstrate expertise while generating leads.
66. **Virtual Summits** - Multi-speaker online events attract audiences through varied perspectives.
67. **Roadshows** - Take your product on the road to meet customers directly.
68. **Local Meetups** - Host or attend local meetups in key markets.
69. **Meetup Sponsorship** - Sponsor relevant meetups to reach engaged local audiences.
70. **Conference Speaking** - Speak at industry conferences to reach engaged audiences.
71. **Conferences** - Host your own conference to become the center of your industry.
72. **Conference Sponsorship** - Sponsor relevant conferences for brand visibility.
---
## PR & Media (73-76)
73. **Media Acquisitions as Marketing** - Acquire newsletters, podcasts, or publications in your space.
74. **Press Coverage** - Pitch newsworthy stories to relevant publications.
75. **Fundraising PR** - Leverage funding announcements for press coverage.
76. **Documentaries** - Create documentary content exploring your industry or customers.
---
## Launches & Promotions (77-86)
77. **Black Friday Promotions** - Annual deals create urgency and acquisition spikes.
78. **Product Hunt Launch** - Structured Product Hunt launches reach early adopters.
79. **Early-Access Referrals** - Reward referrals with earlier access during launches.
80. **New Year Promotions** - New Year brings fresh budgets and goal-setting energy.
81. **Early Access Pricing** - Launch with discounted early access tiers.
82. **Product Hunt Alternatives** - Launch on BetaList, Launching Next, AlternativeTo.
83. **Twitter Giveaways** - Engagement-boosting giveaways that require follows or retweets.
84. **Giveaways** - Strategic giveaways attract attention and capture leads.
85. **Vacation Giveaways** - Grand prize giveaways generate massive engagement.
86. **Lifetime Deals** - One-time payment deals generate cash and users.
---
## Product-Led Growth (87-96)
87. **Powered By Marketing** - "Powered by [Your Product]" badges create free impressions.
88. **Free Migrations** - Offer free migration services from competitors.
89. **Contract Buyouts** - Pay to exit competitor contracts.
90. **One-Click Registration** - Minimize signup friction with OAuth options.
91. **In-App Upsells** - Strategic upgrade prompts within the product experience.
92. **Newsletter Referrals** - Built-in referral programs for newsletters.
93. **Viral Loops** - Product mechanics that naturally encourage sharing.
94. **Offboarding Flows** - Optimize cancellation flows to retain or learn.
95. **Concierge Setup** - White-glove onboarding for high-value accounts.
96. **Onboarding Optimization** - Continuous improvement of new user experience.
---
## Content Formats (97-109)
97. **Playlists as Marketing** - Create Spotify playlists for your audience.
98. **Template Marketing** - Offer free templates users can immediately use.
99. **Graphic Novel Marketing** - Transform complex stories into visual narratives.
100. **Promo Videos** - High-quality promotional videos showcase your product.
101. **Industry Interviews** - Interview customers, experts, and thought leaders.
102. **Social Screenshots** - Design shareable screenshot templates for social proof.
103. **Online Courses** - Educational courses establish authority while generating leads.
104. **Book Marketing** - Author a book establishing expertise in your domain.
105. **Annual Reports** - Publish annual reports showcasing industry data and trends.
106. **End of Year Wraps** - Personalized year-end summaries users want to share.
107. **Podcasts** - Launch a podcast reaching audiences during commutes.
108. **Changelogs** - Public changelogs showcase product momentum.
109. **Public Demos** - Live product demonstrations showing real usage.
---
## Unconventional & Creative (110-122)
110. **Awards as Marketing** - Create industry awards positioning your brand as tastemaker.
111. **Challenges as Marketing** - Launch viral challenges that spread organically.
112. **Reality TV Marketing** - Create reality-show style content following real customers.
113. **Controversy as Marketing** - Strategic positioning against industry norms.
114. **Moneyball Marketing** - Data-driven marketing finding undervalued channels.
115. **Curation as Marketing** - Curate valuable resources for your audience.
116. **Grants as Marketing** - Offer grants to customers or community members.
117. **Product Competitions** - Sponsor competitions using your product.
118. **Cameo Marketing** - Use Cameo celebrities for personalized messages.
119. **OOH Advertising** - Out-of-home advertising—billboards, transit ads.
120. **Marketing Stunts** - Bold, attention-grabbing marketing moments.
121. **Guerrilla Marketing** - Unconventional, low-cost marketing in unexpected places. Deep dive: direct-mail 3-rule framework (relevance / relationship / precision), ROI discipline, and a named case library in [guerrilla-marketing.md](guerrilla-marketing.md).
122. **Humor Marketing** - Use humor to stand out and create memorability.
---
## Platforms & Marketplaces (123-130)
123. **Open Source as Marketing** - Open-source components or tools build developer goodwill.
124. **App Store Optimization** - Optimize app store listings for discoverability.
125. **App Marketplaces** - List in Salesforce AppExchange, Shopify App Store, etc.
126. **YouTube Reviews** - Get YouTubers to review your product.
127. **YouTube Channel** - Build a YouTube presence with tutorials and thought leadership.
128. **Source Platforms** - Submit to G2, Capterra, GetApp, and similar directories.
129. **Review Sites** - Actively manage presence on review platforms.
130. **Live Audio** - Host Twitter Spaces, Clubhouse, or LinkedIn Audio discussions.
---
## International & Localization (131-132)
131. **International Expansion** - Expand to new geographic markets with localization.
132. **Price Localization** - Adjust pricing for local purchasing power.
---
## Developer & Technical (133-136)
133. **Investor Marketing** - Market to investors for portfolio introductions.
134. **Certifications** - Create certification programs validating expertise.
135. **Support as Marketing** - Exceptional support creates stories customers share.
136. **Developer Relations** - Build relationships with developer communities.
---
## Audience-Specific (137-139)
137. **Two-Sided Referrals** - Reward both referrer and referred.
138. **Podcast Tours** - Guest on multiple podcasts reaching your target audience.
139. **Customer Language** - Use the exact words your customers use in marketing.
Khởi chạy N subagent song song trong các git worktree cô lập để cạnh tranh giải cùng một nhiệm vụ của phiên.
---
name: "spawn"
description: "Launch N parallel subagents in isolated git worktrees to compete on the session task."
command: /hub:spawn
---
# /hub:spawn — Launch Parallel Agents
Spawn N subagents that work on the same task in parallel, each in an isolated git worktree.
## Usage
```
/hub:spawn # Spawn agents for the latest session
/hub:spawn 20260317-143022 # Spawn agents for a specific session
/hub:spawn --template optimizer # Use optimizer template for dispatch prompts
/hub:spawn --template refactorer # Use refactorer template
```
## Templates
When `--template <name>` is provided, use the dispatch prompt from `references/agent-templates.md` instead of the default prompt below. Available templates:
| Template | Pattern | Use Case |
|----------|---------|----------|
| `optimizer` | Edit → eval → keep/discard → repeat x10 | Performance, latency, size reduction |
| `refactorer` | Restructure → test → iterate until green | Code quality, tech debt |
| `test-writer` | Write tests → measure coverage → repeat | Test coverage gaps |
| `bug-fixer` | Reproduce → diagnose → fix → verify | Bug fix with competing approaches |
When using a template, replace all `{variables}` with values from the session config. Assign each agent a **different strategy** appropriate to the template and task — diverse strategies maximize the value of parallel exploration.
## What It Does
1. Load session config from `.agenthub/sessions/{session-id}/config.yaml`
2. For each agent 1..N:
- Write task assignment to `.agenthub/board/dispatch/`
- Build agent prompt with task, constraints, and board write instructions
3. Launch ALL agents in a **single message** with multiple Agent tool calls:
```
Agent(
prompt: "You are agent-{i} in hub session {session-id}.
Your task: {task}
Read your full assignment at .agenthub/board/dispatch/{seq}-agent-{i}.md
Instructions:
1. Work in your worktree — make changes, run tests, iterate
2. Commit all changes with descriptive messages
3. Write your result summary to .agenthub/board/results/agent-{i}-result.md
Include: approach taken, files changed, metric if available, confidence level
4. Exit when done
Constraints:
- Do NOT read or modify other agents' work
- Do NOT access .agenthub/board/results/ for other agents
- Commit early and often with descriptive messages
- If you hit a dead end, commit what you have and explain in your result",
isolation: "worktree"
)
```
4. Update session state to `running` via:
```bash
python {skill_path}/scripts/session_manager.py --update {session-id} --state running
```
## Critical Rules
- **All agents in ONE message** — spawn all Agent tool calls simultaneously for true parallelism
- **isolation: "worktree"** is mandatory — each agent needs its own filesystem
- **Never modify session config** after spawn — agents rely on stable configuration
- **Each agent gets a unique board post** — dispatch posts are numbered sequentially
## After Spawn
Tell the user:
- {N} agents launched in parallel
- Each working in an isolated worktree
- Monitor with `/hub:status`
- Evaluate when done with `/hub:eval`
Thiết lập quy trình marketing lặp lại tự chạy do AI agent thực hiện theo chu kỳ hoặc sự kiện kích hoạt, thay vì làm một lần.
---
name: marketing-loops
description: "When the user wants to set up a recurring, self-running marketing workflow — a repeatable loop an AI agent runs on a cadence (weekly, daily, on a trigger) rather than a one-off task. Also use when the user mentions 'marketing loop,' 'recurring marketing workflow,' 'automate my marketing,' 'marketing on autopilot,' 'weekly marketing review,' 'ad fatigue check,' 'content refresh loop,' 'churn watch,' 'ranking drop alert,' 'always-on marketing,' 'marketing automation workflow,' or 'run this every week.' Use this to pick, adapt, and schedule an ongoing marketing loop that orchestrates the other marketing skills. For one-off marketing ideas, see marketing-ideas. For the experimentation loop specifically, see ab-testing."
metadata:
version: 1.2.0
---
# Marketing Loops
You help set up **marketing loops** — repeatable marketing workflows an AI agent runs on a cadence, each with a defined trigger, a bounded set of steps, a self-check, and an explicit stopping condition. A loop turns a marketing task you'd otherwise do manually (and forget) into an always-on system: the weekly SEO opportunity scan, the ad-fatigue refresh, the churn-signal watch.
This is the operational cousin of `marketing-ideas`. Ideas tell you *what to try once*. Loops tell you *what to keep doing on a schedule* — and wire the other marketing skills together to do it.
## How to Use This Skill
**Check for product marketing context first:** if `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md`), read it before asking questions. Use that context and only ask for what's missing.
Then:
1. **Clarify the job.** What outcome should this loop protect or grow? (rankings, ad efficiency, activation, retention, revenue, referrals)
2. **Pick a loop** from the catalog in `references/loop-catalog.md` — or adapt the closest one.
3. **Tune the cadence** to how fast the underlying signal actually changes (see the cadence rule below).
4. **Confirm the human checkpoint.** Decide what the loop does autonomously vs. what it stages for human approval before publishing or spending — see `references/loop-guardrails.md`.
5. **Schedule it** (see "Scheduling a loop" below).
Building more than one loop, or a whole marketing operating system? See `references/loop-orchestration.md` for how loops compose and the order to adopt them (start with tracking + a weekly review; don't build 43 at once).
## Anatomy of a Marketing Loop
Every loop in the catalog has these nine parts. When you author or adapt one, fill all of them — a loop missing a stop condition, a self-check, or its state handling is a liability, not an asset.
| Part | What it defines |
|------|-----------------|
| **Check cadence** | How often the loop *looks* (weekly / daily / on-trigger). Match it to signal speed. |
| **Acts when** | The action condition — what must be true to actually *do* something, vs. just check and skip. Most runs of a good loop are "checked, nothing to do." |
| **Purpose** | The one outcome this loop exists to move. |
| **Skills used** | Which marketing skills the loop orchestrates each iteration. |
| **Loop body** | The ordered steps run each iteration. |
| **Self-check** | The verification done *before* acting — so the loop doesn't act on noise, seasonality, or a tracking bug. |
| **State / idempotency** | What the loop remembers between runs: last-run marker, dedupe key, cooldown window, "already handled" set. Without this, loops double-act, re-nag the same people, or re-alert the same thing. Non-negotiable for anything scheduled — see `references/loop-state.md` for where state lives and the idempotency patterns. |
| **Stop / bail-out** | When the loop skips, halts, escalates to a human, or disables itself — plus what it does on error. Every loop needs one, including heartbeat loops (their stop is "manual disable + error-halt," never "n/a"). |
| **Output** | Where results go: a file, a PR, a staged draft, a notification, a report. |
The **Check cadence / Acts when** split matters: a churn-signal loop might *check* daily but only *act* when an account crosses a risk threshold it hasn't been contacted about inside the cooldown window. Conflating the two produces loops that either miss the window or spam.
## The cadence rule
Match cadence to how fast the signal actually changes — not to how often you'd *like* an update.
| Signal | Realistic cadence | Why |
|--------|-------------------|-----|
| Rankings, backlinks, domain authority | Weekly | Move slowly; daily checks are noise |
| Ad creative fatigue, CPA drift | Every 2–3 days | Meta/Google feedback loops are days, not hours |
| Activation / onboarding funnel | Weekly | Needs enough signups to be significant |
| Churn signals | Daily or on-trigger | Early intervention window is short |
| Content / copy decay | Monthly | Traffic erosion is gradual |
| Competitor changes | Weekly | Pricing/positioning shifts are infrequent but matter |
| Social listening / mentions | Daily | Engagement windows close fast |
Over-frequent loops are the most common failure mode: they generate busywork, burn budget, and train you to ignore the output.
## When NOT to loop
Not everything should be automated on a cadence. Skip a loop — or add a mandatory human checkpoint — when:
- **Strategy or creative direction is the real work.** Loops maintain and optimize; they don't set positioning, invent campaigns, or make brand calls.
- **The action publishes or spends without review.** Auto-*drafting* an ad, email, or post is fine. Auto-*publishing* or auto-*shifting budget* needs a human checkpoint unless the user has explicitly authorized autonomous action and set guardrails (caps, allowlists).
- **The signal is too sparse to be significant.** A weekly conversion-rate loop on 40 visitors/week is measuring noise.
- **It's a vanity loop.** If nobody acts on the output, delete the loop. A loop that emails a dashboard nobody reads is worse than nothing.
For any loop that sends, spends, publishes, or touches personal data, apply `references/loop-guardrails.md` — the two-tier action model (autonomous-safe vs. gated), spend/send caps, CAN-SPAM/GDPR/FTC/ToS rules, the always-escalate list, and a required kill switch.
## Scheduling a loop
These loops are agent-agnostic — the *body* works in any agent. The *scheduling* depends on your environment:
- **Claude Code** — native options: `/loop` (self-paced, until a condition), `ScheduleWakeup` (dynamic pacing that reacts to state), and `CronCreate` (fixed cron schedule). If you have a loop-mechanics skill such as `loopify` installed, use it to choose between them and tune delays; otherwise the guidance below is enough.
- **Any agent + cron** — wrap the loop body as a scheduled prompt/script (`0 9 * * 1` for Mondays 9am, etc.).
- **Manual cadence** — for high-judgment loops, "run this skill every Monday" is a perfectly good loop. The value is the repeatable *body*, not the automation.
Default to time-of-day cron for review-style loops (weekly review, ranking watch) and dynamic pacing for monitor-until-threshold loops (churn watch, launch-day tracking).
## The Catalog
`references/loop-catalog.md` holds the full library — 43 marketing loops with thorough funnel coverage: SEO & Content, Paid, Earned/Social/Partnerships, Activation, Retention, Revenue, Referral & Advocacy, and Ongoing Ops. Each is a complete, adaptable spec. Start there, pick the closest match, and tune it to the user's product, stage, and tooling.
## Authoring a new loop
When nothing in the catalog fits, author a new loop from `references/loop-template.md` — a copy-paste template with fill-in prompts, a worked before/after example, and a ship checklist. Fill all nine anatomy parts; if you can't answer the self-check, state/idempotency, and stop/bail-out concretely, the loop isn't ready to run.
## Anti-patterns
- Looping without a stop condition → runaway spend or infinite churn.
- Same cadence for every loop → most run too often and get ignored.
- No self-check → the loop acts on noise, seasonality, or a tracking bug.
- No human checkpoint on spend/publish actions.
- Building 10 loops at once → start with one, prove it earns its keep, then add the next.
## Banned vocabulary
Avoid: "set it and forget it," "fully autonomous marketing," "AI does everything," "10x on autopilot," "growth hacking machine." Loops are disciplined systems with checkpoints, not magic. Describe them honestly.
## Related Skills
- **marketing-ideas** — one-off tactics and inspiration (what to try). Loops operationalize the ones worth repeating.
- **ab-testing** — the experimentation loop specifically (hypothesis → test → promote winner → repeat).
- **analytics** — most loops read from analytics to decide whether to act.
- Individual channel skills (`ads`, `seo-audit`, `emails`, `social`, `churn-prevention`, `pricing`, `referrals`) — the loop bodies orchestrate these.
FILE:evals/evals.json
{
"skill_name": "marketing-loops",
"evals": [
{
"id": 1,
"prompt": "I want to set up a recurring loop that watches our SEO and tells me what to do each week. We're a B2B SaaS with a content-heavy site.",
"expected_output": "Should check for product-marketing.md first. Should recognize this maps to SEO loops in the catalog (keyword-gap, ranking-drop watch, content-decay) and propose the closest match(es). Should present the loop using the nine-part anatomy: check cadence, acts-when (action condition), purpose, skills used, loop body, self-check, state/idempotency, stop/bail-out, and output. Should set a weekly cadence (matching how slowly rankings move) and explain the cadence rule. Should include a stop/bail-out and note that most runs may find no action. Should reference the seo-audit skill for execution.",
"assertions": [
"Checks for product-marketing.md",
"Selects appropriate SEO loop(s) from the catalog",
"Presents the loop with the nine-part anatomy",
"Includes state/idempotency (last-run marker or dedupe)",
"Includes an explicit stop/bail-out condition",
"Sets a weekly cadence and justifies it via the cadence rule",
"References seo-audit or related skills for execution"
],
"files": []
},
{
"id": 2,
"prompt": "I want to automate my marketing so it basically runs itself. Where do I start?",
"expected_output": "Should clarify the outcome to protect or grow before dumping loops (acquisition, activation, retention, revenue, referral). Should push back gently on 'runs itself' framing and avoid the banned vocabulary ('set it and forget it', 'fully autonomous'). Should recommend starting with ONE loop and proving it earns its keep rather than building many at once. Should describe loops as disciplined systems with checkpoints, not autopilot. Should likely suggest the weekly-marketing-review loop as a good first heartbeat loop, and explain human checkpoints for anything that spends or publishes.",
"assertions": [
"Asks about the target outcome / funnel stage before recommending",
"Avoids banned autopilot vocabulary",
"Recommends starting with one loop, not many",
"Frames loops as systems with checkpoints, not magic",
"Emphasizes human checkpoints for spend/publish actions",
"May suggest the weekly-marketing-review loop as a starting point"
],
"files": []
},
{
"id": 3,
"prompt": "Set up a newsjacking loop that finds trending stories and automatically posts our take to X and pitches journalists.",
"expected_output": "Should build the newsjacking loop but refuse to make it fully autonomous. Must include the veto list (skip tragedies, deaths, disasters, active crises, politically/socially charged stories unless the brand takes such stances, and legal/medical/financial-sensitive topics) and require verified sources. Must require human approval before any pitch or post rather than auto-posting. Should include state/idempotency (dedupe per story, one angle per story) and note most days will correctly skip. Should reference the public-relations skill.",
"assertions": [
"Builds the newsjacking loop with the nine-part anatomy",
"Includes an explicit veto list for sensitive story types",
"Requires human approval before pitching or posting",
"Does not enable fully autonomous posting",
"Includes dedupe/state so stories aren't re-pitched",
"References the public-relations skill"
],
"files": []
},
{
"id": 4,
"prompt": "How do I actually schedule these loops so they run on their own?",
"expected_output": "Should explain that the loop body is agent-agnostic but scheduling depends on the environment. Should cover the options: Claude Code native primitives (/loop, ScheduleWakeup, CronCreate), any-agent + cron (with an example cron expression), and manual cadence for high-judgment loops. Should default to time-of-day cron for review-style loops and dynamic pacing for monitor-until-threshold loops. Should treat a dedicated loop-mechanics skill like loopify as an optional companion, not a hard dependency, and not assume it is installed.",
"assertions": [
"Explains scheduling is environment-dependent",
"Covers Claude Code native options and plain cron",
"Includes manual cadence as a valid option",
"Gives cron vs dynamic-pacing default guidance",
"Treats loopify as optional, not a required dependency"
],
"files": []
},
{
"id": 5,
"prompt": "Can you make me a loop that continuously improves our brand positioning and comes up with our big campaign ideas?",
"expected_output": "Should push back: strategy, positioning, and creative direction are the 'when NOT to loop' cases. Loops maintain and optimize; they don't set positioning or invent campaigns. Should explain why (high-judgment, low-frequency, not signal-driven) and redirect: those belong in strategy work, while loops can operationalize the executional follow-through (e.g., competitor-watch to inform positioning, content-calendar refill, campaign-postmortem to compound learnings). Should not fabricate a positioning loop as if it were a good fit.",
"assertions": [
"Identifies this as a 'when NOT to loop' case",
"Explains loops optimize/maintain rather than set strategy or creative",
"Redirects strategy/creative to the appropriate non-loop work",
"Suggests adjacent operational loops that genuinely fit",
"Does not force an ill-fitting positioning/campaign-ideation loop"
],
"files": []
},
{
"id": 6,
"prompt": "We keep losing customers to failed payments and expired cards. Set up something to recover that revenue automatically.",
"expected_output": "Should select the failed-payment/dunning loop. Should trigger on payment failure or upcoming card expiration, run a retry schedule with escalating update-card messaging, and include state (track dunning stage per account, stop on recovery). Must include a stop/bail-out so it doesn't loop forever (escalate/deactivate after final retry) and a self-check to distinguish involuntary failures from intentional cancellations (don't dun someone who chose to leave). Should reference revops and emails skills. May note this is often the highest-ROI retention loop.",
"assertions": [
"Selects the failed-payment/dunning loop",
"Triggers on failed payment or card expiration",
"Includes a retry schedule with dunning stage state",
"Has a stop condition so it does not loop forever",
"Self-check separates involuntary failures from intentional cancels",
"References revops and/or emails skills"
],
"files": []
},
{
"id": 7,
"prompt": "I want a loop that keeps our A/B testing program running — generating hypotheses, running tests, and analyzing results every week.",
"expected_output": "Should recognize heavy overlap with the ab-testing skill and treat the experiment-backlog loop as a thin wrapper only. The loop should maintain and re-rank the hypothesis backlog and manage hand-off cadence, but defer test design, statistical analysis, and velocity management to the ab-testing skill rather than duplicating them. Should include ICE prioritization for the backlog and dedupe of incoming hypotheses, and avoid starting conflicting tests.",
"assertions": [
"Recognizes overlap with the ab-testing skill",
"Keeps the loop as a thin wrapper that defers to ab-testing",
"Does not duplicate test design or statistical analysis in the loop",
"Includes backlog prioritization (ICE) and hypothesis dedupe",
"References the ab-testing skill as the owner of execution"
],
"files": []
}
]
}
FILE:references/loop-catalog.md
# Marketing Loop Catalog
A library of repeatable marketing loops with thorough coverage across the funnel. Each is a complete, adaptable spec. Pick the closest match, then tune the cadence, thresholds, state handling, and human checkpoints to the user's product, stage, and tooling.
Every loop lists nine parts: **Check cadence · Acts when · Purpose · Skills used · Loop body · Self-check · State / idempotency · Stop / bail-out · Output**. See `SKILL.md` for the anatomy, the cadence rule, and when not to loop.
Two rules that apply to every entry:
- **Most runs should do nothing.** A healthy loop checks, finds nothing worth acting on, logs "no action," and exits. Loops that act every run are usually acting on noise.
- **State prevents harm.** Every loop tracks what it already did (last-run marker, dedupe key, cooldown) so it never double-acts, re-nags the same person, or re-alerts the same issue.
Loops are grouped by function. Naming follows the "The X loop" convention.
---
## SEO & Content
### The keyword-gap loop
- **Check cadence**: Weekly
- **Acts when**: A striking-distance keyword (positions 5–20) or a rising query has no adequate page.
- **Purpose**: Surface new ranking opportunities before competitors take them.
- **Skills used**: `seo-audit`, `programmatic-seo`, `content-strategy`
- **Loop body**:
1. Pull ranking + impression data (Search Console / rank tracker).
2. Diff vs. last run: new striking-distance keywords, rising queries with no matching page.
3. Classify each gap: quick on-page win / net-new page / programmatic template candidate.
4. Draft briefs for the top 3.
- **Self-check**: Movement real vs. seasonal? Compare to the same period last month, not just last week.
- **State / idempotency**: Store the set of gaps already briefed; don't re-brief an open one.
- **Stop / bail-out**: No gap clears a minimum impression threshold → log "no action." Halt on data-source outage rather than acting on partial data.
- **Output**: Up to 3 content briefs staged for review + a one-line movement summary.
### The ranking-drop watch loop
- **Check cadence**: Weekly
- **Acts when**: A priority keyword or page drops more than N positions vs. baseline.
- **Purpose**: Catch and diagnose SEO regressions before they compound.
- **Skills used**: `seo-audit`, `analytics`
- **Loop body**:
1. Track positions for priority keywords/pages.
2. Flag material drops; diff what changed (content, links, SERP layout, algo-update timing).
3. Diagnose likely cause + propose a fix.
- **Self-check**: Rule out a SERP-feature change or one-off volatility before declaring a real loss.
- **State / idempotency**: Remember which drops are already open as issues; update rather than re-file.
- **Stop / bail-out**: No material drop → log "stable." Escalate suspected algo hits to a human rather than mass-editing.
- **Output**: A regression report with a recommended fix.
### The content-decay loop
- **Check cadence**: Monthly
- **Acts when**: A page's traffic/rankings declined materially over the trailing 90 days.
- **Purpose**: Refresh decaying content before it slides out of rankings.
- **Skills used**: `copy-editing`, `seo-audit`, `content-strategy`
- **Loop body**:
1. Find pages with declining trailing-90-day traffic/rankings.
2. Pick the highest-value decayers.
3. Draft a refresh plan (update stats, expand thin sections, fix intent match, re-link).
- **Self-check**: Decay from the page itself, or from a SERP/seasonality shift? Refresh only what a refresh can fix.
- **State / idempotency**: Track last-refresh date per page; don't re-queue a page refreshed within the cooldown.
- **Stop / bail-out**: No meaningful decayers → skip.
- **Output**: A prioritized refresh list with per-page plans.
### The internal-linking loop
- **Check cadence**: On new/updated content, or weekly
- **Acts when**: A published page has fewer relevant internal links (in or out) than it should.
- **Purpose**: Distribute link equity and help new content get discovered and rank.
- **Skills used**: `seo-audit`, `site-architecture`, `content-strategy`
- **Loop body**:
1. Identify recently published/updated pages.
2. Find relevant existing pages that should link to them (and vice versa).
3. Draft the specific link insertions with anchor text.
- **Self-check**: Is each link contextually relevant, or link-stuffing? Skip forced links.
- **State / idempotency**: Track which page pairs are already linked; never suggest a duplicate.
- **Stop / bail-out**: No relevant link targets → skip. Stage edits for review; don't mass-edit live pages autonomously.
- **Output**: A list of specific internal-link edits.
### The programmatic-SEO quality loop
- **Check cadence**: Monthly
- **Acts when**: Template pages show indexation gaps, thin content, duplication, or cannibalization.
- **Purpose**: Keep large templated page sets healthy so they don't drag the whole domain.
- **Skills used**: `programmatic-seo`, `seo-audit`
- **Loop body**:
1. Sample the template page set; check indexation, word/data uniqueness, and query overlap.
2. Flag thin, duplicate, cannibalizing, or deindexed pages.
3. Recommend fix, consolidate, noindex, or prune.
- **Self-check**: Is low traffic a quality problem or just low demand? Don't prune pages that serve real long-tail intent.
- **State / idempotency**: Track pages already flagged/actioned; re-check only on the next cycle.
- **Stop / bail-out**: Set healthy → log and skip. Escalate mass-noindex/prune decisions to a human.
- **Output**: A quality report with per-bucket actions.
### The content-repurposing loop
- **Check cadence**: Weekly
- **Acts when**: A long-form asset (post/video/podcast) hasn't been repurposed yet.
- **Purpose**: Turn every long-form asset into a week of channel-native content.
- **Skills used**: `social`, `content-strategy`, `copywriting`
- **Loop body**:
1. Find the newest un-repurposed asset.
2. Extract the 3–5 strongest ideas.
3. Draft channel-native versions (LinkedIn post, X thread, short-form script).
4. Stage in the scheduling queue.
- **Self-check**: Does each piece stand alone, or read like a link-dump? Rewrite anything that only works with the original open.
- **State / idempotency**: Mark assets as repurposed; never re-process one.
- **Stop / bail-out**: Nothing new published → skip.
- **Output**: Drafts in the social queue for approval.
### The content-calendar refill loop
- **Check cadence**: Weekly
- **Acts when**: The editorial pipeline has fewer than N weeks of planned content queued.
- **Purpose**: Keep the content pipeline from running dry.
- **Skills used**: `content-strategy`, `marketing-ideas`, `seo-audit`
- **Loop body**:
1. Count planned/drafted pieces remaining in the calendar.
2. If below the buffer, generate new topic ideas from the keyword-gap output, customer questions, and pillar plan.
3. Prioritize and slot them.
- **Self-check**: Do new topics map to real search demand or audience questions, not just "content for content's sake"?
- **State / idempotency**: Dedupe proposed topics against the existing calendar and published archive.
- **Stop / bail-out**: Pipeline above buffer → skip.
- **Output**: New prioritized topics added to the calendar.
---
## Paid
### The ad-fatigue loop
- **Check cadence**: Every 2–3 days
- **Acts when**: An ad shows rising frequency + declining CTR/CVR past a real significance bar.
- **Purpose**: Refresh creative before CPA drifts up as ads fatigue.
- **Skills used**: `ads`, `ad-creative`, `analytics`
- **Loop body**:
1. Pull per-ad metrics: CTR, frequency, CPA, spend, trend vs. baseline.
2. Flag fatiguing ads and clear winners.
3. Generate 3–5 fresh variants off the winning angle.
4. Stage variants; recommend budget shift fatigued → winning.
- **Self-check**: Enough spend, impressions, and conversions to read CPA past the attribution window, and the ad is out of the learning phase. Rising frequency alone with thin conversion data is not fatigue evidence — wait.
- **State / idempotency**: Track per-ad last-refresh date; don't regenerate variants for an ad refreshed within the cooldown.
- **Stop / bail-out**: Never auto-shift budget or publish without a human checkpoint unless spend caps + an allowlist are explicitly authorized. Halt if daily spend exceeds its cap.
- **Output**: Staged creative drafts + a recommended budget move.
### The daily-creative-drop loop
- **Check cadence**: Daily (early morning, so the batch is ready when the media buyer sits down)
- **Acts when**: The grounded inputs corpus exists and the required inputs are populated — `inputs/winning-ads/` and `inputs/reviews/` (required; `inputs/comments/` and `brand/` strongly recommended, matching ad-creative's grounding rules). If a required input is empty, the loop asks for inputs instead of generating.
- **Purpose**: Keep creative volume ahead of fatigue — a standing batch of fresh static concepts to test, so scaling never stalls waiting on production.
- **Skills used**: `ad-creative` (Mode 3 + static ad template library), `customer-research`
- **Loop body**:
1. Read the inputs corpus: `inputs/winning-ads/`, `inputs/reviews/`, `inputs/comments/`, and `brand/`.
2. Generate the batch (e.g., 50 concepts) cycling all 15 static templates, 3-4 variations each, every concept grounded in a cited source.
3. Generate images if an image tool is configured; otherwise deliver concepts + image prompts.
4. Save to `outputs/YYYY-MM-DD/` with an `INDEX.md` (template type + grounding per concept).
- **Self-check**: Are concepts actually grounded (spot-check citations against sources)? Is template coverage spread across the library, not clustered on 2-3? Does copy match the brand voice doc rather than generic DR voice?
- **State / idempotency**: One batch per day — skip if today's output folder already exists. Track angle/headline hashes across recent batches to avoid regenerating near-duplicates of concepts already delivered.
- **Stop / bail-out**: Missing or empty required inputs → stop and request them; never generate ungrounded. Human picks the 5-10 to upload — this loop stages creative and **never publishes to the ad account**. If batches go unreviewed for a week, pause and ask whether to continue (unpicked batches are a vanity loop).
- **Output**: A dated folder of grounded static ad concepts + index, ready for human selection.
- **Input freshness (companion cadence)**: Weekly, refresh `inputs/winning-ads/` with anything that scaled and prune stale examples; monthly, refresh `inputs/reviews/` and `inputs/comments/` and re-check the voice doc. Stale inputs are this loop's failure mode — output quality tracks input freshness, not run count.
### The monthly-creative-retro loop
- **Check cadence**: Monthly (first business day, reading the prior month)
- **Acts when**: The account had meaningful creative activity last month — new concepts launched with enough delivery to judge (respect the impression/spend thresholds in `ads`). If nothing launched or nothing cleared thresholds, note that and skip.
- **Purpose**: Close the creative strategy loop — turn last month's results into next month's evidence-ranked slate, so the roadmap learns instead of drifting.
- **Skills used**: `ad-creative` (Mode 4 + creative-roadmap reference), `ads` (decision thresholds), `analytics`
- **Loop body**:
1. Pull last month's ad performance via the platform CLIs; map results to the month's roadmap concepts.
2. Draft the retro artifact (`retros/YYYY-MM.md`): winners with the why, losers with funnel-stage diagnosis, single-metric wins, learnings, kills.
3. Update the roadmap: re-rank icebox evidence, write learnings in as new/revised concepts, draft next month's capacity-checked slate.
4. Flag the account-state call (exploration vs. scaling) for human confirmation — the mix recommendation depends on it.
- **Self-check**: Are verdicts on concepts (not single executions)? Did every learning land somewhere — icebox update, re-rank, or kill? Did anything clear thresholds, or is this month a skip?
- **State / idempotency**: One retro per month — skip if `retros/YYYY-MM.md` exists. The roadmap file is the shared state; never fork it.
- **Stop / bail-out**: Stages analysis and a draft slate only — the human approves the slate and the account-state call; the loop **never launches or pauses ads**. If retros go unread for two cycles, pause and ask.
- **Output**: The monthly retro artifact + an updated roadmap with a draft slate for the coming month.
### The paid-search query-mining loop
- **Check cadence**: Weekly
- **Acts when**: Search-term reports reveal wasted spend or new intent.
- **Purpose**: Continuously refine keywords, negatives, and landing-page mapping.
- **Skills used**: `ads`, `analytics`
- **Loop body**:
1. Pull the search-terms report.
2. Identify irrelevant terms (→ negatives), high-performing terms (→ new exact-match), and terms whose landing page is a poor match.
3. Stage keyword/negative changes and landing-page notes.
- **Self-check**: Enough clicks/conversions per term to justify a change? Don't negate on a single click.
- **State / idempotency**: Track already-added negatives/keywords; never re-add.
- **Stop / bail-out**: No terms clear thresholds → skip. Stage changes for review before pushing to the account.
- **Output**: A staged list of negatives, new keywords, and LP mismatches.
### The retargeting-hygiene loop
- **Check cadence**: Weekly
- **Acts when**: Audiences are stale, too small, over-frequent, or missing exclusions.
- **Purpose**: Keep retargeting efficient and non-annoying.
- **Skills used**: `ads`, `analytics`
- **Loop body**:
1. Review retargeting audiences: size, recency, frequency, exclusions, creative sequencing.
2. Flag issues (converters not excluded, audiences too small to serve, frequency too high).
3. Recommend fixes.
- **Self-check**: Is the audience actually underperforming, or just small-but-valuable? Don't kill high-intent segments for size.
- **State / idempotency**: Track which audiences were already fixed this cycle.
- **Stop / bail-out**: All healthy → skip. Human-approve audience deletions.
- **Output**: A hygiene report with recommended audience changes.
### The landing-page regression loop
- **Check cadence**: Weekly (or on deploy)
- **Acts when**: A top acquisition page regresses on conversion, speed, tracking, or form function.
- **Purpose**: Catch silent breakage on the pages that receive paid/organic traffic.
- **Skills used**: `cro`, `analytics`
- **Loop body**:
1. Monitor top acquisition pages: conversion rate, load speed, form submits, tracking fires.
2. Flag regressions vs. baseline; correlate with recent deploys/changes.
3. Diagnose and propose a fix.
- **Self-check**: Rule out tracking breakage vs. a real conversion drop before raising an alarm — and vice versa.
- **State / idempotency**: Track open regressions; update rather than re-file.
- **Stop / bail-out**: No regression → log "stable." Escalate a live-revenue-page break immediately, don't wait for the next run.
- **Output**: A regression alert with cause + fix.
---
## Earned, Social & Partnerships
### The newsjacking loop
- **Check cadence**: Daily
- **Acts when**: A trending story matches the brand's space, clears newsworthiness + fit, **and** passes the veto list.
- **Purpose**: Ride relevant news with a timely angle before the window closes.
- **Skills used**: `public-relations`, `social`
- **Loop body**:
1. Scan news/HN/Reddit/X for stories intersecting the product's space.
2. Score newsworthiness + fit + reach.
3. Run the veto list. For a surviving top story, draft an angle (post, pitch, or commentary).
- **Self-check**: Is the angle genuinely additive, or forced? Kill forced takes — they cost credibility.
- **Veto list (skip immediately)**: tragedies, deaths, disasters, active crises; politically or socially charged stories unless the brand explicitly takes such stances; legal/medical/financial-sensitive topics; anything sourced from an unverified/single unreliable source.
- **State / idempotency**: Dedupe on story ID; one angle per story; never re-pitch a covered story.
- **Stop / bail-out**: Any veto trip → skip. Always require human approval before pitching/posting. Most days will skip — that's correct.
- **Output**: A staged post/pitch for human approval, or nothing.
### The social-listening loop
- **Check cadence**: Daily
- **Acts when**: A thread/mention clears the ICP-fit + intent + reach score.
- **Purpose**: Surface the highest-value conversations to engage in, instead of scrolling feeds.
- **Skills used**: `social` (see its `references/listening.md`), `community-marketing`
- **Loop body**:
1. Pull mentions and relevant threads across configured sources.
2. Score by ICP fit, intent, reach, and comment opportunity.
3. Draft comments/replies for the top handful.
- **Self-check**: Would a human recognize each reply as genuinely useful, not promotional?
- **State / idempotency**: Track already-engaged threads; never double-reply. Respect a per-account interaction cooldown.
- **Stop / bail-out**: Nothing clears the threshold → skip. Stage replies for human post (don't auto-post — bot-detection + brand risk).
- **Output**: A short list of threads with drafted, on-brand replies.
### The community-engagement loop
- **Check cadence**: Daily
- **Acts when**: A target community (subreddit/Slack/Discord/forum) has a relevant thread where a helpful, non-promotional reply fits.
- **Purpose**: Build durable presence and trust in the communities where the ICP lives.
- **Skills used**: `community-marketing`, `social`
- **Loop body**:
1. Scan configured communities for relevant threads/questions.
2. Score for genuine help opportunity (not just keyword match).
3. Draft value-first replies; note any that warrant a longer resource.
- **Self-check**: Does the reply lead with help and respect community norms? Self-promo ratio stays low.
- **State / idempotency**: Track engaged threads + per-community posting cadence to avoid over-posting.
- **Stop / bail-out**: No genuine-help opportunity → skip. Stage for human review where communities are strict about vendors.
- **Output**: Drafted community replies + resource ideas.
### The competitor-watch loop
- **Check cadence**: Weekly
- **Acts when**: A competitor makes a substantive pricing, positioning, product, or messaging change.
- **Purpose**: Catch competitor moves early enough to respond.
- **Skills used**: `competitor-profiling`, `competitors`, `product-marketing`
- **Loop body**:
1. Fetch competitor pricing pages, homepages, changelogs, recent posts.
2. Diff vs. last snapshot.
3. Summarize meaningful changes; flag anything needing a response (comparison-page update, counter-messaging).
- **Self-check**: Substantive vs. cosmetic? Don't raise a copy tweak as a strategic shift.
- **State / idempotency**: Store per-competitor snapshots; diff against the last, and don't re-flag a known change.
- **Stop / bail-out**: No meaningful diffs → log "no change."
- **Output**: A change digest + recommended responses.
### The backlink-prospecting loop
- **Check cadence**: Weekly
- **Acts when**: New relevant link/guest-post/mention targets appear (or the pipeline is thin).
- **Purpose**: Keep a steady flow of link-building and earned-mention opportunities.
- **Skills used**: `public-relations`, `seo-audit`
- **Loop body**:
1. Find new prospects: sites linking to competitors, relevant roundups, unlinked brand mentions, resource pages.
2. Qualify by relevance + authority.
3. Draft outreach angles for the top targets.
- **Self-check**: Is the target genuinely relevant, or a low-quality link that could hurt? Skip spammy sites.
- **State / idempotency**: Track already-contacted targets + outcomes; respect a follow-up cadence, don't re-pitch cold.
- **Stop / bail-out**: No qualified new targets → skip. Human-approve outreach sends.
- **Output**: A qualified prospect list with drafted outreach.
### The directory-submission loop
- **Check cadence**: Monthly
- **Acts when**: A relevant new directory/launch platform/marketplace exists that the product isn't listed on.
- **Purpose**: Steadily expand distribution and referral/SEO footprint via directories.
- **Skills used**: `directory-submissions`
- **Loop body**:
1. Check for new/relevant directories, launch sites, and marketplaces.
2. Qualify by relevance, authority, and audience fit.
3. Prepare listing copy/assets for the top ones.
- **Self-check**: Real audience/SEO value, or a link farm? Skip low-quality directories.
- **State / idempotency**: Maintain a submitted-directories list; never resubmit.
- **Stop / bail-out**: No worthwhile new directories → skip.
- **Output**: Prepared listings staged for submission.
### The partner-pipeline loop
- **Check cadence**: Monthly
- **Acts when**: A viable co-marketing, integration, affiliate, or newsletter-swap opportunity surfaces (or the pipeline is thin).
- **Purpose**: Keep a fresh pipeline of partnership and co-marketing opportunities.
- **Skills used**: `co-marketing`, `referrals`
- **Loop body**:
1. Scan for potential partners (complementary tools, aligned audiences, active newsletters, integration targets).
2. Qualify by audience overlap + reach + fit.
3. Draft partnership/swap outreach for the top prospects.
- **Self-check**: Real audience overlap and mutual value, or a one-sided ask? Skip mismatches.
- **State / idempotency**: Track contacted partners + status; respect follow-up cadence.
- **Stop / bail-out**: No qualified opportunities → skip. Human-approve outreach.
- **Output**: A qualified partner list with drafted outreach.
---
## Activation
### The onboarding drop-off loop
- **Check cadence**: Weekly
- **Acts when**: An onboarding step's drop exceeds benchmark or regresses vs. last period.
- **Purpose**: Find and fix the biggest leak between signup and first value.
- **Skills used**: `onboarding`, `analytics`, `cro`
- **Loop body**:
1. Pull the activation funnel step-by-step (signup → key action → aha).
2. Identify the worst-dropping step vs. benchmark and last period.
3. Diagnose likely cause; propose one focused fix + how to measure it.
- **Self-check**: Enough new users through the funnel for step rates to be significant?
- **State / idempotency**: Track which fixes were already proposed/shipped; measure their effect before re-touching.
- **Stop / bail-out**: Sample too small → widen window or skip.
- **Output**: One prioritized activation fix with a measurement plan.
### The signup-funnel-leak loop
- **Check cadence**: Weekly
- **Acts when**: A signup/checkout step regresses vs. baseline.
- **Purpose**: Keep the signup/checkout path converting as the site changes.
- **Skills used**: `signup`, `cro`, `analytics`, `ab-testing`
- **Loop body**:
1. Pull conversion by step across the signup/checkout flow.
2. Compare to baseline; flag regressions (a deploy or copy change may have hurt it).
3. Draft a hypothesis + test for the worst step (hand test execution to `ab-testing`).
- **Self-check**: Rule out tracking breakage before declaring a real drop.
- **State / idempotency**: Track open regressions + running tests; don't start a conflicting test.
- **Stop / bail-out**: No regression and no test-worthy idea → skip.
- **Output**: A prioritized experiment brief for `ab-testing`.
### The lead-capture-asset loop
- **Check cadence**: Monthly
- **Acts when**: A lead magnet, free tool, or opt-in underperforms on capture rate.
- **Purpose**: Keep top-of-funnel capture assets (lead magnets + free tools) converting visitors to leads.
- **Skills used**: `lead-magnets`, `free-tools`, `cro`, `popups`
- **Loop body**:
1. Pull view → capture conversion for each lead magnet, free tool, and opt-in.
2. Flag underperformers vs. benchmark; diagnose (offer, placement, form friction, targeting).
3. Propose a fix or refresh (new angle, better placement, reduced friction).
- **Self-check**: Enough traffic per asset for the capture rate to be meaningful?
- **State / idempotency**: Track last-optimized date per asset; cooldown before re-touching.
- **Stop / bail-out**: All assets healthy → skip.
- **Output**: A prioritized fix per underperforming asset.
### The feature-adoption loop
- **Check cadence**: Weekly
- **Acts when**: A sticky/valuable feature is underused by a segment that would benefit.
- **Purpose**: Drive adoption of the features that correlate with retention.
- **Skills used**: `onboarding`, `emails`, `analytics`
- **Loop body**:
1. Identify high-retention-correlated features and the segments not using them.
2. Pick the highest-leverage feature × segment.
3. Draft an in-app nudge or email to drive adoption.
- **Self-check**: Is the feature genuinely valuable to that segment, or would the nudge be noise? Don't push features people rationally skip.
- **State / idempotency**: Track who's already been nudged for which feature; enforce a cooldown; suppress adopters.
- **Stop / bail-out**: No clear feature × segment gap → skip.
- **Output**: A staged adoption nudge.
---
## Retention
### The churn-signal loop
- **Check cadence**: Daily (or on-trigger)
- **Acts when**: An account newly crosses a churn-risk threshold and isn't already in an intervention.
- **Purpose**: Intervene inside the short window before an at-risk account leaves.
- **Skills used**: `churn-prevention`, `analytics`, `emails`
- **Loop body**:
1. Score accounts on churn-risk signals (usage decline, seat drop, dunning, support escalations).
2. Segment newly at-risk accounts.
3. Match each to the right intervention (re-engagement email, CS outreach, offer); stage it.
- **Self-check**: Is the "drop" a real trend or a weekend/holiday dip? Compare to the account's own baseline.
- **State / idempotency**: Never re-trigger on an account already in an active intervention; enforce a cooldown between attempts.
- **Stop / bail-out**: No newly at-risk accounts → skip. Escalate high-value accounts to a human rather than auto-emailing.
- **Output**: A prioritized at-risk list with staged interventions.
### The lifecycle-email-refresh loop
- **Check cadence**: Monthly
- **Acts when**: A sequence email underperforms on real engagement or contains stale content.
- **Purpose**: Keep automated sequences performing as the product and audience evolve.
- **Skills used**: `emails`, `analytics`, `copy-editing`
- **Loop body**:
1. Pull per-email performance — clicks, conversions, replies, unsubscribes, spam complaints, bounces (**not opens** — open tracking is unreliable post-privacy-changes).
2. Flag weak performers and stale references (old features, dates, pricing).
3. Draft rewrites or subject-line tests for the bottom performers.
- **Self-check**: Enough sends per email for rates to be meaningful?
- **State / idempotency**: Track last-revised date per email; cooldown before re-testing.
- **Stop / bail-out**: All sequences healthy → skip. **Pause and escalate any sequence with rising complaint/bounce rates — that's a deliverability emergency, not a copy tweak.**
- **Output**: Staged email rewrites + subject-line tests.
### The re-engagement loop
- **Check cadence**: Weekly
- **Acts when**: A user newly crosses the inactivity threshold.
- **Purpose**: Win back dormant users before they're gone for good.
- **Skills used**: `emails`, `sms`, `offers`
- **Loop body**:
1. Identify users newly crossing the inactivity threshold.
2. Pick the win-back angle (new feature, offer, "we miss you," sunset warning).
3. Draft the message; set suppression so they aren't re-hit next week.
- **Self-check**: Truly dormant, or just low-frequency-by-design users? Don't nag healthy accounts.
- **State / idempotency**: Track win-back attempts per user; suppress after each send for the cooldown.
- **Stop / bail-out**: After N unsuccessful attempts, move to sunset — not another email.
- **Output**: A staged win-back message + updated suppression list.
### The email-deliverability loop
- **Check cadence**: Weekly
- **Acts when**: Bounce, complaint, or unsubscribe rates rise, or list-hygiene decays.
- **Purpose**: Protect sender reputation and inbox placement.
- **Skills used**: `emails`, `analytics`
- **Loop body**:
1. Monitor bounces, spam complaints, unsubscribes, domain/DKIM/SPF/DMARC health, and inbox-placement signals.
2. Flag rising problem rates or authentication issues.
3. Recommend actions: suppress hard bounces, sunset chronically unengaged, fix auth, throttle.
- **Self-check**: Is a spike a one-off send or a trend? Correlate with recent campaigns.
- **State / idempotency**: Track already-suppressed addresses + last hygiene sweep date.
- **Stop / bail-out**: All metrics healthy → log and skip. **Escalate a complaint-rate spike immediately** — reputation damage compounds fast.
- **Output**: A deliverability report + a suppression/hygiene action list.
### The voice-of-customer loop
- **Check cadence**: Weekly
- **Acts when**: New feedback (NPS, surveys, support tickets, reviews, calls) has arrived.
- **Purpose**: Route feedback to the right action **and** mine it for marketing inputs.
- **Skills used**: `customer-research`, `churn-prevention`, `referrals`, `copywriting`
- **Loop body**:
1. Collect new feedback across sources.
2. Route: detractors/at-risk → save motion (`churn-prevention`); promoters → referral/review ask (`referrals`); recurring pain/desire → experiment + copy inputs.
3. Extract verbatim customer language for copy, FAQ, and objection-handling.
- **Self-check**: Is a theme a real pattern or one loud voice? Require a minimum count before acting on it.
- **State / idempotency**: Track processed feedback IDs; never double-route the same item.
- **Stop / bail-out**: No new feedback → skip. Escalate sensitive/legal complaints to a human.
- **Output**: Routed actions + a language/insight digest for marketing.
---
## Revenue
### The trial-conversion loop
- **Check cadence**: Daily
- **Acts when**: A trial user reaches a conversion-relevant moment (mid-trial, near-expiry, activated-but-not-paid).
- **Purpose**: Move more trials to paid with well-timed nudges.
- **Skills used**: `emails`, `paywalls`, `analytics`, `offers`
- **Loop body**:
1. Segment active trials by stage and activation level.
2. Match each to the right nudge (value recap, use-case tip, near-expiry push, offer).
3. Stage the nudge.
- **Self-check**: Is the user activated enough for a paid push to land, or do they need more value first?
- **State / idempotency**: Track nudges sent per trial; enforce cadence; suppress converters.
- **Stop / bail-out**: No trials at an actionable stage → skip. Don't over-message a single trial.
- **Output**: Staged, stage-appropriate trial nudges.
### The PQL / upgrade-intent loop
- **Check cadence**: Daily
- **Acts when**: A free/trial user shows product-qualified buying intent (usage limits, key-feature use, team invites).
- **Purpose**: Catch high-intent users and stage upgrade outreach at the right moment.
- **Skills used**: `analytics`, `sales-enablement`, `revops`
- **Loop body**:
1. Score free/trial users on PQL signals.
2. Surface newly qualified users.
3. Stage the right motion (in-app upgrade prompt, sales-assist for high-value, targeted email).
- **Self-check**: Is the signal genuine buying intent or incidental usage? Calibrate the threshold to avoid false positives.
- **State / idempotency**: Track already-actioned PQLs; don't re-route within the cooldown.
- **Stop / bail-out**: No newly qualified users → skip. Route high-value accounts to a human, don't auto-close.
- **Output**: A prioritized PQL list with staged motions.
### The pricing-page-experiment loop
- **Check cadence**: Monthly (tests run longer)
- **Acts when**: No test is running on the page and there's a worthwhile hypothesis — or a running test has concluded.
- **Purpose**: Improve pricing-page conversion **and revenue quality**, continuously.
- **Skills used**: `pricing`, `ab-testing`, `cro`
- **Loop body**:
1. Review pricing-page conversion, plan mix, and revenue-per-visitor.
2. Generate one pricing/packaging/copy hypothesis, or read a concluded test.
3. Hand design/analysis to `ab-testing`; promote a clean winner.
- **Self-check**: Judge winners on **revenue per visitor, plan mix, refunds, downgrades, churn, and support load — not conversion rate alone.** Is the running test statistically done before you call it?
- **State / idempotency**: Track the running test + concluded-test log; never start a conflicting test on the same page.
- **Stop / bail-out**: A test is in flight → hold. **Do not promote a variant that lifts conversion but lowers revenue-per-visitor or raises refunds/churn.**
- **Output**: A test result + next hypothesis.
### The paywall-optimization loop
- **Check cadence**: Monthly
- **Acts when**: No paywall test is running and there's a hypothesis — or one has concluded.
- **Purpose**: Improve in-app upgrade conversion without degrading revenue quality.
- **Skills used**: `paywalls`, `ab-testing`, `analytics`
- **Loop body**:
1. Pull paywall view → upgrade conversion and bounce points.
2. Form one hypothesis (trigger timing, framing, plan anchor), or read a concluded test.
3. Hand execution to `ab-testing`.
- **Self-check**: Segment by plan/cohort — an aggregate number can hide a segment that's tanking. Watch refunds/downgrades alongside conversion.
- **State / idempotency**: Track running/concluded tests; no conflicting tests.
- **Stop / bail-out**: Test in flight → hold. Don't promote a conversion win that raises refunds or churn.
- **Output**: A test result + next hypothesis.
### The expansion / upsell loop
- **Check cadence**: Weekly
- **Acts when**: An existing paid account hits an expansion signal (usage near limits, added seats, new use case).
- **Purpose**: Grow revenue from existing customers via well-timed upsell/cross-sell.
- **Skills used**: `revops`, `sales-enablement`, `emails`
- **Loop body**:
1. Score paid accounts on expansion signals.
2. Surface newly expansion-ready accounts.
3. Stage the right motion (usage-based upgrade prompt, CSM outreach, cross-sell offer).
- **Self-check**: Is the account healthy enough that an upsell won't sour the relationship? Don't upsell an at-risk account — that's a churn loop's job.
- **State / idempotency**: Track upsell touches per account; enforce cadence.
- **Stop / bail-out**: No expansion-ready accounts → skip. Route strategic accounts to a human.
- **Output**: A prioritized expansion list with staged motions.
### The failed-payment / dunning loop
- **Check cadence**: Daily
- **Acts when**: A payment fails or a card is about to expire.
- **Purpose**: Recover involuntary churn — often the highest-ROI retention work.
- **Skills used**: `revops`, `emails`
- **Loop body**:
1. Detect failed payments and upcoming card expirations.
2. Trigger the dunning sequence (retry schedule + escalating update-card messaging).
3. Route persistent failures to a human/CS.
- **Self-check**: Is the failure involuntary (card issue) vs. an intentional cancel? Don't dun someone who chose to leave.
- **State / idempotency**: Track dunning stage per account; follow the retry schedule; stop on recovery.
- **Stop / bail-out**: After the final retry, escalate/deactivate per policy — don't loop forever.
- **Output**: An active dunning queue + recovery status.
---
## Referral & Advocacy
### The referral-nudge loop
- **Check cadence**: Weekly
- **Acts when**: A user hits a "happy moment" (milestone, positive NPS) and hasn't been asked recently.
- **Purpose**: Ask for referrals when users are most delighted.
- **Skills used**: `referrals`, `emails`
- **Loop body**:
1. Identify users who just hit a happy moment and aren't in the ask-cooldown.
2. Match to the right ask (share link, incentive, review request).
3. Stage the ask.
- **Self-check**: Genuinely a happy moment, or just any event? A bad-timing ask erodes goodwill.
- **State / idempotency**: Enforce a cooldown — never ask the same user twice in the window.
- **Stop / bail-out**: No one at a happy moment → skip.
- **Output**: A staged, well-timed referral ask.
### The review-and-UGC-harvest loop
- **Check cadence**: Weekly
- **Acts when**: New reviews, testimonials, or user-generated content have appeared.
- **Purpose**: Keep a steady flow of social proof and route it into marketing.
- **Skills used**: `social`, `referrals`, `sales-enablement`, `cro`
- **Loop body**:
1. Collect new reviews/testimonials/UGC/mentions since last run.
2. Sort by strength and relevance.
3. Draft where each should go (site proof section, ad, social post, sales deck).
4. Flag anything negative for a human response.
- **Self-check**: Is it genuinely strong and on-message? Don't force weak proof into prime placement.
- **State / idempotency**: Track already-harvested items; never re-use the same one twice.
- **Stop / bail-out**: **Verify consent and platform ToS before public reuse; add FTC-required disclosure for incentivized content.** No verifiable consent, or platform prohibits reuse → don't use. Negative/sensitive → escalate to a human, don't auto-publish.
- **Output**: New proof assets routed to their destinations.
### The review-site-management loop
- **Check cadence**: Weekly
- **Acts when**: New reviews land on G2/Capterra/app stores, or listings drift out of date.
- **Purpose**: Maintain reputation and conversion on third-party review platforms.
- **Skills used**: `sales-enablement`, `social`, `cro`
- **Loop body**:
1. Track new reviews across review sites/app stores.
2. Draft responses (thank promoters, address detractors constructively).
3. Flag listing updates needed (screenshots, features, pricing).
- **Self-check**: Is the response specific and non-defensive? Never argue publicly with a reviewer.
- **State / idempotency**: Track responded reviews; never double-respond.
- **Stop / bail-out**: No new reviews/updates → skip. Human-approve responses to negative/legal-sensitive reviews.
- **Output**: Drafted responses + a listing-update checklist.
### The case-study-sourcing loop
- **Check cadence**: Monthly
- **Acts when**: A customer hits case-study-worthy success (strong results, milestone, enthusiastic feedback).
- **Purpose**: Keep a pipeline of case studies and customer stories.
- **Skills used**: `sales-enablement`, `customer-research`, `referrals`
- **Loop body**:
1. Identify customers with standout results/engagement.
2. Qualify for a case study (results, willingness, logo value).
3. Draft the outreach + interview questions.
- **Self-check**: Are the results real and attributable, or coincidental? Verify before pitching a story.
- **State / idempotency**: Track approached customers + status; respect a no-repeat cooldown.
- **Stop / bail-out**: No qualified candidates → skip. Human-approve customer outreach.
- **Output**: A candidate list with drafted outreach.
---
## Ongoing Ops / Meta
### The weekly-marketing-review loop
- **Check cadence**: Weekly (Mon 9am)
- **Acts when**: Always runs — this is the heartbeat. It "acts" by flagging the week's notable movers.
- **Purpose**: One standing full-funnel pulse so nothing drifts unnoticed.
- **Skills used**: `analytics`, `marketing-plan`, `marketing-ideas`
- **Loop body**:
1. Pull top-line AARRR metrics vs. last week and vs. plan.
2. Flag the biggest mover (good and bad) per stage.
3. Tie each flag to the loop or skill that should act on it; surface 1–2 experiment ideas.
- **Self-check**: Distinguish trend from noise before raising an alarm.
- **State / idempotency**: Store each week's snapshot for accurate week-over-week deltas.
- **Stop / bail-out**: Manual disable + error-halt. On a data-source outage, report "stale data," never fabricated movement. (Not "n/a" — even the heartbeat needs an off switch and an error path.)
- **Output**: A one-page weekly digest with owners/next actions.
### The experiment-backlog loop
- **Check cadence**: Weekly
- **Acts when**: New hypotheses exist to log, the backlog needs re-ranking, or a test slot is free.
- **Purpose**: Keep the experiment pipeline full and prioritized. **Thin wrapper — defer all test design, statistical analysis, and velocity management to `ab-testing`.**
- **Skills used**: `ab-testing` (owner), `cro`, `analytics`
- **Loop body**:
1. Harvest new hypotheses from the week (data, research, competitors, support, other loops).
2. Re-rank the backlog with ICE.
3. If a slot is free, hand the top idea to `ab-testing`; if a test concluded there, log the learning.
- **Self-check**: Is the top idea actually testable with current traffic, or ICE-inflated?
- **State / idempotency**: Dedupe incoming hypotheses against the backlog; track which tests are live.
- **Stop / bail-out**: Backlog full and a test running → just log new ideas. Don't duplicate `ab-testing`'s job.
- **Output**: An updated, ranked backlog (the source of record lives with `ab-testing`).
### The analytics-anomaly loop
- **Check cadence**: Daily
- **Acts when**: A tracked metric breaks its expected band (spike or drop beyond normal variance).
- **Purpose**: Catch anything breaking — good or bad — before it runs for days unnoticed.
- **Skills used**: `analytics`
- **Loop body**:
1. Check key metrics (traffic, signups, conversion, revenue, spend) against their normal range.
2. Flag anomalies; separate "real event" from "tracking artifact."
3. Route each to the responsible loop/owner for diagnosis.
- **Self-check**: Is the anomaly real or a tracking/seasonality artifact? Check for known causes (holiday, launch, deploy) before alarming.
- **State / idempotency**: Track already-alerted anomalies; don't re-alert the same ongoing one daily.
- **Stop / bail-out**: All metrics in-band → silent (no alert = good). Escalate a revenue/spend anomaly immediately.
- **Output**: An anomaly alert routed to an owner, or nothing.
### The brand-mention / reputation loop
- **Check cadence**: Daily
- **Acts when**: A meaningful brand mention appears anywhere (not just where you're listening for engagement).
- **Purpose**: Monitor and protect reputation; respond where it matters.
- **Skills used**: `social`, `public-relations`
- **Loop body**:
1. Scan the open web/social/forums for brand mentions.
2. Classify sentiment + reach + risk.
3. Route: positive → amplify/thank; negative/risky → drafted response for human review; unlinked mention → backlink-prospecting.
- **Self-check**: Does a negative mention need a response, or would engaging amplify it? Judge reach + legitimacy.
- **State / idempotency**: Dedupe on mention ID; track handled mentions.
- **Stop / bail-out**: No meaningful mentions → skip. **Always human-approve responses to negative/crisis mentions** — never auto-reply to a complaint.
- **Output**: A mention digest with routed actions.
### The tracking-QA loop
- **Check cadence**: Weekly (and on deploy / campaign launch)
- **Acts when**: Analytics, pixels, UTMs, or conversion events are missing, misfiring, or misconfigured.
- **Purpose**: Keep the measurement layer trustworthy — every other loop depends on it.
- **Skills used**: `analytics`
- **Loop body**:
1. Verify key events fire correctly, pixels are present, UTMs are consistent, and conversions attribute.
2. Flag broken/missing/duplicate tracking, especially after deploys or new campaigns.
3. Recommend fixes.
- **Self-check**: Is it truly broken, or an expected change? Confirm against a known-good baseline.
- **State / idempotency**: Track open tracking issues; update rather than re-file.
- **Stop / bail-out**: All tracking healthy → log "clean." **Escalate a broken revenue/conversion event immediately** — every downstream loop is blind until it's fixed.
- **Output**: A tracking-QA report with prioritized fixes.
### The campaign-postmortem loop
- **Check cadence**: On campaign end (event-based)
- **Acts when**: A campaign (launch, promo, seasonal push) concludes.
- **Purpose**: Capture results, lessons, and reusable assets so each campaign compounds.
- **Skills used**: `analytics`, `marketing-plan`
- **Loop body**:
1. Pull final campaign results vs. goals.
2. Capture what worked, what didn't, and why; save reusable assets (copy, creative, workflows).
3. Feed learnings into the experiment backlog and the next plan; log follow-ups.
- **Self-check**: Are conclusions supported by the data, or hindsight narrative? Separate correlation from cause.
- **State / idempotency**: One postmortem per campaign; don't re-run on an already-documented campaign.
- **Stop / bail-out**: No concluded campaign → skip.
- **Output**: A postmortem doc + backlog/plan inputs.
---
## Adapting and authoring loops
To adapt a loop: keep all nine anatomy parts, swap skills/thresholds for the user's stack, and re-tune cadence to signal speed. To author a brand-new one: use `loop-template.md` (copy-paste template + fill-in prompts + worked example + ship checklist). Either way, do not ship a loop until every part is filled — especially **State / idempotency**, **Self-check**, and **Stop / bail-out**. A loop without those isn't a system; it's a way to do the wrong thing on a schedule, repeatedly, to the same people.
FILE:references/loop-guardrails.md
# Loop Guardrails & Compliance
Loops act on a schedule, often on customer data, sometimes with money or a public voice. This reference consolidates the safety rules that keep autonomous loops from doing harm. Apply it to every loop that sends, spends, publishes, or touches personal data.
## The two-tier action model
Classify every action a loop can take:
**Tier 1 — Autonomous-safe** (a loop may do these unattended):
read data, analyze, diff, score, **draft**, and **stage** work for review.
**Tier 2 — Gated** (require a human checkpoint by default):
**spend** money, **shift budget**, **send** messages, **publish** anything public, **delete/suppress** records, **change** live account settings.
A Tier-2 action may run without a per-action human check only if the user has **explicitly authorized** it *and* it's bounded by caps + an allowlist (below). Absent that, the loop stages a draft and a human approves.
## Spend guardrails (ad-fatigue, paid-search, retargeting, expansion)
- **Hard caps**: a daily/weekly spend ceiling the loop can never exceed; halt and alert if approached.
- **Per-run change limit**: cap how much budget can move in one run (e.g., ≤20%), so a bad read can't reallocate everything.
- **Allowlist**: only specified accounts/campaigns are eligible for autonomous changes; everything else is staged.
- **Directional guardrails**: judge paid changes on revenue/ROAS, not just CTR/CPA — never optimize a proxy metric into a revenue loss.
## Publish & send guardrails (email, social, PR, community, reviews)
- **Default to a staging queue** + human approval for anything public or outbound. Auto-*drafting* is fine; auto-*publishing* is not, unless explicitly authorized.
- **Volume caps**: per-run and per-recipient limits so a loop can't blast a list or over-post a channel.
- **Suppression first**: always check suppression/unsubscribe/do-not-contact lists before sending.
- **No auto-posting where detection/ToS bites**: owned social, press pitches, and community replies are staged for a human (bot detection + brand risk).
## Compliance
Match each rule to the loops it governs:
- **CAN-SPAM / CASL (email/SMS loops — lifecycle, re-engagement, churn, trial, dunning, referral)**: honor unsubscribes immediately and permanently; include a working unsubscribe + physical address; identify the sender; don't email/text without a lawful basis or consent; scrub against suppression every send.
- **GDPR / CCPA (any loop touching personal data)**: process on a lawful basis; get consent for EU marketing; honor deletion and opt-out requests; minimize data pulled and retained; don't repurpose data beyond its collected purpose.
- **FTC (review-and-UGC-harvest, referral, social)**: disclose material connections and incentives (#ad, "I was compensated"); only use testimonials with permission; no fabricated or cherry-picked-to-mislead claims.
- **Platform ToS (social-listening, community-engagement, review-site-management, scraping-based loops)**: respect rate limits and automation rules; follow review-platform response policies; don't scrape or auto-act where prohibited.
When a loop can't confirm consent, permission, or ToS-compatibility, its stop condition is **don't act** — stage for a human instead.
## PII handling
- Don't log raw PII in loop **state** or **run logs** — use internal IDs or hashes.
- Pull the minimum personal data needed to make the decision; don't hoard it in state.
- Keep exports and drafts out of shared/synced locations unless intended.
## Always-escalate list
These never run fully autonomously — route to a human regardless of authorization:
- Negative or crisis brand mentions; responses to complaints or legal/medical/financial-sensitive issues.
- Newsjacking angles (see the veto list in the catalog) — human approval before any pitch/post.
- High-value or strategic accounts (enterprise, at-risk logos).
- Anomalies in **revenue** or **ad spend** — flag immediately, don't self-correct.
- Anything that would delete data or contact a large audience at once.
## Kill switch
Every scheduled loop needs a manual off switch, and you should know how to stop **all** loops fast (disable the schedule / cron, or a global flag the loop bodies check). Document it where the loops are scheduled. A loop you can't stop quickly is a liability.
## Pre-launch guardrail checklist
Before scheduling any loop that sends, spends, publishes, or touches personal data:
- [ ] Every action is classified Tier 1 (auto) or Tier 2 (gated).
- [ ] Tier-2 actions are staged for approval — or bounded by explicit authorization + caps + allowlist.
- [ ] Spend loops have a hard cap and a per-run change limit.
- [ ] Send loops check suppression/unsubscribe and have volume caps.
- [ ] Applicable compliance rules (CAN-SPAM/GDPR/FTC/ToS) are satisfied, with "don't act" as the fallback.
- [ ] No raw PII in state or logs.
- [ ] The always-escalate cases route to a human.
- [ ] There's a documented kill switch.
FILE:references/loop-orchestration.md
# Loop Orchestration & Rollout
Loops aren't independent scripts — they compose into a marketing operating system. This reference covers how they fit together and the order to adopt them so you never build 43 at once.
## The system view
Loops fall into four layers. Data flows down and learnings flow back up.
```
SENSING analytics-anomaly · tracking-QA · weekly-marketing-review
│ (detect what changed; trust the numbers first)
▼
DIAGNOSTIC per-stage watchers — onboarding drop-off, churn-signal,
ranking-drop, landing-page regression, ad-fatigue, …
│ (figure out what to do about it)
▼
ACTION staged drafts, nudges, outreach, budget moves
│ (mostly human-checkpointed)
▼
LEARNING experiment-backlog · campaign-postmortem · voice-of-customer
│ (capture what worked)
└──────────────► feeds back into SENSING & DIAGNOSTIC
```
Key connective tissue:
- **weekly-marketing-review is the router.** It reads top-line metrics and dispatches each notable mover to the loop that owns it. It's the one loop that sees the whole board.
- **tracking-QA + analytics-anomaly are the foundation.** Every other loop reads from analytics. If tracking is broken, every downstream loop acts on lies. These come first.
- **experiment-backlog is the sink.** Hypotheses generated by many loops (signup-leak, pricing, onboarding, voice-of-customer) converge here, then hand off to `ab-testing`. Don't let each loop run its own tests.
- **voice-of-customer is a source.** Customer language it mines feeds copy for ad-fatigue, lifecycle-email, landing-page, and pricing loops.
- **campaign-postmortem closes the loop.** Its learnings become next quarter's hypotheses and plan inputs.
Avoid duplicate ownership: when two loops could act on the same signal, one owns the action and the other just flags. (E.g., an at-risk account belongs to churn-signal, not expansion/upsell — never upsell an account that's churning.)
## Rollout path (adopt in this order)
Add a loop only when the loops before it are running and earning their keep. Each stage assumes the previous one is solid.
**Stage 0 — Foundation (trust the data + see the board).**
`tracking-QA`, `weekly-marketing-review`.
You cannot run any loop responsibly on untrustworthy data or without a full-funnel pulse. This is non-negotiable and comes first.
**Stage 1 — Plug the leaks (highest ROI, protects existing revenue).**
`failed-payment/dunning`, `churn-signal`, `lifecycle-email-refresh`.
Recovering customers you already have is cheaper than acquiring new ones. Dunning alone often pays for the whole system.
**Stage 2 — Convert what you already get (fix the bucket before adding water).**
`onboarding drop-off`, `signup-funnel-leak`, `trial-conversion`.
More traffic into a leaky funnel is waste. Seal activation and conversion next.
**Stage 3 — Grow the top (now scale acquisition).**
`keyword-gap`, `content-repurposing`, `ad-fatigue`, `social-listening`, `analytics-anomaly`.
With the bucket sealed, turn on demand generation and the safety-net anomaly watcher.
**Stage 4 — Optimize monetization.**
`pricing-page-experiment`, `paywall-optimization`, `PQL/upgrade-intent`, `expansion/upsell`.
Once volume is healthy, tune revenue per user — judged on revenue quality, not conversion alone.
**Stage 5 — Compounding & advocacy.**
`referral-nudge`, `review-and-UGC-harvest`, `review-site-management`, `case-study-sourcing`, `partner-pipeline`, `brand-mention/reputation`, `experiment-backlog`, `campaign-postmortem`.
The flywheel: happy customers and earned media that feed back into acquisition, plus the learning loops that make everything compound.
The remaining catalog loops (content-decay, internal-linking, programmatic-SEO quality, content-calendar refill, paid-search query-mining, retargeting-hygiene, landing-page regression, community-engagement, competitor-watch, backlink-prospecting, directory-submission, feature-adoption, lead-capture-asset, email-deliverability, voice-of-customer) slot into the stage that matches their function as each channel becomes a priority.
## Rollout rules
- **One at a time.** Prove a loop earns its keep (someone acts on its output, it moves its metric) before adding the next.
- **Foundation before growth.** Acquisition loops before solid tracking + retention = pouring water into a leaky bucket.
- **Cap the total.** If you're running more loops than you can review the output of, you have vanity loops. Retire the ones nobody acts on.
- **Re-audit quarterly.** Recalibrate thresholds, kill dead loops, promote the ones that consistently drive action.
FILE:references/loop-state.md
# Loop State & Run Logging
Idempotency is only real if the loop can remember what it already did between runs. This reference defines where that state lives and how to log runs — so loops don't double-act, re-nag the same people, or re-alert the same issue.
## Where state lives
Persist each loop's state in a file under `.agents/loops/` — the same `.agents/` convention this repo uses for `product-marketing.md` and `listening-sources.md`. One state file per loop:
```
.agents/loops/<loop-name>.json # the loop's memory
.agents/loops/<loop-name>.log # append-only run log
```
If your scheduler or platform provides its own dedupe/cursor storage, use that instead — the point is durable state, not the specific file. Never keep state only in memory; a loop that forgets on restart will repeat itself.
## What to store
A state file holds whatever the loop needs to not repeat itself:
```json
{
"loop": "churn-signal",
"last_run": "2026-07-01T09:00:00Z",
"cursor": "2026-06-30T23:59:59Z", // watermark — only process items newer than this
"handled": ["acct_1042", "acct_1077"], // dedupe keys already acted on
"cooldowns": { // entity -> next-eligible timestamp
"acct_1042": "2026-07-15T00:00:00Z"
},
"in_flight": ["exp_pricing_v3"], // actions/tests currently open
"counters": { "acct_1042_attempts": 2 } // e.g. dunning/win-back attempt counts
}
```
- **cursor / watermark** — the high-water mark of what's been processed (a timestamp or last ID). The loop only looks at items past it.
- **handled** — dedupe keys for items already acted on, so re-runs skip them.
- **cooldowns** — per-entity suppression windows so you never re-contact someone inside the window.
- **in_flight** — open items (running tests, active interventions) so the loop doesn't start a conflicting one.
- **counters** — attempt counts that drive stop conditions (e.g., "after 2 win-back emails, stop").
Keep state small and prune it: expire old `handled`/`cooldown` entries once they're past their window.
## Idempotency patterns
- **Watermark**: process only items newer than `cursor`; advance `cursor` at the end of a successful run. Safe to re-run — it won't reprocess.
- **Dedupe set**: before acting on an item, check its key against `handled`; add it after acting.
- **Cooldown map**: before contacting an entity, check `cooldowns[entity]`; set it after contact.
- **In-flight guard**: before starting an action that shouldn't overlap (a test, an intervention), check `in_flight`.
## Run logging
Append one line per run, whether or not it acted. This is the audit trail and the vanity-loop detector.
```
2026-07-01T09:00Z checked=312 acted=2 note="2 accounts newly at-risk, interventions staged"
2026-07-02T09:00Z checked=298 acted=0 note="no action"
2026-07-03T09:00Z checked=305 acted=0 note="no action"
```
Log at minimum: timestamp, how many items checked, how many acted on, and a short note. Use it to answer two questions:
- **Is it a vanity loop?** If every run is `acted=0` for weeks and nobody misses it — or it acts every run (a sign it's chasing noise) — reconsider it.
- **Did it double-act?** Two runs acting on the same entity means the dedupe/cooldown state isn't working.
## Resetting & backfilling safely
- To **reset** a loop, clear its `cursor`/`handled` — but keep `cooldowns` so a reset doesn't spam people who were recently contacted.
- On **first run** (no state yet), set the watermark to "now" rather than processing all history, or you'll blast every historical item. If you genuinely want a backfill, do a dry run first (log what it *would* do, act on nothing) and respect cooldowns.
- Never log raw PII in state or run logs — use IDs or hashes (see `loop-guardrails.md`).
FILE:references/loop-template.md
# Loop Template
A copy-paste template for authoring your own marketing loop. Fill every one of the nine parts — a loop missing its **state/idempotency**, **self-check**, or **stop/bail-out** isn't a system, it's a way to do the wrong thing on a schedule.
Before you start, sanity-check that this *should* be a loop at all (see "When NOT to loop" in `SKILL.md`): it's recurring, signal-driven, and doesn't require human judgment to set strategy or creative direction each run.
---
## Blank template (copy this)
```markdown
### The <name> loop
- **Check cadence**: <how often it looks — match to how fast the signal changes, not how often you'd like an update>
- **Acts when**: <the action condition — what must be true to actually DO something vs. just check and skip. Most runs should skip.>
- **Purpose**: <the ONE outcome this loop exists to move>
- **Skills used**: <which marketing skills the loop orchestrates each run>
- **Loop body**:
1. <step — usually: pull data / diff vs. last run>
2. <step — identify what, if anything, crossed the action condition>
3. <step — draft or stage the response>
- **Self-check**: <the verification done BEFORE acting — is the signal real vs. noise/seasonality/tracking bug? Is the sample big enough to be significant?>
- **State / idempotency**: <what it remembers between runs — last-run marker, dedupe key, cooldown window, "already handled" set — so it doesn't double-act or re-nag the same people>
- **Stop / bail-out**: <when it skips, halts, escalates to a human, or disables itself — plus what it does on error. Include a human checkpoint before anything that spends money or publishes.>
- **Output**: <where results go — a file, a PR, a staged draft, a notification, a report>
```
---
## Fill-in prompts (answer these, in order)
1. **What outcome does this protect or grow?** (rankings, ad efficiency, activation, retention, revenue, referrals) → *Purpose*
2. **How fast does that signal actually change?** (hours / days / weeks / months) → *Check cadence*
3. **What has to be true before it's worth acting?** (a threshold crossed, a new item appeared, a regression vs. baseline) → *Acts when*
4. **What data does it read and what does it produce each run?** → *Loop body* + *Output*
5. **What would make it act on a false signal?** (noise, seasonality, a tracking break, too-small a sample) → *Self-check*
6. **What must it remember so it doesn't repeat itself?** (dedupe key, cooldown, last-run marker) → *State / idempotency*
7. **When should it stop, skip, or hand off to a human?** (no action needed, error, spend/publish decision, N failed attempts) → *Stop / bail-out*
If you can't answer 5, 6, and 7 concretely, the loop isn't ready to run.
---
## Worked example (blank → filled)
Say you sell a freemium API tool and want to stop losing signups who never make their first API call.
```markdown
### The first-call activation loop
- **Check cadence**: Daily
- **Acts when**: A user who signed up 48h ago still hasn't made a successful API call and isn't already in this nudge sequence.
- **Purpose**: Increase the share of new signups that reach first value (first successful API call).
- **Skills used**: `onboarding`, `emails`, `analytics`
- **Loop body**:
1. Pull signups from ~48h ago and their first-call status.
2. Filter to those with zero successful calls and no active nudge.
3. Draft a targeted "get your first call working" email (docs link, common blocker, offer to help).
- **Self-check**: Is "no call" a real activation gap, or a tracking gap (calls firing but not logged)? Confirm against server logs before emailing.
- **State / idempotency**: Track which users have entered this sequence; suppress anyone who has made a call since; one nudge per user per stage.
- **Stop / bail-out**: After 2 nudges with no call, stop and route to the broader re-engagement loop — don't keep emailing. Skip the run entirely if the events pipeline looks stale.
- **Output**: A staged activation email per qualifying user + a daily count of new activations.
```
Notice what makes it safe: the **self-check** guards against a tracking bug emailing active users, the **state** stops it re-nagging, and the **stop** caps attempts and hands off instead of looping forever.
---
## Ship checklist
Before you schedule a new loop, confirm:
- [ ] All nine parts are filled — especially self-check, state, and stop.
- [ ] Cadence matches signal speed (you're not checking daily for a weekly-moving signal).
- [ ] It's designed so **most runs do nothing** — it acts only on a real condition.
- [ ] Anything that **spends money or publishes** has a human checkpoint (unless caps + an allowlist are explicitly authorized).
- [ ] State prevents double-acting and re-nagging the same people.
- [ ] There's an error path (stale data → report "stale," don't fabricate movement) and a manual off switch.
- [ ] For scheduling mechanics, see the "Scheduling a loop" section in `SKILL.md`.
Once it runs, give it a few cycles and ask the "is this a vanity loop?" question: if nobody acts on the output, delete it.
Áp dụng nguyên lý tâm lý, mô hình tư duy và khoa học hành vi vào marketing: thiên kiến nhận thức, thuyết phục, hành vi người tiêu dùng.
---
name: marketing-psychology
description: "When the user wants to apply psychological principles, mental models, or behavioral science to marketing. Also use when the user mentions 'psychology,' 'mental models,' 'cognitive bias,' 'persuasion,' 'behavioral science,' 'why people buy,' 'decision-making,' 'consumer behavior,' 'anchoring,' 'social proof,' 'scarcity,' 'loss aversion,' 'framing,' or 'nudge.' Use this whenever someone wants to understand or leverage how people think and make decisions in a marketing context. For applying psychology to specific pages, see cro; for pricing tactics, see pricing; for copy framing, see copywriting."
metadata:
version: 2.0.0
---
# Marketing Psychology & Mental Models
You are an expert in applying psychological principles and mental models to marketing. Your goal is to help users understand why people buy, how to influence behavior ethically, and how to make better marketing decisions.
## How to Use This Skill
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before applying mental models. Use that context to tailor recommendations to the specific product and audience.
Mental models are thinking tools that help you make better decisions, understand customer behavior, and create more effective marketing. When helping users:
1. Identify which mental models apply to their situation
2. Explain the psychology behind the model
3. Provide specific marketing applications
4. Suggest how to implement ethically
---
## Foundational Thinking Models
These models sharpen your strategy and help you solve the right problems.
### First Principles
Break problems down to basic truths and build solutions from there. Instead of copying competitors, ask "why" repeatedly to find root causes. Use the 5 Whys technique to tunnel down to what really matters.
**Marketing application**: Don't assume you need content marketing because competitors do. Ask why you need it, what problem it solves, and whether there's a better solution.
### Jobs to Be Done
People don't buy products—they "hire" them to get a job done. Focus on the outcome customers want, not features.
**Marketing application**: A drill buyer doesn't want a drill—they want a hole. Frame your product around the job it accomplishes, not its specifications.
### Circle of Competence
Know what you're good at and stay within it. Venture outside only with proper learning or expert help.
**Marketing application**: Don't chase every channel. Double down where you have genuine expertise and competitive advantage.
### Inversion
Instead of asking "How do I succeed?", ask "What would guarantee failure?" Then avoid those things.
**Marketing application**: List everything that would make your campaign fail—confusing messaging, wrong audience, slow landing page—then systematically prevent each.
### Occam's Razor
The simplest explanation is usually correct. Avoid overcomplicating strategies or attributing results to complex causes when simple ones suffice.
**Marketing application**: If conversions dropped, check the obvious first (broken form, page speed) before assuming complex attribution issues.
### Pareto Principle (80/20 Rule)
Roughly 80% of results come from 20% of efforts. Identify and focus on the vital few.
**Marketing application**: Find the 20% of channels, customers, or content driving 80% of results. Cut or reduce the rest.
### Local vs. Global Optima
A local optimum is the best solution nearby, but a global optimum is the best overall. Don't get stuck optimizing the wrong thing.
**Marketing application**: Optimizing email subject lines (local) won't help if email isn't the right channel (global). Zoom out before zooming in.
### Theory of Constraints
Every system has one bottleneck limiting throughput. Find and fix that constraint before optimizing elsewhere.
**Marketing application**: If your funnel converts well but traffic is low, more conversion optimization won't help. Fix the traffic bottleneck first.
### Opportunity Cost
Every choice has a cost—what you give up by not choosing alternatives. Consider what you're saying no to.
**Marketing application**: Time spent on a low-ROI channel is time not spent on high-ROI activities. Always compare against alternatives.
### Law of Diminishing Returns
After a point, additional investment yields progressively smaller gains.
**Marketing application**: The 10th blog post won't have the same impact as the first. Know when to diversify rather than double down.
### Second-Order Thinking
Consider not just immediate effects, but the effects of those effects.
**Marketing application**: A flash sale boosts revenue (first order) but may train customers to wait for discounts (second order).
### Map ≠ Territory
Models and data represent reality but aren't reality itself. Don't confuse your analytics dashboard with actual customer experience.
**Marketing application**: Your customer persona is a useful model, but real customers are more complex. Stay in touch with actual users.
### Probabilistic Thinking
Think in probabilities, not certainties. Estimate likelihoods and plan for multiple outcomes.
**Marketing application**: Don't bet everything on one campaign. Spread risk and plan for scenarios where your primary strategy underperforms.
### Barbell Strategy
Combine extreme safety with small high-risk/high-reward bets. Avoid the mediocre middle.
**Marketing application**: Put 80% of budget into proven channels, 20% into experimental bets. Avoid moderate-risk, moderate-reward middle.
---
## Understanding Buyers & Human Psychology
These models explain how customers think, decide, and behave.
### Fundamental Attribution Error
People attribute others' behavior to character, not circumstances. "They didn't buy because they're not serious" vs. "The checkout was confusing."
**Marketing application**: When customers don't convert, examine your process before blaming them. The problem is usually situational, not personal.
### Mere Exposure Effect
People prefer things they've seen before. Familiarity breeds liking.
**Marketing application**: Consistent brand presence builds preference over time. Repetition across channels creates comfort and trust.
### Availability Heuristic
People judge likelihood by how easily examples come to mind. Recent or vivid events seem more common.
**Marketing application**: Case studies and testimonials make success feel more achievable. Make positive outcomes easy to imagine.
### Confirmation Bias
People seek information confirming existing beliefs and ignore contradictory evidence.
**Marketing application**: Understand what your audience already believes and align messaging accordingly. Fighting beliefs head-on rarely works.
### The Lindy Effect
The longer something has survived, the longer it's likely to continue. Old ideas often outlast new ones.
**Marketing application**: Proven marketing principles (clear value props, social proof) outlast trendy tactics. Don't abandon fundamentals for fads.
### Mimetic Desire
People want things because others want them. Desire is socially contagious.
**Marketing application**: Show that desirable people want your product. Waitlists, exclusivity, and social proof trigger mimetic desire.
### Sunk Cost Fallacy
People continue investing in something because of past investment, even when it's no longer rational.
**Marketing application**: Know when to kill underperforming campaigns. Past spend shouldn't justify future spend if results aren't there.
### Endowment Effect
People value things more once they own them.
**Marketing application**: Free trials, samples, and freemium models let customers "own" the product, making them reluctant to give it up.
### IKEA Effect
People value things more when they've put effort into creating them.
**Marketing application**: Let customers customize, configure, or build something. Their investment increases perceived value and commitment.
### Zero-Price Effect
Free isn't just a low price—it's psychologically different. "Free" triggers irrational preference.
**Marketing application**: Free tiers, free trials, and free shipping have disproportionate appeal. The jump from $1 to $0 is bigger than $2 to $1.
### Hyperbolic Discounting / Present Bias
People strongly prefer immediate rewards over future ones, even when waiting is more rational.
**Marketing application**: Emphasize immediate benefits ("Start saving time today") over future ones ("You'll see ROI in 6 months").
### Status-Quo Bias
People prefer the current state of affairs. Change requires effort and feels risky.
**Marketing application**: Reduce friction to switch. Make the transition feel safe and easy. "Import your data in one click."
### Default Effect
People tend to accept pre-selected options. Defaults are powerful.
**Marketing application**: Pre-select the plan you want customers to choose. Opt-out beats opt-in for subscriptions (ethically applied).
### Paradox of Choice
Too many options overwhelm and paralyze. Fewer choices often lead to more decisions.
**Marketing application**: Limit options. Three pricing tiers beat seven. Recommend a single "best for most" option.
### Goal-Gradient Effect
People accelerate effort as they approach a goal. Progress visualization motivates action.
**Marketing application**: Show progress bars, completion percentages, and "almost there" messaging to drive completion.
### Peak-End Rule
People judge experiences by the peak (best or worst moment) and the end, not the average.
**Marketing application**: Design memorable peaks (surprise upgrades, delightful moments) and strong endings (thank you pages, follow-up emails).
### Zeigarnik Effect
Unfinished tasks occupy the mind more than completed ones. Open loops create tension.
**Marketing application**: "You're 80% done" creates pull to finish. Incomplete profiles, abandoned carts, and cliffhangers leverage this.
### Pratfall Effect
Competent people become more likable when they show a small flaw. Perfection is less relatable.
**Marketing application**: Admitting a weakness ("We're not the cheapest, but...") can increase trust and differentiation.
### Curse of Knowledge
Once you know something, you can't imagine not knowing it. Experts struggle to explain simply.
**Marketing application**: Your product seems obvious to you but confusing to newcomers. Test copy with people unfamiliar with your space.
### Mental Accounting
People treat money differently based on its source or intended use, even though money is fungible.
**Marketing application**: Frame costs in favorable mental accounts. "$3/day" feels different than "$90/month" even though it's the same.
### Regret Aversion
People avoid actions that might cause regret, even if the expected outcome is positive.
**Marketing application**: Address regret directly. Money-back guarantees, free trials, and "no commitment" messaging reduce regret fear.
### Bandwagon Effect / Social Proof
People follow what others are doing. Popularity signals quality and safety.
**Marketing application**: Show customer counts, testimonials, logos, reviews, and "trending" indicators. Numbers create confidence.
---
## Influencing Behavior & Persuasion
These models help you ethically influence customer decisions.
### Reciprocity Principle
People feel obligated to return favors. Give first, and people want to give back.
**Marketing application**: Free content, free tools, and generous free tiers create reciprocal obligation. Give value before asking for anything.
### Commitment & Consistency
Once people commit to something, they want to stay consistent with that commitment.
**Marketing application**: Get small commitments first (email signup, free trial). People who've taken one step are more likely to take the next.
### Authority Bias
People defer to experts and authority figures. Credentials and expertise create trust.
**Marketing application**: Feature expert endorsements, certifications, "featured in" logos, and thought leadership content.
### Liking / Similarity Bias
People say yes to those they like and those similar to themselves.
**Marketing application**: Use relatable spokespeople, founder stories, and community language. "Built by marketers for marketers" signals similarity.
### Unity Principle
Shared identity drives influence. "One of us" is powerful.
**Marketing application**: Position your brand as part of the customer's tribe. Use insider language and shared values.
### Scarcity / Urgency Heuristic
Limited availability increases perceived value. Scarcity signals desirability.
**Marketing application**: Limited-time offers, low-stock warnings, and exclusive access create urgency. Only use when genuine.
### Foot-in-the-Door Technique
Start with a small request, then escalate. Compliance with small requests leads to compliance with larger ones.
**Marketing application**: Free trial → paid plan → annual plan → enterprise. Each step builds on the last.
### Door-in-the-Face Technique
Start with an unreasonably large request, then retreat to what you actually want. The contrast makes the second request seem reasonable.
**Marketing application**: Show enterprise pricing first, then reveal the affordable starter plan. The contrast makes it feel like a deal.
### Loss Aversion / Prospect Theory
Losses feel roughly twice as painful as equivalent gains feel good. People will work harder to avoid losing than to gain.
**Marketing application**: Frame in terms of what they'll lose by not acting. "Don't miss out" beats "You could gain."
### Anchoring Effect
The first number people see heavily influences subsequent judgments.
**Marketing application**: Show the higher price first (original price, competitor price, enterprise tier) to anchor expectations.
### Decoy Effect
Adding a third, inferior option makes one of the original two look better.
**Marketing application**: A "decoy" pricing tier that's clearly worse value makes your preferred tier look like the obvious choice.
### Framing Effect
How something is presented changes how it's perceived. Same facts, different frames.
**Marketing application**: "90% success rate" vs. "10% failure rate" are identical but feel different. Frame positively.
### Contrast Effect
Things seem different depending on what they're compared to.
**Marketing application**: Show the "before" state clearly. The contrast with your "after" makes improvements vivid.
---
## Pricing Psychology
These models specifically address how people perceive and respond to prices.
### Charm Pricing / Left-Digit Effect
Prices ending in 9 seem significantly lower than the next round number. $99 feels much cheaper than $100.
**Marketing application**: Use .99 or .95 endings for value-focused products. The left digit dominates perception.
### Rounded-Price (Fluency) Effect
Round numbers feel premium and are easier to process. $100 signals quality; $99 signals value.
**Marketing application**: Use round prices for premium products ($500/month), charm prices for value products ($497/month).
### Rule of 100
For prices under $100, percentage discounts seem larger ("20% off"). For prices over $100, absolute discounts seem larger ("$50 off").
**Marketing application**: $80 product: "20% off" beats "$16 off." $500 product: "$100 off" beats "20% off."
### Price Relativity / Good-Better-Best
People judge prices relative to options presented. A middle tier seems reasonable between cheap and expensive.
**Marketing application**: Three tiers where the middle is your target. The expensive tier makes it look reasonable; the cheap tier provides an anchor.
### Mental Accounting (Pricing)
Framing the same price differently changes perception.
**Marketing application**: "$1/day" feels cheaper than "$30/month." "Less than your morning coffee" reframes the expense.
---
## Design & Delivery Models
These models help you design effective marketing systems.
### Hick's Law
Decision time increases with the number and complexity of choices. More options = slower decisions = more abandonment.
**Marketing application**: Simplify choices. One clear CTA beats three. Fewer form fields beat more.
### AIDA Funnel
Attention → Interest → Desire → Action. The classic customer journey model.
**Marketing application**: Structure pages and campaigns to move through each stage. Capture attention before building desire.
### Rule of 7
Prospects need roughly 7 touchpoints before converting. One ad rarely converts; sustained presence does.
**Marketing application**: Build multi-touch campaigns across channels. Retargeting, email sequences, and consistent presence compound.
### Nudge Theory / Choice Architecture
Small changes in how choices are presented significantly influence decisions.
**Marketing application**: Default selections, strategic ordering, and friction reduction guide behavior without restricting choice.
### BJ Fogg Behavior Model
Behavior = Motivation × Ability × Prompt. All three must be present for action.
**Marketing application**: High motivation but hard to do = won't happen. Easy to do but no prompt = won't happen. Design for all three.
### EAST Framework
Make desired behaviors: Easy, Attractive, Social, Timely.
**Marketing application**: Reduce friction (easy), make it appealing (attractive), show others doing it (social), ask at the right moment (timely).
### COM-B Model
Behavior requires: Capability, Opportunity, Motivation.
**Marketing application**: Can they do it (capability)? Is the path clear (opportunity)? Do they want to (motivation)? Address all three.
### Activation Energy
The initial energy required to start something. High activation energy prevents action even if the task is easy overall.
**Marketing application**: Reduce starting friction. Pre-fill forms, offer templates, show quick wins. Make the first step trivially easy.
### North Star Metric
One metric that best captures the value you deliver to customers. Focus creates alignment.
**Marketing application**: Identify your North Star (active users, completed projects, revenue per customer) and align all efforts toward it.
### The Cobra Effect
When incentives backfire and produce the opposite of intended results.
**Marketing application**: Test incentive structures. A referral bonus might attract low-quality referrals gaming the system.
---
## Growth & Scaling Models
These models explain how marketing compounds and scales.
### Feedback Loops
Output becomes input, creating cycles. Positive loops accelerate growth; negative loops create decline.
**Marketing application**: Build virtuous cycles: more users → more content → better SEO → more users. Identify and strengthen positive loops.
### Compounding
Small, consistent gains accumulate into large results over time. Early gains matter most.
**Marketing application**: Consistent content, SEO, and brand building compound. Start early; benefits accumulate exponentially.
### Network Effects
A product becomes more valuable as more people use it.
**Marketing application**: Design features that improve with more users: shared workspaces, integrations, marketplaces, communities.
### Flywheel Effect
Sustained effort creates momentum that eventually maintains itself. Hard to start, easy to maintain.
**Marketing application**: Content → traffic → leads → customers → case studies → more content. Each element powers the next.
### Switching Costs
The price (time, money, effort, data) of changing to a competitor. High switching costs create retention.
**Marketing application**: Increase switching costs ethically: integrations, data accumulation, workflow customization, team adoption.
### Exploration vs. Exploitation
Balance trying new things (exploration) with optimizing what works (exploitation).
**Marketing application**: Don't abandon working channels for shiny new ones, but allocate some budget to experiments.
### Critical Mass / Tipping Point
The threshold after which growth becomes self-sustaining.
**Marketing application**: Focus resources on reaching critical mass in one segment before expanding. Depth before breadth.
### Survivorship Bias
Focusing on successes while ignoring failures that aren't visible.
**Marketing application**: Study failed campaigns, not just successful ones. The viral hit you're copying had 99 failures you didn't see.
---
## Quick Reference
When facing a marketing challenge, consider:
| Challenge | Relevant Models |
|-----------|-----------------|
| Low conversions | Hick's Law, Activation Energy, BJ Fogg, Friction |
| Price objections | Anchoring, Framing, Mental Accounting, Loss Aversion |
| Building trust | Authority, Social Proof, Reciprocity, Pratfall Effect |
| Increasing urgency | Scarcity, Loss Aversion, Zeigarnik Effect |
| Retention/churn | Endowment Effect, Switching Costs, Status-Quo Bias |
| Growth stalling | Theory of Constraints, Local vs Global Optima, Compounding |
| Decision paralysis | Paradox of Choice, Default Effect, Nudge Theory |
| Onboarding | Goal-Gradient, IKEA Effect, Commitment & Consistency |
---
## Task-Specific Questions
1. What specific behavior are you trying to influence?
2. What does your customer believe before encountering your marketing?
3. Where in the journey (awareness → consideration → decision) is this?
4. What's currently preventing the desired action?
5. Have you tested this with real customers?
---
## Related Skills
- **cro**: Apply psychology to page optimization
- **copywriting**: Write copy using psychological principles
- **popups**: Use triggers and psychology in popups
- **pricing-page optimization**: See cro for pricing psychology
- **ab-testing**: Test psychological hypotheses
FILE:evals/evals.json
{
"skill_name": "marketing-psychology",
"evals": [
{
"id": 1,
"prompt": "How can I use psychology to increase conversions on our pricing page? We sell a B2B SaaS tool with three tiers ($29, $79, $199/month).",
"expected_output": "Should check for product-marketing.md first. Should apply relevant pricing psychology models: anchoring (show the highest plan first or use a decoy), charm pricing (consider $29 vs $30), Rule of 100 (percentage vs dollar discounts), Good-Better-Best framing, loss aversion (show what they miss on lower tiers). Should also apply broader persuasion models: social proof near pricing, scarcity for limited-time offers, default effect (pre-select recommended plan). Should provide specific, actionable recommendations tied to their price points.",
"assertions": [
"Checks for product-marketing.md",
"Applies pricing psychology models (anchoring, charm pricing, Rule of 100)",
"Applies Good-Better-Best framing",
"Applies loss aversion to tier differentiation",
"Applies social proof near pricing",
"Provides specific recommendations for their price points",
"References specific mental models by name"
],
"files": []
},
{
"id": 2,
"prompt": "Explain the scarcity principle and how to use it ethically in SaaS marketing without being manipulative.",
"expected_output": "Should explain scarcity as a mental model (limited availability increases perceived value). Should provide legitimate SaaS applications: limited beta spots, early-bird pricing with real deadlines, limited-time feature access, cohort-based launches. Should distinguish ethical scarcity (real constraints) from manufactured urgency (fake countdown timers, artificial limits). Should provide specific examples and implementation guidance. Should reference related models (urgency, FOMO, loss aversion).",
"assertions": [
"Explains scarcity principle clearly",
"Provides legitimate SaaS applications",
"Distinguishes ethical from manipulative use",
"Provides specific examples",
"References related mental models",
"Addresses ethical considerations directly"
],
"files": []
},
{
"id": 3,
"prompt": "what psychological principles should I use to write better marketing copy?",
"expected_output": "Should trigger on casual phrasing. Should recommend copy-relevant mental models from the skill's taxonomy: social proof, reciprocity, loss aversion, anchoring, scarcity, IKEA Effect, Endowment Effect, Commitment & Consistency. For each principle, should explain what it is and provide a specific copywriting application. Should reference the quick reference table by challenge. Should organize by where in the copy each principle applies (headlines, body, CTAs, testimonials).",
"assertions": [
"Triggers on casual phrasing",
"Recommends copy-relevant mental models",
"Explains each principle briefly",
"Provides specific copywriting application per principle",
"Organizes by where each applies in copy",
"References multiple model categories"
],
"files": []
},
{
"id": 4,
"prompt": "I'm designing an onboarding flow and want to use behavioral psychology to increase activation. What models should I apply?",
"expected_output": "Should apply design and behavioral models from the skill's taxonomy: Goal-Gradient Effect (motivation increases near goal), Hick's Law (reduce choices), IKEA Effect (let users build something), Endowment Effect (let them experience ownership), Zeigarnik Effect (incomplete tasks drive completion), Commitment & Consistency (small asks first). Should explain how each applies to onboarding specifically. Should provide actionable recommendations for each model.",
"assertions": [
"Applies Goal-Gradient Effect",
"Applies Hick's Law",
"Applies IKEA Effect or Endowment Effect",
"Applies Zeigarnik Effect or commitment principles",
"Explains how each applies to onboarding",
"Provides actionable recommendations per model"
],
"files": []
},
{
"id": 5,
"prompt": "What's the psychology behind why free trials work better than freemium for some products?",
"expected_output": "Should apply relevant mental models: loss aversion (trial users fear losing access), endowment effect (they feel ownership after using), sunk cost (time invested during trial), Zero-Price Effect (free removes psychological barrier to start), status quo bias (inertia to keep what they have). Should explain how these models interact in trial vs freemium contexts. Should note when each model works best (trial for products with high activation effort, freemium for products with network effects).",
"assertions": [
"Applies loss aversion to trial context",
"Applies endowment effect",
"Applies Zero-Price Effect",
"Explains how models interact in trial vs freemium",
"Notes when each approach works best",
"Provides clear, educational explanation"
],
"files": []
},
{
"id": 6,
"prompt": "Help me run an A/B test on which psychological principle works better for our CTA — scarcity vs social proof.",
"expected_output": "Should recognize this is an A/B test setup task, not a psychology task. Should defer to or cross-reference the ab-testing skill for the experiment design. May provide psychological context on both principles to inform the hypothesis, but should make clear that ab-testing is the right skill for designing and running the experiment.",
"assertions": [
"Recognizes this as an A/B test setup task",
"References or defers to ab-testing skill",
"May provide psychological context for hypothesis",
"Does not attempt full test design using psychology patterns"
],
"files": []
}
]
}
Chuyển markdown dài (spec, RFC, báo cáo, kế hoạch) thành tài liệu HTML một file có mục lục, tìm kiếm, nút sao chép code và token thương hiệu.
---
name: md-document
description: Converts long-form markdown (specs, RFCs, reports, plans, explainers) into a single-file, lightly-interactive HTML document with sticky TOC, scrollspy, search filter, code-copy buttons, and design-system-driven brand tokens. Triggers when the markdown-html-orchestrator classifies an input as DOCUMENT, or when invoked directly via /cs:md-document. Reads the design-system config via config_loader.py and inlines the user's 12 derived CSS custom properties; refuses to render if onboarding hasn't run. Single-file output — Google Fonts + Prism.js CDN are the only externals; no framework runtime, no build step. Use after orchestrator routing or after design-system onboarding is confirmed.
version: 2.10.1
author: Alireza Rezvani
license: MIT
tags: [markdown, html, documentation, single-file, toc, scrollspy, search, code-copy, design-system]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# md-document — Long-form Markdown to HTML
The general-purpose converter — handles the 90% case Shihipar describes (specs, plans, RFCs, reports, explainers). Three stdlib tools pipeline together:
```
markdown_parser.py → html_renderer.py → interactivity_injector.py
(md → JSON AST) (AST + tokens → HTML) (HTML + JS behavior)
```
Output is one `.html` file with sticky TOC, search filter, scrollspy, code-copy buttons, and the user's 12 derived brand tokens. Externals limited to Google Fonts CSS + Prism.js CDN.
## When to invoke
| Symptom | Action |
|---|---|
| `markdown-html-orchestrator` routes input as DOCUMENT | Invoke this skill |
| User runs `/cs:md-document <path>.md` directly | Invoke this skill |
| User says "convert this spec/report/RFC/plan to HTML" | Invoke this skill |
| Input is a code review (has ` ```diff ` blocks) | Route to `md-review` instead |
| Input is a slide deck (clear `---` boundaries) | Route to `md-slides` instead |
| Input is < 100 lines | Refuse (Shihipar threshold — markdown still wins) |
| Design-system not onboarded | Refuse, surface `/cs:design-system` |
## Pipeline
```bash
# 1. Parse markdown → JSON AST
python3 markdown-html/skills/md-document/scripts/markdown_parser.py \
--input <path>.md --output sections.json
# 2. Render AST + design-system config → single-file HTML
python3 markdown-html/skills/md-document/scripts/html_renderer.py \
--sections sections.json --output document.html
# 3. Inject lightweight JS (search, copycode, smoothscroll, scrollspy)
python3 markdown-html/skills/md-document/scripts/interactivity_injector.py \
--file document.html \
--features search,copycode,smoothscroll,scrollspy
```
Or all-in-one (sample render):
```bash
python3 markdown-html/skills/md-document/scripts/html_renderer.py --sample \
| python3 markdown-html/skills/md-document/scripts/interactivity_injector.py \
--file /dev/stdin --output document.html
```
## What gets rendered
CommonMark subset sufficient for agent-generated artifacts:
- Headings H1-H6 (every H2+ gets an anchor id and TOC entry)
- Paragraphs with inline **bold** / *italic* / `code` / [links](url) / 
- Fenced code blocks (` ```python `) with Prism.js highlighting on demand
- GFM tables with per-column alignment
- GFM callouts (`> [!NOTE]`, `> [!TIP]`, `> [!IMPORTANT]`, `> [!WARNING]`, `> [!CAUTION]`)
- Blockquotes, ordered + unordered lists (single-level), horizontal rules
Out of scope: nested lists, HTML inlines, footnotes, definition lists, task list checkboxes (rendered as plain text), reference-style links.
## Hard rules
1. **Refuses input < 100 lines.** Markdown wins below the threshold (Shihipar).
2. **Refuses without onboarding.** `config_loader.setup_completed()` must return `True`. Otherwise surface `/cs:design-system`.
3. **Single-file output.** All CSS + JS inline. Only externals are `fonts.googleapis.com` and `cdn.jsdelivr.net` (Prism). Anything else is a regression.
4. **Customization must change behavior.** `design_style=editorial` produces 720px-wide layout with 1.75 line-height; `playful` rounds the callouts and adds shadow; `technical` is dense with 0.875rem code. Smoke-tested.
5. **WCAG-compliant tokens.** Inherits the design-system's WCAG AA palette — body text ≥ 4.5:1 contrast, links iteratively walked to 4.5:1.
6. **Idempotent injection.** Re-injecting interactivity is a no-op (marker check). Re-rendering with a different design_style works cleanly.
## Forcing-question library (Matt Pocock grill discipline)
1. **What's the document for — skim, decide, or deep-read?** Recommended: name it; density follows. Canon: Shihipar; Tufte *Envisioning Information*.
2. **Sticky-sidebar TOC or collapsible-top?** Recommended: sticky-sidebar for > 800 words / 4+ H2s; collapsible-top for shorter mobile-first docs. Canon: NN/g *TOC Best Practices* (2023).
3. **All four interactive features, or a subset?** Recommended: all four — none of them cost more than ~1 KB. Canon: Wattenberger *Why React isn't great for actually building websites*.
4. **Code theme — light, dark, or auto?** Recommended: auto (follows OS `prefers-color-scheme`). Canon: WCAG 2.2 §1.4.3.
5. **Does the document have a clear H1 title?** Recommended: yes — H1 becomes the page `<title>` and is excluded from the TOC.
## Distinct from
- **`md-review`** — that converter renders diff blocks + severity-tagged margin annotations. This one renders prose + tables + code + callouts.
- **`md-slides`** — that converter splits on `---` boundaries into slides. This one renders one continuous document.
- **`marketing/landing/`** — that generates landing pages from scratch (no markdown input). This converts existing markdown.
## Output artifact
`{default_output_dir}/doc-{slug}.html` (path resolved by orchestrator's `output_path_resolver.py`; collision suffix `-2`, `-3`, … by default).
## References
- Shihipar — *Claude Code HTML output* (Medium, 2026)
- Tufte — *Envisioning Information* (1990), ch. 2 "Micro/Macro Readings"
- NN/g — *Table of Contents Best Practices* (2023)
- WCAG 2.2 — §1.4.3 contrast, §2.4.5 multiple ways
- Wattenberger — *Why React isn't great for actually building websites*
- See `references/` for full citations
FILE:assets/md_document_template.html
<!DOCTYPE html>
<!--
md_document_template.html — Reference shape for html_renderer.py output.
This file documents the canonical output structure. The renderer generates
this same shape dynamically from a section AST + design-system config.
Token slots ({{TITLE}}, {{PALETTE}}, etc.) are illustrative — the actual
renderer interpolates Python values directly into the HTML string.
See: html_renderer.py for the live implementation.
-->
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>{{TITLE}}</title>
<!-- Google Fonts CDN — the user's heading + body Google Fonts -->
<link rel="preconnect" href="https://fonts.googleapis.com" crossorigin>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family={{HEADING_FONT}}:wght@400;600&family={{BODY_FONT}}:wght@400;600&display=swap">
<!-- Prism.js CDN — auto-loads per-language plugins on demand -->
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/themes/prism.min.css">
<style>
:root {
/* 12 CSS custom properties derived from the user's brand by
brand_palette_validator.derive_palette() */
--md-bg: {{BG}};
--md-surface: {{SURFACE}};
--md-border: {{BORDER}};
--md-text: {{TEXT}};
--md-text-muted: {{TEXT_MUTED}};
--md-accent: {{ACCENT}};
--md-accent-soft: {{ACCENT_SOFT}};
--md-code-bg: {{CODE_BG}};
--md-link: {{LINK}};
--md-link-hover: {{LINK_HOVER}};
--md-success: {{SUCCESS}};
--md-warn: {{WARN}};
--md-scale: {{SCALE}}; /* e.g. 1.25 */
--md-font-heading: '{{HEADING_FONT}}', system-ui, sans-serif;
--md-font-body: '{{BODY_FONT}}', system-ui, sans-serif;
}
/* ... BASE_CSS ... STYLE_CSS_OVERRIDES[design_style] ... */
</style>
</head>
<!-- Body class: style-{editorial|technical|minimal|playful} and toc-{behavior} -->
<body class="style-{{DESIGN_STYLE}} toc-{{TOC_BEHAVIOR}}">
<!-- TOC (variants: sidebar / collapsible-top / inline / none) -->
<nav class="toc" aria-label="Table of contents">
<ol>
<li><a href="#first-section">First Section</a></li>
<!-- ... -->
</ol>
</nav>
<main>
<!-- Search bar (sticky; hidden when search feature is not injected) -->
<div class="md-search">
<input type="search" id="md-search-input"
placeholder="Filter sections… (Esc to clear)"
aria-label="Filter document sections">
</div>
<!-- Rendered blocks from the section AST -->
<h1>{{TITLE}}</h1>
<h2 id="first-section">First Section</h2>
<p>Paragraph with <strong>bold</strong>, <em>italic</em>, <code>inline code</code>, and <a href="#">link</a>.</p>
<aside class="callout callout-note" role="note">
<div class="callout-label"><span class="callout-icon" aria-hidden="true">i</span>NOTE</div>
<div class="callout-body">Important contextual information.</div>
</aside>
<pre><button class="code-copy" type="button" aria-label="Copy code">Copy</button><code class="language-python">def hello():
return "world"</code></pre>
<table>
<thead><tr><th>Header A</th><th>Header B</th></tr></thead>
<tbody><tr><td>Cell 1</td><td>Cell 2</td></tr></tbody>
</table>
<!-- Footer with company_name + logo (base64-embedded) -->
<footer class="md-footer">
<img src="data:image/png;base64,...{{LOGO_BASE64}}" alt="{{COMPANY_NAME}}">
<span>{{COMPANY_NAME}}</span>
<span style="margin-left:auto">Generated by markdown-html</span>
</footer>
</main>
<!-- Prism.js (deferred; doesn't block first paint) -->
<script defer src="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/components/prism-core.min.js"></script>
<script defer src="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/plugins/autoloader/prism-autoloader.min.js"></script>
<!-- Interactivity script injected by interactivity_injector.py
when features are enabled. Marked with id="md-document-interactivity-v1"
for idempotency. Contains:
- search filter on H2 sections
- code-copy button handlers (navigator.clipboard + execCommand fallback)
- smooth-scroll for TOC anchors
- scrollspy via IntersectionObserver (sets aria-current on TOC links) -->
</body>
</html>
FILE:references/information_density_patterns.md
# Information Density Patterns for Long-form Documents
**Why this exists:** The `md-document` converter renders long-form markdown (specs, RFCs, reports, explainers) — typically 100-2000 lines of prose, code, tables, and callouts. Past 100 lines the linear flow loses orientation. This document codifies the patterns that restore it.
## The four density patterns
### 1. Hierarchy made visible
Linear markdown shows hierarchy through indented `#` characters. HTML shows hierarchy through typography scale, color, weight, spacing, and surface. The renderer uses a modular type scale (`typography.scale_ratio`, default 1.25 = major third) so each heading level is visibly proportional. H2 sections get a hairline `border-bottom` for visual chunking. Callouts get a 4px accent border that signals "stop and read."
### 2. Lateral navigation
Linear reading is one channel — top to bottom. The renderer adds:
- **Sticky-sidebar TOC** (default) — always visible, jumps to any H2/H3 in one click.
- **Scrollspy** — the current section's TOC entry gets `aria-current="location"` as the reader scrolls, so they always know where they are.
- **Anchored headings** — every H2-H6 gets an `id` derived from the heading text, so deep links work without further effort.
- **Smooth scroll** — TOC clicks animate, not jump, so the reader keeps spatial context.
### 3. Lateral structure
Side-by-side comparison is impossible in linear markdown. HTML provides:
- **Tables** — rendered with `<table>`, semantic `<thead>/<tbody>`, per-column alignment from the GFM delimiter row.
- **Collapsible sections** — `<details>` blocks for content the reader can skip on first pass. (TOC variant `collapsible-top` uses this for the TOC itself.)
- **Callouts** — `<aside class="callout">` for NOTE/TIP/IMPORTANT/WARNING/CAUTION. Distinct from paragraphs because they interrupt flow with intent.
### 4. Lateral interaction
Lightweight, no-framework:
- **Search filter** — `<input type="search">` filters H2 sections by heading + body text. Vanilla JS, no debouncing needed because typical documents have under 30 H2 sections.
- **Code-copy buttons** — appear on hover over `<pre>`, copy the entire `<code>` text. `navigator.clipboard` with `document.execCommand` fallback.
- **Smooth scroll** — already covered above.
## What's deliberately excluded
- **Slider/knob controls** — that's Anthropic's official Playground plugin's lane.
- **Real-time collaboration** — documents are read artifacts, not edit surfaces.
- **Multi-page navigation** — single-file is the discipline (see `single_file_html_discipline.md`).
- **Dark mode toggle** — the user picked `code_theme` once; switching mid-document violates the shipped-as-onboarded contract. (`code_theme: auto` does follow `prefers-color-scheme` for syntax highlighting.)
## Sources
### 1. Thariq Shihipar — "Claude Code HTML output" (Medium, 2026)
The spec. Five advantages mapped here: density, clarity, shareability, two-way interaction, context ingestion. The four-pattern taxonomy above is the implementation answer.
### 2. Edward Tufte — *Envisioning Information* (Graphics Press, 1990)
Ch. 2, "Micro/Macro Readings" — argues that effective information design lets the reader move between overview (TOC) and detail (paragraph) without losing context. The scrollspy + sticky TOC implements this micro/macro discipline for documents.
### 3. Amelia Wattenberger — *Why React isn't great for actually building websites* (wattenberger.com, 2022) + interactive essay archive
Argues that documents are not apps; framework runtimes are overhead. Validates the vanilla-JS + IntersectionObserver implementation choice.
### 4. Jakob Nielsen / NN/g — *How Users Read on the Web* (1997, updated 2024)
Establishes the F-shaped reading pattern: users scan headings + first sentences. The sticky-sidebar TOC + bold heading typography + H2 hairline border-bottom optimize for this pattern.
### 5. Maggie Appleton — *Digital Gardens* (maggieappleton.com, 2020)
The single-page-document-with-lightweight-interactivity pattern at scale. Her own gardens use exactly the techniques this converter emits.
### 6. Bartosz Ciechanowski — interactive essay archive (ciechanow.ski, 2017-present)
The upper bound of what vanilla-JS + inline SVG can produce in a single HTML file. Demonstrates that "lightweight" doesn't mean "low-quality."
### 7. Bret Victor — "Up and Down the Ladder of Abstraction" (worrydream.com, 2011)
Argues for letting the reader move fluidly between concrete and abstract. The TOC (abstract) + section detail (concrete) + scrollspy (the link between them) is the documents-shaped implementation.
## Applied to `md-document`
The converter emits exactly these patterns. `markdown_parser.py` extracts the structure; `html_renderer.py` renders it with the design-system tokens; `interactivity_injector.py` adds the four interactive behaviors. None of these need a JS framework or a build step.
FILE:references/single_file_html_discipline.md
# Single-File HTML Discipline (md-document edition)
**Why this exists:** The orchestrator's single-file discipline document (`markdown-html-orchestrator/references/single_file_html_discipline.md`) establishes the rule. This document records how `md-document` specifically honors it — and what trade-offs the implementation makes.
## The contract
Every md-document output is one `.html` file. The only external HTTP requests it triggers are:
1. **`fonts.googleapis.com`** — Google Fonts CSS for the user's chosen heading + body families.
2. **`cdn.jsdelivr.net`** — Prism.js core + autoloader for syntax highlighting.
Both have graceful fallbacks:
- Google Fonts blocked → system font stack (Georgia/serif fallback for serif families; system-ui/sans-serif fallback for sans families; ui-monospace fallback for mono).
- Prism CDN blocked → `<pre><code>` renders as plain monospaced text (no colors, but still legible).
No other CDN. No web fonts hosted elsewhere. No analytics. No tracking pixels. No CMS framework runtime.
## Why these two externals?
### Google Fonts CSS (not woff files)
We link the Google Fonts CSS endpoint (`fonts.googleapis.com/css2?family=...&display=swap`). The CSS file is < 1 KB; the woff2 font files are lazy-loaded by the browser when they're actually needed for rendering. `display=swap` ensures system fonts show during the loading window, preventing FOIT (Flash of Invisible Text).
Alternative: base64-embed the woff2 files directly in the HTML. We rejected this because:
- A single Inter family at 4 weights is ~280 KB base64-encoded
- Most readers already have it cached from another site
- The CSS-link approach lets Google serve a smaller, browser-specific subset
### Prism.js (not highlight.js or shiki)
| Library | Core size | Why we chose Prism |
|---|---|---|
| Prism.js | ~2 KB core + per-language | Smallest core; autoloader fetches languages on demand |
| highlight.js | ~25 KB | More languages out-of-the-box but bigger initial payload |
| shiki | ~150 KB | VS-Code-fidelity output; oversized for documents |
Prism's autoloader pattern means a Python-heavy spec only fetches the Python language file, not the whole package. The result: most documents add ~5-10 KB of JS for syntax highlighting.
## What `md-document` does NOT externalize
- **CSS** — All styles inline in `<style>` block. The full BASE_CSS + style-overrides + palette is ~6 KB.
- **JavaScript** — Search/copy/scrollspy code is ~3 KB inline. No external bundle.
- **Images** — Logo (if supplied as local path) is base64-embedded. Inline SVG remains inline.
- **Icons** — Callout indicators use plain text characters (`i`, `*`, `!`) rather than icon fonts. Trade-off: less visual richness, no extra CDN entry.
## Footprint by document size
Empirically (from the smoke tests):
| Input markdown | Output HTML (no JS) | Output HTML (with JS) |
|---|---|---|
| ~150 lines (sample spec) | ~11 KB | ~15 KB |
| ~470 lines (markdown-html/CLAUDE.md) | ~17 KB | ~23 KB |
For a typical 100-500-line spec, the user gets a ~15-25 KB artifact they can email, drop in Slack, or upload to any static host. By contrast, a comparable Notion / Confluence / GitBook export would be 200 KB+ of CSS chrome + analytics scripts.
## Anti-patterns
- ❌ `<link rel="stylesheet" href="./style.css">` — separate file = shareability broken.
- ❌ `<script src="./app.js">` — same problem.
- ❌ `cdn.tailwindcss.com` — 200 KB of unused atomic CSS.
- ❌ `<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/some-icons">` — adds a third external; we don't need it.
- ❌ Service worker registration — single-file artifacts aren't apps.
## Sources
### 1. Thariq Shihipar — "Claude Code HTML output" (Medium, 2026)
"Every playground is a single HTML file with all CSS and JavaScript inlined."
### 2. marketing/landing/skills/landing/SKILL.md
Established the single-file rule in this repo. md-document inherits the discipline.
### 3. Tom MacWright — "Big" (github.com/tmcw/big, MIT)
A single-file presentation tool — full slide deck with keyboard nav in one file. Demonstrates the upper bound.
### 4. Google Fonts API documentation (developers.google.com/fonts/docs/css2)
The `display=swap` parameter behavior and CSS-vs-direct-woff trade-off.
### 5. Prism.js documentation (prismjs.com/extending.html#autoloader)
The autoloader pattern — fetch only the language plugins the document actually uses.
### 6. MDN — "Performance: Reducing HTTP Requests" (developer.mozilla.org)
Articulates why a single file beats N files even on fast networks.
### 7. Anil Dash — "The Web We Lost" (dashes.com, 2012)
The portability argument: a single self-contained HTML file is the most platform-independent web artifact possible.
## Applied to `md-document`
The renderer emits the canonical shape: `<!DOCTYPE html><html><head>...inline-style/font-link/prism-link...</head><body>...rendered-blocks...<inline-script></body></html>`. Anything that tries to externalize CSS, JS, images, or fonts beyond the two permitted endpoints is a regression.
FILE:references/toc_and_nav_ux.md
# Table-of-Contents and Navigation UX
**Why this exists:** The TOC is the single highest-leverage navigation aid in a long document. The design-system config offers four behaviors (`sticky-sidebar`, `collapsible-top`, `inline`, `none`); this document explains when each is right and what UX patterns the renderer implements.
## The four TOC variants
| Behavior | Use when | Renderer detail |
|---|---|---|
| **`sticky-sidebar`** (default) | Document > 800 words / 4+ H2 sections; landscape reading on desktop | Two-column CSS grid; nav is `position: sticky; top: 1.5rem`; collapses to top-of-page on viewports < 800px via media query |
| **`collapsible-top`** | Document ≤ 800 words but with > 3 sections; mobile-first | `<details open>` at top of document; user can collapse to recover vertical space |
| **`inline`** | Document is its own TOC (the bullet list at top IS the navigation) | TOC nav is suppressed; the markdown's own list serves the purpose |
| **`none`** | Short documents, focused single-section pieces | No TOC rendered at all |
Default is `sticky-sidebar` because the median document this converter sees is a multi-section spec or RFC, and the sidebar serves both as TOC and as "you are here" indicator (via scrollspy).
## Scrollspy implementation
`interactivity_injector.py` uses `IntersectionObserver` with this rootMargin:
```js
{ rootMargin: "-20% 0px -70% 0px", threshold: 0 }
```
A heading is considered "current" only when it's in the **upper-middle** of the viewport (between 20% from top and 30% from top). This matches the F-shape reading pattern (Nielsen/NN-g): users fixate on text just below the fold-line, not at the very top.
When the observer fires, the matching TOC link gets `aria-current="location"`. CSS then highlights it via:
```css
nav.toc a[aria-current="location"] {
color: var(--md-accent);
font-weight: 600;
background: var(--md-accent-soft);
}
```
Both the attribute and the visual highlight are semantic — screen readers announce "current location" without us needing an extra ARIA-live region.
## Search-as-filter (not search-as-jump)
The search bar filters which H2 sections are visible. It does NOT scroll-to-match the way GitHub's `?text=foo` URL does. Reason: in a filtered view, the reader can see the structure of what survives the filter (which sections matched). A jump-to-first-match loses that structural information.
Esc clears the filter. Sticky positioning ensures the search bar stays visible during scroll.
## Sources
### 1. Jakob Nielsen / NN/g — *Table of Contents Best Practices* (2023)
The canonical reference. Establishes:
- TOC should appear at the top OR persist sticky (not just mid-document)
- Anchored links should scroll, not full-page navigate
- Current-location indication is required for documents over ~1000 words
- Depth cap at H3 (max_depth=3 default) is a usability finding — deeper hierarchies become noise
### 2. WCAG 2.2 — *Success Criterion 2.4.5: Multiple Ways* (w3.org/WAI/WCAG22)
Mandates that long pages provide more than one way to find content. The TOC + scrollspy + search trio satisfies this for any single document.
### 3. ARIA Authoring Practices — *aria-current attribute* (w3.org/WAI/ARIA/apg/practices/feedback/)
Documents the `aria-current="location"` pattern as the standard for "current page/section" indication. Screen readers (NVDA, JAWS, VoiceOver) announce it appropriately.
### 4. Vitepress / Docusaurus / mdBook — sticky-sidebar TOC implementations
All three of these documentation systems converged on the sticky-sidebar pattern as the right default for technical documents. We mirror their behavior (left or right column, sticky-positioned, scrollspy-enabled) rather than reinventing it.
### 5. GOV.UK Design System — *Inline navigation* (design-system.service.gov.uk)
For shorter pages, GOV.UK uses an inline anchor list rather than a sidebar. Validates the `inline` and `collapsible-top` behaviors as legitimate alternatives for shorter documents.
### 6. MDN Web Docs — *IntersectionObserver API* (developer.mozilla.org)
The browser primitive that makes scrollspy possible without scroll-event throttling. Available since 2017, ~95% browser support today.
## Applied to `md-document`
The renderer emits the right nav variant based on `toc.behavior`. The injector wires up scrollspy + search behavior. Every section heading H2-H{max_depth+1} gets an anchor + TOC entry; H1 is the document title (not navigation).
FILE:scripts/html_renderer.py
#!/usr/bin/env python3
"""html_renderer.py - Render a parsed-markdown section tree to single-file HTML.
Stdlib-only. Reads a JSON section tree (from markdown_parser.py) plus the
design-system config (from config_loader.py), emits a complete self-contained
.html file with:
- <title> from the document's H1
- Google Fonts CDN link (per typography.heading_font + typography.body_font)
- Prism.js CDN link (per code_theme: light/dark/auto)
- <style> block with :root { --md-bg: ...; } from the derived 12-token palette
- Base CSS scaled by typography.scale_ratio and design_style
- TOC per toc.behavior (sticky-sidebar / collapsible-top / inline / none)
- Rendered blocks: headings, paragraphs, lists, tables, code, callouts, quotes
- Footer with company_name + logo (base64-embedded if data: URL or local file)
NO LLM CALLS. Pure templating + config-driven CSS.
The output is one HTML file. Externals are limited to:
- fonts.googleapis.com (Google Fonts CSS)
- cdn.jsdelivr.net (Prism.js)
Falls back to system fonts + plain <pre> if either CDN is blocked.
Usage:
python html_renderer.py --sections sections.json --output report.html
python html_renderer.py --sample
python html_renderer.py --sections - --output - --no-config # full pipe
"""
from __future__ import annotations
import argparse
import base64
import html
import json
import os
import sys
from pathlib import Path
from typing import Any
_DESIGN_SYSTEM_SCRIPTS = (
Path(__file__).resolve().parent.parent.parent / "design-system" / "scripts"
)
sys.path.insert(0, str(_DESIGN_SYSTEM_SCRIPTS))
try:
import config_loader as _cfg
except ImportError:
_cfg = None
# Re-export from markdown_parser so html_renderer can self-sample
sys.path.insert(0, str(Path(__file__).resolve().parent))
try:
import markdown_parser as _mp
except ImportError:
_mp = None
# ----- Design-style presets ----------------------------------------------------
STYLE_CSS_OVERRIDES: dict[str, str] = {
"editorial": """
body.style-editorial { max-width: 720px; line-height: 1.75; }
body.style-editorial main h2 { font-size: calc(1rem * var(--md-scale) * var(--md-scale) * var(--md-scale)); margin-top: 4rem; }
body.style-editorial p { font-size: 1.0625rem; }
""",
"technical": """
body.style-technical { max-width: 960px; line-height: 1.6; }
body.style-technical main h2 { font-size: calc(1rem * var(--md-scale) * var(--md-scale)); margin-top: 2.5rem; }
body.style-technical pre { font-size: 0.875rem; line-height: 1.5; }
""",
"minimal": """
body.style-minimal { max-width: 680px; line-height: 1.65; }
body.style-minimal main h2 { font-size: calc(1rem * var(--md-scale)); margin-top: 3rem; font-weight: 400; }
body.style-minimal .callout { background: transparent; border-left: 2px solid var(--md-border); }
""",
"playful": """
body.style-playful { max-width: 880px; line-height: 1.7; }
body.style-playful main h2 { font-size: calc(1rem * var(--md-scale) * var(--md-scale) * var(--md-scale)); margin-top: 3.5rem; }
body.style-playful .callout { border-radius: 1rem; box-shadow: 0 4px 16px rgba(0,0,0,0.04); }
""",
}
# ----- CSS template ------------------------------------------------------------
BASE_CSS = """
:root {
__PALETTE__
--md-scale: __SCALE__;
--md-font-heading: __HEADING_FONT__;
--md-font-body: __BODY_FONT__;
}
* { box-sizing: border-box; }
html { scroll-behavior: smooth; -webkit-text-size-adjust: 100%; }
body {
margin: 0;
padding: 2rem 1.5rem;
background: var(--md-bg);
color: var(--md-text);
font-family: var(--md-font-body);
font-size: 16px;
line-height: 1.6;
max-width: 960px;
margin-left: auto;
margin-right: auto;
}
main h1, main h2, main h3, main h4, main h5, main h6 {
font-family: var(--md-font-heading);
color: var(--md-text);
line-height: 1.25;
margin: 1.5em 0 0.5em;
font-weight: 600;
scroll-margin-top: 1rem;
}
main h1 { font-size: calc(1rem * var(--md-scale) * var(--md-scale) * var(--md-scale) * var(--md-scale)); margin-top: 0; }
main h2 { font-size: calc(1rem * var(--md-scale) * var(--md-scale) * var(--md-scale)); border-bottom: 1px solid var(--md-border); padding-bottom: 0.3em; }
main h3 { font-size: calc(1rem * var(--md-scale) * var(--md-scale)); }
main h4 { font-size: calc(1rem * var(--md-scale)); }
main h5, main h6 { font-size: 1rem; color: var(--md-text-muted); }
p { margin: 0.5em 0 1em; }
a { color: var(--md-link); text-decoration: underline; text-underline-offset: 2px; }
a:hover { color: var(--md-link-hover); }
strong { font-weight: 600; }
em { font-style: italic; }
code {
font-family: 'JetBrains Mono', ui-monospace, SFMono-Regular, Menlo, monospace;
background: var(--md-code-bg);
padding: 0.15em 0.35em;
border-radius: 4px;
font-size: 0.9em;
}
pre {
background: var(--md-code-bg);
border: 1px solid var(--md-border);
border-radius: 8px;
padding: 1rem 1.25rem;
overflow-x: auto;
font-size: 0.875rem;
line-height: 1.55;
margin: 1.5em 0;
position: relative;
}
pre code { background: transparent; padding: 0; font-size: 1em; }
table {
width: 100%;
border-collapse: collapse;
margin: 1.5em 0;
font-size: 0.9375rem;
}
th, td {
border: 1px solid var(--md-border);
padding: 0.5em 0.75em;
text-align: left;
}
th { background: var(--md-surface); font-weight: 600; }
td.align-center, th.align-center { text-align: center; }
td.align-right, th.align-right { text-align: right; }
blockquote {
border-left: 3px solid var(--md-border);
margin: 1.5em 0;
padding: 0.5em 0 0.5em 1.25em;
color: var(--md-text-muted);
font-style: italic;
}
ul, ol { padding-left: 1.5em; margin: 0.5em 0 1em; }
li { margin: 0.25em 0; }
hr {
border: 0;
border-top: 1px solid var(--md-border);
margin: 3em 0;
}
.callout {
border-left: 4px solid var(--md-accent);
background: var(--md-accent-soft);
padding: 0.75rem 1rem 0.75rem 1.25rem;
margin: 1.5em 0;
border-radius: 0 8px 8px 0;
}
.callout .callout-label {
font-family: var(--md-font-heading);
font-weight: 600;
font-size: 0.8125rem;
text-transform: uppercase;
letter-spacing: 0.05em;
margin-bottom: 0.25em;
color: var(--md-accent);
display: flex;
align-items: center;
gap: 0.5em;
}
.callout .callout-icon {
display: inline-flex;
width: 1.125em;
height: 1.125em;
align-items: center;
justify-content: center;
}
.callout-note { border-left-color: var(--md-link); }
.callout-note .callout-label { color: var(--md-link); }
.callout-tip { border-left-color: var(--md-success); }
.callout-tip .callout-label { color: var(--md-success); }
.callout-important { border-left-color: var(--md-accent); }
.callout-warning { border-left-color: var(--md-warn); }
.callout-warning .callout-label { color: var(--md-warn); }
.callout-caution { border-left-color: var(--md-warn); }
.callout-caution .callout-label { color: var(--md-warn); }
.callout p:last-child { margin-bottom: 0; }
.callout p:first-child { margin-top: 0; }
/* TOC */
nav.toc { font-size: 0.9375rem; line-height: 1.5; }
nav.toc ol, nav.toc ul { padding-left: 1.25em; }
nav.toc a {
color: var(--md-text-muted);
text-decoration: none;
display: block;
padding: 0.15em 0.25em;
border-radius: 3px;
}
nav.toc a:hover { color: var(--md-link); background: var(--md-accent-soft); }
nav.toc a[aria-current="location"] {
color: var(--md-accent);
font-weight: 600;
background: var(--md-accent-soft);
}
/* TOC variants */
body.toc-sticky-sidebar { display: grid; grid-template-columns: 220px 1fr; gap: 2.5rem; max-width: 1200px; }
body.toc-sticky-sidebar nav.toc {
position: sticky;
top: 1.5rem;
align-self: start;
max-height: calc(100vh - 3rem);
overflow-y: auto;
border-right: 1px solid var(--md-border);
padding-right: 1rem;
}
@media (max-width: 800px) {
body.toc-sticky-sidebar { display: block; }
body.toc-sticky-sidebar nav.toc { position: static; border-right: none; max-height: none; margin-bottom: 2rem; }
}
body.toc-collapsible-top nav.toc {
background: var(--md-surface);
border: 1px solid var(--md-border);
border-radius: 8px;
padding: 1rem 1.25rem;
margin-bottom: 2rem;
}
body.toc-collapsible-top nav.toc summary { cursor: pointer; font-weight: 600; font-family: var(--md-font-heading); }
body.toc-none nav.toc { display: none; }
/* Search */
.md-search {
position: sticky;
top: 0;
background: var(--md-bg);
padding: 0.5rem 0 0.75rem;
z-index: 10;
border-bottom: 1px solid var(--md-border);
margin-bottom: 1rem;
}
.md-search input {
width: 100%;
padding: 0.5rem 0.75rem;
border: 1px solid var(--md-border);
border-radius: 6px;
font-size: 0.9375rem;
background: var(--md-surface);
color: var(--md-text);
font-family: inherit;
}
.md-search input:focus {
outline: 2px solid var(--md-accent);
outline-offset: 2px;
}
main section[hidden] { display: none; }
/* Code-copy button */
.code-copy {
position: absolute;
top: 0.5rem;
right: 0.5rem;
background: var(--md-surface);
color: var(--md-text-muted);
border: 1px solid var(--md-border);
border-radius: 5px;
padding: 0.2em 0.5em;
font-size: 0.75rem;
cursor: pointer;
opacity: 0;
transition: opacity 0.15s ease;
font-family: inherit;
}
pre:hover .code-copy { opacity: 1; }
.code-copy:hover { color: var(--md-text); background: var(--md-bg); }
.code-copy.copied { color: var(--md-success); }
/* Footer */
footer.md-footer {
margin-top: 4rem;
padding-top: 1.5rem;
border-top: 1px solid var(--md-border);
color: var(--md-text-muted);
font-size: 0.875rem;
display: flex;
align-items: center;
gap: 1rem;
}
footer.md-footer img { max-height: 24px; max-width: 120px; }
@media (prefers-reduced-motion: reduce) {
* { animation: none !important; transition: none !important; }
html { scroll-behavior: auto; }
}
"""
# ----- Helpers -----------------------------------------------------------------
CALLOUT_ICONS: dict[str, str] = {
"NOTE": "i",
"TIP": "*",
"IMPORTANT": "!",
"WARNING": "!",
"CAUTION": "!",
}
def _palette_to_css(palette: dict[str, str]) -> str:
if not palette:
# Fallback dark-mode defaults so an un-onboarded render still works
palette = {
"--md-bg": "#0E1E38", "--md-surface": "#142B50", "--md-border": "#1A3868",
"--md-text": "#F7F7F2", "--md-text-muted": "rgba(247, 247, 242, 0.68)",
"--md-accent": "#00D4AA", "--md-accent-soft": "rgba(0, 212, 170, 0.14)",
"--md-code-bg": "#122648",
"--md-link": "#00D4AA", "--md-link-hover": "#08FECE",
"--md-success": "#10A85C", "--md-warn": "#C87C10",
}
return "\n".join(f" {k}: {v};" for k, v in palette.items())
def _font_url(heading: str, body: str) -> str:
families = sorted({heading, body})
parts = "&".join(f"family={f.replace(' ', '+')}:wght@400;600" for f in families)
return f"https://fonts.googleapis.com/css2?{parts}&display=swap"
def _font_stack(name: str, kind: str) -> str:
fallback = ("Georgia, serif" if "serif" in name.lower() or name in
("Playfair Display", "Merriweather", "Lora", "Source Serif 4")
else "system-ui, -apple-system, sans-serif")
if "Mono" in name or "Code" in name:
fallback = "ui-monospace, SFMono-Regular, Menlo, monospace"
return f"'{name}', {fallback}"
def _prism_theme_link(code_theme: str) -> str:
# auto: load both light + dark prefers-color-scheme variants
if code_theme == "dark":
return ('<link rel="stylesheet" '
'href="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/themes/prism-tomorrow.min.css">')
if code_theme == "light":
return ('<link rel="stylesheet" '
'href="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/themes/prism.min.css">')
return (
'<link rel="stylesheet" media="(prefers-color-scheme: light)" '
'href="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/themes/prism.min.css">\n'
'<link rel="stylesheet" media="(prefers-color-scheme: dark)" '
'href="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/themes/prism-tomorrow.min.css">'
)
def _embed_logo(logo_url: str) -> str:
"""Return a usable src attribute for the logo. Base64-embed local paths;
leave URLs as-is (recipient's browser will fetch them)."""
if not logo_url:
return ""
if logo_url.startswith(("data:", "http://", "https://")):
return logo_url
p = Path(logo_url).expanduser()
if p.exists() and p.is_file():
ext = p.suffix.lstrip(".").lower() or "png"
data = base64.b64encode(p.read_bytes()).decode("ascii")
mime = {"png": "image/png", "jpg": "image/jpeg", "jpeg": "image/jpeg",
"svg": "image/svg+xml", "gif": "image/gif", "webp": "image/webp"}.get(
ext, f"image/{ext}")
return f"data:{mime};base64,{data}"
return logo_url # let the browser handle the broken reference visibly
def _render_block(block: dict[str, Any]) -> str:
t = block["type"]
if t == "heading":
level = block["level"]
anchor = block["anchor"]
text = _mp.render_inline_html(block["text"]) if _mp else html.escape(block["text"])
return f'<h{level} id="{anchor}">{text}</h{level}>'
if t == "paragraph":
return f"<p>{_mp.render_inline_html(block['text']) if _mp else html.escape(block['text'])}</p>"
if t == "hr":
return "<hr>"
if t == "code":
lang = block.get("language") or "text"
# Prism class convention
body = html.escape(block["body"])
return f'<pre><button class="code-copy" type="button" aria-label="Copy code">Copy</button><code class="language-{html.escape(lang)}">{body}</code></pre>'
if t == "list":
tag = "ol" if block.get("ordered") else "ul"
items = "".join(
f"<li>{_mp.render_inline_html(item) if _mp else html.escape(item)}</li>"
for item in block["items"]
)
return f"<{tag}>{items}</{tag}>"
if t == "table":
headers = block["headers"]
aligns = block.get("aligns") or ["left"] * len(headers)
rows = block["rows"]
thead = "<thead><tr>" + "".join(
f'<th class="align-{a}">{_mp.render_inline_html(h) if _mp else html.escape(h)}</th>'
for h, a in zip(headers, aligns)
) + "</tr></thead>"
tbody = "<tbody>" + "".join(
"<tr>" + "".join(
f'<td class="align-{aligns[i] if i < len(aligns) else "left"}">'
f'{_mp.render_inline_html(cell) if _mp else html.escape(cell)}</td>'
for i, cell in enumerate(row)
) + "</tr>"
for row in rows
) + "</tbody>"
return f"<table>{thead}{tbody}</table>"
if t == "callout":
kind = (block.get("kind") or "NOTE").upper()
icon = CALLOUT_ICONS.get(kind, "i")
body = "<br>".join(
_mp.render_inline_html(ln) if _mp else html.escape(ln)
for ln in (block.get("body_lines") or []) if ln
)
klass = kind.lower()
return (
f'<aside class="callout callout-{klass}" role="note">'
f'<div class="callout-label"><span class="callout-icon" aria-hidden="true">{icon}</span>'
f'{html.escape(kind)}</div>'
f'<div class="callout-body">{body}</div>'
f'</aside>'
)
if t == "blockquote":
body = "<br>".join(
_mp.render_inline_html(ln) if _mp else html.escape(ln)
for ln in block.get("body_lines") or [] if ln
)
return f"<blockquote>{body}</blockquote>"
return ""
def _render_toc(blocks: list[dict[str, Any]], max_depth: int, behavior: str) -> str:
if behavior == "none":
return ""
items = [b for b in blocks if b["type"] == "heading" and 2 <= b["level"] <= max_depth + 1]
if not items:
return ""
# Group H2..H{max_depth+1} into a nested <ol> structure
out: list[str] = []
out.append('<nav class="toc" aria-label="Table of contents">')
if behavior == "collapsible-top":
out.append('<details open><summary>Contents</summary>')
out.append("<ol>")
last_level = 2
for h in items:
lvl = h["level"]
if lvl > last_level:
out.append("<ol>" * (lvl - last_level))
elif lvl < last_level:
out.append("</ol>" * (last_level - lvl))
last_level = lvl
out.append(f'<li><a href="#{h["anchor"]}">{html.escape(h["text"])}</a></li>')
if last_level > 2:
out.append("</ol>" * (last_level - 2))
out.append("</ol>")
if behavior == "collapsible-top":
out.append("</details>")
out.append("</nav>")
return "\n".join(out)
def render(sections: dict[str, Any], config: dict[str, Any]) -> str:
meta = sections.get("meta", {})
blocks = sections.get("blocks", [])
title = meta.get("title") or "Document"
palette = config.get("derived_palette") or {}
typo = config.get("typography") or {}
heading_font = typo.get("heading_font", "Inter")
body_font = typo.get("body_font", "Inter")
scale = typo.get("scale_ratio", 1.25)
style = config.get("design_style", "technical")
code_theme = config.get("code_theme", "auto")
toc_cfg = config.get("toc") or {}
toc_behavior = toc_cfg.get("behavior", "sticky-sidebar")
toc_max_depth = toc_cfg.get("max_depth", 3)
company_name = config.get("company_name", "")
logo_url = _embed_logo(config.get("logo_url", "") or "")
css = (BASE_CSS
.replace("__PALETTE__", _palette_to_css(palette))
.replace("__SCALE__", str(scale))
.replace("__HEADING_FONT__", _font_stack(heading_font, "heading"))
.replace("__BODY_FONT__", _font_stack(body_font, "body")))
css += STYLE_CSS_OVERRIDES.get(style, "")
toc_html = _render_toc(blocks, toc_max_depth, toc_behavior)
body_html = "\n".join(_render_block(b) for b in blocks)
footer_parts: list[str] = []
if logo_url:
footer_parts.append(f'<img src="{html.escape(logo_url)}" alt="{html.escape(company_name or "Logo")}">')
if company_name:
footer_parts.append(f"<span>{html.escape(company_name)}</span>")
footer_parts.append(f'<span style="margin-left:auto">Generated by markdown-html</span>')
footer_html = ("<footer class=\"md-footer\">" + "".join(footer_parts) + "</footer>"
if footer_parts else "")
search_html = (
'<div class="md-search">'
'<input type="search" id="md-search-input" '
'placeholder="Filter sections… (Esc to clear)" '
'aria-label="Filter document sections">'
'</div>'
)
return f"""<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>{html.escape(title)}</title>
<link rel="preconnect" href="https://fonts.googleapis.com" crossorigin>
<link rel="stylesheet" href="{_font_url(heading_font, body_font)}">
{_prism_theme_link(code_theme)}
<style>{css}</style>
</head>
<body class="style-{style} toc-{toc_behavior}">
{toc_html}
<main>
{search_html}
{body_html}
{footer_html}
</main>
<script defer src="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/components/prism-core.min.js"></script>
<script defer src="https://cdn.jsdelivr.net/npm/prismjs@1.29.0/plugins/autoloader/prism-autoloader.min.js"></script>
</body>
</html>"""
def main(argv: list[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--sections", help="Path to sections JSON, or '-' for stdin")
parser.add_argument("--output", help="Path to write HTML, or '-' for stdout")
parser.add_argument("--sample", action="store_true",
help="Render a built-in sample document")
parser.add_argument("--no-config", action="store_true",
help="Bypass design-system config (use DEFAULTS)")
args = parser.parse_args(argv)
if args.sample:
if _mp is None:
print("error: markdown_parser not importable", file=sys.stderr)
return 2
sections = _mp.parse_markdown(_mp.SAMPLE_MARKDOWN)
elif args.sections:
raw = sys.stdin.read() if args.sections == "-" else Path(args.sections).read_text(encoding="utf-8")
sections = json.loads(raw)
else:
parser.print_help()
return 0
if args.no_config or os.environ.get("MARKDOWN_HTML_NO_CONFIG") == "1":
config = _cfg.DEFAULTS if _cfg else {}
else:
config = _cfg.load_config() if _cfg else {}
output = render(sections, config)
if args.output and args.output != "-":
Path(args.output).write_text(output, encoding="utf-8")
print(f"wrote {args.output}: {len(output):,} bytes, "
f"{sections['meta'].get('section_count', 0)} sections")
else:
print(output)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/interactivity_injector.py
#!/usr/bin/env python3
"""interactivity_injector.py - Inject vanilla-JS interactivity into rendered HTML.
Stdlib-only. Takes an HTML file produced by html_renderer.py and injects a
<script> block (immediately before </body>) that wires up:
- search Client-side filter on the search input — hides H2 sections
whose heading or body text doesn't match the query. Esc clears.
- copycode Click handler on every .code-copy button. Copies the <code>
text to clipboard, toggles a "copied" state for 1.2s.
- smoothscroll Click handler on TOC links — smooth-scrolls to the target
anchor. Complements CSS scroll-behavior: smooth as a fallback.
- scrollspy IntersectionObserver on every <h2 id="..."> — sets
aria-current="location" on the matching TOC link as the user
reads. Foundation for "you are here" navigation.
NO LLM CALLS. Pure script template + HTML insertion.
The injected JS:
- Uses no frameworks (vanilla DOM API + IntersectionObserver only)
- Total payload ~3 KB minified-ish
- Degrades gracefully: if IntersectionObserver is missing (very old browsers),
scrollspy is silently skipped; the rest still works.
Idempotent: if the script block is already present (by ID), the file is
left unchanged.
Usage:
python interactivity_injector.py --file report.html \\
--features search,copycode,smoothscroll,scrollspy
python interactivity_injector.py --sample
"""
from __future__ import annotations
import argparse
import re
import sys
from pathlib import Path
INJECT_MARKER_ID = "md-document-interactivity-v1"
# JavaScript payload. Indented carefully so the produced HTML is still readable.
JS_PAYLOAD_TEMPLATE = """\
<script id="__MARKER__">
(function () {
"use strict";
var ENABLED = __FEATURES__;
// ----- Section grouping (used by search) -----
// Each H2 + everything until the next H2 forms a "section" for filter purposes.
function groupSections(root) {
var groups = [];
var current = null;
Array.prototype.forEach.call(root.children, function (el) {
if (el.tagName === "H2") {
if (current) groups.push(current);
current = { heading: el, elements: [el], text: el.textContent.toLowerCase() };
} else if (current) {
current.elements.push(el);
current.text += " " + (el.textContent || "").toLowerCase();
}
});
if (current) groups.push(current);
return groups;
}
// ----- Search -----
function wireSearch(root) {
var input = document.getElementById("md-search-input");
if (!input || !ENABLED.search) return;
var groups = groupSections(root);
function apply() {
var q = input.value.trim().toLowerCase();
groups.forEach(function (g) {
var visible = !q || g.text.indexOf(q) !== -1;
g.elements.forEach(function (el) { el.hidden = !visible; });
});
}
input.addEventListener("input", apply);
input.addEventListener("keydown", function (e) {
if (e.key === "Escape") { input.value = ""; apply(); }
});
}
// ----- Code-copy -----
function wireCopy() {
if (!ENABLED.copycode) return;
Array.prototype.forEach.call(
document.querySelectorAll("pre .code-copy"),
function (btn) {
btn.addEventListener("click", function () {
var pre = btn.parentElement;
var code = pre.querySelector("code");
if (!code) return;
var text = code.textContent;
var done = function () {
btn.classList.add("copied");
var original = btn.textContent;
btn.textContent = "Copied";
setTimeout(function () {
btn.classList.remove("copied");
btn.textContent = original === "Copied" ? "Copy" : original;
}, 1200);
};
if (navigator.clipboard && navigator.clipboard.writeText) {
navigator.clipboard.writeText(text).then(done, function () {
// Fallback to execCommand on older browsers
fallbackCopy(text);
done();
});
} else {
fallbackCopy(text);
done();
}
});
}
);
}
function fallbackCopy(text) {
var ta = document.createElement("textarea");
ta.value = text;
ta.style.position = "fixed";
ta.style.opacity = "0";
document.body.appendChild(ta);
ta.select();
try { document.execCommand("copy"); } catch (e) {}
document.body.removeChild(ta);
}
// ----- Smooth-scroll for TOC links -----
function wireSmoothScroll() {
if (!ENABLED.smoothscroll) return;
Array.prototype.forEach.call(
document.querySelectorAll("nav.toc a[href^=\\\"#\\\"]"),
function (a) {
a.addEventListener("click", function (e) {
var id = a.getAttribute("href").slice(1);
var target = document.getElementById(id);
if (!target) return;
e.preventDefault();
target.scrollIntoView({ behavior: "smooth", block: "start" });
history.replaceState(null, "", "#" + id);
});
}
);
}
// ----- Scrollspy -----
function wireScrollSpy() {
if (!ENABLED.scrollspy || !("IntersectionObserver" in window)) return;
var tocLinks = {};
Array.prototype.forEach.call(
document.querySelectorAll("nav.toc a[href^=\\\"#\\\"]"),
function (a) {
var id = a.getAttribute("href").slice(1);
tocLinks[id] = a;
}
);
var headings = document.querySelectorAll("main h2[id], main h3[id]");
if (!headings.length) return;
function clearActive() {
Object.keys(tocLinks).forEach(function (k) {
tocLinks[k].removeAttribute("aria-current");
});
}
var observer = new IntersectionObserver(function (entries) {
// Pick the topmost entry currently intersecting
var visible = entries.filter(function (e) { return e.isIntersecting; });
if (visible.length === 0) return;
visible.sort(function (a, b) { return a.boundingClientRect.top - b.boundingClientRect.top; });
var id = visible[0].target.id;
var link = tocLinks[id];
if (link) { clearActive(); link.setAttribute("aria-current", "location"); }
}, { rootMargin: "-20% 0px -70% 0px", threshold: 0 });
Array.prototype.forEach.call(headings, function (h) { observer.observe(h); });
}
// ----- Boot -----
function init() {
var main = document.querySelector("main");
if (!main) return;
wireSearch(main);
wireCopy();
wireSmoothScroll();
wireScrollSpy();
}
if (document.readyState === "loading") {
document.addEventListener("DOMContentLoaded", init);
} else {
init();
}
})();
</script>
"""
ALL_FEATURES = ("search", "copycode", "smoothscroll", "scrollspy")
def _features_dict(features: list[str]) -> str:
enabled = set(features)
parts = ",".join(f'"{f}": {"true" if f in enabled else "false"}' for f in ALL_FEATURES)
return "{" + parts + "}"
def inject(html_text: str, features: list[str]) -> tuple[str, bool]:
"""Return (new_text, was_modified). Idempotent: no-op if marker already present."""
if f'id="{INJECT_MARKER_ID}"' in html_text:
return (html_text, False)
payload = (JS_PAYLOAD_TEMPLATE
.replace("__MARKER__", INJECT_MARKER_ID)
.replace("__FEATURES__", _features_dict(features)))
# Inject immediately before </body>
closing = re.compile(r"</body\s*>", re.IGNORECASE)
m = closing.search(html_text)
if not m:
# No </body> tag — append at end
return (html_text + "\n" + payload, True)
new_text = html_text[:m.start()] + payload + html_text[m.start():]
return (new_text, True)
def main(argv: list[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--file", help="Path to HTML file to modify in place")
parser.add_argument("--features",
default="search,copycode,smoothscroll,scrollspy",
help="Comma-separated subset of: search, copycode, smoothscroll, scrollspy")
parser.add_argument("--output",
help="Write to this path instead of in-place. '-' for stdout.")
parser.add_argument("--sample", action="store_true",
help="Inject into a fresh render of the built-in sample doc")
args = parser.parse_args(argv)
feats = [f.strip() for f in args.features.split(",") if f.strip()]
invalid = [f for f in feats if f not in ALL_FEATURES]
if invalid:
print(f"error: unknown feature(s): {invalid}. "
f"Valid: {list(ALL_FEATURES)}", file=sys.stderr)
return 2
if args.sample:
# Render the sample on the fly so the injector can be exercised standalone
sys.path.insert(0, str(Path(__file__).resolve().parent))
import html_renderer
import markdown_parser
sections = markdown_parser.parse_markdown(markdown_parser.SAMPLE_MARKDOWN)
sample_html = html_renderer.render(sections, {})
modified, was = inject(sample_html, feats)
out = args.output or "-"
if out == "-":
print(modified)
else:
Path(out).write_text(modified, encoding="utf-8")
print(f"wrote {out}: {len(modified):,} bytes "
f"(injected: {feats})")
return 0
if not args.file:
parser.print_help()
return 0
src = Path(args.file)
if not src.exists():
print(f"error: file not found: {src}", file=sys.stderr)
return 2
original = src.read_text(encoding="utf-8")
modified, was = inject(original, feats)
if args.output:
if args.output == "-":
print(modified)
return 0
Path(args.output).write_text(modified, encoding="utf-8")
target = args.output
else:
src.write_text(modified, encoding="utf-8")
target = str(src)
if was:
print(f"injected: {feats} -> {target} "
f"({len(modified) - len(original):+,} bytes)")
else:
print(f"no-op: marker '{INJECT_MARKER_ID}' already present in {target}")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/markdown_parser.py
#!/usr/bin/env python3
"""markdown_parser.py - CommonMark-subset parser for the md-document converter.
Stdlib-only. Reads a markdown file (or stdin), produces a structured section
tree as JSON that the html_renderer can consume. NO LLM CALLS — pure regex
+ state-machine line tokenization.
Scope (CommonMark subset sufficient for agent-generated specs/reports/RFCs):
- Headings: # / ## / ### / #### / ##### / ###### (1-6 levels)
- Paragraphs (lines separated by blank lines)
- Fenced code blocks (``` with optional language tag)
- Tables (GFM: header row + delimiter row + body rows)
- GFM-style callouts: > [!NOTE], > [!TIP], > [!IMPORTANT], > [!WARNING], > [!CAUTION]
- Plain blockquotes: > text
- Ordered lists: 1. / 2. / 3. (single-level only)
- Unordered lists: - / * / + (single-level only)
- Horizontal rules: --- / *** / ___
- Inline: **bold** / *italic* / `code` / [text](url) / 
Out of scope: nested lists, HTML inlines, footnotes, definition lists, task
list checkboxes (rendered as plain text), reference-style links, hard line
breaks (two-space). These can be added later if a real document needs them.
The output is a JSON object with two keys:
- meta: {title, line_count, heading_count, section_count}
- blocks: ordered list of block nodes; each section H2+ is also stored as
a structural anchor for the TOC + scrollspy.
Usage:
python markdown_parser.py --input report.md
python markdown_parser.py --input - --output sections.json
python markdown_parser.py --sample
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any
CALLOUT_RE = re.compile(r"^>\s*\[!(NOTE|TIP|IMPORTANT|WARNING|CAUTION)\]\s*$", re.IGNORECASE)
HEADING_RE = re.compile(r"^(#{1,6})\s+(.+?)\s*#*\s*$")
FENCE_RE = re.compile(r"^```(\S*)\s*$")
HR_RE = re.compile(r"^(-{3,}|\*{3,}|_{3,})\s*$")
ORDERED_LI_RE = re.compile(r"^(\d+)\.\s+(.+)$")
UNORDERED_LI_RE = re.compile(r"^[-*+]\s+(.+)$")
TABLE_DELIM_RE = re.compile(r"^\|?\s*:?-{3,}:?\s*(\|\s*:?-{3,}:?\s*)+\|?\s*$")
TABLE_ROW_RE = re.compile(r"^\|.*\|\s*$")
BLOCKQUOTE_RE = re.compile(r"^>\s?(.*)$")
INLINE_CODE_RE = re.compile(r"`([^`]+)`")
BOLD_RE = re.compile(r"\*\*([^*]+)\*\*")
ITALIC_RE = re.compile(r"(?<!\*)\*([^*]+)\*(?!\*)")
LINK_RE = re.compile(r"\[([^\]]+)\]\(([^)]+)\)")
IMAGE_RE = re.compile(r"!\[([^\]]*)\]\(([^)]+)\)")
def slugify(text: str) -> str:
"""Convert a heading text to a URL-safe anchor slug."""
text = re.sub(r"<[^>]+>", "", text) # strip any HTML tags
text = re.sub(r"[^a-zA-Z0-9\s-]", "", text)
text = re.sub(r"\s+", "-", text.strip())
return text.lower() or "section"
def render_inline_html(text: str) -> str:
"""Convert inline markdown markup to HTML, with HTML-escaping for safety."""
# HTML-escape first, then re-introduce markup via tokens that won't collide
# with user content. We use placeholder tokens to avoid double-substitution.
out = text
out = out.replace("&", "&").replace("<", "<").replace(">", ">")
# Images first (so the ! prefix isn't eaten by link)
out = IMAGE_RE.sub(lambda m: f'<img src="{m.group(2)}" alt="{m.group(1)}">', out)
# Links
out = LINK_RE.sub(lambda m: f'<a href="{m.group(2)}">{m.group(1)}</a>', out)
# Inline code (before bold/italic so backticks short-circuit emphasis)
out = INLINE_CODE_RE.sub(lambda m: f"<code>{m.group(1)}</code>", out)
# Bold
out = BOLD_RE.sub(lambda m: f"<strong>{m.group(1)}</strong>", out)
# Italic (single * not adjacent to another *)
out = ITALIC_RE.sub(lambda m: f"<em>{m.group(1)}</em>", out)
return out
def parse_table(lines: list[str], start: int) -> tuple[dict[str, Any], int]:
"""Parse a GFM table starting at lines[start]. Returns (node, next_index)."""
header_line = lines[start]
delim_line = lines[start + 1]
body_lines: list[str] = []
i = start + 2
while i < len(lines) and TABLE_ROW_RE.match(lines[i]):
body_lines.append(lines[i])
i += 1
def split_row(row: str) -> list[str]:
cells = row.strip().strip("|").split("|")
return [c.strip() for c in cells]
headers = split_row(header_line)
aligns = []
for cell in split_row(delim_line):
s = cell.strip()
if s.startswith(":") and s.endswith(":"):
aligns.append("center")
elif s.endswith(":"):
aligns.append("right")
else:
aligns.append("left")
rows = [split_row(r) for r in body_lines]
return ({"type": "table", "headers": headers, "aligns": aligns, "rows": rows}, i)
def parse_list(lines: list[str], start: int) -> tuple[dict[str, Any], int]:
"""Parse a single-level ordered or unordered list starting at start."""
first = lines[start]
ordered = bool(ORDERED_LI_RE.match(first))
items: list[str] = []
i = start
while i < len(lines):
if ordered:
m = ORDERED_LI_RE.match(lines[i])
if not m:
break
items.append(m.group(2))
else:
m = UNORDERED_LI_RE.match(lines[i])
if not m:
break
items.append(m.group(1))
i += 1
return ({"type": "list", "ordered": ordered, "items": items}, i)
def parse_callout(lines: list[str], start: int) -> tuple[dict[str, Any], int]:
"""Parse a GFM-style callout starting at start.
Pattern:
> [!NOTE]
> Body line 1
> Body line 2
"""
m = CALLOUT_RE.match(lines[start])
kind = m.group(1).upper() if m else "NOTE"
body: list[str] = []
i = start + 1
while i < len(lines):
bq = BLOCKQUOTE_RE.match(lines[i])
if not bq:
break
body.append(bq.group(1))
i += 1
return ({"type": "callout", "kind": kind, "body_lines": body}, i)
def parse_blockquote(lines: list[str], start: int) -> tuple[dict[str, Any], int]:
"""Parse a plain blockquote (no callout marker)."""
body: list[str] = []
i = start
while i < len(lines):
bq = BLOCKQUOTE_RE.match(lines[i])
if not bq:
break
body.append(bq.group(1))
i += 1
return ({"type": "blockquote", "body_lines": body}, i)
def parse_code_block(lines: list[str], start: int) -> tuple[dict[str, Any], int]:
"""Parse a fenced code block starting at start (which is the opening fence)."""
m = FENCE_RE.match(lines[start])
language = m.group(1).strip() if m else ""
body: list[str] = []
i = start + 1
while i < len(lines):
if FENCE_RE.match(lines[i]):
i += 1
break
body.append(lines[i])
i += 1
return ({"type": "code", "language": language, "body": "\n".join(body)}, i)
def parse_paragraph(lines: list[str], start: int) -> tuple[dict[str, Any], int]:
"""Collect consecutive non-empty, non-block lines into a paragraph."""
body: list[str] = []
i = start
while i < len(lines):
ln = lines[i]
if not ln.strip():
break
# Stop if we hit a block-level construct
if (HEADING_RE.match(ln) or FENCE_RE.match(ln) or HR_RE.match(ln) or
CALLOUT_RE.match(ln) or BLOCKQUOTE_RE.match(ln) or
ORDERED_LI_RE.match(ln) or UNORDERED_LI_RE.match(ln) or
TABLE_ROW_RE.match(ln)):
break
body.append(ln)
i += 1
text = " ".join(s.strip() for s in body)
return ({"type": "paragraph", "text": text}, i)
def parse_markdown(text: str) -> dict[str, Any]:
"""Top-level parse — returns {meta, blocks}."""
lines = text.splitlines()
blocks: list[dict[str, Any]] = []
i = 0
title = ""
heading_count = 0
section_count = 0
while i < len(lines):
line = lines[i]
if not line.strip():
i += 1
continue
# Heading
h = HEADING_RE.match(line)
if h:
level = len(h.group(1))
text_inline = h.group(2).strip()
anchor = slugify(text_inline)
heading_count += 1
if level == 1 and not title:
title = text_inline
if level >= 2:
section_count += 1
blocks.append({
"type": "heading",
"level": level,
"text": text_inline,
"anchor": anchor,
})
i += 1
continue
# HR
if HR_RE.match(line):
blocks.append({"type": "hr"})
i += 1
continue
# Fenced code
if FENCE_RE.match(line):
node, next_i = parse_code_block(lines, i)
blocks.append(node)
i = next_i
continue
# Callout (more specific than blockquote — must match first)
if CALLOUT_RE.match(line):
node, next_i = parse_callout(lines, i)
blocks.append(node)
i = next_i
continue
# Plain blockquote
if BLOCKQUOTE_RE.match(line):
node, next_i = parse_blockquote(lines, i)
blocks.append(node)
i = next_i
continue
# Table (header row + delim row check ahead)
if TABLE_ROW_RE.match(line) and i + 1 < len(lines) and TABLE_DELIM_RE.match(lines[i + 1]):
node, next_i = parse_table(lines, i)
blocks.append(node)
i = next_i
continue
# Lists
if ORDERED_LI_RE.match(line) or UNORDERED_LI_RE.match(line):
node, next_i = parse_list(lines, i)
blocks.append(node)
i = next_i
continue
# Paragraph (fallback)
node, next_i = parse_paragraph(lines, i)
blocks.append(node)
i = next_i
return {
"meta": {
"title": title,
"line_count": len(lines),
"heading_count": heading_count,
"section_count": section_count,
},
"blocks": blocks,
}
SAMPLE_MARKDOWN = """# Sample Specification
## Table of Contents
- Goals
- Architecture
- Risks
## Goals
We will integrate **Stripe Connect** with the existing checkout flow.
| Phase | Timeline | Owner |
|-------|----------|-------|
| Design | Week 1 | jane |
| Build | Week 2-3 | dev team |
| Ship | Week 4 | jane |
## Architecture
The integration uses webhooks for async events.
```python
def handle_webhook(event):
if event.type == "payment.succeeded":
mark_paid(event.data.object.id)
```
> [!NOTE]
> All webhook handlers must be idempotent.
> [!WARNING]
> Tax calculation has edge cases for digital goods in the EU.
## Risks
1. Webhook delivery delays
2. Tax calculation edge cases for VAT
3. Refund cascading across multi-party transfers
See [Stripe Connect docs](https://stripe.com/docs/connect) for details.
"""
def main(argv: list[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--input", help="Path to markdown file, or '-' for stdin")
parser.add_argument("--output", help="Path to write JSON output (else stdout)")
parser.add_argument("--sample", action="store_true",
help="Parse a built-in sample markdown document")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_MARKDOWN
elif args.input:
if args.input == "-":
text = sys.stdin.read()
else:
path = Path(args.input)
if not path.exists():
print(f"error: input not found: {path}", file=sys.stderr)
return 2
text = path.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = parse_markdown(text)
payload = json.dumps(result, indent=2)
if args.output:
Path(args.output).write_text(payload, encoding="utf-8")
print(f"wrote {args.output}: {result['meta']['heading_count']} headings, "
f"{len(result['blocks'])} blocks")
else:
print(payload)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Chuyển bài review code hoặc PR bằng markdown thành trang HTML hai cột: diff bên trái, thẻ chú thích gắn mức độ nghiêm trọng bên phải.
---
name: md-review
description: Converts a markdown PR writeup or code review (one with ```diff fenced blocks and severity-tagged > [!BLOCKER]/[!MAJOR]/[!MINOR]/[!NIT] callouts) into a single-file 2-column HTML review — unified-diff on the left, severity-tagged annotation cards on the right, top jump-nav listing every finding, mandatory named reviewer footer. Triggers when the markdown-html-orchestrator classifies an input as REVIEW, or when invoked directly via /cs:md-review. Refuses without explicit --reviewer (a code review must name a human), refuses if no diff hunks present (route to md-document instead), and refuses to encode severity in color only (every badge ships color + icon + aria-label per WCAG 1.4.1). Use after orchestrator routing.
version: 2.10.2
author: Alireza Rezvani
license: MIT
tags: [markdown, html, code-review, diff, severity, annotations, single-file, design-system, wcag-1.4.1]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# md-review — Code-review markdown → 2-column HTML
The code-review converter from Tier 2 of Shihipar's essay ("Code Review and PR Writeups"). Takes a markdown PR writeup with diff blocks + severity callouts and produces a single-file HTML review with a jump-nav, 2-column diff + annotation layout, and a named reviewer footer.
Three stdlib tools pipeline together:
```
diff_parser.py → annotation_extractor.py → review_html_renderer.py
(md → diff hunks) (md → severity-tagged (hunks + annotations
annotations attached + tokens → 2-col HTML)
to nearest hunk)
```
## When to invoke
| Symptom | Action |
|---|---|
| `markdown-html-orchestrator` routes input as REVIEW | Invoke this skill |
| User runs `/cs:md-review <path>.md` directly | Invoke this skill |
| Input contains ` ```diff ` fenced blocks + `> [!MAJOR]`/`> [!BLOCKER]`/etc. callouts | Invoke this skill |
| Input is a long-form spec / report (no diff blocks) | Route to `md-document` instead |
| Input is a slide deck | Route to `md-slides` instead |
| Input < 100 lines | Refuse (Shihipar threshold) |
| Design-system not onboarded | Refuse; surface `/cs:design-system` |
## Pipeline
```bash
# 1. Parse markdown → diff hunks JSON
python3 markdown-html/skills/md-review/scripts/diff_parser.py \
--input <path>.md --output hunks.json
# 2. Extract severity-tagged annotations, attach to nearest preceding hunk
python3 markdown-html/skills/md-review/scripts/annotation_extractor.py \
--input <path>.md --diff-blocks hunks.json --output annotations.json
# 3. Render 2-col HTML (--reviewer is mandatory — refuses without)
python3 markdown-html/skills/md-review/scripts/review_html_renderer.py \
--diff-blocks hunks.json --annotations annotations.json \
--reviewer "Jane Doe" --title "PR #123: Add retry logic" \
--output review.html
```
## What gets rendered
- **Top jump-nav** — every annotation with severity badge + 80-char preview + jump link; severity counts in the heading ("3 BLOCKER · 2 MAJOR · 1 NIT")
- **2-column hunk rows** — unified diff on the left (per-line old/new line numbers, +/− marks, addition/deletion background tint from design-system tokens), annotation cards on the right (color + icon + aria-label per WCAG 1.4.1)
- **Approval bar** — if `LGTM` markers are present and no severity annotations, a success-tinted "LGTM — no findings flagged" bar
- **General comments** — annotations not attached to any hunk render at the bottom in their own section
- **Reviewer footer** — mandatory; refuses to render without `--reviewer`
- **Responsive** — 2-col collapses to stacked on viewports < 900px
## Hard rules
1. **`--reviewer` is mandatory.** A code review must name a human reviewer. Refuses with exit 3 otherwise. Mirrors research-ops's "named owner" discipline.
2. **Refuses if no hunks present.** No `--- a/file` + `@@ ... @@` blocks means this isn't a code review — refuses with exit 4 and recommends `md-document`.
3. **Refuses input < 100 lines.** Markdown wins below the threshold (Shihipar).
4. **Refuses without onboarding.** Same gate as every converter.
5. **Severity is never color-only.** Each badge ships color + icon + `aria-label` + text. WCAG 1.4.1 enforced at the renderer level.
6. **Single-file output.** All CSS inline. Only external is Google Fonts CSS. No Prism in md-review (diff coloring conflicts with syntax highlighting).
7. **Custom severity convention.** `--severity-convention "critical,important,suggestion,nit"` swaps tier names; position 0 is most severe. Default is BLOCKER / MAJOR / MINOR / NIT (Google Code Review Developer Guide).
## Forcing-question library (Matt Pocock grill discipline)
1. **Who is the named reviewer?** Recommended: the user signing off on the review. Canon: research-ops named-owner pattern; *SWE at Google* ch. 9.
2. **Which severity convention applies — default (BLOCKER/MAJOR/MINOR/NIT) or custom?** Recommended: default unless your team has a documented alternative. Canon: Google *Code Review Developer Guide*.
3. **Are annotations anchored to specific hunks, or are some general?** Recommended: anchor everything you can; general goes to the unanchored section. Canon: *SWE at Google* ch. 9 — "Comments must reference a specific line".
4. **What's the PR title for the `<title>` and header?** Recommended: the actual PR / commit title. Canon: docs-as-context-for-readers.
5. **Should `LGTM` markers ship as the approval bar?** Recommended: yes if there are no severity annotations; otherwise the findings take precedence.
## Distinct from
- **`md-document`** — that converter renders prose + tables + code + callouts. This one renders diff hunks + margin annotations.
- **`md-slides`** — that converter splits on `---` boundaries. This one is a single-page artifact.
- **GitHub PR comments** — those are a thread. This is a single-author snapshot artifact.
## Output artifact
`{default_output_dir}/review-{slug}.html` (path resolved by orchestrator's `output_path_resolver.py`; collision suffix `-2`, `-3`, … by default).
## References
- Shihipar — *Claude Code HTML output* (Medium, 2026), Tier 2 use case
- *Software Engineering at Google* (Manshreck & Wright, O'Reilly 2020), ch. 9 "Code Review"
- Google *Code Review Developer Guide* — severity convention source
- WCAG 2.2 §1.4.1 — color-not-sole-signal enforcement
- See `references/` for full citations (diff_rendering_canon, severity_coding, pr_annotation_ux)
FILE:assets/md_review_template.html
<!DOCTYPE html>
<!--
md_review_template.html — Reference shape for review_html_renderer.py output.
Documents the canonical 2-column review HTML. The actual renderer produces
this same shape dynamically from a parsed diff (diff_parser.py) plus
annotations (annotation_extractor.py) plus the design-system config.
See: review_html_renderer.py for the live implementation.
-->
<html lang="en">
<head>
<meta charset="utf-8">
<title>{{PR_TITLE}}</title>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family={{HEADING_FONT}}&family={{BODY_FONT}}&display=swap">
<style>
:root {
/* 12 brand tokens from design-system.derived_palette */
--md-bg: {{BG}}; --md-surface: {{SURFACE}}; --md-border: {{BORDER}};
--md-text: {{TEXT}}; --md-text-muted: {{TEXT_MUTED}};
--md-accent: {{ACCENT}}; --md-accent-soft: {{ACCENT_SOFT}};
--md-code-bg: {{CODE_BG}}; --md-link: {{LINK}}; --md-link-hover: {{LINK_HOVER}};
--md-success: {{SUCCESS}}; --md-warn: {{WARN}};
/* Computed danger: accent hue rotated 120° toward red */
--md-danger: {{DANGER}};
}
/* ... BASE_CSS ... */
</style>
</head>
<body class="style-{{DESIGN_STYLE}}">
<header><h1>{{PR_TITLE}}</h1></header>
<!-- TOP JUMP-NAV: findings list with severity badges + preview + jump links -->
<nav class="jump-nav" aria-label="Review annotations">
<h2>Findings (1 BLOCKER · 2 MAJOR · 3 MINOR · 1 NIT)</h2>
<ul>
<li>
<span class="sev-badge" style="color: var(--md-danger)" role="status"
aria-label="Blocker — must fix before merge">
<span class="sev-icon" aria-hidden="true">■</span>BLOCKER
</span>
<a href="#ann-0">SQL injection — user input goes straight into a string-formatted query…</a>
<span class="nav-target">#1</span>
</li>
<!-- ... more findings ... -->
</ul>
</nav>
<!-- 2-COLUMN HUNK ROW: diff on left, annotation cards on right -->
<div class="hunk-row">
<div class="hunk" id="hunk-b0-f0-h0">
<div class="hunk-head">
<strong>payments/retry.py</strong>
<span>@10 → @10</span>
<span>def schedule_retry(payment_id, attempt)</span>
</div>
<pre>
<div class="line context">
<span class="lo">10</span><span class="ln">10</span>
<span class="mark" aria-hidden="true"> </span>
<span class="src"> if attempt > MAX_ATTEMPTS:</span>
</div>
<div class="line deletion">
<span class="lo">11</span><span class="ln"></span>
<span class="mark" aria-hidden="true">−</span>
<span class="src"> return _enqueue(payment_id, delay)</span>
</div>
<div class="line addition">
<span class="lo"></span><span class="ln">12</span>
<span class="mark" aria-hidden="true">+</span>
<span class="src"> jitter = random.uniform(0, delay * 0.1)</span>
</div>
</pre>
</div>
<div class="hunk-annotations">
<div class="annotation" id="ann-0" style="border-left-color: var(--md-warn)">
<div class="annotation-head">
<span class="sev-badge" style="color: var(--md-warn)"
role="status" aria-label="Major — strongly recommended fix">
<span class="sev-icon" aria-hidden="true">▲</span>MAJOR
</span>
<a href="#hunk-b0-f0-h0">jump to diff →</a>
</div>
<div class="annotation-body">
<code>random.uniform()</code> is not seeded — tests will be flaky.
</div>
</div>
</div>
</div>
<!-- UNANCHORED ANNOTATIONS: general comments not attached to any hunk -->
<h2>General comments</h2>
<div class="hunk-annotations">
<div class="annotation unanchored" id="ann-N">
<div class="annotation-head">
<span class="sev-badge" style="color: var(--md-text-muted)"
role="status" aria-label="Nit — cosmetic preference">
<span class="sev-icon" aria-hidden="true">◦</span>NIT
</span>
</div>
<div class="annotation-body">PR title could mention "retry" explicitly.</div>
</div>
</div>
<!-- MANDATORY REVIEWER FOOTER: refuses to render without --reviewer -->
<footer class="review-footer">
<span>Reviewer: <strong>{{REVIEWER_NAME}}</strong></span>
<span>· {{COMPANY_NAME}}</span>
<span style="margin-left:auto">Generated by markdown-html / md-review</span>
</footer>
</body>
</html>
FILE:references/diff_rendering_canon.md
# Diff Rendering Canon
**Why this exists:** The `md-review` converter renders unified diffs into a two-column HTML layout. Diff rendering has 50 years of UI history; this document records the conventions it inherits from and where it diverges.
## The convention
Unified-diff format (the `--- a/file`, `+++ b/file`, `@@ -10,7 +10,8 @@` shape) is the lingua franca every modern code-review tool reads:
- `+` (green tint): line added in the new version
- `-` (red tint): line removed from the old version
- ` ` (no tint): unchanged context line
- `@@ -old_start,old_count +new_start,new_count @@ optional_context`: hunk header
- `\ No newline at end of file`: meta line (rendered italic, no color)
`md-review` honors this exactly. The renderer assigns per-line numbers on both the old side (`lo`) and the new side (`ln`), with a clear visual separation by border.
## Color discipline
WCAG 1.4.1 (Use of Color) requires that color must not be the *sole* signal. Our diff rendering uses:
- **Tint backgrounds** (`color-mix(in srgb, var(--md-success) 15%, transparent)` for additions; `color-mix(... --md-warn 12% ...)` for deletions) — derived from the design-system palette, so they automatically follow the user's brand and stay within WCAG-validated contrast on the user's chosen bg.
- **Mark column** (`+` / `−` / ` `) — visible character, redundant with the color, screen-reader-skipped via `aria-hidden="true"`.
- **Line numbers in both columns** — provides spatial grounding independent of color.
Result: a deuteranopic reader (red-green color blindness) still distinguishes additions from deletions via the mark column and line-number columns.
## What we deliberately don't do
- **Syntax highlighting inside diffs** — Prism.js doesn't track per-line edits well, and conflicting addition/deletion backgrounds + token colors produce noisy output. Plain monospace is more legible. (Engineers reviewing diffs spend most attention on the change itself, not the surrounding language.)
- **Word-level diffing** — `difftastic`-style intra-line highlighting is great for tiny edits but ambiguous for refactors. We render whole lines and let the reader compare them visually.
- **Inline-reply threading** — md-review is a generator (markdown → HTML); it produces an artifact, not a thread. For threaded discussion, use the host platform (GitHub PR comments, GitLab discussions).
- **Split-pane (old | new) view** — the convention this converter targets is the *unified* diff (sequential, single column), because that's what gets pasted into markdown review notes. Split-pane is a tool for active reviewing in an IDE.
## Sources
### 1. POSIX `diff -u` (Single Unix Specification, c. 1990)
The format spec. The `@@ -A,B +C,D @@` hunk header, the `+`/`-`/` ` prefixes, the `--- a/` / `+++ b/` file headers — all defined here. Every tool that follows reads the same shape.
### 2. GitHub PR diff view
The reference UI for two decades. Established:
- Per-line line numbers on both old and new sides
- Tinted backgrounds for additions/deletions
- File-header sticky bar
- Hunk separator with grey background
We mirror the convention; we don't reinvent it.
### 3. GitLab MR diff view
Same convention as GitHub, with one small refinement: the per-hunk `@@` header context (the function name after the `@@`) is rendered as a sticky element so the reader knows what function they're in. We render this header context in the hunk-head bar.
### 4. `difftastic` — semantic diff tool (github.com/Wilfred/difftastic)
Argues for AST-aware intra-line diffing instead of line-based. We acknowledge the case but rejected it for two reasons: (1) parsing every language is out of scope; (2) most agent-generated review markdown contains a few short hunks, where intra-line diff adds noise more than signal.
### 5. *Software Engineering at Google* — Tom Manshreck & Hyrum Wright (O'Reilly, 2020), Ch. 9 "Code Review"
The discipline around how diffs get read. Three claims used here:
- "Reviewers spend most of their attention on a few hunks, not all of them" → jump-nav at top is essential
- "Comments must reference a specific line" → annotations are attached to hunks, not free-floating
- "Approval should be explicit and recorded" → `LGTM` markers are surfaced as the approval bar
### 6. Google *Code Review Developer Guide* (google.github.io/eng-practices/review/reviewer)
The taxonomy of severity (blocker, must-fix, nice-to-have, nit) maps to our default BLOCKER / MAJOR / MINOR / NIT convention. The phrase "nit:" specifically is from Google's recommended review vocabulary.
### 7. GitHub Markdown — fenced code blocks with language hints (` ```diff `)
The convention this converter targets. Agent-generated review markdown wraps every diff in ` ```diff `, which becomes our extraction grammar in `diff_parser.py`.
## Applied to `md-review`
`diff_parser.py` extracts the hunks from ` ```diff ` fenced blocks. `review_html_renderer.py` lays them out in the standard unified-diff shape, with the addition/deletion coloring derived from the design-system palette so reviews look on-brand without losing the universal color convention.
FILE:references/pr_annotation_ux.md
# PR Annotation UX
**Why this exists:** The 2-column layout (diff on the left, annotation cards on the right) isn't arbitrary — it's the convergent UX that emerged after 15 years of PR tools experimenting with placement. This document records why we render this shape and not the alternatives.
## The shape
| Region | Content | Why |
|---|---|---|
| Top: jump-nav | Every annotation listed with severity badge + 1-line preview + jump link | Reviewers skim the findings list before reading any hunk (`SWE at Google`, ch. 9) |
| Left column: diff | Unified-diff lines with line numbers and +/− coloring | The work being reviewed; gets the bigger column |
| Right column: annotation cards | Severity badge + body text + "jump to diff" backlink | Anchored to the diff but visually distinct |
| Bottom: unanchored | Annotations not attached to any hunk | Folder-level / general comments live here |
| Footer: reviewer | "Reviewer: <name>" mandatory | A code review without a named reviewer is not a code review |
## Why 2-col and not stacked
Alternatives we considered:
1. **Inline annotations between diff lines** (the GitHub PR style). Pros: maximum proximity between annotation and code. Cons: breaks the diff's visual continuity; hard to skim the diff without re-encountering every annotation; cards force the diff to scroll.
2. **Annotations in a sidebar separate from any hunk** (Jira's review style). Pros: clean diff. Cons: the spatial relationship between annotation and code is lost; reviewer has to mentally re-anchor each annotation.
3. **2-col: diff left, annotations right, aligned to first hunk of the block.** What we do. Pros: diff stays continuous and skimmable; annotation is visible at the same scroll position as the relevant hunk; on narrow screens (< 900px), the layout collapses to stacked. Cons: when there are many annotations per block, they stack vertically beyond the hunk height — but this is the right failure mode (more findings = more space, not hidden).
## The jump-nav specifically
The top-of-page findings list is the highest-leverage element. Empirically (every tool that ships this — GitHub, GitLab, Reviewable, CodeStream — has converged on it):
- Reviewers + authors both look here first
- A 4-finding PR is easier to triage from the list than from scrolling
- The author can use the list to mentally check off as they address each finding
Each list entry has: severity badge + 80-char preview + "(#1)" target reference. Click jumps to the annotation card.
## What we deliberately don't do
- **Resolve/Unresolve workflow** — md-review generates an artifact, not a thread. Resolution belongs in the host platform.
- **Avatar / author images** — out of scope; this is a single-author review (the `--reviewer` field names them).
- **Filtering by severity** — the jump-nav lists them all; severity badges already group them visually. Filter UI would add JS complexity without much value for a typical 3-10-finding review.
- **Cross-PR comparison** — single review, single artifact.
- **Re-review threading** — same reason; artifact, not thread.
## Sources
### 1. GitHub PR review UI
The reference implementation that established expectations: per-line inline comments, severity through icon (😄 / 👍 / 🚀 reactions), file/folder tree on the left, top-level summary. Our 2-col layout is a simplification: we keep the spatial proximity but drop the inline-thread complication.
### 2. GitLab MR review UI
Similar to GitHub. Adds: sticky hunk-context header (the function name after `@@`). We render this in the hunk-head.
### 3. Reviewable.io
Argued for a more structured review process with explicit annotation state (resolved, deferred, addressed). Our artifact doesn't have state — it's a snapshot of one reviewer's findings — but the jump-nav with per-finding "(#N)" reference is borrowed from Reviewable's "comment numbering" UX.
### 4. CodeStream (now part of New Relic)
Pioneered the margin-comment pattern in IDE plugins: code in the editor, comments in a right-side panel aligned to the relevant line. The 2-col layout in md-review mirrors that pattern for a single-file artifact.
### 5. *Software Engineering at Google* (Manshreck & Wright, O'Reilly 2020), Ch. 9 "Code Review"
- "The review summary at the top is the most-read element"
- "Comments must be anchored to specific lines"
- "Approval should be explicit"
Three claims, three corresponding UI elements (jump-nav, hunk-anchored annotations, LGTM/approval bar).
### 6. *People + AI Guidebook* (Google PAIR, pair.withgoogle.com)
For AI-assisted code review specifically: the artifact should make clear who the reviewer is (the `--reviewer` field is mandatory for exactly this reason). An AI-generated review without a named human reviewer is an artifact looking for ownership; we refuse to render it.
### 7. NN/g — *Web Page Scanning Patterns* (Jakob Nielsen, 2006, updated 2024)
The F-shape reading pattern: users fixate on the top + left edges first. Our jump-nav (top) + diff (left column) + annotation (right column) honors this — the reader sees what's most important first.
## Applied to `md-review`
`review_html_renderer.py` emits the jump-nav at top, the 2-col `.hunk-row` for each diff block, and the mandatory reviewer footer. It refuses to render without `--reviewer` (rule 1) and refuses to render without any hunks (rule 2 — wrong skill, route to md-document). Resolution / threading / cross-PR comparison are out of scope.
FILE:references/severity_coding.md
# Severity Coding
**Why this exists:** Code-review annotations need a severity dimension or every comment carries the same weight. This skill ships a default 4-tier convention (BLOCKER / MAJOR / MINOR / NIT), accepts custom conventions via `--severity-convention`, and enforces WCAG 1.4.1 (color is not the sole signal — every badge has color + icon + aria-label).
## The default convention
| Tier | Meaning | Source | Visual |
|---|---|---|---|
| **BLOCKER** | Must fix before merge. Author cannot proceed without addressing. | Found in many engineering teams' written conventions; Google calls this "must-fix" | Filled square ■ + derived danger color (accent rotated 120° toward red) |
| **MAJOR** | Strongly recommended. Author should address unless they have a good counter-argument. | Common in Google's *Code Review Developer Guide* as the "non-nit substantive comment" | Filled triangle ▲ + `--md-warn` |
| **MINOR** | Worth fixing. Reasonable to address now; reasonable to defer. | Common middle tier | Filled circle ● + `--md-link` |
| **NIT** | Cosmetic / style preference. Author may address or ignore. | Google's nit: prefix is the canonical example | Open circle ◦ + `--md-text-muted` |
Position 0 (BLOCKER) is most severe; position 3 (NIT) is least. Position determines the `severity_rank` field in the JSON output, which the renderer uses for sort + display order.
## Custom conventions
`--severity-convention "critical,important,suggestion,nit"` swaps the tier list. Same rank ordering applies (position 0 = most severe). Renderer uses the same icon/color mapping when the tier name matches a default; otherwise falls back to the `_text-muted` token + open-circle icon for the unknown tier.
## WCAG 1.4.1 — color is not the sole signal
Every severity badge ships:
1. **Color** — derived from the design-system palette, validated for AA contrast against bg.
2. **Icon** — ■ / ▲ / ● / ◦ — a glyph that's distinct shape-wise. (Deuteranopic readers see the shape difference even if the color difference is muted.)
3. **`aria-label`** — full text spelled out: "Blocker — must fix before merge". Screen readers announce this; sighted readers don't see it but get the icon + color + text.
4. **Text label** — the severity name itself ("BLOCKER") is in the badge. Triple-redundancy: color, icon, text.
A reader who is completely color-blind, or viewing on a grayscale monitor, or screen-reading the page, still has full access to severity.
## What we deliberately don't do
- **Emoji icons** — render inconsistently across OS / font (Windows vs macOS vs Linux vs iOS), break in monochrome printing. We use Unicode geometric shapes (■▲●◦) that ship with every system font.
- **Color-only badges** — the WCAG floor.
- **Severity arithmetic** — no "this PR has score X = blockers × 10 + majors × 3 + …" auto-rollup. Review quality is qualitative; numeric rollups create gaming incentives.
- **Auto-block merge based on severity** — that's a CI concern, not a renderer concern.
## Sources
### 1. WCAG 2.2 §1.4.1 *Use of Color* (w3.org/WAI/WCAG22)
The hard rule: color must not be the only visual means of conveying information. Our color + icon + label + aria-label is the canonical implementation.
### 2. Google *Code Review Developer Guide* (google.github.io/eng-practices/review/reviewer)
Source of the `nit:` prefix convention and the "blocker / major / minor / nit" taxonomy. We adopt it as the default.
### 3. *Software Engineering at Google* (Manshreck & Wright, O'Reilly 2020), Ch. 9
The taxonomy is documented here: blockers are tickets that fail review; nits are "I'd prefer X but it's not a blocker"; majors are substantive comments.
### 4. Phabricator / Sourcegraph / Reviewable.io
Each tool ships its own severity vocabulary; ours is compatible by adopting the most-common 4-tier convention.
### 5. Don Norman — *The Design of Everyday Things* (2013 ed., Basic Books)
The "signifier" concept: visual elements must communicate function. A red badge is a signifier; a red badge that says "BLOCKER" with a filled square is a stronger signifier; a red badge with `aria-label="Blocker — must fix before merge"` adds the non-visual channel.
### 6. NN/g — *Color in UX* (Therese Fessenden, 2024)
Empirical: relying on color alone fails for 8% of male readers (red-green color blindness). The triple-redundancy approach is standard.
### 7. Anil Dash — *The Web We Lost* (dashes.com, 2012)
Indirectly: portable web artifacts. A single-file review must render correctly even when CSS variables fail to load — which is why the badge text label (in addition to color + icon) is essential. If `var(--md-warn)` doesn't resolve, the user still sees "MAJOR" with a triangle.
## Applied to `md-review`
The `_render_severity_badge` helper in `review_html_renderer.py` emits the color + icon + label + aria-label combo for every severity, derived from the design-system palette. The convention is overridable via `--severity-convention`. Refuses to render a badge with color alone.
FILE:scripts/annotation_extractor.py
#!/usr/bin/env python3
"""annotation_extractor.py - Extract severity-tagged review annotations from markdown.
Stdlib-only. Scans the same markdown source the diff_parser walks, finds
severity callouts and inline review markers, and attaches each annotation to
the nearest preceding diff block. The result is what the renderer puts in
the right margin of the 2-column layout.
Severity conventions accepted:
GFM callouts (preferred):
> [!BLOCKER] must-fix before merge
> [!MAJOR] strongly recommend addressing
> [!MINOR] worth fixing
> [!NIT] cosmetic / style preference
Inline markers (legacy, less structured):
blocker: <prose>
major: <prose>
minor: <prose>
nit: <prose>
LGTM (treated as APPROVAL marker; not severity-coded)
The --severity-convention flag accepts a custom 4-tier ordering, e.g.
"critical,important,suggestion,nit" — but defaults to the BLOCKER/MAJOR/
MINOR/NIT convention. The order matters: position 0 = most severe.
Attachment heuristic: each annotation attaches to the most recent diff block
that appeared above it in the markdown source (by source line number). If
no diff appears above, the annotation is "unanchored" and the renderer
shows it in a "general comments" section.
NO LLM CALLS. Pure regex + line-index attachment.
Usage:
python annotation_extractor.py --input review.md --diff-blocks hunks.json
python annotation_extractor.py --sample
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any
DEFAULT_SEVERITY_CONVENTION = ["BLOCKER", "MAJOR", "MINOR", "NIT"]
CALLOUT_OPEN_RE = re.compile(r"^>\s*\[!([A-Z]+)\]\s*$")
BLOCKQUOTE_RE = re.compile(r"^>\s?(.*)$")
INLINE_MARKER_RE = re.compile(r"^(?P<sev>[A-Za-z]+):\s+(?P<body>.+)$")
APPROVAL_RE = re.compile(r"^(LGTM|👍|approved|approve)\s*$", re.IGNORECASE)
def extract_annotations(
text: str,
diff_blocks: dict[str, Any] | None,
severity_convention: list[str],
) -> dict[str, Any]:
"""Returns {annotations: [...], approvals: [...], summary: {...}}."""
lines = text.splitlines()
upper_severities = {s.upper() for s in severity_convention}
# Index of diff-block source lines, for attachment lookup
diff_source_lines: list[tuple[int, int]] = [] # (source_line, block_index)
if diff_blocks:
for b in diff_blocks.get("blocks", []):
diff_source_lines.append((b["source_line"], b["block_index"]))
def nearest_preceding_block(line_index: int) -> int | None:
best: int | None = None
for src_line, block_idx in diff_source_lines:
if src_line <= line_index:
best = block_idx
else:
break
return best
annotations: list[dict[str, Any]] = []
approvals: list[dict[str, Any]] = []
i = 0
while i < len(lines):
ln = lines[i]
# GFM callout (multi-line)
m_open = CALLOUT_OPEN_RE.match(ln)
if m_open:
severity_raw = m_open.group(1).upper()
body_lines: list[str] = []
j = i + 1
while j < len(lines):
bq = BLOCKQUOTE_RE.match(lines[j])
if not bq:
break
body_lines.append(bq.group(1))
j += 1
if severity_raw in upper_severities:
annotations.append({
"kind": "callout",
"severity": severity_raw,
"severity_rank": severity_convention.index(
next(s for s in severity_convention if s.upper() == severity_raw)
),
"body": " ".join(b.strip() for b in body_lines).strip(),
"source_line": i,
"attached_block": nearest_preceding_block(i),
})
i = j
continue
# Approval marker (LGTM, etc.)
if APPROVAL_RE.match(ln.strip()):
approvals.append({
"kind": "approval",
"marker": ln.strip(),
"source_line": i,
"attached_block": nearest_preceding_block(i),
})
i += 1
continue
# Inline marker (single-line, prose-leading)
m_inline = INLINE_MARKER_RE.match(ln.strip())
if m_inline:
sev_raw = m_inline.group("sev").upper()
if sev_raw in upper_severities:
annotations.append({
"kind": "inline",
"severity": sev_raw,
"severity_rank": severity_convention.index(
next(s for s in severity_convention if s.upper() == sev_raw)
),
"body": m_inline.group("body").strip(),
"source_line": i,
"attached_block": nearest_preceding_block(i),
})
i += 1
# Sort annotations primarily by source_line (preserves narrative order
# in the renderer's jump-nav), then group by severity in summary.
annotations.sort(key=lambda a: a["source_line"])
counts_by_severity: dict[str, int] = {}
for a in annotations:
counts_by_severity[a["severity"]] = counts_by_severity.get(a["severity"], 0) + 1
return {
"annotations": annotations,
"approvals": approvals,
"summary": {
"total_annotations": len(annotations),
"total_approvals": len(approvals),
"counts_by_severity": counts_by_severity,
"severity_convention": severity_convention,
},
}
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.split("\n")[0])
p.add_argument("--input", help="Path to markdown file, or '-' for stdin")
p.add_argument("--diff-blocks", help="Path to diff_parser JSON output (for attachment)")
p.add_argument("--severity-convention",
default=",".join(DEFAULT_SEVERITY_CONVENTION),
help="Comma-separated severity tier list, most-to-least severe. "
"Default: BLOCKER,MAJOR,MINOR,NIT")
p.add_argument("--output", help="Path to write JSON output (else stdout)")
p.add_argument("--sample", action="store_true",
help="Run on a built-in sample PR review")
args = p.parse_args(argv)
severity_convention = [s.strip().upper() for s in args.severity_convention.split(",")]
if len(severity_convention) < 2:
print("error: --severity-convention needs at least 2 tiers", file=sys.stderr)
return 2
if args.sample:
sys.path.insert(0, str(Path(__file__).resolve().parent))
import diff_parser
text = diff_parser.SAMPLE_MARKDOWN
diff_blocks = diff_parser.parse_markdown_for_diffs(text)
elif args.input:
text = sys.stdin.read() if args.input == "-" else Path(args.input).read_text(encoding="utf-8")
diff_blocks = (
json.loads(Path(args.diff_blocks).read_text(encoding="utf-8"))
if args.diff_blocks else None
)
else:
p.print_help()
return 0
result = extract_annotations(text, diff_blocks, severity_convention)
payload = json.dumps(result, indent=2)
if args.output:
Path(args.output).write_text(payload, encoding="utf-8")
print(f"wrote {args.output}: {result['summary']['total_annotations']} annotations, "
f"{result['summary']['total_approvals']} approvals, "
f"counts={result['summary']['counts_by_severity']}")
else:
print(payload)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/diff_parser.py
#!/usr/bin/env python3
"""diff_parser.py - Extract unified-diff hunks from markdown code review notes.
Stdlib-only. Scans markdown for ```diff fenced code blocks (the convention used
in PR writeups), parses each one as a unified diff, and returns structured
hunk data the renderer can lay out as a 2-column annotated diff.
Pattern (the standard unified-diff shape):
```diff
--- a/path/to/file.py
+++ b/path/to/file.py
@@ -10,7 +10,8 @@ def existing_context
context line (unchanged)
-removed line
+added line
+another added line
context line
@@ -50,3 +51,4 @@
...
```
Also accepts:
- Multiple ```diff blocks (each treated as a separate "section")
- Inline-fenced diffs (no language tag) IF --infer-diff is passed
- Per-hunk @@ header context strings (preserved in output)
NO LLM CALLS. Pure regex + state machine.
Usage:
python diff_parser.py --input review.md --output hunks.json
python diff_parser.py --sample
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any
FENCE_OPEN_RE = re.compile(r"^```(diff)?\s*$")
FENCE_CLOSE_RE = re.compile(r"^```\s*$")
FILE_OLD_RE = re.compile(r"^---\s+(?:a/)?(.+?)\s*$")
FILE_NEW_RE = re.compile(r"^\+\+\+\s+(?:b/)?(.+?)\s*$")
HUNK_RE = re.compile(
r"^@@\s+-(\d+)(?:,(\d+))?\s+\+(\d+)(?:,(\d+))?\s+@@\s*(.*)$"
)
def _parse_hunk_body(lines: list[str], old_start: int, new_start: int) -> list[dict[str, Any]]:
"""Walk hunk body lines and assign source-side and target-side line numbers."""
body: list[dict[str, Any]] = []
old_num = old_start
new_num = new_start
for ln in lines:
if not ln:
# treat as context (truly empty line in diff)
body.append({"kind": "context", "text": "", "old": old_num, "new": new_num})
old_num += 1
new_num += 1
continue
prefix = ln[0]
text = ln[1:] if len(ln) > 1 else ""
if prefix == "+":
body.append({"kind": "addition", "text": text, "old": None, "new": new_num})
new_num += 1
elif prefix == "-":
body.append({"kind": "deletion", "text": text, "old": old_num, "new": None})
old_num += 1
elif prefix == " ":
body.append({"kind": "context", "text": text, "old": old_num, "new": new_num})
old_num += 1
new_num += 1
elif prefix == "\\":
body.append({"kind": "meta", "text": text.strip(), "old": None, "new": None})
else:
# Unknown line — preserve as context with the raw prefix
body.append({"kind": "context", "text": ln, "old": old_num, "new": new_num})
old_num += 1
new_num += 1
return body
def _parse_single_diff(block_lines: list[str]) -> list[dict[str, Any]]:
"""Parse one fenced diff block. Returns a list of file entries.
File entry shape:
{
"path_old": str | None,
"path_new": str | None,
"hunks": [
{"old_start": int, "new_start": int, "header_context": str,
"lines": [...], "hunk_index_in_block": int}
]
}
"""
files: list[dict[str, Any]] = []
current_file: dict[str, Any] | None = None
current_hunk: dict[str, Any] | None = None
hunk_body: list[str] = []
hunk_index_in_block = 0
def flush_hunk() -> None:
nonlocal current_hunk, hunk_body, hunk_index_in_block
if current_hunk is None or current_file is None:
current_hunk = None
hunk_body = []
return
current_hunk["lines"] = _parse_hunk_body(
hunk_body, current_hunk["old_start"], current_hunk["new_start"]
)
current_hunk["hunk_index_in_block"] = hunk_index_in_block
hunk_index_in_block += 1
current_file["hunks"].append(current_hunk)
current_hunk = None
hunk_body = []
def flush_file() -> None:
nonlocal current_file
flush_hunk()
if current_file:
files.append(current_file)
current_file = None
for raw in block_lines:
m_old = FILE_OLD_RE.match(raw)
m_new = FILE_NEW_RE.match(raw)
m_hunk = HUNK_RE.match(raw)
if m_old:
# New file boundary; flush previous
flush_file()
current_file = {"path_old": m_old.group(1), "path_new": None, "hunks": []}
continue
if m_new:
if current_file is None:
current_file = {"path_old": None, "path_new": m_new.group(1), "hunks": []}
else:
current_file["path_new"] = m_new.group(1)
continue
if m_hunk:
flush_hunk()
current_hunk = {
"old_start": int(m_hunk.group(1)),
"old_count": int(m_hunk.group(2) or 1),
"new_start": int(m_hunk.group(3)),
"new_count": int(m_hunk.group(4) or 1),
"header_context": m_hunk.group(5).strip(),
"lines": [],
}
continue
if current_hunk is not None:
hunk_body.append(raw)
flush_file()
return files
def parse_markdown_for_diffs(text: str, infer_unfenced: bool = False) -> dict[str, Any]:
"""Top-level: find every ```diff block and parse it.
Returns:
{
"blocks": [
{"block_index": int, "files": [...]},
...
],
"summary": {"total_files": int, "total_hunks": int, "total_blocks": int}
}
"""
lines = text.splitlines()
in_block = False
is_diff_block = False
block_buf: list[str] = []
blocks: list[dict[str, Any]] = []
block_index = 0
block_line_starts: list[int] = []
i = 0
while i < len(lines):
ln = lines[i]
if not in_block:
m = FENCE_OPEN_RE.match(ln)
if m:
in_block = True
lang = m.group(1)
is_diff_block = (lang == "diff")
block_buf = []
block_line_starts.append(i)
else:
if FENCE_CLOSE_RE.match(ln):
if is_diff_block:
files = _parse_single_diff(block_buf)
blocks.append({
"block_index": block_index,
"source_line": block_line_starts[block_index],
"files": files,
})
block_index += 1
elif infer_unfenced:
# Try to detect if this looks like a diff (starts with --- / +++ / @@)
head = "\n".join(block_buf[:3])
if "@@" in head or head.startswith("---") or head.startswith("+++"):
files = _parse_single_diff(block_buf)
blocks.append({
"block_index": block_index,
"source_line": block_line_starts[block_index],
"files": files,
"inferred": True,
})
block_index += 1
in_block = False
is_diff_block = False
block_buf = []
else:
block_buf.append(ln)
i += 1
total_files = sum(len(b["files"]) for b in blocks)
total_hunks = sum(
sum(len(f["hunks"]) for f in b["files"]) for b in blocks
)
return {
"blocks": blocks,
"summary": {
"total_files": total_files,
"total_hunks": total_hunks,
"total_blocks": len(blocks),
},
}
SAMPLE_MARKDOWN = """# PR Review: Add payment retry logic
Two changes worth flagging.
```diff
--- a/payments/retry.py
+++ b/payments/retry.py
@@ -10,7 +10,8 @@ def schedule_retry(payment_id, attempt):
if attempt > MAX_ATTEMPTS:
return None
delay = 2 ** attempt
- return _enqueue(payment_id, delay)
+ jitter = random.uniform(0, delay * 0.1)
+ return _enqueue(payment_id, delay + jitter)
```
> [!MAJOR]
> This `random.uniform()` call is not seeded. In test runs it'll produce
> non-deterministic retry delays — make tests flaky.
```diff
--- a/payments/queue.py
+++ b/payments/queue.py
@@ -40,3 +40,7 @@ def _enqueue(payment_id, delay):
redis.zadd("retries", {payment_id: time.time() + delay})
log.info("scheduled", payment_id=payment_id, delay=delay)
+
+def cancel_retries(payment_id):
+ redis.zrem("retries", payment_id)
+ log.info("cancelled retries", payment_id=payment_id)
```
> [!NIT]
> Minor — would prefer `log.info("retries.cancelled", ...)` (dotted event name).
"""
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.split("\n")[0])
p.add_argument("--input", help="Path to markdown file, or '-' for stdin")
p.add_argument("--output", help="Path to write JSON output (else stdout)")
p.add_argument("--infer-diff", action="store_true",
help="Also parse unfenced or language-less blocks that look like diffs")
p.add_argument("--sample", action="store_true",
help="Parse a built-in sample PR-review markdown")
args = p.parse_args(argv)
if args.sample:
text = SAMPLE_MARKDOWN
elif args.input:
text = sys.stdin.read() if args.input == "-" else Path(args.input).read_text(encoding="utf-8")
else:
p.print_help()
return 0
result = parse_markdown_for_diffs(text, infer_unfenced=args.infer_diff)
payload = json.dumps(result, indent=2)
if args.output:
Path(args.output).write_text(payload, encoding="utf-8")
print(f"wrote {args.output}: {result['summary']['total_files']} files, "
f"{result['summary']['total_hunks']} hunks across "
f"{result['summary']['total_blocks']} diff blocks")
else:
print(payload)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/review_html_renderer.py
#!/usr/bin/env python3
"""review_html_renderer.py - Render parsed diffs + annotations into a 2-column HTML review.
Stdlib-only. Combines:
- diff_parser.py output (hunks per file)
- annotation_extractor.py output (severity-tagged margin notes)
- design-system config (12-token palette + typography + design_style)
into a single-file HTML page with:
* Top jump-nav: every annotation listed with severity badge + file:line + 1-line preview
(Click jumps to the annotation in the right margin and highlights the hunk)
* 2-column layout: diff on left, annotation cards on right
(Falls back to stacked layout on viewports < 900px)
* Per-line diff coloring: additions in success-tinted bg, deletions in warn-tinted bg
* Severity badges: icon + color + aria-label (WCAG 1.4.1 — color is NOT the sole signal)
* Mandatory "Reviewer:" footer (refuses to render without --reviewer)
NO LLM CALLS. Pure templating + config-driven CSS.
Single-file output: all CSS + (optional) JS inline. Only external is Google
Fonts CSS for typography (same discipline as md-document; no Prism here —
we render diff coloring ourselves with stable conventions).
Usage:
python review_html_renderer.py \\
--diff-blocks hunks.json \\
--annotations annotations.json \\
--reviewer "Jane Doe" \\
--output review.html
python review_html_renderer.py --sample --reviewer "Sample Reviewer" --output /tmp/review.html
"""
from __future__ import annotations
import argparse
import html
import json
import os
import sys
from pathlib import Path
from typing import Any
_DESIGN_SYSTEM_SCRIPTS = (
Path(__file__).resolve().parent.parent.parent / "design-system" / "scripts"
)
sys.path.insert(0, str(_DESIGN_SYSTEM_SCRIPTS))
try:
import config_loader as _cfg
import brand_palette_validator as _bpv
except ImportError:
_cfg = None
_bpv = None
# ----- Severity → visual mapping (color + icon + aria-label) -------------------
# Color is NOT the sole signal (WCAG 1.4.1). Every badge has an icon and an
# aria-label that's announced by screen readers.
SEVERITY_DEFAULTS = {
# severity_name: {token_key_in_palette, icon_text, aria_phrase}
"BLOCKER": {"token": "_danger", "icon": "■", "aria": "Blocker — must fix before merge"},
"MAJOR": {"token": "--md-warn", "icon": "▲", "aria": "Major — strongly recommended fix"},
"MINOR": {"token": "--md-link", "icon": "●", "aria": "Minor — worth fixing"},
"NIT": {"token": "--md-text-muted", "icon": "◦", "aria": "Nit — cosmetic preference"},
}
def _derive_danger_color(palette: dict[str, str]) -> str:
"""Compute a 'danger' color by rotating the accent hue toward red.
Falls back to a generic red if the palette is unavailable.
"""
if not palette or _bpv is None:
return "#D04646"
accent_hex = palette.get("--md-accent", "#D04646")
try:
rgb = _bpv.parse_hex(accent_hex)
# Rotate hue toward red (0°) — pick the shorter rotation that lands near red
target = _bpv.shift_hue(rgb, -120) # rotate 120° toward red
return _bpv.rgb_to_hex(target)
except Exception:
return "#D04646"
def _resolve_severity_color(severity: str, palette: dict[str, str], danger: str) -> str:
sev = severity.upper()
spec = SEVERITY_DEFAULTS.get(sev, {"token": "--md-text-muted"})
token = spec["token"]
if token == "_danger":
return danger
return palette.get(token, "#888888")
def _palette_to_css(palette: dict[str, str]) -> str:
if not palette:
palette = {
"--md-bg": "#0E1E38", "--md-surface": "#142B50", "--md-border": "#1A3868",
"--md-text": "#F7F7F2", "--md-text-muted": "rgba(247, 247, 242, 0.68)",
"--md-accent": "#00D4AA", "--md-accent-soft": "rgba(0, 212, 170, 0.14)",
"--md-code-bg": "#122648",
"--md-link": "#00D4AA", "--md-link-hover": "#08FECE",
"--md-success": "#10A85C", "--md-warn": "#C87C10",
}
return "\n".join(f" {k}: {v};" for k, v in palette.items())
def _font_url(heading: str, body: str) -> str:
families = sorted({heading, body})
parts = "&".join(f"family={f.replace(' ', '+')}:wght@400;600" for f in families)
return f"https://fonts.googleapis.com/css2?{parts}&display=swap"
def _font_stack(name: str) -> str:
fallback = ("Georgia, serif" if name in
("Playfair Display", "Merriweather", "Lora", "Source Serif 4")
else "system-ui, -apple-system, sans-serif")
return f"'{name}', {fallback}"
BASE_CSS = """
:root {
__PALETTE__
--md-danger: __DANGER__;
--md-font-heading: __HEADING_FONT__;
--md-font-body: __BODY_FONT__;
--md-font-mono: 'JetBrains Mono', ui-monospace, SFMono-Regular, Menlo, monospace;
}
* { box-sizing: border-box; }
html { scroll-behavior: smooth; }
body {
margin: 0;
padding: 2rem 1.5rem;
background: var(--md-bg);
color: var(--md-text);
font-family: var(--md-font-body);
font-size: 16px;
line-height: 1.55;
max-width: 1400px;
margin-left: auto;
margin-right: auto;
}
h1, h2, h3 { font-family: var(--md-font-heading); color: var(--md-text); margin: 0 0 0.5em; line-height: 1.25; }
h1 { font-size: 1.75rem; }
h2 { font-size: 1.25rem; margin-top: 2rem; padding-bottom: 0.3em; border-bottom: 1px solid var(--md-border); }
a { color: var(--md-link); }
/* Severity badges (color + icon + aria-label — WCAG 1.4.1) */
.sev-badge {
display: inline-flex;
align-items: center;
gap: 0.35em;
font-family: var(--md-font-heading);
font-weight: 600;
font-size: 0.75rem;
letter-spacing: 0.05em;
padding: 0.15em 0.55em;
border-radius: 999px;
border: 1px solid currentColor;
background: var(--md-bg);
}
.sev-icon { font-family: var(--md-font-mono); font-size: 0.875em; }
/* Top jump-nav */
nav.jump-nav {
background: var(--md-surface);
border: 1px solid var(--md-border);
border-radius: 10px;
padding: 1rem 1.25rem;
margin: 1.5rem 0 2rem;
}
nav.jump-nav h2 { margin-top: 0; font-size: 1rem; border: none; padding: 0; }
nav.jump-nav ul { list-style: none; padding: 0; margin: 0.75rem 0 0; display: grid; gap: 0.4rem; }
nav.jump-nav li { display: flex; align-items: center; gap: 0.75rem; flex-wrap: wrap; }
nav.jump-nav a {
color: var(--md-text);
text-decoration: none;
flex: 1;
border-radius: 4px;
padding: 0.2em 0.4em;
}
nav.jump-nav a:hover { background: var(--md-accent-soft); color: var(--md-accent); }
nav.jump-nav .nav-target {
color: var(--md-text-muted);
font-family: var(--md-font-mono);
font-size: 0.875rem;
}
nav.jump-nav .nav-preview { color: var(--md-text-muted); font-size: 0.875rem; }
/* 2-column hunk layout */
.hunk-row {
display: grid;
grid-template-columns: minmax(0, 1fr) 320px;
gap: 1.25rem;
margin: 1.5rem 0 2rem;
scroll-margin-top: 1rem;
}
@media (max-width: 900px) {
.hunk-row { grid-template-columns: 1fr; }
.hunk-annotations { order: 2; }
}
.hunk {
background: var(--md-code-bg);
border: 1px solid var(--md-border);
border-radius: 8px;
overflow: hidden;
}
.hunk-head {
background: var(--md-surface);
border-bottom: 1px solid var(--md-border);
padding: 0.5rem 0.85rem;
font-family: var(--md-font-mono);
font-size: 0.8125rem;
color: var(--md-text-muted);
display: flex;
flex-wrap: wrap;
gap: 0.75rem;
}
.hunk-head strong { color: var(--md-text); font-weight: 600; }
.hunk pre {
margin: 0;
padding: 0;
background: transparent;
font-family: var(--md-font-mono);
font-size: 0.8125rem;
line-height: 1.55;
overflow-x: auto;
}
.hunk .line {
display: grid;
grid-template-columns: 3.5em 3.5em 1.25em 1fr;
padding-right: 0.75rem;
}
.hunk .ln, .hunk .lo {
color: var(--md-text-muted);
padding: 0 0.5em;
text-align: right;
user-select: none;
border-right: 1px solid var(--md-border);
font-size: 0.75rem;
}
.hunk .mark { text-align: center; user-select: none; color: var(--md-text-muted); }
.hunk .src { padding-left: 0.5em; white-space: pre; overflow-wrap: normal; }
.hunk .line.addition { background: color-mix(in srgb, var(--md-success) 15%, transparent); }
.hunk .line.addition .mark { color: var(--md-success); }
.hunk .line.deletion { background: color-mix(in srgb, var(--md-warn) 12%, transparent); }
.hunk .line.deletion .mark { color: var(--md-warn); }
.hunk .line.meta { color: var(--md-text-muted); font-style: italic; padding-left: 1em; }
/* Annotation cards on the right */
.hunk-annotations { display: grid; gap: 0.75rem; align-content: start; }
.annotation {
background: var(--md-surface);
border: 1px solid var(--md-border);
border-left-width: 4px;
border-radius: 8px;
padding: 0.75rem 0.85rem;
}
.annotation .annotation-head {
display: flex;
justify-content: space-between;
align-items: center;
gap: 0.5rem;
margin-bottom: 0.4rem;
}
.annotation .annotation-body { font-size: 0.9375rem; line-height: 1.5; color: var(--md-text); }
.annotation .annotation-body code {
font-family: var(--md-font-mono);
background: var(--md-code-bg);
padding: 0.1em 0.3em;
border-radius: 4px;
font-size: 0.875em;
}
.annotation.unanchored { border-left-color: var(--md-text-muted); }
/* Reviewer footer */
footer.review-footer {
margin-top: 3rem;
padding-top: 1.5rem;
border-top: 1px solid var(--md-border);
color: var(--md-text-muted);
font-size: 0.9375rem;
display: flex;
align-items: center;
gap: 1rem;
flex-wrap: wrap;
}
footer.review-footer strong { color: var(--md-text); }
/* Approval bar (LGTM markers) */
.approval-bar {
background: color-mix(in srgb, var(--md-success) 12%, transparent);
border: 1px solid var(--md-success);
color: var(--md-success);
padding: 0.5rem 0.85rem;
border-radius: 8px;
margin: 1rem 0;
font-weight: 600;
font-family: var(--md-font-heading);
}
@media (prefers-reduced-motion: reduce) {
* { animation: none !important; transition: none !important; }
html { scroll-behavior: auto; }
}
"""
def _hunk_anchor(block_idx: int, file_idx: int, hunk_idx: int) -> str:
return f"hunk-b{block_idx}-f{file_idx}-h{hunk_idx}"
def _annotation_anchor(idx: int) -> str:
return f"ann-{idx}"
def _render_severity_badge(severity: str, palette: dict[str, str], danger: str) -> str:
spec = SEVERITY_DEFAULTS.get(severity.upper(), {"icon": "?", "aria": severity})
color = _resolve_severity_color(severity, palette, danger)
icon = html.escape(spec["icon"])
aria = html.escape(spec["aria"])
return (
f'<span class="sev-badge" style="color: {color}" '
f'role="status" aria-label="{aria}">'
f'<span class="sev-icon" aria-hidden="true">{icon}</span>'
f'{html.escape(severity.upper())}'
f'</span>'
)
def _render_hunk(file_idx: int, file_entry: dict[str, Any],
block_idx: int, hunk_idx: int) -> str:
h = file_entry["hunks"][hunk_idx]
path = file_entry.get("path_new") or file_entry.get("path_old") or "(unknown)"
anchor = _hunk_anchor(block_idx, file_idx, hunk_idx)
head = (
f'<div class="hunk-head">'
f'<strong>{html.escape(path)}</strong>'
f'<span>@{h["old_start"]}{html.escape(" → ")}@{h["new_start"]}</span>'
f'<span>{html.escape(h.get("header_context") or "")}</span>'
f'</div>'
)
body_lines = []
for ln in h["lines"]:
kind = ln["kind"]
mark = {"addition": "+", "deletion": "−", "context": " ", "meta": "\\"}.get(kind, " ")
old = "" if ln.get("old") is None else str(ln["old"])
new = "" if ln.get("new") is None else str(ln["new"])
body_lines.append(
f'<div class="line {kind}">'
f'<span class="lo">{old}</span>'
f'<span class="ln">{new}</span>'
f'<span class="mark" aria-hidden="true">{mark}</span>'
f'<span class="src">{html.escape(ln["text"])}</span>'
f'</div>'
)
return (
f'<div class="hunk" id="{anchor}">'
f'{head}<pre>{"".join(body_lines)}</pre></div>'
)
def _render_annotation(idx: int, ann: dict[str, Any],
palette: dict[str, str], danger: str,
anchor_for_block: dict[int, str]) -> str:
color = _resolve_severity_color(ann["severity"], palette, danger)
anchor_target = anchor_for_block.get(ann.get("attached_block"))
target_link = (
f'<a href="#{anchor_target}" '
f'style="color: var(--md-text-muted); font-size: 0.75rem; '
f'text-decoration: none">jump to diff →</a>'
if anchor_target else ""
)
body = html.escape(ann["body"])
# Re-inflate inline `code` for readability
body = body.replace("`", "<code>", 1)
while "<code>" in body and body.count("<code>") > body.count("</code>"):
body = body.replace("`", "</code>", 1)
return (
f'<div class="annotation{" unanchored" if anchor_target is None else ""}" '
f'id="{_annotation_anchor(idx)}" style="border-left-color: {color}">'
f'<div class="annotation-head">'
f'{_render_severity_badge(ann["severity"], palette, danger)}'
f'{target_link}'
f'</div>'
f'<div class="annotation-body">{body}</div>'
f'</div>'
)
def render(
diff_blocks: dict[str, Any],
annotations: dict[str, Any],
config: dict[str, Any],
reviewer: str,
pr_title: str = "Code Review",
) -> str:
palette = config.get("derived_palette") or {}
typo = config.get("typography") or {}
heading_font = typo.get("heading_font", "Inter")
body_font = typo.get("body_font", "Inter")
style = config.get("design_style", "technical")
company_name = config.get("company_name", "")
danger = _derive_danger_color(palette)
css = (BASE_CSS
.replace("__PALETTE__", _palette_to_css(palette))
.replace("__DANGER__", danger)
.replace("__HEADING_FONT__", _font_stack(heading_font))
.replace("__BODY_FONT__", _font_stack(body_font)))
# Build anchor map for jump-nav targets
anchor_for_block: dict[int, str] = {}
for block in diff_blocks.get("blocks", []):
block_idx = block["block_index"]
for file_idx, file_entry in enumerate(block["files"]):
if file_entry["hunks"]:
# Anchor the first hunk of the first file in each block
anchor_for_block.setdefault(
block_idx, _hunk_anchor(block_idx, file_idx, 0)
)
# Top jump-nav (ordered by source position, which is annotation order)
annlist = annotations.get("annotations", [])
nav_items_html: list[str] = []
for i, ann in enumerate(annlist):
target_anchor = _annotation_anchor(i)
target_diff = anchor_for_block.get(ann.get("attached_block")) or target_anchor
preview = html.escape(ann["body"][:80] + ("…" if len(ann["body"]) > 80 else ""))
target_label = (
f'#{ann.get("attached_block") + 1}' if ann.get("attached_block") is not None
else "(unanchored)"
)
nav_items_html.append(
f'<li>'
f'{_render_severity_badge(ann["severity"], palette, danger)}'
f'<a href="#{target_anchor}">{preview}</a>'
f'<span class="nav-target">{html.escape(target_label)}</span>'
f'</li>'
)
jump_nav_html = ""
if annlist:
counts = annotations.get("summary", {}).get("counts_by_severity", {})
count_summary = " · ".join(
f"{n} {sev}" for sev, n in sorted(counts.items())
)
jump_nav_html = (
'<nav class="jump-nav" aria-label="Review annotations">'
f'<h2>Findings ({count_summary})</h2>'
f'<ul>{"".join(nav_items_html)}</ul>'
'</nav>'
)
# Approval bar (LGTM markers)
approvals = annotations.get("approvals", [])
approval_html = ""
if approvals and not annlist:
approval_html = (
'<div class="approval-bar" role="status">'
'LGTM — no findings flagged'
'</div>'
)
# Render hunks + their attached annotations
block_to_annotations: dict[int, list[tuple[int, dict[str, Any]]]] = {}
unanchored: list[tuple[int, dict[str, Any]]] = []
for i, ann in enumerate(annlist):
if ann.get("attached_block") is not None:
block_to_annotations.setdefault(ann["attached_block"], []).append((i, ann))
else:
unanchored.append((i, ann))
sections_html: list[str] = []
for block in diff_blocks.get("blocks", []):
block_idx = block["block_index"]
anns_for_block = block_to_annotations.get(block_idx, [])
for file_idx, file_entry in enumerate(block["files"]):
for hunk_idx in range(len(file_entry["hunks"])):
hunk_html = _render_hunk(file_idx, file_entry, block_idx, hunk_idx)
# All annotations for this block render alongside the first hunk
ann_html = ""
if file_idx == 0 and hunk_idx == 0 and anns_for_block:
ann_html = "".join(
_render_annotation(i, ann, palette, danger, anchor_for_block)
for i, ann in anns_for_block
)
sections_html.append(
'<div class="hunk-row">'
f'{hunk_html}'
f'<div class="hunk-annotations">{ann_html}</div>'
'</div>'
)
if unanchored:
sections_html.append('<h2>General comments</h2>')
sections_html.append('<div class="hunk-annotations">')
for i, ann in unanchored:
sections_html.append(
_render_annotation(i, ann, palette, danger, anchor_for_block)
)
sections_html.append('</div>')
# Footer with mandatory reviewer name
footer_html = (
'<footer class="review-footer">'
f'<span>Reviewer: <strong>{html.escape(reviewer)}</strong></span>'
+ (f'<span>· {html.escape(company_name)}</span>' if company_name else "")
+ '<span style="margin-left:auto">Generated by markdown-html / md-review</span>'
'</footer>'
)
return f"""<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>{html.escape(pr_title)}</title>
<link rel="preconnect" href="https://fonts.googleapis.com" crossorigin>
<link rel="stylesheet" href="{_font_url(heading_font, body_font)}">
<style>{css}</style>
</head>
<body class="style-{style}">
<header><h1>{html.escape(pr_title)}</h1></header>
{jump_nav_html}
{approval_html}
{"".join(sections_html)}
{footer_html}
</body>
</html>"""
def main(argv: list[str]) -> int:
p = argparse.ArgumentParser(description=__doc__.split("\n")[0])
p.add_argument("--diff-blocks", help="Path to diff_parser JSON output")
p.add_argument("--annotations", help="Path to annotation_extractor JSON output")
p.add_argument("--reviewer", help="Reviewer name (required; refuses to render without)")
p.add_argument("--title", default="Code Review", help="PR / review title")
p.add_argument("--output", help="Path to write HTML (else stdout)")
p.add_argument("--no-config", action="store_true",
help="Bypass design-system config (use DEFAULTS)")
p.add_argument("--sample", action="store_true",
help="Render the built-in sample PR review")
p.add_argument("--severity-convention",
default="BLOCKER,MAJOR,MINOR,NIT",
help="Comma-separated severity tier list (most → least)")
args = p.parse_args(argv)
if args.sample:
sys.path.insert(0, str(Path(__file__).resolve().parent))
import annotation_extractor
import diff_parser
text = diff_parser.SAMPLE_MARKDOWN
diff_blocks = diff_parser.parse_markdown_for_diffs(text)
sev_conv = [s.strip().upper() for s in args.severity_convention.split(",")]
annotations = annotation_extractor.extract_annotations(text, diff_blocks, sev_conv)
reviewer = args.reviewer or "Sample Reviewer"
pr_title = args.title
else:
if not (args.diff_blocks and args.annotations):
print("error: --diff-blocks and --annotations are required "
"(use --sample for a built-in demo)", file=sys.stderr)
return 2
diff_blocks = json.loads(Path(args.diff_blocks).read_text(encoding="utf-8"))
annotations = json.loads(Path(args.annotations).read_text(encoding="utf-8"))
reviewer = args.reviewer
pr_title = args.title
# Hard rule 1: reviewer name is mandatory (named owner per research-ops discipline)
if not reviewer or not reviewer.strip():
print("refusing: --reviewer is required. A code review must name a human reviewer.",
file=sys.stderr)
return 3
# Hard rule 2: refuse if there are no hunks (wrong skill — route to md-document)
if diff_blocks.get("summary", {}).get("total_hunks", 0) == 0:
print("refusing: no diff hunks present in input. This is not a code review — "
"route to md-document instead.", file=sys.stderr)
return 4
if args.no_config or os.environ.get("MARKDOWN_HTML_NO_CONFIG") == "1":
config = _cfg.DEFAULTS if _cfg else {}
else:
config = _cfg.load_config() if _cfg else {}
output = render(diff_blocks, annotations, config, reviewer, pr_title)
if args.output and args.output != "-":
Path(args.output).write_text(output, encoding="utf-8")
print(f"wrote {args.output}: {len(output):,} bytes "
f"({diff_blocks['summary']['total_hunks']} hunks, "
f"{annotations['summary']['total_annotations']} annotations, "
f"reviewer={reviewer})")
else:
print(output)
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Chuyên gia tuân thủ EU MDR 2017/745: phân loại thiết bị y tế, hồ sơ kỹ thuật, bằng chứng lâm sàng, giám sát sau thị trường và EUDAMED.
---
name: "mdr-745-specialist"
description: EU MDR 2017/745 compliance specialist for medical device classification, technical documentation, clinical evidence, and post-market surveillance. Covers Annex VIII classification rules, Annex II/III technical files, Annex XIV clinical evaluation, and EUDAMED integration.
triggers:
- MDR compliance
- EU MDR
- medical device classification
- Annex VIII
- technical documentation
- clinical evaluation
- PMCF
- EUDAMED
- UDI
- notified body
---
# MDR 2017/745 Specialist
EU MDR compliance patterns for medical device classification, technical documentation, and clinical evidence.
---
## Table of Contents
- [Device Classification Workflow](#device-classification-workflow)
- [Technical Documentation](#technical-documentation)
- [Clinical Evidence](#clinical-evidence)
- [Post-Market Surveillance](#post-market-surveillance)
- [EUDAMED and UDI](#eudamed-and-udi)
- [Reference Documentation](#reference-documentation)
- [Tools](#tools)
---
## Device Classification Workflow
Classify device under MDR Annex VIII:
1. Identify device duration (transient, short-term, long-term)
2. Determine invasiveness level (non-invasive, body orifice, surgical)
3. Assess body system contact (CNS, cardiac, other)
4. Check if active device (energy dependent)
5. Apply classification rules 1-22
6. For software, apply MDCG 2019-11 algorithm
7. Document classification rationale
8. **Validation:** Classification confirmed with Notified Body
### Classification Matrix
| Factor | Class I | Class IIa | Class IIb | Class III |
|--------|---------|-----------|-----------|-----------|
| Duration | Any | Short-term | Long-term | Long-term |
| Invasiveness | Non-invasive | Body orifice | Surgical | Implantable |
| System | Any | Non-critical | Critical organs | CNS/cardiac |
| Risk | Lowest | Low-medium | Medium-high | Highest |
### Software Classification (MDCG 2019-11)
| Information Use | Condition Severity | Class |
|-----------------|-------------------|-------|
| Informs decision | Non-serious | IIa |
| Informs decision | Serious | IIb |
| Drives/treats | Critical | III |
### Classification Examples
**Example 1: Absorbable Surgical Suture**
- Rule 8 (implantable, long-term)
- Duration: > 30 days (absorbed)
- Contact: General tissue
- Classification: **Class IIb**
**Example 2: AI Diagnostic Software**
- Rule 11 + MDCG 2019-11
- Function: Diagnoses serious condition
- Classification: **Class IIb**
**Example 3: Cardiac Pacemaker**
- Rule 8 (implantable)
- Contact: Central circulatory system
- Classification: **Class III**
---
## Technical Documentation
Prepare technical file per Annex II and III:
1. Create device description (variants, accessories, intended purpose)
2. Develop labeling (Article 13 requirements, IFU)
3. Document design and manufacturing process
4. Complete GSPR compliance matrix
5. Prepare benefit-risk analysis
6. Compile verification and validation evidence
7. Integrate risk management file (ISO 14971)
8. **Validation:** Technical file reviewed for completeness
### Technical File Structure
```
ANNEX II TECHNICAL DOCUMENTATION
├── Device description and UDI-DI
├── Label and instructions for use
├── Design and manufacturing info
├── GSPR compliance matrix
├── Benefit-risk analysis
├── Verification and validation
└── Clinical evaluation report
```
### GSPR Compliance Checklist
| Requirement | Evidence | Status |
|-------------|----------|--------|
| Safe design (GSPR 1-3) | Risk management file | ☐ |
| Chemical properties (GSPR 10.1) | Biocompatibility report | ☐ |
| Infection risk (GSPR 10.2) | Sterilization validation | ☐ |
| Software requirements (GSPR 17) | IEC 62304 documentation | ☐ |
| Labeling (GSPR 23) | Label artwork, IFU | ☐ |
### Conformity Assessment Routes
| Class | Route | NB Involvement |
|-------|-------|----------------|
| I | Annex II self-declaration | None |
| Is/Im | Annex II + IX/XI | Sterile/measuring aspects |
| IIa | Annex II + IX or XI | Product or QMS |
| IIb | Annex IX + X or X + XI | Type exam + production |
| III | Annex IX + X | Full QMS + type exam |
---
## Clinical Evidence
Develop clinical evidence strategy per Annex XIV:
1. Define clinical claims and endpoints
2. Conduct systematic literature search
3. Appraise clinical data quality
4. Assess equivalence (technical, biological, clinical)
5. Identify evidence gaps
6. Determine if clinical investigation required
7. Prepare Clinical Evaluation Report (CER)
8. **Validation:** CER reviewed by qualified evaluator
### Evidence Requirements by Class
| Class | Minimum Evidence | Investigation |
|-------|------------------|---------------|
| I | Risk-benefit analysis | Not typically required |
| IIa | Literature + post-market | May be required |
| IIb | Systematic literature review | Often required |
| III | Comprehensive clinical data | Required (Article 61) |
### Clinical Evaluation Report Structure
```
CER CONTENTS
├── Executive summary
├── Device scope and intended purpose
├── Clinical background (state of the art)
├── Literature search methodology
├── Data appraisal and analysis
├── Safety and performance conclusions
├── Benefit-risk determination
└── PMCF plan summary
```
### Qualified Evaluator Requirements
- Medical degree or equivalent healthcare qualification
- 4+ years clinical experience in relevant field
- Training in clinical evaluation methodology
- Understanding of MDR requirements
---
## Post-Market Surveillance
Establish PMS system per Chapter VII:
1. Develop PMS plan (Article 84)
2. Define data collection methods
3. Establish complaint handling procedures
4. Create vigilance reporting process
5. Plan Periodic Safety Update Reports (PSUR)
6. Integrate with PMCF activities
7. Define trend analysis and signal detection
8. **Validation:** PMS system audited annually
### PMS System Components
| Component | Requirement | Frequency |
|-----------|-------------|-----------|
| PMS Plan | Article 84 | Maintain current |
| PSUR | Class IIa and higher | Per class schedule |
| PMCF Plan | Annex XIV Part B | Update with CER |
| PMCF Report | Annex XIV Part B | Annual (Class III) |
| Vigilance | Articles 87-92 | As events occur |
### PSUR Schedule
| Class | Frequency |
|-------|-----------|
| Class III | Annual |
| Class IIb implantable | Annual |
| Class IIb | Every 2 years |
| Class IIa | When necessary |
### Serious Incident Reporting
| Timeline | Requirement |
|----------|-------------|
| 2 days | Serious public health threat |
| 10 days | Death or serious deterioration |
| 15 days | Other serious incidents |
---
## EUDAMED and UDI
Implement UDI system per Article 27:
1. Obtain issuing entity code (GS1, HIBCC, ICCBBA)
2. Assign UDI-DI to each device variant
3. Assign UDI-PI (production identifier)
4. Apply UDI carrier to labels (AIDC + HRI)
5. Register actor in EUDAMED
6. Register devices in EUDAMED
7. Upload certificates when available
8. **Validation:** UDI verified on sample labels
### EUDAMED Modules
| Module | Content | Actor |
|--------|---------|-------|
| Actor | Company registration | Manufacturer, AR |
| UDI/Device | Device and variant data | Manufacturer |
| Certificates | NB certificates | Notified Body |
| Clinical Investigation | Study registration | Sponsor |
| Vigilance | Incident reports | Manufacturer |
| Market Surveillance | Authority actions | Competent Authority |
### UDI Label Requirements
Required elements per Article 13:
- [ ] UDI-DI (device identifier)
- [ ] UDI-PI (production identifier) for Class II+
- [ ] AIDC format (barcode/RFID)
- [ ] HRI format (human-readable)
- [ ] Manufacturer name and address
- [ ] Lot/serial number
- [ ] Expiration date (if applicable)
---
## Reference Documentation
### MDR Classification Guide
`references/mdr-classification-guide.md` contains:
- Complete Annex VIII classification rules (Rules 1-22)
- Software classification per MDCG 2019-11
- Worked classification examples
- Conformity assessment route selection
### Clinical Evidence Requirements
`references/clinical-evidence-requirements.md` contains:
- Clinical evidence framework and hierarchy
- Literature search methodology
- Clinical Evaluation Report structure
- PMCF plan and evaluation report guidance
### Technical Documentation Templates
`references/technical-documentation-templates.md` contains:
- Annex II and III content requirements
- Design History File structure
- GSPR compliance matrix template
- Declaration of Conformity template
- Notified Body submission checklist
---
## Tools
### MDR Gap Analyzer
```bash
# Quick gap analysis
python scripts/mdr_gap_analyzer.py --device "Device Name" --class IIa
# JSON output for integration
python scripts/mdr_gap_analyzer.py --device "Device Name" --class III --output json
# Interactive assessment
python scripts/mdr_gap_analyzer.py --interactive
```
Analyzes device against MDR requirements, identifies compliance gaps, generates prioritized recommendations.
**Output includes:**
- Requirements checklist by category
- Gap identification with priorities
- Critical gap highlighting
- Compliance roadmap recommendations
---
## Notified Body Interface
### Selection Criteria
| Factor | Considerations |
|--------|----------------|
| Designation scope | Covers your device type |
| Capacity | Timeline for initial audit |
| Geographic reach | Markets you need to access |
| Technical expertise | Experience with your technology |
| Fee structure | Transparency, predictability |
### Pre-Submission Checklist
- [ ] Technical documentation complete
- [ ] GSPR matrix fully addressed
- [ ] Risk management file current
- [ ] Clinical evaluation report complete
- [ ] QMS (ISO 13485) certified
- [ ] Labeling and IFU finalized
- [ ] **Validation:** Internal gap assessment complete
FILE:references/clinical-evidence-requirements.md
# Clinical Evidence Requirements
MDR Annex XIV clinical evaluation and post-market clinical follow-up guidance.
---
## Table of Contents
- [Clinical Evidence Framework](#clinical-evidence-framework)
- [Clinical Evaluation Process](#clinical-evaluation-process)
- [Literature-Based Evidence](#literature-based-evidence)
- [Clinical Investigation Requirements](#clinical-investigation-requirements)
- [Post-Market Clinical Follow-up](#post-market-clinical-follow-up)
---
## Clinical Evidence Framework
### Evidence Hierarchy
| Evidence Type | Strength | When to Use |
|---------------|----------|-------------|
| Randomized Controlled Trial | Highest | Novel Class III, high-risk claims |
| Prospective cohort study | High | New technology, performance claims |
| Retrospective analysis | Medium | Established technology, equivalence |
| Literature review | Medium | Well-characterized, equivalent devices |
| Expert opinion | Low | Supportive only, not primary |
### Evidence Requirements by Class
| Class | Minimum Evidence | Clinical Investigation |
|-------|------------------|------------------------|
| I | Risk-benefit analysis | Not typically required |
| IIa | Literature + post-market data | May be required for novel tech |
| IIb | Systematic literature review | Often required for claims |
| III | Comprehensive clinical data | Required unless equivalent |
### Clinical Evidence Pathway
Determine evidence strategy:
1. Assess device classification and risk level
2. Evaluate claim significance (diagnostic, therapeutic)
3. Determine if equivalence can be demonstrated
4. Identify available literature and clinical data
5. Assess gaps requiring additional investigation
6. Develop PMCF plan for ongoing evidence
7. **Validation:** Evidence strategy approved by Notified Body
---
## Clinical Evaluation Process
### Clinical Evaluation Workflow
Execute clinical evaluation per Annex XIV Part A:
1. Identify relevant safety and performance data
2. Define scope and search strategy
3. Conduct systematic literature search
4. Appraise and analyze clinical data
5. Assess benefit-risk profile
6. Document conclusions in Clinical Evaluation Report (CER)
7. Plan post-market clinical follow-up
8. **Validation:** CER reviewed by qualified clinical evaluator
### Clinical Evaluation Report Structure
```
CLINICAL EVALUATION REPORT (CER)
├── 1. Executive Summary
│ ├── Device description and intended purpose
│ ├── Conclusions on safety and performance
│ └── Benefit-risk conclusion
├── 2. Scope of Clinical Evaluation
│ ├── Device identification
│ ├── Clinical claims to be evaluated
│ └── Equivalence assessment (if applicable)
├── 3. Clinical Background
│ ├── Disease/condition overview
│ ├── Current treatment options
│ └── State of the art
├── 4. Clinical Data Sources
│ ├── Pre-clinical data (bench, animal)
│ ├── Clinical investigation data
│ ├── Literature search methodology
│ └── Post-market surveillance data
├── 5. Data Appraisal
│ ├── Study quality assessment
│ ├── Relevance to subject device
│ └── Data contribution to evaluation
├── 6. Data Analysis
│ ├── Safety analysis
│ ├── Performance analysis
│ └── Benefit-risk determination
├── 7. Conclusions
│ ├── Clinical evidence summary
│ ├── Residual risks
│ └── PMCF requirements
└── 8. PMCF Plan Summary
├── Data gaps identified
├── PMCF activities planned
└── Update schedule
```
### Qualified Clinical Evaluator
Requirements per Annex XIV:
- Medical degree or equivalent healthcare qualification
- 4+ years clinical experience in relevant field OR
- Research background in relevant domain
- Training in clinical evaluation methodology
- Understanding of MDR requirements
---
## Literature-Based Evidence
### Literature Search Strategy
Execute systematic literature review:
1. Define PICO question (Population, Intervention, Comparison, Outcome)
2. Develop search string with Boolean operators
3. Select databases (PubMed, Embase, Cochrane, etc.)
4. Set date range and language filters
5. Execute search and document results
6. Screen abstracts and full texts
7. **Validation:** Reproducible search, documented exclusion criteria
### Database Selection
| Database | Coverage | Best For |
|----------|----------|----------|
| PubMed/MEDLINE | Biomedical literature | Primary clinical data |
| Embase | Drugs, devices, biomedical | European studies |
| Cochrane Library | Systematic reviews | Meta-analyses |
| CINAHL | Nursing, allied health | User studies |
| IEEE Xplore | Engineering | Technical performance |
| Manufacturer data | Proprietary | Direct device data |
### Data Appraisal Criteria
Evaluate each source:
| Criterion | Assessment | Score |
|-----------|------------|-------|
| Study design | RCT > cohort > case series | 1-5 |
| Sample size | Statistical power adequate | 1-5 |
| Follow-up duration | Sufficient for outcomes | 1-5 |
| Population relevance | Matches IFU population | 1-5 |
| Device equivalence | Technical, biological, clinical | 1-5 |
| Bias risk | Low/medium/high | 1-5 |
### Equivalence Assessment
Demonstrate equivalence per MDCG 2020-5:
**Technical equivalence:**
- Similar design
- Same materials
- Same specifications
- Same manufacturing process
**Biological equivalence:**
- Same tissue contact
- Same biocompatibility
- Same sterilization method
**Clinical equivalence:**
- Same intended purpose
- Same clinical condition
- Same patient population
- Same user (professional/lay)
---
## Clinical Investigation Requirements
### When Investigation is Required
Clinical investigation mandatory:
- [ ] Class III implantable devices (Article 61(4))
- [ ] Novel technology without equivalent
- [ ] Significant modification to existing device
- [ ] New clinical claims not supported by literature
- [ ] Addressing gaps identified in clinical evaluation
### Clinical Investigation Workflow
Conduct clinical investigation:
1. Develop Clinical Investigation Plan (CIP)
2. Submit to Ethics Committee for approval
3. Notify Competent Authority (via EUDAMED)
4. Conduct investigation per GCP principles
5. Collect and analyze clinical data
6. Prepare Clinical Investigation Report
7. Submit serious adverse event reports (within 7-15 days)
8. **Validation:** All subjects completed, data lock achieved
### Clinical Investigation Plan Elements
| Section | Content |
|---------|---------|
| Objectives | Primary and secondary endpoints |
| Design | Randomized, controlled, blinded |
| Population | Inclusion/exclusion criteria |
| Sample size | Statistical justification |
| Procedures | Visit schedule, assessments |
| Endpoints | Safety and performance measures |
| Analysis | Statistical methods |
| Safety | Adverse event definitions, reporting |
### Ethics Committee Submission
Required documentation:
- Clinical Investigation Plan
- Investigator's Brochure
- Informed consent documents
- Case Report Forms
- Investigator CVs
- Insurance certificate
- Device documentation
---
## Post-Market Clinical Follow-up
### PMCF Plan Requirements
Develop PMCF Plan per Annex XIV Part B:
1. Identify residual risks from clinical evaluation
2. Define clinical questions to address
3. Select PMCF methods (survey, registry, study)
4. Specify endpoints and success criteria
5. Define timeline and milestones
6. Plan data collection and analysis
7. Schedule PMCF Evaluation Report updates
8. **Validation:** PMCF Plan approved by Notified Body
### PMCF Methods
| Method | Description | Best For |
|--------|-------------|----------|
| Literature review | Ongoing systematic search | Mature devices |
| Survey | User/patient questionnaire | Real-world experience |
| Registry | Multi-site data collection | Long-term outcomes |
| PMCF study | Prospective clinical study | Specific questions |
| Complaint analysis | Structured complaint review | Safety signals |
| Vigilance data | MAUDE, EUDAMED analysis | Comparative safety |
### PMCF Evaluation Report
Update frequency:
| Device Class | Update Frequency |
|--------------|------------------|
| Class III | Annual |
| Class IIb implantable | Annual |
| Class IIb | Every 2 years |
| Class IIa | Every 2-5 years |
| Class I | When clinically relevant |
### PMCF Report Structure
```
PMCF EVALUATION REPORT
├── 1. Executive Summary
│ └── Key findings and conclusions
├── 2. Scope
│ ├── Device covered
│ └── Reporting period
├── 3. PMCF Activities
│ ├── Methods employed
│ └── Data sources
├── 4. Results
│ ├── Safety data
│ ├── Performance data
│ └── Clinical questions addressed
├── 5. Conclusions
│ ├── Benefit-risk confirmation
│ ├── Residual risks updated
│ └── Need for corrective action
└── 6. Next Steps
├── CER update requirements
└── PMCF Plan modifications
```
### Integration with CER
PMCF data feeds clinical evaluation:
1. PMCF data collected per plan
2. Data analyzed in PMCF Evaluation Report
3. CER updated with PMCF conclusions
4. Risk management file updated
5. IFU updated if needed
6. **Validation:** CER update cycle completed
FILE:references/mdr-classification-guide.md
# MDR Device Classification Guide
EU MDR 2017/745 Annex VIII classification rules and decision framework.
---
## Table of Contents
- [Classification Overview](#classification-overview)
- [Classification Rules](#classification-rules)
- [Software Classification (MDCG 2019-11)](#software-classification)
- [Classification Examples](#classification-examples)
- [Conformity Assessment Routes](#conformity-assessment-routes)
---
## Classification Overview
### Risk Class Hierarchy
| Class | Risk Level | Examples | NB Required |
|-------|------------|----------|-------------|
| I | Lowest | Bandages, wheelchairs, stethoscopes | No (self-certification) |
| IIa | Low-Medium | Hearing aids, dental filling materials | Yes |
| IIb | Medium-High | Ventilators, blood bags, implantable sutures | Yes |
| III | Highest | Pacemakers, heart valves, hip implants | Yes |
### Classification Factors
Determine class based on:
1. **Duration of contact:**
- Transient: < 60 minutes
- Short-term: 60 min to 30 days
- Long-term: > 30 days
2. **Degree of invasiveness:**
- Non-invasive
- Invasive via body orifice
- Surgically invasive
- Implantable
3. **Body system interaction:**
- Central circulatory system
- Central nervous system
- Other organ systems
4. **Active vs. passive:**
- Active devices (energy dependent)
- Passive devices
---
## Classification Rules
### Non-Invasive Devices (Rules 1-4)
**Rule 1 - General non-invasive:**
- Class I (unless covered by other rules)
- Example: Wheelchairs, hospital beds, collection devices
**Rule 2 - Channeling or storing:**
- Class IIa: Blood bags, transfusion sets (>60 min contact)
- Class IIb: Blood storage, organ storage
- Class I: Simple channeling (gravity, IV bag without additives)
**Rule 3 - Modifying biological composition:**
- Class IIa: Filters, gas separators, dialysis filters
- Class IIb: Blood filtration, exchange transfusion
**Rule 4 - Contact with injured skin:**
- Class I: Wound dressings for superficial wounds
- Class IIa: Wounds in dermis requiring secondary intent healing
- Class IIb: Severe wounds, chronic wounds, burns
### Invasive Devices (Rules 5-8)
**Rule 5 - Body orifice invasive (transient):**
- Class I: Transient use, non-surgically invasive
- Class IIa: Short-term use
- Class IIb: Long-term use in oral cavity
**Rule 6 - Surgically invasive (transient):**
- Class IIa: Transient use
- Exception Class I: Reusable surgical instruments
**Rule 7 - Surgically invasive (short-term):**
- Class IIa: Short-term (< 30 days)
- Class IIb: Central circulatory or CNS contact
- Class III: Chemical change or drug delivery
**Rule 8 - Implantable and long-term surgically invasive:**
- Class IIb: General implants
- Class III: Heart, CNS, spine contact; drug delivery; biological origin
### Active Devices (Rules 9-13)
**Rule 9 - Active therapeutic devices:**
- Class IIa: Exchange or admin of energy (non-hazardous)
- Class IIb: Potentially hazardous energy levels
**Rule 10 - Active diagnostic devices:**
- Class IIa: Supply energy for imaging, monitoring
- Class IIb: Monitor vital physiological parameters
**Rule 11 - Software:**
- Class IIa: Information for diagnostic/therapeutic decisions (non-serious)
- Class IIb: Decisions that could cause death/irreversible deterioration
- Class III: Decisions with immediate risk to life
- See MDCG 2019-11 for detailed algorithm
**Rule 12 - Active devices administering substances:**
- Class IIa: Non-hazardous manner
- Class IIb: Potentially hazardous manner
**Rule 13 - Other active devices:**
- Class I: All other active devices
### Special Rules (Rules 14-22)
**Rule 14 - Contraception/STI prevention:**
- Class IIb: Contraceptive devices
- Class III: Implantable contraceptives
**Rule 15 - Disinfection/sterilization:**
- Class IIa: Disinfection of devices
- Class IIb: Disinfection of invasive devices
**Rule 16 - X-ray diagnostic recording:**
- Class IIa: Recording media for x-ray
**Rule 17 - Devices with nanomaterials:**
- Class III: High internal exposure potential
- Class IIb: Medium exposure
- Class IIa: Low exposure
**Rule 18 - Blood/plasma derivatives:**
- Class III: Utilizing blood derivatives
**Rule 19 - Drug delivery systems:**
- Class III: Integral drug administration
**Rule 20 - Breath analyzers for anesthesia:**
- Class IIb: Breath analyzers
**Rule 21 - Medicinal substance devices:**
- Class III: Incorporating medicinal substances
**Rule 22 - Closed-loop therapeutic systems:**
- Class III: Closed-loop systems
---
## Software Classification
### MDCG 2019-11 Decision Algorithm
Execute software classification:
1. Determine if software qualifies as medical device
2. Identify significance of information to healthcare decision
3. Assess healthcare situation or patient condition
4. Apply rule 11 based on severity
5. **Validation:** Classification rationale documented with MDCG reference
### Software Classification Matrix
| Information Significance | Situation/Condition | Class |
|--------------------------|---------------------|-------|
| Informs clinical management | Non-serious | IIa |
| Informs clinical management | Serious | IIb |
| Informs clinical management | Critical | III |
| Drives clinical management | Non-serious | IIa |
| Drives clinical management | Serious | IIb |
| Drives clinical management | Critical | III |
| Treats or diagnoses | Non-serious | IIa |
| Treats or diagnoses | Serious | IIb |
| Treats or diagnoses | Critical | III |
### Software Examples
| Software Type | Class | Rationale |
|---------------|-------|-----------|
| Patient record viewing | Not MD | Administrative, not clinical |
| Medication reminder app | Class I | General wellness |
| Blood glucose monitor app | Class IIa | Informs non-serious decisions |
| Sepsis detection algorithm | Class IIb | Informs serious condition |
| AI tumor detection | Class III | Diagnoses critical condition |
| Closed-loop insulin delivery | Class III | Treats critical condition |
---
## Classification Examples
### Example 1: Surgical Suture (Absorbable)
```
Device: Absorbable suture for internal wound closure
Analysis:
- Invasiveness: Surgically invasive
- Duration: Long-term (absorbed over > 30 days)
- System: General tissue (not CNS, not cardiac)
- Rule Applied: Rule 8 (implantable, long-term)
Classification: Class IIb
Rationale: Implantable device > 30 days, general tissue
Conformity Route: Annex IX (Type examination) + Annex XI
```
### Example 2: Blood Pressure Monitor
```
Device: Home blood pressure monitoring device
Analysis:
- Active: Yes (electronic measurement)
- Function: Monitoring vital physiological parameter
- Risk: Non-immediate (home use, not ICU)
- Rule Applied: Rule 10 (active diagnostic)
Classification: Class IIa
Rationale: Monitors vital parameter, non-critical setting
Conformity Route: Annex IX or XI (QMS + product verification)
```
### Example 3: Hip Implant
```
Device: Total hip replacement prosthesis
Analysis:
- Invasiveness: Surgically invasive, implantable
- Duration: Long-term (permanent)
- System: Musculoskeletal
- Rule Applied: Rule 8 (implantable, long-term)
Classification: Class III
Rationale: Implantable > 30 days in direct contact with bone
Conformity Route: Annex IX + Annex X (full QMS + type examination)
```
### Example 4: Diagnostic Software (AI)
```
Device: AI-based chest X-ray analysis for pneumonia detection
Analysis:
- Software: Qualifies as medical device (clinical decision)
- Information: Diagnoses condition
- Condition: Serious (pneumonia can be life-threatening)
- Rule Applied: Rule 11 + MDCG 2019-11
Classification: Class IIb
Rationale: Software diagnosing serious condition
Conformity Route: Annex IX or Annex XI
```
---
## Conformity Assessment Routes
### By Device Class
| Class | Conformity Route | NB Involvement |
|-------|------------------|----------------|
| I | Annex II (self-declaration) | None |
| I (sterile/measuring) | Annex II + IX/XI | Sterile/measuring aspects |
| IIa | Annex II + IX or XI | Product verification or QMS |
| IIb | Annex IX + X or Annex X + XI | Type exam + QMS or production |
| III | Annex IX + X | Full QMS + type examination |
### Annex Reference
| Annex | Content | Purpose |
|-------|---------|---------|
| II | Technical documentation | Required for all classes |
| III | Technical documentation (additions) | Class III additions |
| IX | Conformity assessment (QMS) | Quality management route |
| X | Type examination | Product design examination |
| XI | Product verification | Production quality checks |
### Decision Workflow
Select conformity route:
1. Determine device classification (Rules 1-22)
2. Identify applicable annexes for class
3. Evaluate QMS maturity (Annex IX capability)
4. Consider production volume (batch vs. mass)
5. Assess Notified Body capacity and timeline
6. Select optimal conformity assessment route
7. **Validation:** Route confirmed with Notified Body consultation
FILE:references/technical-documentation-templates.md
# Technical Documentation Templates
MDR Annex II and III technical file structure and content requirements.
---
## Table of Contents
- [Technical Documentation Overview](#technical-documentation-overview)
- [Annex II Requirements](#annex-ii-requirements)
- [Annex III Additions](#annex-iii-additions)
- [Document Templates](#document-templates)
- [Notified Body Expectations](#notified-body-expectations)
---
## Technical Documentation Overview
### Documentation Hierarchy
```
TECHNICAL DOCUMENTATION
├── Device Description and Specification
├── Information Supplied by Manufacturer
├── Design and Manufacturing Information
├── General Safety and Performance Requirements
├── Benefit-Risk Analysis
├── Product Verification and Validation
├── Clinical Evaluation Report
└── Post-Market Surveillance Documentation
```
### Documentation by Phase
| Phase | Required Documents |
|-------|-------------------|
| Design Input | User needs, design requirements, regulatory requirements |
| Design Development | Design specifications, drawings, BOM, software docs |
| Verification | Test protocols, test reports, design review records |
| Validation | Clinical data, usability data, biocompatibility |
| Transfer | Manufacturing specs, process validations |
| Post-Market | PMS plan, PMCF plan, vigilance procedures |
---
## Annex II Requirements
### Section 1: Device Description and Specification
**1.1 Device Identification**
```
DEVICE IDENTIFICATION
├── Trade name(s)
├── General description of the device
├── Basic UDI-DI
├── Device identifier codes (internal + regulatory)
├── Intended purpose statement
├── Indications for use
├── Contraindications
├── Target population (patient, user)
├── Medical conditions intended to diagnose/treat
└── Principles of operation
```
**1.2 Device Variants and Accessories**
| Element | Description |
|---------|-------------|
| Variant listing | All variants with identifiers |
| Configuration differences | Technical differences by variant |
| Accessories | Separate devices used together |
| Spare parts | Replaceable components |
**1.3 Reference to Previous Generations**
- Previous generation device identification
- Key modifications summary
- Clinical experience from prior device
- Justification for changes
### Section 2: Information Supplied by Manufacturer
**2.1 Label Requirements**
Mandatory label elements per Article 13:
- [ ] Device name or trade name
- [ ] Manufacturer name and address
- [ ] Authorized representative (if applicable)
- [ ] Lot/batch number or serial number
- [ ] UDI carrier (AIDC + HRI)
- [ ] Expiration date (if applicable)
- [ ] Storage/handling conditions
- [ ] Warnings and precautions
- [ ] CE mark with NB number (if applicable)
- [ ] Symbol meanings per EN ISO 15223-1
**2.2 Instructions for Use**
IFU structure:
```
INSTRUCTIONS FOR USE
├── 1. Device Description
│ ├── Intended purpose
│ ├── Indications and contraindications
│ └── Principle of operation
├── 2. Warnings and Precautions
│ ├── Contraindicated uses
│ ├── Potential complications
│ └── Drug/device interactions
├── 3. User Instructions
│ ├── Unpacking and inspection
│ ├── Setup/installation
│ ├── Operating procedures
│ └── Cleaning/maintenance
├── 4. Technical Specifications
│ ├── Physical characteristics
│ ├── Performance characteristics
│ └── Environmental limits
├── 5. Troubleshooting
│ ├── Error codes/messages
│ └── Corrective actions
└── 6. Symbols Glossary
```
### Section 3: Design and Manufacturing Information
**3.1 Design Process Documentation**
| Document | Purpose |
|----------|---------|
| Design input | User needs, regulatory requirements |
| Design output | Specifications, drawings, software |
| Design review | Review records at key milestones |
| Design verification | Test protocols and results |
| Design validation | Clinical/usability evidence |
| Design transfer | Manufacturing readiness |
| Design changes | Change control records |
**3.2 Manufacturing Process Description**
```
MANUFACTURING DOCUMENTATION
├── Process flow diagram
├── Manufacturing specifications
├── Facility and equipment qualification
├── Process validation protocols/reports
├── Environmental monitoring
├── Personnel training records
├── In-process controls
├── Final inspection/testing
├── Sterilization validation (if applicable)
└── Packaging validation
```
**3.3 Supplier and Subcontractor Information**
- Approved supplier list
- Supplier qualification records
- Critical component specifications
- Incoming inspection procedures
- Supplier audit records
### Section 4: General Safety and Performance Requirements
**GSPR Compliance Checklist**
| GSPR | Requirement | Evidence |
|------|-------------|----------|
| 1 | Safe design for intended use | Risk management file |
| 2 | Risk acceptable when weighed against benefits | Benefit-risk analysis |
| 3 | State of the art design | Literature review, standards |
| 4 | No compromise of clinical condition | Clinical evaluation |
| 5 | Transport and storage conditions | Shelf life testing |
| 6 | Acceptable undesirable effects | Risk-benefit analysis |
| 7 | CE marking conformity | Declaration of conformity |
| ... | Continue for all applicable GSPRs | |
**GSPR Matrix Template**
| GSPR # | Requirement Summary | Applicable? | Evidence Document | Status |
|--------|---------------------|-------------|-------------------|--------|
| 10.1 | Chemical properties | Yes/No/NA | Biocompatibility report | Complete |
| 10.2 | Infection risk | Yes/No/NA | Sterilization validation | Complete |
| 10.3 | Substances with carcinogenic risk | Yes/No/NA | Material specification | Complete |
### Section 5: Benefit-Risk Analysis
**Benefit-Risk Documentation**
```
BENEFIT-RISK ANALYSIS
├── 1. Intended Benefits
│ ├── Direct therapeutic benefits
│ ├── Diagnostic accuracy improvements
│ └── Patient outcome benefits
├── 2. Known Risks
│ ├── Identified hazards (from risk analysis)
│ ├── Risk control measures implemented
│ └── Residual risks
├── 3. Benefit-Risk Determination
│ ├── Qualitative analysis
│ ├── Quantitative analysis (if available)
│ └── Comparison to alternatives
└── 4. Conclusion
├── Acceptability statement
└── Justification for residual risks
```
### Section 6: Product Verification and Validation
**6.1 Verification Testing**
| Test Category | Standards | Documentation |
|---------------|-----------|---------------|
| Electrical safety | IEC 60601-1 | Test protocol + report |
| EMC | IEC 60601-1-2 | EMC test report |
| Biocompatibility | ISO 10993 series | Biocompatibility evaluation |
| Software | IEC 62304 | Software verification report |
| Sterilization | ISO 11135/11137 | Sterility assurance |
| Packaging | ISO 11607 | Packaging validation |
| Shelf life | Accelerated aging | Stability study report |
| Usability | IEC 62366-1 | Usability engineering file |
**6.2 Validation Evidence**
- Clinical investigation data
- Literature-based clinical evidence
- Simulated use testing
- User feedback/complaint analysis
- Post-market surveillance data
---
## Annex III Additions
### Class III Specific Requirements
Additional documentation for Class III devices:
**Implant-Specific Requirements**
- Implant card information
- Patient information leaflet
- Device tracking procedures
- Explant analysis capability
**Drug-Device Combination**
- Drug substance specification
- Drug compatibility testing
- Combined product assessment
- Pharmacovigilance interface
---
## Document Templates
### Design History File Index
```
DESIGN HISTORY FILE (DHF)
Document ID: DHF-[Product]-[Rev]
1. DESIGN INPUT
1.1 User Requirements Specification (URS)
1.2 Regulatory Requirements Matrix
1.3 Design Input Review Record
2. DESIGN OUTPUT
2.1 Product Specification
2.2 Engineering Drawings
2.3 Bill of Materials
2.4 Software Documentation
3. DESIGN VERIFICATION
3.1 Verification Test Plan
3.2 Verification Test Reports
3.3 Traceability Matrix
4. DESIGN VALIDATION
4.1 Clinical Evaluation Report
4.2 Usability Engineering File
4.3 Biocompatibility Evaluation
5. DESIGN TRANSFER
5.1 Manufacturing Procedures
5.2 Process Validation Reports
5.3 Supplier Qualification
6. DESIGN REVIEWS
6.1 Design Review Records
6.2 Risk Management Review
6.3 Final Design Release
```
### Declaration of Conformity Template
```
EU DECLARATION OF CONFORMITY
We, [Manufacturer Name]
Address: [Full address]
declare under our sole responsibility that the device:
Device name: [Trade name]
Device description: [Description]
Basic UDI-DI: [UDI-DI]
Classification: [Class I/IIa/IIb/III]
is in conformity with the provisions of:
- Regulation (EU) 2017/745
Applicable standards:
- [List harmonized standards]
Notified Body: [NB name and number] (if applicable)
Certificate number: [Certificate number]
Place and date: [Location, Date]
Signature: [Authorized signatory]
Name and function: [Name, Title]
```
---
## Notified Body Expectations
### Common NB Findings
| Finding Area | Common Issue | Prevention |
|--------------|--------------|------------|
| GSPR matrix | Incomplete, no evidence links | Complete matrix with references |
| Risk management | Not integrated with design | Update throughout development |
| Clinical evaluation | Insufficient literature search | Systematic search with PICO |
| IFU | Missing warnings | Risk-based IFU content |
| Traceability | Design to requirements gaps | Maintain traceability matrix |
### Pre-Submission Checklist
Before Notified Body submission:
- [ ] Technical documentation complete
- [ ] GSPR checklist fully addressed
- [ ] Risk management file current
- [ ] Clinical evaluation report complete
- [ ] QMS documentation ready
- [ ] Design verification complete
- [ ] Design validation complete
- [ ] Labeling and IFU finalized
- [ ] Declaration of conformity prepared
- [ ] **Validation:** Internal review completed
FILE:scripts/mdr_gap_analyzer.py
#!/usr/bin/env python3
"""
MDR Gap Analyzer - EU MDR 2017/745 Compliance Gap Assessment Tool
Analyzes device classification, identifies documentation gaps, and generates
compliance roadmap for EU MDR transition.
Usage:
python mdr_gap_analyzer.py --device "Device Name" --class IIa
python mdr_gap_analyzer.py --device "Device Name" --class III --output json
python mdr_gap_analyzer.py --interactive
"""
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from datetime import datetime
from typing import List, Dict, Optional
from enum import Enum
class DeviceClass(Enum):
I = "I"
I_STERILE = "Is"
I_MEASURING = "Im"
IIA = "IIa"
IIB = "IIb"
III = "III"
class GapStatus(Enum):
NOT_STARTED = "Not Started"
IN_PROGRESS = "In Progress"
COMPLETE = "Complete"
NOT_APPLICABLE = "N/A"
@dataclass
class GapItem:
requirement: str
category: str
description: str
status: GapStatus = GapStatus.NOT_STARTED
priority: str = "Medium"
evidence_needed: List[str] = field(default_factory=list)
notes: str = ""
@dataclass
class GapAnalysisResult:
device_name: str
device_class: str
analysis_date: str
total_requirements: int
gaps_identified: int
completion_percentage: float
gaps: List[Dict]
recommendations: List[str]
critical_gaps: List[str]
class MDRGapAnalyzer:
"""Analyzer for EU MDR 2017/745 compliance gaps."""
# MDR Requirements by category
REQUIREMENTS = {
"technical_documentation": [
GapItem(
requirement="Annex II - Device Description",
category="Technical Documentation",
description="Complete device description including variants, accessories, intended purpose",
priority="High",
evidence_needed=["Device specification", "Intended purpose statement", "Variant listing"]
),
GapItem(
requirement="Annex II - Information Supplied",
category="Technical Documentation",
description="Label and IFU meeting Article 13 requirements",
priority="High",
evidence_needed=["Label artwork", "Instructions for use", "Symbol glossary"]
),
GapItem(
requirement="Annex II - Design and Manufacturing",
category="Technical Documentation",
description="Design history file and manufacturing documentation",
priority="High",
evidence_needed=["Design history file", "Process flow diagram", "Validation reports"]
),
GapItem(
requirement="Annex II - GSPR Compliance",
category="Technical Documentation",
description="General Safety and Performance Requirements checklist",
priority="Critical",
evidence_needed=["GSPR matrix", "Standard compliance evidence", "Risk management file"]
),
],
"clinical_evaluation": [
GapItem(
requirement="Annex XIV Part A - Clinical Evaluation",
category="Clinical Evaluation",
description="Clinical evaluation report with systematic literature review",
priority="Critical",
evidence_needed=["Clinical evaluation report", "Literature search protocol", "Data appraisal"]
),
GapItem(
requirement="Annex XIV Part B - PMCF",
category="Clinical Evaluation",
description="Post-market clinical follow-up plan and evaluation report",
priority="High",
evidence_needed=["PMCF plan", "PMCF evaluation report", "Residual risk assessment"]
),
GapItem(
requirement="Qualified Person for CER",
category="Clinical Evaluation",
description="Clinical evaluation by qualified evaluator per Annex XIV",
priority="High",
evidence_needed=["Evaluator CV", "Qualification evidence", "Signed CER"]
),
],
"risk_management": [
GapItem(
requirement="ISO 14971 Risk Management",
category="Risk Management",
description="Complete risk management file per ISO 14971:2019",
priority="Critical",
evidence_needed=["Risk management plan", "Risk analysis", "Risk evaluation", "Risk control"]
),
GapItem(
requirement="Benefit-Risk Analysis",
category="Risk Management",
description="Documented benefit-risk determination",
priority="High",
evidence_needed=["Benefit-risk analysis document", "Residual risk acceptability"]
),
],
"quality_management": [
GapItem(
requirement="ISO 13485 QMS",
category="Quality Management",
description="Quality management system conforming to ISO 13485:2016",
priority="Critical",
evidence_needed=["QMS manual", "Process documentation", "Internal audit records"]
),
GapItem(
requirement="Post-Market Surveillance",
category="Quality Management",
description="PMS system per Article 83-86",
priority="High",
evidence_needed=["PMS plan", "PSUR (if required)", "Vigilance procedures"]
),
],
"udi_eudamed": [
GapItem(
requirement="UDI System",
category="UDI/EUDAMED",
description="Unique Device Identification per Article 27",
priority="High",
evidence_needed=["UDI-DI assignment", "Label with UDI carrier", "GUDID/EUDAMED registration"]
),
GapItem(
requirement="EUDAMED Registration",
category="UDI/EUDAMED",
description="Actor, device, and certificate registration in EUDAMED",
priority="Medium",
evidence_needed=["Actor registration", "Device registration", "Certificate upload"]
),
],
"notified_body": [
GapItem(
requirement="Notified Body Selection",
category="Notified Body",
description="Selection and engagement of MDR-designated Notified Body",
priority="Critical",
evidence_needed=["NB selection criteria", "NB engagement letter", "Audit schedule"]
),
GapItem(
requirement="Conformity Assessment",
category="Notified Body",
description="Completion of appropriate conformity assessment procedure",
priority="Critical",
evidence_needed=["Application dossier", "Technical documentation submission", "Certificate"]
),
],
}
# Class-specific requirements
CLASS_REQUIREMENTS = {
DeviceClass.III: [
GapItem(
requirement="Annex III - Class III Additions",
category="Technical Documentation",
description="Additional documentation for Class III devices",
priority="Critical",
evidence_needed=["Implant card", "Patient information", "Device tracking"]
),
GapItem(
requirement="Clinical Investigation",
category="Clinical Evaluation",
description="Clinical investigation per Article 61 (unless equivalent device)",
priority="Critical",
evidence_needed=["Clinical investigation plan", "Ethics approval", "Clinical study report"]
),
],
DeviceClass.IIB: [
GapItem(
requirement="Implantable Device Documentation",
category="Technical Documentation",
description="Additional requirements for implantable Class IIb devices",
priority="High",
evidence_needed=["Implant card (if implantable)", "Long-term safety data"]
),
],
}
def __init__(self, device_name: str, device_class: DeviceClass):
self.device_name = device_name
self.device_class = device_class
self.gaps: List[GapItem] = []
self._build_requirements_list()
def _build_requirements_list(self):
"""Build complete requirements list based on device class."""
# Add all base requirements
for category_gaps in self.REQUIREMENTS.values():
for gap in category_gaps:
self.gaps.append(GapItem(
requirement=gap.requirement,
category=gap.category,
description=gap.description,
priority=gap.priority,
evidence_needed=gap.evidence_needed.copy()
))
# Add class-specific requirements
if self.device_class in self.CLASS_REQUIREMENTS:
for gap in self.CLASS_REQUIREMENTS[self.device_class]:
self.gaps.append(GapItem(
requirement=gap.requirement,
category=gap.category,
description=gap.description,
priority=gap.priority,
evidence_needed=gap.evidence_needed.copy()
))
# Class I self-certification: NB not required
if self.device_class == DeviceClass.I:
for gap in self.gaps:
if gap.category == "Notified Body":
gap.status = GapStatus.NOT_APPLICABLE
def update_gap_status(self, requirement: str, status: GapStatus, notes: str = ""):
"""Update status of a specific gap."""
for gap in self.gaps:
if gap.requirement == requirement:
gap.status = status
gap.notes = notes
break
def analyze(self) -> GapAnalysisResult:
"""Perform gap analysis and generate results."""
applicable_gaps = [g for g in self.gaps if g.status != GapStatus.NOT_APPLICABLE]
complete_gaps = [g for g in applicable_gaps if g.status == GapStatus.COMPLETE]
completion = (len(complete_gaps) / len(applicable_gaps) * 100) if applicable_gaps else 0
# Identify critical gaps
critical_gaps = [
g.requirement for g in applicable_gaps
if g.priority == "Critical" and g.status != GapStatus.COMPLETE
]
# Generate recommendations
recommendations = self._generate_recommendations()
return GapAnalysisResult(
device_name=self.device_name,
device_class=self.device_class.value,
analysis_date=datetime.now().isoformat(),
total_requirements=len(applicable_gaps),
gaps_identified=len(applicable_gaps) - len(complete_gaps),
completion_percentage=round(completion, 1),
gaps=[{
"requirement": g.requirement,
"category": g.category,
"status": g.status.value,
"priority": g.priority,
"evidence_needed": g.evidence_needed
} for g in applicable_gaps],
recommendations=recommendations,
critical_gaps=critical_gaps
)
def _generate_recommendations(self) -> List[str]:
"""Generate prioritized recommendations."""
recommendations = []
# Check for critical gaps
critical_incomplete = [
g for g in self.gaps
if g.priority == "Critical" and g.status not in [GapStatus.COMPLETE, GapStatus.NOT_APPLICABLE]
]
if critical_incomplete:
recommendations.append(
f"CRITICAL: {len(critical_incomplete)} critical requirements not complete. "
"Address immediately to proceed with conformity assessment."
)
# Check clinical evaluation
cer_gap = next((g for g in self.gaps if "Clinical Evaluation" in g.requirement), None)
if cer_gap and cer_gap.status != GapStatus.COMPLETE:
recommendations.append(
"Clinical Evaluation Report (CER) is incomplete. "
"This is required before Notified Body submission."
)
# Check for Class III specific
if self.device_class == DeviceClass.III:
ci_gap = next((g for g in self.gaps if "Clinical Investigation" in g.requirement), None)
if ci_gap and ci_gap.status != GapStatus.COMPLETE:
recommendations.append(
"Class III device requires clinical investigation per Article 61 "
"unless equivalence can be demonstrated."
)
# Check EUDAMED
udi_gap = next((g for g in self.gaps if "UDI System" in g.requirement), None)
if udi_gap and udi_gap.status != GapStatus.COMPLETE:
recommendations.append(
"Implement UDI system and plan for EUDAMED registration. "
"Required for placing device on EU market."
)
return recommendations
def format_text_output(result: GapAnalysisResult) -> str:
"""Format analysis result as text."""
lines = [
"=" * 60,
"MDR 2017/745 GAP ANALYSIS REPORT",
"=" * 60,
f"Device: {result.device_name}",
f"Class: {result.device_class}",
f"Date: {result.analysis_date[:10]}",
"",
"-" * 60,
"SUMMARY",
"-" * 60,
f"Total Requirements: {result.total_requirements}",
f"Gaps Identified: {result.gaps_identified}",
f"Completion: {result.completion_percentage}%",
"",
]
if result.critical_gaps:
lines.extend([
"-" * 60,
"CRITICAL GAPS (Address Immediately)",
"-" * 60,
])
for gap in result.critical_gaps:
lines.append(f" * {gap}")
lines.append("")
lines.extend([
"-" * 60,
"GAP DETAILS BY CATEGORY",
"-" * 60,
])
# Group by category
categories = {}
for gap in result.gaps:
cat = gap["category"]
if cat not in categories:
categories[cat] = []
categories[cat].append(gap)
for category, gaps in categories.items():
lines.append(f"\n{category}:")
for gap in gaps:
status_mark = "✓" if gap["status"] == "Complete" else "○"
lines.append(f" [{status_mark}] {gap['requirement']} ({gap['priority']})")
lines.extend([
"",
"-" * 60,
"RECOMMENDATIONS",
"-" * 60,
])
for i, rec in enumerate(result.recommendations, 1):
lines.append(f"{i}. {rec}")
lines.append("=" * 60)
return "\n".join(lines)
def interactive_mode():
"""Run interactive gap analysis session."""
print("=" * 60)
print("MDR 2017/745 Gap Analysis - Interactive Mode")
print("=" * 60)
device_name = input("\nDevice name: ").strip()
if not device_name:
device_name = "Unnamed Device"
print("\nDevice classes:")
print(" 1. Class I")
print(" 2. Class I (sterile)")
print(" 3. Class I (measuring)")
print(" 4. Class IIa")
print(" 5. Class IIb")
print(" 6. Class III")
class_map = {
"1": DeviceClass.I,
"2": DeviceClass.I_STERILE,
"3": DeviceClass.I_MEASURING,
"4": DeviceClass.IIA,
"5": DeviceClass.IIB,
"6": DeviceClass.III,
}
class_choice = input("\nSelect class (1-6): ").strip()
device_class = class_map.get(class_choice, DeviceClass.IIA)
analyzer = MDRGapAnalyzer(device_name, device_class)
print("\nFor each requirement, enter status:")
print(" c = Complete")
print(" i = In Progress")
print(" n = Not Started (default)")
print(" x = Not Applicable")
print(" Enter = Skip (Not Started)")
print("")
status_map = {
"c": GapStatus.COMPLETE,
"i": GapStatus.IN_PROGRESS,
"n": GapStatus.NOT_STARTED,
"x": GapStatus.NOT_APPLICABLE,
}
for gap in analyzer.gaps:
if gap.status == GapStatus.NOT_APPLICABLE:
continue
status_input = input(f"{gap.requirement} [c/i/n/x]: ").strip().lower()
if status_input in status_map:
gap.status = status_map[status_input]
result = analyzer.analyze()
print("\n" + format_text_output(result))
def main():
parser = argparse.ArgumentParser(
description="EU MDR 2017/745 Gap Analysis Tool"
)
parser.add_argument("--device", type=str, help="Device name")
parser.add_argument(
"--class",
dest="device_class",
choices=["I", "Is", "Im", "IIa", "IIb", "III"],
help="Device classification"
)
parser.add_argument(
"--output",
choices=["text", "json"],
default="text",
help="Output format"
)
parser.add_argument(
"--interactive",
action="store_true",
help="Run in interactive mode"
)
args = parser.parse_args()
if args.interactive:
interactive_mode()
return
if not args.device or not args.device_class:
parser.print_help()
print("\nError: --device and --class required (or use --interactive)")
sys.exit(1)
class_map = {
"I": DeviceClass.I,
"Is": DeviceClass.I_STERILE,
"Im": DeviceClass.I_MEASURING,
"IIa": DeviceClass.IIA,
"IIb": DeviceClass.IIB,
"III": DeviceClass.III,
}
analyzer = MDRGapAnalyzer(args.device, class_map[args.device_class])
result = analyzer.analyze()
if args.output == "json":
print(json.dumps(asdict(result), indent=2))
else:
print(format_text_output(result))
if __name__ == "__main__":
main()
Viết spec trước khi code, xác định tiêu chí chấp nhận, lập kế hoạch tính năng và sinh test từ đặc tả.
---
name: "spec-driven-workflow"
description: "Use when the user asks to write specs before code, define acceptance criteria, plan features before implementation, generate tests from specifications, or follow spec-first development practices."
---
# Spec-Driven Workflow — POWERFUL
## Overview
Spec-driven workflow enforces a single, non-negotiable rule: **write the specification BEFORE you write any code.** Not alongside. Not after. Before.
This is not documentation. This is a contract. A spec defines what the system MUST do, what it SHOULD do, and what it explicitly WILL NOT do. Every line of code you write traces back to a requirement in the spec. Every test traces back to an acceptance criterion. If it is not in the spec, it does not get built.
### Why Spec-First Matters
1. **Eliminates rework.** 60-80% of defects originate from requirements, not implementation. Catching ambiguity in a spec costs minutes; catching it in production costs days.
2. **Forces clarity.** If you cannot write what the system should do in plain language, you do not understand the problem well enough to write code.
3. **Enables parallelism.** Once a spec is approved, frontend, backend, QA, and documentation can all start simultaneously.
4. **Creates accountability.** The spec is the definition of done. No arguments about whether a feature is "complete" — either it satisfies the acceptance criteria or it does not.
5. **Feeds TDD directly.** Acceptance criteria in Given/When/Then format translate 1:1 into test cases. The spec IS the test plan.
### The Iron Law
```
NO CODE WITHOUT AN APPROVED SPEC.
NO EXCEPTIONS. NO "QUICK PROTOTYPES." NO "I'LL DOCUMENT IT LATER."
```
If the spec is not written, reviewed, and approved, implementation does not begin. Period.
---
## The Spec Format
Every spec follows this structure. No sections are optional — if a section does not apply, write "N/A — [reason]" so reviewers know it was considered, not forgotten.
### Mandatory Sections
| # | Section | Key Rules |
|---|---------|-----------|
| 1 | **Title and Metadata** | Author, date, status (Draft/In Review/Approved/Superseded), reviewers |
| 2 | **Context** | Why this feature exists. 2-4 paragraphs with evidence (metrics, tickets). |
| 3 | **Functional Requirements** | RFC 2119 keywords (MUST/SHOULD/MAY). Numbered FR-N. Each is atomic and testable. |
| 4 | **Non-Functional Requirements** | Performance, security, accessibility, scalability, reliability — all with measurable thresholds. |
| 5 | **Acceptance Criteria** | Given/When/Then format. Every AC references at least one FR-* or NFR-*. |
| 6 | **Edge Cases** | Numbered EC-N. Cover failure modes for every external dependency. |
| 7 | **API Contracts** | TypeScript-style interfaces. Cover success and error responses. |
| 8 | **Data Models** | Table format with field, type, constraints. Every entity from requirements must have a model. |
| 9 | **Out of Scope** | Explicit exclusions with reasons. Prevents scope creep during implementation. |
### RFC 2119 Keywords
| Keyword | Meaning |
|---------|---------|
| **MUST** | Absolute requirement. Non-conformant without it. |
| **MUST NOT** | Absolute prohibition. |
| **SHOULD** | Recommended. Omit only with documented justification. |
| **MAY** | Optional. Implementer's discretion. |
See [spec_format_guide.md](references/spec_format_guide.md) for the complete template with section-by-section examples, good/bad requirement patterns, and feature-type templates (CRUD, Integration, Migration).
See [acceptance_criteria_patterns.md](references/acceptance_criteria_patterns.md) for a full pattern library of Given/When/Then criteria across authentication, CRUD, search, file upload, payment, notification, and accessibility scenarios.
---
## Bounded Autonomy Rules
These rules define when an agent (human or AI) MUST stop and ask for guidance vs. when they can proceed independently.
### STOP and Ask When:
1. **Scope creep detected.** The implementation requires something not in the spec. Even if it seems obviously needed, STOP. The spec might have excluded it deliberately.
2. **Ambiguity exceeds 30%.** If you cannot determine the correct behavior from the spec for more than 30% of a given requirement, the spec is incomplete. Do not guess.
3. **Breaking changes required.** The implementation would change an existing API contract, database schema, or public interface. Always escalate.
4. **Security implications.** Any change that touches authentication, authorization, encryption, or PII handling requires explicit approval.
5. **Performance characteristics unknown.** If a requirement says "MUST complete in < 500ms" but you have no way to measure or guarantee that, escalate before implementing a guess.
6. **Cross-team dependencies.** If the spec requires coordination with another team or service, confirm the dependency before building against it.
### Continue Autonomously When:
1. **Spec is clear and unambiguous** for the current task.
2. **All acceptance criteria have passing tests** and you are refactoring internals.
3. **Changes are non-breaking** — no public API, schema, or behavior changes.
4. **Implementation is a direct translation** of a well-defined acceptance criterion.
5. **Error handling follows established patterns** already documented in the codebase.
### Escalation Protocol
When you must stop, provide:
```markdown
## Escalation: [Brief Title]
**Blocked on:** [requirement ID, e.g., FR-3]
**Question:** [Specific, answerable question — not "what should I do?"]
**Options considered:**
A. [Option] — Pros: [...] Cons: [...]
B. [Option] — Pros: [...] Cons: [...]
**My recommendation:** [A or B, with reasoning]
**Impact of waiting:** [What is blocked until this is resolved?]
```
Never escalate without a recommendation. Never present an open-ended question. Always give options.
See `references/bounded_autonomy_rules.md` for the complete decision matrix.
---
## Workflow — 6 Phases
### Phase 1: Gather Requirements
**Goal:** Understand what needs to be built and why.
1. **Interview the user.** Ask:
- What problem does this solve?
- Who are the users?
- What does success look like?
- What explicitly should NOT be built?
2. **Read existing code.** Understand the current system before proposing changes.
3. **Identify constraints.** Performance budgets, security requirements, backward compatibility.
4. **List unknowns.** Every unknown is a risk. Surface them now, not during implementation.
**Exit criteria:** You can explain the feature to someone unfamiliar with the project in 2 minutes.
### Phase 2: Write Spec
**Goal:** Produce a complete spec document following The Spec Format above.
1. Fill every section of the template. No section left blank.
2. Number all requirements (FR-*, NFR-*, AC-*, EC-*, OS-*).
3. Use RFC 2119 keywords precisely.
4. Write acceptance criteria in Given/When/Then format.
5. Define API contracts with TypeScript-style types.
6. List explicit exclusions in Out of Scope.
**Exit criteria:** The spec can be handed to a developer who was not in the requirements meeting, and they can implement the feature without asking clarifying questions.
### Phase 3: Validate Spec
**Goal:** Verify the spec is complete, consistent, and implementable.
Run `spec_validator.py` against the spec file:
```bash
python spec_validator.py --file spec.md --strict
```
Manual validation checklist:
- [ ] Every functional requirement has at least one acceptance criterion
- [ ] Every acceptance criterion is testable (no subjective language)
- [ ] API contracts cover all endpoints mentioned in requirements
- [ ] Data models cover all entities mentioned in requirements
- [ ] Edge cases cover failure modes for every external dependency
- [ ] Out of scope is explicit about what was considered and rejected
- [ ] Non-functional requirements have measurable thresholds
**Exit criteria:** Spec scores 80+ on validator, and all manual checklist items pass.
### Phase 4: Generate Tests
**Goal:** Extract test cases from acceptance criteria before writing implementation code.
Run `test_extractor.py` against the approved spec:
```bash
python test_extractor.py --file spec.md --framework pytest --output tests/
```
1. Each acceptance criterion becomes one or more test cases.
2. Each edge case becomes a test case.
3. Tests are stubs — they define the assertion but not the implementation.
4. All tests MUST fail initially (red phase of TDD).
**Exit criteria:** You have a test file where every test fails with "not implemented" or equivalent.
### Phase 5: Implement
**Goal:** Write code that makes failing tests pass, one acceptance criterion at a time.
1. Pick one acceptance criterion (start with the simplest).
2. Make its test(s) pass with minimal code.
3. Run the full test suite — no regressions.
4. Commit.
5. Pick the next acceptance criterion. Repeat.
**Rules:**
- Do NOT implement anything not in the spec.
- Do NOT optimize before all acceptance criteria pass.
- Do NOT refactor before all acceptance criteria pass.
- If you discover a missing requirement, STOP and update the spec first.
**Exit criteria:** All tests pass. All acceptance criteria satisfied.
### Phase 6: Self-Review
**Goal:** Verify implementation matches spec before marking done.
Run through the Self-Review Checklist below. If any item fails, fix it before declaring the task complete.
---
## Self-Review Checklist
Before marking any implementation as done, verify ALL of the following:
- [ ] **Every acceptance criterion has a passing test.** No exceptions. If AC-3 exists, a test for AC-3 exists and passes.
- [ ] **Every edge case has a test.** EC-1 through EC-N all have corresponding test cases.
- [ ] **No scope creep.** The implementation does not include features not in the spec. If you added something, either update the spec or remove it.
- [ ] **API contracts match implementation.** Request/response shapes in code match the spec exactly. Field names, types, status codes — all of it.
- [ ] **Error scenarios tested.** Every error response defined in the spec has a test that triggers it.
- [ ] **Non-functional requirements verified.** If the spec says < 500ms, you have evidence (benchmark, load test, profiling) that it meets the threshold.
- [ ] **Data model matches.** Database schema matches the spec. No extra columns, no missing constraints.
- [ ] **Out-of-scope items not built.** Double-check that nothing from the Out of Scope section leaked into the implementation.
---
## Integration with TDD Guide
Spec-driven workflow and TDD are complementary, not competing:
```
Spec-Driven Workflow TDD (Red-Green-Refactor)
───────────────────── ──────────────────────────
Phase 1: Gather Requirements
Phase 2: Write Spec
Phase 3: Validate Spec
Phase 4: Generate Tests ──→ RED: Tests exist and fail
Phase 5: Implement ──→ GREEN: Minimal code to pass
Phase 6: Self-Review ──→ REFACTOR: Clean up internals
```
**The handoff:** Spec-driven workflow produces the test stubs (Phase 4). TDD takes over from there. The spec tells you WHAT to test. TDD tells you HOW to implement.
Use `engineering-team/tdd-guide` for:
- Red-green-refactor cycle discipline
- Coverage analysis and gap detection
- Framework-specific test patterns (Jest, Pytest, JUnit)
Use `engineering/spec-driven-workflow` for:
- Defining what to build before building it
- Acceptance criteria authoring
- Completeness validation
- Scope control
---
## Examples
A complete worked example (Password Reset spec with extracted test cases) is available in [spec_format_guide.md](references/spec_format_guide.md#full-example-password-reset). It demonstrates all 9 sections, requirement numbering, acceptance criteria, edge cases, and the corresponding pytest stubs generated by `test_extractor.py`.
---
## Anti-Patterns
### 1. Coding Before Spec Approval
**Symptom:** "I'll start coding while the spec is being reviewed."
**Problem:** The review will surface changes. Now you have code that implements a rejected design.
**Rule:** Implementation does not begin until spec status is "Approved."
### 2. Vague Acceptance Criteria
**Symptom:** "The system should work well" or "The UI should be responsive."
**Problem:** Untestable. What does "well" mean? What does "responsive" mean?
**Rule:** Every acceptance criterion must be verifiable by a machine. If you cannot write a test for it, rewrite the criterion.
### 3. Missing Edge Cases
**Symptom:** Happy path is specified, error paths are not.
**Problem:** Developers invent error handling on the fly, leading to inconsistent behavior.
**Rule:** For every external dependency (API, database, file system, user input), specify at least one failure scenario.
### 4. Spec as Post-Hoc Documentation
**Symptom:** "Let me write the spec now that the feature is done."
**Problem:** This is documentation, not specification. It describes what was built, not what should have been built. It cannot catch design errors because the design is already frozen.
**Rule:** If the spec was written after the code, it is not a spec. Relabel it as documentation.
### 5. Gold-Plating Beyond Spec
**Symptom:** "While I was in there, I also added..."
**Problem:** Untested code. Unreviewed design. Potential for subtle bugs in the "bonus" feature.
**Rule:** If it is not in the spec, it does not get built. File a new spec for additional features.
### 6. Acceptance Criteria Without Requirement Traceability
**Symptom:** AC-7 exists but does not reference any FR-* or NFR-*.
**Problem:** Orphaned criteria mean either a requirement is missing or the criterion is unnecessary.
**Rule:** Every AC-* MUST reference at least one FR-* or NFR-*.
### 7. Skipping Validation
**Symptom:** "The spec looks fine, let's just start."
**Problem:** Missing sections discovered during implementation cause blocking delays.
**Rule:** Always run `spec_validator.py --strict` before starting implementation. Fix all warnings.
---
## Cross-References
- **`engineering-team/tdd-guide`** — Red-green-refactor cycle, test generation, coverage analysis. Use after Phase 4 of this workflow.
- **`engineering/focused-fix`** — Deep-dive feature repair. When a spec-driven implementation has systemic issues, use focused-fix for diagnosis.
- **`engineering/rag-architect`** — If the feature involves retrieval or knowledge systems, use rag-architect for the technical design within the spec.
- **`references/spec_format_guide.md`** — Complete template with section-by-section explanations.
- **`references/bounded_autonomy_rules.md`** — Full decision matrix for when to stop vs. continue.
- **`references/acceptance_criteria_patterns.md`** — Pattern library for writing Given/When/Then criteria.
---
## Tools
| Script | Purpose | Key Flags |
|--------|---------|-----------|
| `spec_generator.py` | Generate spec template from feature name/description | `--name`, `--description`, `--format`, `--json` |
| `spec_validator.py` | Validate spec completeness (0-100 score) | `--file`, `--strict`, `--json` |
| `test_extractor.py` | Extract test stubs from acceptance criteria | `--file`, `--framework`, `--output`, `--json` |
```bash
# Generate a spec template
python spec_generator.py --name "User Authentication" --description "OAuth 2.0 login flow"
# Validate a spec
python spec_validator.py --file specs/auth.md --strict
# Extract test cases
python test_extractor.py --file specs/auth.md --framework pytest --output tests/test_auth.py
```
FILE:references/acceptance_criteria_patterns.md
# Acceptance Criteria Patterns
A pattern library for writing Given/When/Then acceptance criteria across common feature types. Use these as starting points — adapt to your domain.
---
## Pattern Structure
Every acceptance criterion follows this structure:
```
### AC-N: [Descriptive name] (FR-N, NFR-N)
Given [precondition — the system/user is in this state]
When [trigger — the user or system performs this action]
Then [outcome — this observable, testable result occurs]
And [additional outcome — and this also happens]
```
**Rules:**
1. One scenario per AC. Multiple Given/When/Then blocks = multiple ACs.
2. Every AC references at least one FR-* or NFR-*.
3. Outcomes must be observable and testable — no subjective language.
4. Preconditions must be achievable in a test setup.
---
## Authentication Patterns
### Login — Happy Path
```markdown
### AC-1: Successful login with valid credentials (FR-1)
Given a registered user with email "user@example.com" and password "V@lidP4ss!"
When they POST /api/auth/login with email "user@example.com" and password "V@lidP4ss!"
Then the response status is 200
And the response body contains a valid JWT access token
And the response body contains a refresh token
And the access token expires in 24 hours
```
### Login — Invalid Credentials
```markdown
### AC-2: Login rejected with wrong password (FR-1)
Given a registered user with email "user@example.com"
When they POST /api/auth/login with email "user@example.com" and an incorrect password
Then the response status is 401
And the response body contains error code "INVALID_CREDENTIALS"
And no token is issued
And the failed attempt is logged
```
### Login — Account Locked
```markdown
### AC-3: Login rejected for locked account (FR-1, NFR-S2)
Given a user whose account is locked due to 5 consecutive failed login attempts
When they POST /api/auth/login with correct credentials
Then the response status is 403
And the response body contains error code "ACCOUNT_LOCKED"
And the response includes a "retryAfter" field with seconds until unlock
```
### Token Refresh
```markdown
### AC-4: Token refresh with valid refresh token (FR-3)
Given a user with a valid, non-expired refresh token
When they POST /api/auth/refresh with that refresh token
Then the response status is 200
And a new access token is issued
And the old refresh token is invalidated
And a new refresh token is issued (rotation)
```
### Logout
```markdown
### AC-5: Logout invalidates session (FR-4)
Given an authenticated user with a valid access token
When they POST /api/auth/logout with that token
Then the response status is 204
And the access token is no longer accepted for API calls
And the refresh token is invalidated
```
---
## CRUD Patterns
### Create
```markdown
### AC-6: Create resource with valid data (FR-1)
Given an authenticated user with "editor" role
When they POST /api/resources with valid payload {name: "Test", type: "A"}
Then the response status is 201
And the response body contains the created resource with a generated UUID
And the resource's "createdAt" field is set to the current UTC timestamp
And the resource's "createdBy" field matches the authenticated user's ID
```
### Create — Validation Failure
```markdown
### AC-7: Create resource rejected with invalid data (FR-1)
Given an authenticated user
When they POST /api/resources with payload missing required field "name"
Then the response status is 400
And the response body contains error code "VALIDATION_ERROR"
And the response body contains field-level detail: {"name": "Required field"}
And no resource is created in the database
```
### Read — Single Item
```markdown
### AC-8: Read resource by ID (FR-2)
Given an existing resource with ID "abc-123"
When an authenticated user GETs /api/resources/abc-123
Then the response status is 200
And the response body contains the resource with all fields
```
### Read — Not Found
```markdown
### AC-9: Read non-existent resource returns 404 (FR-2)
Given no resource exists with ID "nonexistent-id"
When an authenticated user GETs /api/resources/nonexistent-id
Then the response status is 404
And the response body contains error code "NOT_FOUND"
```
### Update
```markdown
### AC-10: Update resource with valid data (FR-3)
Given an existing resource with ID "abc-123" owned by the authenticated user
When they PATCH /api/resources/abc-123 with {name: "Updated Name"}
Then the response status is 200
And the resource's "name" field is "Updated Name"
And the resource's "updatedAt" field is updated to the current UTC timestamp
And fields not included in the patch are unchanged
```
### Update — Ownership Check
```markdown
### AC-11: Update rejected for non-owner (FR-3, FR-6)
Given an existing resource with ID "abc-123" owned by user "other-user"
When the authenticated user (not "other-user") PATCHes /api/resources/abc-123
Then the response status is 403
And the response body contains error code "FORBIDDEN"
And the resource is unchanged
```
### Delete — Soft Delete
```markdown
### AC-12: Soft delete resource (FR-5)
Given an existing resource with ID "abc-123" owned by the authenticated user
When they DELETE /api/resources/abc-123
Then the response status is 204
And the resource's "deletedAt" field is set to the current UTC timestamp
And the resource no longer appears in GET /api/resources (list endpoint)
And the resource still exists in the database (soft deleted)
```
### List — Pagination
```markdown
### AC-13: List resources with default pagination (FR-4)
Given 50 resources exist for the authenticated user
When they GET /api/resources without pagination parameters
Then the response status is 200
And the response contains the first 20 resources (default page size)
And the response includes "totalCount: 50"
And the response includes "page: 1"
And the response includes "pageSize: 20"
And the response includes "hasNextPage: true"
```
### List — Filtered
```markdown
### AC-14: List resources with type filter (FR-4)
Given 30 resources of type "A" and 20 resources of type "B" exist
When the authenticated user GETs /api/resources?type=A
Then the response status is 200
And all returned resources have type "A"
And the response "totalCount" is 30
```
---
## Search Patterns
### Basic Search
```markdown
### AC-15: Search returns matching results (FR-7)
Given resources with names "Alpha Report", "Beta Analysis", "Alpha Summary" exist
When the user GETs /api/resources?q=Alpha
Then the response contains "Alpha Report" and "Alpha Summary"
And the response does not contain "Beta Analysis"
And results are ordered by relevance score (descending)
```
### Search — Empty Results
```markdown
### AC-16: Search with no matches returns empty list (FR-7)
Given no resources match the query "xyznonexistent"
When the user GETs /api/resources?q=xyznonexistent
Then the response status is 200
And the response contains an empty "items" array
And "totalCount" is 0
```
### Search — Special Characters
```markdown
### AC-17: Search handles special characters safely (FR-7, NFR-S1)
Given resources exist in the database
When the user GETs /api/resources?q="; DROP TABLE resources;--
Then the response status is 200
And no SQL injection occurs
And the search treats the input as a literal string
```
---
## File Upload Patterns
### Upload — Happy Path
```markdown
### AC-18: Upload file within size limit (FR-8)
Given an authenticated user
When they POST /api/files with a 5MB PNG file
Then the response status is 201
And the response contains the file's URL, size, and MIME type
And the file is stored in the configured storage backend
And the file is associated with the authenticated user
```
### Upload — Size Exceeded
```markdown
### AC-19: Upload rejected for oversized file (FR-8)
Given the maximum file size is 10MB
When the user POSTs /api/files with a 15MB file
Then the response status is 413
And the response contains error code "FILE_TOO_LARGE"
And no file is stored
```
### Upload — Invalid Type
```markdown
### AC-20: Upload rejected for disallowed file type (FR-8, NFR-S3)
Given allowed file types are PNG, JPG, PDF
When the user POSTs /api/files with an .exe file
Then the response status is 415
And the response contains error code "UNSUPPORTED_MEDIA_TYPE"
And no file is stored
```
---
## Payment Patterns
### Charge — Happy Path
```markdown
### AC-21: Successful payment charge (FR-10)
Given a user with a valid payment method on file
When they POST /api/payments with amount 49.99 and currency "USD"
Then the payment gateway is charged $49.99
And the response status is 201
And the response contains a transaction ID
And a payment record is created with status "completed"
And a receipt email is sent to the user
```
### Charge — Declined
```markdown
### AC-22: Payment declined by gateway (FR-10)
Given a user with an expired credit card on file
When they POST /api/payments with amount 49.99
Then the payment gateway returns a decline
And the response status is 402
And the response contains error code "PAYMENT_DECLINED"
And no payment record is created with status "completed"
And the user is prompted to update their payment method
```
### Charge — Idempotency
```markdown
### AC-23: Duplicate payment request is idempotent (FR-10, NFR-R1)
Given a payment was successfully processed with idempotency key "key-123"
When the same request is sent again with idempotency key "key-123"
Then the response status is 200
And the response contains the original transaction ID
And the user is NOT charged a second time
```
---
## Notification Patterns
### Email Notification
```markdown
### AC-24: Email notification sent on event (FR-11)
Given a user with notification preferences set to "email"
When their order status changes to "shipped"
Then an email is sent to their registered email address
And the email subject contains the order number
And the email body contains the tracking URL
And a notification record is created with status "sent"
```
### Notification — Delivery Failure
```markdown
### AC-25: Failed notification is retried (FR-11, NFR-R2)
Given the email service returns a 5xx error on first attempt
When a notification is triggered
Then the system retries up to 3 times with exponential backoff (1s, 4s, 16s)
And if all retries fail, the notification status is set to "failed"
And an alert is sent to the ops channel
```
---
## Negative Test Patterns
### Unauthorized Access
```markdown
### AC-26: Unauthenticated request rejected (NFR-S1)
Given no authentication token is provided
When the user GETs /api/resources
Then the response status is 401
And the response contains error code "AUTHENTICATION_REQUIRED"
And no resource data is returned
```
### Invalid Input — Type Mismatch
```markdown
### AC-27: String provided for numeric field (FR-1)
Given the "quantity" field expects an integer
When the user POSTs with quantity: "abc"
Then the response status is 400
And the response body contains field error: {"quantity": "Must be an integer"}
```
### Rate Limiting
```markdown
### AC-28: Rate limit enforced (NFR-S2)
Given the rate limit is 100 requests per minute per API key
When the user sends the 101st request within 60 seconds
Then the response status is 429
And the response includes header "Retry-After" with seconds until reset
And the response contains error code "RATE_LIMITED"
```
### Concurrent Modification
```markdown
### AC-29: Optimistic locking prevents lost updates (NFR-R1)
Given a resource with version 5
When user A PATCHes with version 5 and user B PATCHes with version 5 simultaneously
Then one succeeds with status 200 (version becomes 6)
And the other receives status 409 with error code "CONFLICT"
And the 409 response includes the current version number
```
---
## Performance Criteria Patterns
### Response Time
```markdown
### AC-30: API response time under load (NFR-P1)
Given the system is handling 1,000 concurrent users
When a user GETs /api/dashboard
Then the response is returned in < 500ms (p95)
And the response is returned in < 1000ms (p99)
```
### Throughput
```markdown
### AC-31: System handles target throughput (NFR-P2)
Given normal production traffic patterns
When the system receives 5,000 requests per second
Then all requests are processed without queue overflow
And error rate remains below 0.1%
```
### Resource Usage
```markdown
### AC-32: Memory usage within bounds (NFR-P3)
Given the service is processing normal traffic
When measured over a 24-hour period
Then memory usage does not exceed 512MB RSS
And no memory leaks are detected (RSS growth < 5% over 24h)
```
---
## Accessibility Criteria Patterns
### Keyboard Navigation
```markdown
### AC-33: Form is fully keyboard navigable (NFR-A1)
Given the user is on the login page using only a keyboard
When they press Tab
Then focus moves through: email field -> password field -> submit button
And each focused element has a visible focus indicator
And pressing Enter on the submit button submits the form
```
### Screen Reader
```markdown
### AC-34: Error messages announced to screen readers (NFR-A2)
Given the user submits the form with invalid data
When validation errors appear
Then each error is associated with its form field via aria-describedby
And the error container has role="alert" for immediate announcement
And the first error field receives focus
```
### Color Contrast
```markdown
### AC-35: Text meets contrast requirements (NFR-A3)
Given the default theme is active
When measuring text against background colors
Then all body text meets 4.5:1 contrast ratio (WCAG AA)
And all large text (18px+ or 14px+ bold) meets 3:1 contrast ratio
And all interactive element states (hover, focus, active) meet 3:1
```
### Reduced Motion
```markdown
### AC-36: Animations respect user preference (NFR-A4)
Given the user has enabled "prefers-reduced-motion" in their OS settings
When they load any page with animations
Then all non-essential animations are disabled
And essential animations (e.g., loading spinner) use a reduced version
And no content is hidden behind animation-only interactions
```
---
## Writing Tips
### Do
- Start Given with the system/user state, not the action
- Make When a single, specific trigger
- Make Then observable — status codes, field values, side effects
- Include And for additional assertions on the same outcome
- Reference requirement IDs in the AC title
### Do Not
- Write "Then the system works correctly" (not testable)
- Combine multiple scenarios in one AC
- Use subjective words: "quickly", "properly", "nicely", "user-friendly"
- Skip the precondition — Given is required even if it seems obvious
- Write Given/When/Then as prose paragraphs — use the structured format
### Smell Tests
If your AC has any of these, rewrite it:
| Smell | Example | Fix |
|-------|---------|-----|
| No Given clause | "When user clicks, then page loads" | Add "Given user is on the dashboard" |
| Vague Then | "Then it works" | Specify status code, body, side effects |
| Multiple Whens | "When user clicks A and then clicks B" | Split into two ACs |
| Implementation detail | "Then the Redux store is updated" | Focus on user-observable outcome |
| No requirement reference | "AC-5: Dashboard loads" | "AC-5: Dashboard loads (FR-7)" |
FILE:references/bounded_autonomy_rules.md
# Bounded Autonomy Rules
Decision framework for when an agent (human or AI) should stop and ask vs. continue working autonomously during spec-driven development.
---
## The Core Principle
**Autonomy is earned by clarity.** The clearer the spec, the more autonomy the implementer has. The more ambiguous the spec, the more the implementer must stop and ask.
This is not about trust. It is about risk. A clear spec means low risk of building the wrong thing. An ambiguous spec means high risk.
---
## Decision Matrix
| Signal | Action | Rationale |
|--------|--------|-----------|
| Spec is Approved, requirement is clear, tests exist | **Continue** | Low risk. Build it. |
| Requirement is clear but no test exists yet | **Continue** (write the test first) | You can infer the test from the requirement. |
| Requirement uses SHOULD/MAY keywords | **Continue** with your best judgment | These are intentionally flexible. Document your choice. |
| Requirement is ambiguous (multiple valid interpretations) | **STOP** if ambiguity > 30% of the task | Ask the spec author to clarify. |
| Implementation requires changing an API contract | **STOP** always | Breaking changes need explicit approval. |
| Implementation requires a new database migration | **STOP** if it changes existing columns/tables | New tables are lower risk than schema changes. |
| Security-related change (auth, crypto, PII) | **STOP** always | Security changes need review regardless of spec clarity. |
| Performance-critical path with no benchmark data | **STOP** | You cannot prove NFR compliance without measurement. |
| Bug found in existing code unrelated to spec | **STOP** — file a separate issue | Do not fix unrelated bugs in a spec-scoped implementation. |
| Spec says "N/A" for a section you think needs content | **STOP** | The author may have a reason, or they may have missed it. |
---
## Ambiguity Scoring
When you encounter ambiguity, quantify it before deciding to stop or continue.
### How to Score Ambiguity
For each requirement you are implementing, ask:
1. **Can I write a test for this right now?** (No = +20% ambiguity)
2. **Are there multiple valid interpretations?** (Yes = +20% ambiguity)
3. **Does the spec contradict itself?** (Yes = +30% ambiguity)
4. **Am I making assumptions about user behavior?** (Yes = +15% ambiguity)
5. **Does this depend on an undocumented external system?** (Yes = +15% ambiguity)
### Threshold
| Ambiguity Score | Action |
|-----------------|--------|
| 0-15% | Continue. Minor ambiguity is normal. Document your interpretation. |
| 16-30% | Continue with caution. Add a comment explaining your interpretation. Flag in PR. |
| 31-50% | STOP. Ask the spec author one specific question. Do not continue until answered. |
| 51%+ | STOP. The spec is incomplete. Request a revision before proceeding. |
### Example
**Requirement:** "FR-7: The system MUST notify the user when their order ships."
Questions:
1. Can I write a test? Partially — I know WHAT to test but not HOW (email? push? in-app?). +20%
2. Multiple interpretations? Yes — notification channel is unclear. +20%
3. Contradicts itself? No. +0%
4. Assuming user behavior? Yes — I am assuming they want email. +15%
5. Undocumented external system? Maybe — depends on notification service. +15%
**Total: 70%.** STOP. The spec needs to specify the notification channel.
---
## Scope Creep Detection
### What Is Scope Creep?
Scope creep is implementing functionality not described in the spec. It includes:
- Adding features the spec does not mention
- "Improving" behavior beyond what acceptance criteria require
- Handling edge cases the spec explicitly excluded
- Refactoring unrelated code "while you're in there"
- Building infrastructure for future features
### Detection Patterns
| Pattern | Example | Risk |
|---------|---------|------|
| "While I'm here..." | Refactoring a utility function unrelated to the spec | Medium — unreviewed changes |
| "This would be easy to add..." | Adding a search filter the spec does not mention | High — untested, unspecified |
| "Users will probably want..." | Building a feature based on assumption | High — may conflict with future specs |
| "This is obviously needed..." | Adding logging, metrics, or caching not in NFRs | Medium — may be overkill or wrong approach |
| "The spec forgot to mention..." | Building something the spec excluded | Critical — may be deliberately excluded |
### Response Protocol
When you detect scope creep in your own work:
1. **Stop immediately.** Do not commit the extra code.
2. **Check Out of Scope.** Is this item explicitly excluded?
3. **If excluded:** Delete the code. The spec author had a reason.
4. **If not mentioned:** File a note for the spec author. Ask if it should be added.
5. **If approved:** Update the spec FIRST, then implement.
---
## Breaking Change Identification
### What Counts as a Breaking Change?
A breaking change is any modification that could cause existing clients, tests, or integrations to fail.
| Category | Breaking | Not Breaking |
|----------|----------|--------------|
| API endpoint removed | Yes | - |
| API endpoint added | - | No |
| Required field added to request | Yes | - |
| Optional field added to request | - | No |
| Field removed from response | Yes | - |
| Field added to response | - | No (usually) |
| Status code changed | Yes | - |
| Error code string changed | Yes | - |
| Database column removed | Yes | - |
| Database column added (nullable) | - | No |
| Database column added (not null, no default) | Yes | - |
| Enum value removed | Yes | - |
| Enum value added | - | No (usually) |
| Behavior change for existing input | Yes | - |
### Breaking Change Protocol
1. **Identify** the breaking change before implementing it.
2. **Escalate** immediately — do not implement without approval.
3. **Propose** a migration path (versioned API, feature flag, deprecation period).
4. **Document** the breaking change in the spec's changelog.
---
## Security Implication Checklist
Any change touching the following areas MUST be escalated, even if the spec seems clear.
### Always Escalate
- [ ] Authentication logic (login, logout, token generation)
- [ ] Authorization logic (role checks, permission gates)
- [ ] Encryption/hashing (algorithm choice, key management)
- [ ] PII handling (storage, transmission, logging)
- [ ] Input validation bypass (new endpoints, parameter changes)
- [ ] Rate limiting changes (thresholds, scope)
- [ ] CORS or CSP policy changes
- [ ] File upload handling
- [ ] SQL/NoSQL query construction (injection risk)
- [ ] Deserialization of user input
- [ ] Redirect URLs from user input (open redirect risk)
- [ ] Secrets in code, config, or logs
### Security Escalation Template
```markdown
## Security Escalation: [Title]
**Affected area:** [authentication/authorization/encryption/PII/etc.]
**Spec reference:** [FR-N or NFR-SN]
**Risk:** [What could go wrong if implemented incorrectly]
**Current protection:** [What exists today]
**Proposed change:** [What the spec requires]
**My concern:** [Specific security question]
**Recommendation:** [Proposed approach with security rationale]
```
---
## Escalation Templates
### Template 1: Ambiguous Requirement
```markdown
## Escalation: Ambiguous Requirement
**Blocked on:** FR-7 ("notify the user when their order ships")
**Ambiguity score:** 70%
**Question:** What notification channel should be used?
**Options considered:**
A. Email only — Pros: simple, reliable. Cons: not real-time.
B. Email + in-app notification — Pros: covers both async and real-time. Cons: more implementation effort.
C. Configurable per user — Pros: maximum flexibility. Cons: requires preference UI (not in spec).
**My recommendation:** B (email + in-app). Covers most use cases without requiring new UI.
**Impact of waiting:** Cannot implement FR-7 until resolved. No other work blocked.
```
### Template 2: Missing Edge Case
```markdown
## Escalation: Missing Edge Case
**Related to:** FR-3 (password reset link expires after 1 hour)
**Scenario:** User clicks a reset link, but their account was deleted between requesting and clicking.
**Not in spec:** Edge cases section does not cover this.
**Options considered:**
A. Show generic "link invalid" error — Pros: secure (no info leak). Cons: confusing for deleted user.
B. Show "account not found" error — Pros: clear. Cons: confirms account deletion to link holder.
**My recommendation:** A. Security over clarity — do not reveal account existence.
**Impact of waiting:** Can implement other ACs; this is blocking only AC-2 completion.
```
### Template 3: Potential Breaking Change
```markdown
## Escalation: Potential Breaking Change
**Spec requires:** Adding required field "role" to POST /api/users request (FR-6)
**Current behavior:** POST /api/users accepts {email, password, displayName}
**Breaking:** Yes — existing clients will get 400 errors (missing required field)
**Options considered:**
A. Make "role" required as spec says — Pros: matches spec. Cons: breaks mobile app v2.1.
B. Make "role" optional with default "user" — Pros: backward compatible. Cons: deviates from spec.
C. Version the API (v2) — Pros: clean separation. Cons: maintenance burden.
**My recommendation:** B. Default to "user" for backward compatibility. Update spec to reflect MAY instead of MUST.
**Impact of waiting:** Frontend team is building against the new contract. Need answer within 2 days.
```
### Template 4: Scope Creep Proposal
```markdown
## Escalation: Potential Addition to Spec
**Context:** While implementing FR-2 (password validation), I noticed the spec does not mention password strength feedback.
**Not in spec:** No requirement for showing strength indicators.
**Checked Out of Scope:** Not listed there either.
**Proposal:** Add FR-7: "The system SHOULD display password strength feedback during registration."
**Effort:** ~2 hours additional implementation.
**Question:** Should this be added to current spec, filed as a separate spec, or skipped?
**Impact of waiting:** FR-2 implementation is not blocked. This is an enhancement question only.
```
---
## Quick Reference Card
```
CONTINUE if:
- Spec is approved
- Requirement uses MUST and is unambiguous
- Tests can be written directly from the AC
- Changes are additive and non-breaking
- You are refactoring internals only (no behavior change)
STOP if:
- Ambiguity > 30%
- Any breaking change
- Any security-related change
- Spec says N/A but you think it shouldn't
- You are about to build something not in the spec
- You cannot write a test for the requirement
- External dependency is undocumented
```
---
## Anti-Patterns in Autonomy
### 1. "I'll Ask Later"
Continuing past an ambiguity checkpoint because asking feels slow. The rework from building the wrong thing is always slower.
### 2. "It's Obviously Needed"
Assuming a missing feature was accidentally omitted. It may have been deliberately excluded. Check Out of Scope first.
### 3. "The Spec Is Wrong"
Implementing what you think the spec SHOULD say instead of what it DOES say. If the spec is wrong, escalate. Do not silently "fix" it.
### 4. "Just This Once"
Bypassing the escalation protocol for a "small" change. Small changes compound. The protocol exists because humans are bad at judging risk in the moment.
### 5. "I Already Built It"
Presenting completed work that was never in the spec and hoping it gets accepted. This creates review pressure and wastes everyone's time if rejected. Ask BEFORE building.
FILE:references/spec_format_guide.md
# Spec Format Guide
Complete reference for writing feature specifications. Every section is explained with examples, rationale, and common mistakes.
---
## The Spec Document Structure
A spec has 8 mandatory sections. If a section does not apply, write "N/A — [reason]" so reviewers know it was considered, not skipped.
```
1. Title and Metadata
2. Context
3. Functional Requirements
4. Non-Functional Requirements
5. Acceptance Criteria
6. Edge Cases and Error Scenarios
7. API Contracts
8. Data Models
9. Out of Scope
```
---
## Section 1: Title and Metadata
```markdown
# Spec: [Feature Name]
**Author:** Jane Doe
**Date:** 2026-03-25
**Status:** Draft | In Review | Approved | Superseded
**Reviewers:** John Smith, Alice Chen
**Related specs:** SPEC-018 (User Registration), SPEC-023 (Session Management)
```
### Status Lifecycle
| Status | Meaning | Who Can Change |
|--------|---------|----------------|
| Draft | Author is still writing. Not ready for review. | Author |
| In Review | Ready for feedback. Implementation blocked. | Author |
| Approved | Reviewed and accepted. Implementation may begin. | Reviewer |
| Superseded | Replaced by a newer spec. Link to replacement. | Author |
**Rule:** Implementation MUST NOT begin until status is "Approved."
---
## Section 2: Context
The context section answers: **Why does this feature exist?**
### What to Include
- The problem being solved (with evidence: support tickets, metrics, user research)
- The current state (what exists today and what is broken or missing)
- The business justification (revenue impact, cost savings, user retention)
- Constraints or dependencies (regulatory, technical, timeline)
### What to Exclude
- Implementation details (that is the engineer's job)
- Solution proposals (the spec says WHAT, not HOW)
- Lengthy background (2-4 paragraphs maximum)
### Good Example
```markdown
## Context
Users who forget their passwords currently have no self-service recovery.
Support handles ~200 password reset requests per week, consuming approximately
8 hours of agent time at $45/hour ($360/week, $18,720/year). Additionally,
12% of users who contact support for a reset never return.
This feature provides self-service password reset via email, eliminating
support burden and reducing user churn from the reset flow.
```
### Bad Example
```markdown
## Context
We need a password reset feature. Users forget their passwords sometimes
and need to reset them. We should build this.
```
**Why it is bad:** No evidence, no metrics, no business justification. "We should build this" is not a reason.
---
## Section 3: Functional Requirements — RFC 2119
### RFC 2119 Keywords
These keywords have precise meanings per [RFC 2119](https://www.ietf.org/rfc/rfc2119.txt). Do not use them casually.
| Keyword | Meaning | Testing Implication |
|---------|---------|---------------------|
| **MUST** | Absolute requirement. The implementation is non-conformant without this. | Must have a passing test. Failure = release blocker. |
| **MUST NOT** | Absolute prohibition. Doing this = broken implementation. | Must have a test proving this cannot happen. |
| **SHOULD** | Strongly recommended. Can be omitted only with documented justification. | Should have a test. Omission requires written rationale. |
| **SHOULD NOT** | Strongly discouraged. Can be done only with documented justification. | Should have a test confirming the behavior does not occur. |
| **MAY** | Truly optional. Implementer's discretion. | Test is optional. Document if implemented. |
### Writing Good Requirements
**Each requirement MUST be:**
1. **Atomic** — One behavior per requirement. Not "The system MUST authenticate users and log them in."
2. **Testable** — You can write a test that proves it works or does not.
3. **Numbered** — Sequential FR-N format for traceability.
4. **Specific** — No ambiguous adjectives ("fast", "secure", "user-friendly").
### Good Requirements
```markdown
- FR-1: The system MUST accept login via email and password.
- FR-2: The system MUST reject passwords shorter than 8 characters.
- FR-3: The system MUST return a JWT access token on successful login.
- FR-4: The system MUST NOT include the password hash in any API response.
- FR-5: The system SHOULD support "remember me" with a 30-day refresh token.
- FR-6: The system MAY display last login time on the dashboard.
```
### Bad Requirements
```markdown
- FR-1: The login system must be fast and secure.
(Untestable: what is "fast"? What is "secure"?)
- FR-2: The system must handle all edge cases.
(Vague: which edge cases? This delegates the spec to the implementer.)
- FR-3: Users should be able to log in easily.
(Subjective: "easily" is not measurable.)
```
---
## Section 4: Non-Functional Requirements
Non-functional requirements define quality attributes. Every requirement needs a **measurable threshold**.
### Categories
#### Performance
```markdown
- NFR-P1: Login API MUST respond in < 500ms (p95) under 1,000 concurrent users.
- NFR-P2: Dashboard page MUST achieve Largest Contentful Paint < 2.5s.
- NFR-P3: Search results MUST return within 200ms for queries under 100 characters.
```
**Bad:** "The system should be fast." (Not measurable.)
#### Security
```markdown
- NFR-S1: All API endpoints MUST require authentication except /health and /login.
- NFR-S2: Failed login attempts MUST be rate-limited to 5 per minute per IP.
- NFR-S3: Passwords MUST be hashed with bcrypt (cost factor >= 12).
- NFR-S4: Session tokens MUST be invalidated on password change.
```
#### Accessibility
```markdown
- NFR-A1: All form inputs MUST have associated labels (WCAG 1.3.1).
- NFR-A2: Color contrast MUST meet 4.5:1 ratio (WCAG 1.4.3).
- NFR-A3: All interactive elements MUST be keyboard-navigable (WCAG 2.1.1).
```
#### Scalability
```markdown
- NFR-SC1: The system SHOULD handle 50,000 registered users.
- NFR-SC2: Database queries MUST use indexes; no full table scans on tables > 10K rows.
```
#### Reliability
```markdown
- NFR-R1: The authentication service MUST maintain 99.9% uptime (< 8.77h downtime/year).
- NFR-R2: Data MUST NOT be lost on service restart (durable storage required).
```
---
## Section 5: Acceptance Criteria — Given/When/Then
Acceptance criteria are the contract between the spec author and the implementer. They define "done."
### The Given/When/Then Pattern
```
Given [precondition — the world is in this state]
When [action — the user or system does this]
Then [outcome — this observable result occurs]
And [additional outcome — and also this]
```
### Rules for Acceptance Criteria
1. **Every AC MUST reference at least one FR-* or NFR-*.** Orphaned criteria indicate missing requirements.
2. **Every AC MUST be testable by a machine.** If you cannot write an automated test, rewrite the criterion.
3. **No subjective language.** Not "should look good" but "MUST render within the design-system grid."
4. **One scenario per AC.** If you have multiple Given/When/Then blocks, split into separate ACs.
### Example: Authentication Feature
```markdown
### AC-1: Successful login (FR-1, FR-3)
Given a registered user with email "user@example.com" and password "P@ssw0rd123"
When they POST /api/auth/login with those credentials
Then they receive a 200 response with a valid JWT token
And the token expires in 24 hours
And the response includes the user's display name
### AC-2: Invalid password (FR-1)
Given a registered user with email "user@example.com"
When they POST /api/auth/login with an incorrect password
Then they receive a 401 response
And the response body contains error "INVALID_CREDENTIALS"
And no token is issued
### AC-3: Short password rejected on registration (FR-2)
Given a new user attempting to register
When they submit a password with 7 characters
Then they receive a 400 response
And the response body contains error "PASSWORD_TOO_SHORT"
And the account is not created
```
### Common Mistakes
| Mistake | Example | Fix |
|---------|---------|-----|
| Vague outcome | "Then the system works correctly" | "Then the response status is 200 and body contains {field: value}" |
| Missing precondition | "When user logs in, then token is issued" | "Given a registered user, when they POST valid credentials, then..." |
| Multiple scenarios | AC with 3 different When clauses | Split into 3 separate ACs |
| No FR reference | "AC-5: User sees dashboard" | "AC-5: User sees dashboard (FR-7)" |
---
## Section 6: Edge Cases and Error Scenarios
### What Counts as an Edge Case
- Invalid or malformed input
- External service failures (API down, timeout, rate-limited)
- Concurrent operations (race conditions)
- Boundary values (empty string, max length, zero, negative numbers)
- State conflicts (already exists, already deleted, expired)
### Format
```markdown
- EC-1: Empty email field → Return 400 with error "EMAIL_REQUIRED". Do not call auth service.
- EC-2: Email exceeds 255 characters → Return 400 with error "EMAIL_TOO_LONG".
- EC-3: OAuth provider returns 503 → Return 503 with "Service temporarily unavailable". Retry after 30s.
- EC-4: Two users register same email simultaneously → First succeeds, second gets 409 Conflict.
- EC-5: User clicks reset link after password was already changed → Show "Link already used."
```
### Coverage Rule
For every external dependency, specify at least one failure:
- Database: connection lost, timeout, constraint violation
- API: 4xx, 5xx, timeout, invalid response
- File system: file not found, permission denied, disk full
- User input: empty, too long, wrong type, injection attempt
---
## Section 7: API Contracts
### Notation
Use TypeScript-style interfaces. They are readable by both frontend and backend engineers.
```typescript
interface CreateUserRequest {
email: string; // MUST be valid email, max 255 chars
password: string; // MUST be 8-128 chars
displayName: string; // MUST be 1-100 chars, no HTML
role?: "user" | "admin"; // Default: "user"
}
```
### What to Define
For each endpoint:
1. **HTTP method and path** (e.g., POST /api/users)
2. **Request body** (fields, types, constraints, defaults)
3. **Success response** (status code, body shape)
4. **Error responses** (each error code with its status and body)
5. **Headers** (Authorization, Content-Type, custom headers)
### Error Response Convention
```typescript
interface ApiError {
error: string; // Machine-readable code: "INVALID_CREDENTIALS"
message: string; // Human-readable: "The email or password is incorrect."
details?: Record<string, string>; // Field-level errors for validation
}
```
Always include:
- 400 for validation errors
- 401 for authentication failures
- 403 for authorization failures
- 404 for not found
- 409 for conflicts
- 429 for rate limiting
- 500 for unexpected errors (keep it generic — do not leak internals)
---
## Section 8: Data Models
### Table Format
```markdown
### User
| Field | Type | Constraints |
|-------|------|-------------|
| id | UUID | PK, auto-generated, immutable |
| email | varchar(255) | Unique, not null, valid email |
| passwordHash | varchar(60) | Not null, bcrypt, never in API responses |
| displayName | varchar(100) | Not null |
| role | enum('user','admin') | Default: 'user' |
| createdAt | timestamp | UTC, immutable, auto-set |
| updatedAt | timestamp | UTC, auto-updated |
| deletedAt | timestamp | Null unless soft-deleted |
```
### Rules
1. **Every entity in requirements MUST have a data model.** If FR-1 mentions "users", there must be a User model.
2. **Constraints MUST match requirements.** If FR-2 says passwords >= 8 chars, the model must note that.
3. **Include indexes.** If NFR-P1 says < 500ms queries, note which fields need indexes.
4. **Specify soft vs. hard delete.** State it explicitly.
---
## Section 9: Out of Scope
### Why This Section Matters
Out of Scope prevents scope creep during implementation. When someone says "while you're in there, could you also..." — point them to this section.
### Format
```markdown
- OS-1: Multi-factor authentication — Planned for Q3 (SPEC-045).
- OS-2: Social login beyond Google/GitHub — Insufficient user demand (< 2% requests).
- OS-3: Admin impersonation — Security review pending. Separate spec required.
- OS-4: Password strength meter UI — Nice-to-have, deferred to design sprint 12.
```
### Rules
1. **Every feature discussed and rejected MUST be listed.** This creates a paper trail.
2. **Include the reason.** "Not now" is not a reason. "Insufficient demand (< 2% of requests)" is.
3. **Link to future specs** when the exclusion is a deferral, not a rejection.
---
## Feature-Type Templates
### CRUD Feature
Focus on: all 4 operations, validation rules, authorization, pagination for list endpoints.
```markdown
- FR-1: Users MUST be able to create a [resource] with [required fields].
- FR-2: Users MUST be able to read a [resource] by ID.
- FR-3: Users MUST be able to list [resources] with pagination (default: 20/page).
- FR-4: Users MUST be able to update [mutable fields] of their own [resources].
- FR-5: Users MUST be able to delete their own [resources] (soft delete).
- FR-6: Users MUST NOT be able to modify or delete other users' [resources].
```
### Integration Feature
Focus on: external API contract, retry/fallback behavior, data mapping, error propagation.
```markdown
- FR-1: The system MUST call [external API] to [purpose].
- FR-2: The system MUST retry failed calls up to 3 times with exponential backoff.
- FR-3: The system MUST map [external field] to [internal field].
- FR-4: The system MUST NOT expose external API errors directly to users.
- EC-1: External API returns 5xx → Log error, return cached data if < 1h old, else 503.
- EC-2: External API response schema changes → Log warning, reject unmappable fields.
```
### Migration Feature
Focus on: backward compatibility, rollback plan, data integrity, zero-downtime deployment.
```markdown
- FR-1: The migration MUST transform [old schema] to [new schema].
- FR-2: The migration MUST be reversible (rollback script required).
- FR-3: The migration MUST NOT cause downtime exceeding 30 seconds.
- FR-4: The migration MUST validate data integrity post-run (row count, checksum).
- EC-1: Migration fails mid-way → Automatic rollback, alert ops team.
- EC-2: New schema has stricter constraints → Log invalid rows, quarantine for manual review.
```
---
## Checklist: Is This Spec Ready for Review?
- [ ] Every section is filled (or marked N/A with reason)
- [ ] All requirements use FR-N, NFR-N numbering
- [ ] RFC 2119 keywords are UPPERCASE
- [ ] Every AC references at least one requirement
- [ ] Every AC uses Given/When/Then
- [ ] Edge cases cover each external dependency failure
- [ ] API contracts define success AND error responses
- [ ] Data models include all entities from requirements
- [ ] Out of Scope lists items discussed and rejected
- [ ] No placeholder text remains
- [ ] Context includes evidence (metrics, tickets, research)
- [ ] Status is "In Review" (not still "Draft")
---
## Full Example: Password Reset
A complete spec demonstrating all sections, followed by extracted test stubs.
### The Spec
```markdown
# Spec: Password Reset Flow
**Author:** Engineering Team
**Date:** 2026-03-25
**Status:** Approved
## Context
Users who forget their passwords currently have no self-service recovery option.
Support receives ~200 password reset requests per week, costing approximately
8 hours of support time. This feature eliminates that burden entirely.
## Functional Requirements
- FR-1: The system MUST allow users to request a password reset via email.
- FR-2: The system MUST send a reset link that expires after 1 hour.
- FR-3: The system MUST invalidate all previous reset links when a new one is requested.
- FR-4: The system MUST enforce minimum password length of 8 characters on reset.
- FR-5: The system MUST NOT reveal whether an email exists in the system.
- FR-6: The system SHOULD log all reset attempts for audit purposes.
## Acceptance Criteria
### AC-1: Request reset (FR-1, FR-5)
Given a user on the password reset page
When they enter any email address and submit
Then they see "If an account exists, a reset link has been sent"
And the response is identical whether the email exists or not
### AC-2: Valid reset link (FR-2)
Given a user who received a reset email 30 minutes ago
When they click the reset link
Then they see the password reset form
### AC-3: Expired reset link (FR-2)
Given a user who received a reset email 2 hours ago
When they click the reset link
Then they see "This link has expired. Please request a new one."
### AC-4: Previous links invalidated (FR-3)
Given a user who requested two reset emails
When they click the link from the first email
Then they see "This link is no longer valid."
## Edge Cases
- EC-1: User submits reset for non-existent email → Same success message (FR-5).
- EC-2: User clicks reset link twice → Second click shows "already used" if password was changed.
- EC-3: Email delivery fails → Log error, do not retry automatically.
- EC-4: User requests reset while already logged in → Allow it, do not force logout.
## Out of Scope
- OS-1: Security questions as alternative reset method.
- OS-2: SMS-based password reset.
- OS-3: Admin-initiated password reset (separate spec).
```
### Extracted Test Cases
Generated by `test_extractor.py --framework pytest`:
```python
class TestPasswordReset:
def test_ac1_request_reset_existing_email(self):
"""AC-1: Request reset with existing email shows generic message."""
# Given a user on the password reset page
# When they enter a registered email and submit
# Then they see "If an account exists, a reset link has been sent"
raise NotImplementedError("Implement this test")
def test_ac1_request_reset_nonexistent_email(self):
"""AC-1: Request reset with unknown email shows same generic message."""
# Given a user on the password reset page
# When they enter an unregistered email and submit
# Then they see identical response to existing email case
raise NotImplementedError("Implement this test")
def test_ac2_valid_reset_link(self):
"""AC-2: Reset link works within expiry window."""
raise NotImplementedError("Implement this test")
def test_ac3_expired_reset_link(self):
"""AC-3: Reset link rejected after 1 hour."""
raise NotImplementedError("Implement this test")
def test_ac4_previous_links_invalidated(self):
"""AC-4: Old reset links stop working when new one is requested."""
raise NotImplementedError("Implement this test")
def test_ec1_nonexistent_email_same_response(self):
"""EC-1: Non-existent email produces identical response."""
raise NotImplementedError("Implement this test")
def test_ec2_reset_link_used_twice(self):
"""EC-2: Already-used reset link shows appropriate message."""
raise NotImplementedError("Implement this test")
```
FILE:scripts/spec_generator.py
#!/usr/bin/env python3
"""
Spec Generator - Generates a feature specification template from a name and description.
Produces a complete spec document with all required sections pre-filled with
guidance prompts. Output can be markdown or structured JSON.
No external dependencies - uses only Python standard library.
"""
import argparse
import json
import sys
import textwrap
from datetime import date
from pathlib import Path
from typing import Dict, Any, Optional
SPEC_TEMPLATE = """\
# Spec: {name}
**Author:** [your name]
**Date:** {date}
**Status:** Draft
**Reviewers:** [list reviewers]
**Related specs:** [links to related specs, or "None"]
---
## Context
{context_prompt}
---
## Functional Requirements
_Use RFC 2119 keywords: MUST, MUST NOT, SHOULD, SHOULD NOT, MAY._
_Each requirement is a single, testable statement. Number sequentially._
- FR-1: The system MUST [describe required behavior].
- FR-2: The system MUST [describe another required behavior].
- FR-3: The system SHOULD [describe recommended behavior].
- FR-4: The system MAY [describe optional behavior].
- FR-5: The system MUST NOT [describe prohibited behavior].
---
## Non-Functional Requirements
### Performance
- NFR-P1: [Operation] MUST complete in < [threshold] (p95) under [conditions].
- NFR-P2: [Operation] SHOULD handle [throughput] requests per second.
### Security
- NFR-S1: All data in transit MUST be encrypted via TLS 1.2+.
- NFR-S2: The system MUST rate-limit [operation] to [limit] per [period] per [scope].
### Accessibility
- NFR-A1: [UI component] MUST meet WCAG 2.1 AA standards.
- NFR-A2: Error messages MUST be announced to screen readers.
### Scalability
- NFR-SC1: The system SHOULD handle [number] concurrent [entities].
### Reliability
- NFR-R1: The [service] MUST maintain [percentage]% uptime.
---
## Acceptance Criteria
_Write in Given/When/Then (Gherkin) format._
_Each criterion MUST reference at least one FR-* or NFR-*._
### AC-1: [Descriptive name] (FR-1)
Given [precondition]
When [action]
Then [expected result]
And [additional assertion]
### AC-2: [Descriptive name] (FR-2)
Given [precondition]
When [action]
Then [expected result]
### AC-3: [Descriptive name] (NFR-S2)
Given [precondition]
When [action]
Then [expected result]
And [additional assertion]
---
## Edge Cases
_For every external dependency (API, database, file system, user input), specify at least one failure scenario._
- EC-1: [Input/condition] -> [expected behavior].
- EC-2: [Input/condition] -> [expected behavior].
- EC-3: [External service] is unavailable -> [expected behavior].
- EC-4: [Concurrent/race condition] -> [expected behavior].
- EC-5: [Boundary value] -> [expected behavior].
---
## API Contracts
_Define request/response shapes using TypeScript-style notation._
_Cover all endpoints referenced in functional requirements._
### [METHOD] [endpoint]
Request:
```typescript
interface [Name]Request {{
field: string; // Description, constraints
optional?: number; // Default: [value]
}}
```
Success Response ([status code]):
```typescript
interface [Name]Response {{
id: string;
field: string;
createdAt: string; // ISO 8601
}}
```
Error Response ([status code]):
```typescript
interface [Name]Error {{
error: "[ERROR_CODE]";
message: string;
}}
```
---
## Data Models
_Define all entities referenced in requirements._
### [Entity Name]
| Field | Type | Constraints |
|-------|------|-------------|
| id | UUID | Primary key, auto-generated |
| [field] | [type] | [constraints] |
| createdAt | timestamp | UTC, immutable |
| updatedAt | timestamp | UTC, auto-updated |
---
## Out of Scope
_Explicit exclusions prevent scope creep. If someone asks for these during implementation, point them here._
- OS-1: [Feature/capability] — [reason for exclusion or link to future spec].
- OS-2: [Feature/capability] — [reason for exclusion].
- OS-3: [Feature/capability] — deferred to [version/sprint].
---
## Open Questions
_Track unresolved questions here. Each must be resolved before status moves to "Approved"._
- [ ] Q1: [Question] — Owner: [name], Due: [date]
- [ ] Q2: [Question] — Owner: [name], Due: [date]
"""
def generate_context_prompt(description: str) -> str:
"""Generate a context section prompt based on the provided description."""
if description:
return textwrap.dedent(f"""\
{description}
_Expand this context section to include:_
_- Why does this feature exist? What problem does it solve?_
_- What is the business motivation? (link to user research, support tickets, metrics)_
_- What is the current state? (what exists today, what pain points exist)_
_- 2-4 paragraphs maximum._""")
return textwrap.dedent("""\
_Why does this feature exist? What problem does it solve? What is the business
motivation? Include links to user research, support tickets, or metrics that
justify this work. 2-4 paragraphs maximum._""")
def generate_spec(name: str, description: str) -> str:
"""Generate a spec document from name and description."""
context_prompt = generate_context_prompt(description)
return SPEC_TEMPLATE.format(
name=name,
date=date.today().isoformat(),
context_prompt=context_prompt,
)
def generate_spec_json(name: str, description: str) -> Dict[str, Any]:
"""Generate structured JSON representation of the spec template."""
return {
"spec": {
"title": f"Spec: {name}",
"metadata": {
"author": "[your name]",
"date": date.today().isoformat(),
"status": "Draft",
"reviewers": [],
"related_specs": [],
},
"context": description or "[Describe why this feature exists]",
"functional_requirements": [
{"id": "FR-1", "keyword": "MUST", "description": "[describe required behavior]"},
{"id": "FR-2", "keyword": "MUST", "description": "[describe another required behavior]"},
{"id": "FR-3", "keyword": "SHOULD", "description": "[describe recommended behavior]"},
{"id": "FR-4", "keyword": "MAY", "description": "[describe optional behavior]"},
{"id": "FR-5", "keyword": "MUST NOT", "description": "[describe prohibited behavior]"},
],
"non_functional_requirements": {
"performance": [
{"id": "NFR-P1", "description": "[operation] MUST complete in < [threshold]"},
],
"security": [
{"id": "NFR-S1", "description": "All data in transit MUST be encrypted via TLS 1.2+"},
],
"accessibility": [
{"id": "NFR-A1", "description": "[UI component] MUST meet WCAG 2.1 AA"},
],
"scalability": [
{"id": "NFR-SC1", "description": "[system] SHOULD handle [N] concurrent [entities]"},
],
"reliability": [
{"id": "NFR-R1", "description": "[service] MUST maintain [N]% uptime"},
],
},
"acceptance_criteria": [
{
"id": "AC-1",
"name": "[descriptive name]",
"references": ["FR-1"],
"given": "[precondition]",
"when": "[action]",
"then": "[expected result]",
},
],
"edge_cases": [
{"id": "EC-1", "condition": "[input/condition]", "behavior": "[expected behavior]"},
],
"api_contracts": [
{
"method": "[METHOD]",
"endpoint": "[/api/path]",
"request_fields": [{"name": "field", "type": "string", "constraints": "[description]"}],
"success_response": {"status": 200, "fields": []},
"error_response": {"status": 400, "fields": []},
},
],
"data_models": [
{
"name": "[Entity]",
"fields": [
{"name": "id", "type": "UUID", "constraints": "Primary key, auto-generated"},
],
},
],
"out_of_scope": [
{"id": "OS-1", "description": "[feature/capability]", "reason": "[reason]"},
],
"open_questions": [],
},
"metadata": {
"generated_by": "spec_generator.py",
"feature_name": name,
"feature_description": description,
},
}
def main():
parser = argparse.ArgumentParser(
description="Generate a feature specification template from a name and description.",
epilog="Example: python spec_generator.py --name 'User Auth' --description 'OAuth 2.0 login flow'",
)
parser.add_argument(
"--name",
required=True,
help="Feature name (used as spec title)",
)
parser.add_argument(
"--description",
default="",
help="Brief feature description (used to seed the context section)",
)
parser.add_argument(
"--output",
"-o",
default=None,
help="Output file path (default: stdout)",
)
parser.add_argument(
"--format",
choices=["md", "json"],
default="md",
help="Output format: md (markdown) or json (default: md)",
)
parser.add_argument(
"--json",
action="store_true",
dest="json_flag",
help="Shorthand for --format json",
)
args = parser.parse_args()
output_format = "json" if args.json_flag else args.format
if output_format == "json":
result = generate_spec_json(args.name, args.description)
output = json.dumps(result, indent=2)
else:
output = generate_spec(args.name, args.description)
if args.output:
out_path = Path(args.output)
out_path.parent.mkdir(parents=True, exist_ok=True)
out_path.write_text(output, encoding="utf-8")
print(f"Spec template written to {out_path}", file=sys.stderr)
else:
print(output)
sys.exit(0)
if __name__ == "__main__":
main()
FILE:scripts/spec_validator.py
#!/usr/bin/env python3
"""
Spec Validator - Validates a feature specification for completeness and quality.
Checks that a spec document contains all required sections, uses RFC 2119 keywords
correctly, has acceptance criteria in Given/When/Then format, and scores overall
completeness from 0-100.
Sections checked:
- Context, Functional Requirements, Non-Functional Requirements
- Acceptance Criteria, Edge Cases, API Contracts, Data Models, Out of Scope
Exit codes: 0 = pass, 1 = warnings, 2 = critical (or --strict with score < 80)
No external dependencies - uses only Python standard library.
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Dict, List, Any, Tuple
# Section definitions: (key, display_name, required_header_patterns, weight)
SECTIONS = [
("context", "Context", [r"^##\s+Context"], 10),
("functional_requirements", "Functional Requirements", [r"^##\s+Functional\s+Requirements"], 15),
("non_functional_requirements", "Non-Functional Requirements", [r"^##\s+Non-Functional\s+Requirements"], 10),
("acceptance_criteria", "Acceptance Criteria", [r"^##\s+Acceptance\s+Criteria"], 20),
("edge_cases", "Edge Cases", [r"^##\s+Edge\s+Cases"], 10),
("api_contracts", "API Contracts", [r"^##\s+API\s+Contracts"], 10),
("data_models", "Data Models", [r"^##\s+Data\s+Models"], 10),
("out_of_scope", "Out of Scope", [r"^##\s+Out\s+of\s+Scope"], 10),
("metadata", "Metadata (Author/Date/Status)", [r"\*\*Author:\*\*", r"\*\*Date:\*\*", r"\*\*Status:\*\*"], 5),
]
RFC_KEYWORDS = ["MUST", "MUST NOT", "SHOULD", "SHOULD NOT", "MAY"]
# Patterns that indicate placeholder/unfilled content
PLACEHOLDER_PATTERNS = [
r"\[your\s+name\]",
r"\[list\s+reviewers\]",
r"\[describe\s+",
r"\[input/condition\]",
r"\[precondition\]",
r"\[action\]",
r"\[expected\s+result\]",
r"\[feature/capability\]",
r"\[operation\]",
r"\[threshold\]",
r"\[UI\s+component\]",
r"\[service\]",
r"\[percentage\]",
r"\[number\]",
r"\[METHOD\]",
r"\[endpoint\]",
r"\[Name\]",
r"\[Entity\s+Name\]",
r"\[type\]",
r"\[constraints\]",
r"\[field\]",
r"\[reason\]",
]
class SpecValidator:
"""Validates a spec document for completeness and quality."""
def __init__(self, content: str, file_path: str = ""):
self.content = content
self.file_path = file_path
self.lines = content.split("\n")
self.findings: List[Dict[str, Any]] = []
self.section_scores: Dict[str, Dict[str, Any]] = {}
def validate(self) -> Dict[str, Any]:
"""Run all validation checks and return results."""
self._check_sections_present()
self._check_functional_requirements()
self._check_acceptance_criteria()
self._check_edge_cases()
self._check_rfc_keywords()
self._check_api_contracts()
self._check_data_models()
self._check_out_of_scope()
self._check_placeholders()
self._check_traceability()
total_score = self._calculate_score()
return {
"file": self.file_path,
"score": total_score,
"grade": self._score_to_grade(total_score),
"sections": self.section_scores,
"findings": self.findings,
"summary": self._build_summary(total_score),
}
def _add_finding(self, severity: str, section: str, message: str):
"""Record a validation finding."""
self.findings.append({
"severity": severity, # "error", "warning", "info"
"section": section,
"message": message,
})
def _find_section_content(self, header_pattern: str) -> str:
"""Extract content between a section header and the next ## header."""
in_section = False
section_lines = []
for line in self.lines:
if re.match(header_pattern, line, re.IGNORECASE):
in_section = True
continue
if in_section and re.match(r"^##\s+", line):
break
if in_section:
section_lines.append(line)
return "\n".join(section_lines)
def _check_sections_present(self):
"""Check that all required sections exist."""
for key, name, patterns, weight in SECTIONS:
found = False
for pattern in patterns:
for line in self.lines:
if re.search(pattern, line, re.IGNORECASE):
found = True
break
if found:
break
if found:
self.section_scores[key] = {"name": name, "present": True, "score": weight, "max": weight}
else:
self.section_scores[key] = {"name": name, "present": False, "score": 0, "max": weight}
self._add_finding("error", key, f"Missing section: {name}")
def _check_functional_requirements(self):
"""Validate functional requirements format and content."""
content = self._find_section_content(r"^##\s+Functional\s+Requirements")
if not content.strip():
return
fr_pattern = re.compile(r"-\s+FR-(\d+):")
matches = fr_pattern.findall(content)
if not matches:
self._add_finding("error", "functional_requirements", "No numbered requirements found (expected FR-N: format)")
if "functional_requirements" in self.section_scores:
self.section_scores["functional_requirements"]["score"] = max(
0, self.section_scores["functional_requirements"]["score"] - 10
)
return
fr_count = len(matches)
if fr_count < 3:
self._add_finding("warning", "functional_requirements", f"Only {fr_count} requirements found. Most features need 3+.")
# Check for RFC keywords
has_keyword = False
for kw in RFC_KEYWORDS:
if kw in content:
has_keyword = True
break
if not has_keyword:
self._add_finding("warning", "functional_requirements", "No RFC 2119 keywords (MUST/SHOULD/MAY) found.")
def _check_acceptance_criteria(self):
"""Validate acceptance criteria use Given/When/Then format."""
content = self._find_section_content(r"^##\s+Acceptance\s+Criteria")
if not content.strip():
return
ac_pattern = re.compile(r"###\s+AC-(\d+):")
matches = ac_pattern.findall(content)
if not matches:
self._add_finding("error", "acceptance_criteria", "No numbered acceptance criteria found (expected ### AC-N: format)")
if "acceptance_criteria" in self.section_scores:
self.section_scores["acceptance_criteria"]["score"] = max(
0, self.section_scores["acceptance_criteria"]["score"] - 15
)
return
ac_count = len(matches)
# Check Given/When/Then
given_count = len(re.findall(r"(?i)\bgiven\b", content))
when_count = len(re.findall(r"(?i)\bwhen\b", content))
then_count = len(re.findall(r"(?i)\bthen\b", content))
if given_count < ac_count:
self._add_finding("warning", "acceptance_criteria",
f"Found {ac_count} criteria but only {given_count} 'Given' clauses. Each AC needs Given/When/Then.")
if when_count < ac_count:
self._add_finding("warning", "acceptance_criteria",
f"Found {ac_count} criteria but only {when_count} 'When' clauses.")
if then_count < ac_count:
self._add_finding("warning", "acceptance_criteria",
f"Found {ac_count} criteria but only {then_count} 'Then' clauses.")
# Check for FR references
fr_refs = re.findall(r"\(FR-\d+", content)
if not fr_refs:
self._add_finding("warning", "acceptance_criteria",
"No acceptance criteria reference functional requirements (expected (FR-N) in title).")
def _check_edge_cases(self):
"""Validate edge cases section."""
content = self._find_section_content(r"^##\s+Edge\s+Cases")
if not content.strip():
return
ec_pattern = re.compile(r"-\s+EC-(\d+):")
matches = ec_pattern.findall(content)
if not matches:
self._add_finding("warning", "edge_cases", "No numbered edge cases found (expected EC-N: format)")
elif len(matches) < 3:
self._add_finding("warning", "edge_cases", f"Only {len(matches)} edge cases. Consider failure modes for each external dependency.")
def _check_rfc_keywords(self):
"""Check RFC 2119 keywords are used consistently (capitalized)."""
# Look for lowercase must/should/may that might be intended as RFC keywords
context_content = self._find_section_content(r"^##\s+Functional\s+Requirements")
context_content += self._find_section_content(r"^##\s+Non-Functional\s+Requirements")
for kw in ["must", "should", "may"]:
# Find lowercase usage in requirement-like sentences
pattern = rf"(?:system|service|API|endpoint)\s+{kw}\s+"
if re.search(pattern, context_content):
self._add_finding("warning", "rfc_keywords",
f"Found lowercase '{kw}' in requirements. RFC 2119 keywords should be UPPERCASE: {kw.upper()}")
def _check_api_contracts(self):
"""Validate API contracts section."""
content = self._find_section_content(r"^##\s+API\s+Contracts")
if not content.strip():
return
# Check for at least one endpoint definition
has_endpoint = bool(re.search(r"(GET|POST|PUT|PATCH|DELETE)\s+/", content))
if not has_endpoint:
self._add_finding("warning", "api_contracts", "No HTTP method + path found (expected e.g., POST /api/endpoint)")
# Check for request/response definitions
has_interface = bool(re.search(r"interface\s+\w+", content))
if not has_interface:
self._add_finding("info", "api_contracts", "No TypeScript interfaces found. Consider defining request/response shapes.")
def _check_data_models(self):
"""Validate data models section."""
content = self._find_section_content(r"^##\s+Data\s+Models")
if not content.strip():
return
# Check for table format
has_table = bool(re.search(r"\|.*\|.*\|", content))
if not has_table:
self._add_finding("warning", "data_models", "No table-formatted data models found. Use | Field | Type | Constraints | format.")
def _check_out_of_scope(self):
"""Validate out of scope section."""
content = self._find_section_content(r"^##\s+Out\s+of\s+Scope")
if not content.strip():
return
os_pattern = re.compile(r"-\s+OS-(\d+):")
matches = os_pattern.findall(content)
if not matches:
self._add_finding("warning", "out_of_scope", "No numbered exclusions found (expected OS-N: format)")
elif len(matches) < 2:
self._add_finding("info", "out_of_scope", "Only 1 exclusion listed. Consider what was deliberately left out.")
def _check_placeholders(self):
"""Check for unfilled placeholder text."""
placeholder_count = 0
for pattern in PLACEHOLDER_PATTERNS:
matches = re.findall(pattern, self.content, re.IGNORECASE)
placeholder_count += len(matches)
if placeholder_count > 0:
self._add_finding("warning", "placeholders",
f"Found {placeholder_count} placeholder(s) that need to be filled in (e.g., [your name], [describe ...]).")
# Deduct from overall score proportionally
for key in self.section_scores:
if self.section_scores[key]["present"]:
deduction = min(3, self.section_scores[key]["score"])
self.section_scores[key]["score"] = max(0, self.section_scores[key]["score"] - deduction)
def _check_traceability(self):
"""Check that acceptance criteria reference functional requirements."""
ac_content = self._find_section_content(r"^##\s+Acceptance\s+Criteria")
fr_content = self._find_section_content(r"^##\s+Functional\s+Requirements")
if not ac_content.strip() or not fr_content.strip():
return
# Extract FR IDs
fr_ids = set(re.findall(r"FR-(\d+)", fr_content))
# Extract FR references from AC
ac_fr_refs = set(re.findall(r"FR-(\d+)", ac_content))
unreferenced = fr_ids - ac_fr_refs
if unreferenced:
unreferenced_list = ", ".join(f"FR-{i}" for i in sorted(unreferenced))
self._add_finding("warning", "traceability",
f"Functional requirements without acceptance criteria: {unreferenced_list}")
def _calculate_score(self) -> int:
"""Calculate the total completeness score."""
total = sum(s["score"] for s in self.section_scores.values())
maximum = sum(s["max"] for s in self.section_scores.values())
if maximum == 0:
return 0
# Apply finding-based deductions
error_count = sum(1 for f in self.findings if f["severity"] == "error")
warning_count = sum(1 for f in self.findings if f["severity"] == "warning")
base_score = round((total / maximum) * 100)
deduction = (error_count * 5) + (warning_count * 2)
return max(0, min(100, base_score - deduction))
@staticmethod
def _score_to_grade(score: int) -> str:
"""Convert score to letter grade."""
if score >= 90:
return "A"
if score >= 80:
return "B"
if score >= 70:
return "C"
if score >= 60:
return "D"
return "F"
def _build_summary(self, score: int) -> str:
"""Build human-readable summary."""
errors = [f for f in self.findings if f["severity"] == "error"]
warnings = [f for f in self.findings if f["severity"] == "warning"]
infos = [f for f in self.findings if f["severity"] == "info"]
lines = [
f"Spec Completeness Score: {score}/100 (Grade: {self._score_to_grade(score)})",
f"Errors: {len(errors)}, Warnings: {len(warnings)}, Info: {len(infos)}",
"",
]
if errors:
lines.append("ERRORS (must fix):")
for e in errors:
lines.append(f" [{e['section']}] {e['message']}")
lines.append("")
if warnings:
lines.append("WARNINGS (should fix):")
for w in warnings:
lines.append(f" [{w['section']}] {w['message']}")
lines.append("")
if infos:
lines.append("INFO:")
for i in infos:
lines.append(f" [{i['section']}] {i['message']}")
lines.append("")
# Section breakdown
lines.append("Section Breakdown:")
for key, data in self.section_scores.items():
status = "PRESENT" if data["present"] else "MISSING"
lines.append(f" {data['name']}: {data['score']}/{data['max']} ({status})")
return "\n".join(lines)
def format_human(result: Dict[str, Any]) -> str:
"""Format validation result for human reading."""
lines = [
"=" * 60,
"SPEC VALIDATION REPORT",
"=" * 60,
"",
]
if result["file"]:
lines.append(f"File: {result['file']}")
lines.append("")
lines.append(result["summary"])
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(
description="Validate a feature specification for completeness and quality.",
epilog="Example: python spec_validator.py --file spec.md --strict",
)
parser.add_argument(
"--file",
"-f",
required=True,
help="Path to the spec markdown file",
)
parser.add_argument(
"--strict",
action="store_true",
help="Exit with code 2 if score is below 80",
)
parser.add_argument(
"--json",
action="store_true",
dest="json_flag",
help="Output results as JSON",
)
args = parser.parse_args()
file_path = Path(args.file)
if not file_path.exists():
print(f"Error: File not found: {file_path}", file=sys.stderr)
sys.exit(2)
content = file_path.read_text(encoding="utf-8")
if not content.strip():
print(f"Error: File is empty: {file_path}", file=sys.stderr)
sys.exit(2)
validator = SpecValidator(content, str(file_path))
result = validator.validate()
if args.json_flag:
print(json.dumps(result, indent=2))
else:
print(format_human(result))
# Determine exit code
score = result["score"]
has_errors = any(f["severity"] == "error" for f in result["findings"])
has_warnings = any(f["severity"] == "warning" for f in result["findings"])
if args.strict and score < 80:
sys.exit(2)
elif has_errors:
sys.exit(2)
elif has_warnings:
sys.exit(1)
else:
sys.exit(0)
if __name__ == "__main__":
main()
FILE:scripts/test_extractor.py
#!/usr/bin/env python3
"""
Test Extractor - Extracts test case stubs from a feature specification.
Parses acceptance criteria (Given/When/Then) and edge cases from a spec
document, then generates test stubs for the specified framework.
Supported frameworks: pytest, jest, go-test
Exit codes: 0 = success, 1 = warnings (some criteria unparseable), 2 = critical error
No external dependencies - uses only Python standard library.
"""
import argparse
import json
import re
import sys
import textwrap
from pathlib import Path
from typing import Dict, List, Any, Optional, Tuple
class SpecParser:
"""Parses spec documents to extract testable criteria."""
def __init__(self, content: str):
self.content = content
self.lines = content.split("\n")
def extract_acceptance_criteria(self) -> List[Dict[str, Any]]:
"""Extract AC-N blocks with Given/When/Then clauses."""
criteria = []
ac_pattern = re.compile(r"###\s+AC-(\d+):\s*(.+?)(?:\s*\(([^)]+)\))?\s*$")
in_ac = False
current_ac: Optional[Dict[str, Any]] = None
body_lines: List[str] = []
for line in self.lines:
match = ac_pattern.match(line)
if match:
# Save previous AC
if current_ac is not None:
current_ac["body"] = "\n".join(body_lines).strip()
self._parse_gwt(current_ac)
criteria.append(current_ac)
ac_id = int(match.group(1))
name = match.group(2).strip()
refs = match.group(3).strip() if match.group(3) else ""
current_ac = {
"id": f"AC-{ac_id}",
"name": name,
"references": [r.strip() for r in refs.split(",") if r.strip()] if refs else [],
"given": "",
"when": "",
"then": [],
"body": "",
}
body_lines = []
in_ac = True
elif in_ac:
# Check if we hit another ## section
if re.match(r"^##\s+", line) and not re.match(r"^###\s+", line):
in_ac = False
if current_ac is not None:
current_ac["body"] = "\n".join(body_lines).strip()
self._parse_gwt(current_ac)
criteria.append(current_ac)
current_ac = None
else:
body_lines.append(line)
# Don't forget the last one
if current_ac is not None:
current_ac["body"] = "\n".join(body_lines).strip()
self._parse_gwt(current_ac)
criteria.append(current_ac)
return criteria
def extract_edge_cases(self) -> List[Dict[str, Any]]:
"""Extract EC-N edge case items."""
edge_cases = []
ec_pattern = re.compile(r"-\s+EC-(\d+):\s*(.+?)(?:\s*->\s*|\s*->\s*|\s*→\s*)(.+)")
in_section = False
for line in self.lines:
if re.match(r"^##\s+Edge\s+Cases", line, re.IGNORECASE):
in_section = True
continue
if in_section and re.match(r"^##\s+", line):
break
if in_section:
match = ec_pattern.match(line.strip())
if match:
edge_cases.append({
"id": f"EC-{match.group(1)}",
"condition": match.group(2).strip().rstrip("."),
"behavior": match.group(3).strip().rstrip("."),
})
return edge_cases
def extract_spec_title(self) -> str:
"""Extract the spec title from the first H1."""
for line in self.lines:
match = re.match(r"^#\s+(?:Spec:\s*)?(.+)", line)
if match:
return match.group(1).strip()
return "UnknownFeature"
@staticmethod
def _parse_gwt(ac: Dict[str, Any]):
"""Parse Given/When/Then from the AC body text."""
body = ac["body"]
lines = body.split("\n")
current_section = None
for line in lines:
stripped = line.strip()
if not stripped:
continue
lower = stripped.lower()
if lower.startswith("given "):
current_section = "given"
ac["given"] = stripped[6:].strip()
elif lower.startswith("when "):
current_section = "when"
ac["when"] = stripped[5:].strip()
elif lower.startswith("then "):
current_section = "then"
ac["then"].append(stripped[5:].strip())
elif lower.startswith("and "):
if current_section == "then":
ac["then"].append(stripped[4:].strip())
elif current_section == "given":
ac["given"] += " AND " + stripped[4:].strip()
elif current_section == "when":
ac["when"] += " AND " + stripped[4:].strip()
def _sanitize_name(name: str) -> str:
"""Convert a human-readable name to a valid function/method name."""
# Remove parenthetical references like (FR-1)
name = re.sub(r"\([^)]*\)", "", name)
# Replace non-alphanumeric with underscore
name = re.sub(r"[^a-zA-Z0-9]+", "_", name)
# Remove leading/trailing underscores
name = name.strip("_").lower()
return name or "unnamed"
def _to_pascal_case(name: str) -> str:
"""Convert to PascalCase for Go test names."""
parts = _sanitize_name(name).split("_")
return "".join(p.capitalize() for p in parts if p)
class PytestGenerator:
"""Generates pytest test stubs."""
def generate(self, title: str, criteria: List[Dict], edge_cases: List[Dict]) -> str:
class_name = "Test" + _to_pascal_case(title)
lines = [
'"""',
f"Test suite for: {title}",
f"Auto-generated from spec. {len(criteria)} acceptance criteria, {len(edge_cases)} edge cases.",
"",
"All tests are stubs — implement the test body to make them pass.",
'"""',
"",
"import pytest",
"",
"",
f"class {class_name}:",
f' """Tests for {title}."""',
"",
]
for ac in criteria:
method_name = f"test_{ac['id'].lower().replace('-', '')}_{_sanitize_name(ac['name'])}"
docstring = f'{ac["id"]}: {ac["name"]}'
ref_str = f" [{', '.join(ac['references'])}]" if ac["references"] else ""
lines.append(f" def {method_name}(self):")
lines.append(f' """{docstring}{ref_str}"""')
if ac["given"]:
lines.append(f" # Given {ac['given']}")
if ac["when"]:
lines.append(f" # When {ac['when']}")
for t in ac["then"]:
lines.append(f" # Then {t}")
lines.append(' raise NotImplementedError("Implement this test")')
lines.append("")
if edge_cases:
lines.append(" # --- Edge Cases ---")
lines.append("")
for ec in edge_cases:
method_name = f"test_{ec['id'].lower().replace('-', '')}_{_sanitize_name(ec['condition'])}"
lines.append(f" def {method_name}(self):")
lines.append(f' """{ec["id"]}: {ec["condition"]} -> {ec["behavior"]}"""')
lines.append(f" # Condition: {ec['condition']}")
lines.append(f" # Expected: {ec['behavior']}")
lines.append(' raise NotImplementedError("Implement this test")')
lines.append("")
return "\n".join(lines)
class JestGenerator:
"""Generates Jest/Vitest test stubs (TypeScript)."""
def generate(self, title: str, criteria: List[Dict], edge_cases: List[Dict]) -> str:
lines = [
f"/**",
f" * Test suite for: {title}",
f" * Auto-generated from spec. {len(criteria)} acceptance criteria, {len(edge_cases)} edge cases.",
f" *",
f" * All tests are stubs — implement the test body to make them pass.",
f" */",
"",
f'describe("{title}", () => {{',
]
for ac in criteria:
ref_str = f" [{', '.join(ac['references'])}]" if ac["references"] else ""
test_name = f"{ac['id']}: {ac['name']}{ref_str}"
lines.append(f' it("{test_name}", () => {{')
if ac["given"]:
lines.append(f" // Given {ac['given']}")
if ac["when"]:
lines.append(f" // When {ac['when']}")
for t in ac["then"]:
lines.append(f" // Then {t}")
lines.append("")
lines.append(' throw new Error("Not implemented");')
lines.append(" });")
lines.append("")
if edge_cases:
lines.append(" // --- Edge Cases ---")
lines.append("")
for ec in edge_cases:
test_name = f"{ec['id']}: {ec['condition']}"
lines.append(f' it("{test_name}", () => {{')
lines.append(f" // Condition: {ec['condition']}")
lines.append(f" // Expected: {ec['behavior']}")
lines.append("")
lines.append(' throw new Error("Not implemented");')
lines.append(" });")
lines.append("")
lines.append("});")
lines.append("")
return "\n".join(lines)
class GoTestGenerator:
"""Generates Go test stubs."""
def generate(self, title: str, criteria: List[Dict], edge_cases: List[Dict]) -> str:
package_name = _sanitize_name(title).split("_")[0] or "feature"
lines = [
f"package {package_name}_test",
"",
"import (",
'\t"testing"',
")",
"",
f"// Test suite for: {title}",
f"// Auto-generated from spec. {len(criteria)} acceptance criteria, {len(edge_cases)} edge cases.",
f"// All tests are stubs — implement the test body to make them pass.",
"",
]
for ac in criteria:
func_name = "Test" + _to_pascal_case(ac["id"] + " " + ac["name"])
ref_str = f" [{', '.join(ac['references'])}]" if ac["references"] else ""
lines.append(f"// {ac['id']}: {ac['name']}{ref_str}")
lines.append(f"func {func_name}(t *testing.T) {{")
if ac["given"]:
lines.append(f"\t// Given {ac['given']}")
if ac["when"]:
lines.append(f"\t// When {ac['when']}")
for then_clause in ac["then"]:
lines.append(f"\t// Then {then_clause}")
lines.append("")
lines.append('\tt.Fatal("Not implemented")')
lines.append("}")
lines.append("")
if edge_cases:
lines.append("// --- Edge Cases ---")
lines.append("")
for ec in edge_cases:
func_name = "Test" + _to_pascal_case(ec["id"] + " " + ec["condition"])
lines.append(f"// {ec['id']}: {ec['condition']} -> {ec['behavior']}")
lines.append(f"func {func_name}(t *testing.T) {{")
lines.append(f"\t// Condition: {ec['condition']}")
lines.append(f"\t// Expected: {ec['behavior']}")
lines.append("")
lines.append('\tt.Fatal("Not implemented")')
lines.append("}")
lines.append("")
return "\n".join(lines)
GENERATORS = {
"pytest": PytestGenerator,
"jest": JestGenerator,
"go-test": GoTestGenerator,
}
FILE_EXTENSIONS = {
"pytest": ".py",
"jest": ".test.ts",
"go-test": "_test.go",
}
def main():
parser = argparse.ArgumentParser(
description="Extract test case stubs from a feature specification.",
epilog="Example: python test_extractor.py --file spec.md --framework pytest --output tests/test_feature.py",
)
parser.add_argument(
"--file",
"-f",
required=True,
help="Path to the spec markdown file",
)
parser.add_argument(
"--framework",
choices=list(GENERATORS.keys()),
default="pytest",
help="Target test framework (default: pytest)",
)
parser.add_argument(
"--output",
"-o",
default=None,
help="Output file path (default: stdout)",
)
parser.add_argument(
"--json",
action="store_true",
dest="json_flag",
help="Output extracted criteria as JSON instead of test code",
)
args = parser.parse_args()
file_path = Path(args.file)
if not file_path.exists():
print(f"Error: File not found: {file_path}", file=sys.stderr)
sys.exit(2)
content = file_path.read_text(encoding="utf-8")
if not content.strip():
print(f"Error: File is empty: {file_path}", file=sys.stderr)
sys.exit(2)
spec_parser = SpecParser(content)
title = spec_parser.extract_spec_title()
criteria = spec_parser.extract_acceptance_criteria()
edge_cases = spec_parser.extract_edge_cases()
if not criteria and not edge_cases:
print("Error: No acceptance criteria or edge cases found in spec.", file=sys.stderr)
sys.exit(2)
warnings = []
for ac in criteria:
if not ac["given"] and not ac["when"]:
warnings.append(f"{ac['id']}: Could not parse Given/When/Then — check format.")
if args.json_flag:
result = {
"spec_title": title,
"framework": args.framework,
"acceptance_criteria": criteria,
"edge_cases": edge_cases,
"warnings": warnings,
"counts": {
"acceptance_criteria": len(criteria),
"edge_cases": len(edge_cases),
"total_test_cases": len(criteria) + len(edge_cases),
},
}
output = json.dumps(result, indent=2)
else:
generator_class = GENERATORS[args.framework]
generator = generator_class()
output = generator.generate(title, criteria, edge_cases)
if args.output:
out_path = Path(args.output)
out_path.parent.mkdir(parents=True, exist_ok=True)
out_path.write_text(output, encoding="utf-8")
total = len(criteria) + len(edge_cases)
print(f"Generated {total} test stubs -> {out_path}", file=sys.stderr)
else:
print(output)
if warnings:
for w in warnings:
print(f"Warning: {w}", file=sys.stderr)
sys.exit(1)
sys.exit(0)
if __name__ == "__main__":
main()
Phát triển năng lực lãnh đạo cho nhà sáng lập và CEO lần đầu: ủy quyền, quản lý năng lượng, điểm mù, hội chứng kẻ giả mạo, kế nhiệm.
---
name: "founder-coach"
description: "Personal leadership development for founders and first-time CEOs. Covers founder archetype identification, delegation frameworks, energy management, CEO calendar audits, leadership style evolution, blind spot identification, imposter syndrome, founder mental health, and succession planning. Use when a founder feels like the bottleneck, struggles to delegate, is burning out, transitioning from IC to executive, managing a board, or when user mentions founder mode, CEO growth, leadership development, delegation, burnout, or imposter syndrome."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: founder-development
updated: 2026-03-05
frameworks: leadership-growth, founder-toolkit
---
# Founder Development Coach
Your company can only grow as fast as you do. This skill treats founder development as a strategic priority — not a personal indulgence.
## Keywords
founder, CEO, founder mode, delegation, burnout, imposter syndrome, leadership growth, energy management, calendar audit, executive team, board management, succession planning, IC to manager, leadership style, founder trap, blind spots, personal OKRs, CEO reflection
## Core Truth
The founder is always the constraint. Not intentionally — it's structural. You built the company. You know everything. Decisions flow through you. This works until it doesn't.
At ~15 people, you hit the first ceiling: you can't be in every meeting and still think. At ~50 people, the second: your style starts creating culture problems. At ~150 people, the third: you need a real executive team or you become the reason the company can't scale.
The earlier you address this, the better.
---
## 1. Founder Archetype Identification
Most founders are primarily one archetype. Knowing yours predicts what you'll struggle with.
| Archetype | Strength | Blind spot | What they need |
|-----------|----------|------------|----------------|
| **Builder** | Product, engineering, technical depth | Go-to-market, storytelling, people | A seller / GTM partner |
| **Seller** | Revenue, relationships, vision communication | Operations, follow-through, process | An operator / COO |
| **Operator** | Execution, process, reliability | Vision, product intuition, risk | A visionary / strategic co-founder |
| **Visionary** | Strategy, narrative, pattern-recognition | Execution, details, grounding | An integrator / COO |
**Self-assessment questions:**
- What do you do when you have a free hour?
- What do you procrastinate on most?
- What do your co-founders or early team complain you don't do?
- What's the best feedback you've received about your leadership?
Most founders are Builder or Visionary. Most scaling problems happen because they don't hire their complementary type early enough.
---
## 2. Delegation Framework
Founders fail to delegate for four reasons:
1. "Nobody does it as well as I do" (often true short-term, fatal long-term)
2. "It takes longer to explain than to do it" (true once; not true the 10th time)
3. "I lose control if I don't do it myself" (control is an illusion at scale)
4. "If it fails, it's my fault" (it's your fault if you never let anyone else try)
### The Skill × Will Matrix
| | High Skill | Low Skill |
|---|-----------|----------|
| **High Will** | Delegate fully | Coach and develop |
| **Low Will** | Motivate or reassign | Manage out or redesign role |
**Rules:**
- High skill + high will → Give the work and get out of the way
- High will + low skill → Invest in them. They want to grow.
- High skill + low will → Find out why. Fix the environment or accept the mismatch.
- Low skill + low will → Don't delegate to them. Address the performance issue.
### The Delegation Ladder
Not all delegation is equal. Build up gradually:
1. "Do exactly what I tell you" — not delegation, instruction
2. "Research this and report back" — information gathering
3. "Propose a solution and I'll decide" — thinking delegation
4. "Decide and tell me what you decided" — decision delegation with review
5. "Handle it completely — update me if it's outside these parameters" — full delegation
Start at level 2–3. Move people up as trust is established. Most founders never get past level 3 with their team — that's the bottleneck.
### What to delegate first
**Delegate first (high volume, low stakes):**
- Recurring operational tasks you do the same way every time
- Information gathering and synthesis
- Meeting coordination and scheduling
- Reports and updates you produce regularly
**Delegate next (skill-buildable):**
- Customer interactions (with clear principles)
- Hiring screens (after you've trained judgment)
- Partner relationship management
- Budget management within parameters
**Delegate last (strategic, irreversible):**
- Major strategic pivots
- Executive hires
- Large financial commitments
- M&A decisions
---
## 3. Energy Management
Founders manage energy, not just time. Time is fixed. Energy is renewable — but only if you manage it.
### The Energy Audit
Map your week by energy, not tasks. See `references/founder-toolkit.md` for the full template.
**Categories:**
- 🟢 **Energizing:** Activities that leave you sharper after doing them
- 🟡 **Neutral:** Neither energizing nor draining
- 🔴 **Draining:** Activities that leave you depleted
**Common founder energy patterns:**
- **Builders:** Energized by creating, drained by politics and process
- **Sellers:** Energized by people and wins, drained by detail work and admin
- **Operators:** Energized by solving, drained by ambiguity and indecision
- **Visionaries:** Energized by strategy and ideas, drained by execution and repetition
**The rule:** Maximize green. Eliminate or delegate red. Accept yellow as the price of leadership.
### Energy management practices
**Protect deep work time.** 2–4 hours of uninterrupted thinking time, 3–5 days per week. Schedule it. Defend it. This is where strategy happens.
**Batch shallow work.** Email, Slack, administrative tasks — twice a day maximum.
**Single-task during recovery.** If you're depleted, don't try to do your best work. Do tasks that don't require your best.
**Identify your peak window.** Most people have 4–6 peak hours per day. Schedule your hardest work in those windows.
---
## 4. CEO Calendar Audit
The calendar is the most honest document in a founder's life. It shows what you actually prioritize, not what you say you prioritize.
### Running the audit
Pull the last 4 weeks of calendar data. Categorize every meeting/block:
| Category | Description | Target % |
|----------|-------------|----------|
| Strategy | Thinking, planning, direction-setting | 20–25% |
| People | 1:1s, coaching, recruiting | 20–25% |
| External | Customers, investors, partners | 20% |
| Execution | Direct work, decisions | 15% |
| Admin | Email, scheduling, overhead | < 15% |
| Recovery | Exercise, meals, thinking | 10–15% |
**Red flags in the audit:**
- Admin > 20%: You're a coordinator, not a CEO. Fix your systems.
- Execution > 30%: You're still an IC. Build the team.
- People < 10%: Your team is running on empty. They need more of you.
- No recovery blocks: You're running on adrenaline. It ends badly.
- Strategy < 10%: You're running the company, not leading it.
### The CEO's primary job at each stage
| Stage | CEO should spend most time on... |
|-------|--------------------------------|
| Seed | Product and customers. Directly. |
| Series A | Hiring the executive team. Recruiting is your job. |
| Series B | Culture, strategy, and external (investors/partners/customers) |
| Series C+ | Vision, board, external narrative, executive development |
If you're spending time on things from two stages ago, you haven't made the transition.
---
## 5. Leadership Style Evolution
The job changes at every stage. Most founders don't change with it.
**IC → Manager (0 to ~10 people):**
You need to teach and build trust. People are watching how you treat failure. The skill: give clear context, set expectations, check in frequently.
**Manager → Leader (~10 to ~50 people):**
You can't manage everyone directly. You need people who manage people. The skill: hire managers you trust, let them manage.
**Leader → Executive (~50 to ~200 people):**
You're now setting culture and direction, not managing work. The skill: communicate obsessively, decide at the right altitude, develop your leadership team.
**Executive → Institutional CEO (200+):**
You're a symbol as much as a manager. The skill: build systems that work without you; focus on board, investors, and external narrative.
**The hardest transition:** Manager → Leader. You have to stop doing things yourself and trust people you're still getting to know.
---
## 6. Blind Spot Identification
Everyone has them. Founders more than most — because nobody in the early company had the authority or safety to tell you.
### Common founder blind spots
- **Communication:** "I said it once, they should know" — you said it; they didn't hear it or didn't believe it
- **Decision speed:** Moving so fast that teams can't orient or build on your direction
- **Context hoarding:** Knowing what's happening without sharing it, then being frustrated that teams make bad decisions
- **Optimism bias:** Consistently underestimating timelines, cost, and difficulty
- **Founder exceptionalism:** Rules that apply to everyone don't apply to you
- **Feedback avoidance:** Creating an environment where no one gives you honest feedback
### How to find your blind spots
1. **360 feedback (anonymous):** Once a year. Ask direct reports, peers, board members. Include "What does [name] do that gets in the way of our success?"
2. **Exit interview analysis:** What do departing employees consistently say? Find the pattern.
3. **Failure post-mortems:** What do your worst decisions have in common? What were you assuming that wasn't true?
4. **The energy audit:** Where do you consistently drain the people around you?
---
## 7. Imposter Syndrome Toolkit
It doesn't go away. It evolves. The founder who was scared to pitch to investors is now scared to manage a board. The founder who was scared to hire is now scared to fire.
**The reframe:** Imposter syndrome is proportional to stretch. If you never feel it, you're not growing.
**Practical tools:**
- **Evidence file:** Document wins, compliments, decisions that worked. Read it when the doubt hits.
- **Normalize the feeling:** "I feel underprepared for this" ≠ "I am an imposter." Feeling and fact are different.
- **Do the thing anyway.** Competence comes from doing, not from feeling ready.
- **Name it:** Saying "I'm feeling imposter syndrome about this investor meeting" to a trusted person removes 50% of its power.
---
## 8. Founder Mental Health
Burnout isn't weakness. It's a predictable outcome of high-demand + low-recovery + no control over inputs.
### Burnout signals
Early: Irritability, difficulty sleeping, decisions feel harder than they should, loss of enthusiasm for the mission.
Mid: Physical symptoms (headaches, illness), cynicism about the company, social withdrawal, all tasks feel equally important (priority paralysis).
Late: Can't function, decisions have stopped, team notices before you do.
**If you're in late burnout:** Stop performing. Get support. The company needs a functioning founder more than it needs a martyred one.
### Structural prevention
- **Protect recovery time.** Not weekends — protected time during the week where you're not available.
- **Therapy or coaching.** Not optional for founders. The job is isolating and the stakes are high.
- **Peer group.** Other founders at similar stages. They're the only people who actually understand the job.
- **Clear off-ramps.** Know what "enough for today" looks like. Don't let the work be infinite.
---
## 9. The Founder Mode Trap
Paul Graham's "Founder Mode" essay made the case that great founders stay deeply involved in operations — skip middle management and go direct. It resonated because it's sometimes true.
**When founder mode helps:**
- Crisis recovery (company needs direct leadership)
- Product-market fit search (speed matters more than org health)
- High-value, irreversible decisions (you should be in the room)
- Early stages when the team is small
**When founder mode hurts:**
- When it undermines managers you've hired (they can't lead if you override them)
- When it's driven by distrust rather than strategy
- When it prevents the team from developing judgment
- When you're doing it because you miss doing, not because the company needs you to
**The test:** Are you going deep because the situation requires it, or because you're uncomfortable with the loss of control? The first is leadership. The second is the trap.
---
## 10. Succession Planning
Building a company that works without you is not disloyalty — it's the ultimate expression of leadership.
**Succession is not just about exit.** It's about resilience. What happens if you're sick? On sabbatical? Acquired?
**Succession readiness levels:**
- Level 1: You've documented your key knowledge and processes
- Level 2: At least one person can cover each of your key functions for 2 weeks
- Level 3: Your leadership team can run the company for a quarter without you
- Level 4: You've identified and developed your potential successor
Most founders are at Level 0. Level 2 is a reasonable target. Level 3 is a strategic asset.
---
## Key Questions for Founder Development
- "What decisions did you make last week that someone else could have made?"
- "What are you still doing that you should have delegated 6 months ago?"
- "When did you last get honest, critical feedback? From whom? What did it say?"
- "What would need to be true for the company to run for a week without you?"
- "What's draining your energy that you've accepted as unavoidable?"
## Detailed References
- `references/leadership-growth.md` — Maxwell levels, situational leadership, founder-to-CEO transition
- `references/founder-toolkit.md` — Weekly reflection, energy audit, delegation matrix, 1:1 templates
FILE:references/founder-toolkit.md
# Founder Toolkit
Practical tools for founder self-management and leadership development.
---
## 1. Weekly CEO Reflection Template
**15 minutes. Every Friday. No excuses.**
This is the most important meeting of the week. You with yourself.
```
DATE: _______________
## This Week
**1. What was my most important contribution this week?**
(Not the longest meeting or the hardest problem — the thing that will matter in 90 days.)
_______________________________________________
**2. Where did I add the least value? Why was I involved?**
(Be honest. Where were you in the room out of habit, not necessity?)
_______________________________________________
**3. What should I have delegated but didn't?**
(Name the specific task and the person you could have delegated it to.)
_______________________________________________
**4. What decision am I avoiding? Why?**
(Fear of being wrong? Not enough information? Conflict avoidance?)
_______________________________________________
**5. What would I do differently this week if I could do it over?**
(One thing. Make it specific.)
_______________________________________________
## Next Week
**My one most important outcome for next week:**
_______________________________________________
**What will I stop doing / not start / protect myself from?**
_______________________________________________
```
---
## 2. Energy Audit Template
Map your week by energy, not tasks. Do this for one full work week.
### Step 1: Time block mapping
For each 30-minute block in your week, record:
- What you did
- Energy level: 🟢 Energizing / 🟡 Neutral / 🔴 Draining
```
Monday:
08:00-08:30: __________________ [🟢/🟡/🔴]
08:30-09:00: __________________ [🟢/🟡/🔴]
09:00-09:30: __________________ [🟢/🟡/🔴]
... (continue through the day)
```
### Step 2: Pattern analysis
After one week, categorize activities:
| Activity type | Energy level | Total hours | % of week |
|--------------|-------------|-------------|-----------|
| Customer calls | | | |
| Investor meetings | | | |
| Team 1:1s | | | |
| Product decisions | | | |
| Strategy/planning | | | |
| Email/Slack | | | |
| Recruiting | | | |
| Financial review | | | |
| External talks/events | | | |
| Administrative tasks | | | |
| Deep work/building | | | |
| Recovery/breaks | | | |
### Step 3: Optimization plan
**Green activities to protect (min 40% of week):**
- _______________________________________________
**Red activities to eliminate or delegate (target: < 15% of week):**
- Activity: __________________ → Delegate to: __________________
- Activity: __________________ → Eliminate via: __________________
**Your personal energy peak hours:**
I do my best thinking: _______ to _______
Schedule this time as: Protected deep work (no meetings)
---
## 3. Delegation Matrix
For every task you regularly do, run it through this matrix.
### Assessment
| Task | Skill level needed | My will to keep it | Decision |
|------|-------------------|-------------------|----------|
| | High / Med / Low | High / Med / Low | Keep / Coach / Delegate / Kill |
### Delegation scoring
| My Skill | My Will | Decision |
|----------|---------|----------|
| High | High | Keep — this is your zone of genius |
| High | Low | Delegate — you can do it, but it drains you. Train someone. |
| Low | High | Develop — learn it or hire for it |
| Low | Low | Kill or outsource — why is this on your plate? |
### The 70% rule
If someone can do a task 70% as well as you, delegate it. Trying to get to 100% is a trap:
- Their 70% will grow to 90% with practice
- Your 30% extra effort costs more than the quality gap
- You free up time for things only you can do
---
## 4. 1:1 Template for Direct Reports
Weekly or biweekly. 30 minutes. Their agenda, not yours.
```
DATE: _______________
PERSON: _______________
## Their Section (first 20 min)
**What's on their mind? (open the meeting with this)**
(No agenda from you first — let them lead)
**What are they working on? Where are they stuck?**
**What do they need from me?**
**Anything they wanted to raise but haven't had the chance to?**
## Your Section (last 10 min)
**Context to share (strategy, changes, what they should know):**
**Direct feedback to give (if any):**
- Be specific: "In Tuesday's meeting, when you [did X], the impact was [Y]"
- Make it actionable: "Next time, I'd suggest [Z]"
**Career/growth check-in (monthly, not every meeting):**
- How are they feeling about their growth?
- What do they want to be doing more of?
- What are they interested in that they're not currently doing?
## Follow-ups
| Commitment | Owner | Due |
|------------|-------|-----|
| | | |
```
### Rules for effective 1:1s
- **Their agenda first.** If you dominate with your updates, they stop bringing theirs.
- **No status updates.** That's what tools are for. This time is for their thinking, blockers, and development.
- **Consistent time.** Rescheduled 1:1s signal that they're not a priority.
- **Take notes.** Review them before the next meeting. It signals that you listened.
- **Follow up on commitments.** If you say "I'll get you that answer by Thursday," get it by Thursday.
---
## 5. Personal OKRs for the Founder
Most founders hold their team accountable to goals but have none themselves. Fix that.
### Template: Quarterly Personal OKRs
```
Q[X] YYYY | FOUNDER OKRs
## My One Priority This Quarter
(The single most important thing I personally must accomplish)
_______________________________________________
## Objective 1: [Leadership Development]
What I'm trying to achieve: _______________________________________________
KR 1.1: [Measurable outcome by EoQ]
KR 1.2: [Measurable outcome by EoQ]
KR 1.3: [Measurable outcome by EoQ]
Progress check (mid-quarter): _______________________________________________
## Objective 2: [Delegation / Team Building]
What I'm trying to achieve: _______________________________________________
KR 2.1: [Measurable outcome by EoQ]
KR 2.2: [Measurable outcome by EoQ]
## Objective 3: [External Impact — Investors / Customers / Market]
What I'm trying to achieve: _______________________________________________
KR 3.1: [Measurable outcome by EoQ]
KR 3.2: [Measurable outcome by EoQ]
## The "Stop Doing" List (equally important)
Things I'm committing to stop doing this quarter:
- Stop: _______________________________________________
- Stop: _______________________________________________
- Stop: _______________________________________________
```
### Personal OKR examples
**Objective: Become a better coach, not just a decision-maker**
- KR: 90% of my direct reports can make their top 3 recurring decisions without me by EoQ
- KR: In 1:1 reviews, 80% of team rates me as "helps me think through problems" vs "tells me what to do"
- KR: Conduct quarterly 360 feedback session with all direct reports
**Objective: Build investor trust before I need it**
- KR: Monthly investor updates sent within 5 days of month-end, every month this quarter
- KR: 1:1 calls with each board member, once per quarter, outside of board meetings
- KR: Create and share 3-year financial model with board by EoQ
**Objective: Protect my energy and performance**
- KR: 3+ hours of protected deep work time per day, 4+ days per week
- KR: Complete weekly CEO reflection every Friday (track: 0/13 weeks → 13/13)
- KR: Zero email after 8pm, zero weekends unless explicit crisis
---
## 6. The "Stop Doing" List
The hardest list to make and the most valuable to keep.
Most founders have clear to-do lists. Few have stop-doing lists. The asymmetry is the problem.
### The stop-doing audit
**Things to stop doing immediately (decision you can make today):**
- Attending meetings you don't add value to
- Being the default person for decisions that should be made by others
- Redoing work that your team completed
- Checking email/Slack during deep work blocks
- Starting tasks you know you'll delegate partway through
**Things to stop doing by delegating (need to train someone):**
- _______________________________________________
- _______________________________________________
- _______________________________________________
**Things to stop doing by building systems:**
- Recurring manual tasks → automate
- Recurring decisions → write decision criteria so others can decide
- Recurring explanations → document once, reference always
### The decision filter
Before accepting new responsibilities, run through:
1. Does this require something only I can do?
2. Is this the highest and best use of my time?
3. If I say yes to this, what am I saying no to?
If the answers are no, no, and something important — say no.
---
## 7. Evidence File
For when imposter syndrome hits. Keep a running file of:
**Wins** (monthly minimum)
- Company milestones you led
- Decisions that worked out well
- Feedback you received that was genuinely positive
**Quotes** (capture as they happen)
- Direct quotes from team members, customers, investors about your impact
- Emails or messages that reflect trust or appreciation
**The hard calls that paid off**
- Decisions you were scared to make that turned out well
- Times you said no to something that would have hurt the company
**When to read it:** When you're doubting yourself before a board meeting, a hard conversation, a big pitch. The feeling isn't fact. The evidence file is.
FILE:references/leadership-growth.md
# Leadership Growth Reference
Frameworks for founder and executive leadership development.
---
## 1. The 5 Levels of Leadership (Maxwell)
John Maxwell's model describes leadership development as a ladder. Most founders start at Level 2–3 and need to reach Level 4–5 to scale effectively.
| Level | Name | People follow because... | What it looks like |
|-------|------|--------------------------|-------------------|
| 1 | Position | They have to (title/authority) | "Do this because I'm the CEO" |
| 2 | Permission | They want to (relationship) | People choose to work with you beyond the job requirement |
| 3 | Production | You produce results | Team rallies because you deliver; your track record gives credibility |
| 4 | People Development | You develop others | You're multiplying leaders; your success is measured by others' growth |
| 5 | Pinnacle | Who you are (reputation) | People follow because of what you've built and who you've become |
**Most founders are at Level 3.** They got here by building and shipping. The path to scaling is Level 4: developing other leaders.
**The Level 3 trap:** Production-focused founders attract doers, not leaders. They value results over growth. Their teams are effective but dependent. Every decision still goes through the founder.
**The Level 4 shift:** Measure your success by how well your team succeeds without you. Your job is to make the people around you better.
---
## 2. Situational Leadership Model
Ken Blanchard's model says effective leadership style shifts based on the person and the task — not the leader's preference.
Four styles based on the follower's development level:
| Development Level | Competence | Commitment | Leadership Style | What to do |
|------------------|------------|------------|-----------------|------------|
| D1 — Enthusiastic Beginner | Low | High | S1: Directing | High direction, low support. Tell them what to do. |
| D2 — Disillusioned Learner | Low/Med | Low | S2: Coaching | High direction + high support. Teach and encourage. |
| D3 — Capable but Cautious | Medium/High | Variable | S3: Supporting | Low direction, high support. Collaborate and encourage. |
| D4 — Self-Reliant Achiever | High | High | S4: Delegating | Low direction, low support. Get out of the way. |
**Common founder error:** Using the same leadership style with everyone. The founder who directs a D4 will frustrate them into leaving. The founder who delegates to a D1 will watch them fail.
**Diagnosis before deciding:**
Before determining your style, ask for each person + task:
- How much do they know about this specific task? (Not in general — this task.)
- How much do they want to do this specific task?
These answers may surprise you. A senior engineer may be D4 on architecture and D1 on customer calls.
---
## 3. The Founder → CEO Transition
The hardest leadership change most founders face, and nobody prepares them for it.
### What changes
**As a founder, you were judged on:**
- What you personally built
- How fast you moved
- Your own output
**As a CEO, you're judged on:**
- What your team produced
- How effectively you set direction
- The quality of the people around you
The skills that made you a great founder — doing, deciding, building — can actively work against you as a CEO.
### The transition phases
**Phase 1: Still doing (0–15 people)**
You're right to be deep in the work. Speed requires it. Your personal output matters.
Risk: Staying here too long.
**Phase 2: Building around you (15–50 people)**
You're hiring and starting to delegate. People do work you used to do.
Challenge: Learning to trust output that doesn't look like yours.
Failure mode: Hiring people and then redoing their work.
**Phase 3: Leading through leaders (50–150 people)**
You no longer know everything happening in the company. That's correct.
Challenge: Managing people who manage people — twice removed from the work.
Failure mode: Bypassing your managers to go direct (undermines them, creates chaos).
**Phase 4: Setting the container (150+ people)**
Your job is culture, strategy, and the senior leadership team. You're a CEO, not a senior contributor.
Challenge: Staying relevant and strategic without getting lost in the weeds.
Failure mode: Retreating to execution to feel productive.
### The emotional reality
Most founders describe the transition as:
- A loss of identity ("I used to know everything that was happening")
- A loss of control ("Decisions happen without me")
- A loss of clarity ("Was I more effective before?")
These are real losses, not just discomfort. Acknowledge them. Find identity in what the CEO role is, not what the founder role was.
---
## 4. Building Your Executive Team
### When to hire your first executive
Common question: "When do I need a VP/C-suite?"
**Trigger signs:**
- The function is failing and you can't fix it by working harder
- You can't attract or develop talent in that function because you lack the expertise
- The function is growing faster than you can lead it
- You're making bad decisions in that domain because you don't have deep knowledge
**Order of first executives:**
Most companies hire in this order, but the right order depends on your archetype and what's breaking:
1. First non-founder exec is usually Sales (VP Sales) or Engineering (VP Eng / CTO)
2. Then COO/Operations when coordination becomes the bottleneck
3. Then Finance (CFO) when fundraising or financial complexity demands it
4. Then People/HR when hiring velocity and culture require dedicated ownership
### How to onboard executives
**The 30-60-90 plan:**
- Day 1–30: Listen. Meet everyone. Learn the current state. No major decisions.
- Day 31–60: Diagnose. What's working, what isn't, what's missing. Share findings.
- Day 61–90: Act. Make changes. Start building systems. Establish their leadership presence.
**The trust-building sequence:**
Start with small, visible wins. Let them prove themselves in low-stakes situations before handing over high-stakes decisions.
**The founder's role during exec onboarding:**
- Provide context generously
- Introduce them with genuine authority ("This is the decision-maker for X — go to them, not me")
- Don't override their decisions publicly
- Give feedback privately, not in front of their team
**Failure mode:** Hiring a great executive and then making them feel like a senior employee. If you override every major decision, you don't have an executive — you have an expensive advisor.
---
## 5. Managing Your Board
### The fundamental tension
You work for the board. The board elected you. They can remove you. This is a governance reality, not a threat.
And: You lead the company. The board sets governance and approves major decisions, but they're not running the business day-to-day. You are.
**Healthy dynamic:** Board holds accountability; CEO holds authority. They're not adversarial — they're complementary.
### The founder mistake
Most founders either:
1. **Over-inform:** Share every detail, create noise, invite micro-management
2. **Under-inform:** Share only wins, board is surprised by problems, trust erodes
Neither works. The goal is strategic partnership.
### What the board actually needs
- **Monthly written update:** Financial performance vs plan, key metrics, top 3 issues + proposed solutions, forward-looking risks. 1–2 pages.
- **Quarterly board meeting:** Strategic discussion, not financial recap. They've read the update. Use the time for decisions and input.
- **Real-time alerts:** Big bad news before the meeting. Never let board members be surprised by negative news they should have known earlier.
### Managing board members individually
Invest in 1:1 relationships with each board member between meetings. Understand what they care about. Use their expertise.
Board members who feel informed and useful are your allies. Board members who feel blindsided or sidelined become difficult.
**The pre-meeting call:** Before every board meeting, call each member individually. Preview the agenda, surface concerns, align on decisions. The meeting itself should have no surprises.
### When the board challenges you
"The board doesn't trust my judgment" is often really: "I haven't given them enough information to trust my judgment."
Fix the transparency gap before assuming it's a political problem.
**When the board is actually wrong:** Make the case clearly, once, with data. If they override you on something important and you can't accept it, that's a signal about fit. Founders get removed. It happens. Build board relationships before you need them to trust you on a hard call.
Thiết kế và cải thiện offer: định khung giá trị, xếp bonus, bảo đảm, khan hiếm/khẩn cấp, đặt tên và cấu trúc thanh toán.
---
name: offers
description: "When the user wants to design, construct, or improve an offer — the thing they actually sell — including value framing, bonus stacking, guarantee design, scarcity/urgency, naming, and payment structure. Also use when the user mentions 'offer,' 'offer design,' 'build an offer,' 'grand slam offer,' 'irresistible offer,' 'value stack,' 'bonus stack,' 'guarantee,' 'risk reversal,' 'money-back guarantee,' 'scarcity,' 'urgency,' 'high-ticket offer,' 'productize a service,' 'naming an offer,' 'payment plan,' 'down-sell,' 'upsell offer,' or 'why isn't my offer converting.' Best for services, agencies, courses, coaching, info products, high-ticket B2B, and direct-response. If you run pure self-serve SaaS, read pricing first — tiers and packaging do more work there. For price level itself (tiers, freemium, value metric), see pricing. For the page that presents the offer, see copywriting. For the launch moment, see launch. For sales collateral, see sales-enablement."
metadata:
version: 1.0.1
---
# Offer Design
You are an expert in offer construction. Your goal is to help the user build offers that move — not by writing better copy on a worse offer, but by improving the offer itself.
## Before Starting
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md` filename, in older setups), read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
---
## Core Philosophy
**The offer is the thing, not the page.** Better copy on a weak offer compounds slowly. A stronger offer with average copy converts immediately. Most "we need better copy" requests are actually "we need a better offer" requests in disguise.
This skill exists because the rest of the repo handles the *expression* of an offer — `copywriting` writes the sales page, `cro` optimizes the conversion path, `pricing` sets the tier structure, `launch` orchestrates the moment, `paywalls` shapes the upgrade prompt. None of them ask the deeper question: **is the offer underneath any of that actually good?**
### When this skill matters
You sell:
- **Services** — consulting, freelance, agency retainers, productized services
- **Courses** — async, cohort-based, live
- **Coaching** — 1:1, group, mastermind
- **Info products** — guides, swipe files, templates, communities
- **High-ticket B2B** — $5K+ ACV with a sales conversation
- **Direct-response** — e-com promo offers, infomercial-style, paid-traffic-to-VSL
### When `pricing` does more of the work
You sell:
- **Self-serve SaaS** with tiered subscriptions — the levers are mostly tier structure, value metric, and packaging; offer construction (bonuses, guarantees) is secondary
- **Marketplaces** — the offer is structural, not constructed
Skim this skill in those cases for the value equation framing, then go to `pricing`.
---
## The Value Equation
The single most useful frame for offer design. Originally from Alex Hormozi's *$100M Offers* — internalized broadly across direct-response and creator-economy training since.
```
Dream Outcome × Perceived Likelihood of Achievement
Value = ─────────────────────────────────────────────────────────
Time Delay × Effort & Sacrifice
```
You move the four levers like this:
| Lever | What it means | How to increase value |
|-------|---------------|-----------------------|
| **Dream outcome** ↑ | What the customer actually wants | Connect to the bigger goal behind the surface ask. Specify and name it. |
| **Perceived likelihood** ↑ | Do they believe they'll get it | Proof (case studies, named customers, data), guarantees, methodology specificity |
| **Time delay** ↓ | How long until result | Faster onboarding, faster first win, faster end-to-end timeline |
| **Effort & sacrifice** ↓ | What it costs them in time/work/risk besides money | Done-for-you, simpler process, fewer decisions, lower learning curve |
**Implication for offer construction**: most "lower the price" requests are actually "raise the numerator or lower the denominator" requests. Price is the comparison, not the value.
**For the full framework, examples, and how to diagnose which lever is broken:** see [references/value-equation.md](references/value-equation.md)
---
## The Anatomy of a Complete Offer
A complete offer has six components. Skip any one and conversion suffers.
| # | Component | Question it answers |
|---|-----------|---------------------|
| 1 | **Core deliverable** | What do they get? |
| 2 | **Bonus stack** | What else do they get that makes the core feel undervalued? |
| 3 | **Guarantee** | What happens if it doesn't work? |
| 4 | **Scarcity / urgency** | Why now, not later? |
| 5 | **Name** | What is this thing called? |
| 6 | **Price + payment structure** | What do they pay and how? |
Most weak offers fail on bonuses (none), guarantees (none or wrong type), or scarcity (none, or fake). Most aggressive-to-the-point-of-cringe offers fail on guarantee (over-promising) or scarcity (fake countdown timers).
**For the full anatomy with worked examples:** see [references/offer-anatomy.md](references/offer-anatomy.md)
---
## Reference Library
| Reference | When to read |
|-----------|--------------|
| [value-equation.md](references/value-equation.md) | Diagnosing which lever is broken on a stuck offer |
| [offer-anatomy.md](references/offer-anatomy.md) | Building a complete offer from scratch |
| [guarantee-design.md](references/guarantee-design.md) | Picking the right type of guarantee for your business model |
| [bonus-stacking.md](references/bonus-stacking.md) | Adding bonuses that raise perceived value without devaluing the core |
| [scarcity-urgency.md](references/scarcity-urgency.md) | Creating *real* scarcity (and avoiding the fake patterns that destroy trust) |
| [offer-formats.md](references/offer-formats.md) | Format playbooks by business type — service, course, coaching, info product, SaaS lead magnet, agency retainer, high-ticket B2B |
| [saas-offers.md](references/saas-offers.md) | SaaS specifically — the discount trap (why discounting to acquire backfires) + four SaaS worked offers (AudienceTap, SaberSim, Teachable, Kit) |
| [examples.md](references/examples.md) | Anonymized worked examples — before/after for each business type |
---
## The Diagnostic Loop
When the user says "my offer isn't converting" or "I want to improve my offer":
1. **Identify the business type** — service, course, coaching, info product, SaaS, agency, B2B. The right playbook is type-specific.
2. **State the current offer in plain language** — name, price, what they get, guarantee, deadline. Write it down even if it lives in scattered places now.
3. **Run the value equation** — score each of the four levers 1–10. The lowest is the binding constraint.
4. **Audit the anatomy** — which of the six components is missing or weak?
5. **Pick one lever to fix this iteration** — don't rebuild everything. The biggest lever is usually the one currently scoring lowest.
6. **Draft the changed component** — new bonus, new guarantee, new scarcity, new name, new payment plan
7. **Project the lift, honestly** — most single-component changes deliver 10–40% conversion lift. Anyone promising 5x is selling something. Two consecutive iterations on different levers can stack to 2–3x.
---
## When NOT to Use Offer-Design Tactics
Some offer patterns work but cost more than they're worth:
- **Manipulative scarcity** — fake countdown timers, "only 3 spots left" lies. Short-term lift, long-term trust collapse. Don't.
- **Over-promising guarantees** — "double your revenue or refund + $1,000." Refund risk eats margin; the few cases that fail nuke your reputation publicly.
- **Bonus inflation** — stacking $50K of "bonuses" on a $497 product so it "feels like a steal." Sophisticated buyers see this. Treat bonuses as additive, not exaggerated.
- **Course-bro aesthetic on a serious product** — Gold logos, "secret method," fake urgency. Pattern-matches to scam. Wrong room.
- **Discounting to acquire** — discount-*askers* churn at ~2× the rate of full-price customers, and a coupon anchors the product as cheap. Discount only for upgrades/cross-sells (rewarding existing customers) or real seasonal windows — never to win a new one. Raise value with an offer instead. See [saas-offers.md](references/saas-offers.md).
The repo voice: opinionated, but honest. Building offers well doesn't mean building offers loud.
---
## Banned Vocabulary
When drafting offer language (sales pages, emails, headlines), avoid:
- **"Game-changing," "revolutionary," "disruptive," "next-level," "10x"** — pattern-matches to AI slop / course-bro
- **"Secret," "hidden," "what they don't want you to know"** — clickbait
- **"Limited time" with no actual time limit** — lying
- **"Worth $X" or "$Y value" with no comparable** — inflation
- **"100% guaranteed" without specifying conditions** — legally and brand-wise risky
Use specific numbers, named customers, concrete outcomes, real timelines. Specificity beats superlatives.
---
## Related Skills
- **pricing** — for price levels, tier structure, value metric, packaging, freemium
- **copywriting** — for the page that presents the offer
- **cro** — for optimizing the conversion path the offer travels through
- **launch** — for the moment you ship the offer
- **paywalls** — for in-app upgrade-prompt versions of an offer
- **sales-enablement** — for the deck and one-pager that carry the offer into a sales conversation
- **emails** — for the email sequence that warms up the offer
- **marketing-psychology** — for the cognitive biases that make offers land or bounce
FILE:evals/evals.json
{
"skill_name": "offers",
"evals": [
{
"id": 1,
"prompt": "Our B2B SaaS trial-to-paid conversion is stuck. Our plan is to run a 50% off first 3 months coupon to push more people to convert. It's a $99/month email-tool competitor and the biggest thing prospects say is they're scared to migrate their list and automations off their current provider. Is the discount the right move?",
"expected_output": "Should push back on discounting to acquire and cite that discount-askers churn at roughly 2x the rate of full-price customers, that a coupon anchors the product as cheap, and attracts price-shoppers over value-buyers. Should state the rule: discount only for upgrades/cross-sells or real seasonal windows, never to acquire a new customer. Should reframe 'offers > discounts' — raise value instead of cutting price. Should diagnose that the real objection is switching cost / migration risk, not price, so a discount does nothing about it. Should recommend a real offer that reverses that risk — e.g. a done-for-you 'Painless Switch' style migration offer with a switching-cost guarantee (migration live/verified in N days or it's free) and a bonus that removes the secondary fear (deliverability/re-warm). May reference the value equation (lower effort/sacrifice and raise perceived likelihood rather than lower price). Should tie back to real, honest scarcity (capacity/queue) rather than a fake timer.",
"assertions": [
"Advises against discounting to acquire new customers",
"Cites discount-askers churn at ~2x full-price customers",
"States discount only for upgrades/cross-sells or seasonal moments",
"Frames offers as beating discounts (raise value, not cut price)",
"Identifies migration/switching cost as the real objection, not price",
"Recommends a done-for-you migration offer with a switching-cost guarantee",
"Suggests a bonus or risk-reversal that removes the real fear",
"Uses real scarcity (capacity/queue) rather than a fake timer"
],
"files": []
}
]
}
FILE:references/bonus-stacking.md
# Bonus Stacking
How to add bonuses that raise perceived value without devaluing the core offer.
## What bonuses actually do
Three jobs at once:
1. **Raise perceived value** of the total offer
2. **Lower perceived risk** — even if the core underdelivers, "I still got X for free"
3. **Close specific buying objections** — each bonus can target one objection
The third job is the underrated one. Most weak bonus stacks throw four generic "extras" at the buyer. Strong bonus stacks read the buyer's specific hesitations and close them in order.
---
## The core principle: bonuses-as-objection-handlers
For each major objection your buyer has, add a bonus that closes it.
### Common objections → matching bonus
| Objection | Targeted bonus |
|-----------|---------------|
| "I don't have time to implement this" | Done-for-you setup, week 1 |
| "I don't know which tools to use" | Pre-vetted tool stack with discount codes |
| "What if I get stuck?" | 30-day async Slack support |
| "I'm not sure my team will buy in" | Stakeholder pitch deck |
| "I've tried something like this before and it didn't work" | Case study from someone in your exact situation |
| "What about [edge case in my industry]?" | Industry-specific bonus reference doc |
| "Will I have to learn a bunch of new tools?" | Pre-built templates for the tools we recommend |
| "What if I don't finish it?" | 1:1 accountability check-in at day 30 |
| "My situation is more complex than the average buyer" | 1:1 onboarding call to customize the plan |
| "Will this work in [region/language]?" | Localized version or addendum |
A 4-bonus stack that closes 4 specific objections converts massively better than a 4-bonus stack of generic "extras."
### How to find your buyer's actual objections
1. Read every refund-request email and sales-call transcript from the last 6 months
2. Read your own sales page out loud and write down every doubt that surfaces
3. Ask 3 recent buyers: "What almost made you not buy?"
The answers cluster around 3–6 objections. Build a bonus for each.
---
## The math of bonus value
Each bonus has a stated value (what it would cost if you bought it separately). Bonuses should:
1. **Have a stated value the buyer can verify.** Compare to a comparable product or service. "$497 value — that's what the standalone template pack costs" beats "$5,000 value." (Standalone? Compared to what?)
2. **Total to less than 2x the price** of the core offer. A $1K offer can comfortably have $1.5K in bonuses. A $1K offer with "$25K in bonuses" reads as a scam.
3. **Be things you'd actually sell separately.** If you'd never sell the bonus as a standalone product, the stated value isn't real. Sophisticated buyers can tell.
4. **Each have a specific named outcome.** "Bonus: marketing toolkit" is weak. "Bonus: 12 pre-built Notion templates for your first 90 days, valued at $297 because that's what the standalone template pack sells for at [link]" is strong.
---
## The 4-bonus pattern that works
Most strong offers stack exactly 3–5 bonuses. More starts to feel like padding; fewer leaves objections un-closed.
A common structure:
| # | Type | Purpose | Typical value |
|---|------|---------|---------------|
| 1 | **Speed bonus** | Removes time-delay objection | Templates, swipes, accelerators |
| 2 | **Trust bonus** | Removes likelihood-of-failure objection | Case study, methodology doc, examples library |
| 3 | **Stuck bonus** | Removes "what if I get stuck" objection | Office hours, Slack, on-demand support |
| 4 | **Decision bonus** | Removes "I have to choose between X and Y" objection | Tool stack with discount codes, pre-vetted recommendations |
| 5 (optional) | **Bigger-than-you-asked bonus** | Adds dream-outcome surface area | Adjacent deliverable, related framework, partner offer |
Example for a $2K B2B copywriting course:
| # | Bonus | Closes |
|---|-------|--------|
| 1 | "30 winning sales page templates (last updated 2026-Q2)" — $297 value | "I don't have time to write from scratch" |
| 2 | "9 case studies from agencies that hit $250K MRR using these frameworks" — $0 (proof, not a saleable asset) | "Does this actually work for my situation?" |
| 3 | "60-day Slack access with weekly office hours" — $497 value | "What if I get stuck on a specific project?" |
| 4 | "The tool stack: 5 tools we use + discount codes (saves ~$1,200/yr)" — $1,200 value | "I don't know what to use" |
| 5 | "Bonus session: How to charge $5K+ per project" — $297 value | Pricing confidence (adjacent dream outcome) |
Total stated value: ~$2,300 in bonuses on a $2K core. Math checks out (under 2x). Each bonus closes a real objection.
---
## When bonuses backfire
### Inflated values
"$50,000 in bonuses included today only!" on a $497 product. The asymmetry is the tell — every sophisticated buyer's bullshit detector fires.
Stated values must be defensible. If you can't point to a comparable price, don't quote the value.
### Bonuses that devalue the core
If your core offer is "I'll write your sales page for $5K" and your bonus is "PLUS — bonus sales page edits for free for life!" — the bonus implies the core is incomplete. Now the buyer wonders why they should buy *without* the bonus.
Bonuses should be *additive* to a complete core, not patches on an incomplete one.
### Bonus-stack-as-substitute-for-core
A weak core surrounded by amazing bonuses converts at the moment of sale but produces angry refund requests. The buyer bought the bonuses, got the core, felt cheated.
Order: strong core *first*, then bonuses to address specific objections.
### Stacked too high
5+ bonuses with high stated values starts to read as a course-bro funnel. Premium buyers ignore the bonus list entirely; mid-market buyers feel they're being upsold; new buyers get confused.
3–5 bonuses, each with a specific purpose. Cap it.
### Same-as-everyone-else bonuses
"BONUS! Private community access!" on every course in your category isn't a bonus, it's table stakes. If every competitor offers the same bonuses, none of them are differentiators.
Find bonuses that are specific to your buyer's situation. A SaaS bonus for a SaaS-focused buyer beats a generic "private community" every time.
---
## Bonus delivery: timing matters
A bonus delivered on day 1 closes "what if I never use it?" risk.
A bonus delivered at week 4 maintains momentum.
A bonus delivered at completion rewards finishing.
Mix the timing intentionally:
- Day 1: speed bonuses (templates, swipes, toolkit)
- Week 2–4: support bonuses (Slack, office hours, check-ins)
- Completion: identity bonuses (certificate, alumni access)
A buyer who gets all bonuses up-front is more likely to abandon (they got what they wanted, lost incentive to finish). A buyer who gets some bonuses at completion is more likely to finish (and refer).
---
## The audit
For each existing bonus on a current offer, ask:
1. **What specific buying objection does this close?** If you can't name one, it's filler.
2. **What's its defensible stated value?** If you can't point to a comparable price, drop the dollar amount.
3. **Does the buyer get it on day 1, or at a meaningful point in their journey?** Timing should support the customer outcome, not just the conversion event.
4. **Is this bonus specific to my buyer, or could any competitor offer the same thing?** If it's generic, replace it.
5. **Is the bonus strong enough that the offer would still convert without the core?** If yes, the core is weak. Fix the core, don't lean on the bonus.
Most stuck offers have either zero bonuses or too many generic ones. The right move is usually: cut to 3–5 specific objection-closing bonuses, name each one clearly, and put defensible values on them.
FILE:references/examples.md
# Worked Examples — Before/After Offers
Anonymized examples drawn from real engagements. Each shows the weak version, the diagnostic, and the strong version.
---
## Example 1: Fractional CMO service
### Before
**The offer (as it was):**
> Fractional CMO services. $15K/month. We'll help you grow.
**Diagnostic:**
- Dream outcome: 4 (vague — "grow")
- Perceived likelihood: 3 (no methodology, no case studies)
- Time delay: 4 (no timeline, indefinite engagement)
- Effort & sacrifice: 5 (unclear what the buyer has to do)
- Anatomy: only the core is present. No bonuses, no guarantee, no scarcity, no name.
**Lowest binding constraint:** perceived likelihood. Buyers don't believe an unnamed service will deliver.
### After
| Component | What was added |
|-----------|----------------|
| **Core** | "The 90-Day Marketing Reset" — 8-week audit + 12-week execution plan, delivered by a CMO who's run marketing at 3+ similar-stage companies |
| **Bonuses** | (1) Weekly 1:1s for 12 weeks (~$12K value); (2) Pre-vetted execution-partner intros (priceless); (3) Board-deck marketing strategy section template |
| **Guarantee** | "After the 8-week audit, if you don't have a clear 90-day plan you'd run yourself, you don't pay the audit fee." |
| **Scarcity** | "We take 2 engagements per quarter — next slot opens [date]" |
| **Name** | "The 90-Day Marketing Reset" |
| **Price** | $15K → $5K start, $5K week 8, $5K week 16 |
Same delivery, same person, ~3x close rate, longer engagements (because the buyer is clearer about scope).
**Lesson:** the price didn't move. The structure did.
---
## Example 2: $1,997 cohort-based copywriting course
### Before
**The offer:**
> Learn copywriting. $1,997. Includes 6 modules and Slack access.
**Diagnostic:**
- Dream outcome: 5 ("learn copywriting" — surface ask, not dream outcome)
- Perceived likelihood: 3 (no case studies, no named methodology)
- Time delay: 4 (6-month course, no first-win)
- Effort & sacrifice: 4 (lots of homework, weekly calls, big commitment)
**Lowest binding constraint:** perceived likelihood. Buyers don't believe THEY can do it.
### After
| Component | What changed |
|-----------|--------------|
| **Core** | "Write sales pages clients pay you $5K+ for in 12 weeks" — outcome-framed |
| **Bonuses** | (1) 30 winning sales page templates (last updated Q2 2026) — $297 value; (2) 9 named case studies from copywriters in 6 industries — proof, not pitch; (3) 60-day Slack with weekly office hours — $497 value; (4) The tool stack with discount codes — $1,200 value |
| **Guarantee** | "Complete all 6 modules, submit the final exercise, and if you haven't written a sales page that lands you a $5K+ client within 12 months, refund in full." |
| **Scarcity** | Cohort scarcity — doors close Friday, next cohort in 3 months |
| **Name** | "The $5K Sales Page Bootcamp" |
| **Price** | $1,997 pay-in-full OR $797 × 3 |
Same modules. Same instructor. ~4x conversion. Lower refund rate (conditional guarantee qualifies).
**Lesson:** rename the outcome, add proof, install a real scarcity mechanic.
---
## Example 3: $97 Notion template pack
### Before
**The offer:**
> Notion templates for marketers. $97. 20 templates included.
**Diagnostic:**
- Dream outcome: 6 (clear what you get, less clear what you achieve with it)
- Perceived likelihood: 6 (templates work for some, less for others — no proof)
- Time delay: 8 (instant access)
- Effort & sacrifice: 5 (setup work to customize each template)
**Lowest binding constraint:** perceived likelihood + dream outcome. "Will these actually save me time, for *my* setup?"
### After
| Component | What changed |
|-----------|--------------|
| **Core** | "The Marketing Ops Stack — 20 Notion templates that turn your scattered docs into a working marketing OS in one Saturday" — outcome-framed |
| **Bonuses** | (1) 10-minute "do this first" Loom — speed bonus; (2) "Stack the templates" flowchart (visual setup map); (3) Lifetime updates as templates are added |
| **Guarantee** | "30-day no-questions money-back" — unconditional, fits the price point |
| **Scarcity** | Founding-buyer pricing — $97 for the first 200 buyers, then $147 |
| **Name** | "The Marketing Ops Stack" |
| **Price** | $97 pay-in-full |
Same templates. ~2x close rate from the same traffic. The differentiator was the "in one Saturday" outcome anchor and the Loom that proves the speed claim.
**Lesson:** for low-priced info products, the dream outcome and a fast first-win are the levers. Don't over-engineer the guarantee.
---
## Example 4: $50K B2B SaaS annual contract
### Before
**The offer:**
> Enterprise plan: $50K/year. Includes unlimited users, all features, dedicated support.
**Diagnostic:**
- Dream outcome: 5 (features-listed, not outcome-framed)
- Perceived likelihood: 5 (no roll-out plan, no time-to-value)
- Time delay: 3 (unclear when value starts; sales says "implementation varies")
- Effort & sacrifice: 4 (procurement + security review + IT integration + change management)
**Lowest binding constraint:** time delay. Enterprise buyers can't tolerate "implementation varies."
### After
| Component | What changed |
|-----------|--------------|
| **Core** | "Production-ready in 30 days, ROI by quarter end" — time-anchored |
| **Bonuses** | (1) Dedicated implementation engineer for 30 days; (2) Pre-built integration packs for top 5 platforms; (3) Custom training session for the buyer's team; (4) Quarterly business reviews with the buyer's CSM |
| **Guarantee** | "Not in production by day 30? You don't pay until you are." SLA-based. |
| **Scarcity** | Capacity-based: "We onboard 4 enterprise accounts per quarter. Next slot starts [date]." |
| **Name** | Tier name stayed "Enterprise" but added the engagement name "Strategic Onboarding" |
| **Price** | $50K annual → $50K annual with quarterly billing + paid 30-day pilot |
Same product. ~30% higher close rate, 50% shorter sales cycle. The pilot + SLA combination removed the procurement objection.
**Lesson:** for enterprise B2B, time-to-value IS the offer. Solve it explicitly.
---
## Example 5: $4K group coaching mastermind
### Before
**The offer:**
> Group coaching for founders. $4K/quarter. Includes 12 calls and Slack.
**Diagnostic:**
- Dream outcome: 5 (vague — "be a better founder")
- Perceived likelihood: 4 (one alumni testimonial, no methodology)
- Time delay: 6 (quarterly cadence reasonable)
- Effort & sacrifice: 7 (12 calls is real time)
**Lowest binding constraint:** dream outcome + perceived likelihood.
### After
| Component | What changed |
|-----------|--------------|
| **Core** | "12 founders, 12 weeks, one specific goal each — and a room that's seen it before" — peer-room positioning |
| **Bonuses** | (1) 1:1 onboarding call to set the personal goal; (2) Founder Library — 90 frameworks from past members; (3) 1:1 mid-quarter check-in; (4) Alumni access for 1 year after |
| **Guarantee** | "First two weeks — if it's not the room you wanted, full refund. After that, you're in." |
| **Scarcity** | Cohort size capped at 12 — once full, you're on the waitlist for next quarter |
| **Name** | "The Founders' Quarter" |
| **Price** | $4K/quarter pay-in-full OR $1,500 × 3 |
Same coach, same cadence. Higher close rate. Notably: members renew at ~70% (was ~35% before) because the "alumni access for 1 year" bonus changed the buying decision frame from "quarter" to "year."
**Lesson:** for coaching, the room IS the offer. Position the room, not the curriculum. Renewal-friendly bonuses lock in long-term LTV.
---
## Example 6: Agency retainer — content marketing
### Before
**The offer:**
> Content marketing retainer. $8K/month. 4 articles per month + SEO strategy.
**Diagnostic:**
- Dream outcome: 4 (output-described, not outcome-framed)
- Perceived likelihood: 5 (no case studies linking content to revenue)
- Time delay: 3 (SEO is slow; client expectations misaligned)
- Effort & sacrifice: 6 (interviews, reviews, approvals all on client side)
**Lowest binding constraint:** dream outcome (vague) and time delay (misaligned expectations).
### After
| Component | What changed |
|-----------|--------------|
| **Core** | "We own the content engine. You get organic-driven sales meetings by month 9, with measurable revenue attribution." — outcome + timeline |
| **Bonuses** | (1) Persona research kickoff (one-time); (2) Quarterly content audit + republish list; (3) Pre-vetted freelance writers with QA layer; (4) Quarterly executive readout |
| **Guarantee** | "First 30 days is a paid pilot — 4 published pieces + 3 keyword roadmap. If at the end you don't see a clear 12-month path, we end the engagement, no balance owed." |
| **Scarcity** | Capacity-based: "We take on 3 retainer clients per quarter. Next slot is [date]." |
| **Name** | Tier name: "Growth Retainer"; engagement name: "The 90-Day Content Reset → 9-Month Growth Engine" |
| **Price** | $8K/month, 6-month minimum, OR $7K/month for 12-month commit |
Same writers. Same SEO methodology. ~2x close rate. 60% of pilots convert to 12-month commits.
**Lesson:** for slow-cycle services (SEO, brand, content), the offer has to address the timeline explicitly. "Trust us, results in 6 months" doesn't sell; "paid pilot → milestone at day 30 → ramp" does.
---
## Pattern across all six examples
Look at the changes side-by-side:
| Example | Core change | Most important other change |
|---------|-------------|----------------------------|
| Fractional CMO | Named it, added scope | First-milestone guarantee |
| Copywriting course | Outcome-framed, added proof | Case studies bonus |
| Notion templates | "in one Saturday" anchor | First-step Loom bonus |
| B2B SaaS | Time-to-value commitment | SLA-based guarantee + pilot |
| Coaching mastermind | Positioned the room, not the coach | 1-year alumni access bonus |
| Agency retainer | Outcome + timeline framing | Paid pilot guarantee |
**The pattern:** in every case, the price barely moved (or didn't move at all). What moved was the *structure* of the offer — naming, framing, guaranteeing, sequencing.
The price is the comparison. The value is the offer.
FILE:references/guarantee-design.md
# Guarantee Design
A guarantee directly raises *perceived likelihood of achievement* (the buyer thinks: "they'll only offer this if they're confident") and lowers *effort & sacrifice* (less emotional risk). It's one of the highest-leverage levers in offer design.
The wrong guarantee hurts more than no guarantee. Pick the type that matches your business model.
## The eight guarantee types
| Type | What it promises | When it works | When it backfires |
|------|------------------|---------------|-------------------|
| 1. **Unconditional money-back** | "Refund anytime within X days, no questions" | Low-priced info, high-trust audience | High-priced/high-touch; refund risk eats margin |
| 2. **Conditional money-back** | "Refund if you complete X and still don't see Y" | Courses, programs requiring effort | Sophisticated buyers; harder to honor publicly |
| 3. **Better-than-money-back** | "If it doesn't work, full refund + $X" | Confident delivery, ample margin | If you fail; the few failures explode publicly |
| 4. **Service-level / SLA** | "If we don't deliver X by Y, your money back" | Productized services, agency work | Vague SLAs you can't measure |
| 5. **Performance-based** | "Pay only when X happens" (rev share, results-based) | Sophisticated B2B, high-confidence delivery | Long cycles, hard-to-attribute outcomes |
| 6. **Anti-guarantee** | "No refunds. Make sure you want it." | Premium audiences, mature buyers | Confused / first-time buyers; reads cold |
| 7. **Outcome-or-extension** | "If you don't get X by Y, we continue free" | Coaching, services with extendable time | Open-ended cost; choose with care |
| 8. **Comparison guarantee** | "Beat [competitor]'s result or refund" | When you can credibly compare | When you can't measure the competitor cleanly |
---
## Picking the right one
Decision tree:
1. **What's your buyer's biggest perceived risk?**
- "What if it doesn't work?" → money-back family (1, 2, 3)
- "What if you don't deliver on time?" → SLA (4)
- "What if I pay and get no results?" → performance-based (5) or outcome-or-extension (7)
- "Is this real or scam?" → comparison or specificity-based (8)
2. **What's your refund tolerance?**
- Can absorb refunds at scale → unconditional (1)
- Need to qualify refunders → conditional (2)
- Confident enough to add a bonus on top → better-than-money-back (3)
- Can't afford refunds at all → anti-guarantee (6) or no guarantee + strong proof
3. **What's your buyer sophistication?**
- Premium / mature buyers → anti-guarantee can work; "we don't do refunds" reads as confidence
- First-time-in-category buyers → strong refund guarantee; they need permission to try
- Sophisticated B2B → SLA or performance-based; they expect commercial terms
4. **How measurable is the outcome?**
- Clean and measurable → performance-based, comparison, or outcome-or-extension
- Fuzzy / subjective → money-back family with a conditional gate (you completed the work)
---
## Examples by business type
### Course / cohort
**Strong:** "Complete all six modules within 60 days, submit the final exercise, and if you haven't [specific outcome] we refund in full." Conditional on effort, clear on outcome.
**Weak:** "100% money-back guarantee." No conditions = refund magnet for buyers who never engaged.
### Coaching / consulting
**Strong:** "After the first two sessions, if you don't think the engagement will deliver, we end it and refund the unused balance." Mid-engagement off-ramp builds trust.
**Weak:** "Satisfaction guaranteed." Means nothing.
### Productized service / agency retainer
**Strong:** "First month is a paid pilot. At the end, if you don't see [specific milestone], you don't pay for month 2 and we end on good terms." Clear gate, clear out.
**Weak:** "We'll work until you're happy." Open-ended cost. Don't.
### High-ticket info product (community, mastermind)
**Strong (premium audience):** "No refunds. The application process is rigorous because the value is real. If you're not sure, don't apply yet." Anti-guarantee works here.
**Weak (premium audience):** Generic 30-day refund. Reads cheap.
### Low-ticket info product (template, swipe file)
**Strong:** "30-day no-questions refund." The transactional bar is "I bought it, looked at it, didn't want it." Unconditional fits.
**Weak:** No guarantee. The buyer's risk is too high for the price.
### SaaS
**Strong:** Free trial *or* annual-prepay-with-money-back-in-first-30-days. Reduces friction without locking in unhappy users.
**Weak:** "Cancel anytime" alone — not a guarantee, just standard SaaS terms.
### Direct response / paid traffic
**Strong:** Double-your-money-back or comparable risk inversion. Direct-response buyers expect risk-reversal-heavy offers.
**Weak:** Vanilla 30-day refund. Doesn't differentiate from every other ad on the platform.
---
## Writing the guarantee
The guarantee text matters. Patterns that work:
**Specific terms:**
> If, after completing the first 4 weeks of the program, you can't point to one specific business outcome you've achieved, email us and we'll refund 100%.
**Confident tone:**
> We know this works. If it doesn't for you, we don't want your money.
**Acknowledge the awkwardness:**
> Guarantees feel slimy. Here's ours anyway: if you do the work in modules 1–3 and don't see meaningful traction, we refund.
**Patterns that don't work:**
- "100% satisfaction guaranteed!" — generic, low-trust
- "Lifetime guarantee" — meaningless without conditions
- Multiple stacked guarantees — sophistication-collapsing
- Guarantees full of legalese — buyers skim and assume the worst
---
## Common mistakes
### Promising more than you can deliver
"Double your revenue or your money back + $1,000." If even 1 in 50 buyers fails and gets the bonus refund + writes a public review, the offer is permanently damaged.
Stress-test: what happens if 10% of buyers invoke the guarantee?
### Conditional guarantees with too many conditions
"Refund if you watched all 24 modules, completed the 6 exercises, attended every live call, and posted in the community at least once per week."
Buyers read this as "they made it impossible to actually get a refund." Trust drops.
Two conditions max. Three only if they're closely related (e.g., "completed the course AND submitted the final project AND emailed us a question").
### Hiding the guarantee in the fine print
If your guarantee is your strongest perceived-likelihood lever, *put it on the sales page in 24pt text*. Move it above the buy button.
### Forgetting to test it
Re-read your guarantee text every six months. The wording that worked a year ago may now be undermined by something you've changed about your offer.
### Treating guarantees as a substitute for proof
Strong proof + weak guarantee > strong guarantee + weak proof. Order matters. Build proof first, then layer on the guarantee.
---
## The honest case for NO guarantee
Anti-guarantees ("no refunds, this is final") work when:
- Buyer sophistication is high
- Application or qualification process precedes the sale
- Price is premium-to-luxury
- Brand is established
- Proof is overwhelming
What you're saying: "We don't need a guarantee because the work is real, the buyer has self-qualified, and we won't engage in transactional refund games."
The wrong audience reads this as cold or scammy. The right audience reads it as confidence. Know your buyer.
---
## The diagnostic
When auditing an offer with no guarantee (or a weak one), ask:
1. **What's the buyer's actual risk?** Make it concrete. ("$2K and I might not get more clients.")
2. **What guarantee structure reverses that specific risk?** Match it to one of the eight types.
3. **What's your honest refund tolerance?** Calculate refund rate × refund cost; can you sustain it?
4. **Does the guarantee match your audience sophistication?** Premium buyers want anti-guarantee; first-time buyers want unconditional.
Most offers don't have the wrong guarantee — they have *no* guarantee at all. Adding any guarantee is almost always a lift. Adding the right one is the lever.
FILE:references/offer-anatomy.md
# Offer Anatomy
A complete offer has six components. Skip any one and conversion suffers — usually noticeably.
## The six components
| # | Component | Question it answers | Where it fails |
|---|-----------|---------------------|----------------|
| 1 | **Core deliverable** | What do they get? | Too vague, or pitched as features instead of outcome |
| 2 | **Bonus stack** | What else do they get that makes the core feel undervalued? | Either no bonuses, or inflated/fake bonuses |
| 3 | **Guarantee** | What happens if it doesn't work? | None, wrong type, or over-promising |
| 4 | **Scarcity / urgency** | Why now, not later? | None, fake, or destructively manipulative |
| 5 | **Name** | What is this thing called? | Generic, internal-jargon, or no name at all |
| 6 | **Price + payment structure** | What do they pay and how? | Single number with no payment flexibility |
---
## 1. Core deliverable
The thing they actually get.
### Define it as an outcome, not a feature list
- **Feature-pitched (weak):** "6 modules, 24 lessons, weekly calls, private community."
- **Outcome-pitched (strong):** "A working customer-acquisition system that brings 5 qualified leads per week within 60 days — built with you, not handed to you."
The features still matter — buyers want to know what they're getting — but the *frame* is the outcome. Features support the outcome, they don't replace it.
### Define the scope explicitly
What's in. What's out. What's optional. Buyers buy clarity; ambiguity erodes perceived likelihood.
Example scope statement:
```
Includes:
- 90-day program with weekly live calls (recorded)
- Private Slack with daily founder access
- 12 fill-in-the-blank templates
- 1 90-minute strategy session with a senior strategist
Doesn't include:
- 1:1 calls outside the strategy session
- Implementation of the work (you/your team does this; we coach)
- Tools and software (you provide; we recommend specific stacks)
```
### Match the depth to the buyer's stage of awareness
Sophisticated buyers want the methodology and scope. New-to-category buyers want the dream outcome and proof. Read your audience.
---
## 2. Bonus stack
What you add to make the core feel undervalued at the asking price.
Bonuses do three jobs at once:
1. **Raise perceived value** of the total offer
2. **Lower perceived risk** — even if the core underdelivers, "I got X for free"
3. **Close specific objections** — each bonus can target a different buying objection
### How to construct bonuses
For each major objection your buyer has, add a bonus that closes it:
| Objection | Targeted bonus |
|-----------|---------------|
| "I don't have time to implement this" | Done-for-you setup, day 1 |
| "I don't know which tools to use" | Pre-vetted tool stack with discount codes |
| "What if I get stuck?" | 30-day async support |
| "I'm not sure my team will buy in" | Stakeholder pitch deck for your team |
| "I've tried something like this before and it didn't work" | Case study of someone in your exact situation |
A 4-bonus stack that closes 4 specific objections converts massively better than a 4-bonus stack of generic "extras."
### Don't inflate
"$50,000 in bonuses!" on a $500 offer reads as scam. The asymmetry destroys trust.
Bonuses should:
- Have a stated value the buyer can verify (compare to a comparable product)
- Total to less than 2x the price (e.g., a $1K offer can have ~$1.5K in bonuses comfortably)
- Be things you'd actually sell separately if you wanted
For the full bonus-stacking framework, see [bonus-stacking.md](bonus-stacking.md).
---
## 3. Guarantee
What happens if it doesn't work.
A guarantee directly raises perceived likelihood of achievement (the buyer thinks: "they'll only offer this if they're sure"). It also lowers effort & sacrifice (less emotional risk).
The wrong guarantee can hurt:
- Over-promising guarantees attract refund-seekers
- Generic "100% guaranteed" with no conditions reads as legally unenforceable
- No guarantee at all signals you're not confident
The right type depends on your business model, refund risk tolerance, and buyer sophistication. For the full taxonomy, see [guarantee-design.md](guarantee-design.md).
---
## 4. Scarcity / urgency
The reason to buy now, not later.
Two flavors:
- **Scarcity** — limited *quantity* (cohort size, seats, inventory, batch)
- **Urgency** — limited *time* (cohort deadline, season, bonus expiry)
The bar: **the scarcity has to be real.** Fake countdown timers and "only 3 spots left" lies work once and torch trust permanently. The internet is small; you will be caught.
Common honest scarcity formats:
- Cohort closes Friday (because the cohort actually starts Monday)
- Founding-member pricing for the first 20 customers (because you're capacity-constrained)
- Seasonal product or service (because demand is seasonal)
- Bonus expires at launch end (because the bonus is your time)
- Capacity-based service tier (because you literally can't take more clients)
For full guidance on creating real scarcity, see [scarcity-urgency.md](scarcity-urgency.md).
---
## 5. Name
What this thing is called.
A named offer beats an unnamed offer for three reasons:
1. **Repeatability** — buyers can tell their friend about it
2. **Distinction** — a name makes it a *thing*, not a generic service
3. **Pricing power** — branded offers can charge more than the same delivery sold as a service
### Naming patterns that work
- **Outcome-named:** "The 30-Day Activation Sprint" — names what they get
- **Methodology-named:** "The VAULT Framework" — names how you do it
- **Identity-named:** "Founder Marketing OS" — names who it's for
- **Compression-named:** "5-Day Cohort" — names the timing/structure
### Naming patterns that don't work
- **Generic descriptors:** "Marketing Coaching Program" — forgettable
- **Internal jargon:** "Tier 2 Standard" — buyer can't repeat
- **Course-bro:** "The Money-Making Machine" — pattern-matches to scam
- **Pun-overload:** "GrowthGoGetter" — reads as low-status
### Practical test
Can a buyer text a friend: "I just signed up for *the [name]*. It's $X and you get [one-line outcome]"? If yes, the name works. If no, rename.
---
## 6. Price + payment structure
The price is the obvious part. The structure is the underrated part.
### Price isn't a number, it's a comparison
Buyers compare the price to:
- The dream outcome (does this get me the result I want?)
- The next-best alternative (what else could I buy?)
- The cost of doing nothing (what does the status quo cost me?)
- Other items in your own catalog (anchor pricing)
You can move price perception without changing the number by:
- Showing the cost of doing nothing more vividly
- Anchoring against a higher-priced alternative
- Sequencing other items in your catalog at higher prices first
### Payment structure is its own lever
Same total price, different structures convert very differently:
| Structure | When it works | Trade-off |
|-----------|---------------|-----------|
| **Pay in full** | High-trust buyers, lower price points | Highest perceived commitment, smallest buyer pool |
| **Pay in 2-4 installments** | Mid-range price, hesitant buyers | More buyers, payment defaults |
| **Monthly subscription** | SaaS, ongoing services | Annuity revenue, churn risk |
| **Pay-after-results** | High-confidence delivery, sophisticated buyers | Cash flow lag, fewer disputes |
| **Down payment + balance on delivery** | Services with milestone-based delivery | Balance risk on backend |
| **Free trial → paid** | Low-friction SaaS, info products | Conversion drop-off |
Often the right move isn't lowering price — it's adding a payment plan. Same $6K price, "$6K today" vs "$2K × 3 monthly" converts very differently.
---
## Putting it together: an example
A B2B fractional CMO service.
| Component | Weak version | Strong version |
|-----------|--------------|----------------|
| **Core** | "Fractional CMO services" | "8-week marketing audit + 90-day execution plan, delivered by a CMO who's done it for 3+ similar companies" |
| **Bonuses** | None | (1) 1:1 weekly check-ins for 90 days; (2) pre-vetted execution-partner introductions; (3) board-deck for marketing strategy section |
| **Guarantee** | None | "If after the 8-week audit you don't have a clear 90-day plan you'd run yourself, you don't pay the audit fee" |
| **Scarcity** | None | "We take 2 engagements per quarter — next slot opens [date]" |
| **Name** | "fCMO Consulting" | "The 90-Day Marketing Reset" |
| **Price** | "$15K, paid up front" | "$15K → $5K to start, $5K at week 8, $5K at week 16" |
Same delivery. Same person. Different offer. Different conversion.
The point: most "we need to lower our price" conversations are actually "we have one of six components missing or weak" conversations.
FILE:references/offer-formats.md
# Offer Formats by Business Type
The right offer format depends on what you sell. The same six components (core, bonuses, guarantee, scarcity, name, price) get assembled differently by business type.
This reference is organized by business type. Find yours, then use the format as a starting point — not a fixed recipe.
---
## Service / freelance
You sell your time and skill.
### Default format
| Component | Default |
|-----------|---------|
| **Core** | A scoped engagement with a specific deliverable and timeline |
| **Bonuses** | Templates, frameworks, post-engagement support, tool stack |
| **Guarantee** | First-milestone gate (paid pilot or first-deliverable refund) |
| **Scarcity** | Capacity-based (next slot opens [date]) |
| **Name** | Methodology-named or outcome-named (e.g., "The 30-Day Activation Sprint") |
| **Price** | Project-based with down-payment, or monthly retainer |
### What to watch
- **Naming matters disproportionately** — services without named offers compete on price; named services compete on positioning
- **Scope creep is the offer killer** — define what's in, out, and optional, *in writing*, before the engagement starts
- **Bonuses should compound the deliverable** — templates and frameworks that make the buyer self-sufficient after the engagement, not during
### Productizing the offer
Move from "I sell consulting" to "I sell the 8-Week Marketing Reset." Same delivery, different offer.
The productized version:
- Has a name
- Has a fixed scope and timeline
- Has a fixed price
- Has the same bonuses every time
- Has a defined gate (week 4 check-in, paid pilot, milestone review)
Productizing raises perceived value, simplifies sales, and creates a repeatable case-study factory.
---
## Course (async, cohort-based, live)
You sell structured learning.
### Default format
| Component | Default |
|-----------|---------|
| **Core** | The curriculum + delivery format |
| **Bonuses** | Templates, swipe files, case studies, community access |
| **Guarantee** | Conditional money-back (completion-gated) |
| **Scarcity** | Cohort scarcity (enrollment closes [date]) |
| **Name** | Outcome-named or methodology-named |
| **Price** | Pay-in-full or 2–4 installments |
### Async vs cohort-based
| Decision | Async | Cohort |
|----------|-------|--------|
| **Pricing** | Lower ($297–$1,997) | Higher ($1,497–$5,000+) |
| **Scarcity** | Bonus expiry, price increases | Cohort start date |
| **Guarantee** | Generous unconditional | Conditional on completion |
| **Conversion mechanic** | Email funnel, evergreen webinar | Launch window + cohort deadline |
| **Default bonus** | Templates, swipes | Slack + office hours + 1:1 review |
### What to watch
- **Completion is the marketing asset** — every completer is a case study. Engineer the first win in week 1.
- **Cohort scarcity must be real** — if "doors close Friday" is followed by "doors reopen Monday because we extended the cohort," the trick gets noticed
- **Refund design matters more than refund rate** — a generous-sounding guarantee with smart conditions converts well and refunds rarely
---
## Coaching (1:1, group, mastermind)
You sell access to your expertise applied to their specific situation.
### Default format
| Component | Default |
|-----------|---------|
| **Core** | Sessions + asynchronous access + specific outcome focus |
| **Bonuses** | Resources from your library, intro to network, post-engagement check-ins |
| **Guarantee** | Outcome-or-extension or first-two-sessions out |
| **Scarcity** | Capacity-based (N spots / quarter) |
| **Name** | Identity-named or outcome-named (e.g., "Founder Marketing Mastermind") |
| **Price** | Monthly retainer or 3/6/12-month engagement |
### 1:1 vs group
| Decision | 1:1 | Group / mastermind |
|----------|-----|---------------------|
| **Pricing** | $1,500–$10,000+/mo | $497–$2,500+/mo |
| **Scarcity** | Capacity (4–10 1:1 clients) | Cohort size (8–30 members) |
| **Guarantee** | First-two-sessions out | Trial period, no refunds after |
| **Bonuses** | Custom resources, intros | Group access, peer accountability, library access |
| **Default delivery** | Weekly or biweekly sessions | Monthly group calls + community |
### What to watch
- **Identity is the offer** — group coaching is often more about being in the room with peers than about the coach's instruction. Name the room, not the coach.
- **Onboarding is part of the offer** — a sloppy intake destroys perceived likelihood
- **Renewal is the real conversion event** — design for the 6-month decision, not the first-month decision
---
## Info product (guide, swipe file, template pack, community)
You sell packaged knowledge or assets.
### Default format
| Component | Default |
|-----------|---------|
| **Core** | The asset(s) + lifetime access |
| **Bonuses** | Adjacent assets, walkthroughs, templates |
| **Guarantee** | Generous unconditional (30-day no-questions) |
| **Scarcity** | Bonus expiry, price-increase scheduling |
| **Name** | Outcome-named, often punchy and specific |
| **Price** | $29–$497, pay-in-full |
### What to watch
- **The first-impression matters disproportionately** — the buyer opens it once. If the first 5 minutes don't feel premium, they don't engage with the rest
- **Quick-start is a bonus** — pair the asset with a 10-minute "do this first" walkthrough
- **Lifetime access is implicit pricing** — clarify what "lifetime" means (yours, the product's, until you sunset it)
---
## High-ticket B2B ($5K+ ACV, sales-led)
You sell to companies with a sales conversation.
### Default format
| Component | Default |
|-----------|---------|
| **Core** | A multi-month engagement or annual contract |
| **Bonuses** | Onboarding, training, integration, dedicated CSM |
| **Guarantee** | SLA, performance-based, or pilot-gated |
| **Scarcity** | Quarter-end pricing, capacity (N onboardings/quarter), tier limits |
| **Name** | Internal-stable (e.g., "Enterprise Plan") + named engagement type (e.g., "Strategic Onboarding") |
| **Price** | Annual contract with quarterly payment, often custom |
### What to watch
- **Buying committee, not buyer** — the offer has to land with the champion, the economic buyer, and the influencer simultaneously
- **Procurement is the offer** — your terms (payment, NET 60, security review, MSA) are part of the offer; rigid terms lose deals
- **Pilot offers convert sophisticated buyers** — "30-day paid pilot, decide to continue at end" reduces decision risk
- **The CSM is part of the offer** — buyers consistently rate post-sale relationship as part of the offer-perceived-value
---
## Agency retainer
You sell ongoing service delivery.
### Default format
| Component | Default |
|-----------|---------|
| **Core** | Monthly deliverables + dedicated team + reporting cadence |
| **Bonuses** | Strategy sessions, tool access, audit credits, library access |
| **Guarantee** | Month-1 paid pilot or 90-day out clause |
| **Scarcity** | Capacity (N clients / vertical / quarter) |
| **Name** | Tier-named ("Growth," "Scale," "Enterprise") + service line |
| **Price** | Monthly retainer with discount for annual commit |
### What to watch
- **Onboarding velocity is the offer** — agencies that take 6 weeks to start delivering lose to agencies that deliver something in week 1
- **Reporting is part of the offer** — clean, monthly, action-oriented reporting reduces churn more than additional deliverables
- **Tier upgrades are the easiest revenue** — design tiers with clear value-step-ups so upgrade conversations are obvious
---
## Self-serve SaaS
You sell a tool with tiered subscriptions.
### Default format
| Component | Default |
|-----------|---------|
| **Core** | Tiered subscription with clear feature differentiation |
| **Bonuses** | Free onboarding, templates, integrations, partner discounts |
| **Guarantee** | Free trial OR annual-with-30-day-refund |
| **Scarcity** | Founding-pricing for first N customers, or seasonal launches |
| **Name** | Tier-named ("Starter," "Pro," "Team," "Enterprise") |
| **Price** | Monthly or annual with discount, value-metric-based |
### What to watch
- **Pricing tier > offer construction** — for self-serve SaaS, packaging and value metric do more work than guarantees and bonuses. Use the `pricing` skill.
- **Free trial design IS offer design** — length, gated features, credit-card-required vs not, automatic conversion. Each is an offer decision.
- **Annual prepay is the offer lever** — same product, different commitment, often 20–40% discount. Many SaaS conversion lifts come from improving the annual offer, not the monthly.
For SaaS, this skill is supplemental. Read [`pricing`](../../pricing/SKILL.md) first.
---
## Direct response / paid traffic
You sell from a sales page or VSL to cold traffic.
### Default format
| Component | Default |
|-----------|---------|
| **Core** | The "thing" + clear payoff |
| **Bonuses** | Heavy bonus stack (5–7 bonuses, layered values) |
| **Guarantee** | Aggressive risk reversal (better-than-money-back, double guarantee) |
| **Scarcity** | Real time-bound (launch window, evergreen with hard close) |
| **Name** | Hooky, often pattern-interrupt |
| **Price** | Often single-payment with payment plan offered |
### What to watch
- **Direct-response buyers expect aggression** — quiet, premium-feeling offers convert badly on cold paid traffic. The aesthetic of the page matters as much as the offer.
- **Refund rates can be 10–20%** — bake this into the math. If margin can't survive 15% refunds, restructure.
- **The first 7 seconds determine the rest** — hook, then offer
This format is high-skill. If you're not from a direct-response background, hire someone or partner with someone who is.
---
## Choosing your format
If you're not sure which format applies, pick the closest match and adapt. The biggest mistake is borrowing a format from a different business type (e.g., applying direct-response bonus stacking to a premium B2B service — wrong audience, wrong aesthetic).
Two diagnostic questions:
1. **Who buys it, and how sophisticated are they?** Premium B2B and direct-response cold traffic both buy, but they need different offers.
2. **What's the dominant constraint?** Service businesses are capacity-constrained, SaaS is pricing-tier-constrained, courses are cohort/season constrained. Match the scarcity format to the real constraint.
For worked examples by business type, see [examples.md](examples.md).
FILE:references/saas-offers.md
# SaaS Offers — the discount trap + worked examples
The rest of the offers library skews services, courses, and coaching. This reference covers the SaaS case specifically: why discounting is the wrong acquisition lever, and four worked offers that stack risk-reversal, bonuses, and scarcity for a software business.
Read this alongside [offer-formats.md](offer-formats.md#self-serve-saas). For price level and tier structure, the `pricing` skill still does the heavier lifting — this covers the *offer* wrapped around the price.
---
## The discount trap
The instinct when a SaaS isn't converting is to cut the price — a launch coupon, a "50% off first 3 months," a permanent lower tier. It's the most-reached-for lever and one of the worst.
**Offers beat discounts.** A discount lowers the price. An offer raises the value — a guarantee, a done-for-you migration, a bonus that removes the switching cost. Same net price to the buyer, but one trains them to expect cheap and the other trains them to expect valuable.
### The data
**Discount-*askers* churn at roughly 2× the rate of full-price customers.** The buyer who negotiated their way in is signaling something: price was the reason they bought, not value. When a cheaper option appears — or when the renewal hits full price — they leave. You bought a customer who was never yours.
This compounds. Discounting to acquire also:
- **Anchors the product as cheap** — hard to raise later without churn spikes
- **Attracts the wrong ICP** — price-shoppers, not value-buyers
- **Trains the market to wait** — buyers learn there's always a sale coming, so they never pay full
- Sits next to **trial fatigue and discount fatigue** — the same buyer who's seen a hundred "50% off" banners no longer feels urgency from yours
### The rule
**Never discount to acquire.** Discount only in two moments, where the mechanics actually work for you:
| When | Why it works |
|------|--------------|
| **Upgrades / cross-sells** | The customer already values the product. A discount to move up a tier or add a product rewards commitment instead of buying a stranger. |
| **Seasonal moments** | Black Friday, year-end, an annual-plan push — a real, time-bound, everyone-gets-it window. Not a permanent price cut wearing a costume. |
Everything else is an **offer**, not a discount. If conversion is stuck at the top of the funnel, reverse the risk and stack the value — don't cut the price. See the four examples below.
---
## Four SaaS worked offers
Each shows how to wrap a SaaS price in a real offer — risk-reversal, a switching-cost-killing bonus, and honest scarcity — instead of a coupon.
### 1. AudienceTap — $297 fixed-price pilot
**Business:** audience-analytics SaaS, self-serve, mid-market buyers hesitant to sign an annual before seeing value on their own data.
**The offer instead of a discount:**
| Component | What it is |
|-----------|------------|
| **Core** | A **$297 fixed-price 30-day pilot** — full product, run against the buyer's real audience, not a demo dataset |
| **Risk reversal** | "If the pilot doesn't surface an insight your team acts on, the $297 is refunded — and it credits toward annual if you convert." |
| **Bonus** | A done-with-you setup session (connect sources, first report) so time-to-value is days, not weeks |
| **Scarcity** | Capacity-real: "We onboard 6 pilots a month so each gets the setup session." Not a fake counter. |
**Why it beats a discount:** the pilot is a paid, low-risk *yes* that filters for value-buyers. The $297 isn't a price cut — it's a fee that credits forward, so full-price annual is the default next step, not a negotiation.
### 2. SaberSim — $497 bundle
**Business:** sports-analytics/optimizer SaaS. Individual tools convert fine alone but the full stack is where retention lives.
**The offer instead of a discount:**
| Component | What it is |
|-----------|------------|
| **Core** | A **$497 bundle** of the optimizer + sync + data feed — the tools that only pay off together |
| **Risk reversal** | 14-day "run it on a real slate" guarantee — use it live, full refund if it doesn't beat the buyer's current workflow |
| **Bonus** | Strategy walkthroughs + a starter template library so the bundle produces a result on day one |
| **Scarcity** | Seasonal: bundle priced for the **start of the season**; after kickoff it unbundles to full à-la-carte |
**Why it beats a discount:** the bundle raises perceived value (three tools, one decision) rather than lowering price on one. The season start is *real* seasonal scarcity — an allowed discount moment — not a permanent markdown.
### 3. Teachable — $1,997 accelerator
**Business:** course-platform SaaS. The subscription is cheap; the value (and the switching cost) is getting a course actually launched.
**The offer instead of a discount:**
| Component | What it is |
|-----------|------------|
| **Core** | A **$1,997 "Launch Accelerator"** — the platform plan **plus** a structured 6-week program to ship the buyer's first course |
| **Risk reversal** | Outcome guarantee: "Publish your course in 6 weeks or we work with you free until you do." Tied to a completion condition. |
| **Bonus** | Launch-email templates, a pricing-page teardown, and a cohort Slack — the pieces creators stall on |
| **Scarcity** | Cohort-based: accelerator runs on a start date, capped seat count, next cohort later |
**Why it beats a discount:** the accelerator sells the *outcome* (a launched course) at a price far above the raw subscription — the opposite of discounting. The guarantee de-risks the real fear ("I'll pay and never launch"), and the cohort cap is honest scarcity.
### 4. Kit — "Painless Switch" $997 migration offer
**Business:** email-platform SaaS. The blocker isn't price — it's the terror of migrating a list, sequences, and automations off the incumbent.
**The offer instead of a discount:**
| Component | What it is |
|-----------|------------|
| **Core** | A **$997 "Painless Switch"** — done-for-you migration of list, forms, sequences, and automations |
| **Risk reversal** | "If your migration isn't live and verified in 14 days, it's free." The guarantee is on the *switching cost*, the actual objection |
| **Bonus** | A deliverability audit + a re-warm plan so the switch doesn't tank open rates — removes the second-biggest fear |
| **Scarcity** | Capacity-real: migration engineers handle a fixed number per month; the offer closes when the queue fills |
**Why it beats a discount:** the entire barrier to a SaaS switch is effort and risk, not price. A discount does nothing about migration dread; a done-for-you offer with a switching-cost guarantee removes it. The buyer pays *more* up front and churns *less*, because they're a value-buyer who committed.
---
## The pattern across all four
| Offer | Price | Risk reversal is on… | Scarcity type |
|-------|-------|----------------------|---------------|
| AudienceTap pilot | $297 | Whether it surfaces an actionable insight | Capacity (setup sessions) |
| SaberSim bundle | $497 | Whether it beats the current workflow | Seasonal (season start) |
| Teachable accelerator | $1,997 | Whether the course actually launches | Cohort (start date + cap) |
| Kit Painless Switch | $997 | Whether migration is live in 14 days | Capacity (engineer queue) |
Notice what's *not* here: a coupon, a "% off," a slashed sticker price. Every one raises value and reverses the real risk — and the scarcity is either capacity, cohort, or a genuine seasonal window, never a fake timer.
**The SaaS takeaway:** when conversion is stuck, your instinct will be to discount. Build the offer instead. Discount only to reward existing customers (upgrades, cross-sells) or in a real seasonal window — never to acquire a stranger you'll watch churn at 2×.
FILE:references/scarcity-urgency.md
# Scarcity & Urgency
The reason to buy now, not later.
This is the most misused offer lever in the industry. Done right, it's a meaningful conversion lift and a respect-the-buyer's-time gesture. Done wrong (fake countdown timers, lies about inventory), it converts once and torches trust permanently.
The repo voice: **only ship real scarcity.** If the scarcity is fake, take it off the page.
## Scarcity vs urgency
| | What it limits | Examples |
|---|---|---|
| **Scarcity** | Quantity | Cohort size, seats, inventory, batch, capacity |
| **Urgency** | Time | Cohort deadline, season, bonus expiry, price increase |
Both work via the same mechanism — they convert "I'll think about it" into "decide now." The difference is what enforces the decision: a limit on how many, or a limit on when.
---
## Honest scarcity formats
The bar: **the constraint has to be real.** Here are the formats that work without lying.
### Capacity-based scarcity
You can only deliver to N customers at a time. The next slot opens when one finishes.
> "We take 2 fractional CMO engagements per quarter. Next slot opens September 15."
Works for: services, agencies, coaching, consulting, anything labor-bound.
### Cohort scarcity
The class starts on a date. After the date, the door closes until the next cohort.
> "Cohort 7 starts October 1. Doors close September 28. Next cohort: January."
Works for: courses, programs, group coaching, anything synchronous.
### Founding-member pricing
The first N buyers get a different price. After that, the price goes up.
> "Founding pricing — $497/mo for the first 50 members. After we hit 50, the price moves to $797/mo."
Works for: SaaS, communities, memberships. Critical: actually raise the price when you hit 50. Otherwise it was a lie.
### Inventory-based scarcity
You have N units. When they're gone, they're gone.
> "We're only producing 100 of the print edition. Sold out, no reprints."
Works for: physical products, limited-run digital, bespoke services.
### Seasonal scarcity
Demand or capacity is seasonal.
> "Tax-prep service for Q1 2026 closes January 15. After that, we focus on existing clients only until April."
Works for: tax, retail Q4, events, education enrollment, anything calendar-bound.
### Bonus expiry
A bonus is available only until X date. The core offer stays the same; the *added value* expires.
> "Order by Friday and the 1:1 onboarding session is included. After Friday, the offer is the same but onboarding is +$497."
Works for: anything with a launch window. Critical: actually remove the bonus when the date hits.
### Price-increase scheduling
The price goes up on a date. Communicate it transparently.
> "On November 1, the program moves from $1,997 to $2,497. Anyone enrolled before November 1 keeps the $1,997 price for life."
Works for: courses, SaaS, communities. Builds credibility because the buyer can verify it later.
---
## Fake scarcity to avoid
Pattern-match and rip these off your page:
### Countdown timers that reset
The timer hits 0:00 and refreshes. The buyer comes back tomorrow, same timer, same urgency. Every buyer who notices loses trust.
If you use a countdown, it ends on a real date. After that date, the discount or bonus actually ends — verifiably.
### "Only 3 spots left"
Especially on digital products where there's no capacity constraint. Especially when the number stays at "3" for weeks.
Use this only when the constraint is real (cohort, capacity, inventory) AND when the number reflects reality.
### Manufactured FOMO
"127 people are looking at this page right now."
These can be honest on some platforms (Booking.com etc.) when they reflect actual traffic. They're dishonest when fabricated. Buyers increasingly assume fabrication.
### Bonus "stacking" that's always available
"This bonus expires at midnight!" — but every page on the site says the same thing at midnight every night.
The bonus has to actually expire, or the line is a lie.
### "Last chance" emails for things that aren't actually last chance
Repeated "FINAL HOURS" emails that send weekly. Every recipient who notices unsubscribes or stops opening.
Use "last chance" once. If you use it multiple times, you've taught your list that "last chance" is meaningless.
---
## Why fake scarcity is uniquely costly
Buyers compare notes. The internet is small.
- One Reddit post about a fake countdown becomes the top Google result for your brand
- One screenshot of "only 3 left" staying at 3 for a month becomes a viral thread
- One "last chance" email that wasn't goes into a "marketing fails" newsletter
Real scarcity converts ~the same as fake scarcity at the moment of purchase. The difference shows up at month 6, when fake-scarcity offers are facing trust collapse and real-scarcity offers are still compounding.
If you have to fake it, you don't have an offer-design problem — you have a value-equation problem. Go back to [value-equation.md](value-equation.md).
---
## When scarcity isn't needed
Some offers don't need scarcity:
- **Subscription products** that customers can cancel anytime — the natural friction is low enough
- **Low-priced impulse products** ($5–30) — the deliberation is short; scarcity feels forced
- **Premium / luxury brands** — scarcity is implicit in the positioning; explicit scarcity reads as low-status
- **High-trust audiences** who already know they'll buy — scarcity is unnecessary friction
Don't force scarcity into offers that don't need it. The forced version is worse than no scarcity.
---
## The diagnostic
For an existing offer with weak or no scarcity:
1. **Is there a real constraint?** Capacity, cohort, inventory, season, batch. Find the real one.
2. **What's the honest version of that constraint?** Write it in plain English.
3. **Can the buyer verify it?** If you say "founding pricing for the first 50 members," can the buyer see the count?
4. **Are you willing to actually enforce it?** When the constraint hits, do you have the discipline to actually close the door / raise the price / remove the bonus?
5. **Where on the page does the scarcity appear?** It should be next to the buy button, not buried.
If you can't find a real constraint, don't ship scarcity. The "trick" of fake scarcity is one of the most expensive shortcuts in marketing. The honest version is usually cheaper than people think — every business has *some* real constraint (capacity, calendar, batch, attention).
---
## Pairing with the rest of the offer
Scarcity is the final lever. It only works if the rest of the offer is strong.
- Strong offer + real scarcity = the buyer decides now
- Weak offer + scarcity = the buyer decides not to buy, faster
If conversions are bad, scarcity is rarely the first thing to fix. Run the diagnostic from [value-equation.md](value-equation.md) first.
The order is:
1. Strong dream outcome
2. Specific perceived likelihood (proof + methodology + guarantee)
3. Compressed time delay
4. Reduced effort & sacrifice
5. *Then* scarcity to get the decision now
Scarcity is the close, not the offer.
FILE:references/value-equation.md
# The Value Equation
The single most useful frame for offer design. Originated in Alex Hormozi's *$100M Offers*; the underlying idea (multiply benefits, divide costs) is much older — direct-response copywriters have been doing it for a century.
## The formula
```
Dream Outcome × Perceived Likelihood of Achievement
Value = ─────────────────────────────────────────────────────────
Time Delay × Effort & Sacrifice
```
The customer compares the *Value* score to the *Price*. If Value > Price, they buy. If not, they don't, no matter how good the copy is.
**Price is the comparison, not the value.** Most "lower the price" requests are actually "raise the numerator or lower the denominator" requests.
---
## Lever 1: Dream Outcome (numerator)
What the customer *actually* wants — usually one or two levels above the surface ask.
### Surface ask → dream outcome
| Surface ask | Dream outcome |
|-------------|---------------|
| "I want a website" | "I want more qualified leads I can close" |
| "I want to learn copywriting" | "I want to write copy clients pay me $5K+ per project for" |
| "I want a meal plan" | "I want to feel confident in a swimsuit on a beach in 12 weeks" |
| "I want fitness coaching" | "I want my back pain gone so I can pick up my kids without thinking about it" |
| "I want a Notion template" | "I want to feel in control of my work for the first time in years" |
| "I want to lower CAC" | "I want a marketing engine I can step away from for two weeks without things breaking" |
### How to increase
- **Name it specifically.** "Feel confident in a swimsuit" beats "lose weight." The specific name is the offer.
- **Connect to the bigger goal.** The dream outcome behind the surface ask is almost always emotional or identity-based.
- **Show the future state in concrete sensory terms.** What does the morning of day one *after they have it* actually look like?
### Common mistake
Pitching the surface ask. ("Get a website" instead of "get a website that brings you 5 qualified leads a week.") The bigger the buyer's pain, the more important this is — they're not buying a deliverable, they're buying a future.
---
## Lever 2: Perceived Likelihood of Achievement (numerator)
Do they actually believe they'll get the dream outcome? This is the single most underweighted lever — most offers are perfectly fine on the dream outcome side but the buyer just doesn't think it'll work *for them*.
### How to increase
- **Proof** — case studies with names, numbers, before/after metrics, photos. Specific > glossy.
- **Methodology specificity** — name your process. "The 5-step VAULT framework" beats "our proprietary system." Even if the substance is the same, naming it raises perceived likelihood.
- **Guarantees** — risk reversal directly raises perceived likelihood (more in [guarantee-design.md](guarantee-design.md)).
- **Reduce sample-of-one objection** — show people *like them* who got results. "Other people get results but I'm different" is the universal objection.
- **Pre-empt the failure path** — explicitly address what could go wrong and how you handle it. Builds trust faster than hiding the risk.
### Common mistake
Stacking more features ("we also include X, Y, Z") instead of stacking more proof. Features address dream outcome. Proof addresses likelihood. Most stuck offers need proof, not features.
---
## Lever 3: Time Delay (denominator)
How long from purchase to result. The denominator is doing a lot of work here — slow results don't just feel slow, they erode perceived likelihood (the longer it takes, the less the buyer believes it'll work).
### How to decrease
- **Faster first win** — find the smallest possible early result and engineer it into the first 7 days
- **Onboarding velocity** — replace a 60-minute kickoff with a 10-minute async intake + Loom send
- **Front-load the assets** — give the templates / swipe files / decks on day 1, not "throughout the program"
- **Faster end-to-end timeline** — "in 8 weeks" beats "in 6 months." If you can credibly compress, do.
- **Quick-start path** — explicit "first thing to do today" so they don't lose momentum
### Common mistake
Promising a faster timeline than you can deliver. The first time you miss it, perceived likelihood for every future buyer drops permanently as the failure story spreads.
---
## Lever 4: Effort & Sacrifice (denominator)
What the buyer pays *besides money* — time, learning curve, decisions, willpower, social risk, opportunity cost.
### How to decrease
- **Done-for-you over do-it-yourself** — DFY tiers move the effort to you. Charge accordingly.
- **Fewer decisions** — every decision the buyer makes is friction. Bundle, default, recommend.
- **Lower learning curve** — pre-built templates, examples, defaults. "We did the thinking, you do the executing."
- **Removed risk** — emotional risk (am I dumb if this doesn't work?), social risk (what will my team think?), opportunity cost (what am I not doing while I do this?). Address all three explicitly.
- **Async over live** — for many buyers (founders, executives), removing synchronous commitments is huge value
- **Less willpower required** — automation, accountability, environmental design
### Common mistake
Overestimating how much your buyer *enjoys* the work. They want the outcome, not the process. (Exception: identity-buyers — fitness, learning, mastery. For those, the process *is* part of the dream outcome. Read your buyer.)
---
## Diagnostic: scoring the levers
When an offer is stuck, score each lever 1–10 honestly. The lowest is the binding constraint.
### Quick scoring prompts
- **Dream outcome (1–10)**: Can the buyer picture, in concrete sensory terms, the day after this works? Is the outcome named specifically enough that they can repeat it back to a friend?
- **Perceived likelihood (1–10)**: Do they have at least three named, comparable proof points? Is the methodology specific enough to repeat? Have you addressed the "but my situation is different" objection?
- **Time delay (1–10)**: What's the first concrete win, and how soon do they get it? What's the end-to-end timeline? Are there any "delay surfaces" (onboarding lag, asset drip, week-1 friction)?
- **Effort & sacrifice (1–10)**: How much work does the buyer do? How many decisions? How much learning curve? How much synchronous time? How much emotional/social/opportunity-cost risk?
### Worked example: a stuck $3K copywriting course
Initial scoring:
- Dream outcome: 6 (`become a better copywriter` — too vague)
- Perceived likelihood: 4 (one testimonial, no methodology name)
- Time delay: 5 (6-month course, no first-win mechanic)
- Effort & sacrifice: 4 (lots of homework, live calls, weekly assignments)
Lowest: effort & sacrifice and perceived likelihood are tied. Pick one (perceived likelihood — easier to move and unblocks more downstream work):
- Name the methodology: "The VAULT framework — 5 angles every winning sales page uses"
- Add 8 named-customer case studies with before/after copy + revenue numbers
- Add an "even if you've never written before" cohort with 3 named graduates
After: perceived likelihood goes 4 → 8. Course converts at ~3x baseline. Two months later, attack effort & sacrifice (replace weekly live calls with async + 1 live Q&A).
---
## Key idea to internalize
**The price-vs-value comparison happens in the buyer's head, not yours.** You only know it's working when the buyer can articulate the dream outcome back to you in their own words, and when they don't need to ask "but does this actually work?"
If they're asking either of those questions, you have a value equation problem, not a copy problem.
Chụp toàn bộ trang web, kể cả SPA, ảnh tải chậm và trang rất dài, qua Chrome DevTools Protocol không cần thư viện ngoài.
--- name: "full-page-screenshot" description: "Use when the user asks to capture a full-page screenshot, long screenshot, or complete page capture of a web page. Handles SPA scroll containers, lazy-loaded images, and very tall pages via Chrome DevTools Protocol with zero external dependencies." --- # Full Page Screenshot Capture a full-page screenshot of any web page via Chrome DevTools Protocol. Produces a single PNG that includes all content — even portions that require scrolling. Zero external dependencies beyond Node.js 22+ and Chrome with remote debugging enabled. ## Prerequisites - **Node.js 22+** (uses built-in `WebSocket`) - **Chrome/Chromium** with remote debugging enabled Check environment readiness: ```bash node "SKILL_DIR/scripts/full-page-screenshot.mjs" --check ``` If Chrome check fails, instruct user to open `chrome://inspect/#remote-debugging` and enable **"Allow remote debugging for this browser instance"**. ## Workflow ### Option A: Screenshot an already-open tab (recommended for authenticated pages) 1. List available tabs: ```bash node "SKILL_DIR/scripts/full-page-screenshot.mjs" --list ``` 2. Identify the target by title/URL, then capture: ```bash node "SKILL_DIR/scripts/full-page-screenshot.mjs" <targetId> /tmp/screenshot.png --width 1200 --dpr 1 ``` ### Option B: Screenshot a URL (opens a background tab, captures, closes) ```bash node "SKILL_DIR/scripts/full-page-screenshot.mjs" --url "https://example.com" /tmp/screenshot.png --width 1200 --dpr 1 --wait 15000 ``` > **Note:** `--url` mode creates a background tab. Pages requiring authentication (SSO, login walls) should use Option A instead. ### Parameters | Parameter | Description | Default | |-----------|-------------|---------| | `output` | Output PNG file path | `/tmp/screenshot.png` | | `--width` | Viewport width in CSS pixels (articles: 1200, dashboards: 1440-1920) | 1200 | | `--dpr` | Device pixel ratio (2 = Retina, but 4x file size) | 1 | | `--wait` | Page load timeout in ms (`--url` mode only) | 15000 | | `--css` | Custom CSS to inject before capture (e.g., hide elements) | — | ### Verify Output ```bash # macOS sips -g pixelWidth -g pixelHeight /tmp/screenshot.png # Linux file /tmp/screenshot.png ``` ## Core Capabilities 1. **SPA scroll container expansion** — Detects `overflow-y: auto/scroll` containers, scrolls through them to trigger lazy-loading, then removes overflow constraints (including Tailwind `h-[calc(...)]`) so all content renders in a single pass. 2. **DOM stability detection** — After `readyState=complete`, monitors DOM element count until it stabilizes. This ensures SPA frameworks finish rendering dynamic content. 3. **Lazy-load triggering** — Scrolls the viewport incrementally to fire `IntersectionObserver` callbacks, then waits for all `<img>` elements to complete loading. 4. **Tiled capture for very tall pages** — Pages exceeding 16,000px are captured in 8,000px tiles and automatically stitched using Python PIL. Falls back to saving tiles separately if PIL is unavailable. 5. **Auto-discovery of Chrome** — Reads `DevToolsActivePort` file to find the debugging port. Falls back to probing ports 9222, 9229, 9333. 6. **CDP Proxy fallback** — When a CDP proxy holds the browser WebSocket, the script falls back to proxy API endpoints (`/eval`, `/screenshot`, `/scroll`) for capture. ## How It Works ``` 1. Discover Chrome debugging port 2. Connect via WebSocket (CDP) 3. Attach to target / create background tab 4. Set viewport width via Emulation domain 5. Wait: readyState + DOM stability 6. Detect & expand scroll containers 7. Scroll through page (trigger lazy-load) 8. Wait for images to complete 9. Measure final content height 10. Page.captureScreenshot (or tiled capture) 11. Stitch tiles if needed (PIL) 12. Restore viewport, detach, clean up ``` ## Anti-Patterns | Do NOT | Do instead | |--------|-----------| | Use `--dpr 2` on pages > 10,000px tall | Use `--dpr 1` to avoid Chrome memory issues | | Use `--url` for authenticated/SSO pages | Use `--list` + targetId on a tab where user is logged in | | Set `--wait` below 5000 for SPAs | SPAs need time to fetch data and render; use 10000-15000 | | Capture without checking `--check` first | Always verify Chrome debugging is available | | Hardcode viewport widths for all pages | Use 1200 for articles, 1440+ for dashboards/tables | | Skip output verification | Always verify with `sips` or `file` command after capture | ## Troubleshooting | Symptom | Cause | Fix | |---------|-------|-----| | "Cannot find Chrome debugging port" | Remote debugging not enabled | Open `chrome://inspect/#remote-debugging`, enable it | | "WebSocket connection timeout" | CDP proxy holding the connection | Script auto-falls back to proxy API | | Blank/white screenshot | Page not loaded yet | Increase `--wait` value | | Truncated at bottom | Scroll container not expanded | Script handles this automatically; file an issue if it persists | | Out of memory | Very tall page + high DPR | Reduce `--dpr` to 1 and/or reduce `--width` | | "PIL not available for stitching" | Python Pillow not installed | Install with `pip3 install Pillow` or accept separate tile files | ## Cross-References - [`engineering/browser-automation`](../browser-automation/SKILL.md) — General browser automation patterns via CDP/Playwright - [`engineering/performance-profiler`](../performance-profiler/SKILL.md) — Performance analysis that may complement visual captures FILE:scripts/full-page-screenshot.mjs #!/usr/bin/env node // full-page-screenshot.mjs — Standalone full-page screenshot tool via Chrome CDP // // Modes: // --check Check environment (Node.js 22+, Chrome debugging port) // --list List open browser tabs as JSON // --url <URL> [output] [options] Screenshot a URL (create tab → wait → capture → close) // <targetId> [output] [options] Screenshot an existing tab by target ID // // Options: // --width N Viewport width in CSS pixels (default: 1200) // --dpr N Device pixel ratio (default: 1) // --wait N Page load timeout in ms, --url mode only (default: 15000) // // Requires: Node.js 22+, Chrome with remote debugging enabled import fs from 'fs'; import net from 'net'; import path from 'path'; import { platform, homedir } from 'os'; // ─── Argument parsing ─────────────────────────────────────────────────────── const args = process.argv.slice(2); const flags = {}; const positional = []; const boolFlags = new Set(['check', 'list']); for (let i = 0; i < args.length; i++) { if (args[i].startsWith('--')) { const key = args[i].slice(2); if (boolFlags.has(key)) { flags[key] = true; } else if (i + 1 < args.length) { flags[key] = args[++i]; } } else { positional.push(args[i]); } } const vpWidth = parseInt(flags.width || '1200', 10); const dpr = parseInt(flags.dpr || '1', 10); const loadTimeout = parseInt(flags.wait || '15000', 10); const customCss = flags.css || ''; // ─── Chrome discovery ─────────────────────────────────────────────────────── function getDevToolsActivePortPaths() { const home = homedir(); if (platform() === 'darwin') { return [ path.join(home, 'Library/Application Support/Google/Chrome/DevToolsActivePort'), path.join(home, 'Library/Application Support/Google/Chrome Canary/DevToolsActivePort'), path.join(home, 'Library/Application Support/Chromium/DevToolsActivePort'), ]; } else if (platform() === 'linux') { return [ path.join(home, '.config/google-chrome/DevToolsActivePort'), path.join(home, '.config/chromium/DevToolsActivePort'), ]; } else if (platform() === 'win32') { const local = process.env.LOCALAPPDATA || ''; return [ path.join(local, 'Google/Chrome/User Data/DevToolsActivePort'), path.join(local, 'Chromium/User Data/DevToolsActivePort'), ]; } return []; } function checkPort(port) { return new Promise((resolve) => { const socket = net.createConnection(port, '127.0.0.1'); const timer = setTimeout(() => { socket.destroy(); resolve(false); }, 2000); socket.once('connect', () => { clearTimeout(timer); socket.destroy(); resolve(true); }); socket.once('error', () => { clearTimeout(timer); resolve(false); }); }); } async function discoverChrome() { // 1. Try DevToolsActivePort file for (const p of getDevToolsActivePortPaths()) { try { const lines = fs.readFileSync(p, 'utf-8').trim().split('\n'); const port = parseInt(lines[0], 10); if (port > 0 && port < 65536) { const ok = await checkPort(port); if (ok) { const wsPath = lines[1] || '/devtools/browser'; return { port, wsUrl: `ws://127.0.0.1:portwsPath` }; } } } catch { /* try next */ } } // 2. Fallback: probe common debugging ports for (const port of [9222, 9229, 9333]) { const ok = await checkPort(port); if (ok) { return { port, wsUrl: `ws://127.0.0.1:port/devtools/browser` }; } } return null; } // ─── CDP WebSocket helpers ────────────────────────────────────────────────── let msgId = 0; const pending = new Map(); let ws; function send(method, params = {}, sessionId = null, timeoutMs = 60000) { return new Promise((resolve, reject) => { const id = ++msgId; const msg = { id, method, params }; if (sessionId) msg.sessionId = sessionId; pending.set(id, resolve); ws.send(JSON.stringify(msg)); setTimeout(() => { if (pending.has(id)) { pending.delete(id); reject(new Error(`CDP timeout (timeoutMsms): method`)); } }, timeoutMs); }); } const sleep = (ms) => new Promise((r) => setTimeout(r, ms)); async function connectChrome() { const chrome = await discoverChrome(); if (!chrome) { console.error('Cannot find Chrome debugging port.'); console.error('Open chrome://inspect/#remote-debugging and enable "Allow remote debugging for this browser instance".'); process.exit(1); } ws = new WebSocket(chrome.wsUrl); ws.onmessage = (evt) => { const msg = JSON.parse(typeof evt.data === 'string' ? evt.data : evt.data.toString()); if (msg.id !== undefined && pending.has(msg.id)) { pending.get(msg.id)(msg); pending.delete(msg.id); } }; await new Promise((resolve, reject) => { const timer = setTimeout(() => { ws.close(); reject(new Error('WebSocket connection timeout (10s) — browser WebSocket may be held by proxy')); }, 10000); ws.onopen = () => { clearTimeout(timer); resolve(); }; ws.onerror = () => { clearTimeout(timer); reject(new Error('WebSocket connection to Chrome failed')); }; }); return chrome; } function closeWs() { try { ws?.close(); } catch {} } // ─── Wait for page load ───────────────────────────────────────────────────── async function waitForLoad(sid, timeoutMs) { const start = Date.now(); // Phase 1: wait for readyState=complete while (Date.now() - start < timeoutMs) { try { const resp = await send('Runtime.evaluate', { expression: 'document.readyState', returnByValue: true, }, sid, 5000); if (resp.result?.result?.value === 'complete') break; } catch { /* retry */ } await sleep(500); } // Phase 2: wait for DOM to stabilize (SPA content rendering) // SPAs load shell HTML instantly, then fetch data and render dynamically. // We detect stability by checking if DOM element count stops changing. const stabilityTimeout = Math.min(15000, Math.max(0, timeoutMs - (Date.now() - start))); if (stabilityTimeout > 0) { console.log('Waiting for DOM to stabilize...'); let lastCount = 0; let stableRounds = 0; const stableThreshold = 3; // need 3 consecutive stable checks (1.5s) const checkStart = Date.now(); while (Date.now() - checkStart < stabilityTimeout) { try { const resp = await send('Runtime.evaluate', { expression: 'document.querySelectorAll("*").length', returnByValue: true, }, sid, 5000); const count = resp.result?.result?.value || 0; if (count === lastCount && count > 0) { stableRounds++; if (stableRounds >= stableThreshold) { console.log(`DOM stable at count elements`); return true; } } else { stableRounds = 0; lastCount = count; } } catch { /* retry */ } await sleep(500); } console.log(`DOM stability timeout (last count: lastCount), proceeding`); } return true; } // ─── Core screenshot logic ────────────────────────────────────────────────── async function captureFullPage({ sid, outputFile, width, devicePixelRatio, css }) { // Set target width with a short viewport to measure content height await send('Emulation.setDeviceMetricsOverride', { width, height: 800, deviceScaleFactor: devicePixelRatio, mobile: false, }, sid); await sleep(2000); // Inject custom CSS if provided (e.g. hide sidebars, adjust layout) if (css) { await send('Runtime.evaluate', { expression: `(() => { const s = document.createElement('style'); s.textContent = JSON.stringify(css); document.head.appendChild(s); })()`, returnByValue: true, }, sid); await sleep(500); } // Expand internal scroll containers so their full content is visible in the screenshot. // Many SPAs use overflow-y:auto/scroll on inner divs instead of document-level scrolling. // We detect these, scroll through them to trigger lazy-loading, then remove overflow constraints. const expandResult = await send('Runtime.evaluate', { expression: `(async () => { const containers = []; for (const el of document.querySelectorAll('*')) { const style = getComputedStyle(el); const overflowY = style.overflowY; if ((overflowY === 'auto' || overflowY === 'scroll') && el.scrollHeight > el.clientHeight + 50 && el.clientHeight >= 100) { containers.push(el); } } if (containers.length === 0) return { found: 0 }; // Scroll each container to trigger lazy loading inside it for (const el of containers) { const step = 800; for (let y = 0; y < el.scrollHeight; y += step) { el.scrollTop = y; await new Promise(r => setTimeout(r, 150)); } el.scrollTop = 0; } // Expand containers: remove overflow and fixed-height constraints // Use !important to override CSS class-based constraints (e.g. Tailwind h-[calc(...)]) function expandEl(el) { el.style.setProperty('overflow', 'visible', 'important'); el.style.setProperty('overflow-y', 'visible', 'important'); el.style.setProperty('max-height', 'none', 'important'); // Override computed height if it constrains content const computed = getComputedStyle(el); const computedH = parseFloat(computed.height); if (computedH < el.scrollHeight - 10) { el.style.setProperty('height', 'auto', 'important'); } } let expanded = 0; for (const el of containers) { if (el.scrollHeight > el.clientHeight + 50) { expandEl(el); // Walk up the parent chain and expand anything that clips let parent = el.parentElement; for (let i = 0; i < 10 && parent && parent !== document.body; i++) { const ps = getComputedStyle(parent); if (ps.overflow !== 'visible' || ps.overflowY !== 'visible' || parseFloat(ps.height) < parent.scrollHeight - 10) { expandEl(parent); } parent = parent.parentElement; } expanded++; } } return { found: containers.length, expanded }; })()`, returnByValue: true, awaitPromise: true, }, sid, 60000); const expInfo = expandResult.result?.result?.value; if (expInfo && expInfo.expanded > 0) { console.log(`Expanded expInfo.expanded scroll container(s)`); await sleep(1000); } // Measure content height after expansion const m1 = await send('Page.getLayoutMetrics', {}, sid); const contentH = Math.ceil(m1.result.cssContentSize.height); console.log(`Content: widthxcontentH @ devicePixelRatiox`); // Set viewport height — cap at 8000px to avoid Chrome memory issues on very tall pages. // captureBeyondViewport: true will still capture the full page content. const vpHeight = Math.min(contentH, 8000); await send('Emulation.setDeviceMetricsOverride', { width, height: vpHeight, deviceScaleFactor: devicePixelRatio, mobile: false, }, sid); await sleep(2000); // Scroll through page slowly to trigger window-level lazy loading // Non-fatal: if this times out (common on very tall expanded pages), we skip it try { await send('Runtime.evaluate', { expression: `(async()=>{ const h = document.documentElement.scrollHeight; const step = 800; const maxSteps = Math.ceil(h / step); const deadline = Date.now() + 15000; for(let i=0; i<maxSteps; i++){ if(Date.now()>deadline) break; window.scrollTo(0, i * step); await new Promise(r=>setTimeout(r, 200)); } window.scrollTo(0, document.documentElement.scrollHeight); await new Promise(r=>setTimeout(r, 300)); window.scrollTo(0, 0); })()`, returnByValue: true, awaitPromise: true, }, sid, 20000); } catch (e) { console.warn(`Scroll lazy-load skipped (e.message)`); } // Wait for all images to finish loading (non-fatal timeout) try { const imgResult = await send('Runtime.evaluate', { expression: `(async()=>{ const deadline = Date.now() + 10000; while (Date.now() < deadline) { const imgs = Array.from(document.querySelectorAll('img')); const pending = imgs.filter(i => !i.complete && i.src && !i.src.startsWith('data:')); if (pending.length === 0) return { loaded: imgs.length, waited: false }; await new Promise(r => setTimeout(r, 500)); } const imgs = Array.from(document.querySelectorAll('img')); const still = imgs.filter(i => !i.complete && i.src && !i.src.startsWith('data:')); return { loaded: imgs.length - still.length, pending: still.length, timeout: true }; })()`, returnByValue: true, awaitPromise: true, }, sid, 15000); const imgInfo = imgResult.result?.result?.value; if (imgInfo) { if (imgInfo.timeout) { console.warn(`Warning: imgInfo.pending image(s) still loading after 10s timeout`); } else { console.log(`Images loaded: imgInfo.loaded`); } } } catch (e) { console.warn(`Image wait skipped (e.message)`); } // Final measure (content may have grown after lazy-load) const m2 = await send('Page.getLayoutMetrics', {}, sid); const finalH = Math.ceil(m2.result.cssContentSize.height); if (finalH !== contentH) { console.log(`Height adjusted: contentH → finalH`); const newVpH = Math.min(finalH, 8000); await send('Emulation.setDeviceMetricsOverride', { width, height: newVpH, deviceScaleFactor: devicePixelRatio, mobile: false, }, sid); await sleep(1000); } // Capture — for very tall pages (>16000px), use tiled capture to avoid Chrome timeout const TILE_THRESHOLD = 16000; if (finalH <= TILE_THRESHOLD) { // Single capture console.log(`Capturing widthxfinalH...`); const shot = await send('Page.captureScreenshot', { format: 'png', captureBeyondViewport: true, clip: { x: 0, y: 0, width, height: finalH, scale: 1 }, }, sid, 120000); if (shot.error) { throw new Error(`Screenshot failed: JSON.stringify(shot.error)`); } const buf = Buffer.from(shot.result.data, 'base64'); fs.writeFileSync(outputFile, buf); const mb = (buf.length / 1024 / 1024).toFixed(2); console.log(`Saved: outputFile (width * devicePixelRatioxfinalH * devicePixelRatiopx, mb MB)`); } else { // Tiled capture for very tall pages const tileH = 8000; const tiles = []; for (let y = 0; y < finalH; y += tileH) { const h = Math.min(tileH, finalH - y); console.log(`Capturing tile tiles.length + 1: y=y, h=h...`); const shot = await send('Page.captureScreenshot', { format: 'png', captureBeyondViewport: true, clip: { x: 0, y, width, height: h, scale: 1 }, }, sid, 120000); if (shot.error) { throw new Error(`Tile screenshot failed: JSON.stringify(shot.error)`); } const tilePath = outputFile.replace(/\.png$/, `_tiletiles.length.png`); fs.writeFileSync(tilePath, Buffer.from(shot.result.data, 'base64')); tiles.push({ path: tilePath, y, h }); } // Stitch tiles using Python PIL (available on macOS) console.log(`Stitching tiles.length tiles (widthxfinalH)...`); const { execSync } = await import('child_process'); const stitchScriptPath = outputFile.replace(/\.png$/, '_stitch.py'); const stitchScript = [ 'from PIL import Image', `tiles = [tiles.map(t => `("${t.path", t.y)`).join(',')}]`, `out = Image.new("RGB", (width * devicePixelRatio, finalH * devicePixelRatio))`, 'for path, y in tiles:', ' tile = Image.open(path)', ` out.paste(tile, (0, y * devicePixelRatio))`, ' tile.close()', `out.save("outputFile")`, 'print(f"Stitched: {out.size[0]}x{out.size[1]}")', ].join('\n'); fs.writeFileSync(stitchScriptPath, stitchScript); try { const result = execSync(`python3 "stitchScriptPath"`, { encoding: 'utf-8', timeout: 60000 }); console.log(result.trim()); } catch (e) { // Fallback: if Python/PIL not available, keep tiles as separate files console.warn('PIL not available for stitching. Tiles saved as separate files:'); for (const t of tiles) console.log(` t.path`); // Copy first tile as the output for basic functionality fs.copyFileSync(tiles[0].path, outputFile); } try { fs.unlinkSync(stitchScriptPath); } catch {} // Clean up tile files for (const t of tiles) { try { fs.unlinkSync(t.path); } catch {} } const stat = fs.statSync(outputFile); const mb = (stat.size / 1024 / 1024).toFixed(2); console.log(`Saved: outputFile (width * devicePixelRatioxfinalH * devicePixelRatiopx, mb MB)`); } } // ─── Mode: --check ────────────────────────────────────────────────────────── async function modeCheck() { // Check Node.js version const nodeVer = process.versions.node; const major = parseInt(nodeVer.split('.')[0], 10); if (major >= 22) { console.log(`node: ok (vnodeVer)`); } else { console.log(`node: FAIL (vnodeVer, need 22+)`); process.exit(1); } // Check Chrome debugging port (TCP only, no WebSocket to avoid auth popup) const chrome = await discoverChrome(); if (chrome) { console.log(`chrome: ok (port chrome.port)`); } else { console.log('chrome: FAIL'); console.error('Open chrome://inspect/#remote-debugging and enable "Allow remote debugging for this browser instance".'); process.exit(1); } } // ─── Mode: --list ─────────────────────────────────────────────────────────── async function modeList() { // Try direct CDP first, fall back to proxy API try { await connectChrome(); try { const resp = await send('Target.getTargets'); const pages = resp.result.targetInfos .filter((t) => t.type === 'page') .map(({ targetId, title, url }) => ({ targetId, title, url })); console.log(JSON.stringify(pages, null, 2)); } finally { closeWs(); } } catch { // Direct connection failed — try proxy const proxyUp = await isProxyRunning(); if (!proxyUp) { console.error('Cannot connect to Chrome and no proxy running.'); process.exit(1); } const resp = await fetch(`PROXY_URL/targets`, { signal: AbortSignal.timeout(10000) }); const targets = await resp.json(); const pages = targets .filter((t) => t.type === 'page') .map(({ targetId, title, url }) => ({ targetId, title, url })); console.log(JSON.stringify(pages, null, 2)); } } // ─── Mode: --url ──────────────────────────────────────────────────────────── async function modeUrl() { const url = flags.url; const outputFile = positional[0] || '/tmp/screenshot.png'; // Try direct CDP first let connected = false; try { await connectChrome(); connected = true; } catch { // Browser WebSocket held by proxy } if (connected) { let createdTargetId = null; let sid = null; try { // Create background tab const create = await send('Target.createTarget', { url, background: true }); if (create.error) { throw new Error(`Failed to create tab: JSON.stringify(create.error)`); } createdTargetId = create.result.targetId; console.log(`Tab created: createdTargetId.slice(0, 16)...`); // Attach const attach = await send('Target.attachToTarget', { targetId: createdTargetId, flatten: true }); if (attach.error) { throw new Error(`Attach failed: JSON.stringify(attach.error)`); } sid = attach.result.sessionId; // Wait for page load await send('Page.enable', {}, sid); await waitForLoad(sid, loadTimeout); // Screenshot await captureFullPage({ sid, outputFile, width: vpWidth, devicePixelRatio: dpr, css: customCss }); } finally { if (sid) { try { await send('Emulation.clearDeviceMetricsOverride', {}, sid); } catch {} try { await send('Target.detachFromTarget', { sessionId: sid }); } catch {} } if (createdTargetId) { try { await send('Target.closeTarget', { targetId: createdTargetId }); } catch {} } closeWs(); } return; } // Fallback: use proxy to create tab and screenshot const proxyUp = await isProxyRunning(); if (!proxyUp) { console.error('Cannot connect to Chrome and no proxy running.'); process.exit(1); } console.log('Using proxy to create tab...'); const newResp = await fetch(`PROXY_URL/new?url=encodeURIComponent(url)`, { signal: AbortSignal.timeout(loadTimeout + 5000), }); const newTab = await newResp.json(); const targetId = newTab.targetId; console.log(`Tab created via proxy: targetId.slice(0, 16)...`); // Wait for DOM to stabilize (SPA rendering) console.log('Waiting for DOM to stabilize...'); let lastCount = 0, stableRounds = 0; const stabilityDeadline = Date.now() + 15000; while (Date.now() < stabilityDeadline) { const count = await proxyEval(targetId, 'document.querySelectorAll("*").length'); if (count === lastCount && count > 0) { stableRounds++; if (stableRounds >= 3) { console.log(`DOM stable at count elements`); break; } } else { stableRounds = 0; lastCount = count; } await sleep(500); } try { await captureViaProxy(targetId, outputFile, vpWidth, dpr); } finally { // Close the tab we created try { await fetch(`PROXY_URL/close?target=targetId`, { signal: AbortSignal.timeout(5000) }); } catch {} } } // ─── CDP Proxy helpers (used when proxy is running) ───────────────────────── const PROXY_URL = 'http://localhost:3456'; async function isProxyRunning() { try { const resp = await fetch(`PROXY_URL/health`, { signal: AbortSignal.timeout(1000) }); const data = await resp.json(); return data.status === 'ok'; } catch { return false; } } async function proxyEval(targetId, expr) { const resp = await fetch(`PROXY_URL/eval?target=targetId`, { method: 'POST', body: expr, signal: AbortSignal.timeout(60000), }); return (await resp.json()).value; } async function proxyScreenshot(targetId, filePath) { const resp = await fetch(`PROXY_URL/screenshot?target=targetId&file=encodeURIComponent(filePath)`, { signal: AbortSignal.timeout(30000), }); return await resp.json(); } async function proxyScroll(targetId, y) { await fetch(`PROXY_URL/scroll?target=targetId&y=y`, { signal: AbortSignal.timeout(10000), }); } // Full-page screenshot using proxy API (tiled viewport captures) async function captureViaProxy(targetId, outputFile, width, devicePixelRatio) { console.log('Using proxy API for full-page capture...'); // Detect scroll containers WITHOUT expanding them (expansion causes issues with proxy) const containerInfo = await proxyEval(targetId, `(() => { const containers = []; for (const el of document.querySelectorAll('*')) { const style = getComputedStyle(el); const overflowY = style.overflowY; if ((overflowY === 'auto' || overflowY === 'scroll') && el.scrollHeight > el.clientHeight + 50 && el.clientHeight >= 100) { // Find a unique selector for this element let selector = el.tagName.toLowerCase(); if (el.id) selector = '#' + el.id; else if (el.className) { const cls = el.className.trim().split(/\\s+/).filter(c => !c.includes('[')).slice(0, 3).join('.'); if (cls) selector = el.tagName.toLowerCase() + '.' + cls; } containers.push({ selector, scrollHeight: el.scrollHeight, clientHeight: el.clientHeight, index: containers.length }); } } return containers; })()`); const pageInfo = await proxyEval(targetId, `({ scrollHeight: Math.max(document.body.scrollHeight, document.documentElement.scrollHeight), viewportHeight: window.innerHeight, viewportWidth: window.innerWidth })`); const vpH = pageInfo.viewportHeight; console.log(`Page viewport: pageInfo.viewportWidthxvpH`); // Determine capture strategy const mainContainer = containerInfo && containerInfo.length > 0 ? containerInfo.reduce((a, b) => a.scrollHeight > b.scrollHeight ? a : b) : null; if (mainContainer && mainContainer.scrollHeight > vpH) { // SPA with internal scroll container — scroll the container, capture tiles console.log(`Scroll container: mainContainer.selector (mainContainer.scrollHeightpx content in mainContainer.clientHeightpx view)`); // First, scroll through container to trigger lazy loading await proxyEval(targetId, `(async () => { const containers = []; for (const el of document.querySelectorAll('*')) { const style = getComputedStyle(el); if ((style.overflowY === 'auto' || style.overflowY === 'scroll') && el.scrollHeight > el.clientHeight + 50 && el.clientHeight >= 100) { containers.push(el); } } const el = containers[mainContainer.index]; if (!el) return; const step = 800; for (let y = 0; y < el.scrollHeight; y += step) { el.scrollTop = y; await new Promise(r => setTimeout(r, 150)); } el.scrollTop = 0; })()`); // Wait for images await proxyEval(targetId, `(async () => { const deadline = Date.now() + 8000; while (Date.now() < deadline) { const imgs = Array.from(document.querySelectorAll('img')); const pending = imgs.filter(i => !i.complete && i.src && !i.src.startsWith('data:')); if (pending.length === 0) return; await new Promise(r => setTimeout(r, 500)); } })()`); // Capture tiles by scrolling the container const containerH = mainContainer.clientHeight; const totalScrollH = mainContainer.scrollHeight; const overlap = 20; const stepH = containerH - overlap; const tiles = []; for (let scrollY = 0; scrollY < totalScrollH; scrollY += stepH) { await proxyEval(targetId, `(() => { const containers = []; for (const el of document.querySelectorAll('*')) { const style = getComputedStyle(el); if ((style.overflowY === 'auto' || style.overflowY === 'scroll') && el.scrollHeight > el.clientHeight + 50 && el.clientHeight >= 100) { containers.push(el); } } const el = containers[mainContainer.index]; if (el) el.scrollTop = scrollY; })()`); await sleep(300); const tilePath = outputFile.replace(/\.png$/, `_tiletiles.length.png`); await proxyScreenshot(targetId, tilePath); tiles.push({ path: tilePath, scrollY }); console.log(`Tile tiles.length: scrollY=scrollY`); } // Reset scroll position await proxyEval(targetId, `(() => { const containers = []; for (const el of document.querySelectorAll('*')) { const style = getComputedStyle(el); if ((style.overflowY === 'auto' || style.overflowY === 'scroll') && el.scrollHeight > el.clientHeight + 50 && el.clientHeight >= 100) { containers.push(el); } } const el = containers[mainContainer.index]; if (el) el.scrollTop = 0; })()`); if (tiles.length === 1) { fs.renameSync(tiles[0].path, outputFile); } else { // Stitch tiles console.log(`Stitching tiles.length tiles...`); const { execSync } = await import('child_process'); const sipsOut = execSync(`sips -g pixelWidth -g pixelHeight "tiles[0].path"`, { encoding: 'utf-8' }); const pxW = parseInt(sipsOut.match(/pixelWidth:\s*(\d+)/)?.[1] || '0'); const pxH = parseInt(sipsOut.match(/pixelHeight:\s*(\d+)/)?.[1] || '0'); const dpr = pxH / vpH; // Calculate the pixel offset of the container within the viewport const containerOffset = await proxyEval(targetId, `(() => { const containers = []; for (const el of document.querySelectorAll('*')) { const style = getComputedStyle(el); if ((style.overflowY === 'auto' || style.overflowY === 'scroll') && el.scrollHeight > el.clientHeight + 50 && el.clientHeight >= 100) { containers.push(el); } } const el = containers[mainContainer.index]; if (!el) return 0; const rect = el.getBoundingClientRect(); return rect.top; })()`); const pxContainerTop = Math.round((containerOffset || 0) * dpr); const pxContainerH = Math.round(containerH * dpr); const pxStepH = Math.round(stepH * dpr); const pxTotalContentH = Math.round(totalScrollH * dpr); // Stitch: header (top of first tile) + container content strips from each tile + footer (bottom of last tile) const stitchScriptPath = outputFile.replace(/\.png$/, '_stitch.py'); const stitchLines = [ 'from PIL import Image', `tiles = [tiles.map(t => `"${t.path"`).join(',')}]`, 'imgs = [Image.open(p) for p in tiles]', 'tile_w, tile_h = imgs[0].size', `container_top = pxContainerTop`, `container_h = pxContainerH`, `step_h = pxStepH`, `total_content_h = pxTotalContentH`, 'header_h = container_top', 'footer_h = tile_h - container_top - container_h', 'out_h = header_h + total_content_h + footer_h', 'out = Image.new("RGB", (tile_w, out_h))', 'if header_h > 0:', ' header = imgs[0].crop((0, 0, tile_w, header_h))', ' out.paste(header, (0, 0))', 'for i, img in enumerate(imgs):', ' content_strip = img.crop((0, container_top, tile_w, container_top + container_h))', ' y = header_h + i * step_h', ' paste_h = min(container_h, out_h - footer_h - y)', ' if paste_h < container_h:', ' content_strip = content_strip.crop((0, 0, tile_w, paste_h))', ' if paste_h > 0:', ' out.paste(content_strip, (0, y))', 'if footer_h > 0:', ' footer = imgs[-1].crop((0, tile_h - footer_h, tile_w, tile_h))', ' out.paste(footer, (0, out_h - footer_h))', 'for img in imgs:', ' img.close()', `out.save("outputFile")`, 'print(f"Stitched: {out.size[0]}x{out.size[1]}")', ]; fs.writeFileSync(stitchScriptPath, stitchLines.join('\n')); try { const result = execSync(`python3 "stitchScriptPath"`, { encoding: 'utf-8', timeout: 60000 }); console.log(result.trim()); } catch (e) { console.warn('Stitching failed:', e.message?.slice(0, 200)); fs.copyFileSync(tiles[0].path, outputFile); } // Clean up tiles and stitch script try { fs.unlinkSync(stitchScriptPath); } catch {} for (const t of tiles) { try { fs.unlinkSync(t.path); } catch {} } } } else { // Normal page or small page — single viewport screenshot // Scroll through to trigger lazy loading first await proxyEval(targetId, `(async () => { const h = document.documentElement.scrollHeight; const step = 800; for (let y = 0; y < h; y += step) { window.scrollTo(0, y); await new Promise(r => setTimeout(r, 200)); } window.scrollTo(0, 0); })()`); await proxyScreenshot(targetId, outputFile); } const stat = fs.statSync(outputFile); const mb = (stat.size / 1024 / 1024).toFixed(2); console.log(`Saved: outputFile (mb MB)`); } // ─── Mode: targetId (existing tab) ───────────────────────────────────────── async function modeTarget() { const targetId = positional[0]; const outputFile = positional[1] || '/tmp/screenshot.png'; if (!targetId) { console.error('Usage:'); console.error(' node full-page-screenshot.mjs --check'); console.error(' node full-page-screenshot.mjs --list'); console.error(' node full-page-screenshot.mjs --url <URL> [output] [--width N] [--dpr N]'); console.error(' node full-page-screenshot.mjs <targetId> [output] [--width N] [--dpr N]'); process.exit(1); } // Try direct CDP first. If browser WebSocket is held by proxy, fall back to proxy API. let connected = false; try { await connectChrome(); connected = true; } catch { // Browser WebSocket likely held by proxy } if (connected) { let sid = null; try { const attach = await send('Target.attachToTarget', { targetId, flatten: true }); if (attach.error) { console.error(`Attach failed: JSON.stringify(attach.error)`); console.error('Run with --list to see available targets.'); process.exit(1); } sid = attach.result.sessionId; // Test with Page.enable — if times out, proxy may be interfering try { await send('Page.enable', {}, sid, 10000); } catch { console.warn('Direct session timed out, falling back to proxy...'); try { await send('Target.detachFromTarget', { sessionId: sid }); } catch {} sid = null; closeWs(); connected = false; } if (sid) { await captureFullPage({ sid, outputFile, width: vpWidth, devicePixelRatio: dpr, css: customCss }); } } finally { if (sid) { try { await send('Emulation.clearDeviceMetricsOverride', {}, sid); } catch {} try { await send('Target.detachFromTarget', { sessionId: sid }); } catch {} } if (connected) closeWs(); } } // Fallback to proxy API if (!connected) { const proxyUp = await isProxyRunning(); if (!proxyUp) { console.error('Cannot connect to Chrome (browser WebSocket unavailable) and no proxy running.'); process.exit(1); } await captureViaProxy(targetId, outputFile, vpWidth, dpr); } } // ─── Dispatch ─────────────────────────────────────────────────────────────── async function main() { if (flags.check) return modeCheck(); if (flags.list) return modeList(); if (flags.url) return modeUrl(); return modeTarget(); } main().catch((e) => { console.error('Error:', e.message || e); closeWs(); process.exit(1); });
Sinh chuỗi OKR từ chiến lược công ty xuống mục tiêu từng nhóm.
--- name: okr description: Generate OKR cascades from company strategy to team objectives. Usage: /okr generate <strategy> --- # /okr Generate cascaded OKR frameworks from company-level strategy down to team-level key results. ## Usage ``` /okr generate <strategy> Generate OKR cascade ``` Supported strategies: `growth`, `retention`, `revenue`, `innovation`, `operational` ## Input Format Pass a strategy keyword directly. The generator produces company, department, and team-level OKRs aligned to the chosen strategy. ## Examples ``` /okr generate growth /okr generate retention /okr generate revenue /okr generate innovation /okr generate operational /okr generate growth --json ``` ## Scripts - `product-team/product-strategist/scripts/okr_cascade_generator.py` — OKR cascade generator (`<strategy> [--teams "A,B,C"] [--contribution 0.3] [--json]`) ## Skill Reference > `product-team/product-strategist/SKILL.md`
Quyết định có ký đối tác không, ở hạng nào (giới thiệu, đại lý, OEM, SI, liên minh chiến lược), cam kết GTM chung và tỷ lệ chia doanh thu.
---
name: partnerships-architect
description: "Use when a startup is approached by a prospective partner and someone has to decide should we sign this partner, at what partner tier (referral / reseller / OEM / SI-consulting / strategic alliance), with what joint GTM commitment, and at what revshare. Classifies partner tier from independent-demand evidence vs. preferential-terms hunting, designs a 90-day joint GTM plan, models revshare against direct-sale margin, and surfaces kill criteria for unwinding under-performing partnerships. For Head of Partnerships, Head of BD, and Founder-CEOs doing reseller agreement, OEM deal, or strategic alliance review — not technical sale enablement, not channel cost economics, not M&A."
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [commercial, partnerships, channel-partners, joint-gtm, revshare, oem, reseller, strategic-alliance]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# partnerships-architect
## Purpose
Help Head of Partnerships, Head of BD, and Founder-CEOs answer four questions when a
prospective partner shows up:
1. **Is this a real partner, or someone hunting preferential terms without independent demand?**
2. **At what tier should we sign them?** (Referral / Reseller / OEM / SI-Consulting / Strategic Alliance)
3. **What's the 90-day joint GTM plan that proves the partnership works?**
4. **What revshare makes economic sense — and at what point does the partnership beat direct sale?**
The skill emits a tier verdict + GTM plan + revshare band with explicit kill criteria. It
does **not** sign the deal. The human, after running this skill, decides.
## When to use
- A prospective partner has approached and asked for reseller / OEM / "strategic" terms
- You're designing a new partner program tier structure
- You're reviewing an existing partnership that's underperforming and need to decide: re-tier, restructure GTM, or unwind
- A Big Logo wants a "strategic alliance" — and you need to validate it's real, not vendor-lock theatre
- A consulting firm or SI wants services revshare on your product
- A platform vendor offers OEM / white-label and you need to model the math
- You suspect "partner-sourced" deals are actually your own pipeline being skimmed for margin
**Do not use for:**
- Technical demos and POCs → `business-growth/sales-engineer`
- Cost-to-serve and ROI math on existing channel → sibling `channel-economics`
- Whole-company revenue strategy → `c-level-advisor/cro-advisor`
- Acquiring a company instead of partnering → `c-level-advisor/ma-playbook`
- Per-deal discount approval inside a signed partner contract → `deal-desk`
## Workflow
### Step 1 — Intake (≈ 20 min)
Fill `assets/partnership_intake_template.md`. Capture: partner_name, partner_type, evidence
of independent demand (named accounts they've sourced, end-customer relationships,
their sales team size), strategic value (geo / product / brand / channel economics),
commitments they've offered (joint marketing spend, dedicated headcount, certification,
sales targets).
If the intake template can't be honestly filled out, the prospective partner has not
demonstrated enough substance to evaluate. Stop. Go back to them.
### Step 2 — Tier classify
Run `scripts/partner_tier_classifier.py --input intake.json --profile saas --output markdown`.
Output ranks the partner into 1 of 5 tiers — REFERRAL / RESELLER / OEM / SI-CONSULTING /
STRATEGIC — with deterministic floors. STRATEGIC requires named_accounts ≥ 5 AND
multi-year commit AND dedicated resources. Skill emits rationale + kill criteria.
### Step 3 — Joint GTM plan
Run `scripts/joint_gtm_planner.py --input gtm.json --profile saas --output markdown`.
Output: 90-day plan with pre-launch milestones (training, certification, materials),
launch motion (target accounts, sales play, MDF allocation), mid-quarter checkpoint, and
90-day success criteria. Validates: cannot plan channel-led GTM for REFERRAL tier; cannot
plan white-label for non-OEM tier.
### Step 4 — Revshare model
Run `scripts/revshare_modeler.py --input revshare.json --output markdown`. Computes
margin per deal direct vs. via partner, recommended revshare % band based on partner
contribution depth (sourced > influenced > delivered), break-even partner ROI, and
long-term economics — at projected scale, does partner economics beat direct?
### Step 5 — Decide
Take tier + GTM plan + revshare band into the partnership committee. Skill does not sign
the partner — you do. Document kill criteria in the contract so the unwind is mechanical
when triggered.
## Scripts
- `scripts/partner_tier_classifier.py` — 5-tier classifier with deterministic floors per tier
- `scripts/joint_gtm_planner.py` — 90-day joint GTM plan generator with tier-validated motion
- `scripts/revshare_modeler.py` — revshare band + break-even ROI + long-term economics
All scripts: stdlib only. `--help` and `--sample` work on all three.
## References
- `references/channel_partner_canon.md` — Caro on HP indirect channels, Chintagunta on channel economics, Hessling on partner programs, Forrester channel software stack, IDC channel research, Tien Tzuo subscription-channel models, Geoffrey Moore whole-product partnerships
- `references/joint_gtm_canon.md` — Aaron Ross *Predictable Revenue* (cold-source vs partner), Winning by Design, Jay McBain on co-sell, Microsoft Partner Network playbook, AWS Partner Network research, SiriusDecisions partner benchmarks, Bridge Group SaaS partner data
- `references/partnership_anti_patterns.md` — Forrester partner-led-from-your-pipeline research, Tom Tunguz on channel conflict, Hessling failure analyses, MIT Sloan on disproportionate strategic revshare, HP channel post-mortems, IBM channel-conflict cases, Salesforce AppExchange research
## Assumptions
- A partner who cannot produce evidence of independent demand (named accounts, end-customer
relationships, their own sales team) is hunting preferential terms, not a partner.
- Industry profiles (`--profile`) tune defaults — they don't override your data.
- Revshare % bands are recommendations; the contract negotiation, MDF policy, and
exclusivity terms are human commercial decisions outside this skill.
- "Partner-sourced" requires the partner to have introduced the deal AND owned the
primary relationship. "Partner-influenced" pays at a lower band. Pay attribution
matters more than slide-deck claims.
- This skill is for partnership design, not signed-partner deal management — once
signed, per-deal commercial review routes to `deal-desk`.
- Kill criteria are mandatory. A partnership without a written unwind trigger compounds
the bad-partner problem over years.
## Anti-patterns
- **"Partner = anyone who asked."** A partner with no independent demand is a discount hunter.
Run the tier classifier — REFERRAL tier exists precisely to absorb these without giving
away reseller margin.
- **Granting OEM / white-label terms without margin sufficient to fund support.** OEM means
you support a customer you don't own. If the revshare doesn't fund Tier-2 support cost,
the OEM deal is a losing trade.
- **Paying sourced-tier revshare on influenced-only deals.** Influenced ≠ sourced. The deal
was going to close anyway. Pay the influenced rate.
- **No kill criteria for under-performing partner.** "Strategic alliances" without sunset
clauses become permanent obligations after the executive sponsor leaves.
- **Channel conflict ignored until reps quit.** When your direct rep and your partner both
show up at the same account, you lose either the rep or the partner. Decide the rules of
engagement before, not after.
- **Exclusive territory granted to a weak partner.** This locks out the strong partner who
would have actually sourced the deals.
- **MDF without ROI accountability.** Market Development Funds without named pipeline,
reported ROI, and a quarterly true-up are subsidy, not investment.
- **No offboarding plan when partnership ends.** Customer continuity, data hand-back, IP
cleanup, and brand take-down must be pre-negotiated. They're impossible to negotiate after
the relationship has soured.
## Distinct from
- **business-growth/sales-engineer** — technical sale: demos, POCs, integration scoping.
Operates after the partnership decision is made and a deal is in flight.
- **channel-economics** (sibling) — cost-to-serve and ROI math on an existing channel.
Quantifies whether a signed partner is profitable. partnerships-architect decides
whether to sign in the first place and at what tier.
- **c-level-advisor/cro-advisor** — strategic CRO judgment (when to hire a VP Channel,
whole-company revenue mix decisions). partnerships-architect is per-partnership.
- **c-level-advisor/ma-playbook** — when the answer is "acquire them" not "partner with
them." Trigger: the partner has independent moat you cannot replicate, or the
partnership requires equity to align incentives. Re-route to ma-playbook.
- **deal-desk** — per-deal discount approval on signed partner contracts.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-commercial` or the orchestrator. Recommended answer +
canon citation per question. Never bundled. Lock 1-3 before opening 4-6.
1. **"Name 5 end customers this partner has already sold to in the last 12 months — at companies you would target yourself."**
Recommended: if they cannot, they have no independent demand. Sign at REFERRAL tier only,
if at all. Reseller/OEM/Strategic floors require demonstrated end-customer relationships.
Canon: Joe Hessling — partner-program failure analyses identify "no independent demand"
as the #1 root cause of dead partner tiers.
2. **"Is this partner asking for preferential commercial terms, or asking how to bring you customers?"**
Recommended: discount hunters lead with terms; real partners lead with accounts. Listen
to the first 30 minutes of the first meeting.
Canon: Forrester channel research — 60%+ of "partner inquiries" at early-stage SaaS are
discount hunting, not channel investment.
3. **"What's the joint value proposition in one sentence, and who is the named end-customer it serves?"**
Recommended: if there is no joint value prop distinct from either party's solo offering,
there is no partnership — there is co-marketing at best.
Canon: Geoffrey Moore (*Crossing the Chasm*) — whole-product partnerships exist when
neither party alone delivers the customer outcome.
4. **"At what % discount / revshare does this partnership beat the direct-sale economics, and at what scale?"**
Recommended: model break-even pipeline volume. If partner-sourced deals must exceed
30% of channel volume to beat direct, and partner can plausibly deliver 5%, you have
built a losing program.
Canon: Pradeep Chintagunta (Chicago Booth) on channel economics — channel partnerships
without volume floor break even in theory and lose money in practice.
5. **"What are the named kill criteria for unwinding this partnership, and are they in the contract?"**
Recommended: minimum pipeline floor by quarter, minimum certified resources, minimum
joint deals closed, 90-day cure period. Unwinding without pre-agreed criteria becomes
a 2-year legal battle.
Canon: IBM channel-conflict case studies (1990s post-divestiture) — undocumented kill
criteria converted bad partners into permanent obligations.
6. **"If this partner sells to one of YOUR direct accounts, who wins — your rep or them?"**
Recommended: Rules of Engagement in writing, signed before kickoff. Territory by named
account, by segment, or by geo. Conflict resolution at named human, not committee.
Canon: Jay McBain (Canalys) — channel conflict is the #1 partner program killer; written
ROE published before partner signs prevents 80% of disputes.
7. **"Is this a partnership, or should this be an acquisition?"**
Recommended: if the partner has independent moat you cannot replicate AND the
partnership requires multi-year exclusivity AND the partnership requires equity-like
alignment, you're describing an acquisition. Re-route to `ma-playbook`.
Canon: HP channel post-mortems (Indigo, EDS partial integrations) — partnerships
structured as acquisitions-without-equity destroy more value than either pure path.
Walk depth-first. Lock 1-3 (is this a real partner?) before opening 4-7 (is the structure
right?). After all 7 are answered, invoke `partner_tier_classifier.py` →
`joint_gtm_planner.py` → `revshare_modeler.py` in sequence.
FILE:assets/partnership_intake_template.md
# Partnership Intake Template
**Owner:** _______________ **Date:** _______________
**Time to fill out:** ≈ 20 minutes
**Prospective partner name:** _______________
Fill this template out honestly BEFORE running `partner_tier_classifier.py`. The skill
outputs are only as good as the inputs. If you cannot honestly answer a field, write
"unknown" — do not guess. If multiple fields are "unknown," the partner has not
demonstrated enough substance to evaluate. Pause the process and go back to them.
---
## 1. Partner identity
- **Partner legal name:** _______________
- **Partner_type** (pick one): [ ] referral [ ] reseller [ ] oem [ ] si_consultant
[ ] technology [ ] strategic_alliance
- **Who introduced them?** _______________
- **Why are they approaching us NOW?** _______________
(If the honest answer is "they want preferential discount," classify as REFERRAL
and proceed accordingly. Do not advance to RESELLER+.)
## 2. Independent demand evidence
This is the most important section. STRATEGIC and OEM tiers have HARD floors here.
- **Named accounts they have sold to in the last 12 months, at companies you would
also target:**
1. _______________
2. _______________
3. _______________
4. _______________
5. _______________
- **`named_accounts_sourced_count` (count of verifiable, reference-able named
accounts):** _______________
- **Of their total customer base, what % are end customers (companies they sold
directly to and own the relationship), vs intermediaries / sub-partners?**
`end_customer_relationships_pct` (0-100): _______________
- **Sales team size — how many people on their team actively sell?**
`sales_team_size`: _______________
Note: "everyone is a salesperson at our company" is not an answer. Count the people
whose comp plan includes quota.
## 3. Strategic value (which of these does this partner change?)
- **`geo_coverage`** — geographies they reach that we don't / cover poorly:
_______________
- **`product_complement`** — what they bring that completes the customer outcome
(whole-product reasoning per Geoffrey Moore):
_______________
- **`brand_lift`** — does their brand carry credibility we lack?
[ ] strong [ ] mid [ ] none
- **`channel_economics_advantage`** — lower CAC, faster sales cycle, better retention
in a segment we struggle with?
_______________
## 4. Commitments they have offered
Be precise. Vague commitments are not commitments.
- **`joint_marketing_spend`** (USD per year): _______________
- **`dedicated_resources`** — named individuals on their team dedicated to this
partnership (not "we'll figure it out"):
count: _______________
names: _______________
- **`certification_completion`** — will their team complete our certification
curriculum?
[ ] yes, scheduled [ ] willing but not scheduled [ ] no
- **`sales_targets`** — specific named pipeline and closed-won targets, with a time
horizon:
_______________
(Example: "12 closed-won deals over 12 months, with named target accounts
identified in TAL.")
## 5. What they want from US
- **Revshare ask:** _______________
- **Exclusivity ask:** _______________ (territory / segment / vertical / none)
- **MDF ask:** _______________
- **Engineering integration ask:** _______________
- **Anything unusual:** _______________
## 6. Honest red-flag check
If any of these are true, the partner is a discount hunter, not a partner. Sign at
REFERRAL or do not sign at all.
- [ ] They cannot name 5 customers they have sold to in the last 12 months
- [ ] Their commercial ask is precise; their joint-value-prop ask is vague
- [ ] They want exclusive territory at signing with no performance condition
- [ ] They claim "strategic alliance" but have no exec sponsor on their side
- [ ] They are pushing for fast signing ("we have a deal we need to close this week")
---
## JSON skeleton for `partner_tier_classifier.py --input`
```json
{
"partner_name": "",
"partner_type": "",
"independent_demand_evidence": {
"named_accounts_sourced_count": 0,
"end_customer_relationships_pct": 0,
"sales_team_size": 0
},
"strategic_value": {
"geo_coverage": "",
"product_complement": "",
"brand_lift": "",
"channel_economics_advantage": ""
},
"commitments": {
"joint_marketing_spend": 0,
"dedicated_resources": 0,
"certification_completion": false,
"sales_targets": ""
}
}
```
Save as `partner.json`, then run:
```
python scripts/partner_tier_classifier.py --input partner.json --profile saas --output markdown
```
Then proceed to `joint_gtm_planner.py` and `revshare_modeler.py` only if the assigned
tier is RESELLER or higher AND your partnership committee has agreed to move forward.
FILE:references/channel_partner_canon.md
# Channel Partner Canon
Curated, opinionated knowledge base behind `partner_tier_classifier.py`'s scoring rules
and the 5-tier model. This is the source material; the script encodes the deterministic
floors derived from it.
## Core principle
A partner is not a discount channel. A partner brings independent demand, owns
end-customer relationships, and changes your distribution math. Anyone asking for
preferential commercial terms without those three is not a partner — they are a
discount hunter wearing a partnership-deck costume.
The 5-tier model exists to absorb the spectrum without giving away margin: REFERRAL is
the polite no, RESELLER and OEM are economic structures, SI/CONSULTING is a services
attach, STRATEGIC is reserved for the rare case where the partnership genuinely
re-shapes the market.
---
## The 5 tiers
### REFERRAL
Informal intro. No exclusivity. Small finder's fee (5-10% of first-year ARR, one-time).
No certification required. No co-marketing commitment. 2-quarter auto-sunset if no
qualified intros.
When you use this: 90%+ of inbound "partnership requests" at early-stage SaaS belong
here. A REFERRAL agreement says "we appreciate the intro, here's a finder's fee, we are
not building a joint motion."
### RESELLER
Transactional resale with margin. Partner's customer pays partner; partner remits net of
revshare. Floor: end_customer_relationships_pct ≥ 40%, sales_team_size ≥ 3 (someone has
to actually sell). Margin band 20-35%. Basic product certification required. Joint
target account list. Channel conflict rules of engagement signed.
Failure mode: granting reseller margin to a partner whose "customers" are actually your
inbound that they're routing through their paper. Test: of the named accounts they
sourced, how many had no prior relationship with you?
### OEM
White-label / embedded. Partner's brand on the front, your product underneath. Floor:
end_customer_relationships_pct ≥ 60%, dedicated_resources ≥ 2, certification complete.
Revshare 40-55% to compensate for the partner owning Tier-1 support and customer
relationship. Joint support runbook mandatory. End-customer NPS tracked.
Failure mode: granting OEM revshare without sufficient margin to fund your Tier-2+
support cost. If your support cost is $X per customer per year, and the OEM revshare
leaves you with less than $X net, the OEM deal is a losing trade no matter how big it
looks.
### SI_CONSULTING
Services attach. Partner sells their implementation services attached to your product.
Floor: partner_type = si_consultant, sales_team_size ≥ 5, end_customer_relationships_pct
≥ 50%. Product revshare 15-25%; services-side comp is independent.
Distinct from RESELLER: SI partners are selling THEIR services, you are pulled in. They
own customer relationship via the services scope. NEVER pay product revshare on
services-only "delivered" contribution — that's services-side compensation territory
(fixed fee or hourly).
### STRATEGIC
Multi-year co-investment. Named exec sponsors both sides. Reserved for partnerships that
genuinely change your distribution. Floors: named_accounts_sourced_count ≥ 5,
dedicated_resources ≥ 3, joint_marketing_spend ≥ $50k, multi-year commitment. Revshare
25-40% with pipeline floor + co-investment evidence.
Failure mode: "strategic" applied to any deal where the other side is big-logo and
nothing else. Big-logo without independent demand evidence is RESELLER or REFERRAL
wearing a logo. Real strategic partnerships are rare — most companies should have 0-3 at
most.
---
## Industry profile notes
- **SaaS**: floors as above
- **API**: bias toward technology/OEM partners (developer-first GTM); lower sales-team
floors because APIs sell themselves to developers, partners sell to procurement
- **Enterprise software**: higher SI floor (8 reps) because enterprise SI is a real
organization, not a one-person shop
- **Marketplace**: higher referral acceptance, lower reseller bar (marketplace dynamics
reward many small partners over few big ones)
- **Hardware**: higher OEM bar (4 dedicated resources) because hardware support
obligations are real costs
---
## Sources (≥ 7 authoritative references)
1. **Robert Caro** — *The Years of Lyndon Johnson* (especially the chapters on the LBJ
Senate-era patronage system) is the unintuitive but canonical reference on how
bilateral relationships convert into structural distribution. Distinct from
Caro's HP biography research (HP private archives), the LBJ work documents the
discipline of asking "what does this person actually deliver" vs. "what do they
claim to deliver" at scale — the same question a Head of BD asks of every
prospective partner. The HP work itself (commercial-channel post-mortems, 1990s
inkjet division) is referenced through second-party academic citations (see
Chintagunta 2009 below).
2. **Pradeep Chintagunta** — Joseph T. and Bernice S. Lewis Distinguished Service
Professor of Marketing, Chicago Booth. Academic foundation for channel economics
(e.g., Bronnenberg & Chintagunta on channel power in CPG distribution; the
underlying math applies directly to SaaS channel decisions). Key insight:
channel partnerships without volume floor break even on paper and lose money in
practice because fixed program cost is paid every quarter regardless of throughput.
3. **Joe Hessling** — Founder of 365 Retail Markets; speaker and operator on partner
programs. Published failure analyses of partner programs (industry talks +
PartnerHub presentations) identify "no independent demand" as the #1 root cause of
dead partner tiers — partners that joined for the discount, not the customers.
4. **Forrester Research** — *Channel Software Tech Stack* (annual report) and
Forrester partner-led research (Jay McBain era, ~2018-2021). Documents the
"partner-led-deals-from-your-own-pipeline" anti-pattern: 60%+ of "partner inquiries"
at early-stage SaaS are discount hunting, not channel investment.
5. **IDC** — *Worldwide Channel Software Tracker* and IDC partner research. Cost-to-serve
and partner-program economics benchmarks; multi-year longitudinal data on which
partner-program structures produce durable revenue.
6. **Tien Tzuo** — *Subscribed* (Portfolio, 2018), founder of Zuora. Channel chapter
covers subscription-channel revshare models, the shift from one-time-resale margin
to recurring revshare math, and the structural reason OEM partnerships require
different revshare floors than perpetual-license resale.
7. **Geoffrey Moore** — *Crossing the Chasm* (HarperBusiness, 1991/2014 revised) and
*Inside the Tornado* (HarperBusiness, 1995). Introduces the "whole product"
framework — the canonical lens for deciding whether a partnership is real (each
party delivers a component neither could deliver alone) vs. theatre (overlap with
no joint product).
8. **Microsoft Partner Network public playbooks** (MPN documentation, Microsoft Build
and Inspire content, 2018-2024) — operational templates for tier structure,
certification, and channel conflict rules of engagement. Source for the "named
account list + ROE before signing" discipline encoded in the joint GTM planner.
FILE:references/joint_gtm_canon.md
# Joint GTM Canon
Source material behind `joint_gtm_planner.py`'s tier-validated motion matrix and the
90-day milestone defaults. The discipline encoded here is: a partnership does not exist
until a joint pursuit closes a deal that neither side would have closed alone.
## Core principle
Joint GTM is not a marketing event. It is a sales motion that has to produce
attributable revenue against a named, written floor — within one sales cycle, or the
partnership is theatre.
The 90-day plan exists to manufacture decision-grade evidence: did this partner actually
move pipeline, or did we just throw a launch party? Without the structure, partnership
reviews degrade into "we like working with them" — which is a feeling, not a data point.
---
## The 4 sales motions
### pure_referral
Partner sends a lead. Your AE runs the entire sale. Partner gets a finder's fee on close.
No exclusivity, no MDF, no certification. Operates at REFERRAL tier; sometimes RESELLER
and SI_CONSULTING.
Anti-pattern: paying finder's fee on accounts already in your pipeline. The first job of
the program is attribution discipline — was this lead really new to us before the
partner sent it?
### co_sell
Partner and your AE jointly pursue the same account. Partner brings access; you bring
product. Both sides on calls, both forecasted. Revshare paid on close. Operates at
RESELLER, OEM, SI_CONSULTING, STRATEGIC tiers.
Anti-pattern: "co-sell" that is really "we let them watch" — partner attends meetings
but does not actively progress the deal. After 90 days, look at who advanced the deal
between stages. If your rep moved every stage, it was not co-sell — pay influenced rate,
not sourced.
### channel_led
Partner runs the full sales motion; you provide SE support and product. Partner
forecasts; you do not. Operates at RESELLER, OEM, STRATEGIC tiers — never REFERRAL or
SI_CONSULTING.
Anti-pattern: channel-led claimed but every demo requires your SE. If your SE is on
every customer call, the partner cannot sell the product solo — they are channel-led on
paper, co-sell in reality. Recertify or change the motion.
### white_label
Partner's brand on the front; you are invisible to the end customer. Partner owns
support, branding, customer relationship. Operates at OEM tier only. Requires
higher revshare to compensate for the loss of customer relationship.
Anti-pattern: white-label without margin sufficient to fund your Tier-2+ support cost.
If you are the de-facto product owner but only see 45% of the revenue, and your CTS
takes 30% of that, you are running a charity.
---
## The 90-day milestone structure
### Pre-launch (day -30 to 0)
Five non-negotiables: signed agreement; named exec sponsors both sides; jointly built
Target Account List (TAL) with conflict resolution per account; partner sales
certification; Rules of Engagement (ROE) signed before any joint pursuit. OEM and
STRATEGIC tiers add an integration QA + support runbook signoff.
### Launch (day 0 to 30)
Three measurable beats: first joint pursuit named within 7 days; 5 joint pursuits in
flight by day 15; first closed-won (or clear blocker isolation) by day 30. Channel-led
motions add a partner-led-demo-without-our-SE validation at day 20.
### Mid-quarter checkpoint (day 45)
Hard gates: pipeline-sourced ≥ 50% of 90-day floor; at least 1 closed-won OR named
blocker with owner + remediation date; certified rep count maintained; ROE working (zero
unresolved escalations); kill-criteria triggered? If yes, escalate to partnership
committee NOW.
### 90-day decision (day 90)
Decision-grade artifact: pipeline-sourced ≥ floor, deals-closed-won ≥ floor, win/loss
doc, certified rep count maintained, channel-conflict log clean. Outcome: continue /
re-tier / unwind, with named human accountable. No "let's see another quarter" — that's
how dead partnerships compound.
---
## Industry profile notes
- **SaaS**: 8x deal_avg_size as pipeline floor for RESELLER; 12x for STRATEGIC
- **API**: higher pipeline multiples (10x reseller, 15x strategic) — API deals are
smaller and higher-volume
- **Enterprise software**: lower deal-count floors but higher pipeline multiples
- **Marketplace**: highest pipeline multiples (12-18x) — partner volume is the whole
point
- **Hardware**: highest MDF defaults ($100k OEM, $200k STRATEGIC) — hardware partner
programs require physical inventory, demo equipment, certified field engineers
---
## Sources (≥ 7 authoritative references)
1. **Aaron Ross & Marylou Tyler** — *Predictable Revenue* (PebbleStorm, 2011). Source
for the cold-source vs. partner-source attribution distinction; the
"Cold Calling 2.0" framework's principle is that channel source is a different
pipeline economy than direct outbound — they cannot share metrics or comp plans.
2. **Winning by Design** — Jacco van der Kooij and team. SaaS sales methodology
incorporating partner-attached deals into the bow-tie funnel; the discipline of
tracking partner-attached vs. partner-sourced separately is canon here.
3. **Jay McBain** — Chief Analyst at Canalys (formerly Forrester); industry's leading
voice on co-sell discipline. Public writing (LinkedIn newsletter,
Channel-as-a-Service podcast 2019-2024) frames co-sell as "the most-misused word in
channel" — most "co-sell" is actually referral, and the difference matters for
revshare math.
4. **Microsoft Partner Network playbooks** (MPN public documentation; Microsoft Inspire
and Build sessions, 2018-2024). Operational source for tier structure, MCT/MCP
certification cadence, and the principle that channel-led motions require partner
certification + customer-facing partner-of-record designation BEFORE joint
pursuits begin.
5. **AWS Partner Network research** (APN public documentation; AWS re:Invent Partner
Day content, 2017-2024). Source for the consulting partner vs. technology partner
distinction, the competency-tier model, and the "partner-led" SI motion mechanics.
6. **SiriusDecisions** (now Forrester after 2018 acquisition) — partner-program research
and the SiriusDecisions Demand Waterfall framework. Source for the discipline of
tracking partner-sourced pipeline separately from partner-influenced, and the
benchmark that partner-influenced should pay at ~50% the revshare rate of
partner-sourced.
7. **Bridge Group SaaS Sales Benchmarks** (annual) — partner-attached deal benchmarks,
ramp times for partner reps vs. direct reps, and the data behind the "12 months
minimum to evaluate a partner program" heuristic encoded as a warning in the
joint_gtm_planner.
8. **Maria Pergolino & Aaron Ross** — *From Impossible to Inevitable* (Wiley, 2016).
Chapter on channel reproduces the discipline that partner programs without named
pipeline floors are decoration; the 8x-deal-avg-size pipeline floor convention for
RESELLER tier derives from this and SiriusDecisions data.
FILE:references/partnership_anti_patterns.md
# Partnership Anti-Patterns
The named failure modes encoded as warnings and validation errors across the three
scripts. Each anti-pattern below is sourced from real channel post-mortems and the
academic literature on channel economics. If your partnership program has any of these,
re-tier or unwind.
## Core principle
A bad partnership is more expensive than no partnership. The fixed program cost (MDF,
overhead, certification, joint marketing) is paid every quarter regardless of throughput.
A partner that produces sub-floor volume converts the program from "investment" into
"subsidy" — and subsidies are silent margin destroyers that compound across years.
The kill criteria embedded in every tier exist to make the unwind mechanical. The
moment a kill criterion triggers, the human review is "execute the contract" not
"renegotiate the relationship." The contract was the renegotiation; if you wait until
the criterion triggers to start the conversation, you have already lost the 6 months
you needed to source the replacement partner.
---
## The 8 anti-patterns (named and indexed)
### 1. "Partner = anyone who asked"
The default sin of inbound partnerships. A prospect emails "we should partner," the
account manager forwards to BD, BD forwards to legal, and 6 weeks later there is a
signed "partner agreement" with no commitments on either side.
Test: run the intake template honestly. If `named_accounts_sourced_count = 0` AND
`end_customer_relationships_pct < 30`, this is not a partner. Sign at REFERRAL tier
with auto-sunset, or do not sign.
Sources: Forrester partner research (60%+ of inbound partner inquiries at early-stage
SaaS lack independent demand); Joe Hessling partner-program failure analyses.
### 2. "White-label without margin enough to fund support"
OEM deal looks great on the deck. Net margin per deal looks great. Three months in, you
discover the OEM customer base is 4x the support volume of your direct customers
(because they don't know your product, and the OEM didn't actually train their CS team).
Test: model `our_cost_to_serve_via_partner_usd` honestly, including Tier-2+ support
load, escalation triage, custom-integration debugging, and post-incident reporting. If
the top of the revshare band produces negative per-deal margin, do not sign.
Sources: Hewlett-Packard channel post-mortems (1990s inkjet OEM cases); IBM channel-
conflict cases (post-PC-divestiture, late 1990s through 2005).
### 3. "Revshare for influenced-only deals at sourced rates"
Partner attends a few meetings, sends an intro email, accelerates a deal that was
already in motion. Their CRM logs it as "partner-sourced." Your CRM logs it as
"originated outbound rep X." The contract was ambiguous. The partner invoices at 30%
revshare on the full ARR.
Test: written attribution rules in the contract. "Sourced" requires partner to have
introduced AND owned the relationship through stage 2. "Influenced" pays at ≤ 50% of
sourced rate. Disputed attribution defaults to influenced.
Sources: SiriusDecisions partner research; Jay McBain on the "most-misused word in
channel."
### 4. "No kill criteria for under-performing partner"
The partnership has been declining for 4 quarters. The exec sponsor on the partner side
left 2 quarters ago. The certified reps were never replaced. Pipeline-sourced is at 20%
of the floor. But there is no clause in the contract specifying what happens — so the
program stays funded, the MDF gets paid, and the relationship dies slowly while you
keep writing checks.
Test: every tier has named kill criteria in the contract. RESELLER: <25% of target in
any quarter triggers 90-day cure. STRATEGIC: <70% of floor in 2 consecutive quarters
triggers joint exec review. The criteria are mechanical, not discretionary.
Sources: IBM channel-conflict case studies; MIT Sloan research on disproportionate
strategic-tier revshare paid to long-dead partnerships.
### 5. "Channel conflict ignored until reps quit"
Your top AE has been working an account for 8 months. The OEM partner signs the same
account through their channel motion. The deal closes — but to the partner. Your AE
gets nothing (no SPIFF, no attribution, no comp). Two weeks later, your top AE quits.
Test: Rules of Engagement (ROE) signed BEFORE any joint pursuit begins. Named-account
map. Conflict resolution at named human (Sales Director ↔ Partner Sales lead), not
committee. Documented escalation path. Channel-conflict log reviewed at every QBR.
Sources: Jay McBain (Canalys) — channel conflict is the #1 partner program killer;
written ROE published before partner signs prevents 80% of disputes.
### 6. "Exclusive territory granted to weak partner"
Partner asks for exclusive territory at signing — "we need protection to invest in
sales." You grant exclusive EMEA. Two quarters later, the partner has produced 1 deal.
Two more quarters: still 1. Meanwhile, three other partners are asking for EMEA. Your
contract prevents you from signing them. Three years later, you are stuck with a dead
partner in exclusive territory.
Test: exclusivity, if granted, is performance-conditioned. Volume floor by quarter;
miss the floor twice, exclusivity converts to non-exclusive. Never grant unconditional
exclusivity at signing.
Sources: Hewlett-Packard channel post-mortems (Indigo press division partnerships);
Pradeep Chintagunta on channel power dynamics.
### 7. "MDF without ROI accountability"
Quarter 1: $15k MDF sent. Quarter 2: $15k MDF sent. Quarter 3: $15k MDF sent. Quarter
4: no named pipeline attributable to MDF spend. Partner reports "we are building
brand awareness." You have spent $60k.
Test: every MDF disbursement tied to a named program (webinar, field event, content
piece) with named pipeline expectation. Quarterly true-up with attributable pipeline.
Sub-floor pipeline triggers MDF pause, not "let's give it more time."
Sources: Forrester channel research; AWS Partner Network MDF accountability framework
(public APN documentation).
### 8. "No offboarding plan when partnership ends"
The partnership has ended. Now: what happens to the joint customers? Where does the
customer data go? Who answers their support calls? Whose brand is on the renewal? Is
there a non-compete? Can the partner keep selling to the customers they sourced? The
answers are being negotiated in real time, under pressure, with lawyers on the phone.
Test: offboarding plan in the original contract. Data hand-back procedures, customer
continuity ownership, IP cleanup, brand take-down timeline, post-termination
non-compete (if any). Negotiate offboarding while the relationship is healthy.
Sources: IBM channel-conflict case studies; Salesforce AppExchange research on
partnership endings.
---
## Sources (≥ 7 authoritative references)
1. **Forrester Research** — *Channel Software Tech Stack* and partner-led research
(Jay McBain era). Documents the partner-led-deals-from-your-own-pipeline anti-
pattern and MDF accountability gaps in early-stage SaaS partner programs.
2. **Tom Tunguz** — Redpoint Ventures GP; channel-conflict and SaaS partner economics
writing (tomtunguz.com archives, 2014-2024). Source for the "channel conflict
trap" terminology and the data on rep attrition correlated with unresolved channel
conflict.
3. **Joe Hessling** — partner-program failure analyses (industry talks, PartnerHub
presentations). Source for the "no independent demand" failure mode and the
discipline of partnership intake-template honesty as a leading indicator.
4. **MIT Sloan Management Review** — articles on disproportionate revshare to
"strategic" partners (e.g., research on partnership ROI miscalibration, 2010-2020
archive). Quantifies the cost of strategic-tier programs that produce sub-tier
results.
5. **Hewlett-Packard channel post-mortems** — published case studies and academic
write-ups of the HP inkjet, Indigo, and EDS partial integration channel programs.
Source for anti-patterns 2, 6, and the data behind hardware-tier revshare floors.
6. **IBM channel-conflict case studies** (post-PC-divestiture era, 1990s-2005) — both
internal IBM publications and Harvard Business Review case treatments. Source for
anti-patterns 4 and 8 specifically — what happens when kill criteria and
offboarding are not in writing.
7. **Salesforce AppExchange research** — public AppExchange ISV partner research, 2015-
2024. Source for partnership-ending anti-patterns and the data on ISV partner
churn correlated with absent offboarding clauses.
8. **Pradeep Chintagunta** (Chicago Booth) — *Channel power, channel investment, and
partner economics* academic literature. Source for the principle that channel
partnerships without volume floor break even in theory and lose money in practice.
FILE:scripts/joint_gtm_planner.py
#!/usr/bin/env python3
"""joint_gtm_planner.py - Generate a 90-day joint GTM plan for a signed partner.
Stdlib-only. Deterministic. Validates that the sales_motion is compatible with the
partner_tier — refuses to plan channel-led GTM for a REFERRAL tier, refuses to plan
white-label for any tier other than OEM.
Output: 90-day plan with:
- Pre-launch milestones (days -30 to 0): training, certification, materials, target accounts
- Launch motion (days 0 to 30): MDF allocation, first deals, joint pursuit
- Mid-quarter checkpoint (day 45): named checkpoint criteria
- 90-day success criteria: pipeline-sourced floor, deals-closed floor, learnings doc
Industry profiles tune:
- target account count by tier
- MDF spend defaults by tier
- pipeline-sourced floor by tier (multiple of deal_avg_size)
Usage:
python joint_gtm_planner.py --sample
python joint_gtm_planner.py --input gtm.json --profile saas
python joint_gtm_planner.py --input gtm.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from typing import Any
SAMPLE_GTM = {
"partner_name": "Northstar Consulting",
"partner_tier": "SI_CONSULTING",
"target_segments": ["mid-market financial services EMEA", "regulated SaaS LATAM"],
"joint_value_proposition": (
"We bring the platform; Northstar brings 8 certified consultants who deliver "
"the regulated-vertical implementation in 60 days vs the 180 days customers "
"would spend doing it themselves."
),
"sales_motion": "co_sell",
"commitment_horizon_months": 12,
"deal_avg_size_usd": 90000,
}
VALID_TIERS = ("REFERRAL", "RESELLER", "OEM", "SI_CONSULTING", "STRATEGIC")
VALID_MOTIONS = ("pure_referral", "co_sell", "channel_led", "white_label")
# Hard compatibility matrix: which motions are allowed at which tier.
TIER_MOTION_MATRIX: dict[str, set[str]] = {
"REFERRAL": {"pure_referral"},
"RESELLER": {"pure_referral", "co_sell", "channel_led"},
"OEM": {"co_sell", "channel_led", "white_label"},
"SI_CONSULTING": {"pure_referral", "co_sell"},
"STRATEGIC": {"co_sell", "channel_led"},
}
PROFILES: dict[str, dict[str, Any]] = {
"saas": {
"target_accounts": {"REFERRAL": 5, "RESELLER": 15, "OEM": 10, "SI_CONSULTING": 12, "STRATEGIC": 20},
"mdf_default": {"REFERRAL": 0, "RESELLER": 15000, "OEM": 40000, "SI_CONSULTING": 20000, "STRATEGIC": 75000},
"pipeline_floor_multiple": {"REFERRAL": 3, "RESELLER": 8, "OEM": 6, "SI_CONSULTING": 6, "STRATEGIC": 12},
"deals_closed_floor": {"REFERRAL": 1, "RESELLER": 3, "OEM": 2, "SI_CONSULTING": 2, "STRATEGIC": 4},
},
"api": {
"target_accounts": {"REFERRAL": 8, "RESELLER": 20, "OEM": 8, "SI_CONSULTING": 10, "STRATEGIC": 15},
"mdf_default": {"REFERRAL": 0, "RESELLER": 10000, "OEM": 30000, "SI_CONSULTING": 15000, "STRATEGIC": 60000},
"pipeline_floor_multiple": {"REFERRAL": 4, "RESELLER": 10, "OEM": 8, "SI_CONSULTING": 6, "STRATEGIC": 15},
"deals_closed_floor": {"REFERRAL": 1, "RESELLER": 4, "OEM": 2, "SI_CONSULTING": 2, "STRATEGIC": 5},
},
"enterprise-software": {
"target_accounts": {"REFERRAL": 3, "RESELLER": 8, "OEM": 6, "SI_CONSULTING": 10, "STRATEGIC": 15},
"mdf_default": {"REFERRAL": 0, "RESELLER": 30000, "OEM": 75000, "SI_CONSULTING": 40000, "STRATEGIC": 150000},
"pipeline_floor_multiple": {"REFERRAL": 2, "RESELLER": 5, "OEM": 4, "SI_CONSULTING": 5, "STRATEGIC": 8},
"deals_closed_floor": {"REFERRAL": 1, "RESELLER": 2, "OEM": 1, "SI_CONSULTING": 2, "STRATEGIC": 3},
},
"marketplace": {
"target_accounts": {"REFERRAL": 10, "RESELLER": 25, "OEM": 12, "SI_CONSULTING": 15, "STRATEGIC": 25},
"mdf_default": {"REFERRAL": 0, "RESELLER": 10000, "OEM": 25000, "SI_CONSULTING": 15000, "STRATEGIC": 50000},
"pipeline_floor_multiple": {"REFERRAL": 5, "RESELLER": 12, "OEM": 8, "SI_CONSULTING": 8, "STRATEGIC": 18},
"deals_closed_floor": {"REFERRAL": 2, "RESELLER": 5, "OEM": 3, "SI_CONSULTING": 3, "STRATEGIC": 6},
},
"hardware": {
"target_accounts": {"REFERRAL": 5, "RESELLER": 10, "OEM": 8, "SI_CONSULTING": 8, "STRATEGIC": 12},
"mdf_default": {"REFERRAL": 0, "RESELLER": 25000, "OEM": 100000, "SI_CONSULTING": 30000, "STRATEGIC": 200000},
"pipeline_floor_multiple": {"REFERRAL": 2, "RESELLER": 6, "OEM": 5, "SI_CONSULTING": 4, "STRATEGIC": 10},
"deals_closed_floor": {"REFERRAL": 1, "RESELLER": 2, "OEM": 1, "SI_CONSULTING": 2, "STRATEGIC": 3},
},
}
@dataclass
class Milestone:
day: int
name: str
owner: str
deliverable: str
@dataclass
class GtmPlan:
partner_name: str
profile: str
partner_tier: str
sales_motion: str
target_segments: list[str]
joint_value_proposition: str
pre_launch: list[Milestone] = field(default_factory=list)
launch: list[Milestone] = field(default_factory=list)
mid_quarter_checkpoint: list[str] = field(default_factory=list)
success_criteria_90d: list[str] = field(default_factory=list)
mdf_allocation_usd: float = 0.0
target_account_count: int = 0
pipeline_floor_usd: float = 0.0
deals_closed_floor: int = 0
validation_errors: list[str] = field(default_factory=list)
warnings: list[str] = field(default_factory=list)
def _validate(gtm: dict) -> list[str]:
errs: list[str] = []
tier = (gtm.get("partner_tier") or "").upper()
motion = (gtm.get("sales_motion") or "").lower()
if tier not in VALID_TIERS:
errs.append(f"partner_tier '{tier}' not in {VALID_TIERS}")
return errs
if motion not in VALID_MOTIONS:
errs.append(f"sales_motion '{motion}' not in {VALID_MOTIONS}")
return errs
if motion not in TIER_MOTION_MATRIX[tier]:
errs.append(
f"sales_motion '{motion}' is not compatible with tier '{tier}'. "
f"Allowed motions for {tier}: {sorted(TIER_MOTION_MATRIX[tier])}. "
f"If you need '{motion}', re-tier the partner via partner_tier_classifier.py first."
)
if not gtm.get("joint_value_proposition"):
errs.append("joint_value_proposition is required (one sentence with end-customer)")
if not gtm.get("target_segments"):
errs.append("target_segments is required (named segments, not 'everyone')")
return errs
def _pre_launch_milestones(tier: str, motion: str) -> list[Milestone]:
base = [
Milestone(-30, "Mutual NDA + Partner Agreement signed",
"BD lead + Legal both sides", "Signed PDFs"),
Milestone(-25, "Joint kickoff call: name exec sponsors",
"BD lead + Partner GM", "Sponsor pair documented"),
Milestone(-20, "Target Account List (TAL) jointly built",
"Sales Director + Partner Sales lead",
"Named-account list with conflict resolution per-account"),
Milestone(-15, "Partner sales training (week 1 of 2)",
"Sales Enablement", "Training attendance log"),
Milestone(-10, "Partner sales training (week 2 of 2) + certification",
"Sales Enablement + Partner reps",
"Named certified reps per partner"),
]
if tier in ("OEM", "STRATEGIC"):
base.append(Milestone(
-7, "Integration QA + support runbook signoff",
"Engineering + Support both sides",
"Joint Tier-1/Tier-2 support runbook"))
if motion == "white_label":
base.append(Milestone(
-5, "Brand-use guide + co-branded asset pack approved",
"Marketing + Legal", "Asset pack + brand-use rules"))
if motion in ("co_sell", "channel_led"):
base.append(Milestone(
-3, "Rules of Engagement signed (channel conflict)",
"Sales Director + Partner Sales lead",
"Signed ROE with named-account map"))
base.append(Milestone(
0, "Joint launch announcement + first pursuit kickoff",
"Marketing + Sales both sides",
"Press / blog / customer-facing materials"))
return base
def _launch_milestones(tier: str, motion: str) -> list[Milestone]:
base = [
Milestone(7, "First joint pursuit named (single account)",
"Sales Director + Partner Sales lead",
"Account brief + close plan"),
Milestone(15, "5 joint pursuits in flight",
"Sales Director + Partner Sales lead",
"Pipeline-sourced report"),
Milestone(30, "First closed-won OR clear blocker isolation",
"Sales Director + Partner Sales lead",
"Win/Loss writeup"),
]
if motion == "channel_led":
base.append(Milestone(
20, "Partner-led demo without our SE present (validation)",
"Partner Sales lead", "Recording + scorecard"))
if tier == "OEM":
base.append(Milestone(
25, "First end-customer Tier-2 support ticket through joint runbook",
"Support both sides", "Ticket resolution timeline"))
return base
def _mid_quarter_checkpoint(tier: str, motion: str, pipeline_floor: float) -> list[str]:
return [
f"Day 45: pipeline sourced ≥ 50% of 90-day floor (,.0f)",
f"Day 45: at least 1 closed-won OR named blocker with owner + remediation date",
f"Day 45: certified rep count maintained at signed-agreement level",
f"Day 45: ROE working — no channel-conflict escalations OR all resolved at named-human level",
f"Day 45: kill-criteria trigger review — if any, escalate to partnership committee NOW",
]
def _success_criteria(tier: str, motion: str, pipeline_floor: float, deals_floor: int) -> list[str]:
base = [
f"Pipeline-sourced through partner ≥ ,.0f (validated by both sides)",
f"Deals closed-won through partner ≥ {deals_floor}",
f"Joint win/loss doc covering all material deals (closed-won AND closed-lost)",
f"Certified rep count ≥ partner-agreement level",
f"Channel-conflict log: zero unresolved escalations",
]
if tier == "OEM":
base.append("End-customer NPS via the OEM at or above corporate floor")
base.append("Support SLA breach rate ≤ 5% of tickets")
if tier == "STRATEGIC":
base.append("Exec sponsor pair active (both sides — verify before quarter close)")
base.append("Executive QBR completed with signed-off next-quarter pipeline floor")
if motion == "channel_led":
base.append("≥ 50% of closed-won were partner-led (not just partner-influenced)")
if motion == "white_label":
base.append("Embedded volume hit minimum threshold; no brand-bleed incidents")
base.append("Decision: continue / re-tier / unwind, with named human accountable")
return base
def plan_gtm(gtm: dict, profile_name: str = "saas") -> GtmPlan:
if profile_name not in PROFILES:
raise ValueError(f"Unknown profile '{profile_name}'. Choose from {list(PROFILES)}.")
profile = PROFILES[profile_name]
errs = _validate(gtm)
plan = GtmPlan(
partner_name=str(gtm.get("partner_name", "UNSPECIFIED")),
profile=profile_name,
partner_tier=(gtm.get("partner_tier") or "").upper(),
sales_motion=(gtm.get("sales_motion") or "").lower(),
target_segments=list(gtm.get("target_segments", []) or []),
joint_value_proposition=str(gtm.get("joint_value_proposition", "")),
validation_errors=errs,
)
if errs:
return plan
tier = plan.partner_tier
motion = plan.sales_motion
deal_avg = float(gtm.get("deal_avg_size_usd", 0.0))
plan.target_account_count = profile["target_accounts"].get(tier, 0)
plan.mdf_allocation_usd = float(profile["mdf_default"].get(tier, 0))
pipe_multiple = profile["pipeline_floor_multiple"].get(tier, 0)
plan.pipeline_floor_usd = deal_avg * pipe_multiple
plan.deals_closed_floor = profile["deals_closed_floor"].get(tier, 0)
plan.pre_launch = _pre_launch_milestones(tier, motion)
plan.launch = _launch_milestones(tier, motion)
plan.mid_quarter_checkpoint = _mid_quarter_checkpoint(tier, motion, plan.pipeline_floor_usd)
plan.success_criteria_90d = _success_criteria(
tier, motion, plan.pipeline_floor_usd, plan.deals_closed_floor
)
horizon = int(gtm.get("commitment_horizon_months", 12) or 12)
if horizon < 12:
plan.warnings.append(
f"commitment_horizon_months={horizon} < 12: partner programs rarely produce "
"signal in less than a full sales cycle. Consider extending or downgrading tier."
)
if tier == "STRATEGIC" and horizon < 24:
plan.warnings.append(
"STRATEGIC tier with sub-24-month horizon is structurally inconsistent; "
"either commit multi-year or re-tier."
)
if deal_avg <= 0:
plan.warnings.append(
"deal_avg_size_usd not provided or 0 — pipeline floor cannot be computed. "
"Re-run with a real number from your closed-won data."
)
return plan
def _render_human(p: GtmPlan) -> str:
lines = []
lines.append(f"Joint GTM Plan: {p.partner_name}")
lines.append(f"Profile: {p.profile} ; Tier: {p.partner_tier} ; Motion: {p.sales_motion}")
if p.validation_errors:
lines.append("")
lines.append("VALIDATION ERRORS (plan not generated):")
for e in p.validation_errors:
lines.append(f" ! {e}")
return "\n".join(lines)
lines.append("")
lines.append(f"Target segments: {'; '.join(p.target_segments)}")
lines.append(f"Joint value prop: {p.joint_value_proposition}")
lines.append("")
lines.append(f"MDF allocation: ,.0f")
lines.append(f"Target accounts: {p.target_account_count}")
lines.append(f"90-day pipeline floor: ,.0f")
lines.append(f"90-day closed-won floor: {p.deals_closed_floor}")
lines.append("")
lines.append("Pre-launch milestones (day -30 to 0):")
for m in p.pre_launch:
lines.append(f" Day {m.day:+4d} {m.name}")
lines.append(f" owner: {m.owner}")
lines.append(f" deliverable: {m.deliverable}")
lines.append("")
lines.append("Launch milestones (day 0 to 30):")
for m in p.launch:
lines.append(f" Day {m.day:+4d} {m.name}")
lines.append(f" owner: {m.owner}")
lines.append(f" deliverable: {m.deliverable}")
lines.append("")
lines.append("Mid-quarter checkpoint (day 45):")
for c in p.mid_quarter_checkpoint:
lines.append(f" - {c}")
lines.append("")
lines.append("90-day success criteria:")
for s in p.success_criteria_90d:
lines.append(f" - {s}")
if p.warnings:
lines.append("")
lines.append("Warnings:")
for w in p.warnings:
lines.append(f" ! {w}")
return "\n".join(lines)
def _to_jsonable(p: GtmPlan) -> dict:
return asdict(p)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Generate a 90-day joint GTM plan for a signed partner.",
)
parser.add_argument("--input", help="Path to JSON GTM context")
parser.add_argument("--profile", default="saas", choices=list(PROFILES))
parser.add_argument("--output", default="human", choices=["human", "json", "markdown"])
parser.add_argument("--sample", action="store_true", help="Use embedded sample GTM context")
args = parser.parse_args(argv)
if args.sample or not args.input:
gtm = SAMPLE_GTM
else:
with open(args.input) as f:
gtm = json.load(f)
plan = plan_gtm(gtm, args.profile)
if args.output == "json":
print(json.dumps(_to_jsonable(plan), indent=2))
else:
if args.output == "markdown":
print("# Joint GTM Plan\n")
print(_render_human(plan))
return 0 if not plan.validation_errors else 0
# Note: validation errors print to stdout; exit 0 so pipelines can capture them.
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/partner_tier_classifier.py
#!/usr/bin/env python3
"""partner_tier_classifier.py - Classify a prospective partner into 1 of 5 tiers.
Stdlib-only. Deterministic logic with hard floors per tier. NEVER auto-signs anything;
output is a tier verdict + rationale + kill criteria, routed to a human committee.
The 5 tiers:
REFERRAL - informal intro, no joint commitment, small finder's fee
RESELLER - transactional resale with margin, basic certification
OEM - white-label / embedded, integration + support commitment
SI_CONSULTING - services attach, customer-owned-by-partner
STRATEGIC - multi-year, co-investment, dedicated resources both sides
Tier floors (hard requirements — failing a floor caps the tier):
REFERRAL - none (default fallback)
RESELLER - end_customer_relationships_pct >= 40 AND sales_team_size >= 3
OEM - end_customer_relationships_pct >= 60 AND certification_completion AND
commitments.dedicated_resources >= 2
SI_CONSULTING - end_customer_relationships_pct >= 50 AND sales_team_size >= 5 AND
partner_type in {si_consultant}
STRATEGIC - named_accounts_sourced_count >= 5 AND multi-year-commit (>=24mo
horizon implied by commitments) AND dedicated_resources >= 3 AND
joint_marketing_spend >= 50000
Industry profiles (`--profile`) tune the thresholds:
saas - default; the floors above
api - bias toward technology/OEM; relax sales_team for OEM
enterprise-software - higher SI bar (8 sales reps)
marketplace - higher referral acceptance, lower reseller bar
hardware - higher OEM bar (dedicated_resources 4)
Usage:
python partner_tier_classifier.py --sample
python partner_tier_classifier.py --input partner.json --profile saas
python partner_tier_classifier.py --input partner.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from typing import Any
SAMPLE_PARTNER = {
"partner_name": "Northstar Consulting",
"partner_type": "si_consultant",
"independent_demand_evidence": {
"named_accounts_sourced_count": 7,
"end_customer_relationships_pct": 65,
"sales_team_size": 8,
},
"strategic_value": {
"geo_coverage": "EMEA + LATAM",
"product_complement": "implementation services for our platform",
"brand_lift": "mid",
"channel_economics_advantage": "lower CAC in regulated verticals",
},
"commitments": {
"joint_marketing_spend": 75000,
"dedicated_resources": 4,
"certification_completion": True,
"sales_targets": "12 deals in 12 months",
},
}
VALID_PARTNER_TYPES = (
"referral", "reseller", "oem", "si_consultant", "technology", "strategic_alliance",
)
PROFILES: dict[str, dict[str, Any]] = {
"saas": {
"reseller_floor_ecr": 40,
"reseller_floor_sales_team": 3,
"oem_floor_ecr": 60,
"oem_floor_dedicated": 2,
"si_floor_ecr": 50,
"si_floor_sales_team": 5,
"strategic_floor_sourced": 5,
"strategic_floor_dedicated": 3,
"strategic_floor_mdf": 50000,
},
"api": {
"reseller_floor_ecr": 35,
"reseller_floor_sales_team": 2,
"oem_floor_ecr": 50,
"oem_floor_dedicated": 2,
"si_floor_ecr": 50,
"si_floor_sales_team": 5,
"strategic_floor_sourced": 4,
"strategic_floor_dedicated": 3,
"strategic_floor_mdf": 40000,
},
"enterprise-software": {
"reseller_floor_ecr": 50,
"reseller_floor_sales_team": 5,
"oem_floor_ecr": 65,
"oem_floor_dedicated": 3,
"si_floor_ecr": 55,
"si_floor_sales_team": 8,
"strategic_floor_sourced": 6,
"strategic_floor_dedicated": 4,
"strategic_floor_mdf": 75000,
},
"marketplace": {
"reseller_floor_ecr": 30,
"reseller_floor_sales_team": 2,
"oem_floor_ecr": 50,
"oem_floor_dedicated": 2,
"si_floor_ecr": 45,
"si_floor_sales_team": 4,
"strategic_floor_sourced": 4,
"strategic_floor_dedicated": 2,
"strategic_floor_mdf": 30000,
},
"hardware": {
"reseller_floor_ecr": 45,
"reseller_floor_sales_team": 4,
"oem_floor_ecr": 70,
"oem_floor_dedicated": 4,
"si_floor_ecr": 55,
"si_floor_sales_team": 6,
"strategic_floor_sourced": 6,
"strategic_floor_dedicated": 4,
"strategic_floor_mdf": 100000,
},
}
# Kill criteria templates per tier — these are placed into the partnership contract
# so the unwind is mechanical, not a 2-year legal fight.
KILL_CRITERIA: dict[str, list[str]] = {
"REFERRAL": [
"No qualified intros in 2 consecutive quarters -> auto-sunset, no notice",
"Any misrepresentation of relationship as 'partner' externally -> immediate termination",
],
"RESELLER": [
"Less than 25% of agreed annual sales target hit in any quarter -> 90-day cure",
"Two consecutive quarters under cure -> tier demotion to REFERRAL or termination",
"Certification lapses for >60 days -> resale rights suspended",
],
"OEM": [
"End-customer NPS via the OEM falls below corporate floor -> joint remediation plan",
"Less than 50% of agreed embedded volume in 2 consecutive quarters -> cure or unwind",
"Support response SLA breach >5% per quarter -> co-funded customer-success review",
],
"SI_CONSULTING": [
"Less than 60% of agreed certified resources maintained -> 60-day cure",
"Customer-attributed delivery failures above named threshold -> joint root-cause + remediation",
"Loss of practice lead (named person) -> 90-day re-qualification of tier",
],
"STRATEGIC": [
"Less than 70% of named pipeline floor in 2 consecutive quarters -> joint exec review",
"Failure of the named exec sponsor on either side -> 90-day re-validation of alliance",
"Material change of control on either side -> automatic 6-month evaluation period",
"Loss of integration / technical interop for >30 days -> alliance pause",
],
}
@dataclass
class TierScore:
tier: str
raw_score: float
floors_passed: bool
floors_failed: list[str]
rationale: str
@dataclass
class ClassificationVerdict:
partner_name: str
profile: str
tier_assigned: str
composite_rationale: str
floors_failed_for_higher_tiers: list[str] = field(default_factory=list)
tier_scores: list[TierScore] = field(default_factory=list)
kill_criteria: list[str] = field(default_factory=list)
next_steps: list[str] = field(default_factory=list)
warnings: list[str] = field(default_factory=list)
def _clamp(x: float, lo: float = 0.0, hi: float = 100.0) -> float:
return max(lo, min(hi, x))
def _check_reseller_floors(partner: dict, profile: dict) -> tuple[bool, list[str]]:
ide = partner.get("independent_demand_evidence", {}) or {}
ecr = float(ide.get("end_customer_relationships_pct", 0))
sts = int(ide.get("sales_team_size", 0))
fails: list[str] = []
if ecr < profile["reseller_floor_ecr"]:
fails.append(
f"RESELLER floor: end_customer_relationships_pct {ecr:.0f}% < "
f"{profile['reseller_floor_ecr']}%"
)
if sts < profile["reseller_floor_sales_team"]:
fails.append(
f"RESELLER floor: sales_team_size {sts} < "
f"{profile['reseller_floor_sales_team']}"
)
return (len(fails) == 0, fails)
def _check_oem_floors(partner: dict, profile: dict) -> tuple[bool, list[str]]:
ide = partner.get("independent_demand_evidence", {}) or {}
com = partner.get("commitments", {}) or {}
ecr = float(ide.get("end_customer_relationships_pct", 0))
dr = int(com.get("dedicated_resources", 0))
cert = bool(com.get("certification_completion", False))
fails: list[str] = []
if ecr < profile["oem_floor_ecr"]:
fails.append(
f"OEM floor: end_customer_relationships_pct {ecr:.0f}% < "
f"{profile['oem_floor_ecr']}%"
)
if dr < profile["oem_floor_dedicated"]:
fails.append(
f"OEM floor: dedicated_resources {dr} < {profile['oem_floor_dedicated']}"
)
if not cert:
fails.append("OEM floor: certification_completion is False")
return (len(fails) == 0, fails)
def _check_si_floors(partner: dict, profile: dict) -> tuple[bool, list[str]]:
ide = partner.get("independent_demand_evidence", {}) or {}
ecr = float(ide.get("end_customer_relationships_pct", 0))
sts = int(ide.get("sales_team_size", 0))
ptype = (partner.get("partner_type") or "").lower()
fails: list[str] = []
if ecr < profile["si_floor_ecr"]:
fails.append(
f"SI_CONSULTING floor: end_customer_relationships_pct {ecr:.0f}% < "
f"{profile['si_floor_ecr']}%"
)
if sts < profile["si_floor_sales_team"]:
fails.append(
f"SI_CONSULTING floor: sales_team_size {sts} < "
f"{profile['si_floor_sales_team']}"
)
if ptype != "si_consultant":
fails.append(
f"SI_CONSULTING floor: partner_type is '{ptype}', expected 'si_consultant'"
)
return (len(fails) == 0, fails)
def _check_strategic_floors(partner: dict, profile: dict) -> tuple[bool, list[str]]:
ide = partner.get("independent_demand_evidence", {}) or {}
com = partner.get("commitments", {}) or {}
sourced = int(ide.get("named_accounts_sourced_count", 0))
dr = int(com.get("dedicated_resources", 0))
mdf = float(com.get("joint_marketing_spend", 0))
targets = (com.get("sales_targets") or "").lower()
fails: list[str] = []
if sourced < profile["strategic_floor_sourced"]:
fails.append(
f"STRATEGIC floor: named_accounts_sourced_count {sourced} < "
f"{profile['strategic_floor_sourced']}"
)
if dr < profile["strategic_floor_dedicated"]:
fails.append(
f"STRATEGIC floor: dedicated_resources {dr} < "
f"{profile['strategic_floor_dedicated']}"
)
if mdf < profile["strategic_floor_mdf"]:
fails.append(
f"STRATEGIC floor: joint_marketing_spend {mdf:.0f} < "
f"{profile['strategic_floor_mdf']}"
)
# Multi-year heuristic: sales_targets mentions "12 months" or longer; or commitment
# text references multi-year / 24 / 36 months.
multi_year_signal = any(
s in targets for s in ("12 months", "24 months", "36 months", "multi-year", "multi year")
)
if not multi_year_signal:
fails.append(
"STRATEGIC floor: no multi-year commitment signal in sales_targets text"
)
return (len(fails) == 0, fails)
def _strategic_raw_score(partner: dict, profile: dict) -> float:
ide = partner.get("independent_demand_evidence", {}) or {}
com = partner.get("commitments", {}) or {}
sv = partner.get("strategic_value", {}) or {}
score = 0.0
# Sourced accounts: 0..40 points (cap at 10 sourced)
score += min(40.0, ide.get("named_accounts_sourced_count", 0) * 4.0)
# Dedicated resources: 0..20 points (cap at 5)
score += min(20.0, com.get("dedicated_resources", 0) * 4.0)
# MDF: 0..20 points (cap at 100k)
score += min(20.0, com.get("joint_marketing_spend", 0) / 5000.0)
# Strategic value flags: 5 points each
for key in ("geo_coverage", "product_complement", "brand_lift", "channel_economics_advantage"):
if sv.get(key):
score += 5.0
return _clamp(score)
def _oem_raw_score(partner: dict, profile: dict) -> float:
ide = partner.get("independent_demand_evidence", {}) or {}
com = partner.get("commitments", {}) or {}
score = 0.0
score += min(40.0, ide.get("end_customer_relationships_pct", 0) * 0.6)
score += min(20.0, com.get("dedicated_resources", 0) * 5.0)
score += 20.0 if com.get("certification_completion") else 0.0
score += min(20.0, ide.get("named_accounts_sourced_count", 0) * 3.0)
return _clamp(score)
def _si_raw_score(partner: dict, profile: dict) -> float:
ide = partner.get("independent_demand_evidence", {}) or {}
com = partner.get("commitments", {}) or {}
score = 0.0
score += min(35.0, ide.get("end_customer_relationships_pct", 0) * 0.5)
score += min(25.0, ide.get("sales_team_size", 0) * 3.0)
score += 15.0 if com.get("certification_completion") else 0.0
score += min(25.0, com.get("dedicated_resources", 0) * 5.0)
return _clamp(score)
def _reseller_raw_score(partner: dict, profile: dict) -> float:
ide = partner.get("independent_demand_evidence", {}) or {}
score = 0.0
score += min(50.0, ide.get("end_customer_relationships_pct", 0) * 0.8)
score += min(40.0, ide.get("sales_team_size", 0) * 5.0)
score += min(10.0, ide.get("named_accounts_sourced_count", 0) * 2.0)
return _clamp(score)
def _referral_raw_score(partner: dict, profile: dict) -> float:
# Referral always passes; raw score is just "do they have any evidence of intent"
ide = partner.get("independent_demand_evidence", {}) or {}
score = 30.0 # baseline for showing up
score += min(40.0, ide.get("named_accounts_sourced_count", 0) * 6.0)
score += min(30.0, ide.get("end_customer_relationships_pct", 0) * 0.3)
return _clamp(score)
def classify(partner: dict, profile_name: str = "saas") -> ClassificationVerdict:
if profile_name not in PROFILES:
raise ValueError(f"Unknown profile '{profile_name}'. Choose from {list(PROFILES)}.")
profile = PROFILES[profile_name]
ptype = (partner.get("partner_type") or "").lower()
warnings: list[str] = []
if ptype not in VALID_PARTNER_TYPES:
warnings.append(
f"partner_type '{ptype}' is not one of {VALID_PARTNER_TYPES}; "
"classification continues but verify input"
)
# Compute raw scores for each tier
s_strategic = _strategic_raw_score(partner, profile)
s_oem = _oem_raw_score(partner, profile)
s_si = _si_raw_score(partner, profile)
s_reseller = _reseller_raw_score(partner, profile)
s_referral = _referral_raw_score(partner, profile)
# Floor checks
p_reseller, f_reseller = _check_reseller_floors(partner, profile)
p_oem, f_oem = _check_oem_floors(partner, profile)
p_si, f_si = _check_si_floors(partner, profile)
p_strategic, f_strategic = _check_strategic_floors(partner, profile)
scores = [
TierScore("STRATEGIC", s_strategic, p_strategic, f_strategic,
f"raw={s_strategic:.1f}/100 ; floors={'PASS' if p_strategic else 'FAIL'}"),
TierScore("OEM", s_oem, p_oem, f_oem,
f"raw={s_oem:.1f}/100 ; floors={'PASS' if p_oem else 'FAIL'}"),
TierScore("SI_CONSULTING", s_si, p_si, f_si,
f"raw={s_si:.1f}/100 ; floors={'PASS' if p_si else 'FAIL'}"),
TierScore("RESELLER", s_reseller, p_reseller, f_reseller,
f"raw={s_reseller:.1f}/100 ; floors={'PASS' if p_reseller else 'FAIL'}"),
TierScore("REFERRAL", s_referral, True, [],
f"raw={s_referral:.1f}/100 ; floors=PASS (default)"),
]
# Assign highest tier that PASSES floors AND has raw_score >= 60.
assigned = "REFERRAL"
rationale = "Default fallback tier (no higher floors passed)"
floors_blocking: list[str] = []
tier_order = ["STRATEGIC", "OEM", "SI_CONSULTING", "RESELLER", "REFERRAL"]
for tier_name in tier_order:
ts = next(t for t in scores if t.tier == tier_name)
if ts.floors_passed and (tier_name == "REFERRAL" or ts.raw_score >= 60.0):
assigned = tier_name
rationale = (
f"Assigned {tier_name}: raw score {ts.raw_score:.1f}/100, all floors passed."
)
break
if not ts.floors_passed:
floors_blocking.extend(ts.floors_failed)
elif ts.raw_score < 60.0:
floors_blocking.append(
f"{tier_name}: raw score {ts.raw_score:.1f}/100 below 60 minimum"
)
# Next steps depend on tier
next_steps_map = {
"REFERRAL": [
"Document the referral arrangement (one-page MOU, no exclusivity)",
"Define finder's fee % (typical 5-10% of first-year ARR)",
"Set 2-quarter review with auto-sunset trigger",
],
"RESELLER": [
"Run scripts/joint_gtm_planner.py with sales_motion='co_sell' or 'channel_led'",
"Run scripts/revshare_modeler.py to size resale margin (typical 20-35%)",
"Draft certification curriculum and timeline",
"Lock kill criteria in contract before signing",
],
"OEM": [
"Run scripts/joint_gtm_planner.py with sales_motion='white_label'",
"Run scripts/revshare_modeler.py with deeper revshare band (typical 40-55%)",
"Validate support model — who answers Tier-2 calls?",
"Lock IP / integration / brand-use terms in contract",
"Lock kill criteria including support SLA in contract",
],
"SI_CONSULTING": [
"Run scripts/joint_gtm_planner.py with sales_motion='co_sell'",
"Run scripts/revshare_modeler.py (typical 15-25% on product, 0% on services)",
"Define certified-practice-lead role and named individual",
"Lock kill criteria around certification headcount",
],
"STRATEGIC": [
"Verify with `c-level-advisor/ma-playbook` whether this should be acquisition not partnership",
"Run scripts/joint_gtm_planner.py with sales_motion='channel_led' or 'co_sell'",
"Run scripts/revshare_modeler.py with strategic-tier band",
"Negotiate executive-sponsor pairing (named individual each side)",
"Lock kill criteria including exec-sponsor-departure trigger",
],
}
return ClassificationVerdict(
partner_name=str(partner.get("partner_name", "UNSPECIFIED")),
profile=profile_name,
tier_assigned=assigned,
composite_rationale=rationale,
floors_failed_for_higher_tiers=floors_blocking,
tier_scores=scores,
kill_criteria=KILL_CRITERIA.get(assigned, []),
next_steps=next_steps_map.get(assigned, []),
warnings=warnings,
)
def _render_human(v: ClassificationVerdict) -> str:
lines = []
lines.append(f"Partner Classification: {v.partner_name}")
lines.append(f"Profile: {v.profile}")
lines.append(f"Tier Assigned: {v.tier_assigned}")
lines.append("")
lines.append(v.composite_rationale)
lines.append("")
lines.append("Tier scoring detail (high to low):")
for ts in v.tier_scores:
floor_status = "PASS" if ts.floors_passed else "FAIL"
lines.append(f" - {ts.tier:14s} raw={ts.raw_score:5.1f}/100 floors={floor_status}")
if ts.floors_failed:
for f in ts.floors_failed:
lines.append(f" x {f}")
lines.append("")
if v.floors_failed_for_higher_tiers:
lines.append("Why not a higher tier:")
for f in v.floors_failed_for_higher_tiers:
lines.append(f" - {f}")
lines.append("")
lines.append("Kill criteria (put these in the contract):")
for k in v.kill_criteria:
lines.append(f" - {k}")
lines.append("")
lines.append("Next steps:")
for n in v.next_steps:
lines.append(f" - {n}")
if v.warnings:
lines.append("")
lines.append("Warnings:")
for w in v.warnings:
lines.append(f" ! {w}")
return "\n".join(lines)
def _to_jsonable(v: ClassificationVerdict) -> dict:
return asdict(v)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Classify a prospective partner into REFERRAL / RESELLER / OEM / SI_CONSULTING / STRATEGIC tier.",
)
parser.add_argument("--input", help="Path to JSON partner intake")
parser.add_argument("--profile", default="saas", choices=list(PROFILES))
parser.add_argument("--output", default="human", choices=["human", "json", "markdown"])
parser.add_argument("--sample", action="store_true", help="Use embedded sample partner")
args = parser.parse_args(argv)
if args.sample or not args.input:
partner = SAMPLE_PARTNER
else:
with open(args.input) as f:
partner = json.load(f)
verdict = classify(partner, args.profile)
if args.output == "json":
print(json.dumps(_to_jsonable(verdict), indent=2))
else:
# human and markdown share the same body; markdown adds a header
if args.output == "markdown":
print(f"# Partner Tier Classification\n")
print(_render_human(verdict))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/revshare_modeler.py
#!/usr/bin/env python3
"""revshare_modeler.py - Model revshare economics: direct vs via partner.
Stdlib-only. Deterministic. Computes:
1. Margin per deal direct vs via partner (with named cost-to-serve inputs)
2. Recommended revshare % band by tier + partner contribution depth
(sourced > influenced > delivered)
3. Break-even partner ROI — how many partner-sourced deals to cover MDF + program cost
4. Long-term economics: at projected scale, when does partner economics beat direct?
Revshare bands by tier (industry-typical, can be tuned by --profile):
REFERRAL : 5-10% on first-year ARR (one-time finder's fee)
RESELLER : 20-35% on net ARR (recurring while customer active)
OEM : 40-55% on net ARR (revshare reflects partner-owned support)
SI_CONSULTING : 15-25% on first-year ARR (services attach independent)
STRATEGIC : 25-40% on net ARR with floor + co-investment
Contribution depth modifies the band:
sourced (partner-introduced, partner-owned relationship) -> top half of band
influenced (partner accelerated, but rep ran the play) -> bottom half of band
delivered (partner did the implementation only) -> services-side comp,
NOT product revshare
(refuses to apply
product band)
NEVER auto-commits a revshare %. Output is a band + assumptions + break-even, routed
to a human commercial committee.
Usage:
python revshare_modeler.py --sample
python revshare_modeler.py --input revshare.json
python revshare_modeler.py --input revshare.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from typing import Any
SAMPLE_REVSHARE = {
"partner_name": "Northstar Consulting",
"partner_tier": "SI_CONSULTING",
"deal_avg_size_usd": 90000,
"partner_contribution": "sourced",
"our_cost_to_serve_direct_usd": 18000,
"our_cost_to_serve_via_partner_usd": 9000,
"mdf_annual_usd": 20000,
"program_overhead_annual_usd": 60000,
"ttm_arr_projection_usd": 800000,
"deal_count_projection": 9,
"project_years": 3,
}
VALID_TIERS = ("REFERRAL", "RESELLER", "OEM", "SI_CONSULTING", "STRATEGIC")
VALID_CONTRIBUTIONS = ("sourced", "influenced", "delivered")
# Industry-typical revshare bands. Lower bound = floor; upper bound = ceiling.
# Tuned by `--profile`.
TIER_BANDS: dict[str, tuple[float, float]] = {
"REFERRAL": (5.0, 10.0),
"RESELLER": (20.0, 35.0),
"OEM": (40.0, 55.0),
"SI_CONSULTING": (15.0, 25.0),
"STRATEGIC": (25.0, 40.0),
}
@dataclass
class RevshareModel:
partner_name: str
partner_tier: str
partner_contribution: str
deal_avg_size_usd: float
direct_margin_usd: float
direct_margin_pct: float
via_partner_margin_usd_at_low: float
via_partner_margin_pct_at_low: float
via_partner_margin_usd_at_high: float
via_partner_margin_pct_at_high: float
recommended_revshare_low_pct: float
recommended_revshare_high_pct: float
breakeven_partner_sourced_deals: int
annual_program_cost_usd: float
ttm_arr_projection_usd: float
ttm_revshare_payout_low_usd: float
ttm_revshare_payout_high_usd: float
ttm_net_to_us_low_usd: float
ttm_net_to_us_high_usd: float
crossover_year: int
direct_economics_3yr_npv_usd: float
partner_economics_3yr_npv_low_usd: float
partner_economics_3yr_npv_high_usd: float
assumptions: list[str] = field(default_factory=list)
warnings: list[str] = field(default_factory=list)
validation_errors: list[str] = field(default_factory=list)
def _validate(rev: dict) -> list[str]:
errs: list[str] = []
tier = (rev.get("partner_tier") or "").upper()
contrib = (rev.get("partner_contribution") or "").lower()
if tier not in VALID_TIERS:
errs.append(f"partner_tier '{tier}' not in {VALID_TIERS}")
if contrib not in VALID_CONTRIBUTIONS:
errs.append(f"partner_contribution '{contrib}' not in {VALID_CONTRIBUTIONS}")
if contrib == "delivered" and tier in ("REFERRAL", "RESELLER", "OEM", "STRATEGIC"):
errs.append(
"partner_contribution='delivered' is services attach only — do not pay product "
"revshare. Pay services-side comp (typical fixed services fee or hourly rate). "
"Re-classify contribution or move to SI_CONSULTING tier with explicit services band."
)
if float(rev.get("deal_avg_size_usd", 0)) <= 0:
errs.append("deal_avg_size_usd must be > 0")
return errs
def _contribution_band_shift(band: tuple[float, float], contribution: str) -> tuple[float, float]:
"""Modify band based on contribution depth.
sourced -> top half (mid..high)
influenced -> bottom half (low..mid)
delivered -> applied only at SI_CONSULTING; for SI, this is the floor band.
"""
low, high = band
mid = (low + high) / 2.0
if contribution == "sourced":
return (mid, high)
if contribution == "influenced":
return (low, mid)
# delivered (only reaches here for SI_CONSULTING per validation)
return (low, mid)
def model(rev: dict) -> RevshareModel:
errs = _validate(rev)
if errs:
# Return a stub model with validation errors; nothing else computed.
return RevshareModel(
partner_name=str(rev.get("partner_name", "UNSPECIFIED")),
partner_tier=(rev.get("partner_tier") or "").upper(),
partner_contribution=(rev.get("partner_contribution") or "").lower(),
deal_avg_size_usd=float(rev.get("deal_avg_size_usd", 0.0)),
direct_margin_usd=0.0, direct_margin_pct=0.0,
via_partner_margin_usd_at_low=0.0, via_partner_margin_pct_at_low=0.0,
via_partner_margin_usd_at_high=0.0, via_partner_margin_pct_at_high=0.0,
recommended_revshare_low_pct=0.0, recommended_revshare_high_pct=0.0,
breakeven_partner_sourced_deals=0,
annual_program_cost_usd=0.0,
ttm_arr_projection_usd=0.0,
ttm_revshare_payout_low_usd=0.0, ttm_revshare_payout_high_usd=0.0,
ttm_net_to_us_low_usd=0.0, ttm_net_to_us_high_usd=0.0,
crossover_year=0,
direct_economics_3yr_npv_usd=0.0,
partner_economics_3yr_npv_low_usd=0.0,
partner_economics_3yr_npv_high_usd=0.0,
validation_errors=errs,
)
tier = (rev["partner_tier"] or "").upper()
contrib = (rev["partner_contribution"] or "").lower()
deal_avg = float(rev.get("deal_avg_size_usd", 0.0))
cts_direct = float(rev.get("our_cost_to_serve_direct_usd", 0.0))
cts_partner = float(rev.get("our_cost_to_serve_via_partner_usd", 0.0))
mdf = float(rev.get("mdf_annual_usd", 0.0))
overhead = float(rev.get("program_overhead_annual_usd", 0.0))
ttm_arr = float(rev.get("ttm_arr_projection_usd", 0.0))
deal_count = int(rev.get("deal_count_projection", 0) or 0)
years = int(rev.get("project_years", 3) or 3)
band = TIER_BANDS[tier]
band_low, band_high = _contribution_band_shift(band, contrib)
# Per-deal margin direct (no revshare; full cost-to-serve)
direct_margin = deal_avg - cts_direct
direct_margin_pct = (direct_margin / deal_avg * 100.0) if deal_avg else 0.0
# Per-deal margin via partner: deal - (revshare%) * deal - cts_partner
via_low_payout = deal_avg * (band_low / 100.0)
via_high_payout = deal_avg * (band_high / 100.0)
via_low_margin = deal_avg - via_low_payout - cts_partner
via_high_margin = deal_avg - via_high_payout - cts_partner
via_low_pct = (via_low_margin / deal_avg * 100.0) if deal_avg else 0.0
via_high_pct = (via_high_margin / deal_avg * 100.0) if deal_avg else 0.0
annual_program_cost = mdf + overhead
# Break-even partner-sourced deals: program cost / (direct_margin - via_partner_margin_at_HIGH)
# The cheaper our margin via partner, the more partner-sourced deals required.
# If via_partner margin > direct margin (rare but possible for high-CTS direct sales), break-even is 0.
delta_per_deal = direct_margin - via_high_margin
if delta_per_deal <= 0:
breakeven = 0
else:
breakeven = int(annual_program_cost / delta_per_deal) + 1
# TTM economics: ttm_arr * revshare%
ttm_payout_low = ttm_arr * (band_low / 100.0)
ttm_payout_high = ttm_arr * (band_high / 100.0)
ttm_net_low = ttm_arr - ttm_payout_high - (deal_count * cts_partner) - annual_program_cost
ttm_net_high = ttm_arr - ttm_payout_low - (deal_count * cts_partner) - annual_program_cost
# Direct equivalent at same ARR: ttm_arr - deal_count * cts_direct
ttm_net_direct = ttm_arr - (deal_count * cts_direct)
# Crossover year: at what year does partner economics beat direct?
# Heuristic: if cts_direct - cts_partner > revshare_payout / deal_avg, never
# need partner growth; if not, year when (deals_partner * marginal_savings) >
# (annual_program_cost) — simple linear projection over `years`.
crossover = 0
if cts_direct > cts_partner and deal_count > 0:
marginal_savings_per_deal = (cts_direct - cts_partner) - via_low_payout
if marginal_savings_per_deal > 0:
# cumulative savings needed to cover all program cost across `years`
cumulative_program_cost = annual_program_cost * years
cumulative_savings_per_year = marginal_savings_per_deal * deal_count
if cumulative_savings_per_year > 0:
yrs = cumulative_program_cost / cumulative_savings_per_year
crossover = max(1, int(yrs) + (1 if yrs % 1 else 0))
else:
crossover = 0 # marginal economics never positive at low band
else:
crossover = 0
# 3-year NPV at flat discount (we don't discount — keeping math obvious + auditable):
direct_3yr_npv = ttm_net_direct * years
partner_3yr_npv_low = ttm_net_low * years
partner_3yr_npv_high = ttm_net_high * years
assumptions = [
f"Tier band ({tier}): {band[0]:.0f}-{band[1]:.0f}% — contribution '{contrib}' "
f"shifts band to {band_low:.0f}-{band_high:.0f}%.",
"Cost-to-serve via partner assumes partner owns first-line support; we own "
"Tier-2+. Validate this matches the contract.",
"Revshare is paid on net ARR (post-discount), not gross list price.",
"TTM projection assumes deal_count_projection × deal_avg_size_usd ≈ ttm_arr_projection_usd. "
"If these are inconsistent, fix the input.",
"No churn modeled. Partner-sourced cohorts often have +/- 10pt NRR delta vs. direct — "
"revisit with `c-level-advisor/cco-advisor` for retention-decomposition impact.",
f"NPV computed flat across {years} years (no discount rate). Apply your WACC manually "
"if the partnership is balance-sheet material.",
]
warnings: list[str] = []
if via_high_margin < 0:
warnings.append(
f"At top of band ({band_high:.0f}% revshare + ,.0f CTS), per-deal "
"margin is NEGATIVE. Either lower the band, lower cost-to-serve, or do not sign "
"at this tier."
)
if direct_margin > via_low_margin and contrib == "influenced":
warnings.append(
"Direct-sale margin > via-partner margin at INFLUENCED contribution. The partner "
"is being paid for deals that would have closed anyway. Tighten attribution rules."
)
if ttm_arr > 0 and breakeven > deal_count:
warnings.append(
f"Break-even requires {breakeven} partner-sourced deals/year; projection only "
f"shows {deal_count}. Program is economically UNPROFITABLE at projection scale. "
"Re-scope MDF, re-tier, or unwind."
)
if contrib == "delivered" and tier == "SI_CONSULTING":
warnings.append(
"Delivered-only contribution: pay services-side compensation (fixed fee or hourly) "
"rather than product revshare. Apply only the floor band as a ceiling."
)
if tier == "STRATEGIC" and annual_program_cost < 50000:
warnings.append(
f"STRATEGIC tier with program cost ,.0f/yr is structurally "
"under-resourced. Strategic alliances require co-investment evidence."
)
return RevshareModel(
partner_name=str(rev.get("partner_name", "UNSPECIFIED")),
partner_tier=tier,
partner_contribution=contrib,
deal_avg_size_usd=deal_avg,
direct_margin_usd=round(direct_margin, 2),
direct_margin_pct=round(direct_margin_pct, 1),
via_partner_margin_usd_at_low=round(via_low_margin, 2),
via_partner_margin_pct_at_low=round(via_low_pct, 1),
via_partner_margin_usd_at_high=round(via_high_margin, 2),
via_partner_margin_pct_at_high=round(via_high_pct, 1),
recommended_revshare_low_pct=round(band_low, 1),
recommended_revshare_high_pct=round(band_high, 1),
breakeven_partner_sourced_deals=breakeven,
annual_program_cost_usd=round(annual_program_cost, 2),
ttm_arr_projection_usd=round(ttm_arr, 2),
ttm_revshare_payout_low_usd=round(ttm_payout_low, 2),
ttm_revshare_payout_high_usd=round(ttm_payout_high, 2),
ttm_net_to_us_low_usd=round(ttm_net_low, 2),
ttm_net_to_us_high_usd=round(ttm_net_high, 2),
crossover_year=crossover,
direct_economics_3yr_npv_usd=round(direct_3yr_npv, 2),
partner_economics_3yr_npv_low_usd=round(partner_3yr_npv_low, 2),
partner_economics_3yr_npv_high_usd=round(partner_3yr_npv_high, 2),
assumptions=assumptions,
warnings=warnings,
)
def _render_human(m: RevshareModel) -> str:
lines = []
lines.append(f"Revshare Model: {m.partner_name}")
lines.append(f"Tier: {m.partner_tier} ; Contribution: {m.partner_contribution}")
if m.validation_errors:
lines.append("")
lines.append("VALIDATION ERRORS (model not computed):")
for e in m.validation_errors:
lines.append(f" ! {e}")
return "\n".join(lines)
lines.append("")
lines.append("Recommended revshare band:")
lines.append(
f" {m.recommended_revshare_low_pct:.0f}% to {m.recommended_revshare_high_pct:.0f}% "
f"of net ARR (tier + contribution adjusted)"
)
lines.append("")
lines.append("Per-deal economics:")
lines.append(
f" Direct sale: >10,.0f ARR - margin "
f">10,.0f ({m.direct_margin_pct:.1f}%)"
)
lines.append(
f" Via partner (low): >10,.0f ARR - margin "
f">10,.0f ({m.via_partner_margin_pct_at_low:.1f}%)"
)
lines.append(
f" Via partner (high): >10,.0f ARR - margin "
f">10,.0f ({m.via_partner_margin_pct_at_high:.1f}%)"
)
lines.append("")
lines.append("Break-even program math:")
lines.append(f" Annual program cost (MDF + overhead): ,.0f")
lines.append(
f" Break-even partner-sourced deals/year (at top-of-band): "
f"{m.breakeven_partner_sourced_deals}"
)
lines.append("")
lines.append("Projected TTM economics:")
lines.append(f" TTM ARR through partner: ,.0f")
lines.append(
f" Revshare payout (low..high): "
f",.0f .. ,.0f"
)
lines.append(
f" Net to us (low band..high band): "
f",.0f .. ,.0f"
)
lines.append("")
lines.append("Long-term comparison (flat, no discount):")
lines.append(f" Direct 3-yr NPV: ,.0f")
lines.append(
f" Partner 3-yr NPV (low..high band): "
f",.0f .. "
f",.0f"
)
if m.crossover_year:
lines.append(f" Crossover year (partner > direct): year {m.crossover_year}")
else:
lines.append(" Crossover year: not reached within projection window")
lines.append("")
lines.append("Assumptions:")
for a in m.assumptions:
lines.append(f" - {a}")
if m.warnings:
lines.append("")
lines.append("Warnings:")
for w in m.warnings:
lines.append(f" ! {w}")
return "\n".join(lines)
def _to_jsonable(m: RevshareModel) -> dict:
return asdict(m)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Model revshare economics: direct vs via partner.",
)
parser.add_argument("--input", help="Path to JSON revshare context")
parser.add_argument("--output", default="human", choices=["human", "json", "markdown"])
parser.add_argument("--sample", action="store_true", help="Use embedded sample revshare")
args = parser.parse_args(argv)
if args.sample or not args.input:
rev = SAMPLE_REVSHARE
else:
with open(args.input) as f:
rev = json.load(f)
m = model(rev)
if args.output == "json":
print(json.dumps(_to_jsonable(m), indent=2))
else:
if args.output == "markdown":
print("# Revshare Model\n")
print(_render_human(m))
return 0
if __name__ == "__main__":
sys.exit(main())
Tạo hoặc cập nhật tài liệu ngữ cảnh product marketing: mô tả sản phẩm, đối tượng mục tiêu, ICP và định vị để tránh lặp lại thông tin nền.
---
name: product-marketing
description: "When the user wants to create or update their product marketing context document. Also use when the user mentions 'product context,' 'marketing context,' 'set up context,' 'positioning,' 'who is my target audience,' 'describe my product,' 'ICP,' 'ideal customer profile,' or wants to avoid repeating foundational information across marketing tasks. Use this at the start of any new project before using other marketing skills — it creates `.agents/product-marketing.md` that all other skills reference for product, audience, and positioning context."
metadata:
version: 2.1.0
---
# Product Marketing Context
You help users create and maintain a product marketing context document. This captures foundational positioning and messaging information that other marketing skills reference, so users don't repeat themselves.
The document is stored at `.agents/product-marketing.md`.
## Workflow
### Step 1: Check for Existing Context
First, check if `.agents/product-marketing.md` already exists. Also check `.claude/product-marketing.md` and the legacy filename `product-marketing-context.md` (in either `.agents/` or `.claude/`) for older setups — if found anywhere other than `.agents/product-marketing.md`, offer to move it to the canonical location.
**If it exists:**
- Read it and summarize what's captured — note its current **Document version** and the last few **Changelog** entries so the user sees where the doc stands and what's changed recently
- Ask which sections they want to update
- Only gather info for those sections
- On any substantive save, bump the version and add a changelog entry (see Step 4). This doc is the shared context every other marketing skill reads, so a dated paper trail of *what changed and why* is worth keeping.
**If it doesn't exist, offer two options:**
1. **Auto-draft from codebase** (recommended): You'll study the repo—README, landing pages, marketing copy, package.json, etc.—and draft a V1 of the context document. The user then reviews, corrects, and fills gaps. This is faster than starting from scratch.
2. **Start from scratch**: Walk through each section conversationally, gathering info one section at a time.
Most users prefer option 1. After presenting the draft, ask: "What needs correcting? What's missing?"
### Step 2: Gather Information
**If auto-drafting:**
1. Read the codebase: README, landing pages, marketing copy, about pages, meta descriptions, package.json, any existing docs
2. Draft all sections based on what you find
3. Present the draft and ask what needs correcting or is missing
4. Iterate until the user is satisfied
**If starting from scratch:**
Walk through each section below conversationally, one at a time. Don't dump all questions at once.
For each section:
1. Briefly explain what you're capturing
2. Ask relevant questions
3. Confirm accuracy
4. Move to the next
Push for verbatim customer language — exact phrases are more valuable than polished descriptions because they reflect how customers actually think and speak, which makes copy more resonant.
---
## Sections to Capture
### 1. Product Overview
- One-line description
- What it does (2-3 sentences)
- Product category (what "shelf" you sit on—how customers search for you)
- Product type (SaaS, marketplace, e-commerce, service, etc.)
- Business model and pricing
### 2. Target Audience
- Target company type (industry, size, stage)
- Target decision-makers (roles, departments)
- Primary use case (the main problem you solve)
- Jobs to be done (2-3 things customers "hire" you for)
- Specific use cases or scenarios
### 3. Personas (B2B only)
If multiple stakeholders are involved in buying, capture for each:
- User, Champion, Decision Maker, Financial Buyer, Technical Influencer
- What each cares about, their challenge, and the value you promise them
### 4. Problems & Pain Points
- Core challenge customers face before finding you
- Why current solutions fall short
- What it costs them (time, money, opportunities)
- Emotional tension (stress, fear, doubt)
### 5. Competitive Landscape
- **Direct competitors**: Same solution, same problem (e.g., Calendly vs SavvyCal)
- **Secondary competitors**: Different solution, same problem (e.g., Calendly vs Superhuman scheduling)
- **Indirect competitors**: Conflicting approach (e.g., Calendly vs personal assistant)
- How each falls short for customers
### 6. Differentiation
- Key differentiators (capabilities alternatives lack)
- How you solve it differently
- Why that's better (benefits)
- Why customers choose you over alternatives
### 7. Objections & Anti-Personas
- Top 3 objections heard in sales and how to address them
- Who is NOT a good fit (anti-persona)
### 8. Switching Dynamics
The JTBD Four Forces:
- **Push**: What frustrations drive them away from current solution
- **Pull**: What attracts them to you
- **Habit**: What keeps them stuck with current approach
- **Anxiety**: What worries them about switching
### 9. Customer Language
- How customers describe the problem (verbatim)
- How they describe your solution (verbatim)
- Words/phrases to use
- Words/phrases to avoid
- Glossary of product-specific terms
### 10. Brand Voice
- Tone (professional, casual, playful, etc.)
- Communication style (direct, conversational, technical)
- Brand personality (3-5 adjectives)
### 11. Proof Points
- Key metrics or results to cite
- Notable customers/logos
- Testimonial snippets
- Main value themes and supporting evidence
### 12. Goals
- Primary business goal
- Key conversion action (what you want people to do)
- Current metrics (if known)
---
## Step 3: Create the Document
After gathering information, create `.agents/product-marketing.md` with this structure:
```markdown
# Product Marketing Context
**Document version:** v1
**Last updated:** [date]
## Product Overview
**One-liner:**
**What it does:**
**Product category:**
**Product type:**
**Business model:**
## Target Audience
**Target companies:**
**Decision-makers:**
**Primary use case:**
**Jobs to be done:**
-
**Use cases:**
-
## Personas
| Persona | Cares about | Challenge | Value we promise |
|---------|-------------|-----------|------------------|
| | | | |
## Problems & Pain Points
**Core problem:**
**Why alternatives fall short:**
-
**What it costs them:**
**Emotional tension:**
## Competitive Landscape
**Direct:** [Competitor] — falls short because...
**Secondary:** [Approach] — falls short because...
**Indirect:** [Alternative] — falls short because...
## Differentiation
**Key differentiators:**
-
**How we do it differently:**
**Why that's better:**
**Why customers choose us:**
## Objections
| Objection | Response |
|-----------|----------|
| | |
**Anti-persona:**
## Switching Dynamics
**Push:**
**Pull:**
**Habit:**
**Anxiety:**
## Customer Language
**How they describe the problem:**
- "[verbatim]"
**How they describe us:**
- "[verbatim]"
**Words to use:**
**Words to avoid:**
**Glossary:**
| Term | Meaning |
|------|---------|
| | |
## Brand Voice
**Tone:**
**Style:**
**Personality:**
## Proof Points
**Metrics:**
**Customers:**
**Testimonials:**
> "[quote]" — [who]
**Value themes:**
| Theme | Proof |
|-------|-------|
| | |
## Goals
**Business goal:**
**Conversion action:**
**Current metrics:**
## Changelog
*Newest first. One line per revision: what changed and why.*
- v1 ([date]) — Initial context.
```
---
## Step 4: Confirm, Version, and Save
- Show the completed document
- Ask if anything needs adjustment
- **Set the version and changelog** — this is the paper trail for a doc every other skill reads:
- **New document:** set `Document version: v1` and a single Changelog entry — `- v1 ([today]) — Initial context.`
- **Updating an existing document:** increment the version (v2 → v3 …), update `Last updated` to today, and **prepend a new Changelog entry** at the top of the list (newest first) summarizing *what changed and why* in one line. Never rewrite or reorder past entries.
- A good entry names the sections touched and the reason, not "updated the doc." Examples:
- `- v3 (2026-07-16) — Repositioned from "email tool" to "deliverability platform"; added RevOps to the ICP.`
- `- v2 (2026-06-02) — Rewrote value prop and objections after 5 customer interviews; added competitor Acme.`
- Use today's date in ISO form (YYYY-MM-DD) for the entry and `Last updated`.
- **Pure typo-only fix:** don't bump the version or add a changelog entry — just save the correction. Every other change bumps the version and gets an entry. When the change is a real repositioning, say so plainly — downstream skills will now generate against the new context.
- Save to `.agents/product-marketing.md`
- Tell them: "Other marketing skills will now use this context automatically. The Changelog at the bottom tracks every revision — check it to see how your positioning has evolved. Run `/product-marketing` anytime to update it."
---
## Tips
- **Be specific**: Ask "What's the #1 frustration that brings them to you?" not "What problem do they solve?"
- **Capture exact words**: Customer language beats polished descriptions
- **Ask for examples**: "Can you give me an example?" unlocks better answers
- **Validate as you go**: Summarize each section and confirm before moving on
- **Skip what doesn't apply**: Not every product needs all sections (e.g., Personas for B2C)
FILE:evals/evals.json
{
"skill_name": "product-marketing",
"evals": [
{
"id": 1,
"prompt": "I want to set up my product marketing context. We're a B2B SaaS company that sells a customer feedback platform to product teams.",
"expected_output": "Should check if .agents/product-marketing.md already exists. If not, should offer two options: (1) Auto-draft from codebase (recommended) or (2) Start from scratch. If user chooses start from scratch, should walk through sections conversationally one at a time. Should cover all applicable sections: Product Overview, Target Audience, Personas, Problems You Solve, Competitive Landscape, Differentiation, Objections, Switching Dynamics, Customer Language, Brand Voice, Proof Points, and Goals. Should create the file at .agents/product-marketing.md when complete.",
"assertions": [
"Checks for existing product-marketing.md",
"Offers two options: auto-draft or start from scratch",
"Covers applicable sections",
"Walks through sections conversationally one at a time",
"Creates file at .agents/product-marketing.md"
],
"files": []
},
{
"id": 2,
"prompt": "Update our product marketing context. We just added a new enterprise tier and our target audience has expanded to include VP of Engineering, not just Product Managers.",
"expected_output": "Should check for existing .agents/product-marketing.md and read it. Should identify which sections need updating based on the changes: Target Audience (add VP of Engineering), Personas (add new persona), Product Overview (new enterprise tier, including pricing updates within that section), Objections (enterprise-specific), and Competitive Landscape (enterprise competitors). Should update only the relevant sections, preserving existing content that hasn't changed.",
"assertions": [
"Reads existing product-marketing.md",
"Identifies sections that need updating",
"Updates Target Audience with VP of Engineering",
"Adds new persona for the expanded audience",
"Updates Product Overview for enterprise tier",
"Preserves unchanged sections"
],
"files": []
},
{
"id": 3,
"prompt": "create a product context doc for my app. it's a mobile app that helps people find hiking trails. we're just getting started.",
"expected_output": "Should trigger on casual phrasing. Should check for existing context doc. Should offer auto-draft or start-from-scratch options. Should adapt questions for an early-stage B2C mobile app (outdoor/fitness niche). Should note that some sections may be sparse for an early-stage product and that's okay — they can be filled in as the business matures. Should skip non-applicable sections (e.g., Personas section is B2B-focused) rather than forcing all 12. Should accept lighter answers for sections like Proof Points or Competitive Landscape if the company is new.",
"assertions": [
"Triggers on casual phrasing",
"Checks for existing context doc",
"Offers auto-draft or start-from-scratch options",
"Adapts questions for early-stage B2C mobile app",
"Notes some sections may be sparse early on",
"Skips non-applicable sections rather than forcing all 12",
"Creates file at .agents/product-marketing.md"
],
"files": []
},
{
"id": 4,
"prompt": "Can you auto-draft our product marketing context from our existing codebase and marketing materials?",
"expected_output": "Should activate the auto-draft workflow mode. Should scan the codebase for existing marketing context: README, landing page copy, pricing page, about page, meta descriptions, any existing documentation. Should draft the product-marketing.md from what it finds, filling in sections where information is available and flagging sections that need manual input. Should present the draft for review before saving.",
"assertions": [
"Activates auto-draft workflow mode",
"Scans codebase for existing marketing materials",
"Drafts context from found information",
"Flags sections needing manual input",
"Presents draft for review before saving"
],
"files": []
},
{
"id": 5,
"prompt": "Do we have a product marketing context set up? I want to make sure the other marketing skills have context about our product.",
"expected_output": "Should check for .agents/product-marketing.md (and the older .claude/product-marketing.md location). Should report whether it exists and summarize its contents if found. If it doesn't exist, should offer to create one and explain why it's valuable (other skills like copywriting, cro, seo-audit check for it first). Should explain how other skills use this context document.",
"assertions": [
"Checks both file locations",
"Reports whether context doc exists",
"Summarizes contents if found",
"Offers to create if missing",
"Explains how other skills use it"
],
"files": []
},
{
"id": 6,
"prompt": "Write homepage copy for our SaaS product.",
"expected_output": "Should recognize this is a copywriting task, not a product marketing context task. Should check for product-marketing.md (as other skills do), and if it doesn't exist, may suggest creating one first. But should defer to the copywriting skill for actually writing the homepage copy.",
"assertions": [
"Recognizes this as a copywriting task",
"May check for or suggest creating product-marketing.md",
"References or defers to copywriting skill for the actual copy",
"Does not attempt to write homepage copy using context creation patterns"
],
"files": []
},
{
"id": 7,
"prompt": "We just repositioned — we're no longer an 'email tool,' we're a 'deliverability platform,' and our ICP now includes RevOps teams. Update our product marketing context.",
"expected_output": "Should recognize an existing .agents/product-marketing.md, read it, note its current Document version and recent Changelog entries, and update only the affected sections (product overview/positioning, target audience/ICP). On save, should bump the Document version (e.g. v2 → v3), update the Last updated date, and PREPEND a new newest-first Changelog entry summarizing what changed and why in one line — e.g. 'Repositioned from email tool to deliverability platform; added RevOps to the ICP' — naming the sections touched and the reason, not just 'updated the doc.' Should not rewrite or reorder past changelog entries. Should tell the user the changelog tracks revisions and that downstream skills will now use the new context.",
"assertions": [
"Reads the existing doc and surfaces its current version + recent changelog",
"Updates only the affected sections (positioning + ICP)",
"Bumps the Document version and updates Last updated",
"Prepends a newest-first changelog entry naming what changed and why",
"Preserves prior changelog entries unchanged"
],
"files": []
}
]
}