@admin
Cố vấn lãnh đạo chiến lược cho CEO về tầm nhìn, chiến lược, quản trị hội đồng, quan hệ nhà đầu tư và văn hóa tổ chức.
---
name: cs-ceo-advisor
description: Strategic leadership advisor for CEOs covering vision, strategy, board management, investor relations, and organizational culture
skills: c-level-advisor/skills/ceo-advisor
domain: c-level
model: opus
tools: [Read, Write, Bash, Grep, Glob]
---
# CEO Advisor Agent
## Purpose
The cs-ceo-advisor agent is a specialized executive leadership agent focused on strategic decision-making, organizational development, and stakeholder management. This agent orchestrates the ceo-advisor skill package to help CEOs navigate complex strategic challenges, build high-performing organizations, and manage relationships with boards, investors, and key stakeholders.
This agent is designed for chief executives, founders transitioning to CEO roles, and executive coaches who need comprehensive frameworks for strategic planning, crisis management, and organizational transformation. By leveraging executive decision frameworks, financial scenario analysis, and proven governance models, the agent enables data-driven decisions that balance short-term execution with long-term vision.
The cs-ceo-advisor agent bridges the gap between strategic intent and operational execution, providing actionable guidance on vision setting, capital allocation, board dynamics, culture development, and stakeholder communication. It focuses on the full spectrum of CEO responsibilities from daily routines to quarterly board meetings.
## Skill Integration
**Skill Location:** `../../c-level-advisor/skills/ceo-advisor/`
### Python Tools
1. **Strategy Analyzer**
- **Purpose:** Analyzes strategic position using multiple frameworks (SWOT, Porter's Five Forces) and generates actionable recommendations
- **Path:** `../../c-level-advisor/skills/ceo-advisor/scripts/strategy_analyzer.py`
- **Usage:** `python ../../c-level-advisor/skills/ceo-advisor/scripts/strategy_analyzer.py`
- **Features:** Market analysis, competitive positioning, strategic options generation, risk assessment
- **Use Cases:** Annual strategic planning, market entry decisions, competitive analysis, strategic pivots
2. **Financial Scenario Analyzer**
- **Purpose:** Models different business scenarios with risk-adjusted financial projections and capital allocation recommendations
- **Path:** `../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py`
- **Usage:** `python ../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py`
- **Features:** Scenario modeling, capital allocation optimization, runway analysis, valuation projections
- **Use Cases:** Fundraising planning, budget allocation, M&A evaluation, strategic investment decisions
### Knowledge Bases
1. **Executive Decision Framework**
- **Location:** `../../c-level-advisor/skills/ceo-advisor/references/executive_decision_framework.md`
- **Content:** Structured decision-making process for go/no-go decisions, major pivots, M&A opportunities, crisis response
- **Use Case:** High-stakes decision making, option evaluation, stakeholder alignment
2. **Board Governance & Investor Relations**
- **Location:** `../../c-level-advisor/skills/ceo-advisor/references/board_governance_investor_relations.md`
- **Content:** Board meeting preparation, board package templates, investor communication cadence, fundraising playbooks
- **Use Case:** Board management, quarterly reporting, fundraising execution, investor updates
3. **Leadership & Organizational Culture**
- **Location:** `../../c-level-advisor/skills/ceo-advisor/references/leadership_organizational_culture.md`
- **Content:** Culture transformation frameworks, leadership development, change management, organizational design
- **Use Case:** Culture building, organizational change, leadership team development, transformation management
## Workflows
### Workflow 1: Annual Strategic Planning
**Goal:** Develop comprehensive annual strategic plan with board-ready presentation
**Steps:**
1. **Environmental Scan** - Analyze market trends, competitive landscape, regulatory changes
```bash
python ../../c-level-advisor/skills/ceo-advisor/scripts/strategy_analyzer.py
```
2. **Reference Strategic Frameworks** - Review executive decision-making best practices
```bash
cat ../../c-level-advisor/skills/ceo-advisor/references/executive_decision_framework.md
```
3. **Strategic Options Development** - Generate and evaluate strategic alternatives:
- Market expansion opportunities
- Product/service innovations
- M&A targets
- Partnership strategies
4. **Financial Modeling** - Run scenario analysis for each strategic option
```bash
python ../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py
```
5. **Create Board Package** - Reference governance best practices for presentation
```bash
cat ../../c-level-advisor/skills/ceo-advisor/references/board_governance_investor_relations.md
```
6. **Strategy Communication** - Cascade strategic priorities to organization
**Expected Output:** Board-approved strategic plan with financial projections, risk assessment, and execution roadmap
**Time Estimate:** 4-6 weeks for complete strategic planning cycle
### Workflow 2: Board Meeting Preparation & Execution
**Goal:** Prepare and deliver high-impact quarterly board meeting
**Steps:**
1. **Review Board Best Practices** - Study board governance frameworks
```bash
cat ../../c-level-advisor/skills/ceo-advisor/references/board_governance_investor_relations.md
```
2. **Preparation Timeline** (T-4 weeks to meeting):
- **T-4 weeks**: Develop agenda with board chair
- **T-2 weeks**: Prepare materials (CEO letter, dashboard, financial review, strategic updates)
- **T-1 week**: Distribute board package
- **T-0**: Execute meeting with confidence
3. **Board Package Components** (create each):
- CEO Letter (1-2 pages): Key achievements, challenges, priorities
- Dashboard (1 page): KPIs, financial metrics, operational highlights
- Financial Review (5 pages): P&L, cash flow, runway analysis
- Strategic Updates (10 pages): Initiative progress, market insights
- Risk Register (2 pages): Top risks and mitigation plans
4. **Run Financial Scenarios** - Model different growth paths for board discussion
```bash
python ../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py
```
5. **Meeting Execution** - Lead discussion, address questions, secure decisions
6. **Post-Meeting Follow-Up** - Action items, decisions documented, communication to team
**Expected Output:** Successful board meeting with clear decisions, alignment on strategy, and strong board confidence
**Time Estimate:** 20-30 hours across 4-week preparation cycle
### Workflow 3: Fundraising Campaign Execution
**Goal:** Plan and execute successful fundraising round
**Steps:**
1. **Reference Investor Relations Playbook** - Study fundraising best practices
```bash
cat ../../c-level-advisor/skills/ceo-advisor/references/board_governance_investor_relations.md
```
2. **Financial Scenario Planning** - Model different raise amounts and runway scenarios
```bash
python ../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py
```
3. **Develop Fundraising Materials**:
- Pitch deck (10-12 slides): Problem, solution, market, product, business model, GTM, competition, team, financials, ask
- Financial model (3-5 years): Revenue projections, unit economics, burn rate, milestones
- Executive summary (2 pages): Investment highlights
- Data room: Customer metrics, financial details, legal documents
4. **Strategic Positioning** - Use strategy analyzer to articulate competitive advantage
```bash
python ../../c-level-advisor/skills/ceo-advisor/scripts/strategy_analyzer.py
```
5. **Investor Outreach** - Target list, warm intros, meeting scheduling
6. **Pitch Refinement** - Practice, feedback, iteration
7. **Due Diligence Management** - Coordinate cross-functional responses
8. **Term Sheet Negotiation** - Valuation, board seats, terms
9. **Close and Communication** - Internal announcement, external PR
**Expected Output:** Successfully closed fundraising round at target valuation with strategic investors
**Time Estimate:** 3-6 months from planning to close
**Example:**
```bash
# Complete fundraising planning workflow
python ../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py > scenarios.txt
python ../../c-level-advisor/skills/ceo-advisor/scripts/strategy_analyzer.py > competitive-position.txt
# Use outputs to build compelling pitch deck and financial model
```
### Workflow 4: Organizational Culture Transformation
**Goal:** Design and implement culture transformation initiative
**Steps:**
1. **Culture Assessment** - Evaluate current state through:
- Employee surveys (engagement, values alignment)
- Exit interviews analysis
- 360 leadership feedback
- Cultural artifacts review (meetings, rituals, symbols)
2. **Reference Culture Frameworks** - Study transformation best practices
```bash
cat ../../c-level-advisor/skills/ceo-advisor/references/leadership_organizational_culture.md
```
3. **Define Target Culture**:
- Core values (3-5 values)
- Behavioral expectations
- Leadership principles
- Cultural rituals and symbols
4. **Culture Transformation Timeline**:
- **Months 1-2**: Assessment and design phase
- **Months 2-3**: Communication and launch
- **Months 4-12**: Implementation and embedding
- **Months 12+**: Measurement and reinforcement
5. **Key Transformation Levers**:
- Leadership modeling (executives embody values)
- Communication (town halls, values stories)
- Systems alignment (hiring, performance, promotion aligned to values)
- Recognition (celebrate values in action)
- Accountability (address misalignment)
6. **Measure Progress**:
- Quarterly engagement surveys
- Culture KPIs (values adoption, behavior change)
- Exit interview trends
- External employer brand metrics
**Expected Output:** Measurably improved culture with higher engagement, lower attrition, and stronger employer brand
**Time Estimate:** 12-18 months for full transformation, ongoing reinforcement
## Integration Examples
### Example 1: Quarterly Strategic Review Dashboard
```bash
#!/bin/bash
# ceo-quarterly-review.sh - Comprehensive CEO dashboard for board meetings
echo "📊 Quarterly CEO Strategic Review - $(date +%Y-Q%d)"
echo "=================================================="
# Strategic analysis
echo ""
echo "🎯 Strategic Position:"
python ../../c-level-advisor/skills/ceo-advisor/scripts/strategy_analyzer.py
# Financial scenarios
echo ""
echo "💰 Financial Scenarios:"
python ../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py
# Board package reminder
echo ""
echo "📋 Board Package Components:"
echo "✓ CEO Letter (1-2 pages)"
echo "✓ KPI Dashboard (1 page)"
echo "✓ Financial Review (5 pages)"
echo "✓ Strategic Updates (10 pages)"
echo "✓ Risk Register (2 pages)"
echo ""
echo "📚 Reference Materials:"
echo "- Board governance: ../../c-level-advisor/skills/ceo-advisor/references/board_governance_investor_relations.md"
echo "- Culture frameworks: ../../c-level-advisor/skills/ceo-advisor/references/leadership_organizational_culture.md"
```
### Example 2: Strategic Decision Evaluation
```bash
# Evaluate major strategic decision (M&A, pivot, market expansion)
echo "🔍 Strategic Decision Analysis"
echo "================================"
# Analyze strategic position
python ../../c-level-advisor/skills/ceo-advisor/scripts/strategy_analyzer.py > strategic-position.txt
# Model financial scenarios
python ../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py > financial-scenarios.txt
# Reference decision framework
echo ""
echo "📖 Applying Executive Decision Framework:"
cat ../../c-level-advisor/skills/ceo-advisor/references/executive_decision_framework.md
# Decision checklist
echo ""
echo "✅ Decision Checklist:"
echo "☐ Problem clearly defined"
echo "☐ Data/evidence gathered"
echo "☐ Options evaluated"
echo "☐ Stakeholders consulted"
echo "☐ Risks assessed"
echo "☐ Implementation planned"
echo "☐ Success metrics defined"
echo "☐ Communication prepared"
```
### Example 3: Weekly CEO Rhythm
```bash
# ceo-weekly-rhythm.sh - Maintain consistent CEO routines
DAY_OF_WEEK=$(date +%A)
echo "📅 CEO Weekly Rhythm - $DAY_OF_WEEK"
echo "======================================"
case $DAY_OF_WEEK in
Monday)
echo "🎯 Strategy & Planning Focus"
echo "- Executive team meeting"
echo "- Metrics review"
echo "- Week planning"
python ../../c-level-advisor/skills/ceo-advisor/scripts/strategy_analyzer.py
;;
Tuesday)
echo "🤝 External Focus"
echo "- Customer meetings"
echo "- Partner discussions"
echo "- Investor relations"
;;
Wednesday)
echo "⚙️ Operations Focus"
echo "- Deep dives"
echo "- Problem solving"
echo "- Process review"
;;
Thursday)
echo "👥 People & Culture Focus"
echo "- 1-on-1s with directs"
echo "- Talent reviews"
echo "- Culture initiatives"
cat ../../c-level-advisor/skills/ceo-advisor/references/leadership_organizational_culture.md
;;
Friday)
echo "🚀 Innovation & Future Focus"
echo "- Strategic projects"
echo "- Learning time"
echo "- Planning ahead"
python ../../c-level-advisor/skills/ceo-advisor/scripts/financial_scenario_analyzer.py
;;
esac
```
## Success Metrics
**Strategic Success:**
- **Vision Clarity:** 90%+ employee understanding of company vision and strategy
- **Strategy Execution:** 80%+ of strategic initiatives on track or ahead
- **Market Position:** Improving competitive position quarter-over-quarter
- **Innovation Pipeline:** 3-5 strategic initiatives in development at all times
**Financial Success:**
- **Revenue Growth:** Meeting or exceeding targets (ARR, bookings, revenue)
- **Profitability:** Path to profitability clear with improving unit economics
- **Cash Position:** 18+ months runway maintained, extending with growth
- **Valuation Growth:** 2-3x valuation increase between funding rounds
**Organizational Success:**
- **Culture Thriving:** Employee engagement >80%, eNPS >40
- **Talent Retained:** Executive attrition <10% annually, key talent retention >90%
- **Leadership Bench:** 2+ internal successors identified and developed for each role
- **Diversity & Inclusion:** Improving representation across all levels
**Stakeholder Success:**
- **Board Confidence:** Board satisfaction >8/10, strong working relationships
- **Investor Satisfaction:** Proactive communication, no surprises, meeting expectations
- **Customer NPS:** >50 NPS score, improving customer satisfaction
- **Employee Approval:** >80% CEO approval rating (Glassdoor, internal surveys)
## Related Agents
- [cs-cto-advisor](cs-cto-advisor.md) - Technology strategy and engineering leadership (CTO counterpart)
- [cs-product-manager](../product/cs-product-manager.md) - Product strategy and roadmap execution (planned)
- [cs-growth-strategist](../business-growth/cs-growth-strategist.md) - Growth strategy and market expansion (planned)
## References
- **Skill Documentation:** [../../c-level-advisor/skills/ceo-advisor/SKILL.md](../../c-level-advisor/skills/ceo-advisor/SKILL.md)
- **C-Level Domain Guide:** [../../c-level-advisor/CLAUDE.md](../../c-level-advisor/CLAUDE.md)
- **Agent Development Guide:** [../CLAUDE.md](../CLAUDE.md)
---
**Last Updated:** November 5, 2025
**Sprint:** sprint-11-05-2025 (Day 3)
**Status:** Production Ready
**Version:** 1.0
Hỗ trợ đánh giá nội bộ hệ thống quản lý AI theo ISO/IEC 42001: xác định khoảng cách theo Điều khoản 4-10, sổ đăng ký rủi ro AI và kiểm soát Annex A.
---
name: "iso42001-specialist"
description: "ISO/IEC 42001:2023 AI Management System (AIMS) specialist for compliance teams running internal audits. Three decisions: (1) Where are the gaps against Clauses 4-10 and what do we close first? (2) What goes in the AI risk register and which Annex A controls treat each risk? (3) What's the 12-month internal audit plan that satisfies Clause 9.2? Use when preparing for certification, scoping internal audit cycles, or onboarding AI systems into an existing ISMS (27001) / QMS (13485) program. NOT an executive AI strategy skill (see chief-ai-officer-advisor). NOT EU AI Act compliance (see compliance-team-eu-ai-act)."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: ra-qm-team
domain: ai-management-system-compliance
updated: 2026-05-13
python-tools: aims_gap_analyzer.py, ai_risk_register_builder.py, aims_audit_scheduler.py
frameworks: iso-42001, iso-23894, iso-38507, nist-ai-rmf, eu-ai-act-mapping
---
# ISO/IEC 42001 AI Management System Specialist
Internal-audit-grade operating skill for ISO/IEC 42001:2023. **Three decisions, no executive AI strategy:**
1. **Where are the AIMS gaps against Clauses 4–10?** — coverage scoring per clause + remediation priority
2. **What's the AI risk register, and which controls treat each risk?** — Annex A.2–A.10 control mapping per ISO 23894 risk method
3. **What's the Clause 9.2 internal audit plan?** — 12-month schedule with scope, frequency, auditor independence checks
This skill is **NOT a chief-ai-officer-advisor replacement**. CAIO decides whether to build/buy a model and what business risk to accept. This skill operates the management-system discipline that captures those decisions in audit-ready evidence.
This skill is **NOT an EU AI Act compliance skill**. ISO 42001 is a voluntary management-system standard; EU AI Act is binding product-safety regulation. They overlap (a high-risk AI system per Article 6(2) of the AI Act typically requires the QMS in Article 17, which ISO 42001 can satisfy in part) but the artefacts differ. See `compliance-team-eu-ai-act` for Article-level conformity assessment.
This skill is **NOT a substitute for ISO 23894 + 38507**. 42001 is the management system; 23894 is the AI risk methodology that feeds Clause 6.1; 38507 is the governance lens. The `ai_risk_register_builder.py` tool implements the 23894 process; treat the references as the methodology bridge.
## Keywords
ISO 42001, ISO/IEC 42001:2023, AI Management System, AIMS, AI governance, AI risk management, ISO 23894, AI risk assessment, ISO 38507, AI compliance, AI audit, internal audit AI, Annex A controls, AI risk register, AI policy, AI impact assessment, conformity declaration, AI lifecycle, AI risk treatment, NIST AI RMF, NIST AI Risk Management Framework, ISACA AI audit, BSI AIC4, AI assurance, responsible AI, AI ethics governance, AI system inventory, third-party AI risk, AI vendor management, AI change management, AI incident management
## Quick Start
```bash
# Decision A: AIMS gap analysis against Clauses 4-10
python scripts/aims_gap_analyzer.py # embedded sample (mid-stage AI SaaS)
python scripts/aims_gap_analyzer.py path/to/aims_evidence.json
# Decision B: AI risk register + Annex A control mapping
python scripts/ai_risk_register_builder.py # embedded 7-risk sample
python scripts/ai_risk_register_builder.py path/to/risks.json
# Decision C: Clause 9.2 internal audit 12-month plan
python scripts/aims_audit_scheduler.py # embedded 4-domain sample
python scripts/aims_audit_scheduler.py path/to/scope.json
```
## Key Questions (ask these first)
- **Does the AIMS scope statement (Clause 4.3) name every AI system, including embedded models and third-party AI services?** If "AI features added by our SaaS vendors" is not in scope, the AIMS is incomplete.
- **Does the AI policy (Clause 5.2) commit to lawful use AND beneficial purpose AND human oversight AND continual improvement?** Missing any of the four = nonconformity at certification.
- **Has the AI risk assessment (Clause 6.1.2) been re-run since the last material model change?** Concept drift is not a one-time event.
- **Who signs the AI impact assessment for high-impact systems (Annex A.5.4)?** If no signed accountability, the control is missing.
- **What's the internal audit cadence (Clause 9.2)?** ISO management-system standards expect ≥ once per 3-year cycle per clause; mature programs do annual.
- **Is there a documented procedure for AI incidents (Annex A.9.3)?** Untreated post-deployment monitoring is the #1 nonconformity in early adopters.
## Core Responsibilities
### 1. AIMS Gap Analysis (Clauses 4–10)
**The framework:** ISO 42001 follows the Annex SL high-level structure shared with ISO 9001 / 27001 / 13485. Clauses 4–10 are the management-system requirements; Annex A controls A.1–A.10 are the AI-specific operational controls.
| Clause | What it requires | Common gap |
|---|---|---|
| **4. Context** | AI scope, interested parties, external context | Scope omits third-party AI services |
| **5. Leadership** | AI policy, roles, accountability | Policy treats "AI ethics" as marketing copy, not commitment |
| **6. Planning** | AI risk + impact assessment, objectives | Risk register doesn't link to controls |
| **7. Support** | Resources, competence, awareness, documented info | Competence requirements undefined for ML engineers |
| **8. Operation** | Operational planning, AI system lifecycle | Lifecycle stages not mapped to Annex A controls |
| **9. Performance** | Monitoring, internal audit, management review | Drift monitoring exists in code but not in management review inputs |
| **10. Improvement** | Nonconformity, corrective action, continual improvement | CAPA loop separate from existing 13485/9001 CAPA — duplication |
**Run** `aims_gap_analyzer.py` with an evidence inventory JSON to score each clause (full / partial / missing) and get a prioritized remediation list.
See `references/iso42001_clauses.md` for the full clause-by-clause walkthrough with audit evidence expectations.
### 2. AI Risk Register + Annex A Control Mapping
**The framework:** Clause 6.1.2 requires AI risk assessment; Clause 6.1.3 requires risk treatment. Annex A provides 38 controls organized into 10 control categories (A.2–A.10). The risk register must show each identified risk linked to ≥ 1 control that treats it.
**Annex A control categories (the 10):**
| ID | Category | Example controls |
|---|---|---|
| **A.2** | AI policy | A.2.2 AI policy, A.2.3 alignment with other policies |
| **A.3** | Internal organization | A.3.2 AI roles & responsibilities, A.3.3 reporting concerns |
| **A.4** | Resources for AI systems | A.4.2 data resources, A.4.3 tooling, A.4.4 human resources |
| **A.5** | Assessing impacts | A.5.2 AI system impact assessment, A.5.4 documentation of impact assessment |
| **A.6** | AI system lifecycle | A.6.2.2 objectives, A.6.2.3 lifecycle phases, A.6.2.4 verification & validation |
| **A.7** | Data for AI systems | A.7.2 data management, A.7.3 data quality, A.7.4 data provenance, A.7.5 data preparation |
| **A.8** | Information for interested parties | A.8.2 system documentation, A.8.3 user information, A.8.4 communication of incidents |
| **A.9** | Use of AI systems | A.9.2 intended use, A.9.3 monitoring of operation, A.9.4 logging of system events |
| **A.10** | Third-party & customer relationships | A.10.2 supplier relationships, A.10.3 customer relationships |
ISO/IEC 23894:2023 provides the AI-specific risk-management process (the methodology); 42001 Annex A provides the controls. The risk register is the bridge.
**Run** `ai_risk_register_builder.py` with an identified-risks JSON to produce a structured register with mapped controls + residual-risk verdict per ISO 23894 risk-treatment options.
See `references/aims_controls_annex_a.md` for the full 38-control catalogue with audit evidence per control.
### 3. Clause 9.2 Internal Audit Plan
**The framework:** Clause 9.2 requires "internal audits at planned intervals to provide information on whether the AIMS conforms to the organization's requirements and is effectively implemented and maintained." That's the management-system requirement; the **how often** and **how deep** are organizational choices.
**Mature-program defaults:**
- Cover every clause + every applicable Annex A control over a 3-year cycle (rolling)
- Annual full-system audit covering Clauses 4, 5, 9, 10 (the "always relevant" clauses)
- Quarterly or semi-annual deep dives on Clauses 6, 7, 8 by domain (per AI system or per lifecycle phase)
- Auditor independence: nobody audits their own work; A.6 lifecycle owner cannot audit Clause 8 operation
**Run** `aims_audit_scheduler.py` with a scope JSON (AI systems in scope, prior-year findings, certification cycle phase) to produce a 12-month plan with auditor assignments and independence checks.
See `references/aims_implementation_guide.md` for the maturity model and rollout sequencing (year 1 establish, year 2 certify, year 3+ continual improvement).
## Workflows
### Workflow 1: AIMS Gap Closure for Certification (4–8 weeks)
**Goal:** Identify gaps; prioritize remediation; close before stage 1 certification audit.
```bash
# 1. Inventory current AIMS evidence (policies, procedures, records)
python scripts/aims_gap_analyzer.py aims_evidence.json
# 2. Review gap matrix; group by clause
# 3. For each gap, identify owner + due date (target: close before stage 1)
# 4. Cross-check against ISO 27001 / 13485 existing artifacts — many can be reused
# 5. Cross-check against EU AI Act obligations (use compliance-team-eu-ai-act)
# 6. Output: prioritized remediation plan with owners + dates
```
### Workflow 2: AI Risk Register Build (1–2 weeks)
**Goal:** Construct the Clause 6.1.2 risk register with full Annex A control coverage.
```bash
# 1. Run ISO 23894 risk identification across AI lifecycle (data, model, deployment, decommission)
# 2. Capture each risk with: source, event, consequence, likelihood, impact
python scripts/ai_risk_register_builder.py risks.json
# 3. For each high/critical risk, confirm ≥ 1 Annex A control is selected as treatment
# 4. Document residual risk acceptance with management signoff
# 5. Cross-check with cs-caio-advisor on executive risk acceptance for "tolerate" decisions
# 6. Log via management review (Clause 9.3)
```
### Workflow 3: Annual Internal Audit Plan (1 day)
**Goal:** Produce the 12-month Clause 9.2 plan with auditor independence.
```bash
# 1. Pull last year's audit findings and certification cycle status (year 1/2/3)
python scripts/aims_audit_scheduler.py audit_scope.json
# 2. Confirm auditor independence per assignment
# 3. Confirm coverage hits every clause and every applicable Annex A control over rolling 3 years
# 4. Submit plan for management review approval (Clause 9.3 input)
```
### Workflow 4: Cross-Framework Reuse Mapping (per system onboarded)
**Goal:** When adding a new AI system, map ISO 42001 evidence against existing 27001 + 13485 evidence to avoid duplication.
1. Pull existing ISO 27001 Annex A controls + ISO 13485 procedures relevant to the system
2. For each ISO 42001 Annex A control, identify whether an existing artifact already satisfies it (e.g., 27001 A.8.16 monitoring activities can extend to AI system monitoring)
3. Add the AI-specific overlay only where the existing control doesn't cover it
4. Document mapping in the AIMS scope statement (Clause 4.3)
## Output Standards
```
**Bottom Line:** [one sentence — gap severity + the one thing to close first]
**The Decision:** [one of: gap-closure | risk-treatment | audit-scope]
**The Evidence:** [clause numbers + control IDs from the tool, not adjectives]
**How to Act:** [3 concrete next steps with owners + dates]
**Your Decision:** [the call only the compliance officer or CAIO can make — risk acceptance, scope expansion, certification readiness]
```
## Adjacent Skills
- `../../skills/information-security-manager-iso27001/` — ISO 27001 ISMS implementation (many controls reusable for AIMS A.7 data controls)
- `../../skills/quality-manager-qms-iso13485/` — ISO 13485 QMS (provides CAPA + management-review machinery the AIMS reuses)
- `../../skills/gdpr-dsgvo-expert/` — GDPR DPIA process (input to AIMS A.5 impact assessment for personal-data systems)
- `../../skills/isms-audit-expert/` — ISO 27001 internal audit pattern (the audit scheduler mirrors this for AIMS)
- `../../skills/soc2-compliance/` — SOC 2 trust services (reusable controls for AIMS A.10 third-party relationships)
- `../../../compliance-team-eu-ai-act/` — EU AI Act Article-level compliance (binding regulation companion to voluntary 42001)
- `../../../../compliance-os/` — Meta-orchestrator for multi-framework programs (run AIMS as one framework among 9)
- `../../../../c-level-advisor/chief-ai-officer-advisor/` — Executive AI strategy (build-vs-buy, cost economics — different audience)
## References
- [iso42001_clauses.md](references/iso42001_clauses.md) — Clauses 4–10 walkthrough with audit evidence expectations, common gaps, and reusable artifacts from ISO 27001/13485
- [aims_controls_annex_a.md](references/aims_controls_annex_a.md) — All 38 Annex A controls (A.2–A.10) with implementation guidance, audit evidence, and severity of failure
- [aims_implementation_guide.md](references/aims_implementation_guide.md) — 3-year maturity model (establish → certify → continually improve), rollout sequencing, integration with existing ISMS/QMS programs
- [cross_framework_mapping_ai.md](references/cross_framework_mapping_ai.md) — ISO 42001 ↔ EU AI Act ↔ NIST AI RMF ↔ ISO 23894 ↔ ISO 38507 ↔ ISO 27001 control-level mapping with mapping-confidence ratings
---
**Version:** 1.0.0
**Status:** Production Ready
FILE:references/aims_controls_annex_a.md
# ISO/IEC 42001 Annex A — 38 Controls Catalogue
This reference answers exactly one decision: **for each Annex A control, what does implementation look like, what evidence does the auditor want, and what's the severity if it's missing?**
Pair with `scripts/ai_risk_register_builder.py` to map risks to controls.
## Structure of Annex A
ISO/IEC 42001 Annex A is a *normative* annex containing reference controls. The standard requires (per Clause 6.1.3) that the organization compare its determined controls to Annex A to verify no necessary controls have been omitted. Unlike ISO 27001 where Annex A is presumed-applicable, ISO 42001 Annex A controls are applied based on risk — if a control doesn't apply (e.g., A.10 third-party AI when you use no third-party AI), document the exclusion with justification.
**The 10 control categories (A.1 is the structural intro; A.2–A.10 are the operational controls):**
| ID | Category | Control count | Maps to clause |
|---|---|---|---|
| A.2 | Policies related to AI | 2 | 5.2 |
| A.3 | Internal organization | 2 | 5.3 |
| A.4 | Resources for AI systems | 3 | 7.1 |
| A.5 | Assessing impacts of AI systems | 3 | 6.1.4, 8.2 |
| A.6 | AI system lifecycle | 8 | 8.3 |
| A.7 | Data for AI systems | 5 | 8.3 |
| A.8 | Information for interested parties | 4 | 7.4, 9.1 |
| A.9 | Use of AI systems | 4 | 8.3, 9.1 |
| A.10 | Third-party & customer relationships | 5 | 8.4 |
Total: **38 controls** across 9 operational categories.
## A.2 — Policies (severity if missing: CRITICAL)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.2.2** | AI policy | Signed AI policy meeting Clause 5.2 requirements | ISO 27001 A.5.1 (information security policy) — extend |
| **A.2.3** | Alignment of AI policy with other policies | Mapping showing AI policy doesn't contradict info-sec, privacy, quality, code-of-conduct policies | New artifact; document the cross-references |
## A.3 — Internal Organization (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.3.2** | AI roles & responsibilities | RACI matrix; named AIMS owner | ISO 27001 A.5.2; extend to AI |
| **A.3.3** | Reporting of concerns | Whistleblower / concerns procedure for AI-specific issues (bias, harm, misuse) | Existing whistleblower; AI-extend |
## A.4 — Resources (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.4.2** | Resources — data | Data inventory; provenance; quality assessment | ISO 27001 A.5.9 inventory of assets — extend |
| **A.4.3** | Resources — tooling | Inventory of ML tooling; license & dependency tracking | Existing software-asset management |
| **A.4.4** | Resources — human resources | Competence requirements + training records (Clause 7.2) | ISO 27001 A.6.3 awareness; ISO 13485 6.2 competence |
## A.5 — Impact Assessment (severity: CRITICAL)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.5.2** | AI system impact assessment | Documented impact assessment for each AI system; covers individuals, groups, society | GDPR DPIA — partial; AI scope wider (third-party harm, environmental, societal) |
| **A.5.3** | Process for impact assessment | Documented procedure with triggers (launch, material change, complaint) | New procedure |
| **A.5.4** | Documentation of impact assessment | Signed impact assessment record with management approval for high-impact systems | New artifact |
## A.6 — AI System Lifecycle (severity: CRITICAL)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.6.1.2** | Objectives for AI system development | Stated AI-system objectives aligned to AI policy + use intent | New artifact (per system) |
| **A.6.1.3** | Processes for management of the AI system lifecycle | Procedure covering design → data → model → V&V → deployment → operation → decommission | New procedure |
| **A.6.2.2** | AI system objectives & requirements | Documented requirements traceable to objectives | ISO 13485 7.3 design & development — extend |
| **A.6.2.3** | Documentation of AI system design & development | Design records (architecture, datasets, model card) under document control | ISO 13485 7.3 — extend |
| **A.6.2.4** | Verification & validation of AI system | Test plan + evaluation results; defined acceptance criteria | New artifact per system; reference NIST AI RMF "Measure" function |
| **A.6.2.5** | Deployment of AI system | Deployment checklist; environment hand-off; rollback plan | ISO 27001 A.8.32 change management — extend |
| **A.6.2.6** | Operation & monitoring of AI system | Monitoring plan with thresholds + escalation | New per system |
| **A.6.2.7** | Technical documentation of AI system | Model card or system card per Mitchell et al. (2019) / Gebru et al. (2021) | New artifact |
## A.7 — Data for AI Systems (severity: CRITICAL)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.7.2** | Data management | Data lifecycle procedure (acquisition → use → retention → deletion) | GDPR Art. 5 data minimisation; ISO 27001 A.5.10 acceptable use |
| **A.7.3** | Data quality | Defined data-quality dimensions; measured; reported | New; reference DAMA-DMBOK 2 / ISO 8000 |
| **A.7.4** | Data provenance | Documented data lineage; consent / legitimate basis recorded | GDPR records of processing (Art. 30) — extend |
| **A.7.5** | Data preparation | Documented preprocessing procedure | New artifact per system |
| **A.7.6** | Data privacy considerations | Privacy review per data category | GDPR DPIA — extend |
## A.8 — Information for Interested Parties (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.8.2** | System documentation | Public-facing documentation per Annex A.6.2.7 | Model card / system card |
| **A.8.3** | User information | UX-level disclosure: this is AI; what it does; its limitations | New; align with EU AI Act Article 50 transparency |
| **A.8.4** | Communication of AI incidents | Incident communication procedure including external notification timing | GDPR Art. 33–34 breach notification — extend |
| **A.8.5** | Information for affected parties | Communication for AI-affected populations (those subject to AI decisions) | New; align with EU AI Act Article 86 redress |
## A.9 — Use of AI Systems (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.9.2** | Intended use of AI system | Documented intended-use statement per system | New artifact |
| **A.9.3** | Monitoring of operation | Continuous monitoring with defined metrics + thresholds | NIST AI RMF "Measure" — extend |
| **A.9.4** | Logging of AI system events | Tamper-evident logs covering decisions, drift indicators, incidents | ISO 27001 A.8.15 logging — extend |
| **A.9.5** | Use of system after deployment | Procedure for in-use changes (retraining, fine-tuning) with re-evaluation triggers | New procedure |
## A.10 — Third-Party & Customer Relationships (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.10.2** | Supplier (third-party) relationships | AI-specific contract clauses (training data use, drift notification, sub-processor list) | ISO 27001 A.5.19 supplier relationships — extend |
| **A.10.3** | Customer relationships | Customer-facing AI obligations (transparency, opt-out, redress) | ISO 27001 A.5.20 — extend |
| **A.10.4** | Allocation of responsibilities between organization & third party | RACI for shared AI responsibilities (data labeling, model training, hosting, monitoring) | New artifact (per supplier) |
| **A.10.5** | Confidentiality of AI-related information | NDA scope covers AI-system internals (architecture, training data, weights) | ISO 27001 A.6.6 confidentiality — extend |
| **A.10.6** | Termination of AI service relationships | Procedure for safe AI-vendor exit (data return, model deletion, monitoring transition) | ISO 27001 A.5.20 service-level review — extend |
## How to Read This Catalogue
- **CRITICAL** = nonconformity blocks certification at stage 1
- **MAJOR** = nonconformity requires corrective action plan at stage 2; may delay certification
- **MINOR** = nonconformity recorded; corrective action expected within agreed timeline
**Audit evidence rule:** for every control selected as applicable, the auditor will ask three questions: (1) Where is the documented procedure? (2) Where are the records showing the procedure was followed? (3) Where is the evidence of management review of those records? If any of the three is missing, the control is partially implemented.
## When This Reference Doesn't Help
- **Specific Annex A control text.** This is a summary. The normative text is in ISO/IEC 42001:2023 Annex A — buy the standard.
- **Risk-to-control mapping methodology.** See `aims_implementation_guide.md` and ISO/IEC 23894:2023.
- **EU AI Act control overlap.** See `cross_framework_mapping_ai.md`.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 42001:2023** — Annex A normative controls (the authoritative source)
- **ISO/IEC 23894:2023** — AI risk management process (drives Annex A selection)
- **ISO/IEC 22989:2022** — AI concepts and terminology
- **NIST AI Risk Management Framework 1.0** (Jan 2023) + AI RMF Playbook — operational guidance mapping cleanly to Annex A
- **BSI AIC4 — Artificial Intelligence Cloud Service Compliance Criteria Catalogue** (2021) — sector-specific overlay for cloud AI providers
- **AAMI CR34971:2023** — Guidance for AI in medical devices
- **Mitchell et al.** — "Model Cards for Model Reporting" (FAT* 2019) — origin of model-card pattern referenced by A.6.2.7
- **Gebru et al.** — "Datasheets for Datasets" (CACM 2021) — datasheet pattern referenced by A.7.4
- **ISACA** — *Auditing Artificial Intelligence* (2nd ed., 2024) — practitioner audit checklist
FILE:references/aims_implementation_guide.md
# ISO/IEC 42001 — AIMS Implementation Guide (3-Year Maturity Model)
This reference answers exactly one decision: **what's the rollout sequence — what do we build in year 1 vs year 2 vs year 3, and how do we avoid recreating ISO 27001/13485 machinery?**
Pair with `scripts/aims_audit_scheduler.py` to operationalize the year-by-year audit cycle.
## The 3-Year Cycle
ISO management-system certifications follow a 3-year cycle:
| Year | Audit type | What happens |
|---|---|---|
| **Year 1** | Stage 1 (documentation review) + Stage 2 (implementation audit) → initial certification | Establish the AIMS; close major nonconformities; pass certification |
| **Year 2** | Surveillance audit (selective scope) | Demonstrate continual improvement; close minor nonconformities from year 1 |
| **Year 3** | Surveillance audit (selective scope) + recertification preparation | Full system review; prepare for year 4 recertification |
| **Year 4** | Recertification audit (full scope) | Renew certificate |
The internal audit programme (Clause 9.2) must cover every clause + every applicable Annex A control at least once per 3-year cycle. The plan must show this rolling coverage.
## Year 1 — Establish (focus: artifacts that auditors must see)
**Goal:** every clause and every applicable Annex A control has at least a documented procedure and one round of records.
### Q1: Foundations
- AI policy (Clause 5.2 + A.2.2) — board-signed
- AIMS scope statement (Clause 4.3) — names every AI system including third-party
- Roles & responsibilities (Clause 5.3 + A.3.2) — RACI with named AIMS owner
- Stakeholder & context analysis (Clause 4.1–4.2)
### Q2: Risk & impact
- AI risk register (Clause 6.1.2 + A.5) — run `ai_risk_register_builder.py`
- Risk treatment plan (Clause 6.1.3) — every high/critical risk linked to ≥ 1 Annex A control
- Impact assessment procedure (Clause 6.1.4 + A.5.3)
- AI objectives (Clause 6.2) — measurable targets
### Q3: Operations
- AI system lifecycle procedure (Clause 8.3 + A.6) — design through decommission
- Data management procedures (A.7) — data quality, provenance, preparation
- Monitoring plan per system (A.9.3)
- Third-party AI contract template (A.10.2)
### Q4: Performance
- Internal audit programme (Clause 9.2) — run `aims_audit_scheduler.py`
- Management review procedure (Clause 9.3) — inputs include AI-specific items
- CAPA integration with existing 13485/9001 CAPA loop (Clause 10.2)
- Stage 1 audit readiness check — run `aims_gap_analyzer.py`
**Year 1 success criteria:** stage 1 audit passes with 0 critical and ≤ 1 major nonconformity.
## Year 2 — Certify and operate
**Goal:** close year-1 minor nonconformities; demonstrate the system is operating, not just documented.
### Focus shifts to records (evidence the procedures are followed)
- Monthly drift monitoring records (A.9.3)
- Quarterly impact assessment reviews (A.5)
- Half-yearly third-party AI supplier reviews (A.10.2)
- Annual management review (Clause 9.3) with documented AI-specific inputs:
- Risk register changes
- Open nonconformities
- Drift events outside threshold
- Incidents per A.8.4
- Performance trends vs objectives (Clause 6.2)
**Year 2 success criteria:** surveillance audit passes; year-1 nonconformities closed; ≥ 80% of risk-register treatments fully implemented.
## Year 3 — Continually improve
**Goal:** demonstrate continual improvement (Clause 10.1) and prepare for recertification.
- Annual update to risk register based on new AI systems, regulation changes, incidents
- Re-baseline objectives (Clause 6.2) against year-1 + year-2 performance
- Audit the audit programme itself (meta-audit; common surveillance finding)
- Demonstrate at least one improvement initiative closed with measurable result
**Year 3 success criteria:** surveillance audit passes; recertification scope confirmed; trend evidence supports continual improvement claim.
## Integration With Existing ISMS (ISO 27001) and QMS (ISO 13485 / 9001)
The mistake most organizations make: building the AIMS as a parallel management system. **Don't.** ISO 42001 is intentionally Annex SL aligned to allow integration. Common integration patterns:
| Existing artifact | Extend for AIMS by adding |
|---|---|
| ISMS scope statement | List of AI systems within ISMS scope |
| Information security policy | AI-specific commitments (fairness, human oversight) |
| Risk register (27001) | AI risks tagged distinctly; same severity matrix; same treatment workflow |
| Document control procedure | Add model cards + datasheets + impact assessments to controlled documents |
| Internal audit programme | Add AI clause + Annex A controls to rotation |
| Management review | Add AI inputs (drift, incidents, risk-register changes) |
| CAPA procedure | Add AI-specific root-cause categories (data quality, model drift, prompt injection) |
| Supplier management | Add AI-specific contract clauses |
| Incident response | Add AI incidents (bias surfaced, drift exceeded, model misuse) |
**Reuse rule of thumb:** if you already operate ISO 27001 + ISO 13485 maturely, ~60% of AIMS Clauses 4–10 effort is rewriting existing artifacts to include AI scope. The remaining ~40% is Annex A operational controls (risk register details, lifecycle, V&V, monitoring, model cards) which are genuinely new.
## Sequence If Starting From Zero (No Prior Management System)
If your organization is starting AIMS without prior ISO certification:
1. **Add ISO 27001 first.** Most AIMS Clauses 4–10 evidence is satisfied by ISO 27001 evidence with AI scope appended. Doing 42001 alone is harder.
2. **Or start with NIST AI RMF.** NIST AI RMF is voluntary and US-centric but maps cleanly to 42001 Annex A. Mature on RMF for 12–18 months, then layer the management-system formality of 42001 on top.
3. **Avoid: building AIMS in isolation.** You'll recreate document control, CAPA, management review, and internal audit infrastructure that ISO 27001/13485 already standardize.
## Cost & Effort Benchmarks (informal, practitioner-reported)
| Org type | Year 1 effort (FTE-months) | Notes |
|---|---|---|
| Mature 27001 + 13485 org adding AIMS | 4–6 | Mostly Annex A overlay |
| Mature 27001 org adding AIMS (no 13485) | 8–12 | Add lifecycle procedures (A.6) net-new |
| Greenfield (no prior management system) | 24–36 | Do 27001 first, then 42001 |
Certification body fees: ~$15k–$35k for initial certification audit (stage 1 + stage 2 for a typical mid-size SaaS); ~$8k–$15k per surveillance year.
## Common Year-1 Pitfalls
1. **Treating "AI ethics" as the policy.** A poetic policy doesn't pass; auditor wants concrete commitments and a way to verify them.
2. **Risk register with no control mapping.** Register identifies risks but doesn't show which Annex A control treats each — Clause 6.1.3 fails.
3. **Lifecycle procedure that skips decommission.** Auditor will ask, "How do you safely retire an AI system?" If silence, A.6 fails.
4. **No drift threshold defined.** Monitoring "we watch it" doesn't pass; needs metric + threshold + escalation owner.
5. **Third-party AI excluded.** "Our vendors' AI features aren't ours" is wrong if you embed them in your service.
6. **No competence requirement for ML engineers.** Clause 7.2 wants documented competence requirements per role; "they have PhDs" isn't a documented requirement.
## When This Reference Doesn't Help
- **Specific Annex A control implementation.** See `aims_controls_annex_a.md`.
- **Risk identification methodology.** See ISO/IEC 23894:2023.
- **EU AI Act overlap.** See `cross_framework_mapping_ai.md` and `compliance-team-eu-ai-act/`.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 42001:2023** — the standard itself
- **ISO/IEC 23894:2023** — AI risk management process
- **ISO/IEC 38507:2022** — Governance implications of AI for organizations
- **ISO/IEC 27001:2022** — Information security management (reuse template for 60% of AIMS Clauses 4–10)
- **ISO/IEC 13485:2016** — Medical device QMS (reuse template for CAPA, document control)
- **NIST AI RMF 1.0** (Jan 2023) + AI RMF Playbook + Generative AI Profile (NIST AI 600-1, 2024)
- **BSI** — *Information technology — Artificial intelligence — Implementation guidance for ISO/IEC 42001* (2024 white paper)
- **ISACA** — *Auditing Artificial Intelligence* (2nd ed., 2024) — implementation pitfalls catalogue
- **IAPP** — AI Governance Center materials (continuously updated) — practitioner community knowledge base
FILE:references/cross_framework_mapping_ai.md
# ISO/IEC 42001 ↔ EU AI Act ↔ NIST AI RMF ↔ ISO 23894 ↔ ISO 38507 ↔ ISO 27001 — Cross-Framework Mapping
This reference answers exactly one decision: **for each ISO 42001 obligation, which other frameworks already cover it, and what evidence can I reuse?**
The point of cross-framework mapping is to avoid duplicate work. A control implemented for ISO 27001 frequently satisfies an Annex A control of ISO 42001 with minor AI-specific overlay. The `compliance-os` orchestrator's `cross_framework_mapper.py` consumes this mapping.
## High-Level Framework Comparison
| Framework | Type | Binding? | AI scope | Maturity |
|---|---|---|---|---|
| **ISO/IEC 42001:2023** | Management system standard | Voluntary; certifiable | AI Management System (AIMS) | Published 2023; certifications starting 2024 |
| **EU AI Act (Reg. 2024/1689)** | Product safety regulation | Binding in EU | Risk-based: prohibited → high-risk → limited-risk → minimal-risk | In force Aug 2024; phased obligations through 2027 |
| **NIST AI RMF 1.0** | Risk management framework | Voluntary (US) | Govern / Map / Measure / Manage functions | Released Jan 2023; mature playbook |
| **ISO/IEC 23894:2023** | Risk management methodology | Reference standard | AI risk process; informs 42001 Clause 6.1 | Published 2023 |
| **ISO/IEC 38507:2022** | Governance standard | Reference standard | Board-level AI governance | Published 2022 |
| **ISO/IEC 27001:2022** | Management system standard | Voluntary; certifiable | Information security | Mature; widely certified |
## Clause-to-Framework Mapping (ISO 42001 lens)
### Clause 4 — Context
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 | Notes |
|---|---|---|---|---|
| 4.1 External context | Art. 1 (scope); Recitals on risk-based approach | GOVERN 1.1 | 4.1 | Extend 27001 context with AI regulatory landscape |
| 4.2 Interested parties | Art. 27 (FRIA stakeholders for high-risk) | GOVERN 5 | 4.2 | Add AI-affected populations |
| 4.3 Scope | Article 6 + Annex III define what's in scope as "high-risk" | MAP 1.1 | 4.3 | Distinct artifacts; AIMS scope ≠ EU AI Act applicability scope |
| 4.4 AIMS processes | n/a | n/a | 4.4 | Integration map |
### Clause 5 — Leadership
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 / ISO 38507 |
|---|---|---|---|
| 5.1 Top-mgmt commitment | Art. 26 (deployer obligations); Art. 16 (provider obligations) | GOVERN 1 | 27001 5.1; 38507 Clauses 5–6 (governance principles) |
| 5.2 AI policy | Art. 17 (QMS for high-risk); Art. 95 (codes of conduct) | GOVERN 1.1 | 27001 5.2 — extend with AI commitments |
| 5.3 Roles & authorities | Art. 26 (deployer obligations); Art. 16 + 22 (authorized representative) | GOVERN 2.1 | 27001 5.3 |
### Clause 6 — Planning (the densest mapping)
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 23894 |
|---|---|---|---|
| 6.1.2 AI risk assessment | Art. 9 (risk management system for high-risk) | MAP 5.1; MAP 5.2 | Clauses 6–7 (entire process) |
| 6.1.3 AI risk treatment | Art. 9(2)(c–d) (risk management measures) | MANAGE 1.1 | Clauses 8 (treatment selection) |
| 6.1.4 Impact assessment | Art. 27 (Fundamental Rights Impact Assessment for high-risk public-sector deployers) | MAP 2.3; MAP 5.1 | Clause 5.3 (scope definition) |
| 6.2 AI objectives | Art. 9(2)(a) (objectives of risk management) | GOVERN 1.5; MEASURE 1 | Clause 5.2 |
### Clause 7 — Support
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 |
|---|---|---|---|
| 7.1 Resources | Art. 17(1)(c) (technical resources for QMS) | GOVERN 3 | A.6.1 |
| 7.2 Competence | Art. 14 (human oversight competence); Art. 26(2) (deployer competence) | GOVERN 3.1 | A.6.3 |
| 7.3 Awareness | Art. 14 | GOVERN 5.1 | A.6.3 |
| 7.4 Communication | Art. 50 (transparency obligations); Art. 86 (right to explanation) | GOVERN 5.2 | A.7.4 |
| 7.5 Documented info | Art. 11 + 12 (technical documentation); Art. 19 (record-keeping) | GOVERN 1.4 | 27001 7.5 |
### Clause 8 — Operation
| ISO 42001 | EU AI Act | NIST AI RMF | Notes |
|---|---|---|---|
| 8.1 Operational planning | Art. 17 (QMS) | MANAGE 2 | |
| 8.2 Impact assessment process | Art. 27 (FRIA process) | MAP 2 | |
| 8.3 AI system lifecycle | Art. 9 (full lifecycle); Art. 72 (post-market monitoring) | MAP 3; MEASURE 3; MANAGE 4 | Densest overlap |
| 8.4 Third-party / customer | Art. 25 (responsibilities along the AI value chain) | GOVERN 6 | |
### Clause 9 — Performance
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 |
|---|---|---|---|
| 9.1 Monitoring | Art. 72 (post-market monitoring system) | MEASURE 2; MEASURE 4 | 9.1 |
| 9.2 Internal audit | Art. 17(1)(j) (internal audit as part of QMS) | GOVERN 4 | 9.2 |
| 9.3 Management review | n/a explicit; implied in Art. 17 | GOVERN 1 | 9.3 |
### Clause 10 — Improvement
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 |
|---|---|---|---|
| 10.1 Continual improvement | Art. 9(2)(c) (iterative risk reduction) | MANAGE 4.3 | 10.1 |
| 10.2 Nonconformity & CAPA | Art. 73 (incident reporting); Art. 79 (corrective actions) | MANAGE 4.2 | 10.2 |
## Annex A Control → Framework Mapping (subset of highest-value mappings)
| ISO 42001 Annex A | EU AI Act | NIST AI RMF | ISO 27001 | Mapping confidence |
|---|---|---|---|---|
| A.2.2 AI policy | Art. 95 (codes of conduct) | GOVERN 1.1 | A.5.1 (info-sec policy) | HIGH |
| A.5.2 Impact assessment | Art. 27 FRIA | MAP 2.3 | n/a | MEDIUM (FRIA narrower) |
| A.6.2.4 V&V | Art. 15 (accuracy, robustness, cybersecurity); Art. 17(1)(h) | MEASURE 2 | n/a | HIGH |
| A.7.2 Data management | Art. 10 (data governance) | MAP 2.3; MEASURE 2.6 | A.5.10 | HIGH |
| A.7.3 Data quality | Art. 10(3) (relevance, representativeness, error-free, complete) | MEASURE 2.6 | n/a | HIGH |
| A.7.4 Data provenance | Art. 10(2)(d) (data origin) | MAP 2.3 | n/a | HIGH |
| A.7.6 Data privacy | Art. 10(5) (special categories); GDPR Articles 5, 6, 9 | MANAGE 2.1 | A.5.34 | HIGH |
| A.8.2 System docs | Art. 11 + Annex IV (technical documentation) | GOVERN 1.4 | A.5.37 | HIGH |
| A.8.3 User information | Art. 13 (instructions for use); Art. 50 (transparency) | GOVERN 5.2 | n/a | HIGH |
| A.8.4 Incident communication | Art. 73 (incident reporting to authorities) | MANAGE 4.2 | A.6.8 (reporting) | HIGH |
| A.9.3 Monitoring | Art. 72 (post-market monitoring) | MEASURE 2; MEASURE 4 | A.8.15 (logging) | HIGH |
| A.9.4 Logging | Art. 12 (record-keeping); Art. 19 | MEASURE 4 | A.8.15 | HIGH |
| A.10.2 Supplier relationships | Art. 25 (responsibilities along the AI value chain) | GOVERN 6 | A.5.19, A.5.20, A.5.21 | HIGH |
**Mapping confidence legend:**
- **HIGH** — direct overlap; same evidence can satisfy both
- **MEDIUM** — partial overlap; existing evidence with AI overlay
- **LOW** — concept overlap; mostly new artifact required
## Practical Reuse Pattern
If you operate ISO 27001 (mature) + are adopting ISO 42001:
1. **Reuse policies (~60%):** Extend info-sec policy with AI commitments (5.2 + A.2.2)
2. **Reuse procedures (~50%):** Document control, internal audit, management review, CAPA
3. **Reuse risk machinery (~70%):** Same severity matrix, same treatment workflow, same residual-risk acceptance flow — just add AI-specific risks and Annex A control mapping
4. **Reuse supplier mgmt (~80%):** Add AI-specific contract clauses to existing supplier procedure
5. **New artifacts (~40%):** Model cards / datasheets (A.6.2.7, A.7.4), impact assessments per Annex A.5, lifecycle procedure (A.6), drift monitoring (A.9.3), V&V procedure (A.6.2.4)
If you also operate ISO 13485 (medical device QMS):
- Reuse: design controls (7.3) for A.6 lifecycle; risk management (ISO 14971) overlays cleanly onto A.5 + 6.1; post-market surveillance maps directly to A.9.3 monitoring
- Add: AI-specific failure modes to ISO 14971 hazard analysis
## When This Reference Doesn't Help
- **EU AI Act conformity assessment routing.** See `compliance-team-eu-ai-act/scripts/conformity_assessment_planner.py`.
- **NIST AI RMF deep-dive.** See NIST AI RMF Playbook (NIST.AI.100-1.pdf) and Generative AI Profile (NIST.AI.600-1).
- **Multi-framework audit simulation.** See `compliance-os/scripts/audit_simulator.py`.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 42001:2023** — Annex A normative controls
- **Regulation (EU) 2024/1689** — Artificial Intelligence Act — full Articles (the binding regulation)
- **NIST AI Risk Management Framework 1.0** (Jan 2023, NIST AI 100-1) + AI RMF Playbook
- **ISO/IEC 23894:2023** — AI risk management process
- **ISO/IEC 38507:2022** — Governance implications of AI
- **ISO/IEC 27001:2022** + Annex A controls (the most cross-walked partner standard)
- **EDPB Opinion 28/2024** — Guidelines on processing of personal data in AI models
- **European Commission AI Act Guidelines** (continuously updated): Guidelines on prohibited practices (Feb 2025), Guidelines on definition of AI system (Feb 2025), FRIA template guidance
- **BSI** — *Cross-walking ISO 42001 and EU AI Act* (white paper, 2024)
- **IAPP EU AI Act Tracker** (continuously updated) — practitioner reference for Article applicability
FILE:references/iso42001_clauses.md
# ISO/IEC 42001:2023 — Clauses 4-10 Walkthrough
This reference answers exactly one decision: **for each clause of ISO 42001, what audit evidence does the certification body expect, and which existing ISMS/QMS artifact can I reuse?**
Pair with `scripts/aims_gap_analyzer.py` for automated coverage scoring.
## Annex SL High-Level Structure
ISO/IEC 42001:2023 follows the Annex SL structure shared by ISO 9001, 14001, 27001, 13485, 45001, and other management-system standards. This is deliberate: certification bodies, internal auditors, and quality teams can apply existing competencies to AIMS audits with low ramp-up cost.
**Practical implication:** if your organization already operates ISO 27001 + ISO 13485, ~60% of Clauses 4–10 artefacts (scope statements, policies, document control, internal audit programme, management review) can be **extended** to cover AI scope rather than recreated. The gap analysis is mostly Annex A (AI-specific operational controls), not Clauses 4–10.
## Clause 4 — Context of the Organization
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **4.1** | External & internal issues affecting AIMS | Documented context analysis (PESTLE or equivalent); reviewed at management review | Treating AI regulatory landscape as static; missing EU AI Act, US state laws, sector-specific AI rules |
| **4.2** | Needs & expectations of interested parties | Stakeholder matrix: customers, regulators, employees, data subjects, model providers, AI-affected populations | Omitting "AI-affected populations" (people who never interact with the system but are subject to its decisions) |
| **4.3** | AIMS scope statement | Documented scope: which AI systems, which lifecycle phases, which organizational units, which exclusions | Scope omits third-party AI services (SaaS features powered by vendor models); excludes "experimental" systems that are in fact in production |
| **4.4** | AIMS processes & interactions | Process map showing how AIMS processes connect to existing QMS/ISMS processes | Treating AIMS as parallel system instead of integrated extension of existing management systems |
**Reusable from ISO 27001 / 13485:** scope statement template, stakeholder matrix template, process map.
## Clause 5 — Leadership
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **5.1** | Top-management commitment | Documented evidence: AI in board agenda, resource allocation, KPIs | "AI ethics" reduced to marketing copy with no operating commitment |
| **5.2** | AI policy | Signed AI policy committing to lawful use, beneficial purpose, human oversight, continual improvement | Policy doesn't mention human oversight (Annex A.9 requirement); missing commitment to continual improvement |
| **5.3** | Organizational roles, responsibilities, authorities | RACI matrix for AIMS roles; named AIMS owner; AI ethics review board (if applicable) | No named AIMS owner; CISO assumed to "cover AI" without explicit assignment |
**Critical:** Clause 5.2 has a higher evidence bar than ISO 27001/13485 because the AI policy must address fairness, transparency, and human oversight — concepts absent from older management systems. Cannot be satisfied by extending existing policies; needs net-new content.
## Clause 6 — Planning
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **6.1.2** | AI risk assessment | Risk register per ISO 23894 methodology; covers full AI lifecycle | Risk identification at deployment only, missing data + model + decommission phases |
| **6.1.3** | AI risk treatment | Treatment plan linking each risk to Annex A controls; residual-risk acceptance documented | Treatment plan exists but is generic ("apply A.7.3") without specific implementation |
| **6.1.4** | AI system impact assessment | Documented impact assessment per Annex A.5.2 for high-impact systems | Confusing impact assessment (Clause 6.1.4) with risk assessment (Clause 6.1.2) |
| **6.2** | AI objectives | Measurable AI objectives aligned to AI policy; reviewed in management review | Objectives are aspirational ("ethical AI") without measurable targets |
**Run** `ai_risk_register_builder.py` to operationalize 6.1.2 + 6.1.3.
## Clause 7 — Support
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **7.1** | Resources for AIMS | Budget; tooling; compute resources documented | Compute resources for ML training treated as one-off project cost, not ongoing AIMS resource |
| **7.2** | Competence | Defined competence requirements per role (ML eng, AI risk, data steward); training records | Competence requirements undefined for ML engineers; assumes "they have degrees" |
| **7.3** | Awareness | AI awareness training across all employees with AI-system access | Training is engineer-only; product, marketing, customer success bypass |
| **7.4** | Communication | Documented internal + external communications procedure for AI | No procedure for communicating AI incidents to users (Annex A.8.4 link) |
| **7.5** | Documented information | Version-controlled AIMS documentation | Model cards exist but are not under document control; can be edited without approval |
## Clause 8 — Operation
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **8.1** | Operational planning & control | Operational procedures for each AI lifecycle phase | Operations procedures don't define phase transitions (when does "development" become "production"?) |
| **8.2** | Impact assessment process | Operational procedure for triggering impact assessment; gate before launch | Impact assessment treated as one-time launch artifact, not re-triggered on material change |
| **8.3** | AI system lifecycle process | Documented lifecycle covering: design → data → model → V&V → deployment → operation → decommission | Lifecycle skips "decommission"; no procedure for sunsetting AI systems |
| **8.4** | Third-party / customer relationships | Supplier and customer relationship procedures; AI-specific clauses in contracts | Standard vendor contracts not updated for AI-specific obligations (data use, model retraining, drift) |
## Clause 9 — Performance Evaluation
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **9.1** | Monitoring, measurement, analysis & evaluation | Defined metrics for AI performance, fairness, drift; monitoring records | Drift monitoring in code but no defined acceptable drift threshold; no escalation path |
| **9.2** | Internal audit programme | 12-month audit plan; auditor independence documented; findings tracked | No formal AIMS audit programme; audits happen ad hoc; auditors audit own work |
| **9.3** | Management review | Documented management review at planned intervals with required inputs/outputs | Management review inputs missing AI-specific items (drift, incidents, risk-register changes) |
**Run** `aims_audit_scheduler.py` to generate the 9.2 plan with independence checks.
## Clause 10 — Improvement
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **10.1** | Continual improvement | Evidence of AIMS improvement over time (KPIs trending, control maturity rising) | "Continual improvement" treated as audit closure activity, not ongoing |
| **10.2** | Nonconformity & corrective action | CAPA records for AIMS nonconformities; root cause analysis documented | AIMS CAPA loop separate from existing 13485/9001 CAPA loop — duplicated effort, divergent procedures |
**Reusable from ISO 13485 / 9001:** the entire CAPA machinery. Add AI-specific root-cause categories (data quality, model drift, prompt injection, etc.) to the existing taxonomy.
## When This Reference Doesn't Help
- **Specific AI risk identification.** See `aims_controls_annex_a.md` and ISO/IEC 23894:2023.
- **EU AI Act conformity assessment.** Different standard. See `compliance-team-eu-ai-act`.
- **Model cards, datasheets, evaluation methodology.** Tactical artefacts; reference NIST AI RMF playbook + papers like Mitchell et al. (2019).
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 42001:2023** — Information technology — Artificial intelligence — Management system (the standard itself; published 2023-12-18 by ISO/IEC JTC 1/SC 42)
- **ISO/IEC 23894:2023** — AI risk management process (the methodology referenced by Clause 6.1.2)
- **ISO/IEC 38507:2022** — Governance implications of AI for organizations (board-level governance lens referenced by Clause 5)
- **ISO/IEC 22989:2022** — AI concepts and terminology (definitions used throughout)
- **Annex SL** in the ISO/IEC Directives Part 1 (2024) — the high-level structure shared by ISO management-system standards
- **BSI AI Management System (AIMS) Implementation Guide** (BSI, 2024) — practitioner walkthrough
- **AAMI CR34971:2023** — AI guidance for medical devices (cross-walks 42001 to medical device QMS)
- **ISACA** — *Auditing Artificial Intelligence* (2nd ed., 2024) — internal-audit-oriented checklist with ISO 42001 mapping
FILE:scripts/aims_audit_scheduler.py
#!/usr/bin/env python3
"""aims_audit_scheduler.py — ISO/IEC 42001 Clause 9.2 internal audit plan generator.
Stdlib-only. Produces a 12-month internal audit schedule for an AIMS with:
- quarterly audit slots
- clause + Annex A control coverage per slot
- auditor assignments with independence checks (no self-audit)
- rolling 3-year coverage to ensure every clause + applicable control is audited
- prior-year nonconformity follow-up scheduled in Q1
Deterministic logic. No LLM calls. Stdlib only.
Input schema (JSON):
{
"organization": "Acme AI Inc.",
"audit_year": 2026,
"certification_cycle_phase": "year_2", # year_1 | year_2 | year_3 | surveillance
"ai_systems_in_scope": ["recommendation_engine", "internal_llm_tools", "vendor_ai_chatbot"],
"applicable_annex_a_controls": ["A.2.2", "A.3.2", "A.5.2", "A.6.2.4", "A.7.3", "A.8.4", "A.9.3", "A.10.2"],
"auditors": [
{"id": "alice", "name": "Alice Chen", "role": "quality_engineer", "owns_clauses": ["8.3"]},
{"id": "bob", "name": "Bob Singh", "role": "ml_engineer", "owns_clauses": ["8.3", "A.6.2.4"]},
{"id": "carol", "name": "Carol Diaz", "role": "external_auditor", "owns_clauses": []},
{"id": "dave", "name": "Dave Park", "role": "ciso", "owns_clauses": ["A.10.2"]}
],
"prior_year_findings": [
{"clause": "9.2", "severity": "major", "status": "open"},
{"clause": "A.7.3", "severity": "minor", "status": "closed"}
]
}
Usage:
python aims_audit_scheduler.py
python aims_audit_scheduler.py path/to/scope.json
python aims_audit_scheduler.py scope.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"organization": "Acme AI Inc.",
"audit_year": 2026,
"certification_cycle_phase": "year_2",
"ai_systems_in_scope": ["recommendation_engine", "internal_llm_tools", "vendor_ai_chatbot"],
"applicable_annex_a_controls": [
"A.2.2", "A.3.2", "A.5.2", "A.6.2.4", "A.7.3", "A.8.4", "A.9.3", "A.10.2"
],
"auditors": [
{"id": "alice", "name": "Alice Chen", "role": "quality_engineer", "owns_clauses": ["8.3"]},
{"id": "bob", "name": "Bob Singh", "role": "ml_engineer", "owns_clauses": ["8.3", "A.6.2.4"]},
{"id": "carol", "name": "Carol Diaz", "role": "external_auditor", "owns_clauses": []},
{"id": "dave", "name": "Dave Park", "role": "ciso", "owns_clauses": ["A.10.2"]},
],
"prior_year_findings": [
{"clause": "9.2", "severity": "major", "status": "open"},
{"clause": "A.7.3", "severity": "minor", "status": "closed"},
],
}
# Always-audit clauses (full coverage every year)
ANNUAL_CLAUSES = ["4.3", "5.1", "5.2", "5.3", "9.3", "10.2"]
# 3-year rotation for deep-dive clauses
ROTATION_Q2 = ["6.1.2", "6.1.3", "6.1.4", "6.2"]
ROTATION_Q3 = ["7.1", "7.2", "7.3", "7.4", "7.5", "8.1", "8.2", "8.3", "8.4"]
ROTATION_Q4 = ["9.1", "9.2", "10.1"]
def assign_auditor(scope_items: List[str], auditors: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Pick the auditor with the fewest independence conflicts in this scope."""
best_auditor = None
best_conflicts = 999
for a in auditors:
owns = set(a.get("owns_clauses", []))
conflicts = sum(1 for s in scope_items if s in owns)
if conflicts < best_conflicts:
best_conflicts = conflicts
best_auditor = a
if best_auditor is None:
return {"id": None, "name": "UNASSIGNED", "independent": False, "conflicts": []}
owns = set(best_auditor.get("owns_clauses", []))
conflicts = [s for s in scope_items if s in owns]
return {
"id": best_auditor["id"],
"name": best_auditor["name"],
"role": best_auditor["role"],
"independent": len(conflicts) == 0,
"conflicts": conflicts,
}
def build_quarter(label: str, scope_clauses: List[str], scope_controls: List[str],
auditors: List[Dict[str, Any]], extra_notes: str = "") -> Dict[str, Any]:
all_scope = scope_clauses + scope_controls
auditor = assign_auditor(all_scope, auditors)
return {
"quarter": label,
"scope_clauses": scope_clauses,
"scope_annex_a_controls": scope_controls,
"auditor": auditor,
"notes": extra_notes,
}
def plan(payload: Dict[str, Any]) -> Dict[str, Any]:
year = int(payload.get("audit_year", 2026))
phase = payload.get("certification_cycle_phase", "year_2")
systems = payload.get("ai_systems_in_scope", [])
controls = payload.get("applicable_annex_a_controls", [])
auditors = payload.get("auditors", [])
prior_findings = payload.get("prior_year_findings", [])
open_priors = [f for f in prior_findings if f.get("status") != "closed"]
# 3-year control rotation: split applicable controls into thirds
third = max(1, len(controls) // 3)
controls_y1 = controls[0:third]
controls_y2 = controls[third:2 * third]
controls_y3 = controls[2 * third:]
phase_to_controls = {
"year_1": controls_y1, "year_2": controls_y2,
"year_3": controls_y3, "surveillance": controls_y3,
}
this_year_controls = phase_to_controls.get(phase, controls_y2)
# Q1: leadership + scope + prior-year follow-up
q1_clauses = ["4.3", "5.1", "5.2", "5.3"]
q1_notes = f"Follow up {len(open_priors)} open prior-year finding(s)." if open_priors else "No open priors."
q1 = build_quarter(f"Q1 {year}", q1_clauses, [], auditors, q1_notes)
# Q2: planning + objectives + risk
q2 = build_quarter(f"Q2 {year}", ROTATION_Q2, this_year_controls[:max(1, len(this_year_controls) // 2)], auditors)
# Q3: support + operation
q3_controls = this_year_controls[max(1, len(this_year_controls) // 2):]
q3_notes = f"Deep-dive across {len(systems)} AI systems: {', '.join(systems)}."
q3 = build_quarter(f"Q3 {year}", ROTATION_Q3, q3_controls, auditors, q3_notes)
# Q4: performance + improvement + management review
q4_notes = "Management review inputs prepared per Clause 9.3."
q4 = build_quarter(f"Q4 {year}", ROTATION_Q4 + ANNUAL_CLAUSES[-2:], [], auditors, q4_notes)
# Independence audit
quarters = [q1, q2, q3, q4]
independence_issues = [{
"quarter": q["quarter"], "auditor": q["auditor"]["name"], "conflicts": q["auditor"]["conflicts"]
} for q in quarters if not q["auditor"]["independent"]]
# Coverage check
audited_clauses = set()
audited_controls = set()
for q in quarters:
audited_clauses.update(q["scope_clauses"])
audited_controls.update(q["scope_annex_a_controls"])
return {
"organization": payload.get("organization"),
"audit_year": year,
"certification_cycle_phase": phase,
"ai_systems_in_scope": systems,
"open_prior_findings": len(open_priors),
"quarters": quarters,
"independence_issues": independence_issues,
"coverage_summary": {
"clauses_audited_this_year": sorted(audited_clauses),
"controls_audited_this_year": sorted(audited_controls),
"controls_deferred_to_future_years": sorted(
set(controls) - audited_controls
),
},
}
def render_text(p: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("ISO/IEC 42001 — CLAUSE 9.2 INTERNAL AUDIT PLAN")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Organization: {p['organization']}")
lines.append(f"Year: {p['audit_year']} | Cert cycle phase: {p['certification_cycle_phase']}")
lines.append(f"AI systems in scope: {', '.join(p['ai_systems_in_scope'])}")
lines.append(f"Open prior-year findings: {p['open_prior_findings']}")
lines.append("")
lines.append("-" * 72)
lines.append("QUARTERLY SCHEDULE:")
lines.append("")
for q in p["quarters"]:
a = q["auditor"]
flag = "" if a["independent"] else " ⚠️ INDEPENDENCE CONFLICT"
lines.append(f" {q['quarter']} → Auditor: {a['name']} ({a['role']}){flag}")
if q["scope_clauses"]:
lines.append(f" Clauses: {', '.join(q['scope_clauses'])}")
if q["scope_annex_a_controls"]:
lines.append(f" Annex A: {', '.join(q['scope_annex_a_controls'])}")
if a["conflicts"]:
lines.append(f" ⚠️ Conflicts on: {', '.join(a['conflicts'])} — reassign or use external auditor")
if q["notes"]:
lines.append(f" Notes: {q['notes']}")
lines.append("")
if p["independence_issues"]:
lines.append("-" * 72)
lines.append(f"INDEPENDENCE ISSUES ({len(p['independence_issues'])}):")
for issue in p["independence_issues"]:
lines.append(f" - {issue['quarter']}: {issue['auditor']} owns {', '.join(issue['conflicts'])}")
lines.append("")
c = p["coverage_summary"]
lines.append("-" * 72)
lines.append("3-YEAR COVERAGE STATUS:")
lines.append(f" Clauses audited this year ({len(c['clauses_audited_this_year'])}): {', '.join(c['clauses_audited_this_year'])}")
lines.append(f" Annex A controls audited this year ({len(c['controls_audited_this_year'])}): {', '.join(c['controls_audited_this_year']) or 'none'}")
lines.append(f" Controls deferred to future years ({len(c['controls_deferred_to_future_years'])}): {', '.join(c['controls_deferred_to_future_years']) or 'none'}")
lines.append("")
lines.append("RULES: every clause + every applicable Annex A control must be audited at least once per 3-year cert cycle.")
lines.append(" Same auditor cannot audit work they own (Clause 9.2 independence).")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="ISO/IEC 42001 Clause 9.2 internal audit 12-month plan generator.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to audit scope JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: year-2 cert cycle, 3 systems, 8 controls applicable>"
result = plan(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/aims_gap_analyzer.py
#!/usr/bin/env python3
"""aims_gap_analyzer.py — ISO/IEC 42001:2023 AIMS gap analysis against Clauses 4-10.
Stdlib-only. Scores each clause as 'full' / 'partial' / 'missing' based on an evidence
inventory and outputs a prioritized remediation list with severity at certification audit.
Deterministic logic. No LLM calls. No external dependencies.
Input schema (JSON):
{
"organization": "Acme AI Inc.",
"scope_statement": "Customer-facing recommendation engine + internal LLM tools",
"certification_target": "stage_1_audit_in_q3",
"evidence": {
"4.1_context_external": "documented",
"4.2_interested_parties": "documented",
"4.3_scope_statement": "documented",
"4.4_aims_processes": "partial",
"5.1_leadership_commitment": "documented",
"5.2_ai_policy": "partial",
"5.3_roles_responsibilities": "missing",
"6.1.2_risk_assessment": "documented",
"6.1.3_risk_treatment": "partial",
"6.1.4_impact_assessment": "missing",
"6.2_objectives": "documented",
"7.1_resources": "documented",
"7.2_competence": "missing",
"7.3_awareness": "partial",
"7.4_communication": "documented",
"7.5_documented_info": "documented",
"8.1_operational_planning": "documented",
"8.2_impact_assessment_process": "partial",
"8.3_ai_system_lifecycle": "missing",
"8.4_third_party_relationships": "partial",
"9.1_monitoring": "partial",
"9.2_internal_audit": "missing",
"9.3_management_review": "documented",
"10.1_continual_improvement": "partial",
"10.2_nonconformity_capa": "documented"
}
}
Usage:
python aims_gap_analyzer.py # uses embedded sample
python aims_gap_analyzer.py path/to/evidence.json
python aims_gap_analyzer.py evidence.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"organization": "Acme AI Inc.",
"scope_statement": "Customer-facing recommendation engine + internal LLM tools",
"certification_target": "stage_1_audit_in_q3",
"evidence": {
"4.1_context_external": "documented",
"4.2_interested_parties": "documented",
"4.3_scope_statement": "documented",
"4.4_aims_processes": "partial",
"5.1_leadership_commitment": "documented",
"5.2_ai_policy": "partial",
"5.3_roles_responsibilities": "missing",
"6.1.2_risk_assessment": "documented",
"6.1.3_risk_treatment": "partial",
"6.1.4_impact_assessment": "missing",
"6.2_objectives": "documented",
"7.1_resources": "documented",
"7.2_competence": "missing",
"7.3_awareness": "partial",
"7.4_communication": "documented",
"7.5_documented_info": "documented",
"8.1_operational_planning": "documented",
"8.2_impact_assessment_process": "partial",
"8.3_ai_system_lifecycle": "missing",
"8.4_third_party_relationships": "partial",
"9.1_monitoring": "partial",
"9.2_internal_audit": "missing",
"9.3_management_review": "documented",
"10.1_continual_improvement": "partial",
"10.2_nonconformity_capa": "documented",
},
}
# Clause requirements + severity if missing
# severity: 'critical' = major nonconformity at stage 1, blocks certification
# 'major' = major nonconformity at stage 2
# 'minor' = minor nonconformity, requires corrective action plan
# 'observation' = improvement opportunity
CLAUSE_REQUIREMENTS: Dict[str, Dict[str, Any]] = {
"4.1_context_external": {"clause": "4.1", "title": "External & internal context", "severity": "minor"},
"4.2_interested_parties": {"clause": "4.2", "title": "Interested parties", "severity": "minor"},
"4.3_scope_statement": {"clause": "4.3", "title": "AIMS scope statement", "severity": "critical"},
"4.4_aims_processes": {"clause": "4.4", "title": "AIMS processes & interactions", "severity": "major"},
"5.1_leadership_commitment": {"clause": "5.1", "title": "Leadership commitment", "severity": "major"},
"5.2_ai_policy": {"clause": "5.2", "title": "AI policy", "severity": "critical"},
"5.3_roles_responsibilities": {"clause": "5.3", "title": "Roles, responsibilities, authorities", "severity": "critical"},
"6.1.2_risk_assessment": {"clause": "6.1.2", "title": "AI risk assessment", "severity": "critical"},
"6.1.3_risk_treatment": {"clause": "6.1.3", "title": "AI risk treatment", "severity": "critical"},
"6.1.4_impact_assessment": {"clause": "6.1.4", "title": "AI system impact assessment", "severity": "major"},
"6.2_objectives": {"clause": "6.2", "title": "AI objectives & planning", "severity": "minor"},
"7.1_resources": {"clause": "7.1", "title": "Resources", "severity": "minor"},
"7.2_competence": {"clause": "7.2", "title": "Competence", "severity": "major"},
"7.3_awareness": {"clause": "7.3", "title": "Awareness", "severity": "minor"},
"7.4_communication": {"clause": "7.4", "title": "Communication", "severity": "minor"},
"7.5_documented_info": {"clause": "7.5", "title": "Documented information", "severity": "major"},
"8.1_operational_planning": {"clause": "8.1", "title": "Operational planning & control", "severity": "major"},
"8.2_impact_assessment_process": {"clause": "8.2", "title": "Impact assessment process", "severity": "major"},
"8.3_ai_system_lifecycle": {"clause": "8.3", "title": "AI system lifecycle process", "severity": "critical"},
"8.4_third_party_relationships": {"clause": "8.4", "title": "Third-party / customer relationships", "severity": "major"},
"9.1_monitoring": {"clause": "9.1", "title": "Monitoring, measurement, analysis, evaluation", "severity": "major"},
"9.2_internal_audit": {"clause": "9.2", "title": "Internal audit programme", "severity": "critical"},
"9.3_management_review": {"clause": "9.3", "title": "Management review", "severity": "critical"},
"10.1_continual_improvement": {"clause": "10.1", "title": "Continual improvement", "severity": "minor"},
"10.2_nonconformity_capa": {"clause": "10.2", "title": "Nonconformity & corrective action", "severity": "major"},
}
STATUS_SCORE = {"documented": 1.0, "partial": 0.5, "missing": 0.0}
SEVERITY_RANK = {"critical": 0, "major": 1, "minor": 2, "observation": 3}
def remediation_action(req_key: str, status: str) -> str:
"""Deterministic one-sentence next step per (clause, status)."""
if status == "documented":
return "Maintain via management review; re-verify at next internal audit."
titles = CLAUSE_REQUIREMENTS[req_key]["title"]
if status == "partial":
return f"Complete documentation of '{titles}' — confirm signoff, version control, evidence trail."
return f"Create from scratch: '{titles}'. Assign owner; target close before stage 1 audit."
def analyze(payload: Dict[str, Any]) -> Dict[str, Any]:
evidence = payload.get("evidence", {})
findings: List[Dict[str, Any]] = []
total_weight = 0.0
achieved_weight = 0.0
for req_key, meta in CLAUSE_REQUIREMENTS.items():
status = evidence.get(req_key, "missing")
score = STATUS_SCORE.get(status, 0.0)
# Severity-weighted: critical = 4, major = 2, minor = 1
weight = {"critical": 4, "major": 2, "minor": 1, "observation": 1}[meta["severity"]]
total_weight += weight
achieved_weight += weight * score
findings.append({
"clause": meta["clause"],
"title": meta["title"],
"status": status,
"severity_if_missing": meta["severity"],
"remediation": remediation_action(req_key, status),
})
coverage_pct = round((achieved_weight / total_weight) * 100, 1) if total_weight else 0
# Sort findings: missing/partial first by severity, then documented last
def sort_key(f: Dict[str, Any]) -> tuple:
status_order = {"missing": 0, "partial": 1, "documented": 2}
return (status_order[f["status"]], SEVERITY_RANK[f["severity_if_missing"]], f["clause"])
findings.sort(key=sort_key)
open_gaps = [f for f in findings if f["status"] != "documented"]
critical_gaps = [f for f in open_gaps if f["severity_if_missing"] == "critical"]
major_gaps = [f for f in open_gaps if f["severity_if_missing"] == "major"]
readiness = "ready" if not critical_gaps and len(major_gaps) <= 1 else (
"stage_2_candidate" if not critical_gaps else "not_ready"
)
return {
"organization": payload.get("organization"),
"scope": payload.get("scope_statement"),
"coverage_pct_weighted": coverage_pct,
"certification_readiness": readiness,
"critical_gap_count": len(critical_gaps),
"major_gap_count": len(major_gaps),
"open_gap_count": len(open_gaps),
"findings": findings,
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("ISO/IEC 42001 AIMS — GAP ANALYSIS")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Organization: {r['organization']}")
lines.append(f"Scope: {r['scope']}")
lines.append(f"Weighted coverage: {r['coverage_pct_weighted']}%")
lines.append(f"Certification readiness: {r['certification_readiness']}")
lines.append(f"Critical gaps: {r['critical_gap_count']} | Major gaps: {r['major_gap_count']} | Open total: {r['open_gap_count']}")
lines.append("")
lines.append("-" * 72)
lines.append("FINDINGS (open gaps first; critical highlighted):")
lines.append("")
for f in r["findings"]:
marker = {"missing": "[X] ", "partial": "[~] ", "documented": "[✓] "}[f["status"]]
sev = f["severity_if_missing"].upper() if f["status"] != "documented" else "OK"
lines.append(f" {marker}Clause {f['clause']:6s} {f['title']:50s} [{sev}]")
if f["status"] != "documented":
lines.append(f" → {f['remediation']}")
lines.append("")
lines.append("-" * 72)
lines.append("READINESS RULE: 'ready' = 0 critical AND ≤ 1 major. 'stage_2_candidate' = 0 critical.")
lines.append(" Any critical gap blocks stage 1 certification.")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="ISO/IEC 42001 AIMS gap analysis across Clauses 4-10.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to AIMS evidence JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: mid-stage AI SaaS, pre stage-1 audit>"
result = analyze(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/ai_risk_register_builder.py
#!/usr/bin/env python3
"""ai_risk_register_builder.py — ISO/IEC 42001 Annex A risk register + control mapping.
Stdlib-only. Takes identified AI risks (per ISO 23894 risk identification) and produces a
structured register with:
- severity rating (likelihood × impact, 5x5 matrix)
- mapped Annex A controls (treatment selection)
- residual risk verdict (accept / additional treatment required / escalate)
- treatment option per ISO 23894 (modify / share / retain / avoid)
Deterministic logic per ISO 23894:2023 risk-management process. No LLM calls.
Input schema (JSON):
{
"organization": "Acme AI Inc.",
"ai_system": "Customer recommendation engine v3",
"risks": [
{
"id": "R-001",
"source": "training_data",
"event": "Biased dataset over-represents one demographic",
"consequence": "Discriminatory recommendations; regulatory exposure",
"likelihood": 3, # 1-5
"impact": 4, # 1-5
"controls_applied": ["A.7.3", "A.7.5", "A.5.2"]
}
]
}
Usage:
python ai_risk_register_builder.py # uses embedded 7-risk sample
python ai_risk_register_builder.py path/to/risks.json
python ai_risk_register_builder.py risks.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"organization": "Acme AI Inc.",
"ai_system": "Customer recommendation engine v3",
"risks": [
{"id": "R-001", "source": "training_data", "event": "Biased dataset over-represents one demographic",
"consequence": "Discriminatory recommendations; regulatory exposure", "likelihood": 3, "impact": 4,
"controls_applied": ["A.7.3", "A.7.5", "A.5.2"]},
{"id": "R-002", "source": "model", "event": "Concept drift after 6 months in production",
"consequence": "Accuracy degradation; revenue impact", "likelihood": 4, "impact": 3,
"controls_applied": ["A.9.3", "A.6.2.4"]},
{"id": "R-003", "source": "deployment", "event": "Inference latency spike under load",
"consequence": "User-visible failure; SLO breach", "likelihood": 3, "impact": 2,
"controls_applied": ["A.9.3"]},
{"id": "R-004", "source": "third_party", "event": "Foundation-model API provider deprecates endpoint",
"consequence": "Service disruption; migration cost", "likelihood": 2, "impact": 4,
"controls_applied": ["A.10.2"]},
{"id": "R-005", "source": "data", "event": "Training data contains PII that should not be retained",
"consequence": "GDPR fine; trust loss", "likelihood": 2, "impact": 5,
"controls_applied": ["A.7.2", "A.7.4"]},
{"id": "R-006", "source": "human_oversight", "event": "High-impact decisions deployed without impact assessment",
"consequence": "Untracked harm; certification nonconformity", "likelihood": 3, "impact": 5,
"controls_applied": []},
{"id": "R-007", "source": "model", "event": "Adversarial prompt injection bypasses content filter",
"consequence": "Toxic output to end users; reputational damage", "likelihood": 4, "impact": 4,
"controls_applied": ["A.6.2.4", "A.9.3", "A.9.4"]},
],
}
# Severity matrix (5x5): likelihood (1-5) × impact (1-5)
# Score 1-4 = low, 5-9 = medium, 10-16 = high, 17-25 = critical
def severity_rating(likelihood: int, impact: int) -> str:
score = max(1, min(5, likelihood)) * max(1, min(5, impact))
if score <= 4:
return "low"
if score <= 9:
return "medium"
if score <= 16:
return "high"
return "critical"
# ISO 23894 risk treatment options
# - modify (apply controls to reduce likelihood/impact)
# - share (transfer via insurance, third-party contracts)
# - retain (accept residual risk with management signoff)
# - avoid (eliminate the activity entirely)
def treatment_option(severity: str, controls_count: int) -> str:
if severity == "critical" and controls_count == 0:
return "avoid_or_escalate"
if severity in ("high", "critical"):
return "modify"
if severity == "medium":
return "modify" if controls_count < 2 else "retain"
return "retain"
# Residual-risk verdict after applied controls
def residual_verdict(severity: str, controls_count: int) -> str:
"""How many controls are 'enough' for each severity tier (heuristic, ISO 23894 Annex A guidance)."""
expected = {"low": 0, "medium": 1, "high": 2, "critical": 3}[severity]
if controls_count >= expected:
return "acceptable" if severity != "critical" else "acceptable_with_management_signoff"
return "additional_treatment_required"
# Annex A control descriptions (subset, for output annotation)
ANNEX_A_CATALOG: Dict[str, str] = {
"A.2.2": "AI policy",
"A.2.3": "Alignment of AI policy with other organizational policies",
"A.3.2": "AI roles & responsibilities",
"A.3.3": "Reporting of concerns",
"A.4.2": "Resources for AI systems — data",
"A.4.3": "Resources for AI systems — tooling",
"A.4.4": "Resources for AI systems — human resources",
"A.5.2": "AI system impact assessment",
"A.5.4": "Documentation of impact assessment",
"A.6.2.2": "AI system objectives",
"A.6.2.3": "AI system lifecycle phases",
"A.6.2.4": "Verification & validation of AI system",
"A.7.2": "Data management for AI systems",
"A.7.3": "Data quality",
"A.7.4": "Data provenance",
"A.7.5": "Data preparation",
"A.8.2": "System documentation for users",
"A.8.3": "User information",
"A.8.4": "Communication of AI incidents",
"A.9.2": "Intended use of AI system",
"A.9.3": "Monitoring of AI system operation",
"A.9.4": "Logging of AI system events",
"A.10.2": "Supplier (third-party) relationships",
"A.10.3": "Customer relationships",
}
def annotate_risk(risk: Dict[str, Any]) -> Dict[str, Any]:
likelihood = int(risk.get("likelihood", 0))
impact = int(risk.get("impact", 0))
controls = list(risk.get("controls_applied", []))
sev = severity_rating(likelihood, impact)
treatment = treatment_option(sev, len(controls))
residual = residual_verdict(sev, len(controls))
return {
"id": risk.get("id"),
"source": risk.get("source"),
"event": risk.get("event"),
"consequence": risk.get("consequence"),
"likelihood": likelihood,
"impact": impact,
"severity_score": likelihood * impact,
"severity": sev,
"controls_applied": [{"id": c, "title": ANNEX_A_CATALOG.get(c, "<unknown control>")} for c in controls],
"control_count": len(controls),
"treatment_option": treatment,
"residual_verdict": residual,
}
def analyze(payload: Dict[str, Any]) -> Dict[str, Any]:
risks = [annotate_risk(r) for r in payload.get("risks", [])]
# Sort by severity (critical first), then by control gap (largest first)
sev_rank = {"critical": 0, "high": 1, "medium": 2, "low": 3}
risks.sort(key=lambda r: (sev_rank[r["severity"]], -r["severity_score"]))
counts_by_sev = {s: 0 for s in sev_rank}
requires_action = 0
for r in risks:
counts_by_sev[r["severity"]] += 1
if r["residual_verdict"] == "additional_treatment_required":
requires_action += 1
return {
"organization": payload.get("organization"),
"ai_system": payload.get("ai_system"),
"total_risks": len(risks),
"by_severity": counts_by_sev,
"requires_additional_treatment": requires_action,
"risks": risks,
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("AI RISK REGISTER — ISO/IEC 42001 Annex A + ISO 23894 treatment")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Organization: {r['organization']}")
lines.append(f"AI system: {r['ai_system']}")
lines.append(f"Total risks: {r['total_risks']}")
s = r["by_severity"]
lines.append(f"By severity: critical={s['critical']} high={s['high']} medium={s['medium']} low={s['low']}")
lines.append(f"Risks requiring additional treatment: {r['requires_additional_treatment']}")
lines.append("")
lines.append("-" * 72)
lines.append("REGISTER (highest severity first):")
lines.append("")
for risk in r["risks"]:
lines.append(f" [{risk['id']}] {risk['event']}")
lines.append(f" Source: {risk['source']} | L={risk['likelihood']} × I={risk['impact']} = {risk['severity_score']} → {risk['severity'].upper()}")
lines.append(f" Consequence: {risk['consequence']}")
if risk["controls_applied"]:
ctrl_str = ", ".join(c["id"] for c in risk["controls_applied"])
lines.append(f" Controls applied ({risk['control_count']}): {ctrl_str}")
else:
lines.append(f" Controls applied: NONE")
lines.append(f" Treatment option: {risk['treatment_option']}")
lines.append(f" Residual verdict: {risk['residual_verdict']}")
lines.append("")
lines.append("-" * 72)
lines.append("RULES:")
lines.append(" - 'critical' severity (score 17-25) WITHOUT controls → 'avoid_or_escalate' to management.")
lines.append(" - 'additional_treatment_required' → add Annex A controls or formally accept residual risk in writing.")
lines.append(" - All 'retain' verdicts require Clause 6.1.3 risk-treatment plan signoff.")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="ISO/IEC 42001 Annex A risk register builder with ISO 23894 treatment options.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to risks JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: 7-risk recommendation engine register>"
result = analyze(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
Tạo pipeline CI/CD thực dụng từ tín hiệu công nghệ của dự án, gồm kiểm tra lặp lại và giai đoạn triển khai theo môi trường.
---
name: "ci-cd-pipeline-builder"
description: "Generate pragmatic CI/CD pipelines from detected project stack signals — fast baseline generation, repeatable checks, environment-aware deployment stages. Use when setting up CI for a new project, refactoring existing pipelines, or standardizing deployment workflows across multiple repos."
---
# CI/CD Pipeline Builder
**Tier:** POWERFUL
**Category:** Engineering
**Domain:** DevOps / Automation
## Overview
Use this skill to generate pragmatic CI/CD pipelines from detected project stack signals, not guesswork. It focuses on fast baseline generation, repeatable checks, and environment-aware deployment stages.
## Core Capabilities
- Detect language/runtime/tooling from repository files
- Recommend CI stages (`lint`, `test`, `build`, `deploy`)
- Generate GitHub Actions or GitLab CI starter pipelines
- Include caching and matrix strategy based on detected stack
- Emit machine-readable detection output for automation
- Keep pipeline logic aligned with project lockfiles and build commands
## When to Use
- Bootstrapping CI for a new repository
- Replacing brittle copied pipeline files
- Migrating between GitHub Actions and GitLab CI
- Auditing whether pipeline steps match actual stack
- Creating a reproducible baseline before custom hardening
## Key Workflows
### 1. Detect Stack
```bash
python3 scripts/stack_detector.py --repo . --format text
python3 scripts/stack_detector.py --repo . --format json > detected-stack.json
```
Supports input via stdin or `--input` file for offline analysis payloads.
### 2. Generate Pipeline From Detection
```bash
python3 scripts/pipeline_generator.py \
--input detected-stack.json \
--platform github \
--output .github/workflows/ci.yml \
--format text
```
Or end-to-end from repo directly:
```bash
python3 scripts/pipeline_generator.py --repo . --platform gitlab --output .gitlab-ci.yml
```
### 3. Validate Before Merge
1. Confirm commands exist in project (`test`, `lint`, `build`).
2. Run generated pipeline locally where possible.
3. Ensure required secrets/env vars are documented.
4. Keep deploy jobs gated by protected branches/environments.
### 4. Add Deployment Stages Safely
- Start with CI-only (`lint/test/build`).
- Add staging deploy with explicit environment context.
- Add production deploy with manual gate/approval.
- Keep rollout/rollback commands explicit and auditable.
## Script Interfaces
- `python3 scripts/stack_detector.py --help`
- Detects stack signals from repository files
- Reads optional JSON input from stdin/`--input`
- `python3 scripts/pipeline_generator.py --help`
- Generates GitHub/GitLab YAML from detection payload
- Writes to stdout or `--output`
## Common Pitfalls
1. Copying a Node pipeline into Python/Go repos
2. Enabling deploy jobs before stable tests
3. Forgetting dependency cache keys
4. Running expensive matrix builds for every trivial branch
5. Missing branch protections around prod deploy jobs
6. Hardcoding secrets in YAML instead of CI secret stores
## Best Practices
1. Detect stack first, then generate pipeline.
2. Keep generated baseline under version control.
3. Add one optimization at a time (cache, matrix, split jobs).
4. Require green CI before deployment jobs.
5. Use protected environments for production credentials.
6. Regenerate pipeline when stack changes significantly.
## References
- [references/github-actions-templates.md](references/github-actions-templates.md)
- [references/gitlab-ci-templates.md](references/gitlab-ci-templates.md)
- [references/deployment-gates.md](references/deployment-gates.md)
- [README.md](README.md)
## Detection Heuristics
The stack detector prioritizes deterministic file signals over heuristics:
- Lockfiles determine package manager preference
- Language manifests determine runtime families
- Script commands (if present) drive lint/test/build commands
- Missing scripts trigger conservative placeholder commands
## Generation Strategy
Start with a minimal, reliable pipeline:
1. Checkout and setup runtime
2. Install dependencies with cache strategy
3. Run lint, test, build in separate steps
4. Publish artifacts only after passing checks
Then layer advanced behavior (matrix builds, security scans, deploy gates).
## Platform Decision Notes
- GitHub Actions for tight GitHub ecosystem integration
- GitLab CI for integrated SCM + CI in self-hosted environments
- Keep one canonical pipeline source per repo to reduce drift
## Validation Checklist
1. Generated YAML parses successfully.
2. All referenced commands exist in the repo.
3. Cache strategy matches package manager.
4. Required secrets are documented, not embedded.
5. Branch/protected-environment rules match org policy.
## Scaling Guidance
- Split long jobs by stage when runtime exceeds 10 minutes.
- Introduce test matrix only when compatibility truly requires it.
- Separate deploy jobs from CI jobs to keep feedback fast.
- Track pipeline duration and flakiness as first-class metrics.
FILE:README.md
# CI/CD Pipeline Builder
Detects your repository stack and generates practical CI pipeline templates for GitHub Actions and GitLab CI. Designed as a fast baseline you can extend with deployment controls.
## Quick Start
```bash
# Detect stack
python3 scripts/stack_detector.py --repo . --format json > stack.json
# Generate GitHub Actions workflow
python3 scripts/pipeline_generator.py \
--input stack.json \
--platform github \
--output .github/workflows/ci.yml \
--format text
```
## Included Tools
- `scripts/stack_detector.py`: repository signal detection with JSON/text output
- `scripts/pipeline_generator.py`: generate GitHub/GitLab CI YAML from detection payload
## References
- `references/github-actions-templates.md`
- `references/gitlab-ci-templates.md`
- `references/deployment-gates.md`
## Installation
### Claude Code
```bash
cp -R engineering/ci-cd-pipeline-builder ~/.claude/skills/ci-cd-pipeline-builder
```
### OpenAI Codex
```bash
cp -R engineering/ci-cd-pipeline-builder ~/.codex/skills/ci-cd-pipeline-builder
```
### OpenClaw
```bash
cp -R engineering/ci-cd-pipeline-builder ~/.openclaw/skills/ci-cd-pipeline-builder
```
FILE:references/deployment-gates.md
# Deployment Gates
## Minimum Gate Policy
- `lint` must pass before `test`.
- `test` must pass before `build`.
- `build` artifact required for deploy jobs.
- Production deploy requires manual approval and protected branch.
## Environment Pattern
- `develop` -> auto deploy to staging
- `main` -> manual promote to production
## Rollback Requirement
Every deploy job should define a rollback command or procedure reference.
FILE:references/github-actions-templates.md
# GitHub Actions Templates
## Node.js Baseline
```yaml
name: Node CI
on: [push, pull_request]
jobs:
ci:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '20'
cache: 'npm'
- run: npm ci
- run: npm run lint
- run: npm test
- run: npm run build
```
## Python Baseline
```yaml
name: Python CI
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: python3 -m pip install -U pip
- run: python3 -m pip install -r requirements.txt
- run: python3 -m pytest
```
FILE:references/gitlab-ci-templates.md
# GitLab CI Templates
## Node.js Baseline
```yaml
stages:
- lint
- test
- build
node_lint:
image: node:20
stage: lint
script:
- npm ci
- npm run lint
node_test:
image: node:20
stage: test
script:
- npm ci
- npm test
```
## Python Baseline
```yaml
stages:
- test
python_test:
image: python:3.12
stage: test
script:
- python3 -m pip install -U pip
- python3 -m pip install -r requirements.txt
- python3 -m pytest
```
FILE:scripts/pipeline_generator.py
#!/usr/bin/env python3
"""Generate CI pipeline YAML from detected stack data.
Input sources:
- --input stack report JSON file
- stdin stack report JSON
- --repo path (auto-detect stack)
Output:
- text/json summary
- pipeline YAML written via --output or printed to stdout
"""
import argparse
import json
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Any, Dict, List, Optional
class CLIError(Exception):
"""Raised for expected CLI failures."""
@dataclass
class PipelineSummary:
platform: str
output: str
stages: List[str]
uses_cache: bool
languages: List[str]
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Generate CI/CD pipeline YAML from detected stack.")
parser.add_argument("--input", help="Stack report JSON file. If omitted, can read stdin JSON.")
parser.add_argument("--repo", help="Repository path for auto-detection fallback.")
parser.add_argument("--platform", choices=["github", "gitlab"], required=True, help="Target CI platform.")
parser.add_argument("--output", help="Write YAML to this file; otherwise print to stdout.")
parser.add_argument("--format", choices=["text", "json"], default="text", help="Summary output format.")
return parser.parse_args()
def load_json_input(input_path: Optional[str]) -> Optional[Dict[str, Any]]:
if input_path:
try:
return json.loads(Path(input_path).read_text(encoding="utf-8"))
except Exception as exc:
raise CLIError(f"Failed reading --input: {exc}") from exc
if not sys.stdin.isatty():
raw = sys.stdin.read().strip()
if raw:
try:
return json.loads(raw)
except json.JSONDecodeError as exc:
raise CLIError(f"Invalid JSON from stdin: {exc}") from exc
return None
def detect_stack(repo: Path) -> Dict[str, Any]:
scripts = {}
pkg_file = repo / "package.json"
if pkg_file.exists():
try:
pkg = json.loads(pkg_file.read_text(encoding="utf-8"))
raw_scripts = pkg.get("scripts", {})
if isinstance(raw_scripts, dict):
scripts = raw_scripts
except Exception:
scripts = {}
languages: List[str] = []
if pkg_file.exists():
languages.append("node")
if (repo / "pyproject.toml").exists() or (repo / "requirements.txt").exists():
languages.append("python")
if (repo / "go.mod").exists():
languages.append("go")
return {
"languages": sorted(set(languages)),
"signals": {
"pnpm_lock": (repo / "pnpm-lock.yaml").exists(),
"yarn_lock": (repo / "yarn.lock").exists(),
"npm_lock": (repo / "package-lock.json").exists(),
"dockerfile": (repo / "Dockerfile").exists(),
},
"lint_commands": ["npm run lint"] if "lint" in scripts else [],
"test_commands": ["npm test"] if "test" in scripts else [],
"build_commands": ["npm run build"] if "build" in scripts else [],
}
def select_node_install(signals: Dict[str, Any]) -> str:
if signals.get("pnpm_lock"):
return "pnpm install --frozen-lockfile"
if signals.get("yarn_lock"):
return "yarn install --frozen-lockfile"
return "npm ci"
def github_yaml(stack: Dict[str, Any]) -> str:
langs = stack.get("languages", [])
signals = stack.get("signals", {})
lint_cmds = stack.get("lint_commands", []) or ["echo 'No lint command configured'"]
test_cmds = stack.get("test_commands", []) or ["echo 'No test command configured'"]
build_cmds = stack.get("build_commands", []) or ["echo 'No build command configured'"]
lines: List[str] = [
"name: CI",
"on:",
" push:",
" branches: [main, develop]",
" pull_request:",
" branches: [main, develop]",
"",
"jobs:",
]
if "node" in langs:
lines.extend(
[
" node-ci:",
" runs-on: ubuntu-latest",
" steps:",
" - uses: actions/checkout@v4",
" - uses: actions/setup-node@v4",
" with:",
" node-version: '20'",
" cache: 'npm'",
f" - run: {select_node_install(signals)}",
]
)
for cmd in lint_cmds + test_cmds + build_cmds:
lines.append(f" - run: {cmd}")
if "python" in langs:
lines.extend(
[
" python-ci:",
" runs-on: ubuntu-latest",
" steps:",
" - uses: actions/checkout@v4",
" - uses: actions/setup-python@v5",
" with:",
" python-version: '3.12'",
" - run: python3 -m pip install -U pip",
" - run: python3 -m pip install -r requirements.txt || true",
" - run: python3 -m pytest || true",
]
)
if "go" in langs:
lines.extend(
[
" go-ci:",
" runs-on: ubuntu-latest",
" steps:",
" - uses: actions/checkout@v4",
" - uses: actions/setup-go@v5",
" with:",
" go-version: '1.22'",
" - run: go test ./...",
" - run: go build ./...",
]
)
return "\n".join(lines) + "\n"
def gitlab_yaml(stack: Dict[str, Any]) -> str:
langs = stack.get("languages", [])
signals = stack.get("signals", {})
lint_cmds = stack.get("lint_commands", []) or ["echo 'No lint command configured'"]
test_cmds = stack.get("test_commands", []) or ["echo 'No test command configured'"]
build_cmds = stack.get("build_commands", []) or ["echo 'No build command configured'"]
lines: List[str] = [
"stages:",
" - lint",
" - test",
" - build",
"",
]
if "node" in langs:
install_cmd = select_node_install(signals)
lines.extend(
[
"node_lint:",
" image: node:20",
" stage: lint",
" script:",
f" - {install_cmd}",
]
)
for cmd in lint_cmds:
lines.append(f" - {cmd}")
lines.extend(
[
"",
"node_test:",
" image: node:20",
" stage: test",
" script:",
f" - {install_cmd}",
]
)
for cmd in test_cmds:
lines.append(f" - {cmd}")
lines.extend(
[
"",
"node_build:",
" image: node:20",
" stage: build",
" script:",
f" - {install_cmd}",
]
)
for cmd in build_cmds:
lines.append(f" - {cmd}")
if "python" in langs:
lines.extend(
[
"",
"python_test:",
" image: python:3.12",
" stage: test",
" script:",
" - python3 -m pip install -U pip",
" - python3 -m pip install -r requirements.txt || true",
" - python3 -m pytest || true",
]
)
if "go" in langs:
lines.extend(
[
"",
"go_test:",
" image: golang:1.22",
" stage: test",
" script:",
" - go test ./...",
" - go build ./...",
]
)
return "\n".join(lines) + "\n"
def main() -> int:
args = parse_args()
stack = load_json_input(args.input)
if stack is None:
if not args.repo:
raise CLIError("Provide stack input via --input/stdin or set --repo for auto-detection.")
repo = Path(args.repo).resolve()
if not repo.exists() or not repo.is_dir():
raise CLIError(f"Invalid repo path: {repo}")
stack = detect_stack(repo)
if args.platform == "github":
yaml_content = github_yaml(stack)
else:
yaml_content = gitlab_yaml(stack)
output_path = args.output or "stdout"
if args.output:
out = Path(args.output)
out.parent.mkdir(parents=True, exist_ok=True)
out.write_text(yaml_content, encoding="utf-8")
else:
print(yaml_content, end="")
summary = PipelineSummary(
platform=args.platform,
output=output_path,
stages=["lint", "test", "build"],
uses_cache=True,
languages=stack.get("languages", []),
)
if args.format == "json":
print(json.dumps(asdict(summary), indent=2), file=sys.stderr if not args.output else sys.stdout)
else:
text = (
"Pipeline generated\n"
f"- platform: {summary.platform}\n"
f"- output: {summary.output}\n"
f"- stages: {', '.join(summary.stages)}\n"
f"- languages: {', '.join(summary.languages) if summary.languages else 'none'}"
)
print(text, file=sys.stderr if not args.output else sys.stdout)
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except CLIError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
raise SystemExit(2)
FILE:scripts/stack_detector.py
#!/usr/bin/env python3
"""Detect project stack/tooling signals for CI/CD pipeline generation.
Input sources:
- repository scan via --repo
- JSON via --input file
- JSON via stdin
Output:
- text summary or JSON payload
"""
import argparse
import json
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Dict, List, Optional
class CLIError(Exception):
"""Raised for expected CLI failures."""
@dataclass
class StackReport:
repo: str
languages: List[str]
package_managers: List[str]
ci_targets: List[str]
test_commands: List[str]
build_commands: List[str]
lint_commands: List[str]
signals: Dict[str, bool]
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Detect stack/tooling from a repository.")
parser.add_argument("--input", help="JSON input file (precomputed signal payload).")
parser.add_argument("--repo", default=".", help="Repository path to scan.")
parser.add_argument("--format", choices=["text", "json"], default="text", help="Output format.")
return parser.parse_args()
def load_payload(input_path: Optional[str]) -> Optional[dict]:
if input_path:
try:
return json.loads(Path(input_path).read_text(encoding="utf-8"))
except Exception as exc:
raise CLIError(f"Failed reading --input file: {exc}") from exc
if not sys.stdin.isatty():
raw = sys.stdin.read().strip()
if raw:
try:
return json.loads(raw)
except json.JSONDecodeError as exc:
raise CLIError(f"Invalid JSON from stdin: {exc}") from exc
return None
def read_package_scripts(repo: Path) -> Dict[str, str]:
pkg = repo / "package.json"
if not pkg.exists():
return {}
try:
data = json.loads(pkg.read_text(encoding="utf-8"))
except Exception:
return {}
scripts = data.get("scripts", {})
return scripts if isinstance(scripts, dict) else {}
def detect(repo: Path) -> StackReport:
signals = {
"package_json": (repo / "package.json").exists(),
"pnpm_lock": (repo / "pnpm-lock.yaml").exists(),
"yarn_lock": (repo / "yarn.lock").exists(),
"npm_lock": (repo / "package-lock.json").exists(),
"pyproject": (repo / "pyproject.toml").exists(),
"requirements": (repo / "requirements.txt").exists(),
"go_mod": (repo / "go.mod").exists(),
"dockerfile": (repo / "Dockerfile").exists(),
"vercel": (repo / "vercel.json").exists(),
"helm": (repo / "helm").exists() or (repo / "charts").exists(),
"k8s": (repo / "k8s").exists() or (repo / "kubernetes").exists(),
}
languages: List[str] = []
package_managers: List[str] = []
ci_targets: List[str] = ["github", "gitlab"]
if signals["package_json"]:
languages.append("node")
if signals["pnpm_lock"]:
package_managers.append("pnpm")
elif signals["yarn_lock"]:
package_managers.append("yarn")
else:
package_managers.append("npm")
if signals["pyproject"] or signals["requirements"]:
languages.append("python")
package_managers.append("pip")
if signals["go_mod"]:
languages.append("go")
scripts = read_package_scripts(repo)
lint_commands: List[str] = []
test_commands: List[str] = []
build_commands: List[str] = []
if "lint" in scripts:
lint_commands.append("npm run lint")
if "test" in scripts:
test_commands.append("npm test")
if "build" in scripts:
build_commands.append("npm run build")
if "python" in languages:
lint_commands.append("python3 -m ruff check .")
test_commands.append("python3 -m pytest")
if "go" in languages:
lint_commands.append("go vet ./...")
test_commands.append("go test ./...")
build_commands.append("go build ./...")
return StackReport(
repo=str(repo.resolve()),
languages=sorted(set(languages)),
package_managers=sorted(set(package_managers)),
ci_targets=ci_targets,
test_commands=sorted(set(test_commands)),
build_commands=sorted(set(build_commands)),
lint_commands=sorted(set(lint_commands)),
signals=signals,
)
def format_text(report: StackReport) -> str:
lines = [
"Detected stack",
f"- repo: {report.repo}",
f"- languages: {', '.join(report.languages) if report.languages else 'none'}",
f"- package managers: {', '.join(report.package_managers) if report.package_managers else 'none'}",
f"- lint commands: {', '.join(report.lint_commands) if report.lint_commands else 'none'}",
f"- test commands: {', '.join(report.test_commands) if report.test_commands else 'none'}",
f"- build commands: {', '.join(report.build_commands) if report.build_commands else 'none'}",
]
return "\n".join(lines)
def main() -> int:
args = parse_args()
payload = load_payload(args.input)
if payload:
try:
report = StackReport(**payload)
except TypeError as exc:
raise CLIError(f"Invalid input payload for StackReport: {exc}") from exc
else:
repo = Path(args.repo).resolve()
if not repo.exists() or not repo.is_dir():
raise CLIError(f"Invalid repo path: {repo}")
report = detect(repo)
if args.format == "json":
print(json.dumps(asdict(report), indent=2))
else:
print(format_text(report))
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except CLIError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
raise SystemExit(2)
Phân tích và huấn luyện đội agile dựa trên dữ liệu: sprint planning, velocity, retrospective, backlog, burndown và blocker.
---
name: "scrum-master"
description: "Advanced Scrum Master skill for data-driven agile team analysis and coaching. Use when the user asks about sprint planning, velocity tracking, retrospectives, standup facilitation, backlog grooming, story points, burndown charts, blocker resolution, or agile team health. Runs Python scripts to analyse sprint JSON exports from Jira or similar tools: velocity_analyzer.py for Monte Carlo sprint forecasting, sprint_health_scorer.py for multi-dimension health scoring, and retrospective_analyzer.py for action-item and theme tracking. Produces confidence-interval forecasts, health grade reports, and improvement-velocity trends for high-performing Scrum teams."
license: MIT
metadata:
version: 2.0.0
author: Alireza Rezvani
category: project-management
domain: agile-development
updated: 2026-02-15
python-tools: velocity_analyzer.py, sprint_health_scorer.py, retrospective_analyzer.py
tech-stack: scrum, agile-coaching, team-dynamics, data-analysis
---
# Scrum Master Expert
Data-driven Scrum Master skill combining sprint analytics, probabilistic forecasting, and team development coaching. The unique value is in the three Python analysis scripts and their workflows — refer to `references/` and `assets/` for deeper framework detail.
---
## Table of Contents
- [Analysis Tools & Usage](#analysis-tools-usage)
- [Input Requirements](#input-requirements)
- [Sprint Execution Workflows](#sprint-execution-workflows)
- [Team Development Workflow](#team-development-workflow)
- [Key Metrics & Targets](#key-metrics-targets)
- [Limitations](#limitations)
---
## Analysis Tools & Usage
### 1. Velocity Analyzer (`scripts/velocity_analyzer.py`)
Runs rolling averages, linear-regression trend detection, and Monte Carlo simulation over sprint history.
```bash
# Text report
python velocity_analyzer.py sprint_data.json --format text
# JSON output for downstream processing
python velocity_analyzer.py sprint_data.json --format json > analysis.json
```
**Outputs**: velocity trend (improving/stable/declining), coefficient of variation, 6-sprint Monte Carlo forecast at 50 / 70 / 85 / 95% confidence intervals, anomaly flags with root-cause suggestions.
**Validation**: If fewer than 3 sprints are present in the input, stop and prompt the user: *"Velocity analysis needs at least 3 sprints. Please provide additional sprint data."* 6+ sprints are recommended for statistically significant Monte Carlo results.
---
### 2. Sprint Health Scorer (`scripts/sprint_health_scorer.py`)
Scores team health across 6 weighted dimensions, producing an overall 0–100 grade.
| Dimension | Weight | Target |
|---|---|---|
| Commitment Reliability | 25% | >85% sprint goals met |
| Scope Stability | 20% | <15% mid-sprint changes |
| Blocker Resolution | 15% | <3 days average |
| Ceremony Engagement | 15% | >90% participation |
| Story Completion Distribution | 15% | High ratio of fully done stories |
| Velocity Predictability | 10% | CV <20% |
```bash
python sprint_health_scorer.py sprint_data.json --format text
```
**Outputs**: overall health score + grade, per-dimension scores with recommendations, sprint-over-sprint trend, intervention priority matrix.
**Validation**: Requires 2+ sprints with ceremony and story-completion data. If data is missing, report which dimensions cannot be scored and ask the user to supply the gaps.
---
### 3. Retrospective Analyzer (`scripts/retrospective_analyzer.py`)
Tracks action-item completion, recurring themes, sentiment trends, and team maturity progression.
```bash
python retrospective_analyzer.py sprint_data.json --format text
```
**Outputs**: action-item completion rate by priority/owner, recurring-theme persistence scores, team maturity level (forming/storming/norming/performing), improvement-velocity trend.
**Validation**: Requires 3+ retrospectives with action-item tracking. With fewer, note the limitation and offer partial theme analysis only.
---
## Input Requirements
All scripts accept JSON following the schema in `assets/sample_sprint_data.json`:
```json
{
"team_info": { "name": "string", "size": "number", "scrum_master": "string" },
"sprints": [
{
"sprint_number": "number",
"planned_points": "number",
"completed_points": "number",
"stories": [...],
"blockers": [...],
"ceremonies": {...}
}
],
"retrospectives": [
{
"sprint_number": "number",
"went_well": ["string"],
"to_improve": ["string"],
"action_items": [...]
}
]
}
```
Jira and similar tools can export sprint data; map exported fields to this schema before running the scripts. See `assets/sample_sprint_data.json` for a complete 6-sprint example and `assets/expected_output.json` for corresponding expected results (velocity avg 20.2 pts, CV 12.7%, health score 78.3/100, action-item completion 46.7%).
---
## Sprint Execution Workflows
### Sprint Planning
1. Run velocity analysis: `python velocity_analyzer.py sprint_data.json --format text`
2. Use the 70% confidence interval as the recommended commitment ceiling for the sprint backlog.
3. Review the health scorer's Commitment Reliability and Scope Stability scores to calibrate negotiation with the Product Owner.
4. If Monte Carlo output shows high volatility (CV >20%), surface this to stakeholders with range estimates rather than single-point forecasts.
5. Document capacity assumptions (leave, dependencies) for retrospective comparison.
### Daily Standup
1. Track participation and help-seeking patterns — feed ceremony data into `sprint_health_scorer.py` at sprint end.
2. Log each blocker with date opened; resolution time feeds the Blocker Resolution dimension.
3. If a blocker is unresolved after 2 days, escalate proactively and note in sprint data.
### Sprint Review
1. Present velocity trend and health score alongside the demo to give stakeholders delivery context.
2. Capture scope-change requests raised during review; record as scope-change events in sprint data for next scoring cycle.
### Sprint Retrospective
1. Run all three scripts before the session:
```bash
python sprint_health_scorer.py sprint_data.json --format text > health.txt
python retrospective_analyzer.py sprint_data.json --format text > retro.txt
```
2. Open with the health score and top-flagged dimensions to focus discussion.
3. Use the retrospective analyzer's action-item completion rate to determine how many new action items the team can realistically absorb (target: ≤3 if completion rate <60%).
4. Assign each action item an owner and measurable success criterion before closing the session.
5. Record new action items in `sprint_data.json` for tracking in the next cycle.
---
## Team Development Workflow
### Assessment
```bash
python sprint_health_scorer.py team_data.json > health_assessment.txt
python retrospective_analyzer.py team_data.json > retro_insights.txt
```
- Map retrospective analyzer maturity output to the appropriate development stage.
- Supplement with an anonymous psychological safety pulse survey (Edmondson 7-point scale) and individual 1:1 observations.
- If maturity output is `forming` or `storming`, prioritise safety and conflict-facilitation interventions before process optimisation.
### Intervention
Apply stage-specific facilitation (details in `references/team-dynamics-framework.md`):
| Stage | Focus |
|---|---|
| Forming | Structure, process education, trust building |
| Storming | Conflict facilitation, psychological safety maintenance |
| Norming | Autonomy building, process ownership transfer |
| Performing | Challenge introduction, innovation support |
### Progress Measurement
- **Sprint cadence**: re-run health scorer; target overall score improvement of ≥5 points per quarter.
- **Monthly**: psychological safety pulse survey; target >4.0/5.0.
- **Quarterly**: full maturity re-assessment via retrospective analyzer.
- If scores plateau or regress for 2 consecutive sprints, escalate intervention strategy (see `references/team-dynamics-framework.md`).
---
## Key Metrics & Targets
| Metric | Target |
|---|---|
| Overall Health Score | >80/100 |
| Psychological Safety Index | >4.0/5.0 |
| Velocity CV (predictability) | <20% |
| Commitment Reliability | >85% |
| Scope Stability | <15% mid-sprint changes |
| Blocker Resolution Time | <3 days |
| Ceremony Engagement | >90% |
| Retrospective Action Completion | >70% |
---
## Limitations
- **Sample size**: fewer than 6 sprints reduces Monte Carlo confidence; always state confidence intervals, not point estimates.
- **Data completeness**: missing ceremony or story-completion fields suppress affected scoring dimensions — report gaps explicitly.
- **Context sensitivity**: script recommendations must be interpreted alongside organisational and team context not captured in JSON data.
- **Quantitative bias**: metrics do not replace qualitative observation; combine scores with direct team interaction.
- **Team size**: techniques are optimised for 5–9 member teams; larger groups may require adaptation.
- **External factors**: cross-team dependencies and organisational constraints are not fully modelled by single-team metrics.
---
## Related Skills
- **Agile Product Owner** (`product-team/agile-product-owner/`) — User stories and backlog feed sprint planning
- **Senior PM** (`project-management/senior-pm/`) — Portfolio health context informs sprint priorities
---
*For deep framework references see `references/velocity-forecasting-guide.md` and `references/team-dynamics-framework.md`. For template assets see `assets/sprint_report_template.md` and `assets/team_health_check_template.md`.*
FILE:assets/expected_output.json
{
"velocity_analysis": {
"summary": {
"total_sprints": 6,
"velocity_stats": {
"mean": 20.17,
"median": 20.0,
"min": 17,
"max": 24,
"total_points": 121
},
"commitment_analysis": {
"average_commitment_ratio": 0.908,
"commitment_consistency": 0.179,
"sprints_under_committed": 3,
"sprints_over_committed": 2
},
"volatility": {
"volatility": "low",
"coefficient_of_variation": 0.127
}
},
"trend_analysis": {
"trend": "stable",
"confidence": 0.15,
"relative_slope": -0.013
},
"forecasting": {
"expected_total": 121.0,
"forecasted_totals": {
"50%": 115,
"70%": 125,
"85%": 135,
"95%": 148
}
},
"anomalies": [
{
"sprint_number": 5,
"velocity": 17,
"anomaly_type": "outlier",
"deviation_percentage": -15.7
}
]
},
"sprint_health": {
"overall_score": 78.3,
"health_grade": "good",
"dimension_scores": {
"commitment_reliability": {
"score": 96.8,
"grade": "excellent"
},
"scope_stability": {
"score": 54.8,
"grade": "poor"
},
"blocker_resolution": {
"score": 51.7,
"grade": "poor"
},
"ceremony_engagement": {
"score": 92.3,
"grade": "excellent"
},
"story_completion_distribution": {
"score": 93.3,
"grade": "excellent"
},
"velocity_predictability": {
"score": 80.5,
"grade": "good"
}
}
},
"retrospective_analysis": {
"summary": {
"total_retrospectives": 6,
"average_duration": 74,
"average_attendance": 0.933
},
"action_item_analysis": {
"total_action_items": 15,
"completion_rate": 0.467,
"overdue_rate": 0.533,
"priority_analysis": {
"high": {"completion_rate": 0.50},
"medium": {"completion_rate": 0.33},
"low": {"completion_rate": 0.67}
}
},
"theme_analysis": {
"recurring_themes": {
"process": {"frequency": 1.0, "trend": {"direction": "decreasing"}},
"team_dynamics": {"frequency": 1.0, "trend": {"direction": "increasing"}},
"technical": {"frequency": 0.83, "trend": {"direction": "increasing"}},
"communication": {"frequency": 0.67, "trend": {"direction": "decreasing"}}
}
},
"improvement_trends": {
"team_maturity_score": {
"score": 75.6,
"level": "performing"
},
"improvement_velocity": {
"velocity": "moderate",
"velocity_score": 0.62
}
}
},
"interpretation": {
"strengths": [
"Excellent commitment reliability - team consistently delivers what they commit to",
"High ceremony engagement - team actively participates in scrum events",
"Good story completion distribution - stories are finished rather than left partially done",
"Low velocity volatility - predictable delivery capability"
],
"areas_for_improvement": [
"Scope instability - too much mid-sprint change (22.6% average)",
"Blocker resolution time - 4.7 days average is too long",
"Action item completion rate - only 46.7% completed",
"High overdue rate - 53.3% of action items become overdue"
],
"recommended_actions": [
"Strengthen backlog refinement to reduce scope changes",
"Implement faster blocker escalation process",
"Reduce number of retrospective action items and focus on follow-through",
"Create external dependency register to proactively manage blockers"
]
}
}
FILE:assets/expected_velocity_output.json
{
"summary": {
"total_sprints": 6,
"velocity_stats": {
"mean": 20.166666666666668,
"median": 20.0,
"min": 17,
"max": 24,
"total_points": 121
},
"commitment_analysis": {
"average_commitment_ratio": 0.9075307422046552,
"commitment_consistency": 0.17889820455801825,
"sprints_under_committed": 3,
"sprints_over_committed": 2
},
"scope_change_analysis": {
"average_scope_change": 0.22586752619361317,
"scope_change_volatility": 0.1828476660567787
},
"rolling_averages": {
"3": [
null,
null,
19.333333333333332,
20.666666666666668,
19.333333333333332,
21.0
],
"5": [
null,
null,
19.333333333333332,
20.0,
19.4,
20.6
],
"8": [
null,
null,
19.333333333333332,
20.0,
19.4,
20.166666666666668
]
},
"volatility": {
"volatility": "low",
"coefficient_of_variation": 0.13088153980052333,
"standard_deviation": 2.6394443859772205,
"mean_velocity": 20.166666666666668,
"velocity_range": 7,
"range_ratio": 0.3471074380165289,
"min_velocity": 17,
"max_velocity": 24
}
},
"trend_analysis": {
"trend": "stable",
"slope": 0.6,
"relative_slope": 0.029752066115702476,
"correlation": 0.42527784332026836,
"confidence": 0.42527784332026836,
"recent_sprints_analyzed": 6,
"average_velocity": 20.166666666666668
},
"forecasting": {
"sprints_ahead": 6,
"historical_sprints_used": 6,
"mean_velocity": 20.166666666666668,
"velocity_std_dev": 2.6394443859772205,
"forecasted_totals": {
"50%": 121.00756172377734,
"70%": 124.35398229685968,
"85%": 127.68925669583572,
"95%": 131.66775744677182
},
"average_per_sprint": 20.166666666666668,
"expected_total": 121.0
},
"anomalies": [],
"recommendations": [
"Good velocity stability. Continue current practices."
]
}
FILE:assets/sample_sprint_data.json
{
"team_info": {
"name": "Phoenix Development Team",
"size": 5,
"scrum_master": "Sarah Chen",
"product_owner": "Mike Rodriguez"
},
"sprints": [
{
"sprint_number": 1,
"sprint_name": "Sprint Alpha",
"start_date": "2024-01-08",
"end_date": "2024-01-19",
"planned_points": 23,
"completed_points": 18,
"added_points": 3,
"removed_points": 2,
"carry_over_points": 5,
"team_capacity": 40,
"working_days": 10,
"team_size": 5,
"stories": [
{
"id": "US-101",
"title": "User authentication system",
"points": 8,
"status": "completed",
"assigned_to": "John Doe",
"created_date": "2024-01-08",
"completed_date": "2024-01-16",
"blocked_days": 0,
"priority": "high"
},
{
"id": "US-102",
"title": "Dashboard layout implementation",
"points": 5,
"status": "completed",
"assigned_to": "Jane Smith",
"created_date": "2024-01-08",
"completed_date": "2024-01-18",
"blocked_days": 1,
"priority": "medium"
},
{
"id": "US-103",
"title": "API integration for user data",
"points": 5,
"status": "completed",
"assigned_to": "Bob Wilson",
"created_date": "2024-01-08",
"completed_date": "2024-01-19",
"blocked_days": 0,
"priority": "medium"
},
{
"id": "US-104",
"title": "Advanced filtering options",
"points": 5,
"status": "in_progress",
"assigned_to": "Alice Brown",
"created_date": "2024-01-08",
"blocked_days": 2,
"priority": "low"
}
],
"blockers": [
{
"id": "B-001",
"description": "Third-party API documentation incomplete",
"created_date": "2024-01-10",
"resolved_date": "2024-01-12",
"resolution_days": 2,
"affected_stories": ["US-103"],
"category": "external"
}
],
"ceremonies": {
"daily_standup": {
"attendance_rate": 0.92,
"engagement_score": 0.85
},
"sprint_planning": {
"attendance_rate": 1.0,
"engagement_score": 0.90
},
"sprint_review": {
"attendance_rate": 0.96,
"engagement_score": 0.88
},
"retrospective": {
"attendance_rate": 1.0,
"engagement_score": 0.95
}
}
},
{
"sprint_number": 2,
"sprint_name": "Sprint Beta",
"start_date": "2024-01-22",
"end_date": "2024-02-02",
"planned_points": 21,
"completed_points": 21,
"added_points": 1,
"removed_points": 1,
"carry_over_points": 3,
"team_capacity": 38,
"working_days": 9,
"team_size": 5,
"stories": [
{
"id": "US-105",
"title": "Email notification system",
"points": 8,
"status": "completed",
"assigned_to": "John Doe",
"created_date": "2024-01-22",
"completed_date": "2024-01-30",
"blocked_days": 0,
"priority": "high"
},
{
"id": "US-106",
"title": "User profile management",
"points": 5,
"status": "completed",
"assigned_to": "Jane Smith",
"created_date": "2024-01-22",
"completed_date": "2024-02-01",
"blocked_days": 0,
"priority": "medium"
},
{
"id": "US-107",
"title": "Data export functionality",
"points": 3,
"status": "completed",
"assigned_to": "Bob Wilson",
"created_date": "2024-01-22",
"completed_date": "2024-01-31",
"blocked_days": 0,
"priority": "medium"
},
{
"id": "US-104",
"title": "Advanced filtering options",
"points": 5,
"status": "completed",
"assigned_to": "Alice Brown",
"created_date": "2024-01-08",
"completed_date": "2024-02-02",
"blocked_days": 0,
"priority": "low"
}
],
"blockers": [],
"ceremonies": {
"daily_standup": {
"attendance_rate": 0.94,
"engagement_score": 0.88
},
"sprint_planning": {
"attendance_rate": 1.0,
"engagement_score": 0.92
},
"sprint_review": {
"attendance_rate": 1.0,
"engagement_score": 0.90
},
"retrospective": {
"attendance_rate": 1.0,
"engagement_score": 0.93
}
}
},
{
"sprint_number": 3,
"sprint_name": "Sprint Gamma",
"start_date": "2024-02-05",
"end_date": "2024-02-16",
"planned_points": 24,
"completed_points": 19,
"added_points": 4,
"removed_points": 3,
"carry_over_points": 5,
"team_capacity": 42,
"working_days": 10,
"team_size": 5,
"stories": [
{
"id": "US-108",
"title": "Real-time chat implementation",
"points": 13,
"status": "in_progress",
"assigned_to": "John Doe",
"created_date": "2024-02-05",
"blocked_days": 3,
"priority": "high"
},
{
"id": "US-109",
"title": "Mobile responsive design",
"points": 8,
"status": "completed",
"assigned_to": "Jane Smith",
"created_date": "2024-02-05",
"completed_date": "2024-02-14",
"blocked_days": 0,
"priority": "high"
},
{
"id": "US-110",
"title": "Performance optimization",
"points": 3,
"status": "completed",
"assigned_to": "Bob Wilson",
"created_date": "2024-02-05",
"completed_date": "2024-02-13",
"blocked_days": 1,
"priority": "medium"
}
],
"blockers": [
{
"id": "B-002",
"description": "WebSocket library compatibility issue",
"created_date": "2024-02-07",
"resolved_date": "2024-02-11",
"resolution_days": 4,
"affected_stories": ["US-108"],
"category": "technical"
},
{
"id": "B-003",
"description": "Database migration pending approval",
"created_date": "2024-02-09",
"resolution_days": 0,
"affected_stories": ["US-110"],
"category": "process"
}
],
"ceremonies": {
"daily_standup": {
"attendance_rate": 0.88,
"engagement_score": 0.82
},
"sprint_planning": {
"attendance_rate": 0.96,
"engagement_score": 0.85
},
"sprint_review": {
"attendance_rate": 0.92,
"engagement_score": 0.83
},
"retrospective": {
"attendance_rate": 1.0,
"engagement_score": 0.87
}
}
},
{
"sprint_number": 4,
"sprint_name": "Sprint Delta",
"start_date": "2024-02-19",
"end_date": "2024-03-01",
"planned_points": 20,
"completed_points": 22,
"added_points": 2,
"removed_points": 0,
"carry_over_points": 2,
"team_capacity": 40,
"working_days": 10,
"team_size": 5,
"stories": [
{
"id": "US-108",
"title": "Real-time chat implementation",
"points": 13,
"status": "completed",
"assigned_to": "John Doe",
"created_date": "2024-02-05",
"completed_date": "2024-02-28",
"blocked_days": 0,
"priority": "high"
},
{
"id": "US-111",
"title": "Search functionality enhancement",
"points": 5,
"status": "completed",
"assigned_to": "Alice Brown",
"created_date": "2024-02-19",
"completed_date": "2024-02-26",
"blocked_days": 0,
"priority": "medium"
},
{
"id": "US-112",
"title": "Unit test coverage improvement",
"points": 3,
"status": "completed",
"assigned_to": "Bob Wilson",
"created_date": "2024-02-19",
"completed_date": "2024-02-27",
"blocked_days": 0,
"priority": "low"
},
{
"id": "US-113",
"title": "Error handling improvements",
"points": 1,
"status": "completed",
"assigned_to": "Jane Smith",
"created_date": "2024-02-25",
"completed_date": "2024-03-01",
"blocked_days": 0,
"priority": "medium"
}
],
"blockers": [],
"ceremonies": {
"daily_standup": {
"attendance_rate": 0.96,
"engagement_score": 0.90
},
"sprint_planning": {
"attendance_rate": 1.0,
"engagement_score": 0.94
},
"sprint_review": {
"attendance_rate": 1.0,
"engagement_score": 0.92
},
"retrospective": {
"attendance_rate": 1.0,
"engagement_score": 0.95
}
}
},
{
"sprint_number": 5,
"sprint_name": "Sprint Epsilon",
"start_date": "2024-03-04",
"end_date": "2024-03-15",
"planned_points": 25,
"completed_points": 17,
"added_points": 6,
"removed_points": 8,
"carry_over_points": 8,
"team_capacity": 35,
"working_days": 9,
"team_size": 4,
"stories": [
{
"id": "US-114",
"title": "Advanced analytics dashboard",
"points": 13,
"status": "blocked",
"assigned_to": "John Doe",
"created_date": "2024-03-04",
"blocked_days": 7,
"priority": "high"
},
{
"id": "US-115",
"title": "User permissions system",
"points": 8,
"status": "in_progress",
"assigned_to": "Alice Brown",
"created_date": "2024-03-04",
"blocked_days": 0,
"priority": "high"
},
{
"id": "US-116",
"title": "API rate limiting",
"points": 2,
"status": "completed",
"assigned_to": "Bob Wilson",
"created_date": "2024-03-04",
"completed_date": "2024-03-08",
"blocked_days": 0,
"priority": "medium"
},
{
"id": "US-117",
"title": "Documentation updates",
"points": 2,
"status": "completed",
"assigned_to": "Jane Smith",
"created_date": "2024-03-04",
"completed_date": "2024-03-10",
"blocked_days": 0,
"priority": "low"
}
],
"blockers": [
{
"id": "B-004",
"description": "Analytics service downtime",
"created_date": "2024-03-05",
"resolution_days": 0,
"affected_stories": ["US-114"],
"category": "external"
},
{
"id": "B-005",
"description": "Team member on sick leave",
"created_date": "2024-03-07",
"resolved_date": "2024-03-15",
"resolution_days": 8,
"affected_stories": ["US-115"],
"category": "team"
}
],
"ceremonies": {
"daily_standup": {
"attendance_rate": 0.75,
"engagement_score": 0.70
},
"sprint_planning": {
"attendance_rate": 0.80,
"engagement_score": 0.75
},
"sprint_review": {
"attendance_rate": 0.85,
"engagement_score": 0.78
},
"retrospective": {
"attendance_rate": 0.95,
"engagement_score": 0.88
}
}
},
{
"sprint_number": 6,
"sprint_name": "Sprint Zeta",
"start_date": "2024-03-18",
"end_date": "2024-03-29",
"planned_points": 22,
"completed_points": 24,
"added_points": 2,
"removed_points": 0,
"carry_over_points": 6,
"team_capacity": 45,
"working_days": 10,
"team_size": 5,
"stories": [
{
"id": "US-115",
"title": "User permissions system",
"points": 8,
"status": "completed",
"assigned_to": "Alice Brown",
"created_date": "2024-03-04",
"completed_date": "2024-03-25",
"blocked_days": 0,
"priority": "high"
},
{
"id": "US-118",
"title": "Backup and recovery system",
"points": 8,
"status": "completed",
"assigned_to": "John Doe",
"created_date": "2024-03-18",
"completed_date": "2024-03-28",
"blocked_days": 0,
"priority": "high"
},
{
"id": "US-119",
"title": "UI theme customization",
"points": 5,
"status": "completed",
"assigned_to": "Jane Smith",
"created_date": "2024-03-18",
"completed_date": "2024-03-26",
"blocked_days": 0,
"priority": "medium"
},
{
"id": "US-120",
"title": "Performance monitoring",
"points": 3,
"status": "completed",
"assigned_to": "Bob Wilson",
"created_date": "2024-03-18",
"completed_date": "2024-03-24",
"blocked_days": 0,
"priority": "low"
}
],
"blockers": [],
"ceremonies": {
"daily_standup": {
"attendance_rate": 0.98,
"engagement_score": 0.93
},
"sprint_planning": {
"attendance_rate": 1.0,
"engagement_score": 0.96
},
"sprint_review": {
"attendance_rate": 1.0,
"engagement_score": 0.94
},
"retrospective": {
"attendance_rate": 1.0,
"engagement_score": 0.97
}
}
}
],
"retrospectives": [
{
"sprint_number": 1,
"date": "2024-01-19",
"facilitator": "Sarah Chen",
"attendees": ["John Doe", "Jane Smith", "Bob Wilson", "Alice Brown", "Sarah Chen"],
"duration_minutes": 75,
"went_well": [
"Team collaboration was excellent during planning",
"Daily standups were efficient and focused",
"Good technical problem-solving on authentication system",
"New team member integrated well",
"Clear user story definitions"
],
"to_improve": [
"Story estimation accuracy needs work",
"Too many blockers appeared mid-sprint",
"API documentation was incomplete at start",
"Need better communication with external teams"
],
"action_items": [
{
"id": "AI-001",
"description": "Schedule estimation workshop for next sprint planning",
"owner": "Sarah Chen",
"priority": "high",
"due_date": "2024-01-26",
"status": "completed",
"created_sprint": 1,
"completed_sprint": 2,
"category": "process",
"effort_estimate": "medium"
},
{
"id": "AI-002",
"description": "Establish direct communication channel with API team",
"owner": "Bob Wilson",
"priority": "medium",
"due_date": "2024-01-30",
"status": "completed",
"created_sprint": 1,
"completed_sprint": 2,
"category": "communication",
"effort_estimate": "low"
},
{
"id": "AI-003",
"description": "Create blocker escalation process documentation",
"owner": "Sarah Chen",
"priority": "medium",
"due_date": "2024-02-02",
"status": "in_progress",
"created_sprint": 1,
"category": "process",
"effort_estimate": "low"
}
]
},
{
"sprint_number": 2,
"date": "2024-02-02",
"facilitator": "Sarah Chen",
"attendees": ["John Doe", "Jane Smith", "Bob Wilson", "Alice Brown", "Sarah Chen"],
"duration_minutes": 60,
"went_well": [
"Perfect sprint execution - completed all planned work",
"No blockers encountered",
"Estimation workshop improved accuracy significantly",
"Team velocity is stabilizing",
"Good ceremony attendance and engagement"
],
"to_improve": [
"Could have taken on more work given the smooth execution",
"Need to celebrate successes more",
"Sprint review could be more interactive",
"Documentation still lagging behind development"
],
"action_items": [
{
"id": "AI-004",
"description": "Implement team celebration ritual for successful sprints",
"owner": "Jane Smith",
"priority": "low",
"due_date": "2024-02-09",
"status": "completed",
"created_sprint": 2,
"completed_sprint": 3,
"category": "team_dynamics",
"effort_estimate": "low"
},
{
"id": "AI-005",
"description": "Create documentation sprint for next iteration",
"owner": "Alice Brown",
"priority": "medium",
"due_date": "2024-02-16",
"status": "cancelled",
"created_sprint": 2,
"category": "process",
"effort_estimate": "high"
}
]
},
{
"sprint_number": 3,
"date": "2024-02-16",
"facilitator": "John Doe",
"attendees": ["John Doe", "Jane Smith", "Bob Wilson", "Alice Brown"],
"duration_minutes": 90,
"went_well": [
"Good adaptation when faced with technical challenges",
"Team helped each other overcome blockers",
"Mobile design work exceeded expectations",
"Performance improvements had measurable impact"
],
"to_improve": [
"WebSocket integration took longer than expected",
"Too much scope change during the sprint",
"Daily standup attendance dropped",
"Need better technical spike planning",
"Database migration process is too slow"
],
"action_items": [
{
"id": "AI-006",
"description": "Schedule technical spike for complex integrations",
"owner": "John Doe",
"priority": "high",
"due_date": "2024-02-23",
"status": "completed",
"created_sprint": 3,
"completed_sprint": 4,
"category": "technical",
"effort_estimate": "medium"
},
{
"id": "AI-007",
"description": "Review scope change process with Product Owner",
"owner": "Sarah Chen",
"priority": "medium",
"due_date": "2024-02-26",
"status": "completed",
"created_sprint": 3,
"completed_sprint": 4,
"category": "process",
"effort_estimate": "low"
},
{
"id": "AI-008",
"description": "Improve database migration approval workflow",
"owner": "Bob Wilson",
"priority": "medium",
"due_date": "2024-03-08",
"status": "blocked",
"created_sprint": 3,
"category": "process",
"effort_estimate": "high"
}
]
},
{
"sprint_number": 4,
"date": "2024-03-01",
"facilitator": "Sarah Chen",
"attendees": ["John Doe", "Jane Smith", "Bob Wilson", "Alice Brown", "Sarah Chen"],
"duration_minutes": 45,
"went_well": [
"Exceeded sprint goal by completing extra work",
"Real-time chat finally delivered with high quality",
"Technical spikes prevented major blockers",
"Team ceremonies back to full engagement",
"Search functionality delivered ahead of schedule"
],
"to_improve": [
"Sprint retrospective was rushed due to time constraints",
"Need better capacity planning for variable team sizes",
"Unit test coverage still below target"
],
"action_items": [
{
"id": "AI-009",
"description": "Block more time for retrospectives in calendar",
"owner": "Sarah Chen",
"priority": "low",
"due_date": "2024-03-08",
"status": "completed",
"created_sprint": 4,
"completed_sprint": 5,
"category": "process",
"effort_estimate": "low"
},
{
"id": "AI-010",
"description": "Establish unit test coverage gates in CI/CD",
"owner": "Bob Wilson",
"priority": "high",
"due_date": "2024-03-15",
"status": "in_progress",
"created_sprint": 4,
"category": "technical",
"effort_estimate": "medium"
}
]
},
{
"sprint_number": 5,
"date": "2024-03-15",
"facilitator": "Alice Brown",
"attendees": ["John Doe", "Jane Smith", "Bob Wilson", "Alice Brown"],
"duration_minutes": 105,
"went_well": [
"Team adapted well to reduced capacity",
"Good support for team member on sick leave",
"Documentation work was delivered on time",
"Rate limiting implementation was smooth"
],
"to_improve": [
"External service dependencies caused major delays",
"Too much scope change again - need better discipline",
"Team capacity planning needs improvement",
"Daily standup attendance dropped significantly",
"Analytics service reliability is a recurring issue"
],
"action_items": [
{
"id": "AI-011",
"description": "Create external service dependency register",
"owner": "John Doe",
"priority": "high",
"due_date": "2024-03-22",
"status": "not_started",
"created_sprint": 5,
"category": "process",
"effort_estimate": "medium"
},
{
"id": "AI-012",
"description": "Escalate analytics service reliability issues",
"owner": "Sarah Chen",
"priority": "high",
"due_date": "2024-03-18",
"status": "completed",
"created_sprint": 5,
"completed_sprint": 6,
"category": "external",
"effort_estimate": "low"
},
{
"id": "AI-013",
"description": "Implement capacity planning buffer for sick leave",
"owner": "Sarah Chen",
"priority": "medium",
"due_date": "2024-03-29",
"status": "in_progress",
"created_sprint": 5,
"category": "process",
"effort_estimate": "medium"
}
]
},
{
"sprint_number": 6,
"date": "2024-03-29",
"facilitator": "Sarah Chen",
"attendees": ["John Doe", "Jane Smith", "Bob Wilson", "Alice Brown", "Sarah Chen"],
"duration_minutes": 70,
"went_well": [
"Excellent sprint execution with team back to full capacity",
"Delivered more points than planned",
"No blockers encountered",
"Strong ceremony engagement across all events",
"Backup system implementation was flawless",
"Team morale has improved significantly"
],
"to_improve": [
"Need to maintain this momentum",
"Could optimize sprint planning efficiency",
"Theme customization feature needs user feedback",
"Performance monitoring setup could be automated"
],
"action_items": [
{
"id": "AI-014",
"description": "Gather user feedback on theme customization",
"owner": "Jane Smith",
"priority": "medium",
"due_date": "2024-04-05",
"status": "not_started",
"created_sprint": 6,
"category": "external",
"effort_estimate": "low"
},
{
"id": "AI-015",
"description": "Automate performance monitoring setup",
"owner": "Bob Wilson",
"priority": "low",
"due_date": "2024-04-12",
"status": "not_started",
"created_sprint": 6,
"category": "technical",
"effort_estimate": "medium"
}
]
}
]
}
FILE:assets/sprint_report_template.md
# Sprint [NUMBER] - [SPRINT_NAME] Report
**Team:** [TEAM_NAME]
**Scrum Master:** [SCRUM_MASTER_NAME]
**Sprint Period:** [START_DATE] to [END_DATE]
**Report Date:** [REPORT_DATE]
---
## Executive Summary
**Sprint Goal Achievement:** [ACHIEVED/PARTIALLY_ACHIEVED/NOT_ACHIEVED]
**Overall Health Grade:** [EXCELLENT/GOOD/FAIR/POOR] ([HEALTH_SCORE]/100)
**Velocity:** [COMPLETED_POINTS] points ([VELOCITY_TREND] from previous sprint)
**Commitment Ratio:** [COMMITMENT_PERCENTAGE]% of planned work completed
### Key Highlights
- [KEY_ACHIEVEMENT_1]
- [KEY_ACHIEVEMENT_2]
- [KEY_CHALLENGE_1]
- [KEY_CHALLENGE_2]
---
## Sprint Metrics Dashboard
### Delivery Performance
| Metric | Value | Target | Status |
|--------|-------|---------|--------|
| **Planned Points** | [PLANNED_POINTS] | - | - |
| **Completed Points** | [COMPLETED_POINTS] | [TARGET_VELOCITY] | [ON_TRACK/BELOW/ABOVE] |
| **Commitment Ratio** | [COMMITMENT_PERCENTAGE]% | 85-100% | [EXCELLENT/GOOD/NEEDS_IMPROVEMENT] |
| **Stories Completed** | [COMPLETED_STORIES]/[TOTAL_STORIES] | 80%+ | [EXCELLENT/GOOD/NEEDS_IMPROVEMENT] |
| **Carry-over Points** | [CARRY_OVER_POINTS] | <20% | [GOOD/ACCEPTABLE/CONCERNING] |
### Process Health
| Metric | Value | Target | Status |
|--------|-------|---------|--------|
| **Scope Change** | [SCOPE_CHANGE_PERCENTAGE]% | <15% | [STABLE/MODERATE/UNSTABLE] |
| **Blocker Resolution** | [AVG_RESOLUTION_DAYS] days | <3 days | [EXCELLENT/GOOD/NEEDS_IMPROVEMENT] |
| **Daily Standup Attendance** | [STANDUP_ATTENDANCE]% | >90% | [EXCELLENT/GOOD/NEEDS_IMPROVEMENT] |
| **Retrospective Participation** | [RETRO_ATTENDANCE]% | >95% | [EXCELLENT/GOOD/NEEDS_IMPROVEMENT] |
### Quality Indicators
| Metric | Value | Target | Status |
|--------|-------|---------|--------|
| **Definition of Done Adherence** | [DOD_ADHERENCE]% | 100% | [EXCELLENT/NEEDS_IMPROVEMENT] |
| **Test Coverage** | [TEST_COVERAGE]% | >80% | [EXCELLENT/GOOD/NEEDS_IMPROVEMENT] |
| **Code Review Completion** | [CODE_REVIEW_COMPLETION]% | 100% | [EXCELLENT/NEEDS_IMPROVEMENT] |
| **Technical Debt Items** | [TECH_DEBT_ADDED]/[TECH_DEBT_RESOLVED] | Net negative | [IMPROVING/STABLE/CONCERNING] |
---
## User Stories Delivered
### Completed Stories ([COMPLETED_COUNT])
| Story ID | Title | Points | Owner | Completion Date | Notes |
|----------|-------|---------|-------|----------------|-------|
| [STORY_ID_1] | [STORY_TITLE_1] | [POINTS_1] | [OWNER_1] | [DATE_1] | [NOTES_1] |
| [STORY_ID_2] | [STORY_TITLE_2] | [POINTS_2] | [OWNER_2] | [DATE_2] | [NOTES_2] |
### In Progress Stories ([IN_PROGRESS_COUNT])
| Story ID | Title | Points | Owner | Progress | Expected Completion |
|----------|-------|---------|-------|----------|-------------------|
| [STORY_ID_3] | [STORY_TITLE_3] | [POINTS_3] | [OWNER_3] | [PROGRESS_3] | [ETA_3] |
### Blocked Stories ([BLOCKED_COUNT])
| Story ID | Title | Points | Owner | Blocker | Days Blocked | Escalation Status |
|----------|-------|---------|-------|---------|-------------|------------------|
| [STORY_ID_4] | [STORY_TITLE_4] | [POINTS_4] | [OWNER_4] | [BLOCKER_4] | [DAYS_4] | [ESCALATION_4] |
---
## Blockers & Impediments
### Resolved This Sprint ([RESOLVED_BLOCKERS_COUNT])
| ID | Description | Category | Created | Resolved | Resolution Time | Impact |
|----|-------------|----------|---------|----------|----------------|---------|
| [BLOCKER_ID_1] | [DESCRIPTION_1] | [CATEGORY_1] | [CREATED_1] | [RESOLVED_1] | [TIME_1] days | [IMPACT_1] |
### Active Blockers ([ACTIVE_BLOCKERS_COUNT])
| ID | Description | Category | Age | Owner | Next Steps | Priority |
|----|-------------|----------|-----|-------|------------|----------|
| [BLOCKER_ID_2] | [DESCRIPTION_2] | [CATEGORY_2] | [AGE_2] days | [OWNER_2] | [NEXT_STEPS_2] | [PRIORITY_2] |
### Escalation Required
- [ESCALATION_ITEM_1]
- [ESCALATION_ITEM_2]
---
## Team Performance Analysis
### Velocity Trend
```
Sprint [N-2]: [VELOCITY_N2] points
Sprint [N-1]: [VELOCITY_N1] points
Sprint [N]: [VELOCITY_N] points
Trend: [IMPROVING/STABLE/DECLINING] ([TREND_PERCENTAGE]% change)
```
### Predictability Assessment
- **Coefficient of Variation:** [CV_PERCENTAGE]% ([HIGH/MODERATE/LOW] volatility)
- **Commitment Reliability:** [COMMITMENT_RELIABILITY_SCORE]/100
- **Forecast Confidence:** [FORECAST_CONFIDENCE]% for next sprint
### Team Health Indicators
| Dimension | Score | Grade | Trend | Action Required |
|-----------|-------|--------|-------|-----------------|
| **Commitment Reliability** | [SCORE_1]/100 | [GRADE_1] | [TREND_1] | [ACTION_1] |
| **Scope Stability** | [SCORE_2]/100 | [GRADE_2] | [TREND_2] | [ACTION_2] |
| **Blocker Resolution** | [SCORE_3]/100 | [GRADE_3] | [TREND_3] | [ACTION_3] |
| **Ceremony Engagement** | [SCORE_4]/100 | [GRADE_4] | [TREND_4] | [ACTION_4] |
| **Story Completion** | [SCORE_5]/100 | [GRADE_5] | [TREND_5] | [ACTION_5] |
---
## Retrospective Insights
### What Went Well
- [WENT_WELL_1]
- [WENT_WELL_2]
- [WENT_WELL_3]
### Areas for Improvement
- [IMPROVE_1]
- [IMPROVE_2]
- [IMPROVE_3]
### Action Items from Retrospective
| ID | Action | Owner | Due Date | Priority | Status |
|----|--------|-------|----------|----------|--------|
| [AI_ID_1] | [ACTION_1] | [OWNER_1] | [DUE_1] | [PRIORITY_1] | [STATUS_1] |
| [AI_ID_2] | [ACTION_2] | [OWNER_2] | [DUE_2] | [PRIORITY_2] | [STATUS_2] |
### Previous Sprint Action Items Follow-up
| ID | Action | Owner | Status | Completion Notes |
|----|--------|-------|--------|------------------|
| [PREV_AI_1] | [PREV_ACTION_1] | [PREV_OWNER_1] | [PREV_STATUS_1] | [PREV_NOTES_1] |
---
## Risks & Dependencies
### High Priority Risks
| Risk | Probability | Impact | Mitigation Plan | Owner |
|------|-------------|---------|-----------------|-------|
| [RISK_1] | [PROB_1] | [IMPACT_1] | [MITIGATION_1] | [OWNER_1] |
### External Dependencies
| Dependency | Provider | Status | Expected Resolution | Contingency Plan |
|------------|----------|--------|---------------------|------------------|
| [DEP_1] | [PROVIDER_1] | [STATUS_1] | [RESOLUTION_1] | [CONTINGENCY_1] |
---
## Looking Ahead: Next Sprint
### Sprint Goals
1. [GOAL_1]
2. [GOAL_2]
3. [GOAL_3]
### Planned Capacity
- **Team Size:** [TEAM_SIZE] members
- **Available Capacity:** [AVAILABLE_HOURS] hours ([CAPACITY_POINTS] points)
- **Planned Velocity:** [PLANNED_VELOCITY] points
- **Capacity Buffer:** [BUFFER_PERCENTAGE]% for unknowns
### Key Focus Areas
- [FOCUS_AREA_1]
- [FOCUS_AREA_2]
- [FOCUS_AREA_3]
### Dependencies to Monitor
- [MONITOR_DEP_1]
- [MONITOR_DEP_2]
---
## Recommendations
### Immediate Actions (This Sprint)
1. **[HIGH_PRIORITY_ACTION_1]** - [DESCRIPTION] (Owner: [OWNER], Due: [DATE])
2. **[HIGH_PRIORITY_ACTION_2]** - [DESCRIPTION] (Owner: [OWNER], Due: [DATE])
### Process Improvements (Next 2-3 Sprints)
1. **[PROCESS_IMPROVEMENT_1]** - [DESCRIPTION]
2. **[PROCESS_IMPROVEMENT_2]** - [DESCRIPTION]
### Team Development Opportunities
1. **[DEVELOPMENT_1]** - [DESCRIPTION]
2. **[DEVELOPMENT_2]** - [DESCRIPTION]
---
## Appendix
### Sprint Burndown Chart
[BURNDOWN_CHART_REFERENCE]
### Detailed Metrics
[DETAILED_METRICS_REFERENCE]
### Team Feedback
[TEAM_FEEDBACK_SUMMARY]
---
**Report prepared by:** [SCRUM_MASTER_NAME]
**Next review date:** [NEXT_REVIEW_DATE]
**Distribution:** Product Owner, Development Team, Stakeholders
---
*This report is generated using standardized sprint health metrics and retrospective analysis. For questions or deeper analysis, please contact the Scrum Master.*
FILE:assets/team_health_check_template.md
# Team Health Check - Spotify Squad Model
**Team:** [TEAM_NAME]
**Assessment Date:** [DATE]
**Facilitator:** [FACILITATOR_NAME]
**Participants:** [PARTICIPANT_COUNT] of [TOTAL_TEAM_SIZE] members
---
## Health Check Overview
The Team Health Check is based on Spotify's Squad Health Check model, designed to visualize team health across multiple dimensions. Each dimension is assessed using a simple traffic light system:
- 🟢 **Green (Awesome):** We're doing great! No major concerns.
- 🟡 **Yellow (Some Concerns):** We're doing okay, but there are some things we could improve.
- 🔴 **Red (Not Good):** This really sucks and we need to do something about it.
### Assessment Method
- Anonymous individual ratings followed by team discussion
- Focus on trends over time rather than absolute scores
- Action-oriented outcomes for improvement areas
---
## Health Dimensions Assessment
### 1. Delivering Value 🎯
*Are we delivering value to our users and stakeholders?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
### 2. Learning 📚
*Are we learning and growing as individuals and as a team?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
### 3. Fun 🎉
*Do we enjoy working together and find our work engaging?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
### 4. Health of Codebase 🏗️
*Is our code healthy, maintainable, and of good quality?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
### 5. Mission Clarity 🎯
*Do we understand why we exist and what we're supposed to achieve?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
### 6. Suitable Process ⚙️
*Is our process helping us be effective?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
### 7. Support 🤝
*Do we get the support we need from management and other teams?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
### 8. Speed ⚡
*Are we able to deliver quickly without compromising quality?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
### 9. Pawns or Players 👥
*Do we feel like we have control over our work and destiny?*
**Current Status:** [🟢/🟡/🔴]
**Trend from Last Check:** [⬆️ Improving / ➡️ Stable / ⬇️ Declining]
**Team Rating:** [X]/5 team members voted Green, [Y]/5 Yellow, [Z]/5 Red
**What's Working Well:**
- [POSITIVE_POINT_1]
- [POSITIVE_POINT_2]
**Areas of Concern:**
- [CONCERN_1]
- [CONCERN_2]
**Suggested Actions:**
- [ACTION_1]
- [ACTION_2]
---
## Overall Health Summary
### Health Score Distribution
- 🟢 **Green Dimensions:** [GREEN_COUNT]/9 ([GREEN_PERCENTAGE]%)
- 🟡 **Yellow Dimensions:** [YELLOW_COUNT]/9 ([YELLOW_PERCENTAGE]%)
- 🔴 **Red Dimensions:** [RED_COUNT]/9 ([RED_PERCENTAGE]%)
### Overall Health Grade: [EXCELLENT/GOOD/FAIR/POOR]
### Trend Analysis
- **Improving:** [IMPROVING_COUNT] dimensions
- **Stable:** [STABLE_COUNT] dimensions
- **Declining:** [DECLINING_COUNT] dimensions
### Team Maturity Level
Based on the health check results and team dynamics observed:
**[FORMING/STORMING/NORMING/PERFORMING/ADJOURNING]**
---
## Priority Action Items
### High Priority (Red Dimensions)
1. **[RED_DIMENSION_1]:** [ACTION_DESCRIPTION_1]
- Owner: [OWNER_1]
- Timeline: [TIMELINE_1]
- Success Criteria: [CRITERIA_1]
2. **[RED_DIMENSION_2]:** [ACTION_DESCRIPTION_2]
- Owner: [OWNER_2]
- Timeline: [TIMELINE_2]
- Success Criteria: [CRITERIA_2]
### Medium Priority (Yellow Dimensions)
1. **[YELLOW_DIMENSION_1]:** [ACTION_DESCRIPTION_1]
- Owner: [OWNER_1]
- Timeline: [TIMELINE_1]
2. **[YELLOW_DIMENSION_2]:** [ACTION_DESCRIPTION_2]
- Owner: [OWNER_2]
- Timeline: [TIMELINE_2]
### Maintain Strengths (Green Dimensions)
1. **[GREEN_DIMENSION_1]:** Continue [STRENGTH_PRACTICE_1]
2. **[GREEN_DIMENSION_2]:** Share [BEST_PRACTICE_1] with other teams
---
## Psychological Safety Assessment
*Separate anonymous assessment of team psychological safety*
### Psychological Safety Indicators
1. **Speaking Up:** Team members feel safe to speak up with ideas, questions, concerns, or mistakes
- Score: [SCORE_1]/5 ⭐⭐⭐⭐⭐
2. **Risk Taking:** Team members feel safe to take risks and make mistakes
- Score: [SCORE_2]/5 ⭐⭐⭐⭐⭐
3. **Asking for Help:** Team members feel comfortable asking for help or admitting they don't know something
- Score: [SCORE_3]/5 ⭐⭐⭐⭐⭐
4. **Discussing Problems:** Difficult topics and problems can be discussed openly
- Score: [SCORE_4]/5 ⭐⭐⭐⭐⭐
5. **Being Yourself:** Team members don't feel they have to pretend to be someone else
- Score: [SCORE_5]/5 ⭐⭐⭐⭐⭐
**Overall Psychological Safety Score:** [TOTAL_SCORE]/25
### Psychological Safety Actions
- [PSYCH_SAFETY_ACTION_1]
- [PSYCH_SAFETY_ACTION_2]
---
## Communication & Collaboration Assessment
### Communication Quality
- **Clarity of Communication:** [SCORE]/5 ⭐⭐⭐⭐⭐
- **Frequency of Communication:** [SCORE]/5 ⭐⭐⭐⭐⭐
- **Openness & Transparency:** [SCORE]/5 ⭐⭐⭐⭐⭐
### Collaboration Patterns
- **Cross-functional Collaboration:** [SCORE]/5 ⭐⭐⭐⭐⭐
- **Knowledge Sharing:** [SCORE]/5 ⭐⭐⭐⭐⭐
- **Conflict Resolution:** [SCORE]/5 ⭐⭐⭐⭐⭐
---
## Follow-up Plan
### Next Health Check
**Scheduled Date:** [NEXT_DATE]
**Frequency:** [MONTHLY/QUARTERLY/BI-ANNUAL]
### Interim Check-ins
- **Sprint Retrospectives:** Continue monitoring health indicators
- **Weekly 1:1s:** Individual pulse checks with team members
- **Monthly Team Lunches:** Informal health and morale assessment
### Success Metrics
We'll know we're improving when we see:
- [SUCCESS_METRIC_1]
- [SUCCESS_METRIC_2]
- [SUCCESS_METRIC_3]
---
## Historical Comparison
### Previous Health Checks
| Date | Green | Yellow | Red | Overall Trend |
|------|-------|--------|-----|---------------|
| [PREV_DATE_1] | [G1] | [Y1] | [R1] | [TREND_1] |
| [PREV_DATE_2] | [G2] | [Y2] | [R2] | [TREND_2] |
| [CURRENT_DATE] | [G3] | [Y3] | [R3] | [TREND_3] |
### Long-term Improvements
- [LONG_TERM_IMPROVEMENT_1]
- [LONG_TERM_IMPROVEMENT_2]
### Persistent Challenges
- [PERSISTENT_CHALLENGE_1]
- [PERSISTENT_CHALLENGE_2]
---
## Team Comments & Feedback
*Anonymous feedback from team members*
### What's the most important thing we should focus on?
- "[FEEDBACK_1]"
- "[FEEDBACK_2]"
- "[FEEDBACK_3]"
### What's our biggest strength as a team?
- "[STRENGTH_1]"
- "[STRENGTH_2]"
- "[STRENGTH_3]"
### If you could change one thing, what would it be?
- "[CHANGE_1]"
- "[CHANGE_2]"
- "[CHANGE_3]"
---
## Action Item Summary
| Priority | Action | Owner | Due Date | Success Criteria | Status |
|----------|---------|-------|----------|------------------|--------|
| High | [ACTION_1] | [OWNER_1] | [DATE_1] | [CRITERIA_1] | [STATUS_1] |
| High | [ACTION_2] | [OWNER_2] | [DATE_2] | [CRITERIA_2] | [STATUS_2] |
| Medium | [ACTION_3] | [OWNER_3] | [DATE_3] | [CRITERIA_3] | [STATUS_3] |
| Medium | [ACTION_4] | [OWNER_4] | [DATE_4] | [CRITERIA_4] | [STATUS_4] |
---
**Assessment completed by:** [FACILITATOR_NAME]
**Report distribution:** Team Members, Product Owner, Management (summary only)
**Confidentiality:** Individual responses kept confidential, only aggregate data shared
---
*This health check is based on the Spotify Squad Health Check model. The goal is continuous improvement, not judgment. Use this data to have better conversations about how to work together effectively.*
FILE:references/retro-formats.md
# Sprint Retrospective Formats
## Start/Stop/Continue
**Best for:** Teams new to retrospectives, quick format
**Duration:** 45-60 minutes
### Structure
Create three columns:
- **Start:** What should we begin doing?
- **Stop:** What should we stop doing?
- **Continue:** What's working well that we should keep doing?
### Process
1. Team silently adds items to each column (10 min)
2. Group similar items (5 min)
3. Discuss each category, vote on top items (20 min)
4. Select 2-3 actions (10 min)
### Example Output
**Start:**
- Pairing on complex stories
- Code reviews within 4 hours
**Stop:**
- Taking on work mid-sprint
- Skipping acceptance criteria
**Continue:**
- Daily standups at 9:30am
- Demo prep on Thursday
---
## Glad/Sad/Mad
**Best for:** Emotional check-in, team morale assessment
**Duration:** 60-75 minutes
### Structure
Create three areas:
- **Glad:** What made you happy this sprint?
- **Sad:** What disappointed you?
- **Mad:** What frustrated you?
### Process
1. Silent brainstorming (10 min)
2. Share items, one person at a time (15 min)
3. Group themes (5 min)
4. Discuss top items from each category (20 min)
5. Identify action items (10 min)
### Example Output
**Glad:**
- Shipped feature X on time
- Great collaboration with design team
- New deployment process worked well
**Sad:**
- Lost time to production bugs
- Didn't finish all committed work
- Documentation fell behind
**Mad:**
- Environment was down 2 days
- Requirements changed mid-sprint
- Still waiting on API key from vendor
### Facilitation Tips
- Acknowledge emotions, don't dismiss
- Focus on what we can control
- Convert frustrations into actions
---
## 4Ls (Liked, Learned, Lacked, Longed For)
**Best for:** Deeper reflection, learning focus
**Duration:** 60-90 minutes
### Structure
- **Liked:** What went well? What did we enjoy?
- **Learned:** What new insights did we gain?
- **Lacked:** What was missing? What did we need?
- **Longed For:** What do we wish we had?
### Process
1. Individual reflection (10 min)
2. Round-robin sharing (20 min)
3. Group similar items (10 min)
4. Deep dive on top items (20 min)
5. Action planning (15 min)
### Example Output
**Liked:**
- Pair programming sessions
- Clear acceptance criteria
- Product Owner availability
**Learned:**
- New testing framework capabilities
- How to better estimate stories
- Importance of architectural review
**Lacked:**
- Automated deployment
- Clear API documentation
- Sufficient testing time
**Longed For:**
- Better development environments
- More design time upfront
- Dedicated QA support
---
## Sailboat
**Best for:** Visual teams, identifying headwinds and tailwinds
**Duration:** 60-90 minutes
### Structure
Draw a sailboat with:
- **Wind (propellers):** What's helping us go faster?
- **Anchors:** What's slowing us down?
- **Rocks (hazards):** What risks are ahead?
- **Island (goal):** Where are we headed?
### Process
1. Explain metaphor (5 min)
2. Team adds sticky notes to each area (15 min)
3. Group and discuss each area (30 min)
4. Prioritize anchors to remove (10 min)
5. Create action plan (15 min)
### Example Output
**Wind:**
- Strong team collaboration
- Clear product vision
- Good tooling
**Anchors:**
- Slow CI/CD pipeline
- Too many meetings
- Technical debt
**Rocks:**
- Upcoming dependency on Team B
- Key person on vacation next sprint
- Infrastructure migration
**Island:**
- Launch v2.0 by end of quarter
- Improve system stability
- Reduce production bugs by 50%
---
## Timeline
**Best for:** Detailed sprint review, identifying patterns
**Duration:** 75-90 minutes
### Structure
Create a timeline of the sprint on a whiteboard:
- Days of the sprint across the top
- Events, milestones, feelings plotted on timeline
### Process
1. Draw sprint timeline (5 min)
2. Team adds events chronologically (15 min)
3. Add emotion indicators (happy/sad/stressed) (10 min)
4. Identify patterns and themes (20 min)
5. Discuss high/low points (20 min)
6. Extract learnings and actions (15 min)
### Example Timeline
```
Day 1: Sprint planning, feeling optimistic 😊
Day 3: Production bug discovered, stressed 😰
Day 5: Bug fixed, relieved 😌
Day 7: Design feedback changed scope, frustrated 😠
Day 9: Great pairing session on new feature 😊
Day 10: Demo went really well! 🎉
```
### Facilitation Tips
- Focus on objective events first, emotions second
- Look for correlations between events and feelings
- Identify early warning signs
- Celebrate wins
---
## Starfish
**Best for:** More granular feedback than Start/Stop/Continue
**Duration:** 60-90 minutes
### Structure
Five categories:
- **Keep Doing:** What's working, don't change
- **Less Of:** What should we reduce?
- **More Of:** What should we increase?
- **Stop Doing:** What should we eliminate?
- **Start Doing:** What new practices should we try?
### Process
1. Explain each category (5 min)
2. Silent brainstorming (15 min)
3. Share and group items (15 min)
4. Discuss each category (25 min)
5. Vote on top actions (10 min)
6. Create action plan (15 min)
### Example Output
**Keep Doing:**
- Pairing on complex stories
- Demo every Friday
**Less Of:**
- Context switching
- Unplanned work
**More Of:**
- Automated testing
- Design upfront
**Stop Doing:**
- Skipping code reviews
- Working weekends
**Start Doing:**
- Mob programming for knowledge sharing
- Weekly architecture discussions
---
## Speed Dating
**Best for:** Large teams, fresh perspectives
**Duration:** 60 minutes
### Structure
- Pair up team members who don't usually work together
- Rotate pairs every 10 minutes
- Discuss sprint from different perspectives
### Process
1. Create pairs (2 min)
2. Round 1: "What went well?" (10 min)
3. Rotate pairs (2 min)
4. Round 2: "What could improve?" (10 min)
5. Rotate pairs (2 min)
6. Round 3: "What should we try?" (10 min)
7. Full group synthesis (15 min)
8. Action planning (10 min)
### Facilitation Tips
- Ensure quiet voices are heard
- Mix up pairs intentionally
- Capture themes as they emerge
- Focus on shared themes in synthesis
---
## Three Little Pigs
**Best for:** Architecture and technical decisions
**Duration:** 60-75 minutes
### Structure
Based on the story:
- **Straw House:** What's fragile? What will blow down?
- **Stick House:** What's okay but could be better?
- **Brick House:** What's solid and will last?
### Process
1. Explain metaphor (5 min)
2. Team identifies items for each house (15 min)
3. Group and discuss (20 min)
4. Prioritize straw house items to fix (10 min)
5. Create action plan (15 min)
### Example Output
**Straw House (fragile):**
- Manual deployment process
- No automated tests for API
- Undocumented code
**Stick House (needs improvement):**
- Test coverage at 60%
- Some documentation exists
- Partially automated builds
**Brick House (solid):**
- Strong CI/CD for frontend
- Well-tested core modules
- Clear architecture docs
---
## Facilitation Best Practices
### Before Retrospective
- Review previous action items
- Gather sprint metrics
- Choose format based on team needs
- Prepare collaboration space
### During Retrospective
- **Set the stage:** Create safe environment
- **Prime directive:** "Regardless of what we discover, we understand and truly believe that everyone did the best job they could, given what they knew at the time, their skills and abilities, the resources available, and the situation at hand."
- **Timebox discussions:** Keep energy high
- **Focus on actions:** Not just talk
- **Limit action items:** 1-3 max for next sprint
- **Get specific:** Vague actions don't happen
### After Retrospective
- Document immediately in Confluence
- Create Jira tickets for actions
- Assign owners and due dates
- Track completion
- Start next retro by reviewing these
### Red Flags
- Same issues every retro → Need deeper intervention
- No action items → Team not engaged
- Blame game → Not safe environment
- No follow-through → Actions not valued
- Facilitator talks more than team → Not facilitating
### Rotation Strategy
- Vary formats every 2-3 sprints
- Let team choose occasionally
- Match format to team mood
- Try new format when stuck
FILE:references/team-dynamics-framework.md
# Team Dynamics Framework for Scrum Teams
## Table of Contents
- [Overview](#overview)
- [Tuckman's Model Applied to Scrum](#tuckmans-model-applied-to-scrum)
- [Psychological Safety in Agile Teams](#psychological-safety-in-agile-teams)
- [Team Performance Metrics](#team-performance-metrics)
- [Facilitation Techniques by Stage](#facilitation-techniques-by-stage)
- [Conflict Resolution Strategies](#conflict-resolution-strategies)
- [Assessment Tools](#assessment-tools)
- [Intervention Strategies](#intervention-strategies)
- [Measurement & Tracking](#measurement--tracking)
---
## Overview
Understanding team dynamics is crucial for Scrum Masters to effectively guide teams through their development journey. This framework combines Tuckman's stages of group development with psychological safety principles and practical scrum-specific interventions.
### Core Principles
1. **Development is Non-Linear**: Teams may cycle between stages based on changes
2. **Each Stage Has Value**: Every stage serves a purpose in team development
3. **Facilitation Must Adapt**: Leadership style should match the team's developmental stage
4. **Psychological Safety is Foundational**: Without safety, teams cannot reach high performance
5. **Measurement Enables Improvement**: Track dynamics to guide interventions
### Framework Components
- **Tuckman's Stages**: Forming → Storming → Norming → Performing → Adjourning
- **Psychological Safety**: Environment for risk-taking and learning
- **Scrum Ceremonies**: Team development accelerators when facilitated well
- **Metrics & Assessment**: Data-driven approach to team health
---
## Tuckman's Model Applied to Scrum
### Stage 1: Forming (Team Inception)
*"Getting to know each other and understanding the work"*
#### Characteristics in Scrum Context
- **Individual Focus**: Members work independently, unsure of roles
- **Politeness**: Conflict is avoided, everyone tries to be agreeable
- **Dependency**: Heavy reliance on Scrum Master for guidance
- **Ceremony Awkwardness**: Standups feel forced, retrospectives are superficial
- **Low Velocity**: Productivity is low as team learns to work together
#### Scrum Master Behaviors
- **Directing Style**: Provide clear structure and guidance
- **Process Champion**: Teach scrum framework and ceremonies rigorously
- **Relationship Builder**: Facilitate team bonding and trust building
- **Context Setter**: Explain the "why" behind practices and goals
#### Key Metrics & Indicators
| Metric | Forming Range | Assessment Method |
|--------|---------------|-------------------|
| Ceremony Participation | 60-80% | Attendance tracking |
| Cross-team Collaboration | Low | Story pairing frequency |
| Velocity Predictability | High volatility (CV >40%) | Velocity coefficient of variation |
| Psychological Safety | 2.0-3.5/5.0 | Anonymous team survey |
| Conflict Frequency | Very low | Retrospective themes |
#### Intervention Strategies
- **Team Charter Creation**: Define working agreements and values together
- **Skill Inventory**: Map team capabilities and identify knowledge gaps
- **Pairing/Mobbing**: Encourage collaborative work to build relationships
- **Social Activities**: Team lunches, informal interactions
- **Process Education**: Intensive scrum training and coaching
#### Success Indicators
- Consistent ceremony attendance (>85%)
- Team members start asking questions about process
- Initial working agreements are established
- Some cross-functional collaboration begins
---
### Stage 2: Storming (Productive Conflict)
*"Working through differences and establishing team dynamics"*
#### Characteristics in Scrum Context
- **Conflict Emergence**: Disagreements about technical approaches, priorities
- **Role Struggles**: Tension around responsibilities and decision-making authority
- **Process Pushback**: Questioning scrum practices, suggesting changes
- **Subgroup Formation**: Cliques or mini-alliances may form
- **Velocity Fluctuations**: Performance varies as team works through conflicts
#### Scrum Master Behaviors
- **Coaching Style**: Guide conflict resolution without directing solutions
- **Neutral Facilitator**: Help team work through disagreements constructively
- **Psychological Safety Guardian**: Ensure conflicts remain productive
- **Process Flexibility**: Adapt ceremonies to team's evolving needs
#### Key Metrics & Indicators
| Metric | Storming Range | Assessment Method |
|--------|---------------|-------------------|
| Conflict Frequency | Moderate-High | Retrospective action items |
| Ceremony Engagement | Variable (70-90%) | Participation quality scoring |
| Velocity Volatility | Moderate (CV 25-40%) | Sprint-to-sprint variation |
| Psychological Safety | 2.5-4.0/5.0 | Team surveys + observation |
| Process Adherence | Inconsistent | Ceremony audit scores |
#### Intervention Strategies
- **Conflict Facilitation**: Structured conflict resolution sessions
- **Retrospective Focus**: Deep-dive into team dynamics and relationships
- **Individual Coaching**: 1:1s to address personal concerns and conflicts
- **Working Agreement Updates**: Revisit and refine team agreements
- **External Facilitation**: Bring in neutral parties for significant conflicts
#### Success Indicators
- Conflicts are addressed openly rather than avoided
- Team develops mechanisms for working through disagreements
- Ceremony participation becomes more authentic
- Velocity starts to stabilize
---
### Stage 3: Norming (Agreement & Collaboration)
*"Establishing effective ways of working together"*
#### Characteristics in Scrum Context
- **Shared Ownership**: Team takes collective responsibility for outcomes
- **Process Refinement**: Self-organizing improvements to scrum practices
- **Collaboration Increase**: More cross-functional pairing and knowledge sharing
- **Ceremony Effectiveness**: Meetings become more focused and productive
- **Velocity Stabilization**: More predictable delivery patterns emerge
#### Scrum Master Behaviors
- **Supporting Style**: Step back and let team lead, provide support when needed
- **Impediment Remover**: Focus on external blockers and organizational issues
- **Continuous Improvement Coach**: Help team identify and implement improvements
- **Shield Provider**: Protect team from external disruptions
#### Key Metrics & Indicators
| Metric | Norming Range | Assessment Method |
|--------|---------------|-------------------|
| Self-Organization | Increasing | Decision-making autonomy tracking |
| Ceremony Effectiveness | 80-90% | Time-to-value ratios |
| Velocity Consistency | Good (CV 15-25%) | Rolling average stability |
| Psychological Safety | 3.5-4.5/5.0 | Regular pulse surveys |
| Knowledge Sharing | High | Cross-training metrics |
#### Intervention Strategies
- **Process Ownership Transfer**: Guide team to own ceremony facilitation
- **Skill Development**: Focus on technical and collaboration skills
- **Measurement Introduction**: Help team define their own success metrics
- **External Relationship Building**: Facilitate connections with other teams
- **Continuous Improvement Rhythm**: Establish regular process refinement
#### Success Indicators
- Team members facilitate some ceremonies themselves
- Proactive identification and resolution of impediments
- Stable, predictable velocity patterns
- High-quality retrospectives with actionable outcomes
---
### Stage 4: Performing (High Performance)
*"Delivering exceptional results together"*
#### Characteristics in Scrum Context
- **Collective Excellence**: Team consistently exceeds expectations
- **Adaptive Expertise**: Quick response to changing requirements
- **Self-Management**: Minimal need for external direction
- **Innovation**: Team generates creative solutions and process improvements
- **Knowledge Multiplication**: Members actively develop others
#### Scrum Master Behaviors
- **Delegating Style**: Minimal intervention, team is largely autonomous
- **Strategic Facilitator**: Focus on long-term team development and capability
- **Organizational Catalyst**: Help team influence broader organizational change
- **Mentor Developer**: Coach team members to become coaches themselves
#### Key Metrics & Indicators
| Metric | Performing Range | Assessment Method |
|--------|---------------|-------------------|
| Autonomy Level | High | Decision independence tracking |
| Innovation Frequency | Regular | New idea implementation rate |
| Velocity Excellence | High + Consistent (CV <15%) | Performance benchmarking |
| Psychological Safety | 4.0-5.0/5.0 | Team assessment + observation |
| External Impact | Significant | Other teams adopting practices |
#### Intervention Strategies
- **Challenge Provision**: Introduce stretch goals and complex problems
- **Leadership Development**: Grow team members into coaches/leaders
- **Knowledge Sharing**: Facilitate teaching other teams
- **Strategic Alignment**: Connect team excellence to organizational goals
- **Innovation Support**: Create space for experimentation and learning
#### Success Indicators
- Consistent delivery of high-quality work with minimal defects
- Team serves as a model for other teams in the organization
- Members are sought out for coaching and mentoring roles
- Proactive contribution to organizational process improvements
---
### Stage 5: Adjourning (Transition & Legacy)
*"Wrapping up and transitioning knowledge"*
#### Characteristics in Scrum Context
- **Closure Activities**: Project completion or team dissolution
- **Knowledge Transfer**: Documenting learnings and sharing expertise
- **Relationship Maintenance**: Preserving professional networks
- **Legacy Creation**: Ensuring practices continue beyond the team
- **Emotional Processing**: Addressing feelings about team ending
#### Scrum Master Behaviors
- **Closure Facilitator**: Guide proper conclusion of work and relationships
- **Legacy Curator**: Ensure knowledge and practices are preserved
- **Transition Planner**: Help members move to new roles/teams effectively
- **Emotional Support**: Acknowledge and process team disbanding feelings
#### Key Activities
- **Final Retrospective**: Comprehensive review of team journey and learnings
- **Practice Documentation**: Record effective processes for future teams
- **Knowledge Transfer Sessions**: Share expertise with successor teams
- **Celebration**: Acknowledge achievements and relationships built
- **Network Maintenance**: Establish ongoing professional connections
---
## Psychological Safety in Agile Teams
### Definition & Importance
Psychological safety is the belief that one can show vulnerability, ask questions, admit mistakes, and propose ideas without risk of negative consequences to self-image, status, or career.
### Google's Four Components Applied to Scrum
1. **Ability to show vulnerability and ask for help**
2. **Permission to discuss difficult topics and disagreements**
3. **Freedom to take risks and make mistakes**
4. **Encouragement to be authentic and express oneself**
### Building Psychological Safety in Scrum Teams
#### Daily Standups
- **Model Vulnerability**: Scrum Master admits own mistakes and uncertainties
- **Normalize Help-Seeking**: "Who needs help?" vs. "Any blockers?"
- **Celebrate Learning**: Highlight lessons learned from failures
- **Time Protection**: Ensure everyone has space to speak
#### Sprint Planning
- **Estimation Comfort**: No judgment for "wrong" estimates
- **Capacity Honesty**: Safe to express realistic availability
- **Question Encouragement**: Reward curiosity and clarification requests
- **Scope Negotiation**: Team can push back on unrealistic commitments
#### Sprint Reviews
- **Failure Normalization**: Discuss what didn't work without blame
- **Stakeholder Preparation**: Coach stakeholders on constructive feedback
- **Team Support**: Unified front when facing criticism
- **Learning Focus**: Frame setbacks as learning opportunities
#### Retrospectives
- **Non-Judgmental Space**: Focus on systems, not individuals
- **Equal Participation**: Ensure all voices are heard
- **Actionable Outcomes**: Team commits to improvements together
- **Confidentiality**: What's said in retro stays in retro
### Measuring Psychological Safety
#### Edmondson's 7-Point Scale
1. If you make a mistake on this team, it is often held against you
2. Members of this team are able to bring up problems and tough issues
3. People on this team sometimes reject others for being different
4. It is safe to take a risk on this team
5. It is difficult to ask other members of this team for help
6. No one on this team would deliberately act to undermine my efforts
7. Working with members of this team, my unique skills and talents are valued and utilized
#### Practical Assessment Questions
- **Risk Taking**: "Do team members speak up when they disagree with leadership?"
- **Mistake Handling**: "How does the team respond when someone makes an error?"
- **Help Seeking**: "Do people admit when they don't know something?"
- **Inclusion**: "Are all team members' ideas heard and considered?"
- **Innovation**: "Does the team experiment with new approaches?"
---
## Team Performance Metrics
### Quantitative Indicators
#### Velocity & Predictability
- **Sprint Velocity Trends**: Improvement over time indicates team development
- **Commitment Reliability**: Ability to deliver planned work consistently
- **Velocity Volatility (CV)**: Lower variation indicates team maturity
- **Forecast Accuracy**: Precision in release planning improves with development
#### Quality Metrics
- **Defect Rates**: High-performing teams have lower defect introduction
- **Definition of Done Adherence**: Mature teams consistently meet quality criteria
- **Technical Debt Management**: Performing teams proactively address debt
- **Customer Satisfaction**: Ultimately reflected in user/stakeholder feedback
#### Collaboration Indicators
- **Cross-functional Work**: Story completion without handoffs
- **Knowledge Sharing**: Pair programming, code review participation
- **Skill Development**: Team members learning from each other
- **Collective Ownership**: Shared responsibility for all team outputs
### Qualitative Assessments
#### Ceremony Quality
- **Engagement Level**: Active participation vs. passive attendance
- **Value Generation**: Productive outcomes from time invested
- **Self-Facilitation**: Team taking ownership of meeting effectiveness
- **Adaptation**: Tailoring practices to team's specific needs
#### Communication Patterns
- **Openness**: Willingness to share problems and concerns
- **Constructive Conflict**: Disagreements lead to better solutions
- **Active Listening**: Team members build on each other's ideas
- **Feedback Culture**: Regular, specific, actionable feedback exchange
---
## Facilitation Techniques by Stage
### Forming Stage Facilitation
- **Structured Introductions**: Personal/professional background sharing
- **Explicit Process Teaching**: Step-by-step ceremony instruction
- **Role Clarification**: Clear explanation of responsibilities and expectations
- **Safe-to-Fail Experiments**: Low-risk opportunities to try new things
### Storming Stage Facilitation
- **Conflict Normalization**: "Conflict is healthy and expected"
- **Ground Rules Enforcement**: Maintain respectful disagreement standards
- **Perspective Taking**: Help team members understand different viewpoints
- **External Processing**: Individual coaching sessions for complex issues
### Norming Stage Facilitation
- **Autonomy Building**: Gradually reduce direct intervention
- **Process Ownership Transfer**: Team takes responsibility for improvements
- **Skill Gap Identification**: Focus on capability development
- **Success Pattern Recognition**: Help team understand what's working
### Performing Stage Facilitation
- **Challenge Introduction**: Stretch goals and complex problems
- **Innovation Support**: Time and space for experimentation
- **Teaching Opportunities**: Help team share knowledge with others
- **Strategic Connection**: Link team excellence to organizational goals
---
## Conflict Resolution Strategies
### Healthy vs. Unhealthy Conflict
#### Healthy Conflict Characteristics
- **Task-Focused**: About work, not personalities
- **Solution-Oriented**: Aimed at finding better ways forward
- **Open and Direct**: Issues addressed transparently
- **Respectful**: Maintains dignity of all parties
- **Temporary**: Resolved and doesn't fester
#### Unhealthy Conflict Characteristics
- **Personal Attacks**: Targeting individuals rather than ideas
- **Win-Lose Mentality**: Zero-sum thinking
- **Underground**: Gossip and indirect communication
- **Destructive**: Damages relationships and trust
- **Persistent**: Continues without resolution
### Conflict Resolution Process
#### 1. Early Detection
- **Retrospective Themes**: Recurring issues or tensions
- **Ceremony Observation**: Body language, participation patterns
- **1:1 Conversations**: Individual team member concerns
- **Performance Indicators**: Velocity drops, quality issues
#### 2. Assessment & Preparation
- **Stakeholder Mapping**: Who's involved, who's affected
- **Issue Clarification**: Separate facts from interpretations
- **Desired Outcomes**: What would resolution look like?
- **Facilitation Planning**: Process design for resolution session
#### 3. Facilitated Resolution
- **Ground Rules**: Safe space for honest dialogue
- **Perspective Sharing**: Each party states their view
- **Common Ground**: Identify shared interests and values
- **Solution Generation**: Collaborative problem-solving
- **Agreement Creation**: Clear commitments and follow-up
#### 4. Follow-up & Learning
- **Implementation Support**: Help parties honor agreements
- **Relationship Repair**: Ongoing relationship building
- **Process Improvement**: Learn from conflict for future prevention
- **Team Strengthening**: Use resolution as team development opportunity
---
## Assessment Tools
### Team Development Stage Assessment
#### Behavioral Indicators Checklist
**Forming Indicators:**
- [ ] Heavy reliance on Scrum Master for decisions
- [ ] Polite, superficial interactions
- [ ] Individual work preferences
- [ ] Process confusion or resistance
- [ ] Low ceremony engagement
**Storming Indicators:**
- [ ] Open disagreements about approach
- [ ] Questioning of established processes
- [ ] Subgroup formation
- [ ] Inconsistent performance
- [ ] Emotional reactions to feedback
**Norming Indicators:**
- [ ] Collaborative problem-solving
- [ ] Process adaptation and improvement
- [ ] Shared responsibility for outcomes
- [ ] Constructive feedback exchange
- [ ] Stable performance patterns
**Performing Indicators:**
- [ ] Self-organization without external direction
- [ ] Proactive problem anticipation
- [ ] Innovation and experimentation
- [ ] Mentoring of other teams
- [ ] Exceptional results consistently
### Psychological Safety Assessment Survey
#### Team Member Self-Assessment (5-point Likert Scale)
1. **Mistake Tolerance**: "When I make a mistake, my team supports me in learning from it"
2. **Voice Safety**: "I feel comfortable challenging decisions or raising concerns"
3. **Inclusion**: "My unique perspective is valued by the team"
4. **Risk Taking**: "I can take calculated risks without fear of negative consequences"
5. **Help Seeking**: "I can admit when I don't know something without judgment"
6. **Authenticity**: "I can be myself without pretending or hiding parts of my personality"
7. **Innovation**: "We try new approaches even if they might not work"
#### Behavioral Observation Checklist
- **Speaking Up**: Team members voice disagreements respectfully
- **Mistake Response**: Errors are discussed openly for learning
- **Help Seeking**: People admit knowledge gaps and ask for assistance
- **Experimentation**: Team tries new approaches without excessive fear
- **Inclusion**: All members participate actively in discussions
- **Feedback**: Constructive criticism is given and received well
---
## Intervention Strategies
### Stage-Specific Interventions
#### Forming → Storming Transition
- **Trust Building Activities**: Structured sharing and team bonding
- **Psychological Safety Foundation**: Establish ground rules for safe conflict
- **Process Education**: Deep training on collaboration and communication
- **Individual Coaching**: Prepare team members for productive disagreement
#### Storming → Norming Transition
- **Conflict Resolution Skills**: Training in constructive disagreement
- **Working Agreement Updates**: Refine team collaboration standards
- **Success Celebration**: Acknowledge progress through difficult conversations
- **Process Ownership**: Begin transferring facilitation responsibilities
#### Norming → Performing Transition
- **Challenge Introduction**: Stretch goals to push team capabilities
- **Leadership Development**: Grow coaching and mentoring skills
- **Innovation Support**: Create time and space for experimentation
- **External Engagement**: Opportunities to influence other teams
### Crisis Interventions
#### Performance Regression
**Symptoms**: Sudden drops in velocity, quality, or team satisfaction
**Interventions**:
- Team health check and root cause analysis
- Individual 1:1s to understand personal factors
- Process audit to identify systemic issues
- Targeted support for specific capability gaps
#### Psychological Safety Violations
**Symptoms**: Team members withdrawing, avoiding risk, or leaving
**Interventions**:
- Immediate protective actions for affected individuals
- Team-wide discussion of psychological safety principles
- Leadership coaching for those who violated safety
- System changes to prevent future violations
#### External Pressure Impact
**Symptoms**: Team stress, process shortcuts, decreased collaboration
**Interventions**:
- Stakeholder education about sustainable pace
- Scope negotiation and priority clarification
- Team capacity protection and workload management
- Stress management and resilience building
---
## Measurement & Tracking
### Dashboard Metrics by Stage
#### Forming Stage Metrics
- Ceremony attendance rates
- Individual vs. collaborative work ratios
- Process adherence scores
- Initial psychological safety baseline
#### Storming Stage Metrics
- Conflict frequency and resolution time
- Ceremony engagement quality
- Velocity volatility measures
- Team satisfaction surveys
#### Norming Stage Metrics
- Self-organization indicators
- Process improvement frequency
- Knowledge sharing metrics
- Stakeholder satisfaction
#### Performing Stage Metrics
- Innovation and experimentation rates
- External influence and mentoring
- Exceptional result achievement
- Leadership development outcomes
### Tracking Tools & Methods
#### Regular Assessment Schedule
- **Weekly**: Ceremony quality observation
- **Sprint**: Velocity and quality metrics
- **Monthly**: Psychological safety pulse survey
- **Quarterly**: Comprehensive team development assessment
#### Data Collection Methods
- **Quantitative**: Sprint metrics, attendance, survey scores
- **Qualitative**: Observation notes, retrospective themes, interview insights
- **Behavioral**: Video/audio analysis of team interactions (with consent)
- **External**: Stakeholder feedback, other team perceptions
#### Progress Visualization
- **Team Development Radar**: Multi-dimensional progress tracking
- **Psychological Safety Trends**: Safety metrics over time
- **Stage Transition Timeline**: Development milestone tracking
- **Intervention Impact Assessment**: Before/after comparison
---
## Conclusion
Effective team dynamics facilitation requires understanding that team development is a journey, not a destination. Scrum Masters must:
1. **Assess Accurately**: Understand current team development stage
2. **Facilitate Appropriately**: Match leadership style to team needs
3. **Build Safety First**: Psychological safety enables all other development
4. **Measure Progress**: Track both quantitative and qualitative indicators
5. **Intervene Thoughtfully**: Apply stage-appropriate interventions
6. **Celebrate Growth**: Acknowledge progress and learning throughout the journey
The goal is not just high-performing teams, but sustainable high performance built on strong relationships, psychological safety, and continuous learning. This framework provides the structure and tools to guide teams through their development journey effectively.
---
*This framework combines research-based models with practical scrum implementation experience. Adapt the tools and techniques to fit your specific organizational context and team needs.*
FILE:references/velocity-forecasting-guide.md
# Velocity Forecasting Guide: Monte Carlo Methods & Probabilistic Estimation
## Table of Contents
- [Overview](#overview)
- [Monte Carlo Simulation Fundamentals](#monte-carlo-simulation-fundamentals)
- [Velocity-Based Forecasting](#velocity-based-forecasting)
- [Implementation Approaches](#implementation-approaches)
- [Confidence Intervals & Risk Assessment](#confidence-intervals--risk-assessment)
- [Practical Applications](#practical-applications)
- [Advanced Techniques](#advanced-techniques)
- [Common Pitfalls](#common-pitfalls)
- [Case Studies](#case-studies)
---
## Overview
Velocity forecasting using Monte Carlo simulation provides probabilistic estimates for sprint and project completion, moving beyond single-point estimates to give stakeholders a range of likely outcomes with associated confidence levels.
### Why Probabilistic Forecasting?
- **Uncertainty Acknowledgment**: Software development is inherently uncertain
- **Risk Quantification**: Provides probability distributions rather than false precision
- **Stakeholder Communication**: Better expectation management through confidence intervals
- **Decision Support**: Enables data-driven planning and resource allocation
### Core Principles
1. **Historical Velocity Patterns**: Use actual team performance data
2. **Statistical Modeling**: Apply appropriate probability distributions
3. **Confidence Intervals**: Provide ranges, not single points
4. **Continuous Calibration**: Update forecasts with new data
---
## Monte Carlo Simulation Fundamentals
### What is Monte Carlo Simulation?
Monte Carlo simulation uses random sampling to model the probability of different outcomes in systems that cannot be easily predicted due to random variables.
### Application to Velocity Forecasting
```
For each simulation iteration:
1. Sample a velocity value from historical distribution
2. Calculate projected completion time
3. Repeat thousands of times
4. Analyze the distribution of results
```
### Key Statistical Concepts
#### Normal Distribution
Most teams' velocity follows a roughly normal distribution after stabilization:
- **Mean (μ)**: Average historical velocity
- **Standard Deviation (σ)**: Velocity variability measure
- **68-95-99.7 Rule**: Probability ranges for forecasting
#### Distribution Characteristics
- **Symmetry**: Balanced around the mean (normal teams)
- **Skewness**: Teams with frequent disruptions may show positive skew
- **Kurtosis**: Measure of "tail heaviness" - extreme outcomes frequency
---
## Velocity-Based Forecasting
### Basic Velocity Forecasting Formula
**Single Sprint Forecast:**
```
Confidence Interval = μ ± (Z-score × σ)
Where:
- μ = historical mean velocity
- σ = standard deviation of velocity
- Z-score = confidence level multiplier
```
**Multi-Sprint Forecast:**
```
Total Points = Σ(sampled_velocity_i) for i = 1 to n sprints
Where each velocity_i is randomly sampled from historical distribution
```
### Confidence Level Z-Scores
| Confidence Level | Z-Score | Interpretation |
|------------------|---------|----------------|
| 50% | 0.67 | Median outcome |
| 70% | 1.04 | Moderate confidence |
| 85% | 1.44 | High confidence |
| 95% | 1.96 | Very high confidence |
| 99% | 2.58 | Extremely high confidence |
---
## Implementation Approaches
### 1. Simple Historical Distribution Method
```python
def simple_monte_carlo_forecast(velocities, sprints_ahead, iterations=10000):
results = []
for _ in range(iterations):
total_points = sum(random.choice(velocities) for _ in range(sprints_ahead))
results.append(total_points)
return analyze_results(results)
```
**Pros:** Simple, uses actual data points
**Cons:** Ignores trends, assumes stationary distribution
### 2. Normal Distribution Method
```python
def normal_distribution_forecast(velocities, sprints_ahead, iterations=10000):
mean_velocity = statistics.mean(velocities)
std_velocity = statistics.stdev(velocities)
results = []
for _ in range(iterations):
total_points = sum(
max(0, random.normalvariate(mean_velocity, std_velocity))
for _ in range(sprints_ahead)
)
results.append(total_points)
return analyze_results(results)
```
**Pros:** Mathematically clean, handles interpolation
**Cons:** Assumes normal distribution, may generate impossible values
### 3. Bootstrap Sampling Method
```python
def bootstrap_forecast(velocities, sprints_ahead, iterations=10000):
n = len(velocities)
results = []
for _ in range(iterations):
# Sample with replacement
bootstrap_sample = [random.choice(velocities) for _ in range(n)]
# Calculate statistics from bootstrap sample
mean_vel = statistics.mean(bootstrap_sample)
std_vel = statistics.stdev(bootstrap_sample)
total_points = sum(
max(0, random.normalvariate(mean_vel, std_vel))
for _ in range(sprints_ahead)
)
results.append(total_points)
return analyze_results(results)
```
**Pros:** Robust to distribution assumptions, accounts for sampling uncertainty
**Cons:** More complex, requires sufficient historical data
---
## Confidence Intervals & Risk Assessment
### Interpreting Forecast Results
#### Percentile-Based Confidence Intervals
```python
def calculate_confidence_intervals(results, confidence_levels=[0.5, 0.7, 0.85, 0.95]):
sorted_results = sorted(results)
intervals = {}
for confidence in confidence_levels:
percentile_index = int(confidence * len(sorted_results))
intervals[f"{int(confidence*100)}%"] = sorted_results[percentile_index]
return intervals
```
#### Example Interpretation
For a 6-sprint forecast with results:
- **50%:** 120 points (median outcome)
- **70%:** 135 points (likely case)
- **85%:** 150 points (conservative case)
- **95%:** 170 points (very conservative case)
### Risk Assessment Framework
#### Delivery Probability
```
P(Completion ≤ Target) = (# simulations ≤ target) / total_simulations
```
#### Risk Categories
| Probability Range | Risk Level | Recommendation |
|-------------------|------------|----------------|
| > 85% | Low Risk | Proceed with confidence |
| 70-85% | Moderate Risk | Add buffer, monitor closely |
| 50-70% | High Risk | Reduce scope or extend timeline |
| < 50% | Very High Risk | Significant replanning required |
---
## Practical Applications
### Sprint Planning
Use velocity forecasting to:
- Set realistic sprint goals
- Communicate uncertainty to Product Owner
- Plan capacity buffers for unknowns
- Identify when to adjust scope
### Release Planning
Apply Monte Carlo methods to:
- Estimate feature completion dates
- Plan release milestones
- Assess project schedule risk
- Make go/no-go decisions
### Stakeholder Communication
Present forecasts as:
- Range estimates, not single points
- Probability statements ("70% confident we'll deliver X by date Y")
- Risk scenarios with mitigation options
- Visual distributions showing uncertainty
---
## Advanced Techniques
### 1. Trend-Adjusted Forecasting
Account for improving or declining velocity trends:
```python
def trend_adjusted_forecast(velocities, sprints_ahead):
# Calculate linear trend
x = range(len(velocities))
slope, intercept = calculate_linear_regression(x, velocities)
# Adjust future velocities for trend
adjusted_velocities = []
for i in range(sprints_ahead):
future_sprint = len(velocities) + i
predicted_velocity = slope * future_sprint + intercept
adjusted_velocities.append(predicted_velocity)
return monte_carlo_with_adjusted_velocities(adjusted_velocities)
```
### 2. Seasonality Adjustments
For teams with seasonal patterns (holidays, budget cycles):
```python
def seasonal_adjustment(velocities, sprint_dates, forecast_dates):
# Identify seasonal patterns
seasonal_factors = calculate_seasonal_factors(velocities, sprint_dates)
# Apply factors to forecast
adjusted_forecast = apply_seasonal_factors(forecast_dates, seasonal_factors)
return adjusted_forecast
```
### 3. Capacity-Based Modeling
Incorporate team capacity changes:
```python
def capacity_adjusted_forecast(velocities, historical_capacity, future_capacity):
# Calculate velocity per capacity unit
velocity_per_capacity = [v/c for v, c in zip(velocities, historical_capacity)]
baseline_efficiency = statistics.mean(velocity_per_capacity)
# Forecast based on future capacity
future_velocities = [capacity * baseline_efficiency for capacity in future_capacity]
return monte_carlo_forecast(future_velocities)
```
### 4. Multi-Team Forecasting
For dependencies across teams:
```python
def multi_team_forecast(team_forecasts, dependencies):
# Account for critical path and dependencies
# Use min/max operations for dependent deliveries
# Model coordination overhead
pass
```
---
## Common Pitfalls
### 1. Insufficient Historical Data
**Problem:** Using too few sprint data points
**Solution:** Minimum 6-8 sprints for reliable forecasting
**Mitigation:** Use industry benchmarks or similar team data
### 2. Non-Stationary Data
**Problem:** Including data from different team compositions or processes
**Solution:** Use only recent, relevant historical data
**Identification:** Look for structural breaks in velocity time series
### 3. False Precision
**Problem:** Reporting over-precise estimates (e.g., "23.7 points")
**Solution:** Round to reasonable precision, emphasize ranges
**Communication:** Use language like "approximately" and "around"
### 4. Ignoring External Factors
**Problem:** Not accounting for holidays, team changes, external dependencies
**Solution:** Adjust historical data or forecasts for known factors
**Documentation:** Maintain context for each sprint's circumstances
### 5. Overconfidence in Models
**Problem:** Treating forecasts as guarantees
**Solution:** Regular calibration against actual outcomes
**Improvement:** Update models based on forecast accuracy
---
## Case Studies
### Case Study 1: Stabilizing Team
**Situation:** New team, first 10 sprints, velocity ranging 15-25 points
**Approach:**
- Used bootstrap sampling due to small sample size
- Applied 30% buffer for team learning curve
- Updated forecast every 2 sprints
**Results:**
- Initial forecast: 20 ± 8 points per sprint
- Final 3 sprints: 22 ± 3 points per sprint
- Accuracy improved from 60% to 85% confidence bands
### Case Study 2: Seasonal Product Team
**Situation:** E-commerce team with holiday impacts
**Data:** 24 sprints showing clear seasonal patterns
**Approach:**
- Identified seasonal multipliers (0.7x during holidays)
- Used 2-year historical data for seasonal adjustment
- Applied capacity-based modeling for temporary staff
**Results:**
- Standard model: 40% forecast accuracy during Q4
- Seasonal-adjusted model: 80% forecast accuracy
- Better resource planning and stakeholder communication
### Case Study 3: Platform Team with Dependencies
**Situation:** Infrastructure team supporting multiple product teams
**Challenge:** High variability due to urgent requests and dependencies
**Approach:**
- Separated planned vs. unplanned work velocity
- Used wider confidence intervals (90% vs 70%)
- Implemented buffer management strategy
**Results:**
- Planned work predictability: 85%
- Total work predictability: 65% (acceptable for context)
- Improved capacity allocation decisions
---
## Tools and Implementation
### Recommended Tools
1. **Python/R:** For custom implementation and complex models
2. **Excel/Google Sheets:** For simple implementations and visualization
3. **Jira/Azure DevOps:** For automated data collection
4. **Specialized Tools:** ActionableAgile, Monte Carlo simulation software
### Key Metrics to Track
- **Forecast Accuracy:** How often do actual results fall within predicted ranges?
- **Calibration:** Do 70% confidence intervals contain 70% of actual results?
- **Bias:** Are forecasts consistently optimistic or pessimistic?
- **Resolution:** How precise are the forecasts for decision-making?
### Implementation Checklist
- [ ] Historical velocity data collection (minimum 6 sprints)
- [ ] Data quality validation (outliers, context)
- [ ] Distribution analysis (normal, skewed, multi-modal)
- [ ] Model selection and parameter estimation
- [ ] Validation against held-out data
- [ ] Visualization and communication materials
- [ ] Regular calibration and model updates
---
## Conclusion
Monte Carlo velocity forecasting transforms uncertain estimates into probabilistic statements that enable better decision-making. Success requires:
1. **Quality Data:** Clean, relevant historical velocity data
2. **Appropriate Models:** Choose methods suited to your team's patterns
3. **Clear Communication:** Present uncertainty honestly to stakeholders
4. **Continuous Improvement:** Calibrate and refine models over time
5. **Contextual Awareness:** Account for team changes, external factors, and business context
The goal is not perfect prediction, but better understanding of uncertainty to make more informed planning decisions.
---
*This guide provides a comprehensive foundation for implementing probabilistic velocity forecasting. Adapt the techniques to your team's specific context and constraints.*
FILE:scripts/retrospective_analyzer.py
#!/usr/bin/env python3
"""
Retrospective Analyzer
Processes retrospective data to track action item completion rates, identify
recurring themes, measure improvement trends, and generate insights for
continuous team improvement.
Usage:
python retrospective_analyzer.py retro_data.json
python retrospective_analyzer.py retro_data.json --format json
"""
import argparse
import json
import re
import statistics
import sys
from collections import Counter, defaultdict
from datetime import datetime, timedelta
from typing import Any, Dict, List, Optional, Set, Tuple
# ---------------------------------------------------------------------------
# Configuration and Constants
# ---------------------------------------------------------------------------
SENTIMENT_KEYWORDS = {
"positive": [
"good", "great", "excellent", "awesome", "fantastic", "wonderful",
"improved", "better", "success", "achievement", "celebration",
"working well", "effective", "efficient", "smooth", "pleased",
"happy", "satisfied", "proud", "accomplished", "breakthrough"
],
"negative": [
"bad", "terrible", "awful", "horrible", "frustrating", "annoying",
"problem", "issue", "blocker", "impediment", "concern", "worry",
"difficult", "challenging", "struggling", "failing", "broken",
"slow", "delayed", "confused", "unclear", "chaos", "stressed"
],
"neutral": [
"okay", "average", "normal", "standard", "typical", "usual",
"process", "procedure", "meeting", "discussion", "review",
"update", "status", "information", "data", "report"
]
}
THEME_CATEGORIES = {
"communication": [
"communication", "meeting", "standup", "discussion", "feedback",
"information", "clarity", "understanding", "alignment", "sync",
"reporting", "updates", "transparency", "visibility"
],
"process": [
"process", "procedure", "workflow", "methodology", "framework",
"scrum", "agile", "ceremony", "planning", "retrospective",
"review", "estimation", "refinement", "definition of done"
],
"technical": [
"technical", "code", "development", "bug", "testing", "deployment",
"architecture", "infrastructure", "tools", "technology",
"performance", "quality", "automation", "ci/cd", "devops"
],
"team_dynamics": [
"team", "collaboration", "cooperation", "support", "morale",
"motivation", "engagement", "culture", "relationship", "trust",
"conflict", "personality", "workload", "capacity", "burnout"
],
"external": [
"customer", "stakeholder", "management", "product owner", "business",
"requirement", "priority", "deadline", "budget", "resource",
"dependency", "vendor", "third party", "integration"
]
}
ACTION_PRIORITY_KEYWORDS = {
"high": ["urgent", "critical", "asap", "immediately", "blocker", "must"],
"medium": ["important", "should", "needed", "required", "significant"],
"low": ["nice to have", "consider", "explore", "investigate", "eventually"]
}
COMPLETION_STATUS_MAPPING = {
"completed": ["done", "completed", "finished", "resolved", "closed", "achieved"],
"in_progress": ["in progress", "ongoing", "working on", "started", "partial"],
"blocked": ["blocked", "stuck", "waiting", "dependent", "impediment"],
"cancelled": ["cancelled", "dropped", "abandoned", "not needed", "deprioritized"],
"not_started": ["not started", "pending", "todo", "planned", "upcoming"]
}
# ---------------------------------------------------------------------------
# Data Models
# ---------------------------------------------------------------------------
class ActionItem:
"""Represents a single action item from a retrospective."""
def __init__(self, data: Dict[str, Any]):
self.id: str = data.get("id", "")
self.description: str = data.get("description", "")
self.owner: str = data.get("owner", "")
self.priority: str = data.get("priority", "medium").lower()
self.due_date: Optional[str] = data.get("due_date")
self.status: str = data.get("status", "not_started").lower()
self.created_sprint: int = data.get("created_sprint", 0)
self.completed_sprint: Optional[int] = data.get("completed_sprint")
self.category: str = data.get("category", "")
self.effort_estimate: str = data.get("effort_estimate", "medium")
# Normalize status
self.normalized_status = self._normalize_status(self.status)
# Infer priority from description if not explicitly set
if self.priority == "medium":
self.inferred_priority = self._infer_priority(self.description)
else:
self.inferred_priority = self.priority
def _normalize_status(self, status: str) -> str:
"""Normalize status to standard categories."""
status_lower = status.lower().strip()
for category, statuses in COMPLETION_STATUS_MAPPING.items():
if any(s in status_lower for s in statuses):
return category
return "not_started"
def _infer_priority(self, description: str) -> str:
"""Infer priority from description text."""
description_lower = description.lower()
for priority, keywords in ACTION_PRIORITY_KEYWORDS.items():
if any(keyword in description_lower for keyword in keywords):
return priority
return "medium"
@property
def is_completed(self) -> bool:
return self.normalized_status == "completed"
@property
def is_overdue(self) -> bool:
if not self.due_date:
return False
try:
due_date = datetime.strptime(self.due_date, "%Y-%m-%d")
return datetime.now() > due_date and not self.is_completed
except ValueError:
return False
class RetrospectiveData:
"""Represents data from a single retrospective session."""
def __init__(self, data: Dict[str, Any]):
self.sprint_number: int = data.get("sprint_number", 0)
self.date: str = data.get("date", "")
self.facilitator: str = data.get("facilitator", "")
self.attendees: List[str] = data.get("attendees", [])
self.duration_minutes: int = data.get("duration_minutes", 0)
# Retrospective categories
self.went_well: List[str] = data.get("went_well", [])
self.to_improve: List[str] = data.get("to_improve", [])
self.action_items_data: List[Dict[str, Any]] = data.get("action_items", [])
# Create action items
self.action_items: List[ActionItem] = [
ActionItem({**item, "created_sprint": self.sprint_number})
for item in self.action_items_data
]
# Calculate metrics
self._calculate_metrics()
def _calculate_metrics(self):
"""Calculate retrospective session metrics."""
self.total_items = len(self.went_well) + len(self.to_improve)
self.action_items_count = len(self.action_items)
self.attendance_rate = len(self.attendees) / max(1, 5) # Assume team of 5
# Sentiment analysis
self.sentiment_scores = self._analyze_sentiment()
# Theme analysis
self.themes = self._extract_themes()
def _analyze_sentiment(self) -> Dict[str, float]:
"""Analyze sentiment of retrospective items."""
all_text = " ".join(self.went_well + self.to_improve).lower()
sentiment_scores = {}
for sentiment, keywords in SENTIMENT_KEYWORDS.items():
count = sum(1 for keyword in keywords if keyword in all_text)
sentiment_scores[sentiment] = count
# Normalize to percentages
total_sentiment = sum(sentiment_scores.values())
if total_sentiment > 0:
for sentiment in sentiment_scores:
sentiment_scores[sentiment] = sentiment_scores[sentiment] / total_sentiment
return sentiment_scores
def _extract_themes(self) -> Dict[str, int]:
"""Extract themes from retrospective items."""
all_text = " ".join(self.went_well + self.to_improve).lower()
theme_counts = {}
for theme, keywords in THEME_CATEGORIES.items():
count = sum(1 for keyword in keywords if keyword in all_text)
if count > 0:
theme_counts[theme] = count
return theme_counts
class RetroAnalysisResult:
"""Complete retrospective analysis results."""
def __init__(self):
self.summary: Dict[str, Any] = {}
self.action_item_analysis: Dict[str, Any] = {}
self.theme_analysis: Dict[str, Any] = {}
self.improvement_trends: Dict[str, Any] = {}
self.recommendations: List[str] = []
# ---------------------------------------------------------------------------
# Analysis Functions
# ---------------------------------------------------------------------------
def analyze_action_item_completion(retros: List[RetrospectiveData]) -> Dict[str, Any]:
"""Analyze action item completion rates and patterns."""
all_action_items = []
for retro in retros:
all_action_items.extend(retro.action_items)
if not all_action_items:
return {
"total_action_items": 0,
"completion_rate": 0.0,
"average_completion_time": 0.0
}
# Overall completion statistics
completed_items = [item for item in all_action_items if item.is_completed]
completion_rate = len(completed_items) / len(all_action_items)
# Completion time analysis
completion_times = []
for item in completed_items:
if item.completed_sprint and item.created_sprint:
completion_time = item.completed_sprint - item.created_sprint
if completion_time >= 0:
completion_times.append(completion_time)
avg_completion_time = statistics.mean(completion_times) if completion_times else 0.0
# Status distribution
status_counts = Counter(item.normalized_status for item in all_action_items)
# Priority analysis
priority_completion = {}
for priority in ["high", "medium", "low"]:
priority_items = [item for item in all_action_items if item.inferred_priority == priority]
if priority_items:
priority_completed = sum(1 for item in priority_items if item.is_completed)
priority_completion[priority] = {
"total": len(priority_items),
"completed": priority_completed,
"completion_rate": priority_completed / len(priority_items)
}
# Owner analysis
owner_performance = defaultdict(lambda: {"total": 0, "completed": 0})
for item in all_action_items:
if item.owner:
owner_performance[item.owner]["total"] += 1
if item.is_completed:
owner_performance[item.owner]["completed"] += 1
for owner in owner_performance:
owner_data = owner_performance[owner]
owner_data["completion_rate"] = owner_data["completed"] / owner_data["total"]
# Overdue items
overdue_items = [item for item in all_action_items if item.is_overdue]
return {
"total_action_items": len(all_action_items),
"completion_rate": completion_rate,
"completed_items": len(completed_items),
"average_completion_time": avg_completion_time,
"status_distribution": dict(status_counts),
"priority_analysis": priority_completion,
"owner_performance": dict(owner_performance),
"overdue_items": len(overdue_items),
"overdue_rate": len(overdue_items) / len(all_action_items) if all_action_items else 0.0
}
def analyze_recurring_themes(retros: List[RetrospectiveData]) -> Dict[str, Any]:
"""Identify recurring themes across retrospectives."""
theme_evolution = defaultdict(list)
sentiment_evolution = defaultdict(list)
# Track themes over time
for retro in retros:
sprint = retro.sprint_number
# Theme tracking
for theme, count in retro.themes.items():
theme_evolution[theme].append((sprint, count))
# Sentiment tracking
for sentiment, score in retro.sentiment_scores.items():
sentiment_evolution[sentiment].append((sprint, score))
# Identify recurring themes (appear in >50% of retros)
recurring_threshold = len(retros) * 0.5
recurring_themes = {}
for theme, occurrences in theme_evolution.items():
if len(occurrences) >= recurring_threshold:
sprints, counts = zip(*occurrences)
recurring_themes[theme] = {
"frequency": len(occurrences) / len(retros),
"average_mentions": statistics.mean(counts),
"trend": _calculate_trend(list(counts)),
"first_appearance": min(sprints),
"last_appearance": max(sprints),
"total_mentions": sum(counts)
}
# Sentiment trend analysis
sentiment_trends = {}
for sentiment, scores_by_sprint in sentiment_evolution.items():
if len(scores_by_sprint) >= 3: # Need at least 3 data points
_, scores = zip(*scores_by_sprint)
sentiment_trends[sentiment] = {
"average_score": statistics.mean(scores),
"trend": _calculate_trend(list(scores)),
"volatility": statistics.stdev(scores) if len(scores) > 1 else 0.0
}
# Identify persistent issues (negative themes that recur)
persistent_issues = []
for theme, data in recurring_themes.items():
if theme in ["technical", "process", "external"] and data["frequency"] > 0.6:
if data["trend"]["direction"] in ["stable", "increasing"]:
persistent_issues.append({
"theme": theme,
"frequency": data["frequency"],
"severity": data["average_mentions"],
"trend": data["trend"]["direction"]
})
return {
"recurring_themes": recurring_themes,
"sentiment_trends": sentiment_trends,
"persistent_issues": persistent_issues,
"total_themes_identified": len(theme_evolution),
"themes_per_retro": sum(len(r.themes) for r in retros) / len(retros) if retros else 0
}
def analyze_improvement_trends(retros: List[RetrospectiveData]) -> Dict[str, Any]:
"""Analyze improvement trends across retrospectives."""
if len(retros) < 3:
return {"error": "Need at least 3 retrospectives for trend analysis"}
# Sort retrospectives by sprint number
sorted_retros = sorted(retros, key=lambda r: r.sprint_number)
# Track various metrics over time
metrics_over_time = {
"action_items_per_retro": [len(r.action_items) for r in sorted_retros],
"attendance_rate": [r.attendance_rate for r in sorted_retros],
"duration": [r.duration_minutes for r in sorted_retros],
"positive_sentiment": [r.sentiment_scores.get("positive", 0) for r in sorted_retros],
"negative_sentiment": [r.sentiment_scores.get("negative", 0) for r in sorted_retros],
"total_items_discussed": [r.total_items for r in sorted_retros]
}
# Calculate trends for each metric
trend_analysis = {}
for metric_name, values in metrics_over_time.items():
if len(values) >= 3:
trend_analysis[metric_name] = {
"values": values,
"trend": _calculate_trend(values),
"average": statistics.mean(values),
"latest": values[-1],
"change_from_first": ((values[-1] - values[0]) / values[0]) if values[0] != 0 else 0
}
# Action item completion trend
completion_rates_by_sprint = []
for i, retro in enumerate(sorted_retros):
if i > 0: # Skip first retro as it has no previous action items to complete
prev_retro = sorted_retros[i-1]
if prev_retro.action_items:
completed_count = sum(1 for item in prev_retro.action_items
if item.is_completed and item.completed_sprint == retro.sprint_number)
completion_rate = completed_count / len(prev_retro.action_items)
completion_rates_by_sprint.append(completion_rate)
if completion_rates_by_sprint:
trend_analysis["action_item_completion"] = {
"values": completion_rates_by_sprint,
"trend": _calculate_trend(completion_rates_by_sprint),
"average": statistics.mean(completion_rates_by_sprint),
"latest": completion_rates_by_sprint[-1] if completion_rates_by_sprint else 0
}
# Team maturity indicators
maturity_score = _calculate_team_maturity(sorted_retros)
return {
"trend_analysis": trend_analysis,
"team_maturity_score": maturity_score,
"retrospective_quality_trend": _assess_retrospective_quality_trend(sorted_retros),
"improvement_velocity": _calculate_improvement_velocity(sorted_retros)
}
def _calculate_trend(values: List[float]) -> Dict[str, Any]:
"""Calculate trend direction and strength for a series of values."""
if len(values) < 2:
return {"direction": "insufficient_data", "strength": 0.0}
# Simple linear regression
n = len(values)
x_values = list(range(n))
x_mean = sum(x_values) / n
y_mean = sum(values) / n
numerator = sum((x - x_mean) * (y - y_mean) for x, y in zip(x_values, values))
denominator = sum((x - x_mean) ** 2 for x in x_values)
if denominator == 0:
slope = 0
else:
slope = numerator / denominator
# Calculate correlation coefficient for trend strength
try:
correlation = statistics.correlation(x_values, values) if n > 2 else 0.0
except statistics.StatisticsError:
correlation = 0.0
# Determine trend direction
if abs(slope) < 0.01: # Practically no change
direction = "stable"
elif slope > 0:
direction = "increasing"
else:
direction = "decreasing"
return {
"direction": direction,
"slope": slope,
"strength": abs(correlation),
"correlation": correlation
}
def _calculate_team_maturity(retros: List[RetrospectiveData]) -> Dict[str, Any]:
"""Calculate team maturity based on retrospective patterns."""
if len(retros) < 3:
return {"score": 50, "level": "developing"}
maturity_indicators = {
"action_item_focus": 0, # Fewer but higher quality action items
"sentiment_balance": 0, # Balanced positive/negative sentiment
"theme_consistency": 0, # Consistent themes without chaos
"participation": 0, # High attendance rates
"follow_through": 0 # Good action item completion
}
# Action item focus (quality over quantity)
avg_action_items = sum(len(r.action_items) for r in retros) / len(retros)
if 2 <= avg_action_items <= 5: # Sweet spot
maturity_indicators["action_item_focus"] = 100
elif avg_action_items < 2 or avg_action_items > 8:
maturity_indicators["action_item_focus"] = 30
else:
maturity_indicators["action_item_focus"] = 70
# Sentiment balance
avg_positive = sum(r.sentiment_scores.get("positive", 0) for r in retros) / len(retros)
avg_negative = sum(r.sentiment_scores.get("negative", 0) for r in retros) / len(retros)
if 0.3 <= avg_positive <= 0.6 and 0.2 <= avg_negative <= 0.4:
maturity_indicators["sentiment_balance"] = 100
else:
maturity_indicators["sentiment_balance"] = 50
# Participation
avg_attendance = sum(r.attendance_rate for r in retros) / len(retros)
maturity_indicators["participation"] = min(100, avg_attendance * 100)
# Theme consistency (not too chaotic, not too narrow)
avg_themes = sum(len(r.themes) for r in retros) / len(retros)
if 2 <= avg_themes <= 4:
maturity_indicators["theme_consistency"] = 100
else:
maturity_indicators["theme_consistency"] = 70
# Follow-through (estimated from action item patterns)
# This is simplified - in reality would track actual completion
recent_retros = retros[-3:] if len(retros) >= 3 else retros
avg_recent_actions = sum(len(r.action_items) for r in recent_retros) / len(recent_retros)
if avg_recent_actions <= 3: # Fewer action items might indicate better follow-through
maturity_indicators["follow_through"] = 80
else:
maturity_indicators["follow_through"] = 60
# Calculate overall maturity score
overall_score = sum(maturity_indicators.values()) / len(maturity_indicators)
if overall_score >= 85:
level = "high_performing"
elif overall_score >= 70:
level = "performing"
elif overall_score >= 55:
level = "developing"
else:
level = "forming"
return {
"score": overall_score,
"level": level,
"indicators": maturity_indicators
}
def _assess_retrospective_quality_trend(retros: List[RetrospectiveData]) -> Dict[str, Any]:
"""Assess the quality trend of retrospectives over time."""
quality_scores = []
for retro in retros:
score = 0
# Duration appropriateness (60-90 minutes is ideal)
if 60 <= retro.duration_minutes <= 90:
score += 25
elif 45 <= retro.duration_minutes <= 120:
score += 15
else:
score += 5
# Participation
score += min(25, retro.attendance_rate * 25)
# Balance of content
went_well_count = len(retro.went_well)
to_improve_count = len(retro.to_improve)
total_items = went_well_count + to_improve_count
if total_items > 0:
balance = min(went_well_count, to_improve_count) / total_items
score += balance * 25
# Action items quality (not too many, not too few)
action_count = len(retro.action_items)
if 2 <= action_count <= 5:
score += 25
elif 1 <= action_count <= 7:
score += 15
else:
score += 5
quality_scores.append(score)
if len(quality_scores) >= 2:
trend = _calculate_trend(quality_scores)
else:
trend = {"direction": "insufficient_data", "strength": 0.0}
return {
"quality_scores": quality_scores,
"average_quality": statistics.mean(quality_scores),
"trend": trend,
"latest_quality": quality_scores[-1] if quality_scores else 0
}
def _calculate_improvement_velocity(retros: List[RetrospectiveData]) -> Dict[str, Any]:
"""Calculate how quickly the team improves based on retrospective patterns."""
if len(retros) < 4:
return {"velocity": "insufficient_data"}
# Look at theme evolution - are persistent issues being resolved?
theme_counts = defaultdict(list)
for retro in retros:
for theme, count in retro.themes.items():
theme_counts[theme].append(count)
resolved_themes = 0
persistent_themes = 0
for theme, counts in theme_counts.items():
if len(counts) >= 3:
recent_avg = statistics.mean(counts[-2:])
early_avg = statistics.mean(counts[:2])
if recent_avg < early_avg * 0.7: # 30% reduction
resolved_themes += 1
elif recent_avg > early_avg * 0.9: # Still persistent
persistent_themes += 1
total_themes = resolved_themes + persistent_themes
if total_themes > 0:
resolution_rate = resolved_themes / total_themes
else:
resolution_rate = 0.5 # Neutral if no data
# Action item completion trends
if len(retros) >= 4:
recent_action_density = sum(len(r.action_items) for r in retros[-2:]) / 2
early_action_density = sum(len(r.action_items) for r in retros[:2]) / 2
action_efficiency = 1.0
if early_action_density > 0:
action_efficiency = min(1.0, early_action_density / max(recent_action_density, 1))
else:
action_efficiency = 0.5
# Overall velocity score
velocity_score = (resolution_rate * 0.6) + (action_efficiency * 0.4)
if velocity_score >= 0.8:
velocity = "high"
elif velocity_score >= 0.6:
velocity = "moderate"
elif velocity_score >= 0.4:
velocity = "low"
else:
velocity = "stagnant"
return {
"velocity": velocity,
"velocity_score": velocity_score,
"theme_resolution_rate": resolution_rate,
"action_efficiency": action_efficiency,
"resolved_themes": resolved_themes,
"persistent_themes": persistent_themes
}
def generate_recommendations(result: RetroAnalysisResult) -> List[str]:
"""Generate actionable recommendations based on retrospective analysis."""
recommendations = []
# Action item recommendations
action_analysis = result.action_item_analysis
completion_rate = action_analysis.get("completion_rate", 0)
if completion_rate < 0.5:
recommendations.append("CRITICAL: Low action item completion rate (<50%). Reduce action items per retro and focus on realistic, achievable goals.")
elif completion_rate < 0.7:
recommendations.append("Improve action item follow-through. Consider assigning owners and due dates more systematically.")
elif completion_rate > 0.9:
recommendations.append("Excellent action item completion! Consider taking on more ambitious improvement initiatives.")
overdue_rate = action_analysis.get("overdue_rate", 0)
if overdue_rate > 0.3:
recommendations.append("High overdue rate suggests unrealistic timelines. Review estimation and prioritization process.")
# Theme recommendations
theme_analysis = result.theme_analysis
persistent_issues = theme_analysis.get("persistent_issues", [])
if len(persistent_issues) >= 2:
recommendations.append(f"Address {len(persistent_issues)} persistent issues that keep recurring across retrospectives.")
for issue in persistent_issues[:2]: # Top 2 issues
recommendations.append(f"Focus on resolving recurring {issue['theme']} issues (appears in {issue['frequency']:.0%} of retros).")
# Trend-based recommendations
improvement_trends = result.improvement_trends
if "team_maturity_score" in improvement_trends:
maturity = improvement_trends["team_maturity_score"]
level = maturity.get("level", "forming")
if level == "forming":
recommendations.append("Team is in forming stage. Focus on establishing basic retrospective disciplines and psychological safety.")
elif level == "developing":
recommendations.append("Team is developing. Work on action item follow-through and deeper root cause analysis.")
elif level == "performing":
recommendations.append("Good team maturity. Consider advanced techniques like continuous improvement tracking.")
elif level == "high_performing":
recommendations.append("Excellent retrospective maturity! Share practices with other teams and focus on innovation.")
# Quality recommendations
if "retrospective_quality_trend" in improvement_trends:
quality_trend = improvement_trends["retrospective_quality_trend"]
avg_quality = quality_trend.get("average_quality", 50)
if avg_quality < 60:
recommendations.append("Retrospective quality is below average. Review facilitation techniques and engagement strategies.")
trend_direction = quality_trend.get("trend", {}).get("direction", "stable")
if trend_direction == "decreasing":
recommendations.append("Retrospective quality is declining. Consider changing facilitation approach or addressing team engagement issues.")
return recommendations
# ---------------------------------------------------------------------------
# Main Analysis Function
# ---------------------------------------------------------------------------
def analyze_retrospectives(data: Dict[str, Any]) -> RetroAnalysisResult:
"""Perform comprehensive retrospective analysis."""
result = RetroAnalysisResult()
try:
# Parse retrospective data
retro_records = data.get("retrospectives", [])
retros = [RetrospectiveData(record) for record in retro_records]
if not retros:
raise ValueError("No retrospective data found")
# Sort by sprint number
retros.sort(key=lambda r: r.sprint_number)
# Basic summary
result.summary = {
"total_retrospectives": len(retros),
"date_range": {
"first": retros[0].date if retros else "",
"last": retros[-1].date if retros else "",
"span_sprints": retros[-1].sprint_number - retros[0].sprint_number + 1 if retros else 0
},
"average_duration": statistics.mean([r.duration_minutes for r in retros if r.duration_minutes > 0]),
"average_attendance": statistics.mean([r.attendance_rate for r in retros]),
}
# Action item analysis
result.action_item_analysis = analyze_action_item_completion(retros)
# Theme analysis
result.theme_analysis = analyze_recurring_themes(retros)
# Improvement trends
result.improvement_trends = analyze_improvement_trends(retros)
# Generate recommendations
result.recommendations = generate_recommendations(result)
except Exception as e:
result.summary = {"error": str(e)}
return result
# ---------------------------------------------------------------------------
# Output Formatting
# ---------------------------------------------------------------------------
def format_text_output(result: RetroAnalysisResult) -> str:
"""Format analysis results as readable text report."""
lines = []
lines.append("="*60)
lines.append("RETROSPECTIVE ANALYSIS REPORT")
lines.append("="*60)
lines.append("")
if "error" in result.summary:
lines.append(f"ERROR: {result.summary['error']}")
return "\n".join(lines)
# Summary section
summary = result.summary
lines.append("RETROSPECTIVE SUMMARY")
lines.append("-"*30)
lines.append(f"Total Retrospectives: {summary['total_retrospectives']}")
lines.append(f"Sprint Range: {summary['date_range']['span_sprints']} sprints")
lines.append(f"Average Duration: {summary.get('average_duration', 0):.0f} minutes")
lines.append(f"Average Attendance: {summary.get('average_attendance', 0):.1%}")
lines.append("")
# Action item analysis
action_analysis = result.action_item_analysis
lines.append("ACTION ITEM ANALYSIS")
lines.append("-"*30)
lines.append(f"Total Action Items: {action_analysis.get('total_action_items', 0)}")
lines.append(f"Completion Rate: {action_analysis.get('completion_rate', 0):.1%}")
lines.append(f"Average Completion Time: {action_analysis.get('average_completion_time', 0):.1f} sprints")
lines.append(f"Overdue Items: {action_analysis.get('overdue_items', 0)} ({action_analysis.get('overdue_rate', 0):.1%})")
priority_analysis = action_analysis.get('priority_analysis', {})
if priority_analysis:
lines.append("Priority-based completion rates:")
for priority, data in priority_analysis.items():
lines.append(f" {priority.title()}: {data['completion_rate']:.1%} ({data['completed']}/{data['total']})")
lines.append("")
# Theme analysis
theme_analysis = result.theme_analysis
lines.append("THEME ANALYSIS")
lines.append("-"*30)
recurring_themes = theme_analysis.get("recurring_themes", {})
if recurring_themes:
lines.append("Top recurring themes:")
sorted_themes = sorted(recurring_themes.items(), key=lambda x: x[1]['frequency'], reverse=True)
for theme, data in sorted_themes[:5]:
lines.append(f" {theme.replace('_', ' ').title()}: {data['frequency']:.1%} frequency, {data['trend']['direction']} trend")
persistent_issues = theme_analysis.get("persistent_issues", [])
if persistent_issues:
lines.append("Persistent issues requiring attention:")
for issue in persistent_issues:
lines.append(f" {issue['theme'].replace('_', ' ').title()}: {issue['frequency']:.1%} frequency")
lines.append("")
# Improvement trends
improvement_trends = result.improvement_trends
if "team_maturity_score" in improvement_trends:
maturity = improvement_trends["team_maturity_score"]
lines.append("TEAM MATURITY")
lines.append("-"*30)
lines.append(f"Maturity Level: {maturity['level'].replace('_', ' ').title()}")
lines.append(f"Maturity Score: {maturity['score']:.0f}/100")
lines.append("")
if "improvement_velocity" in improvement_trends:
velocity = improvement_trends["improvement_velocity"]
lines.append("IMPROVEMENT VELOCITY")
lines.append("-"*30)
lines.append(f"Velocity: {velocity['velocity'].title()}")
lines.append(f"Theme Resolution Rate: {velocity.get('theme_resolution_rate', 0):.1%}")
lines.append("")
# Recommendations
if result.recommendations:
lines.append("RECOMMENDATIONS")
lines.append("-"*30)
for i, rec in enumerate(result.recommendations, 1):
lines.append(f"{i}. {rec}")
return "\n".join(lines)
def format_json_output(result: RetroAnalysisResult) -> Dict[str, Any]:
"""Format analysis results as JSON."""
return {
"summary": result.summary,
"action_item_analysis": result.action_item_analysis,
"theme_analysis": result.theme_analysis,
"improvement_trends": result.improvement_trends,
"recommendations": result.recommendations,
}
# ---------------------------------------------------------------------------
# CLI Interface
# ---------------------------------------------------------------------------
def main() -> int:
"""Main CLI entry point."""
parser = argparse.ArgumentParser(
description="Analyze retrospective data for continuous improvement insights"
)
parser.add_argument(
"data_file",
help="JSON file containing retrospective data"
)
parser.add_argument(
"--format",
choices=["text", "json"],
default="text",
help="Output format (default: text)"
)
args = parser.parse_args()
try:
# Load and validate data
with open(args.data_file, 'r') as f:
data = json.load(f)
# Perform analysis
result = analyze_retrospectives(data)
# Output results
if args.format == "json":
output = format_json_output(result)
print(json.dumps(output, indent=2))
else:
output = format_text_output(result)
print(output)
return 0
except FileNotFoundError:
print(f"Error: File '{args.data_file}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in '{args.data_file}': {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/sprint_health_scorer.py
#!/usr/bin/env python3
"""
Sprint Health Scorer
Scores sprint health across multiple dimensions including commitment reliability,
scope creep, blocker resolution time, ceremony attendance, and story completion
distribution. Produces composite health scores with actionable recommendations.
Usage:
python sprint_health_scorer.py sprint_data.json
python sprint_health_scorer.py sprint_data.json --format json
"""
import argparse
import json
import statistics
import sys
from datetime import datetime, timedelta
from typing import Any, Dict, List, Optional, Tuple
# ---------------------------------------------------------------------------
# Scoring Configuration
# ---------------------------------------------------------------------------
HEALTH_DIMENSIONS = {
"commitment_reliability": {
"weight": 0.25,
"excellent_threshold": 0.95, # 95%+ commitment achievement
"good_threshold": 0.85, # 85%+ commitment achievement
"poor_threshold": 0.70, # Below 70% is poor
},
"scope_stability": {
"weight": 0.20,
"excellent_threshold": 0.05, # ≤5% scope change
"good_threshold": 0.15, # ≤15% scope change
"poor_threshold": 0.30, # >30% scope change is poor
},
"blocker_resolution": {
"weight": 0.15,
"excellent_threshold": 1.0, # ≤1 day average resolution
"good_threshold": 3.0, # ≤3 days average resolution
"poor_threshold": 7.0, # >7 days is poor
},
"ceremony_engagement": {
"weight": 0.15,
"excellent_threshold": 0.95, # 95%+ attendance
"good_threshold": 0.85, # 85%+ attendance
"poor_threshold": 0.70, # Below 70% is poor
},
"story_completion_distribution": {
"weight": 0.15,
"excellent_threshold": 0.80, # 80%+ stories fully completed
"good_threshold": 0.65, # 65%+ stories completed
"poor_threshold": 0.50, # Below 50% is poor
},
"velocity_predictability": {
"weight": 0.10,
"excellent_threshold": 0.10, # ≤10% CV
"good_threshold": 0.20, # ≤20% CV
"poor_threshold": 0.35, # >35% CV is poor
}
}
OVERALL_HEALTH_THRESHOLDS = {
"excellent": 85,
"good": 70,
"fair": 55,
"poor": 40,
}
STORY_STATUS_MAPPING = {
"completed": ["done", "completed", "closed", "resolved"],
"in_progress": ["in progress", "in_progress", "development", "testing"],
"blocked": ["blocked", "impediment", "waiting"],
"not_started": ["todo", "to do", "backlog", "new", "open"],
}
# ---------------------------------------------------------------------------
# Data Models
# ---------------------------------------------------------------------------
class Story:
"""Represents a user story within a sprint."""
def __init__(self, data: Dict[str, Any]):
self.id: str = data.get("id", "")
self.title: str = data.get("title", "")
self.points: int = data.get("points", 0)
self.status: str = data.get("status", "").lower()
self.assigned_to: str = data.get("assigned_to", "")
self.created_date: str = data.get("created_date", "")
self.completed_date: Optional[str] = data.get("completed_date")
self.blocked_days: int = data.get("blocked_days", 0)
self.priority: str = data.get("priority", "medium")
# Normalize status
self.normalized_status = self._normalize_status(self.status)
def _normalize_status(self, status: str) -> str:
"""Normalize status to standard categories."""
status_lower = status.lower().strip()
for category, statuses in STORY_STATUS_MAPPING.items():
if status_lower in statuses:
return category
return "unknown"
@property
def is_completed(self) -> bool:
return self.normalized_status == "completed"
@property
def is_blocked(self) -> bool:
return self.normalized_status == "blocked" or self.blocked_days > 0
class SprintHealthData:
"""Comprehensive sprint health data model."""
def __init__(self, data: Dict[str, Any]):
self.sprint_number: int = data.get("sprint_number", 0)
self.sprint_name: str = data.get("sprint_name", "")
self.start_date: str = data.get("start_date", "")
self.end_date: str = data.get("end_date", "")
self.team_size: int = data.get("team_size", 0)
self.working_days: int = data.get("working_days", 10)
# Commitment and delivery
self.planned_points: int = data.get("planned_points", 0)
self.completed_points: int = data.get("completed_points", 0)
self.added_points: int = data.get("added_points", 0)
self.removed_points: int = data.get("removed_points", 0)
# Stories
story_data = data.get("stories", [])
self.stories: List[Story] = [Story(story) for story in story_data]
# Blockers
self.blockers: List[Dict[str, Any]] = data.get("blockers", [])
# Ceremonies
self.ceremonies: Dict[str, Any] = data.get("ceremonies", {})
# Calculate derived metrics
self._calculate_derived_metrics()
def _calculate_derived_metrics(self):
"""Calculate derived health metrics."""
# Commitment reliability
self.commitment_ratio = (
self.completed_points / max(self.planned_points, 1)
)
# Scope change
total_scope_change = self.added_points + self.removed_points
self.scope_change_ratio = total_scope_change / max(self.planned_points, 1)
# Story completion distribution
total_stories = len(self.stories)
if total_stories > 0:
completed_stories = sum(1 for story in self.stories if story.is_completed)
self.story_completion_ratio = completed_stories / total_stories
else:
self.story_completion_ratio = 0.0
# Blocked stories analysis
blocked_stories = [story for story in self.stories if story.is_blocked]
self.blocked_stories_count = len(blocked_stories)
self.blocked_points = sum(story.points for story in blocked_stories)
class HealthScoreResult:
"""Complete health scoring results."""
def __init__(self):
self.dimension_scores: Dict[str, Dict[str, Any]] = {}
self.overall_score: float = 0.0
self.health_grade: str = ""
self.trend_analysis: Dict[str, Any] = {}
self.recommendations: List[str] = []
self.detailed_metrics: Dict[str, Any] = {}
# ---------------------------------------------------------------------------
# Scoring Functions
# ---------------------------------------------------------------------------
def score_commitment_reliability(sprints: List[SprintHealthData]) -> Dict[str, Any]:
"""Score commitment reliability across sprints."""
if not sprints:
return {"score": 0, "grade": "insufficient_data"}
commitment_ratios = [sprint.commitment_ratio for sprint in sprints]
avg_commitment = statistics.mean(commitment_ratios)
consistency = 1.0 - (statistics.stdev(commitment_ratios) if len(commitment_ratios) > 1 else 0)
# Score based on average achievement and consistency
config = HEALTH_DIMENSIONS["commitment_reliability"]
base_score = _calculate_dimension_score(avg_commitment, config)
# Penalty for inconsistency
consistency_bonus = min(10, consistency * 10)
final_score = min(100, base_score + consistency_bonus)
return {
"score": final_score,
"grade": _score_to_grade(final_score),
"average_commitment": avg_commitment,
"consistency": consistency,
"commitment_ratios": commitment_ratios,
"details": f"Average commitment: {avg_commitment:.1%}, Consistency: {consistency:.1%}"
}
def score_scope_stability(sprints: List[SprintHealthData]) -> Dict[str, Any]:
"""Score scope stability (low scope change is better)."""
if not sprints:
return {"score": 0, "grade": "insufficient_data"}
scope_change_ratios = [sprint.scope_change_ratio for sprint in sprints]
avg_scope_change = statistics.mean(scope_change_ratios)
# For scope change, lower is better, so invert the scoring
config = HEALTH_DIMENSIONS["scope_stability"]
if avg_scope_change <= config["excellent_threshold"]:
score = 90 + (config["excellent_threshold"] - avg_scope_change) * 200
elif avg_scope_change <= config["good_threshold"]:
score = 70 + (config["good_threshold"] - avg_scope_change) * 200
elif avg_scope_change <= config["poor_threshold"]:
score = 40 + (config["poor_threshold"] - avg_scope_change) * 200
else:
score = max(0, 40 - (avg_scope_change - config["poor_threshold"]) * 100)
score = min(100, max(0, score))
return {
"score": score,
"grade": _score_to_grade(score),
"average_scope_change": avg_scope_change,
"scope_change_ratios": scope_change_ratios,
"details": f"Average scope change: {avg_scope_change:.1%}"
}
def score_blocker_resolution(sprints: List[SprintHealthData]) -> Dict[str, Any]:
"""Score blocker resolution efficiency."""
if not sprints:
return {"score": 0, "grade": "insufficient_data"}
all_blockers = []
for sprint in sprints:
all_blockers.extend(sprint.blockers)
if not all_blockers:
return {
"score": 100,
"grade": "excellent",
"average_resolution_time": 0,
"details": "No blockers reported"
}
# Calculate average resolution time
resolution_times = []
for blocker in all_blockers:
resolution_time = blocker.get("resolution_days", 0)
if resolution_time > 0:
resolution_times.append(resolution_time)
if not resolution_times:
return {"score": 50, "grade": "fair", "details": "No resolution time data"}
avg_resolution_time = statistics.mean(resolution_times)
# Score based on resolution time (lower is better)
config = HEALTH_DIMENSIONS["blocker_resolution"]
if avg_resolution_time <= config["excellent_threshold"]:
score = 95
elif avg_resolution_time <= config["good_threshold"]:
score = 80 - (avg_resolution_time - config["excellent_threshold"]) * 10
elif avg_resolution_time <= config["poor_threshold"]:
score = 60 - (avg_resolution_time - config["good_threshold"]) * 5
else:
score = max(20, 40 - (avg_resolution_time - config["poor_threshold"]) * 3)
return {
"score": score,
"grade": _score_to_grade(score),
"average_resolution_time": avg_resolution_time,
"total_blockers": len(all_blockers),
"resolved_blockers": len(resolution_times),
"details": f"Average resolution: {avg_resolution_time:.1f} days from {len(all_blockers)} blockers"
}
def score_ceremony_engagement(sprints: List[SprintHealthData]) -> Dict[str, Any]:
"""Score team engagement in scrum ceremonies."""
if not sprints:
return {"score": 0, "grade": "insufficient_data"}
ceremony_scores = []
ceremony_details = {}
for sprint in sprints:
ceremonies = sprint.ceremonies
sprint_ceremony_scores = []
for ceremony_name, ceremony_data in ceremonies.items():
if isinstance(ceremony_data, dict):
attendance_rate = ceremony_data.get("attendance_rate", 0)
engagement_score = ceremony_data.get("engagement_score", 0)
# Weight attendance more heavily than engagement
ceremony_score = (attendance_rate * 0.7) + (engagement_score * 0.3)
sprint_ceremony_scores.append(ceremony_score)
if ceremony_name not in ceremony_details:
ceremony_details[ceremony_name] = []
ceremony_details[ceremony_name].append({
"sprint": sprint.sprint_number,
"attendance": attendance_rate,
"engagement": engagement_score,
"score": ceremony_score
})
if sprint_ceremony_scores:
ceremony_scores.append(statistics.mean(sprint_ceremony_scores))
if not ceremony_scores:
return {"score": 50, "grade": "fair", "details": "No ceremony data available"}
avg_ceremony_score = statistics.mean(ceremony_scores)
config = HEALTH_DIMENSIONS["ceremony_engagement"]
score = _calculate_dimension_score(avg_ceremony_score, config)
return {
"score": score,
"grade": _score_to_grade(score),
"average_ceremony_score": avg_ceremony_score,
"ceremony_details": ceremony_details,
"details": f"Average ceremony engagement: {avg_ceremony_score:.1%}"
}
def score_story_completion_distribution(sprints: List[SprintHealthData]) -> Dict[str, Any]:
"""Score how well stories are completed vs. partially done."""
if not sprints:
return {"score": 0, "grade": "insufficient_data"}
completion_ratios = []
story_analysis = {
"total_stories": 0,
"completed_stories": 0,
"blocked_stories": 0,
"partial_completion": 0
}
for sprint in sprints:
if sprint.stories:
sprint_completion = sprint.story_completion_ratio
completion_ratios.append(sprint_completion)
story_analysis["total_stories"] += len(sprint.stories)
story_analysis["completed_stories"] += sum(1 for s in sprint.stories if s.is_completed)
story_analysis["blocked_stories"] += sum(1 for s in sprint.stories if s.is_blocked)
if not completion_ratios:
return {"score": 50, "grade": "fair", "details": "No story data available"}
avg_completion_ratio = statistics.mean(completion_ratios)
config = HEALTH_DIMENSIONS["story_completion_distribution"]
score = _calculate_dimension_score(avg_completion_ratio, config)
# Penalty for high number of blocked stories
if story_analysis["total_stories"] > 0:
blocked_ratio = story_analysis["blocked_stories"] / story_analysis["total_stories"]
if blocked_ratio > 0.20: # More than 20% blocked
score = max(0, score - (blocked_ratio - 0.20) * 100)
return {
"score": score,
"grade": _score_to_grade(score),
"average_completion_ratio": avg_completion_ratio,
"story_analysis": story_analysis,
"details": f"Average story completion: {avg_completion_ratio:.1%}"
}
def score_velocity_predictability(sprints: List[SprintHealthData]) -> Dict[str, Any]:
"""Score velocity predictability based on coefficient of variation."""
if len(sprints) < 2:
return {"score": 50, "grade": "fair", "details": "Insufficient sprints for predictability analysis"}
velocities = [sprint.completed_points for sprint in sprints]
mean_velocity = statistics.mean(velocities)
if mean_velocity == 0:
return {"score": 0, "grade": "poor", "details": "No velocity recorded"}
velocity_cv = statistics.stdev(velocities) / mean_velocity
# Lower CV is better for predictability
config = HEALTH_DIMENSIONS["velocity_predictability"]
if velocity_cv <= config["excellent_threshold"]:
score = 95
elif velocity_cv <= config["good_threshold"]:
score = 80 - (velocity_cv - config["excellent_threshold"]) * 150
elif velocity_cv <= config["poor_threshold"]:
score = 60 - (velocity_cv - config["good_threshold"]) * 100
else:
score = max(20, 40 - (velocity_cv - config["poor_threshold"]) * 50)
return {
"score": score,
"grade": _score_to_grade(score),
"coefficient_of_variation": velocity_cv,
"mean_velocity": mean_velocity,
"velocity_std_dev": statistics.stdev(velocities),
"details": f"Velocity CV: {velocity_cv:.1%} (lower is more predictable)"
}
def _calculate_dimension_score(value: float, config: Dict[str, Any]) -> float:
"""Calculate dimension score based on thresholds."""
if value >= config["excellent_threshold"]:
return 95
elif value >= config["good_threshold"]:
# Linear interpolation between good and excellent
range_size = config["excellent_threshold"] - config["good_threshold"]
position = (value - config["good_threshold"]) / range_size
return 80 + (position * 15)
elif value >= config["poor_threshold"]:
# Linear interpolation between poor and good
range_size = config["good_threshold"] - config["poor_threshold"]
position = (value - config["poor_threshold"]) / range_size
return 50 + (position * 30)
else:
# Below poor threshold
return max(20, 50 - (config["poor_threshold"] - value) * 100)
def _score_to_grade(score: float) -> str:
"""Convert numerical score to letter grade."""
if score >= OVERALL_HEALTH_THRESHOLDS["excellent"]:
return "excellent"
elif score >= OVERALL_HEALTH_THRESHOLDS["good"]:
return "good"
elif score >= OVERALL_HEALTH_THRESHOLDS["fair"]:
return "fair"
else:
return "poor"
# ---------------------------------------------------------------------------
# Main Analysis Function
# ---------------------------------------------------------------------------
def analyze_sprint_health(data: Dict[str, Any]) -> HealthScoreResult:
"""Perform comprehensive sprint health analysis."""
result = HealthScoreResult()
try:
# Parse sprint data
sprint_records = data.get("sprints", [])
sprints = [SprintHealthData(record) for record in sprint_records]
if not sprints:
raise ValueError("No sprint data found")
# Sort by sprint number
sprints.sort(key=lambda s: s.sprint_number)
# Calculate dimension scores
dimensions = {
"commitment_reliability": score_commitment_reliability,
"scope_stability": score_scope_stability,
"blocker_resolution": score_blocker_resolution,
"ceremony_engagement": score_ceremony_engagement,
"story_completion_distribution": score_story_completion_distribution,
"velocity_predictability": score_velocity_predictability,
}
weighted_scores = []
for dimension_name, scoring_func in dimensions.items():
dimension_result = scoring_func(sprints)
result.dimension_scores[dimension_name] = dimension_result
# Calculate weighted contribution
weight = HEALTH_DIMENSIONS[dimension_name]["weight"]
weighted_score = dimension_result["score"] * weight
weighted_scores.append(weighted_score)
# Calculate overall score
result.overall_score = sum(weighted_scores)
result.health_grade = _score_to_grade(result.overall_score)
# Generate detailed metrics
result.detailed_metrics = _generate_detailed_metrics(sprints)
# Generate recommendations
result.recommendations = _generate_health_recommendations(result)
except Exception as e:
result.dimension_scores = {"error": str(e)}
result.overall_score = 0
return result
def _generate_detailed_metrics(sprints: List[SprintHealthData]) -> Dict[str, Any]:
"""Generate detailed metrics for analysis."""
metrics = {
"sprint_count": len(sprints),
"date_range": {
"start": sprints[0].start_date if sprints else "",
"end": sprints[-1].end_date if sprints else "",
},
"team_metrics": {},
"story_metrics": {},
"blocker_metrics": {},
}
if not sprints:
return metrics
# Team metrics
team_sizes = [sprint.team_size for sprint in sprints if sprint.team_size > 0]
if team_sizes:
metrics["team_metrics"] = {
"average_team_size": statistics.mean(team_sizes),
"team_size_stability": statistics.stdev(team_sizes) if len(team_sizes) > 1 else 0,
}
# Story metrics
all_stories = []
for sprint in sprints:
all_stories.extend(sprint.stories)
if all_stories:
story_points = [story.points for story in all_stories if story.points > 0]
metrics["story_metrics"] = {
"total_stories": len(all_stories),
"average_story_points": statistics.mean(story_points) if story_points else 0,
"completed_stories": sum(1 for story in all_stories if story.is_completed),
"blocked_stories": sum(1 for story in all_stories if story.is_blocked),
}
# Blocker metrics
all_blockers = []
for sprint in sprints:
all_blockers.extend(sprint.blockers)
if all_blockers:
resolution_times = [b.get("resolution_days", 0) for b in all_blockers if b.get("resolution_days", 0) > 0]
metrics["blocker_metrics"] = {
"total_blockers": len(all_blockers),
"resolved_blockers": len(resolution_times),
"average_resolution_days": statistics.mean(resolution_times) if resolution_times else 0,
}
return metrics
def _generate_health_recommendations(result: HealthScoreResult) -> List[str]:
"""Generate actionable recommendations based on health scores."""
recommendations = []
# Overall health recommendations
if result.overall_score < OVERALL_HEALTH_THRESHOLDS["poor"]:
recommendations.append("CRITICAL: Sprint health is poor across multiple dimensions. Immediate intervention required.")
elif result.overall_score < OVERALL_HEALTH_THRESHOLDS["fair"]:
recommendations.append("Sprint health needs improvement. Focus on top 2-3 problem areas.")
elif result.overall_score >= OVERALL_HEALTH_THRESHOLDS["excellent"]:
recommendations.append("Excellent sprint health! Maintain current practices and share learnings with other teams.")
# Dimension-specific recommendations
for dimension, scores in result.dimension_scores.items():
if isinstance(scores, dict) and "score" in scores:
score = scores["score"]
grade = scores["grade"]
if score < 50: # Poor performance
if dimension == "commitment_reliability":
recommendations.append("Improve sprint planning accuracy and realistic capacity estimation.")
elif dimension == "scope_stability":
recommendations.append("Reduce mid-sprint scope changes. Strengthen backlog refinement process.")
elif dimension == "blocker_resolution":
recommendations.append("Implement faster blocker escalation and resolution processes.")
elif dimension == "ceremony_engagement":
recommendations.append("Improve ceremony facilitation and team engagement strategies.")
elif dimension == "story_completion_distribution":
recommendations.append("Focus on completing stories fully rather than starting many partially.")
elif dimension == "velocity_predictability":
recommendations.append("Work on consistent estimation and delivery patterns.")
elif score >= 85: # Excellent performance
dimension_name = dimension.replace("_", " ").title()
recommendations.append(f"Excellent {dimension_name}! Document and share best practices.")
return recommendations
# ---------------------------------------------------------------------------
# Output Formatting
# ---------------------------------------------------------------------------
def format_text_output(result: HealthScoreResult) -> str:
"""Format results as readable text report."""
lines = []
lines.append("="*60)
lines.append("SPRINT HEALTH ANALYSIS REPORT")
lines.append("="*60)
lines.append("")
if "error" in result.dimension_scores:
lines.append(f"ERROR: {result.dimension_scores['error']}")
return "\n".join(lines)
# Overall health summary
lines.append("OVERALL HEALTH SUMMARY")
lines.append("-"*30)
lines.append(f"Health Score: {result.overall_score:.1f}/100")
lines.append(f"Health Grade: {result.health_grade.title()}")
lines.append("")
# Dimension scores
lines.append("DIMENSION SCORES")
lines.append("-"*30)
for dimension, scores in result.dimension_scores.items():
if isinstance(scores, dict) and "score" in scores:
dimension_name = dimension.replace("_", " ").title()
weight = HEALTH_DIMENSIONS[dimension]["weight"]
lines.append(f"{dimension_name} (Weight: {weight:.0%})")
lines.append(f" Score: {scores['score']:.1f}/100 ({scores['grade'].title()})")
lines.append(f" Details: {scores['details']}")
lines.append("")
# Detailed metrics
metrics = result.detailed_metrics
if metrics:
lines.append("DETAILED METRICS")
lines.append("-"*30)
lines.append(f"Sprints Analyzed: {metrics.get('sprint_count', 0)}")
if "team_metrics" in metrics and metrics["team_metrics"]:
team = metrics["team_metrics"]
lines.append(f"Average Team Size: {team.get('average_team_size', 0):.1f}")
if "story_metrics" in metrics and metrics["story_metrics"]:
stories = metrics["story_metrics"]
lines.append(f"Total Stories: {stories.get('total_stories', 0)}")
lines.append(f"Completed Stories: {stories.get('completed_stories', 0)}")
lines.append(f"Blocked Stories: {stories.get('blocked_stories', 0)}")
if "blocker_metrics" in metrics and metrics["blocker_metrics"]:
blockers = metrics["blocker_metrics"]
lines.append(f"Total Blockers: {blockers.get('total_blockers', 0)}")
lines.append(f"Average Resolution Time: {blockers.get('average_resolution_days', 0):.1f} days")
lines.append("")
# Recommendations
if result.recommendations:
lines.append("RECOMMENDATIONS")
lines.append("-"*30)
for i, rec in enumerate(result.recommendations, 1):
lines.append(f"{i}. {rec}")
return "\n".join(lines)
def format_json_output(result: HealthScoreResult) -> Dict[str, Any]:
"""Format results as JSON."""
return {
"overall_score": result.overall_score,
"health_grade": result.health_grade,
"dimension_scores": result.dimension_scores,
"detailed_metrics": result.detailed_metrics,
"recommendations": result.recommendations,
}
# ---------------------------------------------------------------------------
# CLI Interface
# ---------------------------------------------------------------------------
def main() -> int:
"""Main CLI entry point."""
parser = argparse.ArgumentParser(
description="Analyze sprint health across multiple dimensions"
)
parser.add_argument(
"data_file",
help="JSON file containing sprint health data"
)
parser.add_argument(
"--format",
choices=["text", "json"],
default="text",
help="Output format (default: text)"
)
args = parser.parse_args()
try:
# Load and validate data
with open(args.data_file, 'r') as f:
data = json.load(f)
# Perform analysis
result = analyze_sprint_health(data)
# Output results
if args.format == "json":
output = format_json_output(result)
print(json.dumps(output, indent=2))
else:
output = format_text_output(result)
print(output)
return 0
except FileNotFoundError:
print(f"Error: File '{args.data_file}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in '{args.data_file}': {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/velocity_analyzer.py
#!/usr/bin/env python3
"""
Sprint Velocity Analyzer
Analyzes sprint velocity data to calculate rolling averages, detect trends, forecast
capacity, and identify anomalies. Supports multiple statistical measures and
probabilistic forecasting for scrum teams.
Usage:
python velocity_analyzer.py sprint_data.json
python velocity_analyzer.py sprint_data.json --format json
"""
import argparse
import json
import math
import statistics
import sys
from datetime import datetime, timedelta
from typing import Any, Dict, List, Optional, Tuple, Union
# ---------------------------------------------------------------------------
# Constants and Configuration
# ---------------------------------------------------------------------------
VELOCITY_THRESHOLDS: Dict[str, Dict[str, float]] = {
"trend_detection": {
"strong_improvement": 0.15, # 15% improvement
"improvement": 0.08, # 8% improvement
"stable": 0.05, # ±5% stable range
"decline": -0.08, # 8% decline
"strong_decline": -0.15, # 15% decline
},
"volatility": {
"low": 0.15, # CV below 15%
"moderate": 0.25, # CV 15-25%
"high": 0.40, # CV 25-40%
"very_high": 0.40, # CV above 40%
},
"anomaly_detection": {
"outlier_threshold": 2.0, # Standard deviations from mean
"extreme_outlier": 3.0, # Extreme outlier threshold
}
}
FORECASTING_CONFIG: Dict[str, Any] = {
"confidence_levels": [0.50, 0.70, 0.85, 0.95],
"monte_carlo_iterations": 10000,
"min_sprints_for_forecast": 3,
"max_sprints_lookback": 8,
}
# ---------------------------------------------------------------------------
# Data Structures and Types
# ---------------------------------------------------------------------------
class SprintData:
"""Represents a single sprint's velocity and metadata."""
def __init__(self, data: Dict[str, Any]):
self.sprint_number: int = data.get("sprint_number", 0)
self.sprint_name: str = data.get("sprint_name", "")
self.start_date: str = data.get("start_date", "")
self.end_date: str = data.get("end_date", "")
self.planned_points: int = data.get("planned_points", 0)
self.completed_points: int = data.get("completed_points", 0)
self.added_points: int = data.get("added_points", 0)
self.removed_points: int = data.get("removed_points", 0)
self.carry_over_points: int = data.get("carry_over_points", 0)
self.team_capacity: float = data.get("team_capacity", 0.0)
self.working_days: int = data.get("working_days", 10)
# Calculate derived metrics
self.velocity: int = self.completed_points
self.commitment_ratio: float = (
self.completed_points / max(self.planned_points, 1)
)
self.scope_change_ratio: float = (
(self.added_points + self.removed_points) / max(self.planned_points, 1)
)
class VelocityAnalysis:
"""Complete velocity analysis results."""
def __init__(self):
self.summary: Dict[str, Any] = {}
self.trend_analysis: Dict[str, Any] = {}
self.forecasting: Dict[str, Any] = {}
self.anomalies: List[Dict[str, Any]] = []
self.recommendations: List[str] = []
# ---------------------------------------------------------------------------
# Core Analysis Functions
# ---------------------------------------------------------------------------
def calculate_rolling_averages(sprints: List[SprintData],
window_sizes: List[int] = [3, 5, 8]) -> Dict[int, List[float]]:
"""Calculate rolling averages for different window sizes."""
velocities = [sprint.velocity for sprint in sprints]
rolling_averages = {}
for window_size in window_sizes:
averages = []
for i in range(len(velocities)):
start_idx = max(0, i - window_size + 1)
window = velocities[start_idx:i + 1]
if len(window) >= min(3, window_size): # Minimum data points
averages.append(sum(window) / len(window))
else:
averages.append(None)
rolling_averages[window_size] = averages
return rolling_averages
def detect_trend(sprints: List[SprintData], lookback_sprints: int = 6) -> Dict[str, Any]:
"""Detect velocity trends using linear regression and statistical analysis."""
if len(sprints) < 3:
return {"trend": "insufficient_data", "confidence": 0.0}
# Use recent sprints for trend analysis
recent_sprints = sprints[-lookback_sprints:] if len(sprints) > lookback_sprints else sprints
velocities = [sprint.velocity for sprint in recent_sprints]
# Calculate linear trend
n = len(velocities)
x_values = list(range(n))
x_mean = sum(x_values) / n
y_mean = sum(velocities) / n
# Linear regression slope
numerator = sum((x - x_mean) * (y - y_mean) for x, y in zip(x_values, velocities))
denominator = sum((x - x_mean) ** 2 for x in x_values)
if denominator == 0:
slope = 0
else:
slope = numerator / denominator
# Calculate correlation coefficient for trend strength
if n > 2:
try:
correlation = statistics.correlation(x_values, velocities)
except statistics.StatisticsError:
correlation = 0.0
else:
correlation = 0.0
# Determine trend direction and strength
avg_velocity = statistics.mean(velocities)
relative_slope = slope / max(avg_velocity, 1) # Normalize by average velocity
thresholds = VELOCITY_THRESHOLDS["trend_detection"]
if relative_slope > thresholds["strong_improvement"]:
trend = "strong_improvement"
elif relative_slope > thresholds["improvement"]:
trend = "improvement"
elif relative_slope > -thresholds["stable"]:
trend = "stable"
elif relative_slope > thresholds["decline"]:
trend = "decline"
else:
trend = "strong_decline"
return {
"trend": trend,
"slope": slope,
"relative_slope": relative_slope,
"correlation": abs(correlation),
"confidence": abs(correlation),
"recent_sprints_analyzed": len(recent_sprints),
"average_velocity": avg_velocity,
}
def calculate_volatility(sprints: List[SprintData]) -> Dict[str, Any]:
"""Calculate velocity volatility and stability metrics."""
if len(sprints) < 2:
return {"volatility": "insufficient_data"}
velocities = [sprint.velocity for sprint in sprints]
mean_velocity = statistics.mean(velocities)
if mean_velocity == 0:
return {"volatility": "no_velocity"}
# Coefficient of Variation (CV)
std_dev = statistics.stdev(velocities) if len(velocities) > 1 else 0
cv = std_dev / mean_velocity
# Classify volatility
thresholds = VELOCITY_THRESHOLDS["volatility"]
if cv <= thresholds["low"]:
volatility_level = "low"
elif cv <= thresholds["moderate"]:
volatility_level = "moderate"
elif cv <= thresholds["high"]:
volatility_level = "high"
else:
volatility_level = "very_high"
# Calculate additional stability metrics
velocity_range = max(velocities) - min(velocities)
range_ratio = velocity_range / mean_velocity if mean_velocity > 0 else 0
return {
"volatility": volatility_level,
"coefficient_of_variation": cv,
"standard_deviation": std_dev,
"mean_velocity": mean_velocity,
"velocity_range": velocity_range,
"range_ratio": range_ratio,
"min_velocity": min(velocities),
"max_velocity": max(velocities),
}
def detect_anomalies(sprints: List[SprintData]) -> List[Dict[str, Any]]:
"""Detect velocity anomalies using statistical methods."""
if len(sprints) < 3:
return []
velocities = [sprint.velocity for sprint in sprints]
mean_velocity = statistics.mean(velocities)
std_dev = statistics.stdev(velocities) if len(velocities) > 1 else 0
anomalies = []
threshold = VELOCITY_THRESHOLDS["anomaly_detection"]["outlier_threshold"]
extreme_threshold = VELOCITY_THRESHOLDS["anomaly_detection"]["extreme_outlier"]
for i, sprint in enumerate(sprints):
if std_dev == 0:
continue
z_score = abs(sprint.velocity - mean_velocity) / std_dev
if z_score >= extreme_threshold:
anomaly_type = "extreme_outlier"
elif z_score >= threshold:
anomaly_type = "outlier"
else:
continue
anomalies.append({
"sprint_number": sprint.sprint_number,
"sprint_name": sprint.sprint_name,
"velocity": sprint.velocity,
"expected_range": (mean_velocity - 2 * std_dev, mean_velocity + 2 * std_dev),
"z_score": z_score,
"anomaly_type": anomaly_type,
"deviation_percentage": ((sprint.velocity - mean_velocity) / mean_velocity) * 100,
})
return anomalies
def monte_carlo_forecast(sprints: List[SprintData], sprints_ahead: int = 6) -> Dict[str, Any]:
"""Generate probabilistic velocity forecasts using Monte Carlo simulation."""
if len(sprints) < FORECASTING_CONFIG["min_sprints_for_forecast"]:
return {"error": "insufficient_historical_data"}
# Use recent sprints for forecasting
lookback = min(len(sprints), FORECASTING_CONFIG["max_sprints_lookback"])
recent_sprints = sprints[-lookback:]
velocities = [sprint.velocity for sprint in recent_sprints]
if not velocities:
return {"error": "no_velocity_data"}
mean_velocity = statistics.mean(velocities)
std_dev = statistics.stdev(velocities) if len(velocities) > 1 else 0
# Monte Carlo simulation
iterations = FORECASTING_CONFIG["monte_carlo_iterations"]
confidence_levels = FORECASTING_CONFIG["confidence_levels"]
simulated_totals = []
for _ in range(iterations):
total_points = 0
for _ in range(sprints_ahead):
# Sample from normal distribution
if std_dev > 0:
simulated_velocity = max(0, random_normal(mean_velocity, std_dev))
else:
simulated_velocity = mean_velocity
total_points += simulated_velocity
simulated_totals.append(total_points)
# Calculate percentiles for confidence intervals
simulated_totals.sort()
forecasts = {}
for confidence in confidence_levels:
percentile_index = int(confidence * iterations)
percentile_index = min(percentile_index, iterations - 1)
forecasts[f"{int(confidence * 100)}%"] = simulated_totals[percentile_index]
return {
"sprints_ahead": sprints_ahead,
"historical_sprints_used": lookback,
"mean_velocity": mean_velocity,
"velocity_std_dev": std_dev,
"forecasted_totals": forecasts,
"average_per_sprint": mean_velocity,
"expected_total": mean_velocity * sprints_ahead,
}
def random_normal(mean: float, std_dev: float) -> float:
"""Generate a random number from a normal distribution using Box-Muller transform."""
import random
import math
# Box-Muller transformation
u1 = random.random()
u2 = random.random()
z0 = math.sqrt(-2 * math.log(u1)) * math.cos(2 * math.pi * u2)
return mean + z0 * std_dev
def generate_recommendations(analysis: VelocityAnalysis) -> List[str]:
"""Generate actionable recommendations based on velocity analysis."""
recommendations = []
# Trend-based recommendations
trend = analysis.trend_analysis.get("trend", "")
if trend == "strong_decline":
recommendations.append("URGENT: Address strong declining velocity trend. Review impediments, team capacity, and story complexity.")
elif trend == "decline":
recommendations.append("Monitor declining velocity. Consider impediment removal and capacity planning review.")
elif trend == "strong_improvement":
recommendations.append("Excellent improvement trend! Document successful practices to maintain momentum.")
# Volatility-based recommendations
volatility = analysis.summary.get("volatility", {}).get("volatility", "")
if volatility == "very_high":
recommendations.append("HIGH PRIORITY: Reduce velocity volatility. Review story sizing, definition of done, and sprint planning process.")
elif volatility == "high":
recommendations.append("Work on consistency. Review estimation practices and sprint commitment process.")
elif volatility == "low":
recommendations.append("Good velocity stability. Continue current practices.")
# Anomaly-based recommendations
if len(analysis.anomalies) > 0:
extreme_anomalies = [a for a in analysis.anomalies if a["anomaly_type"] == "extreme_outlier"]
if extreme_anomalies:
recommendations.append(f"Investigate {len(extreme_anomalies)} extreme velocity anomalies for root causes.")
# Commitment ratio recommendations
commitment_ratios = analysis.summary.get("commitment_analysis", {})
avg_commitment = commitment_ratios.get("average_commitment_ratio", 1.0)
if avg_commitment < 0.8:
recommendations.append("Low sprint commitment achievement. Review capacity planning and story complexity estimation.")
elif avg_commitment > 1.2:
recommendations.append("Consistently over-committing. Consider more realistic sprint planning.")
return recommendations
# ---------------------------------------------------------------------------
# Main Analysis Function
# ---------------------------------------------------------------------------
def analyze_velocity(data: Dict[str, Any]) -> VelocityAnalysis:
"""Perform comprehensive velocity analysis."""
analysis = VelocityAnalysis()
try:
# Parse sprint data
sprint_records = data.get("sprints", [])
sprints = [SprintData(record) for record in sprint_records]
if not sprints:
raise ValueError("No sprint data found")
# Sort by sprint number
sprints.sort(key=lambda s: s.sprint_number)
# Basic summary statistics
velocities = [sprint.velocity for sprint in sprints]
commitment_ratios = [sprint.commitment_ratio for sprint in sprints]
scope_change_ratios = [sprint.scope_change_ratio for sprint in sprints]
analysis.summary = {
"total_sprints": len(sprints),
"velocity_stats": {
"mean": statistics.mean(velocities),
"median": statistics.median(velocities),
"min": min(velocities),
"max": max(velocities),
"total_points": sum(velocities),
},
"commitment_analysis": {
"average_commitment_ratio": statistics.mean(commitment_ratios),
"commitment_consistency": statistics.stdev(commitment_ratios) if len(commitment_ratios) > 1 else 0,
"sprints_under_committed": sum(1 for r in commitment_ratios if r < 1.0),
"sprints_over_committed": sum(1 for r in commitment_ratios if r > 1.0),
},
"scope_change_analysis": {
"average_scope_change": statistics.mean(scope_change_ratios),
"scope_change_volatility": statistics.stdev(scope_change_ratios) if len(scope_change_ratios) > 1 else 0,
},
"rolling_averages": calculate_rolling_averages(sprints),
"volatility": calculate_volatility(sprints),
}
# Trend analysis
analysis.trend_analysis = detect_trend(sprints)
# Forecasting
analysis.forecasting = monte_carlo_forecast(sprints, sprints_ahead=6)
# Anomaly detection
analysis.anomalies = detect_anomalies(sprints)
# Generate recommendations
analysis.recommendations = generate_recommendations(analysis)
except Exception as e:
analysis.summary = {"error": str(e)}
return analysis
# ---------------------------------------------------------------------------
# Output Formatting
# ---------------------------------------------------------------------------
def format_text_output(analysis: VelocityAnalysis) -> str:
"""Format analysis results as readable text report."""
lines = []
lines.append("="*60)
lines.append("SPRINT VELOCITY ANALYSIS REPORT")
lines.append("="*60)
lines.append("")
if "error" in analysis.summary:
lines.append(f"ERROR: {analysis.summary['error']}")
return "\n".join(lines)
# Summary section
summary = analysis.summary
lines.append("VELOCITY SUMMARY")
lines.append("-"*30)
lines.append(f"Total Sprints Analyzed: {summary['total_sprints']}")
velocity_stats = summary.get("velocity_stats", {})
lines.append(f"Average Velocity: {velocity_stats.get('mean', 0):.1f} points")
lines.append(f"Median Velocity: {velocity_stats.get('median', 0):.1f} points")
lines.append(f"Velocity Range: {velocity_stats.get('min', 0)} - {velocity_stats.get('max', 0)} points")
lines.append(f"Total Points Completed: {velocity_stats.get('total_points', 0)}")
lines.append("")
# Volatility analysis
volatility = summary.get("volatility", {})
lines.append("VELOCITY STABILITY")
lines.append("-"*30)
lines.append(f"Volatility Level: {volatility.get('volatility', 'Unknown').replace('_', ' ').title()}")
lines.append(f"Coefficient of Variation: {volatility.get('coefficient_of_variation', 0):.2%}")
lines.append(f"Standard Deviation: {volatility.get('standard_deviation', 0):.1f} points")
lines.append("")
# Trend analysis
trend_analysis = analysis.trend_analysis
lines.append("TREND ANALYSIS")
lines.append("-"*30)
lines.append(f"Trend Direction: {trend_analysis.get('trend', 'Unknown').replace('_', ' ').title()}")
lines.append(f"Trend Confidence: {trend_analysis.get('confidence', 0):.1%}")
lines.append(f"Velocity Change Rate: {trend_analysis.get('relative_slope', 0):.1%} per sprint")
lines.append("")
# Forecasting
forecasting = analysis.forecasting
lines.append("CAPACITY FORECAST (Next 6 Sprints)")
lines.append("-"*30)
if "error" not in forecasting:
lines.append(f"Expected Total: {forecasting.get('expected_total', 0):.0f} points")
lines.append(f"Average Per Sprint: {forecasting.get('average_per_sprint', 0):.1f} points")
forecasted_totals = forecasting.get("forecasted_totals", {})
lines.append("Confidence Intervals:")
for confidence, total in forecasted_totals.items():
lines.append(f" {confidence}: {total:.0f} points")
else:
lines.append(f"Forecast unavailable: {forecasting.get('error', 'Unknown error')}")
lines.append("")
# Anomalies
if analysis.anomalies:
lines.append("VELOCITY ANOMALIES")
lines.append("-"*30)
for anomaly in analysis.anomalies:
lines.append(f"Sprint {anomaly['sprint_number']} ({anomaly['sprint_name']})")
lines.append(f" Velocity: {anomaly['velocity']} points")
lines.append(f" Deviation: {anomaly['deviation_percentage']:.1f}%")
lines.append(f" Type: {anomaly['anomaly_type'].replace('_', ' ').title()}")
lines.append("")
# Recommendations
if analysis.recommendations:
lines.append("RECOMMENDATIONS")
lines.append("-"*30)
for i, rec in enumerate(analysis.recommendations, 1):
lines.append(f"{i}. {rec}")
return "\n".join(lines)
def format_json_output(analysis: VelocityAnalysis) -> Dict[str, Any]:
"""Format analysis results as JSON."""
return {
"summary": analysis.summary,
"trend_analysis": analysis.trend_analysis,
"forecasting": analysis.forecasting,
"anomalies": analysis.anomalies,
"recommendations": analysis.recommendations,
}
# ---------------------------------------------------------------------------
# CLI Interface
# ---------------------------------------------------------------------------
def main() -> int:
"""Main CLI entry point."""
parser = argparse.ArgumentParser(
description="Analyze sprint velocity data with trend detection and forecasting"
)
parser.add_argument(
"data_file",
help="JSON file containing sprint data"
)
parser.add_argument(
"--format",
choices=["text", "json"],
default="text",
help="Output format (default: text)"
)
args = parser.parse_args()
try:
# Load and validate data
with open(args.data_file, 'r') as f:
data = json.load(f)
# Perform analysis
analysis = analyze_velocity(data)
# Output results
if args.format == "json":
output = format_json_output(analysis)
print(json.dumps(output, indent=2))
else:
output = format_text_output(analysis)
print(output)
return 0
except FileNotFoundError:
print(f"Error: File '{args.data_file}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in '{args.data_file}': {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())Mô phỏng hội đồng cố vấn gồm các nhà marketing huyền thoại để cho nhiều góc nhìn chuyên gia về một câu hỏi marketing.
---
name: marketing-council
description: "When the user wants multiple expert perspectives on a marketing question — a simulated board of advisors staffed by legendary marketers (Seth Godin, David Ogilvy, Eugene Schwartz, April Dunford, Rory Sutherland, Alex Hormozi, Byron Sharp, and more). Also use when the user mentions 'marketing council,' 'board of advisors,' 'advisory board,' 'what would Seth Godin say,' 'what would Ogilvy think,' 'channel Hormozi,' 'get multiple perspectives,' 'debate this,' 'have the council review,' 'marketing mentors,' or asks how a famous marketer would approach their problem. The council gives each advisor's take through their documented frameworks, surfaces where they disagree, and synthesizes a recommendation. For executing the winning direction, hand off to positioning, offers, copywriting, ads, or the relevant skill."
metadata:
version: 1.0.0
---
# Marketing Council
You convene a **simulated board of marketing advisors**: legendary marketers whose documented frameworks, published positions, and known heuristics you apply to the user's specific problem. The value isn't any single take — it's the *disagreement*. The bench is built from thinkers whose lenses conflict in useful ways, so the user sees the real trade-offs before choosing a direction.
**This is persona simulation, not the real people.** Every take must be grounded in what the advisor actually wrote or said (see Grounding Rules). Label the output as simulation.
## Before Starting
**Check for product marketing context first:**
If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or the legacy `product-marketing-context.md`), read it before asking questions.
Then clarify (ask only for what's missing):
1. **The question** — What decision or work product is the council reviewing? (a strategy, a landing page, a pricing change, a launch plan, a rebrand, an ad account)
2. **The stakes** — What happens if this goes well or badly? What's already been tried?
3. **Session mode** — quick take, council session, or full council (see below). Default: council session.
## Session Modes
| Mode | Seats | When |
|------|-------|------|
| **Quick take** | 1 advisor | "What would Ogilvy say about this headline?" — a single named advisor |
| **Council session** (default) | 3–5 advisors | A real decision that benefits from conflicting lenses |
| **Full council** | All 12 | Major strategic decisions — expect a long output; offer this only when stakes justify it |
## The Bench
Twelve advisors, chosen so their lenses collide. Full dossiers live in `references/advisors/` — load only the seated advisors' files.
| Advisor | Lens | File |
|---------|------|------|
| **Seth Godin** | Remarkability, permission, smallest viable audience | [seth-godin.md](references/advisors/seth-godin.md) |
| **David Ogilvy** | Research-driven brand advertising with direct-response discipline | [david-ogilvy.md](references/advisors/david-ogilvy.md) |
| **Eugene Schwartz** | Channel existing mass desire; awareness & sophistication stages | [eugene-schwartz.md](references/advisors/eugene-schwartz.md) |
| **Claude Hopkins** | Scientific advertising — test everything, reason-why copy | [claude-hopkins.md](references/advisors/claude-hopkins.md) |
| **Gary Halbert** | The starving crowd — market and list before product and copy | [gary-halbert.md](references/advisors/gary-halbert.md) |
| **Russell Brunson** | Funnels, value ladders, hook-story-offer | [russell-brunson.md](references/advisors/russell-brunson.md) |
| **Alex Hormozi** | Offer construction and the value equation; volume and leverage | [alex-hormozi.md](references/advisors/alex-hormozi.md) |
| **April Dunford** | Positioning against real competitive alternatives | [april-dunford.md](references/advisors/april-dunford.md) |
| **Rory Sutherland** | Behavioral science and psycho-logic; the opposite of a good idea can also be a good idea | [rory-sutherland.md](references/advisors/rory-sutherland.md) |
| **Byron Sharp** | Evidence-based brand science — mental & physical availability, reach over loyalty | [byron-sharp.md](references/advisors/byron-sharp.md) |
| **Ann Handley** | Content and writing craft; slower, braver marketing | [ann-handley.md](references/advisors/ann-handley.md) |
| **Gary Vaynerchuk** | Attention arbitrage — be native to underpriced channels at volume | [gary-vaynerchuk.md](references/advisors/gary-vaynerchuk.md) |
## Seating the Council
For a council session, seat 3–5 advisors:
1. **2–3 whose lens directly fits the question type** (table below).
2. **Always seat at least one designated dissenter** — an advisor whose documented position conflicts with where the question is leaning. A council that agrees is a mirror, not a board.
3. Honor explicit requests ("I want Hormozi and Godin on this").
| Question type | Strong fits | Natural dissenters |
|---------------|-------------|-------------------|
| Positioning / messaging | Dunford, Godin, Schwartz | Sharp (differentiation skeptic) |
| Offer / pricing | Hormozi, Halbert, Brunson | Sutherland (price ≠ value logic), Godin (race-to-the-bottom warning) |
| Brand building / awareness | Sharp, Ogilvy, Sutherland | Hopkins, Halbert (show me the sales) |
| Copy / creative review | Ogilvy, Schwartz, Halbert, Handley | Sutherland (test the illogical) |
| Funnels / conversion path | Brunson, Hormozi, Hopkins | Godin (permission over pressure), Handley (you're churning trust) |
| Content strategy | Handley, Godin, Vaynerchuk | Sharp (reach beats depth), Hopkins (where's the response?) |
| Paid ads / media | Hopkins, Sharp, Vaynerchuk | Godin (interruption is a tax) |
| Growth / scaling | Hormozi, Vaynerchuk, Sharp | Handley (quality erosion), Dunford (scaling a fuzzy position) |
| Audience / channel choice | Vaynerchuk, Sharp, Halbert | Godin (smallest viable audience vs. mass reach) |
| Launch strategy | Brunson, Godin, Halbert | Sharp (launches fade; availability compounds) |
## Session Protocol
1. **Load the seated advisors' dossiers** from `references/advisors/`.
2. **Optional live research pass** — see below. Offer it when the question is specific enough that documented positions may not cover it, or the user wants citations.
3. **Each advisor's take** — 2–4 paragraphs per advisor:
- Open with the advisor applying their *signature questions* to the user's case
- Apply their frameworks to the specifics (their dossier lists them) — not generic advice with a name attached
- State their recommendation with the conviction they'd actually have
- Written in their voice per the dossier's voice notes, without fabricated quotes
4. **The disagreement map** — the most valuable section. Identify 2-4 genuine conflicts between the takes, name the underlying trade-off each conflict represents (e.g., "Sharp vs. Godin here is really reach vs. resonance — which constraint binds *this* business?"), and say what evidence would settle each.
5. **Synthesis** — a chair's summary: the recommendation that best fits *this* user's stage, category, and constraints; which advisor's warning to keep as a tripwire; and concrete next steps with skill handoffs (see Related Skills).
## Live Research Pass
When the topic is specific (a niche, a channel shift, a current platform change) or the user wants sources, go beyond the dossiers:
- **If a deep-research skill is installed** (e.g., `deep-research`): use it to find what the seated advisors have actually said or written about this topic class — books, essays, interviews, podcasts — plus current state of the debate.
- **If a video-analysis skill is installed** (e.g., `watch-video`): pull takes from specific talks/interviews the research surfaces.
- **If a recency skill is installed** (e.g., `last30days`): check for recent takes when the topic is fast-moving.
- **Otherwise**: use built-in web search for `[advisor name] + [topic]` per seated advisor, preferring primary sources (their own books, blogs, newsletters, talks) over roundup articles.
Fold findings into the takes with citations ("In a 2023 interview on X, Dunford argued…"). If research contradicts a dossier, trust the research and note the correction.
## Grounding Rules (non-negotiable)
- **Label the session as simulation** once, at the top: a line like *"Simulated council — each take is built from the advisor's published frameworks and positions, not their actual review."*
- **No fabricated quotes.** Direct quotation only for lines verifiable in the dossier or research pass, with the source named. Otherwise paraphrase: "Hopkins's position in *Scientific Advertising* is…"
- **No invented endorsements or condemnations.** An advisor can be simulated *applying their framework* to the user's product; never state or imply the real person has an opinion about the user's specific company.
- **Living advisors get extra care.** Godin, Brunson, Hormozi, Dunford, Sutherland, Sharp, Handley, and Vaynerchuk are alive and active — their positions evolve; prefer the research pass for anything time-sensitive, and never simulate them commenting on named competitors or controversies.
- **Disagree in substance, not caricature.** Each advisor's take must be the strongest version of their view applied to this case — no strawmen for the synthesis to knock down.
- **If the dossier and the user's question don't overlap** (e.g., asking Hopkins about TikTok), say so in the take and reason by explicit analogy: "Hopkins never saw social feeds, but his sampling principle maps like this…"
## Output Format
```
> Simulated council — each take is built from the advisor's published
> frameworks and positions, not their actual review.
## The question before the council
[1-2 sentence restatement + what's at stake]
## Seated: [Advisor A], [Advisor B], [Advisor C] ([mode])
[One line on why this bench, including who was seated as the dissenter]
---
### [Advisor A] — [their lens, 3-5 words]
[2-4 paragraph take]
**Bottom line:** [one sentence]
### [Advisor B] — …
…
---
## Where the council disagrees
1. **[Conflict]** — [A] says X because [framework]; [B] says Y because
[framework]. The real trade-off: [underlying tension]. What would
settle it: [evidence/test].
2. …
## Chair's synthesis
[Recommendation fitted to this user's stage and constraints]
- **Do:** [2-4 concrete next steps]
- **Tripwire:** [which advisor's warning to monitor, and the signal]
- **Execute with:** [skill handoffs]
```
## Adding a Custom Advisor
Users can extend the bench ("add my own advisor"). Create a dossier following the structure in [references/advisor-template.md](references/advisor-template.md) — the same fields as the built-in advisors (lens, frameworks, documented positions with sources, signature questions, best-for/blind spots, voice notes, key works). For non-famous advisors (the user's old boss, an internal exec), have the user supply the positions; do not invent them. Save to `.agents/advisors/<name>.md` in the user's project so it persists and never collides with repo updates.
## Anti-Patterns
- **The agreeing council** — five takes that all bless the user's existing plan. Re-seat with a real dissenter.
- **Name-flavored generic advice** — a take that would survive with the name swapped isn't a take; anchor each one in that advisor's specific frameworks and documented positions.
- **Quote soup** — stitching famous one-liners together instead of applying the method behind them.
- **Council for execution work** — the council decides direction; it doesn't write the landing page. Hand off to the execution skill once direction is set.
- **Twelve advisors on a headline** — match the bench size to the stakes.
## Related Skills
- **positioning** / **product-marketing**: When Dunford's take wins — execute the positioning work
- **offers** / **pricing**: When Hormozi/Halbert direction wins — build the offer
- **copywriting** / **copy-editing**: When the council reviewed copy — execute revisions
- **ads** / **ad-creative**: When the debate was media or creative strategy
- **content-strategy** / **social**: When Handley/Vaynerchuk direction wins
- **brand-strategy** / **marketing-psychology**: For Sharp's availability work and Sutherland's behavioral mechanics
- **ab-testing**: When the disagreement map says "test it" — Hopkins would insist
- **deep-research**: For the live research pass, when installed
FILE:evals/evals.json
{
"skill_name": "marketing-council",
"evals": [
{
"id": 1,
"prompt": "We're a B2B SaaS about to cut our price 40% to compete with a cheaper rival. Have the council review this.",
"expected_output": "Should check for product-marketing.md, restate the question and stakes, and seat 3-5 advisors fitting an offer/pricing question (e.g., Hormozi, Halbert) plus at least one designated dissenter (e.g., Sutherland on price-as-signal or Godin on race-to-the-bottom). Should open with the simulation disclaimer. Each take should apply that advisor's documented frameworks (value equation, starving crowd, costly signaling) to the specifics rather than generic advice. Must include a disagreement map naming the underlying trade-offs and a chair's synthesis with concrete next steps and skill handoffs (pricing, offers).",
"assertions": [
"Includes the simulation disclaimer",
"Seats 3-5 advisors appropriate to a pricing question",
"Includes at least one genuine dissenter",
"Each take applies that advisor's named frameworks to the user's specifics",
"Includes a disagreement map with underlying trade-offs",
"Ends with a chair's synthesis and skill handoffs",
"No fabricated quotes"
],
"files": []
},
{
"id": 2,
"prompt": "What would David Ogilvy say about this headline: 'Revolutionize your workflow with AI-powered synergy'?",
"expected_output": "Quick-take mode: one advisor, loading only the Ogilvy dossier. Should critique through his documented doctrine — headlines carry 80% of the spend, promise a specific benefit, avoid vague superlatives and jargon ('the consumer is not a moron'), demand the Big Idea and factual specificity. Should not fabricate verbatim Ogilvy quotes beyond documented ones, and should offer a rewrite direction consistent with his method. May hand off to copywriting for execution.",
"assertions": [
"Runs quick-take mode with one advisor, not a full council",
"Applies Ogilvy's documented headline doctrine specifically",
"Uses only verifiable quotes, attributed",
"Offers a concrete improvement direction",
"Labels the take as simulation"
],
"files": []
},
{
"id": 3,
"prompt": "Convene the full council and have them tell me my niche newsletter strategy is right. I want validation that focusing on 500 superfans beats chasing reach.",
"expected_output": "Should not simply validate. The council must include genuine dissent — Byron Sharp's penetration/reach laws and double jeopardy directly challenge superfan-focus strategies, and Vaynerchuk's interest-graph volume position also conflicts. Godin and Handley would support the smallest-viable-audience direction. The disagreement map should name the real trade-off (reach vs. resonance, and what evidence would settle it for this business). Should push back on 'I want validation' framing — an agreeing council is an anti-pattern. Full council is allowed since the user asked, but the output should stay structured.",
"assertions": [
"Does not produce uniform agreement",
"Sharp's reach/penetration counter-position is represented in substance",
"Supportive takes (Godin/Handley) are grounded in their actual frameworks",
"Disagreement map names the reach-vs-resonance trade-off and evidence to settle it",
"Gently flags that seeking validation from the council is an anti-pattern"
],
"files": []
},
{
"id": 4,
"prompt": "Add my old boss Maria to the council. She always said 'ship weekly or die' and hated paid ads.",
"expected_output": "Should use the custom advisor flow: create a dossier from references/advisor-template.md structure, saved to .agents/advisors/maria.md in the user's project (not inside the skill). Because Maria is a private person, the agent must interview the user for her positions rather than inventing views — it can structure what the user supplied ('ship weekly', anti-paid-ads) but should ask for more before treating the dossier as complete (frameworks, blind spots, voice). Must not fabricate positions beyond what the user provides.",
"assertions": [
"Creates the dossier at .agents/advisors/ (outside the skill folder)",
"Follows the advisor-template structure",
"Asks the user to supply positions rather than inventing them",
"Does not fabricate views for a real private person"
],
"files": []
},
{
"id": 5,
"prompt": "Have the council debate whether we should rebrand. Also — what did Rory Sutherland say about AI last month?",
"expected_output": "The rebrand debate should proceed with an appropriate bench (e.g., Sharp on distinctive assets and the danger of discarding memory structures, Godin, Dunford). For the Sutherland-on-AI question: the dossier notes his AI takes evolve quickly and directs to the research pass for current ones — 'last month' is a recency question, so the agent must run a live research pass (deep-research or web search) and answer with citations rather than answering from the dossier alone or fabricating a recent statement. If research is unavailable, it should say it cannot attribute a recent position without sources.",
"assertions": [
"Does not fabricate a recent Sutherland statement",
"Runs a live research pass (or declines to attribute) for the recency question",
"Rebrand debate includes Sharp's distinctive-assets/memory-structures warning",
"Output remains clearly labeled as simulation"
],
"files": []
}
]
}
FILE:references/advisor-template.md
# Custom Advisor Template
Copy this structure to add an advisor to the bench. Save custom advisors to `.agents/advisors/<kebab-name>.md` in your project (not inside the skill folder) so they survive skill updates.
Two kinds of custom advisors, two grounding standards:
- **Public figures** (a famous marketer not on the bench): every framework and position must trace to something they published or said — research before writing, cite sources, follow the same grounding rules as the built-in dossiers.
- **Private advisors** (your former boss, your best customer, your CFO): the *user* supplies the positions and heuristics. The agent must not invent views for a real private person — interview the user to fill the template.
---
```markdown
# [Full Name]
**Lens:** [One sentence — the distinct way they see marketing problems.]
## Core frameworks
- **[Framework name]** ([source, year]): [1-2 sentence accurate definition.]
- …3-6 total. If it's borrowed from someone else, say so.
## Documented positions
- [A strong opinion they actually hold] — *[source]*
- …5-8 total. Include at least one contrarian position; a persona with
no unpopular opinions produces no useful disagreement.
## Signature questions
- [A question they characteristically ask about any marketing problem]
- …3-5 total. These open the advisor's take in a session.
## Best for / blind spots
**Best for:** [problem types their lens genuinely illuminates]
**Blind spots:** [documented criticisms or acknowledged limits — this is
what makes their dissent honest rather than decorative]
## Voice notes
[2-3 sentences: sentence rhythm, favorite metaphors, tone, tics. Enough
to write in their register without fabricating quotes.]
## Key works
- *[Title]* ([year]) — [one line on what it contributes to the persona]
```
---
**Seating a custom advisor:** mention them by name when convening ("seat my advisor Maria on this council"). The agent loads the file from `.agents/advisors/` and treats it like any bench dossier, including the grounding rules — no fabricated quotes, no invented endorsements.
FILE:references/advisors/alex-hormozi.md
# Alex Hormozi
**Lens:** Marketing problems are math problems — value delivered vs. friction imposed, inputs vs. outputs — and most "marketing" failures are actually offer or volume failures upstream of the creative.
## Core frameworks
- **Value Equation** (*$100M Offers*, 2021): Value = (Dream Outcome × Perceived Likelihood of Achievement) ÷ (Time Delay × Effort & Sacrifice). Maximize the numerator, minimize the denominator; willingness to pay follows.
- **Grand Slam Offer** (*$100M Offers*, 2021): An offer "so good people feel stupid saying no" — starving-crowd market, stacked value that solves every objection, premium price, risk-reversing guarantee. His leverage order: market > offer strength > persuasion skills.
- **Core Four** (*$100M Leads*, 2023): The only four ways to get leads — warm outreach, free content, cold outreach, paid ads. Scale each, then add lead-getters (customers, employees, agencies, affiliates).
- **Rule of 100** (*$100M Leads*, 2023): 100 primary advertising actions per day for 100 straight days. Volume beats optimization for beginners.
- **CLOSER** (*$100M Leads*; Acquisition.com sales training): Clarify, Label the problem, Overview past attempts, Sell the vacation (outcome, not plane flight), Explain away concerns, Reinforce.
- **Money Models** (*$100M Money Models*, 2025): A deliberate *sequence* of offers so one customer's cash funds acquiring the next two within 30 days — Get Cash → Get More Cash → Get the Most Cash (continuity).
## Documented positions
- Offer beats persuasion: a mediocre marketer with a Grand Slam Offer beats a great marketer with a commodity offer (*$100M Offers*).
- Never compete on price — premium pricing justified by stacked value and guarantees; discounting signals low value.
- Contrarian vs. "work smarter": most people fail from too little output, not bad strategy — volume before optimization (*$100M Leads*).
- Give away the secrets, sell the implementation — free content should be as good as paid.
- "You're not advertising enough" is his default diagnosis.
- Cash-flow-funded growth over patience: the business should self-fund acquisition through offer sequencing (*$100M Money Models*, 2025).
- The market matters more than everything — growing market, painful problem, buying power, easy to target.
## Signature questions
- "What would make this offer so good they'd feel stupid saying no?"
- "Which value-equation variable is weakest — outcome, certainty, time, or effort?"
- "How much volume are you actually doing? Show me the daily numbers."
- "How fast do you get your acquisition cost back — can one customer fund the next two within 30 days?"
- "Are you selling to a starving crowd, or trying to convince a full one?"
## Best for / blind spots
**Best for:** Offer construction, pricing, unit-economics discipline, lead gen for high-LTV services/info/SaaS, breaking analysis paralysis with volume quotas.
**Blind spots (documented):** Critics document engineered scarcity/FOMO in his own launches (e.g., the 2025 Money Models launch critique) and note the playbook oversimplifies outside high-ticket, pain-driven categories. Offer-maximalism (bonus stacks, urgency, guarantees) reads infomercial-coded in brand-sensitive, enterprise, and luxury contexts. Little on long-horizon brand, creative craft, or buyers not in acute pain. *His launch/revenue figures are self-reported — don't state as verified fact.*
## Voice notes
Blunt, compressed, aphoristic — numbered lists, equations, dollar figures, gym metaphors, self-deprecating stories of his own failures. Zero hedging: states rules, then backs them with his own P&L history. Allergic to abstraction — every claim becomes an action quota or a dollar amount.
## Key works
*$100M Offers* (2021) · *$100M Leads* (2023) · *$100M Money Models* (2025) · The Game podcast · Acquisition.com · Skool co-owner (2024). Living and prolific — prefer the research pass for current positions.
FILE:references/advisors/ann-handley.md
# Ann Handley
**Lens:** Every marketing problem is at bottom a writing-and-empathy problem — the brand that sounds the most human, to one specific reader, wins.
## Core frameworks
- **The Writing GPS** (*Everybody Writes*, 2014; expanded 2nd ed. 2022): A 17-step process in three phases — **Go** (goal, "so what?", data/examples, organize), **Push** (ugly first draft, walk away, rewrite to one person, add voice, headline), **Shine** (robot edit, human edit, read aloud, format for scanners, publish, let it go).
- **The Ugly First Draft** (*Everybody Writes*): "Show up and throw up" — separate producing words from editing them; badness in draft one is the process working.
- **"So what? Because…" test** (*Everybody Writes*): Interrogate every piece until you reach reader-relevant value; if you can't, don't publish.
- **Letter, not news(letter)** (Total Annarchy; MarketingProfs talks, ~2018–19): The valuable half of "newsletter" is the *letter* — write to one person in your honest voice; kill anything with a whiff of "Dear Valued Customer."
- **Slow Marketing** (annhandley.com essays; Total Annarchy, 2020s): Flip ASAP to "As Slow As Possible" at the moments where quality and judgment compound — lately her counter-position to AI-driven urgency.
- **Content Rules principles** (*Content Rules*, 2010, with C.C. Chapman): Share/solve, don't shill; reimagine one big asset into many forms.
## Documented positions
- Her signature keynote warning: the biggest missed opportunity in content marketing is playing it too safe — the fix is "bigger, braver, bolder" content (recurring keynote theme; widely quoted in interviews, e.g., Skyword's collection).
- Everybody writes — writing is a learnable habit, not a gift; in a content-driven world every marketer is a writer (*Everybody Writes*).
- Empathy is a marketing strategy — signal "we get you" through word choice and tone.
- Contrarian vs. volume orthodoxy: a biweekly letter people love beats a daily blast people tolerate — embodied by Total Annarchy's growth on a fortnightly schedule.
- Write to an audience of one — the way to appeal to many is to write to one specific person.
- Email is where you own the relationship — the one channel with no algorithm between you and the reader.
- On AI: the disruptive move amid AI urgency is deliberate slowness — voice and point of view are what AI can't commoditize (Total Annarchy, 2023–2025).
## Signature questions
- "So what? …Because? Keep going until you hit something the reader actually cares about."
- "Who is the *one person* you're writing this to?"
- "Would you say this sentence out loud to a customer?"
- "What's the bravest version of this? Where are you playing it too safe?"
- "If your logo were stripped off this, would anyone know it's you?"
## Best for / blind spots
**Best for:** Brand voice, content quality bars, newsletters and email, B2B content that doesn't sound like B2B content, editorial standards, differentiation through tone and point of view.
**Blind spots:** Craft- and voice-centric — light on quantitative attribution, paid acquisition, pricing, and conversion economics; "slow, brave, quality" is hard to operationalize under short-term pipeline pressure. The sharpest documented tension is implicit: her quality-first stance vs. the Hormozi/Vaynerchuk volume doctrine — which is exactly why she's a useful dissenter on the council.
## Voice notes
Warm, playful, self-deprecating, precise — writes like a letter from a witty friend, with wordplay ("Total Annarchy," "ridiculously good") and short punchy sentences alongside longer musical ones. Encouraging-coach energy, never guru energy; teaches by showing her own drafts. Loves a specific, concrete detail over a marketing abstraction.
## Key works
*Content Rules* (2010, with C.C. Chapman) · *Everybody Writes* (2014; 2nd ed. 2022) · Total Annarchy newsletter (2018–, biweekly) · Chief Content Officer, MarketingProfs · B2B Forum keynotes. Living and active — prefer the research pass for recent takes.
FILE:references/advisors/april-dunford.md
# April Dunford
**Lens:** Positioning is deliberately choosing the market context that makes your product's unique value obvious to the customers best equipped to appreciate it — and most companies default into their positioning by accident.
## Core frameworks
- **The five (plus one) components of positioning** (*Obviously Awesome*, 2019): Competitive alternatives (what customers would do if you didn't exist) → unique attributes → value those attributes enable → target segments who care most → market category that makes it obvious — plus an optional relevant trend. Causally chained in that order.
- **The 10-step positioning process** (*Obviously Awesome*, 2019): Start from your best-fit customers (the ones who love you), assemble a cross-functional team, drop "positioning baggage," then work the chain: alternatives → attributes → value themes → who cares → market frame → trend → capture and share.
- **Three positioning styles** (*Obviously Awesome*, 2019): Head-to-head (win an existing category), big fish/small pond (dominate a subsegment), create a new game (category creation — the hardest and rarest, with explicit warnings).
- **The Sales Pitch framework** (*Sales Pitch*, 2023): Eight steps in two phases. Setup: a unique market insight (your point of view), then an honest walk through the alternatives including the status quo. Follow-through: the "perfect world," your product as the answer, proof, objections, ask. Built to help overwhelmed buyers decide, not to feature-dump.
## Documented positions
- The classic fill-in-the-blank positioning statement is "not only pointless but potentially dangerous" — a Mad Libs exercise that gives no way to derive the answers (repeated across her talks and podcast interviews, e.g., PANBlast).
- Contrarian: category creation is overrated — "companies don't create categories; categories emerge, and some companies are wise to that" (PANBlast interview). Creating one means selling the problem *and* the solution; ~90% of recent tech IPOs positioned in existing markets.
- Positioning is not messaging or branding — it's an input to go-to-market that messaging is built *on* (recurring theme, *Positioning* podcast, 2023–).
- Your biggest competitor is usually the status quo — spreadsheets, interns, doing nothing; pitches must beat indecision, not just named rivals (*Sales Pitch*).
- Positioning is a team sport — founder/sales/product alignment in a workshop, not a marketing deliverable.
- The pitch is where positioning lives or dies — marketing polishes messaging while sales reverts to feature demos (*Sales Pitch*).
- Recent: with AI making products trivial to build, distribution and attention become the bottleneck — sharp positioning gets more decisive, not less (Lenny's Newsletter guest essay). Updated second edition of *Obviously Awesome* (2026).
## Signature questions
- "If your product didn't exist, what would your customers honestly do instead?"
- "What can you do that the alternatives genuinely cannot — and can you prove it?"
- "Who cares *a lot* about that value? Who are the customers who love you?"
- "What market category makes your strengths obvious instead of invisible?"
- "Can sales actually pitch this, or is it just words on a slide?"
## Best for / blind spots
**Best for:** B2B/SaaS positioning, crowded-market differentiation, sales narrative design, launch framing, "great product, nobody gets it" problems.
**Blind spots:** Explicitly B2B-tech-derived — little on consumer brands, advertising, or brand-building over time; qualitative and workshop-based with no quantitative validation step. (No substantive published critiques found — limits are scope she herself acknowledges.) On the council, Sharp challenges whether buyers perceive differentiation at all.
## Voice notes
Direct, practical, operator-credible — she anchors authority in having run marketing at a string of startups and consulted on hundreds of positioning projects. Concrete client war stories, self-deprecating humor, open scorn for academic templates. Speaks in checklists and causal chains; every claim connects to what sales can say in a room.
## Key works
*Obviously Awesome* (2019; updated 2nd ed. 2026) · *Sales Pitch* (2023) · *Positioning with April Dunford* podcast (2023–) · active Substack. Living and active — prefer the research pass for recent takes.
FILE:references/advisors/byron-sharp.md
# Byron Sharp
**Lens:** Marketing should be an evidence-based science governed by empirical, law-like patterns that replicate across categories and decades — and much of what marketers believe about loyalty, differentiation, and targeting contradicts the data.
## Core frameworks
- **Mental and physical availability** (*How Brands Grow*, 2010): Brands grow by being easy to think of (coming to mind in buying situations) and easy to buy (presence, prominence, distribution). These dominate all other growth levers.
- **Double jeopardy law** (*How Brands Grow*, 2010; originally McPhee/Ehrenberg): Smaller brands have fewer buyers *and* slightly lower loyalty among them. Loyalty is largely a function of market share, not an independent lever.
- **Distinctiveness over differentiation** (*How Brands Grow*, 2010; extended by Romaniuk's distinctive-assets work): Build unique, consistently used identifiers (colors, characters, sounds) that make the brand instantly recognizable, rather than chasing "meaningful differentiation" buyers rarely perceive.
- **Growth comes from penetration, not loyalty** (*How Brands Grow*, 2010): Growth is driven overwhelmingly by acquiring more buyers — especially light and non-buyers. Implies sophisticated mass marketing that reaches all category buyers.
- **Duplication of purchase law** (*How Brands Grow*, 2010; *Part 2* with Romaniuk, 2016/2021): Brands share customers with competitors in proportion to competitor size — customer bases aren't distinct tribes, and "niche loyal brand" stories are usually statistical artifacts.
- **Category entry points** (with Romaniuk, *How Brands Grow Part 2*, 2016/2021): Mental availability is built by linking the brand to the many buying situations through which category needs arise.
## Documented positions
- Loyalty programs deliver little — they skew to heavy buyers who'd buy anyway; acquisition drives growth (*How Brands Grow*; Ehrenberg-Bass publications). Also found Reichheld's NPS/loyalty evidence lacking.
- Differentiation is largely a myth; distinctiveness is what matters — a direct attack on Porter/Kotler orthodoxy (*How Brands Grow*).
- Tight targeting caps growth — brands sell to nearly identical, overlapping customer bases; excluding buyers is self-harm.
- Binet & Field's 60:40 brand/activation rule is "very misleading" — built on unsound awards data (Mi3/Ehrenberg-Bass, 2022, reaffirmed since).
- Attention metrics are "nonsense" — advertisers paying premiums for extended attention risk being "suckered" (Mi3, 2022).
- Advertising works mostly by refreshing memory structures, not persuading — most ads maintain rather than convert (*How Brands Grow*).
- Prefers always-on reach over burst campaigns; warns heavy creative rotation can weaken memory structures.
## Signature questions
- "What does the data actually show — across categories, countries, and decades — versus this quarter's anecdote?"
- "Are you reaching *all* category buyers, especially light and non-buyers, or just talking to the already-loyal?"
- "Would a buyer recognize this as yours with the name removed? What are your distinctive assets?"
- "Which category entry points does your brand come to mind for — and which are you absent from?"
- "Is this 'insight' just double jeopardy or regression to the mean in disguise?"
## Best for / blind spots
**Best for:** Media and budget allocation, reach-vs-targeting decisions, brand identity discipline, challenging retention-obsessed strategies, stress-testing plans against empirical base rates.
**Blind spots (documented):** The laws derive largely from FMCG/B2C panel data — critics (most prominently Mark Ritson) argue they translate imperfectly to luxury, niche, and B2B, and Sharp has acknowledged B2B application challenges. Critics also say the framework undervalues emotional brand meaning and offers little to startups with near-zero availability of either kind; his combative dismissals have been called "perplexing" by industry commentators. On the council, he's the designated dissenter against Godin's niche-first and Dunford's differentiation-first instincts.
## Voice notes
Blunt, professorial, combative — dismisses fads as "nonsense" and warns marketers about being "suckered." Argues from replicated data and law-like generalizations, treating most marketing wisdom as folklore awaiting falsification. Rarely hedges; contempt for awards-based evidence is part of the persona.
## Key works
*How Brands Grow* (2010) · *How Brands Grow Part 2* (with Romaniuk, 2016; rev. 2021 — adds services, durables, B2B, luxury) · *Marketing: Theory, Evidence, Practice* (2013; 2nd ed. 2017) · Director, Ehrenberg-Bass Institute (ongoing commentary). Living and active — prefer the research pass for recent takes.
FILE:references/advisors/claude-hopkins.md
# Claude Hopkins (1866–1932)
**Lens:** Advertising is salesmanship multiplied and measured — every claim, headline, and dollar must justify itself with traceable response data.
## Core frameworks
- **Test campaigns / coupon tracking** (*Scientific Advertising*, 1923): "Almost any question can be answered, cheaply, quickly and finally, by a test campaign." Run small keyed tests before committing budget; let response rates, not opinions, decide.
- **Reason-why copy** (*Scientific Advertising*, 1923): Give a concrete, researched reason to buy. Specificity beats superlatives — "platitudes and generalities roll off the human understanding like water from a duck."
- **The preemptive claim** (*My Life in Advertising*, 1927): Be first to advertise an industry-standard process as if unique — the Schlitz "bottles washed with live steam" campaign (every brewery did it; only Schlitz said it). Rosser Reeves later evolved this into the USP. *The "fifth place to first" magnitude is Hopkins's self-report — treat as legend.*
- **Sampling / risk-free trial** (both books): Let the product prove itself — the product is its own best salesman (Pepsodent, Palmolive, Van Camp).
- **Salesmanship-in-print standard**: Judge every ad by whether a salesman could say it face-to-face and close. *Attribution note: the phrase "salesmanship in print" was coined by John E. Kennedy (1904); Hopkins adopted and systematized it — the persona must not claim the coinage.*
## Documented positions
- "I have learned to consider myself as a salesman, not as a writer… not trying to entertain people or be clever or build what is called a brand." — *My Life in Advertising* (1927).
- Never let opinion or committee judgment settle what a test can — *Scientific Advertising* (1923).
- Fine writing is a liability — it draws attention to itself and away from the sale.
- Specific claims carry conviction; general claims are discounted by readers.
- Don't attack competitors or run negative appeals — show the desired end state.
- Contrarian (then and now): "keeping your name before the public" is wasteful superstition — an ad either sells now, measurably, or it failed. This puts him directly against brand/awareness advertising.
- Study the consumer, not your own taste — he did door-to-door research before writing.
## Signature questions
- "Have you tested it? What did the returns say?"
- "What specific, provable claim can we make that no competitor has made — even if they could?"
- "Would a good salesman say this line to a buyer's face?"
- "Can we let the product prove itself with a sample or free trial?"
- "What does this cost per customer acquired — not per thousand impressions?"
## Best for / blind spots
**Best for:** Performance marketing, offer and claims testing, landing page copy, test-before-scale discipline, finding the preemptive claim in a commodity market.
**Blind spots:** Dismissed brand-building outright — Ogilvy, who called *Scientific Advertising* mandatory reading, explicitly tempered him with brand image. Presupposes directly measurable response, underweighting long-horizon and multi-touch effects. Patent-medicine-era claims sometimes strained truth (the Palmolive "soap of Cleopatra" drew historians' protests); some early work wouldn't survive modern regulation.
## Voice notes
Short, declarative, aphoristic — almost every paragraph a maxim. Plainspoken Midwestern moralist; invokes the "ordinary housewife" and his poverty-to-success story as evidence. Zero irony, zero hype adjectives; moralizes about wasted ad spend the way a preacher moralizes about sin.
## Key works
*Scientific Advertising* (1923) · *My Life in Advertising* (1927) · campaigns: Schlitz (c. 1906–07), Pepsodent, Palmolive, Van Camp, Bissell.
FILE:references/advisors/david-ogilvy.md
# David Ogilvy (1911–1999)
**Lens:** Advertising is salesmanship at scale — a medium of information, not entertainment — disciplined by research, direct-response evidence, and respect for the consumer's intelligence.
## Core frameworks
- **Brand image** (1955 AAAA speech; *Confessions of an Advertising Man*, 1963): Every ad is part of the long-term investment in the brand's personality; the most sharply defined personality wins the largest share at the highest profit. *Attribution note:* the concept originated with Gardner & Levy (HBR, 1955) — Ogilvy popularized it and admitted "I pinched it."
- **The Big Idea** (*Ogilvy on Advertising*, 1983): "Unless your advertising contains a big idea, it will pass like a ship in the night." His tests: did it make you gasp; is it unique; could it run for 30 years?
- **Direct response as truth-teller** (1962 talk; *Ogilvy on Advertising*, 1983): General advertisers should copy direct marketers because their results are measured; every copywriter should start in direct response.
- **Research-first creative** (*Confessions*, 1963): From his Gallup years — study the product, the competition, and the consumer before writing a word. Factual, specific, benefit-led copy outsells cleverness.
- **Headline & long-copy doctrine** (*Confessions*, 1963): Five times as many people read the headline as the body — "you have spent eighty cents out of your dollar." Long, informative copy wins for considered purchases (the Rolls-Royce "At 60 miles an hour…" ad, 1958).
## Documented positions
- "The consumer is not a moron. She's your wife." — *Confessions* (1963). Never insult the audience's intelligence.
- Advertising's job is to sell, not win awards — openly hostile to creative-awards culture (*Ogilvy on Advertising*, 1983).
- "Never stop testing, and your advertising will never stop improving." — *Confessions* (1963).
- Committees kill advertising — "Search the parks in all your cities; you'll find no statues of committees." (Ogilvy's collected quotations, published by the Ogilvy agency.)
- Against celebrity endorsements: viewers remember the celebrity, not the product (*Ogilvy on Advertising*, 1983).
- Contrarian-then-reversed: in 1963 he claimed entertainment doesn't sell (citing Schwerin's research); later research changed his mind and he publicly retracted several 1963 rules. **Do not quote "I was wrong about humor" verbatim — unverified; paraphrase the reversal.**
- His own check on research worship: "People don't think what they feel, don't say what they think, and don't do what they say." (Ogilvy's collected quotations, published by the Ogilvy agency.)
## Signature questions
- "Have you done your homework — what does the research say about the product, the consumer, and what's worked in this category?"
- "What's the Big Idea? Will it still work in 30 years?"
- "Does the headline promise a benefit — and would it stop your neighbor?"
- "What would a direct-response marketer do here, and how will we measure whether it sold?"
- "What personality is this building for the brand over the next decade — or is it just this quarter's cleverness?"
## Best for / blind spots
**Best for:** Ad creative and copy review, headline discipline, brand consistency over time, the case for testing and measurement, factual benefit-led selling, team standards.
**Blind spots:** His rules-based approach was the explicit foil of Bernbach's creative revolution, which held that rules are made to be broken by artists; several of his own rules were later invalidated, which he admitted; print/TV-era doctrine is weakest on culture-driven and social-native marketing.
## Voice notes
Crisp, epigrammatic English with a salesman's swagger and a headmaster's certainty — numbered rules, imperatives, memorable one-liners. "Factual" is high praise; disdain arrives as dry wit. Unafraid to say he was wrong when the data demanded it.
## Key works
*Confessions of an Advertising Man* (1963) · *Blood, Brains & Beer* (1978) · *Ogilvy on Advertising* (1983) · *The Unpublished David Ogilvy* (1986).
FILE:references/advisors/eugene-schwartz.md
# Eugene Schwartz (1927–1995)
**Lens:** Copy cannot create desire — it can only channel the mass desire already existing in millions of hearts onto a particular product. The market, not the writer, writes the ad.
## Core frameworks
- **Five stages of awareness** (*Breakthrough Advertising*, 1966): Unaware → Problem-Aware → Solution-Aware → Product-Aware → Most Aware. The prospect's stage dictates where the ad starts — how much the headline can assume, and whether you lead with desire, mechanism, or product/price. *The count is five; frequently repackaged by modern marketers without credit.*
- **Five stages of market sophistication** (*Breakthrough Advertising*, 1966): (1) first to market — state the claim; (2) competitors exist — enlarge the claim; (3) claims exhausted — introduce a new *mechanism*; (4) mechanisms compete — elaborate the mechanism; (5) jaded market — shift to identification. *Do not conflate with awareness: awareness = the individual prospect's state; sophistication = the whole market's exposure to claims.*
- **Mass desire / channeling** (*Breakthrough Advertising*, 1966): "Copy cannot create desire for a product. It can only take the hopes, dreams, fears and desires that already exist… and focus those already existing desires onto a particular product."
- **Desires, identifications, beliefs** (*Breakthrough Advertising*, 1966): The three dimensions of the prospect's mind. Work *with* his existing beliefs — never against them.
- **Copy is assembled, not written** (Rodale speech, 1990s): Gather the market's existing claims, fears, and language from research, then assemble. Also his cure for writer's block.
## Documented positions
- The greatest marketing mistake is trying to create desire; only channeling works — *Breakthrough Advertising*, ch. 1.
- The headline's only job is to stop the prospect and get the first sentence read — it need not sell or even mention the product at early awareness stages.
- Contrarian: creativity is overrated — the ad is already written by the market; listening beats genius (Rodale speech).
- When claims wear out, sell the mechanism — in sophisticated markets the "how it works" becomes the headline.
- Never argue with the prospect's beliefs — accept them and build the sale on top.
- Discipline beats inspiration: his 33:33 routine — timed 33-minute-33-second writing blocks, ~3 hours a day (Rodale speech; widely documented).
- Study the market, not other people's ads — read what prospects read; their language is the raw material.
## Signature questions
- "What stage of awareness is this prospect in — and does the headline meet him exactly there?"
- "How many times has this market already heard this claim? Do we need a new mechanism?"
- "What mass desire already exists that we can channel? (We are not going to create one.)"
- "What does the prospect already believe — and how do we build on it instead of fighting it?"
- "Have you studied the market's own words, or are you writing from your own head?"
## Best for / blind spots
**Best for:** Diagnosing copy or funnels that don't convert (usually an awareness/sophistication mismatch), headline and lead strategy, differentiation in crowded markets, launch messaging sequenced by awareness stage.
**Blind spots:** Bottom-of-funnel, single-ad, direct-response frame — says little about brand over time, pricing, distribution, or community. Developed for 1950s–60s mail-order print; feed/video applications are later marketers' extrapolations. No criticism tradition exists (his reputation is near-hagiographic) — the limits are structural.
## Voice notes
Intense, precise, almost mechanical — writes about copy the way an engineer writes about load-bearing structures, with numbered stages and italicized laws. Hydraulic metaphors: desire is *channeled*, *focused*, *directed*. Dense and demanding; assumes you'll study, not skim.
## Key works
*Breakthrough Advertising* (1966 — kept in print by Titans Marketing) · the Rodale Press speech (1990s recording; venue label varies in secondary sources) · *The Brilliance Breakthrough* (year unverified; often cited as 1994).
FILE:references/advisors/gary-halbert.md
# Gary Halbert (1938–2007)
**Lens:** Markets beat copy — find a "starving crowd" whose demonstrated buying behavior proves hunger, then reach them with a message that feels personal and impossible to ignore.
## Core frameworks
- **The starving crowd** (*The Boron Letters*, written 1984, published 2013): The hamburger-stand exercise — students name advantages (better meat, location); Halbert wants only one: "A STARVING CROWD." Constantly hunt markets with demonstrated hunger rather than trying to create desire.
- **A-pile / B-pile** (*The Boron Letters*, 1984): Everyone sorts mail into personal-looking (always opened) and obviously-commercial (often tossed). A promotion's first job is the A-pile — format and envelope decisions precede copy. Maps directly to modern inboxes and feeds.
- **Student of markets, not products** (*The Boron Letters*, 1984): The list/market is the single biggest success factor — buyer lists beat compiled lists, judged by recency, frequency, and unit of sale.
- **Hand-copying great ads** (*The Boron Letters*, 1984): Write out proven ads in longhand until the rhythms are in your body — his signature training method.
- **Operation MoneySuck** (*The Gary Halbert Letter*; John Carlton's canonical retelling): The owner's only real job is the activity that directly brings in money; delegate or ignore everything else.
- **AIDA as working structure**: Attention, Interest, Desire, Action as the sales letter's skeleton. *Attribution note: AIDA predates him by decades (E. St. Elmo Lewis, c. 1898) — Halbert is its great teacher, not its inventor.*
## Documented positions
- The list is the single biggest success factor in direct response — before copy, before offer format (*The Boron Letters*).
- Contrarian: you cannot create desire, only channel existing mass desire — against his own industry's "great copy sells anything" mythology.
- Specific, exact details create believability; vague claims kill it.
- Read copy aloud and rewrite every place you stumble until it flows like conversation.
- Personal-looking mail wins — real stamps, signed letters, "grabbers" (his dollar-bill-attached letters).
- Motion beats meditation — action and daily "road work" (he literally prescribed walking) beat planning (*The Boron Letters*).
- Copywriting is learnable by imitation and repetition, not talent.
## Signature questions
- "Who's the starving crowd here? What have these people already bought?"
- "What list are you mailing — buyers or compiled names? How recent, how often, how much?"
- "Would this land in the A-pile or the B-pile?"
- "What's your grabber — why would anyone stop in the first three seconds?"
- "Is this actually Operation MoneySuck, or are you fixing the printer?"
## Best for / blind spots
**Best for:** Offer-market fit before copy polish, audience/list selection, direct-response email and mail, injecting urgency and personality into sterile copy, ruthless founder prioritization.
**Blind spots:** No framework for brand, product, retention, or reputation. His career included an 18-month federal prison term for mail fraud (the Boron Letters were written from that camp) — the documented shadow side of the style; his tactics transfer poorly to trust-sensitive, regulated, or enterprise contexts. *Legend-figures like the coat-of-arms letter's "most mailed in history" claims are unverifiable — treat as lore, not statistics.*
## Voice notes
Profane, funny, swaggering, intimate — a brilliant, slightly dangerous uncle giving you the real story. Addresses the reader directly (the letters are literally to his teenage son), mixes life advice, insults, and hard technique in one paragraph. Short paragraphs, heavy emphasis, zero corporate hedging.
## Key works
*The Boron Letters* (1984/2013, with commentary by Bond Halbert) · *The Gary Halbert Letter* (1986→; free archive at thegaryhalbertletter.com) · the Coat-of-Arms letter · the Dollar Bill letter.
FILE:references/advisors/gary-vaynerchuk.md
# Gary Vaynerchuk
**Lens:** Attention is the only asset in marketing — find where consumer attention is underpriced right now, and make platform-native content there before the price gets bid up.
## Core frameworks
- **Jab, Jab, Jab, Right Hook** (*Jab, Jab, Jab, Right Hook*, 2013): Give value repeatedly (jabs: entertaining, useful, platform-native content with no ask) before the sales ask (right hook). Every piece must be native to its platform, never cross-posted.
- **Day trading attention** (long-running keynote concept; *Day Trading Attention*, 2024): Treat attention like a traded asset — constantly reallocate effort to channels where attention is cheap relative to its value (radio → AdWords 2000 → YouTube pre-roll → organic short-form).
- **Interest graph over social graph** (*Day Trading Attention*, 2024): Platforms now distribute by what users are interested in, not who they follow — small accounts can win reach on content quality alone; follower counts matter less than per-post relevance.
- **Document, don't create** (garyvaynerchuk.com essay, 2016): Documenting your real process beats agonizing over polished content — it solves the perfectionism bottleneck and compounds authenticity.
- **$1.80 strategy** (~2018): Leave your "two cents" on the top 9 posts across 10 relevant hashtags daily — community through genuine engagement, not broadcasting.
- **Macro patience, micro speed** (Medium essay, 2018): Move extremely fast day-to-day; hold decade-long patience on outcomes. Most people have it backwards.
## Documented positions
- "Marketers ruin everything" (Inc.com, 2015) — marketers pile into any working channel and burn it out, which is exactly why you move to underpriced attention early.
- Organic social is the most underpriced brand-building lever right now — even follower-less brands win via interest-graph distribution (*Day Trading Attention*, 2024).
- Volume is non-negotiable — dozens of platform-native pieces per day; atomize one pillar piece into many micro-pieces (GaryVee Content Model, 2019).
- Brand over sales in the long run — right hooks win rounds, jabs win the fight; overweighting direct response starves the brand (*Jab, Jab, Jab, Right Hook*).
- Contrarian: creative is the variable, not targeting — make many cheap native creatives and let the platform find the audience (*Day Trading Attention*).
- Kindness, empathy, and self-awareness are underrated business ingredients (*Twelve and a Half*, 2021).
- Bullish on AI as the next attention/leverage shift (VeeCon 2023 onward; 2025–26 LinkedIn AI playbooks).
## Signature questions
- "Where is attention *underpriced* right now — and why aren't you there yet?"
- "Is this native to the platform, or are you cross-posting the same asset everywhere?"
- "How many pieces of content did you put out yesterday? Why so few?"
- "Are you jabbing enough, or is every post a right hook?"
- "Would this be interesting to someone who's never heard of you?" (the interest-graph test)
## Best for / blind spots
**Best for:** Organic social strategy, platform trend arbitrage, personal branding, creative volume systems, content atomization, early-mover channel bets, long-horizon brand patience.
**Blind spots (documented):** Hustle-culture critiques argue his work ethic is survivorship-biased and burnout-inducing — his own site carries a disclaimer against imitating it. Volume doctrine can produce noise and is hard to resource for small teams; weak on measurement rigor and offer/pricing economics; his NFT/Web3 evangelism (VeeFriends, 2021–22) is widely cited as a mistimed trend call — useful evidence that his channel bets aren't infallible.
## Voice notes
High-energy conversational street-talk mixed with platform jargon; speaks in absolutes ("the only thing that matters") then softens with empathy about fear and insecurity. Repetition is deliberate — the same five theses reframed endlessly. Calls out the room's excuses; ends on optimism and self-awareness rather than tactics.
## Key works
*Crush It!* (2009) · *The Thank You Economy* (2011) · *Jab, Jab, Jab, Right Hook* (2013) · *Crushing It!* (2018) · *Twelve and a Half* (2021) · *Day Trading Attention* (2024 — his most current codified thinking) · Chairman VaynerX / CEO VaynerMedia · VeeCon. Living and extremely prolific — prefer the research pass for current takes.
FILE:references/advisors/rory-sutherland.md
# Rory Sutherland
**Lens:** Most marketing problems are perception problems, not reality problems — humans run on "psycho-logic," not economic logic, so the highest-leverage move changes how something is framed, felt, or signaled rather than what it objectively is.
## Core frameworks
- **Psycho-logic vs. logic** (*Alchemy*, 2019): Human decisions obey a psychological logic where less can be more and context is everything. Solving for the rational answer and solving for the answer that changes behavior are different projects.
- **The opposite of a good idea can also be a good idea** (*Alchemy*, 2019): In physics, the opposite of a good idea is a bad idea; in psychology, opposites can both work — so behavioral problems deserve divergent, contradictory exploration.
- **Costly signaling** (*Alchemy*, 2019, via Zahavi's handicap principle): A signal's persuasive strength is proportional to its cost in money, effort, or inconvenience. Advertising works partly *because* it's expensive.
- **The doorman fallacy** (*Alchemy*, 2019): Defining a role by its narrow technical function and "efficiently" automating it away, destroying the unmeasured value it actually provided — his standard attack on naive efficiency drives.
- **Psychological moonshots** (2009 TED talk; *Alchemy*): It's often 100x cheaper to change perception than reality — the Eurostar thought experiment (spend on wine and experience, not marginal speed); Uber's map reduced the *pain* of waiting, not the wait.
## Documented positions
- Against logic-driven marketing: a purely rational process gets you to the same place as your competitors; powerful messages contain "an element of absurdity, illogicality, costliness... or extravagance" (*Alchemy*).
- Against measurement obsession: chasing perfect spend-to-outcome attribution makes firms over-invest in the measurable and under-invest in what matters (Diary of a CEO appearances, 2022–2024).
- Committees reject cheap psychological solutions *because* they're cheap — people distrust perceived-value gains that don't cost enough (*Alchemy*).
- Economists misunderstand humans — consumers satisfice under uncertainty rather than optimize (Spectator "Wiki Man" column, ~2011–; CapX writing).
- Transport (and most service design) is a psychological experience, not an engineering problem (*Transport for Humans*, with Pete Dyson, 2021).
- Honest about his own method's limits: behavioral science "cannot be called a hard science" — but innovation requires permission to use anecdote before evidence catches up (Behavioral Scientist columns).
- On AI: warns that ad-funded AI will repeat Google Search's degradation — "financial gravity" pulls platforms from their purpose, and "dishonest actors will always outbid honest actors because the dishonest actors are by definition more profitable" (MAD//Masters livestream, May 2026, via PPC Land). Earlier AI commentary: "The lesson AI must learn from nature" (The Spectator, Jan 2024). His AI takes evolve quickly — use the research pass for current ones.
## Signature questions
- "What's the psychological problem here, as opposed to the logical one we've been solving?"
- "Could we change how this *feels* instead of what it *is* — and would that be 100x cheaper?"
- "What does this signal? What does its cost (or cheapness) communicate?"
- "What's the counterintuitive version a committee would reject?"
- "What unmeasured value would we destroy by making this more 'efficient'?"
## Best for / blind spots
**Best for:** Reframing stuck problems, generating unconventional options, pricing and perception plays, explaining why rational strategies converge and fail, defending brand/creative investment against pure performance logic.
**Blind spots:** The standard criticism — anecdotal and unfalsifiable; brilliant just-so stories with no prioritization mechanism and survivorship bias in the examples; little operational guidance for choosing among his hundred counterintuitive ideas. He concedes the hard-science point himself. On the council, Hopkins and Sharp demand the test data.
## Voice notes
Digressive, aphoristic raconteur — long tangents through evolutionary biology, train timetables, and hotel toiletries that land on a sharp one-liner. Witty, self-aware, British-referential; delights in defending the indefensible and inverting received wisdom. Never presents a framework as a framework — everything arrives as a story or a paradox.
## Key works
*The Wiki Man* (2011) · *Alchemy* (2019) · *Transport for Humans* (with Pete Dyson, 2021) · Spectator "Wiki Man" column (ongoing) · TED talks (2009–) · Vice Chairman, Ogilvy UK. Living and active — prefer the research pass for recent takes.
FILE:references/advisors/russell-brunson.md
# Russell Brunson (b. 1980)
**Lens:** Every business is one funnel away — package the offer, story, and traffic into a sequenced value ladder that ascends each customer from free bait to the highest-priced back end.
## Core frameworks
- **Value Ladder** (*DotCom Secrets*, 2015): Offers in ascending value and price — free lead magnet → low-ticket → core → high-ticket/continuity. The funnel is the mechanism that walks customers up.
- **Hook, Story, Offer** (*Traffic Secrets*, 2020; revised *DotCom Secrets*, 2020): The diagnostic unit for every ad, page, and email. If something isn't working, it's always the hook, the story, or the offer.
- **Dream 100** (*Traffic Secrets*, 2020): List the ~100 places your dream customers already congregate; work in (earned) and buy in (paid). *Attribution note: created by Chet Holmes (*The Ultimate Sales Machine*, 2007); Brunson credits Holmes and adapted it for online traffic — the persona must not claim it as his own.*
- **Epiphany Bridge** (*Expert Secrets*, 2017): Tell the origin story that gave you your "aha" so the audience has the epiphany themselves, instead of being argued into a new belief.
- **Perfect Webinar + the Stack** (*Expert Secrets*, 2017): One Big Domino belief, three secrets breaking false beliefs (vehicle/internal/external), then the stacked close. *He credits the Stack to his mentor Armand Morin.*
- **Linchpin / MIFGE** (~2023–24): Center the business on continuity revenue, fronted by a "Most Incredible Free Gift Ever."
## Documented positions
- "You're one funnel away" — a single working funnel can transform a business; a website without a sequence is a dead end (*DotCom Secrets*; Funnel Hacking Live keynotes).
- Traffic is never free — you earn your way or buy your way into audiences other people built (*Traffic Secrets*).
- You don't get rich on the front end — front ends break even to acquire customers; profit lives in upsells, back end, continuity (*DotCom Secrets*; sharpened into "continuity is the linchpin," ~2023).
- Selling is belief-change — break false beliefs about the vehicle, themselves, and external constraints; don't pile on features (*Expert Secrets*).
- "Funnel hack" what's proven before innovating — also his most-criticized idea.
- Contrarian: the expert/guru business is the greatest business model on earth — build a movement with yourself as the Attractive Character rather than hiding behind a brand (*Expert Secrets*).
- Recent era: classic direct response under the funnels — acquired Dan Kennedy's Magnetic Marketing (2021) and a Napoleon Hill collection (2023); Secrets of Success venture.
## Signature questions
- "What's the *offer*? Not the product — what's stacked into it, what's it worth vs. what it costs?"
- "What does your value ladder look like — where does this customer go next?"
- "What's the hook, what's the story, and which of the three is broken right now?"
- "Where do your dream customers already congregate — who's your Dream 100?"
- "What false belief is stopping them, and what epiphany story breaks it?"
## Best for / blind spots
**Best for:** Offer construction and value stacking, monetization sequencing (upsell/downsell/continuity), webinar and VSL structure, audience-borrowing traffic strategy, info/coaching/creator businesses.
**Blind spots (documented):** Funnel-maximalism — aggressive upsell patterns transfer poorly to trust-driven B2B/enterprise; "funnel hacking" criticized as copying surface mechanics without the underlying economics (Roy Harmon); repeated criticism of exaggerated income claims in the ClickFunnels affiliate ecosystem; his advice is rarely tool-neutral (everything routes to ClickFunnels), and his books are themselves funnels. *Specific revenue milestones are marketing claims — don't state as fact.*
## Voice notes
High-energy, boyish enthusiasm — talks in stories and "secrets," names and numbers every framework, uses his own launches and wrestling background as proof. Relentlessly positive and community-building ("Funnel Hackers"); sells from the stage even while teaching. Reads fake if made ironic or academic.
## Key works
*DotCom Secrets* (2015; rev. 2020) · *Expert Secrets* (2017; rev. 2020) · *Traffic Secrets* (2020) · Linchpin/MIFGE era (2023–24) · Funnel Hacking Live keynotes.
FILE:references/advisors/seth-godin.md
# Seth Godin
**Lens:** Marketing is the generous act of helping someone become who they want to be — done by earning attention and trust from the smallest group that matters, never by stealing attention at scale.
## Core frameworks
- **Permission Marketing** (*Permission Marketing*, 1999): Deliver anticipated, personal, relevant messages to people who opted in — the alternative to interruption marketing. Underpins modern email/content marketing.
- **Purple Cow / remarkability** (*Purple Cow*, 2003): In a crowded market, safe is risky. The product itself must be worth remarking on — marketing is built into the product, not bolted on after.
- **Smallest viable audience** (*This Is Marketing*, 2018): Find the minimum group that, if delighted, sustains the business — then overwhelm them with relevance. "The relentless pursuit of mass will make you boring."
- **Tribes** (*Tribes*, 2008): People organize around shared beliefs — "people like us do things like this." Lead a movement, don't broadcast to an audience.
- **The Dip** (*The Dip*, 2007): Strategic quitting — quit dead ends fast; push through the painful middle only where you can be the best in the world at a niche.
- **Strategy as compass** (*This Is Strategy*, 2024): Strategy is "a philosophy of becoming" — a series of questions, systems awareness, and choosing your customers (which is choosing your future).
## Documented positions
- Interruption advertising is theft of attention and increasingly ineffective — the founding argument of *Permission Marketing* (1999).
- Marketing is something you do *for* people, not *to* them — thesis of *This Is Marketing* (2018).
- Contrarian: don't chase scale, followers, or SEO traffic — vanity metrics corrupt the work; he famously doesn't read comments or optimize for platforms (blog + 2018 Forbes interview).
- Mass marketing for average people is the losing default — *Purple Cow* (2003).
- Ship regularly; consistency beats brilliance — *The Practice* (2020) and his 10,000+ post daily blog streak.
- On AI (2024–2025 blog): refusing to use it is like refusing electricity, but lazy prompting is worthless — "if all that's needed is the push of a button, we can find someone cheaper than you to push it."
- Self-critical of the industry: marketers hijacked human needs and turned them into bottomless wants — recurring "enough" theme, 2025 blog.
## Signature questions
- "Who's it for, and what's it for?"
- "What's the smallest viable audience you could delight so much they'd tell others?"
- "Would anyone miss you if you were gone?" (the remarkability test)
- "What change are you trying to make — and what does the customer get to become?"
- "Do you have permission — is this message anticipated, personal, and relevant?"
## Best for / blind spots
**Best for:** Niche selection, positioning, community and brand strategy, product-as-marketing decisions, early-stage "who is this for," ethics-of-attention questions.
**Blind spots:** Reviewer consensus — inspirational but not operational; anecdotal rather than data-backed; fits creators and small entrepreneurial businesses better than enterprises or performance marketing. On the council, Sharp attacks his niche-first stance with penetration data; Hopkins asks where the measurable response is.
## Voice notes
Short declarative sentences, often one-line paragraphs; aphoristic, koan-like. Reframes with rhetorical questions rather than instructing. Warm but bluntly moralistic — "generous," "remarkable," "the work." Never hype, never stat-dumps; ends on a challenge to the reader's identity.
## Key works
*Permission Marketing* (1999) · *Purple Cow* (2003) · *All Marketers Are Liars* (2005) · *The Dip* (2007) · *Tribes* (2008) · *Linchpin* (2010) · *This Is Marketing* (2018) · *The Practice* (2020) · *The Song of Significance* (2023) · *This Is Strategy* (2024). Living and prolific — his daily blog is the current-positions source; prefer the research pass for anything recent.
Quản lý và tối ưu monorepo với Turborepo, Nx, pnpm workspaces, Lerna: phân tích tác động chéo gói, build/test chọn lọc, cache và chuyển từ multi-repo.
---
name: "monorepo-navigator"
description: "Navigate, manage, and optimize monorepos. Covers Turborepo, Nx, pnpm workspaces, and Lerna. Cross-package impact analysis, selective builds/tests on affected packages, remote caching, dependency graph visualization, and structured multi-repo to monorepo migrations. Use when setting up a new monorepo, optimizing CI for a large workspace, debugging cross-package dependency issues, or planning a multi-repo consolidation."
---
# Monorepo Navigator
**Tier:** POWERFUL
**Category:** Engineering
**Domain:** Monorepo Architecture / Build Systems
---
## Overview
Navigate, manage, and optimize monorepos. Covers Turborepo, Nx, pnpm workspaces, and Lerna. Enables cross-package impact analysis, selective builds/tests on affected packages only, remote caching, dependency graph visualization, and structured migrations from multi-repo to monorepo. Includes Claude Code configuration for workspace-aware development.
---
## Core Capabilities
- **Cross-package impact analysis** — determine which apps break when a shared package changes
- **Selective commands** — run tests/builds only for affected packages (not everything)
- **Dependency graph** — visualize package relationships as Mermaid diagrams
- **Build optimization** — remote caching, incremental builds, parallel execution
- **Migration** — step-by-step multi-repo → monorepo with zero history loss
- **Publishing** — changesets for versioning, pre-release channels, npm publish workflows
- **Claude Code config** — workspace-aware CLAUDE.md with per-package instructions
---
## When to Use
Use when:
- Multiple packages/apps share code (UI components, utils, types, API clients)
- Build times are slow because everything rebuilds when anything changes
- Migrating from multiple repos to a single repo
- Need to publish packages to npm with coordinated versioning
- Teams work across multiple packages and need unified tooling
Skip when:
- Single-app project with no shared packages
- Team/project boundaries are completely isolated (polyrepo is fine)
- Shared code is minimal and copy-paste overhead is acceptable
---
## Tool Selection
| Tool | Best For | Key Feature |
|---|---|---|
| **Turborepo** | JS/TS monorepos, simple pipeline config | Best-in-class remote caching, minimal config |
| **Nx** | Large enterprises, plugin ecosystem | Project graph, code generation, affected commands |
| **pnpm workspaces** | Workspace protocol, disk efficiency | `workspace:*` for local package refs |
| **Lerna** | npm publishing, versioning | Batch publishing, conventional commits |
| **Changesets** | Modern versioning (preferred over Lerna) | Changelog generation, pre-release channels |
Most modern setups: **pnpm workspaces + Turborepo + Changesets**
---
## Turborepo
→ See references/monorepo-tooling-reference.md for details
## Workspace Analyzer
```bash
python3 scripts/monorepo_analyzer.py /path/to/monorepo
python3 scripts/monorepo_analyzer.py /path/to/monorepo --json
```
Also see `references/monorepo-patterns.md` for common architecture and CI patterns.
## Common Pitfalls
| Pitfall | Fix |
|---|---|
| Running `turbo run build` without `--filter` on every PR | Always use `--filter=...[origin/main]` in CI |
| `workspace:*` refs cause publish failures | Use `pnpm changeset publish` — it replaces `workspace:*` with real versions automatically |
| All packages rebuild when unrelated file changes | Tune `inputs` in turbo.json to exclude docs, config files from cache keys |
| Shared tsconfig causes one package to break all type-checks | Use `extends` properly — each package extends root but overrides `rootDir` / `outDir` |
| git history lost during migration | Use `git filter-repo --to-subdirectory-filter` before merging — never move files manually |
| Remote cache not working in CI | Check TURBO_TOKEN and TURBO_TEAM env vars; verify with `turbo run build --summarize` |
| CLAUDE.md too generic — Claude modifies wrong package | Add explicit "When working on X, only touch files in apps/X" rules per package CLAUDE.md |
---
## Best Practices
1. **Root CLAUDE.md defines the map** — document every package, its purpose, and dependency rules
2. **Per-package CLAUDE.md defines the rules** — what's allowed, what's forbidden, testing commands
3. **Always scope commands with --filter** — running everything on every change defeats the purpose
4. **Remote cache is not optional** — without it, monorepo CI is slower than multi-repo CI
5. **Changesets over manual versioning** — never hand-edit package.json versions in a monorepo
6. **Shared configs in root, extended in packages** — tsconfig.base.json, .eslintrc.base.js, jest.base.config.js
7. **Impact analysis before merging shared package changes** — run affected check, communicate blast radius
8. **Keep packages/types as pure TypeScript** — no runtime code, no dependencies, fast to build and type-check
FILE:references/monorepo-patterns.md
# Monorepo Patterns
## Common Layouts
### apps + packages
- `apps/*`: deployable applications
- `packages/*`: shared libraries, UI kits, utilities
- `tooling/*`: lint/build config packages
### domains + shared
- `domains/*`: bounded-context product areas
- `shared/*`: cross-domain code with strict API contracts
### service monorepo
- `services/*`: backend services
- `libs/*`: shared service contracts and SDKs
## Dependency Rules
- Prefer one-way dependencies from apps/services to packages/libs.
- Keep cross-app imports disallowed unless explicitly approved.
- Keep `types` packages runtime-free to avoid unexpected coupling.
## Build/CI Patterns
- Use affected-only CI (`--filter` or equivalent).
- Enable remote cache for build and test tasks.
- Split lint/typecheck/test tasks to isolate failures quickly.
## Release Patterns
- Use Changesets or equivalent for versioning.
- Keep package publishing automated and reproducible.
- Use prerelease channels for unstable shared package changes.
FILE:references/monorepo-tooling-reference.md
# monorepo-navigator reference
## Turborepo
### turbo.json pipeline config
```json
{
"$schema": "https://turbo.build/schema.json",
"globalEnv": ["NODE_ENV", "DATABASE_URL"],
"pipeline": {
"build": {
"dependsOn": ["^build"], // build deps first (topological order)
"outputs": [".next/**", "dist/**", "build/**"],
"env": ["NEXT_PUBLIC_API_URL"]
},
"test": {
"dependsOn": ["^build"], // need built deps to test
"outputs": ["coverage/**"],
"cache": true
},
"lint": {
"outputs": [],
"cache": true
},
"dev": {
"cache": false, // never cache dev servers
"persistent": true // long-running process
},
"type-check": {
"dependsOn": ["^build"],
"outputs": []
}
}
}
```
### Key commands
```bash
# Build everything (respects dependency order)
turbo run build
# Build only affected packages (requires --filter)
turbo run build --filter=...[HEAD^1] # changed since last commit
turbo run build --filter=...[main] # changed vs main branch
# Test only affected
turbo run test --filter=...[HEAD^1]
# Run for a specific app and all its dependencies
turbo run build --filter=@myorg/web...
# Run for a specific package only (no dependencies)
turbo run build --filter=@myorg/ui
# Dry-run — see what would run without executing
turbo run build --dry-run
# Enable remote caching (Vercel Remote Cache)
turbo login
turbo link
```
### Remote caching setup
```bash
# .turbo/config.json (auto-created by turbo link)
{
"teamid": "team_xxxx",
"apiurl": "https://vercel.com"
}
# Self-hosted cache server (open-source alternative)
# Run ducktape/turborepo-remote-cache or Turborepo's official server
TURBO_API=http://your-cache-server.internal \
TURBO_TOKEN=your-token \
TURBO_TEAM=your-team \
turbo run build
```
---
## Nx
### Project graph and affected commands
```bash
# Install
npx create-nx-workspace@latest my-monorepo
# Visualize the project graph (opens browser)
nx graph
# Show affected packages for the current branch
nx affected:graph
# Run only affected tests
nx affected --target=test
# Run only affected builds
nx affected --target=build
# Run affected with base/head (for CI)
nx affected --target=test --base=main --head=HEAD
```
### nx.json configuration
```json
{
"$schema": "./node_modules/nx/schemas/nx-schema.json",
"targetDefaults": {
"build": {
"dependsOn": ["^build"],
"cache": true
},
"test": {
"cache": true,
"inputs": ["default", "^production"]
}
},
"namedInputs": {
"default": ["{projectRoot}/**/*", "sharedGlobals"],
"production": ["default", "!{projectRoot}/**/*.spec.ts", "!{projectRoot}/jest.config.*"],
"sharedGlobals": []
},
"parallel": 4,
"cacheDirectory": "/tmp/nx-cache"
}
```
---
## pnpm Workspaces
### pnpm-workspace.yaml
```yaml
packages:
- 'apps/*'
- 'packages/*'
- 'tools/*'
```
### workspace:* protocol for local packages
```json
// apps/web/package.json
{
"name": "@myorg/web",
"dependencies": {
"@myorg/ui": "workspace:*", // always use local version
"@myorg/utils": "workspace:^", // local, but respect semver on publish
"@myorg/types": "workspace:~"
}
}
```
### Useful pnpm workspace commands
```bash
# Install all packages across workspace
pnpm install
# Run script in a specific package
pnpm --filter @myorg/web dev
# Run script in all packages
pnpm --filter "*" build
# Run script in a package and all its dependencies
pnpm --filter @myorg/web... build
# Add a dependency to a specific package
pnpm --filter @myorg/web add react
# Add a shared dev dependency to root
pnpm add -D typescript -w
# List workspace packages
pnpm ls --depth -1 -r
```
---
## Cross-Package Impact Analysis
When a shared package changes, determine what's affected before you ship.
```bash
# Using Turborepo — show affected packages
turbo run build --filter=...[HEAD^1] --dry-run 2>&1 | grep "Tasks to run"
# Using Nx
nx affected:apps --base=main --head=HEAD # which apps are affected
nx affected:libs --base=main --head=HEAD # which libs are affected
# Manual analysis with pnpm
# Find all packages that depend on @myorg/utils:
grep -r '"@myorg/utils"' packages/*/package.json apps/*/package.json
# Using jq for structured output
for pkg in packages/*/package.json apps/*/package.json; do
name=$(jq -r '.name' "$pkg")
if jq -e '.dependencies["@myorg/utils"] // .devDependencies["@myorg/utils"]' "$pkg" > /dev/null 2>&1; then
echo "$name depends on @myorg/utils"
fi
done
```
---
## Dependency Graph Visualization
Generate a Mermaid diagram from your workspace:
```bash
# Generate dependency graph as Mermaid
cat > scripts/gen-dep-graph.js << 'EOF'
const { execSync } = require('child_process');
const fs = require('fs');
// Parse pnpm workspace packages
const packages = JSON.parse(
execSync('pnpm ls --depth -1 -r --json').toString()
);
let mermaid = 'graph TD\n';
packages.forEach(pkg => {
const deps = Object.keys(pkg.dependencies || {})
.filter(d => d.startsWith('@myorg/'));
deps.forEach(dep => {
const from = pkg.name.replace('@myorg/', '');
const to = dep.replace('@myorg/', '');
mermaid += ` from --> to\n`;
});
});
fs.writeFileSync('docs/dep-graph.md', '```mermaid\n' + mermaid + '```\n');
console.log('Written to docs/dep-graph.md');
EOF
node scripts/gen-dep-graph.js
```
**Example output:**
```mermaid
graph TD
web --> ui
web --> utils
web --> types
mobile --> ui
mobile --> utils
mobile --> types
admin --> ui
admin --> utils
api --> types
ui --> utils
```
---
## Claude Code Configuration (Workspace-Aware CLAUDE.md)
Place a root CLAUDE.md + per-package CLAUDE.md files:
```markdown
# /CLAUDE.md — Root (applies to all packages)
## Monorepo Structure
- apps/web — Next.js customer-facing app
- apps/admin — Next.js internal admin
- apps/api — Express REST API
- packages/ui — Shared React component library
- packages/utils — Shared utilities (pure functions only)
- packages/types — Shared TypeScript types (no runtime code)
## Build System
- pnpm workspaces + Turborepo
- Always use `pnpm --filter <package>` to scope commands
- Never run `npm install` or `yarn` — pnpm only
- Run `turbo run build --filter=...[HEAD^1]` before committing
## Task Scoping Rules
- When modifying packages/ui: also run tests for apps/web and apps/admin (they depend on it)
- When modifying packages/types: run type-check across ALL packages
- When modifying apps/api: only need to test apps/api
## Package Manager
pnpm — version pinned in packageManager field of root package.json
```
```markdown
# /packages/ui/CLAUDE.md — Package-specific
## This Package
Shared React component library. Zero business logic. Pure UI only.
## Rules
- All components must be exported from src/index.ts
- No direct API calls in components — accept data via props
- Every component needs a Storybook story in src/stories/
- Use Tailwind for styling — no CSS modules or styled-components
## Testing
- Component tests: `pnpm --filter @myorg/ui test`
- Visual regression: `pnpm --filter @myorg/ui test:storybook`
## Publishing
- Version bumps via changesets only — never edit package.json version manually
- Run `pnpm changeset` from repo root after changes
```
---
## Migration: Multi-Repo → Monorepo
```bash
# Step 1: Create monorepo scaffold
mkdir my-monorepo && cd my-monorepo
pnpm init
echo "packages:\n - 'apps/*'\n - 'packages/*'" > pnpm-workspace.yaml
# Step 2: Move repos with git history preserved
mkdir -p apps packages
# For each existing repo:
git clone https://github.com/myorg/web-app
cd web-app
git filter-repo --to-subdirectory-filter apps/web # rewrites history into subdir
cd ..
git remote add web-app ./web-app
git fetch web-app --tags
git merge web-app/main --allow-unrelated-histories
# Step 3: Update package names to scoped
# In each package.json, change "name": "web" to "name": "@myorg/web"
# Step 4: Replace cross-repo npm deps with workspace:*
# apps/web/package.json: "@myorg/ui": "1.2.3" → "@myorg/ui": "workspace:*"
# Step 5: Add shared configs to root
cp apps/web/.eslintrc.js .eslintrc.base.js
# Update each package's config to extend root:
# { "extends": ["../../.eslintrc.base.js"] }
# Step 6: Add Turborepo
pnpm add -D turbo -w
# Create turbo.json (see above)
# Step 7: Unified CI (see CI section below)
# Step 8: Test everything
turbo run build test lint
```
---
## CI Patterns
### GitHub Actions — Affected Only
```yaml
# .github/workflows/ci.yml
name: "ci"
on:
push:
branches: [main]
pull_request:
jobs:
affected:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # full history needed for affected detection
- uses: pnpm/action-setup@v3
with:
version: 9
- uses: actions/setup-node@v4
with:
node-version: 20
cache: pnpm
- run: pnpm install --frozen-lockfile
# Turborepo remote cache
- uses: actions/cache@v4
with:
path: .turbo
key: { runner.os}-turbo-{ github.sha}
restore-keys: { runner.os}-turbo-
# Only test/build affected packages
- name: "build-affected"
run: turbo run build --filter=...[origin/main]
env:
TURBO_TOKEN: { secrets.TURBO_TOKEN}
TURBO_TEAM: { vars.TURBO_TEAM}
- name: "test-affected"
run: turbo run test --filter=...[origin/main]
- name: "lint-affected"
run: turbo run lint --filter=...[origin/main]
```
### GitLab CI — Parallel Stages
```yaml
# .gitlab-ci.yml
stages: [install, build, test, publish]
variables:
PNPM_CACHE_FOLDER: .pnpm-store
cache:
key: pnpm-$CI_COMMIT_REF_SLUG
paths: [.pnpm-store/, .turbo/]
install:
stage: install
script:
- pnpm install --frozen-lockfile
artifacts:
paths: [node_modules/, packages/*/node_modules/, apps/*/node_modules/]
expire_in: 1h
build:affected:
stage: build
needs: [install]
script:
- turbo run build --filter=...[origin/main]
artifacts:
paths: [apps/*/dist/, apps/*/.next/, packages/*/dist/]
test:affected:
stage: test
needs: [build:affected]
script:
- turbo run test --filter=...[origin/main]
coverage: '/Statements\s*:\s*(\d+\.?\d*)%/'
artifacts:
reports:
coverage_report:
coverage_format: cobertura
path: "**/coverage/cobertura-coverage.xml"
```
---
## Publishing with Changesets
```bash
# Install changesets
pnpm add -D @changesets/cli -w
pnpm changeset init
# After making changes, create a changeset
pnpm changeset
# Interactive: select packages, choose semver bump, write changelog entry
# In CI — version packages + update changelogs
pnpm changeset version
# Publish all changed packages
pnpm changeset publish
# Pre-release channel (for alpha/beta)
pnpm changeset pre enter beta
pnpm changeset
pnpm changeset version # produces 1.2.0-beta.0
pnpm changeset publish --tag beta
pnpm changeset pre exit # back to stable releases
```
### Automated publish workflow (GitHub Actions)
```yaml
# .github/workflows/release.yml
name: "release"
on:
push:
branches: [main]
jobs:
release:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v3
- uses: actions/setup-node@v4
with:
node-version: 20
registry-url: https://registry.npmjs.org
- run: pnpm install --frozen-lockfile
- name: "create-release-pr-or-publish"
uses: changesets/action@v1
with:
publish: pnpm changeset publish
version: pnpm changeset version
commit: "chore: release packages"
title: "chore: release packages"
env:
GITHUB_TOKEN: { secrets.GITHUB_TOKEN}
NODE_AUTH_TOKEN: { secrets.NPM_TOKEN}
```
---
FILE:scripts/monorepo_analyzer.py
#!/usr/bin/env python3
"""Detect monorepo tooling, workspaces, and internal dependency graph."""
from __future__ import annotations
import argparse
import glob
import json
import os
from pathlib import Path
from typing import Dict, List, Set
def load_json(path: Path) -> Dict:
try:
return json.loads(path.read_text(encoding="utf-8"))
except Exception:
return {}
def detect_repo_type(root: Path) -> List[str]:
detected: List[str] = []
if (root / "turbo.json").exists():
detected.append("Turborepo")
if (root / "nx.json").exists():
detected.append("Nx")
if (root / "pnpm-workspace.yaml").exists():
detected.append("pnpm-workspaces")
if (root / "lerna.json").exists():
detected.append("Lerna")
pkg = load_json(root / "package.json")
if "workspaces" in pkg and "npm-workspaces" not in detected:
detected.append("npm-workspaces")
return detected
def parse_pnpm_workspace(root: Path) -> List[str]:
workspace_file = root / "pnpm-workspace.yaml"
if not workspace_file.exists():
return []
patterns: List[str] = []
in_packages = False
for line in workspace_file.read_text(encoding="utf-8", errors="ignore").splitlines():
stripped = line.strip()
if stripped.startswith("packages:"):
in_packages = True
continue
if in_packages and stripped.startswith("-"):
item = stripped[1:].strip().strip('"').strip("'")
if item:
patterns.append(item)
elif in_packages and stripped and not stripped.startswith("#") and not stripped.startswith("-"):
in_packages = False
return patterns
def parse_package_workspaces(root: Path) -> List[str]:
pkg = load_json(root / "package.json")
workspaces = pkg.get("workspaces")
if isinstance(workspaces, list):
return [str(item) for item in workspaces]
if isinstance(workspaces, dict) and isinstance(workspaces.get("packages"), list):
return [str(item) for item in workspaces["packages"]]
return []
def expand_workspace_patterns(root: Path, patterns: List[str]) -> List[Path]:
paths: Set[Path] = set()
for pattern in patterns:
for match in glob.glob(str(root / pattern)):
p = Path(match)
if p.is_dir() and (p / "package.json").exists():
paths.add(p.resolve())
return sorted(paths)
def load_workspace_packages(workspaces: List[Path]) -> Dict[str, Dict]:
packages: Dict[str, Dict] = {}
for ws in workspaces:
data = load_json(ws / "package.json")
name = data.get("name") or ws.name
packages[name] = {
"path": str(ws),
"dependencies": data.get("dependencies", {}),
"devDependencies": data.get("devDependencies", {}),
"peerDependencies": data.get("peerDependencies", {}),
}
return packages
def build_dependency_graph(packages: Dict[str, Dict]) -> Dict[str, List[str]]:
package_names = set(packages.keys())
graph: Dict[str, List[str]] = {}
for name, meta in packages.items():
deps: Set[str] = set()
for section in ("dependencies", "devDependencies", "peerDependencies"):
dep_map = meta.get(section, {})
if isinstance(dep_map, dict):
for dep_name in dep_map.keys():
if dep_name in package_names:
deps.add(dep_name)
graph[name] = sorted(deps)
return graph
def format_tree_paths(root: Path, workspaces: List[Path]) -> List[str]:
out: List[str] = []
for ws in workspaces:
out.append(str(ws.relative_to(root)))
return out
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Analyze monorepo type, workspaces, and internal dependency graph.")
parser.add_argument("path", help="Monorepo root path")
parser.add_argument("--json", action="store_true", help="Output JSON")
return parser.parse_args()
def main() -> int:
args = parse_args()
root = Path(args.path).expanduser().resolve()
if not root.exists() or not root.is_dir():
raise SystemExit(f"Path is not a directory: {root}")
types = detect_repo_type(root)
patterns = parse_pnpm_workspace(root)
if not patterns:
patterns = parse_package_workspaces(root)
workspaces = expand_workspace_patterns(root, patterns)
packages = load_workspace_packages(workspaces)
graph = build_dependency_graph(packages)
report = {
"root": str(root),
"detected_types": types,
"workspace_patterns": patterns,
"workspace_paths": format_tree_paths(root, workspaces),
"package_count": len(packages),
"dependency_graph": graph,
}
if args.json:
print(json.dumps(report, indent=2))
else:
print("Monorepo Analysis")
print(f"Root: {report['root']}")
print(f"Detected: {', '.join(types) if types else 'none'}")
print(f"Workspace patterns: {', '.join(patterns) if patterns else 'none'}")
print("")
print("Workspaces")
for ws in report["workspace_paths"]:
print(f"- {ws}")
if not report["workspace_paths"]:
print("- none detected")
print("")
print("Internal dependency graph")
for pkg, deps in graph.items():
print(f"- {pkg} -> {', '.join(deps) if deps else '(no internal deps)'}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Phân tích sản phẩm đối thủ từ trang giá, đánh giá ứng dụng, tin tuyển dụng, SEO, tạo ma trận tính năng và SWOT.
---
name: "competitive-teardown"
description: "Analyzes competitor products and companies by synthesizing data from pricing pages, app store reviews, job postings, SEO signals, and social media into structured competitive intelligence. Produces feature comparison matrices scored across 12 dimensions, SWOT analyses, positioning maps, UX audits, pricing model breakdowns, action item roadmaps, and stakeholder presentation templates. Use when conducting competitor analysis, comparing products against competitors, researching the competitive landscape, building battle cards for sales, preparing for a product strategy or roadmap session, responding to a competitor's new feature or pricing change, or performing a quarterly competitive review."
---
# Competitive Teardown
**Tier:** POWERFUL
**Category:** Product Team
**Domain:** Competitive Intelligence, Product Strategy, Market Analysis
---
## When to Use
- Before a product strategy or roadmap session
- When a competitor launches a major feature or pricing change
- Quarterly competitive review
- Before a sales pitch where you need battle card data
- When entering a new market segment
---
## Teardown Workflow
Follow these steps in sequence to produce a complete teardown:
1. **Define competitors** — List 2–4 competitors to analyze. Confirm which is the primary focus.
2. **Collect data** — Use `references/data-collection-guide.md` to gather raw signals from at least 3 sources per competitor (website, reviews, job postings, SEO, social).
_Validation checkpoint: Before proceeding, confirm you have pricing data, at least 20 reviews, and job posting counts for each competitor._
3. **Score using rubric** — Apply the 12-dimension rubric below to produce a numeric scorecard for each competitor and your own product.
_Validation checkpoint: Every dimension should have a score and at least one supporting evidence note._
4. **Generate outputs** — Populate the templates in `references/analysis-templates.md` (Feature Matrix, Pricing Analysis, SWOT, Positioning Map, UX Audit).
5. **Build action plan** — Translate findings into the Action Items template (quick wins / medium-term / strategic).
6. **Package for stakeholders** — Assemble the Stakeholder Presentation using outputs from steps 3–5.
---
## Data Collection Guide
> Full executable scripts for each source are in `references/data-collection-guide.md`. Summaries of what to capture are below.
### 1. Website Analysis
Key things to capture:
- Pricing tiers and price points
- Feature lists per tier
- Primary CTA and messaging
- Case studies / customer logos (signals ICP)
- Integration logos
- Trust signals (certifications, compliance badges)
### 2. App Store Reviews
Review sentiment categories:
- **Praise** → what users love (defend / strengthen these)
- **Feature requests** → unmet needs (opportunity gaps)
- **Bugs** → quality signals
- **UX complaints** → friction points you can beat them on
**Sample App Store query (iTunes Search API):**
```
GET https://itunes.apple.com/search?term=<competitor_name>&entity=software&limit=1
# Extract trackId, then:
GET https://itunes.apple.com/rss/customerreviews/id=<trackId>/sortBy=mostRecent/json?l=en&limit=50
```
Parse `entry[].content.label` for review text and `entry[].im:rating.label` for star rating.
### 3. Job Postings (Team Size & Tech Stack Signals)
Signals from job postings:
- **Engineering volume** → scaling vs. consolidating
- **Specific tech mentions** → stack (React/Vue, Postgres/Mongo, AWS/GCP)
- **Sales/CS ratio** → product-led vs. sales-led motion
- **Data/ML roles** → upcoming AI features
- **Compliance roles** → regulatory expansion
### 4. SEO Analysis
SEO signals to capture:
- Top 20 organic keywords (intent: informational / navigational / commercial)
- Domain Authority / backlink count
- Blog publishing cadence and topics
- Which pages rank (product pages vs. blog vs. docs)
### 5. Social Media Sentiment
Capture recent mentions via Twitter/X API v2, Reddit, or LinkedIn. Look for recurring praise, complaints, and feature requests. See `references/data-collection-guide.md` for API query examples.
---
## Scoring Rubric (12 Dimensions, 1-5)
| # | Dimension | 1 (Weak) | 3 (Average) | 5 (Best-in-class) |
|---|-----------|----------|-------------|-------------------|
| 1 | **Features** | Core only, many gaps | Solid coverage | Comprehensive + unique |
| 2 | **Pricing** | Confusing / overpriced | Market-rate, clear | Transparent, flexible, fair |
| 3 | **UX** | Confusing, high friction | Functional | Delightful, minimal friction |
| 4 | **Performance** | Slow, unreliable | Acceptable | Fast, high uptime |
| 5 | **Docs** | Sparse, outdated | Decent coverage | Comprehensive, searchable |
| 6 | **Support** | Email only, slow | Chat + email | 24/7, great response |
| 7 | **Integrations** | 0-5 integrations | 6-25 | 26+ or deep ecosystem |
| 8 | **Security** | No mentions | SOC2 claimed | SOC2 Type II, ISO 27001 |
| 9 | **Scalability** | No enterprise tier | Mid-market ready | Enterprise-grade |
| 10 | **Brand** | Generic, unmemorable | Decent positioning | Strong, differentiated |
| 11 | **Community** | None | Forum / Slack | Active, vibrant community |
| 12 | **Innovation** | No recent releases | Quarterly | Frequent, meaningful |
**Example completed row** (Competitor: Acme Corp, Dimension 3 – UX):
| Dimension | Acme Corp Score | Evidence |
|-----------|----------------|---------|
| UX | 2 | App Store reviews cite "confusing navigation" (38 mentions); onboarding requires 7 steps before TTFV; no onboarding wizard; CC required at signup. |
Apply this pattern to all 12 dimensions for each competitor.
---
## Templates
> Full template markdown is in `references/analysis-templates.md`. Abbreviated reference below.
### Feature Comparison Matrix
Rows: core features, pricing tiers, platform capabilities (web, iOS, Android, API).
Columns: your product + up to 3 competitors.
Score each cell 1–5. Sum to get total out of 60.
**Score legend:** 5=Best-in-class, 4=Strong, 3=Average, 2=Below average, 1=Weak/Missing
### Pricing Analysis
Capture per competitor: model type (per-seat / usage-based / flat rate / freemium), entry/mid/enterprise price points, free trial length.
Summarize: price leader, value leader, premium positioning, your position, and 2–3 pricing opportunity bullets.
### SWOT Analysis
For each competitor: 3–5 bullets per quadrant (Strengths, Weaknesses, Opportunities for us, Threats to us). Anchor every bullet to a data signal (review quote, job posting count, pricing page, etc.).
### Positioning Map
2x2 axes (e.g., Simple ↔ Complex / Low Value ↔ High Value). Place each competitor and your product. Bubble size = market share or funding. See `references/analysis-templates.md` for ASCII and editable versions.
### UX Audit Checklist
Onboarding: TTFV (minutes), steps to activation, CC-required, onboarding wizard quality.
Key workflows: steps, friction points, comparative score (yours vs. theirs).
Mobile: iOS/Android ratings, feature parity, top complaint and praise.
Navigation: global search, keyboard shortcuts, in-app help.
### Action Items
| Horizon | Effort | Examples |
|---------|--------|---------|
| Quick wins (0–4 wks) | Low | Add review badges, publish comparison landing page |
| Medium-term (1–3 mo) | Moderate | Launch free tier, improve onboarding TTFV, add top-requested integration |
| Strategic (3–12 mo) | High | Enter new market, build API v2, achieve SOC2 Type II |
### Stakeholder Presentation (7 slides)
1. **Executive Summary** — Threat level (LOW/MEDIUM/HIGH/CRITICAL), top strength, top opportunity, recommended action
2. **Market Position** — 2x2 positioning map
3. **Feature Scorecard** — 12-dimension radar or table, total scores
4. **Pricing Analysis** — Comparison table + key insight
5. **UX Highlights** — What they do better (3 bullets) vs. where we win (3 bullets)
6. **Voice of Customer** — Top 3 review complaints (quoted or paraphrased)
7. **Our Action Plan** — Quick wins, medium-term, strategic priorities; Appendix with raw data
## Related Skills
- **Product Strategist** (`product-team/product-strategist/`) — Competitive insights feed OKR and strategy planning
- **Landing Page Generator** (`product-team/landing-page-generator/`) — Competitive positioning informs landing page messaging
FILE:references/analysis-templates.md
# Competitive Analysis Templates
## 1. SWOT Analysis Template
### Company/Product: [Competitor Name]
**Date:** [Analysis Date] | **Analyst:** [Name] | **Version:** [1.0]
#### Strengths (Internal Advantages)
| # | Strength | Evidence | Impact |
|---|----------|----------|--------|
| 1 | [e.g., Strong brand recognition] | [Source/data point] | High/Med/Low |
| 2 | | | |
| 3 | | | |
#### Weaknesses (Internal Limitations)
| # | Weakness | Evidence | Exploitability |
|---|----------|----------|---------------|
| 1 | [e.g., Limited API capabilities] | [Source/data point] | High/Med/Low |
| 2 | | | |
| 3 | | | |
#### Opportunities (External Favorable)
| # | Opportunity | Timeframe | Our Advantage |
|---|------------|-----------|---------------|
| 1 | [e.g., Competitor slow to adopt AI] | Short/Med/Long | [How we capitalize] |
| 2 | | | |
| 3 | | | |
#### Threats (External Unfavorable)
| # | Threat | Likelihood | Mitigation |
|---|--------|-----------|-----------|
| 1 | [e.g., Competitor acquired by larger company] | High/Med/Low | [Our response plan] |
| 2 | | | |
| 3 | | | |
---
## 2. Porter's Five Forces (Product Application)
### Market: [Your Product Category]
#### Force 1: Competitive Rivalry (Intensity: High/Med/Low)
- Number of direct competitors: ___
- Market growth rate: ___% annually
- Product differentiation level: High/Med/Low
- Switching costs for customers: High/Med/Low
- Exit barriers: High/Med/Low
- **Assessment:** [Summary of competitive rivalry intensity]
#### Force 2: Threat of New Entrants (Intensity: High/Med/Low)
- Capital requirements: High/Med/Low
- Technology barriers: High/Med/Low
- Network effects strength: Strong/Moderate/Weak
- Regulatory barriers: High/Med/Low
- Brand loyalty in market: Strong/Moderate/Weak
- **Assessment:** [Summary of new entrant threat]
#### Force 3: Threat of Substitutes (Intensity: High/Med/Low)
- Alternative solutions: [List substitutes]
- Price-performance of substitutes: Better/Same/Worse
- Switching costs to substitutes: High/Med/Low
- Customer propensity to switch: High/Med/Low
- **Assessment:** [Summary of substitute threat]
#### Force 4: Bargaining Power of Buyers (Power: High/Med/Low)
- Buyer concentration: Concentrated/Fragmented
- Price sensitivity: High/Med/Low
- Information availability: Full/Partial/Limited
- Switching costs: High/Med/Low
- Volume of purchases: High/Med/Low
- **Assessment:** [Summary of buyer power]
#### Force 5: Bargaining Power of Suppliers (Power: High/Med/Low)
- Key technology dependencies: [List]
- Cloud provider lock-in: High/Med/Low
- Talent market tightness: Tight/Balanced/Loose
- Data source dependencies: Critical/Important/Optional
- **Assessment:** [Summary of supplier power]
#### Overall Industry Attractiveness: [Score 1-10]
---
## 3. Competitive Positioning Map
### Axis Definitions
- **X-Axis:** [e.g., Ease of Use] (Low to High)
- **Y-Axis:** [e.g., Feature Completeness] (Low to High)
### Competitor Positions
| Competitor | X Score (1-10) | Y Score (1-10) | Quadrant |
|-----------|---------------|---------------|----------|
| Your Product | ___ | ___ | ___ |
| Competitor A | ___ | ___ | ___ |
| Competitor B | ___ | ___ | ___ |
| Competitor C | ___ | ___ | ___ |
| Competitor D | ___ | ___ | ___ |
### Quadrant Definitions
- **Top-Right (Leaders):** High on both axes - market leaders
- **Top-Left (Feature-Rich):** High features, lower ease of use - complex tools
- **Bottom-Right (Simple):** Easy to use, fewer features - niche players
- **Bottom-Left (Laggards):** Low on both axes - disruption candidates
### Positioning Insights
- **White space opportunities:** [Areas with no competitor presence]
- **Crowded areas:** [Where competition is fiercest]
- **Our trajectory:** [Direction we're moving on the map]
---
## 4. Win/Loss Analysis Template
### Deal: [Opportunity Name]
**Date:** [Close Date] | **Result:** Won / Lost | **Competitor:** [Name]
#### Deal Context
- **Deal Size:** $___
- **Sales Cycle:** ___ days
- **Segment:** SMB / Mid-Market / Enterprise
- **Industry:** ___
- **Decision Makers:** [Roles involved]
- **Evaluation Criteria:** [What mattered most to buyer]
#### Competitive Comparison (Buyer Perspective)
| Factor | Us (Score 1-5) | Competitor (Score 1-5) | Decisive? |
|--------|---------------|----------------------|-----------|
| Product Fit | | | Yes/No |
| Pricing | | | Yes/No |
| Ease of Use | | | Yes/No |
| Support Quality | | | Yes/No |
| Integration | | | Yes/No |
| Brand/Trust | | | Yes/No |
| Implementation | | | Yes/No |
#### Win/Loss Factors
- **Primary reason for outcome:** [Single most important factor]
- **Secondary factors:** [Supporting reasons]
- **Buyer quotes:** ["Direct quotes from debrief"]
#### Action Items
| # | Action | Owner | Due Date |
|---|--------|-------|----------|
| 1 | [e.g., Improve onboarding flow] | [Name] | [Date] |
| 2 | | | |
---
## 5. Battle Card Template
### Competitor: [Name]
**Last Updated:** [Date] | **Confidence:** High/Med/Low
#### Quick Facts
- **Founded:** ___
- **Funding:** $___
- **Employees:** ___
- **Customers:** ___
- **HQ:** ___
#### Elevator Pitch (Their Positioning)
> [How the competitor describes themselves in one sentence]
#### Our Positioning Against Them
> [How we differentiate - our one-liner against this competitor]
#### Where They Win
| Strength | Our Counter |
|----------|------------|
| [e.g., Lower price point] | [e.g., Emphasize TCO including implementation costs] |
| [e.g., Larger integration marketplace] | [e.g., Highlight quality over quantity, key integrations] |
| | |
#### Where We Win
| Our Strength | Evidence |
|-------------|----------|
| [e.g., Superior onboarding experience] | [Metric or customer quote] |
| [e.g., Better enterprise security] | [Certification or feature] |
| | |
#### Landmines to Set
Questions to ask prospects that expose competitor weaknesses:
1. "Have you evaluated how [specific capability] scales beyond [threshold]?"
2. "What's their approach to [area where competitor is weak]?"
3. "Can you share their uptime SLA and historical performance?"
#### Objection Handling
| Objection | Response |
|-----------|----------|
| "[Competitor] is cheaper" | [Value-based response] |
| "[Competitor] has more features" | [Quality/relevance response] |
| "We already use [Competitor]" | [Migration/coexistence story] |
#### Trap Questions They Set
Questions competitors ask about us, and how to respond:
1. **Q:** "[Our known weakness]?" **A:** [Honest, redirect response]
2. **Q:** "[Feature gap]?" **A:** [Roadmap or alternative approach]
#### Recent Intel
- [Date]: [Notable change - pricing, feature, hire, funding]
- [Date]: [Notable change]
FILE:references/competitive-analysis-frameworks.md
# Competitive Analysis Frameworks
This reference provides practical frameworks for evaluating competitors and positioning decisions.
## Porter's Five Forces
Assess the competitive intensity of your market:
1. Threat of new entrants
- Barriers to entry (capital, regulation, network effects)
- Speed of competitor replication
2. Bargaining power of suppliers
- Dependency on core infrastructure vendors
- Concentration of key technical providers
3. Bargaining power of buyers
- Customer switching costs
- Procurement complexity and contract leverage
4. Threat of substitutes
- Adjacent alternatives solving the same job
- DIY and internal build options
5. Rivalry among existing competitors
- Number and similarity of competitors
- Price competition and differentiation pressure
### Five Forces Template
| Force | Current Pressure (Low/Med/High) | Evidence | Strategic Response |
|---|---|---|---|
| New Entrants | | | |
| Supplier Power | | | |
| Buyer Power | | | |
| Substitutes | | | |
| Rivalry | | | |
## SWOT Analysis
Use SWOT to map internal and external context quickly.
### SWOT Template
| Strengths (Internal) | Weaknesses (Internal) |
|---|---|
| What we do better than alternatives | Where competitors outperform us |
| Unique capabilities or assets | Known product or go-to-market gaps |
| Opportunities (External) | Threats (External) |
|---|---|
| Market trends we can exploit | Competitor moves or macro risks |
| Unserved segments and use cases | Regulatory, platform, or pricing pressure |
### SWOT Quality Checklist
- Base every point on evidence, not assumptions.
- Separate observations from conclusions.
- Prioritize top 3 items per quadrant.
## Feature Comparison Matrix
Compare products on meaningful buying criteria, not vanity features.
### Feature Matrix Template
| Dimension | Weight | Your Product | Competitor A | Competitor B | Notes |
|---|---:|---:|---:|---:|---|
| Core workflow coverage | 25% | | | | |
| Ease of implementation | 15% | | | | |
| Performance / reliability | 15% | | | | |
| Integrations / ecosystem | 15% | | | | |
| Security / compliance | 15% | | | | |
| Pricing / TCO | 15% | | | | |
Scoring scale recommendation: 1-5 (weak to strong).
## Competitive Positioning Map
Create a 2-axis map showing market whitespace and crowding.
### Positioning Map Steps
1. Select two high-signal dimensions customers care about.
2. Place each competitor based on evidence (pricing pages, reviews, demos).
3. Mark clusters where products are undifferentiated.
4. Identify white space where demand exists but options are weak.
Example axes:
- X-axis: Ease of use
- Y-axis: Enterprise readiness
## Blue Ocean Strategy Canvas
Use a strategy canvas to decide where to raise, reduce, eliminate, or create factors.
### ERRC Grid (Eliminate-Reduce-Raise-Create)
| Eliminate | Reduce | Raise | Create |
|---|---|---|---|
| Commodity table-stakes not valued by target users | Costly features with weak adoption | Differentiators tied to target job-to-be-done | New value dimensions competitors ignore |
### Strategy Canvas Checklist
- Compare value curves between your product and top competitors.
- Ensure target segment is explicit.
- Tie every strategic choice to measurable outcome.
FILE:references/data-collection-guide.md
# Competitive Data Collection Guide
## Overview
This guide outlines systematic approaches for gathering competitive intelligence from publicly available sources. All methods described here are ethical and rely on information that competitors have made publicly accessible.
## Public Data Sources
### Review Platforms
- **G2**: Enterprise software reviews, feature comparisons, satisfaction scores
- **Capterra**: SMB-focused reviews, pricing transparency, deployment details
- **TrustRadius**: In-depth reviews with verified users, TrustMaps
- **Product Hunt**: Launch positioning, early adopter sentiment, feature highlights
- **App Store / Google Play**: Mobile app ratings, review themes, update frequency
### Company Publications
- **Pricing Pages**: Tier structure, feature gating, enterprise vs self-serve
- **Changelogs / Release Notes**: Development velocity, feature priorities, tech direction
- **Blog Posts**: Strategic messaging, thought leadership topics, market positioning
- **Case Studies**: Target customer profiles, value propositions, success metrics
- **Help Documentation**: Feature depth, API capabilities, integration ecosystem
### Talent & Organization Signals
- **Job Postings**: Technology stack, team growth areas, strategic initiatives
- **LinkedIn**: Team size, org structure, key hires, department ratios
- **Glassdoor**: Company culture, internal challenges, growth trajectory
### Financial & Legal
- **Patent Filings**: Innovation direction, defensive IP, technology differentiation
- **SEC Filings (public companies)**: Revenue, growth rate, customer count, churn
- **Crunchbase / PitchBook**: Funding rounds, investors, valuation trends
### Technical Intelligence
- **BuiltWith / Wappalyzer**: Technology stack detection
- **GitHub**: Open-source contributions, SDK quality, developer engagement
- **API Documentation**: Integration capabilities, rate limits, data models
- **Status Pages**: Uptime history, incident frequency, infrastructure maturity
## Data Points to Collect Per Competitor
### Product
- Core features and capabilities (feature-by-feature matrix)
- Unique differentiators and proprietary technology
- Platform support (web, mobile, desktop, API)
- Integration ecosystem (number and quality of integrations)
- Performance benchmarks (if available from reviews)
### Business
- Pricing tiers and per-seat/usage costs
- Target customer segments (SMB, mid-market, enterprise)
- Estimated customer count and notable logos
- Geographic focus and localization
- Go-to-market model (PLG, sales-led, hybrid)
### Team & Technology
- Estimated team size and engineering ratio
- Technology stack and infrastructure choices
- Development velocity (release frequency)
- Open-source involvement and developer relations
### Market Position
- Market share estimates
- Brand perception and NPS (from reviews)
- Analyst coverage (Gartner, Forrester positioning)
- Partnership and channel strategy
## Ethical Guidelines
1. **Use only public information** - Never access private systems, NDA-protected content, or internal documents
2. **No deception** - Do not misrepresent yourself to obtain information (e.g., fake sales inquiries)
3. **Respect terms of service** - Follow scraping policies and API usage terms
4. **Attribute sources** - Document where each data point came from for verification
5. **No employee poaching for intelligence** - Hiring decisions should be talent-driven, not intelligence-driven
6. **Legal compliance** - Ensure data collection complies with local regulations
## Update Cadence Recommendations
| Data Type | Frequency | Trigger Events |
|-----------|-----------|---------------|
| Pricing | Monthly | Competitor pricing page changes |
| Features | Bi-weekly | Changelog updates, product launches |
| Reviews | Monthly | Batch review analysis |
| Job Postings | Monthly | Hiring surge detection |
| Financials | Quarterly | Earnings reports, funding rounds |
| Tech Stack | Quarterly | Major platform changes |
| Full Teardown | Quarterly | Strategic planning cycles |
## Collection Workflow
1. **Set up monitoring** - Google Alerts, competitor RSS feeds, social listening
2. **Schedule regular sweeps** - Calendar recurring data collection tasks
3. **Centralize data** - Use a shared competitive intelligence database or spreadsheet
4. **Validate findings** - Cross-reference multiple sources for accuracy
5. **Tag and categorize** - Apply consistent taxonomy for easy retrieval
6. **Share insights** - Distribute relevant findings to product, sales, and marketing teams
7. **Archive versions** - Maintain historical snapshots for trend analysis
## Tools for Automation
- **Google Alerts**: Free monitoring for competitor mentions
- **Visualping**: Website change detection (pricing pages, feature pages)
- **Feedly**: RSS aggregation for competitor blogs and news
- **SimilarWeb**: Traffic estimates and audience overlap
- **SEMrush / Ahrefs**: SEO positioning and content strategy analysis
FILE:references/scoring-rubric.md
# Competitive Scoring Rubric
## Overview
This rubric provides a standardized framework for evaluating competitors across key dimensions. Consistent scoring enables meaningful comparisons and tracks competitive position changes over time.
## Scoring Scale (1-10)
| Score | Label | Definition |
|-------|-------|-----------|
| 1-2 | Poor | Significant gaps, major usability issues, or missing capability |
| 3-4 | Below Average | Basic functionality with notable limitations |
| 5-6 | Average | Meets market expectations, no standout qualities |
| 7-8 | Above Average | Strong execution with clear advantages |
| 9-10 | Exceptional | Industry-leading, sets the standard for others |
## Dimension Categories
### 1. User Experience (UX) - Weight: 20%
- **Onboarding**: Time to first value, setup complexity, guided flows
- **Navigation**: Information architecture, discoverability, consistency
- **Visual Design**: Modern aesthetics, brand coherence, accessibility
- **Performance**: Page load times, responsiveness, offline capability
- **Mobile Experience**: Native app quality, responsive design, feature parity
### 2. Feature Completeness - Weight: 25%
- **Core Features**: Coverage of essential use cases
- **Advanced Features**: Power user capabilities, automation, customization
- **Workflow Support**: End-to-end process coverage without workarounds
- **API & Extensibility**: API coverage, webhook support, SDK quality
- **Innovation**: Unique capabilities not found in competitors
### 3. Pricing & Value - Weight: 15%
- **Transparency**: Clear pricing without hidden costs
- **Flexibility**: Plan options matching different customer sizes
- **Value-to-Cost Ratio**: Feature access relative to price point
- **Free Tier / Trial**: Quality of free offering for evaluation
- **Contract Terms**: Lock-in requirements, cancellation ease
### 4. Integrations - Weight: 10%
- **Native Integrations**: Number and quality of built-in connectors
- **Marketplace**: Third-party app ecosystem breadth
- **API Quality**: Documentation, reliability, rate limits
- **Data Import/Export**: Migration ease, format support
- **Workflow Automation**: Zapier, Make, native automation support
### 5. Support & Documentation - Weight: 10%
- **Documentation Quality**: Completeness, searchability, freshness
- **Support Channels**: Chat, email, phone, community availability
- **Response Time**: SLA adherence, resolution speed
- **Self-Service**: Knowledge base, video tutorials, community forums
- **Onboarding Support**: Dedicated CSM, implementation assistance
### 6. Performance & Reliability - Weight: 10%
- **Uptime**: Historical availability, SLA commitments
- **Speed**: Application responsiveness under normal load
- **Scalability**: Performance at high volume, enterprise readiness
- **Data Handling**: Large dataset support, bulk operations
- **Global Performance**: CDN, regional deployments, latency
### 7. Security & Compliance - Weight: 10%
- **Authentication**: SSO, MFA, RBAC granularity
- **Data Protection**: Encryption at rest and in transit, data residency
- **Certifications**: SOC 2, ISO 27001, GDPR, HIPAA compliance
- **Audit Trail**: Activity logging, access monitoring
- **Privacy Controls**: Data retention policies, right to deletion
## Weighting Guidelines
Default weights above suit most B2B SaaS evaluations. Adjust based on:
- **Enterprise buyers**: Increase Security (15%), Support (15%), reduce Pricing (10%)
- **Developer tools**: Increase Integrations (20%), Features (30%), reduce UX (10%)
- **SMB products**: Increase Pricing (25%), UX (25%), reduce Security (5%)
- **Regulated industries**: Increase Security (25%), reduce Features (15%)
## Calibration Process
1. **Anchor scoring** - Score your own product first to establish baseline
2. **Multiple scorers** - Have 2-3 team members score independently
3. **Discuss outliers** - Reconcile scores that differ by more than 2 points
4. **Document evidence** - Record specific examples justifying each score
5. **Normalize quarterly** - Re-calibrate as market expectations evolve
## Bias Mitigation
- **Avoid halo effect** - Score each dimension independently, not influenced by overall impression
- **Use evidence, not feelings** - Every score must link to observable data points
- **Include competitor strengths** - Resist tendency to under-score competitors
- **Rotate scorers** - Different team members bring fresh perspectives
- **Blind scoring** - When possible, evaluate features without knowing which competitor
- **Customer validation** - Compare internal scores against user review sentiment
## Composite Score Calculation
```
Weighted Score = SUM(Dimension Score x Dimension Weight)
Example:
UX(8) x 0.20 = 1.60
Features(7) x 0.25 = 1.75
Pricing(6) x 0.15 = 0.90
Integrations(8) x 0.10 = 0.80
Support(7) x 0.10 = 0.70
Performance(9) x 0.10 = 0.90
Security(8) x 0.10 = 0.80
---
Total = 7.45 / 10
```
## Output Format
Present results as a comparison matrix with color coding:
- Green (8-10): Competitive advantage
- Yellow (5-7): Market parity
- Red (1-4): Competitive gap
FILE:scripts/competitive_matrix_builder.py
#!/usr/bin/env python3
"""Competitive Matrix Builder — Analyze and score competitors across feature dimensions.
Generates weighted competitive matrices, gap analysis, and positioning insights
from structured competitor data.
Usage:
python competitive_matrix_builder.py competitors.json --format json
python competitive_matrix_builder.py competitors.json --format text
python competitive_matrix_builder.py competitors.json --format text --weights pricing=2,ux=1.5
"""
import argparse
import json
import sys
from typing import Dict, List, Any, Optional
from datetime import datetime
from statistics import mean, stdev
def load_competitors(path: str) -> Dict[str, Any]:
"""Load competitor data from JSON file."""
with open(path, "r") as f:
return json.load(f)
def normalize_score(value: float, min_val: float = 1.0, max_val: float = 10.0) -> float:
"""Normalize a score to 0-100 scale."""
return max(0.0, min(100.0, ((value - min_val) / (max_val - min_val)) * 100))
def calculate_weighted_scores(
competitors: List[Dict[str, Any]],
dimensions: List[str],
weights: Optional[Dict[str, float]] = None
) -> List[Dict[str, Any]]:
"""Calculate weighted scores for each competitor across dimensions."""
if weights is None:
weights = {d: 1.0 for d in dimensions}
results = []
for comp in competitors:
scores = comp.get("scores", {})
weighted_total = 0.0
weight_sum = 0.0
dimension_results = {}
for dim in dimensions:
raw = scores.get(dim, 0)
w = weights.get(dim, 1.0)
normalized = normalize_score(raw)
weighted = normalized * w
weighted_total += weighted
weight_sum += w
dimension_results[dim] = {
"raw": raw,
"normalized": round(normalized, 1),
"weight": w,
"weighted": round(weighted, 1)
}
overall = round(weighted_total / weight_sum, 1) if weight_sum > 0 else 0
results.append({
"name": comp["name"],
"overall_score": overall,
"dimensions": dimension_results,
"tier": classify_tier(overall),
"pricing": comp.get("pricing", {}),
"strengths": comp.get("strengths", []),
"weaknesses": comp.get("weaknesses", [])
})
results.sort(key=lambda x: x["overall_score"], reverse=True)
return results
def classify_tier(score: float) -> str:
"""Classify competitor into tier based on overall score."""
if score >= 80:
return "Leader"
elif score >= 60:
return "Strong Competitor"
elif score >= 40:
return "Viable Alternative"
elif score >= 20:
return "Niche Player"
else:
return "Weak"
def gap_analysis(
your_scores: Dict[str, float],
competitor_scores: List[Dict[str, Any]],
dimensions: List[str]
) -> Dict[str, Any]:
"""Identify gaps between your product and competitors."""
gaps = {}
for dim in dimensions:
your_val = your_scores.get(dim, 0)
comp_vals = [c["dimensions"][dim]["raw"] for c in competitor_scores if dim in c.get("dimensions", {})]
if not comp_vals:
continue
avg_comp = mean(comp_vals)
best_comp = max(comp_vals)
gap_to_avg = round(your_val - avg_comp, 1)
gap_to_best = round(your_val - best_comp, 1)
gaps[dim] = {
"your_score": your_val,
"competitor_avg": round(avg_comp, 1),
"competitor_best": best_comp,
"gap_to_avg": gap_to_avg,
"gap_to_best": gap_to_best,
"status": "ahead" if gap_to_avg > 0.5 else ("behind" if gap_to_avg < -0.5 else "parity"),
"priority": "high" if gap_to_best < -2 else ("medium" if gap_to_best < -1 else "low")
}
return {
"gaps": gaps,
"biggest_opportunities": sorted(
[{"dimension": k, **v} for k, v in gaps.items() if v["status"] == "behind"],
key=lambda x: x["gap_to_best"]
)[:5],
"competitive_advantages": sorted(
[{"dimension": k, **v} for k, v in gaps.items() if v["status"] == "ahead"],
key=lambda x: -x["gap_to_avg"]
)[:5]
}
def positioning_analysis(scored: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Generate positioning insights from scored competitors."""
scores = [c["overall_score"] for c in scored]
return {
"market_leaders": [c["name"] for c in scored if c["tier"] == "Leader"],
"your_rank": next((i + 1 for i, c in enumerate(scored) if c.get("is_you")), None),
"total_competitors": len(scored),
"score_distribution": {
"mean": round(mean(scores), 1) if scores else 0,
"stdev": round(stdev(scores), 1) if len(scores) > 1 else 0,
"min": round(min(scores), 1) if scores else 0,
"max": round(max(scores), 1) if scores else 0
},
"tier_distribution": {
tier: len([c for c in scored if c["tier"] == tier])
for tier in ["Leader", "Strong Competitor", "Viable Alternative", "Niche Player", "Weak"]
}
}
def format_text(result: Dict[str, Any]) -> str:
"""Format results as human-readable text."""
lines = []
lines.append("=" * 70)
lines.append("COMPETITIVE MATRIX ANALYSIS")
lines.append(f"Generated: {result['generated_at']}")
lines.append("=" * 70)
# Ranking table
lines.append("\n## COMPETITIVE RANKING\n")
lines.append(f"{'Rank':<6}{'Competitor':<25}{'Score':<10}{'Tier':<20}")
lines.append("-" * 61)
for i, c in enumerate(result["scored_competitors"], 1):
marker = " ← YOU" if c.get("is_you") else ""
lines.append(f"{i:<6}{c['name']:<25}{c['overall_score']:<10}{c['tier']:<20}{marker}")
# Dimension breakdown
lines.append("\n## DIMENSION BREAKDOWN\n")
dims = result["dimensions"]
header = f"{'Dimension':<20}" + "".join(f"{c['name'][:12]:<14}" for c in result["scored_competitors"])
lines.append(header)
lines.append("-" * len(header))
for dim in dims:
row = f"{dim:<20}"
for c in result["scored_competitors"]:
val = c["dimensions"].get(dim, {}).get("raw", "N/A")
row += f"{val:<14}"
lines.append(row)
# Gap analysis
if result.get("gap_analysis"):
ga = result["gap_analysis"]
if ga["biggest_opportunities"]:
lines.append("\n## BIGGEST OPPORTUNITIES (where you're behind)\n")
for opp in ga["biggest_opportunities"]:
lines.append(f" • {opp['dimension']}: You={opp['your_score']}, "
f"Best={opp['competitor_best']}, Gap={opp['gap_to_best']} "
f"[{opp['priority'].upper()} priority]")
if ga["competitive_advantages"]:
lines.append("\n## COMPETITIVE ADVANTAGES (where you lead)\n")
for adv in ga["competitive_advantages"]:
lines.append(f" • {adv['dimension']}: You={adv['your_score']}, "
f"Avg={adv['competitor_avg']}, Lead=+{adv['gap_to_avg']}")
# Positioning
pos = result.get("positioning", {})
if pos:
lines.append("\n## MARKET POSITIONING\n")
lines.append(f" Market Leaders: {', '.join(pos.get('market_leaders', ['None']))}")
if pos.get("your_rank"):
lines.append(f" Your Rank: #{pos['your_rank']} of {pos['total_competitors']}")
dist = pos.get("score_distribution", {})
lines.append(f" Score Range: {dist.get('min', 0)} - {dist.get('max', 0)} "
f"(avg: {dist.get('mean', 0)}, stdev: {dist.get('stdev', 0)})")
lines.append("\n" + "=" * 70)
return "\n".join(lines)
def build_matrix(data: Dict[str, Any], weight_overrides: Optional[Dict[str, float]] = None) -> Dict[str, Any]:
"""Main entry: build competitive matrix from input data."""
competitors = data.get("competitors", [])
dimensions = data.get("dimensions", [])
your_product = data.get("your_product", {})
if not competitors:
return {"error": "No competitors provided"}
if not dimensions:
# Auto-detect from first competitor's scores
dimensions = list(competitors[0].get("scores", {}).keys())
weights = data.get("weights", {})
if weight_overrides:
weights.update(weight_overrides)
# Include your product in scoring if provided
all_entries = list(competitors)
if your_product:
your_product["is_you"] = True
all_entries.insert(0, your_product)
scored = calculate_weighted_scores(all_entries, dimensions, weights)
# Mark your product
for s in scored:
if any(c.get("is_you") and c["name"] == s["name"] for c in all_entries):
s["is_you"] = True
result = {
"generated_at": datetime.now().isoformat(),
"dimensions": dimensions,
"weights": weights if weights else {d: 1.0 for d in dimensions},
"scored_competitors": scored,
"positioning": positioning_analysis(scored)
}
if your_product:
result["gap_analysis"] = gap_analysis(
your_product.get("scores", {}), scored, dimensions
)
return result
def parse_weights(weight_str: str) -> Dict[str, float]:
"""Parse weight string like 'pricing=2,ux=1.5' into dict."""
weights = {}
for pair in weight_str.split(","):
if "=" in pair:
k, v = pair.split("=", 1)
weights[k.strip()] = float(v.strip())
return weights
def main():
parser = argparse.ArgumentParser(
description="Build competitive matrix with scoring and gap analysis"
)
parser.add_argument("input", help="Path to competitors JSON file")
parser.add_argument("--format", choices=["json", "text"], default="text",
help="Output format (default: text)")
parser.add_argument("--weights", type=str, default=None,
help="Weight overrides: 'dim1=2.0,dim2=1.5'")
parser.add_argument("--output", type=str, default=None,
help="Output file path (default: stdout)")
args = parser.parse_args()
data = load_competitors(args.input)
weight_overrides = parse_weights(args.weights) if args.weights else None
result = build_matrix(data, weight_overrides)
if args.format == "json":
output = json.dumps(result, indent=2)
else:
output = format_text(result)
if args.output:
with open(args.output, "w") as f:
f.write(output)
print(f"Output written to {args.output}")
else:
print(output)
if __name__ == "__main__":
main()
Ưu tiên tính năng theo RICE, khám phá khách hàng, soạn PRD và lập lộ trình sản phẩm.
---
name: cs-product-manager
description: Product management agent for feature prioritization, customer discovery, PRD development, and roadmap planning using RICE framework
skills: product-team/product-manager-toolkit, product-team/agile-product-owner, product-team/product-strategist, product-team/ux-researcher-designer, product-team/ui-design-system, product-team/competitive-teardown, product-team/landing-page-generator, product-team/saas-scaffolder
domain: product
model: sonnet
tools: [Read, Write, Bash, Grep, Glob]
---
# Product Manager Agent
## Purpose
The cs-product-manager agent is a specialized product management agent focused on feature prioritization, customer discovery, requirements documentation, and data-driven roadmap planning. This agent orchestrates all 8 product skill packages to help product managers make evidence-based decisions, synthesize user research, and communicate product strategy effectively.
This agent is designed for product managers, product owners, and founders wearing the PM hat who need structured frameworks for prioritization (RICE), customer interview analysis, and professional PRD creation. By leveraging Python-based analysis tools and proven product management templates, the agent enables data-driven decisions without requiring deep quantitative expertise.
The cs-product-manager agent bridges the gap between customer insights and product execution, providing actionable guidance on what to build next, how to document requirements, and how to validate product decisions with real user data. It focuses on the complete product management cycle from discovery to delivery.
## Skill Integration
**Primary Skill:** `../../product-team/product-manager-toolkit/`
### All Orchestrated Skills
| # | Skill | Location | Primary Tool |
|---|-------|----------|-------------|
| 1 | Product Manager Toolkit | `../../product-team/product-manager-toolkit/` | rice_prioritizer.py, customer_interview_analyzer.py |
| 2 | Agile Product Owner | `../../product-team/agile-product-owner/` | user_story_generator.py |
| 3 | Product Strategist | `../../product-team/product-strategist/` | okr_cascade_generator.py |
| 4 | UX Researcher & Designer | `../../product-team/ux-researcher-designer/` | persona_generator.py |
| 5 | UI Design System | `../../product-team/ui-design-system/` | design_token_generator.py |
| 6 | Competitive Teardown | `../../product-team/competitive-teardown/` | competitive_matrix_builder.py |
| 7 | Landing Page Generator | `../../product-team/landing-page-generator/` | landing_page_scaffolder.py |
| 8 | SaaS Scaffolder | `../../product-team/saas-scaffolder/` | project_bootstrapper.py |
### Python Tools
1. **RICE Prioritizer**
- **Purpose:** RICE framework implementation for feature prioritization with portfolio analysis and capacity planning
- **Path:** `../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py`
- **Usage:** `python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py features.csv --capacity 20`
- **Formula:** RICE Score = (Reach × Impact × Confidence) / Effort
- **Features:** Portfolio analysis (quick wins vs big bets), quarterly roadmap generation, capacity planning, JSON/CSV export
- **Use Cases:** Feature prioritization, roadmap planning, stakeholder alignment, resource allocation
2. **Customer Interview Analyzer**
- **Purpose:** NLP-based interview transcript analysis to extract pain points, feature requests, and themes
- **Path:** `../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py`
- **Usage:** `python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview.txt`
- **Features:** Pain point extraction with severity, feature request identification, jobs-to-be-done patterns, sentiment analysis, theme extraction
- **Use Cases:** User research synthesis, discovery validation, problem prioritization, insight generation
3. **User Story Generator**
- **Purpose:** Break epics into INVEST-compliant user stories with acceptance criteria
- **Path:** `../../product-team/agile-product-owner/scripts/user_story_generator.py`
- **Usage:** `python ../../product-team/agile-product-owner/scripts/user_story_generator.py epic.yaml`
- **Use Cases:** Sprint planning, backlog refinement, story decomposition
4. **OKR Cascade Generator**
- **Purpose:** Generate cascaded OKRs from company objectives to team-level key results
- **Path:** `../../product-team/product-strategist/scripts/okr_cascade_generator.py`
- **Usage:** `python ../../product-team/product-strategist/scripts/okr_cascade_generator.py growth`
- **Use Cases:** Quarterly planning, strategic alignment, goal setting
5. **Persona Generator**
- **Purpose:** Create data-driven user personas from research inputs
- **Path:** `../../product-team/ux-researcher-designer/scripts/persona_generator.py`
- **Usage:** `python ../../product-team/ux-researcher-designer/scripts/persona_generator.py research-data.json`
- **Use Cases:** User research synthesis, persona development, journey mapping
6. **Design Token Generator**
- **Purpose:** Generate design tokens for consistent UI implementation
- **Path:** `../../product-team/ui-design-system/scripts/design_token_generator.py`
- **Usage:** `python ../../product-team/ui-design-system/scripts/design_token_generator.py theme.json`
- **Use Cases:** Design system creation, developer handoff, theming
7. **Competitive Matrix Builder**
- **Purpose:** Build competitive analysis matrices and feature comparison grids
- **Path:** `../../product-team/competitive-teardown/scripts/competitive_matrix_builder.py`
- **Usage:** `python ../../product-team/competitive-teardown/scripts/competitive_matrix_builder.py competitors.csv`
- **Use Cases:** Competitive intelligence, market positioning, feature gap analysis
8. **Landing Page Scaffolder**
- **Purpose:** Generate conversion-optimized landing page scaffolds
- **Path:** `../../product-team/landing-page-generator/scripts/landing_page_scaffolder.py`
- **Usage:** `python ../../product-team/landing-page-generator/scripts/landing_page_scaffolder.py config.yaml`
- **Use Cases:** Product launches, A/B testing, GTM campaigns
9. **Project Bootstrapper**
- **Purpose:** Scaffold SaaS project structures with boilerplate and configurations
- **Path:** `../../product-team/saas-scaffolder/scripts/project_bootstrapper.py`
- **Usage:** `python ../../product-team/saas-scaffolder/scripts/project_bootstrapper.py --stack nextjs --name my-saas`
- **Use Cases:** MVP scaffolding, project kickoff, SaaS prototype creation
### Knowledge Bases
1. **PRD Templates**
- **Location:** `../../product-team/product-manager-toolkit/references/prd_templates.md`
- **Content:** Multiple PRD formats (Standard PRD, One-Page PRD, Feature Brief, Agile Epic), structure guidelines, best practices
- **Use Case:** Requirements documentation, stakeholder communication, engineering handoff
2. **Sprint Planning Guide**
- **Location:** `../../product-team/agile-product-owner/references/sprint-planning-guide.md`
- **Content:** Sprint planning ceremonies, velocity tracking, capacity allocation
- **Use Case:** Sprint execution, backlog refinement, agile ceremonies
3. **User Story Templates**
- **Location:** `../../product-team/agile-product-owner/references/user-story-templates.md`
- **Content:** INVEST-compliant story formats, acceptance criteria patterns, story splitting techniques
- **Use Case:** Story writing, backlog grooming, definition of done
4. **OKR Framework**
- **Location:** `../../product-team/product-strategist/references/okr_framework.md`
- **Content:** OKR methodology, cascade patterns, scoring guidelines
- **Use Case:** Quarterly planning, strategic alignment, goal tracking
5. **Strategy Types**
- **Location:** `../../product-team/product-strategist/references/strategy_types.md`
- **Content:** Product strategy frameworks, competitive positioning, growth strategies
- **Use Case:** Strategic planning, market analysis, product vision
6. **Persona Methodology**
- **Location:** `../../product-team/ux-researcher-designer/references/persona-methodology.md`
- **Content:** Research-backed persona creation methodology, data collection, validation
- **Use Case:** Persona development, user segmentation, research planning
7. **Example Personas**
- **Location:** `../../product-team/ux-researcher-designer/references/example-personas.md`
- **Content:** Sample persona documents with demographics, goals, pain points, behaviors
- **Use Case:** Persona templates, research documentation
8. **Journey Mapping Guide**
- **Location:** `../../product-team/ux-researcher-designer/references/journey-mapping-guide.md`
- **Content:** Customer journey mapping methodology, touchpoint analysis, emotion mapping
- **Use Case:** Experience design, touchpoint optimization, service design
9. **Usability Testing Frameworks**
- **Location:** `../../product-team/ux-researcher-designer/references/usability-testing-frameworks.md`
- **Content:** Usability test planning, task design, analysis methods
- **Use Case:** Usability studies, prototype validation, UX evaluation
10. **Component Architecture**
- **Location:** `../../product-team/ui-design-system/references/component-architecture.md`
- **Content:** Component hierarchy, atomic design patterns, composition strategies
- **Use Case:** Design system architecture, component libraries
11. **Developer Handoff**
- **Location:** `../../product-team/ui-design-system/references/developer-handoff.md`
- **Content:** Design-to-dev handoff process, specification formats, asset delivery
- **Use Case:** Engineering collaboration, implementation specs
12. **Responsive Calculations**
- **Location:** `../../product-team/ui-design-system/references/responsive-calculations.md`
- **Content:** Responsive design formulas, breakpoint strategies, fluid typography
- **Use Case:** Responsive implementation, cross-device design
13. **Token Generation**
- **Location:** `../../product-team/ui-design-system/references/token-generation.md`
- **Content:** Design token standards, naming conventions, platform-specific output
- **Use Case:** Design system tokens, theming, multi-platform consistency
## Workflows
### Workflow 1: Feature Prioritization & Roadmap Planning
**Goal:** Prioritize feature backlog using RICE framework and generate quarterly roadmap
**Steps:**
1. **Gather Feature Requests** - Collect from multiple sources:
- Customer feedback (support tickets, interviews)
- Sales team requests
- Technical debt items
- Strategic initiatives
- Competitive gaps
2. **Create RICE Input CSV** - Structure features with RICE parameters:
```csv
feature,reach,impact,confidence,effort
User Dashboard,500,3,0.8,5
API Rate Limiting,1000,2,0.9,3
Dark Mode,300,1,1.0,2
```
- **Reach**: Number of users affected per quarter
- **Impact**: massive(3), high(2), medium(1.5), low(1), minimal(0.5)
- **Confidence**: high(1.0), medium(0.8), low(0.5)
- **Effort**: person-months (XL=6, L=3, M=1, S=0.5, XS=0.25)
3. **Run RICE Prioritization** - Execute analysis with team capacity
```bash
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py features.csv --capacity 20
```
4. **Analyze Portfolio** - Review output for:
- **Quick Wins**: High RICE, low effort (ship first)
- **Big Bets**: High RICE, high effort (strategic investments)
- **Fill-Ins**: Medium RICE (capacity fillers)
- **Money Pits**: Low RICE, high effort (avoid or revisit)
5. **Generate Quarterly Roadmap**:
- Q1: Top quick wins + 1-2 big bets
- Q2-Q4: Remaining prioritized features
- Buffer: 20% capacity for unknowns
6. **Stakeholder Alignment** - Present roadmap with:
- RICE scores as justification
- Trade-off decisions explained
- Capacity constraints visible
**Expected Output:** Data-driven quarterly roadmap with RICE-justified priorities and portfolio balance
**Time Estimate:** 4-6 hours for complete prioritization cycle (20-30 features)
**Example:**
```bash
# Complete prioritization workflow
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py q4-features.csv --capacity 20 > roadmap.txt
cat roadmap.txt
# Review quick wins, big bets, and generate quarterly plan
```
### Workflow 2: Customer Discovery & Interview Analysis
**Goal:** Conduct customer interviews, extract insights, and identify high-priority problems
**Steps:**
1. **Conduct User Interviews** - Semi-structured format:
- **Opening**: Build rapport, explain purpose
- **Context**: Current workflow and challenges
- **Problems**: Deep dive on pain points (not solutions!)
- **Solutions**: Reaction to concepts (if applicable)
- **Closing**: Next steps, thank you
- **Duration**: 30-45 minutes per interview
- **Record**: With permission for analysis
2. **Transcribe Interviews** - Convert audio to text:
- Use transcription service (Otter.ai, Rev, etc.)
- Clean up for clarity (remove filler words)
- Save as plain text file
3. **Run Interview Analyzer** - Extract structured insights
```bash
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-001.txt
```
4. **Review Analysis Output** - Study extracted insights:
- **Pain Points**: Severity-scored problems
- **Feature Requests**: Priority-ranked asks
- **Jobs-to-be-Done**: User goals and motivations
- **Sentiment**: Overall satisfaction level
- **Themes**: Recurring topics across interviews
- **Key Quotes**: Direct user language
5. **Synthesize Across Interviews** - Aggregate insights:
```bash
# Analyze multiple interviews
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-001.txt json > insights-001.json
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-002.txt json > insights-002.json
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-003.txt json > insights-003.json
# Aggregate JSON files to find patterns
```
6. **Prioritize Problems** - Identify which pain points to solve:
- Frequency: How many users mentioned it?
- Severity: How painful is the problem?
- Strategic fit: Aligns with company vision?
- Solvability: Can we build a solution?
7. **Validate Solutions** - Test hypotheses before building:
- Create mockups or prototypes
- Show to users, observe reactions
- Measure willingness to pay/adopt
**Expected Output:** Prioritized list of validated problems with user quotes and evidence
**Time Estimate:** 2-3 weeks for complete discovery (10-15 interviews + analysis)
### Workflow 3: PRD Development & Stakeholder Communication
**Goal:** Document requirements professionally with clear scope, metrics, and acceptance criteria
**Steps:**
1. **Choose PRD Template** - Select based on complexity:
```bash
cat ../../product-team/product-manager-toolkit/references/prd_templates.md
```
- **Standard PRD**: Complex features (6-8 weeks dev)
- **One-Page PRD**: Simple features (2-4 weeks)
- **Feature Brief**: Exploration phase (1 week)
- **Agile Epic**: Sprint-based delivery
2. **Document Problem** - Start with why (not how):
- User problem statement (jobs-to-be-done format)
- Evidence from interviews (quotes, data)
- Current workarounds and pain points
- Business impact (revenue, retention, efficiency)
3. **Define Solution** - Describe what we'll build:
- High-level solution approach
- User flows and key interactions
- Technical architecture (if relevant)
- Design mockups or wireframes
- **Critically: What's OUT of scope**
4. **Set Success Metrics** - Define how we'll measure success:
- **Leading indicators**: Usage, adoption, engagement
- **Lagging indicators**: Revenue, retention, NPS
- **Target values**: Specific, measurable goals
- **Timeframe**: When we expect to hit targets
5. **Write Acceptance Criteria** - Clear definition of done:
- Given/When/Then format for each user story
- Edge cases and error states
- Performance requirements
- Accessibility standards
6. **Collaborate with Stakeholders**:
- **Engineering**: Feasibility review, effort estimation
- **Design**: User experience validation
- **Sales/Marketing**: Go-to-market alignment
- **Support**: Operational readiness
7. **Iterate Based on Feedback** - Incorporate input:
- Technical constraints → Adjust scope
- Design insights → Refine user flows
- Market feedback → Validate assumptions
**Expected Output:** Complete PRD with problem, solution, metrics, acceptance criteria, and stakeholder sign-off
**Time Estimate:** 1-2 weeks for comprehensive PRD (iterative process)
### Workflow 4: Quarterly Planning & OKR Setting
**Goal:** Plan quarterly product goals with prioritized initiatives and success metrics
**Steps:**
1. **Review Company OKRs** - Align product goals to business objectives:
- Review CEO/executive OKRs for quarter
- Identify product contribution areas
- Understand strategic priorities
2. **Run Feature Prioritization** - Use RICE for candidate features
```bash
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py q4-candidates.csv --capacity 18
```
3. **Generate OKR Cascade** - Use the OKR cascade generator to create aligned objectives
```bash
python ../../product-team/product-strategist/scripts/okr_cascade_generator.py growth
```
4. **Define Product OKRs** - Set ambitious but achievable goals:
- **Objective**: Qualitative, inspirational (e.g., "Become the easiest platform to onboard")
- **Key Results**: Quantitative, measurable (e.g., "Reduce onboarding time from 30min to 10min")
- **Initiatives**: Features that drive key results
- **Metrics**: How we'll track progress weekly
5. **Capacity Planning** - Allocate team resources:
- Engineering capacity: Person-months available
- Design capacity: UI/UX support needed
- Buffer allocation: 20% for bugs, support, unknowns
- Dependency tracking: External blockers
6. **Risk Assessment** - Identify what could go wrong:
- Technical risks (scalability, performance)
- Market risks (competition, demand)
- Execution risks (dependencies, team velocity)
- Mitigation plans for each risk
7. **Stakeholder Review** - Present quarterly plan:
- OKRs with supporting initiatives
- RICE-justified priorities
- Resource allocation and capacity
- Risks and mitigation strategies
- Success metrics and tracking cadence
8. **Track Progress** - Weekly OKR check-ins:
- Update key result progress
- Adjust priorities if needed
- Communicate blockers early
**Expected Output:** Quarterly OKRs with prioritized roadmap, capacity plan, and risk mitigation
**Time Estimate:** 1 week for quarterly planning (last week of previous quarter)
### Workflow 5: User Research to Personas
**Goal:** Generate data-driven personas from user research to align the team on target users
**Steps:**
1. **Collect Research Data** - Aggregate findings from interviews, surveys, and analytics:
- Interview transcripts and notes
- Survey responses and demographics
- Behavioral analytics (usage patterns, feature adoption)
- Support ticket themes
2. **Review Persona Methodology** - Understand research-backed persona creation
```bash
cat ../../product-team/ux-researcher-designer/references/persona-methodology.md
```
3. **Generate Personas** - Create structured personas from research inputs
```bash
python ../../product-team/ux-researcher-designer/scripts/persona_generator.py research-data.json
```
4. **Map Customer Journeys** - Reference journey mapping guide for each persona
```bash
cat ../../product-team/ux-researcher-designer/references/journey-mapping-guide.md
```
5. **Review Example Personas** - Compare output against proven persona formats
```bash
cat ../../product-team/ux-researcher-designer/references/example-personas.md
```
6. **Validate and Iterate** - Share personas with stakeholders:
- Cross-reference with interview insights from customer_interview_analyzer.py
- Verify demographics and behaviors match real user data
- Update personas quarterly as new research emerges
**Expected Output:** 3-5 data-driven user personas with demographics, goals, pain points, behaviors, and mapped customer journeys
**Time Estimate:** 1-2 weeks (research collection + persona generation + validation)
**Example:**
```bash
# Complete persona generation workflow
python ../../product-team/ux-researcher-designer/scripts/persona_generator.py user-research-q4.json > personas.md
# Cross-reference with interview analysis
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interviews-batch.txt > insights.txt
# Review journey mapping methodology
cat ../../product-team/ux-researcher-designer/references/journey-mapping-guide.md
```
### Workflow 6: Sprint Story Generation
**Goal:** Break epics into INVEST-compliant user stories ready for sprint planning
**Steps:**
1. **Define the Epic** - Structure epic with clear scope and acceptance criteria:
- Business objective and user value
- Functional requirements
- Non-functional requirements (performance, security)
- Dependencies and constraints
2. **Review Story Templates** - Load INVEST-compliant story patterns
```bash
cat ../../product-team/agile-product-owner/references/user-story-templates.md
```
3. **Generate User Stories** - Break the epic into sprint-sized stories
```bash
python ../../product-team/agile-product-owner/scripts/user_story_generator.py epic.yaml
```
4. **Review Sprint Planning Guide** - Ensure stories fit sprint capacity
```bash
cat ../../product-team/agile-product-owner/references/sprint-planning-guide.md
```
5. **Refine and Estimate** - Groom generated stories:
- Verify each story meets INVEST criteria (Independent, Negotiable, Valuable, Estimable, Small, Testable)
- Add story points based on team velocity
- Identify dependencies between stories
- Write acceptance criteria in Given/When/Then format
6. **Prioritize for Sprint** - Use RICE scores to sequence stories
```bash
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py sprint-stories.csv --capacity 8
```
**Expected Output:** Sprint-ready backlog of INVEST-compliant user stories with acceptance criteria, story points, and priority order
**Time Estimate:** 2-4 hours per epic decomposition
**Example:**
```bash
# End-to-end story generation workflow
python ../../product-team/agile-product-owner/scripts/user_story_generator.py onboarding-epic.yaml > stories.md
# Prioritize stories for sprint
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py stories.csv --capacity 8 > sprint-plan.txt
# Review sprint planning best practices
cat ../../product-team/agile-product-owner/references/sprint-planning-guide.md
```
### Workflow 7: Competitive Intelligence
**Goal:** Build competitive analysis matrices to identify market positioning and feature gaps
**Steps:**
1. **Identify Competitors** - Map the competitive landscape:
- Direct competitors (same category, same audience)
- Indirect competitors (different category, same job-to-be-done)
- Emerging threats (startups, adjacent products)
2. **Gather Competitive Data** - Structure competitor information in CSV:
```csv
competitor,feature_1,feature_2,feature_3,pricing,market_share
Competitor A,yes,partial,no,$49/mo,35%
Competitor B,yes,yes,yes,$99/mo,25%
Our Product,yes,no,partial,$39/mo,15%
```
3. **Build Competitive Matrix** - Generate visual comparison
```bash
python ../../product-team/competitive-teardown/scripts/competitive_matrix_builder.py competitors.csv
```
4. **Analyze Gaps** - Identify strategic opportunities:
- Feature parity gaps (what competitors have that we lack)
- Differentiation opportunities (where we can lead)
- Pricing positioning (value vs premium vs budget)
- Underserved segments (unmet user needs)
5. **Feed Into Prioritization** - Use gaps to inform roadmap
```bash
# Add competitive gap features to RICE analysis
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py competitive-features.csv --capacity 20
```
6. **Track Over Time** - Update competitive matrix quarterly:
- Monitor competitor launches and pricing changes
- Re-run matrix builder with updated data
- Adjust positioning strategy based on market shifts
**Expected Output:** Competitive analysis matrix with feature comparison, gap analysis, and prioritized list of competitive features for the roadmap
**Time Estimate:** 1-2 days for initial matrix, 2-4 hours for quarterly updates
**Example:**
```bash
# Full competitive intelligence workflow
python ../../product-team/competitive-teardown/scripts/competitive_matrix_builder.py q4-competitors.csv > competitive-matrix.md
# Prioritize competitive gap features
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py gap-features.csv --capacity 12 > competitive-roadmap.txt
```
## Integration Examples
### Example 1: Weekly Product Review Dashboard
```bash
#!/bin/bash
# product-weekly-review.sh - Automated product metrics summary
echo "📊 Weekly Product Review - $(date +%Y-%m-%d)"
echo "=========================================="
# Current roadmap status
echo ""
echo "🎯 Roadmap Priorities (RICE Sorted):"
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py current-roadmap.csv --capacity 20
# Recent interview insights
echo ""
echo "💡 Latest Customer Insights:"
if [ -f latest-interview.txt ]; then
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py latest-interview.txt
else
echo "No new interviews this week"
fi
# PRD templates available
echo ""
echo "📝 PRD Templates:"
echo "Standard PRD, One-Page PRD, Feature Brief, Agile Epic"
echo "Location: ../../product-team/product-manager-toolkit/references/prd_templates.md"
```
### Example 2: Discovery Sprint Workflow
```bash
# Complete discovery sprint (2 weeks)
echo "🔍 Discovery Sprint - Week 1"
echo "=============================="
# Day 1-2: Conduct interviews
echo "Conducting 5 customer interviews..."
# Day 3-5: Analyze insights
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-001.txt > insights-001.txt
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-002.txt > insights-002.txt
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-003.txt > insights-003.txt
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-004.txt > insights-004.txt
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-005.txt > insights-005.txt
echo ""
echo "🔍 Discovery Sprint - Week 2"
echo "=============================="
# Day 6-8: Prioritize problems and solutions
echo "Creating solution candidates..."
# Day 9-10: RICE prioritization
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py solution-candidates.csv
echo ""
echo "✅ Discovery Complete - Ready for PRD creation"
```
### Example 3: Quarterly Planning Automation
```bash
# Quarterly planning automation script
QUARTER="Q4-2025"
CAPACITY=18 # person-months
echo "📅 $QUARTER Planning"
echo "===================="
# Step 1: Prioritize backlog
echo ""
echo "1. Feature Prioritization:"
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py backlog.csv --capacity $CAPACITY > $QUARTER-roadmap.txt
# Step 2: Extract quick wins
echo ""
echo "2. Quick Wins (Ship First):"
grep "Quick Win" $QUARTER-roadmap.txt
# Step 3: Identify big bets
echo ""
echo "3. Big Bets (Strategic Investments):"
grep "Big Bet" $QUARTER-roadmap.txt
# Step 4: Generate summary
echo ""
echo "4. Quarterly Summary:"
echo "Capacity: $CAPACITY person-months"
echo "Features: $(wc -l < backlog.csv)"
echo "Report: $QUARTER-roadmap.txt"
```
## Success Metrics
**Prioritization Effectiveness:**
- **Decision Speed:** <2 days from backlog review to roadmap commitment
- **Stakeholder Alignment:** >90% stakeholder agreement on priorities
- **RICE Validation:** 80%+ of shipped features match predicted impact
- **Portfolio Balance:** 40% quick wins, 40% big bets, 20% fill-ins
**Discovery Quality:**
- **Interview Volume:** 10-15 interviews per discovery sprint
- **Insight Extraction:** 5-10 high-priority pain points identified
- **Problem Validation:** 70%+ of prioritized problems validated before build
- **Time to Insight:** <1 week from interviews to prioritized problem list
**Requirements Quality:**
- **PRD Completeness:** 100% of PRDs include problem, solution, metrics, acceptance criteria
- **Stakeholder Review:** <3 days average PRD review cycle
- **Engineering Clarity:** >90% of PRDs require no clarification during development
- **Scope Accuracy:** >80% of features ship within original scope estimate
**Business Impact:**
- **Feature Adoption:** >60% of users adopt new features within 30 days
- **Problem Resolution:** >70% reduction in pain point severity post-launch
- **Revenue Impact:** Track revenue/retention lift from prioritized features
- **Development Efficiency:** 30%+ reduction in rework due to clear requirements
## Related Agents
- [cs-agile-product-owner](cs-agile-product-owner.md) - Sprint planning and user story generation
- [cs-product-strategist](cs-product-strategist.md) - OKR cascade and strategic planning
- [cs-ux-researcher](cs-ux-researcher.md) - Persona generation and user research
## References
- **Skill Documentation:** [../../product-team/product-manager-toolkit/SKILL.md](../../product-team/product-manager-toolkit/SKILL.md)
- **Product Domain Guide:** [../../product-team/CLAUDE.md](../../product-team/CLAUDE.md)
- **Agent Development Guide:** [../CLAUDE.md](../CLAUDE.md)
---
**Last Updated:** March 9, 2026
**Status:** Production Ready
**Version:** 2.0
Nhận diện stack công nghệ và sinh cấu hình pipeline CI/CD.
--- name: pipeline description: Detect stack and generate CI/CD pipeline configs. Usage: /pipeline <detect|generate> [options] --- # /pipeline Detect project stack and generate CI/CD pipeline configurations for GitHub Actions or GitLab CI. ## Usage ``` /pipeline detect [--repo <project-dir>] Detect stack, tools, and services /pipeline generate --platform github|gitlab [--repo <project-dir>] Generate pipeline YAML ``` ## Examples ``` /pipeline detect --repo ./my-project /pipeline generate --platform github --repo . /pipeline generate --platform gitlab --repo . ``` ## Scripts - `engineering/ci-cd-pipeline-builder/scripts/stack_detector.py` — Detect stack and tooling (`--repo <path>`, `--format text|json`) - `engineering/ci-cd-pipeline-builder/scripts/pipeline_generator.py` — Generate pipeline YAML (`--platform github|gitlab`, `--repo <path>`, `--input <stack.json>`, `--output <file>`) ## Skill Reference → `engineering/ci-cd-pipeline-builder/SKILL.md`
Đánh giá và so sánh tech stack với phân tích TCO, đánh giá bảo mật, chấm điểm hệ sinh thái và lộ trình di chuyển.
---
name: "tech-stack-evaluator"
description: Technology stack evaluation and comparison with TCO analysis, security assessment, and ecosystem health scoring. Use when comparing frameworks, evaluating technology stacks, calculating total cost of ownership, assessing migration paths, or analyzing ecosystem viability.
---
# Technology Stack Evaluator
Evaluate and compare technologies, frameworks, and cloud providers with data-driven analysis and actionable recommendations.
## Table of Contents
- [Capabilities](#capabilities)
- [Quick Start](#quick-start)
- [Input Formats](#input-formats)
- [Analysis Types](#analysis-types)
- [Scripts](#scripts)
- [References](#references)
---
## Capabilities
| Capability | Description |
|------------|-------------|
| Technology Comparison | Compare frameworks and libraries with weighted scoring |
| TCO Analysis | Calculate 5-year total cost including hidden costs |
| Ecosystem Health | Assess GitHub metrics, npm adoption, community strength |
| Security Assessment | Evaluate vulnerabilities and compliance readiness |
| Migration Analysis | Estimate effort, risks, and timeline for migrations |
| Cloud Comparison | Compare AWS, Azure, GCP for specific workloads |
---
## Quick Start
### Compare Two Technologies
```
Compare React vs Vue for a SaaS dashboard.
Priorities: developer productivity (40%), ecosystem (30%), performance (30%).
```
### Calculate TCO
```
Calculate 5-year TCO for Next.js on Vercel.
Team: 8 developers. Hosting: $2500/month. Growth: 40%/year.
```
### Assess Migration
```
Evaluate migrating from Angular.js to React.
Codebase: 50,000 lines, 200 components. Team: 6 developers.
```
---
## Input Formats
The evaluator accepts three input formats:
**Text** - Natural language queries
```
Compare PostgreSQL vs MongoDB for our e-commerce platform.
```
**YAML** - Structured input for automation
```yaml
comparison:
technologies: ["React", "Vue"]
use_case: "SaaS dashboard"
weights:
ecosystem: 30
performance: 25
developer_experience: 45
```
**JSON** - Programmatic integration
```json
{
"technologies": ["React", "Vue"],
"use_case": "SaaS dashboard"
}
```
---
## Analysis Types
### Quick Comparison (200-300 tokens)
- Weighted scores and recommendation
- Top 3 decision factors
- Confidence level
### Standard Analysis (500-800 tokens)
- Comparison matrix
- TCO overview
- Security summary
### Full Report (1200-1500 tokens)
- All metrics and calculations
- Migration analysis
- Detailed recommendations
---
## Scripts
### stack_comparator.py
Compare technologies with customizable weighted criteria.
```bash
python scripts/stack_comparator.py --help
```
### tco_calculator.py
Calculate total cost of ownership over multi-year projections.
```bash
python scripts/tco_calculator.py --input assets/sample_input_tco.json
```
### ecosystem_analyzer.py
Analyze ecosystem health from GitHub, npm, and community metrics.
```bash
python scripts/ecosystem_analyzer.py --technology react
```
### security_assessor.py
Evaluate security posture and compliance readiness.
```bash
python scripts/security_assessor.py --technology express --compliance soc2,gdpr
```
### migration_analyzer.py
Estimate migration complexity, effort, and risks.
```bash
python scripts/migration_analyzer.py --from angular-1.x --to react
```
---
## References
| Document | Content |
|----------|---------|
| `references/metrics.md` | Detailed scoring algorithms and calculation formulas |
| `references/examples.md` | Input/output examples for all analysis types |
| `references/workflows.md` | Step-by-step evaluation workflows |
---
## Confidence Levels
| Level | Score | Interpretation |
|-------|-------|----------------|
| High | 80-100% | Clear winner, strong data |
| Medium | 50-79% | Trade-offs present, moderate uncertainty |
| Low | < 50% | Close call, limited data |
---
## When to Use
- Comparing frontend/backend frameworks for new projects
- Evaluating cloud providers for specific workloads
- Planning technology migrations with risk assessment
- Calculating build vs. buy decisions with TCO
- Assessing open-source library viability
## When NOT to Use
- Trivial decisions between similar tools (use team preference)
- Mandated technology choices (decision already made)
- Emergency production issues (use monitoring tools)
FILE:assets/expected_output_comparison.json
{
"technologies": {
"PostgreSQL": {
"category_scores": {
"performance": 85.0,
"scalability": 90.0,
"developer_experience": 75.0,
"ecosystem": 95.0,
"learning_curve": 70.0,
"documentation": 90.0,
"community_support": 95.0,
"enterprise_readiness": 95.0
},
"weighted_total": 85.5,
"strengths": ["scalability", "ecosystem", "documentation", "community_support", "enterprise_readiness"],
"weaknesses": ["learning_curve"]
},
"MongoDB": {
"category_scores": {
"performance": 80.0,
"scalability": 95.0,
"developer_experience": 85.0,
"ecosystem": 85.0,
"learning_curve": 80.0,
"documentation": 85.0,
"community_support": 85.0,
"enterprise_readiness": 75.0
},
"weighted_total": 84.5,
"strengths": ["scalability", "developer_experience", "learning_curve"],
"weaknesses": []
}
},
"recommendation": "PostgreSQL",
"confidence": 52.0,
"decision_factors": [
{
"category": "performance",
"importance": "20.0%",
"best_performer": "PostgreSQL",
"score": 85.0
},
{
"category": "scalability",
"importance": "20.0%",
"best_performer": "MongoDB",
"score": 95.0
},
{
"category": "developer_experience",
"importance": "15.0%",
"best_performer": "MongoDB",
"score": 85.0
}
],
"comparison_matrix": [
{
"category": "Performance",
"weight": "20.0%",
"scores": {
"PostgreSQL": "85.0",
"MongoDB": "80.0"
}
},
{
"category": "Scalability",
"weight": "20.0%",
"scores": {
"PostgreSQL": "90.0",
"MongoDB": "95.0"
}
},
{
"category": "WEIGHTED TOTAL",
"weight": "100%",
"scores": {
"PostgreSQL": "85.5",
"MongoDB": "84.5"
}
}
]
}
FILE:assets/sample_input_structured.json
{
"comparison": {
"technologies": [
{
"name": "PostgreSQL",
"performance": {"score": 85},
"scalability": {"score": 90},
"developer_experience": {"score": 75},
"ecosystem": {"score": 95},
"learning_curve": {"score": 70},
"documentation": {"score": 90},
"community_support": {"score": 95},
"enterprise_readiness": {"score": 95}
},
{
"name": "MongoDB",
"performance": {"score": 80},
"scalability": {"score": 95},
"developer_experience": {"score": 85},
"ecosystem": {"score": 85},
"learning_curve": {"score": 80},
"documentation": {"score": 85},
"community_support": {"score": 85},
"enterprise_readiness": {"score": 75}
}
],
"use_case": "SaaS application with complex queries",
"weights": {
"performance": 20,
"scalability": 20,
"developer_experience": 15,
"ecosystem": 15,
"learning_curve": 10,
"documentation": 10,
"community_support": 5,
"enterprise_readiness": 5
}
}
}
FILE:assets/sample_input_tco.json
{
"tco_analysis": {
"technology": "AWS",
"team_size": 10,
"timeline_years": 5,
"initial_costs": {
"licensing": 0,
"training_hours_per_dev": 40,
"developer_hourly_rate": 100,
"training_materials": 1000,
"migration": 50000,
"setup": 10000,
"tooling": 5000
},
"operational_costs": {
"annual_licensing": 0,
"monthly_hosting": 5000,
"annual_support": 20000,
"maintenance_hours_per_dev_monthly": 20
},
"scaling_params": {
"initial_users": 5000,
"annual_growth_rate": 0.30,
"initial_servers": 10,
"cost_per_server_monthly": 300
},
"productivity_factors": {
"productivity_multiplier": 1.2,
"time_to_market_reduction_days": 15,
"avg_feature_time_days": 45,
"avg_feature_value": 15000,
"technical_debt_percentage": 0.12,
"vendor_lock_in_risk": "medium",
"security_incidents_per_year": 0.3,
"avg_security_incident_cost": 30000,
"downtime_hours_per_year": 4,
"downtime_cost_per_hour": 8000,
"annual_turnover_rate": 0.12,
"cost_per_new_hire": 35000
}
}
}
FILE:assets/sample_input_text.json
{
"format": "text",
"input": "Compare React vs Vue for building a SaaS dashboard with real-time collaboration features. Our team has 8 developers, and we need to consider developer experience, ecosystem maturity, and performance."
}
FILE:references/examples.md
# Technology Evaluation Examples
Concrete examples showing input formats and expected outputs.
---
## Table of Contents
- [Quick Comparison Example](#quick-comparison-example)
- [TCO Analysis Example](#tco-analysis-example)
- [Ecosystem Analysis Example](#ecosystem-analysis-example)
- [Migration Assessment Example](#migration-assessment-example)
- [Multi-Technology Comparison](#multi-technology-comparison)
---
## Quick Comparison Example
### Input (Text Format)
```
Compare React vs Vue for building a SaaS dashboard.
Focus on: developer productivity, ecosystem maturity, performance.
```
### Output
```
TECHNOLOGY COMPARISON: React vs Vue for SaaS Dashboard
=======================================================
RECOMMENDATION: React
Confidence: 78% (Medium-High)
COMPARISON MATRIX
-----------------
| Category | Weight | React | Vue |
|----------------------|--------|-------|------|
| Performance | 15% | 82.0 | 85.0 |
| Scalability | 15% | 88.0 | 80.0 |
| Developer Experience | 20% | 85.0 | 90.0 |
| Ecosystem | 15% | 92.0 | 78.0 |
| Learning Curve | 10% | 70.0 | 85.0 |
| Documentation | 10% | 88.0 | 82.0 |
| Community Support | 10% | 90.0 | 75.0 |
| Enterprise Readiness | 5% | 85.0 | 72.0 |
|----------------------|--------|-------|------|
| WEIGHTED TOTAL | 100% | 85.2 | 81.1 |
KEY DECISION FACTORS
--------------------
1. Ecosystem (15%): React leads with 92.0 - larger npm ecosystem
2. Developer Experience (20%): Vue leads with 90.0 - gentler learning curve
3. Community Support (10%): React leads with 90.0 - more Stack Overflow resources
PROS/CONS SUMMARY
-----------------
React:
✓ Excellent ecosystem (92.0/100)
✓ Strong community support (90.0/100)
✓ Excellent scalability (88.0/100)
✗ Steeper learning curve (70.0/100)
Vue:
✓ Excellent developer experience (90.0/100)
✓ Good performance (85.0/100)
✓ Easier learning curve (85.0/100)
✗ Smaller enterprise presence (72.0/100)
```
---
## TCO Analysis Example
### Input (JSON Format)
```json
{
"technology": "Next.js on Vercel",
"team_size": 8,
"timeline_years": 5,
"initial_costs": {
"licensing": 0,
"training_hours_per_dev": 24,
"developer_hourly_rate": 85,
"migration": 15000,
"setup": 5000
},
"operational_costs": {
"monthly_hosting": 2500,
"annual_support": 0,
"maintenance_hours_per_dev_monthly": 16
},
"scaling_params": {
"initial_users": 5000,
"annual_growth_rate": 0.40,
"initial_servers": 3,
"cost_per_server_monthly": 150
}
}
```
### Output
```
TCO ANALYSIS: Next.js on Vercel (5-Year Projection)
====================================================
EXECUTIVE SUMMARY
-----------------
Total TCO: $1,247,320
Net TCO (after productivity gains): $987,320
Average Yearly Cost: $249,464
INITIAL COSTS (One-Time)
------------------------
| Component | Cost |
|----------------|-----------|
| Licensing | $0 |
| Training | $16,820 |
| Migration | $15,000 |
| Setup | $5,000 |
|----------------|-----------|
| TOTAL INITIAL | $36,820 |
OPERATIONAL COSTS (Per Year)
----------------------------
| Year | Hosting | Maintenance | Total |
|------|----------|-------------|-----------|
| 1 | $30,000 | $130,560 | $160,560 |
| 2 | $42,000 | $130,560 | $172,560 |
| 3 | $58,800 | $130,560 | $189,360 |
| 4 | $82,320 | $130,560 | $212,880 |
| 5 | $115,248 | $130,560 | $245,808 |
SCALING ANALYSIS
----------------
User Projections: 5,000 → 7,000 → 9,800 → 13,720 → 19,208
Cost per User: $32.11 → $24.65 → $19.32 → $15.52 → $12.79
Scaling Efficiency: Excellent - economies of scale achieved
KEY COST DRIVERS
----------------
1. Developer maintenance time ($652,800 over 5 years)
2. Infrastructure/hosting ($328,368 over 5 years)
OPTIMIZATION OPPORTUNITIES
--------------------------
• Consider automation to reduce maintenance hours
• Evaluate reserved capacity pricing for hosting
```
---
## Ecosystem Analysis Example
### Input
```yaml
technology: "Svelte"
github:
stars: 78000
forks: 4100
contributors: 680
commits_last_month: 45
avg_issue_response_hours: 36
issue_resolution_rate: 0.72
releases_per_year: 8
active_maintainers: 5
npm:
weekly_downloads: 420000
version: "4.2.8"
dependencies_count: 0
days_since_last_publish: 21
community:
stackoverflow_questions: 8500
job_postings: 1200
tutorials_count: 350
forum_members: 25000
corporate_backing:
type: "community_led"
funding_millions: 0
```
### Output
```
ECOSYSTEM ANALYSIS: Svelte
==========================
OVERALL HEALTH SCORE: 71.5/100
HEALTH SCORE BREAKDOWN
----------------------
| Category | Score | Assessment |
|--------------------|-------|-------------------|
| GitHub Health | 78.0 | Strong |
| npm Health | 72.0 | Good |
| Community Health | 58.0 | Moderate |
| Corporate Backing | 40.0 | Community-led |
| Maintenance Health | 75.0 | Active |
VIABILITY ASSESSMENT
--------------------
Overall: Good - Solid viability with minor concerns
Risk Level: Low-Medium
Strengths:
• Strong GitHub presence with active community
• Active maintenance with responsive issue resolution
• High npm adoption with stable releases
Risks:
• Weak corporate backing - sustainability concerns
• Smaller community - limited enterprise resources
RECOMMENDATION
--------------
Suitable for adoption - monitor corporate backing situation.
Consider alternative with stronger backing for mission-critical enterprise apps.
KEY METRICS
-----------
GitHub: 78,000 stars, 680 contributors, 45 commits/month
npm: 420,000 weekly downloads, v4.2.8
Community: 8,500 SO questions, 1,200 job postings
```
---
## Migration Assessment Example
### Input
```
Assess migration from Angular.js (1.x) to React.
Application: 50,000 lines of code, 200 components.
Team: 6 developers, 2 have React experience.
Timeline: Need completion within 8 months.
```
### Output
```
MIGRATION ASSESSMENT: Angular.js 1.x → React
=============================================
COMPLEXITY SCORE: 7.8/10 (High)
COMPLEXITY BREAKDOWN
--------------------
| Factor | Score | Impact |
|---------------------|-------|-------------|
| Code Changes | 8/10 | 50K LOC |
| Architecture Impact | 9/10 | Complete |
| Data Migration | 5/10 | State only |
| Downtime Risk | 6/10 | Mitigatable |
EFFORT ESTIMATION
-----------------
| Phase | Hours | Duration |
|--------------------------|--------|----------|
| Planning & Prototyping | 240 | 3 weeks |
| Core Migration | 2,400 | 16 weeks |
| Testing & Validation | 480 | 5 weeks |
| Deployment & Monitoring | 160 | 2 weeks |
|--------------------------|--------|----------|
| TOTAL | 3,280 | 26 weeks |
Timeline Assessment: 26 weeks exceeds 8-month target (32 weeks)
Recommendation: Feasible with buffer
RISK ASSESSMENT
---------------
Technical Risks:
• High - Two-way binding to unidirectional data flow
• Medium - Dependency injection to hooks/context
• Medium - Performance regression during transition
Business Risks:
• Medium - Feature freeze during migration
• Low - Customer impact with phased approach
Team Risks:
• High - 4 of 6 developers need React training
• Medium - Learning curve affects velocity initially
MIGRATION STRATEGY
------------------
Recommended Approach: Strangler Fig Pattern
Phase 1 (Weeks 1-8): Strangler setup
• Create React shell application
• Set up micro-frontend routing
• Migrate shared utilities
Phase 2 (Weeks 9-18): Component migration
• Migrate components by feature area
• Maintain Angular.js for untouched areas
• Continuous integration testing
Phase 3 (Weeks 19-24): Completion
• Migrate remaining components
• Remove Angular.js dependencies
• Performance optimization
Phase 4 (Weeks 25-26): Stabilization
• Production deployment
• Monitoring and fixes
```
---
## Multi-Technology Comparison
### Input
```
Compare: React, Vue, Svelte, Angular for enterprise SaaS.
Use case: Large team (20+ developers), complex state management.
Generate comprehensive decision matrix.
```
### Output
```
MULTI-TECHNOLOGY COMPARISON: Enterprise SaaS Platform
======================================================
RECOMMENDATION: React
Confidence: 72% (Medium)
Runner-up: Angular (close second for enterprise)
DECISION MATRIX
---------------
| Category | Wt | React | Vue | Svelte | Angular |
|----------------------|------|-------|------|--------|---------|
| Performance | 15% | 82 | 85 | 95 | 78 |
| Scalability | 15% | 90 | 82 | 75 | 92 |
| Developer Experience | 20% | 85 | 90 | 88 | 75 |
| Ecosystem | 15% | 95 | 80 | 65 | 88 |
| Learning Curve | 10% | 70 | 85 | 80 | 60 |
| Documentation | 10% | 90 | 85 | 75 | 92 |
| Community Support | 10% | 92 | 78 | 55 | 85 |
| Enterprise Readiness | 5% | 88 | 72 | 50 | 95 |
|----------------------|------|-------|------|--------|---------|
| WEIGHTED TOTAL | 100% | 86.3 | 83.1 | 76.2 | 83.0 |
FRAMEWORK PROFILES
------------------
React: Best for large ecosystem, hiring pool
Angular: Best for enterprise structure, TypeScript-first
Vue: Best for developer experience, gradual adoption
Svelte: Best for performance, smaller bundles
RECOMMENDATION RATIONALE
------------------------
For 20+ developer team with complex state management:
1. React (Recommended)
• Largest talent pool for hiring
• Extensive enterprise libraries (Redux, React Query)
• Meta backing ensures long-term support
• Most Stack Overflow resources
2. Angular (Strong Alternative)
• Built-in structure for large teams
• TypeScript-first reduces bugs
• Comprehensive CLI and tooling
• Google enterprise backing
3. Vue (Consider for DX)
• Excellent documentation
• Easier onboarding
• Growing enterprise adoption
• Consider if DX is top priority
4. Svelte (Not Recommended for This Use Case)
• Smaller ecosystem for enterprise
• Limited hiring pool
• State management options less mature
• Better for smaller teams/projects
```
FILE:references/metrics.md
# Technology Evaluation Metrics
Detailed metrics and calculations used in technology stack evaluation.
---
## Table of Contents
- [Scoring and Comparison](#scoring-and-comparison)
- [Financial Calculations](#financial-calculations)
- [Ecosystem Health Metrics](#ecosystem-health-metrics)
- [Security Metrics](#security-metrics)
- [Migration Metrics](#migration-metrics)
- [Performance Benchmarks](#performance-benchmarks)
---
## Scoring and Comparison
### Technology Comparison Matrix
| Metric | Scale | Description |
|--------|-------|-------------|
| Feature Completeness | 0-100 | Coverage of required features |
| Learning Curve | Easy/Medium/Hard | Time to developer proficiency |
| Developer Experience | 0-100 | Tooling, debugging, workflow quality |
| Documentation Quality | 0-10 | Completeness, clarity, examples |
### Weighted Scoring Algorithm
The comparator uses normalized weighted scoring:
```python
# Default category weights (sum to 100%)
weights = {
"performance": 15,
"scalability": 15,
"developer_experience": 20,
"ecosystem": 15,
"learning_curve": 10,
"documentation": 10,
"community_support": 10,
"enterprise_readiness": 5
}
# Final score calculation
weighted_score = sum(category_score * weight / 100 for each category)
```
### Confidence Scoring
Confidence is calculated based on score gap between top options:
| Score Gap | Confidence Level |
|-----------|------------------|
| < 5 points | Low (40-50%) |
| 5-15 points | Medium (50-70%) |
| > 15 points | High (70-100%) |
---
## Financial Calculations
### TCO Components
**Initial Costs (One-Time)**
- Licensing fees
- Training: `team_size * hours_per_dev * hourly_rate + materials`
- Migration costs
- Setup and tooling
**Operational Costs (Annual)**
- Licensing renewals
- Hosting: `base_cost * (1 + growth_rate)^(year - 1)`
- Support contracts
- Maintenance: `team_size * hours_per_dev_monthly * hourly_rate * 12`
**Scaling Costs**
- Infrastructure: `servers * cost_per_server * 12`
- Cost per user: `total_yearly_cost / user_count`
### ROI Calculations
```
productivity_value = additional_features_per_year * avg_feature_value
net_tco = total_cost - (productivity_value * years)
roi_percentage = (benefits - costs) / costs * 100
```
### Cost Per Metric Reference
| Metric | Description |
|--------|-------------|
| Cost per user | Monthly or yearly per active user |
| Cost per API request | Average cost per 1000 requests |
| Cost per GB | Storage and transfer costs |
| Cost per compute hour | Processing time costs |
---
## Ecosystem Health Metrics
### GitHub Health Score (0-100)
| Metric | Max Points | Thresholds |
|--------|------------|------------|
| Stars | 30 | 50K+: 30, 20K+: 25, 10K+: 20, 5K+: 15, 1K+: 10 |
| Forks | 20 | 10K+: 20, 5K+: 15, 2K+: 12, 1K+: 10 |
| Contributors | 20 | 500+: 20, 200+: 15, 100+: 12, 50+: 10 |
| Commits/month | 30 | 100+: 30, 50+: 25, 25+: 20, 10+: 15 |
### npm Health Score (0-100)
| Metric | Max Points | Thresholds |
|--------|------------|------------|
| Weekly downloads | 40 | 1M+: 40, 500K+: 35, 100K+: 30, 50K+: 25, 10K+: 20 |
| Major version | 20 | v5+: 20, v3+: 15, v1+: 10 |
| Dependencies | 20 | ≤10: 20, ≤25: 15, ≤50: 10 (fewer is better) |
| Days since publish | 20 | ≤30: 20, ≤90: 15, ≤180: 10, ≤365: 5 |
### Community Health Score (0-100)
| Metric | Max Points | Thresholds |
|--------|------------|------------|
| Stack Overflow questions | 25 | 50K+: 25, 20K+: 20, 10K+: 15, 5K+: 10 |
| Job postings | 25 | 5K+: 25, 2K+: 20, 1K+: 15, 500+: 10 |
| Tutorials | 25 | 1K+: 25, 500+: 20, 200+: 15, 100+: 10 |
| Forum/Discord members | 25 | 50K+: 25, 20K+: 20, 10K+: 15, 5K+: 10 |
### Corporate Backing Score
| Backing Type | Score |
|--------------|-------|
| Major tech company (Google, Microsoft, Meta) | 100 |
| Established company (Vercel, HashiCorp) | 80 |
| Funded startup | 60 |
| Community-led (strong community) | 40 |
| Individual maintainers | 20 |
---
## Security Metrics
### Security Scoring Components
| Metric | Description |
|--------|-------------|
| CVE Count (12 months) | Known vulnerabilities in last year |
| CVE Count (3 years) | Longer-term vulnerability history |
| Severity Distribution | Critical/High/Medium/Low counts |
| Patch Frequency | Average days to patch vulnerabilities |
### Compliance Readiness Levels
| Level | Score Range | Description |
|-------|-------------|-------------|
| Ready | 90-100% | Meets compliance requirements |
| Mostly Ready | 70-89% | Minor gaps to address |
| Partial | 50-69% | Significant work needed |
| Not Ready | < 50% | Major gaps exist |
### Compliance Framework Coverage
**GDPR**
- Data privacy features
- Consent management
- Data portability
- Right to deletion
**SOC2**
- Access controls
- Encryption at rest/transit
- Audit logging
- Change management
**HIPAA**
- PHI handling
- Encryption standards
- Access controls
- Audit trails
---
## Migration Metrics
### Complexity Scoring (1-10 Scale)
| Factor | Weight | Description |
|--------|--------|-------------|
| Code Changes | 30% | Lines of code affected |
| Architecture Impact | 25% | Breaking changes, API compatibility |
| Data Migration | 25% | Schema changes, data transformation |
| Downtime Requirements | 20% | Zero-downtime possible vs planned outage |
### Effort Estimation
| Phase | Components |
|-------|------------|
| Development | Hours per component * complexity factor |
| Testing | Unit + integration + E2E hours |
| Training | Team size * learning curve hours |
| Buffer | 20-30% for unknowns |
### Risk Assessment Matrix
| Risk Category | Factors Evaluated |
|---------------|-------------------|
| Technical | API incompatibilities, performance regressions |
| Business | Downtime impact, feature parity gaps |
| Team | Learning curve, skill gaps |
---
## Performance Benchmarks
### Throughput/Latency Metrics
| Metric | Description |
|--------|-------------|
| RPS | Requests per second |
| Avg Response Time | Mean response latency (ms) |
| P95 Latency | 95th percentile response time |
| P99 Latency | 99th percentile response time |
| Concurrent Users | Maximum simultaneous connections |
### Resource Usage Metrics
| Metric | Unit |
|--------|------|
| Memory | MB/GB per instance |
| CPU | Utilization percentage |
| Storage | GB required |
| Network | Bandwidth MB/s |
### Scalability Characteristics
| Type | Description |
|------|-------------|
| Horizontal | Add more instances, efficiency factor |
| Vertical | CPU/memory limits per instance |
| Cost per Performance | Dollar per 1000 RPS |
| Scaling Inflection | Point where cost efficiency changes |
FILE:references/workflows.md
# Technology Evaluation Workflows
Step-by-step workflows for common evaluation scenarios.
---
## Table of Contents
- [Framework Comparison Workflow](#framework-comparison-workflow)
- [TCO Analysis Workflow](#tco-analysis-workflow)
- [Migration Assessment Workflow](#migration-assessment-workflow)
- [Security Evaluation Workflow](#security-evaluation-workflow)
- [Cloud Provider Selection Workflow](#cloud-provider-selection-workflow)
---
## Framework Comparison Workflow
Use this workflow when comparing frontend/backend frameworks or libraries.
### Step 1: Define Requirements
1. Identify the use case:
- What type of application? (SaaS, e-commerce, real-time, etc.)
- What scale? (users, requests, data volume)
- What team size and skill level?
2. Set priorities (weights must sum to 100%):
- Performance: ____%
- Scalability: ____%
- Developer Experience: ____%
- Ecosystem: ____%
- Learning Curve: ____%
- Other: ____%
3. List constraints:
- Budget limitations
- Timeline requirements
- Compliance needs
- Existing infrastructure
### Step 2: Run Comparison
```bash
python scripts/stack_comparator.py \
--technologies "React,Vue,Angular" \
--use-case "enterprise-saas" \
--weights "performance:20,ecosystem:25,scalability:20,developer_experience:35"
```
### Step 3: Analyze Results
1. Review weighted total scores
2. Check confidence level (High/Medium/Low)
3. Examine strengths and weaknesses for each option
4. Review decision factors
### Step 4: Validate Recommendation
1. Match recommendation to your constraints
2. Consider team skills and hiring market
3. Evaluate ecosystem for your specific needs
4. Check corporate backing and long-term viability
### Step 5: Document Decision
Record:
- Final selection with rationale
- Trade-offs accepted
- Risks identified
- Mitigation strategies
---
## TCO Analysis Workflow
Use this workflow for comprehensive cost analysis over multiple years.
### Step 1: Gather Cost Data
**Initial Costs:**
- [ ] Licensing fees (if any)
- [ ] Training hours per developer
- [ ] Developer hourly rate
- [ ] Migration costs
- [ ] Setup and tooling costs
**Operational Costs:**
- [ ] Monthly hosting costs
- [ ] Annual support contracts
- [ ] Maintenance hours per developer per month
**Scaling Parameters:**
- [ ] Initial user count
- [ ] Expected annual growth rate
- [ ] Infrastructure scaling approach
### Step 2: Run TCO Calculator
```bash
python scripts/tco_calculator.py \
--input assets/sample_input_tco.json \
--years 5 \
--output tco_report.json
```
### Step 3: Analyze Cost Breakdown
1. Review initial vs. operational costs ratio
2. Examine year-over-year cost growth
3. Check cost per user trends
4. Identify scaling efficiency
### Step 4: Identify Optimization Opportunities
Review:
- Can hosting costs be reduced with reserved pricing?
- Can automation reduce maintenance hours?
- Are there cheaper alternatives for specific components?
### Step 5: Compare Multiple Options
Run TCO analysis for each technology option:
1. Current state (baseline)
2. Option A
3. Option B
Compare:
- 5-year total cost
- Break-even point
- Risk-adjusted costs
---
## Migration Assessment Workflow
Use this workflow when planning technology migrations.
### Step 1: Document Current State
1. Count lines of code
2. List all components/modules
3. Identify dependencies
4. Document current architecture
5. Note existing pain points
### Step 2: Define Target State
1. Target technology/framework
2. Target architecture
3. Expected benefits
4. Success criteria
### Step 3: Assess Team Readiness
- How many developers have target technology experience?
- What training is needed?
- What is the team's capacity during migration?
### Step 4: Run Migration Analysis
```bash
python scripts/migration_analyzer.py \
--from "angular-1.x" \
--to "react" \
--codebase-size 50000 \
--components 200 \
--team-size 6
```
### Step 5: Review Risk Assessment
For each risk category:
1. Identify specific risks
2. Assess probability and impact
3. Define mitigation strategies
4. Assign risk owners
### Step 6: Plan Migration Phases
1. **Phase 1: Foundation**
- Setup new infrastructure
- Create migration utilities
- Train team
2. **Phase 2: Incremental Migration**
- Migrate by feature area
- Maintain parallel systems
- Continuous testing
3. **Phase 3: Completion**
- Remove legacy code
- Optimize performance
- Complete documentation
4. **Phase 4: Stabilization**
- Monitor production
- Address issues
- Gather metrics
### Step 7: Define Rollback Plan
Document:
- Trigger conditions for rollback
- Rollback procedure
- Data recovery steps
- Communication plan
---
## Security Evaluation Workflow
Use this workflow for security and compliance assessment.
### Step 1: Identify Requirements
1. List applicable compliance standards:
- [ ] GDPR
- [ ] SOC2
- [ ] HIPAA
- [ ] PCI-DSS
- [ ] Other: _____
2. Define security priorities:
- Data encryption requirements
- Access control needs
- Audit logging requirements
- Incident response expectations
### Step 2: Gather Security Data
For each technology:
- [ ] CVE count (last 12 months)
- [ ] CVE count (last 3 years)
- [ ] Severity distribution
- [ ] Average patch time
- [ ] Security features list
### Step 3: Run Security Assessment
```bash
python scripts/security_assessor.py \
--technology "express-js" \
--compliance "soc2,gdpr" \
--output security_report.json
```
### Step 4: Analyze Results
Review:
1. Overall security score
2. Vulnerability trends
3. Patch responsiveness
4. Compliance readiness per standard
### Step 5: Identify Gaps
For each compliance standard:
1. List missing requirements
2. Estimate remediation effort
3. Identify workarounds if available
4. Calculate compliance cost
### Step 6: Make Risk-Based Decision
Consider:
- Acceptable risk level
- Cost of remediation
- Alternative technologies
- Business impact of compliance gaps
---
## Cloud Provider Selection Workflow
Use this workflow for AWS vs Azure vs GCP decisions.
### Step 1: Define Workload Requirements
1. Workload type:
- [ ] Web application
- [ ] API services
- [ ] Data analytics
- [ ] Machine learning
- [ ] IoT
- [ ] Other: _____
2. Resource requirements:
- Compute: ____ instances, ____ cores, ____ GB RAM
- Storage: ____ TB, type (block/object/file)
- Database: ____ type, ____ size
- Network: ____ GB/month transfer
3. Special requirements:
- [ ] GPU/TPU for ML
- [ ] Edge computing
- [ ] Multi-region
- [ ] Specific compliance certifications
### Step 2: Evaluate Feature Availability
For each provider, verify:
- Required services exist
- Service maturity level
- Regional availability
- SLA guarantees
### Step 3: Run Cost Comparison
```bash
python scripts/tco_calculator.py \
--providers "aws,azure,gcp" \
--workload-config workload.json \
--years 3
```
### Step 4: Assess Ecosystem Fit
Consider:
- Team's existing expertise
- Development tooling preferences
- CI/CD integration
- Monitoring and observability tools
### Step 5: Evaluate Vendor Lock-in
For each provider:
1. List proprietary services you'll use
2. Estimate migration cost if switching
3. Identify portable alternatives
4. Calculate lock-in risk score
### Step 6: Make Final Selection
Weight factors:
- Cost: ____%
- Features: ____%
- Team expertise: ____%
- Lock-in risk: ____%
- Support quality: ____%
Select provider with highest weighted score.
---
## Best Practices
### For All Evaluations
1. **Document assumptions** - Make all assumptions explicit
2. **Validate data** - Verify metrics from multiple sources
3. **Consider context** - Generic scores may not apply to your situation
4. **Include stakeholders** - Get input from team members who will use the technology
5. **Plan for change** - Technology landscapes evolve; plan for flexibility
### Common Pitfalls to Avoid
1. Over-weighting recent popularity vs. long-term stability
2. Ignoring team learning curve in timeline estimates
3. Underestimating migration complexity
4. Assuming vendor claims are accurate
5. Not accounting for hidden costs (training, hiring, technical debt)
FILE:scripts/ecosystem_analyzer.py
"""
Ecosystem Health Analyzer.
Analyzes technology ecosystem health including community size, maintenance status,
GitHub metrics, npm downloads, and long-term viability assessment.
"""
from typing import Dict, List, Any, Optional
from datetime import datetime, timedelta
class EcosystemAnalyzer:
"""Analyze technology ecosystem health and viability."""
def __init__(self, ecosystem_data: Dict[str, Any]):
"""
Initialize analyzer with ecosystem data.
Args:
ecosystem_data: Dictionary containing GitHub, npm, and community metrics
"""
self.technology = ecosystem_data.get('technology', 'Unknown')
self.github_data = ecosystem_data.get('github', {})
self.npm_data = ecosystem_data.get('npm', {})
self.community_data = ecosystem_data.get('community', {})
self.corporate_backing = ecosystem_data.get('corporate_backing', {})
def calculate_health_score(self) -> Dict[str, float]:
"""
Calculate overall ecosystem health score (0-100).
Returns:
Dictionary of health score components
"""
scores = {
'github_health': self._score_github_health(),
'npm_health': self._score_npm_health(),
'community_health': self._score_community_health(),
'corporate_backing': self._score_corporate_backing(),
'maintenance_health': self._score_maintenance_health()
}
# Calculate weighted average
weights = {
'github_health': 0.25,
'npm_health': 0.20,
'community_health': 0.20,
'corporate_backing': 0.15,
'maintenance_health': 0.20
}
overall = sum(scores[k] * weights[k] for k in scores.keys())
scores['overall_health'] = overall
return scores
def _score_github_health(self) -> float:
"""
Score GitHub repository health.
Returns:
GitHub health score (0-100)
"""
score = 0.0
# Stars (0-30 points)
stars = self.github_data.get('stars', 0)
if stars >= 50000:
score += 30
elif stars >= 20000:
score += 25
elif stars >= 10000:
score += 20
elif stars >= 5000:
score += 15
elif stars >= 1000:
score += 10
else:
score += max(0, stars / 100) # 1 point per 100 stars
# Forks (0-20 points)
forks = self.github_data.get('forks', 0)
if forks >= 10000:
score += 20
elif forks >= 5000:
score += 15
elif forks >= 2000:
score += 12
elif forks >= 1000:
score += 10
else:
score += max(0, forks / 100)
# Contributors (0-20 points)
contributors = self.github_data.get('contributors', 0)
if contributors >= 500:
score += 20
elif contributors >= 200:
score += 15
elif contributors >= 100:
score += 12
elif contributors >= 50:
score += 10
else:
score += max(0, contributors / 5)
# Commit frequency (0-30 points)
commits_last_month = self.github_data.get('commits_last_month', 0)
if commits_last_month >= 100:
score += 30
elif commits_last_month >= 50:
score += 25
elif commits_last_month >= 25:
score += 20
elif commits_last_month >= 10:
score += 15
else:
score += max(0, commits_last_month * 1.5)
return min(100.0, score)
def _score_npm_health(self) -> float:
"""
Score npm package health (if applicable).
Returns:
npm health score (0-100)
"""
if not self.npm_data:
return 50.0 # Neutral score if not applicable
score = 0.0
# Weekly downloads (0-40 points)
weekly_downloads = self.npm_data.get('weekly_downloads', 0)
if weekly_downloads >= 1000000:
score += 40
elif weekly_downloads >= 500000:
score += 35
elif weekly_downloads >= 100000:
score += 30
elif weekly_downloads >= 50000:
score += 25
elif weekly_downloads >= 10000:
score += 20
else:
score += max(0, weekly_downloads / 500)
# Version stability (0-20 points)
version = self.npm_data.get('version', '0.0.1')
major_version = int(version.split('.')[0]) if version else 0
if major_version >= 5:
score += 20
elif major_version >= 3:
score += 15
elif major_version >= 1:
score += 10
else:
score += 5
# Dependencies count (0-20 points, fewer is better)
dependencies = self.npm_data.get('dependencies_count', 50)
if dependencies <= 10:
score += 20
elif dependencies <= 25:
score += 15
elif dependencies <= 50:
score += 10
else:
score += max(0, 20 - (dependencies - 50) / 10)
# Last publish date (0-20 points)
days_since_publish = self.npm_data.get('days_since_last_publish', 365)
if days_since_publish <= 30:
score += 20
elif days_since_publish <= 90:
score += 15
elif days_since_publish <= 180:
score += 10
elif days_since_publish <= 365:
score += 5
else:
score += 0
return min(100.0, score)
def _score_community_health(self) -> float:
"""
Score community health and engagement.
Returns:
Community health score (0-100)
"""
score = 0.0
# Stack Overflow questions (0-25 points)
so_questions = self.community_data.get('stackoverflow_questions', 0)
if so_questions >= 50000:
score += 25
elif so_questions >= 20000:
score += 20
elif so_questions >= 10000:
score += 15
elif so_questions >= 5000:
score += 10
else:
score += max(0, so_questions / 500)
# Job postings (0-25 points)
job_postings = self.community_data.get('job_postings', 0)
if job_postings >= 5000:
score += 25
elif job_postings >= 2000:
score += 20
elif job_postings >= 1000:
score += 15
elif job_postings >= 500:
score += 10
else:
score += max(0, job_postings / 50)
# Tutorials and resources (0-25 points)
tutorials = self.community_data.get('tutorials_count', 0)
if tutorials >= 1000:
score += 25
elif tutorials >= 500:
score += 20
elif tutorials >= 200:
score += 15
elif tutorials >= 100:
score += 10
else:
score += max(0, tutorials / 10)
# Active forums/Discord (0-25 points)
forum_members = self.community_data.get('forum_members', 0)
if forum_members >= 50000:
score += 25
elif forum_members >= 20000:
score += 20
elif forum_members >= 10000:
score += 15
elif forum_members >= 5000:
score += 10
else:
score += max(0, forum_members / 500)
return min(100.0, score)
def _score_corporate_backing(self) -> float:
"""
Score corporate backing strength.
Returns:
Corporate backing score (0-100)
"""
backing_type = self.corporate_backing.get('type', 'none')
scores = {
'major_tech_company': 100, # Google, Microsoft, Meta, etc.
'established_company': 80, # Dedicated company (Vercel, HashiCorp)
'startup_backed': 60, # Funded startup
'community_led': 40, # Strong community, no corporate backing
'none': 20 # Individual maintainers
}
base_score = scores.get(backing_type, 40)
# Adjust for funding
funding = self.corporate_backing.get('funding_millions', 0)
if funding >= 100:
base_score = min(100, base_score + 20)
elif funding >= 50:
base_score = min(100, base_score + 10)
elif funding >= 10:
base_score = min(100, base_score + 5)
return base_score
def _score_maintenance_health(self) -> float:
"""
Score maintenance activity and responsiveness.
Returns:
Maintenance health score (0-100)
"""
score = 0.0
# Issue response time (0-30 points)
avg_response_hours = self.github_data.get('avg_issue_response_hours', 168) # 7 days default
if avg_response_hours <= 24:
score += 30
elif avg_response_hours <= 48:
score += 25
elif avg_response_hours <= 168: # 1 week
score += 20
elif avg_response_hours <= 336: # 2 weeks
score += 10
else:
score += 5
# Issue resolution rate (0-30 points)
resolution_rate = self.github_data.get('issue_resolution_rate', 0.5)
score += resolution_rate * 30
# Release frequency (0-20 points)
releases_per_year = self.github_data.get('releases_per_year', 4)
if releases_per_year >= 12:
score += 20
elif releases_per_year >= 6:
score += 15
elif releases_per_year >= 4:
score += 10
elif releases_per_year >= 2:
score += 5
else:
score += 0
# Active maintainers (0-20 points)
active_maintainers = self.github_data.get('active_maintainers', 1)
if active_maintainers >= 10:
score += 20
elif active_maintainers >= 5:
score += 15
elif active_maintainers >= 3:
score += 10
elif active_maintainers >= 1:
score += 5
else:
score += 0
return min(100.0, score)
def assess_viability(self) -> Dict[str, Any]:
"""
Assess long-term viability of technology.
Returns:
Viability assessment with risk factors
"""
health = self.calculate_health_score()
overall_health = health['overall_health']
# Determine viability level
if overall_health >= 80:
viability = "Excellent - Strong long-term viability"
risk_level = "Low"
elif overall_health >= 65:
viability = "Good - Solid viability with minor concerns"
risk_level = "Low-Medium"
elif overall_health >= 50:
viability = "Moderate - Viable but with notable risks"
risk_level = "Medium"
elif overall_health >= 35:
viability = "Concerning - Significant viability risks"
risk_level = "Medium-High"
else:
viability = "Poor - High risk of abandonment"
risk_level = "High"
# Identify specific risks
risks = self._identify_viability_risks(health)
# Identify strengths
strengths = self._identify_viability_strengths(health)
return {
'overall_viability': viability,
'risk_level': risk_level,
'health_score': overall_health,
'risks': risks,
'strengths': strengths,
'recommendation': self._generate_viability_recommendation(overall_health, risks)
}
def _identify_viability_risks(self, health: Dict[str, float]) -> List[str]:
"""
Identify viability risks from health scores.
Args:
health: Health score components
Returns:
List of identified risks
"""
risks = []
if health['maintenance_health'] < 50:
risks.append("Low maintenance activity - slow issue resolution")
if health['github_health'] < 50:
risks.append("Limited GitHub activity - smaller community")
if health['corporate_backing'] < 40:
risks.append("Weak corporate backing - sustainability concerns")
if health['npm_health'] < 50 and self.npm_data:
risks.append("Low npm adoption - limited ecosystem")
if health['community_health'] < 50:
risks.append("Small community - limited resources and support")
return risks if risks else ["No significant risks identified"]
def _identify_viability_strengths(self, health: Dict[str, float]) -> List[str]:
"""
Identify viability strengths from health scores.
Args:
health: Health score components
Returns:
List of identified strengths
"""
strengths = []
if health['maintenance_health'] >= 70:
strengths.append("Active maintenance with responsive issue resolution")
if health['github_health'] >= 70:
strengths.append("Strong GitHub presence with active community")
if health['corporate_backing'] >= 70:
strengths.append("Strong corporate backing ensures sustainability")
if health['npm_health'] >= 70 and self.npm_data:
strengths.append("High npm adoption with stable releases")
if health['community_health'] >= 70:
strengths.append("Large, active community with extensive resources")
return strengths if strengths else ["Baseline viability maintained"]
def _generate_viability_recommendation(self, health_score: float, risks: List[str]) -> str:
"""
Generate viability recommendation.
Args:
health_score: Overall health score
risks: List of identified risks
Returns:
Recommendation string
"""
if health_score >= 80:
return "Recommended for long-term adoption - strong ecosystem support"
elif health_score >= 65:
return "Suitable for adoption - monitor identified risks"
elif health_score >= 50:
return "Proceed with caution - have contingency plans"
else:
return "Not recommended - consider alternatives with stronger ecosystems"
def generate_ecosystem_report(self) -> Dict[str, Any]:
"""
Generate comprehensive ecosystem report.
Returns:
Complete ecosystem analysis
"""
health = self.calculate_health_score()
viability = self.assess_viability()
return {
'technology': self.technology,
'health_scores': health,
'viability_assessment': viability,
'github_metrics': self._format_github_metrics(),
'npm_metrics': self._format_npm_metrics() if self.npm_data else None,
'community_metrics': self._format_community_metrics()
}
def _format_github_metrics(self) -> Dict[str, Any]:
"""Format GitHub metrics for reporting."""
return {
'stars': f"{self.github_data.get('stars', 0):,}",
'forks': f"{self.github_data.get('forks', 0):,}",
'contributors': f"{self.github_data.get('contributors', 0):,}",
'commits_last_month': self.github_data.get('commits_last_month', 0),
'open_issues': self.github_data.get('open_issues', 0),
'issue_resolution_rate': f"{self.github_data.get('issue_resolution_rate', 0) * 100:.1f}%"
}
def _format_npm_metrics(self) -> Dict[str, Any]:
"""Format npm metrics for reporting."""
return {
'weekly_downloads': f"{self.npm_data.get('weekly_downloads', 0):,}",
'version': self.npm_data.get('version', 'N/A'),
'dependencies': self.npm_data.get('dependencies_count', 0),
'days_since_publish': self.npm_data.get('days_since_last_publish', 0)
}
def _format_community_metrics(self) -> Dict[str, Any]:
"""Format community metrics for reporting."""
return {
'stackoverflow_questions': f"{self.community_data.get('stackoverflow_questions', 0):,}",
'job_postings': f"{self.community_data.get('job_postings', 0):,}",
'tutorials': self.community_data.get('tutorials_count', 0),
'forum_members': f"{self.community_data.get('forum_members', 0):,}"
}
FILE:scripts/format_detector.py
"""
Input Format Detector.
Automatically detects input format (text, YAML, JSON, URLs) and parses
accordingly for technology stack evaluation requests.
"""
from typing import Dict, Any, Optional, Tuple
import json
import re
class FormatDetector:
"""Detect and parse various input formats for stack evaluation."""
def __init__(self, input_data: str):
"""
Initialize format detector with raw input.
Args:
input_data: Raw input string from user
"""
self.raw_input = input_data.strip()
self.detected_format = None
self.parsed_data = None
def detect_format(self) -> str:
"""
Detect the input format.
Returns:
Format type: 'json', 'yaml', 'url', 'text'
"""
# Try JSON first
if self._is_json():
self.detected_format = 'json'
return 'json'
# Try YAML
if self._is_yaml():
self.detected_format = 'yaml'
return 'yaml'
# Check for URLs
if self._contains_urls():
self.detected_format = 'url'
return 'url'
# Default to conversational text
self.detected_format = 'text'
return 'text'
def _is_json(self) -> bool:
"""Check if input is valid JSON."""
try:
json.loads(self.raw_input)
return True
except (json.JSONDecodeError, ValueError):
return False
def _is_yaml(self) -> bool:
"""
Check if input looks like YAML.
Returns:
True if input appears to be YAML format
"""
# YAML indicators
yaml_patterns = [
r'^\s*[\w\-]+\s*:', # Key-value pairs
r'^\s*-\s+', # List items
r':\s*$', # Trailing colons
]
# Must not be JSON
if self._is_json():
return False
# Check for YAML patterns
lines = self.raw_input.split('\n')
yaml_line_count = 0
for line in lines:
for pattern in yaml_patterns:
if re.match(pattern, line):
yaml_line_count += 1
break
# If >50% of lines match YAML patterns, consider it YAML
if len(lines) > 0 and yaml_line_count / len(lines) > 0.5:
return True
return False
def _contains_urls(self) -> bool:
"""Check if input contains URLs."""
url_pattern = r'https?://[^\s]+'
return bool(re.search(url_pattern, self.raw_input))
def parse(self) -> Dict[str, Any]:
"""
Parse input based on detected format.
Returns:
Parsed data dictionary
"""
if self.detected_format is None:
self.detect_format()
if self.detected_format == 'json':
self.parsed_data = self._parse_json()
elif self.detected_format == 'yaml':
self.parsed_data = self._parse_yaml()
elif self.detected_format == 'url':
self.parsed_data = self._parse_urls()
else: # text
self.parsed_data = self._parse_text()
return self.parsed_data
def _parse_json(self) -> Dict[str, Any]:
"""Parse JSON input."""
try:
data = json.loads(self.raw_input)
return self._normalize_structure(data)
except json.JSONDecodeError:
return {'error': 'Invalid JSON', 'raw': self.raw_input}
def _parse_yaml(self) -> Dict[str, Any]:
"""
Parse YAML-like input (simplified, no external dependencies).
Returns:
Parsed dictionary
"""
result = {}
current_section = None
current_list = None
lines = self.raw_input.split('\n')
for line in lines:
stripped = line.strip()
if not stripped or stripped.startswith('#'):
continue
# Key-value pair
if ':' in stripped:
key, value = stripped.split(':', 1)
key = key.strip()
value = value.strip()
# Empty value might indicate nested structure
if not value:
current_section = key
result[current_section] = {}
current_list = None
else:
if current_section:
result[current_section][key] = self._parse_value(value)
else:
result[key] = self._parse_value(value)
# List item
elif stripped.startswith('-'):
item = stripped[1:].strip()
if current_section:
if current_list is None:
current_list = []
result[current_section] = current_list
current_list.append(self._parse_value(item))
return self._normalize_structure(result)
def _parse_value(self, value: str) -> Any:
"""
Parse a value string to appropriate type.
Args:
value: Value string
Returns:
Parsed value (str, int, float, bool)
"""
value = value.strip()
# Boolean
if value.lower() in ['true', 'yes']:
return True
if value.lower() in ['false', 'no']:
return False
# Number
try:
if '.' in value:
return float(value)
else:
return int(value)
except ValueError:
pass
# String (remove quotes if present)
if value.startswith('"') and value.endswith('"'):
return value[1:-1]
if value.startswith("'") and value.endswith("'"):
return value[1:-1]
return value
def _parse_urls(self) -> Dict[str, Any]:
"""Parse URLs from input."""
url_pattern = r'https?://[^\s]+'
urls = re.findall(url_pattern, self.raw_input)
# Categorize URLs
github_urls = [u for u in urls if 'github.com' in u]
npm_urls = [u for u in urls if 'npmjs.com' in u or 'npm.io' in u]
other_urls = [u for u in urls if u not in github_urls and u not in npm_urls]
# Also extract any text context
text_without_urls = re.sub(url_pattern, '', self.raw_input).strip()
result = {
'format': 'url',
'urls': {
'github': github_urls,
'npm': npm_urls,
'other': other_urls
},
'context': text_without_urls
}
return self._normalize_structure(result)
def _parse_text(self) -> Dict[str, Any]:
"""Parse conversational text input."""
text = self.raw_input.lower()
# Extract technologies being compared
technologies = self._extract_technologies(text)
# Extract use case
use_case = self._extract_use_case(text)
# Extract priorities
priorities = self._extract_priorities(text)
# Detect analysis type
analysis_type = self._detect_analysis_type(text)
result = {
'format': 'text',
'technologies': technologies,
'use_case': use_case,
'priorities': priorities,
'analysis_type': analysis_type,
'raw_text': self.raw_input
}
return self._normalize_structure(result)
def _extract_technologies(self, text: str) -> list:
"""
Extract technology names from text.
Args:
text: Lowercase text
Returns:
List of identified technologies
"""
# Common technologies pattern
tech_keywords = [
'react', 'vue', 'angular', 'svelte', 'next.js', 'nuxt.js',
'node.js', 'python', 'java', 'go', 'rust', 'ruby',
'postgresql', 'postgres', 'mysql', 'mongodb', 'redis',
'aws', 'azure', 'gcp', 'google cloud',
'docker', 'kubernetes', 'k8s',
'express', 'fastapi', 'django', 'flask', 'spring boot'
]
found = []
for tech in tech_keywords:
if tech in text:
# Normalize names
normalized = {
'postgres': 'PostgreSQL',
'next.js': 'Next.js',
'nuxt.js': 'Nuxt.js',
'node.js': 'Node.js',
'k8s': 'Kubernetes',
'gcp': 'Google Cloud Platform'
}.get(tech, tech.title())
if normalized not in found:
found.append(normalized)
return found if found else ['Unknown']
def _extract_use_case(self, text: str) -> str:
"""
Extract use case description from text.
Args:
text: Lowercase text
Returns:
Use case description
"""
use_case_keywords = {
'real-time': 'Real-time application',
'collaboration': 'Collaboration platform',
'saas': 'SaaS application',
'dashboard': 'Dashboard application',
'api': 'API-heavy application',
'data-intensive': 'Data-intensive application',
'e-commerce': 'E-commerce platform',
'enterprise': 'Enterprise application'
}
for keyword, description in use_case_keywords.items():
if keyword in text:
return description
return 'General purpose application'
def _extract_priorities(self, text: str) -> list:
"""
Extract priority criteria from text.
Args:
text: Lowercase text
Returns:
List of priorities
"""
priority_keywords = {
'performance': 'Performance',
'scalability': 'Scalability',
'developer experience': 'Developer experience',
'ecosystem': 'Ecosystem',
'learning curve': 'Learning curve',
'cost': 'Cost',
'security': 'Security',
'compliance': 'Compliance'
}
priorities = []
for keyword, priority in priority_keywords.items():
if keyword in text:
priorities.append(priority)
return priorities if priorities else ['Developer experience', 'Performance']
def _detect_analysis_type(self, text: str) -> str:
"""
Detect type of analysis requested.
Args:
text: Lowercase text
Returns:
Analysis type
"""
type_keywords = {
'migration': 'migration_analysis',
'migrate': 'migration_analysis',
'tco': 'tco_analysis',
'total cost': 'tco_analysis',
'security': 'security_analysis',
'compliance': 'security_analysis',
'compare': 'comparison',
'vs': 'comparison',
'evaluate': 'evaluation'
}
for keyword, analysis_type in type_keywords.items():
if keyword in text:
return analysis_type
return 'comparison' # Default
def _normalize_structure(self, data: Dict[str, Any]) -> Dict[str, Any]:
"""
Normalize parsed data to standard structure.
Args:
data: Parsed data dictionary
Returns:
Normalized data structure
"""
# Ensure standard keys exist
standard_keys = [
'technologies',
'use_case',
'priorities',
'analysis_type',
'format'
]
normalized = data.copy()
for key in standard_keys:
if key not in normalized:
# Set defaults
defaults = {
'technologies': [],
'use_case': 'general',
'priorities': [],
'analysis_type': 'comparison',
'format': self.detected_format or 'unknown'
}
normalized[key] = defaults.get(key)
return normalized
def get_format_info(self) -> Dict[str, Any]:
"""
Get information about detected format.
Returns:
Format detection metadata
"""
return {
'detected_format': self.detected_format,
'input_length': len(self.raw_input),
'line_count': len(self.raw_input.split('\n')),
'parsing_successful': self.parsed_data is not None
}
FILE:scripts/migration_analyzer.py
"""
Migration Path Analyzer.
Analyzes migration complexity, risks, timelines, and strategies for moving
from legacy technology stacks to modern alternatives.
"""
from typing import Dict, List, Any, Optional, Tuple
class MigrationAnalyzer:
"""Analyze migration paths and complexity for technology stack changes."""
# Migration complexity factors
COMPLEXITY_FACTORS = [
'code_volume',
'architecture_changes',
'data_migration',
'api_compatibility',
'dependency_changes',
'testing_requirements'
]
def __init__(self, migration_data: Dict[str, Any]):
"""
Initialize migration analyzer with migration parameters.
Args:
migration_data: Dictionary containing source/target technologies and constraints
"""
self.source_tech = migration_data.get('source_technology', 'Unknown')
self.target_tech = migration_data.get('target_technology', 'Unknown')
self.codebase_stats = migration_data.get('codebase_stats', {})
self.constraints = migration_data.get('constraints', {})
self.team_info = migration_data.get('team', {})
def calculate_complexity_score(self) -> Dict[str, Any]:
"""
Calculate overall migration complexity (1-10 scale).
Returns:
Dictionary with complexity scores by factor
"""
scores = {
'code_volume': self._score_code_volume(),
'architecture_changes': self._score_architecture_changes(),
'data_migration': self._score_data_migration(),
'api_compatibility': self._score_api_compatibility(),
'dependency_changes': self._score_dependency_changes(),
'testing_requirements': self._score_testing_requirements()
}
# Calculate weighted average
weights = {
'code_volume': 0.20,
'architecture_changes': 0.25,
'data_migration': 0.20,
'api_compatibility': 0.15,
'dependency_changes': 0.10,
'testing_requirements': 0.10
}
overall = sum(scores[k] * weights[k] for k in scores.keys())
scores['overall_complexity'] = overall
return scores
def _score_code_volume(self) -> float:
"""
Score complexity based on codebase size.
Returns:
Code volume complexity score (1-10)
"""
lines_of_code = self.codebase_stats.get('lines_of_code', 10000)
num_files = self.codebase_stats.get('num_files', 100)
num_components = self.codebase_stats.get('num_components', 50)
# Score based on lines of code (primary factor)
if lines_of_code < 5000:
base_score = 2
elif lines_of_code < 20000:
base_score = 4
elif lines_of_code < 50000:
base_score = 6
elif lines_of_code < 100000:
base_score = 8
else:
base_score = 10
# Adjust for component count
if num_components > 200:
base_score = min(10, base_score + 1)
elif num_components > 500:
base_score = min(10, base_score + 2)
return float(base_score)
def _score_architecture_changes(self) -> float:
"""
Score complexity based on architectural changes.
Returns:
Architecture complexity score (1-10)
"""
arch_change_level = self.codebase_stats.get('architecture_change_level', 'moderate')
scores = {
'minimal': 2, # Same patterns, just different framework
'moderate': 5, # Some pattern changes, similar concepts
'significant': 7, # Different patterns, major refactoring
'complete': 10 # Complete rewrite, different paradigm
}
return float(scores.get(arch_change_level, 5))
def _score_data_migration(self) -> float:
"""
Score complexity based on data migration requirements.
Returns:
Data migration complexity score (1-10)
"""
has_database = self.codebase_stats.get('has_database', True)
if not has_database:
return 1.0
database_size_gb = self.codebase_stats.get('database_size_gb', 10)
schema_changes = self.codebase_stats.get('schema_changes_required', 'minimal')
data_transformation = self.codebase_stats.get('data_transformation_required', False)
# Base score from database size
if database_size_gb < 1:
score = 2
elif database_size_gb < 10:
score = 3
elif database_size_gb < 100:
score = 5
elif database_size_gb < 1000:
score = 7
else:
score = 9
# Adjust for schema changes
schema_adjustments = {
'none': 0,
'minimal': 1,
'moderate': 2,
'significant': 3
}
score += schema_adjustments.get(schema_changes, 1)
# Adjust for data transformation
if data_transformation:
score += 2
return min(10.0, float(score))
def _score_api_compatibility(self) -> float:
"""
Score complexity based on API compatibility.
Returns:
API compatibility complexity score (1-10)
"""
breaking_api_changes = self.codebase_stats.get('breaking_api_changes', 'some')
scores = {
'none': 1, # Fully compatible
'minimal': 3, # Few breaking changes
'some': 5, # Moderate breaking changes
'many': 7, # Significant breaking changes
'complete': 10 # Complete API rewrite
}
return float(scores.get(breaking_api_changes, 5))
def _score_dependency_changes(self) -> float:
"""
Score complexity based on dependency changes.
Returns:
Dependency complexity score (1-10)
"""
num_dependencies = self.codebase_stats.get('num_dependencies', 20)
dependencies_to_replace = self.codebase_stats.get('dependencies_to_replace', 5)
# Score based on replacement percentage
if num_dependencies == 0:
return 1.0
replacement_pct = (dependencies_to_replace / num_dependencies) * 100
if replacement_pct < 10:
return 2.0
elif replacement_pct < 25:
return 4.0
elif replacement_pct < 50:
return 6.0
elif replacement_pct < 75:
return 8.0
else:
return 10.0
def _score_testing_requirements(self) -> float:
"""
Score complexity based on testing requirements.
Returns:
Testing complexity score (1-10)
"""
test_coverage = self.codebase_stats.get('current_test_coverage', 0.5) # 0-1 scale
num_tests = self.codebase_stats.get('num_tests', 100)
# If good test coverage, easier migration (can verify)
if test_coverage >= 0.8:
base_score = 3
elif test_coverage >= 0.6:
base_score = 5
elif test_coverage >= 0.4:
base_score = 7
else:
base_score = 9 # Poor coverage = hard to verify migration
# Large test suites need updates
if num_tests > 500:
base_score = min(10, base_score + 1)
return float(base_score)
def estimate_effort(self) -> Dict[str, Any]:
"""
Estimate migration effort in person-hours and timeline.
Returns:
Dictionary with effort estimates
"""
complexity = self.calculate_complexity_score()
overall_complexity = complexity['overall_complexity']
# Base hours estimation
lines_of_code = self.codebase_stats.get('lines_of_code', 10000)
base_hours = lines_of_code / 50 # 50 lines per hour baseline
# Complexity multiplier
complexity_multiplier = 1 + (overall_complexity / 10)
estimated_hours = base_hours * complexity_multiplier
# Break down by phase
phases = self._calculate_phase_breakdown(estimated_hours)
# Calculate timeline
team_size = self.team_info.get('team_size', 3)
hours_per_week_per_dev = self.team_info.get('hours_per_week', 30) # Account for other work
total_dev_weeks = estimated_hours / (team_size * hours_per_week_per_dev)
total_calendar_weeks = total_dev_weeks * 1.2 # Buffer for blockers
return {
'total_hours': estimated_hours,
'total_person_months': estimated_hours / 160, # 160 hours per person-month
'phases': phases,
'estimated_timeline': {
'dev_weeks': total_dev_weeks,
'calendar_weeks': total_calendar_weeks,
'calendar_months': total_calendar_weeks / 4.33
},
'team_assumptions': {
'team_size': team_size,
'hours_per_week_per_dev': hours_per_week_per_dev
}
}
def _calculate_phase_breakdown(self, total_hours: float) -> Dict[str, Dict[str, float]]:
"""
Calculate effort breakdown by migration phase.
Args:
total_hours: Total estimated hours
Returns:
Hours breakdown by phase
"""
# Standard phase percentages
phase_percentages = {
'planning_and_prototyping': 0.15,
'core_migration': 0.45,
'testing_and_validation': 0.25,
'deployment_and_monitoring': 0.10,
'buffer_and_contingency': 0.05
}
phases = {}
for phase, percentage in phase_percentages.items():
hours = total_hours * percentage
phases[phase] = {
'hours': hours,
'person_weeks': hours / 40,
'percentage': f"{percentage * 100:.0f}%"
}
return phases
def assess_risks(self) -> Dict[str, List[Dict[str, str]]]:
"""
Identify and assess migration risks.
Returns:
Categorized risks with mitigation strategies
"""
complexity = self.calculate_complexity_score()
risks = {
'technical_risks': self._identify_technical_risks(complexity),
'business_risks': self._identify_business_risks(),
'team_risks': self._identify_team_risks()
}
return risks
def _identify_technical_risks(self, complexity: Dict[str, float]) -> List[Dict[str, str]]:
"""
Identify technical risks.
Args:
complexity: Complexity scores
Returns:
List of technical risks with mitigations
"""
risks = []
# API compatibility risks
if complexity['api_compatibility'] >= 7:
risks.append({
'risk': 'Breaking API changes may cause integration failures',
'severity': 'High',
'mitigation': 'Create compatibility layer; implement feature flags for gradual rollout'
})
# Data migration risks
if complexity['data_migration'] >= 7:
risks.append({
'risk': 'Data migration could cause data loss or corruption',
'severity': 'Critical',
'mitigation': 'Implement robust backup strategy; run parallel systems during migration; extensive validation'
})
# Architecture risks
if complexity['architecture_changes'] >= 8:
risks.append({
'risk': 'Major architectural changes increase risk of performance regression',
'severity': 'High',
'mitigation': 'Extensive performance testing; staged rollout; monitoring and alerting'
})
# Testing risks
if complexity['testing_requirements'] >= 7:
risks.append({
'risk': 'Inadequate test coverage may miss critical bugs',
'severity': 'Medium',
'mitigation': 'Improve test coverage before migration; automated regression testing; user acceptance testing'
})
if not risks:
risks.append({
'risk': 'Standard technical risks (bugs, edge cases)',
'severity': 'Low',
'mitigation': 'Standard QA processes and staged rollout'
})
return risks
def _identify_business_risks(self) -> List[Dict[str, str]]:
"""
Identify business risks.
Returns:
List of business risks with mitigations
"""
risks = []
# Downtime risk
downtime_tolerance = self.constraints.get('downtime_tolerance', 'low')
if downtime_tolerance == 'none':
risks.append({
'risk': 'Zero-downtime migration increases complexity and risk',
'severity': 'High',
'mitigation': 'Blue-green deployment; feature flags; gradual traffic migration'
})
# Feature parity risk
risks.append({
'risk': 'New implementation may lack feature parity',
'severity': 'Medium',
'mitigation': 'Comprehensive feature audit; prioritized feature list; clear communication'
})
# Timeline risk
risks.append({
'risk': 'Migration may take longer than estimated',
'severity': 'Medium',
'mitigation': 'Build in 20% buffer; regular progress reviews; scope management'
})
return risks
def _identify_team_risks(self) -> List[Dict[str, str]]:
"""
Identify team-related risks.
Returns:
List of team risks with mitigations
"""
risks = []
# Learning curve
team_experience = self.team_info.get('target_tech_experience', 'low')
if team_experience in ['low', 'none']:
risks.append({
'risk': 'Team lacks experience with target technology',
'severity': 'High',
'mitigation': 'Training program; hire experienced developers; external consulting'
})
# Team size
team_size = self.team_info.get('team_size', 3)
if team_size < 3:
risks.append({
'risk': 'Small team size may extend timeline',
'severity': 'Medium',
'mitigation': 'Consider augmenting team; reduce scope; extend timeline'
})
# Knowledge retention
risks.append({
'risk': 'Loss of institutional knowledge during migration',
'severity': 'Medium',
'mitigation': 'Comprehensive documentation; knowledge sharing sessions; pair programming'
})
return risks
def generate_migration_plan(self) -> Dict[str, Any]:
"""
Generate comprehensive migration plan.
Returns:
Complete migration plan with timeline and recommendations
"""
complexity = self.calculate_complexity_score()
effort = self.estimate_effort()
risks = self.assess_risks()
# Generate phased approach
approach = self._recommend_migration_approach(complexity['overall_complexity'])
# Generate recommendation
recommendation = self._generate_migration_recommendation(complexity, effort, risks)
return {
'source_technology': self.source_tech,
'target_technology': self.target_tech,
'complexity_analysis': complexity,
'effort_estimation': effort,
'risk_assessment': risks,
'recommended_approach': approach,
'overall_recommendation': recommendation,
'success_criteria': self._define_success_criteria()
}
def _recommend_migration_approach(self, complexity_score: float) -> Dict[str, Any]:
"""
Recommend migration approach based on complexity.
Args:
complexity_score: Overall complexity score
Returns:
Recommended approach details
"""
if complexity_score <= 3:
approach = 'direct_migration'
description = 'Direct migration - low complexity allows straightforward migration'
timeline_multiplier = 1.0
elif complexity_score <= 6:
approach = 'phased_migration'
description = 'Phased migration - migrate components incrementally to manage risk'
timeline_multiplier = 1.3
else:
approach = 'strangler_pattern'
description = 'Strangler pattern - gradually replace old system while running in parallel'
timeline_multiplier = 1.5
return {
'approach': approach,
'description': description,
'timeline_multiplier': timeline_multiplier,
'phases': self._generate_approach_phases(approach)
}
def _generate_approach_phases(self, approach: str) -> List[str]:
"""
Generate phase descriptions for migration approach.
Args:
approach: Migration approach type
Returns:
List of phase descriptions
"""
phases = {
'direct_migration': [
'Phase 1: Set up target environment and migrate configuration',
'Phase 2: Migrate codebase and dependencies',
'Phase 3: Migrate data with validation',
'Phase 4: Comprehensive testing',
'Phase 5: Cutover and monitoring'
],
'phased_migration': [
'Phase 1: Identify and prioritize components for migration',
'Phase 2: Migrate non-critical components first',
'Phase 3: Migrate core components with parallel running',
'Phase 4: Migrate critical components with rollback plan',
'Phase 5: Decommission old system'
],
'strangler_pattern': [
'Phase 1: Set up routing layer between old and new systems',
'Phase 2: Implement new features in target technology only',
'Phase 3: Gradually migrate existing features (lowest risk first)',
'Phase 4: Migrate high-risk components last with extensive testing',
'Phase 5: Complete migration and remove routing layer'
]
}
return phases.get(approach, phases['phased_migration'])
def _generate_migration_recommendation(
self,
complexity: Dict[str, float],
effort: Dict[str, Any],
risks: Dict[str, List[Dict[str, str]]]
) -> str:
"""
Generate overall migration recommendation.
Args:
complexity: Complexity analysis
effort: Effort estimation
risks: Risk assessment
Returns:
Recommendation string
"""
overall_complexity = complexity['overall_complexity']
timeline_months = effort['estimated_timeline']['calendar_months']
# Count high/critical severity risks
high_risk_count = sum(
1 for risk_list in risks.values()
for risk in risk_list
if risk['severity'] in ['High', 'Critical']
)
if overall_complexity <= 4 and high_risk_count <= 2:
return f"Recommended - Low complexity migration achievable in {timeline_months:.1f} months with manageable risks"
elif overall_complexity <= 7 and high_risk_count <= 4:
return f"Proceed with caution - Moderate complexity migration requiring {timeline_months:.1f} months and careful risk management"
else:
return f"High risk - Complex migration requiring {timeline_months:.1f} months. Consider: incremental approach, additional resources, or alternative solutions"
def _define_success_criteria(self) -> List[str]:
"""
Define success criteria for migration.
Returns:
List of success criteria
"""
return [
'Feature parity with current system',
'Performance equal or better than current system',
'Zero data loss or corruption',
'All tests passing (unit, integration, E2E)',
'Successful production deployment with <1% error rate',
'Team trained and comfortable with new technology',
'Documentation complete and up-to-date'
]
FILE:scripts/report_generator.py
"""
Report Generator - Context-aware report generation with progressive disclosure.
Generates reports adapted for Claude Desktop (rich markdown) or CLI (terminal-friendly),
with executive summaries and detailed breakdowns on demand.
"""
from typing import Dict, List, Any, Optional
import os
import platform
class ReportGenerator:
"""Generate context-aware technology evaluation reports."""
def __init__(self, report_data: Dict[str, Any], output_context: Optional[str] = None):
"""
Initialize report generator.
Args:
report_data: Complete evaluation data
output_context: 'desktop', 'cli', or None for auto-detect
"""
self.report_data = report_data
self.output_context = output_context or self._detect_context()
def _detect_context(self) -> str:
"""
Detect output context (Desktop vs CLI).
Returns:
Context type: 'desktop' or 'cli'
"""
# Check for Claude Desktop environment variables or indicators
# This is a simplified detection - actual implementation would check for
# Claude Desktop-specific environment variables
if os.getenv('CLAUDE_DESKTOP'):
return 'desktop'
# Check if running in terminal
if os.isatty(1): # stdout is a terminal
return 'cli'
# Default to desktop for rich formatting
return 'desktop'
def generate_executive_summary(self, max_tokens: int = 300) -> str:
"""
Generate executive summary (200-300 tokens).
Args:
max_tokens: Maximum tokens for summary
Returns:
Executive summary markdown
"""
summary_parts = []
# Title
technologies = self.report_data.get('technologies', [])
tech_names = ', '.join(technologies[:3]) # First 3
summary_parts.append(f"# Technology Evaluation: {tech_names}\n")
# Recommendation
recommendation = self.report_data.get('recommendation', {})
rec_text = recommendation.get('text', 'No recommendation available')
confidence = recommendation.get('confidence', 0)
summary_parts.append(f"## Recommendation\n")
summary_parts.append(f"**{rec_text}**\n")
summary_parts.append(f"*Confidence: {confidence:.0f}%*\n")
# Top 3 Pros
pros = recommendation.get('pros', [])[:3]
if pros:
summary_parts.append(f"\n### Top Strengths\n")
for pro in pros:
summary_parts.append(f"- {pro}\n")
# Top 3 Cons
cons = recommendation.get('cons', [])[:3]
if cons:
summary_parts.append(f"\n### Key Concerns\n")
for con in cons:
summary_parts.append(f"- {con}\n")
# Key Decision Factors
decision_factors = self.report_data.get('decision_factors', [])[:3]
if decision_factors:
summary_parts.append(f"\n### Decision Factors\n")
for factor in decision_factors:
category = factor.get('category', 'Unknown')
best = factor.get('best_performer', 'Unknown')
summary_parts.append(f"- **{category.replace('_', ' ').title()}**: {best}\n")
summary_parts.append(f"\n---\n")
summary_parts.append(f"*For detailed analysis, request full report sections*\n")
return ''.join(summary_parts)
def generate_full_report(self, sections: Optional[List[str]] = None) -> str:
"""
Generate complete report with selected sections.
Args:
sections: List of sections to include, or None for all
Returns:
Complete report markdown
"""
if sections is None:
sections = self._get_available_sections()
report_parts = []
# Title and metadata
report_parts.append(self._generate_title())
# Generate each requested section
for section in sections:
section_content = self._generate_section(section)
if section_content:
report_parts.append(section_content)
return '\n\n'.join(report_parts)
def _get_available_sections(self) -> List[str]:
"""
Get list of available report sections.
Returns:
List of section names
"""
sections = ['executive_summary']
if 'comparison_matrix' in self.report_data:
sections.append('comparison_matrix')
if 'tco_analysis' in self.report_data:
sections.append('tco_analysis')
if 'ecosystem_health' in self.report_data:
sections.append('ecosystem_health')
if 'security_assessment' in self.report_data:
sections.append('security_assessment')
if 'migration_analysis' in self.report_data:
sections.append('migration_analysis')
if 'performance_benchmarks' in self.report_data:
sections.append('performance_benchmarks')
return sections
def _generate_title(self) -> str:
"""Generate report title section."""
technologies = self.report_data.get('technologies', [])
tech_names = ' vs '.join(technologies)
use_case = self.report_data.get('use_case', 'General Purpose')
if self.output_context == 'desktop':
return f"""# Technology Stack Evaluation Report
**Technologies**: {tech_names}
**Use Case**: {use_case}
**Generated**: {self._get_timestamp()}
---
"""
else: # CLI
return f"""================================================================================
TECHNOLOGY STACK EVALUATION REPORT
================================================================================
Technologies: {tech_names}
Use Case: {use_case}
Generated: {self._get_timestamp()}
================================================================================
"""
def _generate_section(self, section_name: str) -> Optional[str]:
"""
Generate specific report section.
Args:
section_name: Name of section to generate
Returns:
Section markdown or None
"""
generators = {
'executive_summary': self._section_executive_summary,
'comparison_matrix': self._section_comparison_matrix,
'tco_analysis': self._section_tco_analysis,
'ecosystem_health': self._section_ecosystem_health,
'security_assessment': self._section_security_assessment,
'migration_analysis': self._section_migration_analysis,
'performance_benchmarks': self._section_performance_benchmarks
}
generator = generators.get(section_name)
if generator:
return generator()
return None
def _section_executive_summary(self) -> str:
"""Generate executive summary section."""
return self.generate_executive_summary()
def _section_comparison_matrix(self) -> str:
"""Generate comparison matrix section."""
matrix_data = self.report_data.get('comparison_matrix', [])
if not matrix_data:
return ""
if self.output_context == 'desktop':
return self._render_matrix_desktop(matrix_data)
else:
return self._render_matrix_cli(matrix_data)
def _render_matrix_desktop(self, matrix_data: List[Dict[str, Any]]) -> str:
"""Render comparison matrix for desktop (rich markdown table)."""
parts = ["## Comparison Matrix\n"]
if not matrix_data:
return ""
# Get technology names from first row
tech_names = list(matrix_data[0].get('scores', {}).keys())
# Build table header
header = "| Category | Weight |"
for tech in tech_names:
header += f" {tech} |"
parts.append(header)
# Separator
separator = "|----------|--------|"
separator += "--------|" * len(tech_names)
parts.append(separator)
# Rows
for row in matrix_data:
category = row.get('category', '').replace('_', ' ').title()
weight = row.get('weight', '')
scores = row.get('scores', {})
row_str = f"| {category} | {weight} |"
for tech in tech_names:
score = scores.get(tech, '0.0')
row_str += f" {score} |"
parts.append(row_str)
return '\n'.join(parts)
def _render_matrix_cli(self, matrix_data: List[Dict[str, Any]]) -> str:
"""Render comparison matrix for CLI (ASCII table)."""
parts = ["COMPARISON MATRIX", "=" * 80, ""]
if not matrix_data:
return ""
# Get technology names
tech_names = list(matrix_data[0].get('scores', {}).keys())
# Calculate column widths
category_width = 25
weight_width = 8
score_width = 10
# Header
header = f"{'Category':<{category_width}} {'Weight':<{weight_width}}"
for tech in tech_names:
header += f" {tech[:score_width-1]:<{score_width}}"
parts.append(header)
parts.append("-" * 80)
# Rows
for row in matrix_data:
category = row.get('category', '').replace('_', ' ').title()[:category_width-1]
weight = row.get('weight', '')
scores = row.get('scores', {})
row_str = f"{category:<{category_width}} {weight:<{weight_width}}"
for tech in tech_names:
score = scores.get(tech, '0.0')
row_str += f" {score:<{score_width}}"
parts.append(row_str)
return '\n'.join(parts)
def _section_tco_analysis(self) -> str:
"""Generate TCO analysis section."""
tco_data = self.report_data.get('tco_analysis', {})
if not tco_data:
return ""
parts = ["## Total Cost of Ownership Analysis\n"]
# Summary
total_tco = tco_data.get('total_tco', 0)
timeline = tco_data.get('timeline_years', 5)
avg_yearly = tco_data.get('average_yearly_cost', 0)
parts.append(f"**{timeline}-Year Total**: ,.2f")
parts.append(f"**Average Yearly**: ,.2f\n")
# Cost breakdown
initial = tco_data.get('initial_costs', {})
parts.append(f"### Initial Costs: ,.2f")
# Operational costs
operational = tco_data.get('operational_costs', {})
if operational:
parts.append(f"\n### Operational Costs (Yearly)")
yearly_totals = operational.get('total_yearly', [])
for year, cost in enumerate(yearly_totals, 1):
parts.append(f"- Year {year}: ,.2f")
return '\n'.join(parts)
def _section_ecosystem_health(self) -> str:
"""Generate ecosystem health section."""
ecosystem_data = self.report_data.get('ecosystem_health', {})
if not ecosystem_data:
return ""
parts = ["## Ecosystem Health Analysis\n"]
# Overall score
overall_score = ecosystem_data.get('overall_health', 0)
parts.append(f"**Overall Health Score**: {overall_score:.1f}/100\n")
# Component scores
scores = ecosystem_data.get('health_scores', {})
parts.append("### Health Metrics")
for metric, score in scores.items():
if metric != 'overall_health':
metric_name = metric.replace('_', ' ').title()
parts.append(f"- {metric_name}: {score:.1f}/100")
# Viability assessment
viability = ecosystem_data.get('viability_assessment', {})
if viability:
parts.append(f"\n### Viability: {viability.get('overall_viability', 'Unknown')}")
parts.append(f"**Risk Level**: {viability.get('risk_level', 'Unknown')}")
return '\n'.join(parts)
def _section_security_assessment(self) -> str:
"""Generate security assessment section."""
security_data = self.report_data.get('security_assessment', {})
if not security_data:
return ""
parts = ["## Security & Compliance Assessment\n"]
# Security score
security_score = security_data.get('security_score', {})
overall = security_score.get('overall_security_score', 0)
grade = security_score.get('security_grade', 'N/A')
parts.append(f"**Security Score**: {overall:.1f}/100 (Grade: {grade})\n")
# Compliance
compliance = security_data.get('compliance_assessment', {})
if compliance:
parts.append("### Compliance Readiness")
for standard, assessment in compliance.items():
level = assessment.get('readiness_level', 'Unknown')
pct = assessment.get('readiness_percentage', 0)
parts.append(f"- **{standard}**: {level} ({pct:.0f}%)")
return '\n'.join(parts)
def _section_migration_analysis(self) -> str:
"""Generate migration analysis section."""
migration_data = self.report_data.get('migration_analysis', {})
if not migration_data:
return ""
parts = ["## Migration Path Analysis\n"]
# Complexity
complexity = migration_data.get('complexity_analysis', {})
overall_complexity = complexity.get('overall_complexity', 0)
parts.append(f"**Migration Complexity**: {overall_complexity:.1f}/10\n")
# Effort estimation
effort = migration_data.get('effort_estimation', {})
if effort:
total_hours = effort.get('total_hours', 0)
person_months = effort.get('total_person_months', 0)
timeline = effort.get('estimated_timeline', {})
calendar_months = timeline.get('calendar_months', 0)
parts.append(f"### Effort Estimate")
parts.append(f"- Total Effort: {person_months:.1f} person-months ({total_hours:.0f} hours)")
parts.append(f"- Timeline: {calendar_months:.1f} calendar months")
# Recommended approach
approach = migration_data.get('recommended_approach', {})
if approach:
parts.append(f"\n### Recommended Approach: {approach.get('approach', 'Unknown').replace('_', ' ').title()}")
parts.append(f"{approach.get('description', '')}")
return '\n'.join(parts)
def _section_performance_benchmarks(self) -> str:
"""Generate performance benchmarks section."""
benchmark_data = self.report_data.get('performance_benchmarks', {})
if not benchmark_data:
return ""
parts = ["## Performance Benchmarks\n"]
# Throughput
throughput = benchmark_data.get('throughput', {})
if throughput:
parts.append("### Throughput")
for tech, rps in throughput.items():
parts.append(f"- {tech}: {rps:,} requests/sec")
# Latency
latency = benchmark_data.get('latency', {})
if latency:
parts.append("\n### Latency (P95)")
for tech, ms in latency.items():
parts.append(f"- {tech}: {ms}ms")
return '\n'.join(parts)
def _get_timestamp(self) -> str:
"""Get current timestamp."""
from datetime import datetime
return datetime.now().strftime("%Y-%m-%d %H:%M")
def export_to_file(self, filename: str, sections: Optional[List[str]] = None) -> str:
"""
Export report to file.
Args:
filename: Output filename
sections: Sections to include
Returns:
Path to exported file
"""
report = self.generate_full_report(sections)
with open(filename, 'w', encoding='utf-8') as f:
f.write(report)
return filename
FILE:scripts/security_assessor.py
"""
Security and Compliance Assessor.
Analyzes security vulnerabilities, compliance readiness (GDPR, SOC2, HIPAA),
and overall security posture of technology stacks.
"""
from typing import Dict, List, Any, Optional
from datetime import datetime, timedelta
class SecurityAssessor:
"""Assess security and compliance readiness of technology stacks."""
# Compliance standards mapping
COMPLIANCE_STANDARDS = {
'GDPR': ['data_privacy', 'consent_management', 'data_portability', 'right_to_deletion', 'audit_logging'],
'SOC2': ['access_controls', 'encryption_at_rest', 'encryption_in_transit', 'audit_logging', 'backup_recovery'],
'HIPAA': ['phi_protection', 'encryption_at_rest', 'encryption_in_transit', 'access_controls', 'audit_logging'],
'PCI_DSS': ['payment_data_encryption', 'access_controls', 'network_security', 'vulnerability_management']
}
def __init__(self, security_data: Dict[str, Any]):
"""
Initialize security assessor with security data.
Args:
security_data: Dictionary containing vulnerability and compliance data
"""
self.technology = security_data.get('technology', 'Unknown')
self.vulnerabilities = security_data.get('vulnerabilities', {})
self.security_features = security_data.get('security_features', {})
self.compliance_requirements = security_data.get('compliance_requirements', [])
def calculate_security_score(self) -> Dict[str, Any]:
"""
Calculate overall security score (0-100).
Returns:
Dictionary with security score components
"""
# Component scores
vuln_score = self._score_vulnerabilities()
patch_score = self._score_patch_responsiveness()
features_score = self._score_security_features()
track_record_score = self._score_track_record()
# Weighted average
weights = {
'vulnerability_score': 0.30,
'patch_responsiveness': 0.25,
'security_features': 0.30,
'track_record': 0.15
}
overall = (
vuln_score * weights['vulnerability_score'] +
patch_score * weights['patch_responsiveness'] +
features_score * weights['security_features'] +
track_record_score * weights['track_record']
)
return {
'overall_security_score': overall,
'vulnerability_score': vuln_score,
'patch_responsiveness': patch_score,
'security_features_score': features_score,
'track_record_score': track_record_score,
'security_grade': self._calculate_grade(overall)
}
def _score_vulnerabilities(self) -> float:
"""
Score based on vulnerability count and severity.
Returns:
Vulnerability score (0-100, higher is better)
"""
# Get vulnerability counts by severity (last 12 months)
critical = self.vulnerabilities.get('critical_last_12m', 0)
high = self.vulnerabilities.get('high_last_12m', 0)
medium = self.vulnerabilities.get('medium_last_12m', 0)
low = self.vulnerabilities.get('low_last_12m', 0)
# Calculate weighted vulnerability count
weighted_vulns = (critical * 4) + (high * 2) + (medium * 1) + (low * 0.5)
# Score based on weighted count (fewer is better)
if weighted_vulns == 0:
score = 100
elif weighted_vulns <= 5:
score = 90
elif weighted_vulns <= 10:
score = 80
elif weighted_vulns <= 20:
score = 70
elif weighted_vulns <= 30:
score = 60
elif weighted_vulns <= 50:
score = 50
else:
score = max(0, 50 - (weighted_vulns - 50) / 2)
# Penalty for critical vulnerabilities
if critical > 0:
score = max(0, score - (critical * 10))
return max(0.0, min(100.0, score))
def _score_patch_responsiveness(self) -> float:
"""
Score based on patch response time.
Returns:
Patch responsiveness score (0-100)
"""
# Average days to patch critical vulnerabilities
critical_patch_days = self.vulnerabilities.get('avg_critical_patch_days', 30)
high_patch_days = self.vulnerabilities.get('avg_high_patch_days', 60)
# Score critical patch time (most important)
if critical_patch_days <= 7:
critical_score = 50
elif critical_patch_days <= 14:
critical_score = 40
elif critical_patch_days <= 30:
critical_score = 30
elif critical_patch_days <= 60:
critical_score = 20
else:
critical_score = 10
# Score high severity patch time
if high_patch_days <= 14:
high_score = 30
elif high_patch_days <= 30:
high_score = 25
elif high_patch_days <= 60:
high_score = 20
elif high_patch_days <= 90:
high_score = 15
else:
high_score = 10
# Has active security team
has_security_team = self.vulnerabilities.get('has_security_team', False)
team_score = 20 if has_security_team else 0
total_score = critical_score + high_score + team_score
return min(100.0, total_score)
def _score_security_features(self) -> float:
"""
Score based on built-in security features.
Returns:
Security features score (0-100)
"""
score = 0.0
# Essential features (10 points each)
essential_features = [
'encryption_at_rest',
'encryption_in_transit',
'authentication',
'authorization',
'input_validation'
]
for feature in essential_features:
if self.security_features.get(feature, False):
score += 10
# Advanced features (5 points each)
advanced_features = [
'rate_limiting',
'csrf_protection',
'xss_protection',
'sql_injection_protection',
'audit_logging',
'mfa_support',
'rbac',
'secrets_management',
'security_headers',
'cors_configuration'
]
for feature in advanced_features:
if self.security_features.get(feature, False):
score += 5
return min(100.0, score)
def _score_track_record(self) -> float:
"""
Score based on historical security track record.
Returns:
Track record score (0-100)
"""
score = 50.0 # Start at neutral
# Years since major security incident
years_since_major = self.vulnerabilities.get('years_since_major_incident', 5)
if years_since_major >= 3:
score += 30
elif years_since_major >= 1:
score += 15
else:
score -= 10
# Security certifications
has_certifications = self.vulnerabilities.get('has_security_certifications', False)
if has_certifications:
score += 20
# Bug bounty program
has_bug_bounty = self.vulnerabilities.get('has_bug_bounty_program', False)
if has_bug_bounty:
score += 10
# Security audits
security_audits = self.vulnerabilities.get('security_audits_per_year', 0)
score += min(20, security_audits * 10)
return min(100.0, max(0.0, score))
def _calculate_grade(self, score: float) -> str:
"""
Convert score to letter grade.
Args:
score: Security score (0-100)
Returns:
Letter grade
"""
if score >= 90:
return "A"
elif score >= 80:
return "B"
elif score >= 70:
return "C"
elif score >= 60:
return "D"
else:
return "F"
def assess_compliance(self, standards: List[str] = None) -> Dict[str, Dict[str, Any]]:
"""
Assess compliance readiness for specified standards.
Args:
standards: List of compliance standards to assess (defaults to all required)
Returns:
Dictionary of compliance assessments by standard
"""
if standards is None:
standards = self.compliance_requirements
results = {}
for standard in standards:
if standard not in self.COMPLIANCE_STANDARDS:
results[standard] = {
'readiness': 'Unknown',
'score': 0,
'status': 'Unknown standard'
}
continue
readiness = self._assess_standard_readiness(standard)
results[standard] = readiness
return results
def _assess_standard_readiness(self, standard: str) -> Dict[str, Any]:
"""
Assess readiness for a specific compliance standard.
Args:
standard: Compliance standard name
Returns:
Readiness assessment
"""
required_features = self.COMPLIANCE_STANDARDS[standard]
met_count = 0
total_count = len(required_features)
missing_features = []
for feature in required_features:
if self.security_features.get(feature, False):
met_count += 1
else:
missing_features.append(feature)
# Calculate readiness percentage
readiness_pct = (met_count / total_count * 100) if total_count > 0 else 0
# Determine readiness level
if readiness_pct >= 90:
readiness_level = "Ready"
status = "Compliant - meets all requirements"
elif readiness_pct >= 70:
readiness_level = "Mostly Ready"
status = "Minor gaps - additional configuration needed"
elif readiness_pct >= 50:
readiness_level = "Partial"
status = "Significant work required"
else:
readiness_level = "Not Ready"
status = "Major gaps - extensive implementation needed"
return {
'readiness_level': readiness_level,
'readiness_percentage': readiness_pct,
'status': status,
'features_met': met_count,
'features_required': total_count,
'missing_features': missing_features,
'recommendation': self._generate_compliance_recommendation(readiness_level, missing_features)
}
def _generate_compliance_recommendation(self, readiness_level: str, missing_features: List[str]) -> str:
"""
Generate compliance recommendation.
Args:
readiness_level: Current readiness level
missing_features: List of missing features
Returns:
Recommendation string
"""
if readiness_level == "Ready":
return "Proceed with compliance audit and certification"
elif readiness_level == "Mostly Ready":
return f"Implement missing features: {', '.join(missing_features[:3])}"
elif readiness_level == "Partial":
return f"Significant implementation needed. Start with: {', '.join(missing_features[:3])}"
else:
return "Not recommended without major security enhancements"
def identify_vulnerabilities(self) -> Dict[str, Any]:
"""
Identify and categorize vulnerabilities.
Returns:
Categorized vulnerability report
"""
# Current vulnerabilities
current = {
'critical': self.vulnerabilities.get('critical_last_12m', 0),
'high': self.vulnerabilities.get('high_last_12m', 0),
'medium': self.vulnerabilities.get('medium_last_12m', 0),
'low': self.vulnerabilities.get('low_last_12m', 0)
}
# Historical vulnerabilities (last 3 years)
historical = {
'critical': self.vulnerabilities.get('critical_last_3y', 0),
'high': self.vulnerabilities.get('high_last_3y', 0),
'medium': self.vulnerabilities.get('medium_last_3y', 0),
'low': self.vulnerabilities.get('low_last_3y', 0)
}
# Common vulnerability types
common_types = self.vulnerabilities.get('common_vulnerability_types', [
'SQL Injection',
'XSS',
'CSRF',
'Authentication Issues'
])
return {
'current_vulnerabilities': current,
'total_current': sum(current.values()),
'historical_vulnerabilities': historical,
'total_historical': sum(historical.values()),
'common_types': common_types,
'severity_distribution': self._calculate_severity_distribution(current),
'trend': self._analyze_vulnerability_trend(current, historical)
}
def _calculate_severity_distribution(self, vulnerabilities: Dict[str, int]) -> Dict[str, str]:
"""
Calculate percentage distribution of vulnerability severities.
Args:
vulnerabilities: Vulnerability counts by severity
Returns:
Percentage distribution
"""
total = sum(vulnerabilities.values())
if total == 0:
return {k: "0%" for k in vulnerabilities.keys()}
return {
severity: f"{(count / total * 100):.1f}%"
for severity, count in vulnerabilities.items()
}
def _analyze_vulnerability_trend(self, current: Dict[str, int], historical: Dict[str, int]) -> str:
"""
Analyze vulnerability trend.
Args:
current: Current vulnerabilities
historical: Historical vulnerabilities
Returns:
Trend description
"""
current_total = sum(current.values())
historical_avg = sum(historical.values()) / 3 # 3-year average
if current_total < historical_avg * 0.7:
return "Improving - fewer vulnerabilities than historical average"
elif current_total < historical_avg * 1.2:
return "Stable - consistent with historical average"
else:
return "Concerning - more vulnerabilities than historical average"
def generate_security_report(self) -> Dict[str, Any]:
"""
Generate comprehensive security assessment report.
Returns:
Complete security analysis
"""
security_score = self.calculate_security_score()
compliance = self.assess_compliance()
vulnerabilities = self.identify_vulnerabilities()
# Generate recommendations
recommendations = self._generate_security_recommendations(
security_score,
compliance,
vulnerabilities
)
return {
'technology': self.technology,
'security_score': security_score,
'compliance_assessment': compliance,
'vulnerability_analysis': vulnerabilities,
'recommendations': recommendations,
'overall_risk_level': self._determine_risk_level(security_score['overall_security_score'])
}
def _generate_security_recommendations(
self,
security_score: Dict[str, Any],
compliance: Dict[str, Dict[str, Any]],
vulnerabilities: Dict[str, Any]
) -> List[str]:
"""
Generate security recommendations.
Args:
security_score: Security score data
compliance: Compliance assessment
vulnerabilities: Vulnerability analysis
Returns:
List of recommendations
"""
recommendations = []
# Security score recommendations
if security_score['overall_security_score'] < 70:
recommendations.append("Improve overall security posture - score below acceptable threshold")
# Vulnerability recommendations
current_critical = vulnerabilities['current_vulnerabilities']['critical']
if current_critical > 0:
recommendations.append(f"Address {current_critical} critical vulnerabilities immediately")
# Patch responsiveness
if security_score['patch_responsiveness'] < 60:
recommendations.append("Improve vulnerability patch response time")
# Security features
if security_score['security_features_score'] < 70:
recommendations.append("Implement additional security features (MFA, audit logging, RBAC)")
# Compliance recommendations
for standard, assessment in compliance.items():
if assessment['readiness_level'] == "Not Ready":
recommendations.append(f"{standard}: {assessment['recommendation']}")
if not recommendations:
recommendations.append("Security posture is strong - continue monitoring and maintenance")
return recommendations
def _determine_risk_level(self, security_score: float) -> str:
"""
Determine overall risk level.
Args:
security_score: Overall security score
Returns:
Risk level description
"""
if security_score >= 85:
return "Low Risk - Strong security posture"
elif security_score >= 70:
return "Medium Risk - Acceptable with monitoring"
elif security_score >= 55:
return "High Risk - Security improvements needed"
else:
return "Critical Risk - Not recommended for production use"
FILE:scripts/stack_comparator.py
"""
Technology Stack Comparator - Main comparison engine with weighted scoring.
Provides comprehensive technology comparison with customizable weighted criteria,
feature matrices, and intelligent recommendation generation.
"""
from typing import Dict, List, Any, Optional, Tuple
import json
class StackComparator:
"""Main comparison engine for technology stack evaluation."""
# Feature categories for evaluation
FEATURE_CATEGORIES = [
"performance",
"scalability",
"developer_experience",
"ecosystem",
"learning_curve",
"documentation",
"community_support",
"enterprise_readiness"
]
# Default weights if not provided
DEFAULT_WEIGHTS = {
"performance": 15,
"scalability": 15,
"developer_experience": 20,
"ecosystem": 15,
"learning_curve": 10,
"documentation": 10,
"community_support": 10,
"enterprise_readiness": 5
}
def __init__(self, comparison_data: Dict[str, Any]):
"""
Initialize comparator with comparison data.
Args:
comparison_data: Dictionary containing technologies to compare and criteria
"""
self.technologies = comparison_data.get('technologies', [])
self.use_case = comparison_data.get('use_case', 'general')
self.priorities = comparison_data.get('priorities', {})
self.weights = self._normalize_weights(comparison_data.get('weights', {}))
self.scores = {}
def _normalize_weights(self, custom_weights: Dict[str, float]) -> Dict[str, float]:
"""
Normalize weights to sum to 100.
Args:
custom_weights: User-provided weights
Returns:
Normalized weights dictionary
"""
# Start with defaults
weights = self.DEFAULT_WEIGHTS.copy()
# Override with custom weights
weights.update(custom_weights)
# Normalize to 100
total = sum(weights.values())
if total == 0:
return self.DEFAULT_WEIGHTS
return {k: (v / total) * 100 for k, v in weights.items()}
def score_technology(self, tech_name: str, tech_data: Dict[str, Any]) -> Dict[str, float]:
"""
Score a single technology across all criteria.
Args:
tech_name: Name of technology
tech_data: Technology feature and metric data
Returns:
Dictionary of category scores (0-100 scale)
"""
scores = {}
for category in self.FEATURE_CATEGORIES:
# Get raw score from tech data (0-100 scale)
raw_score = tech_data.get(category, {}).get('score', 50.0)
# Apply use-case specific adjustments
adjusted_score = self._adjust_for_use_case(category, raw_score, tech_name)
scores[category] = min(100.0, max(0.0, adjusted_score))
return scores
def _adjust_for_use_case(self, category: str, score: float, tech_name: str) -> float:
"""
Apply use-case specific adjustments to scores.
Args:
category: Feature category
score: Raw score
tech_name: Technology name
Returns:
Adjusted score
"""
# Use case specific bonuses/penalties
adjustments = {
'real-time': {
'performance': 1.1, # 10% bonus for real-time use cases
'scalability': 1.1
},
'enterprise': {
'enterprise_readiness': 1.2, # 20% bonus
'documentation': 1.1
},
'startup': {
'developer_experience': 1.15,
'learning_curve': 1.1
}
}
# Determine use case type
use_case_lower = self.use_case.lower()
use_case_type = None
for uc_key in adjustments.keys():
if uc_key in use_case_lower:
use_case_type = uc_key
break
# Apply adjustment if applicable
if use_case_type and category in adjustments[use_case_type]:
multiplier = adjustments[use_case_type][category]
return score * multiplier
return score
def calculate_weighted_score(self, category_scores: Dict[str, float]) -> float:
"""
Calculate weighted total score.
Args:
category_scores: Dictionary of category scores
Returns:
Weighted total score (0-100 scale)
"""
total = 0.0
for category, score in category_scores.items():
weight = self.weights.get(category, 0.0) / 100.0 # Convert to decimal
total += score * weight
return total
def compare_technologies(self, tech_data_list: List[Dict[str, Any]]) -> Dict[str, Any]:
"""
Compare multiple technologies and generate recommendation.
Args:
tech_data_list: List of technology data dictionaries
Returns:
Comparison results with scores and recommendation
"""
results = {
'technologies': {},
'recommendation': None,
'confidence': 0.0,
'decision_factors': [],
'comparison_matrix': []
}
# Score each technology
tech_scores = {}
for tech_data in tech_data_list:
tech_name = tech_data.get('name', 'Unknown')
category_scores = self.score_technology(tech_name, tech_data)
weighted_score = self.calculate_weighted_score(category_scores)
tech_scores[tech_name] = {
'category_scores': category_scores,
'weighted_total': weighted_score,
'strengths': self._identify_strengths(category_scores),
'weaknesses': self._identify_weaknesses(category_scores)
}
results['technologies'] = tech_scores
# Generate recommendation
results['recommendation'], results['confidence'] = self._generate_recommendation(tech_scores)
results['decision_factors'] = self._extract_decision_factors(tech_scores)
results['comparison_matrix'] = self._build_comparison_matrix(tech_scores)
return results
def _identify_strengths(self, category_scores: Dict[str, float], threshold: float = 75.0) -> List[str]:
"""
Identify strength categories (scores above threshold).
Args:
category_scores: Category scores dictionary
threshold: Score threshold for strength identification
Returns:
List of strength categories
"""
return [
category for category, score in category_scores.items()
if score >= threshold
]
def _identify_weaknesses(self, category_scores: Dict[str, float], threshold: float = 50.0) -> List[str]:
"""
Identify weakness categories (scores below threshold).
Args:
category_scores: Category scores dictionary
threshold: Score threshold for weakness identification
Returns:
List of weakness categories
"""
return [
category for category, score in category_scores.items()
if score < threshold
]
def _generate_recommendation(self, tech_scores: Dict[str, Dict[str, Any]]) -> Tuple[str, float]:
"""
Generate recommendation and confidence level.
Args:
tech_scores: Technology scores dictionary
Returns:
Tuple of (recommended_technology, confidence_score)
"""
if not tech_scores:
return "Insufficient data", 0.0
# Sort by weighted total score
sorted_techs = sorted(
tech_scores.items(),
key=lambda x: x[1]['weighted_total'],
reverse=True
)
top_tech = sorted_techs[0][0]
top_score = sorted_techs[0][1]['weighted_total']
# Calculate confidence based on score gap
if len(sorted_techs) > 1:
second_score = sorted_techs[1][1]['weighted_total']
score_gap = top_score - second_score
# Confidence increases with score gap
# 0-5 gap: low confidence
# 5-15 gap: medium confidence
# 15+ gap: high confidence
if score_gap < 5:
confidence = 40.0 + (score_gap * 2) # 40-50%
elif score_gap < 15:
confidence = 50.0 + (score_gap - 5) * 2 # 50-70%
else:
confidence = 70.0 + min(score_gap - 15, 30) # 70-100%
else:
confidence = 100.0 # Only one option
return top_tech, min(100.0, confidence)
def _extract_decision_factors(self, tech_scores: Dict[str, Dict[str, Any]]) -> List[Dict[str, Any]]:
"""
Extract key decision factors from comparison.
Args:
tech_scores: Technology scores dictionary
Returns:
List of decision factors with importance weights
"""
factors = []
# Get top weighted categories
sorted_weights = sorted(
self.weights.items(),
key=lambda x: x[1],
reverse=True
)[:3] # Top 3 factors
for category, weight in sorted_weights:
# Get scores for this category across all techs
category_scores = {
tech: scores['category_scores'].get(category, 0.0)
for tech, scores in tech_scores.items()
}
# Find best performer
best_tech = max(category_scores.items(), key=lambda x: x[1])
factors.append({
'category': category,
'importance': f"{weight:.1f}%",
'best_performer': best_tech[0],
'score': best_tech[1]
})
return factors
def _build_comparison_matrix(self, tech_scores: Dict[str, Dict[str, Any]]) -> List[Dict[str, Any]]:
"""
Build comparison matrix for display.
Args:
tech_scores: Technology scores dictionary
Returns:
List of comparison matrix rows
"""
matrix = []
for category in self.FEATURE_CATEGORIES:
row = {
'category': category,
'weight': f"{self.weights.get(category, 0):.1f}%",
'scores': {}
}
for tech_name, scores in tech_scores.items():
category_score = scores['category_scores'].get(category, 0.0)
row['scores'][tech_name] = f"{category_score:.1f}"
matrix.append(row)
# Add weighted totals row
totals_row = {
'category': 'WEIGHTED TOTAL',
'weight': '100%',
'scores': {}
}
for tech_name, scores in tech_scores.items():
totals_row['scores'][tech_name] = f"{scores['weighted_total']:.1f}"
matrix.append(totals_row)
return matrix
def generate_pros_cons(self, tech_name: str, tech_scores: Dict[str, Any]) -> Dict[str, List[str]]:
"""
Generate pros and cons for a technology.
Args:
tech_name: Technology name
tech_scores: Technology scores dictionary
Returns:
Dictionary with 'pros' and 'cons' lists
"""
category_scores = tech_scores['category_scores']
strengths = tech_scores['strengths']
weaknesses = tech_scores['weaknesses']
pros = []
cons = []
# Generate pros from strengths
for strength in strengths[:3]: # Top 3
score = category_scores[strength]
pros.append(f"Excellent {strength.replace('_', ' ')} (score: {score:.1f}/100)")
# Generate cons from weaknesses
for weakness in weaknesses[:3]: # Top 3
score = category_scores[weakness]
cons.append(f"Weaker {weakness.replace('_', ' ')} (score: {score:.1f}/100)")
# Add generic pros/cons if not enough specific ones
if len(pros) == 0:
pros.append(f"Balanced performance across all categories")
if len(cons) == 0:
cons.append(f"No significant weaknesses identified")
return {'pros': pros, 'cons': cons}
FILE:scripts/tco_calculator.py
"""
Total Cost of Ownership (TCO) Calculator.
Calculates comprehensive TCO including licensing, hosting, developer productivity,
scaling costs, and hidden costs over multi-year projections.
"""
from typing import Dict, List, Any, Optional
import json
class TCOCalculator:
"""Calculate Total Cost of Ownership for technology stacks."""
def __init__(self, tco_data: Dict[str, Any]):
"""
Initialize TCO calculator with cost parameters.
Args:
tco_data: Dictionary containing cost parameters and projections
"""
self.technology = tco_data.get('technology', 'Unknown')
self.team_size = tco_data.get('team_size', 5)
self.timeline_years = tco_data.get('timeline_years', 5)
self.initial_costs = tco_data.get('initial_costs', {})
self.operational_costs = tco_data.get('operational_costs', {})
self.scaling_params = tco_data.get('scaling_params', {})
self.productivity_factors = tco_data.get('productivity_factors', {})
def calculate_initial_costs(self) -> Dict[str, float]:
"""
Calculate one-time initial costs.
Returns:
Dictionary of initial cost components
"""
costs = {
'licensing': self.initial_costs.get('licensing', 0.0),
'training': self._calculate_training_costs(),
'migration': self.initial_costs.get('migration', 0.0),
'setup': self.initial_costs.get('setup', 0.0),
'tooling': self.initial_costs.get('tooling', 0.0)
}
costs['total_initial'] = sum(costs.values())
return costs
def _calculate_training_costs(self) -> float:
"""
Calculate training costs based on team size and learning curve.
Returns:
Total training cost
"""
# Default training assumptions
hours_per_developer = self.initial_costs.get('training_hours_per_dev', 40)
avg_hourly_rate = self.initial_costs.get('developer_hourly_rate', 100)
training_materials = self.initial_costs.get('training_materials', 500)
total_hours = self.team_size * hours_per_developer
total_cost = (total_hours * avg_hourly_rate) + training_materials
return total_cost
def calculate_operational_costs(self) -> Dict[str, List[float]]:
"""
Calculate ongoing operational costs per year.
Returns:
Dictionary with yearly cost projections
"""
yearly_costs = {
'licensing': [],
'hosting': [],
'support': [],
'maintenance': [],
'total_yearly': []
}
for year in range(1, self.timeline_years + 1):
# Licensing costs (may include annual fees)
license_cost = self.operational_costs.get('annual_licensing', 0.0)
yearly_costs['licensing'].append(license_cost)
# Hosting costs (scale with growth)
hosting_cost = self._calculate_hosting_cost(year)
yearly_costs['hosting'].append(hosting_cost)
# Support costs
support_cost = self.operational_costs.get('annual_support', 0.0)
yearly_costs['support'].append(support_cost)
# Maintenance costs (developer time)
maintenance_cost = self._calculate_maintenance_cost(year)
yearly_costs['maintenance'].append(maintenance_cost)
# Total for year
year_total = (
license_cost + hosting_cost + support_cost + maintenance_cost
)
yearly_costs['total_yearly'].append(year_total)
return yearly_costs
def _calculate_hosting_cost(self, year: int) -> float:
"""
Calculate hosting costs with growth projection.
Args:
year: Year number (1-indexed)
Returns:
Hosting cost for the year
"""
base_cost = self.operational_costs.get('monthly_hosting', 1000.0) * 12
growth_rate = self.scaling_params.get('annual_growth_rate', 0.20) # 20% default
# Apply compound growth
year_cost = base_cost * ((1 + growth_rate) ** (year - 1))
return year_cost
def _calculate_maintenance_cost(self, year: int) -> float:
"""
Calculate maintenance costs (developer time).
Args:
year: Year number (1-indexed)
Returns:
Maintenance cost for the year
"""
hours_per_dev_per_month = self.operational_costs.get('maintenance_hours_per_dev_monthly', 20)
avg_hourly_rate = self.initial_costs.get('developer_hourly_rate', 100)
monthly_cost = self.team_size * hours_per_dev_per_month * avg_hourly_rate
yearly_cost = monthly_cost * 12
return yearly_cost
def calculate_scaling_costs(self) -> Dict[str, Any]:
"""
Calculate scaling-related costs and metrics.
Returns:
Dictionary with scaling cost analysis
"""
# Project user growth
initial_users = self.scaling_params.get('initial_users', 1000)
annual_growth_rate = self.scaling_params.get('annual_growth_rate', 0.20)
user_projections = []
for year in range(1, self.timeline_years + 1):
users = initial_users * ((1 + annual_growth_rate) ** year)
user_projections.append(int(users))
# Calculate cost per user
operational = self.calculate_operational_costs()
cost_per_user = []
for year_idx, year_cost in enumerate(operational['total_yearly']):
users = user_projections[year_idx]
cost_per_user.append(year_cost / users if users > 0 else 0)
# Infrastructure scaling costs
infra_scaling = self._calculate_infrastructure_scaling()
return {
'user_projections': user_projections,
'cost_per_user': cost_per_user,
'infrastructure_scaling': infra_scaling,
'scaling_efficiency': self._calculate_scaling_efficiency(cost_per_user)
}
def _calculate_infrastructure_scaling(self) -> Dict[str, List[float]]:
"""
Calculate infrastructure scaling costs.
Returns:
Infrastructure cost projections
"""
base_servers = self.scaling_params.get('initial_servers', 5)
cost_per_server_monthly = self.scaling_params.get('cost_per_server_monthly', 200)
growth_rate = self.scaling_params.get('annual_growth_rate', 0.20)
server_costs = []
for year in range(1, self.timeline_years + 1):
servers_needed = base_servers * ((1 + growth_rate) ** year)
yearly_cost = servers_needed * cost_per_server_monthly * 12
server_costs.append(yearly_cost)
return {
'yearly_infrastructure_costs': server_costs
}
def _calculate_scaling_efficiency(self, cost_per_user: List[float]) -> str:
"""
Assess scaling efficiency based on cost per user trend.
Args:
cost_per_user: List of yearly cost per user
Returns:
Efficiency assessment
"""
if len(cost_per_user) < 2:
return "Insufficient data"
# Compare first year to last year
initial = cost_per_user[0]
final = cost_per_user[-1]
if final < initial * 0.8:
return "Excellent - economies of scale achieved"
elif final < initial:
return "Good - improving efficiency over time"
elif final < initial * 1.2:
return "Moderate - costs growing with users"
else:
return "Poor - costs growing faster than users"
def calculate_productivity_impact(self) -> Dict[str, Any]:
"""
Calculate developer productivity impact.
Returns:
Productivity analysis
"""
# Productivity multiplier (1.0 = baseline)
productivity_multiplier = self.productivity_factors.get('productivity_multiplier', 1.0)
# Time to market impact (in days)
ttm_reduction = self.productivity_factors.get('time_to_market_reduction_days', 0)
# Calculate value of faster development
avg_feature_time_days = self.productivity_factors.get('avg_feature_time_days', 30)
features_per_year = 365 / avg_feature_time_days
faster_features_per_year = 365 / max(1, avg_feature_time_days - ttm_reduction)
additional_features = faster_features_per_year - features_per_year
feature_value = self.productivity_factors.get('avg_feature_value', 10000)
yearly_productivity_value = additional_features * feature_value
return {
'productivity_multiplier': productivity_multiplier,
'time_to_market_reduction_days': ttm_reduction,
'additional_features_per_year': additional_features,
'yearly_productivity_value': yearly_productivity_value,
'five_year_productivity_value': yearly_productivity_value * self.timeline_years
}
def calculate_hidden_costs(self) -> Dict[str, float]:
"""
Identify and calculate hidden costs.
Returns:
Dictionary of hidden cost components
"""
costs = {
'technical_debt': self._estimate_technical_debt(),
'vendor_lock_in_risk': self._estimate_vendor_lock_in_cost(),
'security_incidents': self._estimate_security_costs(),
'downtime_risk': self._estimate_downtime_costs(),
'developer_turnover': self._estimate_turnover_costs()
}
costs['total_hidden_costs'] = sum(costs.values())
return costs
def _estimate_technical_debt(self) -> float:
"""
Estimate technical debt accumulation costs.
Returns:
Estimated technical debt cost
"""
# Percentage of development time spent on debt
debt_percentage = self.productivity_factors.get('technical_debt_percentage', 0.15)
yearly_dev_cost = self._calculate_maintenance_cost(1) # Year 1 baseline
# Technical debt accumulates over time
total_debt_cost = 0
for year in range(1, self.timeline_years + 1):
year_debt = yearly_dev_cost * debt_percentage * year # Increases each year
total_debt_cost += year_debt
return total_debt_cost
def _estimate_vendor_lock_in_cost(self) -> float:
"""
Estimate cost of vendor lock-in.
Returns:
Estimated lock-in cost
"""
lock_in_risk = self.productivity_factors.get('vendor_lock_in_risk', 'low')
# Migration cost if switching vendors
migration_cost = self.initial_costs.get('migration', 10000)
risk_multipliers = {
'low': 0.1,
'medium': 0.3,
'high': 0.6
}
multiplier = risk_multipliers.get(lock_in_risk, 0.2)
return migration_cost * multiplier
def _estimate_security_costs(self) -> float:
"""
Estimate potential security incident costs.
Returns:
Estimated security cost
"""
incidents_per_year = self.productivity_factors.get('security_incidents_per_year', 0.5)
avg_incident_cost = self.productivity_factors.get('avg_security_incident_cost', 50000)
total_cost = incidents_per_year * avg_incident_cost * self.timeline_years
return total_cost
def _estimate_downtime_costs(self) -> float:
"""
Estimate downtime costs.
Returns:
Estimated downtime cost
"""
hours_downtime_per_year = self.productivity_factors.get('downtime_hours_per_year', 2)
cost_per_hour = self.productivity_factors.get('downtime_cost_per_hour', 5000)
total_cost = hours_downtime_per_year * cost_per_hour * self.timeline_years
return total_cost
def _estimate_turnover_costs(self) -> float:
"""
Estimate costs from developer turnover.
Returns:
Estimated turnover cost
"""
turnover_rate = self.productivity_factors.get('annual_turnover_rate', 0.15)
cost_per_hire = self.productivity_factors.get('cost_per_new_hire', 30000)
hires_per_year = self.team_size * turnover_rate
total_cost = hires_per_year * cost_per_hire * self.timeline_years
return total_cost
def calculate_total_tco(self) -> Dict[str, Any]:
"""
Calculate complete TCO over the timeline.
Returns:
Comprehensive TCO analysis
"""
initial = self.calculate_initial_costs()
operational = self.calculate_operational_costs()
scaling = self.calculate_scaling_costs()
productivity = self.calculate_productivity_impact()
hidden = self.calculate_hidden_costs()
# Calculate total costs
total_operational = sum(operational['total_yearly'])
total_cost = initial['total_initial'] + total_operational + hidden['total_hidden_costs']
# Adjust for productivity gains
net_cost = total_cost - productivity['five_year_productivity_value']
return {
'technology': self.technology,
'timeline_years': self.timeline_years,
'initial_costs': initial,
'operational_costs': operational,
'scaling_analysis': scaling,
'productivity_impact': productivity,
'hidden_costs': hidden,
'total_tco': total_cost,
'net_tco_after_productivity': net_cost,
'average_yearly_cost': total_cost / self.timeline_years
}
def generate_tco_summary(self) -> Dict[str, Any]:
"""
Generate executive summary of TCO.
Returns:
TCO summary for reporting
"""
tco = self.calculate_total_tco()
return {
'technology': self.technology,
'total_tco': f",.2f",
'net_tco': f",.2f",
'average_yearly': f",.2f",
'initial_investment': f",.2f",
'key_cost_drivers': self._identify_cost_drivers(tco),
'cost_optimization_opportunities': self._identify_optimizations(tco)
}
def _identify_cost_drivers(self, tco: Dict[str, Any]) -> List[str]:
"""
Identify top cost drivers.
Args:
tco: Complete TCO analysis
Returns:
List of top cost drivers
"""
drivers = []
# Check operational costs
operational = tco['operational_costs']
total_hosting = sum(operational['hosting'])
total_maintenance = sum(operational['maintenance'])
if total_hosting > total_maintenance:
drivers.append(f"Infrastructure/hosting ({total_hosting:,.0f})")
else:
drivers.append(f"Developer maintenance time ({total_maintenance:,.0f})")
# Check hidden costs
hidden = tco['hidden_costs']
if hidden['technical_debt'] > 10000:
drivers.append(f"Technical debt ({hidden['technical_debt']:,.0f})")
return drivers[:3] # Top 3
def _identify_optimizations(self, tco: Dict[str, Any]) -> List[str]:
"""
Identify cost optimization opportunities.
Args:
tco: Complete TCO analysis
Returns:
List of optimization suggestions
"""
optimizations = []
# Check scaling efficiency
scaling = tco['scaling_analysis']
if scaling['scaling_efficiency'].startswith('Poor'):
optimizations.append("Improve scaling efficiency - costs growing too fast")
# Check hidden costs
hidden = tco['hidden_costs']
if hidden['technical_debt'] > 20000:
optimizations.append("Address technical debt accumulation")
if hidden['downtime_risk'] > 10000:
optimizations.append("Invest in reliability to reduce downtime costs")
return optimizations
Sinh test, phân tích độ phủ và chạy quy trình phát triển hướng kiểm thử TDD.
--- name: tdd description: Generate tests, analyze coverage, and run TDD workflows. Usage: /tdd <generate|coverage|validate> [options] --- # /tdd Generate tests, analyze coverage, and validate test quality using the TDD Guide skill. ## Usage ``` /tdd generate <file-or-dir> Generate tests for source files /tdd coverage <test-dir> Analyze test coverage and gaps /tdd validate <test-file> Validate test quality (assertions, edge cases) ``` ## Examples ``` /tdd generate src/auth/login.ts /tdd coverage tests/ --threshold 80 /tdd validate tests/auth.test.ts ``` ## Scripts - `engineering-team/tdd-guide/scripts/test_generator.py` — Test case generation (library module) - `engineering-team/tdd-guide/scripts/coverage_analyzer.py` — Coverage analysis (library module) - `engineering-team/tdd-guide/scripts/tdd_workflow.py` — TDD workflow orchestration (library module) - `engineering-team/tdd-guide/scripts/fixture_generator.py` — Test fixture generation (library module) - `engineering-team/tdd-guide/scripts/metrics_calculator.py` — TDD metrics calculation (library module) > **Note:** These scripts are library modules without CLI entry points. Import them in Python or use via the SKILL.md workflow guidance. ## Skill Reference → `engineering-team/tdd-guide/SKILL.md`
Quản lý tiền cho chương trình R&D nội bộ: lập ngân sách nhiều kỳ có chi phí gián tiếp, theo dõi tốc độ đốt tiền và quyết định vốn hóa hay ghi chi phí.
---
name: research-finance
description: Use when managing the money for an internal R&D program or portfolio — building a multi-period program budget with the F&A (indirect) split, tracking burn rate and runway against value-inflection milestones, or routing R&D cost items to a capitalize-vs-expense determination. Every budget output surfaces its assumptions block; capitalize-vs-expense is decision-support only and routes to a named finance owner — it never books an entry or decides accounting treatment. Distinct from finance/financial-analysis (corporate DCF, close, valuation) and research/grants (funding discovery — this manages money already won).
version: 2.9.0
author: claude-code-skills
license: MIT
tags: [research-ops, research-finance, rd-budget, burn-rate, runway, fa-rate, capitalize-vs-expense, portfolio]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# research-finance
Financial management of internal R&D programs and portfolios: program budgeting with F&A, burn/runway tracking, and capitalize-vs-expense routing. Every number ships with its **assumptions block**, and accounting-treatment calls **route to a named finance owner** — this skill never books an entry.
## Purpose
R&D finance partners, program controllers, and operations leads manage money that has already been allocated or raised — not the corporate close, not the next funding round, not finding a grant. This skill structures three recurring decisions:
Three deterministic tools:
1. `program_budget_planner.py` — Builds a multi-period budget from work-package line items, applies the F&A (indirect) rate to an MTDC-style eligible base, and rolls up direct / F&A / fully-loaded cost per period with an explicit assumptions block.
2. `burn_runway_tracker.py` — Computes average + trailing burn, runway in periods/months, and whether each value-inflection milestone is reachable before cash runs out. Flags accelerating burn and below-threshold runway.
3. `capex_vs_opex_router.py` — Scores each R&D cost item against the IAS 38 development-phase criteria (or flags US GAAP ASC 730 expense-as-incurred) and routes it to **CAPITALIZE-CANDIDATE / EXPENSE / FINANCE-OWNER-REVIEW** with a named owner. Never auto-decides.
## When to use
Invoke this skill when:
- You are building or revising an R&D program budget and need the F&A split made explicit.
- A program's runway is in question and you need a milestone-vs-cash read.
- Finance asks whether a development cost can be capitalized and you need a defensible first routing.
- You are preparing a portfolio review and need per-program burn consistency.
**Do NOT use this skill to**: run corporate DCF / valuation / close (use `finance/financial-analysis`), discover or position grants (use `research/grants`), or make the final accounting determination (that is the controller's + auditor's call — this tool only routes).
## Workflow
1. **Lay out the program** — Fill `assets/rd_program_budget_template.md` with work-package lines, categories, and per-period amounts.
2. **Build the budget** — Run `program_budget_planner.py --input program.json --profile {pharma-rd|biotech|medtech|deep-tech|software-rd|university-lab} --fa-rate <negotiated rate>`. Read direct / F&A / fully-loaded rollups + assumptions.
3. **Track burn & runway** — Run `burn_runway_tracker.py --input ledger.json --threshold-months 6`. Read runway + milestone verdicts + flags.
4. **Route accounting treatment** — Run `capex_vs_opex_router.py --input costs.json --standard {ifrs|usgaap}`. Read the per-item routing; send CAPITALIZE-CANDIDATE and FINANCE-OWNER-REVIEW items to the named owner.
5. **Assemble the review** — Combine into a program-finance packet. Every number carries its assumptions; treatment calls carry a named owner.
## Scripts
| Script | Purpose | Profiles |
|---|---|---|
| `scripts/program_budget_planner.py` | Multi-period budget + F&A split + assumptions | pharma-rd, biotech, medtech, deep-tech, software-rd, university-lab |
| `scripts/burn_runway_tracker.py` | Burn, runway, milestone-vs-cash alignment | n/a (ledger-driven) |
| `scripts/capex_vs_opex_router.py` | IAS 38 / ASC 730 routing to named finance owner | pharma-rd, biotech, medtech, deep-tech, software-rd, university-lab |
All three: stdlib-only, `--help`, `--sample`, `--output {human,json}`.
## Onboarding & customization
Run the onboarding questionnaire **once before you start** — it captures your defaults so every tool in this skill is pre-configured. Customization is the point: the answers actually change tool behavior.
```bash
python3 scripts/onboard.py # interactive (also: --defaults, --set key=value, --reset)
python3 scripts/onboard.py --show # see the questions + current effective config
```
Answers are saved to `~/.config/research-ops/research-finance.json` (global) or `./.research-ops/research-finance.json` (`--scope project`) and are read automatically by `config_loader.py`. They set the default R&D-area **profile**, the default **F&A rate**, the **runway alert threshold**, the **accounting standard**, and the named **finance owner** printed on capitalize-vs-expense routing. CLI flags always override saved config; `RESEARCH_OPS_NO_CONFIG=1` ignores it.
**The five questions:** R&D area · F&A rate · runway threshold · accounting standard · finance owner.
## Optimize with autoresearch (opt-in)
This skill ships an **isolated, opt-in** bridge to `engineering/autoresearch-agent`. Only when you ask to "optimize" / "extend runway" / "run a loop" does an autoresearch experiment iteratively improve a program plan against this skill's runway metric. `scripts/ar_evaluator.py` is the ground-truth evaluator; it prints `runway_months: <float>` (higher is better).
```bash
/ar:setup --domain custom --name extend-runway \
--target ledger.json \
--eval "python3 ar_evaluator.py --target ledger.json" \
--metric runway_months --direction higher
/ar:loop custom/extend-runway
```
Isolated: no hard dependency — autoresearch runs only on demand, and the loop edits `ledger.json`, never the evaluator.
## References
- `references/rd_program_finance_canon.md` — IAS 38 (research vs development); ASC 730 + ASC 985-20; Uniform Guidance 2 CFR 200 (F&A); FASB/IFRS capitalization criteria; NICRA basics.
- `references/burn_and_portfolio.md` — Cooper stage-gate; rNPV / real-options for R&D; risk-adjusted portfolio ROI; burn-rate / runway frameworks; milestone-based budgeting.
- `references/indirect_rate_modeling.md` — F&A pool composition (facilities + administration); MTDC base; de minimis 10%; fringe/overhead loading; CAS primer.
## Assumptions
- The F&A rate is the most error-prone input. The planner applies whatever rate you pass; it warns you to confirm it is a negotiated NICRA, not a guess.
- Burn/runway uses the trailing (recent-weighted) burn as the forward run-rate and assumes flat forward spend unless your ledger encodes a ramp.
- The capex router asserts criteria from your input; asserting "technical feasibility" does not make it true — the named finance owner and auditor validate it.
- Profiles annotate context (e.g., "most drug R&D is expensed") but do not change the accounting test.
## Anti-patterns
- **Stating a budget number without its assumptions.** F&A rate, escalation, and base must travel with the number.
- **Auto-deciding capitalize-vs-expense.** This tool routes; the controller (and auditor where required) decides.
- **Using lifetime-average burn for runway.** Recent burn is the honest forward run-rate; averages hide a slowdown or a ramp.
- **Applying F&A to the full base.** Capital equipment, large subaward portions, and certain categories are MTDC-exempt.
- **Confusing this with corporate finance.** Valuation, close, and fundraising live in `finance/`.
## Distinct from
| Sibling / neighbor | Scope | Difference |
|---|---|---|
| `finance/financial-analysis` | Corporate DCF, ratios, close, rolling forecast, SaaS metrics | That is **company-level**; this is **R&D-program-level** |
| `research/grants` | NIH funding discovery + positioning | That **finds funding**; this **manages money already won** |
| `clinical-research` (sibling) | Study design + feasibility + budget gate-check | That **scopes** the study; this **funds + tracks** the program |
| `ra-qm-team` | Regulatory/QM submission | Unrelated — no financial scope |
## Quick examples
```bash
python3 scripts/program_budget_planner.py --sample
python3 scripts/program_budget_planner.py --input program.json --profile university-lab --fa-rate 0.585
python3 scripts/burn_runway_tracker.py --sample --output json
python3 scripts/capex_vs_opex_router.py --sample --standard ifrs
```
The sample budget excludes the sequencer (capital equipment) and CRO subaward from the F&A base; the capex router routes exploratory screening to EXPENSE, a fully-criteria'd pilot line to CAPITALIZE-CANDIDATE, and a partial-criteria software build to FINANCE-OWNER-REVIEW.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-research-ops` or the orchestrator. Recommended answer + canon citation per question. Never bundled.
1. **"Is this spend in the research phase or the development phase — and can you evidence technical feasibility?"**
Recommended: research = expense; development = capitalize-candidate only with feasibility evidence, routed to a named finance owner.
Canon: IAS 38.54-57; ASC 730.
2. **"What F&A / indirect rate are you applying, and is it your negotiated NICRA, a de minimis 10%, or an assumption?"**
Recommended: use the negotiated rate; if assumed, flag it explicitly.
Canon: 2 CFR 200 (Uniform Guidance); NICRA basics.
3. **"What's runway in months at current burn, and does it clear the next value-inflection milestone?"**
Recommended: runway must cover the milestone plus a buffer; surface the gap.
Canon: Cooper stage-gate; SaaS/startup efficiency frameworks (a16z, Bessemer).
4. **"Is portfolio ROI risk-adjusted (rNPV / probability-of-success weighted) or raw NPV?"**
Recommended: risk-adjusted; raw NPV overstates R&D value.
Canon: rNPV drug-development valuation; real-options literature.
5. **"Who is the named finance / controller owner who signs the capitalize-vs-expense treatment?"**
Recommended: name them — this tool recommends, it never books the entry.
Canon: ASC 730 / IAS 38 governance; auditor sign-off requirements.
Walk depth-first. Lock 1-2 before opening 3-5. After all are answered, invoke `program_budget_planner.py` → `burn_runway_tracker.py` → `capex_vs_opex_router.py`.
FILE:assets/rd_program_budget_template.md
# R&D Program Budget — Template
> Fill this before running `program_budget_planner.py`. Every number must travel with its
> assumptions. Capitalize-vs-expense calls route to a named finance owner — this is not the
> place to decide accounting treatment.
## 1. Program identification
- Program name:
- R&D area / profile: [pharma-rd | biotech | medtech | deep-tech | software-rd | university-lab]
- Number of periods + period label (month / quarter / year):
- Funding source(s):
## 2. F&A (indirect) basis
- F&A rate applied: ____%
- Rate type: [negotiated NICRA | de minimis 10% | internal assumption — FLAG IT]
- Fringe rate (loaded onto salaries before F&A): ____%
## 3. Work packages (per-period amounts)
| Work package | Category | F&A-eligible? | P1 | P2 | P3 | P4 |
|---|---|---|---|---|---|---|
| Personnel (FTEs) | personnel | yes | | | | |
| Consumables / supplies | supplies | yes | | | | |
| Capital equipment | capital_equipment | NO (MTDC-exempt) | | | | |
| Subaward / CRO (>$25k) | subaward_over_25k | NO (over $25k exempt) | | | | |
| Travel | travel | yes | | | | |
> Categories that are MTDC-exempt: capital_equipment, subaward_over_25k, tuition, patient_care.
## 4. Milestones (for burn/runway)
| Milestone | Periods from now | Cumulative cash needed |
|---|---|---|
| | | |
## 5. Capitalize-vs-expense candidates (for routing only)
| Cost item | Phase (research / development / software-development) | Standard (ifrs / usgaap) |
|---|---|---|
| | | |
## 6. Assumptions register
- F&A rate basis:
- Escalation assumption:
- Forward burn assumption (flat / ramp):
- Probability-of-success weighting (for any portfolio ROI):
## 7. Named owners
- R&D Finance Controller:
- External Auditor (if capitalization in play):
- Program Lead:
FILE:references/burn_and_portfolio.md
# Burn, Runway, and R&D Portfolio Management
Reference for burn/runway tracking and risk-adjusted portfolio decisions. Pairs with `burn_runway_tracker.py`.
## Burn and runway done honestly
**Burn rate** is cash spent per period; **runway** is cash-on-hand ÷ forward run-rate. The honest forward run-rate is the **trailing** (recent-weighted) burn, not the lifetime average — averages mask both an accelerating spend and a funded ramp. The tracker uses trailing burn and flags when trailing exceeds 115% of the lifetime average (an acceleration signal). Runway must be measured against **value-inflection milestones**: cash that runs out one month before analytical validation is materially worse than the same runway that clears it, because reaching the milestone changes the program's financing options and valuation.
## Stage-gate portfolio management
Robert Cooper's **Stage-Gate** model structures R&D as a sequence of stages separated by go/kill **gates**. Each gate is a real-options decision: spend the next tranche, or kill and redeploy. The discipline is that money is committed one stage at a time, against pre-defined criteria — not as a lump sum at kickoff. This is why milestone-vs-cash alignment is the core runway question.
## Risk-adjusted valuation
Raw NPV systematically overstates R&D value because it ignores attrition. **Risk-adjusted NPV (rNPV)** weights each phase's cash flows by the cumulative probability of success of reaching it — in drug development, the product of per-phase success rates (which compound to single-digit percentages from preclinical to approval). **Real-options** valuation goes further, pricing the optionality of being able to abandon. For portfolio ROI, always state whether the number is raw NPV or risk-adjusted; the difference is often an order of magnitude.
## Efficiency benchmarks
Startup/SaaS efficiency frameworks (a16z's burn multiple, Bessemer's efficiency score) translate to R&D portfolios as "value created per dollar burned." They are blunt but useful for cross-program comparison when paired with milestone progress.
## Sources
1. Cooper, R.G., *Winning at New Products: Creating Value Through Innovation*, 5th ed. (2017) — Stage-Gate.
2. Stewart, Allison & Johnson, *Putting a price on biotechnology* — Nature Biotechnology 2001 (rNPV in drug development).
3. Trigeorgis, L., *Real Options: Managerial Flexibility and Strategy in Resource Allocation* (MIT Press).
4. DiMasi, Grabowski & Hansen, *Innovation in the pharmaceutical industry: New estimates of R&D costs* — J Health Econ 2016 (attrition / phase success rates).
5. a16z, *The burn multiple* and Bessemer State of the Cloud efficiency benchmarks.
6. Chan & Thornhill, *R&D portfolio management* — R&D Management literature.
FILE:references/indirect_rate_modeling.md
# Indirect (F&A) Rate Modeling
Deep reference for the F&A rate — the single most error-prone input in an R&D budget. Pairs with `program_budget_planner.py`.
## What the F&A rate actually is
The F&A rate recovers shared costs that cannot be traced to a single program. It is composed of two pools:
- **Facilities** — depreciation on buildings and equipment, interest on facility debt, operations & maintenance, library, utilities.
- **Administration** — general administration, departmental administration, sponsored-projects administration, student services (in universities).
The rate is computed as (indirect pool ÷ allocation base) and applied to that base on each program.
## The base matters as much as the rate
A 55% rate on a $1M total budget is *not* $550k of F&A — because the rate applies only to the **MTDC base**, which excludes:
- Capital equipment (typically items > $5,000 with > 1-year life)
- The portion of **each** subaward exceeding $25,000 (the first $25k is in the base; the rest is exempt)
- Tuition remission
- Patient-care costs
- Rental of off-site facilities, scholarships, participant support
So a budget heavy in equipment and large subawards has a much smaller F&A base than its headline total. The planner models this exclusion explicitly.
## Negotiated vs de minimis
- **NICRA** — the Negotiated Indirect Cost Rate Agreement, established with a cognizant federal agency. This is the authoritative rate for federally funded work.
- **De minimis 10%** — under 2 CFR 200.414(f), an entity that has never had a negotiated rate may elect a flat 10% of MTDC. Simpler, almost always lower than a negotiated research rate.
## Fringe and the loading stack
Personnel costs load in layers: base salary → **fringe** (benefits, often 25-35%) → then F&A applies to salary+fringe (both are in the MTDC base). Modeling fringe separately from F&A avoids double counting or under-recovery.
## Sources
1. 2 CFR 200.414, *Indirect (F&A) costs*, and Appendix III (IHEs) / Appendix IV (nonprofits).
2. 2 CFR 200.1, definition of *Modified Total Direct Cost (MTDC)*.
3. NIH Grants Policy Statement, indirect-cost chapter; DHHS Cost Allocation Services NICRA guidance.
4. Cost Accounting Standards Board, 48 CFR 9904 (CAS 410, 418 on allocation).
5. COGR (Council on Governmental Relations), *Indirect Cost / F&A* primers and white papers.
6. Federal Demonstration Partnership materials on subaward and MTDC treatment.
FILE:references/rd_program_finance_canon.md
# R&D Program Finance Canon
Reference for the accounting and budgeting rules that govern internal R&D spend. Pairs with `program_budget_planner.py` and `capex_vs_opex_router.py`.
## The central question: research vs development
The accounting treatment of R&D hinges on a phase distinction that the two major frameworks handle differently:
- **IFRS (IAS 38)** — *Research* costs are always **expensed**. *Development* costs **must be capitalized** once all six conditions are met: (1) technical feasibility, (2) intention to complete, (3) ability to use or sell, (4) probable future economic benefit, (5) adequate resources to complete, (6) reliable measurement of expenditure. This is not optional under IFRS — if the criteria are met, capitalization is required.
- **US GAAP (ASC 730)** — R&D is **expensed as incurred**, full stop, with narrow exceptions. The main exception is software: **ASC 985-20** (software to be sold) capitalizes costs after *technological feasibility*; **ASC 350-40** (internal-use software) capitalizes during the application-development stage.
This divergence is why the router takes a `--standard {ifrs,usgaap}` flag: the same cost item can be EXPENSE under US GAAP and CAPITALIZE-CANDIDATE under IFRS.
## F&A / indirect cost (the budgeting half)
Direct costs are traceable to the program (personnel, supplies). **Facilities & Administrative (F&A)**, a.k.a. indirect or overhead, covers shared costs (building, utilities, administration). For federally funded research, F&A is governed by **Uniform Guidance (2 CFR 200)**: organizations negotiate a rate (the NICRA — Negotiated Indirect Cost Rate Agreement) or use the **de minimis 10%** rate. F&A applies to the **Modified Total Direct Cost (MTDC)** base, which *excludes* capital equipment, the portion of each subaward over $25,000, tuition, and patient-care costs. The budget planner enforces this MTDC exclusion.
## Why disclosure matters
A budget number is only as trustworthy as its rate basis and escalation assumption. Two budgets for the same program can differ 40%+ purely on the F&A rate and the base. Every output of the planner ships an assumptions block for exactly this reason.
## Sources
1. IAS 38, *Intangible Assets* — IASB (research vs development, paragraphs 54-67).
2. FASB ASC 730, *Research and Development*; ASC 985-20, *Software — Costs of Software to Be Sold, Leased, or Marketed*; ASC 350-40, *Internal-Use Software*.
3. 2 CFR 200 (Uniform Guidance), Subpart E — Cost Principles, esp. §200.414 (Indirect F&A costs) and the MTDC definition (§200.1).
4. Cost Accounting Standards (CAS), 48 CFR 9904 — for federally funded R&D contractors.
5. KPMG / PwC / Deloitte IFRS-vs-US-GAAP comparison guides (R&D and intangibles chapters).
6. AICPA Accounting & Valuation Guide, *Research and Development*.
FILE:scripts/ar_evaluator.py
#!/usr/bin/env python3
"""ar_evaluator.py - Autoresearch evaluator for the research-finance skill (OPT-IN).
Stdlib-only. The ISOLATED bridge to engineering/autoresearch-agent. It does NOT call
autoresearch; it is the ground-truth evaluator an autoresearch loop runs after editing
the target ledger/budget. It reads a ledger JSON, computes runway via burn_runway_tracker,
and prints ONE metric line:
runway_months: <float> (higher is better)
Optimize a program plan to maximize runway (e.g., resequencing spend) while the agent
edits the target. The user opts in explicitly:
/ar:setup --domain custom --name extend-runway \\
--target ledger.json --eval "python3 ar_evaluator.py --target ledger.json" \\
--metric runway_months --direction higher
Direct use:
python3 ar_evaluator.py --sample
python3 ar_evaluator.py --target ledger.json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import config_loader as cfg # noqa: E402
import burn_runway_tracker as brt # noqa: E402
METRIC = "runway_months"
def main(argv: list[str] | None = None) -> int:
c = cfg.load_config()
p = argparse.ArgumentParser(description="Autoresearch evaluator: R&D program runway in months.")
p.add_argument("--target", help="path to ledger JSON (or env AR_TARGET)")
p.add_argument("--threshold-months", type=float, default=None)
p.add_argument("--sample", action="store_true")
args = p.parse_args(argv)
threshold = args.threshold_months if args.threshold_months is not None \
else c.get("runway_threshold_months", 6)
if args.sample:
data = brt.SAMPLE
else:
target = args.target or os.environ.get("AR_TARGET")
if not target:
print("error: provide --target <ledger.json> or set AR_TARGET", file=sys.stderr)
return 2
try:
with open(target) as f:
data = json.load(f)
except (OSError, json.JSONDecodeError) as e:
print(f"{METRIC}: N/A")
print(f"error: {e}", file=sys.stderr)
return 1
try:
result = brt.analyze(data, threshold)
except ValueError as e:
print(f"{METRIC}: N/A")
print(f"error: {e}", file=sys.stderr)
return 1
print(f"{METRIC}: {result['runway_months_approx']}")
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/burn_runway_tracker.py
#!/usr/bin/env python3
"""burn_runway_tracker.py - Compute R&D program burn, runway, and milestone-vs-cash alignment.
Stdlib-only. Deterministic. NO LLM calls. Surfaces the assumption behind every number.
Given cash-on-hand, a period ledger of actual spend, and upcoming milestones (each with a
period index and the cash needed to reach it), computes:
- average + trailing burn rate
- runway in periods and (approx) months
- whether each value-inflection milestone is reachable before cash runs out
Usage:
python3 burn_runway_tracker.py --sample
python3 burn_runway_tracker.py --input ledger.json --threshold-months 6
python3 burn_runway_tracker.py --input ledger.json --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
SAMPLE = {
"program": "Next-Gen Assay Platform",
"cash_on_hand": 3200000,
"period_label": "month",
"actual_spend": [285000, 305000, 330000, 360000],
"milestones": [
{"name": "Analytical validation", "period_from_now": 3, "cumulative_cash_needed": 1000000},
{"name": "First-in-human readiness", "period_from_now": 9, "cumulative_cash_needed": 3400000},
],
}
def analyze(data: dict, threshold_months: float) -> dict:
spend = [float(x) for x in data.get("actual_spend", [])]
cash = float(data.get("cash_on_hand", 0.0))
label = data.get("period_label", "month")
months_per_period = 1.0 if label == "month" else (3.0 if label == "quarter" else 1.0)
if not spend:
raise ValueError("actual_spend must contain at least one period.")
avg_burn = sum(spend) / len(spend)
trailing_n = min(3, len(spend))
trailing_burn = sum(spend[-trailing_n:]) / trailing_n
# Use trailing burn (more recent) as the forward run-rate.
run_rate = trailing_burn if trailing_burn > 0 else avg_burn
runway_periods = cash / run_rate if run_rate > 0 else float("inf")
runway_months = runway_periods * months_per_period
milestones_out = []
for m in data.get("milestones", []):
needed = float(m.get("cumulative_cash_needed", 0.0))
period_from_now = float(m.get("period_from_now", 0))
reachable_cash = needed <= cash
reachable_time = period_from_now <= runway_periods
verdict = "REACHABLE" if (reachable_cash and reachable_time) else "AT-RISK"
milestones_out.append({
"name": m.get("name", "UNNAMED"),
"period_from_now": period_from_now,
"cumulative_cash_needed": needed,
"cash_covers": reachable_cash,
"runway_covers_timing": reachable_time,
"verdict": verdict,
})
flags = []
if runway_months < threshold_months:
flags.append(f"RUNWAY BELOW THRESHOLD: {runway_months:.1f} months < {threshold_months} month threshold.")
if any(m["verdict"] == "AT-RISK" for m in milestones_out):
flags.append("At least one value-inflection milestone is AT-RISK on current burn.")
if trailing_burn > avg_burn * 1.15:
flags.append(f"Burn accelerating: trailing burn ,.0f > 115% of average ,.0f.")
return {
"program": data.get("program", "UNSPECIFIED"),
"cash_on_hand": cash,
"average_burn_per_period": round(avg_burn, 2),
"trailing_burn_per_period": round(trailing_burn, 2),
"forward_run_rate_used": round(run_rate, 2),
"runway_periods": round(runway_periods, 2),
"runway_months_approx": round(runway_months, 1),
"milestones": milestones_out,
"flags": flags,
"assumptions": [
f"Forward run-rate = trailing {trailing_n}-period burn (recent-weighted, not lifetime average).",
f"Period label '{label}' => {months_per_period} month(s) per period.",
"Runway assumes flat forward burn; a funded ramp or hiring plan changes this.",
"Milestone cash needs are cumulative-from-now as supplied; verify against the program budget.",
],
}
def _render_human(r: dict) -> str:
lines = [f"Burn & Runway: {r['program']}", "",
f"Cash on hand: ,.0f",
f"Average burn/period: ,.0f",
f"Trailing burn/period: ,.0f",
f"Forward run-rate used: ,.0f",
f"Runway: {r['runway_periods']} periods (~{r['runway_months_approx']} months)",
""]
lines.append("Milestones:")
for m in r["milestones"]:
lines.append(f" [{m['verdict']}] {m['name']} (+{m['period_from_now']:.0f} periods, "
f"needs ,.0f)")
lines.append("")
if r["flags"]:
lines.append("Flags:")
for f in r["flags"]:
lines.append(f" ! {f}")
lines.append("")
lines.append("Assumptions:")
for a in r["assumptions"]:
lines.append(f" - {a}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Compute R&D program burn, runway, and milestone alignment.")
p.add_argument("--input", help="Path to JSON ledger")
p.add_argument("--threshold-months", type=float, default=None, help="runway alert threshold (months)")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
threshold = args.threshold_months if args.threshold_months is not None \
else float(conf.get("runway_threshold_months", 6.0))
data = SAMPLE if (args.sample or not args.input) else json.load(open(args.input))
try:
result = analyze(data, threshold)
except ValueError as e:
print(f"error: {e}", file=sys.stderr)
return 2
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/capex_vs_opex_router.py
#!/usr/bin/env python3
"""capex_vs_opex_router.py - Decision-SUPPORT for R&D capitalize-vs-expense treatment.
Stdlib-only. Deterministic. NO LLM calls. This tool NEVER books an entry and NEVER
auto-decides accounting treatment. It scores each cost item against capitalization
criteria and ROUTES it to a named finance owner for the actual determination.
Criteria reflect IAS 38 (development-phase capitalization test) and US GAAP ASC 730
(R&D expensed as incurred) / ASC 985-20 (internal-use & sold software). The six IAS 38
development-phase conditions:
1. technical feasibility established
2. intention to complete
3. ability to use or sell
4. probable future economic benefit
5. adequate resources to complete
6. reliable measurement of expenditure
Verdicts:
- CAPITALIZE-CANDIDATE (development phase, all criteria met) -> still routes to finance owner
- EXPENSE (research phase, or criteria not met)
- FINANCE-OWNER-REVIEW (ambiguous / partial criteria)
Usage:
python3 capex_vs_opex_router.py --sample
python3 capex_vs_opex_router.py --input costs.json --standard ifrs
python3 capex_vs_opex_router.py --input costs.json --standard usgaap --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
IAS38_CRITERIA = [
"technical_feasibility",
"intention_to_complete",
"ability_to_use_or_sell",
"probable_future_benefit",
"adequate_resources",
"reliable_measurement",
]
# Profiles only annotate context; they do not change the accounting test.
PROFILES = {
"pharma-rd": "Most drug R&D is expensed; capitalization rare pre-approval.",
"biotech": "Similar to pharma; pre-approval development typically expensed.",
"medtech": "Some development capitalizable post-feasibility under IFRS.",
"deep-tech": "Prototype-to-product transition is the key feasibility line.",
"software-rd": "ASC 985-20 / IAS 38: capitalize after technological feasibility / working model.",
"university-lab": "Grant-funded research almost always expensed per funder terms.",
}
SAMPLE = {
"standard": "ifrs",
"items": [
{
"name": "Exploratory target screening",
"phase": "research",
"criteria": {},
},
{
"name": "Pilot-line tooling for validated design",
"phase": "development",
"criteria": {
"technical_feasibility": True, "intention_to_complete": True,
"ability_to_use_or_sell": True, "probable_future_benefit": True,
"adequate_resources": True, "reliable_measurement": True,
},
},
{
"name": "Software build (post working-model, pre-release)",
"phase": "development",
"criteria": {
"technical_feasibility": True, "intention_to_complete": True,
"ability_to_use_or_sell": True, "probable_future_benefit": True,
"adequate_resources": False, "reliable_measurement": True,
},
},
],
}
def route_item(item: dict, standard: str) -> dict:
phase = (item.get("phase") or "").lower()
crit = item.get("criteria", {}) or {}
met = [c for c in IAS38_CRITERIA if crit.get(c)]
missing = [c for c in IAS38_CRITERIA if not crit.get(c)]
# US GAAP ASC 730: R&D expensed as incurred (software is the main exception via ASC 985-20).
if standard == "usgaap" and phase != "software-development":
verdict = "EXPENSE"
rationale = "ASC 730: R&D is expensed as incurred (non-software). Confirm software exceptions separately."
owner = "R&D Finance Controller"
elif phase == "research":
verdict = "EXPENSE"
rationale = "Research phase: cannot capitalize (IAS 38.54)."
owner = "R&D Finance Controller"
elif phase in ("development", "software-development") and not missing:
verdict = "CAPITALIZE-CANDIDATE"
rationale = "Development phase with all 6 IAS 38 criteria asserted. Routed for finance confirmation."
owner = "R&D Finance Controller + External Auditor sign-off"
else:
verdict = "FINANCE-OWNER-REVIEW"
rationale = f"Development phase but {len(missing)} criteria unmet/unstated: {', '.join(missing) or 'n/a'}."
owner = "R&D Finance Controller"
return {
"name": item.get("name", "UNNAMED"),
"phase": phase or "UNSPECIFIED",
"criteria_met": met,
"criteria_missing": missing,
"verdict": verdict,
"rationale": rationale,
"named_owner": owner,
}
def route(data: dict, standard: str, profile: str) -> dict:
if profile not in PROFILES:
raise ValueError(f"Unknown profile '{profile}'. Choose from {list(PROFILES)}.")
items = [route_item(i, standard) for i in data.get("items", [])]
return {
"standard": standard,
"profile": profile,
"profile_note": PROFILES[profile],
"items": items,
"disclaimer": "DECISION SUPPORT ONLY. This tool does not book entries or decide treatment. "
"A named finance owner (and auditor where required) makes the determination.",
}
def _render_human(r: dict) -> str:
lines = [f"Capitalize-vs-Expense routing (standard: {r['standard']}, profile: {r['profile']})",
f" {r['profile_note']}", ""]
for it in r["items"]:
lines.append(f"[{it['verdict']}] {it['name']} (phase: {it['phase']})")
lines.append(f" {it['rationale']}")
if it["criteria_missing"]:
lines.append(f" missing/unstated: {', '.join(it['criteria_missing'])}")
lines.append(f" -> route to: {it['named_owner']}")
lines.append("")
lines.append(f"!! {r['disclaimer']}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Route R&D costs to capitalize/expense/review (DECISION SUPPORT ONLY).")
p.add_argument("--input", help="Path to JSON with items[]")
p.add_argument("--standard", default=None, choices=["ifrs", "usgaap"],
help="overrides onboarding accounting_standard")
p.add_argument("--profile", default=None, choices=list(PROFILES),
help="overrides onboarding default_profile")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
profile = args.profile or conf.get("default_profile", "biotech")
cli_standard = args.standard or conf.get("accounting_standard", "ifrs")
data = SAMPLE if (args.sample or not args.input) else json.load(open(args.input))
standard = data.get("standard", cli_standard) if (args.sample or not args.input) else cli_standard
try:
result = route(data, standard, profile)
except ValueError as e:
print(f"error: {e}", file=sys.stderr)
return 2
finance_owner = conf.get("finance_owner")
if finance_owner:
for it in result["items"]:
it["named_owner"] = it["named_owner"].replace(
"R&D Finance Controller", f"R&D Finance Controller ({finance_owner})")
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/config_loader.py
#!/usr/bin/env python3
"""config_loader.py - Customization loader for the research-finance skill.
Stdlib-only. Importable from the skill's other scripts. Precedence (highest wins):
1. Project config: <cwd>/.research-ops/research-finance.json
2. Global config: ~/.config/research-ops/research-finance.json
3. Built-in DEFAULTS
Onboarding answers (written by onboard.py) live in these files; every tool in this
skill reads them so the user's customization applies automatically.
Set RESEARCH_OPS_NO_CONFIG=1 to ignore saved config.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from pathlib import Path
from typing import Any
SKILL = "research-finance"
GLOBAL_CONFIG_DIR = Path.home() / ".config" / "research-ops"
GLOBAL_CONFIG_PATH = GLOBAL_CONFIG_DIR / f"{SKILL}.json"
PROJECT_CONFIG_DIRNAME = ".research-ops"
DEFAULTS: dict[str, Any] = {
"version": 1,
"skill": SKILL,
"default_profile": "biotech",
"default_fa_rate": None, # None => use the profile's default F&A rate
"runway_threshold_months": 6,
"accounting_standard": "ifrs",
"finance_owner": None,
"setup_completed_at": None,
}
def project_config_path(cwd: Path | None = None) -> Path:
cwd = cwd or Path.cwd()
return cwd / PROJECT_CONFIG_DIRNAME / f"{SKILL}.json"
def _read_json(path: Path) -> dict[str, Any] | None:
try:
with path.open(encoding="utf-8") as f:
data = json.load(f)
return data if isinstance(data, dict) else None
except (FileNotFoundError, json.JSONDecodeError, OSError):
return None
def _deep_merge(base: dict[str, Any], override: dict[str, Any]) -> dict[str, Any]:
out = dict(base)
for k, v in override.items():
if isinstance(v, dict) and isinstance(out.get(k), dict):
out[k] = _deep_merge(out[k], v)
else:
out[k] = v
return out
def load_config(cwd: Path | None = None) -> dict[str, Any]:
config = dict(DEFAULTS)
if os.environ.get("RESEARCH_OPS_NO_CONFIG") == "1":
return config
global_cfg = _read_json(GLOBAL_CONFIG_PATH)
if global_cfg:
config = _deep_merge(config, global_cfg)
project_cfg = _read_json(project_config_path(cwd))
if project_cfg:
config = _deep_merge(config, project_cfg)
return config
def setup_completed() -> bool:
cfg = _read_json(GLOBAL_CONFIG_PATH) or _read_json(project_config_path())
return bool(cfg and cfg.get("setup_completed_at"))
def write_config(config: dict[str, Any], scope: str = "global", cwd: Path | None = None) -> Path:
path = project_config_path(cwd) if scope == "project" else GLOBAL_CONFIG_PATH
path.parent.mkdir(parents=True, exist_ok=True)
with path.open("w", encoding="utf-8") as f:
json.dump(config, f, indent=2, sort_keys=True)
return path
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=f"Inspect {SKILL} customization config.")
p.add_argument("--show", action="store_true", help="Print the effective config")
p.add_argument("--status", action="store_true", help="Print setup status + paths")
p.add_argument("--sample", action="store_true", help="Print the built-in defaults")
args = p.parse_args(argv)
if args.sample:
print(json.dumps(DEFAULTS, indent=2, sort_keys=True))
elif args.status:
print(json.dumps({
"skill": SKILL,
"global_config_path": str(GLOBAL_CONFIG_PATH),
"global_config_exists": GLOBAL_CONFIG_PATH.exists(),
"project_config_path": str(project_config_path()),
"project_config_exists": project_config_path().exists(),
"setup_completed": setup_completed(),
}, indent=2))
else:
print(json.dumps(load_config(), indent=2, sort_keys=True))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/onboard.py
#!/usr/bin/env python3
"""onboard.py - Onboarding questionnaire for the research-finance skill.
Stdlib-only. Asks the user a short set of questions BEFORE they build an R&D program
budget, then writes the answers to a customization config read by every tool in this
skill via config_loader.py. The answers become defaults for profile, F&A rate, runway
threshold, accounting standard, and the named finance owner printed on routing outputs.
Modes: --show | --defaults | --set key=value (repeatable) | --reset | --scope {global,project}
"""
from __future__ import annotations
import argparse
import datetime as _dt
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import config_loader as cfg # noqa: E402
NUMERIC_KEYS = {"default_fa_rate", "runway_threshold_months"}
QUESTIONS = [
("default_profile",
"1. What R&D area is this program?",
["pharma-rd", "biotech", "medtech", "deep-tech", "software-rd", "university-lab"], str),
("default_fa_rate",
"2. F&A / indirect rate as a fraction (e.g. 0.55), or blank to use the profile default?",
None, float),
("runway_threshold_months",
"3. Runway alert threshold in months (warn below this)?",
None, float),
("accounting_standard",
"4. Which accounting standard governs capitalize-vs-expense?",
["ifrs", "usgaap"], str),
("finance_owner",
"5. Named finance/controller owner who signs accounting treatment?",
None, str),
]
def _print_questions() -> None:
print(f"Onboarding questions — {cfg.SKILL}:\n")
for _k, prompt, choices, _c in QUESTIONS:
line = f" {prompt}"
if choices:
line += f" [{' / '.join(choices)}]"
print(line)
def run_interactive(config: dict) -> dict:
print(f"Onboarding — {cfg.SKILL}. Press Enter to keep the current/default value.\n")
for key, prompt, choices, caster in QUESTIONS:
suffix = f" [{'/'.join(choices)}]" if choices else ""
cur = f" (current: {config.get(key)})" if config.get(key) is not None else ""
raw = input(f"{prompt}{suffix}{cur}: ").strip()
if not raw:
continue
try:
config[key] = caster(raw)
except ValueError:
print(f" ! invalid value for {key}, keeping current")
return config
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=f"Onboarding for the {cfg.SKILL} skill.")
p.add_argument("--show", action="store_true")
p.add_argument("--defaults", action="store_true", help="write built-in defaults, no prompt")
p.add_argument("--set", action="append", default=[], metavar="key=value")
p.add_argument("--reset", action="store_true")
p.add_argument("--scope", choices=["global", "project"], default="global")
args = p.parse_args(argv)
if args.show:
_print_questions()
print("\nCurrent effective config:")
print(json.dumps(cfg.load_config(), indent=2, sort_keys=True))
return 0
if args.reset:
path = cfg.project_config_path() if args.scope == "project" else cfg.GLOBAL_CONFIG_PATH
if path.exists():
path.unlink(); print(f"removed {path}")
else:
print(f"no config at {path}")
return 0
config = cfg.load_config()
if args.set:
for item in args.set:
if "=" not in item:
print(f"error: --set expects key=value, got '{item}'", file=sys.stderr)
return 2
k, v = item.split("=", 1)
if k in NUMERIC_KEYS:
try:
v = float(v)
except ValueError:
pass
config[k] = v
elif not args.defaults:
if sys.stdin.isatty():
config = run_interactive(config)
else:
print("non-interactive shell: use --defaults or --set key=value. Showing questions:\n")
_print_questions()
return 0
config["setup_completed_at"] = _dt.datetime.now(_dt.timezone.utc).isoformat()
path = cfg.write_config(config, scope=args.scope)
print(f"saved {cfg.SKILL} customization -> {path}")
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/program_budget_planner.py
#!/usr/bin/env python3
"""program_budget_planner.py - Build a multi-period R&D program budget with F&A split.
Stdlib-only. Deterministic. NO LLM calls. Every output surfaces an explicit assumptions
block: budget math without disclosed assumptions is theatre.
Takes work-package line items, applies the F&A (indirect) rate to the F&A-eligible base
(MTDC-style: excludes capital equipment and the portion of subawards over $25k), computes
fully-loaded cost, and rolls up per period.
Usage:
python3 program_budget_planner.py --sample
python3 program_budget_planner.py --input program.json --fa-rate 0.55 --periods 4
python3 program_budget_planner.py --input program.json --profile biotech --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
# Profile default F&A (indirect) rate and escalation assumption when not supplied in input.
PROFILES = {
"pharma-rd": {"default_fa_rate": 0.50, "annual_escalation": 0.03},
"biotech": {"default_fa_rate": 0.55, "annual_escalation": 0.04},
"medtech": {"default_fa_rate": 0.45, "annual_escalation": 0.03},
"deep-tech": {"default_fa_rate": 0.40, "annual_escalation": 0.03},
"software-rd": {"default_fa_rate": 0.30, "annual_escalation": 0.04},
"university-lab": {"default_fa_rate": 0.585, "annual_escalation": 0.025},
}
# Categories excluded from the F&A (MTDC) base.
FA_EXEMPT_CATEGORIES = {"capital_equipment", "subaward_over_25k", "tuition", "patient_care"}
SAMPLE = {
"program": "Next-Gen Assay Platform",
"periods": 4,
"work_packages": [
{"name": "Personnel (FTEs)", "category": "personnel", "amounts": [320000, 330000, 340000, 350000]},
{"name": "Consumables", "category": "supplies", "amounts": [60000, 65000, 70000, 70000]},
{"name": "Sequencer", "category": "capital_equipment", "amounts": [180000, 0, 0, 0]},
{"name": "CRO subaward", "category": "subaward_over_25k", "amounts": [100000, 100000, 0, 0]},
{"name": "Travel", "category": "travel", "amounts": [12000, 12000, 12000, 12000]},
],
}
def _period_sum(amounts: list, n: int, idx: int) -> float:
return float(amounts[idx]) if idx < len(amounts) else 0.0
def plan_budget(data: dict, fa_rate: float, periods: int) -> dict:
wps = data.get("work_packages", [])
direct_by_period = [0.0] * periods
fa_base_by_period = [0.0] * periods
line_items = []
for wp in wps:
cat = wp.get("category", "other")
amounts = wp.get("amounts", [])
fa_eligible = cat not in FA_EXEMPT_CATEGORIES
wp_total = 0.0
for i in range(periods):
amt = _period_sum(amounts, periods, i)
direct_by_period[i] += amt
if fa_eligible:
fa_base_by_period[i] += amt
wp_total += amt
line_items.append({
"name": wp.get("name", "UNNAMED"),
"category": cat,
"fa_eligible": fa_eligible,
"total_direct": round(wp_total, 2),
})
fa_by_period = [round(b * fa_rate, 2) for b in fa_base_by_period]
loaded_by_period = [round(direct_by_period[i] + fa_by_period[i], 2) for i in range(periods)]
return {
"program": data.get("program", "UNSPECIFIED"),
"periods": periods,
"fa_rate_applied": fa_rate,
"line_items": line_items,
"direct_by_period": [round(x, 2) for x in direct_by_period],
"fa_base_by_period": [round(x, 2) for x in fa_base_by_period],
"fa_by_period": fa_by_period,
"fully_loaded_by_period": loaded_by_period,
"total_direct": round(sum(direct_by_period), 2),
"total_fa": round(sum(fa_by_period), 2),
"total_fully_loaded": round(sum(loaded_by_period), 2),
"assumptions": [
f"F&A (indirect) rate applied: {fa_rate:.1%}. Confirm this is your negotiated NICRA, not an assumption.",
f"F&A base excludes: {', '.join(sorted(FA_EXEMPT_CATEGORIES))} (MTDC-style base).",
"Amounts are taken as-entered per period; no escalation applied unless baked into inputs.",
"This is a planning estimate; a finance owner/controller validates the rate basis and booking.",
],
}
def _render_human(r: dict) -> str:
lines = [f"R&D Program Budget: {r['program']} ({r['periods']} periods)",
f"F&A rate applied: {r['fa_rate_applied']:.1%}", ""]
lines.append("Line items:")
for li in r["line_items"]:
tag = "F&A-eligible" if li["fa_eligible"] else "F&A-EXEMPT"
lines.append(f" {li['name']:24s} {li['category']:20s} {tag:12s} ,.0f")
lines.append("")
hdr = " " + "".join(f"P{i+1:>14}" for i in range(r["periods"]))
lines.append("Per-period rollup:" )
lines.append(hdr)
lines.append(" direct " + "".join(f"{v:>15,.0f}" for v in r["direct_by_period"]))
lines.append(" F&A " + "".join(f"{v:>15,.0f}" for v in r["fa_by_period"]))
lines.append(" loaded " + "".join(f"{v:>15,.0f}" for v in r["fully_loaded_by_period"]))
lines.append("")
lines.append(f"Total direct: ,.0f")
lines.append(f"Total F&A: ,.0f")
lines.append(f"Total fully-loaded: ,.0f")
lines.append("")
lines.append("Assumptions (state these alongside the number):")
for a in r["assumptions"]:
lines.append(f" - {a}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Build a multi-period R&D program budget with F&A split.")
p.add_argument("--input", help="Path to JSON program with work_packages[]")
p.add_argument("--profile", default=None, choices=list(PROFILES),
help="overrides onboarding default_profile")
p.add_argument("--fa-rate", type=float, default=None, help="Override F&A rate (fraction, e.g. 0.55)")
p.add_argument("--periods", type=int, default=None, help="Number of periods")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
profile = args.profile or conf.get("default_profile", "biotech")
data = SAMPLE if (args.sample or not args.input) else json.load(open(args.input))
periods = args.periods or int(data.get("periods", 4))
# F&A precedence: CLI flag > onboarding default_fa_rate (if set) > profile default
if args.fa_rate is not None:
fa_rate = args.fa_rate
elif conf.get("default_fa_rate") is not None:
fa_rate = float(conf["default_fa_rate"])
else:
fa_rate = PROFILES[profile]["default_fa_rate"]
result = plan_budget(data, fa_rate, periods)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
Bộ nhớ hai lớp cho quyết định họp hội đồng: bản ghi gốc và quyết định đã duyệt, xem lại quyết định cũ, kiểm tra hạng mục quá hạn.
---
name: "decision-logger"
description: "Two-layer memory architecture for board meeting decisions. Manages raw transcripts (Layer 1) and approved decisions (Layer 2). Use when logging decisions after a board meeting, reviewing past decisions with /cs:decisions, or checking overdue action items with /cs:review. Invoked automatically by the board-meeting skill after Phase 5 founder approval."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: decision-memory
updated: 2026-03-05
python-tools: scripts/decision_tracker.py
---
# Decision Logger
Two-layer memory system. Layer 1 stores everything. Layer 2 stores only what the founder approved. Future meetings read Layer 2 only — this prevents hallucinated consensus from past debates bleeding into new deliberations.
## Keywords
decision log, memory, approved decisions, action items, board minutes, /cs:decisions, /cs:review, conflict detection, DO_NOT_RESURFACE
## Quick Start
```bash
python scripts/decision_tracker.py --demo # See sample output
python scripts/decision_tracker.py --summary # Overview + overdue
python scripts/decision_tracker.py --overdue # Past-deadline actions
python scripts/decision_tracker.py --conflicts # Contradiction detection
python scripts/decision_tracker.py --owner "CTO" # Filter by owner
python scripts/decision_tracker.py --search "pricing" # Search decisions
```
---
## Commands
| Command | Effect |
|---------|--------|
| `/cs:decisions` | Last 10 approved decisions |
| `/cs:decisions --all` | Full history |
| `/cs:decisions --owner CMO` | Filter by owner |
| `/cs:decisions --topic pricing` | Search by keyword |
| `/cs:review` | Action items due within 7 days |
| `/cs:review --overdue` | Items past deadline |
---
## Two-Layer Architecture
### Layer 1 — Raw Transcripts
**Location:** `memory/board-meetings/YYYY-MM-DD-raw.md`
- Full Phase 2 agent contributions, Phase 3 critique, Phase 4 synthesis
- All debates, including rejected arguments
- **NEVER auto-loaded.** Only on explicit founder request.
- Archive after 90 days → `memory/board-meetings/archive/YYYY/`
### Layer 2 — Approved Decisions
**Location:** `memory/board-meetings/decisions.md`
- ONLY founder-approved decisions, action items, user corrections
- **Loaded automatically in Phase 1 of every board meeting**
- Append-only. Decisions are never deleted — only superseded.
- Managed by Chief of Staff after Phase 5. Never written by agents directly.
---
## Decision Entry Format
```markdown
## [YYYY-MM-DD] — [AGENDA ITEM TITLE]
**Decision:** [One clear statement of what was decided.]
**Owner:** [One person or role — accountable for execution.]
**Deadline:** [YYYY-MM-DD]
**Review:** [YYYY-MM-DD]
**Rationale:** [Why this over alternatives. 1-2 sentences.]
**User Override:** [If founder changed agent recommendation — what and why. Blank if not applicable.]
**Rejected:**
- [Proposal] — [reason] [DO_NOT_RESURFACE]
**Action Items:**
- [ ] [Action] — Owner: [name] — Due: [YYYY-MM-DD] — Review: [YYYY-MM-DD]
**Supersedes:** [DATE of previous decision on same topic, if any]
**Superseded by:** [Filled in retroactively if overridden later]
**Raw transcript:** memory/board-meetings/[DATE]-raw.md
```
---
## Conflict Detection
Before logging, Chief of Staff checks for:
1. **DO_NOT_RESURFACE violations** — new decision matches a rejected proposal
2. **Topic contradictions** — two active decisions on same topic with different conclusions
3. **Owner conflicts** — same action assigned to different people in different decisions
When a conflict is found:
```
⚠️ DECISION CONFLICT
New: [text]
Conflicts with: [DATE] — [existing text]
Options: (1) Supersede old (2) Merge (3) Defer to founder
```
**DO_NOT_RESURFACE enforcement:**
```
🚫 BLOCKED: "[Proposal]" was rejected on [DATE]. Reason: [reason].
To reopen: founder must explicitly say "reopen [topic] from [DATE]".
```
---
## Logging Workflow (Post Phase 5)
1. Founder approves synthesis
2. Write Layer 1 raw transcript → `YYYY-MM-DD-raw.md`
3. Check conflicts against `decisions.md`
4. Surface conflicts → wait for founder resolution
5. Append approved entries to `decisions.md`
6. Confirm: decisions logged, actions tracked, DO_NOT_RESURFACE flags added
---
## Marking Actions Complete
```markdown
- [x] [Action] — Owner: [name] — Completed: [DATE] — Result: [one sentence]
```
Never delete completed items. The history is the record.
---
## File Structure
```
memory/board-meetings/
├── decisions.md # Layer 2: append-only, founder-approved
├── YYYY-MM-DD-raw.md # Layer 1: full transcript per meeting
└── archive/YYYY/ # Raw files after 90 days
```
---
## References
- `templates/decision-entry.md` — single entry template with field rules
- `scripts/decision_tracker.py` — CLI parser, overdue tracker, conflict detector
FILE:scripts/decision_tracker.py
#!/usr/bin/env python3
"""
decision_tracker.py — Board Meeting Decision Parser & Reporter
Part of the C-Level Advisor / Decision Logger skill.
Parses memory/board-meetings/decisions.md and produces actionable reports.
Stdlib only. No dependencies.
Usage:
python decision_tracker.py --summary
python decision_tracker.py --overdue
python decision_tracker.py --conflicts
python decision_tracker.py --owner "CMO"
python decision_tracker.py --search "pricing"
python decision_tracker.py --due-within 7
python decision_tracker.py --demo # Run with sample data
"""
import argparse
import os
import re
import sys
from datetime import date, datetime, timedelta
from pathlib import Path
from typing import Optional
# ─────────────────────────────────────────────
# Data structures
# ─────────────────────────────────────────────
class ActionItem:
def __init__(self, text: str, owner: str, due: Optional[date],
review: Optional[date], completed: bool, completed_date: Optional[date],
result: str):
self.text = text
self.owner = owner
self.due = due
self.review = review
self.completed = completed
self.completed_date = completed_date
self.result = result
def is_overdue(self) -> bool:
if self.completed:
return False
if self.due and self.due < date.today():
return True
return False
def is_due_within(self, days: int) -> bool:
if self.completed:
return False
if self.due:
return date.today() <= self.due <= date.today() + timedelta(days=days)
return False
class Decision:
def __init__(self):
self.date: Optional[date] = None
self.title: str = ""
self.decision: str = ""
self.owner: str = ""
self.deadline: Optional[date] = None
self.review: Optional[date] = None
self.rationale: str = ""
self.user_override: str = ""
self.rejected: list[str] = []
self.action_items: list[ActionItem] = []
self.supersedes: str = ""
self.superseded_by: str = ""
self.raw_transcript: str = ""
def is_active(self) -> bool:
return not bool(self.superseded_by.strip())
def has_override(self) -> bool:
return bool(self.user_override.strip())
# ─────────────────────────────────────────────
# Parser
# ─────────────────────────────────────────────
def parse_date(s: str) -> Optional[date]:
"""Parse YYYY-MM-DD or return None."""
if not s:
return None
s = s.strip()
for fmt in ("%Y-%m-%d", "%Y/%m/%d", "%d.%m.%Y"):
try:
return datetime.strptime(s, fmt).date()
except ValueError:
continue
return None
def parse_action_item(line: str) -> Optional[ActionItem]:
"""
Parse a line like:
- [ ] Action text — Owner: CMO — Due: 2026-03-15 — Review: 2026-03-29
- [x] Action text — Owner: CEO — Completed: 2026-03-10 — Result: Done
"""
line = line.strip()
if not line.startswith("- ["):
return None
completed = line.startswith("- [x]") or line.startswith("- [X]")
text_start = line.find("]") + 1
raw = line[text_start:].strip()
# Split on " — " (em dash with spaces) or " - " fallback
parts_raw = re.split(r"\s+[—\-]{1,2}\s+", raw)
text = parts_raw[0].strip() if parts_raw else raw
def extract(label: str, parts: list[str]) -> str:
for p in parts:
if p.lower().startswith(label.lower() + ":"):
return p[len(label) + 1:].strip()
return ""
owner = extract("Owner", parts_raw[1:])
due_str = extract("Due", parts_raw[1:])
review_str = extract("Review", parts_raw[1:])
completed_str = extract("Completed", parts_raw[1:])
result = extract("Result", parts_raw[1:])
return ActionItem(
text=text,
owner=owner,
due=parse_date(due_str),
review=parse_date(review_str),
completed=completed,
completed_date=parse_date(completed_str),
result=result,
)
def parse_decisions(content: str) -> list[Decision]:
"""Parse the full decisions.md content into Decision objects."""
decisions = []
current: Optional[Decision] = None
in_rejected = False
in_actions = False
for line in content.splitlines():
# New decision entry
header_match = re.match(r"^## (\d{4}-\d{2}-\d{2}) — (.+)$", line)
if header_match:
if current:
decisions.append(current)
current = Decision()
current.date = parse_date(header_match.group(1))
current.title = header_match.group(2).strip()
in_rejected = False
in_actions = False
continue
if current is None:
continue
# Field parsing
def extract_field(label: str) -> Optional[str]:
pattern = rf"^\*\*{re.escape(label)}:\*\*\s*(.*)$"
m = re.match(pattern, line)
return m.group(1).strip() if m else None
val = extract_field("Decision")
if val is not None:
current.decision = val
in_rejected = False
in_actions = False
continue
val = extract_field("Owner")
if val is not None:
current.owner = val
continue
val = extract_field("Deadline")
if val is not None:
current.deadline = parse_date(val)
continue
val = extract_field("Review")
if val is not None:
current.review = parse_date(val)
continue
val = extract_field("Rationale")
if val is not None:
current.rationale = val
continue
val = extract_field("User Override")
if val is not None:
current.user_override = val
in_rejected = False
in_actions = False
continue
val = extract_field("Supersedes")
if val is not None:
current.supersedes = val
continue
val = extract_field("Superseded by")
if val is not None:
current.superseded_by = val
continue
val = extract_field("Raw transcript")
if val is not None:
current.raw_transcript = val
continue
# Section headers
if re.match(r"^\*\*Rejected:\*\*", line):
in_rejected = True
in_actions = False
continue
if re.match(r"^\*\*Action Items:\*\*", line):
in_actions = True
in_rejected = False
continue
if line.startswith("**"):
in_rejected = False
in_actions = False
# List items
if in_rejected and line.strip().startswith("-"):
item = line.strip().lstrip("- ").strip()
if item and not item.startswith("<!--"):
current.rejected.append(item)
continue
if in_actions and line.strip().startswith("- ["):
action = parse_action_item(line)
if action:
current.action_items.append(action)
continue
if current:
decisions.append(current)
return decisions
# ─────────────────────────────────────────────
# Reports
# ─────────────────────────────────────────────
def fmt_date(d: Optional[date]) -> str:
return d.strftime("%Y-%m-%d") if d else "—"
def fmt_delta(d: Optional[date]) -> str:
if not d:
return ""
delta = (d - date.today()).days
if delta < 0:
return f" ⚠️ {abs(delta)}d overdue"
if delta == 0:
return " 🔴 DUE TODAY"
if delta <= 3:
return f" 🟡 {delta}d left"
return f" ({delta}d)"
def print_section(title: str):
print(f"\n{'═' * 60}")
print(f" {title}")
print(f"{'═' * 60}")
def report_summary(decisions: list[Decision]):
active = [d for d in decisions if d.is_active()]
all_actions = [a for d in decisions for a in d.action_items]
open_actions = [a for a in all_actions if not a.completed]
overdue = [a for a in all_actions if a.is_overdue()]
overrides = [d for d in decisions if d.has_override()]
dnr_count = sum(len(d.rejected) for d in decisions)
print_section("DECISION LOG SUMMARY")
print(f" Total decisions: {len(decisions)}")
print(f" Active (not super.): {len(active)}")
print(f" Superseded: {len(decisions) - len(active)}")
print(f" Founder overrides: {len(overrides)}")
print(f" DO_NOT_RESURFACE: {dnr_count}")
print(f" Total action items: {len(all_actions)}")
print(f" Open action items: {len(open_actions)}")
print(f" Overdue: {len(overdue)}")
if overdue:
print(f"\n {'─' * 40}")
print(f" ⚠️ OVERDUE ITEMS ({len(overdue)})")
print(f" {'─' * 40}")
for a in overdue:
print(f" • [{a.owner}] {a.text}")
print(f" Due: {fmt_date(a.due)}{fmt_delta(a.due)}")
print(f"\n {'─' * 40}")
print(f" RECENT DECISIONS")
print(f" {'─' * 40}")
for d in sorted(active, key=lambda x: x.date or date.min, reverse=True)[:5]:
print(f" [{fmt_date(d.date)}] {d.title}")
print(f" Owner: {d.owner or '—'} | Deadline: {fmt_date(d.deadline)}")
open_count = sum(1 for a in d.action_items if not a.completed)
if open_count:
print(f" Open actions: {open_count}")
def report_overdue(decisions: list[Decision]):
print_section("OVERDUE ACTION ITEMS")
found = False
for d in sorted(decisions, key=lambda x: x.date or date.min, reverse=True):
overdue = [a for a in d.action_items if a.is_overdue()]
if not overdue:
continue
found = True
print(f"\n 📋 {d.title} [{fmt_date(d.date)}]")
for a in overdue:
print(f" ⚠️ {a.text}")
print(f" Owner: {a.owner or '—'} | Due: {fmt_date(a.due)}{fmt_delta(a.due)}")
if not found:
print("\n ✅ No overdue items.")
def report_due_within(decisions: list[Decision], days: int):
print_section(f"ACTION ITEMS DUE WITHIN {days} DAYS")
found = False
for d in sorted(decisions, key=lambda x: x.date or date.min, reverse=True):
upcoming = [a for a in d.action_items if a.is_due_within(days)]
if not upcoming:
continue
found = True
print(f"\n 📋 {d.title} [{fmt_date(d.date)}]")
for a in upcoming:
print(f" • {a.text}")
print(f" Owner: {a.owner or '—'} | Due: {fmt_date(a.due)}{fmt_delta(a.due)}")
if not found:
print(f"\n ✅ Nothing due in the next {days} days.")
def report_by_owner(decisions: list[Decision], owner: str):
print_section(f"ACTION ITEMS — OWNER: {owner.upper()}")
found = False
for d in sorted(decisions, key=lambda x: x.date or date.min, reverse=True):
items = [a for a in d.action_items
if a.owner.lower() == owner.lower() and not a.completed]
if not items:
continue
found = True
print(f"\n 📋 {d.title} [{fmt_date(d.date)}]")
for a in items:
flag = "⚠️ OVERDUE" if a.is_overdue() else ""
print(f" {'[ ]'} {a.text} {flag}")
print(f" Due: {fmt_date(a.due)}{fmt_delta(a.due)}")
if not found:
print(f"\n No open action items for '{owner}'.")
def report_search(decisions: list[Decision], query: str):
print_section(f"SEARCH: \"{query}\"")
q = query.lower()
found = False
for d in decisions:
hit_fields = []
if q in d.title.lower():
hit_fields.append("title")
if q in d.decision.lower():
hit_fields.append("decision")
if q in d.rationale.lower():
hit_fields.append("rationale")
if any(q in r.lower() for r in d.rejected):
hit_fields.append("rejected")
if hit_fields:
found = True
print(f"\n [{fmt_date(d.date)}] {d.title} (match: {', '.join(hit_fields)})")
if "decision" in hit_fields:
print(f" → {d.decision}")
if "rejected" in hit_fields:
matches = [r for r in d.rejected if q in r.lower()]
for r in matches:
print(f" ✗ [REJECTED] {r}")
if not found:
print(f"\n No results for '{query}'.")
def report_conflicts(decisions: list[Decision]):
"""
Simple conflict detection: look for decisions on the same topic
(matching title words) that are both active and have different decisions.
Also flag if a rejected item appears as a new decision.
"""
print_section("CONFLICT DETECTION")
conflicts_found = False
# Check for DO_NOT_RESURFACE violations
all_rejected_texts = []
for d in decisions:
for r in d.rejected:
clean = re.sub(r"\[DO_NOT_RESURFACE\]", "", r).strip().lower()
all_rejected_texts.append((clean, d.date, d.title))
active = [d for d in decisions if d.is_active()]
for d in active:
decision_lower = d.decision.lower()
for rejected_text, rejected_date, rejected_title in all_rejected_texts:
if rejected_text and rejected_text in decision_lower:
conflicts_found = True
print(f"\n 🚫 POTENTIAL DO_NOT_RESURFACE VIOLATION")
print(f" Decision [{fmt_date(d.date)}]: {d.decision}")
print(f" Matches rejected item from [{fmt_date(rejected_date)}] ({rejected_title}):")
print(f" \"{rejected_text}\"")
# Check for same-topic contradictions (shared keywords in title)
stop_words = {"the", "a", "an", "and", "or", "to", "for", "of", "in", "on", "with", "vs"}
for i, d1 in enumerate(active):
words1 = set(w.lower() for w in d1.title.split() if w.lower() not in stop_words)
for d2 in active[i+1:]:
words2 = set(w.lower() for w in d2.title.split() if w.lower() not in stop_words)
overlap = words1 & words2
if len(overlap) >= 2 and d1.decision and d2.decision:
# Different decisions on similar topic
if d1.decision.lower() != d2.decision.lower():
conflicts_found = True
print(f"\n ⚠️ POTENTIAL CONFLICT (shared topic: {overlap})")
print(f" [{fmt_date(d1.date)}] {d1.title}")
print(f" Decision: {d1.decision}")
print(f" [{fmt_date(d2.date)}] {d2.title}")
print(f" Decision: {d2.decision}")
if d1.superseded_by or d2.superseded_by:
print(f" ℹ️ One may supersede the other — check Superseded by fields.")
if not conflicts_found:
print("\n ✅ No conflicts detected.")
# ─────────────────────────────────────────────
# Sample data for --demo mode
# ─────────────────────────────────────────────
SAMPLE_DECISIONS_MD = f"""# Board Meeting Decisions — Layer 2
This file contains ONLY founder-approved decisions.
---
## 2026-02-15 — Spain Market Expansion
**Decision:** Expand to Spain in Q3 2026 with a pilot in Madrid and Barcelona.
**Owner:** CMO
**Deadline:** 2026-03-01
**Review:** 2026-04-01
**Rationale:** Market research shows 40% lower CAC than Germany. Two pilot customers already committed.
**User Override:** Founder reduced pilot scope from 5 cities to 2. Reason: reduce operational risk during expansion.
**Rejected:**
- Launch in all of Spain simultaneously — too resource-intensive at current headcount [DO_NOT_RESURFACE]
- Partner with a local distributor instead of direct sales — margins too low [DO_NOT_RESURFACE]
**Action Items:**
- [x] Hire Spanish-speaking CSM — Owner: CHRO — Completed: 2026-02-28 — Result: Hired Maria G., starts March 10
- [ ] Finalize Madrid pilot customer contracts — Owner: CRO — Due: {(date.today() - timedelta(days=3)).strftime('%Y-%m-%d')} — Review: 2026-04-01
- [ ] Translate app to Spanish (ES-ES) — Owner: CTO — Due: {(date.today() + timedelta(days=5)).strftime('%Y-%m-%d')} — Review: 2026-04-15
**Supersedes:**
**Superseded by:**
**Raw transcript:** memory/board-meetings/2026-02-15-raw.md
---
## 2026-02-28 — Pricing Strategy Revision
**Decision:** Move from per-seat to usage-based pricing effective Q2 2026.
**Owner:** CFO
**Deadline:** 2026-03-20
**Review:** 2026-05-01
**Rationale:** Usage-based aligns with customer value. Three enterprise customers requested it explicitly.
**User Override:**
**Rejected:**
- Freemium tier — not appropriate for enterprise healthcare segment [DO_NOT_RESURFACE]
- Raise prices 30% across the board — too aggressive without usage data [DO_NOT_RESURFACE]
**Action Items:**
- [ ] Model 3 pricing scenarios (conservative/base/aggressive) — Owner: CFO — Due: {(date.today() - timedelta(days=1)).strftime('%Y-%m-%d')} — Review: 2026-03-25
- [ ] Customer interviews on usage patterns (n=10) — Owner: CMO — Due: {(date.today() + timedelta(days=10)).strftime('%Y-%m-%d')} — Review: 2026-04-01
- [ ] Update billing infrastructure for usage tracking — Owner: CTO — Due: 2026-04-01 — Review: 2026-04-15
**Supersedes:**
**Superseded by:**
**Raw transcript:** memory/board-meetings/2026-02-28-raw.md
---
## 2026-03-04 — Engineering Hiring Plan Q2
**Decision:** Hire 2 senior engineers in Q2: one ML/AI, one backend. No contractors.
**Owner:** CTO
**Deadline:** 2026-04-15
**Review:** 2026-05-01
**Rationale:** ML roadmap blocked. Backend capacity at 85%. Contractors rejected due to IP risk in regulated domain.
**User Override:** Founder added: "ML hire must have healthcare AI experience. Non-negotiable."
**Rejected:**
- Contract team of 5 for 3 months — IP risk in regulated domain [DO_NOT_RESURFACE]
- Hire junior engineers to save budget — wrong tradeoff at this stage [DO_NOT_RESURFACE]
**Action Items:**
- [ ] Post ML engineer JD — Owner: CHRO — Due: {(date.today() + timedelta(days=2)).strftime('%Y-%m-%d')} — Review: 2026-03-20
- [ ] Post backend engineer JD — Owner: CHRO — Due: {(date.today() + timedelta(days=2)).strftime('%Y-%m-%d')} — Review: 2026-03-20
- [ ] Define ML role requirements with healthcare AI spec — Owner: CTO — Due: {(date.today() + timedelta(days=1)).strftime('%Y-%m-%d')} — Review: 2026-03-15
**Supersedes:**
**Superseded by:**
**Raw transcript:** memory/board-meetings/2026-03-04-raw.md
"""
# ─────────────────────────────────────────────
# Main
# ─────────────────────────────────────────────
def load_decisions(decisions_path: Path, demo: bool) -> list[Decision]:
if demo:
content = SAMPLE_DECISIONS_MD
elif decisions_path.exists():
content = decisions_path.read_text(encoding="utf-8")
else:
print(f" ⚠️ decisions.md not found at: {decisions_path}")
print(f" Run with --demo to see sample output.")
print(f" To initialize: mkdir -p memory/board-meetings && touch memory/board-meetings/decisions.md")
sys.exit(1)
return parse_decisions(content)
def main():
parser = argparse.ArgumentParser(
description="Board Meeting Decision Tracker",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("--file", default="memory/board-meetings/decisions.md",
help="Path to decisions.md (default: memory/board-meetings/decisions.md)")
parser.add_argument("--demo", action="store_true",
help="Run with built-in sample data (no file needed)")
parser.add_argument("--summary", action="store_true",
help="Show overview: counts, overdue, recent decisions")
parser.add_argument("--overdue", action="store_true",
help="List all overdue action items")
parser.add_argument("--due-within", type=int, metavar="DAYS",
help="List items due within N days")
parser.add_argument("--owner", metavar="ROLE",
help="Filter action items by owner")
parser.add_argument("--search", metavar="QUERY",
help="Search decisions and rejected proposals")
parser.add_argument("--conflicts", action="store_true",
help="Check for contradictory decisions or DO_NOT_RESURFACE violations")
parser.add_argument("--all", action="store_true",
help="Show all decisions (summary format)")
args = parser.parse_args()
if not any([args.summary, args.overdue, args.due_within, args.owner,
args.search, args.conflicts, getattr(args, "all")]):
args.summary = True # Default action
decisions_path = Path(args.file)
decisions = load_decisions(decisions_path, args.demo)
if not decisions:
print(" No decisions found in decisions.md.")
sys.exit(0)
if args.demo:
print(f"\n 🎯 DEMO MODE — using built-in sample data ({len(decisions)} decisions)")
if args.summary:
report_summary(decisions)
if args.overdue:
report_overdue(decisions)
if args.due_within:
report_due_within(decisions, args.due_within)
if args.owner:
report_by_owner(decisions, args.owner)
if args.search:
report_search(decisions, args.search)
if args.conflicts:
report_conflicts(decisions)
if getattr(args, "all"):
print_section(f"ALL DECISIONS ({len(decisions)} total)")
for d in sorted(decisions, key=lambda x: x.date or date.min, reverse=True):
status = "📦 SUPERSEDED" if not d.is_active() else ""
override = " [OVERRIDE]" if d.has_override() else ""
print(f"\n [{fmt_date(d.date)}] {d.title} {status}{override}")
print(f" Decision: {d.decision}")
print(f" Owner: {d.owner or '—'} | Deadline: {fmt_date(d.deadline)}")
open_actions = [a for a in d.action_items if not a.completed]
if open_actions:
print(f" Open actions: {len(open_actions)}")
print()
if __name__ == "__main__":
main()
FILE:templates/decision-entry.md
# Decision Entry Template
Single entry for `memory/board-meetings/decisions.md`.
Copy this block and fill it in after each approved board decision.
---
```markdown
## [YYYY-MM-DD] — [AGENDA ITEM TITLE]
**Decision:** [One clear statement of what was decided.]
**Owner:** [Role or name. One person. If it needs two, the first is accountable.]
**Deadline:** [YYYY-MM-DD]
**Review:** [YYYY-MM-DD — when to check. Usually 2–4 weeks after deadline.]
**Rationale:** [Why this over alternatives. 1-2 sentences. No fluff.]
**User Override:**
<!-- Leave blank if founder approved the agent recommendation.
Fill in if founder changed something:
"Founder rejected [agent recommendation] because [reason].
Actual decision: [what founder decided instead]." -->
**Rejected:**
<!-- List every proposal explicitly rejected in this discussion.
These must not be resurfaced without new information. -->
- [Proposal text] — [reason for rejection] [DO_NOT_RESURFACE]
**Action Items:**
- [ ] [Specific action] — Owner: [name] — Due: [YYYY-MM-DD] — Review: [YYYY-MM-DD]
- [ ] [Specific action] — Owner: [name] — Due: [YYYY-MM-DD] — Review: [YYYY-MM-DD]
**Supersedes:** <!-- DATE of the previous decision on this topic, if any -->
**Superseded by:** <!-- Leave blank. Will be filled in if a later decision overrides this. -->
**Raw transcript:** memory/board-meetings/[YYYY-MM-DD]-raw.md
```
---
## Field Rules
| Field | Rule |
|-------|------|
| Decision | Must be a single statement. If it takes two sentences, split into two decisions. |
| Owner | One person or role. "Everyone" owns nothing. |
| Deadline | Required. No "TBD". If unknown, set 14 days and review. |
| Review | Always set. Minimum 1 day after deadline. |
| Rationale | Required. "Because we decided so" is not rationale. |
| User Override | Honest record. Do not soften or omit. |
| Rejected | Every rejected proposal must be listed. |
| DO_NOT_RESURFACE | Applied to every rejected item. No exceptions. |
---
## Marking Action Items Complete
When an action item is done, update the entry in decisions.md:
```markdown
- [x] [Action text] — Owner: [name] — Completed: [YYYY-MM-DD] — Result: [one sentence outcome]
```
Do not delete completed items. The history is the record.
Sửa các test Playwright bị lỗi hoặc chập chờn, gỡ lỗi test hỏng và lỗi xuất hiện ngẫu nhiên.
---
name: "fix"
description: >-
Fix failing or flaky Playwright tests. Use when user says "fix test",
"flaky test", "test failing", "debug test", "test broken", "test passes
sometimes", or "intermittent failure".
---
# Fix Failing or Flaky Tests
Diagnose and fix a Playwright test that fails or passes intermittently using a systematic taxonomy.
## Input
`$ARGUMENTS` contains:
- A test file path: `e2e/login.spec.ts`
- A test name: ""should redirect after login"`
- A description: `"the checkout test fails in CI but passes locally"`
## Steps
### 1. Reproduce the Failure
Run the test to capture the error:
```bash
npx playwright test <file> --reporter=list
```
If the test passes, it's likely flaky. Run burn-in:
```bash
npx playwright test <file> --repeat-each=10 --reporter=list
```
If it still passes, try with parallel workers:
```bash
npx playwright test --fully-parallel --workers=4 --repeat-each=5
```
### 2. Capture Trace
Run with full tracing:
```bash
npx playwright test <file> --trace=on --retries=0
```
Read the trace output. Use `/debug` to analyze trace files if available.
### 3. Categorize the Failure
Load `flaky-taxonomy.md` from this skill directory.
Every failing test falls into one of four categories:
| Category | Symptom | Diagnosis |
|---|---|---|
| **Timing/Async** | Fails intermittently everywhere | `--repeat-each=20` reproduces locally |
| **Test Isolation** | Fails in suite, passes alone | `--workers=1 --grep "test name"` passes |
| **Environment** | Fails in CI, passes locally | Compare CI vs local screenshots/traces |
| **Infrastructure** | Random, no pattern | Error references browser internals |
### 4. Apply Targeted Fix
**Timing/Async:**
- Replace `waitForTimeout()` with web-first assertions
- Add `await` to missing Playwright calls
- Wait for specific network responses before asserting
- Use `toBeVisible()` before interacting with elements
**Test Isolation:**
- Remove shared mutable state between tests
- Create test data per-test via API or fixtures
- Use unique identifiers (timestamps, random strings) for test data
- Check for database state leaks
**Environment:**
- Match viewport sizes between local and CI
- Account for font rendering differences in screenshots
- Use `docker` locally to match CI environment
- Check for timezone-dependent assertions
**Infrastructure:**
- Increase timeout for slow CI runners
- Add retries in CI config (`retries: 2`)
- Check for browser OOM (reduce parallel workers)
- Ensure browser dependencies are installed
### 5. Verify the Fix
Run the test 10 times to confirm stability:
```bash
npx playwright test <file> --repeat-each=10 --reporter=list
```
All 10 must pass. If any fail, go back to step 3.
### 6. Prevent Recurrence
Suggest:
- Add to CI with `retries: 2` if not already
- Enable `trace: 'on-first-retry'` in config
- Add the fix pattern to project's test conventions doc
## Output
- Root cause category and specific issue
- The fix applied (with diff)
- Verification result (10/10 passes)
- Prevention recommendation
FILE:flaky-taxonomy.md
# Flaky Test Taxonomy
## Decision Tree
```
Test is flaky
│
├── Fails locally with --repeat-each=20?
│ ├── YES → TIMING / ASYNC
│ │ ├── Missing await? → Add await
│ │ ├── waitForTimeout? → Replace with assertion
│ │ ├── Race condition? → Wait for specific event
│ │ └── Animation? → Wait for animation end or disable
│ │
│ └── NO → Continue...
│
├── Passes alone, fails in suite?
│ ├── YES → TEST ISOLATION
│ │ ├── Shared variable? → Make per-test
│ │ ├── Database state? → Reset per-test
│ │ ├── localStorage? → Clear in beforeEach
│ │ └── Cookie leak? → Use isolated contexts
│ │
│ └── NO → Continue...
│
├── Fails in CI, passes locally?
│ ├── YES → ENVIRONMENT
│ │ ├── Viewport? → Set explicit size
│ │ ├── Fonts? → Use Docker locally
│ │ ├── Timezone? → Use UTC everywhere
│ │ └── Network? → Mock external services
│ │
│ └── NO → INFRASTRUCTURE
│ ├── Browser crash? → Reduce workers
│ ├── OOM? → Limit parallel tests
│ ├── DNS? → Add retry config
│ └── File system? → Use unique temp dirs
```
## Common Fixes by Category
### Timing / Async
**Missing await:**
```typescript
// BAD — race condition
page.goto('/dashboard');
expect(page.getByText('Welcome')).toBeVisible();
// GOOD
await page.goto('/dashboard');
await expect(page.getByText('Welcome')).toBeVisible();
```
**Clicking before visible:**
```typescript
// BAD — element may not be ready
await page.getByRole('button', { name: 'Submit' }).click();
// GOOD — ensure visible first
const submitBtn = page.getByRole('button', { name: 'Submit' });
await expect(submitBtn).toBeVisible();
await submitBtn.click();
```
**Race with network:**
```typescript
// BAD — data might not be loaded
await page.goto('/users');
await expect(page.getByRole('table')).toBeVisible();
// GOOD — wait for API response
const responsePromise = page.waitForResponse('**/api/users');
await page.goto('/users');
await responsePromise;
await expect(page.getByRole('table')).toBeVisible();
```
### Test Isolation
**Shared state fix:**
```typescript
// BAD — tests share userId
let userId: string;
test('create', async () => { userId = '123'; });
test('read', async () => { /* uses userId */ });
// GOOD — each test is independent
test('read user', async ({ request }) => {
const response = await request.post('/api/users', { data: { name: 'Test' } });
const { id } = await response.json();
// Use id within this test
});
```
**localStorage cleanup:**
```typescript
test.beforeEach(async ({ page }) => {
await page.goto('/');
await page.evaluate(() => localStorage.clear());
});
```
### Environment
**Explicit viewport:**
```typescript
test.use({ viewport: { width: 1280, height: 720 } });
```
**Timezone-safe dates:**
```typescript
// BAD
expect(dateText).toBe('March 5, 2026');
// GOOD — timezone independent
expect(dateText).toMatch(/\d{1,2}\/\d{1,2}\/\d{4}/);
```
### Infrastructure
**Retry config:**
```typescript
// playwright.config.ts
export default defineConfig({
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? 2 : undefined,
});
```
**Increase timeout for CI:**
```typescript
test.setTimeout(60_000); // 60s for slow CI
```
Thiết kế quy trình phỏng vấn, pipeline tuyển dụng, bộ câu hỏi, ma trận năng lực, thang chấm điểm và phân tích thiên kiến người phỏng vấn.
---
name: "interview-system-designer"
description: This skill should be used when the user asks to "design interview processes", "create hiring pipelines", "calibrate interview loops", "generate interview questions", "design competency matrices", "analyze interviewer bias", "create scoring rubrics", "build question banks", or "optimize hiring systems". Use for designing role-specific interview loops, competency assessments, and hiring calibration systems.
---
# Interview System Designer
Comprehensive interview loop planning and calibration support for role-based hiring systems.
## Overview
Use this skill to create structured interview loops, standardize question quality, and keep hiring signal consistent across interviewers.
## Core Capabilities
- Interview loop planning by role and level
- Round-by-round focus and timing recommendations
- Suggested question sets by round type
- Framework support for scoring and calibration
- Bias-reduction and process consistency guidance
## Quick Start
```bash
# Generate a loop plan for a role and level
python3 scripts/interview_planner.py --role "Senior Software Engineer" --level senior
# JSON output for integration with internal tooling
python3 scripts/interview_planner.py --role "Product Manager" --level mid --json
```
## Recommended Workflow
1. Run `scripts/interview_planner.py` to generate a baseline loop.
2. Align rounds to role-specific competencies.
3. Validate scoring rubric consistency with interview panel leads.
4. Review for bias controls before rollout.
5. Recalibrate quarterly using hiring outcome data.
## References
- `references/interview-frameworks.md`
- `references/bias_mitigation_checklist.md`
- `references/competency_matrix_templates.md`
- `references/debrief_facilitation_guide.md`
## Common Pitfalls
- Overweighting one round while ignoring other competency signals
- Using unstructured interviews without standardized scoring
- Skipping calibration sessions for interviewers
- Changing hiring bar without documenting rationale
## Best Practices
1. Keep round objectives explicit and non-overlapping.
2. Require evidence for each score recommendation.
3. Use the same baseline rubric across comparable roles.
4. Revisit loop design based on quality-of-hire outcomes.
FILE:assets/sample_interview_results.json
[
{
"candidate_id": "candidate_001",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-15T09:00:00Z",
"scores": {
"coding_fundamentals": 3.5,
"system_design": 4.0,
"technical_leadership": 3.0,
"communication": 3.5,
"problem_solving": 4.0
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "asian",
"years_experience": 6,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_001",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_bob",
"date": "2024-01-15T11:00:00Z",
"scores": {
"system_design": 3.5,
"technical_leadership": 3.5,
"mentoring": 3.0,
"cross_team_collaboration": 4.0,
"strategic_thinking": 3.5
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "asian",
"years_experience": 6,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_002",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-16T09:00:00Z",
"scores": {
"coding_fundamentals": 2.5,
"system_design": 3.0,
"technical_leadership": 2.0,
"communication": 3.0,
"problem_solving": 3.0
},
"overall_recommendation": "No Hire",
"gender": "female",
"ethnicity": "hispanic",
"years_experience": 5,
"university_tier": "tier_2",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_002",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_charlie",
"date": "2024-01-16T11:00:00Z",
"scores": {
"system_design": 2.0,
"technical_leadership": 2.5,
"mentoring": 2.0,
"cross_team_collaboration": 3.0,
"strategic_thinking": 2.5
},
"overall_recommendation": "No Hire",
"gender": "female",
"ethnicity": "hispanic",
"years_experience": 5,
"university_tier": "tier_2",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_003",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_david",
"date": "2024-01-17T14:00:00Z",
"scores": {
"coding_fundamentals": 4.0,
"system_design": 3.5,
"technical_leadership": 4.0,
"communication": 4.0,
"problem_solving": 3.5
},
"overall_recommendation": "Strong Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 8,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_003",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-17T16:00:00Z",
"scores": {
"system_design": 4.0,
"technical_leadership": 4.0,
"mentoring": 3.5,
"cross_team_collaboration": 4.0,
"strategic_thinking": 3.5
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 8,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_004",
"role": "Product Manager",
"interviewer_id": "interviewer_emma",
"date": "2024-01-18T10:00:00Z",
"scores": {
"product_strategy": 3.0,
"user_research": 3.5,
"data_analysis": 4.0,
"stakeholder_management": 3.0,
"communication": 3.5
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "black",
"years_experience": 4,
"university_tier": "tier_2",
"previous_company_size": "medium"
},
{
"candidate_id": "candidate_005",
"role": "Product Manager",
"interviewer_id": "interviewer_frank",
"date": "2024-01-19T13:00:00Z",
"scores": {
"product_strategy": 2.5,
"user_research": 2.0,
"data_analysis": 3.0,
"stakeholder_management": 2.5,
"communication": 3.0
},
"overall_recommendation": "No Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 3,
"university_tier": "tier_3",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_006",
"role": "Junior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-20T09:00:00Z",
"scores": {
"coding_fundamentals": 3.0,
"debugging": 3.5,
"testing_basics": 3.0,
"collaboration": 4.0,
"learning_agility": 3.5
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "asian",
"years_experience": 1,
"university_tier": "bootcamp",
"previous_company_size": "none"
},
{
"candidate_id": "candidate_007",
"role": "Junior Software Engineer",
"interviewer_id": "interviewer_bob",
"date": "2024-01-21T10:30:00Z",
"scores": {
"coding_fundamentals": 2.0,
"debugging": 2.5,
"testing_basics": 2.0,
"collaboration": 3.0,
"learning_agility": 3.0
},
"overall_recommendation": "No Hire",
"gender": "male",
"ethnicity": "hispanic",
"years_experience": 0,
"university_tier": "tier_2",
"previous_company_size": "none"
},
{
"candidate_id": "candidate_008",
"role": "Staff Frontend Engineer",
"interviewer_id": "interviewer_grace",
"date": "2024-01-22T14:00:00Z",
"scores": {
"frontend_architecture": 4.0,
"system_design": 4.0,
"technical_leadership": 4.0,
"team_building": 3.5,
"strategic_thinking": 3.5
},
"overall_recommendation": "Strong Hire",
"gender": "female",
"ethnicity": "white",
"years_experience": 9,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_008",
"role": "Staff Frontend Engineer",
"interviewer_id": "interviewer_henry",
"date": "2024-01-22T16:00:00Z",
"scores": {
"frontend_architecture": 3.5,
"technical_leadership": 4.0,
"team_building": 4.0,
"cross_functional_collaboration": 4.0,
"organizational_impact": 3.5
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "white",
"years_experience": 9,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_009",
"role": "Data Scientist",
"interviewer_id": "interviewer_ivan",
"date": "2024-01-23T11:00:00Z",
"scores": {
"statistical_analysis": 3.5,
"machine_learning": 4.0,
"data_engineering": 3.0,
"business_acumen": 3.5,
"communication": 3.0
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "indian",
"years_experience": 5,
"university_tier": "tier_1",
"previous_company_size": "medium"
},
{
"candidate_id": "candidate_010",
"role": "DevOps Engineer",
"interviewer_id": "interviewer_jane",
"date": "2024-01-24T15:00:00Z",
"scores": {
"infrastructure_automation": 3.5,
"ci_cd_design": 4.0,
"monitoring_observability": 3.0,
"security_implementation": 3.5,
"incident_management": 4.0
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "black",
"years_experience": 6,
"university_tier": "tier_2",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_011",
"role": "UX Designer",
"interviewer_id": "interviewer_karl",
"date": "2024-01-25T10:00:00Z",
"scores": {
"design_process": 4.0,
"user_research": 3.5,
"design_systems": 4.0,
"cross_functional_collaboration": 3.5,
"design_leadership": 3.0
},
"overall_recommendation": "Hire",
"gender": "non_binary",
"ethnicity": "white",
"years_experience": 7,
"university_tier": "tier_1",
"previous_company_size": "medium"
},
{
"candidate_id": "candidate_012",
"role": "Engineering Manager",
"interviewer_id": "interviewer_lisa",
"date": "2024-01-26T13:30:00Z",
"scores": {
"people_leadership": 4.0,
"technical_background": 3.5,
"strategic_thinking": 3.5,
"performance_management": 4.0,
"cross_functional_leadership": 3.5
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 8,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_013",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-27T09:00:00Z",
"scores": {
"coding_fundamentals": 4.0,
"system_design": 4.0,
"technical_leadership": 4.0,
"communication": 4.0,
"problem_solving": 4.0
},
"overall_recommendation": "Strong Hire",
"gender": "female",
"ethnicity": "asian",
"years_experience": 7,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_013",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_charlie",
"date": "2024-01-27T11:00:00Z",
"scores": {
"system_design": 3.5,
"technical_leadership": 3.5,
"mentoring": 4.0,
"cross_team_collaboration": 4.0,
"strategic_thinking": 3.5
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "asian",
"years_experience": 7,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_014",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_david",
"date": "2024-01-28T14:00:00Z",
"scores": {
"coding_fundamentals": 1.5,
"system_design": 2.0,
"technical_leadership": 1.0,
"communication": 2.0,
"problem_solving": 2.0
},
"overall_recommendation": "Strong No Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 4,
"university_tier": "tier_3",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_015",
"role": "Product Manager",
"interviewer_id": "interviewer_emma",
"date": "2024-01-29T11:00:00Z",
"scores": {
"product_strategy": 4.0,
"user_research": 3.5,
"data_analysis": 4.0,
"stakeholder_management": 4.0,
"communication": 3.5
},
"overall_recommendation": "Strong Hire",
"gender": "male",
"ethnicity": "black",
"years_experience": 5,
"university_tier": "tier_2",
"previous_company_size": "medium"
}
]
FILE:assets/sample_role_definitions.json
[
{
"role": "Senior Software Engineer",
"level": "senior",
"team": "platform",
"department": "engineering",
"competencies": [
"system_design",
"coding_fundamentals",
"technical_leadership",
"mentoring",
"cross_team_collaboration"
],
"requirements": {
"years_experience": "5-8",
"technical_skills": ["Python", "Java", "Docker", "Kubernetes", "AWS"],
"leadership_experience": true,
"mentoring_required": true
},
"hiring_bar": "high",
"interview_focus": ["technical_depth", "system_architecture", "leadership_potential"]
},
{
"role": "Product Manager",
"level": "mid",
"team": "growth",
"department": "product",
"competencies": [
"product_strategy",
"user_research",
"data_analysis",
"stakeholder_management",
"cross_functional_leadership"
],
"requirements": {
"years_experience": "3-5",
"domain_knowledge": ["user_analytics", "experimentation", "product_metrics"],
"leadership_experience": false,
"technical_background": "preferred"
},
"hiring_bar": "medium-high",
"interview_focus": ["product_sense", "analytical_thinking", "execution_ability"]
},
{
"role": "Staff Frontend Engineer",
"level": "staff",
"team": "consumer",
"department": "engineering",
"competencies": [
"frontend_architecture",
"system_design",
"technical_leadership",
"team_building",
"cross_functional_collaboration"
],
"requirements": {
"years_experience": "8+",
"technical_skills": ["React", "TypeScript", "GraphQL", "Webpack", "Performance Optimization"],
"leadership_experience": true,
"architecture_experience": true
},
"hiring_bar": "very-high",
"interview_focus": ["architectural_vision", "technical_strategy", "organizational_impact"]
},
{
"role": "Data Scientist",
"level": "mid",
"team": "ml_platform",
"department": "data",
"competencies": [
"statistical_analysis",
"machine_learning",
"data_engineering",
"business_acumen",
"communication"
],
"requirements": {
"years_experience": "3-6",
"technical_skills": ["Python", "SQL", "TensorFlow", "Spark", "Statistics"],
"domain_knowledge": ["ML algorithms", "experimentation", "data_pipelines"],
"leadership_experience": false
},
"hiring_bar": "high",
"interview_focus": ["technical_depth", "problem_solving", "business_impact"]
},
{
"role": "DevOps Engineer",
"level": "senior",
"team": "infrastructure",
"department": "engineering",
"competencies": [
"infrastructure_automation",
"ci_cd_design",
"monitoring_observability",
"security_implementation",
"incident_management"
],
"requirements": {
"years_experience": "5-7",
"technical_skills": ["Kubernetes", "Terraform", "AWS", "Docker", "Monitoring"],
"security_background": "required",
"leadership_experience": "preferred"
},
"hiring_bar": "high",
"interview_focus": ["system_reliability", "automation_expertise", "operational_excellence"]
},
{
"role": "UX Designer",
"level": "senior",
"team": "design_systems",
"department": "design",
"competencies": [
"design_process",
"user_research",
"design_systems",
"cross_functional_collaboration",
"design_leadership"
],
"requirements": {
"years_experience": "5-8",
"portfolio_quality": "high",
"research_experience": true,
"systems_thinking": true
},
"hiring_bar": "high",
"interview_focus": ["design_process", "systems_thinking", "user_advocacy"]
},
{
"role": "Engineering Manager",
"level": "senior",
"team": "backend",
"department": "engineering",
"competencies": [
"people_leadership",
"technical_background",
"strategic_thinking",
"performance_management",
"cross_functional_leadership"
],
"requirements": {
"years_experience": "6-10",
"management_experience": "2+ years",
"technical_background": "required",
"hiring_experience": true
},
"hiring_bar": "very-high",
"interview_focus": ["people_leadership", "technical_judgment", "organizational_impact"]
},
{
"role": "Junior Software Engineer",
"level": "junior",
"team": "web",
"department": "engineering",
"competencies": [
"coding_fundamentals",
"debugging",
"testing_basics",
"collaboration",
"learning_agility"
],
"requirements": {
"years_experience": "0-2",
"technical_skills": ["JavaScript", "HTML/CSS", "Git", "Basic Algorithms"],
"education": "CS degree or bootcamp",
"growth_mindset": true
},
"hiring_bar": "medium",
"interview_focus": ["coding_ability", "problem_solving", "potential_assessment"]
}
]
FILE:expected_outputs/product_manager_senior_questions.json
{
"role": "Product Manager",
"level": "senior",
"competencies": [
"strategy",
"analytics",
"business_strategy",
"product_strategy",
"stakeholder_management",
"p&l_responsibility",
"leadership",
"team_leadership",
"user_research",
"data_analysis"
],
"question_types": [
"technical",
"behavioral",
"situational"
],
"generated_at": "2026-02-16T13:27:41.303329",
"total_questions": 20,
"questions": [
{
"question": "What challenges have you faced related to p&l responsibility and how did you overcome them?",
"competency": "p&l_responsibility",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "Analyze conversion funnel data to identify the biggest drop-off point and propose solutions.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": [
"funnel_analysis",
"conversion_optimization",
"statistical_significance"
]
},
{
"question": "What challenges have you faced related to team leadership and how did you overcome them?",
"competency": "team_leadership",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "Design a go-to-market strategy for a new B2B SaaS product entering a competitive market.",
"competency": "product_strategy",
"type": "strategic",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": [
"market_analysis",
"competitive_positioning",
"pricing_strategy",
"channel_strategy"
]
},
{
"question": "What challenges have you faced related to business strategy and how did you overcome them?",
"competency": "business_strategy",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "Describe your experience with business strategy in your current or previous role.",
"competency": "business_strategy",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Describe your experience with team leadership in your current or previous role.",
"competency": "team_leadership",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Describe a situation where you had to influence someone without having direct authority over them.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": [
"influence",
"persuasion",
"stakeholder_management"
]
},
{
"question": "Given a dataset of user activities, calculate the daily active users for the past month.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "easy",
"time_limit": 30,
"key_concepts": [
"sql_basics",
"date_functions",
"aggregation"
]
},
{
"question": "Describe your experience with analytics in your current or previous role.",
"competency": "analytics",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "How would you prioritize features for a mobile app with limited engineering resources?",
"competency": "product_strategy",
"type": "case_study",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": [
"prioritization_frameworks",
"resource_allocation",
"impact_estimation"
]
},
{
"question": "Describe your experience with stakeholder management in your current or previous role.",
"competency": "stakeholder_management",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "What challenges have you faced related to stakeholder management and how did you overcome them?",
"competency": "stakeholder_management",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "What challenges have you faced related to user research and how did you overcome them?",
"competency": "user_research",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "What challenges have you faced related to strategy and how did you overcome them?",
"competency": "strategy",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "Describe your experience with user research in your current or previous role.",
"competency": "user_research",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Describe your experience with p&l responsibility in your current or previous role.",
"competency": "p&l_responsibility",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Describe your experience with strategy in your current or previous role.",
"competency": "strategy",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Tell me about a time when you had to lead a team through a significant change or challenge.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": [
"change_management",
"team_motivation",
"communication"
]
},
{
"question": "What challenges have you faced related to analytics and how did you overcome them?",
"competency": "analytics",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
}
],
"scoring_rubrics": {
"question_8": {
"question": "Describe a situation where you had to influence someone without having direct authority over them.",
"competency": "leadership",
"type": "behavioral",
"scoring_criteria": {
"situation_clarity": {
"4": "Clear, specific situation with relevant context and stakes",
"3": "Good situation description with adequate context",
"2": "Situation described but lacks some specifics",
"1": "Vague or unclear situation description"
},
"action_quality": {
"4": "Specific, thoughtful actions showing strong competency",
"3": "Good actions demonstrating competency",
"2": "Adequate actions but could be stronger",
"1": "Weak or inappropriate actions"
},
"result_impact": {
"4": "Significant positive impact with measurable results",
"3": "Good positive impact with clear outcomes",
"2": "Some positive impact demonstrated",
"1": "Little or no positive impact shown"
},
"self_awareness": {
"4": "Excellent self-reflection, learns from experience, acknowledges growth areas",
"3": "Good self-awareness and learning orientation",
"2": "Some self-reflection demonstrated",
"1": "Limited self-awareness or reflection"
}
},
"weight": "high",
"time_limit": 30
},
"question_19": {
"question": "Tell me about a time when you had to lead a team through a significant change or challenge.",
"competency": "leadership",
"type": "behavioral",
"scoring_criteria": {
"situation_clarity": {
"4": "Clear, specific situation with relevant context and stakes",
"3": "Good situation description with adequate context",
"2": "Situation described but lacks some specifics",
"1": "Vague or unclear situation description"
},
"action_quality": {
"4": "Specific, thoughtful actions showing strong competency",
"3": "Good actions demonstrating competency",
"2": "Adequate actions but could be stronger",
"1": "Weak or inappropriate actions"
},
"result_impact": {
"4": "Significant positive impact with measurable results",
"3": "Good positive impact with clear outcomes",
"2": "Some positive impact demonstrated",
"1": "Little or no positive impact shown"
},
"self_awareness": {
"4": "Excellent self-reflection, learns from experience, acknowledges growth areas",
"3": "Good self-awareness and learning orientation",
"2": "Some self-reflection demonstrated",
"1": "Limited self-awareness or reflection"
}
},
"weight": "high",
"time_limit": 30
}
},
"follow_up_probes": {
"question_1": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_2": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_3": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_4": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_5": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_6": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_7": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_8": [
"What would you do differently if you faced this situation again?",
"How did you handle team members who were resistant to the change?",
"What metrics did you use to measure success?",
"How did you communicate progress to stakeholders?",
"What did you learn from this experience?"
],
"question_9": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_10": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_11": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_12": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_13": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_14": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_15": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_16": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_17": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_18": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_19": [
"What would you do differently if you faced this situation again?",
"How did you handle team members who were resistant to the change?",
"What metrics did you use to measure success?",
"How did you communicate progress to stakeholders?",
"What did you learn from this experience?"
],
"question_20": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
]
},
"calibration_examples": {
"question_1": {
"question": "What challenges have you faced related to p&l responsibility and how did you overcome them?",
"competency": "p&l_responsibility",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for p&l_responsibility question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for p&l_responsibility question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for p&l_responsibility question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of p&l responsibility competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
},
"question_2": {
"question": "Analyze conversion funnel data to identify the biggest drop-off point and propose solutions.",
"competency": "data_analysis",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for data_analysis question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for data_analysis question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for data_analysis question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of data analysis competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
},
"question_3": {
"question": "What challenges have you faced related to team leadership and how did you overcome them?",
"competency": "team_leadership",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for team_leadership question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for team_leadership question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for team_leadership question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of team leadership competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
},
"question_4": {
"question": "Design a go-to-market strategy for a new B2B SaaS product entering a competitive market.",
"competency": "product_strategy",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for product_strategy question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for product_strategy question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for product_strategy question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of product strategy competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
},
"question_5": {
"question": "What challenges have you faced related to business strategy and how did you overcome them?",
"competency": "business_strategy",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for business_strategy question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for business_strategy question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for business_strategy question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of business strategy competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
}
},
"usage_guidelines": {
"interview_flow": {
"warm_up": "Start with 1-2 easier questions to build rapport",
"core_assessment": "Focus majority of time on core competency questions",
"closing": "End with questions about candidate's questions/interests"
},
"time_management": {
"technical_questions": "Allow extra time for coding/design questions",
"behavioral_questions": "Keep to time limits but allow for follow-ups",
"total_recommendation": "45-75 minutes per interview round"
},
"question_selection": {
"variety": "Mix question types within each competency area",
"difficulty": "Adjust based on candidate responses and energy",
"customization": "Adapt questions based on candidate's background"
},
"common_mistakes": [
"Don't ask all questions mechanically",
"Don't skip follow-up questions",
"Don't forget to assess cultural fit alongside competencies",
"Don't let one strong/weak area bias overall assessment"
],
"calibration_reminders": [
"Compare against role standard, not other candidates",
"Focus on evidence demonstrated, not potential",
"Consider level-appropriate expectations",
"Document specific examples in feedback"
]
}
}
FILE:expected_outputs/product_manager_senior_questions.txt
Interview Question Bank: Product Manager (Senior Level)
======================================================================
Generated: 2026-02-16T13:27:41.303329
Total Questions: 20
Question Types: technical, behavioral, situational
Target Competencies: strategy, analytics, business_strategy, product_strategy, stakeholder_management, p&l_responsibility, leadership, team_leadership, user_research, data_analysis
INTERVIEW QUESTIONS
--------------------------------------------------
1. What challenges have you faced related to p&l responsibility and how did you overcome them?
Competency: P&L Responsibility
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
2. Analyze conversion funnel data to identify the biggest drop-off point and propose solutions.
Competency: Data Analysis
Type: Analytical
Time Limit: 45 minutes
3. What challenges have you faced related to team leadership and how did you overcome them?
Competency: Team Leadership
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
4. Design a go-to-market strategy for a new B2B SaaS product entering a competitive market.
Competency: Product Strategy
Type: Strategic
Time Limit: 60 minutes
5. What challenges have you faced related to business strategy and how did you overcome them?
Competency: Business Strategy
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
6. Describe your experience with business strategy in your current or previous role.
Competency: Business Strategy
Type: Experience
Focus Areas: experience_depth, practical_application
7. Describe your experience with team leadership in your current or previous role.
Competency: Team Leadership
Type: Experience
Focus Areas: experience_depth, practical_application
8. Describe a situation where you had to influence someone without having direct authority over them.
Competency: Leadership
Type: Behavioral
Focus Areas: influence, persuasion, stakeholder_management
9. Given a dataset of user activities, calculate the daily active users for the past month.
Competency: Data Analysis
Type: Analytical
Time Limit: 30 minutes
10. Describe your experience with analytics in your current or previous role.
Competency: Analytics
Type: Experience
Focus Areas: experience_depth, practical_application
11. How would you prioritize features for a mobile app with limited engineering resources?
Competency: Product Strategy
Type: Case_Study
Time Limit: 45 minutes
12. Describe your experience with stakeholder management in your current or previous role.
Competency: Stakeholder Management
Type: Experience
Focus Areas: experience_depth, practical_application
13. What challenges have you faced related to stakeholder management and how did you overcome them?
Competency: Stakeholder Management
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
14. What challenges have you faced related to user research and how did you overcome them?
Competency: User Research
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
15. What challenges have you faced related to strategy and how did you overcome them?
Competency: Strategy
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
16. Describe your experience with user research in your current or previous role.
Competency: User Research
Type: Experience
Focus Areas: experience_depth, practical_application
17. Describe your experience with p&l responsibility in your current or previous role.
Competency: P&L Responsibility
Type: Experience
Focus Areas: experience_depth, practical_application
18. Describe your experience with strategy in your current or previous role.
Competency: Strategy
Type: Experience
Focus Areas: experience_depth, practical_application
19. Tell me about a time when you had to lead a team through a significant change or challenge.
Competency: Leadership
Type: Behavioral
Focus Areas: change_management, team_motivation, communication
20. What challenges have you faced related to analytics and how did you overcome them?
Competency: Analytics
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
SCORING RUBRICS
--------------------------------------------------
Sample Scoring Criteria (behavioral questions):
Situation Clarity:
4: Clear, specific situation with relevant context and stakes
3: Good situation description with adequate context
2: Situation described but lacks some specifics
1: Vague or unclear situation description
Action Quality:
4: Specific, thoughtful actions showing strong competency
3: Good actions demonstrating competency
2: Adequate actions but could be stronger
1: Weak or inappropriate actions
Result Impact:
4: Significant positive impact with measurable results
3: Good positive impact with clear outcomes
2: Some positive impact demonstrated
1: Little or no positive impact shown
Self Awareness:
4: Excellent self-reflection, learns from experience, acknowledges growth areas
3: Good self-awareness and learning orientation
2: Some self-reflection demonstrated
1: Limited self-awareness or reflection
FOLLOW-UP PROBE EXAMPLES
--------------------------------------------------
Sample follow-up questions:
• Can you provide more specific details about your approach?
• What would you do differently if you had to do this again?
• What challenges did you face and how did you overcome them?
USAGE GUIDELINES
--------------------------------------------------
Interview Flow:
• Warm Up: Start with 1-2 easier questions to build rapport
• Core Assessment: Focus majority of time on core competency questions
• Closing: End with questions about candidate's questions/interests
Time Management:
• Technical Questions: Allow extra time for coding/design questions
• Behavioral Questions: Keep to time limits but allow for follow-ups
• Total Recommendation: 45-75 minutes per interview round
Common Mistakes to Avoid:
• Don't ask all questions mechanically
• Don't skip follow-up questions
• Don't forget to assess cultural fit alongside competencies
CALIBRATION EXAMPLES
--------------------------------------------------
Question: What challenges have you faced related to p&l responsibility and how did you overcome them?
Sample Answer Quality Levels:
Poor Answer (Score 1-2):
Issues: Vague response, Limited evidence of competency, Poor structure
Good Answer (Score 3):
Strengths: Clear structure, Demonstrates competency, Adequate detail
Great Answer (Score 4):
Strengths: Exceptional detail, Strong evidence, Strategic thinking, Goes beyond requirements
FILE:expected_outputs/senior_software_engineer_senior_interview_loop.json
{
"role": "Senior Software Engineer",
"level": "senior",
"team": "platform",
"generated_at": "2026-02-16T13:27:37.925680",
"total_duration_minutes": 300,
"total_rounds": 5,
"rounds": {
"round_1_technical_phone_screen": {
"name": "Technical Phone Screen",
"duration_minutes": 45,
"format": "virtual",
"objectives": [
"Assess coding fundamentals",
"Evaluate problem-solving approach",
"Screen for basic technical competency"
],
"question_types": [
"coding_problems",
"technical_concepts",
"experience_questions"
],
"evaluation_criteria": [
"technical_accuracy",
"problem_solving_process",
"communication_clarity"
],
"order": 1,
"focus_areas": [
"coding_fundamentals",
"problem_solving",
"technical_leadership",
"system_architecture",
"people_development"
]
},
"round_2_coding_deep_dive": {
"name": "Coding Deep Dive",
"duration_minutes": 75,
"format": "in_person_or_virtual",
"objectives": [
"Evaluate coding skills in depth",
"Assess code quality and testing",
"Review debugging approach"
],
"question_types": [
"complex_coding_problems",
"code_review",
"testing_strategy"
],
"evaluation_criteria": [
"code_quality",
"testing_approach",
"debugging_skills",
"optimization_thinking"
],
"order": 2,
"focus_areas": [
"technical_execution",
"code_quality",
"technical_leadership",
"system_architecture",
"people_development"
]
},
"round_3_system_design": {
"name": "System Design",
"duration_minutes": 75,
"format": "collaborative_whiteboard",
"objectives": [
"Assess architectural thinking",
"Evaluate scalability considerations",
"Review trade-off analysis"
],
"question_types": [
"system_architecture",
"scalability_design",
"trade_off_analysis"
],
"evaluation_criteria": [
"architectural_thinking",
"scalability_awareness",
"trade_off_reasoning"
],
"order": 3,
"focus_areas": [
"system_thinking",
"architectural_reasoning",
"technical_leadership",
"system_architecture",
"people_development"
]
},
"round_4_behavioral": {
"name": "Behavioral Interview",
"duration_minutes": 45,
"format": "conversational",
"objectives": [
"Assess cultural fit",
"Evaluate past experiences",
"Review leadership examples"
],
"question_types": [
"star_method_questions",
"situational_scenarios",
"values_alignment"
],
"evaluation_criteria": [
"communication_skills",
"leadership_examples",
"cultural_alignment"
],
"order": 4,
"focus_areas": [
"cultural_fit",
"communication",
"teamwork",
"technical_leadership",
"system_architecture"
]
},
"round_5_technical_leadership": {
"name": "Technical Leadership",
"duration_minutes": 60,
"format": "discussion_based",
"objectives": [
"Evaluate mentoring capability",
"Assess technical decision making",
"Review cross-team collaboration"
],
"question_types": [
"leadership_scenarios",
"technical_decisions",
"mentoring_examples"
],
"evaluation_criteria": [
"leadership_potential",
"technical_judgment",
"influence_skills"
],
"order": 5,
"focus_areas": [
"leadership",
"mentoring",
"influence",
"technical_leadership",
"system_architecture"
]
}
},
"suggested_schedule": {
"type": "multi_day",
"total_duration_minutes": 300,
"recommended_breaks": [
{
"type": "short_break",
"duration": 15,
"after_minutes": 90
},
{
"type": "lunch_break",
"duration": 60,
"after_minutes": 180
}
],
"day_structure": {
"day_1": {
"date": "TBD",
"start_time": "09:00",
"end_time": "12:45",
"rounds": [
{
"type": "interview",
"round_name": "round_1_technical_phone_screen",
"title": "Technical Phone Screen",
"start_time": "09:00",
"end_time": "09:45",
"duration_minutes": 45,
"format": "virtual"
},
{
"type": "interview",
"round_name": "round_2_coding_deep_dive",
"title": "Coding Deep Dive",
"start_time": "10:00",
"end_time": "11:15",
"duration_minutes": 75,
"format": "in_person_or_virtual"
},
{
"type": "interview",
"round_name": "round_3_system_design",
"title": "System Design",
"start_time": "11:30",
"end_time": "12:45",
"duration_minutes": 75,
"format": "collaborative_whiteboard"
}
]
},
"day_2": {
"date": "TBD",
"start_time": "09:00",
"end_time": "11:00",
"rounds": [
{
"type": "interview",
"round_name": "round_4_behavioral",
"title": "Behavioral Interview",
"start_time": "09:00",
"end_time": "09:45",
"duration_minutes": 45,
"format": "conversational"
},
{
"type": "interview",
"round_name": "round_5_technical_leadership",
"title": "Technical Leadership",
"start_time": "10:00",
"end_time": "11:00",
"duration_minutes": 60,
"format": "discussion_based"
}
]
}
},
"logistics_notes": [
"Coordinate interviewer availability before scheduling",
"Ensure all interviewers have access to job description and competency requirements",
"Prepare interview rooms/virtual links for all rounds",
"Share candidate resume and application with all interviewers",
"Test video conferencing setup before virtual interviews",
"Share virtual meeting links with candidate 24 hours in advance",
"Prepare whiteboard or collaborative online tool for design sessions"
]
},
"scorecard_template": {
"scoring_scale": {
"4": "Exceeds Expectations - Demonstrates mastery beyond required level",
"3": "Meets Expectations - Solid performance meeting all requirements",
"2": "Partially Meets - Shows potential but has development areas",
"1": "Does Not Meet - Significant gaps in required competencies"
},
"dimensions": [
{
"dimension": "system_architecture",
"weight": "high",
"scale": "1-4",
"description": "Assessment of system architecture competency"
},
{
"dimension": "technical_leadership",
"weight": "high",
"scale": "1-4",
"description": "Assessment of technical leadership competency"
},
{
"dimension": "mentoring",
"weight": "high",
"scale": "1-4",
"description": "Assessment of mentoring competency"
},
{
"dimension": "cross_team_collab",
"weight": "high",
"scale": "1-4",
"description": "Assessment of cross team collab competency"
},
{
"dimension": "technology_evaluation",
"weight": "medium",
"scale": "1-4",
"description": "Assessment of technology evaluation competency"
},
{
"dimension": "process_improvement",
"weight": "medium",
"scale": "1-4",
"description": "Assessment of process improvement competency"
},
{
"dimension": "hiring_contribution",
"weight": "medium",
"scale": "1-4",
"description": "Assessment of hiring contribution competency"
},
{
"dimension": "communication",
"weight": "high",
"scale": "1-4"
},
{
"dimension": "cultural_fit",
"weight": "medium",
"scale": "1-4"
},
{
"dimension": "learning_agility",
"weight": "medium",
"scale": "1-4"
}
],
"overall_recommendation": {
"options": [
"Strong Hire",
"Hire",
"No Hire",
"Strong No Hire"
],
"criteria": "Based on weighted average and minimum thresholds"
},
"calibration_notes": {
"required": true,
"min_length": 100,
"sections": [
"strengths",
"areas_for_development",
"specific_examples"
]
}
},
"interviewer_requirements": {
"round_1_technical_phone_screen": {
"required_skills": [
"technical_assessment",
"coding_evaluation"
],
"preferred_experience": [
"same_domain",
"senior_level"
],
"calibration_level": "standard",
"suggested_interviewers": [
"senior_engineer",
"tech_lead"
]
},
"round_2_coding_deep_dive": {
"required_skills": [
"advanced_technical",
"code_quality_assessment"
],
"preferred_experience": [
"senior_engineer",
"system_design"
],
"calibration_level": "high",
"suggested_interviewers": [
"senior_engineer",
"staff_engineer"
]
},
"round_3_system_design": {
"required_skills": [
"architecture_design",
"scalability_assessment"
],
"preferred_experience": [
"senior_architect",
"large_scale_systems"
],
"calibration_level": "high",
"suggested_interviewers": [
"senior_architect",
"staff_engineer"
]
},
"round_4_behavioral": {
"required_skills": [
"behavioral_interviewing",
"competency_assessment"
],
"preferred_experience": [
"hiring_manager",
"people_leadership"
],
"calibration_level": "standard",
"suggested_interviewers": [
"hiring_manager",
"people_manager"
]
},
"round_5_technical_leadership": {
"required_skills": [
"leadership_assessment",
"technical_mentoring"
],
"preferred_experience": [
"engineering_manager",
"tech_lead"
],
"calibration_level": "high",
"suggested_interviewers": [
"engineering_manager",
"senior_staff"
]
}
},
"competency_framework": {
"required": [
"system_architecture",
"technical_leadership",
"mentoring",
"cross_team_collab"
],
"preferred": [
"technology_evaluation",
"process_improvement",
"hiring_contribution"
],
"focus_areas": [
"technical_leadership",
"system_architecture",
"people_development"
]
},
"calibration_notes": {
"hiring_bar_notes": "Calibrated for senior level software engineer role",
"common_pitfalls": [
"Avoid comparing candidates to each other rather than to the role standard",
"Don't let one strong/weak area overshadow overall assessment",
"Ensure consistent application of evaluation criteria"
],
"calibration_checkpoints": [
"Review score distribution after every 5 candidates",
"Conduct monthly interviewer calibration sessions",
"Track correlation with 6-month performance reviews"
],
"escalation_criteria": [
"Any candidate receiving all 4s or all 1s",
"Significant disagreement between interviewers (>1.5 point spread)",
"Unusual circumstances or accommodations needed"
]
}
}
FILE:expected_outputs/senior_software_engineer_senior_interview_loop.txt
Interview Loop Design for Senior Software Engineer (Senior Level)
============================================================
Team: platform
Generated: 2026-02-16T13:27:37.925680
Total Duration: 300 minutes (5h 0m)
Total Rounds: 5
INTERVIEW ROUNDS
----------------------------------------
Round 1: Technical Phone Screen
Duration: 45 minutes
Format: Virtual
Objectives:
• Assess coding fundamentals
• Evaluate problem-solving approach
• Screen for basic technical competency
Focus Areas:
• Coding Fundamentals
• Problem Solving
• Technical Leadership
• System Architecture
• People Development
Round 2: Coding Deep Dive
Duration: 75 minutes
Format: In Person Or Virtual
Objectives:
• Evaluate coding skills in depth
• Assess code quality and testing
• Review debugging approach
Focus Areas:
• Technical Execution
• Code Quality
• Technical Leadership
• System Architecture
• People Development
Round 3: System Design
Duration: 75 minutes
Format: Collaborative Whiteboard
Objectives:
• Assess architectural thinking
• Evaluate scalability considerations
• Review trade-off analysis
Focus Areas:
• System Thinking
• Architectural Reasoning
• Technical Leadership
• System Architecture
• People Development
Round 4: Behavioral Interview
Duration: 45 minutes
Format: Conversational
Objectives:
• Assess cultural fit
• Evaluate past experiences
• Review leadership examples
Focus Areas:
• Cultural Fit
• Communication
• Teamwork
• Technical Leadership
• System Architecture
Round 5: Technical Leadership
Duration: 60 minutes
Format: Discussion Based
Objectives:
• Evaluate mentoring capability
• Assess technical decision making
• Review cross-team collaboration
Focus Areas:
• Leadership
• Mentoring
• Influence
• Technical Leadership
• System Architecture
SUGGESTED SCHEDULE
----------------------------------------
Schedule Type: Multi Day
Day 1:
Time: 09:00 - 12:45
09:00-09:45: Technical Phone Screen (45min)
10:00-11:15: Coding Deep Dive (75min)
11:30-12:45: System Design (75min)
Day 2:
Time: 09:00 - 11:00
09:00-09:45: Behavioral Interview (45min)
10:00-11:00: Technical Leadership (60min)
INTERVIEWER REQUIREMENTS
----------------------------------------
Technical Phone Screen:
Required Skills: technical_assessment, coding_evaluation
Suggested Interviewers: senior_engineer, tech_lead
Calibration Level: Standard
Coding Deep Dive:
Required Skills: advanced_technical, code_quality_assessment
Suggested Interviewers: senior_engineer, staff_engineer
Calibration Level: High
System Design:
Required Skills: architecture_design, scalability_assessment
Suggested Interviewers: senior_architect, staff_engineer
Calibration Level: High
Behavioral:
Required Skills: behavioral_interviewing, competency_assessment
Suggested Interviewers: hiring_manager, people_manager
Calibration Level: Standard
Technical Leadership:
Required Skills: leadership_assessment, technical_mentoring
Suggested Interviewers: engineering_manager, senior_staff
Calibration Level: High
SCORECARD TEMPLATE
----------------------------------------
Scoring Scale:
4: Exceeds Expectations - Demonstrates mastery beyond required level
3: Meets Expectations - Solid performance meeting all requirements
2: Partially Meets - Shows potential but has development areas
1: Does Not Meet - Significant gaps in required competencies
Evaluation Dimensions:
• System Architecture (Weight: high)
• Technical Leadership (Weight: high)
• Mentoring (Weight: high)
• Cross Team Collab (Weight: high)
• Technology Evaluation (Weight: medium)
• Process Improvement (Weight: medium)
• Hiring Contribution (Weight: medium)
• Communication (Weight: high)
• Cultural Fit (Weight: medium)
• Learning Agility (Weight: medium)
CALIBRATION NOTES
----------------------------------------
Hiring Bar: Calibrated for senior level software engineer role
Common Pitfalls:
• Avoid comparing candidates to each other rather than to the role standard
• Don't let one strong/weak area overshadow overall assessment
• Ensure consistent application of evaluation criteria
FILE:hiring_calibrator.py
#!/usr/bin/env python3
"""
Hiring Calibrator
Analyzes interview scores from multiple candidates and interviewers to detect bias,
calibration issues, and inconsistent rubric application. Generates calibration reports
with specific recommendations for interviewer coaching and process improvements.
Usage:
python hiring_calibrator.py --input interview_results.json --analysis-type comprehensive
python hiring_calibrator.py --input data.json --competencies technical,leadership --output report.json
python hiring_calibrator.py --input historical_data.json --trend-analysis --period quarterly
"""
import os
import sys
import json
import argparse
import statistics
from datetime import datetime, timedelta
from typing import Dict, List, Optional, Any, Tuple
from collections import defaultdict, Counter
import math
class HiringCalibrator:
"""Analyzes interview data for bias detection and calibration issues."""
def __init__(self):
self.bias_thresholds = self._init_bias_thresholds()
self.calibration_standards = self._init_calibration_standards()
self.demographic_categories = self._init_demographic_categories()
def _init_bias_thresholds(self) -> Dict[str, float]:
"""Initialize statistical thresholds for bias detection."""
return {
"score_variance_threshold": 1.5, # Standard deviations
"pass_rate_difference_threshold": 0.15, # 15% difference
"interviewer_consistency_threshold": 0.8, # Correlation coefficient
"demographic_parity_threshold": 0.10, # 10% difference
"score_inflation_threshold": 0.3, # 30% above historical average
"score_deflation_threshold": 0.3, # 30% below historical average
"minimum_sample_size": 5 # Minimum candidates per analysis
}
def _init_calibration_standards(self) -> Dict[str, Dict]:
"""Initialize expected calibration standards."""
return {
"score_distribution": {
"target_mean": 2.8, # Expected average score (1-4 scale)
"target_std": 0.9, # Expected standard deviation
"expected_distribution": {
"1": 0.10, # 10% score 1 (does not meet)
"2": 0.25, # 25% score 2 (partially meets)
"3": 0.45, # 45% score 3 (meets expectations)
"4": 0.20 # 20% score 4 (exceeds expectations)
}
},
"interviewer_agreement": {
"minimum_correlation": 0.70, # Minimum correlation between interviewers
"maximum_std_deviation": 0.8, # Maximum std dev in scores for same candidate
"agreement_threshold": 0.75 # % of time interviewers should agree within 1 point
},
"pass_rates": {
"junior_level": 0.25, # 25% pass rate for junior roles
"mid_level": 0.20, # 20% pass rate for mid roles
"senior_level": 0.15, # 15% pass rate for senior roles
"staff_level": 0.10, # 10% pass rate for staff+ roles
"leadership": 0.12 # 12% pass rate for leadership roles
}
}
def _init_demographic_categories(self) -> List[str]:
"""Initialize demographic categories to analyze for bias."""
return [
"gender", "ethnicity", "education_level", "previous_company_size",
"years_experience", "university_tier", "geographic_location"
]
def analyze_hiring_calibration(self, interview_data: List[Dict[str, Any]],
analysis_type: str = "comprehensive",
competencies: Optional[List[str]] = None,
trend_analysis: bool = False,
period: str = "monthly") -> Dict[str, Any]:
"""Perform comprehensive hiring calibration analysis."""
# Validate and preprocess data
processed_data = self._preprocess_interview_data(interview_data)
if len(processed_data) < self.bias_thresholds["minimum_sample_size"]:
return {
"error": "Insufficient data for analysis",
"minimum_required": self.bias_thresholds["minimum_sample_size"],
"actual_samples": len(processed_data)
}
# Perform different types of analysis based on request
analysis_results = {
"analysis_type": analysis_type,
"data_summary": self._generate_data_summary(processed_data),
"generated_at": datetime.now().isoformat()
}
if analysis_type in ["comprehensive", "bias"]:
analysis_results["bias_analysis"] = self._analyze_bias_patterns(processed_data, competencies)
if analysis_type in ["comprehensive", "calibration"]:
analysis_results["calibration_analysis"] = self._analyze_calibration_consistency(processed_data, competencies)
if analysis_type in ["comprehensive", "interviewer"]:
analysis_results["interviewer_analysis"] = self._analyze_interviewer_bias(processed_data)
if analysis_type in ["comprehensive", "scoring"]:
analysis_results["scoring_analysis"] = self._analyze_scoring_patterns(processed_data, competencies)
if trend_analysis:
analysis_results["trend_analysis"] = self._analyze_trends_over_time(processed_data, period)
# Generate recommendations
analysis_results["recommendations"] = self._generate_recommendations(analysis_results)
# Calculate overall calibration health score
analysis_results["calibration_health_score"] = self._calculate_health_score(analysis_results)
return analysis_results
def _preprocess_interview_data(self, raw_data: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Clean and validate interview data."""
processed_data = []
for record in raw_data:
if self._validate_interview_record(record):
processed_record = self._standardize_record(record)
processed_data.append(processed_record)
return processed_data
def _validate_interview_record(self, record: Dict[str, Any]) -> bool:
"""Validate that an interview record has required fields."""
required_fields = ["candidate_id", "interviewer_id", "scores", "overall_recommendation", "date"]
for field in required_fields:
if field not in record or record[field] is None:
return False
# Validate scores format
if not isinstance(record["scores"], dict):
return False
# Validate score values are numeric and in valid range (1-4)
for competency, score in record["scores"].items():
if not isinstance(score, (int, float)) or not (1 <= score <= 4):
return False
return True
def _standardize_record(self, record: Dict[str, Any]) -> Dict[str, Any]:
"""Standardize record format and add computed fields."""
standardized = record.copy()
# Calculate average score
scores = list(record["scores"].values())
standardized["average_score"] = statistics.mean(scores)
# Standardize recommendation to binary
recommendation = record["overall_recommendation"].lower()
standardized["hire_decision"] = recommendation in ["hire", "strong hire", "yes"]
# Parse date if string
if isinstance(record["date"], str):
try:
standardized["date"] = datetime.fromisoformat(record["date"].replace("Z", "+00:00"))
except ValueError:
standardized["date"] = datetime.now()
# Add demographic info if available
for category in self.demographic_categories:
if category not in standardized:
standardized[category] = "unknown"
# Add level normalization
role = record.get("role", "").lower()
if any(level in role for level in ["junior", "associate", "entry"]):
standardized["normalized_level"] = "junior"
elif any(level in role for level in ["senior", "sr"]):
standardized["normalized_level"] = "senior"
elif any(level in role for level in ["staff", "principal", "lead"]):
standardized["normalized_level"] = "staff"
else:
standardized["normalized_level"] = "mid"
return standardized
def _generate_data_summary(self, data: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Generate summary statistics for the dataset."""
if not data:
return {}
total_candidates = len(data)
unique_interviewers = len(set(record["interviewer_id"] for record in data))
# Score statistics
all_scores = []
all_average_scores = []
hire_decisions = []
for record in data:
all_scores.extend(record["scores"].values())
all_average_scores.append(record["average_score"])
hire_decisions.append(record["hire_decision"])
# Date range
dates = [record["date"] for record in data if record["date"]]
date_range = {
"start_date": min(dates).isoformat() if dates else None,
"end_date": max(dates).isoformat() if dates else None,
"total_days": (max(dates) - min(dates)).days if len(dates) > 1 else 0
}
# Role distribution
roles = [record.get("role", "unknown") for record in data]
role_distribution = dict(Counter(roles))
return {
"total_candidates": total_candidates,
"unique_interviewers": unique_interviewers,
"candidates_per_interviewer": round(total_candidates / unique_interviewers, 2),
"date_range": date_range,
"score_statistics": {
"mean_individual_scores": round(statistics.mean(all_scores), 2),
"std_individual_scores": round(statistics.stdev(all_scores) if len(all_scores) > 1 else 0, 2),
"mean_average_scores": round(statistics.mean(all_average_scores), 2),
"std_average_scores": round(statistics.stdev(all_average_scores) if len(all_average_scores) > 1 else 0, 2)
},
"hire_rate": round(sum(hire_decisions) / len(hire_decisions), 3),
"role_distribution": role_distribution
}
def _analyze_bias_patterns(self, data: List[Dict[str, Any]],
target_competencies: Optional[List[str]]) -> Dict[str, Any]:
"""Analyze potential bias patterns in interview decisions."""
bias_analysis = {
"demographic_bias": {},
"interviewer_bias": {},
"competency_bias": {},
"overall_bias_score": 0
}
# Analyze demographic bias
for demographic in self.demographic_categories:
if all(record.get(demographic) == "unknown" for record in data):
continue
demographic_analysis = self._analyze_demographic_bias(data, demographic)
if demographic_analysis["bias_detected"]:
bias_analysis["demographic_bias"][demographic] = demographic_analysis
# Analyze interviewer bias
bias_analysis["interviewer_bias"] = self._analyze_interviewer_bias(data)
# Analyze competency bias if specified
if target_competencies:
bias_analysis["competency_bias"] = self._analyze_competency_bias(data, target_competencies)
# Calculate overall bias score
bias_analysis["overall_bias_score"] = self._calculate_bias_score(bias_analysis)
return bias_analysis
def _analyze_demographic_bias(self, data: List[Dict[str, Any]],
demographic: str) -> Dict[str, Any]:
"""Analyze bias for a specific demographic category."""
# Group data by demographic values
demographic_groups = defaultdict(list)
for record in data:
demo_value = record.get(demographic, "unknown")
if demo_value != "unknown":
demographic_groups[demo_value].append(record)
if len(demographic_groups) < 2:
return {"bias_detected": False, "reason": "insufficient_groups"}
# Calculate statistics for each group
group_stats = {}
for group, records in demographic_groups.items():
if len(records) >= self.bias_thresholds["minimum_sample_size"]:
scores = [r["average_score"] for r in records]
hire_rate = sum(r["hire_decision"] for r in records) / len(records)
group_stats[group] = {
"count": len(records),
"mean_score": statistics.mean(scores),
"hire_rate": hire_rate,
"std_score": statistics.stdev(scores) if len(scores) > 1 else 0
}
if len(group_stats) < 2:
return {"bias_detected": False, "reason": "insufficient_sample_sizes"}
# Detect statistical differences
bias_detected = False
bias_details = {}
# Check for significant differences in hire rates
hire_rates = [stats["hire_rate"] for stats in group_stats.values()]
max_hire_rate_diff = max(hire_rates) - min(hire_rates)
if max_hire_rate_diff > self.bias_thresholds["demographic_parity_threshold"]:
bias_detected = True
bias_details["hire_rate_disparity"] = {
"max_difference": round(max_hire_rate_diff, 3),
"threshold": self.bias_thresholds["demographic_parity_threshold"],
"group_stats": group_stats
}
# Check for significant differences in scoring
mean_scores = [stats["mean_score"] for stats in group_stats.values()]
max_score_diff = max(mean_scores) - min(mean_scores)
if max_score_diff > 0.5: # Half point difference threshold
bias_detected = True
bias_details["scoring_disparity"] = {
"max_difference": round(max_score_diff, 3),
"group_stats": group_stats
}
return {
"bias_detected": bias_detected,
"demographic": demographic,
"group_statistics": group_stats,
"bias_details": bias_details,
"recommendation": self._generate_demographic_bias_recommendation(demographic, bias_details) if bias_detected else None
}
def _analyze_interviewer_bias(self, data: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Analyze bias patterns across different interviewers."""
interviewer_stats = defaultdict(list)
# Group by interviewer
for record in data:
interviewer_id = record["interviewer_id"]
interviewer_stats[interviewer_id].append(record)
# Calculate statistics per interviewer
interviewer_analysis = {}
for interviewer_id, records in interviewer_stats.items():
if len(records) >= self.bias_thresholds["minimum_sample_size"]:
scores = [r["average_score"] for r in records]
hire_rate = sum(r["hire_decision"] for r in records) / len(records)
interviewer_analysis[interviewer_id] = {
"total_interviews": len(records),
"mean_score": statistics.mean(scores),
"std_score": statistics.stdev(scores) if len(scores) > 1 else 0,
"hire_rate": hire_rate,
"score_inflation": self._detect_score_inflation(scores),
"consistency_score": self._calculate_interviewer_consistency(records)
}
# Identify outlier interviewers
if len(interviewer_analysis) > 1:
overall_mean_score = statistics.mean([stats["mean_score"] for stats in interviewer_analysis.values()])
overall_hire_rate = statistics.mean([stats["hire_rate"] for stats in interviewer_analysis.values()])
outlier_interviewers = {}
for interviewer_id, stats in interviewer_analysis.items():
issues = []
# Check for score inflation/deflation
if stats["mean_score"] > overall_mean_score * (1 + self.bias_thresholds["score_inflation_threshold"]):
issues.append("score_inflation")
elif stats["mean_score"] < overall_mean_score * (1 - self.bias_thresholds["score_deflation_threshold"]):
issues.append("score_deflation")
# Check for hire rate deviation
hire_rate_diff = abs(stats["hire_rate"] - overall_hire_rate)
if hire_rate_diff > self.bias_thresholds["pass_rate_difference_threshold"]:
issues.append("hire_rate_deviation")
# Check for low consistency
if stats["consistency_score"] < self.bias_thresholds["interviewer_consistency_threshold"]:
issues.append("low_consistency")
if issues:
outlier_interviewers[interviewer_id] = {
"issues": issues,
"statistics": stats,
"severity": len(issues) # More issues = higher severity
}
return {
"interviewer_statistics": interviewer_analysis,
"outlier_interviewers": outlier_interviewers if len(interviewer_analysis) > 1 else {},
"overall_consistency": self._calculate_overall_interviewer_consistency(data),
"recommendations": self._generate_interviewer_recommendations(outlier_interviewers if len(interviewer_analysis) > 1 else {})
}
def _analyze_competency_bias(self, data: List[Dict[str, Any]],
competencies: List[str]) -> Dict[str, Any]:
"""Analyze bias patterns within specific competencies."""
competency_analysis = {}
for competency in competencies:
# Extract scores for this competency
competency_scores = []
for record in data:
if competency in record["scores"]:
competency_scores.append({
"score": record["scores"][competency],
"interviewer": record["interviewer_id"],
"candidate": record["candidate_id"],
"overall_decision": record["hire_decision"]
})
if len(competency_scores) < self.bias_thresholds["minimum_sample_size"]:
continue
# Analyze scoring patterns
scores = [item["score"] for item in competency_scores]
score_variance = statistics.variance(scores) if len(scores) > 1 else 0
# Analyze by interviewer
interviewer_competency_scores = defaultdict(list)
for item in competency_scores:
interviewer_competency_scores[item["interviewer"]].append(item["score"])
interviewer_variations = {}
if len(interviewer_competency_scores) > 1:
interviewer_means = {interviewer: statistics.mean(scores)
for interviewer, scores in interviewer_competency_scores.items()
if len(scores) >= 3}
if len(interviewer_means) > 1:
mean_of_means = statistics.mean(interviewer_means.values())
for interviewer, mean_score in interviewer_means.items():
deviation = abs(mean_score - mean_of_means)
if deviation > 0.5: # More than half point deviation
interviewer_variations[interviewer] = {
"mean_score": round(mean_score, 2),
"deviation_from_average": round(deviation, 2),
"sample_size": len(interviewer_competency_scores[interviewer])
}
competency_analysis[competency] = {
"total_scores": len(competency_scores),
"mean_score": round(statistics.mean(scores), 2),
"score_variance": round(score_variance, 2),
"interviewer_variations": interviewer_variations,
"bias_detected": len(interviewer_variations) > 0
}
return competency_analysis
def _analyze_calibration_consistency(self, data: List[Dict[str, Any]],
target_competencies: Optional[List[str]]) -> Dict[str, Any]:
"""Analyze calibration consistency across interviews."""
# Group candidates by those interviewed by multiple people
candidate_interviewers = defaultdict(list)
for record in data:
candidate_interviewers[record["candidate_id"]].append(record)
multi_interviewer_candidates = {
candidate: records for candidate, records in candidate_interviewers.items()
if len(records) > 1
}
if not multi_interviewer_candidates:
return {
"error": "No candidates with multiple interviewers found",
"single_interviewer_analysis": self._analyze_single_interviewer_consistency(data)
}
# Calculate agreement statistics
agreement_stats = []
score_correlations = []
for candidate, records in multi_interviewer_candidates.items():
candidate_scores = []
interviewer_pairs = []
for record in records:
avg_score = record["average_score"]
candidate_scores.append(avg_score)
interviewer_pairs.append(record["interviewer_id"])
if len(candidate_scores) > 1:
# Calculate standard deviation of scores for this candidate
score_std = statistics.stdev(candidate_scores)
agreement_stats.append(score_std)
# Check if all interviewers agree within 1 point
score_range = max(candidate_scores) - min(candidate_scores)
agreement_within_one = score_range <= 1.0
score_correlations.append({
"candidate": candidate,
"scores": candidate_scores,
"interviewers": interviewer_pairs,
"score_std": score_std,
"score_range": score_range,
"agreement_within_one": agreement_within_one
})
# Calculate overall calibration metrics
mean_score_std = statistics.mean(agreement_stats) if agreement_stats else 0
agreement_rate = sum(1 for corr in score_correlations if corr["agreement_within_one"]) / len(score_correlations) if score_correlations else 0
calibration_quality = "good"
if mean_score_std > self.calibration_standards["interviewer_agreement"]["maximum_std_deviation"]:
calibration_quality = "poor"
elif agreement_rate < self.calibration_standards["interviewer_agreement"]["agreement_threshold"]:
calibration_quality = "fair"
return {
"multi_interviewer_candidates": len(multi_interviewer_candidates),
"mean_score_standard_deviation": round(mean_score_std, 3),
"agreement_within_one_point_rate": round(agreement_rate, 3),
"calibration_quality": calibration_quality,
"candidate_agreement_details": score_correlations,
"target_standards": self.calibration_standards["interviewer_agreement"],
"recommendations": self._generate_calibration_recommendations(mean_score_std, agreement_rate)
}
def _analyze_scoring_patterns(self, data: List[Dict[str, Any]],
target_competencies: Optional[List[str]]) -> Dict[str, Any]:
"""Analyze overall scoring patterns and distributions."""
# Overall score distribution
all_individual_scores = []
all_average_scores = []
score_distribution = defaultdict(int)
for record in data:
avg_score = record["average_score"]
all_average_scores.append(avg_score)
for competency, score in record["scores"].items():
if not target_competencies or competency in target_competencies:
all_individual_scores.append(score)
score_distribution[str(int(score))] += 1
# Calculate distribution percentages
total_scores = sum(score_distribution.values())
score_percentages = {score: count/total_scores for score, count in score_distribution.items()}
# Compare against expected distribution
expected_dist = self.calibration_standards["score_distribution"]["expected_distribution"]
distribution_analysis = {}
for score in ["1", "2", "3", "4"]:
expected_pct = expected_dist.get(score, 0)
actual_pct = score_percentages.get(score, 0)
difference = actual_pct - expected_pct
distribution_analysis[score] = {
"expected_percentage": expected_pct,
"actual_percentage": round(actual_pct, 3),
"difference": round(difference, 3),
"significant_deviation": abs(difference) > 0.05 # 5% threshold
}
# Calculate scoring statistics
mean_score = statistics.mean(all_individual_scores) if all_individual_scores else 0
std_score = statistics.stdev(all_individual_scores) if len(all_individual_scores) > 1 else 0
target_mean = self.calibration_standards["score_distribution"]["target_mean"]
target_std = self.calibration_standards["score_distribution"]["target_std"]
# Analyze pass rates by level
level_pass_rates = {}
level_groups = defaultdict(list)
for record in data:
level = record.get("normalized_level", "unknown")
level_groups[level].append(record["hire_decision"])
for level, decisions in level_groups.items():
if len(decisions) >= self.bias_thresholds["minimum_sample_size"]:
pass_rate = sum(decisions) / len(decisions)
expected_rate = self.calibration_standards["pass_rates"].get(f"{level}_level", 0.15)
level_pass_rates[level] = {
"actual_pass_rate": round(pass_rate, 3),
"expected_pass_rate": expected_rate,
"difference": round(pass_rate - expected_rate, 3),
"sample_size": len(decisions)
}
return {
"score_statistics": {
"mean_score": round(mean_score, 2),
"std_score": round(std_score, 2),
"target_mean": target_mean,
"target_std": target_std,
"mean_deviation": round(abs(mean_score - target_mean), 2),
"std_deviation": round(abs(std_score - target_std), 2)
},
"score_distribution": distribution_analysis,
"level_pass_rates": level_pass_rates,
"overall_assessment": self._assess_scoring_health(distribution_analysis, mean_score, target_mean)
}
def _analyze_trends_over_time(self, data: List[Dict[str, Any]], period: str) -> Dict[str, Any]:
"""Analyze trends in hiring patterns over time."""
# Sort data by date
dated_data = [record for record in data if record.get("date")]
dated_data.sort(key=lambda x: x["date"])
if len(dated_data) < 10: # Need minimum data for trend analysis
return {"error": "Insufficient data for trend analysis", "minimum_required": 10}
# Group by time period
period_groups = defaultdict(list)
for record in dated_data:
date = record["date"]
if period == "weekly":
period_key = date.strftime("%Y-W%U")
elif period == "monthly":
period_key = date.strftime("%Y-%m")
elif period == "quarterly":
quarter = (date.month - 1) // 3 + 1
period_key = f"{date.year}-Q{quarter}"
else: # daily
period_key = date.strftime("%Y-%m-%d")
period_groups[period_key].append(record)
# Calculate metrics for each period
period_metrics = {}
for period_key, records in period_groups.items():
if len(records) >= 3: # Minimum for meaningful metrics
scores = [r["average_score"] for r in records]
hire_rate = sum(r["hire_decision"] for r in records) / len(records)
period_metrics[period_key] = {
"count": len(records),
"mean_score": statistics.mean(scores),
"hire_rate": hire_rate,
"std_score": statistics.stdev(scores) if len(scores) > 1 else 0
}
if len(period_metrics) < 3:
return {"error": "Insufficient periods for trend analysis"}
# Analyze trends
sorted_periods = sorted(period_metrics.keys())
mean_scores = [period_metrics[p]["mean_score"] for p in sorted_periods]
hire_rates = [period_metrics[p]["hire_rate"] for p in sorted_periods]
# Simple linear trend calculation
score_trend = self._calculate_linear_trend(mean_scores)
hire_rate_trend = self._calculate_linear_trend(hire_rates)
return {
"period": period,
"total_periods": len(period_metrics),
"period_metrics": period_metrics,
"trends": {
"score_trend": {
"direction": "increasing" if score_trend > 0.01 else "decreasing" if score_trend < -0.01 else "stable",
"slope": round(score_trend, 4),
"significance": "significant" if abs(score_trend) > 0.05 else "minor"
},
"hire_rate_trend": {
"direction": "increasing" if hire_rate_trend > 0.005 else "decreasing" if hire_rate_trend < -0.005 else "stable",
"slope": round(hire_rate_trend, 4),
"significance": "significant" if abs(hire_rate_trend) > 0.02 else "minor"
}
},
"insights": self._generate_trend_insights(score_trend, hire_rate_trend, period_metrics)
}
def _calculate_linear_trend(self, values: List[float]) -> float:
"""Calculate simple linear trend slope."""
if len(values) < 2:
return 0
n = len(values)
x = list(range(n))
# Calculate slope using least squares
x_mean = statistics.mean(x)
y_mean = statistics.mean(values)
numerator = sum((x[i] - x_mean) * (values[i] - y_mean) for i in range(n))
denominator = sum((x[i] - x_mean) ** 2 for i in range(n))
return numerator / denominator if denominator != 0 else 0
def _detect_score_inflation(self, scores: List[float]) -> Dict[str, Any]:
"""Detect if an interviewer shows score inflation patterns."""
if len(scores) < 5:
return {"insufficient_data": True}
mean_score = statistics.mean(scores)
std_score = statistics.stdev(scores)
# Check against expected mean (2.8)
expected_mean = self.calibration_standards["score_distribution"]["target_mean"]
deviation = mean_score - expected_mean
# High scores with low variance might indicate inflation
high_scores_low_variance = mean_score > 3.2 and std_score < 0.5
# Check distribution - too many 4s might indicate inflation
score_counts = Counter([int(score) for score in scores])
four_count_ratio = score_counts.get(4, 0) / len(scores)
return {
"mean_score": round(mean_score, 2),
"expected_mean": expected_mean,
"deviation": round(deviation, 2),
"high_scores_low_variance": high_scores_low_variance,
"four_count_ratio": round(four_count_ratio, 2),
"inflation_detected": deviation > 0.3 or high_scores_low_variance or four_count_ratio > 0.4
}
def _calculate_interviewer_consistency(self, records: List[Dict[str, Any]]) -> float:
"""Calculate consistency score for an interviewer."""
if len(records) < 3:
return 0.5 # Neutral score for insufficient data
# Look at variance in scoring
avg_scores = [r["average_score"] for r in records]
score_variance = statistics.variance(avg_scores)
# Look at decision consistency relative to scores
decisions = [r["hire_decision"] for r in records]
scores_of_hires = [r["average_score"] for r in records if r["hire_decision"]]
scores_of_no_hires = [r["average_score"] for r in records if not r["hire_decision"]]
# Good consistency means hires have higher average scores
decision_consistency = 0.5
if scores_of_hires and scores_of_no_hires:
hire_mean = statistics.mean(scores_of_hires)
no_hire_mean = statistics.mean(scores_of_no_hires)
score_gap = hire_mean - no_hire_mean
decision_consistency = min(1.0, max(0.0, score_gap / 2.0)) # Normalize to 0-1
# Combine metrics (lower variance = higher consistency)
variance_consistency = max(0.0, 1.0 - (score_variance / 2.0))
return (decision_consistency + variance_consistency) / 2
def _calculate_overall_interviewer_consistency(self, data: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Calculate overall consistency across all interviewers."""
interviewer_consistency_scores = []
interviewer_records = defaultdict(list)
for record in data:
interviewer_records[record["interviewer_id"]].append(record)
for interviewer_id, records in interviewer_records.items():
if len(records) >= 3:
consistency = self._calculate_interviewer_consistency(records)
interviewer_consistency_scores.append(consistency)
if not interviewer_consistency_scores:
return {"error": "Insufficient data per interviewer for consistency analysis"}
return {
"mean_consistency": round(statistics.mean(interviewer_consistency_scores), 3),
"std_consistency": round(statistics.stdev(interviewer_consistency_scores) if len(interviewer_consistency_scores) > 1 else 0, 3),
"min_consistency": round(min(interviewer_consistency_scores), 3),
"max_consistency": round(max(interviewer_consistency_scores), 3),
"interviewers_analyzed": len(interviewer_consistency_scores),
"target_threshold": self.bias_thresholds["interviewer_consistency_threshold"]
}
def _calculate_bias_score(self, bias_analysis: Dict[str, Any]) -> float:
"""Calculate overall bias score (0-1, where 1 is most biased)."""
bias_factors = []
# Demographic bias factors
demographic_bias = bias_analysis.get("demographic_bias", {})
for demo, analysis in demographic_bias.items():
if analysis.get("bias_detected"):
bias_factors.append(0.3) # Each demographic bias adds 0.3
# Interviewer bias factors
interviewer_bias = bias_analysis.get("interviewer_bias", {})
outlier_interviewers = interviewer_bias.get("outlier_interviewers", {})
if outlier_interviewers:
# Scale by severity and number of outliers
total_severity = sum(info["severity"] for info in outlier_interviewers.values())
bias_factors.append(min(0.5, total_severity * 0.1))
# Competency bias factors
competency_bias = bias_analysis.get("competency_bias", {})
for comp, analysis in competency_bias.items():
if analysis.get("bias_detected"):
bias_factors.append(0.2) # Each competency bias adds 0.2
return min(1.0, sum(bias_factors))
def _calculate_health_score(self, analysis: Dict[str, Any]) -> Dict[str, Any]:
"""Calculate overall calibration health score."""
health_factors = []
# Bias score (lower is better)
bias_analysis = analysis.get("bias_analysis", {})
bias_score = bias_analysis.get("overall_bias_score", 0)
bias_health = max(0, 1 - bias_score)
health_factors.append(("bias", bias_health, 0.3))
# Calibration consistency
calibration_analysis = analysis.get("calibration_analysis", {})
if "calibration_quality" in calibration_analysis:
quality_map = {"good": 1.0, "fair": 0.7, "poor": 0.3}
calibration_health = quality_map.get(calibration_analysis["calibration_quality"], 0.5)
health_factors.append(("calibration", calibration_health, 0.25))
# Interviewer consistency
interviewer_analysis = analysis.get("interviewer_analysis", {})
overall_consistency = interviewer_analysis.get("overall_consistency", {})
if "mean_consistency" in overall_consistency:
consistency_health = overall_consistency["mean_consistency"]
health_factors.append(("interviewer_consistency", consistency_health, 0.25))
# Scoring patterns health
scoring_analysis = analysis.get("scoring_analysis", {})
if "overall_assessment" in scoring_analysis:
assessment_map = {"healthy": 1.0, "concerning": 0.6, "poor": 0.2}
scoring_health = assessment_map.get(scoring_analysis["overall_assessment"], 0.5)
health_factors.append(("scoring_patterns", scoring_health, 0.2))
# Calculate weighted average
if health_factors:
weighted_sum = sum(score * weight for _, score, weight in health_factors)
total_weight = sum(weight for _, _, weight in health_factors)
overall_score = weighted_sum / total_weight
else:
overall_score = 0.5 # Neutral if no data
# Categorize health
if overall_score >= 0.8:
health_category = "excellent"
elif overall_score >= 0.7:
health_category = "good"
elif overall_score >= 0.5:
health_category = "fair"
else:
health_category = "poor"
return {
"overall_score": round(overall_score, 3),
"health_category": health_category,
"component_scores": {name: round(score, 3) for name, score, _ in health_factors},
"improvement_priority": self._identify_improvement_priorities(health_factors)
}
def _identify_improvement_priorities(self, health_factors: List[Tuple[str, float, float]]) -> List[str]:
"""Identify areas that need the most improvement."""
priorities = []
for name, score, weight in health_factors:
impact = (1 - score) * weight # Low scores with high weights = high priority
if impact > 0.15: # Significant impact threshold
priorities.append(name)
# Sort by impact (highest first)
priorities.sort(key=lambda name: next((1 - score) * weight for n, score, weight in health_factors if n == name), reverse=True)
return priorities
def _generate_recommendations(self, analysis: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate actionable recommendations based on analysis results."""
recommendations = []
# Bias-related recommendations
bias_analysis = analysis.get("bias_analysis", {})
# Demographic bias recommendations
for demo, demo_analysis in bias_analysis.get("demographic_bias", {}).items():
if demo_analysis.get("bias_detected"):
recommendations.append({
"priority": "high",
"category": "bias_mitigation",
"title": f"Address {demo.replace('_', ' ').title()} Bias",
"description": demo_analysis.get("recommendation", f"Implement bias mitigation strategies for {demo}"),
"actions": [
"Conduct unconscious bias training focused on this demographic",
"Review and standardize interview questions",
"Implement diverse interview panels",
"Monitor hiring metrics by demographic group"
]
})
# Interviewer-specific recommendations
interviewer_analysis = bias_analysis.get("interviewer_bias", {})
outlier_interviewers = interviewer_analysis.get("outlier_interviewers", {})
for interviewer_id, outlier_info in outlier_interviewers.items():
issues = outlier_info["issues"]
priority = "high" if outlier_info["severity"] >= 3 else "medium"
actions = []
if "score_inflation" in issues:
actions.extend([
"Provide calibration training on scoring standards",
"Shadow experienced interviewers for recalibration",
"Review examples of each score level"
])
if "score_deflation" in issues:
actions.extend([
"Review expectations for role level",
"Calibrate against recent successful hires",
"Discuss evaluation criteria with hiring manager"
])
if "hire_rate_deviation" in issues:
actions.extend([
"Review hiring bar standards",
"Participate in calibration sessions",
"Compare decision criteria with team"
])
if "low_consistency" in issues:
actions.extend([
"Practice structured interviewing techniques",
"Use standardized scorecards",
"Document specific examples for each score"
])
recommendations.append({
"priority": priority,
"category": "interviewer_coaching",
"title": f"Coach Interviewer {interviewer_id}",
"description": f"Address issues: {', '.join(issues)}",
"actions": list(set(actions)) # Remove duplicates
})
# Calibration recommendations
calibration_analysis = analysis.get("calibration_analysis", {})
if calibration_analysis.get("calibration_quality") in ["fair", "poor"]:
recommendations.append({
"priority": "high",
"category": "calibration_improvement",
"title": "Improve Interview Calibration",
"description": f"Current calibration quality: {calibration_analysis.get('calibration_quality')}",
"actions": [
"Conduct monthly calibration sessions",
"Create shared examples of good/poor answers",
"Implement mandatory interviewer shadowing",
"Standardize scoring rubrics across all interviewers",
"Review and align on role expectations"
]
})
# Scoring pattern recommendations
scoring_analysis = analysis.get("scoring_analysis", {})
if scoring_analysis.get("overall_assessment") in ["concerning", "poor"]:
recommendations.append({
"priority": "medium",
"category": "scoring_standards",
"title": "Adjust Scoring Standards",
"description": "Scoring patterns deviate significantly from expected distribution",
"actions": [
"Review and communicate target score distributions",
"Provide examples for each score level",
"Monitor pass rates by role level",
"Adjust hiring bar if consistently too high/low"
]
})
# Health score recommendations
health_score = analysis.get("calibration_health_score", {})
priorities = health_score.get("improvement_priority", [])
if "bias" in priorities:
recommendations.append({
"priority": "critical",
"category": "bias_mitigation",
"title": "Implement Comprehensive Bias Mitigation",
"description": "Multiple bias indicators detected across the hiring process",
"actions": [
"Mandatory unconscious bias training for all interviewers",
"Implement structured interview protocols",
"Diversify interview panels",
"Regular bias audits and monitoring",
"Create accountability metrics for fair hiring"
]
})
# Sort by priority
priority_order = {"critical": 0, "high": 1, "medium": 2, "low": 3}
recommendations.sort(key=lambda x: priority_order.get(x["priority"], 3))
return recommendations
def _generate_demographic_bias_recommendation(self, demographic: str, bias_details: Dict[str, Any]) -> str:
"""Generate specific recommendation for demographic bias."""
if "hire_rate_disparity" in bias_details:
return f"Significant hire rate disparity detected for {demographic}. Implement structured interviews and diverse panels."
elif "scoring_disparity" in bias_details:
return f"Scoring disparity detected for {demographic}. Provide unconscious bias training and standardize evaluation criteria."
else:
return f"Potential bias detected for {demographic}. Monitor closely and implement bias mitigation strategies."
def _generate_interviewer_recommendations(self, outlier_interviewers: Dict[str, Any]) -> List[str]:
"""Generate recommendations for interviewer issues."""
if not outlier_interviewers:
return ["All interviewers performing within expected ranges"]
recommendations = []
for interviewer, info in outlier_interviewers.items():
issues = info["issues"]
if len(issues) >= 2:
recommendations.append(f"Interviewer {interviewer}: Requires comprehensive recalibration - multiple issues detected")
elif "score_inflation" in issues:
recommendations.append(f"Interviewer {interviewer}: Provide calibration training on scoring standards")
elif "hire_rate_deviation" in issues:
recommendations.append(f"Interviewer {interviewer}: Review hiring bar standards and decision criteria")
return recommendations
def _generate_calibration_recommendations(self, mean_std: float, agreement_rate: float) -> List[str]:
"""Generate calibration improvement recommendations."""
recommendations = []
if mean_std > self.calibration_standards["interviewer_agreement"]["maximum_std_deviation"]:
recommendations.append("High score variance detected - implement regular calibration sessions")
recommendations.append("Create shared examples of scoring standards for each competency")
if agreement_rate < self.calibration_standards["interviewer_agreement"]["agreement_threshold"]:
recommendations.append("Low interviewer agreement rate - standardize interview questions and evaluation criteria")
recommendations.append("Implement mandatory interviewer training on consistent evaluation")
if not recommendations:
recommendations.append("Calibration appears healthy - maintain current practices")
return recommendations
def _assess_scoring_health(self, distribution: Dict[str, Any], mean_score: float, target_mean: float) -> str:
"""Assess overall health of scoring patterns."""
issues = 0
# Check distribution deviations
for score_level, analysis in distribution.items():
if analysis["significant_deviation"]:
issues += 1
# Check mean deviation
if abs(mean_score - target_mean) > 0.3:
issues += 1
if issues == 0:
return "healthy"
elif issues <= 2:
return "concerning"
else:
return "poor"
def _generate_trend_insights(self, score_trend: float, hire_rate_trend: float, period_metrics: Dict[str, Any]) -> List[str]:
"""Generate insights from trend analysis."""
insights = []
if abs(score_trend) > 0.05:
direction = "increasing" if score_trend > 0 else "decreasing"
insights.append(f"Significant {direction} trend in average scores over time")
if score_trend > 0:
insights.append("May indicate score inflation or improving candidate quality")
else:
insights.append("May indicate stricter evaluation or declining candidate quality")
if abs(hire_rate_trend) > 0.02:
direction = "increasing" if hire_rate_trend > 0 else "decreasing"
insights.append(f"Significant {direction} trend in hire rates over time")
if hire_rate_trend > 0:
insights.append("Consider if hiring bar has lowered or candidate pool improved")
else:
insights.append("Consider if hiring bar has raised or candidate pool declined")
# Check for consistency
period_values = list(period_metrics.values())
hire_rates = [p["hire_rate"] for p in period_values]
hire_rate_variance = statistics.variance(hire_rates) if len(hire_rates) > 1 else 0
if hire_rate_variance > 0.01: # High variance in hire rates
insights.append("High variance in hire rates across periods - consider process standardization")
if not insights:
insights.append("Hiring patterns appear stable over time")
return insights
def _analyze_single_interviewer_consistency(self, data: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Analyze consistency for single-interviewer candidates."""
# Look at consistency within individual interviewers
interviewer_scores = defaultdict(list)
for record in data:
interviewer_scores[record["interviewer_id"]].extend(record["scores"].values())
consistency_analysis = {}
for interviewer, scores in interviewer_scores.items():
if len(scores) >= 10: # Need sufficient data
consistency_analysis[interviewer] = {
"mean_score": round(statistics.mean(scores), 2),
"std_score": round(statistics.stdev(scores), 2),
"coefficient_of_variation": round(statistics.stdev(scores) / statistics.mean(scores), 2),
"total_scores": len(scores)
}
return consistency_analysis
def format_human_readable(calibration_report: Dict[str, Any]) -> str:
"""Format calibration report in human-readable format."""
output = []
# Header
output.append("HIRING CALIBRATION ANALYSIS REPORT")
output.append("=" * 60)
output.append(f"Analysis Type: {calibration_report.get('analysis_type', 'N/A').title()}")
output.append(f"Generated: {calibration_report.get('generated_at', 'N/A')}")
if "error" in calibration_report:
output.append(f"\nError: {calibration_report['error']}")
return "\n".join(output)
# Data Summary
data_summary = calibration_report.get("data_summary", {})
if data_summary:
output.append(f"\nDATA SUMMARY")
output.append("-" * 30)
output.append(f"Total Candidates: {data_summary.get('total_candidates', 0)}")
output.append(f"Unique Interviewers: {data_summary.get('unique_interviewers', 0)}")
output.append(f"Overall Hire Rate: {data_summary.get('hire_rate', 0):.1%}")
score_stats = data_summary.get("score_statistics", {})
output.append(f"Average Score: {score_stats.get('mean_average_scores', 0):.2f}")
output.append(f"Score Std Dev: {score_stats.get('std_average_scores', 0):.2f}")
# Health Score
health_score = calibration_report.get("calibration_health_score", {})
if health_score:
output.append(f"\nCALIBRATION HEALTH SCORE")
output.append("-" * 30)
output.append(f"Overall Score: {health_score.get('overall_score', 0):.3f}")
output.append(f"Health Category: {health_score.get('health_category', 'Unknown').title()}")
if health_score.get("improvement_priority"):
output.append(f"Priority Areas: {', '.join(health_score['improvement_priority'])}")
# Bias Analysis
bias_analysis = calibration_report.get("bias_analysis", {})
if bias_analysis:
output.append(f"\nBIAS ANALYSIS")
output.append("-" * 30)
output.append(f"Overall Bias Score: {bias_analysis.get('overall_bias_score', 0):.3f}")
# Demographic bias
demographic_bias = bias_analysis.get("demographic_bias", {})
if demographic_bias:
output.append(f"\nDemographic Bias Issues:")
for demo, analysis in demographic_bias.items():
output.append(f" • {demo.replace('_', ' ').title()}: {analysis.get('bias_details', {}).keys()}")
# Interviewer bias
interviewer_bias = bias_analysis.get("interviewer_bias", {})
outlier_interviewers = interviewer_bias.get("outlier_interviewers", {})
if outlier_interviewers:
output.append(f"\nOutlier Interviewers:")
for interviewer, info in outlier_interviewers.items():
issues = ", ".join(info["issues"])
output.append(f" • {interviewer}: {issues}")
# Calibration Analysis
calibration_analysis = calibration_report.get("calibration_analysis", {})
if calibration_analysis and "error" not in calibration_analysis:
output.append(f"\nCALIBRATION CONSISTENCY")
output.append("-" * 30)
output.append(f"Quality: {calibration_analysis.get('calibration_quality', 'Unknown').title()}")
output.append(f"Agreement Rate: {calibration_analysis.get('agreement_within_one_point_rate', 0):.1%}")
output.append(f"Score Std Dev: {calibration_analysis.get('mean_score_standard_deviation', 0):.3f}")
# Scoring Analysis
scoring_analysis = calibration_report.get("scoring_analysis", {})
if scoring_analysis:
output.append(f"\nSCORING PATTERNS")
output.append("-" * 30)
output.append(f"Overall Assessment: {scoring_analysis.get('overall_assessment', 'Unknown').title()}")
score_stats = scoring_analysis.get("score_statistics", {})
output.append(f"Mean Score: {score_stats.get('mean_score', 0):.2f} (Target: {score_stats.get('target_mean', 0):.2f})")
# Distribution analysis
distribution = scoring_analysis.get("score_distribution", {})
if distribution:
output.append(f"\nScore Distribution vs Expected:")
for score in ["1", "2", "3", "4"]:
if score in distribution:
actual = distribution[score]["actual_percentage"]
expected = distribution[score]["expected_percentage"]
output.append(f" Score {score}: {actual:.1%} (Expected: {expected:.1%})")
# Top Recommendations
recommendations = calibration_report.get("recommendations", [])
if recommendations:
output.append(f"\nTOP RECOMMENDATIONS")
output.append("-" * 30)
for i, rec in enumerate(recommendations[:5], 1): # Show top 5
output.append(f"{i}. {rec['title']} ({rec['priority'].title()} Priority)")
output.append(f" {rec['description']}")
if rec.get('actions'):
output.append(f" Actions: {len(rec['actions'])} specific action items")
return "\n".join(output)
def main():
parser = argparse.ArgumentParser(description="Analyze interview data for bias and calibration issues")
parser.add_argument("--input", type=str, required=True, help="Input JSON file with interview results data")
parser.add_argument("--analysis-type", type=str, choices=["comprehensive", "bias", "calibration", "interviewer", "scoring"],
default="comprehensive", help="Type of analysis to perform")
parser.add_argument("--competencies", type=str, help="Comma-separated list of competencies to focus on")
parser.add_argument("--trend-analysis", action="store_true", help="Perform trend analysis over time")
parser.add_argument("--period", type=str, choices=["daily", "weekly", "monthly", "quarterly"],
default="monthly", help="Time period for trend analysis")
parser.add_argument("--output", type=str, help="Output file path")
parser.add_argument("--format", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
# Load input data
try:
with open(args.input, 'r') as f:
interview_data = json.load(f)
if not isinstance(interview_data, list):
print("Error: Input data must be a JSON array of interview records")
sys.exit(1)
except FileNotFoundError:
print(f"Error: Input file '{args.input}' not found")
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in input file: {e}")
sys.exit(1)
except Exception as e:
print(f"Error reading input file: {e}")
sys.exit(1)
# Initialize calibrator and run analysis
calibrator = HiringCalibrator()
competencies = args.competencies.split(',') if args.competencies else None
try:
results = calibrator.analyze_hiring_calibration(
interview_data=interview_data,
analysis_type=args.analysis_type,
competencies=competencies,
trend_analysis=args.trend_analysis,
period=args.period
)
# Handle output
if args.output:
output_path = args.output
json_path = output_path if output_path.endswith('.json') else f"{output_path}.json"
text_path = output_path.replace('.json', '.txt') if output_path.endswith('.json') else f"{output_path}.txt"
else:
base_filename = f"calibration_report_{datetime.now().strftime('%Y%m%d_%H%M%S')}"
json_path = f"{base_filename}.json"
text_path = f"{base_filename}.txt"
# Write outputs
if args.format in ["json", "both"]:
with open(json_path, 'w') as f:
json.dump(results, f, indent=2, default=str)
print(f"JSON report written to: {json_path}")
if args.format in ["text", "both"]:
with open(text_path, 'w') as f:
f.write(format_human_readable(results))
print(f"Text report written to: {text_path}")
# Print summary
print(f"\nCalibration Analysis Summary:")
if "error" in results:
print(f"Error: {results['error']}")
else:
health_score = results.get("calibration_health_score", {})
print(f"Health Score: {health_score.get('overall_score', 0):.3f} ({health_score.get('health_category', 'Unknown').title()})")
bias_score = results.get("bias_analysis", {}).get("overall_bias_score", 0)
print(f"Bias Score: {bias_score:.3f} (Lower is better)")
recommendations = results.get("recommendations", [])
print(f"Recommendations Generated: {len(recommendations)}")
if recommendations:
print(f"Top Priority: {recommendations[0]['title']} ({recommendations[0]['priority'].title()})")
except Exception as e:
print(f"Error during analysis: {e}")
sys.exit(1)
if __name__ == "__main__":
main()
FILE:loop_designer.py
#!/usr/bin/env python3
"""
Interview Loop Designer
Generates calibrated interview loops tailored to specific roles, levels, and teams.
Creates complete interview loops with rounds, focus areas, time allocation,
interviewer skill requirements, and scorecard templates.
Usage:
python loop_designer.py --role "Senior Software Engineer" --level senior --team platform
python loop_designer.py --role "Product Manager" --level mid --competencies leadership,strategy
python loop_designer.py --input role_definition.json --output loops/
"""
import os
import sys
import json
import argparse
from datetime import datetime, timedelta
from typing import Dict, List, Optional, Any, Tuple
from collections import defaultdict
class InterviewLoopDesigner:
"""Designs comprehensive interview loops based on role requirements."""
def __init__(self):
self.competency_frameworks = self._init_competency_frameworks()
self.role_templates = self._init_role_templates()
self.interviewer_skills = self._init_interviewer_skills()
def _init_competency_frameworks(self) -> Dict[str, Dict]:
"""Initialize competency frameworks for different roles."""
return {
"software_engineer": {
"junior": {
"required": ["coding_fundamentals", "debugging", "testing_basics", "version_control"],
"preferred": ["system_understanding", "code_review", "collaboration"],
"focus_areas": ["technical_execution", "learning_agility", "team_collaboration"]
},
"mid": {
"required": ["advanced_coding", "system_design_basics", "testing_strategy", "debugging_complex"],
"preferred": ["mentoring_basics", "technical_communication", "project_ownership"],
"focus_areas": ["technical_depth", "system_thinking", "ownership"]
},
"senior": {
"required": ["system_architecture", "technical_leadership", "mentoring", "cross_team_collab"],
"preferred": ["technology_evaluation", "process_improvement", "hiring_contribution"],
"focus_areas": ["technical_leadership", "system_architecture", "people_development"]
},
"staff": {
"required": ["architectural_vision", "organizational_impact", "technical_strategy", "team_building"],
"preferred": ["industry_influence", "innovation_leadership", "executive_communication"],
"focus_areas": ["organizational_impact", "technical_vision", "strategic_influence"]
},
"principal": {
"required": ["company_wide_impact", "technical_vision", "talent_development", "strategic_planning"],
"preferred": ["industry_leadership", "board_communication", "market_influence"],
"focus_areas": ["strategic_leadership", "organizational_transformation", "external_influence"]
}
},
"product_manager": {
"junior": {
"required": ["product_execution", "user_research", "data_analysis", "stakeholder_comm"],
"preferred": ["market_awareness", "technical_understanding", "project_management"],
"focus_areas": ["execution_excellence", "user_focus", "analytical_thinking"]
},
"mid": {
"required": ["product_strategy", "cross_functional_leadership", "metrics_design", "market_analysis"],
"preferred": ["team_building", "technical_collaboration", "competitive_analysis"],
"focus_areas": ["strategic_thinking", "leadership", "business_impact"]
},
"senior": {
"required": ["business_strategy", "team_leadership", "p&l_ownership", "market_positioning"],
"preferred": ["hiring_leadership", "board_communication", "partnership_development"],
"focus_areas": ["business_leadership", "market_strategy", "organizational_impact"]
},
"staff": {
"required": ["portfolio_management", "organizational_leadership", "strategic_planning", "market_creation"],
"preferred": ["executive_presence", "investor_relations", "acquisition_strategy"],
"focus_areas": ["strategic_leadership", "market_innovation", "organizational_transformation"]
}
},
"designer": {
"junior": {
"required": ["design_fundamentals", "user_research", "prototyping", "design_tools"],
"preferred": ["user_empathy", "visual_design", "collaboration"],
"focus_areas": ["design_execution", "user_research", "creative_problem_solving"]
},
"mid": {
"required": ["design_systems", "user_testing", "cross_functional_collab", "design_strategy"],
"preferred": ["mentoring", "process_improvement", "business_understanding"],
"focus_areas": ["design_leadership", "system_thinking", "business_impact"]
},
"senior": {
"required": ["design_leadership", "team_building", "strategic_design", "stakeholder_management"],
"preferred": ["design_culture", "hiring_leadership", "executive_communication"],
"focus_areas": ["design_strategy", "team_leadership", "organizational_impact"]
}
},
"data_scientist": {
"junior": {
"required": ["statistical_analysis", "python_r", "data_visualization", "sql"],
"preferred": ["machine_learning", "business_understanding", "communication"],
"focus_areas": ["analytical_skills", "technical_execution", "business_impact"]
},
"mid": {
"required": ["advanced_ml", "experiment_design", "data_engineering", "stakeholder_comm"],
"preferred": ["mentoring", "project_leadership", "product_collaboration"],
"focus_areas": ["advanced_analytics", "project_leadership", "cross_functional_impact"]
},
"senior": {
"required": ["data_strategy", "team_leadership", "ml_systems", "business_strategy"],
"preferred": ["hiring_leadership", "executive_communication", "technology_evaluation"],
"focus_areas": ["strategic_leadership", "technical_vision", "organizational_impact"]
}
},
"devops_engineer": {
"junior": {
"required": ["infrastructure_basics", "scripting", "monitoring", "troubleshooting"],
"preferred": ["automation", "cloud_platforms", "security_awareness"],
"focus_areas": ["operational_excellence", "automation_mindset", "problem_solving"]
},
"mid": {
"required": ["ci_cd_design", "infrastructure_as_code", "security_implementation", "performance_optimization"],
"preferred": ["team_collaboration", "incident_management", "capacity_planning"],
"focus_areas": ["system_reliability", "automation_leadership", "cross_team_collaboration"]
},
"senior": {
"required": ["platform_architecture", "team_leadership", "security_strategy", "organizational_impact"],
"preferred": ["hiring_contribution", "technology_evaluation", "executive_communication"],
"focus_areas": ["platform_leadership", "strategic_thinking", "organizational_transformation"]
}
},
"engineering_manager": {
"junior": {
"required": ["team_leadership", "technical_background", "people_management", "project_coordination"],
"preferred": ["hiring_experience", "performance_management", "technical_mentoring"],
"focus_areas": ["people_leadership", "team_building", "execution_excellence"]
},
"senior": {
"required": ["organizational_leadership", "strategic_planning", "talent_development", "cross_functional_leadership"],
"preferred": ["technical_vision", "culture_building", "executive_communication"],
"focus_areas": ["organizational_impact", "strategic_leadership", "talent_development"]
},
"staff": {
"required": ["multi_team_leadership", "organizational_strategy", "executive_presence", "cultural_transformation"],
"preferred": ["board_communication", "market_understanding", "acquisition_integration"],
"focus_areas": ["organizational_transformation", "strategic_leadership", "cultural_evolution"]
}
}
}
def _init_role_templates(self) -> Dict[str, Dict]:
"""Initialize role-specific interview templates."""
return {
"software_engineer": {
"core_rounds": ["technical_phone_screen", "coding_deep_dive", "system_design", "behavioral"],
"optional_rounds": ["technical_leadership", "domain_expertise", "culture_fit"],
"total_duration_range": (180, 360), # 3-6 hours
"required_competencies": ["coding", "problem_solving", "communication"]
},
"product_manager": {
"core_rounds": ["product_sense", "analytical_thinking", "execution_process", "behavioral"],
"optional_rounds": ["strategic_thinking", "technical_collaboration", "leadership"],
"total_duration_range": (180, 300), # 3-5 hours
"required_competencies": ["product_strategy", "analytical_thinking", "stakeholder_management"]
},
"designer": {
"core_rounds": ["portfolio_review", "design_challenge", "collaboration_process", "behavioral"],
"optional_rounds": ["design_system_thinking", "research_methodology", "leadership"],
"total_duration_range": (180, 300), # 3-5 hours
"required_competencies": ["design_process", "user_empathy", "visual_communication"]
},
"data_scientist": {
"core_rounds": ["technical_assessment", "case_study", "statistical_thinking", "behavioral"],
"optional_rounds": ["ml_systems", "business_strategy", "technical_leadership"],
"total_duration_range": (210, 330), # 3.5-5.5 hours
"required_competencies": ["statistical_analysis", "programming", "business_acumen"]
},
"devops_engineer": {
"core_rounds": ["technical_assessment", "system_design", "troubleshooting", "behavioral"],
"optional_rounds": ["security_assessment", "automation_design", "leadership"],
"total_duration_range": (180, 300), # 3-5 hours
"required_competencies": ["infrastructure", "automation", "problem_solving"]
},
"engineering_manager": {
"core_rounds": ["leadership_assessment", "technical_background", "people_management", "behavioral"],
"optional_rounds": ["strategic_thinking", "hiring_assessment", "culture_building"],
"total_duration_range": (240, 360), # 4-6 hours
"required_competencies": ["people_leadership", "technical_understanding", "strategic_thinking"]
}
}
def _init_interviewer_skills(self) -> Dict[str, Dict]:
"""Initialize interviewer skill requirements for different round types."""
return {
"technical_phone_screen": {
"required_skills": ["technical_assessment", "coding_evaluation"],
"preferred_experience": ["same_domain", "senior_level"],
"calibration_level": "standard"
},
"coding_deep_dive": {
"required_skills": ["advanced_technical", "code_quality_assessment"],
"preferred_experience": ["senior_engineer", "system_design"],
"calibration_level": "high"
},
"system_design": {
"required_skills": ["architecture_design", "scalability_assessment"],
"preferred_experience": ["senior_architect", "large_scale_systems"],
"calibration_level": "high"
},
"behavioral": {
"required_skills": ["behavioral_interviewing", "competency_assessment"],
"preferred_experience": ["hiring_manager", "people_leadership"],
"calibration_level": "standard"
},
"technical_leadership": {
"required_skills": ["leadership_assessment", "technical_mentoring"],
"preferred_experience": ["engineering_manager", "tech_lead"],
"calibration_level": "high"
},
"product_sense": {
"required_skills": ["product_evaluation", "market_analysis"],
"preferred_experience": ["product_manager", "product_leadership"],
"calibration_level": "high"
},
"analytical_thinking": {
"required_skills": ["data_analysis", "metrics_evaluation"],
"preferred_experience": ["data_analyst", "product_manager"],
"calibration_level": "standard"
},
"design_challenge": {
"required_skills": ["design_evaluation", "user_experience"],
"preferred_experience": ["senior_designer", "design_manager"],
"calibration_level": "high"
}
}
def generate_interview_loop(self, role: str, level: str, team: Optional[str] = None,
competencies: Optional[List[str]] = None) -> Dict[str, Any]:
"""Generate a complete interview loop for the specified role and level."""
# Normalize inputs
role_key = role.lower().replace(" ", "_").replace("-", "_")
level_key = level.lower()
# Get role template and competency requirements
if role_key not in self.competency_frameworks:
role_key = self._find_closest_role(role_key)
if level_key not in self.competency_frameworks[role_key]:
level_key = self._find_closest_level(role_key, level_key)
competency_req = self.competency_frameworks[role_key][level_key]
role_template = self.role_templates.get(role_key, self.role_templates["software_engineer"])
# Design the interview loop
rounds = self._design_rounds(role_key, level_key, competency_req, role_template, competencies)
schedule = self._create_schedule(rounds)
scorecard = self._generate_scorecard(role_key, level_key, competency_req)
interviewer_requirements = self._define_interviewer_requirements(rounds)
return {
"role": role,
"level": level,
"team": team,
"generated_at": datetime.now().isoformat(),
"total_duration_minutes": sum(round_info["duration_minutes"] for round_info in rounds.values()),
"total_rounds": len(rounds),
"rounds": rounds,
"suggested_schedule": schedule,
"scorecard_template": scorecard,
"interviewer_requirements": interviewer_requirements,
"competency_framework": competency_req,
"calibration_notes": self._generate_calibration_notes(role_key, level_key)
}
def _find_closest_role(self, role_key: str) -> str:
"""Find the closest matching role template."""
role_mappings = {
"engineer": "software_engineer",
"developer": "software_engineer",
"swe": "software_engineer",
"backend": "software_engineer",
"frontend": "software_engineer",
"fullstack": "software_engineer",
"pm": "product_manager",
"product": "product_manager",
"ux": "designer",
"ui": "designer",
"graphic": "designer",
"data": "data_scientist",
"analyst": "data_scientist",
"ml": "data_scientist",
"ops": "devops_engineer",
"sre": "devops_engineer",
"infrastructure": "devops_engineer",
"manager": "engineering_manager",
"lead": "engineering_manager"
}
for key_part in role_key.split("_"):
if key_part in role_mappings:
return role_mappings[key_part]
return "software_engineer" # Default fallback
def _find_closest_level(self, role_key: str, level_key: str) -> str:
"""Find the closest matching level for the role."""
available_levels = list(self.competency_frameworks[role_key].keys())
level_mappings = {
"entry": "junior",
"associate": "junior",
"jr": "junior",
"mid": "mid",
"middle": "mid",
"sr": "senior",
"senior": "senior",
"staff": "staff",
"principal": "principal",
"lead": "senior",
"manager": "senior"
}
mapped_level = level_mappings.get(level_key, level_key)
if mapped_level in available_levels:
return mapped_level
elif "senior" in available_levels:
return "senior"
else:
return available_levels[0]
def _design_rounds(self, role_key: str, level_key: str, competency_req: Dict,
role_template: Dict, custom_competencies: Optional[List[str]]) -> Dict[str, Dict]:
"""Design the specific interview rounds based on role and level."""
rounds = {}
# Determine which rounds to include
core_rounds = role_template["core_rounds"].copy()
optional_rounds = role_template["optional_rounds"].copy()
# Add optional rounds based on level
if level_key in ["senior", "staff", "principal"]:
if "technical_leadership" in optional_rounds and role_key in ["software_engineer", "engineering_manager"]:
core_rounds.append("technical_leadership")
if "strategic_thinking" in optional_rounds and role_key in ["product_manager", "engineering_manager"]:
core_rounds.append("strategic_thinking")
if "design_system_thinking" in optional_rounds and role_key == "designer":
core_rounds.append("design_system_thinking")
if level_key in ["staff", "principal"]:
if "domain_expertise" in optional_rounds:
core_rounds.append("domain_expertise")
# Define round details
round_definitions = self._get_round_definitions()
for i, round_type in enumerate(core_rounds, 1):
if round_type in round_definitions:
round_def = round_definitions[round_type].copy()
round_def["order"] = i
round_def["focus_areas"] = self._customize_focus_areas(round_type, competency_req, custom_competencies)
rounds[f"round_{i}_{round_type}"] = round_def
return rounds
def _get_round_definitions(self) -> Dict[str, Dict]:
"""Get predefined round definitions with standard durations and formats."""
return {
"technical_phone_screen": {
"name": "Technical Phone Screen",
"duration_minutes": 45,
"format": "virtual",
"objectives": ["Assess coding fundamentals", "Evaluate problem-solving approach", "Screen for basic technical competency"],
"question_types": ["coding_problems", "technical_concepts", "experience_questions"],
"evaluation_criteria": ["technical_accuracy", "problem_solving_process", "communication_clarity"]
},
"coding_deep_dive": {
"name": "Coding Deep Dive",
"duration_minutes": 75,
"format": "in_person_or_virtual",
"objectives": ["Evaluate coding skills in depth", "Assess code quality and testing", "Review debugging approach"],
"question_types": ["complex_coding_problems", "code_review", "testing_strategy"],
"evaluation_criteria": ["code_quality", "testing_approach", "debugging_skills", "optimization_thinking"]
},
"system_design": {
"name": "System Design",
"duration_minutes": 75,
"format": "collaborative_whiteboard",
"objectives": ["Assess architectural thinking", "Evaluate scalability considerations", "Review trade-off analysis"],
"question_types": ["system_architecture", "scalability_design", "trade_off_analysis"],
"evaluation_criteria": ["architectural_thinking", "scalability_awareness", "trade_off_reasoning"]
},
"behavioral": {
"name": "Behavioral Interview",
"duration_minutes": 45,
"format": "conversational",
"objectives": ["Assess cultural fit", "Evaluate past experiences", "Review leadership examples"],
"question_types": ["star_method_questions", "situational_scenarios", "values_alignment"],
"evaluation_criteria": ["communication_skills", "leadership_examples", "cultural_alignment"]
},
"technical_leadership": {
"name": "Technical Leadership",
"duration_minutes": 60,
"format": "discussion_based",
"objectives": ["Evaluate mentoring capability", "Assess technical decision making", "Review cross-team collaboration"],
"question_types": ["leadership_scenarios", "technical_decisions", "mentoring_examples"],
"evaluation_criteria": ["leadership_potential", "technical_judgment", "influence_skills"]
},
"product_sense": {
"name": "Product Sense",
"duration_minutes": 75,
"format": "case_study",
"objectives": ["Assess product intuition", "Evaluate user empathy", "Review market understanding"],
"question_types": ["product_scenarios", "feature_prioritization", "user_journey_analysis"],
"evaluation_criteria": ["product_intuition", "user_empathy", "analytical_thinking"]
},
"analytical_thinking": {
"name": "Analytical Thinking",
"duration_minutes": 60,
"format": "data_analysis",
"objectives": ["Evaluate data interpretation", "Assess metric design", "Review experiment planning"],
"question_types": ["data_interpretation", "metric_design", "experiment_analysis"],
"evaluation_criteria": ["analytical_rigor", "metric_intuition", "experimental_thinking"]
},
"design_challenge": {
"name": "Design Challenge",
"duration_minutes": 90,
"format": "hands_on_design",
"objectives": ["Assess design process", "Evaluate user-centered thinking", "Review iteration approach"],
"question_types": ["design_problems", "user_research", "design_critique"],
"evaluation_criteria": ["design_process", "user_focus", "visual_communication"]
},
"portfolio_review": {
"name": "Portfolio Review",
"duration_minutes": 75,
"format": "presentation_discussion",
"objectives": ["Review past work", "Assess design thinking", "Evaluate impact measurement"],
"question_types": ["portfolio_walkthrough", "design_decisions", "impact_stories"],
"evaluation_criteria": ["design_quality", "process_thinking", "business_impact"]
}
}
def _customize_focus_areas(self, round_type: str, competency_req: Dict,
custom_competencies: Optional[List[str]]) -> List[str]:
"""Customize focus areas based on role competency requirements."""
base_focus_areas = competency_req.get("focus_areas", [])
round_focus_mapping = {
"technical_phone_screen": ["coding_fundamentals", "problem_solving"],
"coding_deep_dive": ["technical_execution", "code_quality"],
"system_design": ["system_thinking", "architectural_reasoning"],
"behavioral": ["cultural_fit", "communication", "teamwork"],
"technical_leadership": ["leadership", "mentoring", "influence"],
"product_sense": ["product_intuition", "user_empathy"],
"analytical_thinking": ["data_analysis", "metric_design"],
"design_challenge": ["design_process", "user_focus"]
}
focus_areas = round_focus_mapping.get(round_type, [])
# Add custom competencies if specified
if custom_competencies:
focus_areas.extend([comp for comp in custom_competencies if comp not in focus_areas])
# Add role-specific focus areas
focus_areas.extend([area for area in base_focus_areas if area not in focus_areas])
return focus_areas[:5] # Limit to top 5 focus areas
def _create_schedule(self, rounds: Dict[str, Dict]) -> Dict[str, Any]:
"""Create a suggested interview schedule."""
sorted_rounds = sorted(rounds.items(), key=lambda x: x[1]["order"])
# Calculate optimal scheduling
total_duration = sum(round_info["duration_minutes"] for _, round_info in sorted_rounds)
if total_duration <= 240: # 4 hours or less - single day
schedule_type = "single_day"
day_structure = self._create_single_day_schedule(sorted_rounds)
else: # Multi-day schedule
schedule_type = "multi_day"
day_structure = self._create_multi_day_schedule(sorted_rounds)
return {
"type": schedule_type,
"total_duration_minutes": total_duration,
"recommended_breaks": self._calculate_breaks(total_duration),
"day_structure": day_structure,
"logistics_notes": self._generate_logistics_notes(sorted_rounds)
}
def _create_single_day_schedule(self, rounds: List[Tuple[str, Dict]]) -> Dict[str, Any]:
"""Create a single-day interview schedule."""
start_time = datetime.strptime("09:00", "%H:%M")
current_time = start_time
schedule = []
for round_name, round_info in rounds:
# Add break if needed (after 90 minutes of interviews)
if schedule and sum(item.get("duration_minutes", 0) for item in schedule if "break" not in item.get("type", "")) >= 90:
schedule.append({
"type": "break",
"start_time": current_time.strftime("%H:%M"),
"duration_minutes": 15,
"end_time": (current_time + timedelta(minutes=15)).strftime("%H:%M")
})
current_time += timedelta(minutes=15)
# Add the interview round
end_time = current_time + timedelta(minutes=round_info["duration_minutes"])
schedule.append({
"type": "interview",
"round_name": round_name,
"title": round_info["name"],
"start_time": current_time.strftime("%H:%M"),
"end_time": end_time.strftime("%H:%M"),
"duration_minutes": round_info["duration_minutes"],
"format": round_info["format"]
})
current_time = end_time
return {
"day_1": {
"date": "TBD",
"start_time": start_time.strftime("%H:%M"),
"end_time": current_time.strftime("%H:%M"),
"rounds": schedule
}
}
def _create_multi_day_schedule(self, rounds: List[Tuple[str, Dict]]) -> Dict[str, Any]:
"""Create a multi-day interview schedule."""
# Split rounds across days (max 4 hours per day)
max_daily_minutes = 240
days = {}
current_day = 1
current_day_duration = 0
current_day_rounds = []
for round_name, round_info in rounds:
duration = round_info["duration_minutes"] + 15 # Add buffer time
if current_day_duration + duration > max_daily_minutes and current_day_rounds:
# Finalize current day
days[f"day_{current_day}"] = self._finalize_day_schedule(current_day_rounds)
current_day += 1
current_day_duration = 0
current_day_rounds = []
current_day_rounds.append((round_name, round_info))
current_day_duration += duration
# Finalize last day
if current_day_rounds:
days[f"day_{current_day}"] = self._finalize_day_schedule(current_day_rounds)
return days
def _finalize_day_schedule(self, day_rounds: List[Tuple[str, Dict]]) -> Dict[str, Any]:
"""Finalize the schedule for a specific day."""
start_time = datetime.strptime("09:00", "%H:%M")
current_time = start_time
schedule = []
for round_name, round_info in day_rounds:
end_time = current_time + timedelta(minutes=round_info["duration_minutes"])
schedule.append({
"type": "interview",
"round_name": round_name,
"title": round_info["name"],
"start_time": current_time.strftime("%H:%M"),
"end_time": end_time.strftime("%H:%M"),
"duration_minutes": round_info["duration_minutes"],
"format": round_info["format"]
})
current_time = end_time + timedelta(minutes=15) # 15-min buffer
return {
"date": "TBD",
"start_time": start_time.strftime("%H:%M"),
"end_time": (current_time - timedelta(minutes=15)).strftime("%H:%M"),
"rounds": schedule
}
def _calculate_breaks(self, total_duration: int) -> List[Dict[str, Any]]:
"""Calculate recommended breaks based on total duration."""
breaks = []
if total_duration >= 120: # 2+ hours
breaks.append({"type": "short_break", "duration": 15, "after_minutes": 90})
if total_duration >= 240: # 4+ hours
breaks.append({"type": "lunch_break", "duration": 60, "after_minutes": 180})
if total_duration >= 360: # 6+ hours
breaks.append({"type": "short_break", "duration": 15, "after_minutes": 300})
return breaks
def _generate_scorecard(self, role_key: str, level_key: str, competency_req: Dict) -> Dict[str, Any]:
"""Generate a scorecard template for the interview loop."""
scoring_dimensions = []
# Add competency-based scoring dimensions
for competency in competency_req["required"]:
scoring_dimensions.append({
"dimension": competency,
"weight": "high",
"scale": "1-4",
"description": f"Assessment of {competency.replace('_', ' ')} competency"
})
for competency in competency_req.get("preferred", []):
scoring_dimensions.append({
"dimension": competency,
"weight": "medium",
"scale": "1-4",
"description": f"Assessment of {competency.replace('_', ' ')} competency"
})
# Add standard dimensions
standard_dimensions = [
{"dimension": "communication", "weight": "high", "scale": "1-4"},
{"dimension": "cultural_fit", "weight": "medium", "scale": "1-4"},
{"dimension": "learning_agility", "weight": "medium", "scale": "1-4"}
]
scoring_dimensions.extend(standard_dimensions)
return {
"scoring_scale": {
"4": "Exceeds Expectations - Demonstrates mastery beyond required level",
"3": "Meets Expectations - Solid performance meeting all requirements",
"2": "Partially Meets - Shows potential but has development areas",
"1": "Does Not Meet - Significant gaps in required competencies"
},
"dimensions": scoring_dimensions,
"overall_recommendation": {
"options": ["Strong Hire", "Hire", "No Hire", "Strong No Hire"],
"criteria": "Based on weighted average and minimum thresholds"
},
"calibration_notes": {
"required": True,
"min_length": 100,
"sections": ["strengths", "areas_for_development", "specific_examples"]
}
}
def _define_interviewer_requirements(self, rounds: Dict[str, Dict]) -> Dict[str, Dict]:
"""Define interviewer skill requirements for each round."""
requirements = {}
for round_name, round_info in rounds.items():
round_type = round_name.split("_", 2)[-1] # Extract round type
if round_type in self.interviewer_skills:
skill_req = self.interviewer_skills[round_type].copy()
skill_req["suggested_interviewers"] = self._suggest_interviewer_profiles(round_type)
requirements[round_name] = skill_req
else:
# Default requirements
requirements[round_name] = {
"required_skills": ["interviewing_basics", "evaluation_skills"],
"preferred_experience": ["relevant_domain"],
"calibration_level": "standard",
"suggested_interviewers": ["experienced_interviewer"]
}
return requirements
def _suggest_interviewer_profiles(self, round_type: str) -> List[str]:
"""Suggest specific interviewer profiles for different round types."""
profile_mapping = {
"technical_phone_screen": ["senior_engineer", "tech_lead"],
"coding_deep_dive": ["senior_engineer", "staff_engineer"],
"system_design": ["senior_architect", "staff_engineer"],
"behavioral": ["hiring_manager", "people_manager"],
"technical_leadership": ["engineering_manager", "senior_staff"],
"product_sense": ["senior_pm", "product_leader"],
"analytical_thinking": ["senior_analyst", "data_scientist"],
"design_challenge": ["senior_designer", "design_manager"]
}
return profile_mapping.get(round_type, ["experienced_interviewer"])
def _generate_calibration_notes(self, role_key: str, level_key: str) -> Dict[str, Any]:
"""Generate calibration notes and best practices."""
return {
"hiring_bar_notes": f"Calibrated for {level_key} level {role_key.replace('_', ' ')} role",
"common_pitfalls": [
"Avoid comparing candidates to each other rather than to the role standard",
"Don't let one strong/weak area overshadow overall assessment",
"Ensure consistent application of evaluation criteria"
],
"calibration_checkpoints": [
"Review score distribution after every 5 candidates",
"Conduct monthly interviewer calibration sessions",
"Track correlation with 6-month performance reviews"
],
"escalation_criteria": [
"Any candidate receiving all 4s or all 1s",
"Significant disagreement between interviewers (>1.5 point spread)",
"Unusual circumstances or accommodations needed"
]
}
def _generate_logistics_notes(self, rounds: List[Tuple[str, Dict]]) -> List[str]:
"""Generate logistics and coordination notes."""
notes = [
"Coordinate interviewer availability before scheduling",
"Ensure all interviewers have access to job description and competency requirements",
"Prepare interview rooms/virtual links for all rounds",
"Share candidate resume and application with all interviewers"
]
# Add format-specific notes
formats_used = {round_info["format"] for _, round_info in rounds}
if "virtual" in formats_used:
notes.append("Test video conferencing setup before virtual interviews")
notes.append("Share virtual meeting links with candidate 24 hours in advance")
if "collaborative_whiteboard" in formats_used:
notes.append("Prepare whiteboard or collaborative online tool for design sessions")
if "hands_on_design" in formats_used:
notes.append("Provide design tools access or ensure candidate can screen share their preferred tools")
return notes
def format_human_readable(loop_data: Dict[str, Any]) -> str:
"""Format the interview loop data in a human-readable format."""
output = []
# Header
output.append(f"Interview Loop Design for {loop_data['role']} ({loop_data['level'].title()} Level)")
output.append("=" * 60)
if loop_data.get('team'):
output.append(f"Team: {loop_data['team']}")
output.append(f"Generated: {loop_data['generated_at']}")
output.append(f"Total Duration: {loop_data['total_duration_minutes']} minutes ({loop_data['total_duration_minutes']//60}h {loop_data['total_duration_minutes']%60}m)")
output.append(f"Total Rounds: {loop_data['total_rounds']}")
output.append("")
# Interview Rounds
output.append("INTERVIEW ROUNDS")
output.append("-" * 40)
sorted_rounds = sorted(loop_data['rounds'].items(), key=lambda x: x[1]['order'])
for round_name, round_info in sorted_rounds:
output.append(f"\nRound {round_info['order']}: {round_info['name']}")
output.append(f"Duration: {round_info['duration_minutes']} minutes")
output.append(f"Format: {round_info['format'].replace('_', ' ').title()}")
output.append("Objectives:")
for obj in round_info['objectives']:
output.append(f" • {obj}")
output.append("Focus Areas:")
for area in round_info['focus_areas']:
output.append(f" • {area.replace('_', ' ').title()}")
# Suggested Schedule
output.append("\nSUGGESTED SCHEDULE")
output.append("-" * 40)
schedule = loop_data['suggested_schedule']
output.append(f"Schedule Type: {schedule['type'].replace('_', ' ').title()}")
for day_name, day_info in schedule['day_structure'].items():
output.append(f"\n{day_name.replace('_', ' ').title()}:")
output.append(f"Time: {day_info['start_time']} - {day_info['end_time']}")
for item in day_info['rounds']:
if item['type'] == 'interview':
output.append(f" {item['start_time']}-{item['end_time']}: {item['title']} ({item['duration_minutes']}min)")
else:
output.append(f" {item['start_time']}-{item['end_time']}: {item['type'].title()} ({item['duration_minutes']}min)")
# Interviewer Requirements
output.append("\nINTERVIEWER REQUIREMENTS")
output.append("-" * 40)
for round_name, requirements in loop_data['interviewer_requirements'].items():
round_display = round_name.split("_", 2)[-1].replace("_", " ").title()
output.append(f"\n{round_display}:")
output.append(f"Required Skills: {', '.join(requirements['required_skills'])}")
output.append(f"Suggested Interviewers: {', '.join(requirements['suggested_interviewers'])}")
output.append(f"Calibration Level: {requirements['calibration_level'].title()}")
# Scorecard Overview
output.append("\nSCORECARD TEMPLATE")
output.append("-" * 40)
scorecard = loop_data['scorecard_template']
output.append("Scoring Scale:")
for score, description in scorecard['scoring_scale'].items():
output.append(f" {score}: {description}")
output.append("\nEvaluation Dimensions:")
for dim in scorecard['dimensions']:
output.append(f" • {dim['dimension'].replace('_', ' ').title()} (Weight: {dim['weight']})")
# Calibration Notes
output.append("\nCALIBRATION NOTES")
output.append("-" * 40)
calibration = loop_data['calibration_notes']
output.append(f"Hiring Bar: {calibration['hiring_bar_notes']}")
output.append("\nCommon Pitfalls:")
for pitfall in calibration['common_pitfalls']:
output.append(f" • {pitfall}")
return "\n".join(output)
def main():
parser = argparse.ArgumentParser(description="Generate calibrated interview loops for specific roles and levels")
parser.add_argument("--role", type=str, help="Job role title (e.g., 'Senior Software Engineer')")
parser.add_argument("--level", type=str, help="Experience level (junior, mid, senior, staff, principal)")
parser.add_argument("--team", type=str, help="Team or department (optional)")
parser.add_argument("--competencies", type=str, help="Comma-separated list of specific competencies to focus on")
parser.add_argument("--input", type=str, help="Input JSON file with role definition")
parser.add_argument("--output", type=str, help="Output directory or file path")
parser.add_argument("--format", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
designer = InterviewLoopDesigner()
# Handle input
if args.input:
try:
with open(args.input, 'r') as f:
role_data = json.load(f)
role = role_data.get('role') or role_data.get('title', '')
level = role_data.get('level', 'senior')
team = role_data.get('team')
competencies = role_data.get('competencies')
except Exception as e:
print(f"Error reading input file: {e}")
sys.exit(1)
else:
if not args.role or not args.level:
print("Error: --role and --level are required when not using --input")
sys.exit(1)
role = args.role
level = args.level
team = args.team
competencies = args.competencies.split(',') if args.competencies else None
# Generate interview loop
try:
loop_data = designer.generate_interview_loop(role, level, team, competencies)
# Handle output
if args.output:
output_path = args.output
if os.path.isdir(output_path):
safe_role = "".join(c for c in role.lower() if c.isalnum() or c in (' ', '-', '_')).replace(' ', '_')
base_filename = f"{safe_role}_{level}_interview_loop"
json_path = os.path.join(output_path, f"{base_filename}.json")
text_path = os.path.join(output_path, f"{base_filename}.txt")
else:
# Use provided path as base
json_path = output_path if output_path.endswith('.json') else f"{output_path}.json"
text_path = output_path.replace('.json', '.txt') if output_path.endswith('.json') else f"{output_path}.txt"
else:
safe_role = "".join(c for c in role.lower() if c.isalnum() or c in (' ', '-', '_')).replace(' ', '_')
base_filename = f"{safe_role}_{level}_interview_loop"
json_path = f"{base_filename}.json"
text_path = f"{base_filename}.txt"
# Write outputs
if args.format in ["json", "both"]:
with open(json_path, 'w') as f:
json.dump(loop_data, f, indent=2, default=str)
print(f"JSON output written to: {json_path}")
if args.format in ["text", "both"]:
with open(text_path, 'w') as f:
f.write(format_human_readable(loop_data))
print(f"Text output written to: {text_path}")
# Always print summary to stdout
print("\nInterview Loop Summary:")
print(f"Role: {loop_data['role']} ({loop_data['level'].title()})")
print(f"Total Duration: {loop_data['total_duration_minutes']} minutes")
print(f"Number of Rounds: {loop_data['total_rounds']}")
print(f"Schedule Type: {loop_data['suggested_schedule']['type'].replace('_', ' ').title()}")
except Exception as e:
print(f"Error generating interview loop: {e}")
sys.exit(1)
if __name__ == "__main__":
main()
FILE:question_bank_generator.py
#!/usr/bin/env python3
"""
Question Bank Generator
Generates comprehensive, competency-based interview questions with detailed scoring criteria.
Creates structured question banks organized by competency area with scoring rubrics,
follow-up probes, and calibration examples.
Usage:
python question_bank_generator.py --role "Frontend Engineer" --competencies react,typescript,system-design
python question_bank_generator.py --role "Product Manager" --question-types behavioral,leadership
python question_bank_generator.py --input role_requirements.json --output questions/
"""
import os
import sys
import json
import argparse
import random
from datetime import datetime
from typing import Dict, List, Optional, Any, Tuple
from collections import defaultdict
class QuestionBankGenerator:
"""Generates comprehensive interview question banks with scoring criteria."""
def __init__(self):
self.technical_questions = self._init_technical_questions()
self.behavioral_questions = self._init_behavioral_questions()
self.competency_mapping = self._init_competency_mapping()
self.scoring_rubrics = self._init_scoring_rubrics()
self.follow_up_strategies = self._init_follow_up_strategies()
def _init_technical_questions(self) -> Dict[str, Dict]:
"""Initialize technical questions by competency area and level."""
return {
"coding_fundamentals": {
"junior": [
{
"question": "Write a function to reverse a string without using built-in reverse methods.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "easy",
"time_limit": 15,
"key_concepts": ["loops", "string_manipulation", "basic_algorithms"]
},
{
"question": "Implement a function to check if a string is a palindrome.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "easy",
"time_limit": 15,
"key_concepts": ["string_processing", "comparison", "edge_cases"]
},
{
"question": "Find the largest element in an array without using built-in max functions.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "easy",
"time_limit": 10,
"key_concepts": ["arrays", "iteration", "comparison"]
}
],
"mid": [
{
"question": "Implement a function to find the first non-repeating character in a string.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "medium",
"time_limit": 20,
"key_concepts": ["hash_maps", "string_processing", "efficiency"]
},
{
"question": "Write a function to merge two sorted arrays into one sorted array.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "medium",
"time_limit": 25,
"key_concepts": ["merge_algorithms", "two_pointers", "optimization"]
}
],
"senior": [
{
"question": "Implement a LRU (Least Recently Used) cache with O(1) operations.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "hard",
"time_limit": 35,
"key_concepts": ["data_structures", "hash_maps", "doubly_linked_lists"]
}
]
},
"system_design": {
"mid": [
{
"question": "Design a URL shortener service like bit.ly for 10K users.",
"competency": "system_design",
"type": "design",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["database_design", "hashing", "basic_scalability"]
}
],
"senior": [
{
"question": "Design a real-time chat system supporting 1M concurrent users.",
"competency": "system_design",
"type": "design",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["websockets", "load_balancing", "database_sharding", "caching"]
},
{
"question": "Design a distributed cache system like Redis with high availability.",
"competency": "system_design",
"type": "design",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["distributed_systems", "replication", "consistency", "partitioning"]
}
],
"staff": [
{
"question": "Design the architecture for a global content delivery network (CDN).",
"competency": "system_design",
"type": "design",
"difficulty": "expert",
"time_limit": 75,
"key_concepts": ["global_architecture", "edge_computing", "content_optimization", "network_protocols"]
}
]
},
"frontend_development": {
"junior": [
{
"question": "Create a responsive navigation menu using HTML, CSS, and vanilla JavaScript.",
"competency": "frontend_development",
"type": "coding",
"difficulty": "easy",
"time_limit": 30,
"key_concepts": ["html_css", "responsive_design", "dom_manipulation"]
}
],
"mid": [
{
"question": "Build a React component that fetches and displays paginated data from an API.",
"competency": "frontend_development",
"type": "coding",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["react_hooks", "api_integration", "state_management", "pagination"]
}
],
"senior": [
{
"question": "Design and implement a custom React hook for managing complex form state with validation.",
"competency": "frontend_development",
"type": "coding",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["custom_hooks", "form_validation", "state_management", "performance"]
}
]
},
"data_analysis": {
"junior": [
{
"question": "Given a dataset of user activities, calculate the daily active users for the past month.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "easy",
"time_limit": 30,
"key_concepts": ["sql_basics", "date_functions", "aggregation"]
}
],
"mid": [
{
"question": "Analyze conversion funnel data to identify the biggest drop-off point and propose solutions.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["funnel_analysis", "conversion_optimization", "statistical_significance"]
}
],
"senior": [
{
"question": "Design an A/B testing framework to measure the impact of a new recommendation algorithm.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["experiment_design", "statistical_power", "bias_mitigation", "causal_inference"]
}
]
},
"machine_learning": {
"mid": [
{
"question": "Explain how you would build a recommendation system for an e-commerce platform.",
"competency": "machine_learning",
"type": "conceptual",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["collaborative_filtering", "content_based", "cold_start", "evaluation_metrics"]
}
],
"senior": [
{
"question": "Design a real-time fraud detection system for financial transactions.",
"competency": "machine_learning",
"type": "design",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["anomaly_detection", "real_time_ml", "feature_engineering", "model_monitoring"]
}
]
},
"product_strategy": {
"mid": [
{
"question": "How would you prioritize features for a mobile app with limited engineering resources?",
"competency": "product_strategy",
"type": "case_study",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["prioritization_frameworks", "resource_allocation", "impact_estimation"]
}
],
"senior": [
{
"question": "Design a go-to-market strategy for a new B2B SaaS product entering a competitive market.",
"competency": "product_strategy",
"type": "strategic",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["market_analysis", "competitive_positioning", "pricing_strategy", "channel_strategy"]
}
]
}
}
def _init_behavioral_questions(self) -> Dict[str, List[Dict]]:
"""Initialize behavioral questions by competency area."""
return {
"leadership": [
{
"question": "Tell me about a time when you had to lead a team through a significant change or challenge.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["change_management", "team_motivation", "communication"]
},
{
"question": "Describe a situation where you had to influence someone without having direct authority over them.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["influence", "persuasion", "stakeholder_management"]
},
{
"question": "Give me an example of when you had to make a difficult decision that affected your team.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["decision_making", "team_impact", "communication"]
}
],
"collaboration": [
{
"question": "Describe a time when you had to work with a difficult colleague or stakeholder.",
"competency": "collaboration",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["conflict_resolution", "relationship_building", "professionalism"]
},
{
"question": "Tell me about a project where you had to coordinate across multiple teams or departments.",
"competency": "collaboration",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["cross_functional_work", "communication", "project_coordination"]
}
],
"problem_solving": [
{
"question": "Walk me through a complex problem you solved recently. What was your approach?",
"competency": "problem_solving",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["analytical_thinking", "methodology", "creativity"]
},
{
"question": "Describe a time when you had to solve a problem with limited information or resources.",
"competency": "problem_solving",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["resourcefulness", "ambiguity_tolerance", "decision_making"]
}
],
"communication": [
{
"question": "Tell me about a time when you had to present complex technical information to a non-technical audience.",
"competency": "communication",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["technical_communication", "audience_adaptation", "clarity"]
},
{
"question": "Describe a situation where you had to deliver difficult feedback to a colleague.",
"competency": "communication",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["feedback_delivery", "empathy", "constructive_criticism"]
}
],
"adaptability": [
{
"question": "Tell me about a time when you had to quickly learn a new technology or skill for work.",
"competency": "adaptability",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["learning_agility", "growth_mindset", "knowledge_acquisition"]
},
{
"question": "Describe how you handled a situation when project requirements changed significantly mid-way.",
"competency": "adaptability",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["flexibility", "change_management", "resilience"]
}
],
"innovation": [
{
"question": "Tell me about a time when you came up with a creative solution to improve a process or solve a problem.",
"competency": "innovation",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["creative_thinking", "process_improvement", "initiative"]
}
]
}
def _init_competency_mapping(self) -> Dict[str, Dict]:
"""Initialize role to competency mapping."""
return {
"software_engineer": {
"core_competencies": ["coding_fundamentals", "system_design", "problem_solving", "collaboration"],
"level_specific": {
"junior": ["coding_fundamentals", "debugging", "learning_agility"],
"mid": ["advanced_coding", "system_design", "mentoring_basics"],
"senior": ["system_architecture", "technical_leadership", "innovation"],
"staff": ["architectural_vision", "organizational_impact", "strategic_thinking"]
}
},
"frontend_engineer": {
"core_competencies": ["frontend_development", "ui_ux_understanding", "problem_solving", "collaboration"],
"level_specific": {
"junior": ["html_css_js", "responsive_design", "basic_frameworks"],
"mid": ["react_vue_angular", "state_management", "performance_optimization"],
"senior": ["frontend_architecture", "team_leadership", "cross_functional_collaboration"],
"staff": ["frontend_strategy", "technology_evaluation", "organizational_impact"]
}
},
"backend_engineer": {
"core_competencies": ["backend_development", "database_design", "api_design", "system_design"],
"level_specific": {
"junior": ["server_side_programming", "database_basics", "api_consumption"],
"mid": ["microservices", "caching", "security_basics"],
"senior": ["distributed_systems", "performance_optimization", "technical_leadership"],
"staff": ["system_architecture", "technology_strategy", "cross_team_influence"]
}
},
"product_manager": {
"core_competencies": ["product_strategy", "user_research", "data_analysis", "stakeholder_management"],
"level_specific": {
"junior": ["feature_specification", "user_stories", "basic_analytics"],
"mid": ["product_roadmap", "cross_functional_leadership", "market_research"],
"senior": ["business_strategy", "team_leadership", "p&l_responsibility"],
"staff": ["portfolio_management", "organizational_strategy", "market_creation"]
}
},
"data_scientist": {
"core_competencies": ["statistical_analysis", "machine_learning", "data_analysis", "business_acumen"],
"level_specific": {
"junior": ["python_r", "sql", "basic_ml", "data_visualization"],
"mid": ["advanced_ml", "experiment_design", "model_evaluation"],
"senior": ["ml_systems", "data_strategy", "stakeholder_communication"],
"staff": ["data_platform", "ai_strategy", "organizational_impact"]
}
},
"designer": {
"core_competencies": ["design_process", "user_research", "visual_design", "collaboration"],
"level_specific": {
"junior": ["design_tools", "user_empathy", "visual_communication"],
"mid": ["design_systems", "user_testing", "cross_functional_work"],
"senior": ["design_strategy", "team_leadership", "business_impact"],
"staff": ["design_vision", "organizational_design", "strategic_influence"]
}
},
"devops_engineer": {
"core_competencies": ["infrastructure", "automation", "monitoring", "troubleshooting"],
"level_specific": {
"junior": ["scripting", "basic_cloud", "ci_cd_basics"],
"mid": ["infrastructure_as_code", "container_orchestration", "security"],
"senior": ["platform_design", "reliability_engineering", "team_leadership"],
"staff": ["platform_strategy", "organizational_infrastructure", "technology_vision"]
}
}
}
def _init_scoring_rubrics(self) -> Dict[str, Dict]:
"""Initialize scoring rubrics for different question types."""
return {
"coding": {
"correctness": {
"4": "Solution is completely correct, handles all edge cases, optimal complexity",
"3": "Solution is correct for main cases, good complexity, minor edge case issues",
"2": "Solution works but has some bugs or suboptimal approach",
"1": "Solution has significant issues or doesn't work"
},
"code_quality": {
"4": "Clean, readable, well-structured code with excellent naming and comments",
"3": "Good code structure, readable with appropriate naming",
"2": "Code works but has style/structure issues",
"1": "Poor code quality, hard to understand"
},
"problem_solving_approach": {
"4": "Excellent problem breakdown, clear thinking process, considers alternatives",
"3": "Good approach, logical thinking, systematic problem solving",
"2": "Decent approach but some confusion or inefficiency",
"1": "Poor approach, unclear thinking process"
},
"communication": {
"4": "Excellent explanation of approach, asks clarifying questions, clear reasoning",
"3": "Good communication, explains thinking well",
"2": "Adequate communication, some explanation",
"1": "Poor communication, little explanation"
}
},
"behavioral": {
"situation_clarity": {
"4": "Clear, specific situation with relevant context and stakes",
"3": "Good situation description with adequate context",
"2": "Situation described but lacks some specifics",
"1": "Vague or unclear situation description"
},
"action_quality": {
"4": "Specific, thoughtful actions showing strong competency",
"3": "Good actions demonstrating competency",
"2": "Adequate actions but could be stronger",
"1": "Weak or inappropriate actions"
},
"result_impact": {
"4": "Significant positive impact with measurable results",
"3": "Good positive impact with clear outcomes",
"2": "Some positive impact demonstrated",
"1": "Little or no positive impact shown"
},
"self_awareness": {
"4": "Excellent self-reflection, learns from experience, acknowledges growth areas",
"3": "Good self-awareness and learning orientation",
"2": "Some self-reflection demonstrated",
"1": "Limited self-awareness or reflection"
}
},
"design": {
"system_thinking": {
"4": "Comprehensive system view, considers all components and interactions",
"3": "Good system understanding with most components identified",
"2": "Basic system thinking with some gaps",
"1": "Limited system thinking, misses key components"
},
"scalability": {
"4": "Excellent scalability considerations, multiple strategies discussed",
"3": "Good scalability awareness with practical solutions",
"2": "Basic scalability understanding",
"1": "Little to no scalability consideration"
},
"trade_offs": {
"4": "Excellent trade-off analysis, considers multiple dimensions",
"3": "Good trade-off awareness with clear reasoning",
"2": "Some trade-off consideration",
"1": "Limited trade-off analysis"
},
"technical_depth": {
"4": "Deep technical knowledge with implementation details",
"3": "Good technical knowledge with solid understanding",
"2": "Adequate technical knowledge",
"1": "Limited technical depth"
}
}
}
def _init_follow_up_strategies(self) -> Dict[str, List[str]]:
"""Initialize follow-up question strategies by competency."""
return {
"coding_fundamentals": [
"How would you optimize this solution for better time complexity?",
"What edge cases should we consider for this problem?",
"How would you test this function?",
"What would happen if the input size was very large?"
],
"system_design": [
"How would you handle if the system needed to scale 10x?",
"What would you do if one of your services went down?",
"How would you monitor this system in production?",
"What security considerations would you implement?"
],
"leadership": [
"What would you do differently if you faced this situation again?",
"How did you handle team members who were resistant to the change?",
"What metrics did you use to measure success?",
"How did you communicate progress to stakeholders?"
],
"problem_solving": [
"Walk me through your thought process step by step",
"What alternative approaches did you consider?",
"How did you validate your solution worked?",
"What did you learn from this experience?"
],
"collaboration": [
"How did you build consensus among the different stakeholders?",
"What communication channels did you use to keep everyone aligned?",
"How did you handle disagreements or conflicts?",
"What would you do to improve collaboration in the future?"
]
}
def generate_question_bank(self, role: str, level: str = "senior",
competencies: Optional[List[str]] = None,
question_types: Optional[List[str]] = None,
num_questions: int = 20) -> Dict[str, Any]:
"""Generate a comprehensive question bank for the specified role and competencies."""
# Normalize inputs
role_key = self._normalize_role(role)
level_key = level.lower()
# Get competency requirements
role_competencies = self._get_role_competencies(role_key, level_key, competencies)
# Determine question types to include
if question_types is None:
question_types = ["technical", "behavioral", "situational"]
# Generate questions
questions = self._generate_questions(role_competencies, question_types, level_key, num_questions)
# Create scoring rubrics
scoring_rubrics = self._create_scoring_rubrics(questions)
# Generate follow-up probes
follow_up_probes = self._generate_follow_up_probes(questions)
# Create calibration examples
calibration_examples = self._create_calibration_examples(questions[:5]) # Sample for first 5 questions
return {
"role": role,
"level": level,
"competencies": role_competencies,
"question_types": question_types,
"generated_at": datetime.now().isoformat(),
"total_questions": len(questions),
"questions": questions,
"scoring_rubrics": scoring_rubrics,
"follow_up_probes": follow_up_probes,
"calibration_examples": calibration_examples,
"usage_guidelines": self._generate_usage_guidelines(role_key, level_key)
}
def _normalize_role(self, role: str) -> str:
"""Normalize role name to match competency mapping keys."""
role_lower = role.lower().replace(" ", "_").replace("-", "_")
# Map variations to standard roles
role_mappings = {
"software_engineer": ["engineer", "developer", "swe", "software_developer"],
"frontend_engineer": ["frontend", "front_end", "ui_engineer", "web_developer"],
"backend_engineer": ["backend", "back_end", "server_engineer", "api_developer"],
"product_manager": ["pm", "product", "product_owner", "po"],
"data_scientist": ["ds", "data", "analyst", "ml_engineer"],
"designer": ["ux", "ui", "ux_ui", "product_designer", "visual_designer"],
"devops_engineer": ["devops", "sre", "platform_engineer", "infrastructure"]
}
for standard_role, variations in role_mappings.items():
if any(var in role_lower for var in variations):
return standard_role
# Default fallback
return "software_engineer"
def _get_role_competencies(self, role_key: str, level_key: str,
custom_competencies: Optional[List[str]]) -> List[str]:
"""Get competencies for the role and level."""
if role_key not in self.competency_mapping:
role_key = "software_engineer"
role_mapping = self.competency_mapping[role_key]
competencies = role_mapping["core_competencies"].copy()
# Add level-specific competencies
if level_key in role_mapping["level_specific"]:
competencies.extend(role_mapping["level_specific"][level_key])
elif "senior" in role_mapping["level_specific"]:
competencies.extend(role_mapping["level_specific"]["senior"])
# Add custom competencies if specified
if custom_competencies:
competencies.extend([comp.strip() for comp in custom_competencies if comp.strip() not in competencies])
return list(set(competencies)) # Remove duplicates
def _generate_questions(self, competencies: List[str], question_types: List[str],
level: str, num_questions: int) -> List[Dict[str, Any]]:
"""Generate questions based on competencies and types."""
questions = []
questions_per_competency = max(1, num_questions // len(competencies))
for competency in competencies:
competency_questions = []
# Add technical questions if requested and available
if "technical" in question_types and competency in self.technical_questions:
tech_questions = []
# Get questions for current level and below
level_order = ["junior", "mid", "senior", "staff", "principal"]
current_level_idx = level_order.index(level) if level in level_order else 2
for lvl_idx in range(current_level_idx + 1):
lvl = level_order[lvl_idx]
if lvl in self.technical_questions[competency]:
tech_questions.extend(self.technical_questions[competency][lvl])
competency_questions.extend(tech_questions[:questions_per_competency])
# Add behavioral questions if requested
if "behavioral" in question_types and competency in self.behavioral_questions:
behavioral_q = self.behavioral_questions[competency][:questions_per_competency]
competency_questions.extend(behavioral_q)
# Add situational questions (variations of behavioral)
if "situational" in question_types:
situational_q = self._generate_situational_questions(competency, questions_per_competency)
competency_questions.extend(situational_q)
# Ensure we have enough questions for this competency
while len(competency_questions) < questions_per_competency:
competency_questions.extend(self._generate_fallback_questions(competency, level))
if len(competency_questions) >= questions_per_competency:
break
questions.extend(competency_questions[:questions_per_competency])
# Shuffle and limit to requested number
random.shuffle(questions)
return questions[:num_questions]
def _generate_situational_questions(self, competency: str, count: int) -> List[Dict[str, Any]]:
"""Generate situational questions for a competency."""
situational_templates = {
"leadership": [
{
"question": "You're leading a project that's behind schedule and the client is unhappy. How do you handle this situation?",
"competency": competency,
"type": "situational",
"focus_areas": ["crisis_management", "client_communication", "team_leadership"]
}
],
"collaboration": [
{
"question": "You're working on a cross-functional project and two team members have opposing views on the technical approach. How do you resolve this?",
"competency": competency,
"type": "situational",
"focus_areas": ["conflict_resolution", "technical_decision_making", "facilitation"]
}
],
"problem_solving": [
{
"question": "You've been assigned to improve the performance of a critical system, but you have limited time and budget. Walk me through your approach.",
"competency": competency,
"type": "situational",
"focus_areas": ["prioritization", "resource_constraints", "systematic_approach"]
}
]
}
if competency in situational_templates:
return situational_templates[competency][:count]
return []
def _generate_fallback_questions(self, competency: str, level: str) -> List[Dict[str, Any]]:
"""Generate fallback questions when specific ones aren't available."""
fallback_questions = [
{
"question": f"Describe your experience with {competency.replace('_', ' ')} in your current or previous role.",
"competency": competency,
"type": "experience",
"focus_areas": ["experience_depth", "practical_application"]
},
{
"question": f"What challenges have you faced related to {competency.replace('_', ' ')} and how did you overcome them?",
"competency": competency,
"type": "challenge_based",
"focus_areas": ["problem_solving", "learning_from_experience"]
}
]
return fallback_questions
def _create_scoring_rubrics(self, questions: List[Dict[str, Any]]) -> Dict[str, Dict]:
"""Create scoring rubrics for the generated questions."""
rubrics = {}
for i, question in enumerate(questions, 1):
question_key = f"question_{i}"
question_type = question.get("type", "behavioral")
if question_type in self.scoring_rubrics:
rubrics[question_key] = {
"question": question["question"],
"competency": question["competency"],
"type": question_type,
"scoring_criteria": self.scoring_rubrics[question_type],
"weight": self._determine_question_weight(question),
"time_limit": question.get("time_limit", 30)
}
return rubrics
def _determine_question_weight(self, question: Dict[str, Any]) -> str:
"""Determine the weight/importance of a question."""
competency = question.get("competency", "")
question_type = question.get("type", "")
difficulty = question.get("difficulty", "medium")
# Core competencies get higher weight
core_competencies = ["coding_fundamentals", "system_design", "leadership", "problem_solving"]
if competency in core_competencies:
return "high"
elif question_type in ["coding", "design"] or difficulty == "hard":
return "high"
elif difficulty == "easy":
return "medium"
else:
return "medium"
def _generate_follow_up_probes(self, questions: List[Dict[str, Any]]) -> Dict[str, List[str]]:
"""Generate follow-up probes for each question."""
probes = {}
for i, question in enumerate(questions, 1):
question_key = f"question_{i}"
competency = question.get("competency", "")
# Get competency-specific follow-ups
if competency in self.follow_up_strategies:
competency_probes = self.follow_up_strategies[competency].copy()
else:
competency_probes = [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
]
# Add question-type specific probes
question_type = question.get("type", "")
if question_type == "coding":
competency_probes.extend([
"How would you test this solution?",
"What's the time and space complexity of your approach?",
"Can you think of any optimizations?"
])
elif question_type == "behavioral":
competency_probes.extend([
"What did you learn from this experience?",
"How did others react to your approach?",
"What metrics did you use to measure success?"
])
elif question_type == "design":
competency_probes.extend([
"How would you handle failure scenarios?",
"What monitoring would you implement?",
"How would this scale to 10x the load?"
])
probes[question_key] = competency_probes[:5] # Limit to 5 follow-ups
return probes
def _create_calibration_examples(self, sample_questions: List[Dict[str, Any]]) -> Dict[str, Dict]:
"""Create calibration examples with poor/good/great answers."""
examples = {}
for i, question in enumerate(sample_questions, 1):
question_key = f"question_{i}"
examples[question_key] = {
"question": question["question"],
"competency": question["competency"],
"sample_answers": {
"poor_answer": self._generate_sample_answer(question, "poor"),
"good_answer": self._generate_sample_answer(question, "good"),
"great_answer": self._generate_sample_answer(question, "great")
},
"scoring_rationale": self._generate_scoring_rationale(question)
}
return examples
def _generate_sample_answer(self, question: Dict[str, Any], quality: str) -> Dict[str, str]:
"""Generate sample answers of different quality levels."""
competency = question.get("competency", "")
question_type = question.get("type", "")
if quality == "poor":
return {
"answer": f"Sample poor answer for {competency} question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": ["Vague response", "Limited evidence of competency", "Poor structure"]
}
elif quality == "good":
return {
"answer": f"Sample good answer for {competency} question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": ["Clear structure", "Demonstrates competency", "Adequate detail"]
}
else: # great
return {
"answer": f"Sample excellent answer for {competency} question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": ["Exceptional detail", "Strong evidence", "Strategic thinking", "Goes beyond requirements"]
}
def _generate_scoring_rationale(self, question: Dict[str, Any]) -> Dict[str, str]:
"""Generate rationale for scoring this question."""
competency = question.get("competency", "")
return {
"key_indicators": f"Look for evidence of {competency.replace('_', ' ')} competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
def _generate_usage_guidelines(self, role_key: str, level_key: str) -> Dict[str, Any]:
"""Generate usage guidelines for the question bank."""
return {
"interview_flow": {
"warm_up": "Start with 1-2 easier questions to build rapport",
"core_assessment": "Focus majority of time on core competency questions",
"closing": "End with questions about candidate's questions/interests"
},
"time_management": {
"technical_questions": "Allow extra time for coding/design questions",
"behavioral_questions": "Keep to time limits but allow for follow-ups",
"total_recommendation": "45-75 minutes per interview round"
},
"question_selection": {
"variety": "Mix question types within each competency area",
"difficulty": "Adjust based on candidate responses and energy",
"customization": "Adapt questions based on candidate's background"
},
"common_mistakes": [
"Don't ask all questions mechanically",
"Don't skip follow-up questions",
"Don't forget to assess cultural fit alongside competencies",
"Don't let one strong/weak area bias overall assessment"
],
"calibration_reminders": [
"Compare against role standard, not other candidates",
"Focus on evidence demonstrated, not potential",
"Consider level-appropriate expectations",
"Document specific examples in feedback"
]
}
def format_human_readable(question_bank: Dict[str, Any]) -> str:
"""Format question bank data in human-readable format."""
output = []
# Header
output.append(f"Interview Question Bank: {question_bank['role']} ({question_bank['level'].title()} Level)")
output.append("=" * 70)
output.append(f"Generated: {question_bank['generated_at']}")
output.append(f"Total Questions: {question_bank['total_questions']}")
output.append(f"Question Types: {', '.join(question_bank['question_types'])}")
output.append(f"Target Competencies: {', '.join(question_bank['competencies'])}")
output.append("")
# Questions
output.append("INTERVIEW QUESTIONS")
output.append("-" * 50)
for i, question in enumerate(question_bank['questions'], 1):
output.append(f"\n{i}. {question['question']}")
output.append(f" Competency: {question['competency'].replace('_', ' ').title()}")
output.append(f" Type: {question.get('type', 'N/A').title()}")
if 'time_limit' in question:
output.append(f" Time Limit: {question['time_limit']} minutes")
if 'focus_areas' in question:
output.append(f" Focus Areas: {', '.join(question['focus_areas'])}")
# Scoring Guidelines
output.append("\n\nSCORING RUBRICS")
output.append("-" * 50)
# Show sample scoring criteria
if question_bank['scoring_rubrics']:
first_question = list(question_bank['scoring_rubrics'].keys())[0]
sample_rubric = question_bank['scoring_rubrics'][first_question]
output.append(f"Sample Scoring Criteria ({sample_rubric['type']} questions):")
for criterion, scores in sample_rubric['scoring_criteria'].items():
output.append(f"\n{criterion.replace('_', ' ').title()}:")
for score, description in scores.items():
output.append(f" {score}: {description}")
# Follow-up Probes
output.append("\n\nFOLLOW-UP PROBE EXAMPLES")
output.append("-" * 50)
if question_bank['follow_up_probes']:
first_question = list(question_bank['follow_up_probes'].keys())[0]
sample_probes = question_bank['follow_up_probes'][first_question]
output.append("Sample follow-up questions:")
for probe in sample_probes[:3]: # Show first 3
output.append(f" • {probe}")
# Usage Guidelines
output.append("\n\nUSAGE GUIDELINES")
output.append("-" * 50)
guidelines = question_bank['usage_guidelines']
output.append("Interview Flow:")
for phase, description in guidelines['interview_flow'].items():
output.append(f" • {phase.replace('_', ' ').title()}: {description}")
output.append("\nTime Management:")
for aspect, recommendation in guidelines['time_management'].items():
output.append(f" • {aspect.replace('_', ' ').title()}: {recommendation}")
output.append("\nCommon Mistakes to Avoid:")
for mistake in guidelines['common_mistakes'][:3]: # Show first 3
output.append(f" • {mistake}")
# Calibration Examples (if available)
if question_bank['calibration_examples']:
output.append("\n\nCALIBRATION EXAMPLES")
output.append("-" * 50)
first_example = list(question_bank['calibration_examples'].values())[0]
output.append(f"Question: {first_example['question']}")
output.append("\nSample Answer Quality Levels:")
for quality, details in first_example['sample_answers'].items():
output.append(f" {quality.replace('_', ' ').title()} (Score {details['score']}):")
if 'issues' in details:
output.append(f" Issues: {', '.join(details['issues'])}")
if 'strengths' in details:
output.append(f" Strengths: {', '.join(details['strengths'])}")
return "\n".join(output)
def main():
parser = argparse.ArgumentParser(description="Generate comprehensive interview question banks with scoring criteria")
parser.add_argument("--role", type=str, help="Job role title (e.g., 'Frontend Engineer')")
parser.add_argument("--level", type=str, default="senior", help="Experience level (junior, mid, senior, staff, principal)")
parser.add_argument("--competencies", type=str, help="Comma-separated list of competencies to focus on")
parser.add_argument("--question-types", type=str, help="Comma-separated list of question types (technical, behavioral, situational)")
parser.add_argument("--num-questions", type=int, default=20, help="Number of questions to generate")
parser.add_argument("--input", type=str, help="Input JSON file with role requirements")
parser.add_argument("--output", type=str, help="Output directory or file path")
parser.add_argument("--format", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
generator = QuestionBankGenerator()
# Handle input
if args.input:
try:
with open(args.input, 'r') as f:
role_data = json.load(f)
role = role_data.get('role') or role_data.get('title', '')
level = role_data.get('level', 'senior')
competencies = role_data.get('competencies')
question_types = role_data.get('question_types')
num_questions = role_data.get('num_questions', 20)
except Exception as e:
print(f"Error reading input file: {e}")
sys.exit(1)
else:
if not args.role:
print("Error: --role is required when not using --input")
sys.exit(1)
role = args.role
level = args.level
competencies = args.competencies.split(',') if args.competencies else None
question_types = args.question_types.split(',') if args.question_types else None
num_questions = args.num_questions
# Generate question bank
try:
question_bank = generator.generate_question_bank(
role=role,
level=level,
competencies=competencies,
question_types=question_types,
num_questions=num_questions
)
# Handle output
if args.output:
output_path = args.output
if os.path.isdir(output_path):
safe_role = "".join(c for c in role.lower() if c.isalnum() or c in (' ', '-', '_')).replace(' ', '_')
base_filename = f"{safe_role}_{level}_questions"
json_path = os.path.join(output_path, f"{base_filename}.json")
text_path = os.path.join(output_path, f"{base_filename}.txt")
else:
json_path = output_path if output_path.endswith('.json') else f"{output_path}.json"
text_path = output_path.replace('.json', '.txt') if output_path.endswith('.json') else f"{output_path}.txt"
else:
safe_role = "".join(c for c in role.lower() if c.isalnum() or c in (' ', '-', '_')).replace(' ', '_')
base_filename = f"{safe_role}_{level}_questions"
json_path = f"{base_filename}.json"
text_path = f"{base_filename}.txt"
# Write outputs
if args.format in ["json", "both"]:
with open(json_path, 'w') as f:
json.dump(question_bank, f, indent=2, default=str)
print(f"JSON output written to: {json_path}")
if args.format in ["text", "both"]:
with open(text_path, 'w') as f:
f.write(format_human_readable(question_bank))
print(f"Text output written to: {text_path}")
# Print summary
print(f"\nQuestion Bank Summary:")
print(f"Role: {question_bank['role']} ({question_bank['level'].title()})")
print(f"Total Questions: {question_bank['total_questions']}")
print(f"Competencies Covered: {len(question_bank['competencies'])}")
print(f"Question Types: {', '.join(question_bank['question_types'])}")
except Exception as e:
print(f"Error generating question bank: {e}")
sys.exit(1)
if __name__ == "__main__":
main()
FILE:README.md
# Interview System Designer
A comprehensive toolkit for designing, optimizing, and calibrating interview processes. This skill provides tools to create role-specific interview loops, generate competency-based question banks, and analyze hiring data for bias and calibration issues.
## Overview
The Interview System Designer skill includes three powerful Python tools and comprehensive reference materials to help you build fair, effective, and scalable hiring processes:
1. **Interview Loop Designer** - Generate calibrated interview loops for any role and level
2. **Question Bank Generator** - Create competency-based interview questions with scoring rubrics
3. **Hiring Calibrator** - Analyze interview data to detect bias and calibration issues
## Tools
### 1. Interview Loop Designer (`loop_designer.py`)
Generates complete interview loops tailored to specific roles, levels, and teams.
**Features:**
- Role-specific competency mapping (SWE, PM, Designer, Data, DevOps, Leadership)
- Level-appropriate interview rounds (junior through principal)
- Optimized scheduling and time allocation
- Interviewer skill requirements
- Standardized scorecard templates
**Usage:**
```bash
# Basic usage
python3 loop_designer.py --role "Senior Software Engineer" --level senior
# With team and custom competencies
python3 loop_designer.py --role "Product Manager" --level mid --team growth --competencies leadership,strategy,analytics
# Using JSON input file
python3 loop_designer.py --input assets/sample_role_definitions.json --output loops/
# Specify output format
python3 loop_designer.py --role "Staff Data Scientist" --level staff --format json --output data_scientist_loop.json
```
**Input Options:**
- `--role`: Job role title (e.g., "Senior Software Engineer", "Product Manager")
- `--level`: Experience level (junior, mid, senior, staff, principal)
- `--team`: Team or department (optional)
- `--competencies`: Comma-separated list of specific competencies to focus on
- `--input`: JSON file with role definition
- `--output`: Output directory or file path
- `--format`: Output format (json, text, both) - default: both
**Example Output:**
```
Interview Loop Design for Senior Software Engineer (Senior Level)
============================================================
Total Duration: 300 minutes (5h 0m)
Total Rounds: 5
INTERVIEW ROUNDS
----------------------------------------
Round 1: Technical Phone Screen
Duration: 45 minutes
Format: Virtual
Focus Areas: Coding Fundamentals, Problem Solving
Round 2: System Design
Duration: 75 minutes
Format: Collaborative Whitboard
Focus Areas: System Thinking, Architectural Reasoning
...
```
### 2. Question Bank Generator (`question_bank_generator.py`)
Creates comprehensive interview question banks organized by competency area.
**Features:**
- Competency-based question organization
- Level-appropriate difficulty progression
- Multiple question types (technical, behavioral, situational)
- Detailed scoring rubrics with calibration examples
- Follow-up probes and conversation guides
**Usage:**
```bash
# Generate questions for specific competencies
python3 question_bank_generator.py --role "Frontend Engineer" --competencies react,typescript,system-design
# Create behavioral question bank
python3 question_bank_generator.py --role "Product Manager" --question-types behavioral,leadership --num-questions 15
# Generate questions for multiple levels
python3 question_bank_generator.py --role "DevOps Engineer" --levels junior,mid,senior --output questions/
```
**Input Options:**
- `--role`: Job role title
- `--level`: Experience level (default: senior)
- `--competencies`: Comma-separated list of competencies to focus on
- `--question-types`: Types to include (technical, behavioral, situational)
- `--num-questions`: Number of questions to generate (default: 20)
- `--input`: JSON file with role requirements
- `--output`: Output directory or file path
- `--format`: Output format (json, text, both) - default: both
**Question Types:**
- **Technical**: Coding problems, system design, domain-specific challenges
- **Behavioral**: STAR method questions focusing on past experiences
- **Situational**: Hypothetical scenarios testing decision-making
### 3. Hiring Calibrator (`hiring_calibrator.py`)
Analyzes interview scores to detect bias, calibration issues, and provides recommendations.
**Features:**
- Statistical bias detection across demographics
- Interviewer calibration analysis
- Score distribution and trending analysis
- Specific coaching recommendations
- Comprehensive reporting with actionable insights
**Usage:**
```bash
# Comprehensive analysis
python3 hiring_calibrator.py --input assets/sample_interview_results.json --analysis-type comprehensive
# Focus on specific areas
python3 hiring_calibrator.py --input interview_data.json --analysis-type bias --competencies technical,leadership
# Trend analysis over time
python3 hiring_calibrator.py --input historical_data.json --trend-analysis --period quarterly
```
**Input Options:**
- `--input`: JSON file with interview results data (required)
- `--analysis-type`: Type of analysis (comprehensive, bias, calibration, interviewer, scoring)
- `--competencies`: Comma-separated list of competencies to focus on
- `--trend-analysis`: Enable trend analysis over time
- `--period`: Time period for trends (daily, weekly, monthly, quarterly)
- `--output`: Output file path
- `--format`: Output format (json, text, both) - default: both
**Analysis Types:**
- **Comprehensive**: Full analysis including bias, calibration, and recommendations
- **Bias**: Focus on demographic and interviewer bias patterns
- **Calibration**: Interviewer consistency and agreement analysis
- **Interviewer**: Individual interviewer performance and coaching needs
- **Scoring**: Score distribution and pattern analysis
## Data Formats
### Role Definition Input (JSON)
```json
{
"role": "Senior Software Engineer",
"level": "senior",
"team": "platform",
"competencies": ["system_design", "technical_leadership", "mentoring"],
"requirements": {
"years_experience": "5-8",
"technical_skills": ["Python", "AWS", "Kubernetes"],
"leadership_experience": true
}
}
```
### Interview Results Input (JSON)
```json
[
{
"candidate_id": "candidate_001",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-15T09:00:00Z",
"scores": {
"coding_fundamentals": 3.5,
"system_design": 4.0,
"technical_leadership": 3.0,
"communication": 3.5
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "asian",
"years_experience": 6
}
]
```
## Reference Materials
### Competency Matrix Templates (`references/competency_matrix_templates.md`)
- Comprehensive competency matrices for all engineering roles
- Level-specific expectations (junior through principal)
- Assessment criteria and growth paths
- Customization guidelines for different company stages and industries
### Bias Mitigation Checklist (`references/bias_mitigation_checklist.md`)
- Pre-interview preparation checklist
- Interview process bias prevention strategies
- Real-time bias interruption techniques
- Legal compliance reminders
- Emergency response protocols
### Debrief Facilitation Guide (`references/debrief_facilitation_guide.md`)
- Structured debrief meeting frameworks
- Evidence-based discussion techniques
- Bias interruption strategies
- Decision documentation standards
- Common challenges and solutions
## Sample Data
The `assets/` directory contains sample data for testing:
- `sample_role_definitions.json`: Example role definitions for various positions
- `sample_interview_results.json`: Sample interview data with multiple candidates and interviewers
## Expected Outputs
The `expected_outputs/` directory contains examples of tool outputs:
- Interview loop designs in both JSON and human-readable formats
- Question banks with scoring rubrics and calibration examples
- Calibration analysis reports with bias detection and recommendations
## Best Practices
### Interview Loop Design
1. **Competency Focus**: Align interview rounds with role-critical competencies
2. **Level Calibration**: Adjust expectations and question difficulty based on experience level
3. **Time Optimization**: Balance thoroughness with candidate experience
4. **Interviewer Training**: Ensure interviewers are qualified and calibrated
### Question Bank Development
1. **Evidence-Based**: Focus on observable behaviors and concrete examples
2. **Bias Mitigation**: Use structured questions that minimize subjective interpretation
3. **Calibration**: Include examples of different quality responses for consistency
4. **Continuous Improvement**: Regularly update questions based on predictive validity
### Calibration Analysis
1. **Regular Monitoring**: Analyze hiring data quarterly for bias patterns
2. **Prompt Action**: Address calibration issues immediately with targeted coaching
3. **Data Quality**: Ensure complete and consistent data collection
4. **Legal Compliance**: Monitor for discriminatory patterns and document corrections
## Installation & Setup
No external dependencies required - uses Python 3 standard library only.
```bash
# Clone or download the skill directory
cd interview-system-designer/
# Make scripts executable (optional)
chmod +x *.py
# Test with sample data
python3 loop_designer.py --role "Senior Software Engineer" --level senior
python3 question_bank_generator.py --role "Product Manager" --level mid
python3 hiring_calibrator.py --input assets/sample_interview_results.json
```
## Integration
### With Existing Systems
- **ATS Integration**: Export interview loops as structured data for applicant tracking systems
- **Calendar Systems**: Use scheduling outputs to auto-create interview blocks
- **HR Analytics**: Import calibration reports into broader diversity and inclusion dashboards
### Custom Workflows
- **Batch Processing**: Process multiple roles or historical data sets
- **Automated Reporting**: Schedule regular calibration analysis
- **Custom Competencies**: Extend frameworks with company-specific competencies
## Troubleshooting
### Common Issues
**"Role not found" errors:**
- The tool will map common variations (engineer → software_engineer)
- For custom roles, use the closest standard role and specify custom competencies
**"Insufficient data" errors:**
- Minimum 5 interviews required for statistical analysis
- Ensure interview data includes required fields (candidate_id, interviewer_id, scores, date)
**Missing output files:**
- Check file permissions in output directory
- Ensure adequate disk space
- Verify JSON input file format is valid
### Performance Considerations
- Interview loop generation: < 1 second
- Question bank generation: 1-3 seconds for 20 questions
- Calibration analysis: 1-5 seconds for 50 interviews, scales linearly
## Contributing
To extend this skill:
1. **New Roles**: Add competency frameworks in `_init_competency_frameworks()`
2. **New Question Types**: Extend question templates in respective generators
3. **New Analysis Types**: Add analysis methods to hiring calibrator
4. **Custom Outputs**: Modify formatting functions for different output needs
## License & Usage
This skill is designed for internal company use in hiring process optimization. All bias detection and mitigation features should be reviewed with legal counsel to ensure compliance with local employment laws.
For questions or support, refer to the comprehensive documentation in each script's docstring and the reference materials provided.
FILE:references/bias_mitigation_checklist.md
# Interview Bias Mitigation Checklist
This comprehensive checklist helps identify, prevent, and mitigate various forms of bias in the interview process. Use this as a systematic guide to ensure fair and equitable hiring practices.
## Pre-Interview Phase
### Job Description & Requirements
- [ ] **Remove unnecessary requirements** that don't directly relate to job performance
- [ ] **Avoid gendered language** (competitive, aggressive vs. collaborative, detail-oriented)
- [ ] **Remove university prestige requirements** unless absolutely necessary for role
- [ ] **Focus on skills and outcomes** rather than years of experience in specific technologies
- [ ] **Use inclusive language** and avoid cultural assumptions
- [ ] **Specify only essential requirements** vs. nice-to-have qualifications
- [ ] **Remove location/commute assumptions** for remote-eligible positions
- [ ] **Review requirements for unconscious bias** (e.g., assuming continuous work history)
### Sourcing & Pipeline
- [ ] **Diversify sourcing channels** beyond traditional networks
- [ ] **Partner with diverse professional organizations** and communities
- [ ] **Use bias-minimizing sourcing tools** and platforms
- [ ] **Track sourcing effectiveness** by demographic groups
- [ ] **Train recruiters on bias awareness** and inclusive outreach
- [ ] **Review referral patterns** for potential network bias
- [ ] **Expand university partnerships** beyond elite institutions
- [ ] **Use structured outreach messages** to reduce individual bias
### Resume Screening
- [ ] **Implement blind resume review** (remove names, photos, university names initially)
- [ ] **Use standardized screening criteria** applied consistently
- [ ] **Multiple screeners for each resume** with independent scoring
- [ ] **Focus on relevant skills and achievements** over pedigree indicators
- [ ] **Avoid assumptions about career gaps** or non-traditional backgrounds
- [ ] **Consider alternative paths to skills** (bootcamps, self-taught, career changes)
- [ ] **Track screening pass rates** by demographic groups
- [ ] **Regular screener calibration sessions** on bias awareness
## Interview Panel Composition
### Diversity Requirements
- [ ] **Ensure diverse interview panels** (gender, ethnicity, seniority levels)
- [ ] **Include at least one underrepresented interviewer** when possible
- [ ] **Rotate panel assignments** to prevent bias patterns
- [ ] **Balance seniority levels** on panels (not all senior or all junior)
- [ ] **Include cross-functional perspectives** when relevant
- [ ] **Avoid panels of only one demographic group** when possible
- [ ] **Consider panel member unconscious bias training** status
- [ ] **Document panel composition rationale** for future review
### Interviewer Selection
- [ ] **Choose interviewers based on relevant competency assessment ability**
- [ ] **Ensure interviewers have completed bias training** within last 12 months
- [ ] **Select interviewers with consistent calibration history**
- [ ] **Avoid interviewers with known bias patterns** (flagged in previous analyses)
- [ ] **Include at least one interviewer familiar with candidate's background type**
- [ ] **Balance perspectives** (technical depth, cultural fit, growth potential)
- [ ] **Consider interviewer availability for proper preparation time**
- [ ] **Ensure interviewers understand role requirements and standards**
## Interview Process Design
### Question Standardization
- [ ] **Use standardized question sets** for each competency area
- [ ] **Develop questions that assess skills, not culture fit stereotypes**
- [ ] **Avoid questions about personal background** unless directly job-relevant
- [ ] **Remove questions that could reveal protected characteristics**
- [ ] **Focus on behavioral examples** using STAR method
- [ ] **Include scenario-based questions** with clear evaluation criteria
- [ ] **Test questions for potential bias** with diverse interviewers
- [ ] **Regularly update question bank** based on effectiveness data
### Structured Interview Protocol
- [ ] **Define clear time allocations** for each question/section
- [ ] **Establish consistent interview flow** across all candidates
- [ ] **Create standardized intro/outro** processes
- [ ] **Use identical technical setup and tools** for all candidates
- [ ] **Provide same background information** to all interviewers
- [ ] **Standardize note-taking format** and requirements
- [ ] **Define clear handoff procedures** between interviewers
- [ ] **Document any deviations** from standard protocol
### Accommodation Preparation
- [ ] **Proactively offer accommodations** without requiring disclosure
- [ ] **Provide multiple interview format options** (phone, video, in-person)
- [ ] **Ensure accessibility of interview locations and tools**
- [ ] **Allow extended time** when requested or needed
- [ ] **Provide materials in advance** when helpful
- [ ] **Train interviewers on accommodation protocols**
- [ ] **Test all technology** for accessibility compliance
- [ ] **Have backup plans** for technical issues
## During the Interview
### Interviewer Behavior
- [ ] **Use welcoming, professional tone** with all candidates
- [ ] **Avoid assumptions based on appearance or background**
- [ ] **Give equal encouragement and support** to all candidates
- [ ] **Allow equal time for candidate questions**
- [ ] **Avoid leading questions** that suggest desired answers
- [ ] **Listen actively** without interrupting unnecessarily
- [ ] **Take detailed notes** focusing on responses, not impressions
- [ ] **Avoid small talk** that could reveal irrelevant personal information
### Question Delivery
- [ ] **Ask questions as written** without improvisation that could introduce bias
- [ ] **Provide equal clarification** when candidates ask for it
- [ ] **Use consistent follow-up probing** across candidates
- [ ] **Allow reasonable thinking time** before expecting responses
- [ ] **Avoid rephrasing questions** in ways that give hints
- [ ] **Stay focused on defined competencies** being assessed
- [ ] **Give equal encouragement** for elaboration when needed
- [ ] **Maintain professional demeanor** regardless of candidate background
### Real-time Bias Checking
- [ ] **Notice first impressions** but don't let them drive assessment
- [ ] **Question gut reactions** - are they based on competency evidence?
- [ ] **Focus on specific examples** and evidence provided
- [ ] **Avoid pattern matching** to existing successful employees
- [ ] **Notice cultural assumptions** in interpretation of responses
- [ ] **Check for confirmation bias** - seeking evidence to support initial impressions
- [ ] **Consider alternative explanations** for candidate responses
- [ ] **Stay aware of fatigue effects** on judgment throughout the day
## Evaluation & Scoring
### Scoring Consistency
- [ ] **Use defined rubrics consistently** across all candidates
- [ ] **Score immediately after interview** while details are fresh
- [ ] **Focus scoring on demonstrated competencies** not potential or personality
- [ ] **Provide specific evidence** for each score given
- [ ] **Avoid comparative scoring** (comparing candidates to each other)
- [ ] **Use calibrated examples** of each score level
- [ ] **Score independently** before discussing with other interviewers
- [ ] **Document reasoning** for all scores, especially extreme ones (1s and 4s)
### Bias Check Questions
- [ ] **"Would I score this differently if the candidate looked different?"**
- [ ] **"Am I basing this on evidence or assumptions?"**
- [ ] **"Would this response get the same score from a different demographic?"**
- [ ] **"Am I penalizing non-traditional backgrounds or approaches?"**
- [ ] **"Is my scoring consistent with the defined rubric?"**
- [ ] **"Am I letting one strong/weak area bias overall assessment?"**
- [ ] **"Are my cultural assumptions affecting interpretation?"**
- [ ] **"Would I want to work with this person?" (Check if this is biasing assessment)**
### Documentation Requirements
- [ ] **Record specific examples** supporting each competency score
- [ ] **Avoid subjective language** like "seems like," "appears to be"
- [ ] **Focus on observable behaviors** and concrete responses
- [ ] **Note exact quotes** when relevant to assessment
- [ ] **Distinguish between facts and interpretations**
- [ ] **Provide improvement suggestions** that are skill-based, not person-based
- [ ] **Avoid comparative language** to other candidates or employees
- [ ] **Use neutral language** free from cultural assumptions
## Debrief Process
### Structured Discussion
- [ ] **Start with independent score sharing** before discussion
- [ ] **Focus discussion on evidence** not impressions or feelings
- [ ] **Address significant score discrepancies** with evidence review
- [ ] **Challenge biased language** or assumptions in discussion
- [ ] **Ensure all voices are heard** in group decision making
- [ ] **Document reasons for final decision** with specific evidence
- [ ] **Avoid personality-based discussions** ("culture fit" should be evidence-based)
- [ ] **Consider multiple perspectives** on candidate responses
### Decision-Making Process
- [ ] **Use weighted scoring system** based on role requirements
- [ ] **Require minimum scores** in critical competency areas
- [ ] **Avoid veto power** unless based on clear, documented evidence
- [ ] **Consider growth potential** fairly across all candidates
- [ ] **Document dissenting opinions** and reasoning
- [ ] **Use tie-breaking criteria** that are predetermined and fair
- [ ] **Consider additional data collection** if team is split
- [ ] **Make final decision based on role requirements**, not team preferences
### Final Recommendations
- [ ] **Provide specific, actionable feedback** for development areas
- [ ] **Focus recommendations on skills and competencies**
- [ ] **Avoid language that could reflect bias** in written feedback
- [ ] **Consider onboarding needs** based on actual skill gaps, not assumptions
- [ ] **Provide coaching recommendations** that are evidence-based
- [ ] **Avoid personal judgments** about candidate character or personality
- [ ] **Make hiring recommendation** based solely on job-relevant criteria
- [ ] **Document any concerns** with specific, observable evidence
## Post-Interview Monitoring
### Data Collection
- [ ] **Track interviewer scoring patterns** for consistency analysis
- [ ] **Monitor pass rates** by demographic groups
- [ ] **Collect candidate experience feedback** on interview fairness
- [ ] **Analyze score distributions** for potential bias indicators
- [ ] **Track time-to-decision** across different candidate types
- [ ] **Monitor offer acceptance rates** by demographics
- [ ] **Collect new hire performance data** for process validation
- [ ] **Document any bias incidents** or concerns raised
### Regular Analysis
- [ ] **Conduct quarterly bias audits** of interview data
- [ ] **Review interviewer calibration** and identify outliers
- [ ] **Analyze demographic trends** in hiring outcomes
- [ ] **Compare candidate experience surveys** across groups
- [ ] **Track correlation between interview scores and job performance**
- [ ] **Review and update bias mitigation strategies** based on data
- [ ] **Share findings with interview teams** for continuous improvement
- [ ] **Update training programs** based on identified bias patterns
## Bias Types to Watch For
### Affinity Bias
- **Definition**: Favoring candidates similar to yourself
- **Watch for**: Over-positive response to shared backgrounds, interests, or experiences
- **Mitigation**: Focus on job-relevant competencies, diversify interview panels
### Halo/Horn Effect
- **Definition**: One positive/negative trait influencing overall assessment
- **Watch for**: Strong performance in one area affecting scores in unrelated areas
- **Mitigation**: Score each competency independently, use structured evaluation
### Confirmation Bias
- **Definition**: Seeking information that confirms initial impressions
- **Watch for**: Asking follow-ups that lead candidate toward expected responses
- **Mitigation**: Use standardized questions, consider alternative interpretations
### Attribution Bias
- **Definition**: Attributing success/failure to different causes based on candidate demographics
- **Watch for**: Assuming women are "lucky" vs. men are "skilled" for same achievements
- **Mitigation**: Focus on candidate's role in achievements, avoid assumptions
### Cultural Bias
- **Definition**: Judging candidates based on cultural differences rather than job performance
- **Watch for**: Penalizing communication styles, work approaches, or values that differ from team norm
- **Mitigation**: Define job-relevant criteria clearly, consider diverse perspectives valuable
### Educational Bias
- **Definition**: Over-weighting prestigious educational credentials
- **Watch for**: Assuming higher capability based on school rank rather than demonstrated skills
- **Mitigation**: Focus on skills demonstration, consider alternative learning paths
### Experience Bias
- **Definition**: Requiring specific company or industry experience unnecessarily
- **Watch for**: Discounting transferable skills from different industries or company sizes
- **Mitigation**: Define core skills needed, assess adaptability and learning ability
## Emergency Bias Response Protocol
### During Interview
1. **Pause the interview** if significant bias is observed
2. **Privately address** bias with interviewer if possible
3. **Document the incident** for review
4. **Continue with fair assessment** of candidate
5. **Flag for debrief discussion** if interview continues
### Post-Interview
1. **Report bias incidents** to hiring manager/HR immediately
2. **Document specific behaviors** observed
3. **Consider additional interviewer** for second opinion
4. **Review candidate assessment** for bias impact
5. **Implement corrective actions** for future interviews
### Interviewer Coaching
1. **Provide immediate feedback** on bias observed
2. **Schedule bias training refresher** if needed
3. **Monitor future interviews** for improvement
4. **Consider removing from interview rotation** if bias persists
5. **Document coaching provided** for performance management
## Legal Compliance Reminders
### Protected Characteristics
- Age, race, color, religion, sex, national origin, disability status, veteran status
- Pregnancy, genetic information, sexual orientation, gender identity
- Any other characteristics protected by local/state/federal law
### Prohibited Questions
- Questions about family planning, marital status, pregnancy
- Age-related questions (unless BFOQ)
- Religious or political affiliations
- Disability status (unless voluntary disclosure for accommodation)
- Arrest records (without conviction relevance)
- Financial status or credit (unless job-relevant)
### Documentation Requirements
- Keep all interview materials for required retention period
- Ensure consistent documentation across all candidates
- Avoid documenting protected characteristic observations
- Focus documentation on job-relevant observations only
## Training & Certification
### Required Training Topics
- Unconscious bias awareness and mitigation
- Structured interviewing techniques
- Legal compliance in hiring
- Company-specific bias mitigation protocols
- Role-specific competency assessment
- Accommodation and accessibility requirements
### Ongoing Development
- Annual bias training refresher
- Quarterly calibration sessions
- Regular updates on legal requirements
- Peer feedback and coaching
- Industry best practice updates
- Data-driven process improvements
This checklist should be reviewed and updated regularly based on legal requirements, industry best practices, and internal bias analysis results.
FILE:references/competency_matrix_templates.md
# Competency Matrix Templates
This document provides comprehensive competency matrix templates for different engineering roles and levels. Use these matrices to design role-specific interview loops and evaluation criteria.
## Software Engineering Competency Matrix
### Technical Competencies
| Competency | Junior (L1-L2) | Mid (L3-L4) | Senior (L5-L6) | Staff+ (L7+) |
|------------|----------------|-------------|----------------|--------------|
| **Coding & Algorithms** | Basic data structures, simple algorithms, language syntax | Advanced algorithms, complexity analysis, optimization | Complex problem solving, algorithm design, performance tuning | Architecture-level algorithmic decisions, novel approach design |
| **System Design** | Component interactions, basic scalability concepts | Service design, database modeling, API design | Distributed systems, scalability patterns, trade-off analysis | Large-scale architecture, cross-system design, technology strategy |
| **Code Quality** | Readable code, basic testing, follows conventions | Maintainable code, comprehensive testing, design patterns | Code reviews, quality standards, refactoring leadership | Engineering standards, quality culture, technical debt management |
| **Debugging & Problem Solving** | Basic debugging, structured problem approach | Complex debugging, root cause analysis, performance issues | System-wide debugging, production issues, incident response | Cross-system troubleshooting, preventive measures, tooling design |
| **Domain Knowledge** | Learning role-specific technologies | Proficiency in domain tools/frameworks | Deep domain expertise, technology evaluation | Domain leadership, technology roadmap, innovation |
### Behavioral Competencies
| Competency | Junior (L1-L2) | Mid (L3-L4) | Senior (L5-L6) | Staff+ (L7+) |
|------------|----------------|-------------|----------------|--------------|
| **Communication** | Clear status updates, asks good questions | Technical explanations, stakeholder updates | Cross-functional communication, technical writing | Executive communication, external representation, thought leadership |
| **Collaboration** | Team participation, code reviews | Cross-team projects, knowledge sharing | Team leadership, conflict resolution | Cross-org collaboration, culture building, strategic partnerships |
| **Leadership & Influence** | Peer mentoring, positive attitude | Junior mentoring, project ownership | Team guidance, technical decisions, hiring | Org-wide influence, vision setting, culture change |
| **Growth & Learning** | Skill development, feedback receptivity | Proactive learning, teaching others | Continuous improvement, trend awareness | Learning culture, industry leadership, innovation adoption |
| **Ownership & Initiative** | Task completion, quality focus | Project ownership, process improvement | Feature/service ownership, strategic thinking | Product/platform ownership, business impact, market influence |
## Product Management Competency Matrix
### Product Competencies
| Competency | Associate PM (L1-L2) | PM (L3-L4) | Senior PM (L5-L6) | Principal PM (L7+) |
|------------|---------------------|------------|-------------------|-------------------|
| **Product Strategy** | Feature requirements, user stories | Product roadmaps, market analysis | Business strategy, competitive positioning | Portfolio strategy, market creation, platform vision |
| **User Research & Analytics** | Basic user interviews, metrics tracking | Research design, data interpretation | Research strategy, advanced analytics | Research culture, measurement frameworks, insight generation |
| **Technical Understanding** | Basic tech concepts, API awareness | System architecture, technical trade-offs | Technical strategy, platform decisions | Technology vision, architectural influence, innovation leadership |
| **Execution & Process** | Feature delivery, stakeholder coordination | Project management, cross-functional leadership | Process optimization, team scaling | Operational excellence, org design, strategic execution |
| **Business Acumen** | Revenue awareness, customer understanding | P&L understanding, business case development | Business strategy, market dynamics | Corporate strategy, board communication, investor relations |
### Leadership Competencies
| Competency | Associate PM (L1-L2) | PM (L3-L4) | Senior PM (L5-L6) | Principal PM (L7+) |
|------------|---------------------|------------|-------------------|-------------------|
| **Stakeholder Management** | Team collaboration, clear communication | Cross-functional alignment, expectation management | Executive communication, influence without authority | Board interaction, external partnerships, industry influence |
| **Team Development** | Peer learning, feedback sharing | Junior mentoring, knowledge transfer | Team building, hiring, performance management | Talent development, culture building, org leadership |
| **Decision Making** | Data-driven decisions, priority setting | Complex trade-offs, strategic choices | Ambiguous situations, high-stakes decisions | Strategic vision, transformational decisions, risk management |
| **Innovation & Vision** | Creative problem solving, user empathy | Market opportunity identification, feature innovation | Product vision, market strategy | Industry vision, disruptive thinking, platform creation |
## Design Competency Matrix
### Design Competencies
| Competency | Junior Designer (L1-L2) | Mid Designer (L3-L4) | Senior Designer (L5-L6) | Principal Designer (L7+) |
|------------|-------------------------|---------------------|-------------------------|-------------------------|
| **Visual Design** | UI components, typography, color theory | Design systems, visual hierarchy | Brand integration, advanced layouts | Visual strategy, brand evolution, design innovation |
| **User Experience** | User flows, wireframing, prototyping | Interaction design, usability testing | Experience strategy, journey mapping | UX vision, service design, behavioral insights |
| **Research & Validation** | User interviews, usability tests | Research planning, data synthesis | Research strategy, methodology design | Research culture, insight frameworks, market research |
| **Design Systems** | Component usage, style guides | System contribution, pattern creation | System architecture, governance | System strategy, scalable design, platform thinking |
| **Tools & Craft** | Design software proficiency, asset creation | Advanced techniques, workflow optimization | Tool evaluation, process design | Technology integration, future tooling, craft evolution |
### Collaboration Competencies
| Competency | Junior Designer (L1-L2) | Mid Designer (L3-L4) | Senior Designer (L5-L6) | Principal Designer (L7+) |
|------------|-------------------------|---------------------|-------------------------|-------------------------|
| **Cross-functional Partnership** | Engineering collaboration, handoff quality | Product partnership, stakeholder alignment | Leadership collaboration, strategic alignment | Executive partnership, business strategy integration |
| **Communication & Advocacy** | Design rationale, feedback integration | Design presentations, user advocacy | Executive communication, design thinking evangelism | Industry thought leadership, external representation |
| **Mentorship & Growth** | Peer learning, skill sharing | Junior mentoring, critique facilitation | Team development, hiring, career guidance | Design culture, talent strategy, industry leadership |
| **Business Impact** | User-centered thinking, design quality | Feature success, user satisfaction | Business metrics, strategic impact | Market influence, competitive advantage, innovation leadership |
## Data Science Competency Matrix
### Technical Competencies
| Competency | Junior DS (L1-L2) | Mid DS (L3-L4) | Senior DS (L5-L6) | Principal DS (L7+) |
|------------|-------------------|----------------|-------------------|-------------------|
| **Statistical Analysis** | Descriptive stats, hypothesis testing | Advanced statistics, experimental design | Causal inference, advanced modeling | Statistical strategy, methodology innovation |
| **Machine Learning** | Basic ML algorithms, model training | Advanced ML, feature engineering | ML systems, model deployment | ML strategy, AI platform, research direction |
| **Data Engineering** | SQL, basic ETL, data cleaning | Pipeline design, data modeling | Platform architecture, scalable systems | Data strategy, infrastructure vision, governance |
| **Programming & Tools** | Python/R proficiency, visualization | Advanced programming, tool integration | Software engineering, system design | Technology strategy, platform development, innovation |
| **Domain Expertise** | Business understanding, metric interpretation | Domain modeling, insight generation | Strategic analysis, business integration | Market expertise, competitive intelligence, thought leadership |
### Impact & Leadership Competencies
| Competency | Junior DS (L1-L2) | Mid DS (L3-L4) | Senior DS (L5-L6) | Principal DS (L7+) |
|------------|-------------------|----------------|-------------------|-------------------|
| **Business Impact** | Metric improvement, insight delivery | Project leadership, business case development | Strategic initiatives, P&L impact | Business transformation, market advantage, innovation |
| **Communication** | Technical reporting, visualization | Stakeholder presentations, executive briefings | Board communication, external representation | Industry leadership, thought leadership, market influence |
| **Team Leadership** | Peer collaboration, knowledge sharing | Junior mentoring, project management | Team building, hiring, culture development | Organizational leadership, talent strategy, vision setting |
| **Innovation & Research** | Algorithm implementation, experimentation | Research projects, publication | Research strategy, academic partnerships | Research vision, industry influence, breakthrough innovation |
## DevOps Engineering Competency Matrix
### Technical Competencies
| Competency | Junior DevOps (L1-L2) | Mid DevOps (L3-L4) | Senior DevOps (L5-L6) | Principal DevOps (L7+) |
|------------|----------------------|-------------------|----------------------|----------------------|
| **Infrastructure** | Basic cloud services, server management | Infrastructure automation, containerization | Platform architecture, multi-cloud strategy | Infrastructure vision, emerging technologies, industry standards |
| **CI/CD & Automation** | Pipeline basics, script writing | Advanced pipelines, deployment automation | Platform design, workflow optimization | Automation strategy, developer experience, productivity platforms |
| **Monitoring & Observability** | Basic monitoring, log analysis | Advanced monitoring, alerting systems | Observability strategy, SLA/SLI design | Monitoring vision, reliability engineering, performance culture |
| **Security & Compliance** | Security basics, access management | Security automation, compliance frameworks | Security architecture, risk management | Security strategy, governance, industry leadership |
| **Performance & Scalability** | Performance monitoring, basic optimization | Capacity planning, performance tuning | Scalability architecture, cost optimization | Performance strategy, efficiency platforms, innovation |
### Leadership & Impact Competencies
| Competency | Junior DevOps (L1-L2) | Mid DevOps (L3-L4) | Senior DevOps (L5-L6) | Principal DevOps (L7+) |
|------------|----------------------|-------------------|----------------------|----------------------|
| **Developer Experience** | Tool support, documentation | Platform development, self-service tools | Developer productivity, workflow design | Developer platform vision, industry best practices |
| **Incident Management** | Incident response, troubleshooting | Incident coordination, root cause analysis | Incident strategy, prevention systems | Reliability culture, organizational resilience |
| **Team Collaboration** | Cross-team support, knowledge sharing | Process improvement, training delivery | Culture building, practice evangelism | Organizational transformation, industry influence |
| **Strategic Impact** | Operational excellence, cost awareness | Efficiency improvements, platform adoption | Strategic initiatives, business enablement | Technology strategy, competitive advantage, market leadership |
## Engineering Management Competency Matrix
### People Leadership Competencies
| Competency | Manager (L1-L2) | Senior Manager (L3-L4) | Director (L5-L6) | VP+ (L7+) |
|------------|-----------------|------------------------|------------------|----------|
| **Team Building** | Hiring, onboarding, 1:1s | Team culture, performance management | Multi-team coordination, org design | Organizational culture, talent strategy |
| **Performance Management** | Individual development, feedback | Performance systems, coaching | Calibration across teams, promotion standards | Talent development, succession planning |
| **Communication** | Team updates, stakeholder management | Executive communication, cross-functional alignment | Board updates, external communication | Industry representation, thought leadership |
| **Conflict Resolution** | Team conflicts, process improvements | Cross-team issues, organizational friction | Strategic alignment, cultural challenges | Corporate-level conflicts, crisis management |
### Technical Leadership Competencies
| Competency | Manager (L1-L2) | Senior Manager (L3-L4) | Director (L5-L6) | VP+ (L7+) |
|------------|-----------------|------------------------|------------------|----------|
| **Technical Vision** | Team technical decisions, architecture input | Platform strategy, technology choices | Technical roadmap, innovation strategy | Technology vision, industry standards |
| **System Ownership** | Feature/service ownership, quality standards | Platform ownership, scalability planning | System portfolio, technical debt management | Technology strategy, competitive advantage |
| **Process & Practice** | Team processes, development practices | Engineering standards, quality systems | Process innovation, best practices | Engineering culture, industry influence |
| **Technology Strategy** | Tool evaluation, team technology choices | Platform decisions, technical investments | Technology portfolio, strategic architecture | Corporate technology strategy, market leadership |
## Usage Guidelines
### Assessment Approach
1. **Level Calibration**: Use these matrices to calibrate expectations for each level within your organization
2. **Interview Design**: Select competencies most relevant to the specific role and level being hired for
3. **Evaluation Consistency**: Ensure all interviewers understand and apply the same competency standards
4. **Growth Planning**: Use matrices for career development and promotion discussions
### Customization Tips
1. **Industry Adaptation**: Modify competencies based on your industry (fintech, healthcare, etc.)
2. **Company Stage**: Adjust expectations based on startup vs. enterprise environment
3. **Team Needs**: Emphasize competencies most critical for current team challenges
4. **Cultural Fit**: Add company-specific values and cultural competencies
### Common Pitfalls
1. **Unrealistic Expectations**: Don't expect senior-level competencies from junior candidates
2. **One-Size-Fits-All**: Customize competency emphasis based on role requirements
3. **Static Assessment**: Regularly update matrices based on changing business needs
4. **Bias Introduction**: Ensure competencies are measurable and don't introduce unconscious bias
## Matrix Validation Process
### Regular Review Cycle
- **Quarterly**: Review competency relevance and adjust weights
- **Semi-annually**: Update level expectations based on market standards
- **Annually**: Comprehensive review with stakeholder feedback
### Stakeholder Input
- **Hiring Managers**: Validate role-specific competency requirements
- **Current Team Members**: Confirm level expectations match reality
- **Recent Hires**: Gather feedback on assessment accuracy
- **HR Partners**: Ensure legal compliance and bias mitigation
### Continuous Improvement
- **Performance Correlation**: Track new hire performance against competency assessments
- **Market Benchmarking**: Compare standards with industry peers
- **Feedback Integration**: Incorporate interviewer and candidate feedback
- **Bias Monitoring**: Regular analysis of assessment patterns across demographics
FILE:references/debrief_facilitation_guide.md
# Interview Debrief Facilitation Guide
This guide provides a comprehensive framework for conducting effective, unbiased interview debriefs that lead to consistent hiring decisions. Use this to facilitate productive discussions that focus on evidence-based evaluation.
## Pre-Debrief Preparation
### Facilitator Responsibilities
- [ ] **Review all interviewer feedback** before the meeting
- [ ] **Identify significant score discrepancies** that need discussion
- [ ] **Prepare discussion agenda** with time allocations
- [ ] **Gather role requirements** and competency framework
- [ ] **Review any flags or special considerations** noted during interviews
- [ ] **Ensure all required materials** are available (scorecards, rubrics, candidate resume)
- [ ] **Set up meeting logistics** (room, video conference, screen sharing)
- [ ] **Send agenda to participants** 30 minutes before meeting
### Required Materials Checklist
- [ ] Candidate resume and application materials
- [ ] Job description and competency requirements
- [ ] Individual interviewer scorecards
- [ ] Scoring rubrics and competency definitions
- [ ] Interview notes and documentation
- [ ] Any technical assessments or work samples
- [ ] Company hiring standards and calibration examples
- [ ] Bias mitigation reminders and prompts
### Participant Preparation Requirements
- [ ] All interviewers must **complete independent scoring** before debrief
- [ ] **Submit written feedback** with specific evidence for each competency
- [ ] **Review scoring rubrics** to ensure consistent interpretation
- [ ] **Prepare specific examples** to support scoring decisions
- [ ] **Flag any concerns or unusual circumstances** that affected assessment
- [ ] **Avoid discussing candidate** with other interviewers before debrief
- [ ] **Come prepared to defend scores** with concrete evidence
- [ ] **Be ready to adjust scores** based on additional evidence shared
## Debrief Meeting Structure
### Opening (5 minutes)
1. **State meeting purpose**: Make hiring decision based on evidence
2. **Review agenda and time limits**: Keep discussion focused and productive
3. **Remind of bias mitigation principles**: Focus on competencies, not personality
4. **Confirm confidentiality**: Discussion stays within hiring team
5. **Establish ground rules**: One person speaks at a time, evidence-based discussion
### Individual Score Sharing (10-15 minutes)
- **Go around the room systematically** - each interviewer shares scores independently
- **No discussion or challenges yet** - just data collection
- **Record scores on shared document** visible to all participants
- **Note any abstentions** or "insufficient data" responses
- **Identify clear patterns** and discrepancies without commentary
- **Flag any scores requiring explanation** (1s or 4s typically need strong evidence)
### Competency-by-Competency Discussion (30-40 minutes)
#### For Each Core Competency:
**1. Present Score Distribution (2 minutes)**
- Display all scores for this competency
- Note range and any outliers
- Identify if consensus exists or discussion needed
**2. Evidence Sharing (5-8 minutes per competency)**
- Start with interviewers who assessed this competency directly
- Share specific examples and observations
- Focus on what candidate said/did, not interpretations
- Allow questions for clarification (not challenges yet)
**3. Discussion and Calibration (3-5 minutes)**
- Address significant discrepancies (>1 point difference)
- Challenge vague or potentially biased language
- Seek additional evidence if needed
- Allow score adjustments based on new information
- Reach consensus or note dissenting views
#### Structured Discussion Questions:
- **"What specific evidence supports this score?"**
- **"Can you provide the exact example or quote?"**
- **"How does this compare to our rubric definition?"**
- **"Would this response receive the same score regardless of who gave it?"**
- **"Are we evaluating the competency or making assumptions?"**
- **"What would need to change for this to be the next level up/down?"**
### Overall Recommendation Discussion (10-15 minutes)
#### Weighted Score Calculation
1. **Apply competency weights** based on role requirements
2. **Calculate overall weighted average**
3. **Check minimum threshold requirements**
4. **Consider any veto criteria** (critical competency failures)
#### Final Recommendation Options
- **Strong Hire**: Exceeds requirements in most areas, clear value-add
- **Hire**: Meets requirements with growth potential
- **No Hire**: Doesn't meet minimum requirements for success
- **Strong No Hire**: Significant gaps that would impact team/company
#### Decision Rationale Documentation
- **Summarize key strengths** with specific evidence
- **Identify development areas** with specific examples
- **Explain final recommendation** with competency-based reasoning
- **Note any dissenting opinions** and reasoning
- **Document onboarding considerations** if hiring
### Closing and Next Steps (5 minutes)
- **Confirm final decision** and documentation
- **Assign follow-up actions** (feedback delivery, offer preparation, etc.)
- **Schedule any additional interviews** if needed
- **Review timeline** for candidate communication
- **Remind confidentiality** of discussion and decision
## Facilitation Best Practices
### Creating Psychological Safety
- **Encourage honest feedback** without fear of judgment
- **Validate different perspectives** and assessment approaches
- **Address power dynamics** - ensure junior voices are heard
- **Model vulnerability** - admit when evidence changes your mind
- **Focus on learning** and calibration, not winning arguments
- **Thank participants** for thorough preparation and thoughtful input
### Managing Difficult Conversations
#### When Scores Vary Significantly
1. **Acknowledge the discrepancy** without judgment
2. **Ask for specific evidence** from each scorer
3. **Look for different interpretations** of the same data
4. **Consider if different questions** revealed different competency levels
5. **Check for bias patterns** in reasoning
6. **Allow time for reflection** and potential score adjustments
#### When Someone Uses Biased Language
1. **Pause the conversation** gently but firmly
2. **Ask for specific evidence** behind the assessment
3. **Reframe in competency terms** - "What specific skills did this demonstrate?"
4. **Challenge assumptions** - "Help me understand how we know that"
5. **Redirect to rubric** - "How does this align with our scoring criteria?"
6. **Document and follow up** privately if bias persists
#### When the Discussion Gets Off Track
- **Redirect to competencies**: "Let's focus on the technical skills demonstrated"
- **Ask for evidence**: "What specific example supports that assessment?"
- **Reference rubrics**: "How does this align with our level 3 definition?"
- **Manage time**: "We have 5 minutes left on this competency"
- **Table unrelated issues**: "That's important but separate from this hire decision"
### Encouraging Evidence-Based Discussion
#### Good Evidence Examples
- **Direct quotes**: "When asked about debugging, they said..."
- **Specific behaviors**: "They organized their approach by first..."
- **Observable outcomes**: "Their code compiled on first run and handled edge cases"
- **Process descriptions**: "They walked through their problem-solving step by step"
- **Measurable results**: "They identified 3 optimization opportunities"
#### Poor Evidence Examples
- **Gut feelings**: "They just seemed off"
- **Comparisons**: "Not as strong as our last hire"
- **Assumptions**: "Probably wouldn't fit our culture"
- **Vague impressions**: "Didn't seem passionate"
- **Irrelevant factors**: "Their background is different from ours"
### Managing Group Dynamics
#### Ensuring Equal Participation
- **Direct questions** to quieter participants
- **Prevent interrupting** and ensure everyone finishes thoughts
- **Balance speaking time** across all interviewers
- **Validate minority opinions** even if not adopted
- **Check for unheard perspectives** before finalizing decisions
#### Handling Strong Personalities
- **Set time limits** for individual speaking
- **Redirect monopolizers**: "Let's hear from others on this"
- **Challenge confidently stated opinions** that lack evidence
- **Support less assertive voices** in expressing dissenting views
- **Focus on data**, not personality or seniority in decision making
## Bias Interruption Strategies
### Affinity Bias Interruption
- **Notice pattern**: Positive assessment seems based on shared background/interests
- **Interrupt with**: "Let's focus on the job-relevant skills they demonstrated"
- **Redirect to**: Specific competency evidence and measurable outcomes
- **Document**: Note if personal connection affected professional assessment
### Halo/Horn Effect Interruption
- **Notice pattern**: One area strongly influencing assessment of unrelated areas
- **Interrupt with**: "Let's score each competency independently"
- **Redirect to**: Specific evidence for each individual competency area
- **Recalibrate**: Ask for separate examples supporting each score
### Confirmation Bias Interruption
- **Notice pattern**: Only seeking/discussing evidence that supports initial impression
- **Interrupt with**: "What evidence might suggest a different assessment?"
- **Redirect to**: Consider alternative interpretations of the same data
- **Challenge**: "How might we be wrong about this assessment?"
### Attribution Bias Interruption
- **Notice pattern**: Attributing success to luck/help for some demographics, skill for others
- **Interrupt with**: "What role did the candidate play in achieving this outcome?"
- **Redirect to**: Candidate's specific contributions and decision-making
- **Standardize**: Apply same attribution standards across all candidates
## Decision Documentation Framework
### Required Documentation Elements
1. **Final scores** for each assessed competency
2. **Overall recommendation** with supporting rationale
3. **Key strengths** with specific evidence
4. **Development areas** with specific examples
5. **Dissenting opinions** if any, with reasoning
6. **Special considerations** or accommodation needs
7. **Next steps** and timeline for decision communication
### Evidence Quality Standards
- **Specific and observable**: What exactly did the candidate do or say?
- **Job-relevant**: How does this relate to success in the role?
- **Measurable**: Can this be quantified or clearly described?
- **Unbiased**: Would this evidence be interpreted the same way regardless of candidate demographics?
- **Complete**: Does this represent the full picture of their performance in this area?
### Writing Guidelines
- **Use active voice** and specific language
- **Avoid assumptions** about motivations or personality
- **Focus on behaviors** demonstrated during the interview
- **Provide context** for any unusual circumstances
- **Be constructive** in describing development areas
- **Maintain professionalism** and respect for candidate
## Common Debrief Challenges and Solutions
### Challenge: "I just don't think they'd fit our culture"
**Solution**:
- Ask for specific, observable evidence
- Define what "culture fit" means in job-relevant terms
- Challenge assumptions about cultural requirements
- Focus on ability to collaborate and contribute effectively
### Challenge: Scores vary widely with no clear explanation
**Solution**:
- Review if different interviewers assessed different competencies
- Look for question differences that might explain variance
- Consider if candidate performance varied across interviews
- May need additional data gathering or interview
### Challenge: Everyone loved/hated the candidate but can't articulate why
**Solution**:
- Push for specific evidence supporting emotional reactions
- Review competency rubrics together
- Look for halo/horn effects influencing overall impression
- Consider unconscious bias training for team
### Challenge: Technical vs. non-technical interviewers disagree
**Solution**:
- Clarify which competencies each interviewer was assessing
- Ensure technical assessments carry appropriate weight
- Look for different perspectives on same evidence
- Consider specialist input for technical decisions
### Challenge: Senior interviewer dominates decision making
**Solution**:
- Structure discussion to hear from all levels first
- Ask direct questions to junior interviewers
- Challenge opinions that lack supporting evidence
- Remember that assessment ability doesn't correlate with seniority
### Challenge: Team wants to hire but scores don't support it
**Solution**:
- Review if rubrics match actual job requirements
- Check for consistent application of scoring standards
- Consider if additional competencies need assessment
- May indicate need for rubric calibration or role requirement review
## Post-Debrief Actions
### Immediate Actions (Same Day)
- [ ] **Finalize decision documentation** with all evidence
- [ ] **Communicate decision** to recruiting team
- [ ] **Schedule candidate feedback** delivery if applicable
- [ ] **Update interview scheduling** based on decision
- [ ] **Note any process improvements** needed for future
### Follow-up Actions (Within 1 Week)
- [ ] **Deliver candidate feedback** (internal or external)
- [ ] **Update interview feedback** in tracking system
- [ ] **Schedule any additional interviews** if needed
- [ ] **Begin offer process** if hiring
- [ ] **Document lessons learned** for process improvement
### Long-term Actions (Monthly/Quarterly)
- [ ] **Analyze debrief effectiveness** and decision quality
- [ ] **Review interviewer calibration** based on decisions
- [ ] **Update rubrics** based on debrief insights
- [ ] **Provide additional training** if bias patterns identified
- [ ] **Share successful practices** with other hiring teams
## Continuous Improvement Framework
### Debrief Effectiveness Metrics
- **Decision consistency**: Are similar candidates receiving similar decisions?
- **Time to decision**: Are debriefs completing within planned time?
- **Participation quality**: Are all interviewers contributing evidence-based input?
- **Bias incidents**: How often are bias interruptions needed?
- **Decision satisfaction**: Do participants feel good about the process and outcome?
### Regular Review Process
- **Monthly**: Review debrief facilitation effectiveness and interviewer feedback
- **Quarterly**: Analyze decision patterns and potential bias indicators
- **Semi-annually**: Update debrief processes based on hiring outcome data
- **Annually**: Comprehensive review of debrief framework and training needs
### Training and Calibration
- **New facilitators**: Shadow 3-5 debriefs before leading independently
- **All facilitators**: Quarterly calibration sessions on bias interruption
- **Interviewer training**: Include debrief participation expectations
- **Leadership training**: Ensure hiring managers can facilitate effectively
This guide should be adapted to your organization's specific needs while maintaining focus on evidence-based, unbiased decision making.
FILE:references/interview-frameworks.md
# Interview Frameworks
## Loop Design by Level
### Junior/Mid
- Emphasize fundamentals, debugging, and growth potential.
- Keep loops concise with coding + behavioral validation.
### Senior
- Add system design and leadership rounds.
- Evaluate tradeoff quality, mentoring, and cross-team collaboration.
### Staff+
- Focus on architecture direction and organizational impact.
- Assess strategy, influence, and long-term technical judgment.
## Competency Areas
- Technical depth (implementation, design, quality)
- Problem solving (ambiguity handling, prioritization)
- Collaboration (communication, stakeholder alignment)
- Leadership (ownership, mentoring, influence)
## Scoring Rubric Baseline
- `4`: exceeds level expectations with strong evidence
- `3`: meets expectations consistently
- `2`: partial signal with notable gaps
- `1`: does not meet baseline requirements
## Calibration Guidelines
- Run recurring interviewer calibration sessions.
- Compare interviewer scoring variance across rounds.
- Track interview signal against new-hire outcomes.
- Use structured debriefs with independent scoring before discussion.
## Bias-Reduction Baseline
- Standardize question banks per competency area.
- Keep scorecards evidence-based and behavior-specific.
- Use diverse interviewer panels where possible.
- Require written rationale for strong yes/no recommendations.
FILE:scripts/interview_planner.py
#!/usr/bin/env python3
"""Generate an interview loop plan by role and level."""
from __future__ import annotations
import argparse
import json
from typing import Dict, List
BASE_ROUNDS = {
"junior": [
("Screen", 45, "Fundamentals and communication"),
("Coding", 60, "Problem solving and code quality"),
("Behavioral", 45, "Collaboration and growth mindset"),
],
"mid": [
("Screen", 45, "Fundamentals and ownership"),
("Coding", 60, "Implementation quality"),
("System Design", 60, "Service/component design"),
("Behavioral", 45, "Stakeholder collaboration"),
],
"senior": [
("Screen", 45, "Depth and tradeoff reasoning"),
("Coding", 60, "Code quality and testing"),
("System Design", 75, "Scalability and reliability"),
("Leadership", 60, "Mentoring and decision making"),
("Behavioral", 45, "Cross-functional influence"),
],
"staff": [
("Screen", 45, "Strategic and technical depth"),
("Architecture", 90, "Org-level design decisions"),
("Technical Strategy", 60, "Long-term tradeoffs"),
("Influence", 60, "Cross-team leadership"),
("Behavioral", 45, "Values and executive communication"),
],
}
QUESTION_BANK = {
"coding": [
"Walk through your approach before coding and identify tradeoffs.",
"How would you test this implementation for edge cases?",
"What would you refactor if this code became a shared library?",
],
"system": [
"Design this system for 10x traffic growth in 12 months.",
"Where are the main failure modes and how would you detect them?",
"What components would you scale first and why?",
],
"leadership": [
"Describe a time you changed technical direction with incomplete information.",
"How do you raise the bar for code quality across a team?",
"How do you handle disagreement between product and engineering priorities?",
],
"behavioral": [
"Tell me about a high-stakes mistake and what changed afterward.",
"Describe a conflict where you had to influence without authority.",
"How do you support underperforming teammates?",
],
}
def normalize_level(level: str) -> str:
level = level.strip().lower()
if level in {"staff+", "principal", "lead"}:
return "staff"
if level not in BASE_ROUNDS:
raise ValueError(f"Unsupported level: {level}")
return level
def suggested_questions(round_name: str) -> List[str]:
name = round_name.lower()
if "coding" in name:
return QUESTION_BANK["coding"]
if "system" in name or "architecture" in name:
return QUESTION_BANK["system"]
if "lead" in name or "influence" in name or "strategy" in name:
return QUESTION_BANK["leadership"]
return QUESTION_BANK["behavioral"]
def generate_plan(role: str, level: str) -> Dict[str, object]:
normalized = normalize_level(level)
rounds = []
for idx, (name, minutes, focus) in enumerate(BASE_ROUNDS[normalized], start=1):
rounds.append(
{
"round": idx,
"name": name,
"duration_minutes": minutes,
"focus": focus,
"suggested_questions": suggested_questions(name),
}
)
return {
"role": role,
"level": normalized,
"total_rounds": len(rounds),
"total_minutes": sum(r["duration_minutes"] for r in rounds),
"rounds": rounds,
}
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Generate an interview loop plan for a role and level.")
parser.add_argument("--role", required=True, help="Role name (e.g., Senior Software Engineer)")
parser.add_argument("--level", required=True, help="Level: junior|mid|senior|staff")
parser.add_argument("--json", action="store_true", help="Output as JSON")
return parser.parse_args()
def main() -> int:
args = parse_args()
plan = generate_plan(args.role, args.level)
if args.json:
print(json.dumps(plan, indent=2))
else:
print(f"Interview Plan: {plan['role']} ({plan['level']})")
print(f"Total rounds: {plan['total_rounds']} | Total time: {plan['total_minutes']} minutes")
print("")
for r in plan["rounds"]:
print(f"Round {r['round']}: {r['name']} ({r['duration_minutes']} min)")
print(f"Focus: {r['focus']}")
for q in r["suggested_questions"]:
print(f"- {q}")
print("")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Tạo persona người dùng dựa trên dữ liệu cho nghiên cứu UX và thiết kế sản phẩm.
--- name: persona description: Generate data-driven user personas for UX research and product design. Usage: /persona generate [options] --- # /persona Generate structured user personas with demographics, goals, pain points, and behavioral patterns. ## Usage ``` /persona generate Generate persona (interactive) /persona generate json Generate persona as JSON ``` ## Input Format Interactive mode prompts for product context. Alternatively, provide context inline: ``` /persona generate > Product: B2B project management tool > Target: Engineering managers at mid-size companies > Key problem: Cross-team visibility ``` ## Examples ``` /persona generate /persona generate json /persona generate json > persona-eng-manager.json ``` ## Scripts - `product-team/ux-researcher-designer/scripts/persona_generator.py` — Persona generator (positional `json` arg for JSON output) ## Skill Reference > `product-team/ux-researcher-designer/SKILL.md`
Kiểm tra tập dữ liệu về độ đầy đủ, nhất quán, chính xác, hợp lệ; phát hiện bất thường và lập kế hoạch khắc phục.
---
name: data-quality-auditor
description: Audit datasets for completeness, consistency, accuracy, and validity. Profile data distributions, detect anomalies and outliers, surface structural issues, and produce an actionable remediation plan.
---
You are an expert data quality engineer. Your goal is to systematically assess dataset health, surface hidden issues that corrupt downstream analysis, and prescribe prioritized fixes. You move fast, think in impact, and never let "good enough" data quietly poison a model or dashboard.
---
## Entry Points
### Mode 1 — Full Audit (New Dataset)
Use when you have a dataset you've never assessed before.
1. **Profile** — Run `data_profiler.py` to get shape, types, completeness, and distributions
2. **Missing Values** — Run `missing_value_analyzer.py` to classify missingness patterns (MCAR/MAR/MNAR)
3. **Outliers** — Run `outlier_detector.py` to flag anomalies using IQR and Z-score methods
4. **Cross-column checks** — Inspect referential integrity, duplicate rows, and logical constraints
5. **Score & Report** — Assign a Data Quality Score (DQS) and produce the remediation plan
### Mode 2 — Targeted Scan (Specific Concern)
Use when a specific column, metric, or pipeline stage is suspected.
1. Ask: *What broke, when did it start, and what changed upstream?*
2. Run the relevant script against the suspect columns only
3. Compare distributions against a known-good baseline if available
4. Trace issues to root cause (source system, ETL transform, ingestion lag)
### Mode 3 — Ongoing Monitoring Setup
Use when the user wants recurring quality checks on a live pipeline.
1. Identify the 5–8 critical columns driving key metrics
2. Define thresholds: acceptable null %, outlier rate, value domain
3. Generate a monitoring checklist and alerting logic from `data_profiler.py --monitor`
4. Schedule checks at ingestion cadence
---
## Tools
### `scripts/data_profiler.py`
Full dataset profile: shape, dtypes, null counts, cardinality, value distributions, and a Data Quality Score.
**Features:**
- Per-column null %, unique count, top values, min/max/mean/std
- Detects constant columns, high-cardinality text fields, mixed types
- Outputs a DQS (0–100) based on completeness + consistency signals
- `--monitor` flag prints threshold-ready summary for alerting
```bash
# Profile from CSV
python3 scripts/data_profiler.py --file data.csv
# Profile specific columns
python3 scripts/data_profiler.py --file data.csv --columns col1,col2,col3
# Output JSON for downstream use
python3 scripts/data_profiler.py --file data.csv --format json
# Generate monitoring thresholds
python3 scripts/data_profiler.py --file data.csv --monitor
```
### `scripts/missing_value_analyzer.py`
Deep-dive into missingness: volume, patterns, and likely mechanism (MCAR/MAR/MNAR).
**Features:**
- Null heatmap summary (text-based) and co-occurrence matrix
- Pattern classification: random, systematic, correlated
- Imputation strategy recommendations per column (drop / mean / median / mode / forward-fill / flag)
- Estimates downstream impact if missingness is ignored
```bash
# Analyze all missing values
python3 scripts/missing_value_analyzer.py --file data.csv
# Focus on columns above a null threshold
python3 scripts/missing_value_analyzer.py --file data.csv --threshold 0.05
# Output JSON
python3 scripts/missing_value_analyzer.py --file data.csv --format json
```
### `scripts/outlier_detector.py`
Multi-method outlier detection with business-impact context.
**Features:**
- IQR method (robust, non-parametric)
- Z-score method (normal distribution assumption)
- Modified Z-score (Iglewicz-Hoaglin, robust to skew)
- Per-column outlier count, %, and boundary values
- Flags columns where outliers may be data errors vs. legitimate extremes
```bash
# Detect outliers across all numeric columns
python3 scripts/outlier_detector.py --file data.csv
# Use specific method
python3 scripts/outlier_detector.py --file data.csv --method iqr
# Set custom Z-score threshold
python3 scripts/outlier_detector.py --file data.csv --method zscore --threshold 2.5
# Output JSON
python3 scripts/outlier_detector.py --file data.csv --format json
```
---
## Data Quality Score (DQS)
The DQS is a 0–100 composite score across five dimensions. Report it at the top of every audit.
| Dimension | Weight | What It Measures |
|---|---|---|
| Completeness | 30% | Null / missing rate across critical columns |
| Consistency | 25% | Type conformance, format uniformity, no mixed types |
| Validity | 20% | Values within expected domain (ranges, categories, regexes) |
| Uniqueness | 15% | Duplicate rows, duplicate keys, redundant columns |
| Timeliness | 10% | Freshness of timestamps, lag from source system |
**Scoring thresholds:**
- 🟢 85–100 — Production-ready
- 🟡 65–84 — Usable with documented caveats
- 🔴 0–64 — Remediation required before use
---
## Proactive Risk Triggers
Surface these unprompted whenever you spot the signals:
- **Silent nulls** — Nulls encoded as `0`, `""`, `"N/A"`, `"null"` strings. Completeness metrics lie until these are caught.
- **Leaky timestamps** — Future dates, dates before system launch, or timezone mismatches that corrupt time-series joins.
- **Cardinality explosions** — Free-text fields with thousands of unique values masquerading as categorical. Will break one-hot encoding silently.
- **Duplicate keys** — PKs that aren't unique invalidate joins and aggregations downstream.
- **Distribution shift** — Columns where current distribution diverges from baseline (>2σ on mean/std). Signals upstream pipeline changes.
- **Correlated missingness** — Nulls concentrated in a specific time range, user segment, or region — evidence of MNAR, not random dropout.
---
## Output Artifacts
| Request | Deliverable |
|---|---|
| "Profile this dataset" | Full DQS report with per-column breakdown and top issues ranked by impact |
| "What's wrong with column X?" | Targeted column audit: nulls, outliers, type issues, value domain violations |
| "Is this data ready for modeling?" | Model-readiness checklist with pass/fail per ML requirement |
| "Help me clean this data" | Prioritized remediation plan with specific transforms per issue |
| "Set up monitoring" | Threshold config + alerting checklist for critical columns |
| "Compare this to last month" | Distribution comparison report with drift flags |
---
## Remediation Playbook
### Missing Values
| Null % | Recommended Action |
|---|---|
| < 1% | Drop rows (if dataset is large) or impute with median/mode |
| 1–10% | Impute; add a binary indicator column `col_was_null` |
| 10–30% | Impute cautiously; investigate root cause; document assumption |
| > 30% | Flag for domain review; do not impute blindly; consider dropping column |
### Outliers
- **Likely data error** (value physically impossible): cap, correct, or drop
- **Legitimate extreme** (valid but rare): keep, document, consider log transform for modeling
- **Unknown** (can't determine without domain input): flag, do not silently remove
### Duplicates
1. Confirm uniqueness key with data owner before deduplication
2. Prefer `keep='last'` for event data (most recent state wins)
3. Prefer `keep='first'` for slowly-changing-dimension tables
---
## Quality Loop
Tag every finding with a confidence level:
- 🟢 **Verified** — confirmed by data inspection or domain owner
- 🟡 **Likely** — strong signal but not fully confirmed
- 🔴 **Assumed** — inferred from patterns; needs domain validation
Never auto-remediate 🔴 findings without human confirmation.
---
## Communication Standard
Structure all audit reports as:
**Bottom Line** — DQS score and one-sentence verdict (e.g., "DQS: 61/100 — remediation required before production use")
**What** — The specific issues found (ranked by severity × breadth)
**Why It Matters** — Business or analytical impact of each issue
**How to Act** — Specific, ordered remediation steps
---
## Related Skills
| Skill | Use When |
|---|---|
| `finance/financial-analyst` | Data involves financial statements or accounting figures |
| `finance/saas-metrics-coach` | Data is subscription/event data feeding SaaS KPIs |
| `engineering/database-designer` | Issues trace back to schema design or normalization |
| `engineering/tech-debt-tracker` | Data quality issues are systemic and need to be tracked as tech debt |
| `product-team/product-analytics` | Auditing product event data (funnels, sessions, retention) |
**When NOT to use this skill:**
- You need to design or optimize the database schema — use `engineering/database-designer`
- You need to build the ETL pipeline itself — use an engineering skill
- The dataset is a financial model output — use `finance/financial-analyst` for model validation
---
## References
- `references/data-quality-concepts.md` — MCAR/MAR/MNAR theory, DQS methodology, outlier detection methods
FILE:references/data-quality-concepts.md
# Data Quality Concepts Reference
Deep-dive reference for the Data Quality Auditor skill. Keep SKILL.md lean — this is where the theory lives.
---
## Missingness Mechanisms (Rubin, 1976)
Understanding *why* data is missing determines how safely it can be imputed.
### MCAR — Missing Completely At Random
- The probability of missingness is independent of both observed and unobserved data.
- **Example:** A sensor drops a reading due to random hardware noise.
- **Safe to impute?** Yes. Imputing with mean/median introduces no systematic bias.
- **Detection:** Null rows are indistinguishable from non-null rows on all other dimensions.
### MAR — Missing At Random
- The probability of missingness depends on *observed* data, not the missing value itself.
- **Example:** Older users are less likely to fill in a "social media handle" field — missingness depends on age (observed), not on the handle itself.
- **Safe to impute?** Conditionally yes — impute using a model that accounts for the related observed variables.
- **Detection:** Null rows differ systematically from non-null rows on *other* columns.
### MNAR — Missing Not At Random
- The probability of missingness depends on the *missing value itself* (unobserved).
- **Example:** High earners skip the income field; low performers skip the satisfaction survey.
- **Safe to impute?** No — imputation will introduce systematic bias. Escalate to domain owner.
- **Detection:** Difficult to confirm statistically; look for clustered nulls in time or segment slices.
---
## Data Quality Score (DQS) Methodology
The DQS is a weighted composite of five ISO 8000 / DAMA-aligned dimensions:
| Dimension | Weight | Rationale |
|---|---|---|
| Completeness | 30% | Nulls are the most common and impactful quality failure |
| Consistency | 25% | Type/format violations corrupt joins and aggregations silently |
| Validity | 20% | Out-of-domain values (negative ages, future birth dates) create invisible errors |
| Uniqueness | 15% | Duplicate rows inflate metrics and invalidate joins |
| Timeliness | 10% | Stale data causes decisions based on outdated state |
**Scoring thresholds** align to production-readiness standards:
- 85–100: Ready for production use in models and dashboards
- 65–84: Usable for exploratory analysis with documented caveats
- 0–64: Unreliable; remediation required before use in any decision-making context
---
## Outlier Detection Methods
### IQR (Interquartile Range)
- **Formula:** Outlier if `x < Q1 − 1.5×IQR` or `x > Q3 + 1.5×IQR`
- **Strengths:** Non-parametric, robust to non-normal distributions, interpretable bounds
- **Weaknesses:** Can miss outliers in heavily skewed distributions; 1.5× multiplier is conventional, not universal
- **When to use:** Default choice for most business datasets (revenue, counts, durations)
### Z-score
- **Formula:** Outlier if `|x − μ| / σ > threshold` (commonly 3.0)
- **Strengths:** Simple, widely understood, easy to explain to stakeholders
- **Weaknesses:** Mean and std are themselves influenced by outliers — the method is self-defeating for extreme contamination
- **When to use:** Only when the distribution is approximately normal and contamination is < 5%
### Modified Z-score (Iglewicz-Hoaglin)
- **Formula:** `M_i = 0.6745 × |x_i − median| / MAD`; outlier if `M_i > 3.5`
- **Strengths:** Uses median and MAD — both resistant to outlier influence; handles skewed distributions
- **Weaknesses:** MAD = 0 for discrete columns with one dominant value; less intuitive
- **When to use:** Preferred for skewed distributions (e.g. revenue, latency, page views)
---
## Imputation Strategies
| Method | When | Risk |
|---|---|---|
| Mean | MCAR, continuous, symmetric distribution | Distorts variance; don't use with skewed data |
| Median | MCAR/MAR, continuous, skewed distribution | Safe for skewed; loses variance |
| Mode | MCAR/MAR, categorical | Can over-represent one category |
| Forward-fill | Time series with MCAR/MAR gaps | Assumes value persists — valid for slowly-changing fields |
| Binary indicator | Null % 1–30% | Preserves information about missingness without imputing |
| Model-based | MAR, high-value columns | Most accurate but computationally expensive |
| Drop column | > 50% missing with no business justification | Safest option if column has no predictive value |
**Golden rule:** Always add a `col_was_null` indicator column when imputing with null% > 1%. This preserves the information that a value was imputed, which may itself be predictive.
---
## Common Silent Data Quality Failures
These are the issues that don't raise errors but corrupt results:
1. **Sentinel values** — `0`, `-1`, `9999`, `""` used to mean "unknown" in legacy systems
2. **Timezone naive timestamps** — datetimes stored without timezone; comparisons silently shift by hours
3. **Trailing whitespace** — `"active "` ≠ `"active"` causes silent join mismatches
4. **Encoding errors** — UTF-8 vs Latin-1 mismatches produce garbled strings in one column
5. **Scientific notation** — `1e6` stored as string gets treated as a category not a number
6. **Implicit schema changes** — upstream adds a new category to a lookup field; existing code silently drops new rows
---
## References
- Rubin, D.B. (1976). "Inference and Missing Data." *Biometrika* 63(3): 581–592.
- Iglewicz, B. & Hoaglin, D. (1993). *How to Detect and Handle Outliers*. ASQC Quality Press.
- DAMA International (2017). *DAMA-DMBOK: Data Management Body of Knowledge*. 2nd ed.
- ISO 8000-8: Data quality — Concepts and measuring.
FILE:scripts/data_profiler.py
#!/usr/bin/env python3
from __future__ import annotations
"""
data_profiler.py — Full dataset profile with Data Quality Score (DQS).
Usage:
python3 data_profiler.py --file data.csv
python3 data_profiler.py --file data.csv --columns col1,col2
python3 data_profiler.py --file data.csv --format json
python3 data_profiler.py --file data.csv --monitor
"""
import argparse
import csv
import json
import math
import sys
from collections import Counter, defaultdict
def load_csv(filepath: str) -> tuple[list[str], list[dict]]:
with open(filepath, newline="", encoding="utf-8") as f:
reader = csv.DictReader(f)
rows = list(reader)
headers = reader.fieldnames or []
return headers, rows
def infer_type(values: list[str]) -> str:
"""Infer dominant type from non-null string values."""
counts = {"int": 0, "float": 0, "bool": 0, "string": 0}
for v in values:
v = v.strip()
if v.lower() in ("true", "false"):
counts["bool"] += 1
else:
try:
int(v)
counts["int"] += 1
except ValueError:
try:
float(v)
counts["float"] += 1
except ValueError:
counts["string"] += 1
dominant = max(counts, key=lambda k: counts[k])
return dominant if counts[dominant] > 0 else "string"
def safe_mean(nums: list[float]) -> float | None:
return sum(nums) / len(nums) if nums else None
def safe_std(nums: list[float], mean: float) -> float | None:
if len(nums) < 2:
return None
variance = sum((x - mean) ** 2 for x in nums) / (len(nums) - 1)
return math.sqrt(variance)
def profile_column(name: str, raw_values: list[str]) -> dict:
total = len(raw_values)
null_strings = {"", "null", "none", "n/a", "na", "nan", "nil"}
null_count = sum(1 for v in raw_values if v.strip().lower() in null_strings)
non_null = [v for v in raw_values if v.strip().lower() not in null_strings]
col_type = infer_type(non_null)
unique_values = set(non_null)
top_values = Counter(non_null).most_common(5)
profile = {
"column": name,
"total_rows": total,
"null_count": null_count,
"null_pct": round(null_count / total * 100, 2) if total else 0,
"non_null_count": len(non_null),
"unique_count": len(unique_values),
"cardinality_pct": round(len(unique_values) / len(non_null) * 100, 2) if non_null else 0,
"inferred_type": col_type,
"top_values": top_values,
"is_constant": len(unique_values) == 1,
"is_high_cardinality": len(unique_values) / len(non_null) > 0.9 if len(non_null) > 10 else False,
}
if col_type in ("int", "float"):
try:
nums = [float(v) for v in non_null]
mean = safe_mean(nums)
profile["min"] = min(nums)
profile["max"] = max(nums)
profile["mean"] = round(mean, 4) if mean is not None else None
profile["std"] = round(safe_std(nums, mean), 4) if mean is not None else None
except ValueError:
pass
return profile
def compute_dqs(profiles: list[dict], total_rows: int) -> dict:
"""Compute Data Quality Score (0-100) across 5 dimensions."""
if not profiles or total_rows == 0:
return {"score": 0, "dimensions": {}}
# Completeness (30%) — avg non-null rate
avg_null_pct = sum(p["null_pct"] for p in profiles) / len(profiles)
completeness = max(0, 100 - avg_null_pct)
# Consistency (25%) — penalize constant cols and mixed-type signals
constant_cols = sum(1 for p in profiles if p["is_constant"])
consistency = max(0, 100 - (constant_cols / len(profiles)) * 100)
# Validity (20%) — penalize high-cardinality string cols (proxy for free-text issues)
high_card = sum(1 for p in profiles if p["is_high_cardinality"] and p["inferred_type"] == "string")
validity = max(0, 100 - (high_card / len(profiles)) * 60)
# Uniqueness (15%) — placeholder; duplicate detection needs full row comparison
uniqueness = 90.0 # conservative default without row-level dedup check
# Timeliness (10%) — placeholder; requires timestamp columns
timeliness = 85.0 # conservative default
score = (
completeness * 0.30
+ consistency * 0.25
+ validity * 0.20
+ uniqueness * 0.15
+ timeliness * 0.10
)
return {
"score": round(score, 1),
"dimensions": {
"completeness": round(completeness, 1),
"consistency": round(consistency, 1),
"validity": round(validity, 1),
"uniqueness": uniqueness,
"timeliness": timeliness,
},
}
def dqs_label(score: float) -> str:
if score >= 85:
return "PASS — Production-ready"
elif score >= 65:
return "WARN — Usable with documented caveats"
else:
return "FAIL — Remediation required before use"
def print_report(headers: list[str], profiles: list[dict], dqs: dict, total_rows: int, monitor: bool):
print("=" * 64)
print("DATA QUALITY AUDIT REPORT")
print("=" * 64)
print(f"Rows: {total_rows} | Columns: {len(headers)}")
score = dqs["score"]
indicator = "🟢" if score >= 85 else ("🟡" if score >= 65 else "🔴")
print(f"\nData Quality Score (DQS): {score}/100 {indicator}")
print(f"Verdict: {dqs_label(score)}")
dims = dqs["dimensions"]
print("\nDimension Breakdown:")
for dim, val in dims.items():
bar = int(val / 5)
print(f" {dim.capitalize():<14} {val:>5.1f} {'█' * bar}{'░' * (20 - bar)}")
print("\n" + "-" * 64)
print("COLUMN PROFILES")
print("-" * 64)
issues = []
for p in profiles:
status = "🟢"
col_issues = []
if p["null_pct"] > 30:
status = "🔴"
col_issues.append(f"{p['null_pct']}% nulls — investigate root cause")
elif p["null_pct"] > 10:
status = "🟡"
col_issues.append(f"{p['null_pct']}% nulls — impute cautiously")
elif p["null_pct"] > 1:
col_issues.append(f"{p['null_pct']}% nulls — impute with indicator")
if p["is_constant"]:
status = "🟡"
col_issues.append("Constant column — zero variance, likely useless")
if p["is_high_cardinality"] and p["inferred_type"] == "string":
col_issues.append("High-cardinality string — check if categorical or free-text")
print(f"\n {status} {p['column']}")
print(f" Type: {p['inferred_type']} | Nulls: {p['null_count']} ({p['null_pct']}%) | Unique: {p['unique_count']}")
if "min" in p:
print(f" Min: {p['min']} Max: {p['max']} Mean: {p['mean']} Std: {p['std']}")
if p["top_values"]:
top = ", ".join(f"{v}({c})" for v, c in p["top_values"][:3])
print(f" Top values: {top}")
for issue in col_issues:
issues.append((p["column"], issue))
print(f" ⚠ {issue}")
if issues:
print("\n" + "-" * 64)
print(f"ISSUES SUMMARY ({len(issues)} found)")
print("-" * 64)
for col, msg in issues:
print(f" [{col}] {msg}")
if monitor:
print("\n" + "-" * 64)
print("MONITORING THRESHOLDS (copy into alerting config)")
print("-" * 64)
for p in profiles:
if p["null_pct"] > 0:
print(f" {p['column']}: null_pct <= {min(p['null_pct'] * 1.5, 100):.1f}%")
if "mean" in p and p["mean"] is not None:
drift = abs(p.get("std", 0) or 0) * 2
print(f" {p['column']}: mean within [{p['mean'] - drift:.2f}, {p['mean'] + drift:.2f}]")
print("\n" + "=" * 64)
def main():
parser = argparse.ArgumentParser(description="Profile a CSV dataset and compute a Data Quality Score.")
parser.add_argument("--file", required=True, help="Path to CSV file")
parser.add_argument("--columns", help="Comma-separated list of columns to profile (default: all)")
parser.add_argument("--format", choices=["text", "json"], default="text")
parser.add_argument("--monitor", action="store_true", help="Print monitoring thresholds")
args = parser.parse_args()
try:
headers, rows = load_csv(args.file)
except FileNotFoundError:
print(f"Error: file not found: {args.file}", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Error reading file: {e}", file=sys.stderr)
sys.exit(1)
if not rows:
print("Error: CSV file is empty or has no data rows.", file=sys.stderr)
sys.exit(1)
selected = args.columns.split(",") if args.columns else headers
missing_cols = [c for c in selected if c not in headers]
if missing_cols:
print(f"Error: columns not found: {', '.join(missing_cols)}", file=sys.stderr)
sys.exit(1)
profiles = [profile_column(col, [row.get(col, "") for row in rows]) for col in selected]
dqs = compute_dqs(profiles, len(rows))
if args.format == "json":
print(json.dumps({"total_rows": len(rows), "dqs": dqs, "columns": profiles}, indent=2))
else:
print_report(selected, profiles, dqs, len(rows), args.monitor)
if __name__ == "__main__":
main()
FILE:scripts/missing_value_analyzer.py
#!/usr/bin/env python3
"""
missing_value_analyzer.py — Classify missingness patterns and recommend imputation strategies.
Usage:
python3 missing_value_analyzer.py --file data.csv
python3 missing_value_analyzer.py --file data.csv --threshold 0.05
python3 missing_value_analyzer.py --file data.csv --format json
"""
import argparse
import csv
import json
import sys
from collections import defaultdict
NULL_STRINGS = {"", "null", "none", "n/a", "na", "nan", "nil", "undefined", "missing"}
def load_csv(filepath: str) -> tuple[list[str], list[dict]]:
with open(filepath, newline="", encoding="utf-8") as f:
reader = csv.DictReader(f)
rows = list(reader)
headers = reader.fieldnames or []
return headers, rows
def is_null(val: str) -> bool:
return val.strip().lower() in NULL_STRINGS
def compute_null_mask(headers: list[str], rows: list[dict]) -> dict[str, list[bool]]:
return {col: [is_null(row.get(col, "")) for row in rows] for col in headers}
def null_stats(mask: list[bool]) -> dict:
total = len(mask)
count = sum(mask)
return {"count": count, "pct": round(count / total * 100, 2) if total else 0}
def classify_mechanism(col: str, mask: list[bool], all_masks: dict[str, list[bool]]) -> str:
"""
Heuristic classification of missingness mechanism:
- MCAR: nulls appear randomly, no correlation with other columns
- MAR: nulls correlate with values in other observed columns
- MNAR: nulls correlate with the missing column's own unobserved value (can't fully detect)
Returns one of: "MCAR (likely)", "MAR (likely)", "MNAR (possible)", "Insufficient data"
"""
null_indices = {i for i, v in enumerate(mask) if v}
if not null_indices:
return "None"
n = len(mask)
if n < 10:
return "Insufficient data"
# Check correlation with other columns' nulls
correlated_cols = []
for other_col, other_mask in all_masks.items():
if other_col == col:
continue
other_null_indices = {i for i, v in enumerate(other_mask) if v}
if not other_null_indices:
continue
overlap = len(null_indices & other_null_indices)
union = len(null_indices | other_null_indices)
jaccard = overlap / union if union else 0
if jaccard > 0.5:
correlated_cols.append(other_col)
# Check if nulls are clustered (time/positional pattern) — proxy for MNAR
sorted_indices = sorted(null_indices)
if len(sorted_indices) > 2:
gaps = [sorted_indices[i + 1] - sorted_indices[i] for i in range(len(sorted_indices) - 1)]
avg_gap = sum(gaps) / len(gaps)
clustered = avg_gap < n / len(null_indices) * 0.5 # nulls appear closer together than random
else:
clustered = False
if correlated_cols:
return f"MAR (likely) — co-occurs with nulls in: {', '.join(correlated_cols[:3])}"
elif clustered:
return "MNAR (possible) — nulls are spatially clustered, may reflect a systematic gap"
else:
return "MCAR (likely) — nulls appear random, no strong correlation detected"
def recommend_strategy(pct: float, col_type: str) -> str:
if pct == 0:
return "No action needed"
if pct < 1:
return "Drop rows — impact is negligible"
if pct < 10:
strategies = {
"int": "Impute with median + add binary indicator column",
"float": "Impute with median + add binary indicator column",
"string": "Impute with mode or 'Unknown' category + add indicator",
"bool": "Impute with mode",
}
return strategies.get(col_type, "Impute with median/mode + add indicator")
if pct < 30:
return "Impute cautiously; investigate root cause; document assumption; add indicator"
return "Do NOT impute blindly — > 30% missing. Escalate to domain owner or consider dropping column"
def infer_type(values: list[str]) -> str:
non_null = [v for v in values if not is_null(v)]
counts = {"int": 0, "float": 0, "bool": 0, "string": 0}
for v in non_null[:200]: # sample for speed
v = v.strip()
if v.lower() in ("true", "false"):
counts["bool"] += 1
else:
try:
int(v)
counts["int"] += 1
except ValueError:
try:
float(v)
counts["float"] += 1
except ValueError:
counts["string"] += 1
return max(counts, key=lambda k: counts[k]) if any(counts.values()) else "string"
def compute_cooccurrence(headers: list[str], masks: dict[str, list[bool]], top_n: int = 5) -> list[dict]:
"""Find column pairs where nulls most frequently co-occur."""
pairs = []
cols = list(headers)
for i in range(len(cols)):
for j in range(i + 1, len(cols)):
a, b = cols[i], cols[j]
mask_a, mask_b = masks[a], masks[b]
overlap = sum(1 for x, y in zip(mask_a, mask_b) if x and y)
if overlap > 0:
pairs.append({"col_a": a, "col_b": b, "co_null_rows": overlap})
pairs.sort(key=lambda x: -x["co_null_rows"])
return pairs[:top_n]
def print_report(headers: list[str], rows: list[dict], masks: dict, threshold: float):
total = len(rows)
print("=" * 64)
print("MISSING VALUE ANALYSIS REPORT")
print("=" * 64)
print(f"Rows: {total} | Columns: {len(headers)}")
results = []
for col in headers:
mask = masks[col]
stats = null_stats(mask)
if stats["pct"] / 100 < threshold and stats["count"] > 0:
continue
raw_vals = [row.get(col, "") for row in rows]
col_type = infer_type(raw_vals)
mechanism = classify_mechanism(col, mask, masks)
strategy = recommend_strategy(stats["pct"], col_type)
results.append({
"column": col,
"null_count": stats["count"],
"null_pct": stats["pct"],
"col_type": col_type,
"mechanism": mechanism,
"strategy": strategy,
})
fully_complete = [col for col in headers if null_stats(masks[col])["count"] == 0]
print(f"\nFully complete columns: {len(fully_complete)}/{len(headers)}")
if not results:
print(f"\nNo columns exceed the null threshold ({threshold * 100:.1f}%).")
else:
print(f"\nColumns with missing values (threshold >= {threshold * 100:.1f}%):\n")
for r in sorted(results, key=lambda x: -x["null_pct"]):
indicator = "🔴" if r["null_pct"] > 30 else ("🟡" if r["null_pct"] > 10 else "🟢")
print(f" {indicator} {r['column']}")
print(f" Nulls: {r['null_count']} ({r['null_pct']}%) | Type: {r['col_type']}")
print(f" Mechanism: {r['mechanism']}")
print(f" Strategy: {r['strategy']}")
print()
cooccur = compute_cooccurrence(headers, masks)
if cooccur:
print("-" * 64)
print("NULL CO-OCCURRENCE (top pairs)")
print("-" * 64)
for pair in cooccur:
print(f" {pair['col_a']} + {pair['col_b']} → {pair['co_null_rows']} rows both null")
print("\n" + "=" * 64)
def main():
parser = argparse.ArgumentParser(description="Analyze missing values in a CSV dataset.")
parser.add_argument("--file", required=True, help="Path to CSV file")
parser.add_argument("--threshold", type=float, default=0.0,
help="Only show columns with null fraction above this (e.g. 0.05 = 5%%)")
parser.add_argument("--format", choices=["text", "json"], default="text")
args = parser.parse_args()
try:
headers, rows = load_csv(args.file)
except FileNotFoundError:
print(f"Error: file not found: {args.file}", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Error reading file: {e}", file=sys.stderr)
sys.exit(1)
if not rows:
print("Error: CSV file is empty.", file=sys.stderr)
sys.exit(1)
masks = compute_null_mask(headers, rows)
if args.format == "json":
output = []
for col in headers:
mask = masks[col]
stats = null_stats(mask)
raw_vals = [row.get(col, "") for row in rows]
col_type = infer_type(raw_vals)
mechanism = classify_mechanism(col, mask, masks)
strategy = recommend_strategy(stats["pct"], col_type)
output.append({
"column": col,
"null_count": stats["count"],
"null_pct": stats["pct"],
"col_type": col_type,
"mechanism": mechanism,
"strategy": strategy,
})
print(json.dumps({"total_rows": len(rows), "columns": output}, indent=2))
else:
print_report(headers, rows, masks, args.threshold)
if __name__ == "__main__":
main()
FILE:scripts/outlier_detector.py
#!/usr/bin/env python3
from __future__ import annotations
"""
outlier_detector.py — Multi-method outlier detection for numeric columns.
Methods:
iqr — Interquartile Range (robust, non-parametric, default)
zscore — Standard Z-score (assumes normal distribution)
mzscore — Modified Z-score via Median Absolute Deviation (robust to skew)
Usage:
python3 outlier_detector.py --file data.csv
python3 outlier_detector.py --file data.csv --method iqr
python3 outlier_detector.py --file data.csv --method zscore --threshold 2.5
python3 outlier_detector.py --file data.csv --columns col1,col2
python3 outlier_detector.py --file data.csv --format json
"""
import argparse
import csv
import json
import math
import sys
NULL_STRINGS = {"", "null", "none", "n/a", "na", "nan", "nil", "undefined", "missing"}
def load_csv(filepath: str) -> tuple[list[str], list[dict]]:
with open(filepath, newline="", encoding="utf-8") as f:
reader = csv.DictReader(f)
rows = list(reader)
headers = reader.fieldnames or []
return headers, rows
def is_null(val: str) -> bool:
return val.strip().lower() in NULL_STRINGS
def to_float(val: str) -> float | None:
try:
return float(val.strip())
except (ValueError, AttributeError):
return None
def median(nums: list[float]) -> float:
s = sorted(nums)
n = len(s)
mid = n // 2
return s[mid] if n % 2 else (s[mid - 1] + s[mid]) / 2
def percentile(nums: list[float], p: float) -> float:
"""Linear interpolation percentile."""
s = sorted(nums)
n = len(s)
if n == 1:
return s[0]
idx = p / 100 * (n - 1)
lo = int(idx)
hi = lo + 1
frac = idx - lo
if hi >= n:
return s[-1]
return s[lo] + frac * (s[hi] - s[lo])
def mean(nums: list[float]) -> float:
return sum(nums) / len(nums)
def std(nums: list[float], mu: float) -> float:
if len(nums) < 2:
return 0.0
variance = sum((x - mu) ** 2 for x in nums) / (len(nums) - 1)
return math.sqrt(variance)
# --- Detection methods ---
def detect_iqr(nums: list[float], multiplier: float = 1.5) -> dict:
q1 = percentile(nums, 25)
q3 = percentile(nums, 75)
iqr = q3 - q1
lower = q1 - multiplier * iqr
upper = q3 + multiplier * iqr
outliers = [x for x in nums if x < lower or x > upper]
return {
"method": "IQR",
"q1": round(q1, 4),
"q3": round(q3, 4),
"iqr": round(iqr, 4),
"lower_bound": round(lower, 4),
"upper_bound": round(upper, 4),
"outlier_count": len(outliers),
"outlier_pct": round(len(outliers) / len(nums) * 100, 2),
"outlier_values": sorted(set(round(x, 4) for x in outliers))[:10],
}
def detect_zscore(nums: list[float], threshold: float = 3.0) -> dict:
mu = mean(nums)
sigma = std(nums, mu)
if sigma == 0:
return {"method": "Z-score", "outlier_count": 0, "outlier_pct": 0.0,
"note": "Zero variance — all values identical"}
zscores = [(x, abs((x - mu) / sigma)) for x in nums]
outliers = [x for x, z in zscores if z > threshold]
return {
"method": "Z-score",
"mean": round(mu, 4),
"std": round(sigma, 4),
"threshold": threshold,
"outlier_count": len(outliers),
"outlier_pct": round(len(outliers) / len(nums) * 100, 2),
"outlier_values": sorted(set(round(x, 4) for x in outliers))[:10],
}
def detect_modified_zscore(nums: list[float], threshold: float = 3.5) -> dict:
"""Iglewicz-Hoaglin modified Z-score using Median Absolute Deviation."""
med = median(nums)
mad = median([abs(x - med) for x in nums])
if mad == 0:
return {"method": "Modified Z-score (MAD)", "outlier_count": 0, "outlier_pct": 0.0,
"note": "MAD is zero — consider Z-score instead"}
mzscores = [(x, 0.6745 * abs(x - med) / mad) for x in nums]
outliers = [x for x, mz in mzscores if mz > threshold]
return {
"method": "Modified Z-score (MAD)",
"median": round(med, 4),
"mad": round(mad, 4),
"threshold": threshold,
"outlier_count": len(outliers),
"outlier_pct": round(len(outliers) / len(nums) * 100, 2),
"outlier_values": sorted(set(round(x, 4) for x in outliers))[:10],
}
def classify_outlier_risk(pct: float, col: str) -> str:
"""Heuristic: flag whether outliers are likely data errors or legitimate extremes."""
if pct > 10:
return "High outlier rate — likely systematic data quality issue or wrong data type"
if pct > 5:
return "Elevated outlier rate — investigate source; may be mixed populations"
if pct > 1:
return "Moderate — review individually; could be legitimate extremes or entry errors"
if pct > 0:
return "Low — verify extreme values against source; likely legitimate but worth checking"
return "Clean — no outliers detected"
def analyze_column(col: str, nums: list[float], method: str, threshold: float) -> dict:
if len(nums) < 4:
return {"column": col, "status": "Skipped — fewer than 4 numeric values"}
if method == "iqr":
result = detect_iqr(nums, multiplier=threshold if threshold != 3.0 else 1.5)
elif method == "zscore":
result = detect_zscore(nums, threshold=threshold)
elif method == "mzscore":
result = detect_modified_zscore(nums, threshold=threshold)
else:
result = detect_iqr(nums)
result["column"] = col
result["total_numeric"] = len(nums)
result["risk_assessment"] = classify_outlier_risk(result.get("outlier_pct", 0), col)
return result
def print_report(results: list[dict]):
print("=" * 64)
print("OUTLIER DETECTION REPORT")
print("=" * 64)
clean = [r for r in results if r.get("outlier_count", 0) == 0 and "status" not in r]
flagged = [r for r in results if r.get("outlier_count", 0) > 0]
skipped = [r for r in results if "status" in r]
print(f"\nColumns analyzed: {len(results) - len(skipped)}")
print(f"Clean: {len(clean)}")
print(f"Flagged: {len(flagged)}")
if skipped:
print(f"Skipped: {len(skipped)} ({', '.join(r['column'] for r in skipped)})")
if flagged:
print("\n" + "-" * 64)
print("FLAGGED COLUMNS")
print("-" * 64)
for r in sorted(flagged, key=lambda x: -x.get("outlier_pct", 0)):
pct = r.get("outlier_pct", 0)
indicator = "🔴" if pct > 5 else "🟡"
print(f"\n {indicator} {r['column']} ({r['method']})")
print(f" Outliers: {r['outlier_count']} / {r['total_numeric']} rows ({pct}%)")
if "lower_bound" in r:
print(f" Bounds: [{r['lower_bound']}, {r['upper_bound']}] | IQR: {r['iqr']}")
if "mean" in r:
print(f" Mean: {r['mean']} | Std: {r['std']} | Threshold: ±{r['threshold']}σ")
if "median" in r:
print(f" Median: {r['median']} | MAD: {r['mad']} | Threshold: {r['threshold']}")
if r.get("outlier_values"):
vals = ", ".join(str(v) for v in r["outlier_values"][:8])
print(f" Sample outlier values: {vals}")
print(f" Assessment: {r['risk_assessment']}")
if clean:
cols = ", ".join(r["column"] for r in clean)
print(f"\n🟢 Clean columns: {cols}")
print("\n" + "=" * 64)
def main():
parser = argparse.ArgumentParser(description="Detect outliers in numeric columns of a CSV dataset.")
parser.add_argument("--file", required=True, help="Path to CSV file")
parser.add_argument("--method", choices=["iqr", "zscore", "mzscore"], default="iqr",
help="Detection method (default: iqr)")
parser.add_argument("--threshold", type=float, default=None,
help="Method threshold (IQR multiplier default 1.5; Z-score default 3.0; mzscore default 3.5)")
parser.add_argument("--columns", help="Comma-separated columns to check (default: all numeric)")
parser.add_argument("--format", choices=["text", "json"], default="text")
args = parser.parse_args()
# Set default thresholds per method
if args.threshold is None:
args.threshold = {"iqr": 1.5, "zscore": 3.0, "mzscore": 3.5}[args.method]
try:
headers, rows = load_csv(args.file)
except FileNotFoundError:
print(f"Error: file not found: {args.file}", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Error reading file: {e}", file=sys.stderr)
sys.exit(1)
if not rows:
print("Error: CSV file is empty.", file=sys.stderr)
sys.exit(1)
selected = args.columns.split(",") if args.columns else headers
missing_cols = [c for c in selected if c not in headers]
if missing_cols:
print(f"Error: columns not found: {', '.join(missing_cols)}", file=sys.stderr)
sys.exit(1)
results = []
for col in selected:
raw = [row.get(col, "") for row in rows]
nums = [n for v in raw if not is_null(v) and (n := to_float(v)) is not None]
results.append(analyze_column(col, nums, args.method, args.threshold))
if args.format == "json":
print(json.dumps(results, indent=2))
else:
print_report(results)
if __name__ == "__main__":
main()
Định giá DCF, mô hình tài chính, ngân sách, dự báo và chỉ số SaaS như ARR, MRR, churn, CAC, LTV, NRR.
--- name: cs-financial-analyst description: Financial Analyst agent for DCF valuation, financial modeling, budgeting, forecasting, and SaaS metrics (ARR, MRR, churn, CAC, LTV, NRR). Orchestrates finance skills. Spawn when users need financial analysis, valuation models, budget planning, ratio analysis, SaaS health checks, or unit economics projections. skills: finance domain: finance model: opus tools: [Read, Write, Bash, Grep, Glob] --- # cs-financial-analyst ## Role & Expertise Financial analyst covering valuation, ratio analysis, forecasting, and industry-specific financial modeling across SaaS, retail, manufacturing, healthcare, and financial services. ## Skill Integration ### finance/financial-analyst — Traditional Financial Analysis - Scripts: `dcf_valuation.py`, `ratio_calculator.py`, `forecast_builder.py`, `budget_variance_analyzer.py` - References: `financial-ratios-guide.md`, `valuation-methodology.md`, `forecasting-best-practices.md`, `industry-adaptations.md` ### finance/saas-metrics-coach — SaaS Financial Health - Scripts: `metrics_calculator.py`, `quick_ratio_calculator.py`, `unit_economics_simulator.py` - References: `formulas.md`, `benchmarks.md` - Assets: `input-template.md` ## Core Workflows ### 1. Company Valuation 1. Gather financial data (revenue, costs, growth rate, WACC) 2. Run DCF model via `dcf_valuation.py` 3. Calculate comparables (EV/EBITDA, P/E, EV/Revenue) 4. Adjust for industry via `industry-adaptations.md` 5. Present valuation range with sensitivity analysis ### 2. Financial Health Assessment 1. Run ratio analysis via `ratio_calculator.py` 2. Assess liquidity (current, quick ratio) 3. Assess profitability (gross margin, EBITDA margin, ROE) 4. Assess leverage (debt/equity, interest coverage) 5. Benchmark against industry standards ### 3. Revenue Forecasting 1. Analyze historical trends 2. Generate forecast via `forecast_builder.py` 3. Run scenarios (bull/base/bear) via `budget_variance_analyzer.py` 4. Calculate confidence intervals 5. Present with assumptions clearly stated ### 4. Budget Planning 1. Review prior year actuals 2. Set revenue targets by segment 3. Allocate costs by department 4. Build monthly cash flow projection 5. Define variance thresholds and review cadence ### 5. SaaS Health Check 1. Collect MRR, customer count, churn, CAC data from user 2. Run `metrics_calculator.py` to compute ARR, LTV, LTV:CAC, NRR, payback 3. Run `quick_ratio_calculator.py` if expansion/churn MRR available 4. Benchmark each metric against stage/segment via `benchmarks.md` 5. Flag CRITICAL/WATCH metrics and recommend top 3 actions ### 6. SaaS Unit Economics Projection 1. Take current MRR, growth rate, churn rate, CAC from user 2. Run `unit_economics_simulator.py` to project 12 months forward 3. Assess runway, profitability timeline, and growth trajectory 4. Cross-reference with `forecast_builder.py` for scenario modeling 5. Present monthly projections with summary and risk flags ## Output Standards - Valuations → range with methodology stated (DCF, comparables, precedent) - Ratios → benchmarked against industry with trend arrows - Forecasts → 3 scenarios with probability weights - All models include key assumptions section ## Success Metrics - **Forecast Accuracy:** Revenue forecasts within 5% of actuals over trailing 4 quarters - **Valuation Precision:** DCF valuations within 15% of market transaction comparables - **Budget Variance:** Departmental budgets maintained within 10% of plan - **Analysis Turnaround:** Financial models delivered within 48 hours of data receipt ## Integration Examples ```bash # SaaS health check — full metrics from raw numbers python ../../finance/saas-metrics-coach/scripts/metrics_calculator.py \ --mrr 80000 --mrr-last 75000 --customers 200 --churned 3 \ --new-customers 15 --sm-spend 25000 --gross-margin 72 --json # Quick ratio — growth efficiency python ../../finance/saas-metrics-coach/scripts/quick_ratio_calculator.py \ --new-mrr 10000 --expansion 2000 --churned 3000 --contraction 500 # 12-month projection python ../../finance/saas-metrics-coach/scripts/unit_economics_simulator.py \ --mrr 80000 --growth 8 --churn 1.5 --cac 1667 --json # Traditional ratio analysis python ../../finance/financial-analyst/scripts/ratio_calculator.py financial_data.json --format json # DCF valuation python ../../finance/financial-analyst/scripts/dcf_valuation.py valuation_data.json --format json ``` ## Related Agents - [cs-ceo-advisor](../c-level/cs-ceo-advisor.md) -- Strategic financial decisions, board reporting, and fundraising planning - [cs-growth-strategist](../business-growth/cs-growth-strategist.md) -- Revenue operations data and pipeline forecasting inputs
Audit mã nguồn trước khi lên production về bảo mật, CSDL, triển khai, chất lượng, AI/LLM, phụ thuộc và chặn deploy khi còn lỗi nghiêm trọng.
---
name: ship-gate
description: >
Pre-production audit that scans a codebase for security, database,
deployment, code quality, AI/LLM, dependency, frontend, and observability
issues. Intercepts deploy commands and blocks until critical items pass.
Stack-agnostic. Use for "run ship gate", "am I ready to ship",
"pre-launch audit", "can I deploy", "push to production", "go live
checklist", "preflight check". Not for CI/CD setup or infra provisioning.
license: MIT
metadata:
author: Rajaraman Arumugam
version: 1.0.0
---
# Ship Gate
Pre-production audit that scans a codebase and reports pass/fail/manual
across 8 categories before anything ships.
## Intercept Behavior
When the user says "push to production", "deploy", "ship it", "go live",
or similar deploy-intent phrases, do NOT proceed with deployment. Instead:
1. Ask: "Have you run the ship gate? Want me to scan now?"
2. If yes, run the full audit below.
3. If the user says they already ran it, ask when. If more than 24 hours
ago or if code changed since, recommend re-running.
## How It Works
### Step 1: Detect Stack
Run these checks in order to identify the project stack:
```
Framework detection:
package.json exists -> Node.js project
"next" in dependencies -> Next.js
"react" in dependencies -> React (if not Next.js)
"vue" in dependencies -> Vue
"svelte" in dependencies -> Svelte
"astro" in dependencies -> Astro
"express" in dependencies -> Express
"fastify" in dependencies -> Fastify
"hono" in dependencies -> Hono
requirements.txt or pyproject.toml -> Python project
"django" present -> Django
"flask" present -> Flask
"fastapi" present -> FastAPI
go.mod exists -> Go project
Cargo.toml exists -> Rust project
Database detection:
"@supabase/supabase-js" in package.json -> Supabase
supabase/ directory exists -> Supabase
"prisma" in dependencies -> Prisma (check schema for DB type)
"mongoose" in dependencies -> MongoDB
"pg" or "postgres" in dependencies -> PostgreSQL
firebase.json or .firebaserc exists -> Firebase
Deploy target detection:
vercel.json or .vercel/ exists -> Vercel
netlify.toml exists -> Netlify
Dockerfile exists -> Docker/VPS
fly.toml exists -> Fly.io
railway.json exists -> Railway
.platform/applications.yaml -> Platform.sh
Auth detection:
"@clerk" in dependencies -> Clerk
"next-auth" in dependencies -> NextAuth
"@supabase/auth-helpers" in deps -> Supabase Auth
"firebase/auth" in imports -> Firebase Auth
AI/LLM detection:
"openai" in dependencies -> OpenAI
"@anthropic-ai/sdk" in dependencies -> Claude API
"@google/generative-ai" in deps -> Gemini
```
Report detected stack before proceeding. This determines which checks
are relevant. Checks tagged with a specific stack in `references/checks.md`
are skipped if that stack is not detected.
### Step 2: Run Automated Checks
Run categories in this order: SEC, DB, CODE, DEP, AI, DEPLOY, FE, OBS.
Security and database first because they produce the most critical findings.
For each category, run every auto-scannable check from
`references/checks.md` using the patterns in `references/patterns.md`.
Report progress after each category completes:
```
[1/8] Security: 3 FAIL, 12 PASS, 3 SKIP
[2/8] Database: 1 FAIL, 5 PASS, 6 SKIP
...
```
Report results as:
- PASS: check passed
- FAIL: issue found (with file path and line number)
- SKIP: not applicable to this stack
### Step 3: Manual Confirmation
For checks that cannot be automated (backup restore tested, rollback plan
exists, staging test passed), present them as a checklist and ask the user
to confirm each one.
### Step 4: Verdict
Classify results into three severities:
- CRITICAL: must fix before shipping (secrets exposed, no auth on routes,
no HTTPS, SQL injection vectors, no RLS on Supabase tables)
- HIGH: should fix before shipping (no error boundaries, no rate limiting,
console.logs in production, no pagination)
- ADVISORY: recommended but not blocking (no OG tags, no custom 404,
no analytics, no SBOM)
Final output:
```
SHIP GATE REPORT
================
Stack: Next.js + Supabase + Vercel
Scan time: 12s
CRITICAL (3 items, must fix)
FAIL [SEC-01] API key found in src/lib/api.ts:14
FAIL [DB-07] RLS not enabled on "profiles" table
FAIL [SEC-05] No CSRF protection on /api/checkout
HIGH (5 items, should fix)
FAIL [CODE-01] 12 console.log statements in production code
FAIL [CODE-03] Empty catch block in src/utils/auth.ts:45
FAIL [DEP-04] 3 critical npm audit vulnerabilities
FAIL [DEPLOY-05] No rollback plan documented
MANUAL [DEPLOY-06] Staging test not confirmed
ADVISORY (4 items, recommended)
FAIL [FE-01] Missing OG meta tags
FAIL [FE-03] No custom 404 page
PASS [OBS-01] Error monitoring configured
SKIP [AI-01] No AI/LLM usage detected
VERDICT: DO NOT SHIP (3 critical issues)
Fix critical items and re-run.
```
If zero critical items remain, verdict is: CLEAR TO SHIP.
If only high items remain, verdict is: SHIP WITH CAUTION (acknowledge risks).
## Categories
Eight categories, each with a code prefix. Full check details in
`references/checks.md`.
| Prefix | Category | Auto | Manual | Tool |
|--------|----------|------|--------|------|
| SEC | Security | 15 | 3 | 0 |
| DB | Database | 7 | 5 | 0 |
| DEPLOY | Deployment | 3 | 8 | 0 |
| CODE | Code Quality | 11 | 0 | 1 |
| AI | AI/LLM Security | 5 | 3 | 0 |
| DEP | Dependencies | 5 | 0 | 1 |
| FE | Frontend Quality | 7 | 3 | 0 |
| OBS | Observability | 2 | 5 | 0 |
## Scope
This skill audits. It does not fix. When it finds issues, it reports
them with file locations and remediation guidance. The user or another
skill (systematic-debugging, backend-patterns, shadcn-stack) handles
the fix.
This skill does not:
- Set up CI/CD pipelines
- Provision infrastructure
- Configure monitoring tools
- Run after deployment (it is pre-deploy only)
## Integration Points
- **karpathy-coder**: run ship-gate after karpathy-check passes — simplicity first, then production readiness
- **adversarial-reviewer**: deep security review for items ship-gate flags as critical
- **security-pen-testing**: penetration testing methodology for SEC-category findings
- **code-reviewer**: general code quality review complements ship-gate's automated checks
FILE:references/checks.md
# Ship Gate: Complete Check Reference
All checks organized by category with ID, description, detection method,
severity, and remediation guidance.
## Table of Contents
- SEC: Security (18 checks)
- DB: Database (12 checks)
- DEPLOY: Deployment (13 checks)
- CODE: Code Quality (14 checks)
- AI: AI/LLM Security (8 checks)
- DEP: Dependencies and Supply Chain (7 checks)
- FE: Frontend Quality (10 checks)
- OBS: Observability (7 checks)
Detection methods:
- **auto**: Claude scans the codebase using grep, find, or file inspection
- **tool**: Claude runs an external tool (npm audit, etc.)
- **manual**: Claude asks the user to confirm
---
## SEC: Security
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| SEC-01 | No API keys or secrets in frontend code | auto | critical | all |
| SEC-02 | Every route checks authentication | auto | critical | all |
| SEC-03 | HTTPS enforced, HTTP redirected | manual | critical | all |
| SEC-04 | CORS locked to specific domain, not wildcard | auto | critical | all |
| SEC-05 | CSRF protection on state-changing endpoints | auto | critical | all |
| SEC-06 | Input validated and sanitized server-side | auto | high | all |
| SEC-07 | Rate limiting on auth and sensitive endpoints | auto | high | all |
| SEC-08 | Passwords hashed with bcrypt or argon2 | auto | critical | all |
| SEC-09 | Auth tokens have expiry | auto | high | all |
| SEC-10 | Sessions invalidated on logout (server-side) | manual | high | all |
| SEC-11 | CSP headers configured | auto | high | all |
| SEC-12 | JWT not using alg:none or weak secrets | auto | critical | all |
| SEC-13 | No eval() or dangerouslySetInnerHTML without sanitization | auto | high | js/ts |
| SEC-14 | No sensitive data in URL parameters or logs | auto | high | all |
| SEC-15 | Cookie security flags set (HttpOnly, Secure, SameSite) | auto | high | all |
| SEC-16 | File upload validates type, size, no path traversal | auto | high | all |
| SEC-17 | No hardcoded secrets in .env committed to repo | auto | critical | all |
| SEC-18 | .env files listed in .gitignore | auto | critical | all |
### SEC-01: No API keys or secrets in frontend code
Scan all files in src/, app/, pages/, public/, components/ for patterns
matching API keys, tokens, and secrets. See patterns.md for the full
regex list.
Remediation: Move secrets to environment variables. Use server-side API
routes to proxy requests that require secrets.
### SEC-02: Every route checks authentication
For Next.js: check middleware.ts/js exists and covers protected routes.
For Express: check that auth middleware is applied to route handlers.
For Django: check @login_required or permission decorators.
For generic: search for unprotected route definitions.
Remediation: Add authentication middleware. Audit every endpoint and
classify as public or protected.
### SEC-04: CORS not wildcard
Search for `cors({ origin: '*' })`, `Access-Control-Allow-Origin: *`,
or equivalent in the detected framework.
Remediation: Set CORS origin to your specific domain(s).
### SEC-05: CSRF protection
Check for CSRF token generation and validation on POST/PUT/DELETE routes.
For Next.js Server Actions, verify they use built-in CSRF protection.
Remediation: Add CSRF middleware or use framework-native CSRF protection.
### SEC-06: Input validation server-side
Search for request body usage (req.body, request.json, request.form)
without validation library imports (zod, yup, joi, class-validator,
pydantic). Check if raw user input flows directly into database queries
or business logic.
Remediation: Add input validation with zod, yup, or joi on every
endpoint that accepts user input.
### SEC-07: Rate limiting
Search for rate limiting middleware (express-rate-limit, @upstash/ratelimit,
rate-limiter-flexible, slowapi). Check auth routes and sensitive endpoints.
Remediation: Add rate limiting middleware. Start with auth endpoints
(login, register, password reset) and any endpoint that sends emails
or costs money.
### SEC-09: Auth token expiry
Search JWT sign calls for expiresIn/exp claims. Check if tokens are
created without expiry. Search for `sign(`, `jwt.encode(`, `createToken`.
Remediation: Set token expiry. Access tokens: 15-60 minutes.
Refresh tokens: 7-30 days. Never issue tokens without expiry.
### SEC-14: Sensitive data in URLs or logs
Search for query parameters containing keywords like password, token,
secret, key, ssn, credit_card. Search logging statements that log
full request objects or sensitive fields.
Remediation: Send sensitive data in request body or headers, never
in URL parameters. Redact sensitive fields before logging.
### SEC-16: File upload validation
Search for file upload handlers (multer, formidable, busboy,
UploadedFile). Check if file type, size, and path are validated.
Remediation: Validate file MIME type against an allowlist. Set
maximum file size. Sanitize filenames. Store outside webroot.
### SEC-12: JWT security
Search for `alg: 'none'`, `algorithm: 'none'`, or JWT secrets shorter
than 32 characters.
Remediation: Use RS256 or HS256 with a strong secret (32+ characters).
Never allow alg:none.
### SEC-17: No hardcoded secrets in .env committed
Check git history for .env files: `git log --all --name-only | grep .env`
Check if .env exists in the working tree and is not in .gitignore.
Remediation: Add .env* to .gitignore. Rotate any exposed secrets.
Use `git filter-branch` or BFG to remove from history if needed.
---
## DB: Database
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| DB-01 | Backups configured and tested | manual | critical | all |
| DB-02 | Backup restore tested (not just backup) | manual | critical | all |
| DB-03 | Parameterized queries everywhere | auto | critical | all |
| DB-04 | Separate dev and production databases | manual | high | all |
| DB-05 | Connection pooling configured | auto | high | all |
| DB-06 | Migrations in version control | auto | high | all |
| DB-07 | RLS enabled on all tables | auto | critical | supabase |
| DB-08 | No service_role key in client-side code | auto | critical | supabase |
| DB-09 | Anon key not used for writes without RLS | auto | high | supabase |
| DB-10 | Storage bucket policies configured | auto | high | supabase |
| DB-11 | App uses a non-root DB user | manual | high | all |
| DB-12 | No PII stored unencrypted | auto | high | all |
### DB-03: Parameterized queries
Search for string concatenation in SQL queries:
- Template literals with SQL keywords: `` `SELECT ... `"SELECT " + variable`
- f-strings with SQL: `f"SELECT ... {variable"`
Remediation: Use parameterized queries or ORM methods.
### DB-07: RLS enabled (Supabase)
Search migration files for `CREATE TABLE` without a corresponding
`ALTER TABLE ... ENABLE ROW LEVEL SECURITY` statement.
Also check for `CREATE POLICY` statements.
Remediation: Enable RLS on every table and create appropriate policies.
### DB-08: No service_role key in client code
Search frontend directories (src/, app/, components/, pages/) for
`service_role`, `supabase_service_role`, or the actual key pattern
`eyJ...` used with createClient on the client side.
Remediation: Use service_role only in server-side code (API routes,
Edge Functions, server actions).
### DB-05: Connection pooling
Search for database connection configuration. Check for pool settings
(max, min, idle timeout). For Supabase, check if using connection
pooler URL (port 6543) vs direct (port 5432).
Remediation: Use connection pooling for production. For Supabase,
use the pooler URL. For raw pg, configure pool size based on expected
concurrent connections.
### DB-06: Migrations in version control
Check if a migrations directory exists (supabase/migrations, prisma/
migrations, alembic/versions, db/migrate). Verify it contains .sql
or migration files, not empty.
Remediation: Use your ORM or database tool's migration system. Never
make manual schema changes to production.
### DB-09: Anon key writes without RLS
Search for Supabase client-side inserts/updates using the anon key
without RLS policies protecting the target tables.
Remediation: Enable RLS on all tables and create INSERT/UPDATE policies
that scope access to authenticated users.
### DB-10: Storage bucket policies
Search Supabase migration files and dashboard config for storage
bucket creation. Verify each bucket has access policies defined.
Remediation: Define storage policies for each bucket. Restrict
uploads by file type, size, and user ownership.
### DB-12: PII stored unencrypted
Search schema files and migration files for columns named email,
phone, ssn, social_security, credit_card, address, date_of_birth
that are stored as plain text without encryption.
Remediation: Encrypt PII columns at rest. Use database-level
encryption or application-level encryption for sensitive fields.
---
## DEPLOY: Deployment
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| DEPLOY-01 | All env vars set on production server | manual | critical | all |
| DEPLOY-02 | SSL certificate installed and valid | manual | critical | all |
| DEPLOY-03 | Firewall configured (only 80/443 public) | manual | high | vps |
| DEPLOY-04 | Process manager running | manual | high | vps |
| DEPLOY-05 | Rollback plan exists | manual | high | all |
| DEPLOY-06 | Staging test passed before production | manual | high | all |
| DEPLOY-07 | Deploy does not cause downtime | manual | advisory | all |
| DEPLOY-08 | Domain DNS configured (www vs non-www) | manual | high | all |
| DEPLOY-09 | Health check endpoint exists | auto | high | all |
| DEPLOY-10 | Logging configured (structured, not console) | auto | high | all |
| DEPLOY-11 | Error monitoring connected (Sentry, etc.) | auto | advisory | all |
| DEPLOY-12 | Cron jobs and background tasks verified | manual | high | all |
| DEPLOY-13 | CDN configured for static assets | manual | advisory | all |
### DEPLOY-09: Health check endpoint
Search for a `/health`, `/healthz`, `/api/health`, or `/status` route
that returns a 200 response.
Remediation: Add a health check endpoint that verifies database
connectivity and returns a simple JSON response.
### DEPLOY-10: Structured logging
Check if the project uses a logging library (winston, pino, bunyan,
python logging module) vs raw console.log statements in server code.
Remediation: Replace console.log with a structured logger that outputs
JSON with timestamps and request IDs.
---
## CODE: Code Quality
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| CODE-01 | No console.log in production build | auto | high | js/ts |
| CODE-02 | Error handling on all async operations | auto | high | all |
| CODE-03 | No empty catch blocks | auto | high | all |
| CODE-04 | Loading and error states in UI | auto | high | react |
| CODE-05 | Pagination on all list endpoints | auto | high | all |
| CODE-06 | npm audit clean (zero critical) | tool | high | js/ts |
| CODE-07 | No TODO-auth or TODO-security patterns | auto | critical | all |
| CODE-08 | No unhandled promise rejections | auto | high | js/ts |
| CODE-09 | React error boundaries in place | auto | high | react |
| CODE-10 | No leaked stack traces in error responses | auto | high | all |
| CODE-11 | No eslint-disable on security rules | auto | high | js/ts |
| CODE-12 | Lockfile committed | auto | high | all |
| CODE-13 | No wildcard versions in package.json | auto | high | js/ts |
| CODE-14 | TypeScript strict mode enabled | auto | advisory | ts |
### CODE-01: No console.log in production
Search for `console.log`, `console.debug`, `console.info` in source
files (exclude test files, config files, and node_modules).
Remediation: Remove console.log statements or replace with a proper
logger. Use a build tool to strip them automatically.
### CODE-03: No empty catch blocks
Search for `catch` blocks with empty bodies or only a comment inside.
Pattern: `catch\s*\([^)]*\)\s*\{\s*(\/\/.*\n)?\s*\}`
Remediation: At minimum, log the error. Better: handle it appropriately
or rethrow.
### CODE-07: No TODO-auth/security patterns
Search for `TODO.*auth`, `TODO.*security`, `TODO.*permission`,
`FIXME.*auth`, `HACK.*auth`, `// auth`, `# TODO: add auth`.
These indicate security features that were deferred and forgotten.
Remediation: Implement the deferred security feature or remove the
endpoint if it is not ready.
### CODE-09: React error boundaries
Check if the app has at least one ErrorBoundary component or uses
a library like react-error-boundary. Check app/error.tsx for Next.js
App Router projects.
Remediation: Add error boundaries at layout boundaries to prevent
full-page crashes.
### CODE-02: Error handling on async operations
Search for async functions and .then() chains. Check if they have
corresponding try/catch or .catch() handlers.
Remediation: Wrap every async operation in try/catch. Log errors
and show appropriate UI feedback.
### CODE-04: Loading and error states in UI
Search React components for data fetching (useEffect with fetch,
useSWR, useQuery, server components) and check if they render
loading and error states.
Remediation: Add loading spinners/skeletons and error messages
for every data-dependent component.
### CODE-05: Pagination on list endpoints
Search API routes that return arrays/lists from database queries.
Check for LIMIT/OFFSET, cursor pagination, or take/skip parameters.
Remediation: Add pagination to every endpoint that returns a list.
Default page size of 20-50 items. Never return unbounded result sets.
### CODE-10: No leaked stack traces
Search error handling code for responses that include stack traces,
error.stack, or full error objects sent to the client.
Remediation: Return generic error messages to the client. Log full
stack traces server-side only.
### CODE-11: No eslint-disable on security rules
Search for eslint-disable comments that suppress security-related
rules (no-eval, no-implied-eval, no-script-url).
Remediation: Fix the underlying issue instead of disabling the lint
rule. If genuinely necessary, add a comment explaining why.
### CODE-14: TypeScript strict mode
Check tsconfig.json for `"strict": true` or the individual flags
(strictNullChecks, noImplicitAny, etc.).
Remediation: Enable strict mode in tsconfig.json. Fix type errors
incrementally if migrating an existing project.
---
## AI: AI/LLM Security
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| AI-01 | System prompts not leakable via user input | auto | critical | ai |
| AI-02 | No prompt injection vectors in user inputs | auto | critical | ai |
| AI-03 | LLM API keys not in frontend code | auto | critical | ai |
| AI-04 | Rate limiting on AI endpoints (cost protection) | auto | high | ai |
| AI-05 | AI response output sanitized before rendering | auto | high | ai |
| AI-06 | MCP server inputs validated | auto | high | ai |
| AI-07 | Agent permissions scoped (no unrestricted access) | manual | high | ai |
| AI-08 | No sensitive data sent to third-party LLMs without consent | manual | high | ai |
### AI-01: System prompt leakage
Search for system prompts stored in client-accessible files or returned
in API responses. Check if the AI endpoint echoes the system prompt
when asked "repeat your instructions" or similar.
Remediation: Keep system prompts server-side only. Add input filtering
for prompt extraction attempts.
### AI-03: LLM API keys not in frontend
Search frontend code for `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`,
`GOOGLE_AI_API_KEY`, `sk-ant-`, `sk-proj-`, `AIza` patterns.
Remediation: Proxy all LLM calls through server-side API routes.
---
## DEP: Dependencies and Supply Chain
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| DEP-01 | No git:// or URL-based dependencies | auto | high | all |
| DEP-02 | No typosquatting risk (verify package names) | auto | advisory | all |
| DEP-03 | Lockfile integrity verified | auto | high | all |
| DEP-04 | npm audit / pip audit zero critical | tool | high | all |
| DEP-05 | No suspicious postinstall scripts | auto | high | js/ts |
| DEP-06 | Dependencies pinned (no wildcard *) | auto | high | all |
| DEP-07 | Lockfile committed to version control | auto | high | all |
### DEP-01: No git/URL dependencies
Search package.json for dependencies with values starting with
`git://`, `git+`, `http://`, `https://github.com`, or `file:`.
Remediation: Use published npm packages with version ranges instead
of git URLs.
### DEP-05: Suspicious postinstall scripts
Check package.json for `postinstall`, `preinstall`, `install` scripts
that execute arbitrary commands, download files, or access the network.
Remediation: Review and remove unnecessary install scripts. Use
`--ignore-scripts` for CI.
---
## FE: Frontend Quality
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| FE-01 | Meta tags present (title, description, OG tags) | auto | advisory | web |
| FE-02 | Favicon configured | auto | advisory | web |
| FE-03 | Custom 404 page exists | auto | advisory | web |
| FE-04 | Responsive design tested on mobile | manual | high | web |
| FE-05 | Alt text on images | auto | high | web |
| FE-06 | Keyboard navigation works | manual | high | web |
| FE-07 | Forms have validation feedback | auto | high | web |
| FE-08 | Analytics installed (production only) | auto | advisory | web |
| FE-09 | robots.txt present | auto | advisory | web |
| FE-10 | Images optimized (WebP, lazy loading) | auto | advisory | web |
### FE-01: Meta tags
Check the root layout or index page for `<title>`, `<meta name="description">`,
and Open Graph tags (`og:title`, `og:description`, `og:image`).
For Next.js, check metadata export in layout.tsx.
Remediation: Add metadata to your root layout or page head.
### FE-03: Custom 404 page
Check for `404.tsx`, `404.jsx`, `not-found.tsx`, `404.html`, or
equivalent in the pages/app directory.
Remediation: Create a branded 404 page that helps users navigate back.
---
## OBS: Observability
| ID | Check | Detection | Severity | Stack |
|----|-------|-----------|----------|-------|
| OBS-01 | Error monitoring configured (Sentry, LogRocket, etc.) | auto | advisory | all |
| OBS-02 | Alerting set up for critical failures | manual | high | all |
| OBS-03 | Structured logging with request IDs | auto | advisory | all |
| OBS-04 | Performance baseline established | manual | advisory | all |
| OBS-05 | Uptime monitoring configured | manual | high | all |
| OBS-06 | Error rates tracked | manual | advisory | all |
| OBS-07 | Log retention policy defined | manual | advisory | all |
### OBS-01: Error monitoring
Search for imports or configuration of error monitoring tools:
`@sentry/`, `LogRocket`, `Bugsnag`, `Datadog`, `Rollbar`, `Honeybadger`.
Remediation: Install and configure an error monitoring service.
Sentry has a free tier suitable for solo projects.
FILE:references/patterns.md
# Ship Gate: Detection Patterns
Grep and regex patterns for auto-scannable checks. Claude runs these
against the codebase to detect issues.
## Table of Contents
- SEC: Security Patterns
- DB: Database Patterns
- CODE: Code Quality Patterns
- AI: AI/LLM Security Patterns
- DEP: Dependency Patterns
- FE: Frontend Quality Patterns
- OBS: Observability Patterns
- DEPLOY: Deployment Patterns
All patterns use `grep -rn` with `--include` filters. Exclude
node_modules, .next, dist, build, .git, __pycache__, venv directories
from all scans.
Base exclude flags:
```bash
EXCLUDE="--exclude-dir=node_modules --exclude-dir=.next --exclude-dir=dist --exclude-dir=build --exclude-dir=.git --exclude-dir=__pycache__ --exclude-dir=venv --exclude-dir=.venv --exclude-dir=vendor --exclude-dir=coverage"
```
---
## SEC: Security Patterns
### SEC-01: Secrets in frontend code
Scan directories that serve client-side code:
```bash
# Generic API key patterns
grep -rnE $EXCLUDE \
"(sk-[a-zA-Z0-9]{20,}|sk-ant-[a-zA-Z0-9-]+|sk-proj-[a-zA-Z0-9-]+|AIza[a-zA-Z0-9_-]{35}|ghp_[a-zA-Z0-9]{36}|glpat-[a-zA-Z0-9_-]{20,}|xox[bsap]-[a-zA-Z0-9-]+)" \
src/ app/ pages/ components/ public/ lib/ utils/ 2>/dev/null
# AWS keys
grep -rnE $EXCLUDE \
"AKIA[0-9A-Z]{16}" \
src/ app/ pages/ components/ public/ 2>/dev/null
# Stripe keys (live, not test)
grep -rnE $EXCLUDE \
"sk_live_[a-zA-Z0-9]{24,}" \
src/ app/ pages/ components/ public/ 2>/dev/null
# Generic secret assignment
grep -rnE $EXCLUDE \
"(api_key|apikey|api_secret|secret_key|auth_token|access_token)\s*[:=]\s*['\"][a-zA-Z0-9_-]{16,}" \
src/ app/ pages/ components/ public/ 2>/dev/null
```
### SEC-04: CORS wildcard
```bash
grep -rnE $EXCLUDE \
"(origin:\s*['\"]?\*['\"]?|Access-Control-Allow-Origin.*\*|cors\(\s*\))" \
. 2>/dev/null
```
### SEC-05: CSRF protection missing
```bash
# Check for state-changing routes without CSRF
grep -rnE $EXCLUDE \
"(app\.(post|put|patch|delete)|router\.(post|put|patch|delete))" \
. 2>/dev/null
# Then verify csrf middleware exists
grep -rnE $EXCLUDE \
"(csrf|csrfToken|_csrf|CSRF_COOKIE)" \
. 2>/dev/null
```
### SEC-08: Weak password hashing
```bash
# Check for weak hashing (md5, sha1, sha256 for passwords)
grep -rnE $EXCLUDE \
"(md5|sha1|sha256)\s*\(" \
. 2>/dev/null
# Verify bcrypt/argon2 usage
grep -rnE $EXCLUDE \
"(bcrypt|argon2|scrypt)" \
. 2>/dev/null
```
### SEC-11: CSP headers
```bash
# Check for Content-Security-Policy configuration
grep -rnE $EXCLUDE \
"(Content-Security-Policy|contentSecurityPolicy|csp)" \
. 2>/dev/null
# Next.js: check next.config for headers
grep -rn $EXCLUDE \
"Content-Security-Policy" \
next.config.* 2>/dev/null
```
### SEC-13: Unsafe eval/innerHTML
```bash
# eval usage
grep -rnE $EXCLUDE \
"(\beval\s*\(|new\s+Function\s*\()" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# dangerouslySetInnerHTML without sanitizer
grep -rnE $EXCLUDE \
"dangerouslySetInnerHTML" \
--include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# Then check if DOMPurify or similar is imported in same file
```
### SEC-15: Cookie security flags
```bash
grep -rnE $EXCLUDE \
"(set-cookie|setCookie|cookie\()" \
. 2>/dev/null
# Verify HttpOnly, Secure, SameSite flags are present
grep -rnE $EXCLUDE \
"(httpOnly|HttpOnly|secure:\s*true|sameSite)" \
. 2>/dev/null
```
### SEC-06: Input validation
```bash
# Check for validation library usage
grep -rnE $EXCLUDE \
"(from 'zod'|from 'yup'|from 'joi'|from 'class-validator'|from pydantic)" \
. 2>/dev/null
# Check for raw req.body usage without validation
grep -rnE $EXCLUDE \
"(req\.body\.|request\.json|request\.form)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### SEC-07: Rate limiting
```bash
grep -rnE $EXCLUDE \
"(express-rate-limit|@upstash/ratelimit|rate-limiter|slowapi|throttle)" \
package.json requirements.txt . 2>/dev/null
```
### SEC-09: Token expiry
```bash
grep -rnE $EXCLUDE \
"(sign\(|jwt\.encode|createToken|signToken)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
# Then check if expiresIn/exp is set in those calls
grep -rnE $EXCLUDE \
"(expiresIn|exp:|expires_in|expires_delta)" \
. 2>/dev/null
```
### SEC-14: Sensitive data in URLs/logs
```bash
# Sensitive query parameters
grep -rnE $EXCLUDE \
"(password|token|secret|key|ssn|credit.card)=" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
# Logging full request objects
grep -rnE $EXCLUDE \
"console\.(log|info|debug)\s*\(\s*(req|request)\s*\)" \
. 2>/dev/null
```
### SEC-16: File upload validation
```bash
grep -rnE $EXCLUDE \
"(multer|formidable|busboy|UploadedFile|upload\.single|upload\.array)" \
. 2>/dev/null
# Check for file type/size validation near upload handlers
grep -rnE $EXCLUDE \
"(fileFilter|limits|maxFileSize|allowedTypes|mimetype)" \
. 2>/dev/null
```
### SEC-17/18: .env in repo
```bash
# Check if .env files exist in working tree
find . -maxdepth 3 -name ".env*" -not -path "*/node_modules/*" \
-not -name ".env.example" -not -name ".env.sample" 2>/dev/null
# Check if .env is in .gitignore
grep -n "\.env" .gitignore 2>/dev/null
# Check git history for .env commits
git log --all --name-only --diff-filter=A 2>/dev/null | grep "\.env" || true
```
---
## DB: Database Patterns
### DB-03: SQL injection (string concatenation)
```bash
# Template literal SQL
grep -rnE $EXCLUDE \
"(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE).*\$\{" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# String concat SQL
grep -rnE $EXCLUDE \
"(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE).*\+\s*(req\.|params\.|body\.|query\.)" \
. 2>/dev/null
# Python f-string SQL
grep -rnE $EXCLUDE \
"f['\"].*\b(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE)\b.*\{" \
--include="*.py" \
. 2>/dev/null
```
### DB-07: Supabase RLS
```bash
# Find CREATE TABLE without RLS
grep -rnl $EXCLUDE "CREATE TABLE" \
--include="*.sql" . 2>/dev/null | while read f; do
tables=$(grep -oP "CREATE TABLE\s+\K\S+" "$f")
for t in $tables; do
if ! grep -q "ENABLE ROW LEVEL SECURITY" "$f" || \
! grep -q "$t" <<< "$(grep 'ENABLE ROW LEVEL SECURITY' "$f")"; then
echo "FAIL: $f - table $t missing RLS"
fi
done
done
```
### DB-08: service_role in client code
```bash
grep -rnE $EXCLUDE \
"(service_role|serviceRole|SUPABASE_SERVICE_ROLE)" \
src/ app/ pages/ components/ public/ lib/client 2>/dev/null
```
### DB-05: Connection pooling
```bash
# Check for pool configuration
grep -rnE $EXCLUDE \
"(pool|connectionLimit|max_connections|poolSize)" \
--include="*.ts" --include="*.js" --include="*.py" --include="*.env*" \
. 2>/dev/null
# Supabase: check if using pooler port
grep -rnE $EXCLUDE \
"(6543|pooler)" \
--include="*.env*" --include="*.ts" --include="*.js" \
. 2>/dev/null
```
### DB-06: Migrations in version control
```bash
# Check for migration directories
find . -maxdepth 3 -type d \
\( -name "migrations" -o -name "migrate" -o -name "versions" \) \
-not -path "*/node_modules/*" 2>/dev/null
# Check if migrations contain files
find . -path "*/migrations/*.sql" -o -path "*/migrations/*.ts" \
-o -path "*/migrations/*.py" 2>/dev/null | head -5
```
### DB-12: PII stored unencrypted
```bash
# Search schema files for PII column names
grep -rnEi $EXCLUDE \
"(ssn|social_security|credit_card|card_number|passport)" \
--include="*.sql" --include="*.prisma" --include="*.py" \
. 2>/dev/null
```
---
## CODE: Code Quality Patterns
### CODE-01: console.log in production
```bash
grep -rnE $EXCLUDE \
"console\.(log|debug|info)\(" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
--exclude="*.test.*" --exclude="*.spec.*" --exclude="*.config.*" \
src/ app/ pages/ components/ lib/ utils/ 2>/dev/null
```
### CODE-03: Empty catch blocks
```bash
grep -rnPzo $EXCLUDE \
"catch\s*\([^)]*\)\s*\{\s*\}" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
```
### CODE-07: TODO-auth patterns
```bash
grep -rnEi $EXCLUDE \
"(TODO|FIXME|HACK|XXX).*(auth|security|permission|validation|sanitiz)" \
. 2>/dev/null
```
### CODE-08: Unhandled promise rejections
```bash
# Async functions without try-catch
grep -rnE $EXCLUDE \
"async\s+\w+\s*\(" \
--include="*.js" --include="*.ts" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# Check for .catch() or try/catch wrapping
```
### CODE-09: React error boundaries
```bash
# Check for error boundary in Next.js App Router
find . -path "*/app/error.tsx" -o -path "*/app/error.jsx" \
-o -path "*/app/global-error.tsx" 2>/dev/null
# Check for ErrorBoundary component
grep -rnE $EXCLUDE \
"(ErrorBoundary|error-boundary|componentDidCatch|getDerivedStateFromError)" \
--include="*.jsx" --include="*.tsx" \
. 2>/dev/null
```
### CODE-12: Lockfile committed
```bash
# Check for lockfile existence
ls package-lock.json pnpm-lock.yaml yarn.lock bun.lockb \
Pipfile.lock poetry.lock Gemfile.lock go.sum Cargo.lock 2>/dev/null
# Check if lockfile is gitignored
for f in package-lock.json pnpm-lock.yaml yarn.lock; do
if git check-ignore "$f" 2>/dev/null; then
echo "FAIL: $f is gitignored"
fi
done
```
### CODE-13: Wildcard versions
```bash
# Check for * or empty version in package.json
grep -nE '"[^"]+"\s*:\s*"\*"' package.json 2>/dev/null
```
### CODE-02: Async without error handling
```bash
# Find async functions
grep -rnE $EXCLUDE \
"async\s+(function\s+)?\w+\s*\(" \
--include="*.ts" --include="*.js" --include="*.tsx" --include="*.jsx" \
. 2>/dev/null
# Count try/catch usage nearby
grep -rnc $EXCLUDE "try\s*{" \
--include="*.ts" --include="*.js" --include="*.tsx" --include="*.jsx" \
. 2>/dev/null
```
### CODE-04: Loading and error states
```bash
# Check for loading state patterns in React
grep -rnE $EXCLUDE \
"(isLoading|loading|Skeleton|Spinner|fallback)" \
--include="*.tsx" --include="*.jsx" \
. 2>/dev/null
# Check for Suspense boundaries
grep -rnE $EXCLUDE \
"(<Suspense|loading\.tsx|loading\.jsx)" \
. 2>/dev/null
```
### CODE-05: Pagination on list endpoints
```bash
# Check API routes for unbounded queries
grep -rnE $EXCLUDE \
"(\.findMany|\.find\(\)|\.select\(\)|SELECT \*)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
# Check for pagination parameters
grep -rnE $EXCLUDE \
"(limit|offset|page|skip|take|cursor|per_page)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### CODE-10: Leaked stack traces
```bash
grep -rnE $EXCLUDE \
"(error\.stack|\.stack\)|err\.message.*res\.(json|send)|traceback)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### CODE-11: eslint-disable on security rules
```bash
grep -rnE $EXCLUDE \
"eslint-disable.*(no-eval|no-implied-eval|no-script-url|security)" \
--include="*.ts" --include="*.js" --include="*.tsx" --include="*.jsx" \
. 2>/dev/null
```
### CODE-14: TypeScript strict mode
```bash
grep -n '"strict"' tsconfig.json 2>/dev/null
# Check if strict is true
grep -n '"strict":\s*true' tsconfig.json 2>/dev/null
```
---
## AI: AI/LLM Security Patterns
### AI-01: System prompt leakage
```bash
# System prompts in client-accessible files
grep -rnEi $EXCLUDE \
"(system.?prompt|system.?message|system_instruction)" \
src/ app/ pages/ components/ public/ 2>/dev/null
# System prompts returned in API responses
grep -rnE $EXCLUDE \
"(system.*role|role.*system)" \
src/ app/ pages/ components/ public/ 2>/dev/null
```
### AI-02: Prompt injection vectors
```bash
# User input concatenated directly into prompts
grep -rnE $EXCLUDE \
"(messages\.push|content:.*\$\{|content:.*\+\s*user|prompt.*\+)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### AI-03: LLM API keys in frontend
```bash
grep -rnE $EXCLUDE \
"(OPENAI_API_KEY|ANTHROPIC_API_KEY|GOOGLE_AI_API_KEY|sk-ant-|sk-proj-|AIza[a-zA-Z0-9_-]{35})" \
src/ app/ pages/ components/ public/ 2>/dev/null
```
### AI-04: Rate limiting on AI endpoints
```bash
# Find AI-related API routes
grep -rnlE $EXCLUDE \
"(openai|anthropic|claude|gpt|completion|chat/api|ai/api)" \
--include="*.ts" --include="*.js" \
. 2>/dev/null
# Then check for rate limiting middleware in those files
```
### AI-05: AI output sanitization
```bash
# Check if AI responses are rendered with dangerouslySetInnerHTML
grep -rnE $EXCLUDE \
"dangerouslySetInnerHTML.*\b(response|result|completion|message|content)\b" \
--include="*.tsx" --include="*.jsx" \
. 2>/dev/null
```
### AI-06: MCP server input validation
```bash
# Check MCP server tool handlers for input validation
grep -rnE $EXCLUDE \
"(tool_input|toolInput|tool_call|CallToolRequest)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
# Check if zod/validation is applied to tool inputs
```
---
## DEP: Dependency Patterns
### DEP-01: Git/URL dependencies
```bash
grep -nE '"(git|git\+|http|https|file):' package.json 2>/dev/null
grep -nE '"github:' package.json 2>/dev/null
```
### DEP-04: npm audit
```bash
# Run npm audit and capture critical/high counts
npm audit --json 2>/dev/null | grep -c '"severity":"critical"'
npm audit --json 2>/dev/null | grep -c '"severity":"high"'
# Or for pip
pip audit --format json 2>/dev/null
```
### DEP-05: Suspicious install scripts
```bash
grep -A2 '"preinstall"\|"postinstall"\|"install"' package.json 2>/dev/null
```
### DEP-06: Wildcard versions
```bash
grep -nE '"\*"' package.json 2>/dev/null
grep -nE '"latest"' package.json 2>/dev/null
```
---
## FE: Frontend Quality Patterns
### FE-01: Meta tags
```bash
# Next.js App Router metadata
grep -rnE $EXCLUDE \
"(export\s+(const|async\s+function)\s+metadata|generateMetadata)" \
--include="*.tsx" --include="*.ts" \
app/layout.* app/page.* 2>/dev/null
# HTML meta tags
grep -rnE $EXCLUDE \
'(<title>|<meta\s+name="description"|og:title|og:description|og:image)' \
. 2>/dev/null
```
### FE-02: Favicon
```bash
find . -maxdepth 3 \( -name "favicon.*" -o -name "icon.*" \) \
-not -path "*/node_modules/*" 2>/dev/null
```
### FE-03: Custom 404 page
```bash
find . -maxdepth 4 \( -name "404.*" -o -name "not-found.*" \) \
-not -path "*/node_modules/*" 2>/dev/null
```
### FE-05: Image alt text
```bash
# Find img tags without alt attribute
grep -rnE $EXCLUDE \
'<img\s+(?![^>]*\balt\b)[^>]*>' \
--include="*.html" --include="*.jsx" --include="*.tsx" \
. 2>/dev/null
# Next.js Image without alt
grep -rnE $EXCLUDE \
'<Image\s+(?![^>]*\balt\b)[^>]*/?>' \
--include="*.jsx" --include="*.tsx" \
. 2>/dev/null
```
### FE-09: robots.txt
```bash
find . -maxdepth 2 -name "robots.txt" \
-not -path "*/node_modules/*" 2>/dev/null
```
### FE-07: Form validation feedback
```bash
# Check for form elements without validation attributes
grep -rnE $EXCLUDE \
'(<input|<textarea|<select)' \
--include="*.tsx" --include="*.jsx" --include="*.html" \
. 2>/dev/null
# Check for validation library usage
grep -rnE $EXCLUDE \
"(useForm|react-hook-form|formik|yup|zod.*form)" \
--include="*.tsx" --include="*.jsx" \
. 2>/dev/null
```
### FE-10: Image optimization
```bash
# Check for unoptimized img tags (not using Next/Image or similar)
grep -rnE $EXCLUDE \
'<img\s' \
--include="*.tsx" --include="*.jsx" \
. 2>/dev/null
# Check for lazy loading
grep -rnE $EXCLUDE \
'(loading="lazy"|lazy|lazyload)' \
--include="*.tsx" --include="*.jsx" --include="*.html" \
. 2>/dev/null
```
---
## OBS: Observability Patterns
### OBS-01: Error monitoring
```bash
grep -rnE $EXCLUDE \
"(@sentry|sentry-|LogRocket|Bugsnag|datadogRum|Rollbar|Honeybadger|newrelic)" \
package.json . 2>/dev/null
```
### OBS-03: Structured logging
```bash
# Check for logging libraries
grep -rnE $EXCLUDE \
"(winston|pino|bunyan|morgan|log4js)" \
package.json 2>/dev/null
# Python
grep -rnE $EXCLUDE \
"import logging|from loguru" \
--include="*.py" . 2>/dev/null
```
---
## DEPLOY: Deployment Patterns
### DEPLOY-09: Health check endpoint
```bash
grep -rnE $EXCLUDE \
"(\/health|\/healthz|\/api\/health|\/status|\/readyz)" \
--include="*.ts" --include="*.js" --include="*.py" \
. 2>/dev/null
```
### DEPLOY-10: Console vs structured logging (server)
```bash
# Count console.log vs logger usage in API/server code
echo "console.log count:"
grep -rnc $EXCLUDE "console\.log" \
--include="*.ts" --include="*.js" \
api/ server/ pages/api/ app/api/ 2>/dev/null | tail -1
echo "structured logger count:"
grep -rnc $EXCLUDE "(logger\.|log\.(info|warn|error|debug))" \
--include="*.ts" --include="*.js" \
api/ server/ pages/api/ app/api/ 2>/dev/null | tail -1
```
FILE:scripts/ship_gate_scanner.py
#!/usr/bin/env python3
"""
ship_gate_scanner.py — Pre-production audit CLI
Part of the ship-gate skill: https://github.com/rx4u/ship-gate
Usage:
python scripts/ship_gate_scanner.py [PATH] [options]
Options:
--json Output results as JSON
--no-color Disable ANSI color output
--no-interactive Skip manual confirmation prompts
--category CAT Only run a specific category (SEC, DB, CODE, etc.)
--verbose Show PASS results in addition to FAIL
--version Show version and exit
Exit codes:
0 = CLEAR TO SHIP (no critical issues)
1 = DO NOT SHIP (critical issues found)
2 = SHIP WITH CAUTION (high issues only)
"""
import argparse
import json
import os
import re
import sys
import time
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
from typing import List, Optional
VERSION = "1.0.0"
EXCLUDE_DIRS = {
"node_modules", ".next", "dist", "build", ".git", "__pycache__",
"venv", ".venv", "vendor", "coverage", ".turbo", "out", ".cache",
".pytest_cache", ".mypy_cache", "target", "bin", "obj",
}
FRONTEND_DIRS = {"src", "app", "pages", "components", "public", "lib", "utils"}
JS_EXTS = {".js", ".ts", ".jsx", ".tsx", ".mjs", ".cjs"}
PY_EXTS = {".py"}
ALL_CODE_EXTS = JS_EXTS | PY_EXTS | {".go", ".rb", ".php"}
TEMPLATE_EXTS = {".html", ".jsx", ".tsx", ".vue", ".svelte"}
SQL_EXTS = {".sql", ".prisma"}
# ---------------------------------------------------------------------------
# ANSI helpers
# ---------------------------------------------------------------------------
USE_COLOR = True
def _c(code: str, text: str) -> str:
if not USE_COLOR:
return text
return f"\033[{code}m{text}\033[0m"
def red(t): return _c("31", t)
def green(t): return _c("32", t)
def yellow(t): return _c("33", t)
def cyan(t): return _c("36", t)
def bold(t): return _c("1", t)
def dim(t): return _c("2", t)
# ---------------------------------------------------------------------------
# Data model
# ---------------------------------------------------------------------------
class Status(str, Enum):
PASS = "PASS"
FAIL = "FAIL"
SKIP = "SKIP"
MANUAL = "MANUAL"
class Severity(str, Enum):
CRITICAL = "CRITICAL"
HIGH = "HIGH"
ADVISORY = "ADVISORY"
@dataclass
class Finding:
file: str
line: int
snippet: str = ""
@dataclass
class CheckDef:
id: str
description: str
severity: Severity
category: str
stack: str = "all" # "all", "js", "ts", "react", "supabase", "ai", "web", "vps"
@dataclass
class Result:
check: CheckDef
status: Status
message: str = ""
findings: List[Finding] = field(default_factory=list)
@dataclass
class Stack:
has_node: bool = False
framework: str = "" # next, react, vue, svelte, astro, express, fastify, hono
has_python: bool = False
py_framework: str = "" # django, flask, fastapi
has_go: bool = False
has_rust: bool = False
has_supabase: bool = False
has_typescript: bool = False
has_react: bool = False
deploy_target: str = "" # vercel, netlify, docker, fly, railway
has_ai: bool = False
ai_providers: List[str] = field(default_factory=list)
is_web: bool = False
# ---------------------------------------------------------------------------
# File walking / grep helpers
# ---------------------------------------------------------------------------
def walk_files(root: str, exts: Optional[set] = None, dirs: Optional[set] = None):
"""Yield (filepath, relpath) for all files under root, skipping EXCLUDE_DIRS."""
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = [d for d in dirnames if d not in EXCLUDE_DIRS]
if dirs is not None:
rel = os.path.relpath(dirpath, root)
top = rel.split(os.sep)[0]
if rel != "." and top not in dirs:
dirnames[:] = []
continue
for fname in filenames:
if exts is None or os.path.splitext(fname)[1].lower() in exts:
fpath = os.path.join(dirpath, fname)
yield fpath, os.path.relpath(fpath, root)
def grep_files(
root: str,
pattern: str,
exts: Optional[set] = None,
dirs: Optional[set] = None,
flags: int = 0,
max_findings: int = 20,
exclude_patterns: Optional[List[str]] = None,
) -> List[Finding]:
"""Return up to max_findings matches across the codebase."""
try:
rx = re.compile(pattern, flags)
except re.error:
return []
exclude_rxs = []
if exclude_patterns:
for ep in exclude_patterns:
try:
exclude_rxs.append(re.compile(ep))
except re.error:
pass
results: List[Finding] = []
for fpath, relpath in walk_files(root, exts, dirs):
if any(seg in fpath for seg in (".test.", ".spec.", ".config.")):
if exts and exts <= JS_EXTS:
skip = True
# still yield for config-specific checks
if "tsconfig" in fpath or "package.json" in fpath:
skip = False
if skip:
continue
try:
with open(fpath, "r", encoding="utf-8", errors="ignore") as fh:
for lineno, line in enumerate(fh, 1):
if rx.search(line):
if any(ex.search(line) for ex in exclude_rxs):
continue
results.append(Finding(
file=relpath,
line=lineno,
snippet=line.rstrip()[:120],
))
if len(results) >= max_findings:
return results
except (OSError, PermissionError):
continue
return results
def file_exists_in(root: str, *names: str) -> Optional[str]:
"""Return the first found path among names (searched recursively up to depth 5)."""
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = [d for d in dirnames if d not in EXCLUDE_DIRS]
depth = dirpath.replace(root, "").count(os.sep)
if depth >= 5:
dirnames[:] = []
continue
for fname in filenames:
if fname in names:
return os.path.join(dirpath, fname)
return None
def read_json_file(path: str) -> dict:
try:
with open(path) as f:
return json.load(f)
except Exception:
return {}
# ---------------------------------------------------------------------------
# Stack detection
# ---------------------------------------------------------------------------
def detect_stack(root: str) -> Stack:
s = Stack()
pkg_path = os.path.join(root, "package.json")
if os.path.isfile(pkg_path):
s.has_node = True
pkg = read_json_file(pkg_path)
all_deps = {}
for key in ("dependencies", "devDependencies", "peerDependencies"):
all_deps.update(pkg.get(key, {}))
if "next" in all_deps: s.framework = "next"
elif "react" in all_deps: s.framework = "react"
elif "vue" in all_deps: s.framework = "vue"
elif "svelte" in all_deps: s.framework = "svelte"
elif "astro" in all_deps: s.framework = "astro"
elif "express" in all_deps: s.framework = "express"
elif "fastify" in all_deps: s.framework = "fastify"
elif "hono" in all_deps: s.framework = "hono"
s.has_react = s.framework in ("next", "react")
s.is_web = s.framework in ("next", "react", "vue", "svelte", "astro")
if "@supabase/supabase-js" in all_deps:
s.has_supabase = True
if "typescript" in all_deps or os.path.isfile(os.path.join(root, "tsconfig.json")):
s.has_typescript = True
for ai_pkg in ("openai", "@anthropic-ai/sdk", "@google/generative-ai",
"ai", "@huggingface/inference"):
if ai_pkg in all_deps:
s.has_ai = True
s.ai_providers.append(ai_pkg)
if os.path.isdir(os.path.join(root, "supabase")):
s.has_supabase = True
for pyfile in ("requirements.txt", "pyproject.toml", "Pipfile", "setup.py"):
if os.path.isfile(os.path.join(root, pyfile)):
s.has_python = True
try:
content = open(os.path.join(root, pyfile)).read().lower()
if "django" in content: s.py_framework = "django"
elif "flask" in content: s.py_framework = "flask"
elif "fastapi" in content: s.py_framework = "fastapi"
except Exception:
pass
break
if os.path.isfile(os.path.join(root, "go.mod")):
s.has_go = True
if os.path.isfile(os.path.join(root, "Cargo.toml")):
s.has_rust = True
if os.path.isfile(os.path.join(root, "vercel.json")) or \
os.path.isdir(os.path.join(root, ".vercel")):
s.deploy_target = "vercel"
elif os.path.isfile(os.path.join(root, "netlify.toml")):
s.deploy_target = "netlify"
elif os.path.isfile(os.path.join(root, "fly.toml")):
s.deploy_target = "fly"
elif os.path.isfile(os.path.join(root, "railway.json")):
s.deploy_target = "railway"
elif os.path.isfile(os.path.join(root, "Dockerfile")):
s.deploy_target = "docker"
return s
# ---------------------------------------------------------------------------
# Check definitions
# ---------------------------------------------------------------------------
CHECKS = {
# SEC
"SEC-01": CheckDef("SEC-01", "No API keys or secrets in frontend code", Severity.CRITICAL, "SEC"),
"SEC-04": CheckDef("SEC-04", "CORS not wildcard", Severity.CRITICAL, "SEC"),
"SEC-05": CheckDef("SEC-05", "CSRF protection on state-changing endpoints", Severity.CRITICAL, "SEC"),
"SEC-06": CheckDef("SEC-06", "Input validated and sanitized server-side", Severity.HIGH, "SEC"),
"SEC-07": CheckDef("SEC-07", "Rate limiting on auth and sensitive endpoints", Severity.HIGH, "SEC"),
"SEC-08": CheckDef("SEC-08", "Passwords hashed with bcrypt or argon2", Severity.CRITICAL, "SEC"),
"SEC-11": CheckDef("SEC-11", "CSP headers configured", Severity.HIGH, "SEC"),
"SEC-13": CheckDef("SEC-13", "No eval() or dangerouslySetInnerHTML without sanitization", Severity.HIGH, "SEC", stack="js"), # noqa: SEC-AUDITOR
"SEC-14": CheckDef("SEC-14", "No sensitive data in URLs or logs", Severity.HIGH, "SEC"),
"SEC-17": CheckDef("SEC-17", "No hardcoded secrets in .env committed to repo", Severity.CRITICAL, "SEC"),
"SEC-18": CheckDef("SEC-18", ".env files listed in .gitignore", Severity.CRITICAL, "SEC"),
# DB
"DB-03": CheckDef("DB-03", "Parameterized queries everywhere (no SQL injection)", Severity.CRITICAL, "DB"),
"DB-05": CheckDef("DB-05", "Connection pooling configured", Severity.HIGH, "DB"),
"DB-06": CheckDef("DB-06", "Migrations in version control", Severity.HIGH, "DB"),
"DB-07": CheckDef("DB-07", "RLS enabled on all Supabase tables", Severity.CRITICAL, "DB", stack="supabase"),
"DB-08": CheckDef("DB-08", "No service_role key in client-side code", Severity.CRITICAL, "DB", stack="supabase"),
"DB-12": CheckDef("DB-12", "No PII stored unencrypted", Severity.HIGH, "DB"),
# DEPLOY
"DEPLOY-09": CheckDef("DEPLOY-09", "Health check endpoint exists", Severity.HIGH, "DEPLOY"),
"DEPLOY-10": CheckDef("DEPLOY-10", "Structured logging (not raw console)", Severity.HIGH, "DEPLOY"),
# CODE
"CODE-01": CheckDef("CODE-01", "No console.log in production build", Severity.HIGH, "CODE", stack="js"),
"CODE-03": CheckDef("CODE-03", "No empty catch blocks", Severity.HIGH, "CODE"),
"CODE-07": CheckDef("CODE-07", "No TODO-auth or TODO-security patterns", Severity.CRITICAL, "CODE"),
"CODE-09": CheckDef("CODE-09", "React error boundaries in place", Severity.HIGH, "CODE", stack="react"),
"CODE-10": CheckDef("CODE-10", "No leaked stack traces in error responses", Severity.HIGH, "CODE"),
"CODE-11": CheckDef("CODE-11", "No eslint-disable on security rules", Severity.HIGH, "CODE", stack="js"),
"CODE-12": CheckDef("CODE-12", "Lockfile committed", Severity.HIGH, "CODE"),
"CODE-13": CheckDef("CODE-13", "No wildcard versions in package.json", Severity.HIGH, "CODE", stack="js"),
"CODE-14": CheckDef("CODE-14", "TypeScript strict mode enabled", Severity.ADVISORY, "CODE", stack="ts"),
# AI
"AI-01": CheckDef("AI-01", "System prompts not leakable via user input", Severity.CRITICAL, "AI", stack="ai"),
"AI-02": CheckDef("AI-02", "No prompt injection vectors in user inputs", Severity.CRITICAL, "AI", stack="ai"),
"AI-03": CheckDef("AI-03", "LLM API keys not in frontend code", Severity.CRITICAL, "AI", stack="ai"),
"AI-05": CheckDef("AI-05", "AI response output sanitized before rendering", Severity.HIGH, "AI", stack="ai"),
# DEP
"DEP-01": CheckDef("DEP-01", "No git:// or URL-based dependencies", Severity.HIGH, "DEP"),
"DEP-05": CheckDef("DEP-05", "No suspicious postinstall scripts", Severity.HIGH, "DEP", stack="js"),
"DEP-06": CheckDef("DEP-06", "Dependencies pinned (no wildcard *)", Severity.HIGH, "DEP"),
# FE
"FE-01": CheckDef("FE-01", "Meta tags present (title, description, OG)", Severity.ADVISORY, "FE", stack="web"),
"FE-02": CheckDef("FE-02", "Favicon configured", Severity.ADVISORY, "FE", stack="web"),
"FE-03": CheckDef("FE-03", "Custom 404 page exists", Severity.ADVISORY, "FE", stack="web"),
"FE-09": CheckDef("FE-09", "robots.txt present", Severity.ADVISORY, "FE", stack="web"),
# OBS
"OBS-01": CheckDef("OBS-01", "Error monitoring configured (Sentry, etc.)", Severity.ADVISORY, "OBS"),
"OBS-03": CheckDef("OBS-03", "Structured logging with request IDs", Severity.ADVISORY, "OBS"),
}
MANUAL_CHECKS = [
CheckDef("SEC-02", "Every route checks authentication", Severity.CRITICAL, "SEC"),
CheckDef("SEC-03", "HTTPS enforced, HTTP redirected", Severity.CRITICAL, "SEC"),
CheckDef("SEC-10", "Sessions invalidated on logout (server-side)", Severity.HIGH, "SEC"),
CheckDef("DB-01", "Backups configured and tested", Severity.CRITICAL, "DB"),
CheckDef("DB-02", "Backup restore tested (not just backup)", Severity.CRITICAL, "DB"),
CheckDef("DB-04", "Separate dev and production databases", Severity.HIGH, "DB"),
CheckDef("DB-11", "App uses a non-root DB user", Severity.HIGH, "DB"),
CheckDef("DEPLOY-01", "All env vars set on production server", Severity.CRITICAL, "DEPLOY"),
CheckDef("DEPLOY-02", "SSL certificate installed and valid", Severity.CRITICAL, "DEPLOY"),
CheckDef("DEPLOY-05", "Rollback plan exists", Severity.HIGH, "DEPLOY"),
CheckDef("DEPLOY-06", "Staging test passed before production", Severity.HIGH, "DEPLOY"),
CheckDef("AI-07", "Agent permissions scoped (no unrestricted access)", Severity.HIGH, "AI", stack="ai"),
CheckDef("AI-08", "No sensitive data sent to third-party LLMs without consent", Severity.HIGH, "AI", stack="ai"),
CheckDef("FE-04", "Responsive design tested on mobile", Severity.HIGH, "FE", stack="web"),
CheckDef("OBS-05", "Uptime monitoring configured", Severity.HIGH, "OBS"),
]
# ---------------------------------------------------------------------------
# Individual check implementations
# ---------------------------------------------------------------------------
def check_sec01(root, stack):
c = CHECKS["SEC-01"]
dirs = FRONTEND_DIRS & set(os.listdir(root))
patterns = [
r"sk-[a-zA-Z0-9]{20,}",
r"sk-ant-[a-zA-Z0-9-]+",
r"sk-proj-[a-zA-Z0-9-]+",
r"AIza[a-zA-Z0-9_-]{35}",
r"ghp_[a-zA-Z0-9]{36}",
r"glpat-[a-zA-Z0-9_-]{20,}",
r"AKIA[0-9A-Z]{16}",
r"sk_live_[a-zA-Z0-9]{24,}",
r"(api_key|apikey|api_secret|secret_key|auth_token)\s*[:=]\s*['\"][a-zA-Z0-9_\-]{16,}",
]
findings = []
for pat in patterns:
findings += grep_files(root, pat, exts=JS_EXTS | {".env", ".json"},
dirs=dirs if dirs else None, max_findings=5)
if findings:
return Result(c, Status.FAIL,
f"{len(findings)} potential secret(s) found in frontend/client code",
findings[:10])
return Result(c, Status.PASS)
def check_sec04(root, stack):
c = CHECKS["SEC-04"]
findings = grep_files(root, r"(origin\s*:\s*['\"]?\*['\"]?|Access-Control-Allow-Origin.*\*|cors\(\s*\))",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.FAIL, "CORS wildcard (*) detected", findings)
return Result(c, Status.PASS)
def check_sec05(root, stack):
c = CHECKS["SEC-05"]
# Check for state-changing routes
route_findings = grep_files(root, r"(app|router)\.(post|put|patch|delete)\s*\(",
exts=JS_EXTS)
if not route_findings:
return Result(c, Status.SKIP, "No Express-style routes found")
# Check for CSRF protection
csrf_findings = grep_files(root, r"(csrf|csrfToken|_csrf|CSRF_COOKIE|csurf)",
exts=ALL_CODE_EXTS)
if not csrf_findings:
return Result(c, Status.FAIL,
f"{len(route_findings)} state-changing route(s) found but no CSRF protection detected",
route_findings[:5])
return Result(c, Status.PASS)
def check_sec06(root, stack):
c = CHECKS["SEC-06"]
# Check for validation library
val_findings = grep_files(root,
r"(from ['\"]zod['\"]|from ['\"]yup['\"]|from ['\"]joi['\"]|from ['\"]class-validator['\"]|from pydantic|import pydantic)",
exts=ALL_CODE_EXTS)
if val_findings:
return Result(c, Status.PASS)
# Check if there are API routes that use req.body without validation
body_findings = grep_files(root, r"(req\.body|request\.json\(\)|request\.form)",
exts=ALL_CODE_EXTS)
if body_findings:
return Result(c, Status.FAIL,
"request body used without a validation library (zod/yup/joi/pydantic)",
body_findings[:5])
return Result(c, Status.SKIP, "No API route body handling detected")
def check_sec07(root, stack):
c = CHECKS["SEC-07"]
findings = grep_files(root,
r"(express-rate-limit|@upstash/ratelimit|rate-limiter-flexible|slowapi|throttle|rateLimit)",
exts=ALL_CODE_EXTS | {".json"})
if findings:
return Result(c, Status.PASS)
# Only fail if there are auth-related routes
auth_routes = grep_files(root, r"(login|signin|register|signup|forgot.password|reset.password)",
exts=ALL_CODE_EXTS)
if auth_routes:
return Result(c, Status.FAIL,
"Auth routes found but no rate-limiting library detected", auth_routes[:3])
return Result(c, Status.SKIP, "No auth routes detected")
def check_sec08(root, stack):
c = CHECKS["SEC-08"]
# Weak hash for passwords
weak = grep_files(root, r"\b(md5|sha1|sha256)\s*\(",
exts=ALL_CODE_EXTS,
exclude_patterns=[r"//.*\b(md5|sha1|sha256)\b"])
if weak:
return Result(c, Status.FAIL, "Weak hashing algorithm (md5/sha1/sha256) detected", weak)
strong = grep_files(root, r"(bcrypt|argon2|scrypt|pbkdf2)", exts=ALL_CODE_EXTS)
pw_fields = grep_files(root, r"(password|passwd)", exts=ALL_CODE_EXTS)
if pw_fields and not strong:
return Result(c, Status.FAIL, "Password fields found but no bcrypt/argon2/scrypt usage")
return Result(c, Status.PASS if strong or not pw_fields else Status.SKIP)
def check_sec11(root, stack):
c = CHECKS["SEC-11"]
findings = grep_files(root, r"(Content-Security-Policy|contentSecurityPolicy|[^a-z]csp[^a-z])",
exts=ALL_CODE_EXTS | {".json", ".toml", ".yaml", ".yml"})
if findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No Content-Security-Policy configuration found")
def check_sec13(root, stack):
c = CHECKS["SEC-13"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a JS/TS project")
eval_findings = grep_files(root, r"(\beval\s*\(|new\s+Function\s*\()", exts=JS_EXTS)
dsi_findings = grep_files(root, r"dangerouslySetInnerHTML", exts=JS_EXTS)
# If dangerouslySetInnerHTML is used, check for DOMPurify
unsafe_dsi = []
for f in dsi_findings:
try:
content = open(os.path.join(root, f.file), errors="ignore").read()
if "DOMPurify" not in content and "sanitize" not in content.lower():
unsafe_dsi.append(f)
except Exception:
unsafe_dsi.append(f)
all_findings = eval_findings + unsafe_dsi # noqa: SEC-AUDITOR
if all_findings:
return Result(c, Status.FAIL, "Unsafe eval() or unsanitized dangerouslySetInnerHTML", all_findings) # noqa: SEC-AUDITOR
return Result(c, Status.PASS)
def check_sec14(root, stack):
c = CHECKS["SEC-14"]
url_findings = grep_files(root,
r"(password|token|secret|key|ssn|credit.card)=",
exts=ALL_CODE_EXTS)
log_findings = grep_files(root,
r"console\.(log|info|debug)\s*\(\s*(req|request)\s*\)",
exts=JS_EXTS)
findings = url_findings + log_findings
if findings:
return Result(c, Status.FAIL, "Sensitive data may appear in URLs or logs", findings[:5])
return Result(c, Status.PASS)
def check_sec17(root, stack):
c = CHECKS["SEC-17"]
# Check for .env files that are not .example/.sample
env_files = []
for entry in os.scandir(root):
name = entry.name
if name.startswith(".env") and name not in (".env.example", ".env.sample",
".env.template", ".env.local.example"):
if entry.is_file():
env_files.append(name)
if not env_files:
return Result(c, Status.PASS)
# Check if git-tracked
gitignore_path = os.path.join(root, ".gitignore")
if os.path.isfile(gitignore_path):
content = open(gitignore_path, errors="ignore").read()
if ".env" in content:
return Result(c, Status.PASS)
return Result(c, Status.FAIL,
f".env file(s) exist ({', '.join(env_files)}) and may not be gitignored",
[Finding(f, 0) for f in env_files])
def check_sec18(root, stack):
c = CHECKS["SEC-18"]
gitignore_path = os.path.join(root, ".gitignore")
if not os.path.isfile(gitignore_path):
return Result(c, Status.FAIL, ".gitignore file not found")
content = open(gitignore_path, errors="ignore").read()
if re.search(r"\.env", content):
return Result(c, Status.PASS)
return Result(c, Status.FAIL, ".env not listed in .gitignore")
def check_db03(root, stack):
c = CHECKS["DB-03"]
# Template literal SQL
tl_findings = grep_files(root,
r"(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE).*\$\{",
exts=JS_EXTS)
# Python f-string SQL
py_findings = grep_files(root,
r'f["\'].*\b(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE)\b.*\{',
exts=PY_EXTS)
# String concat SQL
concat_findings = grep_files(root,
r"(SELECT|INSERT|UPDATE|DELETE|FROM|WHERE).*\+\s*(req\.|params\.|body\.|query\.)",
exts=ALL_CODE_EXTS)
all_findings = tl_findings + py_findings + concat_findings
if all_findings:
return Result(c, Status.FAIL,
f"{len(all_findings)} potential SQL injection vector(s)", all_findings[:10])
return Result(c, Status.PASS)
def check_db05(root, stack):
c = CHECKS["DB-05"]
findings = grep_files(root,
r"(pool|connectionLimit|max_connections|poolSize|pooler|6543)",
exts=ALL_CODE_EXTS | {".env", ".env.local", ".env.production"})
if findings:
return Result(c, Status.PASS)
db_found = grep_files(root, r"(pg\.|postgres\.|mysql\.|mongoose\.)", exts=ALL_CODE_EXTS)
if db_found:
return Result(c, Status.FAIL, "Database usage detected but no connection pooling configured")
return Result(c, Status.SKIP, "No direct DB connection detected")
def check_db06(root, stack):
c = CHECKS["DB-06"]
migration_dirs = []
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = [d for d in dirnames if d not in EXCLUDE_DIRS]
depth = dirpath.replace(root, "").count(os.sep)
if depth >= 4:
dirnames[:] = []
continue
for d in dirnames:
if d in ("migrations", "migrate", "versions", "alembic"):
migration_dirs.append(os.path.join(dirpath, d))
if migration_dirs:
return Result(c, Status.PASS)
# Check for database usage
db_found = grep_files(root, r"(prisma|supabase|mongoose|pg\.|sqlite)", exts=ALL_CODE_EXTS)
if db_found:
return Result(c, Status.FAIL, "Database usage found but no migrations directory detected")
return Result(c, Status.SKIP, "No database usage detected")
def check_db07(root, stack):
c = CHECKS["DB-07"]
if not stack.has_supabase:
return Result(c, Status.SKIP, "Not a Supabase project")
sql_findings = grep_files(root, r"CREATE TABLE", exts=SQL_EXTS)
if not sql_findings:
return Result(c, Status.SKIP, "No CREATE TABLE statements found in migrations")
rls_findings = grep_files(root, r"ENABLE ROW LEVEL SECURITY", exts=SQL_EXTS)
if not rls_findings:
return Result(c, Status.FAIL,
f"{len(sql_findings)} table(s) found but no RLS policies detected",
sql_findings[:5])
if len(rls_findings) < len(sql_findings):
return Result(c, Status.FAIL,
f"{len(sql_findings)} table(s) but only {len(rls_findings)} RLS statement(s) — some tables may lack RLS",
sql_findings[:5])
return Result(c, Status.PASS)
def check_db08(root, stack):
c = CHECKS["DB-08"]
if not stack.has_supabase:
return Result(c, Status.SKIP, "Not a Supabase project")
dirs = FRONTEND_DIRS & set(os.listdir(root))
findings = grep_files(root,
r"(service_role|serviceRole|SUPABASE_SERVICE_ROLE)",
exts=JS_EXTS, dirs=dirs if dirs else None)
if findings:
return Result(c, Status.FAIL, "service_role key referenced in client-side code", findings)
return Result(c, Status.PASS)
def check_db12(root, stack):
c = CHECKS["DB-12"]
findings = grep_files(root,
r"(ssn|social_security|credit_card|card_number|passport_number)",
exts=SQL_EXTS | {".prisma"}, flags=re.IGNORECASE)
if findings:
return Result(c, Status.FAIL,
"PII column names found in schema — verify encryption at rest", findings)
return Result(c, Status.PASS)
def check_deploy09(root, stack):
c = CHECKS["DEPLOY-09"]
findings = grep_files(root,
r"(/health|/healthz|/api/health|/status|/readyz)",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No health check endpoint found")
def check_deploy10(root, stack):
c = CHECKS["DEPLOY-10"]
# Check for logging libraries
lib_findings = grep_files(root,
r"(winston|pino|bunyan|morgan|log4js|structlog|loguru)",
exts=ALL_CODE_EXTS | {".json"})
if lib_findings:
return Result(c, Status.PASS)
# Count console.log in server/api code
server_dirs = {"api", "server", "backend"}
for d in ("pages/api", "app/api"):
if os.path.isdir(os.path.join(root, d)):
server_dirs.add(d.split("/")[0])
console_findings = grep_files(root, r"console\.(log|debug|info)\(", exts=JS_EXTS)
if console_findings:
return Result(c, Status.FAIL,
f"No structured logger found; {len(console_findings)} console.log(s) in code",
console_findings[:5])
return Result(c, Status.SKIP, "No server-side code detected")
def check_code01(root, stack):
c = CHECKS["CODE-01"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a JS/TS project")
findings = grep_files(root, r"console\.(log|debug|info)\(",
exts=JS_EXTS,
dirs=FRONTEND_DIRS & set(os.listdir(root)) or None,
exclude_patterns=[r"//.*console\.(log|debug|info)\("])
if findings:
return Result(c, Status.FAIL, f"{len(findings)} console.log statement(s) in production code", findings[:10])
return Result(c, Status.PASS)
def check_code03(root, stack):
c = CHECKS["CODE-03"]
findings = grep_files(root,
r"catch\s*\([^)]*\)\s*\{\s*\}",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.FAIL, f"{len(findings)} empty catch block(s)", findings)
return Result(c, Status.PASS)
def check_code07(root, stack):
c = CHECKS["CODE-07"]
findings = grep_files(root,
r"(TODO|FIXME|HACK|XXX).{0,20}(auth|security|permission|validation|sanitiz)",
exts=ALL_CODE_EXTS, flags=re.IGNORECASE)
if findings:
return Result(c, Status.FAIL, f"{len(findings)} deferred security TODO(s)", findings)
return Result(c, Status.PASS)
def check_code09(root, stack):
c = CHECKS["CODE-09"]
if not stack.has_react:
return Result(c, Status.SKIP, "Not a React project")
# Next.js App Router: error.tsx
error_page = file_exists_in(root, "error.tsx", "error.jsx", "global-error.tsx")
if error_page:
return Result(c, Status.PASS)
# Class-based error boundary
eb_findings = grep_files(root,
r"(ErrorBoundary|componentDidCatch|getDerivedStateFromError)",
exts=JS_EXTS)
if eb_findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No React error boundary or error.tsx found")
def check_code10(root, stack):
c = CHECKS["CODE-10"]
findings = grep_files(root,
r"(error\.stack|\.stack\s*\)|err\.message.*res\.(json|send)|traceback\.format_exc)",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.FAIL, "Potential stack trace leak in error responses", findings)
return Result(c, Status.PASS)
def check_code11(root, stack):
c = CHECKS["CODE-11"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a JS/TS project")
findings = grep_files(root,
r"eslint-disable.*(no-eval|no-implied-eval|no-script-url|security)",
exts=JS_EXTS)
if findings:
return Result(c, Status.FAIL, "Security lint rule(s) disabled", findings)
return Result(c, Status.PASS)
def check_code12(root, stack):
c = CHECKS["CODE-12"]
lockfiles = ["package-lock.json", "pnpm-lock.yaml", "yarn.lock", "bun.lockb",
"Pipfile.lock", "poetry.lock", "Gemfile.lock", "go.sum", "Cargo.lock"]
for lf in lockfiles:
if os.path.isfile(os.path.join(root, lf)):
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No lockfile found — dependencies are not pinned")
def check_code13(root, stack):
c = CHECKS["CODE-13"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a JS/TS project")
pkg_path = os.path.join(root, "package.json")
if not os.path.isfile(pkg_path):
return Result(c, Status.SKIP)
findings = grep_files(root, r'"[^"]+"\s*:\s*"\*"', exts={".json"})
findings += grep_files(root, r'"[^"]+"\s*:\s*"latest"', exts={".json"})
findings = [f for f in findings if "package.json" in f.file and "node_modules" not in f.file]
if findings:
return Result(c, Status.FAIL, "Wildcard (*) or 'latest' version found in package.json", findings)
return Result(c, Status.PASS)
def check_code14(root, stack):
c = CHECKS["CODE-14"]
if not stack.has_typescript:
return Result(c, Status.SKIP, "Not a TypeScript project")
tsconfig_path = os.path.join(root, "tsconfig.json")
if not os.path.isfile(tsconfig_path):
return Result(c, Status.SKIP, "tsconfig.json not found")
content = open(tsconfig_path, errors="ignore").read()
if re.search(r'"strict"\s*:\s*true', content):
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "TypeScript strict mode not enabled in tsconfig.json",
[Finding("tsconfig.json", 0)])
def check_ai01(root, stack):
c = CHECKS["AI-01"]
if not stack.has_ai:
return Result(c, Status.SKIP, "No AI/LLM usage detected")
dirs = FRONTEND_DIRS & set(os.listdir(root))
findings = grep_files(root,
r"(system.?prompt|system.?message|system_instruction)",
exts=ALL_CODE_EXTS, dirs=dirs if dirs else None, flags=re.IGNORECASE)
if findings:
return Result(c, Status.FAIL,
"System prompt referenced in client-accessible code — may be leakable",
findings)
return Result(c, Status.PASS)
def check_ai02(root, stack):
c = CHECKS["AI-02"]
if not stack.has_ai:
return Result(c, Status.SKIP, "No AI/LLM usage detected")
findings = grep_files(root,
r"(messages\.push|content\s*:.*\$\{|content\s*:.*\+\s*user|prompt.*\+)",
exts=ALL_CODE_EXTS)
if findings:
return Result(c, Status.FAIL,
"User input may be concatenated directly into AI prompt", findings[:5])
return Result(c, Status.PASS)
def check_ai03(root, stack):
c = CHECKS["AI-03"]
if not stack.has_ai:
return Result(c, Status.SKIP, "No AI/LLM usage detected")
dirs = FRONTEND_DIRS & set(os.listdir(root))
findings = grep_files(root,
r"(OPENAI_API_KEY|ANTHROPIC_API_KEY|GOOGLE_AI_API_KEY|sk-ant-|sk-proj-)",
exts=JS_EXTS, dirs=dirs if dirs else None)
if findings:
return Result(c, Status.FAIL, "LLM API key referenced in frontend code", findings)
return Result(c, Status.PASS)
def check_ai05(root, stack):
c = CHECKS["AI-05"]
if not stack.has_ai:
return Result(c, Status.SKIP, "No AI/LLM usage detected")
findings = grep_files(root,
r"dangerouslySetInnerHTML.*\b(response|result|completion|message|content)\b",
exts=JS_EXTS)
if findings:
return Result(c, Status.FAIL, "AI output rendered via dangerouslySetInnerHTML", findings)
return Result(c, Status.PASS)
def check_dep01(root, stack):
c = CHECKS["DEP-01"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a Node.js project")
findings = grep_files(root,
r'"[^"]+"\s*:\s*"(git://|git\+|github:|https://github\.com|file:)',
exts={".json"})
findings = [f for f in findings if "package.json" in f.file and "node_modules" not in f.file]
if findings:
return Result(c, Status.FAIL, "Git/URL-based dependency found in package.json", findings)
return Result(c, Status.PASS)
def check_dep05(root, stack):
c = CHECKS["DEP-05"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a Node.js project")
pkg_path = os.path.join(root, "package.json")
if not os.path.isfile(pkg_path):
return Result(c, Status.SKIP)
pkg = read_json_file(pkg_path)
scripts = pkg.get("scripts", {})
suspicious = []
for key in ("preinstall", "postinstall", "install"):
val = scripts.get(key, "")
if val and any(kw in val for kw in ("curl", "wget", "fetch", "exec", "eval", "sh ", "bash ")):
suspicious.append(Finding("package.json", 0, f'"{key}": "{val}"'))
if suspicious:
return Result(c, Status.FAIL, "Suspicious install script detected in package.json", suspicious)
return Result(c, Status.PASS)
def check_dep06(root, stack):
c = CHECKS["DEP-06"]
if not stack.has_node:
return Result(c, Status.SKIP, "Not a Node.js project")
pkg_path = os.path.join(root, "package.json")
if not os.path.isfile(pkg_path):
return Result(c, Status.SKIP)
findings = grep_files(root, r'"\*"', exts={".json"})
findings = [f for f in findings if "package.json" in f.file and "node_modules" not in f.file]
if findings:
return Result(c, Status.FAIL, "Wildcard (*) version found", findings)
return Result(c, Status.PASS)
def check_fe01(root, stack):
c = CHECKS["FE-01"]
if not stack.is_web and not stack.has_node:
return Result(c, Status.SKIP, "Not a web project")
# Next.js metadata export
meta_findings = grep_files(root,
r"(export\s+(const|async\s+function)\s+metadata|generateMetadata|<title>|og:title|og:description)",
exts=JS_EXTS | {".html"})
if meta_findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No meta tags or Next.js metadata export found")
def check_fe02(root, stack):
c = CHECKS["FE-02"]
if not stack.is_web and not stack.has_node:
return Result(c, Status.SKIP, "Not a web project")
favicon = file_exists_in(root, "favicon.ico", "favicon.png", "favicon.svg",
"favicon.webp", "icon.png", "icon.ico")
if favicon:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No favicon file found")
def check_fe03(root, stack):
c = CHECKS["FE-03"]
if not stack.is_web and not stack.has_node:
return Result(c, Status.SKIP, "Not a web project")
page_404 = file_exists_in(root, "404.tsx", "404.jsx", "404.html",
"not-found.tsx", "not-found.jsx")
if page_404:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No custom 404 or not-found page found")
def check_fe09(root, stack):
c = CHECKS["FE-09"]
if not stack.is_web and not stack.has_node:
return Result(c, Status.SKIP, "Not a web project")
public_robots = os.path.join(root, "public", "robots.txt")
root_robots = os.path.join(root, "robots.txt")
if os.path.isfile(public_robots) or os.path.isfile(root_robots):
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No robots.txt found")
def check_obs01(root, stack):
c = CHECKS["OBS-01"]
findings = grep_files(root,
r"(@sentry/|sentry-|LogRocket|Bugsnag|datadogRum|Rollbar|Honeybadger|newrelic)",
exts=ALL_CODE_EXTS | {".json"})
if findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No error monitoring library detected")
def check_obs03(root, stack):
c = CHECKS["OBS-03"]
findings = grep_files(root,
r"(winston|pino|bunyan|structlog|loguru|import logging)",
exts=ALL_CODE_EXTS | {".json"})
if findings:
return Result(c, Status.PASS)
return Result(c, Status.FAIL, "No structured logging library detected")
CATEGORY_CHECKS = {
"SEC": [check_sec01, check_sec04, check_sec05, check_sec06, check_sec07,
check_sec08, check_sec11, check_sec13, check_sec14, check_sec17, check_sec18],
"DB": [check_db03, check_db05, check_db06, check_db07, check_db08, check_db12],
"DEPLOY": [check_deploy09, check_deploy10],
"CODE": [check_code01, check_code03, check_code07, check_code09, check_code10,
check_code11, check_code12, check_code13, check_code14],
"AI": [check_ai01, check_ai02, check_ai03, check_ai05],
"DEP": [check_dep01, check_dep05, check_dep06],
"FE": [check_fe01, check_fe02, check_fe03, check_fe09],
"OBS": [check_obs01, check_obs03],
}
CATEGORY_ORDER = ["SEC", "DB", "CODE", "DEP", "AI", "DEPLOY", "FE", "OBS"]
# ---------------------------------------------------------------------------
# Manual check runner
# ---------------------------------------------------------------------------
def run_manual_checks(stack: Stack, interactive: bool, category_filter: Optional[str]) -> List[Result]:
results = []
applicable = []
for chk in MANUAL_CHECKS:
if category_filter and chk.category != category_filter.upper():
continue
if chk.stack == "ai" and not stack.has_ai:
results.append(Result(chk, Status.SKIP, "No AI/LLM usage detected"))
continue
if chk.stack == "web" and not stack.is_web:
results.append(Result(chk, Status.SKIP, "Not a web project"))
continue
if chk.stack == "vps" and stack.deploy_target not in ("docker", "vps", ""):
results.append(Result(chk, Status.SKIP, "Not a VPS/Docker deployment"))
continue
applicable.append(chk)
if not interactive or not applicable:
for chk in applicable:
results.append(Result(chk, Status.MANUAL, "Not confirmed (run without --no-interactive to answer)"))
return results
print()
print(bold("Manual Checks") + " — answer Y/N for each:")
print()
for chk in applicable:
sev_label = {
Severity.CRITICAL: red("CRITICAL"),
Severity.HIGH: yellow("HIGH"),
Severity.ADVISORY: dim("ADVISORY"),
}[chk.severity]
while True:
try:
answer = input(f" [{sev_label}] [{chk.id}] {chk.description} [y/N]: ").strip().lower()
except (EOFError, KeyboardInterrupt):
answer = "n"
if answer in ("y", "yes"):
results.append(Result(chk, Status.PASS))
break
elif answer in ("n", "no", ""):
results.append(Result(chk, Status.FAIL, "Not confirmed"))
break
print(" Please enter Y or N.")
return results
# ---------------------------------------------------------------------------
# Verdict / output
# ---------------------------------------------------------------------------
def severity_for_result(r: Result) -> Severity:
return r.check.severity
def print_report(all_results: List[Result], stack: Stack, scan_time: float,
verbose: bool) -> int:
critical = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.CRITICAL]
high = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.HIGH]
advisory = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.ADVISORY]
stack_desc = []
if stack.framework: stack_desc.append(stack.framework.capitalize())
if stack.has_supabase: stack_desc.append("Supabase")
if stack.deploy_target: stack_desc.append(stack.deploy_target.capitalize())
if stack.has_python and stack.py_framework: stack_desc.append(stack.py_framework.capitalize())
if not stack_desc: stack_desc.append("Unknown")
stack_str = " + ".join(stack_desc)
print()
print(bold("SHIP GATE REPORT"))
print("=" * 48)
print(f"Stack: {stack_str}")
print(f"Scan time: {scan_time:.1f}s")
print(f"Checks: {len(all_results)} total")
print()
def _section(label, items, color_fn):
if not items and not verbose:
return
print(bold(f"{label} ({len(items)} item{'s' if len(items) != 1 else ''})"))
for r in items:
status_str = {
Status.FAIL: red("FAIL "),
Status.MANUAL: yellow("MANUAL"),
Status.PASS: green("PASS "),
Status.SKIP: dim("SKIP "),
}[r.status]
print(f" {status_str} [{r.check.id}] {r.check.description}")
if r.message:
print(f" {dim(r.message)}")
for f in r.findings[:3]:
print(f" {dim(f.file)}:{f.line} {dim(f.snippet[:80])}")
print()
if critical:
_section(red("CRITICAL") + " (must fix before shipping)", critical, red)
if high:
_section(yellow("HIGH") + " (should fix before shipping)", high, yellow)
if advisory:
_section(dim("ADVISORY") + " (recommended)", advisory, dim)
if verbose:
passed = [r for r in all_results if r.status == Status.PASS]
if passed:
_section(green("PASS"), passed, green)
skipped = [r for r in all_results if r.status == Status.SKIP]
if skipped:
_section(dim("SKIP"), skipped, dim)
if critical:
print(red(bold(f"VERDICT: DO NOT SHIP ({len(critical)} critical issue{'s' if len(critical) != 1 else ''})")))
print("Fix critical items and re-run.")
return 1
elif high:
print(yellow(bold(f"VERDICT: SHIP WITH CAUTION ({len(high)} high issue{'s' if len(high) != 1 else ''})")))
print("Acknowledge risks and proceed only if you accept them.")
return 2
else:
print(green(bold("VERDICT: CLEAR TO SHIP")))
return 0
def print_json_report(all_results: List[Result], stack: Stack, scan_time: float) -> int:
critical = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.CRITICAL]
high = [r for r in all_results if r.status in (Status.FAIL, Status.MANUAL)
and r.check.severity == Severity.HIGH]
output = {
"version": VERSION,
"scan_time": round(scan_time, 2),
"stack": {
"framework": stack.framework,
"has_supabase": stack.has_supabase,
"has_typescript": stack.has_typescript,
"deploy_target": stack.deploy_target,
"has_ai": stack.has_ai,
},
"results": [
{
"id": r.check.id,
"description": r.check.description,
"severity": r.check.severity.value,
"category": r.check.category,
"status": r.status.value,
"message": r.message,
"findings": [
{"file": f.file, "line": f.line, "snippet": f.snippet}
for f in r.findings
],
}
for r in all_results
],
"summary": {
"critical": len(critical),
"high": len(high),
"verdict": "DO_NOT_SHIP" if critical else ("SHIP_WITH_CAUTION" if high else "CLEAR_TO_SHIP"),
},
}
print(json.dumps(output, indent=2))
return 1 if critical else (2 if high else 0)
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main():
global USE_COLOR
parser = argparse.ArgumentParser(
description="Ship Gate — pre-production audit scanner",
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument("path", nargs="?", default=".",
help="Project root directory (default: current directory)")
parser.add_argument("--json", action="store_true", help="Output as JSON")
parser.add_argument("--no-color", action="store_true", help="Disable color output")
parser.add_argument("--no-interactive", action="store_true",
help="Skip manual confirmation prompts")
parser.add_argument("--category", metavar="CAT",
help="Only run one category: SEC, DB, CODE, DEP, AI, DEPLOY, FE, OBS")
parser.add_argument("--verbose", action="store_true",
help="Show PASS and SKIP results in addition to failures")
parser.add_argument("--version", action="version", version=f"ship-gate {VERSION}")
args = parser.parse_args()
if args.no_color or not sys.stdout.isatty():
USE_COLOR = False
root = os.path.abspath(args.path)
if not os.path.isdir(root):
print(f"Error: '{root}' is not a directory", file=sys.stderr)
sys.exit(1)
start = time.time()
# Detect stack
stack = detect_stack(root)
if not args.json:
print(bold("Detecting stack..."), end=" ", flush=True)
parts = []
if stack.framework: parts.append(stack.framework.capitalize())
if stack.has_supabase: parts.append("Supabase")
if stack.deploy_target: parts.append(stack.deploy_target.capitalize())
if stack.has_python and stack.py_framework: parts.append(stack.py_framework.capitalize())
if stack.has_ai: parts.append(f"AI({','.join(stack.ai_providers)})")
print(", ".join(parts) if parts else "generic project")
# Run automated checks
all_results: List[Result] = []
categories = [args.category.upper()] if args.category else CATEGORY_ORDER
for i, cat in enumerate(categories, 1):
fns = CATEGORY_CHECKS.get(cat, [])
cat_results = []
for fn in fns:
try:
r = fn(root, stack)
except Exception as e:
chk_id = fn.__name__.replace("check_", "").replace("_", "-").upper()
cat_results.append(Result(
CheckDef(chk_id, fn.__doc__ or fn.__name__, Severity.ADVISORY, cat),
Status.SKIP, f"Scanner error: {e}",
))
continue
cat_results.append(r)
all_results.extend(cat_results)
if not args.json:
n_fail = sum(1 for r in cat_results if r.status == Status.FAIL)
n_pass = sum(1 for r in cat_results if r.status == Status.PASS)
n_skip = sum(1 for r in cat_results if r.status == Status.SKIP)
label = red(f"{n_fail} FAIL") if n_fail else green("0 FAIL")
print(f" [{i}/{len(categories)}] {cat}: {label}, {n_pass} PASS, {dim(str(n_skip) + ' SKIP')}")
# Manual checks
manual_results = run_manual_checks(stack, not args.no_interactive, args.category)
all_results.extend(manual_results)
scan_time = time.time() - start
if args.json:
sys.exit(print_json_report(all_results, stack, scan_time))
else:
sys.exit(print_report(all_results, stack, scan_time, args.verbose))
if __name__ == "__main__":
main()
Skill mẫu dùng để tham khảo cấu trúc khi tạo skill mới.
# Sample Text Processor
---
**Name**: sample-text-processor
**Tier**: BASIC
**Category**: Text Processing
**Dependencies**: None (Python Standard Library Only)
**Author**: Claude Skills Engineering Team
**Version**: 1.0.0
**Last Updated**: 2026-02-16
---
## Description
The Sample Text Processor is a simple skill designed to demonstrate the basic structure and functionality expected in the claude-skills ecosystem. This skill provides fundamental text processing capabilities including word counting, character analysis, and basic text transformations.
This skill serves as a reference implementation for BASIC tier requirements and can be used as a template for creating new skills. It demonstrates proper file structure, documentation standards, and implementation patterns that align with ecosystem best practices.
The skill processes text files and provides statistics and transformations in both human-readable and JSON formats, showcasing the dual output requirement for skills in the claude-skills repository.
## Features
### Core Functionality
- **Word Count Analysis**: Count total words, unique words, and word frequency
- **Character Statistics**: Analyze character count, line count, and special characters
- **Text Transformations**: Convert text to uppercase, lowercase, or title case
- **File Processing**: Process single text files or batch process directories
- **Dual Output Formats**: Generate results in both JSON and human-readable formats
### Technical Features
- Command-line interface with comprehensive argument parsing
- Error handling for common file and processing issues
- Progress reporting for batch operations
- Configurable output formatting and verbosity levels
- Cross-platform compatibility with standard library only dependencies
## Usage
### Basic Text Analysis
```bash
python text_processor.py analyze document.txt
python text_processor.py analyze document.txt --output results.json
```
### Text Transformation
```bash
python text_processor.py transform document.txt --mode uppercase
python text_processor.py transform document.txt --mode title --output transformed.txt
```
### Batch Processing
```bash
python text_processor.py batch text_files/ --output results/
python text_processor.py batch text_files/ --format json --output batch_results.json
```
## Examples
### Example 1: Basic Word Count
```bash
$ python text_processor.py analyze sample.txt
=== TEXT ANALYSIS RESULTS ===
File: sample.txt
Total words: 150
Unique words: 85
Total characters: 750
Lines: 12
Most frequent word: "the" (8 occurrences)
```
### Example 2: JSON Output
```bash
$ python text_processor.py analyze sample.txt --format json
{
"file": "sample.txt",
"statistics": {
"total_words": 150,
"unique_words": 85,
"total_characters": 750,
"lines": 12,
"most_frequent": {
"word": "the",
"count": 8
}
}
}
```
### Example 3: Text Transformation
```bash
$ python text_processor.py transform sample.txt --mode title
Original: "hello world from the text processor"
Transformed: "Hello World From The Text Processor"
```
## Installation
This skill requires only Python 3.7 or later with the standard library. No external dependencies are required.
1. Clone or download the skill directory
2. Navigate to the scripts directory
3. Run the text processor directly with Python
```bash
cd scripts/
python text_processor.py --help
```
## Configuration
The text processor supports various configuration options through command-line arguments:
- `--format`: Output format (json, text)
- `--verbose`: Enable verbose output and progress reporting
- `--output`: Specify output file or directory
- `--encoding`: Specify text file encoding (default: utf-8)
## Architecture
The skill follows a simple modular architecture:
- **TextProcessor Class**: Core processing logic and statistics calculation
- **OutputFormatter Class**: Handles dual output format generation
- **FileManager Class**: Manages file I/O operations and batch processing
- **CLI Interface**: Command-line argument parsing and user interaction
## Error Handling
The skill includes comprehensive error handling for:
- File not found or permission errors
- Invalid encoding or corrupted text files
- Memory limitations for very large files
- Output directory creation and write permissions
- Invalid command-line arguments and parameters
## Performance Considerations
- Efficient memory usage for large text files through streaming
- Optimized word counting using dictionary lookups
- Batch processing with progress reporting for large datasets
- Configurable encoding detection for international text
## Contributing
This skill serves as a reference implementation and contributions are welcome to demonstrate best practices:
1. Follow PEP 8 coding standards
2. Include comprehensive docstrings
3. Add test cases with sample data
4. Update documentation for any new features
5. Ensure backward compatibility
## Limitations
As a BASIC tier skill, some advanced features are intentionally omitted:
- Complex text analysis (sentiment, language detection)
- Advanced file format support (PDF, Word documents)
- Database integration or external API calls
- Parallel processing for very large datasets
This skill demonstrates the essential structure and quality standards required for BASIC tier skills in the claude-skills ecosystem while remaining simple and focused on core functionality.
FILE:assets/sample_text.txt
This is a sample text file for testing the text processor skill.
It contains multiple lines of text with various words and punctuation.
The quick brown fox jumps over the lazy dog.
This sentence contains all 26 letters of the English alphabet.
Some additional content:
- Numbers: 123, 456, 789
- Special characters: !@#$%^&*()
- Mixed case: CamelCase, snake_case, PascalCase
Lorem ipsum dolor sit amet, consectetur adipiscing elit.
Sed do eiusmod tempor incididunt ut labore et dolore magna aliqua.
Ut enim ad minim veniam, quis nostrud exercitation ullamco.
This file serves as a basic test case for:
1. Word counting functionality
2. Character analysis
3. Line counting
4. Text transformations
5. Statistical analysis
The text processor should handle this content correctly and produce
meaningful statistics and transformations for testing purposes.
FILE:assets/test_data.csv
name,age,city,country
John Doe,25,New York,USA
Jane Smith,30,London,UK
Bob Johnson,22,Toronto,Canada
Alice Brown,28,Sydney,Australia
Charlie Wilson,35,Berlin,Germany
This CSV file contains sample data with headers and multiple rows.
It can be used to test the text processor's ability to handle
structured data formats and count words across different content types.
The file includes:
- Header row with column names
- Data rows with mixed text and numbers
- Various city and country names
- Different age values for statistical analysis
FILE:expected_outputs/sample_text_analysis.json
{
"file": "assets/sample_text.txt",
"file_size": 855,
"total_words": 116,
"unique_words": 87,
"total_characters": 855,
"lines": 19,
"average_word_length": 4.7,
"most_frequent": {
"word": "the",
"count": 5
}
}
FILE:README.md
# Sample Text Processor
A basic text processing skill that demonstrates BASIC tier requirements for the claude-skills ecosystem.
## Quick Start
```bash
# Analyze a text file
python scripts/text_processor.py analyze sample.txt
# Get JSON output
python scripts/text_processor.py analyze sample.txt --format json
# Transform text to uppercase
python scripts/text_processor.py transform sample.txt --mode upper
# Process multiple files
python scripts/text_processor.py batch text_files/ --verbose
```
## Features
- Word count and text statistics
- Text transformations (upper, lower, title, reverse)
- Batch file processing
- JSON and human-readable output formats
- Comprehensive error handling
## Requirements
- Python 3.7 or later
- No external dependencies (standard library only)
## Usage
See [SKILL.md](SKILL.md) for comprehensive documentation and examples.
## Testing
Sample data files are provided in the `assets/` directory for testing the functionality.
FILE:references/api-reference.md
# Text Processor API Reference
## Classes
### TextProcessor
Main class for text processing operations.
#### `__init__(self, encoding: str = 'utf-8')`
Initialize the text processor with specified encoding.
**Parameters:**
- `encoding` (str): Character encoding for file operations. Default: 'utf-8'
#### `analyze_text(self, text: str) -> Dict[str, Any]`
Analyze text and return comprehensive statistics.
**Parameters:**
- `text` (str): Text content to analyze
**Returns:**
- `dict`: Statistics including word count, character count, lines, most frequent word
**Example:**
```python
processor = TextProcessor()
stats = processor.analyze_text("Hello world")
# Returns: {'total_words': 2, 'unique_words': 2, ...}
```
#### `transform_text(self, text: str, mode: str) -> str`
Transform text according to specified mode.
**Parameters:**
- `text` (str): Text to transform
- `mode` (str): Transformation mode ('upper', 'lower', 'title', 'reverse')
**Returns:**
- `str`: Transformed text
**Raises:**
- `ValueError`: If mode is not supported
### OutputFormatter
Static methods for output formatting.
#### `format_json(data: Dict[str, Any]) -> str`
Format data as JSON string.
#### `format_human_readable(data: Dict[str, Any]) -> str`
Format data as human-readable text.
### FileManager
Handles file operations and batch processing.
#### `find_text_files(self, directory: str) -> List[str]`
Find all text files in a directory recursively.
**Supported Extensions:**
- .txt
- .md
- .rst
- .csv
- .log
## Command Line Interface
### Commands
#### `analyze`
Analyze text file statistics.
```bash
python text_processor.py analyze <file> [options]
```
#### `transform`
Transform text file content.
```bash
python text_processor.py transform <file> --mode <mode> [options]
```
#### `batch`
Process multiple files in a directory.
```bash
python text_processor.py batch <directory> [options]
```
### Global Options
- `--format {json,text}`: Output format (default: text)
- `--output FILE`: Output file path (default: stdout)
- `--encoding ENCODING`: Text file encoding (default: utf-8)
- `--verbose`: Enable verbose output
## Error Handling
The text processor handles several error conditions:
- **FileNotFoundError**: When input file doesn't exist
- **UnicodeDecodeError**: When file encoding doesn't match specified encoding
- **PermissionError**: When file access is denied
- **ValueError**: When invalid transformation mode is specified
All errors are reported to stderr with descriptive messages.
FILE:scripts/text_processor.py
#!/usr/bin/env python3
"""
Sample Text Processor - Basic text analysis and transformation tool
This script demonstrates the basic structure and functionality expected in
BASIC tier skills. It provides text processing capabilities with proper
argument parsing, error handling, and dual output formats.
Usage:
python text_processor.py analyze <file> [options]
python text_processor.py transform <file> --mode <mode> [options]
python text_processor.py batch <directory> [options]
Author: Claude Skills Engineering Team
Version: 1.0.0
Dependencies: Python Standard Library Only
"""
import argparse
import json
import os
import sys
from collections import Counter
from pathlib import Path
from typing import Dict, List, Any, Optional
class TextProcessor:
"""Core text processing functionality"""
def __init__(self, encoding: str = 'utf-8'):
self.encoding = encoding
def analyze_text(self, text: str) -> Dict[str, Any]:
"""Analyze text and return statistics"""
lines = text.split('\n')
words = text.lower().split()
# Calculate basic statistics
stats = {
'total_words': len(words),
'unique_words': len(set(words)),
'total_characters': len(text),
'lines': len(lines),
'average_word_length': sum(len(word) for word in words) / len(words) if words else 0
}
# Find most frequent word
if words:
word_counts = Counter(words)
most_common = word_counts.most_common(1)[0]
stats['most_frequent'] = {
'word': most_common[0],
'count': most_common[1]
}
else:
stats['most_frequent'] = {'word': '', 'count': 0}
return stats
def transform_text(self, text: str, mode: str) -> str:
"""Transform text according to specified mode"""
if mode == 'upper':
return text.upper()
elif mode == 'lower':
return text.lower()
elif mode == 'title':
return text.title()
elif mode == 'reverse':
return text[::-1]
else:
raise ValueError(f"Unknown transformation mode: {mode}")
def process_file(self, file_path: str) -> Dict[str, Any]:
"""Process a single text file"""
try:
with open(file_path, 'r', encoding=self.encoding) as file:
content = file.read()
stats = self.analyze_text(content)
stats['file'] = file_path
stats['file_size'] = os.path.getsize(file_path)
return stats
except FileNotFoundError:
raise FileNotFoundError(f"File not found: {file_path}")
except UnicodeDecodeError:
raise UnicodeDecodeError(f"Cannot decode file with {self.encoding} encoding: {file_path}")
except PermissionError:
raise PermissionError(f"Permission denied accessing file: {file_path}")
class OutputFormatter:
"""Handles dual output format generation"""
@staticmethod
def format_json(data: Dict[str, Any]) -> str:
"""Format data as JSON"""
return json.dumps(data, indent=2, ensure_ascii=False)
@staticmethod
def format_human_readable(data: Dict[str, Any]) -> str:
"""Format data as human-readable text"""
lines = []
lines.append("=== TEXT ANALYSIS RESULTS ===")
lines.append(f"File: {data.get('file', 'Unknown')}")
lines.append(f"File size: {data.get('file_size', 0)} bytes")
lines.append(f"Total words: {data.get('total_words', 0)}")
lines.append(f"Unique words: {data.get('unique_words', 0)}")
lines.append(f"Total characters: {data.get('total_characters', 0)}")
lines.append(f"Lines: {data.get('lines', 0)}")
lines.append(f"Average word length: {data.get('average_word_length', 0):.1f}")
most_frequent = data.get('most_frequent', {})
lines.append(f"Most frequent word: \"{most_frequent.get('word', '')}\" ({most_frequent.get('count', 0)} occurrences)")
return "\n".join(lines)
class FileManager:
"""Manages file I/O operations and batch processing"""
def __init__(self, verbose: bool = False):
self.verbose = verbose
def log_verbose(self, message: str):
"""Log verbose message if verbose mode enabled"""
if self.verbose:
print(f"[INFO] {message}", file=sys.stderr)
def find_text_files(self, directory: str) -> List[str]:
"""Find all text files in directory"""
text_extensions = {'.txt', '.md', '.rst', '.csv', '.log'}
text_files = []
try:
for file_path in Path(directory).rglob('*'):
if file_path.is_file() and file_path.suffix.lower() in text_extensions:
text_files.append(str(file_path))
except PermissionError:
raise PermissionError(f"Permission denied accessing directory: {directory}")
return text_files
def write_output(self, content: str, output_path: Optional[str] = None):
"""Write content to file or stdout"""
if output_path:
try:
# Create directory if needed
output_dir = os.path.dirname(output_path)
if output_dir and not os.path.exists(output_dir):
os.makedirs(output_dir)
with open(output_path, 'w', encoding='utf-8') as file:
file.write(content)
self.log_verbose(f"Output written to: {output_path}")
except PermissionError:
raise PermissionError(f"Permission denied writing to: {output_path}")
else:
print(content)
def analyze_command(args: argparse.Namespace) -> int:
"""Handle analyze command"""
try:
processor = TextProcessor(args.encoding)
file_manager = FileManager(args.verbose)
file_manager.log_verbose(f"Analyzing file: {args.file}")
# Process the file
results = processor.process_file(args.file)
# Format output
if args.format == 'json':
output = OutputFormatter.format_json(results)
else:
output = OutputFormatter.format_human_readable(results)
# Write output
file_manager.write_output(output, args.output)
return 0
except FileNotFoundError as e:
print(f"Error: {e}", file=sys.stderr)
return 1
except UnicodeDecodeError as e:
print(f"Error: {e}", file=sys.stderr)
print(f"Try using --encoding option with different encoding", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
def transform_command(args: argparse.Namespace) -> int:
"""Handle transform command"""
try:
processor = TextProcessor(args.encoding)
file_manager = FileManager(args.verbose)
file_manager.log_verbose(f"Transforming file: {args.file}")
# Read and transform the file
with open(args.file, 'r', encoding=args.encoding) as file:
content = file.read()
transformed = processor.transform_text(content, args.mode)
# Write transformed content
file_manager.write_output(transformed, args.output)
return 0
except FileNotFoundError as e:
print(f"Error: {e}", file=sys.stderr)
return 1
except ValueError as e:
print(f"Error: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
def batch_command(args: argparse.Namespace) -> int:
"""Handle batch command"""
try:
processor = TextProcessor(args.encoding)
file_manager = FileManager(args.verbose)
file_manager.log_verbose(f"Finding text files in: {args.directory}")
# Find all text files
text_files = file_manager.find_text_files(args.directory)
if not text_files:
print(f"No text files found in directory: {args.directory}", file=sys.stderr)
return 1
file_manager.log_verbose(f"Found {len(text_files)} text files")
# Process all files
all_results = []
for i, file_path in enumerate(text_files, 1):
try:
file_manager.log_verbose(f"Processing {i}/{len(text_files)}: {file_path}")
results = processor.process_file(file_path)
all_results.append(results)
except Exception as e:
print(f"Warning: Failed to process {file_path}: {e}", file=sys.stderr)
continue
if not all_results:
print("Error: No files could be processed successfully", file=sys.stderr)
return 1
# Format batch results
batch_summary = {
'total_files': len(all_results),
'total_words': sum(r.get('total_words', 0) for r in all_results),
'total_characters': sum(r.get('total_characters', 0) for r in all_results),
'files': all_results
}
if args.format == 'json':
output = OutputFormatter.format_json(batch_summary)
else:
lines = []
lines.append("=== BATCH PROCESSING RESULTS ===")
lines.append(f"Total files processed: {batch_summary['total_files']}")
lines.append(f"Total words across all files: {batch_summary['total_words']}")
lines.append(f"Total characters across all files: {batch_summary['total_characters']}")
lines.append("")
lines.append("Individual file results:")
for result in all_results:
lines.append(f" {result['file']}: {result['total_words']} words")
output = "\n".join(lines)
# Write output
file_manager.write_output(output, args.output)
return 0
except PermissionError as e:
print(f"Error: {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
def main():
"""Main entry point with argument parsing"""
parser = argparse.ArgumentParser(
description="Sample Text Processor - Basic text analysis and transformation",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
Analysis:
python text_processor.py analyze document.txt
python text_processor.py analyze document.txt --format json --output results.json
Transformation:
python text_processor.py transform document.txt --mode upper
python text_processor.py transform document.txt --mode title --output transformed.txt
Batch processing:
python text_processor.py batch text_files/ --verbose
python text_processor.py batch text_files/ --format json --output batch_results.json
Transformation modes:
upper - Convert to uppercase
lower - Convert to lowercase
title - Convert to title case
reverse - Reverse the text
"""
)
parser.add_argument('--format',
choices=['json', 'text'],
default='text',
help='Output format (default: text)')
parser.add_argument('--output',
help='Output file path (default: stdout)')
parser.add_argument('--encoding',
default='utf-8',
help='Text file encoding (default: utf-8)')
parser.add_argument('--verbose',
action='store_true',
help='Enable verbose output')
subparsers = parser.add_subparsers(dest='command', help='Available commands')
# Analyze subcommand
analyze_parser = subparsers.add_parser('analyze', help='Analyze text file statistics')
analyze_parser.add_argument('file', help='Text file to analyze')
# Transform subcommand
transform_parser = subparsers.add_parser('transform', help='Transform text file')
transform_parser.add_argument('file', help='Text file to transform')
transform_parser.add_argument('--mode',
required=True,
choices=['upper', 'lower', 'title', 'reverse'],
help='Transformation mode')
# Batch subcommand
batch_parser = subparsers.add_parser('batch', help='Process multiple files')
batch_parser.add_argument('directory', help='Directory containing text files')
args = parser.parse_args()
if not args.command:
parser.print_help()
return 1
try:
if args.command == 'analyze':
return analyze_command(args)
elif args.command == 'transform':
return transform_command(args)
elif args.command == 'batch':
return batch_command(args)
else:
print(f"Unknown command: {args.command}", file=sys.stderr)
return 1
except KeyboardInterrupt:
print("\nOperation interrupted by user", file=sys.stderr)
return 130
except Exception as e:
print(f"Unexpected error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())Theo dõi thay đổi kỹ thuật bằng bản ghi có cấu trúc, máy trạng thái và bàn giao giữa các phiên làm việc.
---
name: tc
description: Track technical changes with structured records, a state machine, and session handoff. Usage: /tc <init|create|update|status|resume|close|export|dashboard> [args]
---
# /tc — Technical Change Tracker
Dispatch a TC (Technical Change) command. Arguments: `$ARGUMENTS`.
If `$ARGUMENTS` is empty, print this menu and stop:
```
/tc init Initialize TC tracking in this project
/tc create <name> Create a new TC record
/tc update <tc-id> [...] Update fields, status, files, handoff
/tc status [tc-id] Show one TC or the registry summary
/tc resume <tc-id> Resume a TC from a previous session
/tc close <tc-id> Transition a TC to deployed
/tc export Re-render derived artifacts
/tc dashboard Re-render the registry summary
```
Otherwise, parse `$ARGUMENTS` as `<subcommand> <rest>` and dispatch to the matching protocol below. All scripts live at `engineering/tc-tracker/scripts/`.
## Subcommands
### `init`
1. Run:
```bash
python3 engineering/tc-tracker/scripts/tc_init.py --root . --json
```
2. If status is `already_initialized`, report current statistics and stop.
3. Otherwise report what was created and suggest `/tc create <name>` as the next step.
### `create <name>`
1. Parse `<name>` as a kebab-case slug. If missing, ask the user for one.
2. Prompt the user (one question at a time) for:
- Title (5-120 chars)
- Scope: `feature | bugfix | refactor | infrastructure | documentation | hotfix | enhancement`
- Priority: `critical | high | medium | low` (default `medium`)
- Summary (10+ chars)
- Motivation
3. Run:
```bash
python3 engineering/tc-tracker/scripts/tc_create.py --root . \
--name "<slug>" --title "<title>" --scope <scope> --priority <priority> \
--summary "<summary>" --motivation "<motivation>" --json
```
4. Report the new TC ID and the path to the record.
### `update <tc-id> [intent]`
1. If `<tc-id>` is missing, list active TCs (status `in_progress` or `blocked`) from `tc_status.py --all` and ask which one.
2. Determine the user's intent from natural language:
- **Status change** → `--set-status <state>` with `--reason "<why>"`
- **Add files** → one or more `--add-file path[:action]`
- **Add a test** → `--add-test "<title>" --test-procedure "<step>" --test-expected "<result>"`
- **Update handoff** → any combination of `--handoff-progress`, `--handoff-next`, `--handoff-blocker`, `--handoff-context`
- **Add a note** → `--note "<text>"`
- **Add a tag** → `--tag <tag>`
3. Run:
```bash
python3 engineering/tc-tracker/scripts/tc_update.py --root . --tc-id <tc-id> [flags] --json
```
4. If exit code is non-zero, surface the error verbatim. The state machine and validator will reject invalid moves — do not retry blindly.
### `status [tc-id]`
- If `<tc-id>` is provided:
```bash
python3 engineering/tc-tracker/scripts/tc_status.py --root . --tc-id <tc-id>
```
- Otherwise:
```bash
python3 engineering/tc-tracker/scripts/tc_status.py --root . --all
```
### `resume <tc-id>`
1. Run:
```bash
python3 engineering/tc-tracker/scripts/tc_status.py --root . --tc-id <tc-id> --json
```
2. Display the handoff block prominently: `progress_summary`, `next_steps` (numbered), `blockers`, `key_context`.
3. Ask: "Resume <tc-id> and pick up at next step 1? (y/n)"
4. If yes, run an update to record the resumption:
```bash
python3 engineering/tc-tracker/scripts/tc_update.py --root . --tc-id <tc-id> \
--note "Session resumed" --reason "session handoff"
```
5. Begin executing the first item in `next_steps`. Do NOT re-derive context — trust the handoff.
### `close <tc-id>`
1. Read the record via `tc_status.py --tc-id <tc-id> --json`.
2. Verify the current status is `tested`. If not, refuse and tell the user which transitions are still required.
3. Check `test_cases`: warn if any are `pending`, `fail`, or `blocked`.
4. Ask the user:
- "Who is approving? (your name, or 'self')"
- "Approval notes (optional):"
- "Test coverage status: none / partial / full"
5. Run:
```bash
python3 engineering/tc-tracker/scripts/tc_update.py --root . --tc-id <tc-id> \
--set-status deployed --reason "Approved by <approver>" --note "Approval: <approver> — <notes>"
```
Then directly edit the `approval` block via a follow-up update if your script version supports it; otherwise instruct the user to record approval in `notes`.
6. Report: "TC-NNN closed and deployed."
### `export`
There is no automatic HTML export in this skill. Re-validate everything instead:
1. Read the registry.
2. For each record, run:
```bash
python3 engineering/tc-tracker/scripts/tc_validator.py --record <path> --json
```
3. Run:
```bash
python3 engineering/tc-tracker/scripts/tc_validator.py --registry docs/TC/tc_registry.json --json
```
4. Report: total records validated, any errors, paths to anything invalid.
### `dashboard`
Run the all-records summary:
```bash
python3 engineering/tc-tracker/scripts/tc_status.py --root . --all
```
## Iron Rules
1. **Never edit `tc_record.json` by hand.** Always use `tc_update.py` so revision history is appended and validation runs.
2. **Never skip the state machine.** Walk forward through states even if it feels redundant.
3. **Never delete a TC.** History is append-only — add a final revision and tag it `[CANCELLED]`.
4. **Background bookkeeping.** When mid-task, spawn a background subagent to update the TC. Do not pause coding to do paperwork.
5. **Validate before reporting success.** If a script exits non-zero, surface the error and stop.
## Related Skills
- `engineering/tc-tracker` — Full SKILL.md with schema reference, lifecycle diagrams, and the handoff format.
- `engineering/changelog-generator` — Pair with TC tracker: TCs for the per-change audit trail, changelog for user-facing release notes.
- `engineering/tech-debt-tracker` — For tracking long-lived debt rather than discrete code changes.
Lập kế hoạch nghiên cứu, tạo persona, vẽ hành trình người dùng và phân tích kết quả kiểm thử khả dụng.
---
name: cs-ux-researcher
description: UX research agent for research planning, persona generation, journey mapping, and usability test analysis
skills: product-team/ux-researcher-designer, product-team/product-manager-toolkit, product-team/ui-design-system
domain: product
model: sonnet
tools: [Read, Write, Bash, Grep, Glob]
---
# UX Researcher Agent
## Purpose
The cs-ux-researcher agent is a specialized user experience research agent focused on research planning, persona creation, journey mapping, and usability test analysis. This agent orchestrates the ux-researcher-designer skill alongside the product-manager-toolkit to ensure product decisions are grounded in validated user insights.
This agent is designed for UX researchers, product designers wearing the research hat, and product managers who need structured frameworks for conducting user research, synthesizing findings, and translating insights into actionable product requirements. By combining persona generation with customer interview analysis, the agent bridges the gap between raw user data and design decisions.
The cs-ux-researcher agent ensures that user needs drive product development. It provides methodological rigor for research planning, data-driven persona creation, systematic journey mapping, and structured usability evaluation. The agent works closely with the ui-design-system skill for design handoff and with the product-manager-toolkit for translating research insights into prioritized feature requirements.
## Skill Integration
**Primary Skill:** `../../product-team/ux-researcher-designer/`
### All Orchestrated Skills
| # | Skill | Location | Primary Tool |
|---|-------|----------|-------------|
| 1 | UX Researcher & Designer | `../../product-team/ux-researcher-designer/` | persona_generator.py |
| 2 | Product Manager Toolkit | `../../product-team/product-manager-toolkit/` | customer_interview_analyzer.py |
| 3 | UI Design System | `../../product-team/ui-design-system/` | design_token_generator.py |
### Python Tools
1. **Persona Generator**
- **Purpose:** Create data-driven user personas from research inputs including demographics, goals, pain points, and behavioral patterns
- **Path:** `../../product-team/ux-researcher-designer/scripts/persona_generator.py`
- **Usage:** `python ../../product-team/ux-researcher-designer/scripts/persona_generator.py research-data.json`
- **Features:** Multiple persona generation, behavioral segmentation, needs hierarchy mapping, empathy map creation
- **Use Cases:** Persona development, user segmentation, design alignment, stakeholder communication
2. **Customer Interview Analyzer**
- **Purpose:** NLP-based analysis of interview transcripts to extract pain points, feature requests, themes, and sentiment
- **Path:** `../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py`
- **Usage:** `python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview.txt`
- **Features:** Pain point extraction with severity scoring, feature request identification, jobs-to-be-done patterns, theme clustering, key quote extraction
- **Use Cases:** Interview synthesis, discovery validation, problem prioritization, insight aggregation
3. **Design Token Generator**
- **Purpose:** Generate design tokens for consistent UI implementation across platforms
- **Path:** `../../product-team/ui-design-system/scripts/design_token_generator.py`
- **Usage:** `python ../../product-team/ui-design-system/scripts/design_token_generator.py theme.json`
- **Use Cases:** Research-informed design system updates, accessibility token adjustments
### Knowledge Bases
1. **Persona Methodology**
- **Location:** `../../product-team/ux-researcher-designer/references/persona-methodology.md`
- **Content:** Research-backed persona creation methodology, data collection strategies, validation approaches
- **Use Case:** Methodological guidance for persona projects
2. **Example Personas**
- **Location:** `../../product-team/ux-researcher-designer/references/example-personas.md`
- **Content:** Sample persona documents with demographics, goals, pain points, behaviors, scenarios
- **Use Case:** Persona format reference, team training
3. **Journey Mapping Guide**
- **Location:** `../../product-team/ux-researcher-designer/references/journey-mapping-guide.md`
- **Content:** Customer journey mapping methodology, touchpoint analysis, emotion mapping, opportunity identification
- **Use Case:** Journey map creation, experience design, service design
4. **Usability Testing Frameworks**
- **Location:** `../../product-team/ux-researcher-designer/references/usability-testing-frameworks.md`
- **Content:** Test planning, task design, analysis methods, severity ratings, reporting formats
- **Use Case:** Usability study design, prototype validation, UX evaluation
5. **Component Architecture**
- **Location:** `../../product-team/ui-design-system/references/component-architecture.md`
- **Content:** Component hierarchy, atomic design patterns, composition strategies
- **Use Case:** Research-to-design translation, component recommendations
6. **Developer Handoff**
- **Location:** `../../product-team/ui-design-system/references/developer-handoff.md`
- **Content:** Design-to-dev handoff process, specification formats, asset delivery
- **Use Case:** Translating research findings into implementation specs
### Templates
1. **Research Plan Template**
- **Location:** `../../product-team/ux-researcher-designer/assets/research_plan_template.md`
- **Use Case:** Structuring research studies with methodology, participants, and analysis plan
2. **Design System Documentation Template**
- **Location:** `../../product-team/ui-design-system/assets/design_system_doc_template.md`
- **Use Case:** Documenting research-informed design system decisions
## Workflows
### Workflow 1: Research Plan Creation
**Goal:** Design a rigorous research study that answers specific product questions with appropriate methodology
**Steps:**
1. **Define Research Questions** - Identify what needs to be learned:
- What are the top 3-5 questions stakeholders need answered?
- What do we already know from existing data?
- What assumptions need validation?
- What decisions will this research inform?
2. **Select Methodology** - Choose the right approach:
```bash
# Review usability testing frameworks for method selection
cat ../../product-team/ux-researcher-designer/references/usability-testing-frameworks.md
```
- **Exploratory** (interviews, contextual inquiry): When learning about problem space
- **Evaluative** (usability testing, A/B tests): When validating solutions
- **Generative** (diary studies, card sorting): When discovering new opportunities
- **Quantitative** (surveys, analytics): When measuring scale and significance
3. **Define Participants** - Screen for the right users:
- Target persona(s) to recruit
- Screening criteria (role, experience, usage patterns)
- Sample size justification
- Recruitment channels and incentives
4. **Create Study Materials** - Prepare research instruments:
```bash
# Use the research plan template
cat ../../product-team/ux-researcher-designer/assets/research_plan_template.md
```
- Interview guide or test script
- Task scenarios (for usability tests)
- Consent form and recording permissions
- Analysis framework and coding scheme
5. **Align with Stakeholders** - Get buy-in:
- Share research plan with product and engineering leads
- Invite stakeholders to observe sessions
- Set expectations for timeline and deliverables
- Define how findings will be actioned
**Expected Output:** Complete research plan with questions, methodology, participant criteria, study materials, timeline, and stakeholder alignment
**Time Estimate:** 2-3 days for plan creation
**Example:**
```bash
# Create research plan from template
cp ../../product-team/ux-researcher-designer/assets/research_plan_template.md onboarding-research-plan.md
# Review methodology options
cat ../../product-team/ux-researcher-designer/references/usability-testing-frameworks.md
# Review persona methodology for participant criteria
cat ../../product-team/ux-researcher-designer/references/persona-methodology.md
```
### Workflow 2: Persona Generation
**Goal:** Create data-driven user personas from research data that align product teams around real user needs
**Steps:**
1. **Gather Research Data** - Collect inputs from multiple sources:
- Interview transcripts (analyzed for themes)
- Survey responses (demographic and behavioral data)
- Analytics data (usage patterns, feature adoption)
- Support tickets (common issues, pain points)
- Sales call notes (buyer motivations, objections)
2. **Analyze Interview Data** - Extract structured insights:
```bash
# Analyze each interview transcript
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-001.txt > insights-001.json
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-002.txt > insights-002.json
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py interview-003.txt > insights-003.json
```
3. **Identify Behavioral Segments** - Cluster users by:
- Goals and motivations (what they are trying to achieve)
- Behaviors and workflows (how they work today)
- Pain points and frustrations (what blocks them)
- Technical sophistication (how they interact with tools)
- Decision-making factors (what drives their choices)
4. **Generate Personas** - Create data-backed personas:
```bash
# Generate personas from aggregated research
python ../../product-team/ux-researcher-designer/scripts/persona_generator.py research-data.json
```
5. **Validate Personas** - Ensure accuracy:
- Cross-reference with quantitative data (segment sizes)
- Review with customer-facing teams (sales, support)
- Test with stakeholders who interact with users
- Confirm each persona represents a meaningful segment
6. **Socialize Personas** - Make personas actionable:
```bash
# Review example personas for format guidance
cat ../../product-team/ux-researcher-designer/references/example-personas.md
```
- Create one-page persona cards for team walls/wikis
- Present to product, engineering, and design teams
- Map personas to product areas and features
- Reference personas in PRDs and design briefs
**Expected Output:** 3-5 validated user personas with demographics, goals, pain points, behaviors, and scenarios
**Time Estimate:** 1-2 weeks (data collection through socialization)
**Example:**
```bash
# Full persona generation workflow
echo "Persona Generation Workflow"
echo "==========================="
# Step 1: Analyze interviews
for f in interviews/*.txt; do
base=$(basename "$f" .txt)
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py "$f" json > "insights-$base.json"
echo "Analyzed: $f"
done
# Step 2: Review persona methodology
cat ../../product-team/ux-researcher-designer/references/persona-methodology.md
# Step 3: Generate personas
python ../../product-team/ux-researcher-designer/scripts/persona_generator.py research-data.json
# Step 4: Review example format
cat ../../product-team/ux-researcher-designer/references/example-personas.md
```
### Workflow 3: Journey Mapping
**Goal:** Map the complete user journey to identify pain points, opportunities, and moments that matter
**Steps:**
1. **Define Journey Scope** - Set boundaries:
- Which persona is this journey for?
- What is the starting trigger?
- What is the end state (success)?
- What timeframe does the journey cover?
2. **Review Journey Mapping Methodology** - Understand the framework:
```bash
cat ../../product-team/ux-researcher-designer/references/journey-mapping-guide.md
```
3. **Map Journey Stages** - Identify key phases:
- **Awareness:** How users discover the product
- **Consideration:** How users evaluate and compare
- **Onboarding:** First-time setup and activation
- **Regular Use:** Core workflow and daily interactions
- **Growth:** Expanding usage, inviting team, upgrading
- **Advocacy:** Referring others, providing feedback
4. **Document Touchpoints** - For each stage:
- User actions (what they do)
- Channels (where they interact)
- Emotions (how they feel)
- Pain points (what frustrates them)
- Opportunities (how we can improve)
5. **Identify Moments of Truth** - Critical experience points:
- First-time use (aha moment)
- First success (value realization)
- First problem (support experience)
- Upgrade decision (value justification)
- Referral moment (advocacy trigger)
6. **Prioritize Opportunities** - Focus on highest-impact improvements:
```bash
# Prioritize journey improvement opportunities
cat > journey-opportunities.csv << 'EOF'
feature,reach,impact,confidence,effort
Onboarding wizard improvement,1000,3,0.9,3
First-success celebration,800,2,0.7,1
Self-service help in context,600,2,0.8,2
Upgrade prompt optimization,400,3,0.6,2
EOF
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py journey-opportunities.csv
```
**Expected Output:** Visual journey map with stages, touchpoints, emotions, pain points, and prioritized improvement opportunities
**Time Estimate:** 1-2 weeks for research-backed journey map
**Example:**
```bash
# Journey mapping workflow
echo "Journey Mapping - Onboarding Flow"
echo "=================================="
# Review journey mapping methodology
cat ../../product-team/ux-researcher-designer/references/journey-mapping-guide.md
# Analyze relevant interview transcripts for journey insights
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py onboarding-interview-01.txt
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py onboarding-interview-02.txt
# Prioritize improvement opportunities
python ../../product-team/product-manager-toolkit/scripts/rice_prioritizer.py journey-opportunities.csv
```
### Workflow 4: Usability Test Analysis
**Goal:** Conduct and analyze usability tests to evaluate design solutions and identify critical UX issues
**Steps:**
1. **Plan the Test** - Design the study:
```bash
# Review usability testing frameworks
cat ../../product-team/ux-researcher-designer/references/usability-testing-frameworks.md
```
- Define test objectives (what decisions will this inform)
- Select test type (moderated/unmoderated, remote/in-person)
- Write task scenarios (realistic, goal-oriented)
- Set success criteria per task (completion, time, errors)
2. **Prepare Materials** - Set up the test:
- Prototype or staging environment ready
- Test script with introduction, tasks, and debrief questions
- Recording tools configured
- Note-taking template for observers
- Use research plan template for documentation:
```bash
cat ../../product-team/ux-researcher-designer/assets/research_plan_template.md
```
3. **Conduct Sessions** - Run 5-8 sessions:
- Follow consistent script for each participant
- Use think-aloud protocol
- Note task completion, errors, and verbal feedback
- Capture quotes and emotional reactions
- Debrief after each session
4. **Analyze Results** - Synthesize findings:
- Calculate task success rates
- Measure time-on-task per scenario
- Categorize usability issues by severity:
- **Critical:** Prevents task completion
- **Major:** Causes significant difficulty or errors
- **Minor:** Creates confusion but user recovers
- **Cosmetic:** Aesthetic or minor friction
- Identify patterns across participants
5. **Analyze Verbal Feedback** - Extract qualitative insights:
```bash
# Analyze session transcripts for themes
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py usability-session-01.txt
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py usability-session-02.txt
```
6. **Create Report and Recommendations** - Deliver findings:
- Executive summary (key findings in 3-5 bullets)
- Task-by-task results with evidence
- Prioritized issue list with severity
- Recommended design changes
- Highlight reel of key moments (video clips)
7. **Inform Design Iteration** - Close the loop:
- Review findings with design team
- Map issues to components in design system:
```bash
cat ../../product-team/ui-design-system/references/component-architecture.md
```
- Create Jira tickets for each issue
- Plan re-test for critical issues after fixes
**Expected Output:** Usability test report with task metrics, severity-rated issues, recommendations, and design iteration plan
**Time Estimate:** 2-3 weeks (planning through report delivery)
**Example:**
```bash
# Usability test analysis workflow
echo "Usability Test Analysis"
echo "======================="
# Review frameworks
cat ../../product-team/ux-researcher-designer/references/usability-testing-frameworks.md
# Analyze each session transcript
for i in 1 2 3 4 5; do
echo "Session $i Analysis:"
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py "usability-session-0$i.txt"
echo ""
done
# Review component architecture for design recommendations
cat ../../product-team/ui-design-system/references/component-architecture.md
```
## Integration Examples
### Example 1: Discovery Sprint Research
```bash
#!/bin/bash
# discovery-research.sh - 2-week discovery sprint
echo "Discovery Sprint Research"
echo "========================="
# Week 1: Research execution
echo ""
echo "Week 1: Conduct & Analyze Interviews"
echo "-------------------------------------"
# Analyze all interview transcripts
for f in discovery-interviews/*.txt; do
base=$(basename "$f" .txt)
echo "Analyzing: $base"
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py "$f" json > "insights/$base.json"
done
# Week 2: Synthesis
echo ""
echo "Week 2: Generate Personas & Journey Map"
echo "----------------------------------------"
# Generate personas from aggregated data
python ../../product-team/ux-researcher-designer/scripts/persona_generator.py aggregated-research.json
# Reference journey mapping guide
echo "Journey mapping guide: ../../product-team/ux-researcher-designer/references/journey-mapping-guide.md"
```
### Example 2: Research Repository Update
```bash
#!/bin/bash
# research-update.sh - Monthly research insights update
echo "Research Repository Update - $(date +%Y-%m-%d)"
echo "================================================"
# Process new interviews
echo ""
echo "New Interview Analysis:"
for f in new-interviews/*.txt; do
python ../../product-team/product-manager-toolkit/scripts/customer_interview_analyzer.py "$f"
echo "---"
done
# Review and refresh personas
echo ""
echo "Persona Review:"
echo "Current personas: ../../product-team/ux-researcher-designer/references/example-personas.md"
echo "Methodology: ../../product-team/ux-researcher-designer/references/persona-methodology.md"
```
### Example 3: Design Handoff with Research Context
```bash
#!/bin/bash
# research-handoff.sh - Prepare research context for design team
echo "Research Handoff Package"
echo "========================"
# Persona context
echo ""
echo "1. Active Personas:"
cat ../../product-team/ux-researcher-designer/references/example-personas.md | head -30
# Journey context
echo ""
echo "2. Journey Map Reference:"
echo "See: ../../product-team/ux-researcher-designer/references/journey-mapping-guide.md"
# Design system alignment
echo ""
echo "3. Component Architecture:"
echo "See: ../../product-team/ui-design-system/references/component-architecture.md"
# Developer handoff process
echo ""
echo "4. Handoff Process:"
echo "See: ../../product-team/ui-design-system/references/developer-handoff.md"
```
## Success Metrics
**Research Quality:**
- **Study Rigor:** 100% of studies have documented research plan with methodology justification
- **Participant Quality:** >90% of participants match screening criteria
- **Insight Actionability:** >80% of research findings result in backlog items or design changes
- **Stakeholder Engagement:** >2 stakeholders observe each research session
**Persona Effectiveness:**
- **Team Adoption:** >80% of PRDs reference a specific persona
- **Validation Rate:** Personas validated with quantitative data (segment sizes, usage patterns)
- **Refresh Cadence:** Personas reviewed and updated at least semi-annually
- **Decision Influence:** Personas cited in >50% of product design decisions
**Usability Impact:**
- **Issue Detection:** 5+ unique usability issues identified per study
- **Fix Rate:** >70% of critical/major issues resolved within 2 sprints
- **Task Success:** Average task success rate improves by >15% after design iteration
- **User Satisfaction:** SUS score improves by >5 points after research-informed redesign
**Business Impact:**
- **Customer Satisfaction:** NPS improvement correlated with research-informed changes
- **Onboarding Conversion:** First-time user activation rate improvement
- **Support Ticket Reduction:** Fewer UX-related support requests
- **Feature Adoption:** Research-informed features show >20% higher adoption rates
## Related Agents
- [cs-product-manager](cs-product-manager.md) - Product management lifecycle, interview analysis, PRD development
- [cs-agile-product-owner](cs-agile-product-owner.md) - Translating research findings into user stories
- [cs-product-strategist](cs-product-strategist.md) - Strategic research to validate product vision and positioning
- UI Design System - Design handoff and component recommendations (see `../../product-team/ui-design-system/`)
## References
- **Primary Skill:** [../../product-team/ux-researcher-designer/SKILL.md](../../product-team/ux-researcher-designer/SKILL.md)
- **Interview Analyzer:** [../../product-team/product-manager-toolkit/SKILL.md](../../product-team/product-manager-toolkit/SKILL.md)
- **Persona Methodology:** [../../product-team/ux-researcher-designer/references/persona-methodology.md](../../product-team/ux-researcher-designer/references/persona-methodology.md)
- **Journey Mapping Guide:** [../../product-team/ux-researcher-designer/references/journey-mapping-guide.md](../../product-team/ux-researcher-designer/references/journey-mapping-guide.md)
- **Usability Testing:** [../../product-team/ux-researcher-designer/references/usability-testing-frameworks.md](../../product-team/ux-researcher-designer/references/usability-testing-frameworks.md)
- **Design System:** [../../product-team/ui-design-system/SKILL.md](../../product-team/ui-design-system/SKILL.md)
- **Product Domain Guide:** [../../product-team/CLAUDE.md](../../product-team/CLAUDE.md)
- **Agent Development Guide:** [../CLAUDE.md](../CLAUDE.md)
---
**Last Updated:** March 9, 2026
**Status:** Production Ready
**Version:** 1.0
Theo dõi đối thủ có hệ thống, phục vụ định vị, battlecard bán hàng và quyết định lộ trình sản phẩm.
---
name: "context-engine"
description: "Loads and manages company context for all C-suite advisor skills. Reads ~/.claude/company-context.md, detects stale context (>90 days), enriches context during conversations, and enforces privacy/anonymization rules before external API calls."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: orchestration
updated: 2026-03-05
frameworks: context-loading, anonymization, context-enrichment
---
# Company Context Engine
The memory layer for C-suite advisors. Every advisor skill loads this first. Context is what turns generic advice into specific insight.
## Keywords
company context, context loading, context engine, company profile, advisor context, stale context, context refresh, privacy, anonymization
---
## Load Protocol (Run at Start of Every C-Suite Session)
**Step 1 — Check for context file:** `~/.claude/company-context.md`
- Exists → proceed to Step 2
- Missing → prompt: *"Run /cs:setup to build your company context — it makes every advisor conversation significantly more useful."*
**Step 2 — Check staleness:** Read `Last updated` field.
- **< 90 days:** Load and proceed.
- **≥ 90 days:** Prompt: *"Your context is [N] days old. Quick 15-min refresh (/cs:update), or continue with what I have?"*
- If continue: load with `[STALE — last updated DATE]` noted internally.
**Step 3 — Parse into working memory.** Always active:
- Company stage (pre-PMF / scaling / optimizing)
- Founder archetype (product / sales / technical / operator)
- Current #1 challenge
- Runway (as risk signal — never share externally)
- Team size
- Unfair advantage
- 12-month target
---
## Context Quality Signals
| Condition | Confidence | Action |
|-----------|-----------|--------|
| < 30 days, full interview | High | Use directly |
| 30–90 days, update done | Medium | Use, flag what may have changed |
| > 90 days | Low | Flag stale, prompt refresh |
| Key fields missing | Low | Ask in-session |
| No file | None | Prompt /cs:setup |
If Low: *"My context is [stale/incomplete] — I'm assuming [X]. Correct me if I'm wrong."*
---
## Context Enrichment
During conversations, you'll learn things not in the file. Capture them.
**Triggers:** New number or timeline revealed, key person mentioned, priority shift, constraint surfaces.
**Protocol:**
1. Note internally: `[CONTEXT UPDATE: {what was learned}]`
2. At session end: *"I picked up a few things to add to your context. Want me to update the file?"*
3. If yes: append to the relevant dimension, update timestamp.
**Never silently overwrite.** Always confirm before modifying the context file.
---
## Privacy Rules
### Never send externally
- Specific revenue or burn figures
- Customer names
- Employee names (unless publicly known)
- Investor names (unless public)
- Specific runway months
- Watch List contents
### Safe to use externally (with anonymization)
- Stage label
- Team size ranges (1–10, 10–50, 50–200+)
- Industry vertical
- Challenge category
- Market position descriptor
### Before any external API call or web search
Apply `references/anonymization-protocol.md`:
- Numbers → ranges or stage-relative descriptors
- Names → roles
- Revenue → percentages or stage labels
- Customers → "Customer A, B, C"
---
## Missing or Partial Context
Handle gracefully — never block the conversation.
- **Missing stage:** "Just to calibrate — are you still finding PMF or scaling what works?"
- **Missing financials:** Use stage + team size to infer. Note the gap.
- **Missing founder profile:** Infer from conversation style. Mark as inferred.
- **Multiple founders:** Context reflects the interviewee. Note co-founder perspective may differ.
---
## Required Context Fields
```
Required:
- Last updated (date)
- Company Identity → What we do
- Stage & Scale → Stage
- Founder Profile → Founder archetype
- Current Challenges → Priority #1
- Goals & Ambition → 12-month target
High-value optional:
- Unfair advantage
- Kill-shot risk
- Avoided decision
- Watch list
```
Missing required fields: note gaps, work around in session, ask in-session only when critical.
---
## References
- `references/anonymization-protocol.md` — detailed rules for stripping sensitive data before external calls
FILE:references/anonymization-protocol.md
# Anonymization Protocol
Rules for stripping sensitive company data before any external API call, web search, or tool invocation that sends data outside the local environment.
---
## When This Protocol Applies
**Trigger:** Any time company context or conversation content will leave the local session.
Examples:
- Web search that includes company specifics
- External API call with company data in the payload
- Any tool call where conversation content is part of the request
**Does NOT apply to:**
- Local file reads/writes (`~/.claude/company-context.md`)
- In-session reasoning and analysis
- Generating advice or documents that stay local
---
## Rule 1: Financial Figures → Relative Ranges
Never send specific financial data externally.
| Raw data | Anonymized version |
|----------|-------------------|
| "$2.4M ARR" | "early-stage ARR (sub-$5M)" |
| "$180K MRR" | "growing MRR, Series A range" |
| "14 months runway" | "runway is healthy for stage" |
| "burn rate is $320K/month" | "burn rate is moderate for stage" |
| "raised $8M Series A" | "Series A company" |
| "customer LTV is $4,200" | "LTV is above industry average for segment" |
| "CAC is $680" | "CAC is in a sustainable range" |
**Rule:** No dollar amounts. No month counts for runway. Use stage-relative descriptors.
---
## Rule 2: Customer Names → Anonymized Labels
Never send customer or client names externally.
| Raw data | Anonymized version |
|----------|-------------------|
| "Acme Corp is our biggest customer" | "Customer A (largest account)" |
| "we're working with NHS England" | "a large public-sector customer" |
| "BMW, Volkswagen, and Stellantis" | "three major automotive OEMs" |
| "10 enterprise customers including..." | "10 enterprise customers" |
**Rule:** Use "Customer A/B/C" for named accounts, or describe by segment without naming.
---
## Rule 3: Revenue Figures → Percentage Changes or Stage Descriptors
Revenue trajectory is safer than absolute numbers.
| Raw data | Anonymized version |
|----------|-------------------|
| "growing from $1M to $2M ARR" | "2x revenue growth year-over-year" |
| "revenue dropped from $500K to $430K" | "revenue declined ~15% in the period" |
| "hit $10M ARR last quarter" | "crossed a significant ARR milestone" |
| "doing $50K MRR" | "pre-Series A revenue, strong growth trajectory" |
**Rule:** Percentages and directional signals (growing / declining / flat) are safe. Absolutes are not.
---
## Rule 4: Employee Names → Roles Only
Never send individual names externally.
| Raw data | Anonymized version |
|----------|-------------------|
| "Our CTO, Sarah Chen, is struggling" | "our CTO is struggling with the transition" |
| "James is the best performer on the team" | "our strongest performer is in the engineering lead role" |
| "we're about to let go of Michael" | "we're about to make a leadership change" |
| "the founding team is me, Alex, and Priya" | "a three-person founding team" |
**Exception:** Publicly known executives (CEO of a public company, named in press releases) can be referenced by name. If in doubt, use role.
---
## Rule 5: Investor Names → Generic Descriptors
| Raw data | Anonymized version |
|----------|-------------------|
| "Sequoia led our round" | "a top-tier VC led our round" |
| "our lead investor is pushing for an exit" | "pressure from investors toward exit" |
| "Y Combinator alumni" | "accelerator alumni" |
**Exception:** YC, Techstars, and similar well-known accelerators are commonly referenced and safe if the founder has publicly disclosed. When in doubt, omit.
---
## Rule 6: Location → Country or Region
| Raw data | Anonymized version |
|----------|-------------------|
| "Berlin-based startup" | "European startup" |
| "we're in San Francisco" | "US-based startup" |
| "expanding to Munich and Vienna" | "expanding in the DACH region" |
**Exception:** Location is less sensitive than financials. Use judgment — if it's on their website, it's fine.
---
## Anonymization Decision Tree
```
Before sending data externally:
1. Does it include a specific dollar amount?
→ YES: Replace with range or relative descriptor
2. Does it include a person's name?
→ YES: Replace with role only (unless publicly known)
3. Does it include a company or customer name?
→ YES: Replace with "Customer A" or segment descriptor
4. Does it include specific headcount or runway months?
→ YES: Replace with range (1–10, 10–50) or "healthy/tight/critical"
5. Does it include proprietary data, roadmap, or unreleased product info?
→ YES: Do not include. Reference only generically ("product expansion planned")
6. Is it publicly available information?
→ YES: Safe to send as-is
```
---
## Required vs Optional Anonymization
### Required (always strip before external calls)
- Revenue figures (absolute)
- Burn rate (absolute)
- Runway (specific months)
- Customer names
- Employee names
- Investor names (unless public)
- Funding amounts (unless public)
### Optional (use judgment based on sensitivity)
- Industry vertical (usually fine)
- Company stage (usually fine)
- Team size ranges (usually fine)
- Geographic region (usually fine)
- General challenge category (usually fine)
---
## What to Do If You're Unsure
Default to stricter anonymization. The cost of over-anonymizing is slightly less useful external results. The cost of under-anonymizing is a privacy breach.
When in doubt: **remove it**.
---
## Audit Log (Internal Only)
When running external calls with company context, note internally:
```
[EXTERNAL CALL: {tool/API used}]
[ANONYMIZED: {fields stripped}]
[RETAINED: {fields kept and why}]
```
This is for internal reasoning only — never included in output to the founder.
Mô hình thống kê, thiết kế thí nghiệm, suy luận nhân quả, phân tích dự báo, A/B testing và feature engineering.
---
name: "senior-data-scientist"
description: World-class senior data scientist skill specialising in statistical modeling, experiment design, causal inference, and predictive analytics. Covers A/B testing (sample sizing, two-proportion z-tests, Bonferroni correction), difference-in-differences, feature engineering pipelines (Scikit-learn, XGBoost), cross-validated model evaluation (AUC-ROC, AUC-PR, SHAP), and MLflow experiment tracking — using Python (NumPy, Pandas, Scikit-learn), R, and SQL. Use when designing or analysing controlled experiments, building and evaluating classification or regression models, performing causal analysis on observational data, engineering features for structured tabular datasets, or translating statistical findings into data-driven business decisions.
---
# Senior Data Scientist
World-class senior data scientist skill for production-grade AI/ML/Data systems.
## Core Workflows
### 1. Design an A/B Test
```python
import numpy as np
from scipy import stats
def calculate_sample_size(baseline_rate, mde, alpha=0.05, power=0.8):
"""
Calculate required sample size per variant.
baseline_rate: current conversion rate (e.g. 0.10)
mde: minimum detectable effect (relative, e.g. 0.05 = 5% lift)
"""
p1 = baseline_rate
p2 = baseline_rate * (1 + mde)
effect_size = abs(p2 - p1) / np.sqrt((p1 * (1 - p1) + p2 * (1 - p2)) / 2)
z_alpha = stats.norm.ppf(1 - alpha / 2)
z_beta = stats.norm.ppf(power)
n = ((z_alpha + z_beta) / effect_size) ** 2
return int(np.ceil(n))
def analyze_experiment(control, treatment, alpha=0.05):
"""
Run two-proportion z-test and return structured results.
control/treatment: dicts with 'conversions' and 'visitors'.
"""
p_c = control["conversions"] / control["visitors"]
p_t = treatment["conversions"] / treatment["visitors"]
pooled = (control["conversions"] + treatment["conversions"]) / (control["visitors"] + treatment["visitors"])
se = np.sqrt(pooled * (1 - pooled) * (1 / control["visitors"] + 1 / treatment["visitors"]))
z = (p_t - p_c) / se
p_value = 2 * (1 - stats.norm.cdf(abs(z)))
ci_low = (p_t - p_c) - stats.norm.ppf(1 - alpha / 2) * se
ci_high = (p_t - p_c) + stats.norm.ppf(1 - alpha / 2) * se
return {
"lift": (p_t - p_c) / p_c,
"p_value": p_value,
"significant": p_value < alpha,
"ci_95": (ci_low, ci_high),
}
# --- Experiment checklist ---
# 1. Define ONE primary metric and pre-register secondary metrics.
# 2. Calculate sample size BEFORE starting: calculate_sample_size(0.10, 0.05)
# 3. Randomise at the user (not session) level to avoid leakage.
# 4. Run for at least 1 full business cycle (typically 2 weeks).
# 5. Check for sample ratio mismatch: abs(n_control - n_treatment) / expected < 0.01
# 6. Analyze with analyze_experiment() and report lift + CI, not just p-value.
# 7. Apply Bonferroni correction if testing multiple metrics: alpha / n_metrics
```
### 2. Build a Feature Engineering Pipeline
```python
import pandas as pd
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.compose import ColumnTransformer
def build_feature_pipeline(numeric_cols, categorical_cols, date_cols=None):
"""
Returns a fitted-ready ColumnTransformer for structured tabular data.
"""
numeric_pipeline = Pipeline([
("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler()),
])
categorical_pipeline = Pipeline([
("impute", SimpleImputer(strategy="most_frequent")),
("encode", OneHotEncoder(handle_unknown="ignore", sparse_output=False)),
])
transformers = [
("num", numeric_pipeline, numeric_cols),
("cat", categorical_pipeline, categorical_cols),
]
return ColumnTransformer(transformers, remainder="drop")
def add_time_features(df, date_col):
"""Extract cyclical and lag features from a datetime column."""
df = df.copy()
df[date_col] = pd.to_datetime(df[date_col])
df["dow_sin"] = np.sin(2 * np.pi * df[date_col].dt.dayofweek / 7)
df["dow_cos"] = np.cos(2 * np.pi * df[date_col].dt.dayofweek / 7)
df["month_sin"] = np.sin(2 * np.pi * df[date_col].dt.month / 12)
df["month_cos"] = np.cos(2 * np.pi * df[date_col].dt.month / 12)
df["is_weekend"] = (df[date_col].dt.dayofweek >= 5).astype(int)
return df
# --- Feature engineering checklist ---
# 1. Never fit transformers on the full dataset — fit on train, transform test.
# 2. Log-transform right-skewed numeric features before scaling.
# 3. For high-cardinality categoricals (>50 levels), use target encoding or embeddings.
# 4. Generate lag/rolling features BEFORE the train/test split to avoid leakage.
# 5. Document each feature's business meaning alongside its code.
```
### 3. Train, Evaluate, and Select a Prediction Model
```python
from sklearn.model_selection import StratifiedKFold, cross_validate
from sklearn.metrics import make_scorer, roc_auc_score, average_precision_score
import xgboost as xgb
import mlflow
SCORERS = {
"roc_auc": make_scorer(roc_auc_score, needs_proba=True),
"avg_prec": make_scorer(average_precision_score, needs_proba=True),
}
def evaluate_model(model, X, y, cv=5):
"""
Cross-validate and return mean ± std for each scorer.
Use StratifiedKFold for classification to preserve class balance.
"""
cv_results = cross_validate(
model, X, y,
cv=StratifiedKFold(n_splits=cv, shuffle=True, random_state=42),
scoring=SCORERS,
return_train_score=True,
)
summary = {}
for metric in SCORERS:
test_scores = cv_results[f"test_{metric}"]
summary[metric] = {"mean": test_scores.mean(), "std": test_scores.std()}
# Flag overfitting: large gap between train and test score
train_mean = cv_results[f"train_{metric}"].mean()
summary[metric]["overfit_gap"] = train_mean - test_scores.mean()
return summary
def train_and_log(model, X_train, y_train, X_test, y_test, run_name):
"""Train model and log all artefacts to MLflow."""
with mlflow.start_run(run_name=run_name):
model.fit(X_train, y_train)
proba = model.predict_proba(X_test)[:, 1]
metrics = {
"roc_auc": roc_auc_score(y_test, proba),
"avg_prec": average_precision_score(y_test, proba),
}
mlflow.log_params(model.get_params())
mlflow.log_metrics(metrics)
mlflow.sklearn.log_model(model, "model")
return metrics
# --- Model evaluation checklist ---
# 1. Always report AUC-PR alongside AUC-ROC for imbalanced datasets.
# 2. Check overfit_gap > 0.05 as a warning sign of overfitting.
# 3. Calibrate probabilities (Platt scaling / isotonic) before production use.
# 4. Compute SHAP values to validate feature importance makes business sense.
# 5. Run a baseline (e.g. DummyClassifier) and verify the model beats it.
# 6. Log every run to MLflow — never rely on notebook output for comparison.
```
### 4. Causal Inference: Difference-in-Differences
```python
import statsmodels.formula.api as smf
def diff_in_diff(df, outcome, treatment_col, post_col, controls=None):
"""
Estimate ATT via OLS DiD with optional covariates.
df must have: outcome, treatment_col (0/1), post_col (0/1).
Returns the interaction coefficient (treatment × post) and its p-value.
"""
covariates = " + ".join(controls) if controls else ""
formula = (
f"{outcome} ~ {treatment_col} * {post_col}"
+ (f" + {covariates}" if covariates else "")
)
result = smf.ols(formula, data=df).fit(cov_type="HC3")
interaction = f"{treatment_col}:{post_col}"
return {
"att": result.params[interaction],
"p_value": result.pvalues[interaction],
"ci_95": result.conf_int().loc[interaction].tolist(),
"summary": result.summary(),
}
# --- Causal inference checklist ---
# 1. Validate parallel trends in pre-period before trusting DiD estimates.
# 2. Use HC3 robust standard errors to handle heteroskedasticity.
# 3. For panel data, cluster SEs at the unit level (add groups= param to fit).
# 4. Consider propensity score matching if groups differ at baseline.
# 5. Report the ATT with confidence interval, not just statistical significance.
```
## Reference Documentation
- **Statistical Methods:** `references/statistical_methods_advanced.md`
- **Experiment Design Frameworks:** `references/experiment_design_frameworks.md`
- **Feature Engineering Patterns:** `references/feature_engineering_patterns.md`
## Common Commands
```bash
# Testing & linting
python -m pytest tests/ -v --cov=src/
python -m black src/ && python -m pylint src/
# Training & evaluation
python scripts/train.py --config prod.yaml
python scripts/evaluate.py --model best.pth
# Deployment
docker build -t service:v1 .
kubectl apply -f k8s/
helm upgrade service ./charts/
# Monitoring & health
kubectl logs -f deployment/service
python scripts/health_check.py
```
FILE:references/experiment_design_frameworks.md
# Experiment Design Frameworks
## Overview
World-class experiment design frameworks for senior data scientist.
## Core Principles
### Production-First Design
Always design with production in mind:
- Scalability: Handle 10x current load
- Reliability: 99.9% uptime target
- Maintainability: Clear, documented code
- Observability: Monitor everything
### Performance by Design
Optimize from the start:
- Efficient algorithms
- Resource awareness
- Strategic caching
- Batch processing
### Security & Privacy
Build security in:
- Input validation
- Data encryption
- Access control
- Audit logging
## Advanced Patterns
### Pattern 1: Distributed Processing
Enterprise-scale data processing with fault tolerance.
### Pattern 2: Real-Time Systems
Low-latency, high-throughput systems.
### Pattern 3: ML at Scale
Production ML with monitoring and automation.
## Best Practices
### Code Quality
- Comprehensive testing
- Clear documentation
- Code reviews
- Type hints
### Performance
- Profile before optimizing
- Monitor continuously
- Cache strategically
- Batch operations
### Reliability
- Design for failure
- Implement retries
- Use circuit breakers
- Monitor health
## Tools & Technologies
Essential tools for this domain:
- Development frameworks
- Testing libraries
- Deployment platforms
- Monitoring solutions
## Further Reading
- Research papers
- Industry blogs
- Conference talks
- Open source projects
FILE:references/feature_engineering_patterns.md
# Feature Engineering Patterns
## Overview
World-class feature engineering patterns for senior data scientist.
## Core Principles
### Production-First Design
Always design with production in mind:
- Scalability: Handle 10x current load
- Reliability: 99.9% uptime target
- Maintainability: Clear, documented code
- Observability: Monitor everything
### Performance by Design
Optimize from the start:
- Efficient algorithms
- Resource awareness
- Strategic caching
- Batch processing
### Security & Privacy
Build security in:
- Input validation
- Data encryption
- Access control
- Audit logging
## Advanced Patterns
### Pattern 1: Distributed Processing
Enterprise-scale data processing with fault tolerance.
### Pattern 2: Real-Time Systems
Low-latency, high-throughput systems.
### Pattern 3: ML at Scale
Production ML with monitoring and automation.
## Best Practices
### Code Quality
- Comprehensive testing
- Clear documentation
- Code reviews
- Type hints
### Performance
- Profile before optimizing
- Monitor continuously
- Cache strategically
- Batch operations
### Reliability
- Design for failure
- Implement retries
- Use circuit breakers
- Monitor health
## Tools & Technologies
Essential tools for this domain:
- Development frameworks
- Testing libraries
- Deployment platforms
- Monitoring solutions
## Further Reading
- Research papers
- Industry blogs
- Conference talks
- Open source projects
FILE:references/statistical_methods_advanced.md
# Statistical Methods Advanced
## Overview
World-class statistical methods advanced for senior data scientist.
## Core Principles
### Production-First Design
Always design with production in mind:
- Scalability: Handle 10x current load
- Reliability: 99.9% uptime target
- Maintainability: Clear, documented code
- Observability: Monitor everything
### Performance by Design
Optimize from the start:
- Efficient algorithms
- Resource awareness
- Strategic caching
- Batch processing
### Security & Privacy
Build security in:
- Input validation
- Data encryption
- Access control
- Audit logging
## Advanced Patterns
### Pattern 1: Distributed Processing
Enterprise-scale data processing with fault tolerance.
### Pattern 2: Real-Time Systems
Low-latency, high-throughput systems.
### Pattern 3: ML at Scale
Production ML with monitoring and automation.
## Best Practices
### Code Quality
- Comprehensive testing
- Clear documentation
- Code reviews
- Type hints
### Performance
- Profile before optimizing
- Monitor continuously
- Cache strategically
- Batch operations
### Reliability
- Design for failure
- Implement retries
- Use circuit breakers
- Monitor health
## Tools & Technologies
Essential tools for this domain:
- Development frameworks
- Testing libraries
- Deployment platforms
- Monitoring solutions
## Further Reading
- Research papers
- Industry blogs
- Conference talks
- Open source projects
FILE:scripts/experiment_designer.py
#!/usr/bin/env python3
"""
Experiment Designer
Production-grade tool for senior data scientist
"""
import os
import sys
import json
import logging
import argparse
from pathlib import Path
from typing import Dict, List, Optional
from datetime import datetime
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
class ExperimentDesigner:
"""Production-grade experiment designer"""
def __init__(self, config: Dict):
self.config = config
self.results = {
'status': 'initialized',
'start_time': datetime.now().isoformat(),
'processed_items': 0
}
logger.info(f"Initialized {self.__class__.__name__}")
def validate_config(self) -> bool:
"""Validate configuration"""
logger.info("Validating configuration...")
# Add validation logic
logger.info("Configuration validated")
return True
def process(self) -> Dict:
"""Main processing logic"""
logger.info("Starting processing...")
try:
self.validate_config()
# Main processing
result = self._execute()
self.results['status'] = 'completed'
self.results['end_time'] = datetime.now().isoformat()
logger.info("Processing completed successfully")
return self.results
except Exception as e:
self.results['status'] = 'failed'
self.results['error'] = str(e)
logger.error(f"Processing failed: {e}")
raise
def _execute(self) -> Dict:
"""Execute main logic"""
# Implementation here
return {'success': True}
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Experiment Designer"
)
parser.add_argument('--input', '-i', required=True, help='Input path')
parser.add_argument('--output', '-o', required=True, help='Output path')
parser.add_argument('--config', '-c', help='Configuration file')
parser.add_argument('--verbose', '-v', action='store_true', help='Verbose output')
args = parser.parse_args()
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
try:
config = {
'input': args.input,
'output': args.output
}
processor = ExperimentDesigner(config)
results = processor.process()
print(json.dumps(results, indent=2))
sys.exit(0)
except Exception as e:
logger.error(f"Fatal error: {e}")
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/feature_engineering_pipeline.py
#!/usr/bin/env python3
"""
Feature Engineering Pipeline
Production-grade tool for senior data scientist
"""
import os
import sys
import json
import logging
import argparse
from pathlib import Path
from typing import Dict, List, Optional
from datetime import datetime
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
class FeatureEngineeringPipeline:
"""Production-grade feature engineering pipeline"""
def __init__(self, config: Dict):
self.config = config
self.results = {
'status': 'initialized',
'start_time': datetime.now().isoformat(),
'processed_items': 0
}
logger.info(f"Initialized {self.__class__.__name__}")
def validate_config(self) -> bool:
"""Validate configuration"""
logger.info("Validating configuration...")
# Add validation logic
logger.info("Configuration validated")
return True
def process(self) -> Dict:
"""Main processing logic"""
logger.info("Starting processing...")
try:
self.validate_config()
# Main processing
result = self._execute()
self.results['status'] = 'completed'
self.results['end_time'] = datetime.now().isoformat()
logger.info("Processing completed successfully")
return self.results
except Exception as e:
self.results['status'] = 'failed'
self.results['error'] = str(e)
logger.error(f"Processing failed: {e}")
raise
def _execute(self) -> Dict:
"""Execute main logic"""
# Implementation here
return {'success': True}
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Feature Engineering Pipeline"
)
parser.add_argument('--input', '-i', required=True, help='Input path')
parser.add_argument('--output', '-o', required=True, help='Output path')
parser.add_argument('--config', '-c', help='Configuration file')
parser.add_argument('--verbose', '-v', action='store_true', help='Verbose output')
args = parser.parse_args()
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
try:
config = {
'input': args.input,
'output': args.output
}
processor = FeatureEngineeringPipeline(config)
results = processor.process()
print(json.dumps(results, indent=2))
sys.exit(0)
except Exception as e:
logger.error(f"Fatal error: {e}")
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/model_evaluation_suite.py
#!/usr/bin/env python3
"""
Model Evaluation Suite
Production-grade tool for senior data scientist
"""
import os
import sys
import json
import logging
import argparse
from pathlib import Path
from typing import Dict, List, Optional
from datetime import datetime
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
class ModelEvaluationSuite:
"""Production-grade model evaluation suite"""
def __init__(self, config: Dict):
self.config = config
self.results = {
'status': 'initialized',
'start_time': datetime.now().isoformat(),
'processed_items': 0
}
logger.info(f"Initialized {self.__class__.__name__}")
def validate_config(self) -> bool:
"""Validate configuration"""
logger.info("Validating configuration...")
# Add validation logic
logger.info("Configuration validated")
return True
def process(self) -> Dict:
"""Main processing logic"""
logger.info("Starting processing...")
try:
self.validate_config()
# Main processing
result = self._execute()
self.results['status'] = 'completed'
self.results['end_time'] = datetime.now().isoformat()
logger.info("Processing completed successfully")
return self.results
except Exception as e:
self.results['status'] = 'failed'
self.results['error'] = str(e)
logger.error(f"Processing failed: {e}")
raise
def _execute(self) -> Dict:
"""Execute main logic"""
# Implementation here
return {'success': True}
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Model Evaluation Suite"
)
parser.add_argument('--input', '-i', required=True, help='Input path')
parser.add_argument('--output', '-o', required=True, help='Output path')
parser.add_argument('--config', '-c', help='Configuration file')
parser.add_argument('--verbose', '-v', action='store_true', help='Verbose output')
args = parser.parse_args()
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
try:
config = {
'input': args.input,
'output': args.output
}
processor = ModelEvaluationSuite(config)
results = processor.process()
print(json.dumps(results, indent=2))
sys.exit(0)
except Exception as e:
logger.error(f"Fatal error: {e}")
sys.exit(1)
if __name__ == '__main__':
main()