Tuân thủ EU AI Act (Quy định 2024/1689): phân loại mức rủi ro hệ thống AI và xác định nghĩa vụ theo từng điều luật.
---
name: "eu-ai-act-specialist"
description: "EU AI Act (Regulation (EU) 2024/1689) operational compliance for compliance teams. Three Article-level decisions: (1) What's the risk tier of this AI system — prohibited (Art. 5), high-risk (Art. 6 + Annex III), limited-risk (Art. 50), or minimal-risk? (2) For high-risk systems, what's the Article 43 conformity assessment route (Module A internal control vs Module H full QMS + notified body) and what goes in the Annex IV technical documentation? (3) Per organizational role (provider / deployer / importer / distributor / authorized representative), what are the active obligations and deadlines? Use during AI system intake review, when planning conformity assessment, or when scoping deployer obligations. Cites Articles + Annexes for every output. NOT executive AI strategy (see chief-ai-officer-advisor). NOT a legal substitute."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: ra-qm-team
domain: eu-ai-act-compliance
updated: 2026-05-13
python-tools: ai_system_risk_classifier.py, conformity_assessment_planner.py, ai_act_obligation_tracker.py
frameworks: eu-ai-act, gdpr-overlap, iso-42001-mapping, nist-ai-rmf-mapping
---
# EU AI Act Compliance Specialist
Article-cited operational skill for Regulation (EU) 2024/1689. **Three decisions, no executive AI strategy:**
1. **What tier is this AI system?** — prohibited (Article 5) / high-risk (Article 6 + Annex III) / limited-risk transparency (Article 50) / minimal-risk
2. **For high-risk systems, what's the conformity assessment route + documentation pack?** — Article 43 Module A vs Module H + Annex IV technical documentation
3. **Per organizational role, what are the obligations?** — provider / deployer / importer / distributor / authorized representative matrix per Article 16, 22, 25, 26
This skill is **NOT chief-ai-officer-advisor**. CAIO decides whether to ship the AI feature at all and accepts business risk. This skill operates the conformity work that turns "we'll ship it" into Article-compliant artefacts.
This skill is **NOT a legal substitute**. The Act is binding regulation. For novel cases (Is this a GPAI model? Does Article 6(2) carve-out apply? Is fine-tuning a foundation model "substantial modification"?), engage qualified outside counsel. The skill cites Articles + Annexes and uses Commission/EDPB published interpretation but does not provide binding legal opinion.
This skill is **NOT GDPR**. Many AI systems also trigger GDPR (training data, output processing). See `ra-qm-team/skills/gdpr-dsgvo-expert/` for DPIA + lawful basis work. The Acts interact (Recital 10, Article 10 for high-risk training data).
## Keywords
EU AI Act, EU AI Regulation, Regulation 2024/1689, AI Act, AI regulation Europe, high-risk AI, prohibited AI, Article 5 AI Act, Article 6 AI Act, Article 9 AI Act, Article 50 AI Act, Annex III, Annex IV, conformity assessment, CE marking AI, notified body AI, Module A, Module H, technical documentation AI, post-market monitoring AI, fundamental rights impact assessment, FRIA, GPAI, general-purpose AI model, systemic risk GPAI, AI Office, ENISA AI, EDPB AI, AI Act timeline, AI Act penalties, EU AI Act provider, EU AI Act deployer, EU AI Act importer, EU AI Act distributor, EU AI Act fines, AI literacy
## Quick Start
```bash
# Decision A: Classify an AI system per the Act
python scripts/ai_system_risk_classifier.py # embedded 5-system sample
python scripts/ai_system_risk_classifier.py path/to/systems.json
# Decision B: Conformity assessment plan for a high-risk system
python scripts/conformity_assessment_planner.py # embedded high-risk sample
python scripts/conformity_assessment_planner.py path/to/system.json
# Decision C: Obligation tracker per organizational role
python scripts/ai_act_obligation_tracker.py # embedded sample (provider + deployer)
python scripts/ai_act_obligation_tracker.py path/to/roles.json
```
## Key Questions (ask these first)
- **Does this AI system fall under Article 5 (prohibited practices)?** Social scoring, emotion recognition in workplace/education, manipulative subliminal techniques, real-time remote biometric identification in public — any of these are flat-out prohibited.
- **Does it fall under Annex III (high-risk categories)?** 8 categories: biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration, justice. Triggering Annex III triggers Article 6(2) — unless the Article 6(3) carve-outs apply.
- **What organizational role does the company play?** Provider (placed on market), deployer (uses under own authority), importer (places third-country system on EU market), distributor (makes available in supply chain). Many companies are BOTH provider AND deployer simultaneously.
- **Is this a general-purpose AI model?** GPAI has its own track (Articles 51–55) with stricter rules above 10²⁵ FLOPs training compute (Article 51 systemic risk).
- **For high-risk: have we run Article 9 risk management AND Article 27 FRIA?** Article 9 is the lifecycle risk management; Article 27 is the Fundamental Rights Impact Assessment for public-sector deployers + essential services.
- **What's the conformity assessment Module per Article 43?** Module A (internal control, possible for most Annex III systems) vs Module H (full QMS + notified body, required for biometrics + sometimes others).
## Core Responsibilities
### 1. AI System Risk Classification
**The framework:** The Act takes a risk-based approach (Recital 26). Each AI system falls into exactly one of four tiers:
| Tier | Source | Examples | Obligations |
|---|---|---|---|
| **Prohibited** | Article 5 | Social scoring; emotion recognition in workplace/education; subliminal manipulation; real-time public biometrics by law enforcement (with narrow exceptions) | Cannot be placed on market or used (penalties up to EUR 35M / 7% turnover) |
| **High-risk** | Article 6 + Annex III; Article 6(1) + Annex I | CV-screening, credit scoring, biometric categorisation, safety components of regulated products | Articles 8–17 (provider) + Article 26 (deployer); conformity assessment; CE marking |
| **Limited-risk (transparency)** | Article 50 | Chatbots, deepfakes, emotion recognition outside Article 5 contexts | Transparency disclosures to natural persons |
| **Minimal-risk** | Default | Spam filters, video-game AI, inventory forecasters | None under the Act (voluntary codes of conduct, Article 95) |
**Critical carve-outs (Article 6(3)):** an Annex III system is NOT high-risk if it (a) performs a narrow procedural task, (b) improves the result of previously completed human activity, (c) detects decision-making patterns without replacing human assessment, (d) performs a preparatory task. Caveat: profiling of natural persons is always Annex III high-risk regardless of carve-outs.
**Run** `ai_system_risk_classifier.py` with system characteristics. The tool checks Article 5 prohibitions first, then Annex III categories, then Article 6(3) carve-outs, then Article 50 transparency, then minimal-risk default.
See `references/eu_ai_act_titles.md` for the full Article-by-Article walkthrough.
### 2. Conformity Assessment + Annex IV Technical Documentation
**The framework (Article 43 + Annex VI/VII):** for high-risk AI systems, the provider must demonstrate conformity before placing on market. Two routes:
- **Module A — Internal control** (Annex VI): provider self-assesses against the requirements. Applies to most Annex III systems where the provider has implemented harmonised standards.
- **Module H — Full quality management system + technical documentation** (Annex VII): notified body involvement. Required for biometrics systems (Article 43(1)).
**Required artifacts per Annex IV — Technical Documentation:**
1. General description of the AI system (intended purpose, identification, version)
2. Detailed description of system elements (architecture, training data, validation procedures)
3. Information about monitoring, functioning and control
4. Description of risk management system (Article 9)
5. Description of changes after placing on market
6. List of harmonised standards applied (or alternative)
7. EU declaration of conformity (Article 47)
8. Description of the post-market monitoring system (Article 72)
**Run** `conformity_assessment_planner.py` to select the Module and produce the Annex IV checklist for a given high-risk system.
See `references/high_risk_systems_annex_iii.md` for which systems require which conformity route.
### 3. Per-Role Obligation Tracker
**The framework (Articles 16, 22, 23, 24, 25, 26):** the Act distinguishes provider obligations (most) from downstream-actor obligations (deployer, importer, distributor, authorized representative). A single company can play multiple roles simultaneously.
| Role | Primary Articles | Key obligations |
|---|---|---|
| **Provider** (Article 3(3)) | 8–17, 47, 49, 72 | Conformity assessment; CE marking; risk management; data governance; technical documentation; post-market monitoring; serious incident reporting (Article 73) |
| **Deployer** (Article 3(4)) | 26 | Use according to instructions; human oversight; input data quality; record-keeping (Article 19); inform workers (Article 26(7)); FRIA if public-sector/essential-services (Article 27) |
| **Importer** (Article 3(6)) | 23 | Verify conformity; affixed CE marking; technical documentation availability |
| **Distributor** (Article 3(7)) | 24 | Verify CE marking + documentation before making available |
| **Authorized representative** (Article 22) | 22 | Non-EU providers must appoint one; representative liable for provider obligations |
**Important:** under Article 25, a deployer who substantially modifies a high-risk AI system, or places it on the market under their own name, becomes a **provider** and inherits provider obligations.
**Run** `ai_act_obligation_tracker.py` with the roles JSON to produce a deadline-sorted obligation matrix.
See `references/gpai_obligations.md` for the separate GPAI Articles 51–55 track.
## Workflows
### Workflow 1: AI System Intake Review (per system, ~2 hours)
**Goal:** classify, identify obligations, scope the conformity work.
```bash
# 1. Document system characteristics: purpose, users, data, autonomy, deployment context
# 2. Run classifier
python scripts/ai_system_risk_classifier.py systems.json
# 3. If high-risk: run planner
python scripts/conformity_assessment_planner.py system.json
# 4. Identify org roles played (provider / deployer / both)
python scripts/ai_act_obligation_tracker.py roles.json
# 5. Cross-check with GDPR DPIA (gdpr-dsgvo-expert) if personal data
# 6. Cross-check with ISO 42001 AIMS evidence (compliance-team-iso42001)
# 7. Output: classification memo + conformity plan + obligation list
```
### Workflow 2: Annex IV Technical Documentation Build (per high-risk system, 2–4 weeks)
**Goal:** assemble the Annex IV pack before conformity assessment.
```bash
# 1. Run conformity assessment planner to get the checklist
python scripts/conformity_assessment_planner.py system.json
# 2. Assemble: system description, architecture, training data, validation, risk management
# 3. Reference ISO 42001 evidence where it satisfies Annex IV items
# 4. Reference ISO 27001 evidence for security controls
# 5. Run Article 9 risk management lifecycle
# 6. Sign EU declaration of conformity (Article 47) AFTER assessment passes
# 7. Affix CE marking (Article 48)
# 8. Register in EU database (Article 71) — high-risk Annex III systems
```
### Workflow 3: Pre-Deployment Obligation Audit (per system, before launch)
**Goal:** confirm all active obligations are in place before EU placement.
```bash
# 1. Confirm classification still correct (re-run classifier if system changed)
# 2. Confirm conformity assessment completed (if high-risk)
# 3. Confirm transparency requirements (Article 50) — for chatbots, deepfakes, emotion detection
# 4. Confirm post-market monitoring system (Article 72) is live
# 5. Confirm serious-incident reporting procedure (Article 73) is documented
# 6. For deployers: FRIA done (Article 27, if applicable); workers informed (Article 26(7))
# 7. For GPAI: Articles 51-55 obligations met if applicable
```
### Workflow 4: Annual Compliance Refresh (per organization, yearly)
**Goal:** re-verify classifications + obligations as the Act phases in.
1. List all AI systems on or planned for EU market
2. Run classifier for each — Article 5 prohibited list may expand via delegated acts
3. Run obligation tracker — deadlines shift as Title III phases in (2025 → 2026 → 2027)
4. For each high-risk system: verify post-market monitoring data flow + serious incident reporting capacity
5. Update Annex IV technical documentation per Article 11 ongoing requirement
6. Pair with ISO 42001 management review (Clause 9.3) if both operate
## Output Standards
```
**Bottom Line:** [one sentence — classification + most-significant obligation]
**Article Citation:** [Article + paragraph number; do not paraphrase without cite]
**The Decision:** [one of: classify | conformity-route | obligation-scope]
**The Evidence:** [Article + Annex references; classification confidence]
**How to Act:** [3 concrete next steps with owner + deadline aligned to phasing]
**Your Decision:** [the call for compliance officer or legal counsel — risk-class disputes, novel cases, GPAI threshold determinations]
```
## Adjacent Skills
- `../../skills/gdpr-dsgvo-expert/` — GDPR DPIA + lawful basis (most AI systems also trigger GDPR)
- `../../../compliance-team-iso42001/` — ISO 42001 AIMS (voluntary management system that satisfies parts of Article 17 QMS for providers)
- `../../skills/information-security-manager-iso27001/` — ISO 27001 for cybersecurity requirements (Article 15)
- `../../skills/risk-management-specialist/` — ISO 14971 risk management (referenced for safety-component AI under Article 6(1))
- `../../skills/mdr-745-specialist/` — MDR 2017/745 (medical-device AI overlap)
- `../../../../compliance-os/` — Meta-orchestrator for multi-framework programs
- `../../../../c-level-advisor/chief-ai-officer-advisor/` — Executive AI strategy
## References
- [eu_ai_act_titles.md](references/eu_ai_act_titles.md) — Titles I–XII Article-by-Article walkthrough with deployer/provider/importer/distributor obligation breakdown
- [high_risk_systems_annex_iii.md](references/high_risk_systems_annex_iii.md) — Annex III 8 categories detailed + Article 6(2)–(3) interaction + carve-out test
- [gpai_obligations.md](references/gpai_obligations.md) — Articles 51–55 GPAI track + systemic-risk threshold + transparency rules + Code of Practice status
- [cross_framework_mapping_ai_act.md](references/cross_framework_mapping_ai_act.md) — AI Act ↔ ISO 42001 ↔ NIST AI RMF ↔ GDPR control-level mapping
---
**Version:** 1.0.0
**Status:** Production Ready
FILE:references/cross_framework_mapping_ai_act.md
# EU AI Act ↔ ISO 42001 ↔ NIST AI RMF ↔ GDPR — Cross-Framework Mapping
This reference answers exactly one decision: **for each EU AI Act obligation, what existing framework evidence can I reuse?**
The point: minimize duplicate work. EU AI Act compliance for high-risk systems requires significant artefacts (Annex IV technical documentation, Article 9 risk management, Article 17 QMS, Article 72 post-market monitoring). Most of these can be satisfied — partly or fully — by evidence from existing ISO 42001 / ISO 27001 / GDPR programs.
## Framework Reuse Cheat Sheet
| EU AI Act requirement | Best reuse source | Reuse confidence |
|---|---|---|
| Article 9 Risk management system | ISO 42001 Clause 6.1 + ISO 23894 process | HIGH |
| Article 10 Data governance | ISO 42001 Annex A.7 + GDPR Art. 5 + Records of Processing (Art. 30) | HIGH |
| Article 11 Technical documentation (Annex IV) | ISO 42001 documented information (Clause 7.5) + Annex A.6.2.7 model cards | HIGH |
| Article 12 Logging | ISO 27001 A.8.15 + ISO 42001 A.9.4 | HIGH |
| Article 13 Instructions for use | ISO 42001 A.8.3 user information | HIGH |
| Article 14 Human oversight | ISO 42001 A.9 use of AI systems | MEDIUM (AI Act more prescriptive) |
| Article 15 Accuracy, robustness, cybersecurity | ISO 27001 (cybersecurity) + ISO 42001 A.6.2.4 V&V + NIST AI RMF MEASURE 2 | HIGH |
| Article 16 Provider obligations | ISO 42001 Clauses 5–6 leadership + responsibilities | MEDIUM |
| Article 17 Quality management system | ISO 42001 entire AIMS satisfies this in large part | HIGH (subject to Article 17(1) item-by-item check) |
| Article 26 Deployer obligations | ISO 42001 Annex A.9 + own operating discipline | MEDIUM |
| Article 27 FRIA (public sector) | ISO 42001 A.5 impact assessment + GDPR DPIA — both inputs | MEDIUM |
| Article 50 Transparency | New artifacts (Article 50 specific) — limited reuse | LOW |
| Article 72 Post-market monitoring | ISO 42001 A.9.3 monitoring + ISO 13485 PMS pattern | HIGH |
| Article 73 Serious-incident reporting | ISO 27001 A.6.8 information security event reporting + GDPR Art. 33 breach notification — extend | MEDIUM |
## Article-by-Article Detailed Mapping
### Article 9 — Risk Management System
**EU AI Act requirement:** establish, implement, document, maintain a risk management system across the AI lifecycle.
**Best reuse:**
- ISO/IEC 42001 Clause 6.1 + Annex A.5: provides the management-system framing
- ISO/IEC 23894:2023: provides the AI-specific risk methodology
- NIST AI RMF "MAP" + "MANAGE" functions: provides operational guidance
**Gap to fill:**
- Article 9(2)(c) requires "iterative" application across full lifecycle — operational discipline, not just artifact
- Article 9(5) requires testing of high-risk systems in real-world conditions or in test environments
### Article 10 — Data Governance
**EU AI Act requirement:** training, validation, test datasets meet quality criteria including:
- Article 10(3): "relevant, sufficiently representative, free of errors, complete"
- Article 10(2)(d): documentation of data origin and provenance
- Article 10(5): processing of special categories permissible if strictly necessary for bias detection
**Best reuse:**
- ISO 42001 Annex A.7.2 data management + A.7.3 data quality + A.7.4 data provenance + A.7.5 data preparation: direct overlap
- GDPR Article 5 (data minimisation), Article 6 (lawful basis), Article 30 (records of processing): for personal data
- ISO 8000 + DAMA-DMBOK 2: data-quality framework
**Gap to fill:**
- Article 10(5) bias-detection-specific processing of special categories — explicit DPIA + ISO 23894 risk treatment combination
### Article 11 — Technical Documentation (Annex IV)
**EU AI Act requirement:** maintain technical documentation per Annex IV (8 items).
**Best reuse per Annex IV item:**
| Annex IV item | Reuse source |
|---|---|
| 1. General description | ISO 42001 SKILL scope statement; ISO 27001 system documentation |
| 2. System elements (architecture, training data, validation, human oversight) | ISO 42001 Annex A.6 + A.7 + model card pattern (Mitchell 2019) |
| 3. Monitoring, functioning, control | ISO 42001 Annex A.9 + ISO 27001 A.8.15 logging |
| 4. Risk management | ISO 42001 Clause 6.1 + Annex A.5 |
| 5. Changes after market | ISO 27001 A.8.32 change management + ISO 42001 A.6.2.5 |
| 6. Harmonised standards applied | Standards register |
| 7. EU declaration of conformity | New artifact (signed at end) |
| 8. Post-market monitoring | ISO 42001 A.9.3 + ISO 13485 PMS pattern |
### Article 14 — Human Oversight
**EU AI Act requirement:** design + enable effective human oversight by natural persons to prevent/minimise risks. Including:
- Article 14(4)(a-e): oversight personnel must understand capabilities/limitations, remain aware of automation bias, correctly interpret output, decide not to use the output, intervene/halt operation
**Best reuse:**
- ISO 42001 Annex A.9.2 intended use + A.9 use of AI systems: partial coverage
- ISO 42001 Clause 7.2 competence (define competence for oversight personnel)
- Workplace operating discipline (procedure for halting + escalating)
**Gap to fill:**
- Article 14 is more prescriptive than ISO 42001 — requires explicit design for the 5 oversight capabilities. Build the design artefact net-new.
### Article 17 — Quality Management System
**EU AI Act requirement:** providers shall put in place QMS ensuring compliance. Article 17(1)(a)–(m) lists 13 items the QMS must include.
**Best reuse:**
- ISO 42001 AIMS: satisfies most Article 17(1) items
- ISO 9001 / ISO 13485 (if already operated): satisfies the "general QMS" framing
- Map each Article 17(1) item against ISO 42001 evidence to identify any remaining gap
**Article 17(1) item-by-item mapping to ISO 42001:**
| Article 17(1) item | ISO 42001 reference |
|---|---|
| (a) Compliance strategy | Clause 5.2 AI policy |
| (b) Techniques for design/development/QA | Annex A.6 lifecycle |
| (c) Examination, testing, validation procedures | Annex A.6.2.4 V&V |
| (d) Technical specs + standards applied | Clause 7.5 documented information |
| (e) Data management procedures | Annex A.7 |
| (f) Risk management system | Clause 6.1 + Annex A.5 |
| (g) Post-market monitoring | Annex A.9.3 |
| (h) Reporting of serious incidents | Annex A.8.4 |
| (i) Communication w/ authorities, notified bodies, suppliers | Annex A.10 + Clause 7.4 |
| (j) Internal record-keeping system | Clause 7.5 + Annex A.9.4 logging |
| (k) Resource management including supply security | Annex A.4 |
| (l) Accountability framework | Annex A.3 |
| (m) Internal audit + management review | Clause 9.2 + 9.3 |
This is the closest framework alignment in the entire mapping — ISO 42001 is essentially the AI-specific operating model for Article 17.
### Article 26 — Deployer Obligations
**EU AI Act requirement:** use AI per provider's instructions, assign human oversight, ensure input data quality, monitor + cease use if Article 79 risk, retain logs ≥ 6 months, inform workers.
**Best reuse:**
- ISO 42001 Annex A.9 use of AI systems: partial
- Existing operational procedures (HR notification for workforce-impacting AI)
**Gap to fill:** Article 26 is operationally specific; build deployer-procedure net-new with reuse cross-references.
### Article 50 — Transparency
**EU AI Act requirement:** disclose AI interaction; mark synthetic content; disclose emotion/biometric categorisation; disclose deepfakes.
**Best reuse:** none direct. New UX/disclosure artefacts required.
**Cross-reference:** ISO 42001 Annex A.8 information for interested parties (overlap on framing only).
### Article 72 — Post-Market Monitoring
**EU AI Act requirement:** establish + document post-market monitoring system collecting, documenting, analysing data on performance throughout lifetime.
**Best reuse:**
- ISO 42001 Annex A.9.3 monitoring: direct overlap
- ISO 13485 post-market surveillance pattern (for medical-device AI providers): proven operational template
- NIST AI RMF MEASURE 4 + MANAGE 4: methodology
### Article 73 — Serious-Incident Reporting
**EU AI Act requirement:** report serious incidents (Article 3(49)) to market surveillance authority — 15 days general, 2 days for critical infrastructure.
**Best reuse:**
- ISO 27001 A.6.8 information security event reporting: process framework
- GDPR Article 33 personal data breach notification: 72-hour pattern
- ISO 13485 vigilance reporting (medical devices)
**Gap to fill:** Article 73 has its own serious-incident definition + report content; align reporting template with the regulation specifically.
## NIST AI RMF ↔ EU AI Act Cross-Walk
NIST AI RMF is voluntary US guidance but maps cleanly to EU AI Act provisions:
| NIST AI RMF function | EU AI Act articles satisfied (partial) |
|---|---|
| GOVERN | Articles 16, 17, 26 (broad governance) |
| MAP | Articles 9 (risk identification), 10 (data) |
| MEASURE | Articles 15 (accuracy/robustness/cybersecurity), 9 (risk evaluation) |
| MANAGE | Articles 9 (risk treatment), 26 (deployer monitoring) |
A mature NIST AI RMF program covers ~70% of EU AI Act high-risk system obligations operationally.
## GDPR ↔ EU AI Act Interaction
The two regulations interact heavily. Recital 10 + Article 10 of the AI Act + EDPB Opinion 28/2024 (Dec 2024) establish:
1. **AI Act does not modify GDPR.** GDPR continues to apply in full to personal data processing in AI systems.
2. **Article 10(5) AI Act** permits processing of special categories of personal data strictly necessary for bias detection — but only with safeguards (e.g., effective anonymisation after use).
3. **DPIA + FRIA overlap (Article 27 AI Act).** Both can be integrated into a single impact-assessment artefact for public-sector deployers of high-risk AI systems.
4. **Right to explanation (Article 86 AI Act + Article 22 GDPR).** Article 86 strengthens individual rights for high-risk AI decisions.
## When This Reference Doesn't Help
- **ISO 42001 deep-dive.** See `compliance-team-iso42001/`.
- **Single-framework audit simulation.** See `compliance-os/scripts/audit_simulator.py`.
- **Specific NIST AI RMF Playbook entries.** Refer to NIST AI 100-1 directly.
---
**Source authorities (non-exhaustive):**
- **Regulation (EU) 2024/1689** — the AI Act
- **ISO/IEC 42001:2023** — AI Management System
- **ISO/IEC 23894:2023** — AI risk management process
- **ISO/IEC 27001:2022** — Information security management
- **NIST AI Risk Management Framework 1.0** (Jan 2023) + Generative AI Profile (NIST AI 600-1, July 2024)
- **General Data Protection Regulation (EU) 2016/679** — GDPR
- **EDPB Opinion 28/2024** — AI models and personal data (December 2024)
- **EDPS** — interpretive opinions on AI Act ↔ GDPR interaction
- **European Commission** — Article 17 implementing guidance (continuously updated)
- **BSI** — Cross-walking ISO 42001 and EU AI Act (white paper 2024)
- **IAPP** — EU AI Act Tracker + AI Governance Center materials
FILE:references/eu_ai_act_titles.md
# EU AI Act (Regulation (EU) 2024/1689) — Titles I–XII Walkthrough
This reference answers exactly one decision: **what does each Title of the Act actually require, and which Articles do I cite in compliance artifacts?**
Pair with `scripts/ai_system_risk_classifier.py` to map a system to obligations.
## Structure of the Regulation
The Act has 13 Titles + 13 Annexes. Adopted as Regulation (EU) 2024/1689 (the "AI Act"); published in OJEU L on 12 July 2024; entered into force 1 August 2024 (Article 113).
## Title I — General Provisions (Articles 1–4)
| Article | Topic | Key requirement |
|---|---|---|
| **1** | Subject matter | Establishes harmonised rules for AI systems placed on EU market, used or put into service |
| **2** | Scope | Applies to providers, deployers, importers, distributors, authorized representatives. Extraterritorial: applies to non-EU providers placing systems on EU market. Excludes military / national security / pure scientific research |
| **3** | Definitions | "AI system" (Article 3(1)): a machine-based system designed to operate with varying levels of autonomy that may exhibit adaptiveness after deployment; infers from input how to generate outputs (predictions, content, recommendations, decisions). Per Commission Feb 2025 Guidelines, excludes simple rule-based systems with no adaptiveness |
| **4** | AI literacy | **In force from 2 Feb 2025.** Organizations must ensure staff dealing with AI systems have AI literacy proportionate to their roles |
## Title II — Prohibited AI Practices (Article 5)
**In force from 2 Feb 2025.** Penalty: up to EUR 35M or 7% worldwide annual turnover (Article 99).
8 prohibited categories per Article 5(1):
- **(a)** Subliminal techniques beyond awareness causing harm
- **(b)** Exploitation of vulnerabilities (age, disability, socioeconomic situation)
- **(c)** Social scoring by public authorities causing detrimental treatment
- **(d)** Predictive policing based solely on profiling natural persons (with narrow law-enforcement exceptions per Article 5(2))
- **(e)** Untargeted scraping of facial images for facial recognition databases
- **(f)** Emotion recognition in workplace and educational institutions
- **(g)** Biometric categorisation by sensitive attributes (race, religion, political opinions, sexual orientation, etc.)
- **(h)** Real-time remote biometric identification in publicly accessible spaces for law-enforcement purposes (with narrow Article 5(2)(d)–(h) exceptions)
## Title III — High-Risk AI Systems (Articles 6–49)
The densest part of the regulation. **Title III general high-risk obligations in force 2 Aug 2026; Annex I sectoral 2 Aug 2027.**
### Chapter 1 — Classification (Articles 6–7)
- **Article 6(1)** + Annex I: AI systems that are safety components of products covered by sectoral law (machinery, toys, medical devices, etc.) are high-risk
- **Article 6(2)** + Annex III: AI systems in 8 categories (biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration, justice) are high-risk
- **Article 6(3)**: carve-out — a system in Annex III is NOT high-risk if it performs a narrow procedural task, improves a previously completed human activity, detects decision-making patterns without replacing human assessment, or performs a preparatory task. **Profiling overrides the carve-out** (Article 6(3) last sentence)
See `high_risk_systems_annex_iii.md` for the detailed Annex III walkthrough.
### Chapter 2 — Requirements for High-Risk Systems (Articles 8–17)
| Article | Requirement |
|---|---|
| **8** | Compliance with all Section 2 requirements |
| **9** | Risk management system across full lifecycle |
| **10** | Data governance: training/validation/test datasets quality + bias examination |
| **11** | Technical documentation per Annex IV |
| **12** | Automatic event logging |
| **13** | Transparency + instructions for use to deployers |
| **14** | Human oversight design |
| **15** | Accuracy, robustness, cybersecurity |
| **16** | General provider obligations + named contact person |
| **17** | Quality management system (provider) |
### Chapter 3 — Obligations of Actors (Articles 22–27)
| Article | Topic | Applies to |
|---|---|---|
| **22** | Authorized representative | Non-EU providers must appoint one |
| **23** | Importer obligations | Verify provider conformity assessment before import |
| **24** | Distributor obligations | Verify CE marking before making available |
| **25** | Responsibilities along the value chain | Substantial modification turns deployer into provider |
| **26** | Deployer obligations | Use per instructions; human oversight; input data; logs; transparency |
| **27** | Fundamental Rights Impact Assessment (FRIA) | Public-sector deployers + essential-services deployers of high-risk |
### Chapter 4 — Notified Bodies (Articles 28–39)
Procedures for designating + monitoring notified bodies (involved in Module H conformity assessment per Annex VII).
### Chapter 5 — Standards, Conformity Assessment, Certificates, Registration (Articles 40–49)
| Article | Topic |
|---|---|
| **40** | Harmonised standards — presumption of conformity |
| **41** | Common specifications (where standards lacking) |
| **43** | Conformity assessment procedure (Module A internal control vs Module H notified body) |
| **47** | EU declaration of conformity (provider signs; 10-year retention) |
| **48** | CE marking |
| **49** | Registration in EU database (Article 71) for Annex III systems |
## Title IV — Transparency Obligations (Article 50)
**In force from 2 Aug 2025.**
| Article 50 paragraph | Requirement |
|---|---|
| **50(1)** | Disclose AI interaction (chatbots): natural persons must be informed |
| **50(2)** | Mark synthetic content (machine-readable) as AI-generated |
| **50(3)** | Disclose emotion recognition / biometric categorisation to subjects (outside Article 5 prohibition) |
| **50(4)** | Disclose deepfakes (image/audio/video) — exception for art, satire, security |
## Title V — General-Purpose AI Models (Articles 51–55)
**In force from 2 Aug 2025.** See `gpai_obligations.md` for the detailed walkthrough.
| Article | Topic |
|---|---|
| **51** | Classification of GPAI with systemic risk (training compute ≥ 10²⁵ FLOPs) |
| **52** | Procedure for adding/removing systemic-risk designation |
| **53** | Obligations for ALL GPAI providers (technical docs, transparency to downstream, copyright policy, training data summary) |
| **54** | Authorized representative for non-EU GPAI providers |
| **55** | Additional obligations for systemic-risk GPAI (model evaluations, adversarial testing, incident reporting, cybersecurity) |
## Title VI — Measures in Support of Innovation (Articles 57–63)
| Article | Topic |
|---|---|
| **57** | AI regulatory sandboxes by Member States |
| **58** | Modalities for sandboxes |
| **59** | Further processing of personal data for AI development in sandboxes |
| **60** | Real-world testing of high-risk systems outside sandboxes |
| **62** | SME / start-up specific measures |
## Title VII — Governance (Articles 64–70)
| Article | Body |
|---|---|
| **64** | European Artificial Intelligence Office (the "AI Office") |
| **65** | European AI Board |
| **66** | Member State national competent authorities |
| **67** | Advisory Forum (industry + civil society) |
| **68** | Scientific Panel of independent experts |
## Title VIII — EU Database (Article 71)
EU-wide database of stand-alone high-risk Annex III AI systems. Provider registration before placing on market.
## Title IX — Post-Market Monitoring, Information Sharing, Market Surveillance (Articles 72–84)
| Article | Topic |
|---|---|
| **72** | Provider post-market monitoring system |
| **73** | Serious-incident reporting (provider) — 15 days general; 2 days for critical infrastructure |
| **74** | Market surveillance + AI Office cooperation |
| **75–84** | Market surveillance powers, enforcement, mutual assistance |
## Title X — Codes of Conduct and Guidelines (Articles 95–96)
Voluntary codes of conduct extending Title III principles to non-high-risk systems. Commission may issue guidelines.
## Title XI — Delegated and Implementing Acts (Articles 97–98)
Commission powers to update Annexes (notably Annex III categories) via delegated acts.
## Title XII — Final Provisions (Articles 99–113)
| Article | Topic |
|---|---|
| **99** | Penalties: up to EUR 35M / 7% turnover (Article 5); EUR 15M / 3% (most high-risk); EUR 7.5M / 1% (incorrect info) |
| **102** | Amendments to other regulations (medical devices, etc.) |
| **113** | Entry into force + application phasing |
## Annexes — At a Glance
| Annex | Topic |
|---|---|
| **I** | List of EU sectoral product legislation (machinery, toys, MDR, IVDR, etc.) — Article 6(1) trigger |
| **II** | List of Union harmonisation legislation |
| **III** | High-risk AI systems referred to in Article 6(2) — 8 categories |
| **IV** | Technical documentation referred to in Article 11 (8 items) |
| **V** | EU declaration of conformity (Article 47) |
| **VI** | Conformity assessment Module A — Internal Control |
| **VII** | Conformity assessment Module H — Full Quality Assurance |
| **VIII** | Information to be submitted upon registration in EU database (Article 71) |
| **IX** | Information for testing in real-world conditions (Article 60) |
| **X** | Union legislative acts on large-scale IT systems |
| **XI** | Technical documentation for GPAI providers (Article 53) |
| **XII** | Transparency information for downstream providers (Article 53(1)(b)) |
| **XIII** | Designation of GPAI with systemic risk (Article 51 criteria) |
## When This Reference Doesn't Help
- **Specific Annex III high-risk system classification.** See `high_risk_systems_annex_iii.md`.
- **GPAI obligations detail.** See `gpai_obligations.md`.
- **Cross-walking to ISO 42001 / NIST AI RMF.** See `cross_framework_mapping_ai_act.md`.
---
**Source authorities (non-exhaustive):**
- **Regulation (EU) 2024/1689** — the AI Act (the binding regulation; published in OJEU L on 12 July 2024)
- **European Commission** — Guidelines on the definition of an AI system (Feb 2025)
- **European Commission** — Guidelines on prohibited AI practices (Feb 2025)
- **European Commission Q&A** — AI Act explanatory materials (continuously updated)
- **European Data Protection Board (EDPB)** — Opinion 28/2024 (Dec 2024) on personal-data processing in AI models
- **European Data Protection Supervisor (EDPS)** — AI Act commentary + GDPR-AI Act interaction
- **ENISA** — Multilayer Framework for Good Cybersecurity Practices for AI (Mar 2023)
- **IAPP** — EU AI Act Tracker (continuously updated practitioner reference)
- **CEN-CENELEC JTC 21** — harmonised standards work programme (Article 40 reference)
FILE:references/gpai_obligations.md
# GPAI Obligations — Articles 51–55 + Annex XI–XIII
This reference answers exactly one decision: **is a foundation model a GPAI, does it have systemic risk, and what obligations apply?**
## What is GPAI?
Per **Article 3(63)**, a "general-purpose AI model" is:
> an AI model, including where such an AI model is trained with a large amount of data using self-supervision at scale, that displays significant generality and is capable of competently performing a wide range of distinct tasks regardless of the way the model is placed on the market and that can be integrated into a variety of downstream systems or applications.
In practice: foundation models such as large language models, multimodal models, diffusion models for image/video generation. The distinguishing characteristic is generality + integration into downstream systems.
GPAI is governed by **Title V** (Articles 51–55), separate from the high-risk AI system regime in Title III. A given application can simultaneously be a GPAI provider AND a high-risk system provider (e.g., a downstream provider fine-tuning a foundation model for credit scoring).
## Systemic-Risk GPAI Designation (Article 51)
A GPAI model is presumed to have systemic risk if **either**:
- **Article 51(1)(a):** trained with compute > 10²⁵ floating-point operations (FLOPs), OR
- **Article 51(1)(b):** designated by Commission decision based on Annex XIII criteria
**Article 51(3)** provides a list of Annex XIII criteria for designation: model capabilities, parameter count, dataset size + quality, autonomy, modalities, scalability, reach to internal market, registered business users.
A provider may contest a presumption (Article 52) by submitting evidence to Commission. Commission may also designate a model with systemic risk even if below the FLOPs threshold.
## Article 53 — Obligations for ALL GPAI Providers
In force from 2 Aug 2025.
| Article | Obligation |
|---|---|
| **53(1)(a)** | Draw up and keep up-to-date technical documentation of the model (per Annex XI) — model architecture, training process, training compute, energy consumption, evaluation results, limitations |
| **53(1)(b)** | Make information available to downstream providers integrating the model (per Annex XII) — intended uses, technical means for integration, computational + hardware requirements |
| **53(1)(c)** | Put in place policy to comply with EU copyright law (training data + outputs) |
| **53(1)(d)** | Draw up and publicly publish a sufficiently detailed summary about content used for training |
**Annex XI items (technical documentation for GPAI):**
1. General description of GPAI model (intended tasks, architecture, integration paradigm)
2. Detailed description (training process, design choices, training data sources, energy consumption)
3. Training process (compute, data, methodology)
4. Information for downstream providers
**Annex XII items (transparency to downstream providers):**
1. General description (capabilities, modalities, intended uses)
2. Acceptable use policy
3. Technical means + computational requirements
4. Evaluation results + limitations
## Article 54 — Authorized Representative for Non-EU GPAI Providers
GPAI providers established outside the EU must appoint, by written mandate, an authorized representative established in the EU. The representative:
- Holds the technical documentation (Annex XI)
- Holds the information for downstream providers (Annex XII)
- Cooperates with AI Office and national authorities
- May terminate the mandate if provider refuses to cooperate with Article 53 obligations
This parallels the Article 22 representative obligation for non-EU providers of high-risk AI systems.
## Article 55 — Additional Obligations for Systemic-Risk GPAI
Applies only to GPAI designated under Article 51.
| Obligation | Detail |
|---|---|
| **Model evaluations** | Including adversarial testing — identify + mitigate systemic risks |
| **Systemic risk assessment** | Track risk along entire lifecycle |
| **Serious incident reporting** | Document + report serious incidents and possible corrective measures to AI Office without undue delay |
| **Cybersecurity** | Ensure adequate level of cybersecurity protection for the model + the physical infrastructure |
Penalties for systemic-risk GPAI non-compliance: up to EUR 15M or 3% of worldwide annual turnover per Article 101.
## Code of Practice (Article 56) — Bridging Instrument
The AI Office facilitates a **Code of Practice** for GPAI providers covering Article 53 and 55 obligations. The Code is voluntary but provides a presumption of compliance. The first Code is expected to be finalised by 2 Aug 2025 (with iteration thereafter).
**Practical implication:** until harmonised standards are published under Article 40 for GPAI (not yet available as of mid-2026), the Code of Practice is the primary "what does compliance look like" reference.
## Provider-of-System vs Provider-of-Model Boundaries
A common ambiguity: when does a downstream provider become a GPAI provider in their own right?
Per **Article 25(3)**: a downstream provider that **substantially modifies** a GPAI model (e.g., extensive fine-tuning that changes the model's intended purpose) becomes a GPAI provider with its own Article 53 obligations.
Per **Article 25(1)**: if a downstream provider integrates a GPAI model into a high-risk AI system, the downstream provider remains the high-risk AI system's provider with Title III obligations; the GPAI model's provider retains its Article 53 + (if applicable) Article 55 obligations.
The Commission Q&A and emerging Code of Practice provide more detail on "substantial modification" boundary.
## Practical Decision Tree
```
Is the model a GPAI per Article 3(63)?
├─ No → Not GPAI. Apply standard high-risk rules if applicable.
└─ Yes → Article 53 obligations apply.
└─ Training compute > 10^25 FLOPs OR Commission-designated?
├─ No → Article 53 only.
└─ Yes → Article 53 + Article 55 (systemic-risk additional obligations).
```
## When This Reference Doesn't Help
- **Article 5 prohibitions applied to GPAI use cases.** See `eu_ai_act_titles.md` Title II.
- **Article 40 harmonised standards for GPAI.** Not published as of mid-2026; CEN-CENELEC JTC 21 work in progress.
- **Open-source GPAI carve-out (Article 53(2)).** GPAI models released under free + open-source license can be exempt from some Article 53 obligations IF they do not have systemic risk. Article 53(2) specifies the exact exemption scope.
---
**Source authorities (non-exhaustive):**
- **Regulation (EU) 2024/1689** — Articles 3(63), 51–55, Annex XI–XIII (binding)
- **European AI Office** — GPAI Code of Practice (published in drafts during 2024–2025)
- **European Commission** — GPAI guidance Q&A
- **NIST** — Generative AI Profile (NIST AI 600-1, July 2024) — voluntary US guidance with conceptual overlap to Article 55 model-evaluation requirements
- **Stanford CRFM** — Foundation Model Transparency Index (2023–) — practitioner benchmark of GPAI disclosure practices
- **MIT** — AI Risk Repository (continuously updated)
- **IAPP** — GPAI Tracker section of EU AI Act Tracker
- **Open Future / Knowledge Rights 21** — Code of Practice + copyright analysis (civil society input)
- **Mozilla / Hugging Face / GitHub** — open-source GPAI submissions to the Commission consultation on Article 53(2) exemption
FILE:references/high_risk_systems_annex_iii.md
# Annex III High-Risk AI Categories + Article 6(2)–(3) Decision Tree
This reference answers exactly one decision: **for a given AI system, is it Annex III high-risk, and does any Article 6(3) carve-out apply?**
Pair with `scripts/ai_system_risk_classifier.py` for the decision-tree implementation.
## The Article 6 Decision Order
```
1. Article 5 — prohibited? → YES: STOP. Prohibited. Cannot place on market.
2. Article 6(1) + Annex I product? → YES: high-risk per sectoral law (e.g., MDR 745 medical device with AI safety component)
3. Article 6(2) + Annex III? → YES: enter Article 6(3) carve-out check
4. Article 6(3) carve-out applies? → YES (and no profiling): NOT high-risk
→ NO (or profiling present): high-risk
5. Article 50 transparency trigger? → YES: limited-risk
6. Default → minimal-risk
```
## Annex III — The 8 Categories (Article 6(2))
### §1 — Biometrics (the heaviest category)
- Remote biometric identification systems
- Biometric categorisation according to sensitive or protected attributes (where not prohibited under Article 5)
- Emotion recognition (where not prohibited under Article 5)
**Conformity assessment:** Module H (notified body required) per Article 43(1).
**Carve-out applicability:** Article 6(3) carve-outs do NOT apply to biometric ID systems performing biometric verification. Carve-out can apply to other Annex III §1 systems only if profiling is absent.
### §2 — Critical Infrastructure
- AI used as safety component in management/operation of road, rail, air, water, gas, electricity, heating
**Carve-out applicability:** rarely satisfied — safety components by definition affect critical operation.
### §3 — Education and Vocational Training
- Determining access, admission, or assignment to educational institutions
- Evaluating learning outcomes including in steering learning process
- Assessing appropriate level of education for an individual
- Monitoring and detecting prohibited behaviour during tests
**Carve-out applicability:** narrow procedural tasks (e.g., automatic answer-sheet OCR) may carve out; substantive evaluation does not.
### §4 — Employment, Workers Management, Self-Employment Access (a frequent trigger)
- Recruitment / selection (e.g., placing targeted job ads, screening applications, evaluating candidates)
- Decisions about promotion, termination, task allocation based on individual behaviour or traits
- Monitoring/evaluating performance + behaviour
**Carve-out applicability:** profiling of natural persons is always present in employment AI by definition (Article 6(3) last sentence overrides carve-out claim).
### §5 — Access to Essential Private and Public Services
- Public benefits and services (eligibility evaluation)
- Credit scoring of natural persons (with limited exception for fraud detection)
- Risk assessment + pricing of life and health insurance
- Emergency dispatch services (police, fire, ambulance) prioritisation
**Carve-out applicability:** profiling typically present; carve-out rare.
### §6 — Law Enforcement (high political sensitivity)
- Risk assessment of natural persons becoming offender or victim
- Polygraphs and similar
- Reliability evaluation of evidence
- Predictive policing (subject to Article 5 prohibition limits)
- Profiling of natural persons under Article 3(4) GDPR
**Carve-out applicability:** rarely applicable; political bar high.
### §7 — Migration, Asylum, Border Control Management
- Polygraphs and similar
- Risk assessment of natural persons crossing borders
- Examination of applications for asylum, visa, residence permits
- Identifying / verifying natural persons at borders (except routine document checks)
**Carve-out applicability:** rarely applicable.
### §8 — Administration of Justice and Democratic Processes
- Assisting judicial authority in interpretation of facts and law and applying law to facts
- Influencing the outcome of elections or referendums or natural persons' voting behaviour (excludes purely organizational/logistical uses)
**Carve-out applicability:** rarely applicable in substantive use; logistical electoral systems may carve out.
## Article 6(3) Carve-Out Test
Per Article 6(3), an Annex III AI system is NOT high-risk if **at least one** of these conditions is met AND no profiling occurs:
| Carve-out | Description | Example |
|---|---|---|
| **(a)** | Performs a narrow procedural task | Automatic spell-check on application forms |
| **(b)** | Improves the result of a previously completed human activity | Polish-up tool applied after human-drafted decision |
| **(c)** | Detects decision-making patterns or deviations from prior decision-making patterns without replacing or influencing the human assessment | Auditing tool that flags inconsistency in past human decisions but does not generate decisions |
| **(d)** | Performs a preparatory task to an assessment relevant for the purposes referred to in Annex III | Organizing applications by submission date before human review |
**Critical override (last sentence of Article 6(3)):** if the AI system performs **profiling of natural persons**, it remains high-risk regardless of carve-out claim. Profiling is defined by Article 4(4) of GDPR: any form of automated processing of personal data consisting of using personal data to evaluate certain personal aspects relating to a natural person.
In practice: most decision-support / decision-making AI involving natural persons performs profiling. Carve-out works for narrow procedural / preparatory / aggregation tools, not for substantive evaluation.
## Provider's Article 6(4) Documentation Duty
If a provider claims Article 6(3) carve-out for an Annex III system, the provider must:
1. Document the rationale before placing on market
2. Register the system in the EU database (Article 71)
3. Make documentation available to national competent authorities on request
Failure to document the carve-out claim properly is itself a compliance failure subject to Article 99 penalties.
## Real-World Decision Heuristic
For each AI system, ask in order:
1. **Does it touch hiring, credit, insurance, education, law enforcement, migration, justice, or critical infrastructure?** If yes, continue. If no, skip to step 4.
2. **Does it influence (not just inform) decisions about natural persons?** If yes → high-risk per Annex III. Conformity assessment required.
3. **If it only informs / does narrow procedural work AND there's no profiling:** carve-out may apply. Document thoroughly. Still register if Annex III §1 / §6 / §7.
4. **Does it interact directly with natural persons, generate synthetic content, or do emotion recognition outside Article 5?** Article 50 transparency applies (limited-risk).
5. **Otherwise:** minimal-risk.
## When This Reference Doesn't Help
- **Whether a system is "an AI system" at all (Article 3(1)).** See Commission Guidelines Feb 2025.
- **Annex I sectoral product law overlap.** See sectoral regulation (MDR 745, machinery, toys, etc.).
- **GPAI separate track.** See `gpai_obligations.md`.
---
**Source authorities (non-exhaustive):**
- **Regulation (EU) 2024/1689** — Articles 5, 6, 7 and Annex III (binding text)
- **European Commission** — Guidelines on prohibited AI practices (Feb 2025)
- **European Commission** — Article 6(3) implementing guidelines (expected; check current Commission communications)
- **European Data Protection Board** — Opinion 28/2024 (Article 6 GDPR + AI Act interaction)
- **EDPS** — interpretive guidance on biometric and profiling provisions
- **Future of Life Institute** — Annex III decision tree (community reference)
- **IAPP EU AI Act Tracker** — running practitioner interpretation
- **National AI authorities** (per Article 70) — emerging Member State guidance: BfDI (Germany), CNIL (France), AEPD (Spain) AI position papers
FILE:scripts/ai_act_obligation_tracker.py
#!/usr/bin/env python3
"""ai_act_obligation_tracker.py — EU AI Act per-role obligation matrix.
Stdlib-only. Given an organization's role(s) per Article 25 (provider, deployer,
importer, distributor, authorized representative) and AI system tier(s), produces
a deadline-sorted obligation matrix tied to the Act's phased application:
- 2 Feb 2025: Article 5 prohibitions + Article 4 AI literacy
- 2 Aug 2025: GPAI Articles 51-55 + governance + penalties
- 2 Aug 2026: Title III high-risk obligations
- 2 Aug 2027: Annex I sectoral high-risk obligations
Deterministic logic referencing Articles 16, 22, 23, 24, 25, 26, 27, 50,
51-55, 72, 73 + phasing per Article 113.
Input schema (JSON):
{
"organization": "Acme AI Inc.",
"establishment": "non_eu", # eu | non_eu
"roles": [
{"role": "provider", "systems_tier": "high_risk"},
{"role": "deployer", "systems_tier": "high_risk", "public_sector": false},
{"role": "deployer", "systems_tier": "limited_risk"}
],
"deploys_gpai": true,
"gpai_systemic_risk": false
}
Usage:
python ai_act_obligation_tracker.py
python ai_act_obligation_tracker.py path/to/roles.json
python ai_act_obligation_tracker.py roles.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"organization": "Acme AI Inc.",
"establishment": "non_eu",
"roles": [
{"role": "provider", "systems_tier": "high_risk"},
{"role": "deployer", "systems_tier": "high_risk", "public_sector": False},
{"role": "deployer", "systems_tier": "limited_risk"},
],
"deploys_gpai": True,
"gpai_systemic_risk": False,
}
# Phasing reference (per Article 113)
PHASE_DATES = {
"article_5_prohibitions": "2025-02-02",
"article_4_ai_literacy": "2025-02-02",
"gpai_articles_51_55": "2025-08-02",
"governance_penalties": "2025-08-02",
"title_iii_high_risk_general": "2026-08-02",
"title_iii_annex_i_sectoral": "2027-08-02",
}
# Obligations per role + tier
PROVIDER_HIGH_RISK = [
("Article 9 — Establish risk management system across the full AI lifecycle", "title_iii_high_risk_general"),
("Article 10 — Data governance: training/validation/test data quality + bias mitigation", "title_iii_high_risk_general"),
("Article 11 — Maintain technical documentation per Annex IV", "title_iii_high_risk_general"),
("Article 12 — Implement automatic event logging", "title_iii_high_risk_general"),
("Article 13 — Provide instructions for use to deployers", "title_iii_high_risk_general"),
("Article 14 — Design for human oversight", "title_iii_high_risk_general"),
("Article 15 — Accuracy, robustness, cybersecurity", "title_iii_high_risk_general"),
("Article 16 — General provider obligations + named contact person", "title_iii_high_risk_general"),
("Article 17 — Establish quality management system (QMS)", "title_iii_high_risk_general"),
("Article 43 — Undertake conformity assessment before placing on market", "title_iii_high_risk_general"),
("Article 47 — Sign EU declaration of conformity (10-year retention per Article 18)", "title_iii_high_risk_general"),
("Article 48 — Affix CE marking", "title_iii_high_risk_general"),
("Article 49 — Register in EU database (Article 71) for Annex III systems", "title_iii_high_risk_general"),
("Article 72 — Establish post-market monitoring system", "title_iii_high_risk_general"),
("Article 73 — Report serious incidents to market surveillance authority within 15 days (or 2 days for critical-infrastructure incidents)", "title_iii_high_risk_general"),
]
DEPLOYER_HIGH_RISK = [
("Article 26(1) — Use the AI system according to provider's instructions for use", "title_iii_high_risk_general"),
("Article 26(2) — Assign human oversight to natural persons with necessary competence + authority + support", "title_iii_high_risk_general"),
("Article 26(3) — Ensure input data is relevant + sufficiently representative", "title_iii_high_risk_general"),
("Article 26(4) — Monitor operation; cease use if it presents Article 79 risk", "title_iii_high_risk_general"),
("Article 26(5) — Maintain automatically generated logs (Article 12) for ≥ 6 months", "title_iii_high_risk_general"),
("Article 26(7) — Inform workers + their representatives before putting the system into use in workplace", "title_iii_high_risk_general"),
("Article 26(8) — Cooperate with national competent authorities + AI Office", "title_iii_high_risk_general"),
("Article 50 — Inform natural persons subject to AI-decisions (transparency)", "title_iii_high_risk_general"),
("Article 86 — Right to explanation of individual decision", "title_iii_high_risk_general"),
]
DEPLOYER_PUBLIC_SECTOR = [
("Article 27 — Conduct Fundamental Rights Impact Assessment (FRIA) before deploying", "title_iii_high_risk_general"),
]
DEPLOYER_LIMITED_RISK = [
("Article 50(1) — Inform natural persons they are interacting with an AI system", "governance_penalties"),
("Article 50(4) — Disclose deepfakes (image, audio, video) as AI-generated; mark machine-readable", "governance_penalties"),
]
IMPORTER = [
("Article 23 — Verify provider completed conformity assessment + has technical docs", "title_iii_high_risk_general"),
("Article 23(3) — Indicate name, contact, address on the AI system or accompanying docs", "title_iii_high_risk_general"),
]
DISTRIBUTOR = [
("Article 24 — Verify CE marking + documentation before making the system available", "title_iii_high_risk_general"),
]
AUTH_REP_NON_EU_PROVIDER = [
("Article 22 — Non-EU providers MUST appoint an authorized representative established in the EU", "title_iii_high_risk_general"),
("Article 22(3) — Representative keeps technical docs available + liable for provider obligations", "title_iii_high_risk_general"),
]
GPAI_ALL = [
("Article 53 — Maintain up-to-date technical documentation of GPAI model", "gpai_articles_51_55"),
("Article 53 — Provide information to downstream providers integrating the model", "gpai_articles_51_55"),
("Article 53(1)(c) — Establish policy to comply with EU copyright law", "gpai_articles_51_55"),
("Article 53(1)(d) — Publish detailed summary about training data", "gpai_articles_51_55"),
]
GPAI_SYSTEMIC_RISK = [
("Article 55 — Perform model evaluations including adversarial testing", "gpai_articles_51_55"),
("Article 55 — Assess + mitigate systemic risks", "gpai_articles_51_55"),
("Article 55 — Track + report serious incidents to AI Office", "gpai_articles_51_55"),
("Article 55 — Ensure cybersecurity protection of the model + physical infrastructure", "gpai_articles_51_55"),
]
UNIVERSAL = [
("Article 4 — Ensure AI literacy of staff dealing with AI systems", "article_4_ai_literacy"),
("Article 5 — No prohibited AI practices", "article_5_prohibitions"),
]
def _make_obs(items: List[tuple], role_label: str) -> List[Dict[str, Any]]:
return [{"role": role_label, "obligation": ob, "deadline_phase": phase,
"deadline_date": PHASE_DATES[phase]} for ob, phase in items]
def _role_obligations(role: Dict[str, Any]) -> List[Dict[str, Any]]:
r_type = role.get("role")
tier = role.get("systems_tier")
if r_type == "provider" and tier == "high_risk":
return _make_obs(PROVIDER_HIGH_RISK, "provider/high-risk")
if r_type == "deployer" and tier == "high_risk":
out = _make_obs(DEPLOYER_HIGH_RISK, "deployer/high-risk")
if role.get("public_sector"):
out += _make_obs(DEPLOYER_PUBLIC_SECTOR, "deployer/public-sector")
return out
if r_type == "deployer" and tier == "limited_risk":
return _make_obs(DEPLOYER_LIMITED_RISK, "deployer/limited-risk")
if r_type == "importer":
return _make_obs(IMPORTER, "importer")
if r_type == "distributor":
return _make_obs(DISTRIBUTOR, "distributor")
return []
def gather_obligations(payload: Dict[str, Any]) -> List[Dict[str, Any]]:
obligations: List[Dict[str, Any]] = []
obligations += _make_obs(UNIVERSAL, "any")
roles = payload.get("roles", [])
for role in roles:
obligations += _role_obligations(role)
if payload.get("establishment") == "non_eu":
provider_role = any(r.get("role") == "provider" for r in roles)
if provider_role:
obligations += _make_obs(AUTH_REP_NON_EU_PROVIDER, "non-EU provider")
if payload.get("deploys_gpai"):
obligations += _make_obs(GPAI_ALL, "GPAI provider")
if payload.get("gpai_systemic_risk"):
obligations += _make_obs(GPAI_SYSTEMIC_RISK, "GPAI systemic risk")
obligations.sort(key=lambda x: (x["deadline_date"], x["role"]))
return obligations
def analyze(payload: Dict[str, Any]) -> Dict[str, Any]:
obs = gather_obligations(payload)
by_phase: Dict[str, int] = {}
by_role: Dict[str, int] = {}
for o in obs:
by_phase[o["deadline_phase"]] = by_phase.get(o["deadline_phase"], 0) + 1
by_role[o["role"]] = by_role.get(o["role"], 0) + 1
return {
"organization": payload.get("organization"),
"establishment": payload.get("establishment"),
"total_obligations": len(obs),
"by_phase": by_phase,
"by_role": by_role,
"obligations": obs,
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("EU AI ACT — OBLIGATION MATRIX (deadline-sorted)")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Organization: {r['organization']}")
lines.append(f"Establishment: {r['establishment']}")
lines.append(f"Total obligations: {r['total_obligations']}")
lines.append("")
lines.append("By deadline phase:")
for phase, n in sorted(r["by_phase"].items(), key=lambda x: PHASE_DATES.get(x[0], "")):
lines.append(f" {PHASE_DATES.get(phase, '?')} {phase:35s} {n} obligations")
lines.append("")
lines.append("By role:")
for role, n in sorted(r["by_role"].items()):
lines.append(f" {role:30s} {n} obligations")
lines.append("")
lines.append("-" * 72)
lines.append("FULL LIST (deadline order):")
lines.append("")
current_date = None
for o in r["obligations"]:
if o["deadline_date"] != current_date:
current_date = o["deadline_date"]
lines.append(f" >> Deadline {current_date} — {o['deadline_phase']}")
lines.append(f" [{o['role']:25s}] {o['obligation']}")
lines.append("")
lines.append("-" * 72)
lines.append("PHASING (Article 113):")
lines.append(" 2025-02-02: Article 5 prohibitions + Article 4 AI literacy")
lines.append(" 2025-08-02: GPAI (Art. 51-55) + governance + penalties")
lines.append(" 2026-08-02: Title III high-risk (general)")
lines.append(" 2027-08-02: Annex I sectoral high-risk")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="EU AI Act per-role obligation matrix with phasing deadlines.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to roles JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: non-EU provider + deployer high-risk + GPAI>"
result = analyze(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/ai_system_risk_classifier.py
#!/usr/bin/env python3
"""ai_system_risk_classifier.py — EU AI Act (2024/1689) risk-tier classifier.
Stdlib-only. Takes AI-system characteristics and classifies into one of:
- prohibited (Article 5)
- high-risk (Article 6 + Annex III, OR Article 6(1) + Annex I)
- limited-risk transparency (Article 50)
- minimal-risk (default)
Deterministic decision tree following the regulation's risk-based architecture
(Recital 26 + Articles 5, 6, 50). Article 6(3) carve-outs applied.
Input schema (JSON):
{
"systems": [
{
"name": "Resume screening AI",
"intended_purpose": "Filter and rank candidates for hiring",
"users": "internal_hr",
"data_processes_natural_persons": true,
"annex_iii_category": "employment",
"performs_profiling": true,
"article_5_practice": null,
"article_6_1_safety_component": false,
"article_6_3_carveout_applies": false,
"interacts_with_natural_persons_directly": false,
"is_general_purpose_ai_model": false,
"training_compute_flops": null
}
]
}
Usage:
python ai_system_risk_classifier.py # uses embedded 5-system sample
python ai_system_risk_classifier.py path/to/systems.json
python ai_system_risk_classifier.py systems.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Optional
SAMPLE: Dict[str, Any] = {
"systems": [
{
"name": "Emotion recognition in retail store CCTV",
"intended_purpose": "Detect emotions of shoppers to optimize layout",
"users": "store_managers",
"data_processes_natural_persons": True,
"annex_iii_category": None,
"performs_profiling": False,
"article_5_practice": "emotion_recognition_in_workplace_or_education",
"article_6_1_safety_component": False,
"article_6_3_carveout_applies": False,
"interacts_with_natural_persons_directly": False,
"is_general_purpose_ai_model": False,
"training_compute_flops": None,
},
{
"name": "CV-screening AI for job applications",
"intended_purpose": "Filter and rank candidates for shortlist",
"users": "internal_hr",
"data_processes_natural_persons": True,
"annex_iii_category": "employment",
"performs_profiling": True,
"article_5_practice": None,
"article_6_1_safety_component": False,
"article_6_3_carveout_applies": False,
"interacts_with_natural_persons_directly": False,
"is_general_purpose_ai_model": False,
"training_compute_flops": None,
},
{
"name": "Customer support chatbot",
"intended_purpose": "Answer support questions; route to human agents",
"users": "customers",
"data_processes_natural_persons": True,
"annex_iii_category": None,
"performs_profiling": False,
"article_5_practice": None,
"article_6_1_safety_component": False,
"article_6_3_carveout_applies": False,
"interacts_with_natural_persons_directly": True,
"is_general_purpose_ai_model": False,
"training_compute_flops": None,
},
{
"name": "Spam email filter",
"intended_purpose": "Classify inbound email as spam or not",
"users": "all_employees",
"data_processes_natural_persons": False,
"annex_iii_category": None,
"performs_profiling": False,
"article_5_practice": None,
"article_6_1_safety_component": False,
"article_6_3_carveout_applies": False,
"interacts_with_natural_persons_directly": False,
"is_general_purpose_ai_model": False,
"training_compute_flops": None,
},
{
"name": "Foundation model deployed via API",
"intended_purpose": "General-purpose text generation",
"users": "developers",
"data_processes_natural_persons": True,
"annex_iii_category": None,
"performs_profiling": False,
"article_5_practice": None,
"article_6_1_safety_component": False,
"article_6_3_carveout_applies": False,
"interacts_with_natural_persons_directly": False,
"is_general_purpose_ai_model": True,
"training_compute_flops": 5e25,
},
]
}
# Article 5 prohibited practices (per the binding regulation text)
ARTICLE_5_PRACTICES = {
"subliminal_manipulation": "Article 5(1)(a) — Subliminal techniques beyond awareness causing harm",
"exploitation_of_vulnerabilities": "Article 5(1)(b) — Exploiting vulnerabilities of age/disability/socioeconomic situation",
"social_scoring": "Article 5(1)(c) — Social scoring by public authorities causing detrimental treatment",
"predictive_policing_individual": "Article 5(1)(d) — Predictive policing based solely on profiling",
"untargeted_facial_scraping": "Article 5(1)(e) — Untargeted scraping of facial images for facial recognition databases",
"emotion_recognition_in_workplace_or_education": "Article 5(1)(f) — Emotion recognition in workplace and educational institutions",
"biometric_categorisation_sensitive": "Article 5(1)(g) — Biometric categorisation by sensitive attributes",
"real_time_remote_biometric_id_public_law_enforcement": "Article 5(1)(h) — Real-time remote biometric ID in publicly accessible spaces for law enforcement",
}
# Annex III high-risk categories (the 8 — Article 6(2))
ANNEX_III_CATEGORIES = {
"biometrics": "Annex III §1 — Biometrics including biometric ID and categorisation",
"critical_infrastructure": "Annex III §2 — Critical infrastructure (safety components)",
"education": "Annex III §3 — Education and vocational training",
"employment": "Annex III §4 — Employment, workers management, self-employment access",
"essential_services": "Annex III §5 — Access to essential private/public services and benefits (including credit scoring, emergency dispatch, insurance pricing)",
"law_enforcement": "Annex III §6 — Law enforcement",
"migration_asylum": "Annex III §7 — Migration, asylum, border control",
"justice_democratic_processes": "Annex III §8 — Administration of justice and democratic processes",
}
def classify(system: Dict[str, Any]) -> Dict[str, Any]:
"""Deterministic classification per Articles 5, 6, 50 + Annex III."""
name = system.get("name", "<unnamed>")
article_5 = system.get("article_5_practice")
annex_iii = system.get("annex_iii_category")
safety_component = system.get("article_6_1_safety_component", False)
carveout = system.get("article_6_3_carveout_applies", False)
profiling = system.get("performs_profiling", False)
interacts = system.get("interacts_with_natural_persons_directly", False)
is_gpai = system.get("is_general_purpose_ai_model", False)
flops = system.get("training_compute_flops")
# Step 1: Article 5 prohibitions (binary, no carve-out)
if article_5 and article_5 in ARTICLE_5_PRACTICES:
return {
"name": name,
"tier": "prohibited",
"primary_citation": ARTICLE_5_PRACTICES[article_5],
"rationale": "Listed Article 5 practice. Cannot be placed on EU market or used (penalty up to EUR 35M / 7% turnover).",
"is_gpai": is_gpai,
"gpai_systemic_risk": False,
}
# Step 2: Article 6(1) — safety component of regulated product per Annex I
if safety_component:
return {
"name": name,
"tier": "high_risk",
"primary_citation": "Article 6(1) — Safety component of Annex I product",
"rationale": "Safety component subject to third-party conformity assessment under sectoral law (Annex I).",
"is_gpai": is_gpai,
"gpai_systemic_risk": False,
}
# Step 3: Article 6(2) + Annex III — high-risk by category
if annex_iii and annex_iii in ANNEX_III_CATEGORIES:
# Article 6(3) carve-out check
if carveout and not profiling:
# Carve-out applies AND no profiling — drops to limited or minimal
tier = "limited_risk" if interacts else "minimal_risk"
return {
"name": name,
"tier": tier,
"primary_citation": "Article 6(3) carve-out from Annex III — narrow procedural task / preparatory / human-result improvement",
"rationale": "Annex III category triggered but Article 6(3) carve-out applies and no profiling.",
"is_gpai": is_gpai,
"gpai_systemic_risk": False,
}
if carveout and profiling:
# Profiling overrides carve-out — Article 6(3) last sentence
return {
"name": name,
"tier": "high_risk",
"primary_citation": f"Article 6(2) + {ANNEX_III_CATEGORIES[annex_iii]}",
"rationale": "Carve-out claimed but profiling of natural persons keeps it high-risk per Article 6(3) last sentence.",
"is_gpai": is_gpai,
"gpai_systemic_risk": False,
}
return {
"name": name,
"tier": "high_risk",
"primary_citation": f"Article 6(2) + {ANNEX_III_CATEGORIES[annex_iii]}",
"rationale": "Falls in Annex III high-risk category; no Article 6(3) carve-out applied.",
"is_gpai": is_gpai,
"gpai_systemic_risk": False,
}
# Step 4: Article 50 transparency (limited-risk)
if interacts:
return {
"name": name,
"tier": "limited_risk",
"primary_citation": "Article 50(1) — Transparency for AI systems interacting with natural persons",
"rationale": "Direct interaction with natural persons requires disclosure that they are interacting with AI.",
"is_gpai": is_gpai,
"gpai_systemic_risk": _gpai_systemic_risk(is_gpai, flops),
}
# Step 5: Default — minimal-risk
return {
"name": name,
"tier": "minimal_risk",
"primary_citation": "No Article 5, Annex III, or Article 50 trigger",
"rationale": "Minimal-risk default. No obligations under the Act (Article 95 voluntary codes of conduct only).",
"is_gpai": is_gpai,
"gpai_systemic_risk": _gpai_systemic_risk(is_gpai, flops),
}
def _gpai_systemic_risk(is_gpai: bool, flops: Optional[float]) -> bool:
"""Article 51 — systemic-risk GPAI threshold: training compute ≥ 10^25 FLOPs."""
if not is_gpai or flops is None:
return False
return flops >= 1e25
def annotate_all(payload: Dict[str, Any]) -> Dict[str, Any]:
classified = [classify(s) for s in payload.get("systems", [])]
tier_counts: Dict[str, int] = {}
for c in classified:
tier_counts[c["tier"]] = tier_counts.get(c["tier"], 0) + 1
gpai_systems = [c["name"] for c in classified if c["is_gpai"]]
systemic_risk = [c["name"] for c in classified if c["gpai_systemic_risk"]]
return {
"total_systems": len(classified),
"by_tier": tier_counts,
"gpai_systems": gpai_systems,
"gpai_systemic_risk_systems": systemic_risk,
"systems": classified,
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("EU AI ACT (Reg. 2024/1689) — RISK CLASSIFICATION")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Total systems: {r['total_systems']}")
lines.append(f"By tier: {r['by_tier']}")
if r["gpai_systems"]:
lines.append(f"GPAI systems: {', '.join(r['gpai_systems'])}")
if r["gpai_systemic_risk_systems"]:
lines.append(f"GPAI with systemic risk (Article 51): {', '.join(r['gpai_systemic_risk_systems'])}")
lines.append("")
lines.append("-" * 72)
for s in r["systems"]:
tier_label = s["tier"].replace("_", "-").upper()
gpai_flag = " [GPAI]" if s["is_gpai"] else ""
sysrisk_flag = " [SYSTEMIC RISK]" if s["gpai_systemic_risk"] else ""
lines.append(f" {s['name']}{gpai_flag}{sysrisk_flag}")
lines.append(f" Tier: {tier_label}")
lines.append(f" Citation: {s['primary_citation']}")
lines.append(f" Rationale: {s['rationale']}")
lines.append("")
lines.append("-" * 72)
lines.append("DECISION ORDER: Article 5 prohibitions → Article 6(1) Annex I → Article 6(2) Annex III")
lines.append(" → Article 6(3) carve-outs (overridden by profiling) → Article 50 transparency → minimal-risk default")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="EU AI Act risk tier classifier per Articles 5/6/50 + Annex III.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to systems JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: 5 systems across all 4 tiers + 1 GPAI>"
result = annotate_all(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/conformity_assessment_planner.py
#!/usr/bin/env python3
"""conformity_assessment_planner.py — EU AI Act Article 43 conformity routing + Annex IV checklist.
Stdlib-only. For a high-risk AI system, selects the conformity assessment Module
(A internal control vs H full QMS + notified body) per Article 43 and produces the
Annex IV technical documentation checklist.
Decision rule (Article 43):
- Biometrics (Annex III §1) → Module H (notified body required) by default
- All other Annex III categories → Module A (internal control) is permissible
where harmonised standards are applied (Article 40)
- Annex I products (safety components) → follow sectoral law's existing procedure
Input schema (JSON):
{
"system_name": "CV-screening AI",
"annex_iii_category": "employment",
"applies_harmonised_standards": true,
"harmonised_standards_referenced": ["EN ISO/IEC 42001", "EN ISO/IEC 23894"],
"annex_i_product": false,
"annex_i_sectoral_law": null,
"existing_iso_42001_certification": false,
"existing_iso_27001_certification": true
}
Usage:
python conformity_assessment_planner.py # embedded sample
python conformity_assessment_planner.py path/to/system.json
python conformity_assessment_planner.py system.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"system_name": "CV-screening AI for hiring",
"annex_iii_category": "employment",
"applies_harmonised_standards": True,
"harmonised_standards_referenced": ["EN ISO/IEC 42001", "EN ISO/IEC 23894"],
"annex_i_product": False,
"annex_i_sectoral_law": None,
"existing_iso_42001_certification": False,
"existing_iso_27001_certification": True,
}
# Annex IV — Technical Documentation requirements (per Article 11(1))
ANNEX_IV_ITEMS = [
{
"id": "iv.1",
"title": "General description of the AI system",
"subitems": [
"intended purpose",
"name & version of provider",
"system architecture overview",
"instructions for use (Article 13)",
],
"reusable_from": "ISO 42001 SKILL scope statement; ISO 27001 system documentation",
},
{
"id": "iv.2",
"title": "Detailed description of system elements",
"subitems": [
"methods used (ML, rule-based, etc.)",
"training, validation, test datasets (provenance + quality + bias mitigation per Article 10)",
"human oversight measures (Article 14)",
"key design choices including assumptions",
"computational resources used",
],
"reusable_from": "ISO 42001 A.6 lifecycle documentation; ISO 42001 A.7 data evidence; model cards",
},
{
"id": "iv.3",
"title": "Information about monitoring, functioning, control",
"subitems": [
"performance metrics & expected accuracy",
"logging capabilities (Article 12)",
"input data specifications",
"human-in-the-loop and oversight (Article 14)",
],
"reusable_from": "ISO 42001 A.9.3 monitoring; ISO 42001 A.9.4 logging",
},
{
"id": "iv.4",
"title": "Description of risk management system",
"subitems": [
"Article 9 risk management process",
"identified risks + mitigation measures",
"residual risk acceptance",
"testing methodology",
],
"reusable_from": "ISO 42001 Clause 6.1 + Annex A.5 + Annex A.6.2.4; ISO 23894 process",
},
{
"id": "iv.5",
"title": "Description of changes to the system after placing on market",
"subitems": [
"change-management procedure",
"version control of model + data",
"re-evaluation triggers (concept drift, fine-tuning)",
],
"reusable_from": "ISO 27001 A.8.32 change management; ISO 42001 A.6.2.5 deployment",
},
{
"id": "iv.6",
"title": "List of harmonised standards applied",
"subitems": [
"presumption of conformity per Article 40",
"alternative solutions documented where standards not applied",
],
"reusable_from": "Standards register",
},
{
"id": "iv.7",
"title": "EU declaration of conformity",
"subitems": [
"Article 47 — provider declares conformity, signed by authorized signatory",
"kept for 10 years post-market (Article 18)",
],
"reusable_from": "Template only — signed at end of process",
},
{
"id": "iv.8",
"title": "Post-market monitoring system",
"subitems": [
"Article 72 — proactive collection of performance + incident data",
"serious incident reporting procedure (Article 73)",
"feedback loop into risk management (Article 9)",
],
"reusable_from": "ISO 42001 A.9.3 monitoring + ISO 13485 post-market surveillance pattern",
},
]
def select_module(payload: Dict[str, Any]) -> Dict[str, Any]:
"""Select conformity assessment Module per Article 43."""
annex_iii = payload.get("annex_iii_category")
applies_standards = payload.get("applies_harmonised_standards", False)
annex_i = payload.get("annex_i_product", False)
sectoral_law = payload.get("annex_i_sectoral_law")
if annex_i and sectoral_law:
return {
"module": "sectoral",
"citation": "Article 43(3) — Annex I product follows existing sectoral conformity procedure",
"notified_body_required": "depends_on_sectoral_law",
"rationale": f"Follow {sectoral_law} existing procedure; AI Act layered on top.",
}
if annex_iii == "biometrics":
return {
"module": "H",
"citation": "Article 43(1) + Annex VII — Full QMS + Notified Body for biometrics",
"notified_body_required": "yes",
"rationale": "Biometrics under Annex III §1 require notified-body involvement by default.",
}
if annex_iii and applies_standards:
return {
"module": "A",
"citation": "Article 43(2) + Annex VI — Internal control with presumption of conformity",
"notified_body_required": "no",
"rationale": "Annex III system applying harmonised standards (Article 40) may use internal control.",
}
if annex_iii and not applies_standards:
return {
"module": "A_with_caveats",
"citation": "Article 43(2) + Annex VI — Internal control without harmonised standards",
"notified_body_required": "optional_but_recommended",
"rationale": "Internal control still permitted but without presumption of conformity; document alternative compliance evidence in full.",
}
return {
"module": "not_applicable",
"citation": "System not classified as high-risk; conformity assessment not required",
"notified_body_required": "no",
"rationale": "Re-run ai_system_risk_classifier.py to confirm tier.",
}
def reuse_summary(payload: Dict[str, Any]) -> List[str]:
"""What evidence can be reused from existing certifications."""
notes = []
if payload.get("existing_iso_42001_certification"):
notes.append("ISO 42001 certification: reuse AIMS Clause 6.1 risk evidence (Annex IV item 4)")
notes.append("ISO 42001 certification: reuse Annex A.6 lifecycle evidence (Annex IV items 1-3)")
notes.append("ISO 42001 certification: reuse Annex A.9 monitoring evidence (Annex IV item 8)")
if payload.get("existing_iso_27001_certification"):
notes.append("ISO 27001 certification: reuse cybersecurity evidence for Article 15 cybersecurity requirement")
notes.append("ISO 27001 certification: reuse A.5.19 supplier mgmt for Article 25 value-chain responsibilities")
notes.append("ISO 27001 certification: reuse A.8.15 logging for Annex IV item 3 logging")
if not notes:
notes.append("No prior certifications declared; build all Annex IV evidence from scratch")
return notes
def plan(payload: Dict[str, Any]) -> Dict[str, Any]:
module = select_module(payload)
return {
"system_name": payload.get("system_name"),
"annex_iii_category": payload.get("annex_iii_category"),
"conformity_assessment": module,
"annex_iv_checklist": ANNEX_IV_ITEMS,
"reuse_from_existing_certifications": reuse_summary(payload),
"next_steps": _next_steps(module["module"]),
}
def _next_steps(module: str) -> List[str]:
base = [
"Assemble Annex IV pack per the checklist (see Article 11 + Annex IV).",
"Conduct Article 9 risk management lifecycle (input to Annex IV item 4).",
"Implement Article 12 logging capabilities (input to Annex IV item 3).",
"Implement Article 14 human-oversight measures (input to Annex IV items 2-3).",
"Stand up Article 72 post-market monitoring (input to Annex IV item 8).",
]
if module == "H":
base.append("Engage notified body for Module H assessment (Annex VII).")
base.append("Operate full QMS per Article 17 — pair with ISO 42001 AIMS for cross-reuse.")
elif module == "A":
base.append("Verify each harmonised standard referenced is on Article 40 list at decision date.")
base.append("Sign EU declaration of conformity (Article 47) AFTER assembling Annex IV pack.")
base.append("Affix CE marking (Article 48).")
base.append("Register in EU database (Article 71) before placing on market.")
elif module == "A_with_caveats":
base.append("Document equivalent alternative evidence for each requirement without a harmonised standard.")
base.append("Consider voluntary notified-body engagement to reduce regulatory risk.")
return base
def render_text(p: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("EU AI ACT — CONFORMITY ASSESSMENT PLAN")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"System: {p['system_name']}")
lines.append(f"Annex III category: {p['annex_iii_category']}")
lines.append("")
c = p["conformity_assessment"]
lines.append(f"Conformity Module: {c['module']}")
lines.append(f"Citation: {c['citation']}")
lines.append(f"Notified body required: {c['notified_body_required']}")
lines.append(f"Rationale: {c['rationale']}")
lines.append("")
lines.append("-" * 72)
lines.append("ANNEX IV TECHNICAL DOCUMENTATION CHECKLIST (8 items):")
lines.append("")
for item in p["annex_iv_checklist"]:
lines.append(f" [{item['id']}] {item['title']}")
for sub in item["subitems"]:
lines.append(f" - {sub}")
lines.append(f" Reusable: {item['reusable_from']}")
lines.append("")
lines.append("-" * 72)
lines.append("REUSE FROM EXISTING CERTIFICATIONS:")
for note in p["reuse_from_existing_certifications"]:
lines.append(f" - {note}")
lines.append("")
lines.append("-" * 72)
lines.append("NEXT STEPS:")
for step in p["next_steps"]:
lines.append(f" - {step}")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="EU AI Act Article 43 conformity routing + Annex IV technical documentation checklist.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to system JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: CV-screening AI, harmonised standards applied>"
result = plan(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
Đánh giá và xếp hạng kết quả của các agent theo chỉ số hoặc LLM làm giám khảo cho một phiên AgentHub.
---
name: "eval"
description: "Evaluate and rank agent results by metric or LLM judge for an AgentHub session."
command: /hub:eval
---
# /hub:eval — Evaluate Agent Results
Rank all agent results for a session. Supports metric-based evaluation (run a command), LLM judge (compare diffs), or hybrid.
## Usage
```
/hub:eval # Eval latest session using configured criteria
/hub:eval 20260317-143022 # Eval specific session
/hub:eval --judge # Force LLM judge mode (ignore metric config)
```
## What It Does
### Metric Mode (eval command configured)
Run the evaluation command in each agent's worktree:
```bash
python {skill_path}/scripts/result_ranker.py \
--session {session-id} \
--eval-cmd "{eval_cmd}" \
--metric {metric} --direction {direction}
```
Output:
```
RANK AGENT METRIC DELTA FILES
1 agent-2 142ms -38ms 2
2 agent-1 165ms -15ms 3
3 agent-3 190ms +10ms 1
Winner: agent-2 (142ms)
```
### LLM Judge Mode (no eval command, or --judge flag)
For each agent:
1. Get the diff: `git diff {base_branch}...{agent_branch}`
2. Read the agent's result post from `.agenthub/board/results/agent-{i}-result.md`
3. Compare all diffs and rank by:
- **Correctness** — Does it solve the task?
- **Simplicity** — Fewer lines changed is better (when equal correctness)
- **Quality** — Clean execution, good structure, no regressions
Present rankings with justification.
Example LLM judge output for a content task:
```
RANK AGENT VERDICT WORD COUNT
1 agent-1 Strong narrative, clear CTA 1480
2 agent-3 Good data points, weak intro 1520
3 agent-2 Generic tone, no differentiation 1350
Winner: agent-1 (strongest narrative arc and call-to-action)
```
### Hybrid Mode
1. Run metric evaluation first
2. If top agents are within 10% of each other, use LLM judge to break ties
3. Present both metric and qualitative rankings
## After Eval
1. Update session state:
```bash
python {skill_path}/scripts/session_manager.py --update {session-id} --state evaluating
```
2. Tell the user:
- Ranked results with winner highlighted
- Next step: `/hub:merge` to merge the winner
- Or `/hub:merge {session-id} --agent {winner}` to be explicit
Bộ skill phân tích tài chính: phân tích tỷ số, định giá DCF, chênh lệch ngân sách, dự báo cuốn chiếu và 4 công cụ Python.
--- name: "finance-skills" description: "Financial analyst agent skill and plugin for Claude Code, Codex, Gemini CLI, Cursor, OpenClaw. Ratio analysis, DCF valuation, budget variance, rolling forecasts. 4 Python tools (stdlib-only)." version: 2.9.0 author: Alireza Rezvani license: MIT tags: - finance - financial-analysis - dcf - valuation - budgeting agents: - claude-code - codex-cli - openclaw --- # Finance Skills Production-ready financial analysis skill for strategic decision-making. ## Quick Start ### Claude Code ``` /read finance/financial-analyst/SKILL.md ``` ### Codex CLI ```bash npx agent-skills-cli add alirezarezvani/claude-skills/finance ``` ## Skills Overview | Skill | Folder | Focus | |-------|--------|-------| | Financial Analyst | `financial-analyst/` | Ratio analysis, DCF, budget variance, forecasting | ## Python Tools 4 scripts, all stdlib-only: ```bash python3 financial-analyst/scripts/ratio_calculator.py --help python3 financial-analyst/scripts/dcf_valuation.py --help python3 financial-analyst/scripts/budget_variance_analyzer.py --help python3 financial-analyst/scripts/forecast_builder.py --help ``` ## Rules - Load only the specific skill SKILL.md you need - Always validate financial outputs against source data
Chất vấn của Chief Data Officer với kế hoạch về dữ liệu huấn luyện, kiến trúc dữ liệu, sản phẩm hóa dữ liệu và nhân sự.
--- name: "cdo-review" description: "/cs:cdo-review <plan> — Decision-driven Chief Data Officer interrogation of any plan that touches training data, data architecture, data productization, or data team hiring." --- # /cs:cdo-review — CDO Forcing Questions **Command:** `/cs:cdo-review <plan>` The decision-driven CDO pressure-tests any plan that touches data strategy. Six questions before any commitment to a data architecture, AI training run, data productization, or data team hire. ## When to Run - Before approving any new ML model training run that uses customer data - Before signing a multi-year data-infrastructure SaaS contract (Snowflake, Databricks, Fivetran) - Before productizing any customer data (benchmark report, embedding endpoint, license) - Before a major data team hire (head of data, CDO, data PM, ML engineer) - Before M&A diligence — yours or theirs - When the founder uses the word "monetize" near "data" ## The Six CDO Questions ### 1. What decision does this data drive? **If no decision is unblocked, why are we collecting / training on / productizing it?** - "We might need it later" is not a decision. - "It feels like a moat" is not a decision. - A real answer names a specific business call that requires this data. ### 2. What's the consent provenance for every source? **For each data source: origin, consent flow, data class, intended use.** - 1st-party-TOS-only is weaker than 1st-party-explicit-opt-in. - Bundled TOS doesn't cover material new purposes (training on PII for foundation models). - Run `ai_training_data_audit.py` if there's any AI use case in scope. ### 3. Who consumes this internally — and how many distinct functional domains? **Drives the centralize-vs-embed and warehouse-vs-mesh decisions.** - <5 consumers: warehouse-only. - 5-25 consumers: lakehouse. - 25+ consumers + federated culture: mesh. - Premature architecture choice is the #1 cause of data-team burnout. ### 4. What's the M&A diligence impact? **If an acquirer asks about this data corpus tomorrow, are we ready?** - Is there a documented anonymization process? - What % of customers have MSA carve-outs? - Are training-data provenance logs current? - Run `data_asset_valuator.py` quarterly. ### 5. Can the model / decision / report be retrained / re-run / re-published without this source? **Tests how much you depend on a specific data source.** - If yes → low blast radius; you can change consent posture later. - If no → high blast radius; you've structurally committed to the source. Vet harder. ### 6. What role unblocks this — and is it the right next hire? **Wrong hire (data scientist) when right answer (analytics engineer) is a 12-month productivity loss.** - Map the decision being unblocked to the specific role. - Confirm prerequisite roles are in place (data engineer before ML engineer, analyst before data scientist). ## Workflow ```bash # 1. AI training audit (if any ML / AI use case) python ../../../skills/chief-data-officer-advisor/scripts/ai_training_data_audit.py sources.json # 2. Architecture decision (if changing the stack) python ../../../skills/chief-data-officer-advisor/scripts/data_product_strategy_picker.py profile.json # 3. Data asset valuation (if productizing or pre-M&A) python ../../../skills/chief-data-officer-advisor/scripts/data_asset_valuator.py corpus.json ``` ## Output Format ```markdown # CDO Review: <plan> **Date:** YYYY-MM-DD ## The Decision Being Made [one sentence — which of the four CDO decisions: training | architecture | asset | hire] ## Training Audit (if applicable) - NO-GO sources: N - MITIGATE sources: N - GO sources: N - Top remediation: <one line> ## Architecture (if applicable) - Recommended: WAREHOUSE / LAKEHOUSE / MESH - Build-vs-buy summary: <one line> - Kill criteria: <when to revisit> ## Asset Value (if applicable) - Strategic value: X/10 | Moat: STRONG / MEDIUM / WEAK - M&A multiplier: X.Xx – X.Xx ARR - Recommended productization path: <name> ## Org (if applicable) - Next hire: <role> - Why this, not that: <one line> - Prerequisite hires in place: yes/no ## Verdict 🟢 SHIP | 🟡 SHARPEN | 🔴 BLOCK ## Next Steps [3 concrete actions] ``` ## Routing - `/cs:gc-review` — for any productization or licensing path - `/cs:ciso-review` — for any architecture change touching customer data - `/cs:cfo-review` — for build-vs-buy TCO and M&A valuation math - `/cs:chro-review` — for data team hires (comp, ladder, leveling) - `/cs:decide` — log the verdict - `/cs:freeze 90` — on multi-year infrastructure contracts ## Related - Agent: [`cs-cdo-advisor`](../../agents/cs-cdo-advisor.md) - Skill: [`chief-data-officer-advisor`](../../../skills/chief-data-officer-advisor/SKILL.md) - Adjacent: `../../../skills/general-counsel-advisor/` (contractual constraints), `../../../skills/cto-advisor/` (architecture capacity) --- **Version:** 1.0.0
Chạy phân tích tỷ số tài chính, định giá DCF, chênh lệch ngân sách và dự báo cuốn chiếu từ tệp dữ liệu JSON.
--- name: financial-health description: Run financial ratio analysis, DCF valuation, budget variance analysis, and rolling forecasts. Usage: /financial-health <ratios|dcf|budget|forecast> <data.json> --- # /financial-health Analyze financial statements, build valuation models, assess budget variances, and construct forecasts. ## Usage ``` /financial-health ratios <financial_data.json> [--format json|text] /financial-health dcf <valuation_data.json> [--format json|text] /financial-health budget <budget_data.json> [--format json|text] /financial-health forecast <forecast_data.json> [--format json|text] ``` ## Examples ``` /financial-health ratios quarterly_financials.json --format json /financial-health dcf acme_valuation.json /financial-health budget q1_budget.json --format json /financial-health forecast revenue_history.json ``` ## Scripts - `finance/financial-analyst/scripts/ratio_calculator.py` — Profitability, liquidity, leverage, efficiency, valuation ratios - `finance/financial-analyst/scripts/dcf_valuation.py` — DCF enterprise and equity valuation with sensitivity analysis - `finance/financial-analyst/scripts/budget_variance_analyzer.py` — Actual vs budget vs prior year variance analysis - `finance/financial-analyst/scripts/forecast_builder.py` — Driver-based revenue forecasting with scenario modeling ## Skill Reference → `finance/financial-analyst/SKILL.md` ## Related Commands - `/saas-health` — SaaS-specific metrics (ARR, MRR, churn, CAC, LTV, Quick Ratio)
Tư vấn pháp lý cho startup: rà soát hợp đồng (MSA, SaaS, NDA, DPA), chiến lược sở hữu trí tuệ, term sheet và bản đồ quy định.
---
name: "general-counsel-advisor"
description: "General Counsel advisory for startups: contract review (MSA, SaaS, NDA, DPA, employment), IP strategy, term sheet decoding, and regulatory landscape mapping. Use when reviewing any contract or term sheet, deciding when to engage outside counsel, defining IP strategy, evaluating regulatory exposure (HIPAA, GDPR, FDA, fintech), or when user mentions general counsel, GC, legal review, contract risk, term sheet, IP assignment, or regulatory exposure. NOT a substitute for licensed counsel — surfaces questions to bring to qualified attorneys."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: general-counsel-leadership
updated: 2026-05-12
python-tools: contract_risk_scanner.py, term_sheet_analyzer.py
frameworks: contract-review, ip-strategy, term-sheet-decoding, regulatory-mapping
---
# General Counsel Advisor
Strategic legal frameworks for startup General Counsels and founders without one. Contract risk, IP strategy, term sheet decoding, regulatory landscape.
This is **not legal advice**. It surfaces the right questions to bring to qualified outside counsel and catches the obvious traps before they reach a signature. Treat every output as a starting point for a conversation with a licensed attorney, not as a substitute for one.
## Keywords
general counsel, GC, legal review, contract review, MSA, SaaS agreement, NDA, DPA, employment agreement, contractor agreement, IP assignment, invention assignment, open source license, OSS compliance, term sheet, liquidation preference, anti-dilution, option pool, vesting, acceleration, drag-along, pro-rata, board composition, regulatory, HIPAA, GDPR, CCPA, FDA, MDR, fintech, BSA/AML, money transmitter, AI Act, indemnity, liability cap, force majeure, auto-renewal, choice of law, venue, non-compete, non-solicit
## Quick Start
```bash
# Scan a contract for risky clauses (uses bundled sample if no path given)
python scripts/contract_risk_scanner.py
python scripts/contract_risk_scanner.py path/to/contract.txt
# Analyze a term sheet for founder-friendliness
python scripts/term_sheet_analyzer.py
python scripts/term_sheet_analyzer.py path/to/term_sheet.json
```
## Key Questions (ask these first)
- **Who owns the IP being created or shared?** (Founders forget that contractors don't auto-assign IP without a written clause.)
- **What's the liability cap, and what's carved out?** (Standard: 12 months of fees, with carve-outs for IP infringement, data breach, willful misconduct.)
- **Is there a DPA in place if any personal data flows?** (GDPR, CCPA, state laws — non-negotiable if EU/CA data is touched.)
- **What's the termination right, notice period, and auto-renewal trap?** (5-year auto-renew with 60-day notice is a common founder mistake.)
- **Does this contract or product launch trigger a new regulatory regime?** (Healthcare → HIPAA. Fintech → BSA/AML. Medical device → FDA/MDR.)
- **For term sheets: liquidation preference, pre-money option pool, anti-dilution flavor?** (Three places where 5% of founder economics can quietly disappear.)
## Core Responsibilities
### 1. Contract Review
Standard contracts a startup signs in its first 5 years:
- **Vendor MSA** — Master Service Agreement (cloud, tooling, services)
- **Customer SaaS Agreement** — your standard customer paper + customer redlines
- **NDA** — mutual + one-way, with carve-outs for residuals + independent development
- **DPA** — Data Processing Agreement (required when personal data flows)
- **Employment Agreement** — offer letter, IP assignment, non-compete (where enforceable), arbitration
- **Contractor / 1099 Agreement** — IP assignment is critical; misclassification risk
- **Equity Agreements** — option grants, RSU agreements, advisor grants (FAST template, YC SAFE for advisors)
**Run** `contract_risk_scanner.py` on the text. It flags the 12 most common founder-killer clauses.
### 2. IP Strategy
- **Invention assignment** — every employee and contractor signs one. No exceptions.
- **Open source license compliance** — track every OSS dependency's license; AGPL and GPL trigger copyleft obligations.
- **Trade secrets** — define what's protected and how (clean room dev, access controls, NDAs).
- **Patents** — file provisional within 12 months of disclosure; PCT for international.
- **Trademarks** — register the word mark first, design mark second; clear before launch.
- **Copyright** — automatic on creation, but register for statutory damages eligibility.
See `references/ip_and_regulatory.md`.
### 3. Term Sheet Decoding
When a term sheet arrives, the difference between a founder-friendly and founder-hostile sheet often hides in three clauses:
- **Liquidation preference** — 1x non-participating is standard; 1x participating or 2x is hostile
- **Pre-money vs post-money option pool** — pre-money pool dilutes founders; post-money dilutes everyone proportionally
- **Anti-dilution** — broad-based weighted average is standard; full ratchet is hostile
**Run** `term_sheet_analyzer.py` to get a 0-100 founder-friendliness score with flags.
### 4. Regulatory Landscape
When to engage outside counsel **before** committing:
| Trigger | Regime | First Step |
|---|---|---|
| Healthcare data | HIPAA, HITECH, state breach laws | Specialist health-tech counsel |
| Cardholder data | PCI DSS (industry standard, not law, but contractually required) | QSA + counsel |
| Money movement | BSA/AML, state money-transmitter (50-state patchwork) | Fintech specialist |
| Medical device claims | FDA 510(k) / De Novo / PMA, MDR (EU), ISO 13485 | Medical-device specialist |
| EU residents' personal data | GDPR + EU AI Act if AI is deployed | EU privacy counsel |
| California residents | CCPA / CPRA | Privacy generalist |
| Securities (tokens, equity crowdfunding) | SEC rules (Reg D, Reg A+, Reg CF) | Securities counsel |
| Defense / aerospace customers | ITAR, EAR, DFARS, CMMC | Export-control counsel |
| AI in EU | EU AI Act (risk-tiered) | EU privacy + product counsel |
| AI for hiring (NYC, CO, IL) | Local bias-audit laws | Employment counsel |
See `references/ip_and_regulatory.md` for sequencing.
## Workflows
### Workflow 1: Contract Review
1. Save the contract as plain text
2. Run `contract_risk_scanner.py path/to/contract.txt`
3. For each HIGH risk finding, draft a counter-proposal
4. Bring the redline + counter-proposals to outside counsel
5. Log the decision via `/cs:decide`
### Workflow 2: Term Sheet Response
1. Save the term sheet as a JSON file matching the schema in `term_sheet_analyzer.py --help`
2. Run `python scripts/term_sheet_analyzer.py path/to/term_sheet.json`
3. Review the founder-friendliness score and per-clause flags
4. Negotiate the worst 3 clauses (don't try to win all 20)
5. Always have a securities/venture attorney review before signing
6. Log via `/cs:decide` with `/cs:freeze 30` to prevent regret-driven re-opening
### Workflow 3: IP Hygiene Audit
1. Confirm every employee and contractor (past 12 months) signed invention assignment
2. Run an OSS license inventory (`pip-licenses`, `license-checker` for npm)
3. Map AGPL/GPL dependencies and confirm compliance (or remove)
4. File provisional patents on novel inventions (12-month deadline from disclosure)
5. Register word-mark trademarks for the product name
### Workflow 4: Regulatory Trigger Assessment
1. List planned product features for the next 12 months
2. Map each feature to the trigger table in this document
3. For any HIPAA / FDA / fintech trigger, engage a specialist counsel **before** building
4. Document the regulatory roadmap and budget alongside the product roadmap
5. Pair with `cs-ciso-advisor` for ISO 27001 / SOC 2 sequencing
## Output Standard (when invoked via `/cs:gc-review`)
```
**Bottom Line:** [sign / negotiate / do not sign]
**The Risks:** [3 highest-severity issues]
**Counter-Proposals:** [specific language]
**Outside Counsel Action Items:** [what to bring to the attorney]
**Your Decision:** [the call only the founder can make]
```
## Adjacent Skills
- `../ciso-advisor/` — Compliance overlap (SOC 2, ISO 27001, HIPAA technical safeguards)
- `../cfo-advisor/` — Term sheet → dilution math
- `../ma-playbook/` — Acquisition agreements, integration playbooks
- `../../../ra-qm-team/` — ISO 13485, MDR, FDA 510(k), GDPR execution
- `../../c-level-agents/skills/gc-review/SKILL.md` — `/cs:gc-review` slash command
## References
- [contracts_playbook.md](references/contracts_playbook.md) — Standard contracts, clause checklist, common founder traps
- [ip_and_regulatory.md](references/ip_and_regulatory.md) — IP protection + regulatory landscape mapping
- [term_sheet_decoder.md](references/term_sheet_decoder.md) — Term sheet glossary + founder-friendly defaults + pushback strategies
---
**Version:** 1.0.0
**Status:** Production Ready
**Disclaimer:** Not legal advice. Always engage qualified counsel for binding decisions.
FILE:references/contracts_playbook.md
# Contracts Playbook — Standard Startup Agreements
Reference for the 7 contracts every startup signs in its first 5 years and the clause traps to avoid in each. **Not legal advice.** Bring redlines to qualified counsel.
## 1. Master Service Agreement (MSA) — Vendor Side (you signing theirs)
**What it is:** The umbrella contract for an ongoing relationship with a vendor (cloud, tooling, services, agencies). Usually paired with one or more SOWs / Order Forms.
**Top 5 redlines to push:**
1. **Auto-renewal:** Cut notice period to 30 days max. Reject 60/90/180 day notice.
2. **Liability cap:** Insist on 12 months of fees. Reject "fees in the preceding 3 months" (too narrow).
3. **Mutual indemnification:** Reject one-sided. Mirror the scope on both sides.
4. **IP ownership of deliverables:** All work product belongs to you. Vendor retains rights to pre-existing tools / methodologies, granted back to you for use.
5. **Data: DPA + return-or-destroy on termination.** Specifically: vendor cannot use your data to train AI models.
**Bonus catch:** Watch for "Vendor may modify these terms upon notice" — this means the contract you signed isn't the contract you have.
## 2. Customer SaaS Agreement (your paper)
**Standard structure:**
1. License grant (subscription, scope, term)
2. Acceptable use policy (what customer can/can't do)
3. Fees & payment (annual prepay vs. monthly, late fee, currency)
4. Service Level Agreement (uptime %, credits, exclusions)
5. Confidentiality (mutual, residuals carve-out)
6. Data Protection (DPA exhibit, subprocessor list, security commitments)
7. Warranties (limited, disclaim implied)
8. Indemnification (mutual, IP-infringement focused)
9. Limitation of liability (12 months fees, carve-outs for IP/data breach/willful)
10. Term & termination (term, termination for cause, termination for convenience)
**Founder traps when accepting customer redlines:**
- "Most-favored-nation" pricing (means you can never give anyone else a better deal).
- Uncapped liability for data breach with no minimum threshold.
- Customer right to perpetual license-back of "improvements" to your product.
- Customer "ownership" of any custom configuration (often hiding IP creep).
- Source-code escrow with auto-release triggers tied to customer convenience.
## 3. Non-Disclosure Agreement (NDA)
**One-way (you receiving):** Acceptable to sign without redlines for short evaluations.
**Mutual NDA (both directions):** The default for ongoing discussions.
**Critical carve-outs (always include):**
- **Residuals:** Information retained in unaided memory after end of engagement is not confidential.
- **Independent development:** If you build something similar without using their info, it's yours.
- **Public domain:** Information already public is not confidential.
- **Rightfully received:** Information received from a third party without confidentiality obligation.
- **Required by law:** Information disclosed under subpoena (with notice).
**Founder trap:** NDAs that prevent you from "engaging in similar business" — that's a non-compete in disguise. Strip it out.
## 4. Data Processing Agreement (DPA)
**Required when:** Personal data of EU residents flows (GDPR Article 28), or California residents (CCPA / CPRA), or HIPAA-covered data, or biometrics in IL/TX/WA (BIPA).
**Standard structure (GDPR-aligned):**
- Scope of processing (what data, what purpose)
- Controller / Processor designation
- Subprocessor list + flow-down obligations
- Data subject rights (access, deletion, portability)
- Security measures (encryption, access controls, training)
- Breach notification timelines (within 72 hours for GDPR)
- Audit rights (annual, reasonable)
- International transfer mechanism (SCCs, adequacy decision, BCRs)
- Return-or-destroy on termination
**Templates:** Use IAPP, EU Commission SCCs, or vendor-friendly DPA (e.g., Vanta's, Stripe's).
**Founder trap:** Missing DPA when EU/CA data flows = contract may be unenforceable AND regulatory fine exposure.
## 5. Employment Agreement / Offer Letter
**Must-have provisions:**
- **At-will employment** (US most states; not enforceable in MT for example)
- **Compensation:** salary, bonus structure, equity (option grant separately documented)
- **Invention assignment:** all IP created during employment using company resources belongs to company
- **Confidentiality:** ongoing duty, surviving termination
- **Non-solicit:** 12 months post-termination, employees + customers (carve out general advertising)
- **Non-compete:** state-dependent (CA, ND, OK, DC: void; many other states: enforceable if reasonable)
- **Arbitration:** mutual, AAA or JAMS rules, employer pays fees
**Founder traps:**
- Forgetting to require employees to sign **before** starting work (otherwise IP assignment is weak).
- Not including a "previously created inventions" exhibit (lets founders document pre-existing IP brought into the company).
- Skipping background checks for senior hires.
## 6. Contractor / 1099 Agreement
**Critical differences from employment:**
- **IP assignment is NOT automatic.** Without a written clause, the contractor owns what they create (under US law, "work for hire" applies only to specific categories of work).
- **Misclassification risk:** If a contractor functions like an employee (controlled hours, exclusive engagement, supplied equipment), tax authorities can reclassify, triggering back taxes + penalties.
- **No benefits, no withholding, contractor handles their own taxes.**
**Must-have provisions:**
- **Explicit work-for-hire OR written IP assignment** ("Contractor hereby assigns all right, title, and interest...").
- **Independent contractor status:** contractor controls means and methods.
- **Termination:** 30-day notice, immediate for cause.
- **Indemnification:** contractor indemnifies you for misclassification claims if they misrepresent status.
**Tooling:** Use Deel, Remote, or Velocity Global for international contractors to handle classification correctly.
## 7. Equity Agreements (Option Grants, Advisor Grants)
**Employee option grant:**
- **Strike price:** must be ≥ fair market value (FMV) at grant date (409A valuation, refreshed annually).
- **Vesting:** standard 4 years, 1 year cliff, monthly thereafter.
- **Exercise window post-termination:** 90 days standard; 7-10 years is founder-friendly.
- **ISO vs NSO:** ISOs have tax advantages (long-term capital gains if held) but limits ($100K vest/year) and US-citizen-only.
**Advisor grant (FAST template by Founder Institute):**
- 0.1% - 1% equity vested over 1-2 years, depending on level and stage.
- 2-year vesting, no cliff (advisors are tested through engagement, not retention).
- Single trigger acceleration on change of control (rare; double trigger more common).
**Founder trap:**
- Issuing options before completing the 409A valuation — strike price might be challenged by IRS.
- Verbal promises about acceleration — must be in writing.
- Forgetting to issue option grants to early employees within 90 days of hire (loses ISO eligibility).
## Quick Triage Heuristics
When you have 5 minutes to look at a contract:
1. **Find the liability cap.** No cap or > 24 months of fees = red flag.
2. **Find the indemnity clauses.** One-sided = red flag.
3. **Find the IP clause.** Vague or "as agreed" = red flag.
4. **Find the term + termination.** Auto-renewal with > 30 day notice = red flag.
5. **Find the choice of law/venue.** Exclusive in counterparty home jurisdiction = red flag.
Run `scripts/contract_risk_scanner.py` for the automated version.
---
**Final reminder:** This is a triage playbook. Every contract over $100K or longer than 1 year deserves outside counsel review. Every contract that touches personal data deserves a privacy attorney. Every term sheet deserves a securities / venture attorney. Period.
FILE:references/ip_and_regulatory.md
# IP Strategy & Regulatory Landscape
The two areas where startups most often discover legal exposure after it's too late to fix cheaply: IP ownership and regulatory triggers. **Not legal advice.**
## Part 1: IP Strategy
### IP Inventory — The Four Categories
| Type | What it protects | How you get it | How you lose it |
|---|---|---|---|
| **Patents** | Inventions (novel, non-obvious, useful) | File application | Public disclosure > 12 months before filing |
| **Copyright** | Original works of authorship (code, content, designs) | Automatic on fixation | Almost never; can be assigned away |
| **Trademark** | Brand identifiers (names, logos, slogans) | Use in commerce + registration | Not policing infringement; becoming generic |
| **Trade secret** | Confidential business information | Reasonable measures to keep secret | Public disclosure; failure to maintain confidentiality |
### Invention Assignment — The Single Most Important IP Practice
**Rule:** Every person who touches the company's product or systems must sign an invention assignment agreement **before** they start work.
This includes:
- Co-founders (often forgotten — usually fixed via founder restricted-stock purchase agreements)
- Employees (in employment agreement)
- Contractors (in contractor agreement; NOT automatic in US law)
- Interns (often forgotten — use a short standalone IP agreement)
- Advisors (in advisor agreement, scope limited to inventions related to company)
**Why it matters:** Without written assignment, the creator retains ownership. A contractor who built a critical service for 6 months and never signed an assignment can come back years later and demand a license — or assert that competitors can also use what they built.
**The "previously created inventions" exhibit:** Every IP assignment should include an exhibit where the signer lists pre-existing inventions they want to exclude. This protects everyone — the signer's prior work isn't accidentally assigned, and the company has documentation of what came in.
### Open Source License Compliance
**Permissive licenses** (MIT, Apache 2.0, BSD 2/3): Use freely, attribute, no copyleft.
**Weak copyleft** (LGPL, MPL): Can use in proprietary product; modifications to the OSS itself must be released. Distribution model matters.
**Strong copyleft** (GPL v2, GPL v3, AGPL): Distribution / SaaS use of a strong-copyleft component can require releasing your derivative work under the same license. **AGPL is the most aggressive** — it applies even when you only run the software on a server (SaaS / network use).
**Practice:**
1. Maintain an OSS inventory: `pip-licenses`, `license-checker` (npm), `cargo-license`, `go-licenses`.
2. Identify any GPL / AGPL / SSPL dependencies.
3. For each: either (a) comply with the license, (b) replace with a permissively-licensed alternative, or (c) document the carve-out (some companies build internally with GPL but only ship the binary externally — verify with counsel).
4. Run the inventory before any due diligence (acquisition, financing).
### Patents — When to File
**File when:**
- You have a genuinely novel technical invention (algorithm, hardware design, materials, biotech process).
- You face well-funded competitors who could copy without consequence.
- You're in a patent-dense industry (semiconductors, pharma, networking, medical devices).
- Filing strengthens fundraising / acquisition optics (limited weight for software-only startups).
**Don't bother when:**
- Your "invention" is a UX flow or business method (these are extremely hard to patent post-Alice Corp).
- You're in early stage with limited capital and no competitors close enough to copy.
- Defensive only and joining a patent pool (LOT Network, OIN) might be cheaper.
**Process:**
1. **Provisional patent** ($300-500 USPTO fee + $3K-5K attorney). 12 months to file non-provisional.
2. **Non-provisional / utility patent** ($1K USPTO fee + $10K-15K attorney + prosecution costs).
3. **PCT application** for international filings ($5K-10K).
4. **National phase entries** in each country you care about ($5K-15K per country).
Budget $25K-50K total for one well-prosecuted patent family with international coverage.
### Trade Secrets
**Reasonable measures required for legal protection:**
- NDA / confidentiality clauses with everyone who has access.
- Access controls (need-to-know basis, not company-wide).
- Marking documents "Confidential."
- Departure procedures (return of materials, exit interview, deactivation).
- Training employees on what's a trade secret.
**Without these measures, the information may not qualify for trade secret protection if disclosed — even by a thief.**
**Common trade secrets:**
- Customer lists with usage / pricing data
- Algorithms not disclosed in published patents
- Manufacturing processes
- Sales playbooks and pricing models
- Internal financial projections
- Source code (unless OSS)
### Trademark Strategy
**Search before launch:**
- USPTO TESS search (free, but limited; doesn't catch common-law marks).
- Professional search via attorney ($500-2K) catches common-law marks and similar-mark conflicts.
- International searches via WIPO Global Brand Database.
**Register early:**
- US: Intent-to-use application (1B) lets you reserve a mark before launch.
- International: Madrid Protocol filing extends to 100+ countries.
- Word marks first (the brand name itself), design marks second (logos).
**Policing:**
- Set up Google Alerts and USPTO TMNG for your mark.
- Send cease-and-desist letters promptly; failure to police can weaken the mark.
---
## Part 2: Regulatory Landscape — When to Engage Counsel
The startups that survive their first regulatory encounter engage specialist counsel **before** building, not after. The ones that don't usually pivot, retreat, or pay heavy fines.
### Trigger Matrix
| Trigger | Regulatory Regime | Specialist Needed | Earliest Action |
|---|---|---|---|
| Healthcare data (patient records, claims, PHI) | HIPAA, HITECH, state breach laws | Health-tech attorney | Business Associate Agreement, OCR-aligned risk assessment |
| Cardholder data | PCI DSS (industry standard; contractually required) | QSA + counsel | Scope reduction, tokenization, certified processor |
| Money movement (transmitting funds, custody, crypto) | BSA/AML, state money-transmitter (50-state patchwork) | Fintech attorney | Stripe Treasury / Banking as a Service to avoid MT registration |
| Lending | Truth in Lending Act, state usury laws, ECOA | Fintech / consumer-finance attorney | Bank partnership, state licensing analysis |
| Medical device claims | FDA 510(k), De Novo, PMA; EU MDR; ISO 13485 | Medical-device regulatory specialist | Pre-submission meeting with FDA |
| EU residents' personal data | GDPR + ePrivacy + EU AI Act if AI | EU privacy attorney | DPA, SCCs for international transfer, DPIA |
| California residents | CCPA / CPRA | Privacy generalist | Privacy notice, opt-out mechanisms, vendor management |
| Children's data (under 13 US, under 16 in some EU states) | COPPA, GDPR-K | Privacy attorney | Parental consent, no-track defaults |
| Securities (tokens, equity crowdfunding, advisory boards) | SEC rules (Reg D, Reg A+, Reg CF, Howey test) | Securities attorney | Token sale legal opinion, Form D filing |
| Defense / aerospace customers | ITAR, EAR, DFARS, CMMC | Export-control attorney | Export classification, registered with State Dept |
| AI in EU | EU AI Act (risk-tiered: prohibited / high-risk / limited / minimal) | EU privacy + product attorney | Risk assessment, conformity assessment for high-risk |
| AI for hiring | NYC Local Law 144, CO SB 21-169, IL HB 53 | Employment attorney | Bias audit, candidate notice |
| Telehealth / online prescribing | State medical board rules, DEA registration for controlled substances | Telehealth specialist | State-by-state physician licensing strategy |
| Insurance (sale, underwriting, brokerage) | State insurance commissioners | Insurance attorney | State licensing, agency agreement |
### Sequencing: SOC 2 → ISO 27001 → Industry-Specific
For most B2B SaaS, the security/compliance sequence is:
1. **SOC 2 Type 1** (point-in-time audit) — ~$15K-25K, 3-6 months prep
2. **SOC 2 Type 2** (continuous, ~6-12 month audit window) — ~$25K-50K
3. **ISO 27001** if expanding internationally — ~$30K-60K, builds on SOC 2 controls
4. **ISO 42001** if AI is core to product — first AI management system standard
5. **Industry overlays:** HIPAA technical safeguards, FedRAMP (federal customers), PCI DSS (cardholder data)
**Sequencing logic:** SOC 2 unlocks the majority of enterprise sales. ISO 27001 unlocks European and Asia-Pacific. Industry overlays are required for specific verticals.
### When to Get a General Counsel Hire
| Stage | GC need |
|---|---|
| Pre-seed / seed | None. Use outside counsel ad-hoc + Clerky/Stripe Atlas templates |
| Series A | Fractional GC (~$10-20K/month) OR senior associate at firm |
| Series B | Full-time GC if regulated industry, customer contracts are heavy, or fundraising is constant |
| Series C+ | Full-time GC + Deputy/Associate GC if international |
**Signs you need a GC hire:**
- You're spending > $200K/year on outside counsel
- You're signing > 1 enterprise contract per week with customer redlines
- You're in a regulated industry (healthcare, fintech, defense)
- You're preparing for IPO or going-public transaction
- You're acquiring companies
### Cross-Border Considerations
**Hiring international employees:**
- Use Deel / Remote / Velocity Global for first 1-5 contractors per country.
- Establish an entity (subsidiary or EOR-to-entity transition) at 5-10+ employees.
- Tax residency, permanent establishment risk, and equity grants vary significantly.
**International data flows:**
- EU → US: SCCs + Transfer Impact Assessment (TIA); DPF if certified.
- China → outbound: PIPL approval + standard contract + security assessment.
- UK → outside: UK SCCs (similar to EU).
- Schrems / DPF status changes regularly — monitor with privacy counsel.
**International IP:**
- Patent: PCT application within 12 months of first national filing.
- Trademark: Madrid Protocol for multi-country filings.
- Copyright: Berne Convention covers most countries automatically.
---
## Closing: The General Counsel's Three Rules
1. **Get it in writing.** Verbal agreements and "we'll figure it out later" produce 80% of post-engagement disputes.
2. **Identify the regulatory trigger before you build.** It's 10x cheaper to design around a regulation than to retrofit.
3. **Always have outside counsel review anything binding.** This document is triage; real legal review is mandatory.
FILE:references/term_sheet_decoder.md
# Term Sheet Decoder
Glossary + founder-friendly defaults + pushback strategies for every clause in a standard venture term sheet. **Not legal advice.** Always engage venture / securities counsel before responding.
## The Three Clauses That Matter Most
In any term sheet review, focus disproportionately on these three. They drive ~80% of the founder economics impact.
### 1. Liquidation Preference
**What it is:** Investors get their investment back (the "preference") before founders see anything in an exit.
**The dimensions:**
- **Multiple:** 1x (standard) means $1 back per $1 invested. 2x means $2 back. Higher = more hostile.
- **Participating vs Non-participating:**
- **Non-participating (founder-friendly):** Investor chooses preference OR convert to common at exit. Most exits hit the conversion threshold, so preference is effectively just downside protection.
- **Participating ("double-dip"):** Investor gets preference back AND a pro-rata share of remaining proceeds as if converted. Significantly increases investor take in mid-range exits.
- **Cap:** Caps the total return at, say, 2x or 3x of investment for participating preferences. Limits the double-dip.
**Standard (Series A/B):** 1x non-participating.
**Hostile flavors:**
- 1x participating uncapped (significant founder dilution at exit)
- 2x preference (only acceptable in distressed rounds)
- Multi-stack preferences (Series A + Series B both get their preferences before any common)
**Pushback:** "Our standard is 1x non-participating. Participating preferences create misalignment with management at exit."
### 2. Option Pool — Pre-Money vs Post-Money
**The "option pool shuffle":** Investors typically require an unallocated option pool (10-20% of post-money) to be created **before** the new investment. If this comes out of pre-money, founders are diluted; if post-money, all shareholders dilute proportionally.
**Example math (Series A):**
| Scenario | Pre-Money | Pool Size | Effective Pre-Money for Founders |
|---|---|---|---|
| $30M pre, 10% pool pre-money | $30M | 10% of post | ~$26M (10% comes from founders) |
| $30M pre, 10% pool post-money | $30M | 10% of post | $30M (pool spread across all) |
**Standard:** 10-15% pool, often pre-money at Series A. Founder-friendly: smaller pool or post-money.
**Pushback:** "We've modeled our hiring plan and 8% supports the next 18 months. Let's right-size to actual need, not standard percentage." Or: "Pool top-up should come out of post-money so the new investor shares the dilution."
### 3. Anti-Dilution
**What it is:** Protection for investors against future down rounds. If a later round prices below the current, the current investor's price is adjusted retroactively.
**Flavors (least to most hostile):**
- **None:** Rare; only in seed SAFEs sometimes.
- **Broad-based weighted average (standard):** Adjusts using all shares (common, options, warrants). Modest founder dilution in a down round.
- **Narrow-based weighted average:** Uses only preferred. More dilutive than broad-based.
- **Full ratchet (hostile):** Investor's price resets entirely to the new round's price. Massively dilutive to founders.
**Standard:** Broad-based weighted average.
**Pushback:** "Full ratchet is non-starter at this stage. Narrow-based is unusual. We need broad-based weighted average — this is the NVCA standard."
---
## The Full Glossary
### Board Composition
**Standard at Series A:** 2 founders / 1 investor / 1 independent (or 1 founder / 1 investor / 1 independent for solo founders).
**At Series B:** Often 2 / 2 / 1 (balanced with independent tie-breaker).
**At Series C+:** Often investors get majority (signals control transition).
**Founder protection:** Always insist on the independent seat. Independent directors prevent deadlock and provide a neutral voice.
**Pushback on investor-majority boards at A:** "Investor control of the board at Series A is premature. Let's keep founder control with an independent tie-breaker until Series B."
### Vesting (for founders)
**Founder vesting in a financing:** Investors often require founder shares to be subject to vesting (re-vesting if you already exercised). Standard: 4 years, 1-year cliff. Often the cliff is waived if you've been at the company > 1 year.
**Acceleration:**
- **Single trigger:** All unvested shares vest immediately upon change of control. Founder-friendly but rare; investors resist.
- **Double trigger (standard):** Acceleration requires (a) change of control AND (b) involuntary termination of the founder within X months. Industry standard at Series A+.
**Pushback:** "Double-trigger acceleration is industry standard. Without it, founders are exposed to acquirer post-acquisition staffing decisions."
### Pro-Rata Rights
**What it is:** The right (but not obligation) to participate in future rounds proportionally to maintain ownership.
**Standard:** Lead investor + major investors (typically those above some ownership threshold) get pro-rata. Smaller checks often don't.
**Founder impact:** Granting pro-rata is generally fine — it shows investor conviction and aligns long-term. The cost is small dilution in future rounds.
**Pushback:** Only push back if there's a long tail of small investors each demanding pro-rata; cap to "major investors" defined by ownership %.
### Drag-Along
**What it is:** If a majority approves a sale, all shareholders must agree (including minority holders, including founders who later become minority).
**Founder-friendly version:** Drag-along requires founder consent OR a minimum sale price threshold (e.g., > 3x liquidation preference).
**Hostile version:** Drag-along with no founder consent and no price floor. Investors can force a sale at any price over founder objection.
**Pushback:** "Drag-along is standard, but we need founder consent OR a price floor."
### Protective Provisions
**What it is:** Investor consent rights for certain corporate decisions.
**Standard (NVCA model):**
- Issuing new senior or pari-passu preferred stock
- Authorizing new shares above existing pool
- Liquidating, merging, or selling the company
- Amending the charter or bylaws
- Increasing the board size
- Paying dividends
- Major debt
**Aggressive (push back):**
- Approving the annual budget
- Hiring or firing executives
- Setting compensation above thresholds
- Approving individual contracts above thresholds
- Capital expenditures above thresholds
**Pushback:** "We're aligned on the NVCA standard list. Operating decisions like budget and hiring are management's responsibility — protective provisions are for fundamental corporate changes."
### Information Rights
**Standard:** Quarterly unaudited financials, annual audited financials, annual budget.
**Aggressive (push back):** Monthly financials, board observer rights, weekly KPI dashboards, inspection rights at will.
**Pushback:** "Standard quarterly + annual is enough. Monthly creates significant CFO overhead at our stage. We'll commit to ad-hoc updates on material events."
### Dividends
**Standard:** None (default).
**Acceptable:** Non-cumulative dividends "when and if declared by the board" — almost never paid in practice.
**Hostile:** Cumulative dividends accrue every year regardless of declaration and must be paid in cash at exit. This is a creeping liquidation preference.
**Pushback:** "Cumulative dividends create a hidden liquidation preference that accrues over time. Non-cumulative when-declared, or none, is standard."
### Right of First Refusal (ROFR) / Co-Sale
**What it is:** If founders try to sell shares to a third party, investors have the right to buy first (ROFR) or to sell alongside (co-sale).
**Founder-friendly:** Standard ROFR + co-sale for all preferred; founders can still do secondary up to small thresholds without triggering.
**Hostile:** No secondary at all without unanimous investor consent.
**Pushback:** "We need to allow modest founder secondary (e.g., up to $1M aggregate) without investor consent — this is needed for founder financial planning."
### Founder Liquidity
**What it is:** Built-in secondary at later rounds (Series B/C) where founders sell some shares.
**Standard:** Becoming more common; 10-20% of round size as founder secondary.
**Pushback:** Raise this in Series B+ discussions; not typically negotiated at Series A.
### Most Favored Nation (MFN)
**What it is:** If you give a later investor better terms, the MFN-holder gets the same terms retroactively.
**Common in:** Seed SAFEs and convertible notes; rare in priced rounds.
**Founder trap:** MFN provisions can prevent you from offering competitive terms to new lead investors later. Be specific about what's covered (just SAFE terms? all terms?).
### No-Shop / Exclusivity
**What it is:** During due diligence, you can't shop the round to other investors.
**Standard:** 30-45 days. Founder-friendly. Investor-aligned because it shows commitment.
**Pushback only if:** > 60 days, or if it extends post-execution of definitive docs.
---
## Founder-Friendly Defaults (Cheat Sheet)
| Clause | Founder-Friendly Default |
|---|---|
| Liquidation preference | 1x non-participating |
| Anti-dilution | Broad-based weighted average |
| Option pool | 8-12%, post-money |
| Board (Series A) | 2F / 1I / 1Indep |
| Vesting (founder re-vest) | 4yr / 1yr cliff, often with credit for time served |
| Acceleration | Double-trigger |
| Pro-rata | For lead + major investors |
| Drag-along | Requires founder consent or price floor |
| Protective provisions | NVCA standard list only |
| Information rights | Quarterly + annual + budget |
| Dividends | None or non-cumulative when-declared |
| ROFR / co-sale | Standard, with carve-out for modest founder secondary |
| MFN (in notes/SAFEs) | Avoid if possible; if not, narrow scope |
| No-shop | 30-45 days |
---
## Negotiation Strategy
**Pick your battles:** A term sheet has 25-40 clauses. Winning every one is impossible and signals you don't understand priorities.
**Focus on the top 3 mistakes (in order):**
1. Liquidation preference flavor (participating vs non-participating)
2. Option pool pre-money vs post-money + size
3. Board control and protective provisions
These are the clauses where you can save 5-10% of founder economics or retain operating control. Everything else is secondary.
**The "founder-friendly NVCA" framing:** Many investors signal their posture by deviating from the NVCA model (the industry standard documents published by the National Venture Capital Association). Pushing back to "let's use the NVCA standard" is rarely rejected and resolves most issues.
**Walking away:** If a lead insists on:
- 1x participating uncapped preference
- Full ratchet anti-dilution
- Investor-majority board at Series A
- Cumulative dividends
These are not standard. A founder-friendly lead doesn't insist on these. Either walk or get specific written justification (sometimes a distressed cap-table situation justifies one of them, but never all).
---
## After Signing
Once the term sheet is signed:
1. **No-shop is active.** Don't talk to other investors except to officially decline.
2. **Definitive documents (SPA, IRA, Voting Agreement, ROFR Agreement) take 4-6 weeks.** Don't lose energy here; main fight was the term sheet.
3. **Closing conditions:** legal opinion, secretary's certificate, charter filing, capitalization confirmation.
4. **Wire timing:** Investors often wire 1-3 days after charter filing. Plan accordingly.
Run `scripts/term_sheet_analyzer.py` on the structured JSON of the term sheet for an automated scoring + flag analysis.
---
**Final reminder:** This document is a decoder, not a negotiation manual. Real term sheet response always involves your venture / securities counsel + your lead investor's diligence + your board (if any). Use this as a primer before those conversations.
FILE:scripts/contract_risk_scanner.py
#!/usr/bin/env python3
"""contract_risk_scanner.py — Scan a contract for founder-killer clauses.
Stdlib-only. Outputs human-readable or JSON. Detects 12 common risk patterns:
1. Unilateral termination favoring the counterparty
2. Auto-renewal with long notice (60+ days)
3. Uncapped liability or exclusion of standard caps
4. Broad indemnification flowing one direction
5. Non-mutual confidentiality
6. Missing or vague IP ownership clauses
7. Aggressive non-compete / non-solicit
8. Choice of law/venue in counterparty's home jurisdiction (one-sided)
9. Force majeure favoring only the counterparty
10. Missing DPA reference when personal data flows
11. Most-favored-nation pricing clauses
12. Audit rights without reciprocity
NOT legal advice. Use this to triage; bring findings to qualified counsel.
Usage:
python contract_risk_scanner.py # uses embedded sample
python contract_risk_scanner.py path/to/contract.txt
python contract_risk_scanner.py contract.txt --output json
python contract_risk_scanner.py --help
"""
import argparse
import json
import re
import sys
from dataclasses import dataclass, asdict
from typing import List
SAMPLE_CONTRACT = """\
MASTER SERVICES AGREEMENT
This Agreement shall automatically renew for successive one (1) year terms
unless either party provides ninety (90) days written notice of non-renewal.
LIMITATION OF LIABILITY. In no event shall Provider's aggregate liability
arising out of this Agreement exceed the fees paid by Customer in the
twelve (12) months preceding the claim. Notwithstanding the foregoing,
Customer's indemnification obligations under Section 8 shall be uncapped.
INDEMNIFICATION. Customer shall defend, indemnify and hold harmless
Provider, its affiliates, officers, directors and employees from and against
any and all claims, damages, losses and expenses arising out of or relating
to Customer's use of the Services.
INTELLECTUAL PROPERTY. The parties agree that intellectual property created
during the engagement shall belong to the party who develops it.
NON-COMPETE. For a period of three (3) years following termination, Customer
shall not engage with any competitor of Provider in any capacity, in any
geography.
GOVERNING LAW. This Agreement shall be governed by the laws of Delaware,
and any disputes shall be resolved exclusively in the state and federal
courts located in Wilmington, Delaware.
FORCE MAJEURE. Provider shall not be liable for any failure to perform due
to causes beyond its reasonable control.
"""
@dataclass
class Finding:
rule_id: str
severity: str # CRITICAL | HIGH | MEDIUM | LOW
title: str
excerpt: str
why_it_matters: str
suggested_redline: str
RULES = [
{
"id": "AUTO_RENEW_LONG_NOTICE",
"severity": "HIGH",
"title": "Auto-renewal with long notice period",
"pattern": re.compile(
r"automatically renew.{0,200}?(\d+|sixty|ninety|one hundred|180)\s*(\(\d+\))?\s*day",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Auto-renewal with >30 day notice is a classic vendor trap: founders forget the "
"deadline and get locked into another full term. Especially painful on multi-year contracts."
),
"redline": (
"Counter: '...unless either party provides thirty (30) days written notice of non-renewal' "
"OR remove auto-renewal entirely and require affirmative re-signature."
),
},
{
"id": "UNCAPPED_CUSTOMER_INDEMNITY",
"severity": "CRITICAL",
"title": "Customer indemnity carved out from liability cap (uncapped)",
"pattern": re.compile(
r"(customer'?s|your)\s+indemnification.{0,200}?(uncapped|shall be uncapped|excluded from)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Uncapped customer indemnity means a single bad claim can exceed all fees ever paid. "
"Standard practice: mutual indemnity, both sides capped at fees, with narrow carve-outs "
"(IP infringement, data breach, gross negligence)."
),
"redline": (
"Counter: cap customer indemnity at 12 months of fees, mutual indemnity, carve-outs only "
"for willful misconduct and breach of confidentiality."
),
},
{
"id": "ONE_SIDED_INDEMNITY",
"severity": "HIGH",
"title": "Indemnification flows in one direction only",
"pattern": re.compile(
r"(customer|client)\s+shall\s+(defend|indemnify).{0,500}?(provider|company|vendor)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"One-sided indemnity means you take on risk for the counterparty's actions without reciprocity. "
"A balanced contract has mutual indemnification with mirrored carve-outs."
),
"redline": (
"Counter: 'Each party shall defend, indemnify and hold harmless the other party...' with "
"mirrored scope and equal caps."
),
},
{
"id": "VAGUE_IP",
"severity": "CRITICAL",
"title": "Vague IP ownership clause",
"pattern": re.compile(
r"intellectual property.{0,200}?(belong to the party who develops it|jointly owned|to be determined|as agreed)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Vague IP language is the #1 source of post-engagement disputes. Joint ownership often means "
"neither party can license freely without the other's consent. 'As agreed' is unenforceable."
),
"redline": (
"Counter: 'All work product, deliverables, and derivative works created under this Agreement "
"shall be the sole and exclusive property of Customer. Provider hereby assigns all right, title "
"and interest...' Or explicitly carve out Provider's pre-existing IP and tools with a license back."
),
},
{
"id": "AGGRESSIVE_NONCOMPETE",
"severity": "HIGH",
"title": "Aggressive non-compete (long duration or broad geography)",
"pattern": re.compile(
r"non.compete.{0,300}?(two|three|four|five|2|3|4|5)\s*\(?\d*\)?\s*year",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Non-competes >12 months or with unbounded geography are often unenforceable (especially in "
"California, and increasingly federally) but create chilling effects. They also signal the "
"counterparty's overall negotiation posture."
),
"redline": (
"Counter: maximum 12 months, specific competitor list (not 'any competitor'), specific "
"geography. For California-resident counterparties, remove entirely (California labor code "
"voids most non-competes)."
),
},
{
"id": "ONE_SIDED_VENUE",
"severity": "MEDIUM",
"title": "Choice of law/venue exclusively in counterparty jurisdiction",
"pattern": re.compile(
r"(exclusively in|exclusive jurisdiction).{0,300}?(courts? located in|state and federal courts of)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Exclusive venue in counterparty's jurisdiction means you bear travel cost and out-of-state "
"counsel cost for any dispute. For startups this can effectively prevent enforcement."
),
"redline": (
"Counter: neutral venue (Delaware is common), or 'venue in the jurisdiction of the defendant' "
"(forces plaintiff to travel), or arbitration in a neutral location with AAA/JAMS rules."
),
},
{
"id": "ONE_SIDED_FORCE_MAJEURE",
"severity": "MEDIUM",
"title": "Force majeure clause favors one party",
"pattern": re.compile(
r"(provider|company|vendor)\s+shall not be liable.{0,200}?(force majeure|causes beyond)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"If only the vendor gets force-majeure protection, you pay full price during a pandemic / "
"outage / supply chain disruption but receive nothing. Mutual force majeure is standard."
),
"redline": (
"Counter: 'Neither party shall be liable...' with explicit list of qualifying events "
"(pandemic, war, natural disaster, government action) and a termination right after 30 days."
),
},
{
"id": "MISSING_DPA",
"severity": "HIGH",
"title": "Personal data appears to flow but no DPA referenced",
"pattern": re.compile(
r"(personal data|personally identifiable|user data|customer data|PII)(?!.{0,500}(DPA|data processing agreement|GDPR))",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"If personal data of EU residents (or California residents) flows, a DPA is legally required. "
"Missing DPA = GDPR Article 28 violation, potential 4%-of-revenue fine, contract unenforceable "
"with EU customers."
),
"redline": (
"Counter: 'The parties shall execute a Data Processing Agreement substantially in the form "
"of Exhibit X prior to any processing of Personal Data.' Use IAPP or Vendor-friendly DPA template."
),
},
{
"id": "MOST_FAVORED_NATION",
"severity": "MEDIUM",
"title": "Most-favored-nation (MFN) pricing clause",
"pattern": re.compile(
r"(most.favored.nation|MFN|best price|lowest price).{0,200}?(offered to|charged to)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"MFN clauses prevent you from offering volume discounts or strategic pricing to anyone else. "
"If you sign with one customer, every future customer can demand the same price."
),
"redline": (
"Counter: remove the MFN entirely. If kept, narrow to 'similarly situated customers, same "
"tier and volume, excluding strategic / launch / migration discounts.'"
),
},
{
"id": "ONE_SIDED_AUDIT",
"severity": "MEDIUM",
"title": "Audit rights without reciprocity",
"pattern": re.compile(
r"(customer|client).{0,100}?right to audit",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"One-sided audit rights mean the counterparty can demand records on demand, often at your "
"expense. Reciprocity is standard for B2B agreements."
),
"redline": (
"Counter: mutual audit rights, max once per year, at requesting party's expense, with "
"30-day notice, during business hours, narrowed to specific compliance categories."
),
},
{
"id": "BROAD_NON_SOLICIT",
"severity": "MEDIUM",
"title": "Broad non-solicit (employees AND customers, long duration)",
"pattern": re.compile(
r"non.solicit.{0,300}?(employees? and customers?|customers? and employees?)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"Combined employee + customer non-solicits, especially with long duration, can severely "
"limit hiring and business development. Many states limit enforceability."
),
"redline": (
"Counter: split into employee-only (12 months max) and customer-only (12 months max) clauses, "
"with carve-outs for general advertising / open job postings and for customers who initiate "
"contact independently."
),
},
{
"id": "PERPETUAL_LICENSE_BACK",
"severity": "HIGH",
"title": "Perpetual license-back to counterparty of your data or work",
"pattern": re.compile(
r"perpetual.{0,100}?(license|right).{0,300}?(customer data|user data|work product|deliverables)",
re.IGNORECASE | re.DOTALL,
),
"why_it_matters": (
"A perpetual license-back lets the counterparty use your data or deliverables forever, even "
"after termination. This is acceptable for usage analytics, NOT for customer data or core IP."
),
"redline": (
"Counter: time-limited license (for the term of the agreement only), specific purpose "
"(service delivery only, not training AI models, not sharing with third parties), and "
"post-termination return-or-destroy obligation."
),
},
]
def scan(text: str) -> List[Finding]:
findings: List[Finding] = []
for rule in RULES:
for match in rule["pattern"].finditer(text):
excerpt = match.group(0).strip()
# truncate long excerpts
if len(excerpt) > 300:
excerpt = excerpt[:297] + "..."
findings.append(Finding(
rule_id=rule["id"],
severity=rule["severity"],
title=rule["title"],
excerpt=excerpt,
why_it_matters=rule["why_it_matters"],
suggested_redline=rule["redline"],
))
# rank by severity then rule order
severity_order = {"CRITICAL": 0, "HIGH": 1, "MEDIUM": 2, "LOW": 3}
findings.sort(key=lambda f: (severity_order.get(f.severity, 9), f.rule_id))
return findings
def render_text(findings: List[Finding], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("CONTRACT RISK SCAN")
lines.append(f"Source: {source}")
lines.append(f"Findings: {len(findings)}")
lines.append("=" * 72)
lines.append("")
if not findings:
lines.append("No risk patterns matched. (Absence of findings does not mean the contract is safe;")
lines.append("it means the 12 common patterns this scanner checks did not trigger.)")
lines.append("")
lines.append("Always engage qualified counsel before signing.")
return "\n".join(lines)
severity_counts = {}
for f in findings:
severity_counts[f.severity] = severity_counts.get(f.severity, 0) + 1
severity_summary = " ".join(
f"{sev}: {severity_counts.get(sev, 0)}"
for sev in ("CRITICAL", "HIGH", "MEDIUM", "LOW")
if severity_counts.get(sev, 0) > 0
)
lines.append(f"Severity: {severity_summary}")
lines.append("")
for i, f in enumerate(findings, 1):
lines.append(f"[{i}] {f.severity} — {f.title}")
lines.append(f" Rule: {f.rule_id}")
lines.append(f" Excerpt: \"{f.excerpt}\"")
lines.append("")
lines.append(f" Why it matters:")
for line in _wrap(f.why_it_matters, 4):
lines.append(line)
lines.append("")
lines.append(f" Suggested redline:")
for line in _wrap(f.suggested_redline, 4):
lines.append(line)
lines.append("")
lines.append("-" * 72)
lines.append("")
lines.append("REMINDER: This scanner triages obvious traps. Always bring redlines to qualified counsel.")
return "\n".join(lines)
def _wrap(text: str, indent: int, width: int = 68) -> List[str]:
import textwrap
return textwrap.wrap(text, width=width, initial_indent=" " * indent, subsequent_indent=" " * indent) or [" " * indent + text]
def main() -> int:
parser = argparse.ArgumentParser(
description="Scan a contract for the 12 most common founder-killer clauses.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to contract text file (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
text = f.read()
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
else:
text = SAMPLE_CONTRACT
source = "<embedded sample MSA>"
findings = scan(text)
if args.output == "json":
payload = {
"source": source,
"findings_count": len(findings),
"findings": [asdict(f) for f in findings],
}
print(json.dumps(payload, indent=2))
else:
print(render_text(findings, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/term_sheet_analyzer.py
#!/usr/bin/env python3
"""term_sheet_analyzer.py — Score a term sheet on founder-friendliness.
Stdlib-only. Computes a 0-100 score across 12 dimensions and flags
hostile clauses. Outputs human-readable or JSON.
NOT legal advice — surfaces questions for venture / securities counsel.
Input schema (JSON):
{
"round": "Series A",
"pre_money": 30000000,
"raise_amount": 8000000,
"liquidation_preference": {
"multiple": 1.0,
"participating": false,
"cap": null
},
"anti_dilution": "broad_based_weighted_average", // | "narrow_based_weighted_average" | "full_ratchet" | "none"
"option_pool": {
"size_pct": 12.0,
"pre_money": true
},
"board_composition": {
"investor_seats": 1,
"founder_seats": 2,
"independent_seats": 1
},
"vesting": {
"standard_years": 4,
"cliff_months": 12,
"single_trigger_acceleration": false,
"double_trigger_acceleration": true
},
"pro_rata": true,
"drag_along": {
"exists": true,
"founder_consent_required": true
},
"protective_provisions": "standard", // | "standard" | "aggressive"
"information_rights": "standard", // | "standard" | "aggressive"
"dividends": "none" // | "none" | "non_cumulative_when_declared" | "cumulative"
}
Usage:
python term_sheet_analyzer.py # uses embedded sample
python term_sheet_analyzer.py path/to/term_sheet.json
python term_sheet_analyzer.py term_sheet.json --output json
python term_sheet_analyzer.py --help
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Tuple
SAMPLE = {
"round": "Series A",
"pre_money": 30_000_000,
"raise_amount": 8_000_000,
"liquidation_preference": {"multiple": 1.0, "participating": False, "cap": None},
"anti_dilution": "broad_based_weighted_average",
"option_pool": {"size_pct": 12.0, "pre_money": True},
"board_composition": {"investor_seats": 1, "founder_seats": 2, "independent_seats": 1},
"vesting": {
"standard_years": 4,
"cliff_months": 12,
"single_trigger_acceleration": False,
"double_trigger_acceleration": True,
},
"pro_rata": True,
"drag_along": {"exists": True, "founder_consent_required": True},
"protective_provisions": "standard",
"information_rights": "standard",
"dividends": "none",
}
def score(ts: Dict[str, Any]) -> Tuple[int, List[Dict[str, Any]]]:
"""Returns (total_score_0_to_100, list_of_findings).
Each dimension is scored 0-100, then averaged. Findings list contains
per-clause analysis with severity.
"""
findings: List[Dict[str, Any]] = []
scores: List[int] = []
# --- 1. Liquidation Preference (high signal) ---
lp = ts.get("liquidation_preference", {})
lp_mult = lp.get("multiple", 1.0)
lp_part = lp.get("participating", False)
lp_cap = lp.get("cap")
if lp_mult == 1.0 and not lp_part:
lp_score = 100
findings.append(_ok("liquidation_preference", "1x non-participating — founder-friendly standard."))
elif lp_mult == 1.0 and lp_part and lp_cap and lp_cap <= 3:
lp_score = 55
findings.append(_warn("liquidation_preference",
f"1x participating with {lp_cap}x cap. Investor double-dips up to cap. "
"Push for non-participating; if accepted, accept cap < 3x."))
elif lp_mult == 1.0 and lp_part and not lp_cap:
lp_score = 25
findings.append(_crit("liquidation_preference",
"1x PARTICIPATING UNCAPPED. Investor gets their money back AND a pro-rata share of remaining proceeds, "
"forever. Hostile. Push to non-participating or at minimum cap at 2x."))
elif lp_mult > 1.0:
lp_score = 10
findings.append(_crit("liquidation_preference",
f"{lp_mult}x preference. Investor gets {lp_mult}x their money back before founders see a dollar. "
"Hostile; only acceptable in distressed rounds."))
else:
lp_score = 80
findings.append(_ok("liquidation_preference", f"{lp_mult}x configuration acceptable."))
scores.append(lp_score)
# --- 2. Anti-Dilution ---
ad = ts.get("anti_dilution", "broad_based_weighted_average")
if ad == "broad_based_weighted_average":
ad_score = 100
findings.append(_ok("anti_dilution", "Broad-based weighted average — founder-friendly standard."))
elif ad == "narrow_based_weighted_average":
ad_score = 70
findings.append(_warn("anti_dilution",
"Narrow-based weighted average. More dilutive to founders than broad-based in a down round. "
"Push to broad-based."))
elif ad == "full_ratchet":
ad_score = 10
findings.append(_crit("anti_dilution",
"FULL RATCHET. In a down round, investor's price is reset to the new round price entirely, "
"massively diluting founders. Hostile; reject."))
elif ad == "none":
ad_score = 100
findings.append(_ok("anti_dilution", "No anti-dilution provision. Unusual but founder-friendly."))
else:
ad_score = 50
findings.append(_warn("anti_dilution", f"Unrecognized anti-dilution type: {ad}. Verify with counsel."))
scores.append(ad_score)
# --- 3. Option Pool (pre-money vs post-money) ---
op = ts.get("option_pool", {})
op_pre = op.get("pre_money", True)
op_size = op.get("size_pct", 10.0)
if not op_pre:
op_score = 100
findings.append(_ok("option_pool",
f"Pool of {op_size}% sits post-money — dilutes all shareholders proportionally."))
elif op_pre and op_size <= 10.0:
op_score = 70
findings.append(_warn("option_pool",
f"Pool of {op_size}% pre-money — comes out of founders' shares. Reasonable size, but consider "
"negotiating post-money or sharing the pool top-up across the round."))
elif op_pre and op_size > 10.0:
op_score = 30
findings.append(_crit("option_pool",
f"Pool of {op_size}% PRE-MONEY. This is the 'option pool shuffle' — typically reduces pre-money "
f"by ~{op_size}%, diluting founders silently. Negotiate hard: justify the size with a hiring plan "
"or push for post-money."))
else:
op_score = 60
findings.append(_warn("option_pool", "Option pool structure unclear; verify."))
scores.append(op_score)
# --- 4. Board Composition ---
bc = ts.get("board_composition", {})
inv = bc.get("investor_seats", 0)
fnd = bc.get("founder_seats", 0)
ind = bc.get("independent_seats", 0)
total = inv + fnd + ind
if total == 0:
bc_score = 50
findings.append(_warn("board_composition", "Board composition unspecified."))
elif fnd > inv and ind >= 1:
bc_score = 100
findings.append(_ok("board_composition",
f"{fnd} founder / {inv} investor / {ind} independent — founder-friendly; founders retain control "
"with independent tie-breaker."))
elif fnd == inv and ind >= 1:
bc_score = 75
findings.append(_ok("board_composition",
f"{fnd} founder / {inv} investor / {ind} independent — balanced, independent is critical."))
elif inv > fnd:
bc_score = 30
findings.append(_crit("board_composition",
f"{fnd} founder / {inv} investor / {ind} independent — investors control the board at Series A. "
"This is unusually early; investor control typically arrives at Series B or later."))
else:
bc_score = 50
findings.append(_warn("board_composition", f"Composition: {fnd}F/{inv}I/{ind}Ind — verify with counsel."))
scores.append(bc_score)
# --- 5. Vesting & Acceleration ---
vest = ts.get("vesting", {})
years = vest.get("standard_years", 4)
cliff = vest.get("cliff_months", 12)
single = vest.get("single_trigger_acceleration", False)
double = vest.get("double_trigger_acceleration", False)
if years == 4 and cliff == 12 and double and not single:
vest_score = 100
findings.append(_ok("vesting",
"4yr/1yr cliff with double-trigger acceleration — founder-friendly standard. "
"Single-trigger is rare and not recommended by counsel."))
elif years == 4 and cliff == 12 and not double:
vest_score = 60
findings.append(_warn("vesting",
"4yr/1yr cliff WITHOUT acceleration. Push for double-trigger (change of control + termination "
"without cause) to protect founder upside in acquisition scenarios."))
elif years > 4:
vest_score = 20
findings.append(_crit("vesting",
f"{years}-year vesting. Non-standard; reject. 4 years is industry norm."))
else:
vest_score = 70
findings.append(_warn("vesting", f"{years}yr/{cliff}mo cliff — verify acceleration with counsel."))
scores.append(vest_score)
# --- 6. Pro-Rata Rights ---
if ts.get("pro_rata", True):
pr_score = 100
findings.append(_ok("pro_rata", "Pro-rata rights — standard for the lead and major investors."))
else:
pr_score = 60
findings.append(_warn("pro_rata",
"No pro-rata rights. Unusual; if investor is offering this, ask why (signals weak conviction "
"or competitive pressure). Pro-rata is generally fine for founders to grant."))
scores.append(pr_score)
# --- 7. Drag-Along ---
drag = ts.get("drag_along", {})
if drag.get("exists") and drag.get("founder_consent_required"):
drag_score = 100
findings.append(_ok("drag_along",
"Drag-along exists but requires founder consent — balanced."))
elif drag.get("exists") and not drag.get("founder_consent_required"):
drag_score = 40
findings.append(_crit("drag_along",
"Drag-along WITHOUT founder consent. Investors can force a sale over founder objection. "
"Push for founder consent OR a minimum price threshold (e.g., 3x preference) to trigger drag."))
else:
drag_score = 80
findings.append(_ok("drag_along", "No drag-along — neutral; common at early stages."))
scores.append(drag_score)
# --- 8. Protective Provisions ---
pp = ts.get("protective_provisions", "standard")
if pp == "standard":
pp_score = 100
findings.append(_ok("protective_provisions",
"Standard protective provisions (NVCA model) — acceptable."))
elif pp == "aggressive":
pp_score = 40
findings.append(_crit("protective_provisions",
"Aggressive protective provisions can require investor consent for routine operating "
"decisions (hiring execs, budget changes, vendor contracts). Push back to NVCA standard."))
else:
pp_score = 70
findings.append(_warn("protective_provisions", f"Verify scope with counsel: {pp}"))
scores.append(pp_score)
# --- 9. Information Rights ---
ir = ts.get("information_rights", "standard")
if ir == "standard":
ir_score = 100
findings.append(_ok("information_rights",
"Standard information rights (quarterly financials, annual audited, budget) — acceptable."))
elif ir == "aggressive":
ir_score = 60
findings.append(_warn("information_rights",
"Aggressive information rights (monthly financials, board observer rights, inspection rights). "
"Reasonable for lead at Series B+; at Series A, push to quarterly."))
else:
ir_score = 75
findings.append(_warn("information_rights", f"Verify: {ir}"))
scores.append(ir_score)
# --- 10. Dividends ---
div = ts.get("dividends", "none")
if div == "none":
div_score = 100
findings.append(_ok("dividends", "No dividend obligation — founder-friendly standard."))
elif div == "non_cumulative_when_declared":
div_score = 80
findings.append(_ok("dividends",
"Non-cumulative when-declared dividends — acceptable; rare to actually be paid."))
elif div == "cumulative":
div_score = 30
findings.append(_crit("dividends",
"CUMULATIVE dividends accrue every year regardless of declaration and must be paid at exit. "
"Hostile; push to non-cumulative or none."))
else:
div_score = 60
findings.append(_warn("dividends", f"Verify dividend type: {div}"))
scores.append(div_score)
# --- 11. Valuation Sanity ---
pre = ts.get("pre_money", 0)
raise_amt = ts.get("raise_amount", 0)
if pre and raise_amt:
post = pre + raise_amt
dilution = (raise_amt / post) * 100
if dilution > 30:
val_score = 40
findings.append(_crit("valuation",
f"Round dilutes {dilution:.1f}% (raise , on , pre = , post). "
"Over 30% in a single round is heavy; standard is 15-25%."))
elif dilution > 25:
val_score = 70
findings.append(_warn("valuation",
f"Round dilutes {dilution:.1f}%. Acceptable but on the high end. Standard 15-25%."))
else:
val_score = 100
findings.append(_ok("valuation",
f"Round dilutes {dilution:.1f}% — within standard 15-25% range."))
scores.append(val_score)
# --- 12. Holistic posture ---
crit_count = sum(1 for f in findings if f["severity"] == "CRITICAL")
if crit_count >= 3:
findings.append(_crit("holistic",
f"{crit_count} CRITICAL flags. This is a hostile term sheet. Either renegotiate the worst clauses "
"or walk. Do not sign as-is."))
elif crit_count >= 1:
findings.append(_warn("holistic",
f"{crit_count} CRITICAL flag(s). Address before signing; the rest is negotiable but not "
"disqualifying."))
else:
findings.append(_ok("holistic", "No critical flags. Standard founder-friendly term sheet."))
total_score = round(sum(scores) / len(scores)) if scores else 0
return total_score, findings
def _ok(clause: str, msg: str) -> Dict[str, Any]:
return {"clause": clause, "severity": "OK", "message": msg}
def _warn(clause: str, msg: str) -> Dict[str, Any]:
return {"clause": clause, "severity": "WARN", "message": msg}
def _crit(clause: str, msg: str) -> Dict[str, Any]:
return {"clause": clause, "severity": "CRITICAL", "message": msg}
def render_text(score_val: int, findings: List[Dict[str, Any]], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("TERM SHEET ANALYSIS")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
grade = (
"🟢 FOUNDER-FRIENDLY" if score_val >= 85 else
"🟡 NEGOTIATE" if score_val >= 65 else
"🔴 HOSTILE"
)
lines.append(f"Founder-friendliness score: {score_val}/100 {grade}")
lines.append("")
lines.append("-" * 72)
for f in findings:
sev = f["severity"]
marker = {"OK": "✅", "WARN": "⚠️ ", "CRITICAL": "🚨"}.get(sev, "•")
lines.append(f"{marker} [{sev:>8}] {f['clause']}")
lines.append(f" {f['message']}")
lines.append("")
lines.append("-" * 72)
lines.append("REMINDER: This tool is not legal advice. Always engage venture / securities counsel.")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Score a term sheet on founder-friendliness across 12 dimensions.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to term sheet JSON file (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
ts = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
ts = SAMPLE
source = "<embedded sample Series A term sheet>"
score_val, findings = score(ts)
if args.output == "json":
print(json.dumps({
"source": source,
"score": score_val,
"grade": "FOUNDER_FRIENDLY" if score_val >= 85 else "NEGOTIATE" if score_val >= 65 else "HOSTILE",
"findings": findings,
}, indent=2))
else:
print(render_text(score_val, findings, source))
return 0
if __name__ == "__main__":
sys.exit(main())
Tạo test Playwright: viết test cho trang, component hoặc tính năng, gồm test e2e.
---
name: "generate"
description: >-
Generate Playwright tests. Use when user says "write tests", "generate tests",
"add tests for", "test this component", "e2e test", "create test for",
"test this page", or "test this feature".
---
# Generate Playwright Tests
Generate production-ready Playwright tests from a user story, URL, component name, or feature description.
## Input
`$ARGUMENTS` contains what to test. Examples:
- `"user can log in with email and password"`
- `"the checkout flow"`
- `"src/components/UserProfile.tsx"`
- `"the search page with filters"`
## Steps
### 1. Understand the Target
Parse `$ARGUMENTS` to determine:
- **User story**: Extract the behavior to verify
- **Component path**: Read the component source code
- **Page/URL**: Identify the route and its elements
- **Feature name**: Map to relevant app areas
### 2. Explore the Codebase
Use the `Explore` subagent to gather context:
- Read `playwright.config.ts` for `testDir`, `baseURL`, `projects`
- Check existing tests in `testDir` for patterns, fixtures, and conventions
- If a component path is given, read the component to understand its props, states, and interactions
- Check for existing page objects in `pages/`
- Check for existing fixtures in `fixtures/`
- Check for auth setup (`auth.setup.ts` or `storageState` config)
### 3. Select Templates
Check `templates/` in this plugin for matching patterns:
| If testing... | Load template from |
|---|---|
| Login/auth flow | `templates/auth/login.md` |
| CRUD operations | `templates/crud/` |
| Checkout/payment | `templates/checkout/` |
| Search/filter UI | `templates/search/` |
| Form submission | `templates/forms/` |
| Dashboard/data | `templates/dashboard/` |
| Settings page | `templates/settings/` |
| Onboarding flow | `templates/onboarding/` |
| API endpoints | `templates/api/` |
| Accessibility | `templates/accessibility/` |
Adapt the template to the specific app — replace `{{placeholders}}` with actual selectors, URLs, and data.
### 4. Generate the Test
Follow these rules:
**Structure:**
```typescript
import { test, expect } from '@playwright/test';
// Import custom fixtures if the project uses them
test.describe('Feature Name', () => {
// Group related behaviors
test('should <expected behavior>', async ({ page }) => {
// Arrange: navigate, set up state
// Act: perform user action
// Assert: verify outcome
});
});
```
**Locator priority** (use the first that works):
1. `getByRole()` — buttons, links, headings, form elements
2. `getByLabel()` — form fields with labels
3. `getByText()` — non-interactive text content
4. `getByPlaceholder()` — inputs with placeholder text
5. `getByTestId()` — when semantic options aren't available
**Assertions** — always web-first:
```typescript
// GOOD — auto-retries
await expect(page.getByRole('heading')).toBeVisible();
await expect(page.getByRole('alert')).toHaveText('Success');
// BAD — no retry
const text = await page.textContent('.msg');
expect(text).toBe('Success');
```
**Never use:**
- `page.waitForTimeout()`
- `page.$(selector)` or `page.$$(selector)`
- Bare CSS selectors unless absolutely necessary
- `page.evaluate()` for things locators can do
**Always include:**
- Descriptive test names that explain the behavior
- Error/edge case tests alongside happy path
- Proper `await` on every Playwright call
- `baseURL`-relative navigation (`page.goto('/')` not `page.goto('http://...')`)
### 5. Match Project Conventions
- If project uses TypeScript → generate `.spec.ts`
- If project uses JavaScript → generate `.spec.js` with `require()` imports
- If project has page objects → use them instead of inline locators
- If project has custom fixtures → import and use them
- If project has a test data directory → create test data files there
### 6. Generate Supporting Files (If Needed)
- **Page object**: If the test touches 5+ unique locators on one page, create a page object
- **Fixture**: If the test needs shared setup (auth, data), create or extend a fixture
- **Test data**: If the test uses structured data, create a JSON file in `test-data/`
### 7. Verify
Run the generated test:
```bash
npx playwright test <generated-file> --reporter=list
```
If it fails:
1. Read the error
2. Fix the test (not the app)
3. Run again
4. If it's an app issue, report it to the user
## Output
- Generated test file(s) with path
- Any supporting files created (page objects, fixtures, data)
- Test run result
- Coverage note: what behaviors are now tested
FILE:patterns.md
# Test Generation Patterns
## Pattern: Authentication Flow
```typescript
test.describe('Authentication', () => {
test('should login with valid credentials', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill('user@example.com');
await page.getByLabel('Password').fill('password123');
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page).toHaveURL('/dashboard');
await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
});
test('should show error for invalid credentials', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill('wrong@example.com');
await page.getByLabel('Password').fill('wrong');
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page.getByRole('alert')).toHaveText(/invalid/i);
await expect(page).toHaveURL('/login');
});
});
```
## Pattern: CRUD Operations
```typescript
test.describe('Items', () => {
test('should create a new item', async ({ page }) => {
await page.goto('/items');
await page.getByRole('button', { name: 'Add item' }).click();
await page.getByLabel('Name').fill('Test Item');
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByText('Test Item')).toBeVisible();
});
test('should edit an existing item', async ({ page }) => {
await page.goto('/items');
await page.getByRole('row', { name: /Test Item/ })
.getByRole('button', { name: 'Edit' }).click();
await page.getByLabel('Name').clear();
await page.getByLabel('Name').fill('Updated Item');
await page.getByRole('button', { name: 'Save' }).click();
await expect(page.getByText('Updated Item')).toBeVisible();
});
test('should delete an item with confirmation', async ({ page }) => {
await page.goto('/items');
await page.getByRole('row', { name: /Test Item/ })
.getByRole('button', { name: 'Delete' }).click();
await page.getByRole('button', { name: 'Confirm' }).click();
await expect(page.getByText('Test Item')).not.toBeVisible();
});
});
```
## Pattern: Form with Validation
```typescript
test.describe('Contact Form', () => {
test.beforeEach(async ({ page }) => {
await page.goto('/contact');
});
test('should submit valid form', async ({ page }) => {
await page.getByLabel('Name').fill('Jane Doe');
await page.getByLabel('Email').fill('jane@example.com');
await page.getByLabel('Message').fill('Hello, this is a test message.');
await page.getByRole('button', { name: 'Send' }).click();
await expect(page.getByText('Message sent')).toBeVisible();
});
test('should show validation errors for empty required fields', async ({ page }) => {
await page.getByRole('button', { name: 'Send' }).click();
await expect(page.getByText('Name is required')).toBeVisible();
await expect(page.getByText('Email is required')).toBeVisible();
});
test('should validate email format', async ({ page }) => {
await page.getByLabel('Email').fill('not-an-email');
await page.getByRole('button', { name: 'Send' }).click();
await expect(page.getByText('Invalid email')).toBeVisible();
});
});
```
## Pattern: Search and Filter
```typescript
test.describe('Product Search', () => {
test('should return results for valid query', async ({ page }) => {
await page.goto('/products');
await page.getByPlaceholder('Search products').fill('laptop');
await page.getByRole('button', { name: 'Search' }).click();
await expect(page.getByRole('list')).toBeVisible();
const results = page.getByRole('listitem');
await expect(results).not.toHaveCount(0);
});
test('should show empty state for no results', async ({ page }) => {
await page.goto('/products');
await page.getByPlaceholder('Search products').fill('xyznonexistent');
await page.getByRole('button', { name: 'Search' }).click();
await expect(page.getByText('No products found')).toBeVisible();
});
test('should filter by category', async ({ page }) => {
await page.goto('/products');
await page.getByRole('combobox', { name: 'Category' }).selectOption('Electronics');
await expect(page.getByRole('listitem')).not.toHaveCount(0);
});
});
```
## Pattern: Navigation and Layout
```typescript
test.describe('Navigation', () => {
test('should navigate between pages', async ({ page }) => {
await page.goto('/');
await page.getByRole('link', { name: 'About' }).click();
await expect(page).toHaveURL('/about');
await expect(page.getByRole('heading', { level: 1 })).toHaveText('About');
});
test('should show mobile menu on small screens', async ({ page }) => {
await page.setViewportSize({ width: 375, height: 667 });
await page.goto('/');
await expect(page.getByRole('navigation')).not.toBeVisible();
await page.getByRole('button', { name: 'Menu' }).click();
await expect(page.getByRole('navigation')).toBeVisible();
});
});
```
## Pattern: API Mocking
```typescript
test.describe('Dashboard with mocked API', () => {
test('should display data from API', async ({ page }) => {
await page.route('**/api/dashboard', (route) => {
route.fulfill({
status: 200,
contentType: 'application/json',
body: JSON.stringify({ revenue: 50000, users: 1200 }),
});
});
await page.goto('/dashboard');
await expect(page.getByText('$50,000')).toBeVisible();
await expect(page.getByText('1,200')).toBeVisible();
});
test('should handle API errors gracefully', async ({ page }) => {
await page.route('**/api/dashboard', (route) => {
route.fulfill({ status: 500 });
});
await page.goto('/dashboard');
await expect(page.getByText(/error|try again/i)).toBeVisible();
});
});
```
Chạy song song nhiều tính năng với Git worktree: cô lập nhánh, cấp cổng, đồng bộ môi trường và dọn dẹp cho từng worktree.
---
name: "git-worktree-manager"
description: "Run parallel feature work safely with Git worktrees. Standardizes branch isolation, port allocation, environment sync, and cleanup so each worktree behaves like an independent local app. Optimized for multi-agent workflows where each agent or terminal session owns one worktree. Use when running multiple feature branches simultaneously, isolating experimental work, or coordinating multi-agent development across the same repo."
---
# Git Worktree Manager
**Tier:** POWERFUL
**Category:** Engineering
**Domain:** Parallel Development & Branch Isolation
## Overview
Use this skill to run parallel feature work safely with Git worktrees. It standardizes branch isolation, port allocation, environment sync, and cleanup so each worktree behaves like an independent local app without stepping on another branch.
This skill is optimized for multi-agent workflows where each agent or terminal session owns one worktree.
## Core Capabilities
- Create worktrees from new or existing branches with deterministic naming
- Auto-allocate non-conflicting ports per worktree and persist assignments
- Copy local environment files (`.env*`) from main repo to new worktree
- Optionally install dependencies based on lockfile detection
- Detect stale worktrees and uncommitted changes before cleanup
- Identify merged branches and safely remove outdated worktrees
## When to Use
- You need 2+ concurrent branches open locally
- You want isolated dev servers for feature, hotfix, and PR validation
- You are working with multiple agents that must not share a branch
- Your current branch is blocked but you need to ship a quick fix now
- You want repeatable cleanup instead of ad-hoc `rm -rf` operations
## Key Workflows
### 1. Create a Fully-Prepared Worktree
1. Pick a branch name and worktree name.
2. Run the manager script (creates branch if missing).
3. Review generated port map.
4. Start app using allocated ports.
```bash
python scripts/worktree_manager.py \
--repo . \
--branch feature/new-auth \
--name wt-auth \
--base-branch main \
--install-deps \
--format text
```
If you use JSON automation input:
```bash
cat config.json | python scripts/worktree_manager.py --format json
# or
python scripts/worktree_manager.py --input config.json --format json
```
### 2. Run Parallel Sessions
Recommended convention:
- Main repo: integration branch (`main`/`develop`) on default port
- Worktree A: feature branch + offset ports
- Worktree B: hotfix branch + next offset
Each worktree contains `.worktree-ports.json` with assigned ports.
### 3. Cleanup with Safety Checks
1. Scan all worktrees and stale age.
2. Inspect dirty trees and branch merge status.
3. Remove only merged + clean worktrees, or force explicitly.
```bash
python scripts/worktree_cleanup.py --repo . --stale-days 14 --format text
python scripts/worktree_cleanup.py --repo . --remove-merged --format text
```
### 4. Docker Compose Pattern
Use per-worktree override files mapped from allocated ports. The script outputs a deterministic port map; apply it to `docker-compose.worktree.yml`.
See [docker-compose-patterns.md](references/docker-compose-patterns.md) for concrete templates.
### 5. Port Allocation Strategy
Default strategy is `base + (index * stride)` with collision checks:
- App: `3000`
- Postgres: `5432`
- Redis: `6379`
- Stride: `10`
See [port-allocation-strategy.md](references/port-allocation-strategy.md) for the full strategy and edge cases.
## Script Interfaces
- `python scripts/worktree_manager.py --help`
- Create/list worktrees
- Allocate/persist ports
- Copy `.env*` files
- Optional dependency installation
- `python scripts/worktree_cleanup.py --help`
- Stale detection by age
- Dirty-state detection
- Merged-branch detection
- Optional safe removal
Both tools support stdin JSON and `--input` file mode for automation pipelines.
## Common Pitfalls
1. Creating worktrees inside the main repo directory
2. Reusing `localhost:3000` across all branches
3. Sharing one database URL across isolated feature branches
4. Removing a worktree with uncommitted changes
5. Forgetting to prune old metadata after branch deletion
6. Assuming merged status without checking against the target branch
## Best Practices
1. One branch per worktree, one agent per worktree.
2. Keep worktrees short-lived; remove after merge.
3. Use a deterministic naming pattern (`wt-<topic>`).
4. Persist port mappings in file, not memory or terminal notes.
5. Run cleanup scan weekly in active repos.
6. Use `--format json` for machine flows and `--format text` for human review.
7. Never force-remove dirty worktrees unless changes are intentionally discarded.
## Validation Checklist
Before claiming setup complete:
1. `git worktree list` shows expected path + branch.
2. `.worktree-ports.json` exists and contains unique ports.
3. `.env` files copied successfully (if present in source repo).
4. Dependency install command exits with code `0` (if enabled).
5. Cleanup scan reports no unintended stale dirty trees.
## References
- [port-allocation-strategy.md](references/port-allocation-strategy.md)
- [docker-compose-patterns.md](references/docker-compose-patterns.md)
- [README.md](README.md) for quick start and installation details
## Decision Matrix
Use this quick selector before creating a new worktree:
- Need isolated dependencies and server ports -> create a new worktree
- Need only a quick local diff review -> stay on current tree
- Need hotfix while feature branch is dirty -> create dedicated hotfix worktree
- Need ephemeral reproduction branch for bug triage -> create temporary worktree and cleanup same day
## Operational Checklist
### Before Creation
1. Confirm main repo has clean baseline or intentional WIP commits.
2. Confirm target branch naming convention.
3. Confirm required base branch exists (`main`/`develop`).
4. Confirm no reserved local ports are already occupied by non-repo services.
### After Creation
1. Verify `git status` branch matches expected branch.
2. Verify `.worktree-ports.json` exists.
3. Verify app boots on allocated app port.
4. Verify DB and cache endpoints target isolated ports.
### Before Removal
1. Verify branch has upstream and is merged when intended.
2. Verify no uncommitted files remain.
3. Verify no running containers/processes depend on this worktree path.
## CI and Team Integration
- Use worktree path naming that maps to task ID (`wt-1234-auth`).
- Include the worktree path in terminal title to avoid wrong-window commits.
- In automated setups, persist creation metadata in CI artifacts/logs.
- Trigger cleanup report in scheduled jobs and post summary to team channel.
## Failure Recovery
- If `git worktree add` fails due to existing path: inspect path, do not overwrite.
- If dependency install fails: keep worktree created, mark status and continue manual recovery.
- If env copy fails: continue with warning and explicit missing file list.
- If port allocation collides with external service: rerun with adjusted base ports.
FILE:README.md
# Git Worktree Manager
Production workflow for parallel branch development with isolated ports, env sync, and cleanup safety checks. This skill packages practical CLI tooling and operating guidance for multi-worktree teams.
## Quick Start
```bash
# Create + prepare a worktree
python scripts/worktree_manager.py \
--repo . \
--branch feature/api-hardening \
--name wt-api-hardening \
--base-branch main \
--install-deps \
--format text
# Review stale worktrees
python scripts/worktree_cleanup.py --repo . --stale-days 14 --format text
```
## Included Tools
- `scripts/worktree_manager.py`: create/list-prep workflow, deterministic ports, `.env*` sync, optional dependency install
- `scripts/worktree_cleanup.py`: stale/dirty/merged analysis with optional safe removal
Both support `--input <json-file>` and stdin JSON for automation.
## References
- `references/port-allocation-strategy.md`
- `references/docker-compose-patterns.md`
## Installation
### Claude Code
```bash
cp -R engineering/git-worktree-manager ~/.claude/skills/git-worktree-manager
```
### OpenAI Codex
```bash
cp -R engineering/git-worktree-manager ~/.codex/skills/git-worktree-manager
```
### OpenClaw
```bash
cp -R engineering/git-worktree-manager ~/.openclaw/skills/git-worktree-manager
```
FILE:references/docker-compose-patterns.md
# Docker Compose Patterns For Worktrees
## Pattern 1: Override File Per Worktree
Base compose file remains shared; each worktree has a local override.
`docker-compose.worktree.yml`:
```yaml
services:
app:
ports:
- "3010:3000"
db:
ports:
- "5442:5432"
redis:
ports:
- "6389:6379"
```
Run:
```bash
docker compose -f docker-compose.yml -f docker-compose.worktree.yml up -d
```
## Pattern 2: `.env` Driven Ports
Use compose variable substitution and write worktree-specific values into `.env.local`.
`docker-compose.yml` excerpt:
```yaml
services:
app:
ports: ["-3000:3000"]
db:
ports: ["-5432:5432"]
```
Worktree `.env.local`:
```env
APP_PORT=3010
DB_PORT=5442
REDIS_PORT=6389
```
## Pattern 3: Project Name Isolation
Use unique compose project name so container, network, and volume names do not collide.
```bash
docker compose -p myapp_wt_auth up -d
```
## Common Mistakes
- Reusing default `5432` from multiple worktrees simultaneously
- Sharing one database volume across incompatible migration branches
- Forgetting to scope compose project name per worktree
FILE:references/port-allocation-strategy.md
# Port Allocation Strategy
## Objective
Allocate deterministic, non-overlapping local ports for each worktree to avoid collisions across concurrent development sessions.
## Default Mapping
- App HTTP: `3000`
- Postgres: `5432`
- Redis: `6379`
- Stride per worktree: `10`
Formula by slot index `n`:
- `app = 3000 + (10 * n)`
- `db = 5432 + (10 * n)`
- `redis = 6379 + (10 * n)`
Examples:
- Slot 0: `3000/5432/6379`
- Slot 1: `3010/5442/6389`
- Slot 2: `3020/5452/6399`
## Collision Avoidance
1. Read `.worktree-ports.json` from existing worktrees.
2. Skip any slot where one or more ports are already assigned.
3. Persist selected mapping in the new worktree.
## Operational Notes
- Keep stride >= number of services to avoid accidental overlaps when adding ports later.
- For custom service sets, reserve a contiguous block per worktree.
- If you also run local infra outside worktrees, offset bases to avoid global collisions.
## Recommended File Format
```json
{
"app": 3010,
"db": 5442,
"redis": 6389
}
```
FILE:scripts/worktree_cleanup.py
#!/usr/bin/env python3
"""Inspect and clean stale git worktrees with safety checks.
Supports:
- JSON input from stdin or --input file
- Stale age detection
- Dirty working tree detection
- Merged branch detection
- Optional removal of merged, clean stale worktrees
"""
import argparse
import json
import subprocess
import sys
import time
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Any, Dict, List, Optional
class CLIError(Exception):
"""Raised for expected CLI errors."""
@dataclass
class WorktreeInfo:
path: str
branch: str
is_main: bool
age_days: int
stale: bool
dirty: bool
merged_into_base: bool
def run(cmd: List[str], cwd: Optional[Path] = None, check: bool = True) -> subprocess.CompletedProcess[str]:
return subprocess.run(cmd, cwd=cwd, text=True, capture_output=True, check=check)
def load_json_input(input_file: Optional[str]) -> Dict[str, Any]:
if input_file:
try:
return json.loads(Path(input_file).read_text(encoding="utf-8"))
except Exception as exc:
raise CLIError(f"Failed reading --input file: {exc}") from exc
if not sys.stdin.isatty():
raw = sys.stdin.read().strip()
if raw:
try:
return json.loads(raw)
except json.JSONDecodeError as exc:
raise CLIError(f"Invalid JSON from stdin: {exc}") from exc
return {}
def parse_worktrees(repo: Path) -> List[Dict[str, str]]:
proc = run(["git", "worktree", "list", "--porcelain"], cwd=repo)
entries: List[Dict[str, str]] = []
current: Dict[str, str] = {}
for line in proc.stdout.splitlines():
if not line.strip():
if current:
entries.append(current)
current = {}
continue
key, _, value = line.partition(" ")
current[key] = value
if current:
entries.append(current)
return entries
def get_branch(path: Path) -> str:
proc = run(["git", "rev-parse", "--abbrev-ref", "HEAD"], cwd=path)
return proc.stdout.strip()
def get_last_commit_age_days(path: Path) -> int:
proc = run(["git", "log", "-1", "--format=%ct"], cwd=path)
timestamp = int(proc.stdout.strip() or "0")
age_seconds = int(time.time()) - timestamp
return max(0, age_seconds // 86400)
def is_dirty(path: Path) -> bool:
proc = run(["git", "status", "--porcelain"], cwd=path)
return bool(proc.stdout.strip())
def is_merged(repo: Path, branch: str, base_branch: str) -> bool:
if branch in ("HEAD", base_branch):
return False
try:
run(["git", "merge-base", "--is-ancestor", branch, base_branch], cwd=repo)
return True
except subprocess.CalledProcessError:
return False
def format_text(items: List[WorktreeInfo], removed: List[str]) -> str:
lines = ["Worktree cleanup report"]
for item in items:
lines.append(
f"- {item.path} | branch={item.branch} | age={item.age_days}d | "
f"stale={item.stale} dirty={item.dirty} merged={item.merged_into_base}"
)
if removed:
lines.append("Removed:")
for path in removed:
lines.append(f"- {path}")
return "\n".join(lines)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Analyze and optionally cleanup stale git worktrees.")
parser.add_argument("--input", help="Path to JSON input file. If omitted, reads JSON from stdin when piped.")
parser.add_argument("--repo", default=".", help="Repository root path.")
parser.add_argument("--base-branch", default="main", help="Base branch to evaluate merged branches.")
parser.add_argument("--stale-days", type=int, default=14, help="Threshold for stale worktrees.")
parser.add_argument("--remove-merged", action="store_true", help="Remove worktrees that are stale, clean, and merged.")
parser.add_argument("--force", action="store_true", help="Allow removal even if dirty (use carefully).")
parser.add_argument("--format", choices=["text", "json"], default="text", help="Output format.")
return parser.parse_args()
def main() -> int:
args = parse_args()
payload = load_json_input(args.input)
repo = Path(str(payload.get("repo", args.repo))).resolve()
stale_days = int(payload.get("stale_days", args.stale_days))
base_branch = str(payload.get("base_branch", args.base_branch))
remove_merged = bool(payload.get("remove_merged", args.remove_merged))
force = bool(payload.get("force", args.force))
try:
run(["git", "rev-parse", "--is-inside-work-tree"], cwd=repo)
except subprocess.CalledProcessError as exc:
raise CLIError(f"Not a git repository: {repo}") from exc
try:
run(["git", "rev-parse", "--verify", base_branch], cwd=repo)
except subprocess.CalledProcessError as exc:
raise CLIError(f"Base branch not found: {base_branch}") from exc
entries = parse_worktrees(repo)
if not entries:
raise CLIError("No worktrees found.")
main_path = Path(entries[0].get("worktree", "")).resolve()
infos: List[WorktreeInfo] = []
removed: List[str] = []
for entry in entries:
path = Path(entry.get("worktree", "")).resolve()
branch = get_branch(path)
age = get_last_commit_age_days(path)
dirty = is_dirty(path)
stale = age >= stale_days
merged = is_merged(repo, branch, base_branch)
info = WorktreeInfo(
path=str(path),
branch=branch,
is_main=path == main_path,
age_days=age,
stale=stale,
dirty=dirty,
merged_into_base=merged,
)
infos.append(info)
if remove_merged and not info.is_main and info.stale and info.merged_into_base and (force or not info.dirty):
try:
cmd = ["git", "worktree", "remove", str(path)]
if force:
cmd.append("--force")
run(cmd, cwd=repo)
removed.append(str(path))
except subprocess.CalledProcessError as exc:
raise CLIError(f"Failed removing worktree {path}: {exc.stderr}") from exc
if args.format == "json":
print(json.dumps({"worktrees": [asdict(i) for i in infos], "removed": removed}, indent=2))
else:
print(format_text(infos, removed))
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except CLIError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
raise SystemExit(2)
FILE:scripts/worktree_manager.py
#!/usr/bin/env python3
"""Create and prepare git worktrees with deterministic port allocation.
Supports:
- JSON input from stdin or --input file
- Worktree creation from existing/new branch
- .env file sync from main repo
- Optional dependency installation
- JSON or text output
"""
import argparse
import json
import os
import shutil
import subprocess
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Any, Dict, List, Optional
ENV_FILES = [".env", ".env.local", ".env.development", ".envrc"]
LOCKFILE_COMMANDS = [
("pnpm-lock.yaml", ["pnpm", "install"]),
("yarn.lock", ["yarn", "install"]),
("package-lock.json", ["npm", "install"]),
("bun.lockb", ["bun", "install"]),
("requirements.txt", [sys.executable, "-m", "pip", "install", "-r", "requirements.txt"]),
]
@dataclass
class WorktreeResult:
repo: str
worktree_path: str
branch: str
created: bool
ports: Dict[str, int]
copied_env_files: List[str]
dependency_install: str
class CLIError(Exception):
"""Raised for expected CLI errors."""
def run(cmd: List[str], cwd: Optional[Path] = None, check: bool = True) -> subprocess.CompletedProcess[str]:
return subprocess.run(cmd, cwd=cwd, text=True, capture_output=True, check=check)
def load_json_input(input_file: Optional[str]) -> Dict[str, Any]:
if input_file:
try:
return json.loads(Path(input_file).read_text(encoding="utf-8"))
except Exception as exc:
raise CLIError(f"Failed reading --input file: {exc}") from exc
if not sys.stdin.isatty():
data = sys.stdin.read().strip()
if data:
try:
return json.loads(data)
except json.JSONDecodeError as exc:
raise CLIError(f"Invalid JSON from stdin: {exc}") from exc
return {}
def parse_worktree_list(repo: Path) -> List[Dict[str, str]]:
proc = run(["git", "worktree", "list", "--porcelain"], cwd=repo)
entries: List[Dict[str, str]] = []
current: Dict[str, str] = {}
for line in proc.stdout.splitlines():
if not line.strip():
if current:
entries.append(current)
current = {}
continue
key, _, value = line.partition(" ")
current[key] = value
if current:
entries.append(current)
return entries
def find_next_ports(repo: Path, app_base: int, db_base: int, redis_base: int, stride: int) -> Dict[str, int]:
used_ports = set()
for entry in parse_worktree_list(repo):
wt_path = Path(entry.get("worktree", ""))
ports_file = wt_path / ".worktree-ports.json"
if ports_file.exists():
try:
payload = json.loads(ports_file.read_text(encoding="utf-8"))
used_ports.update(int(v) for v in payload.values() if isinstance(v, int))
except Exception:
continue
index = 0
while True:
ports = {
"app": app_base + (index * stride),
"db": db_base + (index * stride),
"redis": redis_base + (index * stride),
}
if all(p not in used_ports for p in ports.values()):
return ports
index += 1
def sync_env_files(src_repo: Path, dest_repo: Path) -> List[str]:
copied = []
for name in ENV_FILES:
src = src_repo / name
if src.exists() and src.is_file():
dst = dest_repo / name
shutil.copy2(src, dst)
copied.append(name)
return copied
def install_dependencies_if_requested(worktree_path: Path, install: bool) -> str:
if not install:
return "skipped"
for lockfile, command in LOCKFILE_COMMANDS:
if (worktree_path / lockfile).exists():
try:
run(command, cwd=worktree_path, check=True)
return f"installed via {' '.join(command)}"
except subprocess.CalledProcessError as exc:
raise CLIError(f"Dependency install failed: {' '.join(command)}\n{exc.stderr}") from exc
return "no known lockfile found"
def ensure_worktree(repo: Path, branch: str, name: str, base_branch: str) -> Path:
wt_parent = repo.parent
wt_path = wt_parent / name
existing_paths = {Path(e.get("worktree", "")) for e in parse_worktree_list(repo)}
if wt_path in existing_paths:
return wt_path
try:
run(["git", "show-ref", "--verify", f"refs/heads/{branch}"], cwd=repo)
run(["git", "worktree", "add", str(wt_path), branch], cwd=repo)
except subprocess.CalledProcessError:
try:
run(["git", "worktree", "add", "-b", branch, str(wt_path), base_branch], cwd=repo)
except subprocess.CalledProcessError as exc:
raise CLIError(f"Failed to create worktree: {exc.stderr}") from exc
return wt_path
def format_text(result: WorktreeResult) -> str:
lines = [
"Worktree prepared",
f"- repo: {result.repo}",
f"- path: {result.worktree_path}",
f"- branch: {result.branch}",
f"- created: {result.created}",
f"- ports: app={result.ports['app']} db={result.ports['db']} redis={result.ports['redis']}",
f"- copied env files: {', '.join(result.copied_env_files) if result.copied_env_files else 'none'}",
f"- dependency install: {result.dependency_install}",
]
return "\n".join(lines)
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Create and prepare a git worktree.")
parser.add_argument("--input", help="Path to JSON input file. If omitted, reads JSON from stdin when piped.")
parser.add_argument("--repo", default=".", help="Path to repository root (default: current directory).")
parser.add_argument("--branch", help="Branch name for the worktree.")
parser.add_argument("--name", help="Worktree directory name (created adjacent to repo).")
parser.add_argument("--base-branch", default="main", help="Base branch when creating a new branch.")
parser.add_argument("--app-base", type=int, default=3000, help="Base app port.")
parser.add_argument("--db-base", type=int, default=5432, help="Base DB port.")
parser.add_argument("--redis-base", type=int, default=6379, help="Base Redis port.")
parser.add_argument("--stride", type=int, default=10, help="Port stride between worktrees.")
parser.add_argument("--install-deps", action="store_true", help="Install dependencies in the new worktree.")
parser.add_argument("--format", choices=["text", "json"], default="text", help="Output format.")
return parser.parse_args()
def main() -> int:
args = parse_args()
payload = load_json_input(args.input)
repo = Path(str(payload.get("repo", args.repo))).resolve()
branch = payload.get("branch", args.branch)
name = payload.get("name", args.name)
base_branch = str(payload.get("base_branch", args.base_branch))
app_base = int(payload.get("app_base", args.app_base))
db_base = int(payload.get("db_base", args.db_base))
redis_base = int(payload.get("redis_base", args.redis_base))
stride = int(payload.get("stride", args.stride))
install_deps = bool(payload.get("install_deps", args.install_deps))
if not branch or not name:
raise CLIError("Missing required values: --branch and --name (or provide via JSON input).")
try:
run(["git", "rev-parse", "--is-inside-work-tree"], cwd=repo)
except subprocess.CalledProcessError as exc:
raise CLIError(f"Not a git repository: {repo}") from exc
wt_path = ensure_worktree(repo, branch, name, base_branch)
created = (wt_path / ".worktree-ports.json").exists() is False
ports = find_next_ports(repo, app_base, db_base, redis_base, stride)
(wt_path / ".worktree-ports.json").write_text(json.dumps(ports, indent=2), encoding="utf-8")
copied = sync_env_files(repo, wt_path)
install_status = install_dependencies_if_requested(wt_path, install_deps)
result = WorktreeResult(
repo=str(repo),
worktree_path=str(wt_path),
branch=branch,
created=created,
ports=ports,
copied_env_files=copied,
dependency_install=install_status,
)
if args.format == "json":
print(json.dumps(asdict(result), indent=2))
else:
print(format_text(result))
return 0
if __name__ == "__main__":
try:
raise SystemExit(main())
except CLIError as exc:
print(f"ERROR: {exc}", file=sys.stderr)
raise SystemExit(2)
Thao tác Google Workspace CLI: chẩn đoán thiết lập, kiểm tra bảo mật, tìm công thức mẫu và phân tích đầu ra.
--- name: google-workspace description: "Google Workspace CLI operations: setup diagnostics, security audit, recipe discovery, and output analysis. Usage: /google-workspace <setup|audit|recipe|analyze> [options]" --- # /google-workspace Google Workspace CLI administration via the `gws` CLI. Run setup diagnostics, security audits, browse and execute recipes, and analyze command output. ## Usage ``` /google-workspace setup [--json] /google-workspace audit [--services gmail,drive,calendar] [--json] /google-workspace recipe list [--persona <role>] [--json] /google-workspace recipe search <keyword> [--json] /google-workspace recipe run <name> [--dry-run] /google-workspace recipe describe <name> /google-workspace analyze [--filter <field=value>] [--group-by <field>] [--stats <field>] [--format table|csv|json] ``` ## Examples ``` /google-workspace setup /google-workspace audit --services gmail,drive --json /google-workspace recipe list --persona pm /google-workspace recipe search "email" /google-workspace recipe run standup-report --dry-run /google-workspace recipe describe morning-briefing /google-workspace analyze --filter "mimeType=pdf" --select "name,size" --format table ``` ## Scripts - `engineering-team/google-workspace-cli/scripts/gws_doctor.py` — Pre-flight diagnostics - `engineering-team/google-workspace-cli/scripts/auth_setup_guide.py` — Auth setup guide - `engineering-team/google-workspace-cli/scripts/gws_recipe_runner.py` — Recipe catalog & runner - `engineering-team/google-workspace-cli/scripts/workspace_audit.py` — Security audit - `engineering-team/google-workspace-cli/scripts/output_analyzer.py` — JSON/NDJSON analyzer ## Subcommands ### setup Run pre-flight diagnostics and auth validation. ```bash python3 engineering-team/google-workspace-cli/scripts/gws_doctor.py [--json] python3 engineering-team/google-workspace-cli/scripts/auth_setup_guide.py --validate [--json] ``` ### audit Run security and configuration audit. ```bash python3 engineering-team/google-workspace-cli/scripts/workspace_audit.py [--services gmail,drive,calendar] [--json] ``` ### recipe Browse, search, and execute the 43 built-in gws recipes. ```bash python3 engineering-team/google-workspace-cli/scripts/gws_recipe_runner.py --list [--persona <role>] [--json] python3 engineering-team/google-workspace-cli/scripts/gws_recipe_runner.py --search <keyword> [--json] python3 engineering-team/google-workspace-cli/scripts/gws_recipe_runner.py --describe <name> python3 engineering-team/google-workspace-cli/scripts/gws_recipe_runner.py --run <name> [--dry-run] ``` ### analyze Parse, filter, and aggregate JSON output from any gws command. ```bash gws <command> --json | python3 engineering-team/google-workspace-cli/scripts/output_analyzer.py [options] python3 engineering-team/google-workspace-cli/scripts/output_analyzer.py --demo --format table ``` ## Skill Reference -> `engineering-team/google-workspace-cli/SKILL.md` ## Related Commands - No direct dependencies (self-contained Google Workspace skill)
Quản trị Google Workspace bằng gws CLI: cài đặt, xác thực, tự động hóa Gmail, Drive, Sheets, Calendar, Docs, Chat, Tasks và kiểm tra bảo mật.
---
name: "google-workspace-cli"
description: "Google Workspace administration via the gws CLI. Install, authenticate, and automate Gmail, Drive, Sheets, Calendar, Docs, Chat, and Tasks. Run security audits, execute 43 built-in recipes, and use 10 persona bundles. Use for Google Workspace admin, gws CLI setup, Gmail automation, Drive management, or Calendar scheduling."
---
# Google Workspace CLI
Expert guidance and automation for Google Workspace administration using the open-source `gws` CLI. Covers installation, authentication, 18+ service APIs, 43 built-in recipes, and 10 persona bundles for role-based workflows.
---
## Quick Start
### Check Installation
```bash
# Verify gws is installed and authenticated
python3 scripts/gws_doctor.py
```
### Send an Email
```bash
gws gmail users.messages send me --to "team@company.com" \
--subject "Weekly Update" --body "Here's this week's summary..."
```
### List Drive Files
```bash
gws drive files list --json --limit 20 | python3 scripts/output_analyzer.py --select "name,mimeType,modifiedTime" --format table
```
---
## Installation
### npm (recommended)
```bash
npm install -g @anthropic/gws
gws --version
```
### Cargo (from source)
```bash
cargo install gws-cli
gws --version
```
### Pre-built Binaries
Download from [github.com/googleworkspace/cli/releases](https://github.com/googleworkspace/cli/releases) for macOS, Linux, or Windows.
### Verify Installation
```bash
python3 scripts/gws_doctor.py
# Checks: PATH, version, auth status, service connectivity
```
---
## Authentication
### OAuth Setup (Interactive)
```bash
# Step 1: Create Google Cloud project and OAuth credentials
python3 scripts/auth_setup_guide.py --guide oauth
# Step 2: Run auth setup
gws auth setup
# Step 3: Validate
gws auth status --json
```
### Service Account (Headless/CI)
```bash
# Generate setup instructions
python3 scripts/auth_setup_guide.py --guide service-account
# Configure with key file
export GWS_SERVICE_ACCOUNT_KEY=/path/to/key.json
export GWS_DELEGATED_USER=admin@company.com
gws auth status
```
### Environment Variables
```bash
# Generate .env template
python3 scripts/auth_setup_guide.py --generate-env
```
| Variable | Purpose |
|----------|---------|
| `GWS_CLIENT_ID` | OAuth client ID |
| `GWS_CLIENT_SECRET` | OAuth client secret |
| `GWS_TOKEN_PATH` | Custom token storage path |
| `GWS_SERVICE_ACCOUNT_KEY` | Service account JSON key path |
| `GWS_DELEGATED_USER` | User to impersonate (service accounts) |
| `GWS_DEFAULT_FORMAT` | Default output format (json/ndjson/table) |
### Validate Authentication
```bash
python3 scripts/auth_setup_guide.py --validate --json
# Tests each service endpoint
```
---
## Workflow 1: Gmail Automation
**Goal:** Automate email operations — send, search, label, and filter management.
### Send and Reply
```bash
# Send a new email
gws gmail users.messages send me --to "client@example.com" \
--subject "Proposal" --body "Please find attached..." \
--attachment proposal.pdf
# Reply to a thread
gws gmail users.messages reply me --thread-id <THREAD_ID> \
--body "Thanks for your feedback..."
# Forward a message
gws gmail users.messages forward me --message-id <MSG_ID> \
--to "manager@company.com"
```
### Search and Filter
```bash
# Search emails
gws gmail users.messages list me --query "from:client@example.com after:2025/01/01" --json \
| python3 scripts/output_analyzer.py --count
# List labels
gws gmail users.labels list me --json
# Create a filter
gws gmail users.settings.filters create me \
--criteria '{"from":"notifications@service.com"}' \
--action '{"addLabelIds":["Label_123"],"removeLabelIds":["INBOX"]}'
```
### Bulk Operations
```bash
# Archive all read emails older than 30 days
gws gmail users.messages list me --query "is:read older_than:30d" --json \
| python3 scripts/output_analyzer.py --select "id" --format json \
| xargs -I {} gws gmail users.messages modify me {} --removeLabelIds INBOX
```
---
## Workflow 2: Drive & Sheets
**Goal:** Manage files, create spreadsheets, configure sharing, and export data.
### File Operations
```bash
# List files
gws drive files list --json --limit 50 \
| python3 scripts/output_analyzer.py --select "name,mimeType,size" --format table
# Upload a file
gws drive files create --name "Q1 Report" --upload report.pdf \
--parents <FOLDER_ID>
# Create a Google Sheet
gws sheets spreadsheets create --title "Budget 2026" --json
# Download/export
gws drive files export <FILE_ID> --mime "application/pdf" --output report.pdf
```
### Sharing
```bash
# Share with user
gws drive permissions create <FILE_ID> \
--type user --role writer --emailAddress "colleague@company.com"
# Share with domain (view only)
gws drive permissions create <FILE_ID> \
--type domain --role reader --domain "company.com"
# List who has access
gws drive permissions list <FILE_ID> --json
```
### Sheets Data
```bash
# Read a range
gws sheets spreadsheets.values get <SHEET_ID> --range "Sheet1!A1:D10" --json
# Write data
gws sheets spreadsheets.values update <SHEET_ID> --range "Sheet1!A1" \
--values '[["Name","Score"],["Alice",95],["Bob",87]]'
# Append rows
gws sheets spreadsheets.values append <SHEET_ID> --range "Sheet1!A1" \
--values '[["Charlie",92]]'
```
---
## Workflow 3: Calendar & Meetings
**Goal:** Schedule events, find available times, and generate standup reports.
### Event Management
```bash
# Create an event
gws calendar events insert primary \
--summary "Sprint Planning" \
--start "2026-03-15T10:00:00" --end "2026-03-15T11:00:00" \
--attendees "team@company.com" \
--location "Conference Room A"
# List upcoming events
gws calendar events list primary --timeMin "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
--maxResults 10 --json
# Quick event (natural language)
gws helpers quick-event "Lunch with Sarah tomorrow at noon"
```
### Find Available Time
```bash
# Check free/busy for multiple people
gws helpers find-time \
--attendees "alice@co.com,bob@co.com,charlie@co.com" \
--duration 60 --within "2026-03-15,2026-03-19" --json
```
### Standup Report
```bash
# Generate daily standup from calendar + tasks
gws recipes standup-report --json \
| python3 scripts/output_analyzer.py --format table
# Meeting prep (agenda + attendee info)
gws recipes meeting-prep --event-id <EVENT_ID>
```
---
## Workflow 4: Security Audit
**Goal:** Audit Google Workspace security configuration and generate remediation commands.
### Run Full Audit
```bash
# Full audit across all services
python3 scripts/workspace_audit.py --json
# Audit specific services
python3 scripts/workspace_audit.py --services gmail,drive,calendar
# Demo mode (no gws required)
python3 scripts/workspace_audit.py --demo
```
### Audit Checks
| Area | Check | Risk |
|------|-------|------|
| Drive | External sharing enabled | Data exfiltration |
| Gmail | Auto-forwarding rules | Data exfiltration |
| Gmail | DMARC/SPF/DKIM records | Email spoofing |
| Calendar | Default sharing visibility | Information leak |
| OAuth | Third-party app grants | Unauthorized access |
| Admin | Super admin count | Privilege escalation |
| Admin | 2-Step verification enforcement | Account takeover |
### Review and Remediate
```bash
# Review findings
python3 scripts/workspace_audit.py --json | python3 scripts/output_analyzer.py \
--filter "status=FAIL" --select "area,check,remediation"
# Execute remediation (example: restrict external sharing)
gws drive about get --json # Check current settings
# Follow remediation commands from audit output
```
---
## Python Tools
| Script | Purpose | Usage |
|--------|---------|-------|
| `gws_doctor.py` | Pre-flight diagnostics | `python3 scripts/gws_doctor.py [--json] [--services gmail,drive]` |
| `auth_setup_guide.py` | Guided auth setup | `python3 scripts/auth_setup_guide.py --guide oauth` |
| `gws_recipe_runner.py` | Recipe catalog & runner | `python3 scripts/gws_recipe_runner.py --list [--persona pm]` |
| `workspace_audit.py` | Security/config audit | `python3 scripts/workspace_audit.py [--json] [--demo]` |
| `output_analyzer.py` | JSON/NDJSON analysis | `gws ... --json \| python3 scripts/output_analyzer.py --count` |
All scripts are stdlib-only, support `--json` output, and include demo mode with embedded sample data.
---
## Best Practices
### Security
1. Use OAuth with minimal scopes — request only what each workflow needs
2. Store tokens in the system keyring, never in plain text files
3. Rotate service account keys every 90 days
4. Audit third-party OAuth app grants quarterly
5. Use `--dry-run` before bulk destructive operations
### Automation
1. Pipe `--json` output through `output_analyzer.py` for filtering and aggregation
2. Use recipes for multi-step operations instead of chaining raw commands
3. Select a persona bundle to scope recipes to your role
4. Use NDJSON format (`--format ndjson`) for streaming large result sets
5. Set `GWS_DEFAULT_FORMAT=json` in your shell profile for scripting
### Performance
1. Use `--fields` to request only needed fields (reduces payload size)
2. Use `--limit` to cap results when browsing
3. Use `--page-all` only when you need complete datasets
4. Batch operations with recipes rather than individual API calls
5. Cache frequently accessed data (e.g., label IDs, folder IDs) in variables
---
## Limitations
| Constraint | Impact |
|------------|--------|
| OAuth tokens expire after 1 hour | Re-auth needed for long-running scripts |
| API rate limits (per-user, per-service) | Bulk operations may hit 429 errors |
| Scope requirements vary by service | Must request correct scopes during auth |
| Pre-v1.0 CLI status | Breaking changes possible between releases |
| Google Cloud project required | Free, but requires setup in Cloud Console |
| Admin API needs admin privileges | Some audit checks require Workspace Admin role |
### Required Scopes by Service
```bash
# List scopes for specific services
python3 scripts/auth_setup_guide.py --scopes gmail,drive,calendar,sheets
```
| Service | Key Scopes |
|---------|-----------|
| Gmail | `gmail.modify`, `gmail.send`, `gmail.labels` |
| Drive | `drive.file`, `drive.metadata.readonly` |
| Sheets | `spreadsheets` |
| Calendar | `calendar`, `calendar.events` |
| Admin | `admin.directory.user.readonly`, `admin.directory.group` |
| Tasks | `tasks` |
FILE:assets/persona-profiles.md
# Google Workspace CLI Persona Profiles
10 role-based bundles that scope recipes and commands to your daily workflow.
---
## 1. Executive Assistant
**Description:** Managing schedules, emails, and communications for executives.
**Top Commands:**
- `gws helpers morning-briefing` — Start the day with schedule + inbox overview
- `gws helpers find-time` — Find available slots for meetings
- `gws helpers meeting-prep --event-id <id>` — Prepare meeting agenda
- `gws gmail users.messages send me` — Send emails on behalf
- `gws helpers eod-wrap` — End of day summary
**Recommended Recipes:** morning-briefing, today-schedule, find-time, send-email, reply-to-thread, meeting-prep, eod-wrap, quick-event, inbox-zero, standup-report
**Daily Workflow:**
1. Run `morning-briefing` at 8:00 AM
2. Process inbox with `inbox-zero`
3. Schedule meetings with `find-time` + `create-event`
4. Prep for meetings with `meeting-prep`
5. Close day with `eod-wrap`
---
## 2. Project Manager
**Description:** Tracking tasks, meetings, and project deliverables.
**Top Commands:**
- `gws recipes standup-report` — Generate standup updates
- `gws helpers find-time` — Schedule sprint ceremonies
- `gws tasks tasks insert` — Create and assign tasks
- `gws sheets spreadsheets.values get` — Read project trackers
- `gws recipes project-status` — Aggregate project status
**Recommended Recipes:** standup-report, create-event, find-time, task-create, task-progress, project-status, weekly-summary, share-folder, sheet-read, morning-briefing
**Daily Workflow:**
1. Run `standup-report` before standup
2. Update project tracker via `sheet-write`
3. Create action items with `task-create`
4. Run `weekly-summary` on Fridays
5. Share updates via `chat-message`
---
## 3. HR
**Description:** Managing people, onboarding, and team communications.
**Top Commands:**
- `gws admin users list` — List all domain users
- `gws admin users get <email>` — Look up employee details
- `gws docs documents create` — Create onboarding docs
- `gws drive permissions create` — Share folders with new hires
- `gws people people.connections list` — Export contact directory
**Recommended Recipes:** list-users, user-info, send-email, create-event, create-doc, share-folder, chat-message, list-groups, export-contacts, today-schedule
**Daily Workflow:**
1. Check new hire onboarding queue
2. Create welcome docs with `create-doc`
3. Set up 1:1s with `create-event`
4. Share team folders with `share-folder`
5. Send announcements via `send-email`
---
## 4. Sales
**Description:** Managing client communications, proposals, and scheduling.
**Top Commands:**
- `gws gmail users.messages send me` — Send proposals and follow-ups
- `gws gmail users.messages list me --query` — Search client conversations
- `gws helpers find-time` — Schedule client meetings
- `gws docs documents create` — Create proposals
- `gws sheets spreadsheets.values update` — Update pipeline tracker
**Recommended Recipes:** send-email, search-emails, create-event, find-time, create-doc, share-file, sheet-read, sheet-write, export-file, morning-briefing
**Daily Workflow:**
1. Run `morning-briefing` for meeting overview
2. Search emails for client updates
3. Update pipeline in Sheets
4. Send proposals via `send-email` + `share-file`
5. Schedule follow-ups with `create-event`
---
## 5. IT Admin
**Description:** Managing Workspace configuration, security, and user administration.
**Top Commands:**
- `gws admin users list --domain` — Audit user accounts
- `gws admin activities list login` — Monitor login activity
- `gws admin groups list` — Manage groups
- `python3 workspace_audit.py` — Run security audit
- `gws drive files list --orderBy "quotaBytesUsed desc"` — Find storage hogs
**Recommended Recipes:** list-users, list-groups, user-info, audit-logins, drive-activity, find-large-files, cleanup-trash, label-manager, filter-setup, share-folder
**Daily Workflow:**
1. Check `audit-logins` for suspicious activity
2. Run `workspace_audit.py` weekly
3. Process user provisioning requests
4. Monitor storage with `find-large-files`
5. Review group memberships
---
## 6. Developer
**Description:** Using Workspace APIs for automation and data integration.
**Top Commands:**
- `gws sheets spreadsheets.values get` — Read config/data from Sheets
- `gws sheets spreadsheets.values update` — Write results to Sheets
- `gws drive files create --upload` — Upload build artifacts
- `gws chat spaces.messages create` — Post deployment notifications
- `gws tasks tasks insert` — Create tasks from CI/CD
**Recommended Recipes:** sheet-read, sheet-write, sheet-append, upload-file, create-doc, chat-message, task-create, list-files, export-file, send-email
**Daily Workflow:**
1. Read config from Sheets API
2. Run automated reports to Sheets
3. Post updates to Chat spaces
4. Upload artifacts to Drive
5. Create tasks for bugs/issues
---
## 7. Marketing
**Description:** Managing campaigns, content creation, and team coordination.
**Top Commands:**
- `gws docs documents create` — Draft blog posts and briefs
- `gws drive files create --upload` — Upload creative assets
- `gws sheets spreadsheets.values append` — Log campaign metrics
- `gws gmail users.messages send me` — Send campaign emails
- `gws chat spaces.messages create` — Coordinate with team
**Recommended Recipes:** send-email, create-doc, share-file, upload-file, create-sheet, sheet-write, chat-message, create-event, email-stats, weekly-summary
**Daily Workflow:**
1. Check `email-stats` for campaign performance
2. Create content in Docs
3. Upload assets to shared Drive folders
4. Update metrics in Sheets
5. Coordinate launches via Chat
---
## 8. Finance
**Description:** Managing spreadsheets, financial reports, and data analysis.
**Top Commands:**
- `gws sheets spreadsheets.values get` — Pull financial data
- `gws sheets spreadsheets.values update` — Update forecasts
- `gws sheets spreadsheets create` — Create new reports
- `gws drive files export` — Export reports as PDF
- `gws drive permissions create` — Share with auditors
**Recommended Recipes:** sheet-read, sheet-write, sheet-append, create-sheet, export-file, share-file, send-email, find-large-files, drive-activity, weekly-summary
**Daily Workflow:**
1. Pull latest data into Sheets
2. Update financial models
3. Generate PDF reports with `export-file`
4. Share reports with stakeholders
5. Weekly summary for leadership
---
## 9. Legal
**Description:** Managing documents, contracts, and compliance.
**Top Commands:**
- `gws docs documents create` — Draft contracts
- `gws drive files export` — Export final versions as PDF
- `gws drive permissions create` — Manage document access
- `gws gmail users.messages list me --query` — Search for compliance emails
- `gws admin activities list` — Audit trail for compliance
**Recommended Recipes:** create-doc, share-file, export-file, search-emails, send-email, upload-file, list-files, drive-activity, audit-logins, find-large-files
**Daily Workflow:**
1. Draft and review documents
2. Search email for contract references
3. Export finalized docs as PDF
4. Set precise sharing permissions
5. Maintain audit trail
---
## 10. Customer Support
**Description:** Managing customer communications and ticket tracking.
**Top Commands:**
- `gws gmail users.messages list me --query` — Search customer emails
- `gws gmail users.messages reply me` — Reply to tickets
- `gws gmail users.labels create` — Organize by ticket status
- `gws tasks tasks insert` — Create follow-up tasks
- `gws chat spaces.messages create` — Escalate to team
**Recommended Recipes:** search-emails, send-email, reply-to-thread, label-manager, filter-setup, task-create, chat-message, unread-digest, inbox-zero, morning-briefing
**Daily Workflow:**
1. Run `morning-briefing` for ticket overview
2. Process inbox with label-based triage
3. Reply to open tickets
4. Escalate via Chat for urgent issues
5. Create follow-up tasks for pending items
FILE:assets/workspace-config.json
{
"_comment": "Google Workspace CLI automation config template. Copy and customize for your environment.",
"auth": {
"method": "oauth",
"client_id": "",
"client_secret": "",
"token_path": "~/.config/gws/token.json",
"service_account_key": "",
"delegated_user": ""
},
"defaults": {
"output_format": "json",
"pagination_limit": 100,
"timeout_ms": 30000,
"log_level": "warn"
},
"persona": "developer",
"scopes": [
"gmail.modify",
"gmail.send",
"drive.file",
"drive.metadata.readonly",
"spreadsheets",
"calendar",
"calendar.events",
"tasks"
],
"scheduled_tasks": [
{
"name": "morning-briefing",
"recipe": "morning-briefing",
"schedule": "0 8 * * 1-5",
"output": "~/workspace-reports/morning-{date}.json"
},
{
"name": "eod-wrap",
"recipe": "eod-wrap",
"schedule": "0 17 * * 1-5",
"output": "~/workspace-reports/eod-{date}.json"
},
{
"name": "weekly-summary",
"recipe": "weekly-summary",
"schedule": "0 9 * * 5",
"output": "~/workspace-reports/weekly-{date}.json"
},
{
"name": "security-audit",
"command": "python3 scripts/workspace_audit.py --json",
"schedule": "0 10 * * 1",
"output": "~/workspace-reports/audit-{date}.json"
}
],
"aliases": {
"inbox": "gws gmail users.messages list me --query 'is:inbox' --limit 20 --json",
"unread": "gws gmail users.messages list me --query 'is:unread' --limit 20 --json",
"files": "gws drive files list --limit 20 --json",
"events": "gws calendar events list primary --timeMin $(date -u +%Y-%m-%dT%H:%M:%SZ) --maxResults 10 --json",
"tasks": "gws tasks tasks list @default --json"
}
}
FILE:references/gws-command-reference.md
# Google Workspace CLI Command Reference
Comprehensive reference for the `gws` CLI covering 18 services, 22 helper commands, global flags, and environment variables.
---
## Global Flags
| Flag | Description |
|------|-------------|
| `--json` | Output as JSON |
| `--format ndjson` | Output as newline-delimited JSON |
| `--dry-run` | Show what would be done without executing |
| `--limit <n>` | Maximum results to return |
| `--page-all` | Fetch all pages of results |
| `--fields <spec>` | Partial response field mask |
| `--quiet` | Suppress non-error output |
| `--verbose` | Verbose debug output |
| `--timeout <ms>` | Request timeout in milliseconds |
---
## Environment Variables
| Variable | Description | Default |
|----------|-------------|---------|
| `GWS_CLIENT_ID` | OAuth client ID | — |
| `GWS_CLIENT_SECRET` | OAuth client secret | — |
| `GWS_TOKEN_PATH` | Token storage location | `~/.config/gws/token.json` |
| `GWS_SERVICE_ACCOUNT_KEY` | Service account JSON key path | — |
| `GWS_DELEGATED_USER` | User to impersonate (service accounts) | — |
| `GWS_DEFAULT_FORMAT` | Default output format | `text` |
| `GWS_PAGINATION_LIMIT` | Default pagination limit | `100` |
| `GWS_LOG_LEVEL` | Logging level (debug/info/warn/error) | `warn` |
---
## Services
### Gmail
```bash
gws gmail users.messages list me --query "<query>" --json
gws gmail users.messages get me <messageId> --json
gws gmail users.messages send me --to <email> --subject <subj> --body <body>
gws gmail users.messages reply me --thread-id <id> --body <body>
gws gmail users.messages forward me --message-id <id> --to <email>
gws gmail users.messages modify me <id> --addLabelIds <label> --removeLabelIds INBOX
gws gmail users.messages trash me <id>
gws gmail users.labels list me --json
gws gmail users.labels create me --name <name>
gws gmail users.settings.filters create me --criteria <json> --action <json>
gws gmail users.settings.forwardingAddresses list me --json
gws gmail users getProfile me --json
```
### Google Drive
```bash
gws drive files list --json --limit <n>
gws drive files list --query "name contains '<term>'" --json
gws drive files list --parents <folderId> --json
gws drive files get <fileId> --json
gws drive files create --name <name> --upload <path> --parents <folderId>
gws drive files create --name <name> --mimeType application/vnd.google-apps.folder
gws drive files update <fileId> --upload <path>
gws drive files delete <fileId>
gws drive files export <fileId> --mime <mimeType> --output <path>
gws drive files copy <fileId> --name <newName>
gws drive permissions list <fileId> --json
gws drive permissions create <fileId> --type <user|group|domain> --role <reader|writer|owner> --emailAddress <email>
gws drive permissions delete <fileId> <permissionId>
gws drive about get --json
gws drive files emptyTrash
```
### Google Sheets
```bash
gws sheets spreadsheets create --title <title> --json
gws sheets spreadsheets get <spreadsheetId> --json
gws sheets spreadsheets.values get <spreadsheetId> --range <range> --json
gws sheets spreadsheets.values update <spreadsheetId> --range <range> --values <json>
gws sheets spreadsheets.values append <spreadsheetId> --range <range> --values <json>
gws sheets spreadsheets.values clear <spreadsheetId> --range <range>
gws sheets spreadsheets.values batchGet <spreadsheetId> --ranges <range1>,<range2> --json
gws sheets spreadsheets.values batchUpdate <spreadsheetId> --data <json>
```
### Google Calendar
```bash
gws calendar calendarList list --json
gws calendar calendarList get <calendarId> --json
gws calendar events list <calendarId> --timeMin <datetime> --timeMax <datetime> --json
gws calendar events get <calendarId> <eventId> --json
gws calendar events insert <calendarId> --summary <title> --start <datetime> --end <datetime> --attendees <emails>
gws calendar events update <calendarId> <eventId> --summary <title>
gws calendar events patch <calendarId> <eventId> --start <datetime> --end <datetime>
gws calendar events delete <calendarId> <eventId>
gws calendar freebusy query --timeMin <start> --timeMax <end> --items <calendarId1>,<calendarId2> --json
```
### Google Docs
```bash
gws docs documents create --title <title> --json
gws docs documents get <documentId> --json
gws docs documents batchUpdate <documentId> --requests <json>
```
### Google Slides
```bash
gws slides presentations create --title <title> --json
gws slides presentations get <presentationId> --json
gws slides presentations.pages get <presentationId> <pageId> --json
gws slides presentations.pages getThumbnail <presentationId> <pageId> --json
```
### Google Chat
```bash
gws chat spaces list --json
gws chat spaces get <spaceName> --json
gws chat spaces.messages create <spaceName> --text <message>
gws chat spaces.messages list <spaceName> --json
gws chat spaces.messages get <messageName> --json
gws chat spaces.members list <spaceName> --json
```
### Google Tasks
```bash
gws tasks tasklists list --json
gws tasks tasklists get <tasklistId> --json
gws tasks tasklists insert --title <title> --json
gws tasks tasks list <tasklistId> --json
gws tasks tasks get <tasklistId> <taskId> --json
gws tasks tasks insert <tasklistId> --title <title> --due <datetime>
gws tasks tasks update <tasklistId> <taskId> --status completed
gws tasks tasks delete <tasklistId> <taskId>
```
### Admin SDK (Directory)
```bash
gws admin users list --domain <domain> --json
gws admin users get <email> --json
gws admin users insert --primaryEmail <email> --name.givenName <first> --name.familyName <last>
gws admin users update <email> --suspended true
gws admin groups list --domain <domain> --json
gws admin groups get <email> --json
gws admin groups insert --email <email> --name <name>
gws admin groups.members list <groupEmail> --json
gws admin groups.members insert <groupEmail> --email <memberEmail> --role MEMBER
gws admin orgunits list --customerId my_customer --json
```
### Google Groups
```bash
gws groups groups list --domain <domain> --json
gws groups groups get <email> --json
gws groups memberships list <groupEmail> --json
```
### Google People (Contacts)
```bash
gws people people.connections list me --personFields names,emailAddresses --json
gws people people get <resourceName> --personFields names,emailAddresses,phoneNumbers --json
gws people people searchContacts --query <term> --readMask names,emailAddresses --json
```
### Google Meet
```bash
gws meet spaces create --json
gws meet spaces get <spaceName> --json
gws meet conferenceRecords list --json
```
### Google Classroom
```bash
gws classroom courses list --json
gws classroom courses get <courseId> --json
gws classroom courses.courseWork list <courseId> --json
gws classroom courses.students list <courseId> --json
```
### Google Forms
```bash
gws forms forms get <formId> --json
gws forms forms.responses list <formId> --json
```
### Google Keep
```bash
gws keep notes list --json
gws keep notes get <noteId> --json
```
### Google Sites
```bash
gws sites sites list --json
gws sites sites get <siteId> --json
```
### Google Vault
```bash
gws vault matters list --json
gws vault matters get <matterId> --json
gws vault matters.holds list <matterId> --json
```
### Admin Reports / Activities
```bash
gws admin activities list <applicationName> --json
gws admin activities list login --json
gws admin activities list drive --json
gws admin activities list admin --json
```
---
## Helper Commands (22)
| Helper | Description | Example |
|--------|-------------|---------|
| `send` | Quick send email | `gws helpers send --to a@b.com --subject Hi --body Hello` |
| `reply` | Quick reply | `gws helpers reply --thread <id> --body Thanks` |
| `forward` | Quick forward | `gws helpers forward --message <id> --to a@b.com` |
| `upload` | Quick upload to Drive | `gws helpers upload file.pdf --folder <id>` |
| `download` | Quick download | `gws helpers download <fileId> --output file.pdf` |
| `share` | Quick share | `gws helpers share <fileId> --with a@b.com --role writer` |
| `quick-event` | Natural language event | `gws helpers quick-event "Lunch tomorrow at noon"` |
| `find-time` | Find free slots | `gws helpers find-time --attendees a,b --duration 60` |
| `standup-report` | Daily standup | `gws helpers standup-report` |
| `meeting-prep` | Prep for meeting | `gws helpers meeting-prep --event <id>` |
| `weekly-summary` | Week summary | `gws helpers weekly-summary` |
| `morning-briefing` | Morning overview | `gws helpers morning-briefing` |
| `eod-wrap` | End of day wrap | `gws helpers eod-wrap` |
| `inbox-zero` | Process inbox | `gws helpers inbox-zero` |
| `search` | Cross-service search | `gws helpers search "quarterly report"` |
| `create-task` | Quick task creation | `gws helpers create-task "Review PR" --due tomorrow` |
| `list-tasks` | Quick task listing | `gws helpers list-tasks` |
| `chat-send` | Quick chat message | `gws helpers chat-send --space <id> --text "Hello"` |
| `export-pdf` | Export as PDF | `gws helpers export-pdf <fileId> --output file.pdf` |
| `trash-old` | Trash old files | `gws helpers trash-old --older-than 365d` |
| `audit-sharing` | Audit file sharing | `gws helpers audit-sharing --folder <id>` |
| `backup-labels` | Backup Gmail labels | `gws helpers backup-labels --output labels.json` |
---
## Schema Introspection
```bash
# View the API schema for any service method
gws schema gmail.users.messages.list
gws schema drive.files.create
gws schema calendar.events.insert
# List all available services
gws schema --list
# List methods for a service
gws schema gmail --methods
```
---
## Authentication Commands
```bash
gws auth setup # Interactive OAuth setup
gws auth setup --service-account # Service account setup
gws auth status # Check current auth
gws auth status --json # JSON auth details
gws auth refresh # Refresh expired token
gws auth revoke # Revoke current token
gws auth switch <profile> # Switch auth profile
gws auth profiles list # List saved profiles
```
---
## Recipe Commands
```bash
gws recipes list # List all 43 recipes
gws recipes list --category email # Filter by category
gws recipes describe <name> # Show recipe details
gws recipes run <name> # Execute a recipe
gws recipes run <name> --dry-run # Preview recipe commands
```
---
## Persona Commands
```bash
gws persona list # List all 10 personas
gws persona select <name> # Activate a persona
gws persona show # Show active persona
gws persona recipes # Show recipes for active persona
```
FILE:references/recipes-cookbook.md
# Google Workspace CLI Recipes Cookbook
Complete catalog of 43 built-in recipes organized by category, with command sequences and persona mapping.
---
## Recipe Categories
| Category | Count | Description |
|----------|-------|-------------|
| Email | 8 | Gmail operations — send, search, label, filter |
| Files | 7 | Drive file management — upload, share, export |
| Calendar | 6 | Events, scheduling, meeting prep |
| Reporting | 5 | Activity summaries and analytics |
| Collaboration | 5 | Chat, Docs, Tasks teamwork |
| Data | 4 | Sheets read/write and contacts |
| Admin | 4 | User and group management |
| Cross-Service | 4 | Multi-service workflows |
---
## Email Recipes (8)
### send-email
Send an email with optional attachments.
```bash
gws gmail users.messages send me --to "recipient@example.com" \
--subject "Subject" --body "Body text" [--attachment file.pdf]
```
### reply-to-thread
Reply to an existing email thread.
```bash
gws gmail users.messages reply me --thread-id <THREAD_ID> --body "Reply text"
```
### forward-email
Forward an email to another recipient.
```bash
gws gmail users.messages forward me --message-id <MSG_ID> --to "forward@example.com"
```
### search-emails
Search emails using Gmail query syntax.
```bash
gws gmail users.messages list me --query "from:sender@example.com after:2025/01/01" --json
```
**Query examples:** `is:unread`, `has:attachment`, `label:important`, `newer_than:7d`
### archive-old
Archive read emails older than N days.
```bash
gws gmail users.messages list me --query "is:read older_than:30d" --json
# Extract IDs, then batch modify to remove INBOX label
```
### label-manager
Create and organize Gmail labels.
```bash
gws gmail users.labels list me --json
gws gmail users.labels create me --name "Projects/Alpha"
```
### filter-setup
Create auto-labeling filters.
```bash
gws gmail users.settings.filters create me \
--criteria '{"from":"notifications@service.com"}' \
--action '{"addLabelIds":["Label_123"],"removeLabelIds":["INBOX"]}'
```
### unread-digest
Get digest of unread emails.
```bash
gws gmail users.messages list me --query "is:unread" --limit 20 --json
```
---
## Files Recipes (7)
### upload-file
Upload a file to Google Drive.
```bash
gws drive files create --name "Report Q1" --upload report.pdf --parents <FOLDER_ID>
```
### create-sheet
Create a new Google Spreadsheet.
```bash
gws sheets spreadsheets create --title "Budget 2026" --json
```
### share-file
Share a Drive file with a user or domain.
```bash
gws drive permissions create <FILE_ID> --type user --role writer --emailAddress "user@example.com"
```
### export-file
Export a Google Doc/Sheet as PDF.
```bash
gws drive files export <FILE_ID> --mime "application/pdf" --output report.pdf
```
### list-files
List files in a Drive folder.
```bash
gws drive files list --parents <FOLDER_ID> --json
```
### find-large-files
Find the largest files in Drive.
```bash
gws drive files list --orderBy "quotaBytesUsed desc" --limit 20 --json
```
### cleanup-trash
Empty Drive trash.
```bash
gws drive files emptyTrash
```
---
## Calendar Recipes (6)
### create-event
Create a calendar event with attendees.
```bash
gws calendar events insert primary \
--summary "Sprint Planning" \
--start "2026-03-15T10:00:00" --end "2026-03-15T11:00:00" \
--attendees "team@company.com" --location "Room A"
```
### quick-event
Create event from natural language.
```bash
gws helpers quick-event "Lunch with Sarah tomorrow at noon"
```
### find-time
Find available time slots for a meeting.
```bash
gws helpers find-time --attendees "alice@co.com,bob@co.com" --duration 60 \
--within "2026-03-15,2026-03-19" --json
```
### today-schedule
Show today's calendar events.
```bash
gws calendar events list primary \
--timeMin "$(date -u +%Y-%m-%dT00:00:00Z)" \
--timeMax "$(date -u +%Y-%m-%dT23:59:59Z)" --json
```
### meeting-prep
Prepare for an upcoming meeting.
```bash
gws recipes meeting-prep --event-id <EVENT_ID>
```
**Output:** Agenda, attendee list, related Drive files, previous meeting notes.
### reschedule
Move an event to a new time.
```bash
gws calendar events patch primary <EVENT_ID> \
--start "2026-03-16T14:00:00" --end "2026-03-16T15:00:00"
```
---
## Reporting Recipes (5)
### standup-report
Generate daily standup from calendar and tasks.
```bash
gws recipes standup-report --json
```
**Output:** Yesterday's events, today's schedule, pending tasks, blockers.
### weekly-summary
Summarize week's emails, events, and tasks.
```bash
gws recipes weekly-summary --json
```
### drive-activity
Report on Drive file activity.
```bash
gws drive activities list --json
```
### email-stats
Email volume statistics for the past 7 days.
```bash
gws gmail users.messages list me --query "newer_than:7d" --json | python3 output_analyzer.py --count
```
### task-progress
Report on task completion.
```bash
gws tasks tasks list <TASKLIST_ID> --json | python3 output_analyzer.py --group-by "status"
```
---
## Collaboration Recipes (5)
### share-folder
Share a Drive folder with a team.
```bash
gws drive permissions create <FOLDER_ID> --type group --role writer --emailAddress "team@company.com"
```
### create-doc
Create a Google Doc with initial content.
```bash
gws docs documents create --title "Meeting Notes - March 15" --json
```
### chat-message
Send a message to a Google Chat space.
```bash
gws chat spaces.messages create <SPACE_NAME> --text "Deployment complete!"
```
### list-spaces
List Google Chat spaces.
```bash
gws chat spaces list --json
```
### task-create
Create a task in Google Tasks.
```bash
gws tasks tasks insert <TASKLIST_ID> --title "Review PR #42" --due "2026-03-16"
```
---
## Data Recipes (4)
### sheet-read
Read data from a spreadsheet range.
```bash
gws sheets spreadsheets.values get <SHEET_ID> --range "Sheet1!A1:D10" --json
```
### sheet-write
Write data to a spreadsheet.
```bash
gws sheets spreadsheets.values update <SHEET_ID> --range "Sheet1!A1" \
--values '[["Name","Score"],["Alice",95],["Bob",87]]'
```
### sheet-append
Append rows to a spreadsheet.
```bash
gws sheets spreadsheets.values append <SHEET_ID> --range "Sheet1!A1" \
--values '[["Charlie",92]]'
```
### export-contacts
Export contacts list.
```bash
gws people people.connections list me --personFields names,emailAddresses --json
```
---
## Admin Recipes (4)
### list-users
List all users in the Workspace domain.
```bash
gws admin users list --domain company.com --json
```
**Prerequisites:** Admin SDK API enabled, `admin.directory.user.readonly` scope.
### list-groups
List all groups in the domain.
```bash
gws admin groups list --domain company.com --json
```
### user-info
Get detailed user information.
```bash
gws admin users get user@company.com --json
```
### audit-logins
Audit recent login activity.
```bash
gws admin activities list login --json
```
---
## Cross-Service Recipes (4)
### morning-briefing
Today's events + unread emails + pending tasks.
```bash
gws recipes morning-briefing --json
```
**Combines:** Calendar events, Gmail unread count, Tasks pending.
### eod-wrap
End-of-day summary: completed, pending, tomorrow's schedule.
```bash
gws recipes eod-wrap --json
```
### project-status
Aggregate project status from Drive, Sheets, Tasks.
```bash
gws recipes project-status --project "Project Alpha" --json
```
### inbox-zero
Process inbox to zero: label, archive, reply, or create task.
```bash
gws recipes inbox-zero --interactive
```
---
## Persona Mapping
| Persona | Top Recipes |
|---------|-------------|
| Executive Assistant | morning-briefing, today-schedule, find-time, send-email, meeting-prep, eod-wrap |
| Project Manager | standup-report, create-event, find-time, task-create, project-status, weekly-summary |
| HR | list-users, user-info, send-email, create-event, create-doc, export-contacts |
| Sales | send-email, search-emails, create-event, find-time, create-doc, share-file |
| IT Admin | list-users, list-groups, audit-logins, drive-activity, find-large-files, cleanup-trash |
| Developer | sheet-read, sheet-write, upload-file, chat-message, task-create, send-email |
| Marketing | send-email, create-doc, share-file, upload-file, create-sheet, chat-message |
| Finance | sheet-read, sheet-write, sheet-append, create-sheet, export-file, share-file |
| Legal | create-doc, share-file, export-file, search-emails, upload-file, audit-logins |
| Customer Support | search-emails, send-email, reply-to-thread, label-manager, task-create, inbox-zero |
FILE:references/troubleshooting.md
# Google Workspace CLI Troubleshooting
Common errors, fixes, and platform-specific guidance for the `gws` CLI.
---
## Installation Issues
### gws not found on PATH
**Error:** `command not found: gws`
**Fixes:**
```bash
# Check if installed
npm list -g @anthropic/gws 2>/dev/null || echo "Not installed via npm"
which gws || echo "Not on PATH"
# Install via npm
npm install -g @anthropic/gws
# If npm global bin not on PATH
export PATH="$(npm config get prefix)/bin:$PATH"
# Add to ~/.zshrc or ~/.bashrc for persistence
```
### npm permission errors
**Error:** `EACCES: permission denied`
**Fixes:**
```bash
# Option 1: Fix npm prefix (recommended)
mkdir -p ~/.npm-global
npm config set prefix '~/.npm-global'
export PATH=~/.npm-global/bin:$PATH
# Option 2: Use npx without installing
npx @anthropic/gws --version
```
### Cargo build failures
**Error:** `error[E0463]: can't find crate`
**Fixes:**
```bash
# Ensure Rust is up to date
rustup update stable
# Clean build
cargo clean && cargo install gws-cli
```
---
## Authentication Errors
### Token expired
**Error:** `401 Unauthorized: Token has been expired or revoked`
**Cause:** OAuth tokens expire after 1 hour.
**Fix:**
```bash
gws auth refresh
# If refresh fails:
gws auth setup # Re-authenticate
```
### Insufficient scopes
**Error:** `403 Forbidden: Request had insufficient authentication scopes`
**Fix:**
```bash
# Check current scopes
gws auth status --json | grep scopes
# Re-auth with additional scopes
gws auth setup --scopes gmail,drive,calendar,sheets,tasks
# Or list required scopes for a service
python3 scripts/auth_setup_guide.py --scopes gmail,drive
```
### Keyring/keychain errors
**Error:** `Failed to access keyring` or `SecKeychainFindGenericPassword failed`
**Fixes:**
```bash
# macOS: Unlock keychain
security unlock-keychain ~/Library/Keychains/login.keychain-db
# Linux: Install keyring backend
sudo apt install gnome-keyring # or libsecret
# Fallback: Use file-based token storage
export GWS_TOKEN_PATH=~/.config/gws/token.json
gws auth setup
```
### Service account delegation errors
**Error:** `403: Not Authorized to access this resource/api`
**Fix:**
1. Verify domain-wide delegation is enabled on the service account
2. Verify client ID is authorized in Admin Console > Security > API Controls
3. Verify scopes match exactly (no trailing slashes)
4. Verify `GWS_DELEGATED_USER` is a valid admin account
```bash
# Debug
echo $GWS_SERVICE_ACCOUNT_KEY # Should point to valid JSON key file
echo $GWS_DELEGATED_USER # Should be admin@yourdomain.com
gws auth status --json # Check auth details
```
---
## API Errors
### Rate limit exceeded (429)
**Error:** `429 Too Many Requests: Rate Limit Exceeded`
**Cause:** Google Workspace APIs have per-user, per-service rate limits.
**Fix:**
```bash
# Add delays between bulk operations
for id in $(cat file_ids.txt); do
gws drive files get $id --json >> results.json
sleep 0.5 # 500ms delay
done
# Use --limit to reduce result size
gws drive files list --limit 100 --json
# For admin operations, batch in groups of 50
```
**Rate limits by service:**
| Service | Limit |
|---------|-------|
| Gmail | 250 quota units/second/user |
| Drive | 1,000 requests/100 seconds/user |
| Sheets | 60 read requests/minute/user |
| Calendar | 500 requests/100 seconds/user |
| Admin SDK | 2,400 requests/minute |
### Permission denied (403)
**Error:** `403 Forbidden: The caller does not have permission`
**Causes and fixes:**
1. **Wrong scope** — Re-auth with correct scopes
2. **Not the file owner** — Request access from the owner
3. **Domain policy** — Check Admin Console sharing policies
4. **API not enabled** — Enable the API in Google Cloud Console
```bash
# Check which APIs are enabled
gws schema --list
# Enable an API
# Go to: console.cloud.google.com > APIs & Services > Library
```
### Not found (404)
**Error:** `404 Not Found: File not found`
**Causes:**
1. File was deleted or moved to trash
2. File ID is incorrect
3. No permission to see the file
```bash
# Check trash
gws drive files list --query "trashed=true and name='filename'" --json
# Verify file ID
gws drive files get <fileId> --json
```
---
## Output Parsing Issues
### NDJSON vs JSON array
**Problem:** Output format varies between commands and versions.
```bash
# Force JSON array output
gws drive files list --json
# Force NDJSON output
gws drive files list --format ndjson
# Handle both in output_analyzer.py (automatic detection)
gws drive files list --json | python3 scripts/output_analyzer.py --count
```
### Pagination
**Problem:** Only partial results returned.
```bash
# Fetch all pages
gws drive files list --page-all --json
# Or set a high limit
gws drive files list --limit 1000 --json
# Check if more pages exist (look for nextPageToken in output)
gws drive files list --limit 100 --json | grep nextPageToken
```
### Empty response
**Problem:** Command returns empty or `{}`.
```bash
# Check auth
gws auth status
# Try with verbose output
gws drive files list --verbose --json
# Check if the service is accessible
gws drive about get --json
```
---
## Platform-Specific Issues
### macOS
**Keychain access prompts:**
```bash
# Allow gws to access keychain without repeated prompts
# In Keychain Access.app, find "gws" entries and set "Allow all applications"
# Or use file-based storage
export GWS_TOKEN_PATH=~/.config/gws/token.json
```
**Browser not opening for OAuth:**
```bash
# If default browser doesn't open
gws auth setup --no-browser
# Copy the URL manually and paste in browser
```
### Linux
**Headless OAuth (no browser):**
```bash
# Use out-of-band flow
gws auth setup --no-browser
# Prints a URL — open on another machine, paste code back
# Or use service account (no browser needed)
export GWS_SERVICE_ACCOUNT_KEY=/path/to/key.json
export GWS_DELEGATED_USER=admin@domain.com
```
**Missing keyring backend:**
```bash
# Install a keyring backend
sudo apt install gnome-keyring libsecret-1-dev
# Or use file-based storage
export GWS_TOKEN_PATH=~/.config/gws/token.json
```
### Windows
**PATH issues:**
```powershell
# Add npm global bin to PATH
$env:PATH += ";$(npm config get prefix)\bin"
# Or use npx
npx @anthropic/gws --version
```
**PowerShell quoting:**
```powershell
# Use single quotes for JSON arguments
gws gmail users.settings.filters create me `
--criteria '{"from":"test@example.com"}' `
--action '{"addLabelIds":["Label_1"]}'
```
---
## Getting Help
```bash
# General help
gws --help
gws <service> --help
gws <service> <resource> --help
# API schema for a method
gws schema gmail.users.messages.send
# Version info
gws --version
# Debug mode
gws --verbose <command>
# Report issues
# https://github.com/googleworkspace/cli/issues
```
FILE:scripts/auth_setup_guide.py
#!/usr/bin/env python3
"""
Google Workspace CLI Auth Setup Guide — Guided authentication configuration.
Prints step-by-step instructions for OAuth and service account setup,
generates .env templates, lists required scopes, and validates auth.
Usage:
python3 auth_setup_guide.py --guide oauth
python3 auth_setup_guide.py --guide service-account
python3 auth_setup_guide.py --scopes gmail,drive,calendar
python3 auth_setup_guide.py --generate-env
python3 auth_setup_guide.py --validate [--json]
python3 auth_setup_guide.py --check [--json]
"""
import argparse
import json
import shutil
import subprocess
import sys
from dataclasses import dataclass, field, asdict
from typing import List, Dict
SERVICE_SCOPES: Dict[str, List[str]] = {
"gmail": [
"https://www.googleapis.com/auth/gmail.modify",
"https://www.googleapis.com/auth/gmail.send",
"https://www.googleapis.com/auth/gmail.labels",
"https://www.googleapis.com/auth/gmail.settings.basic",
],
"drive": [
"https://www.googleapis.com/auth/drive",
"https://www.googleapis.com/auth/drive.file",
"https://www.googleapis.com/auth/drive.metadata.readonly",
],
"sheets": [
"https://www.googleapis.com/auth/spreadsheets",
],
"calendar": [
"https://www.googleapis.com/auth/calendar",
"https://www.googleapis.com/auth/calendar.events",
],
"tasks": [
"https://www.googleapis.com/auth/tasks",
],
"chat": [
"https://www.googleapis.com/auth/chat.spaces.readonly",
"https://www.googleapis.com/auth/chat.messages",
],
"docs": [
"https://www.googleapis.com/auth/documents",
],
"admin": [
"https://www.googleapis.com/auth/admin.directory.user.readonly",
"https://www.googleapis.com/auth/admin.directory.group",
"https://www.googleapis.com/auth/admin.directory.orgunit.readonly",
],
"meet": [
"https://www.googleapis.com/auth/meetings.space.created",
],
}
OAUTH_GUIDE = """
=== Google Workspace CLI: OAuth Setup Guide ===
Step 1: Create a Google Cloud Project
1. Go to https://console.cloud.google.com/
2. Click "Select a project" -> "New Project"
3. Name it (e.g., "gws-cli-access") and click Create
4. Note the Project ID
Step 2: Enable Required APIs
1. Go to APIs & Services -> Library
2. Search and enable each API you need:
- Gmail API
- Google Drive API
- Google Sheets API
- Google Calendar API
- Tasks API
- Admin SDK API (for admin operations)
Step 3: Configure OAuth Consent Screen
1. Go to APIs & Services -> OAuth consent screen
2. Select "Internal" (for Workspace) or "External" (for personal)
3. Fill in app name, support email
4. Add scopes for the services you need
5. Save and continue
Step 4: Create OAuth Credentials
1. Go to APIs & Services -> Credentials
2. Click "Create Credentials" -> "OAuth client ID"
3. Application type: "Desktop app"
4. Name it "gws-cli"
5. Download the JSON file
Step 5: Configure gws CLI
1. Set environment variables:
export GWS_CLIENT_ID=<your-client-id>
export GWS_CLIENT_SECRET=<your-client-secret>
2. Or place the credentials JSON:
mv client_secret_*.json ~/.config/gws/credentials.json
Step 6: Authenticate
gws auth setup
# Opens browser for consent, stores token in system keyring
Step 7: Verify
gws auth status
gws gmail users getProfile me
"""
SERVICE_ACCOUNT_GUIDE = """
=== Google Workspace CLI: Service Account Setup Guide ===
Step 1: Create a Google Cloud Project
(Same as OAuth Step 1)
Step 2: Create a Service Account
1. Go to IAM & Admin -> Service Accounts
2. Click "Create Service Account"
3. Name: "gws-cli-service"
4. Grant roles as needed (no role needed for Workspace API access)
5. Click "Done"
Step 3: Create Key
1. Click on the service account
2. Go to "Keys" tab
3. Add Key -> Create new key -> JSON
4. Download and store securely
Step 4: Enable Domain-Wide Delegation
1. On the service account page, click "Edit"
2. Check "Enable Google Workspace domain-wide delegation"
3. Save
4. Note the Client ID (numeric)
Step 5: Authorize in Google Admin
1. Go to admin.google.com
2. Security -> API Controls -> Domain-wide Delegation
3. Add new:
- Client ID: <numeric client ID from Step 4>
- Scopes: (paste required scopes)
4. Authorize
Step 6: Configure gws CLI
export GWS_SERVICE_ACCOUNT_KEY=/path/to/service-account-key.json
export GWS_DELEGATED_USER=admin@yourdomain.com
Step 7: Verify
gws auth status
gws gmail users getProfile me
"""
ENV_TEMPLATE = """# Google Workspace CLI Configuration
# Copy to .env and fill in values
# OAuth Credentials (for interactive auth)
GWS_CLIENT_ID=
GWS_CLIENT_SECRET=
GWS_TOKEN_PATH=~/.config/gws/token.json
# Service Account (for headless/CI auth)
# GWS_SERVICE_ACCOUNT_KEY=/path/to/key.json
# GWS_DELEGATED_USER=admin@yourdomain.com
# Defaults
GWS_DEFAULT_FORMAT=json
GWS_PAGINATION_LIMIT=100
"""
@dataclass
class ValidationResult:
service: str
status: str # PASS, FAIL
message: str
@dataclass
class ValidationReport:
auth_method: str = ""
user: str = ""
results: List[dict] = field(default_factory=list)
summary: str = ""
demo_mode: bool = False
DEMO_VALIDATION = ValidationReport(
auth_method="oauth",
user="admin@company.com",
results=[
{"service": "gmail", "status": "PASS", "message": "Gmail API accessible"},
{"service": "drive", "status": "PASS", "message": "Drive API accessible"},
{"service": "calendar", "status": "PASS", "message": "Calendar API accessible"},
{"service": "sheets", "status": "PASS", "message": "Sheets API accessible"},
{"service": "tasks", "status": "FAIL", "message": "Scope not authorized"},
],
summary="4/5 services validated (demo mode)",
demo_mode=True,
)
def check_auth_status() -> dict:
"""Check current gws auth status."""
try:
result = subprocess.run(
["gws", "auth", "status", "--json"],
capture_output=True, text=True, timeout=15
)
if result.returncode == 0:
try:
return json.loads(result.stdout)
except json.JSONDecodeError:
return {"status": "authenticated", "raw": result.stdout.strip()}
return {"status": "not_authenticated", "error": result.stderr.strip()[:200]}
except (FileNotFoundError, OSError):
return {"status": "gws_not_found"}
def validate_services(services: List[str]) -> ValidationReport:
"""Validate auth by testing each service."""
report = ValidationReport()
auth = check_auth_status()
if auth.get("status") == "gws_not_found":
report.summary = "gws CLI not installed"
return report
if auth.get("status") == "not_authenticated":
report.auth_method = "none"
report.summary = "Not authenticated"
return report
report.auth_method = auth.get("method", "oauth")
report.user = auth.get("user", auth.get("email", "unknown"))
service_cmds = {
"gmail": ["gws", "gmail", "users", "getProfile", "me", "--json"],
"drive": ["gws", "drive", "files", "list", "--limit", "1", "--json"],
"calendar": ["gws", "calendar", "calendarList", "list", "--limit", "1", "--json"],
"sheets": ["gws", "sheets", "spreadsheets", "get", "test", "--json"],
"tasks": ["gws", "tasks", "tasklists", "list", "--limit", "1", "--json"],
}
for svc in services:
cmd = service_cmds.get(svc)
if not cmd:
report.results.append(asdict(
ValidationResult(svc, "WARN", f"No test available for {svc}")
))
continue
try:
result = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
if result.returncode == 0:
report.results.append(asdict(
ValidationResult(svc, "PASS", f"{svc.title()} API accessible")
))
else:
report.results.append(asdict(
ValidationResult(svc, "FAIL", result.stderr.strip()[:100])
))
except (subprocess.TimeoutExpired, OSError) as e:
report.results.append(asdict(
ValidationResult(svc, "FAIL", str(e)[:100])
))
passed = sum(1 for r in report.results if r["status"] == "PASS")
total = len(report.results)
report.summary = f"{passed}/{total} services validated"
return report
def main():
parser = argparse.ArgumentParser(
description="Guided authentication setup for Google Workspace CLI (gws)",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
%(prog)s --guide oauth # OAuth setup instructions
%(prog)s --guide service-account # Service account setup
%(prog)s --scopes gmail,drive # Show required scopes
%(prog)s --generate-env # Generate .env template
%(prog)s --check # Check current auth status
%(prog)s --validate --json # Validate all services (JSON)
""",
)
parser.add_argument("--guide", choices=["oauth", "service-account"],
help="Print setup guide")
parser.add_argument("--scopes", help="Comma-separated services to show scopes for")
parser.add_argument("--generate-env", action="store_true",
help="Generate .env template")
parser.add_argument("--check", action="store_true",
help="Check current auth status")
parser.add_argument("--validate", action="store_true",
help="Validate auth by testing services")
parser.add_argument("--services", default="gmail,drive,calendar,sheets,tasks",
help="Services to validate (default: gmail,drive,calendar,sheets,tasks)")
parser.add_argument("--json", action="store_true", help="Output JSON")
args = parser.parse_args()
if not any([args.guide, args.scopes, args.generate_env, args.check, args.validate]):
parser.print_help()
return
if args.guide:
if args.guide == "oauth":
print(OAUTH_GUIDE)
else:
print(SERVICE_ACCOUNT_GUIDE)
return
if args.scopes:
services = [s.strip() for s in args.scopes.split(",") if s.strip()]
if args.json:
output = {}
for svc in services:
output[svc] = SERVICE_SCOPES.get(svc, [])
print(json.dumps(output, indent=2))
else:
print(f"\n{'='*60}")
print(f" REQUIRED OAUTH SCOPES")
print(f"{'='*60}\n")
for svc in services:
scopes = SERVICE_SCOPES.get(svc, [])
print(f" {svc.upper()}:")
if scopes:
for scope in scopes:
print(f" - {scope}")
else:
print(f" (no scopes defined for '{svc}')")
print()
# Print combined for easy copy-paste
all_scopes = []
for svc in services:
all_scopes.extend(SERVICE_SCOPES.get(svc, []))
if all_scopes:
print(f" COMBINED (for consent screen):")
print(f" {','.join(all_scopes)}")
print(f"\n{'='*60}\n")
return
if args.generate_env:
print(ENV_TEMPLATE)
return
if args.check:
if shutil.which("gws"):
status = check_auth_status()
else:
status = {"status": "gws_not_found",
"note": "Install gws first: cargo install gws-cli OR https://github.com/googleworkspace/cli/releases"}
if args.json:
print(json.dumps(status, indent=2))
else:
print(f"\nAuth Status: {status.get('status', 'unknown')}")
for k, v in status.items():
if k != "status":
print(f" {k}: {v}")
print()
return
if args.validate:
services = [s.strip() for s in args.services.split(",") if s.strip()]
if not shutil.which("gws"):
report = DEMO_VALIDATION
else:
report = validate_services(services)
if args.json:
print(json.dumps(asdict(report), indent=2))
else:
print(f"\n{'='*60}")
print(f" AUTH VALIDATION REPORT")
if report.demo_mode:
print(f" (DEMO MODE)")
print(f"{'='*60}\n")
if report.user:
print(f" User: {report.user}")
print(f" Method: {report.auth_method}\n")
for r in report.results:
icon = "PASS" if r["status"] == "PASS" else "FAIL"
print(f" [{icon}] {r['service']}: {r['message']}")
print(f"\n {report.summary}")
print(f"\n{'='*60}\n")
if __name__ == "__main__":
main()
FILE:scripts/gws_doctor.py
#!/usr/bin/env python3
"""
Google Workspace CLI Doctor — Pre-flight diagnostics for gws CLI.
Checks installation, version, authentication status, and service
connectivity. Runs in demo mode with embedded sample data when gws
is not installed.
Usage:
python3 gws_doctor.py
python3 gws_doctor.py --json
python3 gws_doctor.py --services gmail,drive,calendar
"""
import argparse
import json
import shutil
import subprocess
import sys
from dataclasses import dataclass, field, asdict
from typing import List, Optional
@dataclass
class Check:
name: str
status: str # PASS, WARN, FAIL
message: str
fix: str = ""
@dataclass
class DiagnosticReport:
gws_installed: bool = False
gws_version: str = ""
auth_status: str = ""
checks: List[dict] = field(default_factory=list)
summary: str = ""
demo_mode: bool = False
DEMO_CHECKS = [
Check("gws-installed", "PASS", "gws v0.9.2 found at /usr/local/bin/gws"),
Check("gws-version", "PASS", "Version 0.9.2 (latest)"),
Check("auth-status", "PASS", "Authenticated as admin@company.com"),
Check("token-expiry", "WARN", "Token expires in 23 minutes",
"Run 'gws auth refresh' to extend token lifetime"),
Check("gmail-access", "PASS", "Gmail API accessible — user profile retrieved"),
Check("drive-access", "PASS", "Drive API accessible — root folder listed"),
Check("calendar-access", "PASS", "Calendar API accessible — primary calendar found"),
Check("sheets-access", "PASS", "Sheets API accessible"),
Check("tasks-access", "FAIL", "Tasks API not authorized",
"Run 'gws auth setup' and add 'tasks' scope"),
]
SERVICE_TEST_COMMANDS = {
"gmail": ["gws", "gmail", "users", "getProfile", "me", "--json"],
"drive": ["gws", "drive", "files", "list", "--limit", "1", "--json"],
"calendar": ["gws", "calendar", "calendarList", "list", "--limit", "1", "--json"],
"sheets": ["gws", "sheets", "spreadsheets", "get", "test", "--json"],
"tasks": ["gws", "tasks", "tasklists", "list", "--limit", "1", "--json"],
"chat": ["gws", "chat", "spaces", "list", "--limit", "1", "--json"],
"docs": ["gws", "docs", "documents", "get", "test", "--json"],
}
def check_installation() -> Check:
"""Check if gws is installed and on PATH."""
path = shutil.which("gws")
if path:
return Check("gws-installed", "PASS", f"gws found at {path}")
return Check("gws-installed", "FAIL", "gws not found on PATH",
"Install via: cargo install gws-cli OR download from https://github.com/googleworkspace/cli/releases")
def check_version() -> Check:
"""Get gws version."""
try:
result = subprocess.run(
["gws", "--version"], capture_output=True, text=True, timeout=10
)
version = result.stdout.strip()
if version:
return Check("gws-version", "PASS", f"Version: {version}")
return Check("gws-version", "WARN", "Could not parse version output")
except (subprocess.TimeoutExpired, FileNotFoundError, OSError) as e:
return Check("gws-version", "FAIL", f"Version check failed: {e}")
def check_auth() -> Check:
"""Check authentication status."""
try:
result = subprocess.run(
["gws", "auth", "status", "--json"],
capture_output=True, text=True, timeout=15
)
if result.returncode == 0:
try:
data = json.loads(result.stdout)
user = data.get("user", data.get("email", "unknown"))
return Check("auth-status", "PASS", f"Authenticated as {user}")
except json.JSONDecodeError:
return Check("auth-status", "PASS", "Authenticated (could not parse details)")
return Check("auth-status", "FAIL", "Not authenticated",
"Run 'gws auth setup' to configure authentication")
except (subprocess.TimeoutExpired, FileNotFoundError, OSError) as e:
return Check("auth-status", "FAIL", f"Auth check failed: {e}",
"Run 'gws auth setup' to configure authentication")
def check_service(service: str) -> Check:
"""Test connectivity to a specific service."""
cmd = SERVICE_TEST_COMMANDS.get(service)
if not cmd:
return Check(f"{service}-access", "WARN", f"No test command for {service}")
try:
result = subprocess.run(cmd, capture_output=True, text=True, timeout=15)
if result.returncode == 0:
return Check(f"{service}-access", "PASS", f"{service.title()} API accessible")
stderr = result.stderr.strip()[:100]
if "403" in stderr or "permission" in stderr.lower():
return Check(f"{service}-access", "FAIL",
f"{service.title()} API permission denied",
f"Add '{service}' scope: gws auth setup --scopes {service}")
return Check(f"{service}-access", "FAIL",
f"{service.title()} API error: {stderr}",
f"Check scope and permissions for {service}")
except (subprocess.TimeoutExpired, FileNotFoundError, OSError) as e:
return Check(f"{service}-access", "FAIL", f"{service.title()} test failed: {e}")
def run_diagnostics(services: List[str]) -> DiagnosticReport:
"""Run all diagnostic checks."""
report = DiagnosticReport()
checks = []
# Installation check
install_check = check_installation()
checks.append(install_check)
report.gws_installed = install_check.status == "PASS"
if not report.gws_installed:
report.checks = [asdict(c) for c in checks]
report.summary = "FAIL: gws is not installed"
return report
# Version check
version_check = check_version()
checks.append(version_check)
if version_check.status == "PASS":
report.gws_version = version_check.message.replace("Version: ", "")
# Auth check
auth_check = check_auth()
checks.append(auth_check)
report.auth_status = auth_check.status
if auth_check.status != "PASS":
report.checks = [asdict(c) for c in checks]
report.summary = "FAIL: Authentication not configured"
return report
# Service checks
for svc in services:
checks.append(check_service(svc))
report.checks = [asdict(c) for c in checks]
# Summary
fails = sum(1 for c in checks if c.status == "FAIL")
warns = sum(1 for c in checks if c.status == "WARN")
passes = sum(1 for c in checks if c.status == "PASS")
if fails > 0:
report.summary = f"ISSUES FOUND: {passes} passed, {warns} warnings, {fails} failures"
elif warns > 0:
report.summary = f"MOSTLY OK: {passes} passed, {warns} warnings"
else:
report.summary = f"ALL CLEAR: {passes}/{passes} checks passed"
return report
def run_demo() -> DiagnosticReport:
"""Return demo report with embedded sample data."""
report = DiagnosticReport(
gws_installed=True,
gws_version="0.9.2",
auth_status="PASS",
checks=[asdict(c) for c in DEMO_CHECKS],
summary="MOSTLY OK: 7 passed, 1 warning, 1 failure (demo mode)",
demo_mode=True,
)
return report
def main():
parser = argparse.ArgumentParser(
description="Pre-flight diagnostics for Google Workspace CLI (gws)",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
%(prog)s # Run all checks
%(prog)s --json # JSON output
%(prog)s --services gmail,drive # Check specific services only
%(prog)s --demo # Demo mode (no gws required)
""",
)
parser.add_argument("--json", action="store_true", help="Output JSON")
parser.add_argument(
"--services", default="gmail,drive,calendar,sheets,tasks",
help="Comma-separated services to check (default: gmail,drive,calendar,sheets,tasks)"
)
parser.add_argument("--demo", action="store_true", help="Run with demo data")
args = parser.parse_args()
services = [s.strip() for s in args.services.split(",") if s.strip()]
# Use demo mode if requested or gws not installed
if args.demo or not shutil.which("gws"):
report = run_demo()
else:
report = run_diagnostics(services)
if args.json:
print(json.dumps(asdict(report), indent=2))
else:
print(f"\n{'='*60}")
print(f" GWS CLI DIAGNOSTIC REPORT")
if report.demo_mode:
print(f" (DEMO MODE — sample data)")
print(f"{'='*60}\n")
for c in report.checks:
icon = {"PASS": "PASS", "WARN": "WARN", "FAIL": "FAIL"}.get(c["status"], "????")
print(f" [{icon}] {c['name']}: {c['message']}")
if c.get("fix") and c["status"] != "PASS":
print(f" -> {c['fix']}")
print(f"\n {'-'*56}")
print(f" {report.summary}")
print(f"\n{'='*60}\n")
if __name__ == "__main__":
main()
FILE:scripts/gws_recipe_runner.py
#!/usr/bin/env python3
"""
Google Workspace CLI Recipe Runner — Catalog, search, and execute gws recipes.
Browse 43 built-in recipes, filter by persona, search by keyword,
and run with dry-run support.
Usage:
python3 gws_recipe_runner.py --list
python3 gws_recipe_runner.py --search "email"
python3 gws_recipe_runner.py --describe standup-report
python3 gws_recipe_runner.py --run standup-report --dry-run
python3 gws_recipe_runner.py --persona pm --list
python3 gws_recipe_runner.py --list --json
"""
import argparse
import json
import subprocess
import sys
from dataclasses import dataclass, field, asdict
from typing import List, Dict, Optional
@dataclass
class Recipe:
name: str
description: str
category: str
services: List[str]
commands: List[str]
prerequisites: str = ""
RECIPES: Dict[str, Recipe] = {
# Email (8)
"send-email": Recipe("send-email", "Send an email with optional attachments", "email",
["gmail"], ["gws gmail users.messages send me --to {to} --subject {subject} --body {body}"]),
"reply-to-thread": Recipe("reply-to-thread", "Reply to an existing email thread", "email",
["gmail"], ["gws gmail users.messages reply me --thread-id {thread_id} --body {body}"]),
"forward-email": Recipe("forward-email", "Forward an email to another recipient", "email",
["gmail"], ["gws gmail users.messages forward me --message-id {msg_id} --to {to}"]),
"search-emails": Recipe("search-emails", "Search emails with Gmail query syntax", "email",
["gmail"], ["gws gmail users.messages list me --query {query} --json"]),
"archive-old": Recipe("archive-old", "Archive read emails older than N days", "email",
["gmail"], [
"gws gmail users.messages list me --query 'is:read older_than:{days}d' --json",
"# Pipe IDs to batch modify to remove INBOX label",
]),
"label-manager": Recipe("label-manager", "Create, list, and organize Gmail labels", "email",
["gmail"], ["gws gmail users.labels list me --json", "gws gmail users.labels create me --name {name}"]),
"filter-setup": Recipe("filter-setup", "Create email filters for auto-labeling", "email",
["gmail"], ["gws gmail users.settings.filters create me --criteria {criteria} --action {action}"]),
"unread-digest": Recipe("unread-digest", "Get digest of unread emails", "email",
["gmail"], ["gws gmail users.messages list me --query 'is:unread' --limit 20 --json"]),
# Files (7)
"upload-file": Recipe("upload-file", "Upload a file to Google Drive", "files",
["drive"], ["gws drive files create --name {name} --upload {path} --parents {folder_id}"]),
"create-sheet": Recipe("create-sheet", "Create a new Google Spreadsheet", "files",
["sheets"], ["gws sheets spreadsheets create --title {title} --json"]),
"share-file": Recipe("share-file", "Share a Drive file with a user or domain", "files",
["drive"], ["gws drive permissions create {file_id} --type user --role writer --emailAddress {email}"]),
"export-file": Recipe("export-file", "Export a Google Doc/Sheet as PDF", "files",
["drive"], ["gws drive files export {file_id} --mime application/pdf --output {output}"]),
"list-files": Recipe("list-files", "List files in a Drive folder", "files",
["drive"], ["gws drive files list --parents {folder_id} --json"]),
"find-large-files": Recipe("find-large-files", "Find largest files in Drive", "files",
["drive"], ["gws drive files list --orderBy 'quotaBytesUsed desc' --limit 20 --json"]),
"cleanup-trash": Recipe("cleanup-trash", "Empty Drive trash", "files",
["drive"], ["gws drive files emptyTrash"]),
# Calendar (6)
"create-event": Recipe("create-event", "Create a calendar event with attendees", "calendar",
["calendar"], [
"gws calendar events insert primary --summary {title} "
"--start {start} --end {end} --attendees {attendees}"
]),
"quick-event": Recipe("quick-event", "Create event from natural language", "calendar",
["calendar"], ["gws helpers quick-event {text}"]),
"find-time": Recipe("find-time", "Find available time slots for a meeting", "calendar",
["calendar"], ["gws helpers find-time --attendees {attendees} --duration {minutes} --within {date_range}"]),
"today-schedule": Recipe("today-schedule", "Show today's calendar events", "calendar",
["calendar"], ["gws calendar events list primary --timeMin {today_start} --timeMax {today_end} --json"]),
"meeting-prep": Recipe("meeting-prep", "Prepare for an upcoming meeting (agenda + attendees)", "calendar",
["calendar"], ["gws recipes meeting-prep --event-id {event_id}"]),
"reschedule": Recipe("reschedule", "Move an event to a new time", "calendar",
["calendar"], ["gws calendar events patch primary {event_id} --start {new_start} --end {new_end}"]),
# Reporting (5)
"standup-report": Recipe("standup-report", "Generate daily standup from calendar and tasks", "reporting",
["calendar", "tasks"], ["gws recipes standup-report --json"]),
"weekly-summary": Recipe("weekly-summary", "Summarize week's emails, events, and tasks", "reporting",
["gmail", "calendar", "tasks"], ["gws recipes weekly-summary --json"]),
"drive-activity": Recipe("drive-activity", "Report on Drive file activity", "reporting",
["drive"], ["gws drive activities list --json"]),
"email-stats": Recipe("email-stats", "Email volume statistics", "reporting",
["gmail"], [
"gws gmail users.messages list me --query 'newer_than:7d' --json",
"# Pipe through output_analyzer.py --count",
]),
"task-progress": Recipe("task-progress", "Report on task completion", "reporting",
["tasks"], ["gws tasks tasks list {tasklist_id} --json"]),
# Collaboration (5)
"share-folder": Recipe("share-folder", "Share a Drive folder with a team", "collaboration",
["drive"], ["gws drive permissions create {folder_id} --type group --role writer --emailAddress {group}"]),
"create-doc": Recipe("create-doc", "Create a Google Doc with initial content", "collaboration",
["docs"], ["gws docs documents create --title {title} --json"]),
"chat-message": Recipe("chat-message", "Send a message to a Google Chat space", "collaboration",
["chat"], ["gws chat spaces.messages create {space} --text {message}"]),
"list-spaces": Recipe("list-spaces", "List Google Chat spaces", "collaboration",
["chat"], ["gws chat spaces list --json"]),
"task-create": Recipe("task-create", "Create a task in Google Tasks", "collaboration",
["tasks"], ["gws tasks tasks insert {tasklist_id} --title {title} --due {due_date}"]),
# Data (4)
"sheet-read": Recipe("sheet-read", "Read data from a spreadsheet range", "data",
["sheets"], ["gws sheets spreadsheets.values get {sheet_id} --range {range} --json"]),
"sheet-write": Recipe("sheet-write", "Write data to a spreadsheet", "data",
["sheets"], ["gws sheets spreadsheets.values update {sheet_id} --range {range} --values {data}"]),
"sheet-append": Recipe("sheet-append", "Append rows to a spreadsheet", "data",
["sheets"], ["gws sheets spreadsheets.values append {sheet_id} --range {range} --values {data}"]),
"export-contacts": Recipe("export-contacts", "Export contacts list", "data",
["people"], ["gws people people.connections list me --personFields names,emailAddresses --json"]),
# Admin (4)
"list-users": Recipe("list-users", "List all users in the Workspace domain", "admin",
["admin"], ["gws admin users list --domain {domain} --json"],
"Requires Admin SDK API and admin.directory.user.readonly scope"),
"list-groups": Recipe("list-groups", "List all groups in the domain", "admin",
["admin"], ["gws admin groups list --domain {domain} --json"]),
"user-info": Recipe("user-info", "Get detailed user information", "admin",
["admin"], ["gws admin users get {email} --json"]),
"audit-logins": Recipe("audit-logins", "Audit recent login activity", "admin",
["admin"], ["gws admin activities list login --json"]),
# Cross-Service (4)
"morning-briefing": Recipe("morning-briefing", "Today's events + unread emails + pending tasks", "cross-service",
["gmail", "calendar", "tasks"], [
"gws calendar events list primary --timeMin {today} --maxResults 10 --json",
"gws gmail users.messages list me --query 'is:unread' --limit 10 --json",
"gws tasks tasks list {default_tasklist} --json",
]),
"eod-wrap": Recipe("eod-wrap", "End-of-day wrap up: summarize completed, pending, tomorrow", "cross-service",
["calendar", "tasks"], [
"gws calendar events list primary --timeMin {today_start} --timeMax {today_end} --json",
"gws tasks tasks list {default_tasklist} --json",
]),
"project-status": Recipe("project-status", "Aggregate project status from Drive, Sheets, Tasks", "cross-service",
["drive", "sheets", "tasks"], [
"gws drive files list --query 'name contains {project}' --json",
"gws tasks tasks list {tasklist_id} --json",
]),
"inbox-zero": Recipe("inbox-zero", "Process inbox to zero: label, archive, reply, task", "cross-service",
["gmail", "tasks"], [
"gws gmail users.messages list me --query 'is:inbox' --json",
"# Process each: label, archive, or create task",
]),
}
PERSONAS: Dict[str, Dict] = {
"executive-assistant": {
"description": "Executive assistant managing schedules, emails, and communications",
"recipes": ["morning-briefing", "today-schedule", "find-time", "send-email", "reply-to-thread",
"standup-report", "meeting-prep", "eod-wrap", "quick-event", "inbox-zero"],
},
"pm": {
"description": "Project manager tracking tasks, meetings, and deliverables",
"recipes": ["standup-report", "create-event", "find-time", "task-create", "task-progress",
"project-status", "weekly-summary", "share-folder", "sheet-read", "morning-briefing"],
},
"hr": {
"description": "HR managing people, onboarding, and communications",
"recipes": ["list-users", "user-info", "send-email", "create-event", "create-doc",
"share-folder", "chat-message", "list-groups", "export-contacts", "today-schedule"],
},
"sales": {
"description": "Sales rep managing client communications and proposals",
"recipes": ["send-email", "search-emails", "create-event", "find-time", "create-doc",
"share-file", "sheet-read", "sheet-write", "export-file", "morning-briefing"],
},
"it-admin": {
"description": "IT administrator managing Workspace configuration and security",
"recipes": ["list-users", "list-groups", "user-info", "audit-logins", "drive-activity",
"find-large-files", "cleanup-trash", "label-manager", "filter-setup", "share-folder"],
},
"developer": {
"description": "Developer using Workspace APIs for automation",
"recipes": ["sheet-read", "sheet-write", "sheet-append", "upload-file", "create-doc",
"chat-message", "task-create", "list-files", "export-file", "send-email"],
},
"marketing": {
"description": "Marketing team member managing campaigns and content",
"recipes": ["send-email", "create-doc", "share-file", "upload-file", "create-sheet",
"sheet-write", "chat-message", "create-event", "email-stats", "weekly-summary"],
},
"finance": {
"description": "Finance team managing spreadsheets and reports",
"recipes": ["sheet-read", "sheet-write", "sheet-append", "create-sheet", "export-file",
"share-file", "send-email", "find-large-files", "drive-activity", "weekly-summary"],
},
"legal": {
"description": "Legal team managing documents and compliance",
"recipes": ["create-doc", "share-file", "export-file", "search-emails", "send-email",
"upload-file", "list-files", "drive-activity", "audit-logins", "find-large-files"],
},
"support": {
"description": "Customer support managing tickets and communications",
"recipes": ["search-emails", "send-email", "reply-to-thread", "label-manager", "filter-setup",
"task-create", "chat-message", "unread-digest", "inbox-zero", "morning-briefing"],
},
}
def list_recipes(persona: Optional[str], output_json: bool):
"""List all recipes, optionally filtered by persona."""
if persona:
if persona not in PERSONAS:
print(f"Unknown persona: {persona}. Available: {', '.join(PERSONAS.keys())}")
sys.exit(1)
recipe_names = PERSONAS[persona]["recipes"]
recipes = {k: v for k, v in RECIPES.items() if k in recipe_names}
title = f"Recipes for {persona.upper()}: {PERSONAS[persona]['description']}"
else:
recipes = RECIPES
title = "All 43 Google Workspace CLI Recipes"
if output_json:
output = []
for name, r in recipes.items():
output.append(asdict(r))
print(json.dumps(output, indent=2))
return
print(f"\n{'='*60}")
print(f" {title}")
print(f"{'='*60}\n")
by_category: Dict[str, list] = {}
for name, r in recipes.items():
by_category.setdefault(r.category, []).append(r)
for cat, cat_recipes in sorted(by_category.items()):
print(f" {cat.upper()} ({len(cat_recipes)})")
for r in cat_recipes:
svcs = ",".join(r.services)
print(f" {r.name:<24} {r.description:<40} [{svcs}]")
print()
print(f" Total: {len(recipes)} recipes")
print(f"\n{'='*60}\n")
def search_recipes(keyword: str, output_json: bool):
"""Search recipes by keyword."""
keyword_lower = keyword.lower()
matches = {k: v for k, v in RECIPES.items()
if keyword_lower in k.lower()
or keyword_lower in v.description.lower()
or keyword_lower in v.category.lower()
or any(keyword_lower in s for s in v.services)}
if output_json:
print(json.dumps([asdict(r) for r in matches.values()], indent=2))
return
print(f"\n Search results for '{keyword}': {len(matches)} matches\n")
for name, r in matches.items():
print(f" {r.name:<24} {r.description}")
print()
def describe_recipe(name: str, output_json: bool):
"""Show full details for a recipe."""
recipe = RECIPES.get(name)
if not recipe:
print(f"Unknown recipe: {name}")
print(f"Use --list to see available recipes")
sys.exit(1)
if output_json:
print(json.dumps(asdict(recipe), indent=2))
return
print(f"\n{'='*60}")
print(f" Recipe: {recipe.name}")
print(f"{'='*60}\n")
print(f" Description: {recipe.description}")
print(f" Category: {recipe.category}")
print(f" Services: {', '.join(recipe.services)}")
if recipe.prerequisites:
print(f" Prerequisites: {recipe.prerequisites}")
print(f"\n Commands:")
for i, cmd in enumerate(recipe.commands, 1):
print(f" {i}. {cmd}")
print(f"\n{'='*60}\n")
def run_recipe(name: str, dry_run: bool):
"""Execute a recipe (or print commands in dry-run mode)."""
recipe = RECIPES.get(name)
if not recipe:
print(f"Unknown recipe: {name}")
sys.exit(1)
if dry_run:
print(f"\n [DRY RUN] Recipe: {recipe.name}\n")
for i, cmd in enumerate(recipe.commands, 1):
print(f" {i}. {cmd}")
print(f"\n (No commands executed)")
return
print(f"\n Executing recipe: {recipe.name}\n")
for cmd in recipe.commands:
if cmd.startswith("#"):
print(f" {cmd}")
continue
print(f" $ {cmd}")
try:
result = subprocess.run(cmd, shell=True, capture_output=True, text=True, timeout=30)
if result.stdout:
print(result.stdout)
if result.returncode != 0 and result.stderr:
print(f" Error: {result.stderr.strip()[:200]}")
except subprocess.TimeoutExpired:
print(f" Timeout after 30s")
except OSError as e:
print(f" Execution error: {e}")
def list_personas(output_json: bool):
"""List all available personas."""
if output_json:
print(json.dumps(PERSONAS, indent=2))
return
print(f"\n{'='*60}")
print(f" 10 PERSONA BUNDLES")
print(f"{'='*60}\n")
for name, p in PERSONAS.items():
print(f" {name:<24} {p['description']}")
print(f" {'':24} Recipes: {', '.join(p['recipes'][:5])}...")
print()
print(f"{'='*60}\n")
def main():
parser = argparse.ArgumentParser(
description="Catalog, search, and execute Google Workspace CLI recipes",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
%(prog)s --list # List all 43 recipes
%(prog)s --list --persona pm # Recipes for project managers
%(prog)s --search "email" # Search by keyword
%(prog)s --describe standup-report # Full recipe details
%(prog)s --run standup-report --dry-run # Preview recipe commands
%(prog)s --personas # List all 10 personas
%(prog)s --list --json # JSON output
""",
)
parser.add_argument("--list", action="store_true", help="List all recipes")
parser.add_argument("--search", help="Search recipes by keyword")
parser.add_argument("--describe", help="Show full details for a recipe")
parser.add_argument("--run", help="Execute a recipe")
parser.add_argument("--dry-run", action="store_true", help="Print commands without executing")
parser.add_argument("--persona", help="Filter recipes by persona")
parser.add_argument("--personas", action="store_true", help="List all personas")
parser.add_argument("--json", action="store_true", help="Output JSON")
args = parser.parse_args()
if not any([args.list, args.search, args.describe, args.run, args.personas]):
parser.print_help()
return
if args.personas:
list_personas(args.json)
return
if args.list:
list_recipes(args.persona, args.json)
return
if args.search:
search_recipes(args.search, args.json)
return
if args.describe:
describe_recipe(args.describe, args.json)
return
if args.run:
run_recipe(args.run, args.dry_run)
return
if __name__ == "__main__":
main()
FILE:scripts/output_analyzer.py
#!/usr/bin/env python3
"""
Google Workspace CLI Output Analyzer — Parse, filter, and aggregate JSON/NDJSON output.
Reads JSON arrays or NDJSON streams from stdin or file, applies filters,
projections, sorting, grouping, and outputs in table/csv/json format.
Usage:
gws drive files list --json | python3 output_analyzer.py --count
gws drive files list --json | python3 output_analyzer.py --filter "mimeType=application/pdf"
gws drive files list --json | python3 output_analyzer.py --select "name,size" --format table
python3 output_analyzer.py --input results.json --group-by "mimeType"
python3 output_analyzer.py --demo --select "name,mimeType,size" --format table
"""
import argparse
import csv
import io
import json
import sys
from dataclasses import dataclass
from typing import List, Dict, Any, Optional
DEMO_DATA = [
{"id": "1", "name": "Q1 Report.pdf", "mimeType": "application/pdf", "size": "245760",
"modifiedTime": "2026-03-10T14:30:00Z", "shared": True, "owners": [{"displayName": "Alice"}]},
{"id": "2", "name": "Budget 2026.xlsx", "mimeType": "application/vnd.google-apps.spreadsheet",
"size": "0", "modifiedTime": "2026-03-09T09:15:00Z", "shared": True,
"owners": [{"displayName": "Bob"}]},
{"id": "3", "name": "Meeting Notes.docx", "mimeType": "application/vnd.google-apps.document",
"size": "0", "modifiedTime": "2026-03-08T16:00:00Z", "shared": False,
"owners": [{"displayName": "Alice"}]},
{"id": "4", "name": "Logo.png", "mimeType": "image/png", "size": "102400",
"modifiedTime": "2026-03-07T11:00:00Z", "shared": False,
"owners": [{"displayName": "Charlie"}]},
{"id": "5", "name": "Presentation.pptx", "mimeType": "application/vnd.google-apps.presentation",
"size": "0", "modifiedTime": "2026-03-06T10:00:00Z", "shared": True,
"owners": [{"displayName": "Alice"}]},
{"id": "6", "name": "Invoice-001.pdf", "mimeType": "application/pdf", "size": "89000",
"modifiedTime": "2026-03-05T08:30:00Z", "shared": False,
"owners": [{"displayName": "Bob"}]},
{"id": "7", "name": "Project Plan.xlsx", "mimeType": "application/vnd.google-apps.spreadsheet",
"size": "0", "modifiedTime": "2026-03-04T13:45:00Z", "shared": True,
"owners": [{"displayName": "Charlie"}]},
{"id": "8", "name": "Contract Draft.docx", "mimeType": "application/vnd.google-apps.document",
"size": "0", "modifiedTime": "2026-03-03T09:00:00Z", "shared": False,
"owners": [{"displayName": "Alice"}]},
]
def read_input(input_file: Optional[str]) -> List[Dict[str, Any]]:
"""Read JSON array or NDJSON from file or stdin."""
if input_file:
with open(input_file, "r") as f:
text = f.read().strip()
else:
if sys.stdin.isatty():
return []
text = sys.stdin.read().strip()
if not text:
return []
# Try JSON array first
try:
data = json.loads(text)
if isinstance(data, list):
return data
if isinstance(data, dict):
# Some gws commands wrap results in a key
for key in ("files", "messages", "events", "items", "results",
"spreadsheets", "spaces", "tasks", "users", "groups"):
if key in data and isinstance(data[key], list):
return data[key]
return [data]
except json.JSONDecodeError:
pass
# Try NDJSON
records = []
for line in text.split("\n"):
line = line.strip()
if line:
try:
records.append(json.loads(line))
except json.JSONDecodeError:
continue
return records
def get_nested(obj: Dict, path: str) -> Any:
"""Get a nested value by dot-separated path."""
parts = path.split(".")
current = obj
for part in parts:
if isinstance(current, dict):
current = current.get(part)
elif isinstance(current, list) and part.isdigit():
idx = int(part)
current = current[idx] if idx < len(current) else None
else:
return None
if current is None:
return None
return current
def apply_filter(records: List[Dict], filter_expr: str) -> List[Dict]:
"""Filter records by field=value expression."""
if "=" not in filter_expr:
return records
field_path, value = filter_expr.split("=", 1)
result = []
for rec in records:
rec_val = get_nested(rec, field_path)
if rec_val is None:
continue
rec_str = str(rec_val).lower()
if rec_str == value.lower() or value.lower() in rec_str:
result.append(rec)
return result
def apply_select(records: List[Dict], fields: str) -> List[Dict]:
"""Project specific fields from records."""
field_list = [f.strip() for f in fields.split(",")]
result = []
for rec in records:
projected = {}
for f in field_list:
projected[f] = get_nested(rec, f)
result.append(projected)
return result
def apply_sort(records: List[Dict], sort_field: str, reverse: bool = False) -> List[Dict]:
"""Sort records by a field."""
def sort_key(rec):
val = get_nested(rec, sort_field)
if val is None:
return ""
if isinstance(val, (int, float)):
return val
try:
return float(val)
except (ValueError, TypeError):
return str(val).lower()
return sorted(records, key=sort_key, reverse=reverse)
def apply_group_by(records: List[Dict], field: str) -> Dict[str, int]:
"""Group records by a field and count."""
groups: Dict[str, int] = {}
for rec in records:
val = get_nested(rec, field)
key = str(val) if val is not None else "(null)"
groups[key] = groups.get(key, 0) + 1
return dict(sorted(groups.items(), key=lambda x: x[1], reverse=True))
def compute_stats(records: List[Dict], field: str) -> Dict[str, Any]:
"""Compute min/max/avg/sum for a numeric field."""
values = []
for rec in records:
val = get_nested(rec, field)
if val is not None:
try:
values.append(float(val))
except (ValueError, TypeError):
continue
if not values:
return {"field": field, "count": 0, "error": "No numeric values found"}
return {
"field": field,
"count": len(values),
"min": min(values),
"max": max(values),
"sum": sum(values),
"avg": sum(values) / len(values),
}
def format_table(records: List[Dict]) -> str:
"""Format records as an aligned text table."""
if not records:
return "(no records)"
headers = list(records[0].keys())
# Calculate column widths
widths = {h: len(h) for h in headers}
for rec in records:
for h in headers:
val = str(rec.get(h, ""))
if len(val) > 60:
val = val[:57] + "..."
widths[h] = max(widths[h], len(val))
# Header
header_line = " ".join(h.ljust(widths[h]) for h in headers)
sep_line = " ".join("-" * widths[h] for h in headers)
lines = [header_line, sep_line]
# Rows
for rec in records:
row = []
for h in headers:
val = str(rec.get(h, ""))
if len(val) > 60:
val = val[:57] + "..."
row.append(val.ljust(widths[h]))
lines.append(" ".join(row))
return "\n".join(lines)
def format_csv_output(records: List[Dict]) -> str:
"""Format records as CSV."""
if not records:
return ""
output = io.StringIO()
writer = csv.DictWriter(output, fieldnames=records[0].keys())
writer.writeheader()
writer.writerows(records)
return output.getvalue()
def main():
parser = argparse.ArgumentParser(
description="Parse, filter, and aggregate JSON/NDJSON from gws CLI output",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
gws drive files list --json | %(prog)s --count
gws drive files list --json | %(prog)s --filter "mimeType=pdf" --select "name,size"
gws drive files list --json | %(prog)s --group-by "mimeType" --format table
gws drive files list --json | %(prog)s --sort "size" --reverse --format table
gws drive files list --json | %(prog)s --stats "size"
%(prog)s --input results.json --select "name,mimeType" --format csv
%(prog)s --demo --select "name,mimeType,size" --format table
""",
)
parser.add_argument("--input", help="Input file (default: stdin)")
parser.add_argument("--demo", action="store_true", help="Use demo data")
parser.add_argument("--count", action="store_true", help="Count records")
parser.add_argument("--filter", help="Filter by field=value")
parser.add_argument("--select", help="Comma-separated fields to project")
parser.add_argument("--sort", help="Sort by field")
parser.add_argument("--reverse", action="store_true", help="Reverse sort order")
parser.add_argument("--group-by", help="Group by field and count")
parser.add_argument("--stats", help="Compute stats for a numeric field")
parser.add_argument("--format", choices=["json", "table", "csv"], default="json",
help="Output format (default: json)")
parser.add_argument("--json", action="store_true",
help="Shorthand for --format json")
args = parser.parse_args()
if args.json:
args.format = "json"
# Read input
if args.demo:
records = DEMO_DATA[:]
else:
records = read_input(args.input)
if not records and not args.demo:
# If no pipe input and no file, use demo
records = DEMO_DATA[:]
print("(No input detected, using demo data)\n", file=sys.stderr)
# Apply operations in order
if args.filter:
records = apply_filter(records, args.filter)
if args.sort:
records = apply_sort(records, args.sort, args.reverse)
# Count
if args.count:
if args.format == "json":
print(json.dumps({"count": len(records)}))
else:
print(f"Count: {len(records)}")
return
# Group by
if args.group_by:
groups = apply_group_by(records, args.group_by)
if args.format == "json":
print(json.dumps(groups, indent=2))
elif args.format == "csv":
print(f"{args.group_by},count")
for k, v in groups.items():
print(f"{k},{v}")
else:
print(f"\n Group by: {args.group_by}\n")
for k, v in groups.items():
print(f" {k:<50} {v}")
print(f"\n Total groups: {len(groups)}")
return
# Stats
if args.stats:
stats = compute_stats(records, args.stats)
if args.format == "json":
print(json.dumps(stats, indent=2))
else:
print(f"\n Stats for '{args.stats}':")
for k, v in stats.items():
if isinstance(v, float):
print(f" {k}: {v:,.2f}")
else:
print(f" {k}: {v}")
return
# Select fields
if args.select:
records = apply_select(records, args.select)
# Output
if args.format == "json":
print(json.dumps(records, indent=2))
elif args.format == "csv":
print(format_csv_output(records))
else:
print(f"\n{format_table(records)}\n")
print(f" ({len(records)} records)\n")
if __name__ == "__main__":
main()
FILE:scripts/workspace_audit.py
#!/usr/bin/env python3
"""
Google Workspace Security Audit — Audit Workspace configuration for security risks.
Checks Drive external sharing, Gmail forwarding rules, OAuth app grants,
Calendar visibility, admin settings, and generates remediation commands.
Runs in demo mode with embedded sample data when gws is not installed.
Usage:
python3 workspace_audit.py
python3 workspace_audit.py --json
python3 workspace_audit.py --services gmail,drive,calendar
python3 workspace_audit.py --demo
"""
import argparse
import json
import shutil
import subprocess
import sys
from dataclasses import dataclass, field, asdict
from typing import List, Dict, Optional
@dataclass
class AuditFinding:
area: str
check: str
status: str # PASS, WARN, FAIL
message: str
risk: str = ""
remediation: str = ""
@dataclass
class AuditReport:
findings: List[dict] = field(default_factory=list)
score: int = 0
max_score: int = 100
grade: str = ""
summary: str = ""
demo_mode: bool = False
DEMO_FINDINGS = [
AuditFinding("drive", "External sharing", "WARN",
"External sharing is enabled for the domain",
"Data exfiltration via shared links",
"Review sharing settings in Admin Console > Apps > Google Workspace > Drive"),
AuditFinding("drive", "Link sharing defaults", "FAIL",
"Default link sharing is set to 'Anyone with the link'",
"Sensitive files accessible without authentication",
"gws admin settings update drive --defaultLinkSharing restricted"),
AuditFinding("gmail", "Auto-forwarding", "PASS",
"No auto-forwarding rules detected for admin accounts"),
AuditFinding("gmail", "SPF record", "PASS",
"SPF record configured correctly"),
AuditFinding("gmail", "DMARC record", "WARN",
"DMARC policy is set to 'none' (monitoring only)",
"Email spoofing not actively blocked",
"Update DMARC DNS record: v=DMARC1; p=quarantine; rua=mailto:dmarc@company.com"),
AuditFinding("gmail", "DKIM signing", "PASS",
"DKIM signing is enabled"),
AuditFinding("calendar", "Default visibility", "WARN",
"Calendar default visibility is 'See all event details'",
"Meeting details visible to all domain users",
"Admin Console > Apps > Calendar > Sharing settings > Set to 'Free/Busy'"),
AuditFinding("calendar", "External sharing", "PASS",
"External calendar sharing is restricted"),
AuditFinding("oauth", "Third-party apps", "FAIL",
"12 third-party OAuth apps with broad access detected",
"Unauthorized data access via OAuth grants",
"Review: Admin Console > Security > API controls > App access control"),
AuditFinding("oauth", "High-risk apps", "WARN",
"3 apps have Drive full access scope",
"Apps can read/modify all Drive files",
"Audit each app: gws admin tokens list --json | filter by scope"),
AuditFinding("admin", "Super admin count", "WARN",
"4 super admin accounts detected (recommended: 2-3)",
"Increased attack surface for privilege escalation",
"Reduce super admins: gws admin users list --query 'isAdmin=true' --json"),
AuditFinding("admin", "2-Step verification", "PASS",
"2-Step verification enforced for all users"),
AuditFinding("admin", "Password policy", "PASS",
"Minimum password length: 12 characters"),
AuditFinding("admin", "Login challenges", "PASS",
"Suspicious login challenges enabled"),
]
def run_gws_command(cmd: List[str]) -> Optional[str]:
"""Run a gws command and return stdout, or None on failure."""
try:
result = subprocess.run(cmd, capture_output=True, text=True, timeout=20)
if result.returncode == 0:
return result.stdout
return None
except (subprocess.TimeoutExpired, FileNotFoundError, OSError):
return None
def audit_drive() -> List[AuditFinding]:
"""Audit Drive sharing and security settings."""
findings = []
# Check sharing settings
output = run_gws_command(["gws", "drive", "about", "get", "--json"])
if output:
try:
data = json.loads(output)
# Check if external sharing is enabled
if data.get("canShareOutsideDomain", True):
findings.append(AuditFinding(
"drive", "External sharing", "WARN",
"External sharing is enabled",
"Data exfiltration via shared links",
"Review Admin Console > Apps > Drive > Sharing settings"
))
else:
findings.append(AuditFinding(
"drive", "External sharing", "PASS",
"External sharing is restricted"
))
except json.JSONDecodeError:
findings.append(AuditFinding(
"drive", "External sharing", "WARN",
"Could not parse Drive settings"
))
else:
findings.append(AuditFinding(
"drive", "External sharing", "WARN",
"Could not retrieve Drive settings"
))
return findings
def audit_gmail() -> List[AuditFinding]:
"""Audit Gmail forwarding and email security."""
findings = []
# Check forwarding rules
output = run_gws_command(["gws", "gmail", "users.settings.forwardingAddresses", "list", "me", "--json"])
if output:
try:
data = json.loads(output)
addrs = data if isinstance(data, list) else data.get("forwardingAddresses", [])
if addrs:
findings.append(AuditFinding(
"gmail", "Auto-forwarding", "WARN",
f"{len(addrs)} forwarding addresses configured",
"Data exfiltration via email forwarding",
"Review: gws gmail users.settings.forwardingAddresses list me --json"
))
else:
findings.append(AuditFinding(
"gmail", "Auto-forwarding", "PASS",
"No forwarding addresses configured"
))
except json.JSONDecodeError:
pass
else:
findings.append(AuditFinding(
"gmail", "Auto-forwarding", "WARN",
"Could not check forwarding settings"
))
return findings
def audit_calendar() -> List[AuditFinding]:
"""Audit Calendar sharing settings."""
findings = []
output = run_gws_command(["gws", "calendar", "calendarList", "get", "primary", "--json"])
if output:
findings.append(AuditFinding(
"calendar", "Primary calendar", "PASS",
"Primary calendar accessible"
))
else:
findings.append(AuditFinding(
"calendar", "Primary calendar", "WARN",
"Could not access primary calendar"
))
return findings
def run_live_audit(services: List[str]) -> AuditReport:
"""Run live audit against actual gws installation."""
report = AuditReport()
all_findings = []
audit_map = {
"drive": audit_drive,
"gmail": audit_gmail,
"calendar": audit_calendar,
}
for svc in services:
fn = audit_map.get(svc)
if fn:
all_findings.extend(fn())
report.findings = [asdict(f) for f in all_findings]
report = calculate_score(report)
return report
def run_demo_audit() -> AuditReport:
"""Return demo audit report with embedded sample data."""
report = AuditReport(
findings=[asdict(f) for f in DEMO_FINDINGS],
demo_mode=True,
)
report = calculate_score(report)
return report
def calculate_score(report: AuditReport) -> AuditReport:
"""Calculate audit score and grade."""
total = len(report.findings)
if total == 0:
report.score = 0
report.grade = "N/A"
report.summary = "No checks performed"
return report
passes = sum(1 for f in report.findings if f["status"] == "PASS")
warns = sum(1 for f in report.findings if f["status"] == "WARN")
fails = sum(1 for f in report.findings if f["status"] == "FAIL")
# Score: PASS=100, WARN=50, FAIL=0
score = int(((passes * 100) + (warns * 50)) / total)
report.score = score
report.max_score = 100
if score >= 90:
report.grade = "A"
elif score >= 75:
report.grade = "B"
elif score >= 60:
report.grade = "C"
elif score >= 40:
report.grade = "D"
else:
report.grade = "F"
report.summary = f"{passes} passed, {warns} warnings, {fails} failures — Score: {score}/100 (Grade: {report.grade})"
return report
def main():
parser = argparse.ArgumentParser(
description="Security and configuration audit for Google Workspace",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples:
%(prog)s # Full audit (or demo if gws not installed)
%(prog)s --json # JSON output
%(prog)s --services gmail,drive # Audit specific services
%(prog)s --demo # Demo mode with sample data
""",
)
parser.add_argument("--json", action="store_true", help="Output JSON")
parser.add_argument("--services", default="gmail,drive,calendar",
help="Comma-separated services to audit (default: gmail,drive,calendar)")
parser.add_argument("--demo", action="store_true", help="Run with demo data")
args = parser.parse_args()
services = [s.strip() for s in args.services.split(",") if s.strip()]
if args.demo or not shutil.which("gws"):
report = run_demo_audit()
else:
report = run_live_audit(services)
if args.json:
print(json.dumps(asdict(report), indent=2))
else:
print(f"\n{'='*60}")
print(f" GOOGLE WORKSPACE SECURITY AUDIT")
if report.demo_mode:
print(f" (DEMO MODE — sample data)")
print(f"{'='*60}\n")
print(f" Score: {report.score}/{report.max_score} (Grade: {report.grade})\n")
current_area = ""
for f in report.findings:
if f["area"] != current_area:
current_area = f["area"]
print(f"\n {current_area.upper()}")
print(f" {'-'*40}")
icon = {"PASS": "PASS", "WARN": "WARN", "FAIL": "FAIL"}.get(f["status"], "????")
print(f" [{icon}] {f['check']}: {f['message']}")
if f.get("risk") and f["status"] != "PASS":
print(f" Risk: {f['risk']}")
if f.get("remediation") and f["status"] != "PASS":
print(f" Fix: {f['remediation']}")
print(f"\n {'='*56}")
print(f" {report.summary}")
print(f"\n{'='*60}\n")
if __name__ == "__main__":
main()
Hỗ trợ nhà nghiên cứu lâm sàng tìm tài trợ NIH: phỏng vấn ý tưởng, giai đoạn sự nghiệp, dữ liệu sơ bộ và định vị chiến lược tài trợ.
---
name: grants
description: "NIH grant research skill for clinical researchers. Grill-me intake (research idea + career stage + preliminary data + environment + submission posture + known institute targets) locks down the funding strategy before any search runs. Runs a 5-facet Consensus positioning analysis (with draft Significance/Innovation language), maps the research to the right NIH institutes and study sections via RePORTER, finds NOSIs and funded overlap, and produces an editable Word document (.docx) with budget/scope-aware mechanism recommendations, submission timelines, and a mandatory program officer recommendation. Triggers: 'grants for [topic]', 'find grants for my research idea', 'what grants match my research', 'help me find NIH funding', 'grant opportunities for my research', or any grant-related request. NIH-only scope — non-NIH funders (PCORI, DOD CDMRP, VA, foundations) are out of scope and flagged at intake."
license: MIT
metadata:
source_spec: "megaprompts/08-grants-megaprompt.md"
build_pattern: "Path B (direct conversion)"
research_pack_convention: "Agent Integrity Rules verbatim per PR #657 audit"
version: 1.0.0
---
# Grants — NIH Funding Intelligence
> **Portability:** Requires `bash_tool` (for RePORTER POST via curl), Node.js with `docx` package, and a Consensus MCP connection. Works in Claude Code CLI natively. In Claude.ai with Code Execution + Consensus MCP, the workflow is supported but slower.
> **Scope: NIH-only.** Non-NIH funders (PCORI, DOD CDMRP, VA, foundations) are out of scope and flagged at intake.
For a clinical researcher with a research idea, produce a strategic NIH funding overview as an editable `.docx`. Output covers research positioning analysis, institute mapping, targeted grant discovery, and strategic recommendations the researcher can edit, copy from, and share with their mentor.
## Agent Integrity Rules (Research-Pack Convention)
Inherited; locked verbatim per PR #657 audit.
- **Execution discipline.** A step isn't complete until result is confirmed received. Consensus calls **sequential with 1+ sec pause**. RePORTER calls sequential.
- **Data sourcing.** Count only what tool calls returned this session. Never supplement with training knowledge. Training knowledge labeled `[Not from Consensus/RePORTER — reference information]` and excluded from counts.
- **Counts & attribution.** Queries sent / results shown / results cited — three separate numbers, never conflate. Every cited paper has retrievable URL from this session.
- **Error handling.** On failure → wait 3s → retry once → log. After 3 consecutive failures across tools: stop, alert researcher, explain what's missing. Never silently skip.
- **Transparency.** Audit Log section in the DOCX. Same standards in chat summary as in document.
See [`references/reporter_post_patterns.md`](references/reporter_post_patterns.md) for the RePORTER POST canon + plan-tier detection.
## Phase 1: Grill-Me Intake (6 forcing questions, one at a time)
### Q1 (root) — Research idea
> **Describe the research idea in 2–3 sentences. What's the question, what's new, and what's the clinical relevance? Vague answers ("AI for healthcare", "biomarkers for disease X") will be rejected — push for specificity.**
>
> *Why I'm asking:* Five Consensus searches (established / stakes / current approaches / adjacent methods / gaps) depend on a precise research idea. Vague ideas produce vague gap quotes and useless positioning narrative.
Refuse mush. Re-ask once with examples if user is too broad.
### Q2 (depends on Q1) — Career stage
> **Career stage — pick one:**
>
> 1. Pre-doctoral (PhD student, T32 trainee)
> 2. Postdoctoral fellow (F32, K99 candidate)
> 3. Early career (K-award candidate, first R01)
> 4. Independent investigator (multiple R01s, established lab)
> 5. Senior PI (R35, P-series, U01 leadership)
>
> *Why I'm asking:* Career stage filters mechanism recommendations. F-series for trainees, K-series for early career, R-series for independent. Picking the wrong stage produces unfundable mechanism suggestions.
Forcing choice.
### Q3 (depends on Q2) — Preliminary data status
> **Preliminary data — pick one:**
>
> 1. None (de novo project, no pilot data yet)
> 2. Pilot data (early findings, single-site)
> 3. Strong preliminary (multi-experiment, ready for R01-scale)
> 4. Validated and ready (multi-site, publication-ready)
>
> *Why I'm asking:* Prelim data status drives mechanism budget. No data → R03 / R21 pilot scope. Strong prelim → R01 / U01 multi-site scale. Mismatch produces uncompetitive applications.
### Q4 (depends on Q2) — Environment
> **Research environment — pick one:**
>
> 1. R01-eligible (research-intensive institution with NIH base funding)
> 2. Mid-tier (regional academic medical center, modest NIH portfolio)
> 3. Resource-constrained (smaller institution, minimal NIH base)
> 4. Industry-collaborative (academic + industry partnership)
>
> *Why I'm asking:* Environment affects scope realism (multi-site U01 requires R01-eligible) and which mechanism categories are competitive (R15 specifically targets resource-constrained).
### Q5 (depends on Q1) — Submission posture
> **Submission posture — pick one:**
>
> 1. New application (first submission, no prior reviews)
> 2. Resubmission (A1 with reviewer responses needed)
> 3. Exploring (haven't decided yet whether to submit)
>
> *Why I'm asking:* Resubmissions need reviewer-response guidance in the DOCX (Section 7). New applications skip that. Exploring shifts emphasis to landscape over strategy.
### Q6 (depends on Q1) — Known institute targets
> **Are you already considering specific NIH institutes? List names (NCI / NHLBI / NIMH / NINDS / NIDDK / etc.) or say "no preference — find the right ones".**
>
> *Why I'm asking:* If you have an institute hypothesis, I'll validate it against RePORTER data. If not, I'll surface the top-3 institutes funding adjacent work from the institute-tally.
Accept "no preference" as the common case.
**Stop condition:** After Q6, commit and start Phase 2A. Never re-open intake after Phase 2A begins.
## Phase 2A: Research Positioning (5 Consensus searches)
Run sequentially at 1 q/sec. Each search corresponds to one positioning facet:
1. **Established** — `"<research idea>" established evidence` — what's known
2. **Stakes** — `"<topic>" mortality OR burden OR cost OR prevalence` — why it matters
3. **Current Approaches** — `"<topic>" current treatment OR standard of care OR approach` — state of the art
4. **Adjacent Methods** — `"<related technique>" applied to <topic>` — methodological possibilities
5. **Gaps** — `"<topic>" limitations OR unanswered OR future directions OR challenge` — gap signals
Use `scripts/citation_tracker.py --action record_consensus_search` for each. Plan-tier detected from first response.
**Synthesis:** for each facet, extract 2-3 quotable findings (becomes Section 2 gap quotes). Draft Significance/Innovation language using "the field has established X (refs), but Y remains unanswered (refs)" pattern.
## Phase 2B: Institute Mapping + Grant Discovery (RePORTER POST)
RePORTER is **POST-only**. Use `bash_tool` + `curl` — never `web_fetch`.
### Dynamic fiscal year window
Compute at runtime via `scripts/fiscal_year_calculator.py`. Default: current FY + 3 prior. Federal FY starts Oct 1, so:
```bash
python ../scripts/fiscal_year_calculator.py --output json
# Returns: {"current_fy": 2026, "window": [2023, 2024, 2025, 2026]}
```
### Narrow (AND) search — finds direct overlap
```bash
curl -X POST 'https://api.reporter.nih.gov/v2/projects/search' \
-H 'Content-Type: application/json' \
-d '{
"criteria": {
"fiscal_years": [2023, 2024, 2025, 2026],
"include_active_projects": true,
"advanced_text_search": {
"operator": "AND",
"search_field": "all",
"search_text": "<key term 1> <key term 2>"
}
},
"limit": 50,
"include_fields": ["project_num", "project_title", "agency_ic_admin", "study_section", "fiscal_year", "principal_investigators", "abstract_text"]
}'
```
### Broad (OR) search — finds adjacent work
```bash
curl -X POST 'https://api.reporter.nih.gov/v2/projects/search' \
-H 'Content-Type: application/json' \
-d '{
"criteria": {
"fiscal_years": [2023, 2024, 2025, 2026],
"advanced_text_search": {
"operator": "OR",
"search_field": "all",
"search_text": "<term> <synonym> <related concept>"
}
},
"limit": 50
}'
```
### Institute tally + study section ranking
After RePORTER responses:
- Tally `agency_ic_admin` (institute code: NCI, NHLBI, NIMH, etc.) → top-3 funding institutes
- Tally `study_section` → top-2 study sections (where applications go for review)
### NOSI discovery
Parse RePORTER responses for `NOT-*` opportunity numbers. For each:
```bash
# NOSIs live at predictable URLs:
# https://grants.nih.gov/grants/guide/notice-files/NOT-<INSTITUTE>-<YEAR>-<NUMBER>.html
web_fetch <url>
```
If fetch fails: log `[NOSI {number} — fetch failed, not included]`, continue.
## Mechanism Matching (Scope-Aware)
NOT career stage alone. Career stage **+** project scope **+** prelim data drive recommendation.
Use `scripts/mechanism_matcher.py`:
```bash
python ../scripts/mechanism_matcher.py \
--career-stage "early_career" \
--prelim-data "pilot" \
--environment "r01_eligible" \
--scope "single_site" \
--output json
# Returns mechanism shortlist with rationale
```
See [`references/nih_mechanism_matching.md`](references/nih_mechanism_matching.md) for the full matrix.
## Phase 3: DOCX Generation
9 sections via Node.js + `docx` library. See [`references/docx_9_sections.md`](references/docx_9_sections.md) for full spec.
1. **Executive Summary** — title + career stage + environment + 3-4 key findings bullets
2. **Research Positioning** — 3-5 gap quotes (italicized, inline Consensus citations) + 2-3 paragraph positioning narrative + supporting evidence table
3. **Target Institutes** — ranking table (institute, project count in window, % match to your idea) + 2-3 sentence interpretation
4. **Grant Opportunities** — bold NOSI callout if any. Top-3 grants table with hyperlinked FOAs + per-grant scope/budget fit paragraph
5. **Funded Overlap** — top-5 projects table (PI, project_num, IC, year, hyperlinked to RePORTER) + differentiation paragraph
6. **Study Sections** — ranking table + best-match interpretation
7. **Strategic Recommendations & Next Steps** — 3-4 numbered recs + **mandatory program officer rec** + submission timeline note + (if resubmission Q5=2) reviewer-response guidance + closing paragraph
8. **References** — numbered bibliography, hyperlinked to Consensus
9. **Audit Log** — Consensus searches table, plan-tier note, RePORTER searches table, NOSI fetches table, summary stats, tool constraints note, failed steps
### Styling
Arial 12pt body, navy headings (#1a3a5c), light blue table headers (#e8f0f8), amber NOSI callout. `ExternalHyperlink` patterns:
- Paper citations: `https://consensus.app/papers/...`
- FOA links: `https://grants.nih.gov/grants/guide/...`
- RePORTER projects: `https://reporter.nih.gov/project-details/<id>`
## Mandatory Program Officer Recommendation
Always include in Section 7:
> **Recommended next step: contact program officer at {top institute}.** Find their staff page at https://www.nih.gov/institutes-nih/list-nih-institutes-centers-offices → {institute} → Program Officers. Prepare: 1-page specific aims + your CV + 3 specific questions about fit. Email subject: "Pre-application inquiry: <topic>".
This is the single most valuable advice for any applicant. Never skip.
## Submission Timeline (Embedded in DOCX Section 7)
| Mechanism | Standard receipt dates |
|---|---|
| R01, R21, R03 | Feb 5, Jun 5, Oct 5 |
| K awards (K01, K08, K23, K99) | Feb 12, Jun 12, Oct 12 |
| R34, R61/R33 | Feb 16, Jun 16, Oct 16 |
| F31, F32 | Apr 8, Aug 8, Dec 8 |
## Phase 4: Deliver
- Save DOCX to `<output-dir>/grants_<topic-slug>_<YYYY-MM-DD>.docx`
- Chat summary: file path + audit counts + plan tier + verdict on institute targets
- Validate: `python scripts/office/validate.py <docx>`
## Tooling
| Script | Role |
|---|---|
| `scripts/citation_tracker.py` | Three-count audit (Consensus sent/shown/cited + RePORTER projects/cited) at `~/.grants_sessions/<session>.json` |
| `scripts/fiscal_year_calculator.py` | Current FY + 3-prior window. Computed at runtime, never hardcoded. |
| `scripts/mechanism_matcher.py` | Career stage × scope × prelim → mechanism recommendation shortlist |
## References
- [`references/nih_mechanism_matching.md`](references/nih_mechanism_matching.md) — career stage × scope × prelim → mechanism canon (7+ sources)
- [`references/reporter_post_patterns.md`](references/reporter_post_patterns.md) — RePORTER curl POST templates + plan-tier detection (7+ sources)
- [`references/docx_9_sections.md`](references/docx_9_sections.md) — 9-section .docx spec + technical requirements (7+ sources)
## Error Handling
| Failure | Behavior |
|---|---|
| Consensus rate-limit hit | Wait 3s, retry once, log; if still failing, alert researcher |
| Consensus returns 0 for a facet | Surface explicitly; never fill with training knowledge |
| Consensus plan-tier cap detected | Log tier, note in audit, surface to researcher |
| RePORTER POST returns error | Retry once after 3s; if still failing, log and continue |
| RePORTER returns <5 on narrow | Document; broad OR should compensate; surface low count |
| NOSI fetch fails | Log `[NOSI {n} — fetch failed]`, continue |
| 3 consecutive tool failures | Stop, alert researcher with what's missing |
| DOCX generation fails | Save raw data as JSON fallback so researcher doesn't lose work |
## Anti-Patterns To Reject
- Parallelizing Consensus calls (will hit rate limit)
- Using `web_fetch` for RePORTER (POST-only — `web_fetch` is GET)
- Hardcoded fiscal year values
- Mechanism recommendations based on career stage alone (must consider scope too)
- Silently filling thin facet results with training knowledge
- Skipping the audit log
- Skipping the program officer recommendation
- Conflating "papers found" with "papers shown" with "papers cited"
- Fabricating NOSI details when fetch fails
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/08-grants-megaprompt.md`](../../../../megaprompts/08-grants-megaprompt.md)
**Build pattern:** Path B (direct conversion). Research-pack sibling of pulse + litreview.
FILE:references/docx_9_sections.md
# DOCX 9-Section Spec — NIH Grants Strategic Overview
This reference answers exactly one decision: **what are the 9 sections of the grants .docx, and what does each need to be useful to a researcher submitting to NIH?**
## The Core Frame
The output is a **strategic overview**, not a complete application draft. The researcher edits, copies sections into their actual application, shares with their mentor. Useful means: actionable, source-attributed, scope-aware, ready for program officer conversation.
## Section 1: Executive Summary
**Length:** Title + metadata + 3-4 bullets. Half a page.
**Contents:**
- Title: "NIH Funding Strategy: {topic}"
- Date generated
- Career stage (from Q2)
- Environment (from Q4)
- 3-4 key findings:
- Top institute(s) funding this area (from RePORTER)
- Top recommended mechanism (from `mechanism_matcher.py`)
- Submission posture insight (from Q5)
- Critical gap or opportunity (from Phase 2A positioning)
**Tone:** Confident, actionable. Reader knows what to do after this section.
## Section 2: Research Positioning
**Length:** 1-1.5 pages.
**Contents:**
### Lead with 3-5 gap quotes
Italicized, with inline Consensus citations. Example:
> *"Existing approaches to sepsis prediction rely on static risk scores that fail to capture dynamic deterioration trajectories"* (Smith et al. 2023, Consensus).
These quotes become the foundation for the Significance section of the actual application.
### Positioning narrative (2-3 paragraphs)
Draft Significance/Innovation tone:
- Paragraph 1: The field has established X (refs from "Established" facet)
- Paragraph 2: Current approaches do Y, but Z remains unanswered (refs from "Current Approaches" + "Gaps" facets)
- Paragraph 3: This proposal addresses Z via {novel approach} (anchored in Q1 research idea)
### Supporting evidence table
| Finding | Source | Year | Cites |
|---|---|---|---|
| ... | Smith et al. | 2023 | 47 |
## Section 3: Target Institutes
**Length:** Half page.
### Ranking table
| Rank | Institute | Projects in window | % of total | Mission alignment |
|---|---|---|---|---|
| 1 | NHLBI | 23 | 38% | High — cardiovascular focus matches |
| 2 | NIDDK | 14 | 23% | Medium — metabolic angle |
| 3 | NCI | 8 | 13% | Low — oncology adjacent |
### 2-3 sentence interpretation
> NHLBI dominates this funding area with 38% of projects in the recent 4-year window. Their mission specifically prioritizes... If your Q1 hypothesis maps to cardiovascular outcomes, NHLBI is the primary target. NIDDK is a viable secondary if metabolic outcomes are involved.
## Section 4: Grant Opportunities
**Length:** 1 page.
### NOSI callout (if any found)
Bold amber box:
> 🔶 **Active NOSI: NOT-HL-25-014** — Special interest in machine learning for cardiovascular risk prediction. Expires: 2027-09-30. URL: https://grants.nih.gov/grants/guide/notice-files/NOT-HL-25-014.html
>
> If your project fits this NOSI, your application is reviewed with knowledge of the institute's specific interest in this area — substantially increases prospects.
### Top 3 grants table
| FOA | Mechanism | Institute | Deadline | Budget | Hyperlink |
|---|---|---|---|---|---|
| PAR-25-XXX | R01 | NHLBI | Feb 5 | $499k × 5 yr | [link to PA] |
| PA-25-YYY | R21 | NHLBI | Jun 16 | $275k × 2 yr | [link] |
| RFA-HL-25-ZZZ | U01 | NHLBI | Oct 5 | varies | [link] |
### Per-grant paragraph
For each: scope/budget fit. Whether the user's career stage + prelim + environment align with this specific FOA.
## Section 5: Funded Overlap
**Length:** 1 page.
### Top 5 funded projects table
| PI | Project | IC | Year | Hyperlink |
|---|---|---|---|---|
| Smith, J. | "AI-driven sepsis prediction..." | NHLBI | 2024 | [RePORTER] |
### Differentiation paragraph
> The closest existing project is Smith et al. (Project #R01HL12345) at Johns Hopkins. They focus on adult ICU patients with sepsis. **Your differentiation:** pediatric population, prospective trial design, real-time deployment vs retrospective benchmarking.
This differentiation paragraph is what the reviewer reads BEFORE the Approach section. Make it sharp.
## Section 6: Study Sections
**Length:** Half page.
### Ranking table
| Rank | Study Section | Projects in window | Specialization |
|---|---|---|---|
| 1 | MEDS (Medical Imaging Study Section) | 12 | Imaging/AI methods |
| 2 | BMIO (Bioinformatics Methods + ML) | 8 | Methods development |
### Best-match interpretation
> MEDS reviews most similar applications. Implications: lean into methods rigor (their reviewers will know the methodology landscape); abstract should make method specifically clear; supplementary methods section should be detailed.
## Section 7: Strategic Recommendations & Next Steps
**Length:** 1-1.5 pages.
### 3-4 numbered recommendations
1. **Target NHLBI as primary** — strongest institute alignment + active NOSI matches your scope
2. **Apply for R21 first if Q3=pilot, R01 if Q3=strong** — scope-aware mechanism (from `mechanism_matcher.py`)
3. **Frame as ML methods + clinical application** — appeals to MEDS reviewers
4. **(If resubmission, Q5=2):** Address prior reviewer concern A by adding aim X; address concern B with prelim data Y
### MANDATORY program officer recommendation
> **Single most valuable next step: contact program officer at NHLBI.**
>
> Staff page: https://www.nhlbi.nih.gov/about/divisions → relevant division → Program Officers.
>
> Prepare:
> 1. 1-page specific aims draft
> 2. NIH biosketch
> 3. 3 specific questions about NOSI fit + mechanism preference + study section recommendation
>
> Email subject: "Pre-application inquiry: <topic>". Mention specific NOSI if applicable.
### Submission timeline note
| Mechanism | Standard receipt dates |
|---|---|
| R01, R21, R03 | Feb 5, Jun 5, Oct 5 |
| K awards | Feb 12, Jun 12, Oct 12 |
| R34, R61/R33 | Feb 16, Jun 16, Oct 16 |
| F31, F32 | Apr 8, Aug 8, Dec 8 |
Work backwards from the deadline: typical writing window is 4-6 months. Pre-application program officer contact 3-4 months before. Internal institutional pre-review 6 weeks before.
### Closing paragraph
> Your strongest path is {top recommendation}. Highest-leverage next action: contact {top institute} program officer this week with the 1-pager. They'll tell you whether to proceed with {mechanism} or pivot.
## Section 8: References
**Length:** As many as cited; numbered + hyperlinked.
Bibliography:
1. Smith, J. et al. (2023). "AI for Sepsis Prediction." *Nature Med* 29(4), 456-468. [View on Consensus](https://consensus.app/papers/...)
2. ...
Discipline:
- Every inline citation in Sections 1-7 appears here
- Every entry hyperlinked to Consensus
- No phantom or orphan entries
## Section 9: Audit Log
**Length:** Half to full page.
### Consensus searches table
| # | Facet | Query | Results returned | Cited |
|---|---|---|---|---|
| 1 | Established | "..." | 10 | 3 |
| 2 | Stakes | "..." | 10 | 2 |
| ... | ... | ... | ... | ... |
### Plan-tier note
> Detected: Free tier (~10/query). Theoretical ceiling: 5 facets × 10 = 50 papers max from positioning. Actual unique papers: 38 (after deduplication).
### RePORTER searches table
| # | Type | Search text | Window | Projects |
|---|---|---|---|---|
| 1 | Narrow (AND) | "..." | FY 2023-2026 | 23 |
| 2 | Broad (OR) | "..." | FY 2023-2026 | 67 |
### NOSI fetches table
| NOSI | Status | URL |
|---|---|---|
| NOT-HL-25-014 | Fetched, included | [link] |
| NOT-DK-24-009 | Fetch failed | (not included) |
### Summary stats
```
Three counts:
- Queries sent: 7 (5 Consensus + 2 RePORTER)
- Results received: 120 (Consensus 50 + RePORTER 67 + NOSI 3)
- Results cited: 28 (Consensus 22 + RePORTER 5 + NOSI 1)
Failed steps: 1 (NOSI NOT-DK-24-009 fetch — included in NOSI table above)
```
### Tool constraints note
> RePORTER queried via POST (web_fetch is GET-only and would have failed silently). Consensus per-query cap detected as 10 (free tier). 3 consecutive failures threshold not reached this run.
## DOCX Technical Requirements
### Styling
- Body: Arial 12pt
- Headings: Navy (#1a3a5c) for H1/H2
- Table headers: Light blue (#e8f0f8) shading
- NOSI callout: Amber (#F5A623) background with bold border
- Italics for gap quotes (Section 2)
### Hyperlink patterns
```js
new ExternalHyperlink({
link: "https://consensus.app/papers/<id>",
children: [new TextRun({ text: paperTitle, style: "Hyperlink" })],
});
new ExternalHyperlink({
link: "https://reporter.nih.gov/project-details/<id>",
children: [new TextRun({ text: projectNum, style: "Hyperlink" })],
});
new ExternalHyperlink({
link: "https://grants.nih.gov/grants/guide/notice-files/<NOSI>.html",
children: [new TextRun({ text: nosiNumber, style: "Hyperlink" })],
});
```
### Tables (dual widths)
```js
new Table({
columnWidths: [3000, 2000, 1500, 2500], // EMU
rows: rows.map(r => new TableRow({
children: r.cells.map(c => new TableCell({
width: { size: c.width, type: WidthType.DXA },
shading: { type: ShadingType.CLEAR, color: "auto", fill: c.fill || "auto" },
children: [new Paragraph(c.text)],
})),
})),
});
```
### Validation
After save:
```bash
python scripts/office/validate.py output.docx
```
If validation fails: unpack DOCX (it's a ZIP), inspect document.xml, fix the offending XML, repack.
## Citations (7 sources)
1. **`docx` Node.js library — github.com/dolanmiu/docx (MIT).** Authoritative API source for Paragraph, Table, ExternalHyperlink patterns.
2. **NIH OER, *Writing the NIH Grant Application: Strategies for Success* (2022 ed.).** Source for the Section 2 "draft Significance/Innovation language" pattern. Mirrors NIH's own application sections.
3. **Russell, S. W. & Morrison, D. C., *The Grant Application Writer's Workbook* (Grant Writers' Seminars, multiple eds.).** Source for the differentiation-paragraph (Section 5) discipline. "Reviewers spend 30 seconds on differentiation; make it sharp."
4. **PRISMA 2020 Statement — Page, M. J. et al., *BMJ* 372, 2021.** Source for audit-log section requirements. Every search query + filter + result count must be reproducible.
5. **NIH RePORTER documentation + portfolios.** Source for the institute mission summaries that anchor Section 3 interpretation. Each institute publishes mission + priority areas.
6. **Heggeness, M. L., "What Makes a Successful Grant Application" — *Nature Human Behaviour* 5, 2021.** Empirical meta-analysis. Source for "program officer contact is #1 predictor of submission success after scientific merit" (basis for the mandatory program officer recommendation in Section 7).
7. **Strunk, W. & White, E. B., *Elements of Style* (Macmillan).** Source for "Section 7 closing paragraph" voice — direct, no hedging, named highest-leverage action. "Omit needless words" applies to grant strategy: every sentence should pass the "what is the actionable" test.
FILE:references/nih_mechanism_matching.md
# NIH Mechanism Matching — Career Stage × Scope × Prelim
This reference answers exactly one decision: **given a researcher's career stage, project scope, and preliminary data status, which NIH mechanism(s) should the skill recommend?**
Pair with `scripts/mechanism_matcher.py` for the deterministic implementation.
## The Core Rule
**Career stage alone does NOT determine mechanism.** Scope and prelim data matter equally. The biggest misalignment is "early career + R01 with pilot data" — review reads as overscoped and goes unfunded.
The matching is a 3-dimensional lookup:
```
(career_stage, project_scope, preliminary_data) → mechanism shortlist
```
## Career Stage Buckets (from Q2)
| Bucket | Examples | Eligible mechanisms |
|---|---|---|
| Pre-doctoral | PhD student, T32 trainee | F31, T32 |
| Postdoctoral | F32, K99 candidate | F32, K99/R00, T32 |
| Early career | First R01 candidate, K-awardee | K01/K08/K23, K99/R00 → R00, R21, R03 |
| Independent | Multiple R01s, established lab | R01, R21, R03, R34, R61/R33 |
| Senior PI | R35, P-series | R35, P01, P30, U01 |
## Project Scope Buckets (inferred or asked)
| Scope | Indicator | Mechanism implication |
|---|---|---|
| Solo / pilot | Single site, single hypothesis, <2 yr | R03, R21 |
| Hypothesis-driven independent | Single PI, multi-aim, 4-5 yr | R01 |
| Multi-site cooperative | Multi-PI, multi-site, coord centers | U01 |
| Program-scale | Multiple aims, multiple PIs, sustained | P01, P30, R35 |
| Early/exploratory | High-risk, high-reward | DP1, DP2, R21 |
## Preliminary Data Buckets (from Q3)
| Status | Indicator | Mechanism budget tier |
|---|---|---|
| None | De novo project, no pilot | R03, R21, F-series |
| Pilot | Single-site early findings | R21, K-series, K99/R00 |
| Strong | Multi-experiment, R01-ready | R01, R34 |
| Validated | Multi-site publication-ready | R01, U01, P-series |
## Matching Matrix
The skill applies this matrix in `scripts/mechanism_matcher.py`:
### Pre-doctoral
- **Solo + None → F31** (NRSA individual fellowship)
- **Solo + Pilot → F31, T32 slot**
- **Larger → not eligible as PI** (work as co-investigator on mentor's grant)
### Postdoctoral
- **Solo + None → F32** (postdoc fellowship)
- **Solo + Pilot → F32, K99 candidate prep**
- **Strong + transitioning → K99/R00** (career-transition mechanism, unique to NIH)
### Early career
- **Solo + None/Pilot → K-series** (K01 / K08 / K23 — career development)
- **Solo + Pilot → R21 candidate** (after K-award completion or as parallel)
- **Independent + Pilot → R03, R21**
- **Independent + Strong → R01** (this is the "qualifying" R01 — most career-defining)
- **Resource-constrained env (Q4=3) → R15** (specifically targets this — fund undergrad-involving research)
### Independent
- **Pilot scope + Strong prelim → R01** (the standard)
- **Multi-aim + Strong → R01** (the standard 5-yr R01)
- **Multi-site + Validated → U01** (cooperative agreement)
- **Pilot/early → R21** (exploratory)
- **Clinical trial planning → R34**
- **Early-phase trial → R61/R33** (phased innovation award)
- **High-risk → DP1, DP2** (Pioneer / New Innovator)
### Senior PI
- **Program scope → R35** (outstanding investigator award, unrestricted by topic)
- **Program scope → P01** (program project, multi-PI)
- **Core facility → P30** (center grant)
- **Multi-site cooperative → U01**
## Critical Anti-Patterns
### Career stage alone
Common error: "Early career → K-award". Misses scope. Early-career researcher with **strong prelim** + **independent scope** should target **R01**, not K. K-award is for protected research time; R01 is for hypothesis-driven research budget.
### Scope/prelim mismatch
- "R01 + No prelim" → unfundable. Reviewers will reject as premature.
- "R03 + Strong prelim" → underscoped. Researcher leaves money + scope on the table.
`mechanism_matcher.py` flags both as warnings.
### Environment-blind recommendations
Resource-constrained institution (Q4=3) → consider **R15** specifically. R15 only goes to non-research-intensive institutions. Recommending R01 to a researcher at a resource-constrained college is malpractice — even with strong prelim, their environment can't support R01-scale costs.
### Skipping multi-PI options
For collaborative-by-design projects, **multi-PI R01** (multiple-PI option) is often better than splitting into two R01s. Don't default to single-PI just because it's the default.
## Mechanism Reference Table (Full)
| Mechanism | Budget (annual DC) | Duration | Best for | Prelim needed |
|---|---|---|---|---|
| F31 | $40-50k stipend + tuition | 2-3 yr | Pre-doc training | None-pilot |
| F32 | $48-58k stipend | 2-3 yr | Postdoc training | None-pilot |
| T32 | Institutional | 5-yr renewable | Pre-doc/postdoc training cohort | Institutional commitment |
| R03 | $50k × 2 yr | 2 yr | Small pilot studies | None-pilot |
| R21 | $275k DC × 2 yr | 2 yr | Pilot/exploratory R&D | None-pilot |
| R34 | $450k × 3 yr | 3 yr | Clinical trial planning | Pilot |
| R61/R33 | Phased: $250k + $500k × 2 yr | Up to 5 yr | Phased innovation | Pilot |
| K01 | $100k × 5 yr | 5 yr | Mentored research scientist | Pilot |
| K08 | $100k × 5 yr | 5 yr | Mentored clinical scientist | Pilot |
| K23 | $100k × 5 yr | 5 yr | Mentored patient-oriented | Pilot |
| K99/R00 | $90k mentored + $250k indep | Up to 5 yr | Postdoc → independence | Strong |
| R01 | $250-499k DC × 4-5 yr | 4-5 yr | Hypothesis-driven research | Strong |
| R15 | $300k total × 3 yr | 3 yr | Resource-constrained institutions | Pilot |
| R35 | $750k × 5-8 yr | 5-8 yr | Senior outstanding investigators | Validated |
| P01 | Multi-PI, $1-2M/yr × 5 yr | 5 yr | Program project (3+ PIs) | Validated |
| P30 | Core facility funding | 5 yr | Multi-investigator core | Validated |
| U01 | Cooperative agreement | 5 yr | Multi-site collaborative | Strong-validated |
| DP1 (Pioneer) | $700k × 5 yr | 5 yr | High-risk individual | None (visionary) |
| DP2 (New Innovator) | $300k × 5 yr | 5 yr | Early-career high-risk | Pilot |
## Program Officer Recommendation (Mandatory Per Skill)
After mechanism shortlist is generated, the skill MUST recommend:
> **Contact program officer at {top institute, top match} BEFORE writing.**
>
> Find them at: https://www.nih.gov/institutes-nih/list-nih-institutes-centers-offices → {institute} → Program Officers.
>
> Prepare:
> 1. 1-page specific aims
> 2. Your CV (NIH biosketch format if available)
> 3. 3 specific questions about institute priorities / mechanism fit
>
> Email subject: "Pre-application inquiry: <topic>"
This is the **single highest-leverage step** in any NIH application. Program officers signal "yes, submit" or "no, not the right institute" before you spend months writing. Skipping this is common; the cost is huge.
## Citations (7 sources)
1. **NIH Office of Extramural Research — *Types of Grant Programs* (https://grants.nih.gov/grants/funding/funding_program.htm).** Authoritative source for mechanism definitions + budget ranges + duration. The skill's mechanism reference table mirrors NIH's published catalog.
2. **Sally Rockey, "Mechanism Selection Guide" — *NIH Extramural Nexus*, 2014-2022.** Former NIH Deputy Director's blog series on mechanism selection. Source for the "career stage alone is wrong" framing.
3. **Robertson, M. et al., "Successful K-to-R Transition" — *Academic Medicine* 92(3), 2017.** Empirical analysis of K-award → R01 transitions. Source for the early-career mechanism sequencing (K → R21 → R01) heuristic.
4. **NIH RePORTER Project Database (https://reporter.nih.gov).** The empirical ground truth for what NIH actually funds — institute portfolios, study section ranges, project sizes. The skill queries this via POST API.
5. **Mehrotra, A. et al., "R01 Funding Patterns Across Career Stages" — *JAMA Internal Medicine*, 2020.** Career-stage-stratified analysis of R01 application + funding rates. Source for the "early career + strong prelim → R01 IS appropriate" guidance.
6. **NIH NRSA Fellowship guidelines (https://grants.nih.gov/training/F_files_index.htm).** Authoritative F31/F32 source. Source for the trainee-stage mechanism shortlist.
7. **Heggeness, M. L., "What Makes a Successful Grant Application" — *Nature Human Behaviour* 5, 2021.** Meta-analysis of grant-writing predictors. Source for the program-officer-contact recommendation (#1 predictor of submission success after scientific merit).
FILE:references/reporter_post_patterns.md
# RePORTER POST Patterns + Plan-Tier Detection
This reference answers exactly one decision: **how does the grants skill query NIH RePORTER, and what plan-tier signals does it detect from Consensus responses?**
## The Critical Constraint
**NIH RePORTER's API v2 is POST-only.** `web_fetch` (which performs GET requests) **will not work**. You MUST use `bash_tool` + `curl`.
This is the #1 anti-pattern for the grants skill. If a future maintainer "simplifies" to web_fetch, RePORTER queries silently fail and the skill produces hollow institute-mapping sections.
## RePORTER API Reference
- **Endpoint:** `https://api.reporter.nih.gov/v2/projects/search`
- **Method:** POST
- **Content-Type:** `application/json`
- **No auth required** for public-data queries
- **Rate limit:** documented as 1 q/sec; the skill applies 1+ sec sequential pause per research-pack convention
## Standard POST Templates
### Narrow (AND) — direct overlap
```bash
curl -X POST 'https://api.reporter.nih.gov/v2/projects/search' \
-H 'Content-Type: application/json' \
-d '{
"criteria": {
"fiscal_years": [2023, 2024, 2025, 2026],
"include_active_projects": true,
"advanced_text_search": {
"operator": "AND",
"search_field": "all",
"search_text": "deep learning electronic health records sepsis prediction"
}
},
"limit": 50,
"offset": 0,
"include_fields": [
"project_num",
"project_title",
"agency_ic_admin",
"study_section",
"fiscal_year",
"principal_investigators",
"abstract_text",
"project_terms"
]
}'
```
### Broad (OR) — adjacent work
```bash
curl -X POST 'https://api.reporter.nih.gov/v2/projects/search' \
-H 'Content-Type: application/json' \
-d '{
"criteria": {
"fiscal_years": [2023, 2024, 2025, 2026],
"advanced_text_search": {
"operator": "OR",
"search_field": "all",
"search_text": "machine learning critical care sepsis early warning"
}
},
"limit": 50
}'
```
## Dynamic Fiscal Year Window
NIH fiscal year runs **Oct 1 → Sep 30**. Current FY = year of next Sep 30.
Use `scripts/fiscal_year_calculator.py`:
```bash
python ../scripts/fiscal_year_calculator.py
# Output:
# Current calendar year: 2026
# Current fiscal year: 2026 (Oct 1 2025 - Sep 30 2026)
# Window (current + 3 prior): [2023, 2024, 2025, 2026]
```
**Never hardcode years.** A skill committed in 2025 with hardcoded `[2022, 2023, 2024, 2025]` produces stale results in 2027.
## Institute Tally + Study Section Ranking
After both narrow + broad responses return, aggregate:
### Institute tally
For each project: extract `agency_ic_admin` (the institute code like NCI, NHLBI, NIMH).
```python
from collections import Counter
institute_counts = Counter()
for project in projects:
institute_counts[project['agency_ic_admin']] += 1
top_institutes = institute_counts.most_common(3)
```
Surface in DOCX Section 3 as ranked table with project counts + brief institute mission.
### Study section ranking
For each project: extract `study_section`.
```python
study_section_counts = Counter()
for project in projects:
section = project.get('study_section', '')
if section: # Some projects unassigned
study_section_counts[section] += 1
top_sections = study_section_counts.most_common(2)
```
Surface in DOCX Section 6.
## NOSI Discovery from RePORTER Results
NOSI (Notice of Special Interest) numbers appear as `NOT-*` in project abstracts, project terms, or related-FOA fields. Parse with regex:
```python
import re
NOSI_RE = re.compile(r'NOT-[A-Z]{2,3}-\d{2}-\d{3}')
nosi_numbers = set()
for project in projects:
abstract = project.get('abstract_text', '')
nosi_numbers.update(NOSI_RE.findall(abstract))
```
For each NOSI number, fetch via `web_fetch` (NOSIs have predictable URLs):
```
https://grants.nih.gov/grants/guide/notice-files/{NOSI_NUMBER}.html
```
If fetch fails: log `[NOSI {number} — fetch failed, not included]`. Never fabricate NOSI details.
## Plan-Tier Detection (Consensus)
Consensus has tiered plans with different per-query result caps. The skill detects from response text patterns:
| Pattern in response | Tier | Per-query cap |
|---|---|---|
| `"Showing top 10"` / `"upgrade for more"` | Free | 10 results |
| Receives 20 results without "showing top" | Pro | 20 results |
| Receives ≤3 results consistently | Unauthenticated / API quota issue | 3 results |
| No response / 401 / 403 | Auth failure | n/a |
Surface at end of Phase 2A in DOCX audit log:
> **Plan tier detected: Free** (Consensus returns ~10 results per query, capped). Total positioning landscape: 5 facets × 10 results = ~50 papers max. For deeper coverage, consider Consensus Pro (20/query).
This calibrates user expectations BEFORE they read the DOCX and wonder why coverage seems thin.
## Sequential Execution Discipline
Per research-pack convention: **1 q/sec, never parallelize.**
- 5 Consensus searches (Phase 2A) sequential — pause 1+ sec between
- 2 RePORTER POST searches (narrow + broad) sequential
- N NOSI `web_fetch` calls sequential
Each call records timestamp via `citation_tracker.py`; second call within 1s is rejected.
Total Phase 2 wall-clock time: ~7-10 sec for searches + however long NOSI fetches take.
## Error Handling
| Failure | Handling |
|---|---|
| Consensus 429 (rate limit) | Wait 3s, retry once, log to audit |
| Consensus 0 results for a facet | Surface explicitly in DOCX positioning section; mark `[no results — verify terminology]` |
| RePORTER 5xx | Retry once after 3s; if still failing, log and continue with what's available |
| RePORTER <5 results on narrow | Document low count; rely on broad OR for coverage |
| NOSI fetch fails | `[NOSI {number} — fetch failed]`; never fabricate |
| 3 consecutive failures across tools | Halt; alert researcher with what's missing |
| Auth failure (401/403) | Halt; tell user to check API key or MCP connection |
## Citations (7 sources)
1. **NIH RePORTER API v2 documentation — https://api.reporter.nih.gov/documents/Data%20Element%20Descriptions.pdf.** Authoritative spec for POST endpoint, field definitions, fiscal-year filter semantics. The skill's curl templates are direct applications.
2. **NIH Office of Extramural Research — *NIH Guide for Grants and Contracts* (https://grants.nih.gov/grants/guide).** Source for NOSI / FOA URL structure. NOSI naming conventions (`NOT-{IC}-{YY}-{NNN}`) are documented here.
3. **`praw` library + Reddit API community guidance.** Source for the "1 q/sec is the polite default" pattern that the skill applies to RePORTER even though RePORTER's documented limits are looser. Politeness with shared infrastructure.
4. **Mike Cohen, "Exponential Backoff and Jitter" — AWS Architecture Blog, 2015.** Source for the "wait 3s + retry once" retry pattern. Research-workflow scale doesn't justify exponential backoff.
5. **`curl` documentation (https://curl.se/docs/manual.html).** Source for POST body + Content-Type header syntax. The skill's curl templates follow `curl --help`.
6. **Maynez et al., "On Faithfulness and Factuality in Abstractive Summarization" — ACL 2020.** Source for the source-discipline rule that justifies refusing to fabricate NOSI details when fetch fails. LLMs hallucinate plausible-looking NIH NOSI numbers; refuse.
7. **Susskind, D., "Show your work" — *Communications of the ACM*, 2024.** Source for the audit-log section's role: transparent surfacing of what was queried, what was returned, what was cited. The audit-log table in DOCX Section 9 is this principle's implementation.
FILE:scripts/citation_tracker.py
#!/usr/bin/env python3
"""citation_tracker.py — JSON-backed three-count audit for grants runs.
Stdlib-only. Mirrors litreview's tracker but extended for grants's
multi-source workflow (Consensus + RePORTER + NOSI fetches).
Tracked counts:
- consensus_searches (5 facets sent)
- consensus_received (papers shown across facets)
- consensus_cited (papers cited in DOCX)
- reporter_searches (typically 2: narrow + broad)
- reporter_projects (projects returned across both)
- reporter_cited (projects cited in DOCX)
- nosi_fetches (NOT-* fetches attempted)
- nosi_succeeded (fetches that returned content)
Enforces 1s sequential gap on Consensus searches (research-pack convention).
Persists at ~/.grants_sessions/<session>.json.
Usage:
python citation_tracker.py --action start --session grants-20260515 --topic "sepsis prediction"
python citation_tracker.py --action record_consensus_search --session ... --facet established --query "..." --tier free
python citation_tracker.py --action record_consensus_received --session ... --count 10
python citation_tracker.py --action record_consensus_cited --session ... --url "https://consensus.app/..."
python citation_tracker.py --action record_reporter_search --session ... --type narrow --query "..." --projects 23
python citation_tracker.py --action record_reporter_cited --session ... --project-num "R01HL12345"
python citation_tracker.py --action record_nosi --session ... --nosi "NOT-HL-25-014" --status fetched
python citation_tracker.py --action status --session ...
python citation_tracker.py --action close --session ...
"""
import argparse
import json
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional
SESSIONS_DIR = Path.home() / ".grants_sessions"
MIN_CONSENSUS_GAP_SECONDS = 1.0
def session_path(name: str) -> Path:
return SESSIONS_DIR / f"{name}.json"
def load_session(name: str) -> Dict[str, Any]:
p = session_path(name)
if not p.exists():
raise FileNotFoundError(f"Session not found: {name}")
return json.loads(p.read_text(encoding="utf-8"))
def save_session(name: str, data: Dict[str, Any]) -> None:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
session_path(name).write_text(json.dumps(data, indent=2), encoding="utf-8")
def now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
def now_ts() -> float:
return datetime.now(timezone.utc).timestamp()
def action_start(name: str, topic: Optional[str]) -> Dict[str, Any]:
if session_path(name).exists():
raise FileExistsError(f"Session already exists: {name}")
data: Dict[str, Any] = {
"session": name,
"topic": topic or "",
"started_at": now_iso(),
"ended_at": None,
"consensus_tier": None,
"consensus_searches": [],
"consensus_received_log": [],
"consensus_cited": [],
"reporter_searches": [],
"reporter_cited": [],
"nosi_fetches": [],
"counts": {
"consensus_searches": 0,
"consensus_received": 0,
"consensus_cited": 0,
"reporter_searches": 0,
"reporter_projects": 0,
"reporter_cited": 0,
"nosi_fetches": 0,
"nosi_succeeded": 0,
},
}
save_session(name, data)
return data
def action_record_consensus_search(name: str, facet: str, query: str, tier: Optional[str]) -> Dict[str, Any]:
data = load_session(name)
if data["consensus_searches"]:
last_ts = data["consensus_searches"][-1].get("ts", 0)
gap = now_ts() - last_ts
if gap < MIN_CONSENSUS_GAP_SECONDS:
raise RuntimeError(
f"Consensus sequential discipline violated: {gap:.2f}s gap (need >= {MIN_CONSENSUS_GAP_SECONDS}s). "
f"Wait {MIN_CONSENSUS_GAP_SECONDS - gap:.2f}s more."
)
if tier and not data["consensus_tier"]:
data["consensus_tier"] = tier
data["consensus_searches"].append({"facet": facet, "query": query, "tier": tier, "at": now_iso(), "ts": now_ts()})
data["counts"]["consensus_searches"] += 1
save_session(name, data)
return data
def action_record_consensus_received(name: str, count: int) -> Dict[str, Any]:
data = load_session(name)
data["consensus_received_log"].append({"count": count, "at": now_iso()})
data["counts"]["consensus_received"] += count
save_session(name, data)
return data
def action_record_consensus_cited(name: str, url: str) -> Dict[str, Any]:
data = load_session(name)
if any(p["url"] == url for p in data["consensus_cited"]):
return data
data["consensus_cited"].append({"url": url, "at": now_iso()})
data["counts"]["consensus_cited"] += 1
save_session(name, data)
return data
def action_record_reporter_search(name: str, search_type: str, query: str, projects: int) -> Dict[str, Any]:
data = load_session(name)
data["reporter_searches"].append({"type": search_type, "query": query, "projects_returned": projects, "at": now_iso()})
data["counts"]["reporter_searches"] += 1
data["counts"]["reporter_projects"] += projects
save_session(name, data)
return data
def action_record_reporter_cited(name: str, project_num: str) -> Dict[str, Any]:
data = load_session(name)
if any(p["project_num"] == project_num for p in data["reporter_cited"]):
return data
data["reporter_cited"].append({"project_num": project_num, "at": now_iso()})
data["counts"]["reporter_cited"] += 1
save_session(name, data)
return data
def action_record_nosi(name: str, nosi: str, status: str) -> Dict[str, Any]:
data = load_session(name)
data["nosi_fetches"].append({"nosi": nosi, "status": status, "at": now_iso()})
data["counts"]["nosi_fetches"] += 1
if status == "fetched" or status == "succeeded":
data["counts"]["nosi_succeeded"] += 1
save_session(name, data)
return data
def action_status(name: str) -> Dict[str, Any]:
return load_session(name)
def action_close(name: str) -> Dict[str, Any]:
data = load_session(name)
if data.get("ended_at") is None:
data["ended_at"] = now_iso()
save_session(name, data)
return data
def action_list() -> List[Dict[str, Any]]:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
out: List[Dict[str, Any]] = []
for p in sorted(SESSIONS_DIR.glob("*.json")):
try:
d = json.loads(p.read_text(encoding="utf-8"))
out.append({
"session": d.get("session", p.stem),
"topic": d.get("topic", ""),
"tier": d.get("consensus_tier"),
"counts": d.get("counts", {}),
"ended_at": d.get("ended_at"),
})
except (OSError, json.JSONDecodeError):
continue
return out
def render_status_human(data: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Session: {data['session']}")
out.append(f"Topic: {data.get('topic', '(unset)')}")
out.append(f"Consensus tier: {data.get('consensus_tier') or '(not detected)'}")
out.append(f"Started: {data['started_at']}")
out.append(f"Ended: {data.get('ended_at') or '(active)'}")
out.append("")
c = data["counts"]
out.append("Counts:")
out.append(f" Consensus searches: {c['consensus_searches']}")
out.append(f" Consensus received: {c['consensus_received']}")
out.append(f" Consensus cited: {c['consensus_cited']}")
out.append(f" RePORTER searches: {c['reporter_searches']}")
out.append(f" RePORTER projects: {c['reporter_projects']}")
out.append(f" RePORTER cited: {c['reporter_cited']}")
out.append(f" NOSI fetches: {c['nosi_fetches']} ({c['nosi_succeeded']} succeeded)")
out.append("")
out.append("Audit block (paste in DOCX Section 9):")
out.append(
f" Three counts — Queries sent: {c['consensus_searches'] + c['reporter_searches']} "
f"(Consensus {c['consensus_searches']}, RePORTER {c['reporter_searches']}). "
f"Results received: {c['consensus_received'] + c['reporter_projects']} "
f"(Consensus {c['consensus_received']} + RePORTER {c['reporter_projects']}). "
f"Results cited: {c['consensus_cited'] + c['reporter_cited']} "
f"(Consensus {c['consensus_cited']} + RePORTER {c['reporter_cited']}). "
f"NOSI fetches: {c['nosi_succeeded']}/{c['nosi_fetches']} succeeded."
)
return "\n".join(out)
def render_list_human(rows: List[Dict[str, Any]]) -> str:
if not rows:
return "(no sessions)"
out: List[str] = []
out.append(f"{'session':<35s} {'tier':<5s} {'C-srch':>6s} {'C-rcvd':>6s} {'C-cit':>5s} {'R-srch':>6s} {'R-prj':>5s} {'R-cit':>5s} {'NOSI':>4s}")
out.append("-" * 90)
for r in rows:
c = r["counts"]
out.append(
f"{r['session']:<35s} {(r.get('tier') or '—'):<5s} "
f"{c.get('consensus_searches', 0):>6d} {c.get('consensus_received', 0):>6d} "
f"{c.get('consensus_cited', 0):>5d} {c.get('reporter_searches', 0):>6d} "
f"{c.get('reporter_projects', 0):>5d} {c.get('reporter_cited', 0):>5d} "
f"{c.get('nosi_succeeded', 0):>4d}"
)
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument(
"--action",
required=True,
choices=[
"start", "record_consensus_search", "record_consensus_received", "record_consensus_cited",
"record_reporter_search", "record_reporter_cited", "record_nosi",
"status", "list", "close",
],
)
parser.add_argument("--session")
parser.add_argument("--topic")
parser.add_argument("--facet")
parser.add_argument("--query")
parser.add_argument("--tier")
parser.add_argument("--count", type=int)
parser.add_argument("--url")
parser.add_argument("--type", dest="search_type")
parser.add_argument("--projects", type=int)
parser.add_argument("--project-num")
parser.add_argument("--nosi")
parser.add_argument("--status")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
try:
if args.action == "start":
result = action_start(args.session, args.topic)
elif args.action == "record_consensus_search":
result = action_record_consensus_search(args.session, args.facet, args.query, args.tier)
elif args.action == "record_consensus_received":
result = action_record_consensus_received(args.session, args.count)
elif args.action == "record_consensus_cited":
result = action_record_consensus_cited(args.session, args.url)
elif args.action == "record_reporter_search":
result = action_record_reporter_search(args.session, args.search_type, args.query, args.projects)
elif args.action == "record_reporter_cited":
result = action_record_reporter_cited(args.session, args.project_num)
elif args.action == "record_nosi":
result = action_record_nosi(args.session, args.nosi, args.status)
elif args.action == "status":
result = action_status(args.session)
elif args.action == "close":
result = action_close(args.session)
else:
result = action_list()
except (FileNotFoundError, FileExistsError, RuntimeError) as e:
print(f"error: {e}", file=sys.stderr); return 2
if args.output == "json":
print(json.dumps(result, indent=2, default=str))
else:
if args.action == "list":
print(render_list_human(result))
else:
print(render_status_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/fiscal_year_calculator.py
#!/usr/bin/env python3
"""fiscal_year_calculator.py — Current NIH fiscal year + lookback window.
Stdlib-only. NIH FY = year of next Sep 30. October starts a new FY.
NIH RePORTER queries need a `fiscal_years` array. Hardcoding values produces
stale skill behavior over time. This script computes them at runtime.
Default window: current FY + 3 prior (4 years total). User can override.
Usage:
python fiscal_year_calculator.py
python fiscal_year_calculator.py --window 4 --output json
python fiscal_year_calculator.py --reference-date 2026-10-15
python fiscal_year_calculator.py --reference-date 2026-09-15
"""
import argparse
import json
import sys
from datetime import date, datetime
from typing import Any, Dict, List, Optional
def fiscal_year(reference: date) -> int:
"""Return the fiscal year that the given date falls within.
NIH FY runs Oct 1 → Sep 30. FY 2026 = Oct 1 2025 → Sep 30 2026.
"""
if reference.month >= 10:
return reference.year + 1
return reference.year
def calculate(reference: date, window_years: int) -> Dict[str, Any]:
if window_years < 1:
raise ValueError(f"--window must be >= 1, got {window_years}")
current_fy = fiscal_year(reference)
years = list(range(current_fy - window_years + 1, current_fy + 1))
fy_start_date = date(current_fy - 1, 10, 1)
fy_end_date = date(current_fy, 9, 30)
return {
"reference_date": reference.isoformat(),
"calendar_year": reference.year,
"current_fiscal_year": current_fy,
"current_fy_start": fy_start_date.isoformat(),
"current_fy_end": fy_end_date.isoformat(),
"window_years": window_years,
"window_fiscal_years": years,
"reporter_payload_snippet": f'"fiscal_years": {json.dumps(years)}',
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Reference date: {result['reference_date']}")
out.append(f"Calendar year: {result['calendar_year']}")
out.append(f"Current fiscal year: FY {result['current_fiscal_year']} ({result['current_fy_start']} → {result['current_fy_end']})")
out.append(f"Window: {result['window_years']} years")
out.append(f"FY values for query: {result['window_fiscal_years']}")
out.append("")
out.append("Use in RePORTER POST body:")
out.append(f" {result['reporter_payload_snippet']}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--reference-date", help="ISO date (default: today)")
parser.add_argument("--window", type=int, default=4, help="Years to include (default: 4 = current + 3 prior)")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.reference_date:
try:
ref = datetime.strptime(args.reference_date, "%Y-%m-%d").date()
except ValueError:
print(f"error: invalid --reference-date '{args.reference_date}', expected YYYY-MM-DD", file=sys.stderr); return 2
else:
ref = date.today()
try:
result = calculate(ref, args.window)
except ValueError as e:
print(f"error: {e}", file=sys.stderr); return 2
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/mechanism_matcher.py
#!/usr/bin/env python3
"""mechanism_matcher.py — NIH mechanism shortlist from career stage + scope + prelim.
Stdlib-only. The skill must NOT recommend mechanisms by career stage alone —
that's the most common mistake. Matching is 3-dimensional:
(career_stage, project_scope, preliminary_data, environment) → mechanism shortlist
See references/nih_mechanism_matching.md for the full matrix.
NO LLM CALLS. Pure rule-based lookup.
Usage:
python mechanism_matcher.py --career-stage early_career --prelim-data pilot \\
--environment r01_eligible --scope single_site
python mechanism_matcher.py --sample
"""
import argparse
import json
import sys
from typing import Any, Dict, List
VALID_CAREER_STAGES = ["pre_doctoral", "postdoctoral", "early_career", "independent", "senior"]
VALID_PRELIM = ["none", "pilot", "strong", "validated"]
VALID_ENVIRONMENTS = ["r01_eligible", "mid_tier", "resource_constrained", "industry_collab"]
VALID_SCOPES = ["solo_pilot", "single_site", "multi_aim", "multi_site", "program_scale", "high_risk"]
MECHANISMS = {
"F31": {"budget": "$40-50k stipend + tuition × 2-3 yr", "prelim": "None-pilot", "best_for": "Pre-doc training"},
"F32": {"budget": "$48-58k stipend × 2-3 yr", "prelim": "None-pilot", "best_for": "Postdoc training"},
"T32": {"budget": "Institutional × 5-yr renewable", "prelim": "Institutional", "best_for": "Pre-doc/postdoc cohort"},
"R03": {"budget": "$50k × 2 yr", "prelim": "None-pilot", "best_for": "Small pilot studies"},
"R21": {"budget": "$275k DC × 2 yr", "prelim": "None-pilot", "best_for": "Pilot/exploratory R&D"},
"R34": {"budget": "$450k × 3 yr", "prelim": "Pilot", "best_for": "Clinical trial planning"},
"R61/R33": {"budget": "Phased ($250k + $500k × 2 yr)", "prelim": "Pilot", "best_for": "Phased innovation"},
"K01": {"budget": "$100k × 5 yr", "prelim": "Pilot", "best_for": "Mentored research scientist"},
"K08": {"budget": "$100k × 5 yr", "prelim": "Pilot", "best_for": "Mentored clinical scientist"},
"K23": {"budget": "$100k × 5 yr", "prelim": "Pilot", "best_for": "Mentored patient-oriented research"},
"K99/R00": {"budget": "$90k + $250k × 3 yr", "prelim": "Strong", "best_for": "Postdoc → independence transition"},
"R01": {"budget": "$250-499k DC × 4-5 yr", "prelim": "Strong", "best_for": "Hypothesis-driven research"},
"R15": {"budget": "$300k total × 3 yr", "prelim": "Pilot", "best_for": "Resource-constrained institutions only"},
"R35": {"budget": "$750k × 5-8 yr", "prelim": "Validated", "best_for": "Senior outstanding investigators"},
"P01": {"budget": "Multi-PI, $1-2M/yr × 5 yr", "prelim": "Validated", "best_for": "Program project (3+ PIs)"},
"P30": {"budget": "Core facility funding × 5 yr", "prelim": "Validated", "best_for": "Multi-investigator core"},
"U01": {"budget": "Cooperative agreement, varies", "prelim": "Strong-validated", "best_for": "Multi-site collaborative"},
"DP1": {"budget": "$700k × 5 yr", "prelim": "None (visionary)", "best_for": "Pioneer Award — high-risk individual"},
"DP2": {"budget": "$300k × 5 yr", "prelim": "Pilot", "best_for": "New Innovator — early-career high-risk"},
}
def match(career_stage: str, prelim_data: str, environment: str, scope: str) -> Dict[str, Any]:
if career_stage not in VALID_CAREER_STAGES:
raise ValueError(f"Invalid --career-stage. Pick from: {VALID_CAREER_STAGES}")
if prelim_data not in VALID_PRELIM:
raise ValueError(f"Invalid --prelim-data. Pick from: {VALID_PRELIM}")
if environment not in VALID_ENVIRONMENTS:
raise ValueError(f"Invalid --environment. Pick from: {VALID_ENVIRONMENTS}")
if scope not in VALID_SCOPES:
raise ValueError(f"Invalid --scope. Pick from: {VALID_SCOPES}")
recommendations: List[Dict[str, Any]] = []
warnings: List[str] = []
# === Pre-doctoral ===
if career_stage == "pre_doctoral":
if prelim_data in ("none", "pilot") and scope in ("solo_pilot", "single_site"):
recommendations.append({"mechanism": "F31", "rationale": "Pre-doc + pilot scope → NRSA individual fellowship"})
recommendations.append({"mechanism": "T32", "rationale": "Pre-doc + institutional context → T32 training slot if available"})
else:
warnings.append("Pre-doctoral PI eligibility is limited. Consider co-investigator role on mentor's grant.")
# === Postdoctoral ===
elif career_stage == "postdoctoral":
if prelim_data == "none":
recommendations.append({"mechanism": "F32", "rationale": "Postdoc + no prelim → NRSA F32 fellowship"})
if prelim_data == "pilot":
recommendations.append({"mechanism": "F32", "rationale": "Postdoc + pilot data → F32"})
recommendations.append({"mechanism": "K99/R00", "rationale": "Postdoc + pilot → K99/R00 candidate prep (top mechanism)"})
if prelim_data == "strong":
recommendations.append({"mechanism": "K99/R00", "rationale": "Strong prelim + postdoc-transitioning → K99/R00 is the highest-value mechanism for this stage"})
# === Early career ===
elif career_stage == "early_career":
if prelim_data in ("none", "pilot"):
if environment == "resource_constrained":
recommendations.append({"mechanism": "R15", "rationale": "Resource-constrained env + early career → R15 (specifically targets this; R01 not competitive without env match)"})
recommendations.append({"mechanism": "K01", "rationale": "Early career + pilot prelim → K-series for career development"})
recommendations.append({"mechanism": "K08", "rationale": "Early career (clinical) + pilot → K08 mentored clinical"})
recommendations.append({"mechanism": "K23", "rationale": "Early career patient-oriented → K23"})
recommendations.append({"mechanism": "R21", "rationale": "Early career + pilot scope → R21 exploratory"})
if prelim_data == "strong" and scope in ("single_site", "multi_aim"):
recommendations.append({"mechanism": "R01", "rationale": "Strong prelim + independent scope → R01 (the qualifying R01)"})
if scope == "multi_aim":
warnings.append("Multi-aim R01 at early career is ambitious; consider mentored R01 with senior co-PI")
if scope == "high_risk":
recommendations.append({"mechanism": "DP2", "rationale": "Early career + high-risk → New Innovator (DP2)"})
# === Independent ===
elif career_stage == "independent":
if prelim_data in ("none", "pilot") and scope == "solo_pilot":
recommendations.append({"mechanism": "R03", "rationale": "Independent + pilot scope → R03 small pilot"})
recommendations.append({"mechanism": "R21", "rationale": "Independent + exploratory → R21"})
warnings.append("R01 NOT recommended without strong prelim — reviewers will reject as premature")
if prelim_data == "strong":
if scope == "multi_aim" or scope == "single_site":
recommendations.append({"mechanism": "R01", "rationale": "Independent + strong prelim + hypothesis-driven → R01 (standard)"})
if scope == "multi_site":
recommendations.append({"mechanism": "U01", "rationale": "Multi-site + strong prelim → U01 cooperative agreement"})
if prelim_data == "validated" and scope == "multi_site":
recommendations.append({"mechanism": "U01", "rationale": "Validated + multi-site → U01"})
recommendations.append({"mechanism": "R01", "rationale": "Validated + multi-site → R01 alternate path"})
if scope == "high_risk":
recommendations.append({"mechanism": "DP1", "rationale": "High-risk + independent → Pioneer Award"})
if scope == "single_site" and prelim_data == "pilot":
recommendations.append({"mechanism": "R34", "rationale": "Clinical trial planning + pilot → R34"})
# === Senior PI ===
elif career_stage == "senior":
if scope == "program_scale":
recommendations.append({"mechanism": "R35", "rationale": "Senior + program scope → R35 outstanding investigator (unrestricted by topic)"})
recommendations.append({"mechanism": "P01", "rationale": "Senior + multi-PI program → P01"})
if scope == "multi_site":
recommendations.append({"mechanism": "U01", "rationale": "Senior + multi-site → U01"})
if "core" in scope or scope == "program_scale":
recommendations.append({"mechanism": "P30", "rationale": "Senior + core facility → P30"})
if scope in ("multi_aim", "single_site") and prelim_data in ("strong", "validated"):
recommendations.append({"mechanism": "R01", "rationale": "Senior PI continuing R01 portfolio"})
if not recommendations:
warnings.append("No mechanism shortlist matched. Likely inputs are inconsistent (e.g., pre-doctoral + senior-scope). Re-check the answer combinations.")
# Enrich with mechanism details
enriched = []
for rec in recommendations:
m = rec["mechanism"]
info = MECHANISMS.get(m, {})
enriched.append({
"mechanism": m,
"rationale": rec["rationale"],
"budget": info.get("budget", ""),
"prelim_needed": info.get("prelim", ""),
"best_for": info.get("best_for", ""),
})
return {
"inputs": {
"career_stage": career_stage,
"prelim_data": prelim_data,
"environment": environment,
"scope": scope,
},
"recommendations": enriched,
"warnings": warnings,
"program_officer_note": "MANDATORY: contact program officer at top institute before writing. Find via https://www.nih.gov/institutes-nih/list-nih-institutes-centers-offices",
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append("Inputs:")
for k, v in result["inputs"].items():
out.append(f" {k}: {v}")
out.append("")
if result["recommendations"]:
out.append(f"Recommended mechanisms ({len(result['recommendations'])}):")
for r in result["recommendations"]:
out.append(f"")
out.append(f" → {r['mechanism']}")
out.append(f" Rationale: {r['rationale']}")
out.append(f" Budget: {r['budget']}")
out.append(f" Prelim: {r['prelim_needed']}")
out.append(f" Best for: {r['best_for']}")
else:
out.append("No mechanisms recommended (see warnings)")
if result["warnings"]:
out.append("")
out.append("Warnings:")
for w in result["warnings"]:
out.append(f" ! {w}")
out.append("")
out.append(result["program_officer_note"])
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--career-stage", choices=VALID_CAREER_STAGES)
parser.add_argument("--prelim-data", choices=VALID_PRELIM)
parser.add_argument("--environment", choices=VALID_ENVIRONMENTS)
parser.add_argument("--scope", choices=VALID_SCOPES)
parser.add_argument("--sample", action="store_true")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = match("early_career", "pilot", "r01_eligible", "single_site")
elif args.career_stage and args.prelim_data and args.environment and args.scope:
try:
result = match(args.career_stage, args.prelim_data, args.environment, args.scope)
except ValueError as e:
print(f"error: {e}", file=sys.stderr); return 2
else:
parser.print_help(); return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Phỏng vấn người dùng liên tục về kế hoạch hoặc thiết kế cho đến khi thống nhất, giải quyết từng nhánh của cây quyết định.
---
name: grill-me
description: Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
license: MIT
metadata:
derived_from: "https://github.com/mattpocock/skills/tree/main/skills/productivity/grill-me"
original_author: "Matt Pocock (@mattpocock)"
original_license: MIT
voice: "Matt Pocock — relentless, one-at-a-time, explores-codebase-first"
version: 1.0.0
---
# Grill Me
> Derived from [Matt Pocock's grill-me](https://github.com/mattpocock/skills/tree/main/skills/productivity/grill-me) (MIT). Matt's interview discipline preserved verbatim. Additions: extraction + question + session tools + references + cs-* wrapper (see [references/companion_tooling.md](references/companion_tooling.md)).
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
Ask the questions one at a time.
If a question can be answered by exploring the codebase, explore the codebase instead.
## Rules (preserved + amplified)
1. **One question per turn.** Never bundle.
2. **Provide a recommended answer with each question.** Defaulting to "what do you think?" is lazy.
3. **Explore the codebase before asking.** If `grep` / `Read` resolves it, do that first. Saves a turn.
4. **Walk the tree depth-first.** Finish a branch before opening another.
5. **Track dependencies.** If decision B depends on decision A, ask A first.
## Workflow
1. User provides a plan or design (or path to one).
2. Run `scripts/decision_tree_extractor.py` to extract branches.
3. Run `scripts/question_generator.py` to produce the question list with recommendations.
4. Start a session: `scripts/grill_session_tracker.py --action start`.
5. Walk the tree, one question at a time, recording answers in the session.
6. When all branches resolved: report "shared understanding reached" + the locked-in decisions.
## Output Pattern
Per question turn:
```
Q[i]/[total]: [question]
Recommended answer: [your call + 1-sentence rationale]
(Or: I explored the codebase and found [evidence]. Confirm?)
```
## Tooling
See [references/companion_tooling.md](references/companion_tooling.md). Tools: extractor + generator + tracker. Agent: `cs-grill-master`. Command: `/cs:grill-me`.
---
**Version:** 1.0.0
**Derived:** Matt Pocock (MIT) + this repo's wrapper
FILE:references/companion_tooling.md
# Companion Tooling
Interrogation tools + cs-* wrapper layered on top of Matt's grill-me skill.
## Validation Tools (stdlib Python)
| Tool | Purpose | Run when |
|---|---|---|
| `scripts/decision_tree_extractor.py` | Scan a plan doc for decision branches (intent / choice / open / tradeoff / dependency / question) | Starting a grill session — see what's there to interrogate |
| `scripts/question_generator.py` | Generate forcing questions from extracted branches with recommended answers + dependency-aware ordering | Producing the question list for a grill session |
| `scripts/grill_session_tracker.py` | JSON-backed session storage in `~/.grill_sessions/` — track answers across turns, resume sessions | Running a multi-turn grill (most real grills) |
All three:
- Stdlib-only
- Run with embedded sample if no input provided
- Output text or JSON (`--output json`)
## Session Storage
`grill_session_tracker.py` persists state to `~/.grill_sessions/<name>.json`. This enables:
- Resume a grill across days
- Switch between concurrent grills (e.g., per project)
- Audit which decisions were resolved when
- Generate a "decisions locked" summary at end
## cs-grill-master Persona Agent
Lives at `../agents/cs-grill-master.md`. Voice: relentless, one-question-at-a-time, codebase-exploration-first.
The persona's hard rule: **never bundle questions**. Even when there are 10 obvious follow-ups, ask one, wait for answer, then ask the next.
## `/cs:grill-me` Slash Command
Lives at `../commands/cs-grill-me.md`. Activation pattern:
1. `/cs:grill-me <path-to-plan>` — start grill session on plan doc
2. Persona asks Q1 with recommended answer
3. User answers
4. Persona asks Q2
5. ...continues until all branches resolved
## Why Wrap Matt's Original
Matt's grill-me skill is intentionally minimal (3 sentences). The wrapper adds:
1. **Automatic branch extraction** — manually identifying decision branches is the slow part; the extractor does it deterministically
2. **Question templating** — consistent question patterns per branch kind (intent / choice / tradeoff)
3. **Session persistence** — grills span days; persistence prevents re-asking + losing context
4. **Recommendation defaults** — every question carries a recommended answer (per Matt's "provide your recommended answer" rule)
## Attribution
Original: [matt-pocock/skills/skills/productivity/grill-me](https://github.com/mattpocock/skills/tree/main/skills/productivity/grill-me) (MIT).
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — grill-me** (https://github.com/mattpocock/skills/, MIT) — the upstream source
- **Socratic Method** (5th-century BC) — interrogation as truth-finding; one-question-at-a-time discipline
- **YC office hours format** (Y Combinator) — forcing questions for founders; "what's blocking this?" + "why this and not Y?"
- **Cockburn, A. — "Writing Effective Use Cases"** (2000) — exploring decision branches in requirements
- **Fournier, C. — "The Manager's Path"** (2017) — interview discipline for hard decisions
- **Larson, W. — "An Elegant Puzzle"** (2019) — engineering manager decision-making patterns
- **5 Whys (Toyota Production System)** — Sakichi Toyoda — sequential interrogation for root cause
FILE:references/forcing_question_patterns.md
# Forcing-Question Patterns for Plan Interrogation
This reference answers exactly one decision: **what makes a question "forcing" vs "soft", and how do we ask forcing questions that resolve decisions?**
Pair with `scripts/question_generator.py` for templated forcing questions.
## What Makes a Question "Forcing"
A forcing question:
1. **Cannot be answered with "yes"/"no"** without follow-up
2. **Names the alternative** — "X or Y" not "is X right?"
3. **Demands evidence** — "what's the kill criterion?" not "what do you think?"
4. **Removes the escape hatch** — asks the trade-off explicitly
Soft questions let the answerer evade. Forcing questions don't.
## Six Forcing-Question Patterns
### Pattern 1: "Why X and not Y?"
When user says "We'll use Postgres" — forcing question: "Why Postgres and not MySQL?"
The forcing element: requires the answerer to articulate the alternative + the rejection reason. Reveals whether the choice was deliberate or default.
**Soft variant (bad):** "Are you sure about Postgres?"
### Pattern 2: "What's the kill criterion?"
When user says "We'll try approach X" — forcing question: "What would convince you X is wrong?"
The forcing element: requires the answerer to commit to falsifiability ahead of time. Prevents motivated reasoning later.
**Soft variant (bad):** "What if it doesn't work?"
### Pattern 3: "What's blocking the decision?"
When user says "TBD" or "open question" — forcing question: "What input is missing, and when does it arrive?"
The forcing element: separates "haven't decided" from "can't decide yet". Most TBDs are decideable now under uncertainty.
**Soft variant (bad):** "Have you thought about that?"
### Pattern 4: "Which side of the trade-off?"
When user says "trade-off between A and B" — forcing question: "Which side are you optimizing for, and what's the deciding constraint?"
The forcing element: requires picking. "Both" is not an option for actual trade-offs.
**Soft variant (bad):** "Have you considered the trade-offs?"
### Pattern 5: "What's the dependency?"
When user says "depends on X" — forcing question: "Is X locked in? If not, that decision comes first."
The forcing element: surfaces dependency chains. Forces depth-first walk of the decision tree.
**Soft variant (bad):** "Have you thought about dependencies?"
### Pattern 6: "Even at 60% confidence — what's your best guess?"
When user hedges — forcing question: "Even uncertain, what would you decide today?"
The forcing element: prevents indefinite deferral. Most decisions can be made under uncertainty + revised later.
**Soft variant (bad):** "When will you decide?"
## The "Recommended Answer" Rule (per Matt)
Every question should carry a recommended answer with rationale. Why:
1. **Models the depth of analysis expected** — answerer sees what "good" looks like
2. **Accelerates the interview** — answerer can agree/disagree faster than constructing from scratch
3. **Surfaces interrogator bias** — if the recommendation is wrong, answerer can correct it explicitly
4. **Prevents "what do you think?" loops** — both sides commit to a position
Format:
> Q: [forcing question]
> Recommended: [position] because [1-sentence reason].
## One-at-a-Time Discipline (per Matt)
> "Ask the questions one at a time."
Why this matters:
1. **Bundled questions get partial answers** — answerer addresses the easiest one; hard ones get skipped
2. **Each answer constrains the next** — the second question often changes after hearing the first answer
3. **Cognitive load** — answerer can focus + give a complete response
4. **Visible progress** — each Q→A pair locks one decision; bundle masks progress
**Anti-pattern:** "Here are 8 questions: [list]". This is a survey, not an interrogation.
## Codebase Exploration > Speculation (per Matt)
> "If a question can be answered by exploring the codebase, explore the codebase instead."
When to explore instead of asking:
| Question | Action |
|---|---|
| "What auth library are we using?" | `grep -r "auth" package.json` — don't ask |
| "Does X already exist?" | `find . -name "X*"` — don't ask |
| "What's the current schema?" | `Read path/to/migrations/latest.sql` — don't ask |
| "Are tests passing?" | Run the test suite — don't ask |
When to ask anyway:
- Intent: "Why this approach?" can't be grepped
- Trade-offs: only the human knows which they value
- Future state: codebase shows current, not desired
## Anti-Patterns
1. **"Are you sure?"** — invites defensive answer; no information value
2. **"Have you thought about ...?"** — implies "no" is acceptable; doesn't force a decision
3. **"What if it fails?"** — speculative; better: "what's the kill criterion?"
4. **"Could you elaborate?"** — passive; better: name the specific gap
5. **Yes/no questions** without follow-up — wastes the turn
6. **Stacking questions** — bundles violate one-at-a-time rule
## How `question_generator.py` Implements This
The tool's question templates map each detected branch kind to a forcing-question pattern:
- `intent` → "Why this approach and not the obvious alternative?" (Pattern 1)
- `choice` → "Which side of the choice, and what's the deciding criterion?" (Pattern 4)
- `open` → "What's blocking this decision?" (Pattern 3)
- `tradeoff` → "Which side of the trade-off are you optimizing for?" (Pattern 4)
- `dependency` → "Is the dependency locked in?" (Pattern 5)
- `question` → "What's your current best answer, even if uncertain?" (Pattern 6)
Each generated question carries a recommended-answer template per Matt's rule.
## When This Reference Doesn't Help
- **Open-ended exploration** — early-stage ideation needs soft questions; grill-me is for plans not yet committed
- **Therapeutic/coaching contexts** — forcing questions can feel adversarial; tone matters
- **Hiring interviews** — different mode; behavioral questions follow different patterns
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — grill-me** (https://github.com/mattpocock/skills/, MIT) — the one-at-a-time + recommended-answer rules
- **Socratic Method** (5th-century BC) — Plato's dialogues — sequential questioning toward truth
- **Y Combinator office-hour format** (Garry Tan + Michael Seibel) — founder interrogation pattern
- **Toyota Production System — 5 Whys** (Sakichi Toyoda) — sequential causal questioning
- **Cockburn, A. — "Writing Effective Use Cases"** (2000) — decision-branch enumeration
- **Popper, K. — "Conjectures and Refutations"** (1963) — falsifiability + kill criteria
- **Galef, J. — "The Scout Mindset"** (2021) — calibrating beliefs under uncertainty
- **Larson, W. — "An Elegant Puzzle"** (2019) — eng decision-making in practice
FILE:references/when_to_stop_grilling.md
# When to Stop Grilling
This reference answers exactly one decision: **when is "shared understanding" actually reached, and how do we know to stop the interrogation?**
Pair with `scripts/grill_session_tracker.py` — the session tracker shows progress and surfaces unanswered branches.
## Matt Pocock's Stopping Condition (Implicit)
> "Interview me relentlessly about every aspect of this plan until we reach a shared understanding."
>
> — Matt Pocock, grill-me SKILL.md
"Shared understanding" is the stopping condition. But what does that mean operationally?
## Three Conditions That Mean "Stop"
### Condition 1: Every decision branch has an answer
Track via `grill_session_tracker.py status`. When `percent_complete = 100%`, every detected branch has a recorded answer. Stop grilling.
**Risk:** The extractor missed branches. Run `decision_tree_extractor.py` once more after answers are in — sometimes answers reveal new branches.
### Condition 2: No new questions arise from the last 3 answers
If the last 3 answers all triggered follow-up questions, grilling continues. If 3 answers in a row resolve cleanly with no new questions, the tree is exhausted.
**Pattern:** count the rate of new-question generation per turn. When it drops to zero for 3+ turns, stop.
### Condition 3: The interrogator can predict the answerer's response
If the interrogator can predict, with high confidence, what the answerer will say to the next question — that question doesn't add information. Skip it or stop entirely.
**Test:** before asking the next question, write down your guess at the answer. If the guess matches, you don't need to ask. Move on.
## Three Conditions That Mean "Keep Going"
### Condition A: The answerer is dodging
Signs:
- "We'll figure that out later" (without a date)
- "It depends" (without naming the dependency)
- Answers a different question than was asked
- Hedges every answer with "probably" / "likely" / "maybe"
Action: re-ask the same question with the same words. If dodged twice, name the dodge: "You said 'we'll figure it out later' — what's the latest moment you can decide and still ship?"
### Condition B: Answers contradict each other
If Q3 answer contradicts Q1 answer, stop the forward progress and reconcile:
> "You said X in Q1 but now Y in Q3. Which is it?"
Reconciliation is a separate grill phase — don't continue forward until resolved.
### Condition C: A new branch surfaces
If the answerer says "but if we do X, then we also need to decide Y" — Y is a new branch. Add to the question queue. Don't stop until Y is resolved.
## The "Recommended Answer Match" Heuristic
When generating questions with `question_generator.py`, each question has a recommended answer. Track:
| Answer matches recommendation? | What it means |
|---|---|
| Yes, with same rationale | Strong signal — both interrogator + answerer converged on the same logic |
| Yes, different rationale | Worth probing — same conclusion via different reasoning could mean one is wrong |
| No, with strong rationale | Healthy disagreement — record the rationale; this is the value of the grill |
| No, weak rationale | Push back — "the recommendation was X because Y; your answer rejects Y — why?" |
When 80%+ of answers match the recommendations cleanly, the grill is over-engineered for this plan — stop.
## The "Diminishing Returns" Test
Each grill question costs ~1 turn. After 10-15 questions on a single plan, returns diminish:
- First 3-5: high value (catches major missing decisions)
- Questions 6-10: medium value (refines edge cases)
- Questions 11-15: lower value (catches rare edge cases)
- Questions 16+: noise (usually the interrogator over-conditioning)
If a plan has 20+ branches, consider splitting into multiple plans rather than one mega-grill.
## When to Stop Even Before Conditions Met
### When the user signals fatigue
> "Can we move on?" / "Let's just decide and revisit if needed" / "Skip ahead"
Stop. Note unresolved branches in the session for later. Don't push through fatigue — answers under fatigue are often wrong.
### When the cost of deciding exceeds the cost of being wrong
For reversible decisions, grilling is overhead. Ship and revisit. For irreversible decisions, grill thoroughly.
Test: "If we're wrong about this, what does it cost to fix?" If the answer is "trivial" or "we just change a flag", stop grilling early.
### When the plan is exploratory
If the plan is "let's try X for a week and see" — don't grill the details. Grill the decision criteria for after the week.
## The Locking-In Pattern
When the grill ends, the session should produce a "decisions locked" summary:
```
Session: my-plan
Started: 2026-05-13
Closed: 2026-05-13
Status: Complete (8/8 branches resolved)
Decisions locked:
1. [L4] Schema-per-tenant chosen for cost reasons; isolation risk accepted.
2. [L8] Okta for SSO. Auth0 rejected (less Workday integration).
3. ...
```
The summary becomes the reference document. The grill session is throwaway; the summary is the artifact.
## Anti-Patterns
1. **Grilling forever** — every plan has 100 decideable details; grill stops at "shared understanding", not "complete certainty"
2. **Grilling reversible decisions** — wasteful; ship + revise
3. **Grilling without producing a summary** — wastes the answers; lock them in
4. **Grilling without exploring codebase first** — wastes turns asking questions the code answers
5. **Re-grilling the same plan** — if the plan was already grilled, don't re-grill the same branches; only grill new branches
## When This Reference Doesn't Help
- **Live-decision grilling in a meeting** — different mode; meetings have time pressure
- **Code review** — different scope; review is post-decision
- **Brainstorming** — wrong tool; grilling is for committed plans, not exploration
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — grill-me** (https://github.com/mattpocock/skills/, MIT) — the "shared understanding" stopping condition
- **Galef, J. — "The Scout Mindset"** (2021) — when to stop seeking more evidence
- **Kahneman, D. — "Thinking, Fast and Slow"** (2011) — decision fatigue + diminishing returns
- **Bezos, J. — Type 1 vs Type 2 decisions** (Amazon shareholder letter, 2015) — reversible vs irreversible decisions
- **YC Founder School — "Decide and move on"** — when grilling becomes procrastination
- **Larson, W. — "An Elegant Puzzle"** (2019) — engineering decision-making sequencing
- **Cynefin framework (Snowden)** — different decision domains require different evidence thresholds
FILE:scripts/decision_tree_extractor.py
#!/usr/bin/env python3
"""decision_tree_extractor.py — Extract decision branches from a plan/design doc.
Stdlib-only. Scans a markdown plan and identifies decision branches by detecting:
1. Modal verbs of intent: "we'll", "we will", "we plan to", "we should", "we could"
2. Open questions: sentences ending in "?"
3. Choices: "X or Y" / "either X or Y" / "vs"
4. TBDs: "TBD", "to be decided", "open question"
5. Trade-off markers: "trade-off", "tradeoff", "pros/cons"
Output: numbered list of decision branches with line refs.
NO LLM CALLS. Pure regex + line walking.
Usage:
python decision_tree_extractor.py # uses embedded sample
python decision_tree_extractor.py path/to/plan.md
python decision_tree_extractor.py plan.md --output json
"""
import argparse
import json
import re
import sys
from typing import Any, Dict, List
# Regex patterns that indicate a decision branch
DECISION_PATTERNS = [
(re.compile(r"\bwe\s*(?:'ll|will|plan\s+to|should|could|might|may)\b", re.IGNORECASE),
"intent"),
(re.compile(r"\b(?:either|or)\b.{0,80}\b(?:or|alternatively)\b", re.IGNORECASE),
"choice"),
(re.compile(r"\bversus\b|\bvs\.?\b", re.IGNORECASE),
"choice"),
(re.compile(r"\bTBD\b|\bto\s+be\s+(?:decided|determined)\b", re.IGNORECASE),
"open"),
(re.compile(r"\bopen\s+question\b", re.IGNORECASE),
"open"),
(re.compile(r"\btrade-?offs?\b", re.IGNORECASE),
"tradeoff"),
(re.compile(r"\bdepends?\s+on\b", re.IGNORECASE),
"dependency"),
(re.compile(r"\?\s*$"),
"question"),
]
SAMPLE_PLAN = """# Plan: Multi-tenant SaaS Migration
## Architecture
We'll move to a single-tenant database per customer. Or maybe we should
do schema-per-tenant for cost. This is a trade-off between isolation and ops cost.
## Auth
TBD: SSO provider — Okta or Auth0?
## Migration sequence
We plan to migrate the largest tenant first. Depends on whether their data fits in 24h.
Open question: rollback strategy?
## Data layer
We could use Postgres logical replication, but we might prefer dual-writes.
Trade-off: complexity vs zero-downtime guarantee.
## Cut-over
Final decision TBD on whether to flip DNS at midnight or use feature flags.
"""
def extract_branches(text: str) -> List[Dict[str, Any]]:
branches: List[Dict[str, Any]] = []
seen_lines = set()
for line_no, line in enumerate(text.splitlines(), start=1):
for pattern, kind in DECISION_PATTERNS:
match = pattern.search(line)
if not match:
continue
if line_no in seen_lines:
continue
seen_lines.add(line_no)
branches.append({
"line": line_no,
"kind": kind,
"trigger": match.group(0),
"context": line.strip()[:160],
})
break
return branches
def analyze(text: str) -> Dict[str, Any]:
branches = extract_branches(text)
by_kind: Dict[str, int] = {}
for b in branches:
by_kind[b["kind"]] = by_kind.get(b["kind"], 0) + 1
return {
"total_branches": len(branches),
"by_kind": by_kind,
"branches": branches,
}
def render_text(r: Dict[str, Any]) -> str:
lines = []
lines.append("=" * 72)
lines.append("DECISION TREE EXTRACTOR")
lines.append("=" * 72)
lines.append("")
lines.append(f"Total decision branches found: {r['total_branches']}")
lines.append(f"By kind: {r['by_kind']}")
lines.append("")
lines.append("-" * 72)
for i, b in enumerate(r["branches"], start=1):
lines.append(f" [{i:2d}] L{b['line']:>4d} ({b['kind']:11s}) {b['context']}")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Extract decision branches from a plan/design document.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to markdown plan (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
text = f.read()
except (IOError, OSError) as e:
print(f"error: {e}", file=sys.stderr)
return 1
else:
text = SAMPLE_PLAN
result = analyze(text)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_text(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/grill_session_tracker.py
#!/usr/bin/env python3
"""grill_session_tracker.py — Track grill-me session state across turns.
Stdlib-only. JSON-backed session storage for the relentless interrogation pattern.
Tracks: questions asked, answers received, recommendations, decisions locked,
remaining branches. Persistence enables resume across sessions.
Storage: ~/.grill_sessions/<session_name>.json
Actions:
- start <session_name>: initialize new session from plan doc
- record <session_name> --question-id N --answer "text": record an answer
- status <session_name>: show progress
- list: list all sessions
- close <session_name>: mark complete + summary
NO LLM CALLS. Stdlib only.
Usage:
python grill_session_tracker.py --action list
python grill_session_tracker.py --action start --session my-plan --plan path/to/plan.md
python grill_session_tracker.py --action record --session my-plan --question-id 1 --answer "we chose X"
python grill_session_tracker.py --action status --session my-plan
python grill_session_tracker.py --action close --session my-plan
"""
import argparse
import json
import os
import sys
from datetime import datetime
from typing import Any, Dict, List
# Import question generator
_HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, _HERE)
from question_generator import analyze as analyze_plan, SAMPLE_PLAN # noqa: E402
SESSIONS_DIR = os.path.expanduser("~/.grill_sessions")
def _ensure_dir() -> None:
os.makedirs(SESSIONS_DIR, exist_ok=True)
def _session_path(name: str) -> str:
return os.path.join(SESSIONS_DIR, f"{name}.json")
def _load(name: str) -> Dict[str, Any]:
path = _session_path(name)
if not os.path.isfile(path):
return {}
with open(path, "r", encoding="utf-8") as f:
return json.load(f)
def _save(name: str, data: Dict[str, Any]) -> None:
_ensure_dir()
with open(_session_path(name), "w", encoding="utf-8") as f:
json.dump(data, f, indent=2)
def start_session(name: str, plan_path: str) -> Dict[str, Any]:
if plan_path:
with open(plan_path, "r", encoding="utf-8") as f:
plan_text = f.read()
else:
plan_text = SAMPLE_PLAN
plan_path = "<embedded sample>"
plan_analysis = analyze_plan(plan_text)
session = {
"name": name,
"started_at": datetime.now().isoformat(timespec="seconds"),
"plan_source": plan_path,
"total_questions": plan_analysis["total_questions"],
"questions": plan_analysis["questions"],
"answers": {}, # question_n -> {"answer": str, "recorded_at": iso}
"status": "active",
}
_save(name, session)
return session
def record_answer(name: str, qid: int, answer: str) -> Dict[str, Any]:
session = _load(name)
if not session:
raise ValueError(f"Session not found: {name}")
session["answers"][str(qid)] = {
"answer": answer,
"recorded_at": datetime.now().isoformat(timespec="seconds"),
}
_save(name, session)
return session
def session_status(name: str) -> Dict[str, Any]:
session = _load(name)
if not session:
return {"error": f"Session not found: {name}"}
answered = len(session.get("answers", {}))
total = session.get("total_questions", 0)
pct = round(100.0 * answered / max(total, 1), 1)
next_q = None
for q in session.get("questions", []):
if str(q["n"]) not in session.get("answers", {}):
next_q = q
break
return {
"name": session["name"],
"status": session.get("status", "active"),
"answered": answered,
"total": total,
"percent_complete": pct,
"next_question": next_q,
"all_answers": session.get("answers", {}),
}
def list_sessions() -> List[str]:
_ensure_dir()
return sorted(
os.path.splitext(f)[0]
for f in os.listdir(SESSIONS_DIR)
if f.endswith(".json")
)
def close_session(name: str) -> Dict[str, Any]:
session = _load(name)
if not session:
raise ValueError(f"Session not found: {name}")
session["status"] = "closed"
session["closed_at"] = datetime.now().isoformat(timespec="seconds")
_save(name, session)
return session
def render_status(r: Dict[str, Any]) -> str:
if "error" in r:
return f"ERROR: {r['error']}"
lines = []
lines.append("=" * 72)
lines.append(f"GRILL SESSION: {r['name']}")
lines.append("=" * 72)
lines.append(f"Status: {r['status']} ({r['answered']} / {r['total']} answered, {r['percent_complete']}%)")
lines.append("")
if r["next_question"]:
q = r["next_question"]
lines.append(f"Next question (Q{q['n']}):")
lines.append(f" {q['question']}")
lines.append(f" Recommended: {q['recommended']}")
else:
lines.append("All questions answered. Run --action close to mark session complete.")
lines.append("")
if r["all_answers"]:
lines.append("Answered:")
for qid, ans in sorted(r["all_answers"].items(), key=lambda x: int(x[0])):
lines.append(f" Q{qid}: {ans['answer'][:100]}")
return "\n".join(lines)
def _build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Track grill-me session state across turns.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
action_choices = ("start", "record", "status", "list", "close")
parser.add_argument("--action", default="status", choices=action_choices, help="Session action")
parser.add_argument("--session", help="Session name")
parser.add_argument("--plan", default="", help="Path to plan markdown (start action)")
parser.add_argument("--question-id", type=int, help="Question number to record")
parser.add_argument("--answer", help="Answer text (record action)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
return parser
def _print_session_list(sessions: List[str], json_output: bool) -> None:
if json_output:
print(json.dumps({"sessions": sessions}, indent=2))
return
print("Sessions:")
items = sessions or ["(none)"]
for s in items:
print(f" - {s}")
def _print_start_summary(session: Dict[str, Any]) -> None:
print(f"Started session: {session['name']}")
print(f" Plan: {session['plan_source']}")
print(f" Total questions: {session['total_questions']}")
questions = session.get("questions") or []
first = questions[0]["question"] if questions else "(none)"
print(f" First question: {first}")
def _action_list(args: argparse.Namespace) -> int:
_print_session_list(list_sessions(), args.output == "json")
return 0
def _action_start(args: argparse.Namespace) -> int:
name = args.session or "sample-session"
session = start_session(name, args.plan)
if args.output == "json":
print(json.dumps(session, indent=2))
else:
_print_start_summary(session)
return 0
def _action_record(args: argparse.Namespace) -> int:
if not args.session or args.question_id is None or not args.answer:
print("error: record requires --session, --question-id, --answer", file=sys.stderr)
return 1
record_answer(args.session, args.question_id, args.answer)
result = session_status(args.session)
output = json.dumps(result, indent=2) if args.output == "json" else render_status(result)
print(output)
return 0
def _action_status(args: argparse.Namespace) -> int:
name = args.session or "sample-session"
result = session_status(name)
output = json.dumps(result, indent=2) if args.output == "json" else render_status(result)
print(output)
return 0
def _action_close(args: argparse.Namespace) -> int:
if not args.session:
print("error: close requires --session", file=sys.stderr)
return 1
session = close_session(args.session)
if args.output == "json":
print(json.dumps(session, indent=2))
else:
print(f"Closed session: {args.session}")
return 0
ACTION_DISPATCH = {
"list": _action_list,
"start": _action_start,
"record": _action_record,
"status": _action_status,
"close": _action_close,
}
def main() -> int:
args = _build_parser().parse_args()
handler = ACTION_DISPATCH.get(args.action)
if handler is None:
return 0
return handler(args)
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/question_generator.py
#!/usr/bin/env python3
"""question_generator.py — Generate forcing questions from extracted decision branches.
Stdlib-only. Takes a plan doc, runs decision_tree_extractor, then generates
forcing questions per Matt Pocock's grill-me discipline:
- Each question maps to one decision branch
- Each question proposes a recommended answer
- Questions ordered by dependency (independent first, dependent last)
- One question per turn (output is a list, not a paragraph)
Template per question:
Q: [forcing question]
Recommended: [recommendation with 1-sentence rationale]
Question templates by branch kind:
- intent -> "You said you'll X. Why X and not Y?"
- choice -> "Between X and Y, which one and why?"
- open -> "X is marked TBD. What's blocking the decision?"
- tradeoff -> "Trade-off between A and B. Which side are you optimizing for?"
- dependency -> "X depends on Y. Is Y locked in? If not, ask about Y first."
- question -> "[original question] — what's your current answer?"
Usage:
python question_generator.py # uses embedded sample
python question_generator.py path/to/plan.md
python question_generator.py plan.md --output json
"""
import argparse
import json
import sys
import os
from typing import Any, Dict, List
# Import extractor as a module
_HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, _HERE)
from decision_tree_extractor import extract_branches, SAMPLE_PLAN # noqa: E402
QUESTION_TEMPLATES = {
"intent": "Why this approach and not the obvious alternative?",
"choice": "Which side of the choice, and what's the deciding criterion?",
"open": "What's blocking this decision? What would unblock it today?",
"tradeoff": "Which side of the trade-off are you optimizing for, and what's the kill criterion?",
"dependency": "Is the dependency locked in? If not, that decision comes first.",
"question": "What's your current best answer, even if uncertain?",
}
RECOMMENDED_TEMPLATES = {
"intent": "State the alternative explicitly + 1 sentence why you rejected it.",
"choice": "Pick the option that aligns with the constraint you can't change (budget, deadline, team).",
"open": "Name the missing input. Estimate when it arrives. Decide now under uncertainty if it won't arrive in time.",
"tradeoff": "Choose the side that's reversible later. Trade-offs are usually one-way; pick the one with the escape hatch.",
"dependency": "Resolve the upstream decision first. Then re-evaluate this one.",
"question": "Even a 60%-confidence answer is better than 'we'll figure it out later'.",
}
def _detect_dependencies(branches: List[Dict[str, Any]]) -> List[int]:
"""Reorder: dependency branches go AFTER what they depend on (best-effort)."""
dep_indices = [i for i, b in enumerate(branches) if b["kind"] == "dependency"]
non_dep_indices = [i for i, b in enumerate(branches) if b["kind"] != "dependency"]
return non_dep_indices + dep_indices
def generate_questions(branches: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
ordered = _detect_dependencies(branches)
questions: List[Dict[str, Any]] = []
for n, idx in enumerate(ordered, start=1):
b = branches[idx]
q_template = QUESTION_TEMPLATES.get(b["kind"], "What's the current state?")
r_template = RECOMMENDED_TEMPLATES.get(b["kind"], "State your best answer.")
questions.append({
"n": n,
"line": b["line"],
"branch_kind": b["kind"],
"context": b["context"],
"question": f"L{b['line']}: {b['context']} -> {q_template}",
"recommended": r_template,
})
return questions
def analyze(text: str) -> Dict[str, Any]:
branches = extract_branches(text)
questions = generate_questions(branches)
return {
"total_questions": len(questions),
"branch_kinds": sorted(set(b["kind"] for b in branches)),
"questions": questions,
}
def render_text(r: Dict[str, Any]) -> str:
lines = []
lines.append("=" * 72)
lines.append("FORCING QUESTION GENERATOR (one at a time, per Matt's grill-me)")
lines.append("=" * 72)
lines.append("")
lines.append(f"Total questions: {r['total_questions']}")
lines.append(f"Branch kinds: {r['branch_kinds']}")
lines.append("")
lines.append("-" * 72)
for q in r["questions"]:
lines.append(f" Q{q['n']:>2d}: {q['question']}")
lines.append(f" Recommended: {q['recommended']}")
lines.append("")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Generate forcing questions from a plan/design document.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to markdown plan (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
text = f.read()
except (IOError, OSError) as e:
print(f"error: {e}", file=sys.stderr)
return 1
else:
text = SAMPLE_PLAN
result = analyze(text)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_text(result))
return 0
if __name__ == "__main__":
sys.exit(main())
Giúp tiếp thu kiến thức mới nhanh, tóm tắt tài liệu dài, ghi nhớ theo kỹ thuật Feynman và tạo bộ câu hỏi ôn tập.
--- name: hoc-tap-nghien-cuu description: Giúp tiếp thu kiến thức mới nhanh hơn, tóm tắt tài liệu dài, ghi nhớ theo kỹ thuật Feynman và tạo bộ câu hỏi ôn tập thực tế. Dùng khi nói "học tập", "tóm tắt tài liệu", "hiểu sâu chủ đề". --- # Học tập & nghiên cứu (Feynman Learning) ## Mục tiêu Giúp tiếp thu kiến thức mới nhanh hơn, ghi nhớ lâu hơn và áp dụng được vào thực tế. ## Khi nào dùng - Cần tóm tắt tài liệu dài - Muốn hiểu sâu một chủ đề mới - Cần tạo flashcard hoặc câu hỏi ôn tập - Muốn kết nối kiến thức mới với thứ đã biết ## Đầu vào cần cung cấp - Tài liệu hoặc chủ đề cần học - Mục tiêu học (hiểu tổng quan / hiểu sâu / áp dụng ngay) - Thời gian có thể dành ra - Kiến thức nền hiện tại ## Quy trình xử lý 1. Xác định khung kiến thức tổng quan (big picture) 2. Chia thành các module nhỏ có thể học trong 25 phút 3. Tóm tắt theo kỹ thuật Feynman: giải thích như cho người không biết nghe 4. Tạo 5–10 câu hỏi kiểm tra mức độ hiểu 5. Kết nối với ví dụ thực tế hoặc kiến thức đã có ## Tiêu chuẩn đầu ra - Tóm tắt ngắn gọn, không quá 500 từ - Có phần "ý chính cần nhớ" (bullet points) - Có ví dụ minh họa thực tế - Có câu hỏi tự kiểm tra ## Tránh - Tóm tắt quá dài dẫn đến không đọc được - Dùng thuật ngữ khó mà không giải thích - Bỏ qua phần ứng dụng thực tế
Phỏng vấn người dùng về thói quen email, bối cảnh, phong cách trả lời và ưu tiên để xây cơ sở tri thức phân loại hộp thư cá nhân hóa.
--- name: inbox-setup description: "One-time setup skill that builds a personalized inbox triage knowledge base via interactive interview. Interviews the user about their email patterns, business context, reply style, and priorities using grill-me discipline (one question at a time, forcing format where possible, dependency-ordered, each question explains why I'm asking), then generates the knowledge base files that power the companion 'inbox-triage' skill. Run this once before using inbox-triage for the first time. Re-run when business, pricing, or priorities change significantly. Triggers: 'set up my inbox', 'configure inbox triage', 'set up my email system', 'configure email triage', 'build my email knowledge base', 'initialize email management', 'set up inbox triage', 'onboard email triage', or any variation where someone wants to get the email triage system running for the first time." license: MIT metadata: source_spec: "megaprompts/06-inbox-setup-megaprompt.md" build_pattern: "Path B (direct conversion)" paired_with: "inbox-triage (shared 7-file KB contract)" version: 1.0.0 --- # Inbox-Setup — Email Triage Onboarding > **Paired with `inbox-triage`.** This skill writes the 7-file knowledge base at `WORKSPACE/Email/` that `inbox-triage` reads on every run. The file contracts (names, sections, fields) MUST match between the two skills exactly. See [`references/kb_file_contract.md`](references/kb_file_contract.md). Run once (or re-run when business/priorities change). Interview the user about their email patterns, business context, reply style, and priorities. Generate the structured knowledge base in `WORKSPACE/Email/` that captures everything `inbox-triage` needs to process the inbox effectively. ## Invocation Triggers - "set up my inbox" - "configure inbox triage" - "set up my email system" - "configure email triage" - "build my email knowledge base" - "initialize email management" - "set up inbox triage" - "onboard email triage" ## Conduct Discipline **Do NOT generate all files at once.** Walk through the 8 sections one at a time. Each section commits its file(s) before moving on. Partial completion (e.g., user drops off mid-interview) still produces a usable partial KB. Grill-me discipline applies throughout: - **One question per turn.** Never bundle. Even across section boundaries. - **"Why I'm asking" on every question** — so users can answer well. - **Forcing format where possible.** Multi-choice > open-ended. - **Dependency-ordered.** Q2 depends on Q1; downstream sections depend on upstream. See [`references/grill_me_section_walk.md`](references/grill_me_section_walk.md) for the 8-section discipline detail. ## Knowledge Base Contract — Files To Produce Exactly these files at `WORKSPACE/Email/`: | File | Purpose | Required? | |---|---|---| | `email-taxonomy.md` | Classification system + report preferences | **Yes** | | `email-patterns.md` | Reply voice, tone, templates, hard rules | **Yes** | | `evaluation-framework.md` | Decision tree for opportunity emails | Only if user receives pitches/opportunities | | `rate-card.md` | Pricing, terms, negotiation posture | Only if user has pricing | | `blocklist.md` | Auto-skip senders + learned decline patterns | **Yes** (seeded, grows over time) | | `tracker.md` | Active follow-ups, overdue items, deadlines | **Yes** (starts mostly empty) | | `triage-log/` | Directory for per-run logs | **Yes** (created empty) | The contract is identical to what `inbox-triage` expects — see [`references/kb_file_contract.md`](references/kb_file_contract.md) for the full spec. ## Stop Condition (Full Interview) ~25–31 questions total across the 8 sections (depending on skip-logic). Hard ceiling: 35 questions including all sub-clarifications. Section 4 (Evaluation Framework) is skipped entirely when Section 1 surfaced no opportunity-email category, dropping the total by 6 questions and the rate-card file. After Section 8's confirmation + handoff message, intake is closed — **never re-open it**. To change preferences later, the user re-runs the skill (which detects existing files and asks per-file: replace / merge / skip). The grill-me one-at-a-time rule applies across section boundaries: do NOT batch questions even when moving from S{n} to S{n+1}. ## Section 1: The Big Picture Six grill-me questions, one at a time: - **S1.Q1:** "What do you do? Give me your role and business in 1–2 sentences. *Why I'm asking:* Context shapes what email patterns to expect — a solo creator's inbox looks nothing like an enterprise PM's." - **S1.Q2:** "What dominates your inbox? Pick the top 1–2: sales pitches / client work / internal team / newsletters / customer support / financial / other. *Why I'm asking:* Dominant categories drive the taxonomy." - **S1.Q3:** "Rough volume split — e.g., '60% business inquiries, 20% ops, 20% noise'. *Why I'm asking:* The split tells me where to focus triage effort." - **S1.Q4:** "Which email address(es) should triage cover? *Why I'm asking:* If multiple, I'll set up per-address taxonomies." - **S1.Q5:** "Run frequency: once daily / 2x daily / 3x daily / on-demand only? *Why I'm asking:* Drives the default search window in triage (9h overlap for 2x/day)." - **S1.Q6:** "Anyone helping manage email — assistant, VA, team — or solo? *Why I'm asking:* Persona handling differs for delegated inboxes." **Action:** Build mental model. Do NOT write files yet. Note whether opportunity emails are a category (drives S4 skip-logic). ## Section 2: Email Categories Propose 5–7 categories based on Section 1 — pre-recommend a subset, not the whole template menu: - New Opportunities - Active Conversations - Action Required - Financial - Important/Personal - Informational - Ignore/Low Priority Then three forcing questions, one at a time: - **S2.Q1:** "Here's my proposed taxonomy: [list]. Does this match your inbox reality — yes / mostly / no? *Why I'm asking:* If 'no', I need to redo the taxonomy before any other section makes sense." - **S2.Q2:** "Missing categories? List them. (Skip if none.) *Why I'm asking:* Missing categories produce uncategorized emails downstream, which hurts triage quality." - **S2.Q3:** "Which category takes the MOST time per email? *Why I'm asking:* That's where draft-reply effort needs to focus most." **Action:** Generate `email-taxonomy.md` with categories, signals (for each: trigger phrases / sender patterns / subject markers), and default actions per category. ## Section 3: Reply Style & Voice Six grill-me questions plus the critical sample request: - **S3.Q1:** "Register: formal / casual / in-between? *Why I'm asking:* Calibrates default voice; we'll refine from samples next." - **S3.Q2:** "Three communication pet peeves — phrases you hate, openings you avoid. *Why I'm asking:* I treat these as forbidden tokens in drafts." - **S3.Q3:** "Phrases or sign-offs you always use — list as many as come to mind. *Why I'm asking:* These are your voice fingerprints." - **S3.Q4:** "Different persona for different contexts — e.g., assistant replies as you? *Why I'm asking:* Persona context changes pronoun + signature handling." - **S3.Q5:** "Typical reply length — one-liner / short paragraph / longer? *Why I'm asking:* Length is the easiest voice signal to get wrong." - **S3.Q6:** "Hard rules — never X / always Y? (E.g., never emojis, always reply within 24h, never take calls without context.) *Why I'm asking:* Hard rules are enforced as non-negotiable in every draft." ### S3.SAMPLES (the critical highest-quality input) > **Paste 3–5 real sent emails from your inbox.** > > *Why I'm asking:* Self-description of voice is unreliable. Real samples are the best signal — I'll analyze them for voice patterns that supplement everything above. Use `scripts/voice_sample_analyzer.py` to extract patterns deterministically. If user runs a business: also ask about media kits, rate sheets, standard pitches, repeated replies. **Action:** Generate `email-patterns.md` with tone description (with do/don't examples), persona rules, templates, signatures, hard rules. See [`references/voice_calibration.md`](references/voice_calibration.md) for the sample-extraction discipline. ## Section 4: Evaluation Framework (Conditional) **Skip-logic:** only run this section if Section 1 surfaced opportunity emails as a meaningful inbox category. Otherwise jump straight to Section 5. Six grill-me questions, one at a time: - **S4.Q1:** "First thing you check when pitched something — give me your gut filter. *Why I'm asking:* That's the top of the decision tree." - **S4.Q2:** "Three instant deal-breakers — things that make you decline immediately. *Why I'm asking:* These become PASS-auto signals." - **S4.Q3:** "Three things that make you immediately interested. *Why I'm asking:* These become TAKE-IT signals." - **S4.Q4:** "Standard pricing / terms — or 'no fixed pricing' if you negotiate every time. *Why I'm asking:* If you have a rate card, I'll generate one; if not, I'll skip." - **S4.Q5:** "Negotiation posture: firm / flexible / depends on context? *Why I'm asking:* Drives draft tone on counter-offers." - **S4.Q6:** "VIP senders or organizations that always get engagement — list names or domains. *Why I'm asking:* VIP list bypasses normal PASS filters." **Action:** Generate `evaluation-framework.md` (decision tree + recommendation categories + VIP list) AND `rate-card.md` if pricing exists. ## Section 5: Blocklist & Patterns Three grill-me questions, one at a time: - **S5.Q1:** "Senders or domains to always skip — list them. (Skip if none.) *Why I'm asking:* Auto-blocklist saves the most time per run." - **S5.Q2:** "Patterns in emails you always delete — e.g., 'unsubscribe' links from specific marketers, recruiter cold outreach, newsletters? *Why I'm asking:* Patterns let triage auto-skip variants without exact-match maintenance." - **S5.Q3:** "Specific companies / recruiters / newsletters wasting time — list any. *Why I'm asking:* These seed the blocklist; triage will add more as you override decisions." **Action:** Generate `blocklist.md` (auto-maintained by triage thereafter). ## Section 6: Current State Three grill-me questions, one at a time: - **S6.Q1:** "Active threads you're tracking — list with one-line context each. (Skip if none.) *Why I'm asking:* These become tracker entries so triage knows existing context." - **S6.Q2:** "Overdue replies — anything you should have responded to but haven't? *Why I'm asking:* Triage flags these as priority every run until resolved." - **S6.Q3:** "Time-sensitive items with deadlines — list with dates. *Why I'm asking:* Tracker enforces deadlines and surfaces them as overdue at the right time." **Action:** Generate `tracker.md` with active follow-ups table, overdue section, resolved section (empty), update log (empty). Also create empty `triage-log/` directory. ## Section 7: Report Preferences Three grill-me questions, one at a time: - **S7.Q1:** "Delivery format — pick one: email draft to self / file in workspace / chat summary only. *Why I'm asking:* The triage report goes here every run." - **S7.Q2:** "Detail level — pick one: 30-second scan / detailed breakdown / both (scan first, expand on request). *Why I'm asking:* Affects report length." - **S7.Q3:** "Anything always shown first — e.g., overdue payments, VIP messages? *Why I'm asking:* Custom 'top-of-report' rules surface what you care about above standard sections." **Action:** Save these preferences into `email-taxonomy.md` under a "Report Preferences" section. ## Section 8: Confirmation & Handoff List every file created with one-sentence summary. Then: > Your triage system is ready. Run the **inbox-triage** skill to process your inbox. First runs need oversight — system learns from your edits and overrides. Remind: re-run this setup anytime business/pricing/priorities change. Run `scripts/kb_validator.py --workspace WORKSPACE` to confirm the 7-file contract is satisfied before final handoff. ## Privacy Boundary **Never persist passwords, full account numbers, SSNs, or other sensitive credentials in knowledge base files.** If the user volunteers such info during the interview, acknowledge it but don't store it; the relevant KB file gets `[stored separately by user]` in its place. ## Re-Run Behavior Re-running on an existing setup: 1. Detect `WORKSPACE/Email/` 2. For each existing file, ask per-file: **replace / merge / skip** 3. Walk only the sections whose files the user chose to update 4. Skip sections whose files the user kept ## Error Handling | Situation | Behavior | |---|---| | Workspace inaccessible | Stop. Tell user where files would go and ask for permission/path | | User refuses to share samples | Use self-description; flag in patterns file that calibration may need iteration | | User says "skip this" mid-interview | Honor it; flag the gap in the file as `[needs follow-up]` | | Sensitive info volunteered | Acknowledge but don't persist; note in file as `[stored separately by user]` | | Re-run on existing setup | Detect existing files; ask user per-file: replace, merge, skip | | User has no pricing / opportunities | Skip Section 4 entirely; don't create empty files | ## Portability - **Claude Code CLI:** Native — writes markdown files directly to filesystem. - **Claude.ai web:** Works with project files / artifacts. Document the alternate path: generate files as artifacts, instruct user to save to their workspace, or use connected file system if available. ## Tooling | Script | Role | |---|---| | `scripts/kb_validator.py` | Validates the 7-file KB output (required files present, conditional files only if their sections ran, headers + structure correct). | | `scripts/section_progress_tracker.py` | JSON-backed walk state at `~/.inbox_setup_sessions/<session>.json`. Tracks active section, answered questions, committed files. | | `scripts/voice_sample_analyzer.py` | Extracts voice patterns from pasted sent-email samples — opening phrases, sign-offs, length distribution, register markers. | ## References - [`references/kb_file_contract.md`](references/kb_file_contract.md) — the canonical 7-file contract (write perspective; mirror lives in `inbox-triage/references/`) - [`references/grill_me_section_walk.md`](references/grill_me_section_walk.md) — 8-section discipline, skip-logic, commit-per-section - [`references/voice_calibration.md`](references/voice_calibration.md) — sample-based voice extraction theory + anti-patterns ## Anti-Patterns To Reject - Generating all files at once instead of walking through sections - Asking all questions in one batch - Hardcoded provider references (Gmail-only thinking) - Persisting sensitive credentials in knowledge base - Skipping the "why this question matters" explanation - Skipping the sample-emails ask for voice (it's the highest-quality input) - Overwriting existing files without consent on re-run - Forcing creation of `rate-card.md` or `evaluation-framework.md` when they don't apply --- **Version:** 1.0.0 **Source spec:** [`megaprompts/06-inbox-setup-megaprompt.md`](../../../../megaprompts/06-inbox-setup-megaprompt.md) **Build pattern:** Path B (direct conversion). Paired with `inbox-triage`. FILE:references/grill_me_section_walk.md # Grill-Me Section Walk Discipline This reference answers exactly one decision: **how does inbox-setup walk 8 sections of ~25-31 questions without violating grill-me discipline, and what makes the discipline survive heavy intake?** ## The Core Tension Capture's grill-me is **max-1 question** per dump (light intake). Inbox-setup is **25-31 questions across 8 sections** (heavy intake). At that scale, the one-question-at-a-time rule is easy to break — the interviewer is tempted to batch, the user is tempted to dump everything at once. The discipline survives because: 1. **Section boundaries** create natural commit points 2. **Skip-logic** removes ~6 questions when irrelevant (Section 4) 3. **Per-section file writes** make partial completion still useful 4. **Forcing format** keeps questions answerable in seconds ## The Four Rules ### Rule 1: One Question Per Turn — Across Section Boundaries The rule does NOT relax when moving between sections. After S2.Q3 commits `email-taxonomy.md`, ask S3.Q1 alone — not "S3.Q1 and S3.Q2 since you already know your voice." **Why:** the user is fatigued by question 18; bundling 3 at once produces shallower answers. Better to be slow than to lose answer quality on the high-leverage voice + framework questions. ### Rule 2: "Why I'm Asking" On Every Single Question Without the rationale, users either: - Skip past the question thinking it's optional - Answer minimally because they don't know what's at stake - Misunderstand the depth needed The rationale is short (1-2 sentences) and concrete ("This becomes a forbidden token in drafts" beats "this helps me understand your style"). ### Rule 3: Forcing Format > Open-Ended | ✅ Forcing | ❌ Open-ended | |---|---| | "Run frequency: once daily / 2x daily / 3x daily / on-demand only?" | "How often should I run?" | | "Does this taxonomy match: yes / mostly / no?" | "What do you think of this taxonomy?" | | "Register: formal / casual / in-between?" | "Describe your tone." | Open-ended works for: pet peeves (S3.Q2), sign-offs (S3.Q3), hard rules (S3.Q6), VIP list (S4.Q6), tracker entries (S6.Q1) — where the answer space is genuinely unbounded and forcing format would harm signal. ### Rule 4: Commit Per Section, Not End-Of-Interview After Section 2's 3 questions: write `email-taxonomy.md`. Do NOT wait until Section 8 to write all files at once. **Why:** if the user drops off after Section 4 (~16 questions in), the user has a useful partial KB (taxonomy + patterns + framework + rate card). If files were batched at the end, drop-off leaves nothing. ## The 8 Sections at a Glance | Section | Questions | Skip-Logic | Files Written at End | |---|---:|---|---| | 1. The Big Picture | 6 | always run | (none — build mental model) | | 2. Email Categories | 3 | always run | `email-taxonomy.md` | | 3. Reply Style & Voice | 6 + samples | always run | `email-patterns.md` | | 4. Evaluation Framework | 6 | skipped if no opportunity category in S1 | `evaluation-framework.md` + `rate-card.md` (cond) | | 5. Blocklist & Patterns | 3 | always run | `blocklist.md` | | 6. Current State | 3 | always run | `tracker.md` + `triage-log/` dir | | 7. Report Preferences | 3 | always run | appended to `email-taxonomy.md` | | 8. Confirmation & Handoff | 0 (summary) | always run | (no file write; handoff message) | **Total: 24 + 6 conditional = 30 max** (or 24 if S4 skipped). Hard ceiling 35 includes sub-clarifications. ## Skip-Logic Detail ### Section 4 Skip After S1.Q2 ("what dominates your inbox?"), if the answer does NOT include: - "sales pitches" / "opportunities" / "client work proposals" Then mark S4 as skipped. State to user: > Skipping Section 4 (Evaluation Framework) since your inbox doesn't include pitches/opportunities. Moving to Section 5. The user CAN override: "Actually I do get opportunity emails — run that section." Honor the override. ### Per-Question Conditional Skips Some individual questions have "(Skip if none)" suffix: - S2.Q2 (missing categories?) — skip if user says all listed - S5.Q1 (skip-senders?) — skip if user has none yet - S6.Q1 (active threads?) — skip if user has none - S6.Q2 (overdue?) — skip if user has none - S6.Q3 (deadlines?) — skip if user has none These skips ALSO commit to the file (with empty section) so triage knows the section was considered, not forgotten. ## Per-Section File Commit Pattern ``` 1. Ask all questions in Section N (one at a time) 2. Synthesize answers into structured file content 3. Write file(s) at WORKSPACE/Email/{filename} 4. Confirm to user: "✓ Section N complete. {file(s)} committed." 5. Record in session tracker: python scripts/section_progress_tracker.py \ --action record_section_done --session NAME \ --section N --files "{filename}" 6. Move to Section N+1's first question. ``` ## Re-Run Mode Detect re-run when `WORKSPACE/Email/email-taxonomy.md` exists. Walk the user through per-file consent: ``` Found email-taxonomy.md from 2026-03-04 (45 days ago). Replace / merge / skip? - replace: rewrite from new interview answers - merge: keep existing categories, add new ones from this run - skip: leave file as-is; move to next file ``` Walk only the sections whose files the user chose to replace or merge. If user chose skip for a file, do NOT re-ask that section's questions. ## Sample-Collection Discipline (S3.SAMPLES) The sample-emails ask is **the highest-quality voice signal** the skill has. It is NOT optional from a quality standpoint, but it IS skippable by user choice. **Discipline:** 1. Ask for 3-5 real sent emails. Frame it as "the best signal I have." 2. If user pastes them: run `scripts/voice_sample_analyzer.py` and incorporate the output into `email-patterns.md` under "Voice Patterns (Extracted from Samples)." 3. If user refuses: use S3.Q1-Q6 self-description only. Flag in `email-patterns.md`: > `[calibration may need iteration — voice samples not collected during setup. First few triage runs will likely produce drafts that need editing; the system learns from your edits.]` 4. Never proceed past Section 3 without either samples OR explicit user-skip + flag. ## Anti-Patterns To Reject - Asking S1.Q1-Q3 in one message ("tell me your role, what dominates your inbox, and rough volume split") - Asking S2.Q1 without "Why I'm asking" - Writing all 7 files at end of S8 (no per-section commit) - Asking S4 questions when no opportunities surfaced in S1 - Asking S5.Q1 again when user already said "I have no blocklist yet" in S1 - Forcing closed-format on genuinely open questions (e.g., "Pet peeves: a) clichés b) emojis c) other" — kills signal) - Skipping the rationale ("Why I'm asking") to "save time" - Skipping the sample ask in S3 - Re-running and overwriting existing files without per-file consent ## Citations The grill-me discipline this reference enforces is canonical in this repo. See: - [`engineering/grill-me/`](../../../../engineering/grill-me/) — the source skill that formalized the discipline - Matt Pocock's original grill-me skill (MIT) - This repo's PR #657 cross-skill consistency audit, which verified the discipline transfers consistently across all intake-having skills (1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13) FILE:references/kb_file_contract.md # Knowledge Base File Contract (Write Perspective) This reference answers exactly one decision: **what 7 files must `inbox-setup` produce, in what structure, so that `inbox-triage` can read them without ambiguity?** This is the integration boundary between the paired skills. Any drift breaks the pair. PR #657's cross-skill consistency audit verified that the 7 KB filenames align verbatim between the two megaprompts; this reference is the canonical write-side spec. A mirror lives at `inbox-triage/references/kb_file_contract.md` (read perspective). ## The 7 Files at `WORKSPACE/Email/` | File | Required? | Triggered by | Triage uses for | |---|---|---|---| | `email-taxonomy.md` | yes | Section 2 + Section 7 | classification + report preferences | | `email-patterns.md` | yes | Section 3 | reply voice + templates + hard rules | | `evaluation-framework.md` | conditional | Section 4 (only if S1 surfaced opportunities) | TAKE-IT / WORTH / PASS / FLAG decisions | | `rate-card.md` | conditional | Section 4 (only if user has pricing) | negotiation posture + counter-offers | | `blocklist.md` | yes (seeded) | Section 5 | auto-skip senders + decline patterns | | `tracker.md` | yes (seeded) | Section 6 | active follow-ups + deadlines | | `triage-log/` | yes (empty dir) | Section 6 | per-run logs (populated by triage) | ## File Specs (Write Side) ### email-taxonomy.md (required) ```markdown # Email Taxonomy ## Categories ### {Category Name} - Signals: {trigger phrases, sender patterns, subject markers} - Default action: {classify / draft-reply / skip / flag-for-review} - Typical volume: {N% of inbox} ### {Category 2} ... ## Report Preferences - Delivery format: {email-draft-to-self | file-in-workspace | chat-summary-only} - Detail level: {30-second-scan | detailed-breakdown | both} - Always-shown-first: {overdue payments | VIP messages | custom rules} ``` **Generated at:** end of Section 2 (categories) + appended at end of Section 7 (Report Preferences). ### email-patterns.md (required) ```markdown # Email Patterns ## Voice Register {formal | casual | in-between} ## Pet Peeves (Forbidden Tokens) - {phrase 1} - {phrase 2} - {phrase 3} ## Sign-Offs (Voice Fingerprints) - {sign-off 1} - {sign-off 2} - ... ## Persona Context {single-user | delegated (assistant replies as user) | multi-persona} ## Typical Reply Length {one-liner | short-paragraph | longer} ## Hard Rules (Non-Negotiable in Every Draft) - Never: {X} - Always: {Y} ## Voice Patterns (Extracted from Samples) - Opening phrases observed: {list} - Sentence length distribution: {short / medium / long mix} - Casual / formal markers: {list} ## Templates (Repeated Replies) - {template 1 name}: {body} - {template 2 name}: {body} ``` **Generated at:** end of Section 3. The "Voice Patterns" subsection comes from `scripts/voice_sample_analyzer.py` if samples were provided; otherwise marked `[calibration may need iteration]`. ### evaluation-framework.md (conditional) ```markdown # Evaluation Framework (Opportunity Emails) ## Gut Filter (First Check) {user's gut filter from S4.Q1} ## TAKE-IT Signals - {signal 1} - {signal 2} - {signal 3} ## PASS Signals (Instant Deal-Breakers) - {deal-breaker 1} - {deal-breaker 2} - {deal-breaker 3} ## Decision Tree 1. If sender in VIP list → TAKE IT (skip filter) 2. If any PASS signal matches → PASS (auto-decline draft) 3. If all TAKE-IT signals match → TAKE IT (auto-engage draft) 4. If partial TAKE-IT match → WORTH CONSIDERING 5. If unusual / ambiguous → FLAG FOR REVIEW ## VIP List (Bypass PASS Filters) - {sender / domain 1} - {sender / domain 2} - ... ## Negotiation Posture {firm | flexible | depends-on-context} ``` **Generated at:** end of Section 4. Skipped entirely if S1 surfaced no opportunity-email category. ### rate-card.md (conditional) ```markdown # Rate Card ## Standard Pricing - {service / offering 1}: {price} - {service / offering 2}: {price} ## Terms - Payment: {net X days | upfront | milestone} - Revisions included: {N} - Rush fee: {Y%} ## Negotiation Posture {firm | flexible | depends-on-context} ## Counter-Offer Patterns - If they offer < {floor}: {how to counter} - If timeline is tight: {how to counter} ``` **Generated at:** end of Section 4. Skipped if user has no fixed pricing (S4.Q4 = "no fixed pricing"). ### blocklist.md (required, seeded) ```markdown # Blocklist ## Sender / Domain Auto-Skip - {sender 1}: {reason} — added {date} - {domain 1}: {reason} — added {date} ## Decline Patterns (Pattern-Match Auto-Skip) - "{pattern phrase 1}": {reason} - "{pattern phrase 2}": {reason} ## Recently Removed (User Overrode) - {sender}: removed on {date} — user override ``` **Generated at:** end of Section 5 (initial seed). `inbox-triage` appends new declines + observed patterns on every run. ### tracker.md (required, seeded) ```markdown # Tracker ## Active Follow-Ups | Item | Context | Deadline | Status | |---|---|---|---| | {thread} | {one-line context} | {date} | pending | | ... | ... | ... | ... | ## Overdue - {thread}: missed deadline {date} — {context} ## Resolved (Recent) ## Update Log - {date}: {what changed} — by {triage run | user} ``` **Generated at:** end of Section 6 (initial seed from S6.Q1-Q3). `inbox-triage` updates on every run. ### triage-log/ (required, empty directory) Empty directory created at end of Section 6. `inbox-triage` writes per-run logs to `triage-log/<YYYY-MM-DD>-<run-label>.md`. ## Validation Run `scripts/kb_validator.py --workspace WORKSPACE` after Section 8 confirmation. It checks: - All required files exist - Conditional files exist iff their triggering section ran - Each file has the expected H1 + section structure - `triage-log/` is a directory (not a file) ## Why This Contract Matters `inbox-triage` halts with a clear error if any required core file is missing. The contract is the integration boundary — both skills can be developed and tested independently, but they must agree on the file shape. When updating either skill: update both sides of the contract simultaneously, or use `/cs:grill-with-docs` to detect drift between the two megaprompts before drift reaches code. FILE:references/voice_calibration.md # Voice Calibration — Extracting Style from Sent-Email Samples This reference answers exactly one decision: **why are real sent-email samples the highest-quality voice signal for inbox-triage's draft generation, and how does the skill extract usable patterns from them deterministically?** Pair with `scripts/voice_sample_analyzer.py` for the deterministic extraction. ## The Core Claim Users describe their own voice unreliably. They say "professional but warm" and their actual emails alternate between three sentences of formal hedging and "lol no" replies to colleagues. They say "I'm pretty casual" and their actual emails open with "I hope this email finds you well." > **What users say about their voice ≠ what their voice actually is.** Real sent emails resolve this gap. They show: - Real opening phrases (not "I hope this email finds you well" if the user doesn't actually say that) - Real sentence length (not "short" if the actual average is 3 paragraphs) - Real sign-offs (not "thanks!" if the actual ratio is 80% "—Alex" and 20% no sign-off) - Real register (the variation across recipient type that self-description misses) ## What S3.SAMPLES Asks For > "Paste 3–5 real sent emails from your inbox." 3-5 is the operational sweet spot: - **<3:** too few to detect patterns vs anomalies - **3-5:** enough variance to detect baseline + adaptations - **>5:** marginal signal, diminishing returns; takes longer to extract The samples should span the user's typical email mix — at least one to a peer, one external, one transactional. If the user pastes 5 identical newsletters, ask for more variety. ## What `voice_sample_analyzer.py` Extracts Deterministic stdlib analysis (no LLM): 1. **Opening phrases** — first 5-10 tokens of each sample's body. Pattern frequency. 2. **Sign-offs** — last 5-10 tokens of each sample. Pattern frequency. 3. **Sentence length distribution** — short (<10 words) / medium (10-25) / long (>25) ratio. 4. **Register markers** — counts of casual indicators ("lol", "yeah", "tbh", "btw") vs formal indicators ("I would like to", "please find", "kindly"). 5. **Hedging frequency** — counts of softeners ("maybe", "I think", "perhaps", "just"). High hedging is a voice fingerprint. 6. **Personal pronouns** — "I" vs "we" frequency. Tells whether user writes as solo or representing a team. 7. **Punctuation patterns** — em-dash usage, exclamation marks, ellipses. Output is a structured patterns block that goes into `email-patterns.md` under "Voice Patterns (Extracted from Samples)." ## How Self-Description (S3.Q1-Q6) Combines With Samples Self-description and samples are **complementary**, not competing: - **Self-description wins for:** hard rules (S3.Q6 — "never emojis"), forbidden tokens (S3.Q2 — "phrases I hate"), explicit sign-offs (S3.Q3 — what the user remembers using). - **Samples win for:** baseline register, actual sentence length, opening phrases, register adaptation across recipient types. In `email-patterns.md`, the two are combined: self-described preferences are stated as hard rules; sample-extracted patterns supplement as baseline behavior. ## When Samples Aren't Available If the user refuses to paste samples (privacy, time, or just "I'd rather not"): 1. Honor the choice. Don't push back twice. 2. Use S3.Q1-Q6 self-description only. 3. Flag in `email-patterns.md`: ```markdown ## Voice Calibration Status [calibration may need iteration — voice samples not collected during setup. First few triage runs will likely produce drafts that need editing; the system learns from your edits and overrides. Re-run inbox-setup with samples when you're ready, OR triage will refine voice from your edit patterns over 5+ runs.] ``` 4. Inbox-triage will produce drafts in a more conservative default register (medium-formal, short-paragraph length). Drafts will need more editing on early runs. ## Common Anti-Patterns ### "I described my voice, that's enough" Self-description has known blind spots (per the "Core Claim" above). Even high-self-awareness users overestimate their formality or underestimate their hedging frequency. Skip the samples and the first 10 triage runs produce drafts that "sound off" in a way users struggle to articulate. ### "I'll paste 5 emails that are similar" 5 emails to peers about the same project don't show register adaptation. The skill needs variance: one to a peer, one to a client/external, one transactional. If user pastes 5 similar emails, ask for one more from a different context. ### "I'll paste from my drafts folder" Drafts may not represent voice the user actually sends — they may include rejected attempts. Ask for sent emails specifically. ### "I'll write 5 example emails for you" Written-for-the-skill emails are self-description in disguise. Reject: > "Examples written for me don't capture your actual voice — they capture how you describe your voice (which has known blind spots). Paste real sent emails, even short/boring ones. The mundane ones often signal voice better than carefully-crafted ones." ### "Forbidden tokens" extracted from samples instead of S3.Q2 Don't pull "forbidden tokens" from sample analysis — if a phrase appeared in a sent email, the user used it at some point. Forbidden tokens ONLY come from S3.Q2 (explicit "phrases I hate"). Voice extraction surfaces what the user DOES say, not what they DON'T. ## Operational Checklist (Per Setup Run) - [ ] S3.Q1-Q6 asked one at a time with "why I'm asking" - [ ] S3.SAMPLES asked AFTER Q1-Q6 (self-description first, samples second — samples calibrate the description, not replace it) - [ ] 3-5 samples collected (or explicit user-skip + flag in patterns file) - [ ] If collected: `scripts/voice_sample_analyzer.py` run; output incorporated into "Voice Patterns" subsection of patterns file - [ ] Self-described hard rules + forbidden tokens preserved as authoritative - [ ] Sample-extracted baseline preserved as descriptive (not authoritative) - [ ] Calibration-status block included in patterns file (states whether samples were collected) ## Why This Reference Exists The S3.SAMPLES step is the SINGLE most important question in the entire 25-31 question interview. Skipping it or doing it poorly compromises every subsequent triage run. This reference exists to make the discipline of "samples first, self-description second" explicit and operationally enforceable. ## Citations Voice analysis canon: 1. **Brian Kernighan & Rob Pike, *The Practice of Programming* (Addison-Wesley, 1999)** — Chapter 1 on Style. The point that "names describe roles, not types" generalizes: a user's voice describes their habits, not their aspirations. Sample-based extraction captures habits. 2. **Steven Pinker, *The Sense of Style* (Viking, 2014)** — Chapter on register and the "Classic Style" trap. Self-described voice often defaults to Classic Style ideals that the user's actual voice doesn't match. 3. **Bryan Garner, *Garner's Modern English Usage* (5th ed., Oxford, 2022)** — Sections on register variation and register-adaptation across contexts. The justification for requiring sample variance (peer / external / transactional). 4. **Geoffrey Pullum, *The Cambridge Grammar of the English Language* (Cambridge, 2002), Chapter 12** — Register theory. Establishes that register is detectable from text features (sentence length, pronoun choice, hedging frequency) more reliably than from speaker self-report. 5. **Stylometric authorship attribution literature** — work by Patrick Juola, José Nilo G. Binongo, and the broader stylometry community. Establishes that text features (function-word frequency, punctuation patterns, sentence-length distribution) are robust voice signals. The features `voice_sample_analyzer.py` extracts are a subset of this canonical set. 6. **John Searle, *Speech Acts* (Cambridge, 1969)** — Performative theory. Useful framing for the "hard rules" (S3.Q6) discipline: hard rules are performatives the user commits to; voice is descriptive. 7. **Email-writing style guides at scale: *The Yahoo! Style Guide* (St. Martin's, 2010), *The Microsoft Manual of Style* (4th ed.).** Real-world style guides establish that register depends heavily on recipient + context, not on a single "professional voice." Justifies asking for sample variance. FILE:scripts/kb_validator.py #!/usr/bin/env python3 """kb_validator.py — Validate the 7-file KB contract at WORKSPACE/Email/. Stdlib-only. Confirms the inbox-setup skill produced the files inbox-triage expects to read on every run. Used at end of Section 8 (Confirmation & Handoff) and any time the user wants to spot-check the KB state. Checks (per `references/kb_file_contract.md`): 1. Required core files exist: - email-taxonomy.md - email-patterns.md - blocklist.md - tracker.md 2. triage-log/ exists as a DIRECTORY (not a file) 3. Conditional files exist iff their triggering section ran: - evaluation-framework.md (only if opportunity emails category) - rate-card.md (only if user has pricing) 4. Each required file has an H1 header 5. email-taxonomy.md has both "## Categories" + "## Report Preferences" 6. email-patterns.md has "## Voice Calibration Status" (samples collected or not) Output: PASS / WARN / FAIL per rule + overall verdict. NO LLM CALLS. Pure filesystem + regex. Usage: python kb_validator.py --workspace /path/to/workspace python kb_validator.py --workspace . --expect-evaluation --expect-rate-card python kb_validator.py --sample """ import argparse import json import re import sys from pathlib import Path from typing import Any, Dict, List, Optional CORE_REQUIRED = ["email-taxonomy.md", "email-patterns.md", "blocklist.md", "tracker.md"] CONDITIONAL = ["evaluation-framework.md", "rate-card.md"] LOG_DIR = "triage-log" SAMPLE_KB: Dict[str, str] = { "email-taxonomy.md": ( "# Email Taxonomy\n\n## Categories\n\n### New Opportunities\n" "- Signals: pitch / proposal / collab\n- Default action: classify + draft\n\n" "### Newsletters\n- Signals: unsubscribe / newsletter / digest\n" "- Default action: skip\n\n## Report Preferences\n\n" "- Delivery format: email-draft-to-self\n- Detail level: 30-second-scan\n" ), "email-patterns.md": ( "# Email Patterns\n\n## Voice Register\nCasual\n\n## Hard Rules\n" "- Never: emojis in client emails\n- Always: reply within 24h\n\n" "## Voice Calibration Status\nSamples collected: 4 emails analyzed.\n" ), "blocklist.md": ( "# Blocklist\n\n## Sender / Domain Auto-Skip\n" "- recruiter@*: cold outreach — added 2026-05-15\n\n" "## Decline Patterns\n- 'looking for backend engineers': cold recruiter\n" ), "tracker.md": ( "# Tracker\n\n## Active Follow-Ups\n\n" "| Item | Context | Deadline | Status |\n|---|---|---|---|\n" "| Q3 contract | renewal due | 2026-06-15 | pending |\n\n## Overdue\n\n" "## Resolved (Recent)\n\n## Update Log\n" ), "evaluation-framework.md": ( "# Evaluation Framework (Opportunity Emails)\n\n## Gut Filter (First Check)\n" "Is the budget realistic for the scope?\n\n## TAKE-IT Signals\n- Clear budget stated\n" "- VIP sender\n- Aligned to stated focus\n\n## PASS Signals (Instant Deal-Breakers)\n" "- Free / unpaid\n- Equity-only\n- Out-of-scope industry\n" ), } def check_file(workspace: Path, filename: str) -> Dict[str, Any]: p = workspace / "Email" / filename return { "filename": filename, "exists": p.exists() and p.is_file(), "path": str(p), "size": p.stat().st_size if p.exists() and p.is_file() else 0, } def check_h1(workspace: Path, filename: str) -> Optional[str]: p = workspace / "Email" / filename if not p.exists() or not p.is_file(): return None try: for line in p.read_text(encoding="utf-8").splitlines(): m = re.match(r"^#\s+(.+?)\s*$", line) if m: return m.group(1).strip() return None except OSError: return None def has_section(workspace: Path, filename: str, section_header: str) -> bool: p = workspace / "Email" / filename if not p.exists() or not p.is_file(): return False try: text = p.read_text(encoding="utf-8") return bool(re.search(rf"^##\s+{re.escape(section_header)}\s*$", text, re.MULTILINE)) except OSError: return False def validate( workspace: Path, expect_evaluation: bool = False, expect_rate_card: bool = False, ) -> Dict[str, Any]: findings: List[Dict[str, str]] = [] def add(rule: str, level: str, message: str) -> None: findings.append({"rule": rule, "level": level, "message": message}) email_dir = workspace / "Email" if not email_dir.exists(): add("workspace-email-dir", "FAIL", f"{email_dir} does not exist. Run inbox-setup first.") return finalize(findings) if not email_dir.is_dir(): add("workspace-email-dir", "FAIL", f"{email_dir} is not a directory.") return finalize(findings) add("workspace-email-dir", "PASS", f"{email_dir} exists.") # Core required files for fn in CORE_REQUIRED: info = check_file(workspace, fn) if not info["exists"]: add(f"core-file:{fn}", "FAIL", f"Required file missing: Email/{fn}") elif info["size"] == 0: add(f"core-file:{fn}", "FAIL", f"Required file is empty: Email/{fn}") else: add(f"core-file:{fn}", "PASS", f"Email/{fn} present ({info['size']} bytes).") # H1 check on core files that exist for fn in CORE_REQUIRED: if not (workspace / "Email" / fn).exists(): continue h1 = check_h1(workspace, fn) if h1: add(f"h1:{fn}", "PASS", f"Email/{fn} H1: '{h1}'") else: add(f"h1:{fn}", "FAIL", f"Email/{fn} has no H1.") # email-taxonomy.md must have both required subsections if (workspace / "Email" / "email-taxonomy.md").exists(): if has_section(workspace, "email-taxonomy.md", "Categories"): add("taxonomy-categories", "PASS", "email-taxonomy.md has '## Categories' section.") else: add("taxonomy-categories", "FAIL", "email-taxonomy.md missing '## Categories' section.") if has_section(workspace, "email-taxonomy.md", "Report Preferences"): add("taxonomy-report-prefs", "PASS", "email-taxonomy.md has '## Report Preferences' section.") else: add("taxonomy-report-prefs", "WARN", "email-taxonomy.md missing '## Report Preferences' section (added at end of S7).") # email-patterns.md must have Voice Calibration Status if (workspace / "Email" / "email-patterns.md").exists(): if has_section(workspace, "email-patterns.md", "Voice Calibration Status"): add("patterns-calibration", "PASS", "email-patterns.md has '## Voice Calibration Status' section.") else: add("patterns-calibration", "WARN", "email-patterns.md missing '## Voice Calibration Status' section (states whether samples were collected).") # Conditional files for fn in CONDITIONAL: info = check_file(workspace, fn) expect = (fn == "evaluation-framework.md" and expect_evaluation) or (fn == "rate-card.md" and expect_rate_card) if expect and not info["exists"]: add(f"conditional-file:{fn}", "FAIL", f"Expected (per --expect flag) but missing: Email/{fn}") elif not expect and info["exists"]: add(f"conditional-file:{fn}", "WARN", f"Email/{fn} exists but neither --expect-evaluation nor --expect-rate-card was set (may be stale from earlier setup).") elif expect and info["exists"]: add(f"conditional-file:{fn}", "PASS", f"Email/{fn} present (expected).") else: add(f"conditional-file:{fn}", "PASS", f"Email/{fn} correctly absent (not expected).") # triage-log/ must be a directory triage_log = workspace / "Email" / LOG_DIR if not triage_log.exists(): add("triage-log-dir", "FAIL", f"Email/{LOG_DIR}/ missing. Must be created as empty directory at end of S6.") elif not triage_log.is_dir(): add("triage-log-dir", "FAIL", f"Email/{LOG_DIR} exists but is not a directory.") else: add("triage-log-dir", "PASS", f"Email/{LOG_DIR}/ exists as directory.") return finalize(findings) def finalize(findings: List[Dict[str, str]]) -> Dict[str, Any]: counts = {"PASS": 0, "WARN": 0, "FAIL": 0} for f in findings: counts[f["level"]] += 1 if counts["FAIL"] > 0: verdict = "FAIL" elif counts["WARN"] > 0: verdict = "WARN" else: verdict = "PASS" return {"verdict": verdict, "counts": counts, "findings": findings} def render_human(result: Dict[str, Any]) -> str: out: List[str] = [] out.append(f"KB contract verdict: {result['verdict']}") counts = result["counts"] out.append(f" PASS: {counts['PASS']} WARN: {counts['WARN']} FAIL: {counts['FAIL']}") out.append("") out.append("Findings:") for f in result["findings"]: marker = {"PASS": "[ok]", "WARN": "[warn]", "FAIL": "[FAIL]"}[f["level"]] out.append(f" {marker} {f['rule']}: {f['message']}") return "\n".join(out) def run_sample() -> Dict[str, Any]: import tempfile with tempfile.TemporaryDirectory() as td: ws = Path(td) email_dir = ws / "Email" email_dir.mkdir(parents=True) for name, content in SAMPLE_KB.items(): (email_dir / name).write_text(content, encoding="utf-8") (email_dir / LOG_DIR).mkdir() return validate(ws, expect_evaluation=True, expect_rate_card=False) def main(argv: List[str]) -> int: parser = argparse.ArgumentParser(description=__doc__.split("\n")[0]) parser.add_argument("--workspace", help="Path to workspace (looks at <workspace>/Email/)") parser.add_argument("--expect-evaluation", action="store_true", help="Expect evaluation-framework.md to exist") parser.add_argument("--expect-rate-card", action="store_true", help="Expect rate-card.md to exist") parser.add_argument("--sample", action="store_true", help="Run on embedded sample KB") parser.add_argument("--output", choices=["human", "json"], default="human") args = parser.parse_args(argv) if args.sample: result = run_sample() elif args.workspace: ws = Path(args.workspace) if not ws.exists(): print(f"error: {args.workspace} not found", file=sys.stderr) return 2 result = validate(ws, args.expect_evaluation, args.expect_rate_card) else: parser.print_help() return 0 if args.output == "json": print(json.dumps(result, indent=2)) else: print(render_human(result)) return 0 if result["verdict"] != "FAIL" else 1 if __name__ == "__main__": sys.exit(main(sys.argv[1:])) FILE:scripts/section_progress_tracker.py #!/usr/bin/env python3 """section_progress_tracker.py — JSON-backed walk state for 8-section setup. Stdlib-only. Tracks the setup interview state at ~/.inbox_setup_sessions/<session>.json so the skill can: - Know which section is currently active - Record each question's answer - Mark each section as done with the file(s) it committed - Detect drop-off and produce useful partial state - Resume later if the user drops off mid-interview Actions: start Create a new session record_q Record an answer to a question record_section_done Mark section complete with files committed status Show current session state list List all sessions close Mark session ended Usage: python section_progress_tracker.py --action start --session inbox-setup-20260515 --user alice python section_progress_tracker.py --action record_q --session ... --section 1 --question 1 --answer "Solo consultant" python section_progress_tracker.py --action record_section_done --session ... --section 2 --files "email-taxonomy.md" python section_progress_tracker.py --action status --session ... python section_progress_tracker.py --action list python section_progress_tracker.py --action close --session ... """ import argparse import json import sys from datetime import datetime, timezone from pathlib import Path from typing import Any, Dict, List, Optional SESSIONS_DIR = Path.home() / ".inbox_setup_sessions" TOTAL_SECTIONS = 8 def session_path(name: str) -> Path: return SESSIONS_DIR / f"{name}.json" def load_session(name: str) -> Dict[str, Any]: p = session_path(name) if not p.exists(): raise FileNotFoundError(f"Session not found: {name}") return json.loads(p.read_text(encoding="utf-8")) def save_session(name: str, data: Dict[str, Any]) -> None: SESSIONS_DIR.mkdir(parents=True, exist_ok=True) session_path(name).write_text(json.dumps(data, indent=2), encoding="utf-8") def now_iso() -> str: return datetime.now(timezone.utc).isoformat() def action_start(name: str, user: Optional[str]) -> Dict[str, Any]: if session_path(name).exists(): raise FileExistsError(f"Session already exists: {name}") data: Dict[str, Any] = { "session": name, "user": user or "(anonymous)", "started_at": now_iso(), "ended_at": None, "active_section": 1, "sections": {str(i): {"status": "pending", "questions_answered": [], "files_committed": []} for i in range(1, TOTAL_SECTIONS + 1)}, "total_questions_answered": 0, "skip_log": [], } save_session(name, data) return data def action_record_q(name: str, section: int, question: int, answer: str) -> Dict[str, Any]: data = load_session(name) key = str(section) if key not in data["sections"]: raise ValueError(f"Invalid section: {section}") sec = data["sections"][key] if sec["status"] == "pending": sec["status"] = "in_progress" sec["started_at"] = now_iso() sec["questions_answered"].append({ "question": question, "answer": answer, "at": now_iso(), }) data["total_questions_answered"] += 1 data["active_section"] = section save_session(name, data) return data def action_record_section_done(name: str, section: int, files: List[str]) -> Dict[str, Any]: data = load_session(name) key = str(section) if key not in data["sections"]: raise ValueError(f"Invalid section: {section}") sec = data["sections"][key] sec["status"] = "done" sec["files_committed"] = files sec["ended_at"] = now_iso() # Advance active section if section < TOTAL_SECTIONS: data["active_section"] = section + 1 save_session(name, data) return data def action_record_skip(name: str, section: int, reason: str) -> Dict[str, Any]: data = load_session(name) key = str(section) sec = data["sections"][key] sec["status"] = "skipped" sec["skip_reason"] = reason sec["ended_at"] = now_iso() data["skip_log"].append({"section": section, "reason": reason, "at": now_iso()}) if section < TOTAL_SECTIONS: data["active_section"] = section + 1 save_session(name, data) return data def action_status(name: str) -> Dict[str, Any]: return load_session(name) def action_close(name: str) -> Dict[str, Any]: data = load_session(name) if data.get("ended_at") is None: data["ended_at"] = now_iso() save_session(name, data) return data def action_list() -> List[Dict[str, Any]]: SESSIONS_DIR.mkdir(parents=True, exist_ok=True) out: List[Dict[str, Any]] = [] for p in sorted(SESSIONS_DIR.glob("*.json")): try: data = json.loads(p.read_text(encoding="utf-8")) done_sections = sum(1 for s in data["sections"].values() if s["status"] == "done") out.append({ "session": data["session"], "user": data["user"], "started_at": data["started_at"], "ended_at": data["ended_at"], "active_section": data["active_section"], "done_sections": done_sections, "total_questions_answered": data["total_questions_answered"], }) except (OSError, json.JSONDecodeError): continue return out def render_status_human(data: Dict[str, Any]) -> str: out: List[str] = [] out.append(f"Session: {data['session']}") out.append(f"User: {data['user']}") out.append(f"Started: {data['started_at']}") out.append(f"Ended: {data.get('ended_at') or '(active)'}") out.append(f"Active section: {data['active_section']}/{TOTAL_SECTIONS}") out.append(f"Total Qs answered:{data['total_questions_answered']}") out.append("") out.append("Per-section state:") for key in sorted(data["sections"].keys(), key=lambda k: int(k)): sec = data["sections"][key] marker = {"pending": " ", "in_progress": "↻ ", "done": "✓ ", "skipped": "→ "}.get(sec["status"], " ") files = ", ".join(sec["files_committed"]) if sec["files_committed"] else "—" out.append(f" {marker}S{key}: {sec['status']:<12s} ({len(sec['questions_answered'])} Q answered, files: {files})") if data["skip_log"]: out.append("") out.append("Skip log:") for s in data["skip_log"]: out.append(f" S{s['section']}: {s['reason']}") return "\n".join(out) def render_list_human(rows: List[Dict[str, Any]]) -> str: if not rows: return "(no sessions)" out: List[str] = [] out.append(f"{'session':<40s} {'user':<15s} {'active':>6s} {'done':>4s} {'Q':>3s} status") out.append("-" * 90) for r in rows: status = "closed" if r["ended_at"] else "active" out.append( f"{r['session']:<40s} {r['user']:<15s} {r['active_section']:>6d} {r['done_sections']:>4d} {r['total_questions_answered']:>3d} {status}" ) return "\n".join(out) def main(argv: List[str]) -> int: parser = argparse.ArgumentParser(description=__doc__.split("\n")[0]) parser.add_argument("--action", required=True, choices=["start", "record_q", "record_section_done", "record_skip", "status", "list", "close"]) parser.add_argument("--session", help="Session name") parser.add_argument("--user", help="(start only) user identifier") parser.add_argument("--section", type=int, help="Section number 1-8") parser.add_argument("--question", type=int, help="(record_q only) question number within section") parser.add_argument("--answer", help="(record_q only) answer text") parser.add_argument("--files", help="(record_section_done only) comma-separated filenames") parser.add_argument("--reason", help="(record_skip only) why section was skipped") parser.add_argument("--output", choices=["human", "json"], default="human") args = parser.parse_args(argv) try: if args.action == "start": if not args.session: print("error: --session required for start", file=sys.stderr); return 2 result = action_start(args.session, args.user) elif args.action == "record_q": if not (args.session and args.section and args.question is not None and args.answer is not None): print("error: --session, --section, --question, --answer required", file=sys.stderr); return 2 result = action_record_q(args.session, args.section, args.question, args.answer) elif args.action == "record_section_done": if not (args.session and args.section and args.files): print("error: --session, --section, --files required", file=sys.stderr); return 2 files = [f.strip() for f in args.files.split(",") if f.strip()] result = action_record_section_done(args.session, args.section, files) elif args.action == "record_skip": if not (args.session and args.section and args.reason): print("error: --session, --section, --reason required", file=sys.stderr); return 2 result = action_record_skip(args.session, args.section, args.reason) elif args.action == "status": if not args.session: print("error: --session required for status", file=sys.stderr); return 2 result = action_status(args.session) elif args.action == "close": if not args.session: print("error: --session required for close", file=sys.stderr); return 2 result = action_close(args.session) else: result = action_list() except (FileNotFoundError, FileExistsError, ValueError) as e: print(f"error: {e}", file=sys.stderr); return 2 if args.output == "json": print(json.dumps(result, indent=2, default=str)) else: if args.action == "list": print(render_list_human(result)) else: print(render_status_human(result)) return 0 if __name__ == "__main__": sys.exit(main(sys.argv[1:])) FILE:scripts/voice_sample_analyzer.py #!/usr/bin/env python3 """voice_sample_analyzer.py — Extract voice patterns from sent-email samples. Stdlib-only. Reads 3-5 sent-email samples (separated by `---` delimiters) and extracts deterministic voice signals: 1. Opening phrases — first 4-6 tokens of each sample body 2. Sign-offs — last 4-6 tokens of each sample 3. Sentence-length distribution — short (<10 words) / medium (10-25) / long (>25) ratio 4. Register markers — counts of casual indicators (lol, yeah, btw, tbh) vs formal (I would like to, please find, kindly) 5. Hedging frequency — counts of softeners (maybe, perhaps, I think, just) 6. Personal pronouns — "I" vs "we" ratio 7. Punctuation patterns — em-dashes, exclamation marks, ellipses per sample Output: a structured patterns block that gets dropped into email-patterns.md under "Voice Patterns (Extracted from Samples)". NO LLM CALLS. Pure regex + frequency counting. Limitations (intentional, stdlib-only): - No semantic understanding (it's surface-feature stylometry) - English-only register markers - Tokenization is whitespace-based (not linguistic) Usage: python voice_sample_analyzer.py --samples-file /path/to/samples.txt python voice_sample_analyzer.py --samples-file /path/to/samples.txt --output json python voice_sample_analyzer.py --sample """ import argparse import json import re import sys from collections import Counter from pathlib import Path from typing import Any, Dict, List, Tuple SAMPLE_DELIMITER_RE = re.compile(r"^\s*---+\s*$", re.MULTILINE) SENTENCE_END_RE = re.compile(r"[.!?]+(?:\s|$)") CASUAL_MARKERS = { "lol", "lmao", "haha", "yeah", "yup", "nope", "tbh", "btw", "fwiw", "imo", "imho", "rn", "btw", "ok", "okay", "cool", "sure", "yep", "gonna", "wanna", "kinda", "sorta", "dunno", } FORMAL_MARKERS_PHRASES = [ "i would like to", "please find", "kindly", "i hope this email finds you", "i am writing to", "as per our", "at your earliest convenience", "thank you for your", "i look forward to hearing", "to whom it may concern", "respectfully", "sincerely", ] HEDGING_MARKERS = { "maybe", "perhaps", "i think", "i guess", "i suppose", "just", "kinda", "sorta", "might", "could", "possibly", "potentially", "i feel", "i believe", } def split_samples(text: str) -> List[str]: """Split combined samples text on `---` delimiters; trim each.""" parts = SAMPLE_DELIMITER_RE.split(text) return [p.strip() for p in parts if p.strip()] def first_n_tokens(text: str, n: int) -> str: tokens = text.split() return " ".join(tokens[:n]) def last_n_tokens(text: str, n: int) -> str: tokens = text.split() return " ".join(tokens[-n:]) def count_phrase_occurrences(text_lower: str, phrases: List[str]) -> int: return sum(text_lower.count(p) for p in phrases) def count_word_occurrences(text_lower: str, words: set) -> int: pattern = re.compile(rf"\b({'|'.join(re.escape(w) for w in words)})\b", re.IGNORECASE) return len(pattern.findall(text_lower)) def split_sentences(text: str) -> List[str]: parts = SENTENCE_END_RE.split(text) return [s.strip() for s in parts if s.strip()] def length_bucket(word_count: int) -> str: if word_count < 10: return "short" if word_count <= 25: return "medium" return "long" def analyze_sample(sample: str) -> Dict[str, Any]: text_lower = sample.lower() sentences = split_sentences(sample) length_dist = Counter() for s in sentences: words = s.split() length_dist[length_bucket(len(words))] += 1 return { "opening": first_n_tokens(sample, 6), "sign_off": last_n_tokens(sample, 6), "sentence_count": len(sentences), "length_distribution": dict(length_dist), "casual_marker_count": count_word_occurrences(text_lower, CASUAL_MARKERS), "formal_marker_count": count_phrase_occurrences(text_lower, FORMAL_MARKERS_PHRASES), "hedging_count": count_word_occurrences(text_lower, HEDGING_MARKERS), "i_count": count_word_occurrences(text_lower, {"i", "i'm", "i've", "i'll", "i'd"}), "we_count": count_word_occurrences(text_lower, {"we", "we're", "we've", "we'll", "we'd", "our", "us"}), "em_dash_count": sample.count("—") + sample.count(" -- "), "exclamation_count": sample.count("!"), "ellipsis_count": sample.count("...") + sample.count("…"), } def aggregate(per_sample: List[Dict[str, Any]]) -> Dict[str, Any]: if not per_sample: return {"error": "no samples"} n = len(per_sample) openings = [s["opening"] for s in per_sample] sign_offs = [s["sign_off"] for s in per_sample] total_sentences = sum(s["sentence_count"] for s in per_sample) total_lengths: Counter = Counter() for s in per_sample: total_lengths.update(s["length_distribution"]) casual = sum(s["casual_marker_count"] for s in per_sample) formal = sum(s["formal_marker_count"] for s in per_sample) hedging = sum(s["hedging_count"] for s in per_sample) i_count = sum(s["i_count"] for s in per_sample) we_count = sum(s["we_count"] for s in per_sample) em_dash = sum(s["em_dash_count"] for s in per_sample) exclamation = sum(s["exclamation_count"] for s in per_sample) ellipsis = sum(s["ellipsis_count"] for s in per_sample) if casual > formal * 2: register_verdict = "casual" elif formal > casual * 2: register_verdict = "formal" else: register_verdict = "in-between" if total_sentences > 0: short_ratio = total_lengths.get("short", 0) / total_sentences medium_ratio = total_lengths.get("medium", 0) / total_sentences long_ratio = total_lengths.get("long", 0) / total_sentences else: short_ratio = medium_ratio = long_ratio = 0.0 if short_ratio > 0.5: length_verdict = "one-liner / short-paragraph" elif long_ratio > 0.3: length_verdict = "longer (multi-paragraph)" else: length_verdict = "short-paragraph (medium average)" return { "sample_count": n, "openings": openings, "sign_offs": sign_offs, "register_verdict": register_verdict, "register_signals": {"casual_markers": casual, "formal_markers": formal}, "length_verdict": length_verdict, "length_distribution": { "short_pct": round(short_ratio * 100, 1), "medium_pct": round(medium_ratio * 100, 1), "long_pct": round(long_ratio * 100, 1), }, "hedging_frequency_per_sample": round(hedging / n, 2), "i_vs_we": { "i_count": i_count, "we_count": we_count, "voice": "individual" if i_count > we_count * 2 else "team" if we_count > i_count * 2 else "mixed", }, "punctuation": { "em_dash_per_sample": round(em_dash / n, 2), "exclamation_per_sample": round(exclamation / n, 2), "ellipsis_per_sample": round(ellipsis / n, 2), }, } def render_human(result: Dict[str, Any]) -> str: out: List[str] = [] out.append(f"Voice analysis ({result['sample_count']} samples)") out.append("") out.append(f"Register verdict: {result['register_verdict']}") out.append(f" Casual markers: {result['register_signals']['casual_markers']}") out.append(f" Formal markers: {result['register_signals']['formal_markers']}") out.append("") out.append(f"Length verdict: {result['length_verdict']}") ld = result['length_distribution'] out.append(f" Short / Medium / Long: {ld['short_pct']}% / {ld['medium_pct']}% / {ld['long_pct']}%") out.append("") out.append(f"Hedging frequency: {result['hedging_frequency_per_sample']} per sample") iw = result['i_vs_we'] out.append(f"I vs We voice: {iw['voice']} (I:{iw['i_count']} We:{iw['we_count']})") out.append("") p = result['punctuation'] out.append(f"Punctuation per sample: em-dash {p['em_dash_per_sample']}, ! {p['exclamation_per_sample']}, ... {p['ellipsis_per_sample']}") out.append("") out.append("Opening phrases (first 6 tokens):") for o in result['openings']: out.append(f" - {o}") out.append("") out.append("Sign-offs (last 6 tokens):") for s in result['sign_offs']: out.append(f" - {s}") out.append("") out.append("Output block for email-patterns.md:") out.append("---") out.append("## Voice Patterns (Extracted from Samples)") out.append("") out.append(f"- Register: {result['register_verdict']}") out.append(f"- Typical reply length: {result['length_verdict']}") out.append(f"- Hedging frequency: {result['hedging_frequency_per_sample']} per email") out.append(f"- Voice perspective: {result['i_vs_we']['voice']}") out.append(f"- Sentence-length distribution: short {ld['short_pct']}% / medium {ld['medium_pct']}% / long {ld['long_pct']}%") out.append("- Observed opening patterns:") for o in result['openings'][:5]: out.append(f" - \"{o}\"") out.append("- Observed sign-off patterns:") for s in result['sign_offs'][:5]: out.append(f" - \"{s}\"") return "\n".join(out) SAMPLE_TEXT = """Hey, just looping back on the Q3 launch — pricing's mostly locked but I want to revisit the bundle option before we ship. Quick call tomorrow? —Alex --- Thanks for the proposal. Honestly, the timeline is tight and our team is heads-down on shipping. We'd need to push to Q4. Open to that? Alex --- Got it — sending the revised draft now. Couple of comments inline, mostly around the auth flow. Let me know what you think. Best, Alex --- I'm going to pass on this one. Scope is too broad for what we can commit to in the next 6 weeks and the budget doesn't match the work involved. Thanks for thinking of us though. —Alex --- Quick update: shipped the migration today, no incidents so far. Will keep an eye on it through the weekend. Lmk if you see anything weird. """ def main(argv: List[str]) -> int: parser = argparse.ArgumentParser(description=__doc__.split("\n")[0]) parser.add_argument("--samples-file", help="Path to file containing sent-email samples separated by ---") parser.add_argument("--sample", action="store_true", help="Analyze embedded sample text") parser.add_argument("--output", choices=["human", "json"], default="human") args = parser.parse_args(argv) if args.sample: text = SAMPLE_TEXT elif args.samples_file: p = Path(args.samples_file) if not p.exists(): print(f"error: {args.samples_file} not found", file=sys.stderr); return 2 text = p.read_text(encoding="utf-8") else: parser.print_help(); return 0 samples = split_samples(text) if not samples: print("error: no samples detected (use --- as delimiter between samples)", file=sys.stderr); return 2 if len(samples) < 3: print(f"warning: only {len(samples)} sample(s) detected; recommend 3-5 for reliable patterns", file=sys.stderr) per_sample = [analyze_sample(s) for s in samples] result = aggregate(per_sample) if args.output == "json": print(json.dumps(result, indent=2)) else: print(render_human(result)) return 0 if __name__ == "__main__": sys.exit(main(sys.argv[1:]))
Triển khai hợp tác với influencer và creator: tìm, thẩm định đối tác, cấu trúc thỏa thuận, brief, tuân thủ công bố và đo ROI.
---
name: influencer-marketing
description: "When the user wants to run influencer, creator, or ambassador partnerships to promote their product — finding and vetting partners, structuring deals, briefing creators, disclosure compliance, and measuring ROI. Also use when the user mentions 'influencer marketing,' 'creator partnerships,' 'sponsorships,' 'YouTube sponsorships,' 'podcast sponsorships,' 'brand ambassador,' 'ambassador program,' 'creator program,' 'UGC creators,' 'tech UGC,' 'UGC creator program,' 'creator network,' 'B2B influencers,' 'thought leader ads,' 'gifting,' 'product seeding,' 'whitelisting creator content,' 'how much to pay an influencer,' or 'FTC disclosure.' For affiliate/referral payout mechanics, see referrals. For community-led advocacy, see community-marketing. For turning creator content into paid ads, see ad-creative."
metadata:
version: 1.1.0
---
# Influencer & Creator Marketing
You are an expert in influencer, creator, and ambassador marketing across B2C (Instagram, TikTok, YouTube) and B2B (LinkedIn, X, newsletters, niche podcasts). Your goal is to help the user pick the right partners, structure fair deals, keep the program compliant, and measure real ROI — not vanity reach.
> Foundation contributed by @Adi29102000-s; compensation benchmarks and run-of-show checklist adapted from @SamSon75's PR; expanded to the repo's standard.
## Before Starting
**Check for product marketing context first.** If `.agents/product-marketing.md` exists (or `.claude/product-marketing.md`, or legacy `product-marketing-context.md`), read it before asking questions — the ICP, positioning, and offer anchor every partner-fit decision. Then gather what's missing: goal (awareness / conversions / content / trust), budget and whether it's cash or product, target platform(s), and any brand-safety redlines.
## The Influencer ↔ Ambassador Spectrum
"Influencer marketing" and "ambassador programs" are points on one spectrum — from a one-off paid post to an unpaid long-term advocate. Pick the model that fits the goal and stage, not the buzzword:
| Model | What it is | Pay | Best for | Home |
|---|---|---|---|---|
| **Paid influencer** | A creator posts sponsored content for a fee | Cash (flat / hybrid) | Reach + a credibility borrow, fast | This skill |
| **Affiliate creator** | A creator promotes for commission on sales | Performance (CPA/rev-share) | Conversion at scale, low upfront risk | This skill + **referrals** (payout mechanics) |
| **Gifted / seeding** | Free product, no obligation to post | Product only | Physical DTC, nano/micro, volume | This skill |
| **Brand ambassador program** | A cohort of ongoing advocates (paid, gifted, or perks) posting over months | Mixed / perks | Sustained presence, community depth | This skill (design below) + **community-marketing** |
| **Organic advocate** | A customer who already recommends you unprompted | None | Authenticity, cheapest trust | **community-marketing** |
The further right you go, the more it's about *relationship* than *transaction* — and the cheaper and more durable the trust, but the slower to scale. Most programs blend several (a few paid macro placements for reach + a gifted micro cohort + an affiliate tier for conversion).
**One more model — the volume UGC creator program ("tech UGC"):** an in-house network of creators posting disclosed native short-form from dedicated brand-affiliated accounts at test volume (10 creators × 3 posts/day ≈ 900 organic tests/month). Content volume, not any creator's audience, is the asset. See [references/ugc-creator-program.md](references/ugc-creator-program.md) for the full system — playbook-first concepts, the four formats, trial-week vetting, account warming, the review loop, the conversion ladder, and the compliance rewrite that makes the viral version of this playbook legal to run.
## 1. Finding & Vetting Partners
Influence is trust and relevance, not follower count.
**The audience-alignment test.** Don't ask "Are they famous?" Ask "Does their *audience* match our ICP?" A 12k-follower creator whose audience is exactly your buyer beats a 500k generalist. Where you can, look at *their* audience (comments, who engages, any media-kit demographics), not just the creator.
**Creator tiers** (reach vs. trust trade-off):
| Tier | Followers | Character |
|---|---|---|
| **Nano** | 1k–10k | Highest engagement, hyper-niche, often works for gifting. High ROI, low reach. |
| **Micro** | 10k–50k | Best balance of reach and trust; usually paid; strong conversion. |
| **Mid** | 50k–500k | Broader reach, more awareness than conversion, pricier. |
| **Macro / celebrity** | 500k+ | Top-of-funnel awareness; lowest conversion rate per follower; expensive. |
| **B2B thought leader** | Any size | LinkedIn creators, newsletter writers, niche podcasters — small audiences, extreme purchasing power. Judge by *who* follows, not how many. |
For most brands, a portfolio of **micro + nano** partners out-converts one macro placement at the same total spend — and produces more content to repurpose.
**Vetting checklist:**
- **Engagement rate**, not follower count (a rough floor: ~1–3% is healthy on IG/TikTok at scale; higher for nano). Suspiciously round numbers, comment pods, or comments that don't match the audience are red flags.
- **Fake-follower / bot check** — a sudden follower spike, generic comments, or engagement wildly out of line with reach. Tools like SparkToro (audience intelligence) help; media kits overstate.
- **Sponsored-content track record** — do their *ads* still get engagement, or does their audience tune out promos? Ask for past campaign results.
- **Brand safety** — scroll their last ~3 months. Controversy, competitor conflicts, or off-brand content that would attach to you.
- **Authenticity of fit** — have they mentioned your category unprompted? A genuine user is worth several cold partners.
## 2. Outreach
Reach out **1:1 and personally** — reference specific content, why *them*, and what's in it for their audience. A generic form blast to 200 creators converts worse than 20 tailored notes. For writing the outreach itself, use **cold-email** (personalization, deliverability, follow-up cadence). Lead with the offer and the fit; don't bury the ask.
## 3. Structuring the Deal
Move beyond "pay for a post."
**Compensation models:**
- **Flat fee** — standard for awareness; you pay for the placement regardless of result.
- **Performance / CPA** — pay per click or conversion. Hard to get larger creators to accept without a baseline; best with affiliate-minded creators (see **referrals** for tracking + payout).
- **Hybrid (flat + CPA)** — usually the best deal: a lower baseline to cover their production time, plus commission for upside. Aligns incentives.
- **Gifting / seeding** — free product, no obligation. Works for physical DTC with nano/micro at volume; expect a low but authentic post rate.
**Rate reality:** published "rates" are wildly variable by niche, geography, and platform, and creators quote high. Treat any benchmark as a *range to negotiate from*, not a price — and anchor on **cost per qualified outcome** (CPA, cost per qualified follower/lead), not cost per post. A cheap post to the wrong audience is the expensive one.
**Starting ranges for a single post** (negotiation anchors, *not* fixed prices — aligned to the tiers above):
| Tier | Single post (rough range) | Notes |
|---|---|---|
| **Nano** (1k–10k) | Free product – $100 | Often product-only |
| **Micro** (10k–50k) | $100 – $1,500 | Widest range; negotiate on engagement, not follower count |
| **Mid** (50k–500k) | $1,500 – $10,000 | Rate cards common at this tier |
| **Macro / celebrity** (500k+) | $10,000 – $30,000+ | Usually has an agent/manager |
| **Video / long-form** (YouTube) | Higher than short-form at the same follower count | More production effort |
| **B2B thought leader** | Priced on audience quality, not size | A 5k-follower niche voice can command more than a 200k generalist |
Ask for their **rate card first** — it sets an anchor you respond to rather than naming a number blind.
**Deliverables to negotiate:**
- **Content usage rights (crucial)** — the right to repurpose their content as **paid ads** (whitelisting / dark posting / "creator ads") for a defined window (commonly 3–6 months). This is often the highest-ROI clause: their content becomes your best-performing ad. Then run it through **ad-creative** (and present variations for sign-off with the creative review page).
- **Exclusivity** — competitor lockout for a set period; costs more, worth it in tight categories.
- **Format & specifics** — dedicated video vs. a 60-second integration; number of posts; stories vs. feed; posting window; approval rights; how long it stays up.
- **Approvals & revisions** — one review round is normal; scripting word-for-word is not (below).
Put it in a simple written agreement: deliverables, timing, usage rights, exclusivity, disclosure obligation (below), payment terms, and a kill/rework clause.
## 4. Disclosure & Compliance (non-negotiable)
Influencer marketing has hard legal requirements — this is the part most brands under-do, and the brand — not just the creator — can be held liable.
- **Any material connection must be disclosed** — payment, free product, commission, a family/employee relationship, even a free trial. Gifting is *not* a loophole; a gifted post still needs disclosure.
- **The disclosure must be clear and hard to miss** — "#ad" or "#sponsored" placed where viewers actually see it (not buried in a wall of hashtags, not below the "more" fold, and spoken aloud in video/audio, not just in the description). "#sp," "#collab," "#ambassador," and "thanks to [brand]" are considered insufficient on their own by the FTC.
- **Use the platform's own tool** — Instagram/TikTok/YouTube "paid partnership" labels *in addition to* the written disclosure, not instead of it.
- **You're responsible for your creators.** Build the disclosure requirement into the brief and the agreement, and check that they actually did it. Non-disclosure exposes the brand to liability, not just the creator — the FTC expects advertisers to have a program to guide, monitor, and remediate disclosure (FTC actions target advertisers).
- **No fabricated claims.** Creators can't say things about the product that aren't true, can't fake results, and can't imply they're a customer if they aren't. Give them what's true and let them speak it in their voice.
- **International + platform rules vary** (e.g., stricter regimes in the UK/EU, category rules for health/finance/alcohol). When the campaign is regulated or cross-border, route to legal.
Disclosure done well doesn't hurt performance — audiences expect it, and the FTC has never found "#ad" to tank a genuinely good integration.
## 5. The Creative Brief
Do **not** script the creator word-for-word — they know their audience better than you, and scripted reads convert worst. Provide:
- **The "why"** — the core problem your product solves (the one sentence).
- **Key talking points (2–3 max)** — the most important benefits; more than three and none land.
- **The CTA** — exactly what to tell the audience to do (a specific vanity link, a unique promo code).
- **Guardrails** — what *not* to say (don't promise features that don't exist), the disclosure requirement, and any brand redlines.
- **Creative freedom** — explicitly grant it. The integration should live inside their normal content style.
Ground the talking points in real proof (reviews, results) — same grounding discipline as **ad-creative**'s inputs. Never hand a creator a claim you can't back.
## 6. Measurement & ROI
Influencer marketing suffers from attribution gaps — fix them upfront, before the campaign runs:
- **Unique promo codes** (e.g., `CREATOR20`) — the easiest direct-conversion tracker, and essential for podcasts/video where links aren't clickable.
- **UTM tracking links** — mandatory on every digital placement; one per creator per placement.
- **Vanity / dedicated landing pages** — `yourdomain.com/creatorname` with a personalized welcome; lifts conversion *and* attributes cleanly.
- **Post-purchase survey** — "How did you hear about us?" catches the halo/branded-search effect that promo codes and last-click miss (much of influencer impact shows up later as branded search and direct — see the attribution blind spot in **ai-seo**'s citations-vs-recommendations).
- **Whitelisting performance** — when you repurpose creator content as ads, that ad's own metrics are a clean read on the creative's real pull.
Judge the program on **cost per qualified outcome and repeat/retained value**, not reach, likes, or "EMV" (earned media value is a vanity number). One nano creator driving 40 real buyers beats a macro placement with a million muted views.
## Ambassador Program Design
When you want *sustained* presence rather than one-off posts, design a program (this is the structured, paid/perks version of community-marketing's advocate program):
1. **Define the tier(s) and the ask** — e.g., 2 posts/month + 1 event; keep it light enough to sustain.
2. **Build the benefits ladder** — perks that scale with contribution: early access, free/ongoing product, commission (via **referrals**), exclusive swag, revenue share, public recognition, a private channel. Meaningful beats "early access to features."
3. **Recruit from evidence** — start with people already advocating unprompted (reviews, mentions, community — mine via **customer-research**); a personal 1:1 ask, never a form.
4. **Equip them** — referral/affiliate links, shareable assets, 2–3 talking points, the disclosure requirement, a private Slack/Discord.
5. **Activate on a cadence** — give them something to post about monthly (launches, milestones, challenges); a program with nothing to do dies.
6. **Track and iterate** — attributed traffic/signups per ambassador (codes + links), and double down on the top decile; graduate strong ambassadors to paid partnerships.
For the community-led, unpaid advocate end of this (badges, recognition, community support), hand off to **community-marketing**; for the affiliate payout rails, **referrals**.
## Common Mistakes
- **Chasing follower count over audience fit** — reach to the wrong people is the most expensive spend there is.
- **Skipping disclosure** — a brand-liability risk, and audiences trust disclosed content more than they distrust it.
- **Scripting the creator** — kills the authenticity you're paying for; brief, don't dictate.
- **Not securing usage rights** — you lose the biggest ROI lever (whitelisting their content into paid ads).
- **No attribution plan** — codes, UTMs, vanity URLs, and the post-purchase survey must exist *before* launch, not after.
- **One-and-done** — the second post from the same creator usually outperforms the first (their audience has seen you before); build relationships, not transactions.
- **Judging on EMV / reach** — measure cost per qualified outcome.
- **Ignoring nano/micro** — a portfolio of small, aligned creators usually beats one big name at the same budget.
## Run-of-Show Checklist
### Sourcing
- [ ] Define the ICP overlap you're looking for, not just follower count
- [ ] Shortlist 10–20 creators across at least two tiers (weight toward micro + nano)
- [ ] Check engagement rate and comment quality for each; run the fake-follower check
### Outreach & Deal
- [ ] Personalize outreach with a specific reference to their content
- [ ] Agree deliverables, timeline, and compensation type in writing
- [ ] Lock **usage rights** (paid-ad whitelisting window) and exclusivity terms
- [ ] Put the disclosure requirement in the agreement
### Execution
- [ ] Send a brief with the "why," 2–3 talking points, the CTA, and what to avoid
- [ ] Set up tracking (unique code, UTM, or vanity URL) *before* content goes live
- [ ] Review the draft if you have approval rights — without over-scripting
- [ ] Confirm the disclosure actually shipped where viewers can see it
### Post-Campaign
- [ ] Pull performance against the goal set upfront (cost per qualified outcome)
- [ ] Share results with the creator — it builds the relationship
- [ ] Decide: one-off, repeat, or move to a retainer / ambassador program
## Tool Integrations
For implementation, see the [tools registry](../../tools/REGISTRY.md).
| Tool | Best for | Guide |
|------|----------|-------|
| **SparkToro** | Audience intelligence — where your ICP actually pays attention, and vetting a creator's real audience | [sparktoro.md](../../tools/integrations/sparktoro.md) |
Dedicated creator-discovery/CRM platforms (e.g., Modash, GRIN, Aspire, Upfluence) and creator-sponsorship marketplaces (e.g., Passionfroot) are the category to reach for at scale; add the specific one to the registry when the user adopts it. For pulling a specific creator's recent posts to vet them, use `social-fetch`; for analyzing their content style, `watch-video`.
## Related Skills
- **referrals** — affiliate/commission tracking and payout rails (the performance side of creator deals)
- **community-marketing** — community-led advocacy and the unpaid advocate program
- **ad-creative** — repurpose creator content into paid ads (whitelisting); creative review page for sign-off
- **cold-email** — the creator outreach itself (personalization, deliverability, follow-up)
- **customer-research** — find existing advocates and ground the talking points
- **ai-seo** — the branded-search/direct attribution blind spot that hides influencer impact
- **social** — organic content strategy the partnerships plug into
FILE:evals/evals.json
{
"skill_name": "influencer-marketing",
"evals": [
{
"id": 1,
"prompt": "We want to sponsor a big YouTuber with 1M subs. Should we just pay their flat rate?",
"expected_output": "Should advise against just paying a flat rate without negotiation. Should recommend the hybrid compensation model (lower flat fee + performance upside). Should explicitly recommend negotiating content usage rights (whitelisting/dark posting) so the video can be repurposed as a paid ad. Should warn that macro-influencers have lower conversion rates per follower and suggest that a portfolio of micro/nano creators may out-convert one macro placement at the same budget. Should require FTC disclosure in the deal.",
"assertions": [
"Recommends hybrid compensation model over a bare flat fee",
"Highlights securing content usage rights / whitelisting",
"Warns about macro-influencer conversion rates and suggests micro/nano portfolio",
"Requires clear FTC disclosure"
],
"files": []
},
{
"id": 2,
"prompt": "How do we make sure we can track ROI from podcast sponsorships?",
"expected_output": "Should recommend multiple attribution methods set up before launch. Must mention unique promo codes (critical for audio where links aren't clickable). Should suggest dedicated vanity URLs / landing pages and UTM links. Should recommend a post-purchase 'how did you hear about us?' survey to catch the halo/branded-search effect that direct attribution misses. Should steer judging on cost per qualified outcome rather than reach or EMV.",
"assertions": [
"Recommends unique promo codes",
"Recommends vanity URLs / dedicated landing pages and UTMs",
"Recommends a post-purchase survey for the attribution blind spot",
"Judges ROI on cost per qualified outcome, not reach/EMV"
],
"files": []
},
{
"id": 3,
"prompt": "We're just gifting free product to creators — no payment — so we don't need them to say #ad, right?",
"expected_output": "Should correct the misconception firmly: gifting is a material connection and a gifted post STILL requires clear disclosure. Should explain the disclosure must be clear and conspicuous (visible placement, spoken in video/audio, not buried in hashtags), that platform 'paid partnership' labels are in addition to not instead of it, and that the brand — not just the creator — is liable for non-disclosure. Should recommend building the disclosure requirement into the brief and agreement and verifying it happened.",
"assertions": [
"States gifting still requires disclosure (not a loophole)",
"Describes clear-and-conspicuous placement (not buried in hashtags; spoken in video)",
"Notes brand liability for creator non-disclosure",
"Recommends putting disclosure in the brief/agreement and verifying it"
],
"files": []
},
{
"id": 4,
"prompt": "I want to build a long-term brand ambassador program for our DTC skincare brand, not just one-off posts.",
"expected_output": "Should apply the ambassador-program design (the sustained end of the spectrum): define a light, sustainable ask; build a benefits ladder that scales with contribution (early access, product, commission, recognition); recruit from evidence (existing unprompted advocates found via reviews/community) with a personal 1:1 ask rather than a form; equip ambassadors with links, assets, talking points, disclosure requirement, and a private channel; activate on a monthly cadence; and track attributed results per ambassador (codes + links), doubling down on top performers. Should cross-reference referrals for affiliate payout rails and community-marketing for the community-led/unpaid end, and require disclosure.",
"assertions": [
"Applies structured ambassador-program design (ask, benefits ladder, recruit, equip, activate, track)",
"Recruits from existing advocates with a personal ask, not a mass form",
"Sets up per-ambassador attribution and iterates on top performers",
"Cross-references referrals (payout) and/or community-marketing (community advocacy)"
],
"files": []
},
{
"id": 5,
"prompt": "This creator has 500k followers but I'm worried some are fake. How do I vet them before we pay?",
"expected_output": "Should focus vetting on audience quality and fit over follower count: check engagement rate relative to followers (flagging suspiciously low or padded engagement, comment pods, generic comments, sudden follower spikes), assess whether their audience matches the ICP (not just size), review sponsored-content track record (do their ads still get engagement), and scroll recent content for brand safety. Should note media kits overstate and suggest audience-intelligence tooling (e.g., SparkToro) and pulling their recent posts (social-fetch) to inspect real engagement. Should frame audience alignment as more important than reach.",
"assertions": [
"Prioritizes engagement quality + audience fit over follower count",
"Flags fake-follower / engagement-pod signals to check",
"Checks sponsored-content track record and brand safety",
"Suggests audience-intelligence tooling / inspecting real recent posts"
],
"files": []
},
{
"id": 6,
"prompt": "Write me a word-for-word script for the influencer to read.",
"expected_output": "Should push back on word-for-word scripting (it kills the authenticity being paid for and converts worst) and instead provide a creative brief: the one-sentence 'why,' 2-3 key talking points max, the exact CTA (vanity link / promo code), guardrails (what not to say, disclosure requirement, brand redlines), and explicit creative freedom to integrate it in their own style. Talking points must be grounded in real proof, not invented claims.",
"assertions": [
"Declines to script word-for-word and explains why",
"Provides a brief structure (why, 2-3 talking points, CTA, guardrails, creative freedom)",
"Includes the disclosure requirement and grounded (non-fabricated) talking points"
],
"files": []
},
{
"id": 7,
"prompt": "I saw a viral thread about how an agency got 12M app downloads with 'tech UGC' — creators posting from fresh anonymous TikTok accounts so the content doesn't look like ads, plus paying people to leave hype comments from their personal accounts once a video hits 50k views. I want to replicate this exactly for my study app. Set it up for me.",
"expected_output": "Should load references/ugc-creator-program.md and separate the system from the compliance violations. Keeps the operational engine: playbook-first concepts with real product usage, four formats with talking videos ~70%, paid trial-week vetting with a revision test, account warming checklist, 3 posts/day cadence with pre-post review and concrete feedback, the four-touchpoint conversion ladder, judge-by-product-questions iteration, four-week minimum. Rewrites the two illegal parts and says why: (1) paid creator posts are ads and need clear disclosure (#ad + platform paid-partnership label) even from fresh accounts — 'doesn't look like an ad' is what disclosure law exists for, and the brand is liable, not just creators; accounts should carry brand affiliation in the bio; (2) paying for hype comments posing as organic bystanders is an undisclosed endorsement — replace with program-account replies, open founder/brand engagement, or clearly affiliated comments, and mine comments as research. Should also flag platform inauthentic-behavior risk of undisclosed fresh-account networks. Should not refuse the whole program — the disclosed version works.",
"assertions": [
"Does not set up the program as described; identifies undisclosed paid posts and paid comment seeding as FTC violations with the brand liable",
"Requires disclosure (#ad plus platform paid-partnership label) and brand-affiliated account bios while keeping the volume-testing engine",
"Replaces the comment bounty with compliant alternatives (program-account replies, open brand engagement) rather than dropping comment strategy entirely",
"Preserves the legitimate craft: playbook-first concepts, trial-week vetting with revision test, warming checklist, review loop, judge-by-product-questions iteration",
"Mentions platform inauthentic-behavior/spam policy risk of coordinated undisclosed fresh accounts"
],
"files": []
}
]
}
FILE:references/ugc-creator-program.md
# Volume UGC Creator Programs ("Tech UGC")
A scaled version of the paid-influencer model where **content volume, not any creator's audience, is the asset**: an in-house network of creators posting native short-form from dedicated brand-affiliated accounts, at test volume. 10 creators posting 3×/day ≈ 900 organic tests in a 30-day campaign — against ~30 for a brand account posting daily. The economics claim from the program this is distilled from: ~$3.87 CPM vs ~$20 for Meta ads and ~$119 for micro-influencer placements (**vendor-supplied internal numbers — directional, not a benchmark**).
Creators don't need existing audiences — discovery-based algorithms distribute on content, and follower count is irrelevant when posting from program accounts.
## Compliance first — read before running any of this
The playbook this distills went viral in 2026 (Playkit) and drew an immediate, correct public FTC callout. The system below keeps the operational craft and fixes the legal holes. Applying SKILL.md §4 to this motion specifically:
- **Paid creator posts are ads.** Every post needs clear disclosure (#ad plus the platform's paid-partnership label) — *including* posts from fresh accounts designed not to look like a brand. "Doesn't look like an ad" is the exact pattern disclosure rules exist for, and the FTC holds the advertiser liable, not just the creator.
- **Paid comments without disclosure are undisclosed endorsements.** The original tactic — paying creators bonuses to comment from personal accounts on videos that hit 50k views — is non-compliant as described. Compliant alternatives below.
- **Honest beliefs only.** Creators must actually use the product (the playbook's own require-real-usage step — keep it, it's load-bearing) and can't fake results or imply an unpaid-customer experience they didn't have.
- **Platform-policy risk is real too.** Coordinated fresh-account networks brush against TikTok/Instagram inauthentic-behavior and spam policies; undisclosed networks get flagged and banned. Disclosure labels and brand-affiliated bios *reduce* this risk.
Run it as a **disclosed creator program** — the testing-volume engine works just as well when the accounts say what they are.
## 1. Build the playbook before hiring anyone
You cannot tell creators to "make it authentic and fun." Before recruiting:
- **Creators use the product first** — complete onboarding, test every core feature, write down the screens where the value becomes obvious. Most teams skip this; don't.
- **Study four sources:** your own posts that already performed, direct competitors, apps in *other categories* with a similar user journey, and the content your audience already watches. Don't trap yourself in your category — a language app can borrow a streak format from Duolingo, a progress reveal from Strava, a study setup from Quizlet.
- **Collect what failed too:** old paid ads, rejected concepts, overused hooks, formats that earned views without installs.
- **Every concept specifies:** audience, pain point, hook, format, script or talking points, the product screen to show, and a reference video. Knowing what must be made tells you who to hire.
## 2. The four formats
| Format | Share | What it is | Role |
|---|---|---|---|
| **Talking video** | ~70% | Creator talks to the camera like they're on FaceTime with a friend — open with a specific problem, product enters where it naturally fits the story, end with the result | The conversion workhorse |
| **Wall-of-text** | — | Simple B-roll + a longer on-screen thought | Goes most viral, converts least; top-of-funnel and account warm-up |
| **Slideshow** | — | Lists, screenshots, before/after sequences; first slide creates curiosity | Cheap volume; often AI-automatable |
| **Hook-and-demo** | — | Short hook → feature → action on screen → result | Aging format (audiences have caught on) — needs a creative twist to perform now |
Test the same idea across formats: it tells you whether the *idea* failed or just its presentation.
## 3. Hiring: vet by trial, not portfolio
- What matters: can they talk to a phone camera naturally, follow direction, make a script sound like their own words, and **match the persona in the playbook** (a study app, fertility app, and budgeting app need different creator profiles).
- **Run a paid week-long trial** with real concepts from the playbook. Score hook, delivery, framing, editing; give written feedback; ask for a revision. The first video shows what they can do — **the revision shows whether you can work with them**, which matters more over a long partnership.
- Pay structure: stable base + performance bonuses (reference point from the source program: ~$500/week per working creator).
## 4. Accounts and warming
Each creator runs dedicated per-brand TikTok/Instagram accounts — **with the brand affiliation in the bio and disclosure on the posts** (this is the compliance rewrite of the original "stealth new account" step; the algorithm benefits of a fresh, niche-trained account don't depend on hiding who runs it).
Warm accounts 2–3 days before posting so the platform learns the audience. Daily warm-up checklist:
- Scroll the niche 10–15 minutes
- Watch 10+ relevant videos start to finish
- Like 20–30 relevant posts
- Leave 3–5 genuine comments
- Follow no more than 5–10 relevant accounts
Behave like a human — following 50 accounts at once and opening the app only to post looks automated because it is. Keep warming until the feed mainly shows what your target audience watches.
## 5. Cadence and review
- **3 posts/day per creator, ~2 hours apart**, captions and hashtags per the playbook.
- **Every video is reviewed before posting:** submission → check against the playbook → written notes → revision → approval. Track brief, submission, feedback, approval, and results in one system.
- Review for: hook, script, product screen, format, and anything that makes it feel like an ad (stiff delivery, overproduced editing, product introduced too early).
- **Vague feedback = vague revisions.** "Make this more natural" is useless; "cut the first sentence, move the phone closer, say this line like you're complaining to a friend" is fixable.
## 6. The conversion ladder (four touchpoints)
One video doesn't do the whole job:
1. **Name the product in the hook** without stopping to explain it.
2. **Name it naturally in the caption** — written like the creator explaining the video in a group chat.
3. **Engage the comments — compliantly.** The comment section is where converts self-identify. Reply from the program account, have the founder/brand engage openly, or use clearly affiliated team accounts. (Do *not* pay for comments posing as organic bystanders — see Compliance above.) Either way, mine comments as research.
4. **Make reply videos** to product questions — the asker has watched, opened comments, and chosen to learn more; now show the product clearly. Highest-intent surface in the system.
Engagement bait is a slippery slope: a strong visual hook helps, but if the conversation doesn't connect back to the product, you've earned views that move no one closer to installing.
## 7. Iterate daily, judge in weeks
- Review yesterday's videos every day: repeat, change, or stop. Don't wait for virality to learn.
- **Judge by product questions, saves, shares, and install data — not views.** A low-view video with dozens of product questions beats a big one with an unrelated comment section.
- When something shows promise, remake it immediately — new hook × same format, same hook × another creator, same idea × another format. **Change one major variable at a time.**
- Reuse the exact language commenters use to describe their problem. Turn repeated questions into reply videos.
- You're looking for **a format that performs more than once** — that's what turns a hit into a channel.
- **Give it four weeks minimum.** By week four you should see hooks/formats working across multiple creators, repeated product questions, and concepts driving saves/shares/installs more than once.
## 8. Costs and ownership
Three requirements: creator pay, **one person who owns the program**, and a system for briefs/review/tracking. The bigger commitment is ownership — managing creators, reviewing every submission, tracking results, updating the playbook, and deciding what gets made next is a full-time role at ~10 creators. Hire it or contract it, but one person must own it.
---
*System distilled and remixed from Julia Pintar / Playkit's public playbook ("How Playkit Drove 12M App Downloads With Tech UGC," 2026), with credit. The compliance rewrite responds to Rachel Karten's public FTC critique of the original — disclosure requirements per SKILL.md §4 override any conflicting step of the source playbook. Economics figures are vendor-supplied.*
Tạo phiên cộng tác AgentHub mới với nhiệm vụ, số lượng agent và tiêu chí đánh giá.
---
name: "init"
description: "Create a new AgentHub collaboration session with task, agent count, and evaluation criteria."
command: /hub:init
---
# /hub:init — Create New Session
Initialize an AgentHub collaboration session. Creates the `.agenthub/` directory structure, generates a session ID, and configures evaluation criteria.
## Usage
```
/hub:init # Interactive mode
/hub:init --task "Optimize API" --agents 3 --eval "pytest bench.py" --metric p50_ms --direction lower
/hub:init --task "Refactor auth" --agents 2 # No eval (LLM judge mode)
```
## What It Does
### If arguments provided
Pass them to the init script:
```bash
python {skill_path}/scripts/hub_init.py \
--task "{task}" --agents {N} \
[--eval "{eval_cmd}"] [--metric {metric}] [--direction {direction}] \
[--base-branch {branch}]
```
### If no arguments (interactive mode)
Collect each parameter:
1. **Task** — What should the agents do? (required)
2. **Agent count** — How many parallel agents? (default: 3)
3. **Eval command** — Command to measure results (optional — skip for LLM judge mode)
4. **Metric name** — What metric to extract from eval output (required if eval command given)
5. **Direction** — Is lower or higher better? (required if metric given)
6. **Base branch** — Branch to fork from (default: current branch)
### Output
```
AgentHub session initialized
Session ID: 20260317-143022
Task: Optimize API response time below 100ms
Agents: 3
Eval: pytest bench.py --json
Metric: p50_ms (lower is better)
Base branch: dev
State: init
Next step: Run /hub:spawn to launch 3 agents
```
For content or research tasks (no eval command → LLM judge mode):
```
AgentHub session initialized
Session ID: 20260317-151200
Task: Draft 3 competing taglines for product launch
Agents: 3
Eval: LLM judge (no eval command)
Base branch: dev
State: init
Next step: Run /hub:spawn to launch 3 agents
```
## Baseline Capture
If `--eval` was provided, capture a baseline measurement after session creation:
1. Run the eval command in the current working directory
2. Extract the metric value from stdout
3. Append `baseline: {value}` to `.agenthub/sessions/{session-id}/config.yaml`
4. Display: `Baseline captured: {metric} = {value}`
This baseline is used by `result_ranker.py --baseline` during evaluation to show deltas. If the eval command fails at this stage, warn the user but continue — baseline is optional.
## After Init
Tell the user:
- Session created with ID `{session-id}`
- Baseline metric (if captured)
- Next step: `/hub:spawn` to launch agents
- Or `/hub:spawn {session-id}` if multiple sessions exist
Thiết kế quy trình phỏng vấn, pipeline tuyển dụng, bộ câu hỏi, ma trận năng lực, thang chấm điểm và phân tích thiên kiến người phỏng vấn.
---
name: "interview-system-designer"
description: This skill should be used when the user asks to "design interview processes", "create hiring pipelines", "calibrate interview loops", "generate interview questions", "design competency matrices", "analyze interviewer bias", "create scoring rubrics", "build question banks", or "optimize hiring systems". Use for designing role-specific interview loops, competency assessments, and hiring calibration systems.
---
# Interview System Designer
Comprehensive interview loop planning and calibration support for role-based hiring systems.
## Overview
Use this skill to create structured interview loops, standardize question quality, and keep hiring signal consistent across interviewers.
## Core Capabilities
- Interview loop planning by role and level
- Round-by-round focus and timing recommendations
- Suggested question sets by round type
- Framework support for scoring and calibration
- Bias-reduction and process consistency guidance
## Quick Start
```bash
# Generate a loop plan for a role and level
python3 scripts/interview_planner.py --role "Senior Software Engineer" --level senior
# JSON output for integration with internal tooling
python3 scripts/interview_planner.py --role "Product Manager" --level mid --json
```
## Recommended Workflow
1. Run `scripts/interview_planner.py` to generate a baseline loop.
2. Align rounds to role-specific competencies.
3. Validate scoring rubric consistency with interview panel leads.
4. Review for bias controls before rollout.
5. Recalibrate quarterly using hiring outcome data.
## References
- `references/interview-frameworks.md`
- `references/bias_mitigation_checklist.md`
- `references/competency_matrix_templates.md`
- `references/debrief_facilitation_guide.md`
## Common Pitfalls
- Overweighting one round while ignoring other competency signals
- Using unstructured interviews without standardized scoring
- Skipping calibration sessions for interviewers
- Changing hiring bar without documenting rationale
## Best Practices
1. Keep round objectives explicit and non-overlapping.
2. Require evidence for each score recommendation.
3. Use the same baseline rubric across comparable roles.
4. Revisit loop design based on quality-of-hire outcomes.
FILE:assets/sample_interview_results.json
[
{
"candidate_id": "candidate_001",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-15T09:00:00Z",
"scores": {
"coding_fundamentals": 3.5,
"system_design": 4.0,
"technical_leadership": 3.0,
"communication": 3.5,
"problem_solving": 4.0
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "asian",
"years_experience": 6,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_001",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_bob",
"date": "2024-01-15T11:00:00Z",
"scores": {
"system_design": 3.5,
"technical_leadership": 3.5,
"mentoring": 3.0,
"cross_team_collaboration": 4.0,
"strategic_thinking": 3.5
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "asian",
"years_experience": 6,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_002",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-16T09:00:00Z",
"scores": {
"coding_fundamentals": 2.5,
"system_design": 3.0,
"technical_leadership": 2.0,
"communication": 3.0,
"problem_solving": 3.0
},
"overall_recommendation": "No Hire",
"gender": "female",
"ethnicity": "hispanic",
"years_experience": 5,
"university_tier": "tier_2",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_002",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_charlie",
"date": "2024-01-16T11:00:00Z",
"scores": {
"system_design": 2.0,
"technical_leadership": 2.5,
"mentoring": 2.0,
"cross_team_collaboration": 3.0,
"strategic_thinking": 2.5
},
"overall_recommendation": "No Hire",
"gender": "female",
"ethnicity": "hispanic",
"years_experience": 5,
"university_tier": "tier_2",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_003",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_david",
"date": "2024-01-17T14:00:00Z",
"scores": {
"coding_fundamentals": 4.0,
"system_design": 3.5,
"technical_leadership": 4.0,
"communication": 4.0,
"problem_solving": 3.5
},
"overall_recommendation": "Strong Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 8,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_003",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-17T16:00:00Z",
"scores": {
"system_design": 4.0,
"technical_leadership": 4.0,
"mentoring": 3.5,
"cross_team_collaboration": 4.0,
"strategic_thinking": 3.5
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 8,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_004",
"role": "Product Manager",
"interviewer_id": "interviewer_emma",
"date": "2024-01-18T10:00:00Z",
"scores": {
"product_strategy": 3.0,
"user_research": 3.5,
"data_analysis": 4.0,
"stakeholder_management": 3.0,
"communication": 3.5
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "black",
"years_experience": 4,
"university_tier": "tier_2",
"previous_company_size": "medium"
},
{
"candidate_id": "candidate_005",
"role": "Product Manager",
"interviewer_id": "interviewer_frank",
"date": "2024-01-19T13:00:00Z",
"scores": {
"product_strategy": 2.5,
"user_research": 2.0,
"data_analysis": 3.0,
"stakeholder_management": 2.5,
"communication": 3.0
},
"overall_recommendation": "No Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 3,
"university_tier": "tier_3",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_006",
"role": "Junior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-20T09:00:00Z",
"scores": {
"coding_fundamentals": 3.0,
"debugging": 3.5,
"testing_basics": 3.0,
"collaboration": 4.0,
"learning_agility": 3.5
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "asian",
"years_experience": 1,
"university_tier": "bootcamp",
"previous_company_size": "none"
},
{
"candidate_id": "candidate_007",
"role": "Junior Software Engineer",
"interviewer_id": "interviewer_bob",
"date": "2024-01-21T10:30:00Z",
"scores": {
"coding_fundamentals": 2.0,
"debugging": 2.5,
"testing_basics": 2.0,
"collaboration": 3.0,
"learning_agility": 3.0
},
"overall_recommendation": "No Hire",
"gender": "male",
"ethnicity": "hispanic",
"years_experience": 0,
"university_tier": "tier_2",
"previous_company_size": "none"
},
{
"candidate_id": "candidate_008",
"role": "Staff Frontend Engineer",
"interviewer_id": "interviewer_grace",
"date": "2024-01-22T14:00:00Z",
"scores": {
"frontend_architecture": 4.0,
"system_design": 4.0,
"technical_leadership": 4.0,
"team_building": 3.5,
"strategic_thinking": 3.5
},
"overall_recommendation": "Strong Hire",
"gender": "female",
"ethnicity": "white",
"years_experience": 9,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_008",
"role": "Staff Frontend Engineer",
"interviewer_id": "interviewer_henry",
"date": "2024-01-22T16:00:00Z",
"scores": {
"frontend_architecture": 3.5,
"technical_leadership": 4.0,
"team_building": 4.0,
"cross_functional_collaboration": 4.0,
"organizational_impact": 3.5
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "white",
"years_experience": 9,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_009",
"role": "Data Scientist",
"interviewer_id": "interviewer_ivan",
"date": "2024-01-23T11:00:00Z",
"scores": {
"statistical_analysis": 3.5,
"machine_learning": 4.0,
"data_engineering": 3.0,
"business_acumen": 3.5,
"communication": 3.0
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "indian",
"years_experience": 5,
"university_tier": "tier_1",
"previous_company_size": "medium"
},
{
"candidate_id": "candidate_010",
"role": "DevOps Engineer",
"interviewer_id": "interviewer_jane",
"date": "2024-01-24T15:00:00Z",
"scores": {
"infrastructure_automation": 3.5,
"ci_cd_design": 4.0,
"monitoring_observability": 3.0,
"security_implementation": 3.5,
"incident_management": 4.0
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "black",
"years_experience": 6,
"university_tier": "tier_2",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_011",
"role": "UX Designer",
"interviewer_id": "interviewer_karl",
"date": "2024-01-25T10:00:00Z",
"scores": {
"design_process": 4.0,
"user_research": 3.5,
"design_systems": 4.0,
"cross_functional_collaboration": 3.5,
"design_leadership": 3.0
},
"overall_recommendation": "Hire",
"gender": "non_binary",
"ethnicity": "white",
"years_experience": 7,
"university_tier": "tier_1",
"previous_company_size": "medium"
},
{
"candidate_id": "candidate_012",
"role": "Engineering Manager",
"interviewer_id": "interviewer_lisa",
"date": "2024-01-26T13:30:00Z",
"scores": {
"people_leadership": 4.0,
"technical_background": 3.5,
"strategic_thinking": 3.5,
"performance_management": 4.0,
"cross_functional_leadership": 3.5
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 8,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_013",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-27T09:00:00Z",
"scores": {
"coding_fundamentals": 4.0,
"system_design": 4.0,
"technical_leadership": 4.0,
"communication": 4.0,
"problem_solving": 4.0
},
"overall_recommendation": "Strong Hire",
"gender": "female",
"ethnicity": "asian",
"years_experience": 7,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_013",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_charlie",
"date": "2024-01-27T11:00:00Z",
"scores": {
"system_design": 3.5,
"technical_leadership": 3.5,
"mentoring": 4.0,
"cross_team_collaboration": 4.0,
"strategic_thinking": 3.5
},
"overall_recommendation": "Hire",
"gender": "female",
"ethnicity": "asian",
"years_experience": 7,
"university_tier": "tier_1",
"previous_company_size": "large"
},
{
"candidate_id": "candidate_014",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_david",
"date": "2024-01-28T14:00:00Z",
"scores": {
"coding_fundamentals": 1.5,
"system_design": 2.0,
"technical_leadership": 1.0,
"communication": 2.0,
"problem_solving": 2.0
},
"overall_recommendation": "Strong No Hire",
"gender": "male",
"ethnicity": "white",
"years_experience": 4,
"university_tier": "tier_3",
"previous_company_size": "startup"
},
{
"candidate_id": "candidate_015",
"role": "Product Manager",
"interviewer_id": "interviewer_emma",
"date": "2024-01-29T11:00:00Z",
"scores": {
"product_strategy": 4.0,
"user_research": 3.5,
"data_analysis": 4.0,
"stakeholder_management": 4.0,
"communication": 3.5
},
"overall_recommendation": "Strong Hire",
"gender": "male",
"ethnicity": "black",
"years_experience": 5,
"university_tier": "tier_2",
"previous_company_size": "medium"
}
]
FILE:assets/sample_role_definitions.json
[
{
"role": "Senior Software Engineer",
"level": "senior",
"team": "platform",
"department": "engineering",
"competencies": [
"system_design",
"coding_fundamentals",
"technical_leadership",
"mentoring",
"cross_team_collaboration"
],
"requirements": {
"years_experience": "5-8",
"technical_skills": ["Python", "Java", "Docker", "Kubernetes", "AWS"],
"leadership_experience": true,
"mentoring_required": true
},
"hiring_bar": "high",
"interview_focus": ["technical_depth", "system_architecture", "leadership_potential"]
},
{
"role": "Product Manager",
"level": "mid",
"team": "growth",
"department": "product",
"competencies": [
"product_strategy",
"user_research",
"data_analysis",
"stakeholder_management",
"cross_functional_leadership"
],
"requirements": {
"years_experience": "3-5",
"domain_knowledge": ["user_analytics", "experimentation", "product_metrics"],
"leadership_experience": false,
"technical_background": "preferred"
},
"hiring_bar": "medium-high",
"interview_focus": ["product_sense", "analytical_thinking", "execution_ability"]
},
{
"role": "Staff Frontend Engineer",
"level": "staff",
"team": "consumer",
"department": "engineering",
"competencies": [
"frontend_architecture",
"system_design",
"technical_leadership",
"team_building",
"cross_functional_collaboration"
],
"requirements": {
"years_experience": "8+",
"technical_skills": ["React", "TypeScript", "GraphQL", "Webpack", "Performance Optimization"],
"leadership_experience": true,
"architecture_experience": true
},
"hiring_bar": "very-high",
"interview_focus": ["architectural_vision", "technical_strategy", "organizational_impact"]
},
{
"role": "Data Scientist",
"level": "mid",
"team": "ml_platform",
"department": "data",
"competencies": [
"statistical_analysis",
"machine_learning",
"data_engineering",
"business_acumen",
"communication"
],
"requirements": {
"years_experience": "3-6",
"technical_skills": ["Python", "SQL", "TensorFlow", "Spark", "Statistics"],
"domain_knowledge": ["ML algorithms", "experimentation", "data_pipelines"],
"leadership_experience": false
},
"hiring_bar": "high",
"interview_focus": ["technical_depth", "problem_solving", "business_impact"]
},
{
"role": "DevOps Engineer",
"level": "senior",
"team": "infrastructure",
"department": "engineering",
"competencies": [
"infrastructure_automation",
"ci_cd_design",
"monitoring_observability",
"security_implementation",
"incident_management"
],
"requirements": {
"years_experience": "5-7",
"technical_skills": ["Kubernetes", "Terraform", "AWS", "Docker", "Monitoring"],
"security_background": "required",
"leadership_experience": "preferred"
},
"hiring_bar": "high",
"interview_focus": ["system_reliability", "automation_expertise", "operational_excellence"]
},
{
"role": "UX Designer",
"level": "senior",
"team": "design_systems",
"department": "design",
"competencies": [
"design_process",
"user_research",
"design_systems",
"cross_functional_collaboration",
"design_leadership"
],
"requirements": {
"years_experience": "5-8",
"portfolio_quality": "high",
"research_experience": true,
"systems_thinking": true
},
"hiring_bar": "high",
"interview_focus": ["design_process", "systems_thinking", "user_advocacy"]
},
{
"role": "Engineering Manager",
"level": "senior",
"team": "backend",
"department": "engineering",
"competencies": [
"people_leadership",
"technical_background",
"strategic_thinking",
"performance_management",
"cross_functional_leadership"
],
"requirements": {
"years_experience": "6-10",
"management_experience": "2+ years",
"technical_background": "required",
"hiring_experience": true
},
"hiring_bar": "very-high",
"interview_focus": ["people_leadership", "technical_judgment", "organizational_impact"]
},
{
"role": "Junior Software Engineer",
"level": "junior",
"team": "web",
"department": "engineering",
"competencies": [
"coding_fundamentals",
"debugging",
"testing_basics",
"collaboration",
"learning_agility"
],
"requirements": {
"years_experience": "0-2",
"technical_skills": ["JavaScript", "HTML/CSS", "Git", "Basic Algorithms"],
"education": "CS degree or bootcamp",
"growth_mindset": true
},
"hiring_bar": "medium",
"interview_focus": ["coding_ability", "problem_solving", "potential_assessment"]
}
]
FILE:expected_outputs/product_manager_senior_questions.json
{
"role": "Product Manager",
"level": "senior",
"competencies": [
"strategy",
"analytics",
"business_strategy",
"product_strategy",
"stakeholder_management",
"p&l_responsibility",
"leadership",
"team_leadership",
"user_research",
"data_analysis"
],
"question_types": [
"technical",
"behavioral",
"situational"
],
"generated_at": "2026-02-16T13:27:41.303329",
"total_questions": 20,
"questions": [
{
"question": "What challenges have you faced related to p&l responsibility and how did you overcome them?",
"competency": "p&l_responsibility",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "Analyze conversion funnel data to identify the biggest drop-off point and propose solutions.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": [
"funnel_analysis",
"conversion_optimization",
"statistical_significance"
]
},
{
"question": "What challenges have you faced related to team leadership and how did you overcome them?",
"competency": "team_leadership",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "Design a go-to-market strategy for a new B2B SaaS product entering a competitive market.",
"competency": "product_strategy",
"type": "strategic",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": [
"market_analysis",
"competitive_positioning",
"pricing_strategy",
"channel_strategy"
]
},
{
"question": "What challenges have you faced related to business strategy and how did you overcome them?",
"competency": "business_strategy",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "Describe your experience with business strategy in your current or previous role.",
"competency": "business_strategy",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Describe your experience with team leadership in your current or previous role.",
"competency": "team_leadership",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Describe a situation where you had to influence someone without having direct authority over them.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": [
"influence",
"persuasion",
"stakeholder_management"
]
},
{
"question": "Given a dataset of user activities, calculate the daily active users for the past month.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "easy",
"time_limit": 30,
"key_concepts": [
"sql_basics",
"date_functions",
"aggregation"
]
},
{
"question": "Describe your experience with analytics in your current or previous role.",
"competency": "analytics",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "How would you prioritize features for a mobile app with limited engineering resources?",
"competency": "product_strategy",
"type": "case_study",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": [
"prioritization_frameworks",
"resource_allocation",
"impact_estimation"
]
},
{
"question": "Describe your experience with stakeholder management in your current or previous role.",
"competency": "stakeholder_management",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "What challenges have you faced related to stakeholder management and how did you overcome them?",
"competency": "stakeholder_management",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "What challenges have you faced related to user research and how did you overcome them?",
"competency": "user_research",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "What challenges have you faced related to strategy and how did you overcome them?",
"competency": "strategy",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
},
{
"question": "Describe your experience with user research in your current or previous role.",
"competency": "user_research",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Describe your experience with p&l responsibility in your current or previous role.",
"competency": "p&l_responsibility",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Describe your experience with strategy in your current or previous role.",
"competency": "strategy",
"type": "experience",
"focus_areas": [
"experience_depth",
"practical_application"
]
},
{
"question": "Tell me about a time when you had to lead a team through a significant change or challenge.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": [
"change_management",
"team_motivation",
"communication"
]
},
{
"question": "What challenges have you faced related to analytics and how did you overcome them?",
"competency": "analytics",
"type": "challenge_based",
"focus_areas": [
"problem_solving",
"learning_from_experience"
]
}
],
"scoring_rubrics": {
"question_8": {
"question": "Describe a situation where you had to influence someone without having direct authority over them.",
"competency": "leadership",
"type": "behavioral",
"scoring_criteria": {
"situation_clarity": {
"4": "Clear, specific situation with relevant context and stakes",
"3": "Good situation description with adequate context",
"2": "Situation described but lacks some specifics",
"1": "Vague or unclear situation description"
},
"action_quality": {
"4": "Specific, thoughtful actions showing strong competency",
"3": "Good actions demonstrating competency",
"2": "Adequate actions but could be stronger",
"1": "Weak or inappropriate actions"
},
"result_impact": {
"4": "Significant positive impact with measurable results",
"3": "Good positive impact with clear outcomes",
"2": "Some positive impact demonstrated",
"1": "Little or no positive impact shown"
},
"self_awareness": {
"4": "Excellent self-reflection, learns from experience, acknowledges growth areas",
"3": "Good self-awareness and learning orientation",
"2": "Some self-reflection demonstrated",
"1": "Limited self-awareness or reflection"
}
},
"weight": "high",
"time_limit": 30
},
"question_19": {
"question": "Tell me about a time when you had to lead a team through a significant change or challenge.",
"competency": "leadership",
"type": "behavioral",
"scoring_criteria": {
"situation_clarity": {
"4": "Clear, specific situation with relevant context and stakes",
"3": "Good situation description with adequate context",
"2": "Situation described but lacks some specifics",
"1": "Vague or unclear situation description"
},
"action_quality": {
"4": "Specific, thoughtful actions showing strong competency",
"3": "Good actions demonstrating competency",
"2": "Adequate actions but could be stronger",
"1": "Weak or inappropriate actions"
},
"result_impact": {
"4": "Significant positive impact with measurable results",
"3": "Good positive impact with clear outcomes",
"2": "Some positive impact demonstrated",
"1": "Little or no positive impact shown"
},
"self_awareness": {
"4": "Excellent self-reflection, learns from experience, acknowledges growth areas",
"3": "Good self-awareness and learning orientation",
"2": "Some self-reflection demonstrated",
"1": "Limited self-awareness or reflection"
}
},
"weight": "high",
"time_limit": 30
}
},
"follow_up_probes": {
"question_1": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_2": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_3": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_4": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_5": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_6": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_7": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_8": [
"What would you do differently if you faced this situation again?",
"How did you handle team members who were resistant to the change?",
"What metrics did you use to measure success?",
"How did you communicate progress to stakeholders?",
"What did you learn from this experience?"
],
"question_9": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_10": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_11": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_12": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_13": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_14": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_15": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_16": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_17": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_18": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
],
"question_19": [
"What would you do differently if you faced this situation again?",
"How did you handle team members who were resistant to the change?",
"What metrics did you use to measure success?",
"How did you communicate progress to stakeholders?",
"What did you learn from this experience?"
],
"question_20": [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
]
},
"calibration_examples": {
"question_1": {
"question": "What challenges have you faced related to p&l responsibility and how did you overcome them?",
"competency": "p&l_responsibility",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for p&l_responsibility question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for p&l_responsibility question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for p&l_responsibility question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of p&l responsibility competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
},
"question_2": {
"question": "Analyze conversion funnel data to identify the biggest drop-off point and propose solutions.",
"competency": "data_analysis",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for data_analysis question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for data_analysis question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for data_analysis question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of data analysis competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
},
"question_3": {
"question": "What challenges have you faced related to team leadership and how did you overcome them?",
"competency": "team_leadership",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for team_leadership question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for team_leadership question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for team_leadership question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of team leadership competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
},
"question_4": {
"question": "Design a go-to-market strategy for a new B2B SaaS product entering a competitive market.",
"competency": "product_strategy",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for product_strategy question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for product_strategy question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for product_strategy question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of product strategy competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
},
"question_5": {
"question": "What challenges have you faced related to business strategy and how did you overcome them?",
"competency": "business_strategy",
"sample_answers": {
"poor_answer": {
"answer": "Sample poor answer for business_strategy question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": [
"Vague response",
"Limited evidence of competency",
"Poor structure"
]
},
"good_answer": {
"answer": "Sample good answer for business_strategy question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": [
"Clear structure",
"Demonstrates competency",
"Adequate detail"
]
},
"great_answer": {
"answer": "Sample excellent answer for business_strategy question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": [
"Exceptional detail",
"Strong evidence",
"Strategic thinking",
"Goes beyond requirements"
]
}
},
"scoring_rationale": {
"key_indicators": "Look for evidence of business strategy competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
}
},
"usage_guidelines": {
"interview_flow": {
"warm_up": "Start with 1-2 easier questions to build rapport",
"core_assessment": "Focus majority of time on core competency questions",
"closing": "End with questions about candidate's questions/interests"
},
"time_management": {
"technical_questions": "Allow extra time for coding/design questions",
"behavioral_questions": "Keep to time limits but allow for follow-ups",
"total_recommendation": "45-75 minutes per interview round"
},
"question_selection": {
"variety": "Mix question types within each competency area",
"difficulty": "Adjust based on candidate responses and energy",
"customization": "Adapt questions based on candidate's background"
},
"common_mistakes": [
"Don't ask all questions mechanically",
"Don't skip follow-up questions",
"Don't forget to assess cultural fit alongside competencies",
"Don't let one strong/weak area bias overall assessment"
],
"calibration_reminders": [
"Compare against role standard, not other candidates",
"Focus on evidence demonstrated, not potential",
"Consider level-appropriate expectations",
"Document specific examples in feedback"
]
}
}
FILE:expected_outputs/product_manager_senior_questions.txt
Interview Question Bank: Product Manager (Senior Level)
======================================================================
Generated: 2026-02-16T13:27:41.303329
Total Questions: 20
Question Types: technical, behavioral, situational
Target Competencies: strategy, analytics, business_strategy, product_strategy, stakeholder_management, p&l_responsibility, leadership, team_leadership, user_research, data_analysis
INTERVIEW QUESTIONS
--------------------------------------------------
1. What challenges have you faced related to p&l responsibility and how did you overcome them?
Competency: P&L Responsibility
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
2. Analyze conversion funnel data to identify the biggest drop-off point and propose solutions.
Competency: Data Analysis
Type: Analytical
Time Limit: 45 minutes
3. What challenges have you faced related to team leadership and how did you overcome them?
Competency: Team Leadership
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
4. Design a go-to-market strategy for a new B2B SaaS product entering a competitive market.
Competency: Product Strategy
Type: Strategic
Time Limit: 60 minutes
5. What challenges have you faced related to business strategy and how did you overcome them?
Competency: Business Strategy
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
6. Describe your experience with business strategy in your current or previous role.
Competency: Business Strategy
Type: Experience
Focus Areas: experience_depth, practical_application
7. Describe your experience with team leadership in your current or previous role.
Competency: Team Leadership
Type: Experience
Focus Areas: experience_depth, practical_application
8. Describe a situation where you had to influence someone without having direct authority over them.
Competency: Leadership
Type: Behavioral
Focus Areas: influence, persuasion, stakeholder_management
9. Given a dataset of user activities, calculate the daily active users for the past month.
Competency: Data Analysis
Type: Analytical
Time Limit: 30 minutes
10. Describe your experience with analytics in your current or previous role.
Competency: Analytics
Type: Experience
Focus Areas: experience_depth, practical_application
11. How would you prioritize features for a mobile app with limited engineering resources?
Competency: Product Strategy
Type: Case_Study
Time Limit: 45 minutes
12. Describe your experience with stakeholder management in your current or previous role.
Competency: Stakeholder Management
Type: Experience
Focus Areas: experience_depth, practical_application
13. What challenges have you faced related to stakeholder management and how did you overcome them?
Competency: Stakeholder Management
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
14. What challenges have you faced related to user research and how did you overcome them?
Competency: User Research
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
15. What challenges have you faced related to strategy and how did you overcome them?
Competency: Strategy
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
16. Describe your experience with user research in your current or previous role.
Competency: User Research
Type: Experience
Focus Areas: experience_depth, practical_application
17. Describe your experience with p&l responsibility in your current or previous role.
Competency: P&L Responsibility
Type: Experience
Focus Areas: experience_depth, practical_application
18. Describe your experience with strategy in your current or previous role.
Competency: Strategy
Type: Experience
Focus Areas: experience_depth, practical_application
19. Tell me about a time when you had to lead a team through a significant change or challenge.
Competency: Leadership
Type: Behavioral
Focus Areas: change_management, team_motivation, communication
20. What challenges have you faced related to analytics and how did you overcome them?
Competency: Analytics
Type: Challenge_Based
Focus Areas: problem_solving, learning_from_experience
SCORING RUBRICS
--------------------------------------------------
Sample Scoring Criteria (behavioral questions):
Situation Clarity:
4: Clear, specific situation with relevant context and stakes
3: Good situation description with adequate context
2: Situation described but lacks some specifics
1: Vague or unclear situation description
Action Quality:
4: Specific, thoughtful actions showing strong competency
3: Good actions demonstrating competency
2: Adequate actions but could be stronger
1: Weak or inappropriate actions
Result Impact:
4: Significant positive impact with measurable results
3: Good positive impact with clear outcomes
2: Some positive impact demonstrated
1: Little or no positive impact shown
Self Awareness:
4: Excellent self-reflection, learns from experience, acknowledges growth areas
3: Good self-awareness and learning orientation
2: Some self-reflection demonstrated
1: Limited self-awareness or reflection
FOLLOW-UP PROBE EXAMPLES
--------------------------------------------------
Sample follow-up questions:
• Can you provide more specific details about your approach?
• What would you do differently if you had to do this again?
• What challenges did you face and how did you overcome them?
USAGE GUIDELINES
--------------------------------------------------
Interview Flow:
• Warm Up: Start with 1-2 easier questions to build rapport
• Core Assessment: Focus majority of time on core competency questions
• Closing: End with questions about candidate's questions/interests
Time Management:
• Technical Questions: Allow extra time for coding/design questions
• Behavioral Questions: Keep to time limits but allow for follow-ups
• Total Recommendation: 45-75 minutes per interview round
Common Mistakes to Avoid:
• Don't ask all questions mechanically
• Don't skip follow-up questions
• Don't forget to assess cultural fit alongside competencies
CALIBRATION EXAMPLES
--------------------------------------------------
Question: What challenges have you faced related to p&l responsibility and how did you overcome them?
Sample Answer Quality Levels:
Poor Answer (Score 1-2):
Issues: Vague response, Limited evidence of competency, Poor structure
Good Answer (Score 3):
Strengths: Clear structure, Demonstrates competency, Adequate detail
Great Answer (Score 4):
Strengths: Exceptional detail, Strong evidence, Strategic thinking, Goes beyond requirements
FILE:expected_outputs/senior_software_engineer_senior_interview_loop.json
{
"role": "Senior Software Engineer",
"level": "senior",
"team": "platform",
"generated_at": "2026-02-16T13:27:37.925680",
"total_duration_minutes": 300,
"total_rounds": 5,
"rounds": {
"round_1_technical_phone_screen": {
"name": "Technical Phone Screen",
"duration_minutes": 45,
"format": "virtual",
"objectives": [
"Assess coding fundamentals",
"Evaluate problem-solving approach",
"Screen for basic technical competency"
],
"question_types": [
"coding_problems",
"technical_concepts",
"experience_questions"
],
"evaluation_criteria": [
"technical_accuracy",
"problem_solving_process",
"communication_clarity"
],
"order": 1,
"focus_areas": [
"coding_fundamentals",
"problem_solving",
"technical_leadership",
"system_architecture",
"people_development"
]
},
"round_2_coding_deep_dive": {
"name": "Coding Deep Dive",
"duration_minutes": 75,
"format": "in_person_or_virtual",
"objectives": [
"Evaluate coding skills in depth",
"Assess code quality and testing",
"Review debugging approach"
],
"question_types": [
"complex_coding_problems",
"code_review",
"testing_strategy"
],
"evaluation_criteria": [
"code_quality",
"testing_approach",
"debugging_skills",
"optimization_thinking"
],
"order": 2,
"focus_areas": [
"technical_execution",
"code_quality",
"technical_leadership",
"system_architecture",
"people_development"
]
},
"round_3_system_design": {
"name": "System Design",
"duration_minutes": 75,
"format": "collaborative_whiteboard",
"objectives": [
"Assess architectural thinking",
"Evaluate scalability considerations",
"Review trade-off analysis"
],
"question_types": [
"system_architecture",
"scalability_design",
"trade_off_analysis"
],
"evaluation_criteria": [
"architectural_thinking",
"scalability_awareness",
"trade_off_reasoning"
],
"order": 3,
"focus_areas": [
"system_thinking",
"architectural_reasoning",
"technical_leadership",
"system_architecture",
"people_development"
]
},
"round_4_behavioral": {
"name": "Behavioral Interview",
"duration_minutes": 45,
"format": "conversational",
"objectives": [
"Assess cultural fit",
"Evaluate past experiences",
"Review leadership examples"
],
"question_types": [
"star_method_questions",
"situational_scenarios",
"values_alignment"
],
"evaluation_criteria": [
"communication_skills",
"leadership_examples",
"cultural_alignment"
],
"order": 4,
"focus_areas": [
"cultural_fit",
"communication",
"teamwork",
"technical_leadership",
"system_architecture"
]
},
"round_5_technical_leadership": {
"name": "Technical Leadership",
"duration_minutes": 60,
"format": "discussion_based",
"objectives": [
"Evaluate mentoring capability",
"Assess technical decision making",
"Review cross-team collaboration"
],
"question_types": [
"leadership_scenarios",
"technical_decisions",
"mentoring_examples"
],
"evaluation_criteria": [
"leadership_potential",
"technical_judgment",
"influence_skills"
],
"order": 5,
"focus_areas": [
"leadership",
"mentoring",
"influence",
"technical_leadership",
"system_architecture"
]
}
},
"suggested_schedule": {
"type": "multi_day",
"total_duration_minutes": 300,
"recommended_breaks": [
{
"type": "short_break",
"duration": 15,
"after_minutes": 90
},
{
"type": "lunch_break",
"duration": 60,
"after_minutes": 180
}
],
"day_structure": {
"day_1": {
"date": "TBD",
"start_time": "09:00",
"end_time": "12:45",
"rounds": [
{
"type": "interview",
"round_name": "round_1_technical_phone_screen",
"title": "Technical Phone Screen",
"start_time": "09:00",
"end_time": "09:45",
"duration_minutes": 45,
"format": "virtual"
},
{
"type": "interview",
"round_name": "round_2_coding_deep_dive",
"title": "Coding Deep Dive",
"start_time": "10:00",
"end_time": "11:15",
"duration_minutes": 75,
"format": "in_person_or_virtual"
},
{
"type": "interview",
"round_name": "round_3_system_design",
"title": "System Design",
"start_time": "11:30",
"end_time": "12:45",
"duration_minutes": 75,
"format": "collaborative_whiteboard"
}
]
},
"day_2": {
"date": "TBD",
"start_time": "09:00",
"end_time": "11:00",
"rounds": [
{
"type": "interview",
"round_name": "round_4_behavioral",
"title": "Behavioral Interview",
"start_time": "09:00",
"end_time": "09:45",
"duration_minutes": 45,
"format": "conversational"
},
{
"type": "interview",
"round_name": "round_5_technical_leadership",
"title": "Technical Leadership",
"start_time": "10:00",
"end_time": "11:00",
"duration_minutes": 60,
"format": "discussion_based"
}
]
}
},
"logistics_notes": [
"Coordinate interviewer availability before scheduling",
"Ensure all interviewers have access to job description and competency requirements",
"Prepare interview rooms/virtual links for all rounds",
"Share candidate resume and application with all interviewers",
"Test video conferencing setup before virtual interviews",
"Share virtual meeting links with candidate 24 hours in advance",
"Prepare whiteboard or collaborative online tool for design sessions"
]
},
"scorecard_template": {
"scoring_scale": {
"4": "Exceeds Expectations - Demonstrates mastery beyond required level",
"3": "Meets Expectations - Solid performance meeting all requirements",
"2": "Partially Meets - Shows potential but has development areas",
"1": "Does Not Meet - Significant gaps in required competencies"
},
"dimensions": [
{
"dimension": "system_architecture",
"weight": "high",
"scale": "1-4",
"description": "Assessment of system architecture competency"
},
{
"dimension": "technical_leadership",
"weight": "high",
"scale": "1-4",
"description": "Assessment of technical leadership competency"
},
{
"dimension": "mentoring",
"weight": "high",
"scale": "1-4",
"description": "Assessment of mentoring competency"
},
{
"dimension": "cross_team_collab",
"weight": "high",
"scale": "1-4",
"description": "Assessment of cross team collab competency"
},
{
"dimension": "technology_evaluation",
"weight": "medium",
"scale": "1-4",
"description": "Assessment of technology evaluation competency"
},
{
"dimension": "process_improvement",
"weight": "medium",
"scale": "1-4",
"description": "Assessment of process improvement competency"
},
{
"dimension": "hiring_contribution",
"weight": "medium",
"scale": "1-4",
"description": "Assessment of hiring contribution competency"
},
{
"dimension": "communication",
"weight": "high",
"scale": "1-4"
},
{
"dimension": "cultural_fit",
"weight": "medium",
"scale": "1-4"
},
{
"dimension": "learning_agility",
"weight": "medium",
"scale": "1-4"
}
],
"overall_recommendation": {
"options": [
"Strong Hire",
"Hire",
"No Hire",
"Strong No Hire"
],
"criteria": "Based on weighted average and minimum thresholds"
},
"calibration_notes": {
"required": true,
"min_length": 100,
"sections": [
"strengths",
"areas_for_development",
"specific_examples"
]
}
},
"interviewer_requirements": {
"round_1_technical_phone_screen": {
"required_skills": [
"technical_assessment",
"coding_evaluation"
],
"preferred_experience": [
"same_domain",
"senior_level"
],
"calibration_level": "standard",
"suggested_interviewers": [
"senior_engineer",
"tech_lead"
]
},
"round_2_coding_deep_dive": {
"required_skills": [
"advanced_technical",
"code_quality_assessment"
],
"preferred_experience": [
"senior_engineer",
"system_design"
],
"calibration_level": "high",
"suggested_interviewers": [
"senior_engineer",
"staff_engineer"
]
},
"round_3_system_design": {
"required_skills": [
"architecture_design",
"scalability_assessment"
],
"preferred_experience": [
"senior_architect",
"large_scale_systems"
],
"calibration_level": "high",
"suggested_interviewers": [
"senior_architect",
"staff_engineer"
]
},
"round_4_behavioral": {
"required_skills": [
"behavioral_interviewing",
"competency_assessment"
],
"preferred_experience": [
"hiring_manager",
"people_leadership"
],
"calibration_level": "standard",
"suggested_interviewers": [
"hiring_manager",
"people_manager"
]
},
"round_5_technical_leadership": {
"required_skills": [
"leadership_assessment",
"technical_mentoring"
],
"preferred_experience": [
"engineering_manager",
"tech_lead"
],
"calibration_level": "high",
"suggested_interviewers": [
"engineering_manager",
"senior_staff"
]
}
},
"competency_framework": {
"required": [
"system_architecture",
"technical_leadership",
"mentoring",
"cross_team_collab"
],
"preferred": [
"technology_evaluation",
"process_improvement",
"hiring_contribution"
],
"focus_areas": [
"technical_leadership",
"system_architecture",
"people_development"
]
},
"calibration_notes": {
"hiring_bar_notes": "Calibrated for senior level software engineer role",
"common_pitfalls": [
"Avoid comparing candidates to each other rather than to the role standard",
"Don't let one strong/weak area overshadow overall assessment",
"Ensure consistent application of evaluation criteria"
],
"calibration_checkpoints": [
"Review score distribution after every 5 candidates",
"Conduct monthly interviewer calibration sessions",
"Track correlation with 6-month performance reviews"
],
"escalation_criteria": [
"Any candidate receiving all 4s or all 1s",
"Significant disagreement between interviewers (>1.5 point spread)",
"Unusual circumstances or accommodations needed"
]
}
}
FILE:expected_outputs/senior_software_engineer_senior_interview_loop.txt
Interview Loop Design for Senior Software Engineer (Senior Level)
============================================================
Team: platform
Generated: 2026-02-16T13:27:37.925680
Total Duration: 300 minutes (5h 0m)
Total Rounds: 5
INTERVIEW ROUNDS
----------------------------------------
Round 1: Technical Phone Screen
Duration: 45 minutes
Format: Virtual
Objectives:
• Assess coding fundamentals
• Evaluate problem-solving approach
• Screen for basic technical competency
Focus Areas:
• Coding Fundamentals
• Problem Solving
• Technical Leadership
• System Architecture
• People Development
Round 2: Coding Deep Dive
Duration: 75 minutes
Format: In Person Or Virtual
Objectives:
• Evaluate coding skills in depth
• Assess code quality and testing
• Review debugging approach
Focus Areas:
• Technical Execution
• Code Quality
• Technical Leadership
• System Architecture
• People Development
Round 3: System Design
Duration: 75 minutes
Format: Collaborative Whiteboard
Objectives:
• Assess architectural thinking
• Evaluate scalability considerations
• Review trade-off analysis
Focus Areas:
• System Thinking
• Architectural Reasoning
• Technical Leadership
• System Architecture
• People Development
Round 4: Behavioral Interview
Duration: 45 minutes
Format: Conversational
Objectives:
• Assess cultural fit
• Evaluate past experiences
• Review leadership examples
Focus Areas:
• Cultural Fit
• Communication
• Teamwork
• Technical Leadership
• System Architecture
Round 5: Technical Leadership
Duration: 60 minutes
Format: Discussion Based
Objectives:
• Evaluate mentoring capability
• Assess technical decision making
• Review cross-team collaboration
Focus Areas:
• Leadership
• Mentoring
• Influence
• Technical Leadership
• System Architecture
SUGGESTED SCHEDULE
----------------------------------------
Schedule Type: Multi Day
Day 1:
Time: 09:00 - 12:45
09:00-09:45: Technical Phone Screen (45min)
10:00-11:15: Coding Deep Dive (75min)
11:30-12:45: System Design (75min)
Day 2:
Time: 09:00 - 11:00
09:00-09:45: Behavioral Interview (45min)
10:00-11:00: Technical Leadership (60min)
INTERVIEWER REQUIREMENTS
----------------------------------------
Technical Phone Screen:
Required Skills: technical_assessment, coding_evaluation
Suggested Interviewers: senior_engineer, tech_lead
Calibration Level: Standard
Coding Deep Dive:
Required Skills: advanced_technical, code_quality_assessment
Suggested Interviewers: senior_engineer, staff_engineer
Calibration Level: High
System Design:
Required Skills: architecture_design, scalability_assessment
Suggested Interviewers: senior_architect, staff_engineer
Calibration Level: High
Behavioral:
Required Skills: behavioral_interviewing, competency_assessment
Suggested Interviewers: hiring_manager, people_manager
Calibration Level: Standard
Technical Leadership:
Required Skills: leadership_assessment, technical_mentoring
Suggested Interviewers: engineering_manager, senior_staff
Calibration Level: High
SCORECARD TEMPLATE
----------------------------------------
Scoring Scale:
4: Exceeds Expectations - Demonstrates mastery beyond required level
3: Meets Expectations - Solid performance meeting all requirements
2: Partially Meets - Shows potential but has development areas
1: Does Not Meet - Significant gaps in required competencies
Evaluation Dimensions:
• System Architecture (Weight: high)
• Technical Leadership (Weight: high)
• Mentoring (Weight: high)
• Cross Team Collab (Weight: high)
• Technology Evaluation (Weight: medium)
• Process Improvement (Weight: medium)
• Hiring Contribution (Weight: medium)
• Communication (Weight: high)
• Cultural Fit (Weight: medium)
• Learning Agility (Weight: medium)
CALIBRATION NOTES
----------------------------------------
Hiring Bar: Calibrated for senior level software engineer role
Common Pitfalls:
• Avoid comparing candidates to each other rather than to the role standard
• Don't let one strong/weak area overshadow overall assessment
• Ensure consistent application of evaluation criteria
FILE:hiring_calibrator.py
#!/usr/bin/env python3
"""
Hiring Calibrator
Analyzes interview scores from multiple candidates and interviewers to detect bias,
calibration issues, and inconsistent rubric application. Generates calibration reports
with specific recommendations for interviewer coaching and process improvements.
Usage:
python hiring_calibrator.py --input interview_results.json --analysis-type comprehensive
python hiring_calibrator.py --input data.json --competencies technical,leadership --output report.json
python hiring_calibrator.py --input historical_data.json --trend-analysis --period quarterly
"""
import os
import sys
import json
import argparse
import statistics
from datetime import datetime, timedelta
from typing import Dict, List, Optional, Any, Tuple
from collections import defaultdict, Counter
import math
class HiringCalibrator:
"""Analyzes interview data for bias detection and calibration issues."""
def __init__(self):
self.bias_thresholds = self._init_bias_thresholds()
self.calibration_standards = self._init_calibration_standards()
self.demographic_categories = self._init_demographic_categories()
def _init_bias_thresholds(self) -> Dict[str, float]:
"""Initialize statistical thresholds for bias detection."""
return {
"score_variance_threshold": 1.5, # Standard deviations
"pass_rate_difference_threshold": 0.15, # 15% difference
"interviewer_consistency_threshold": 0.8, # Correlation coefficient
"demographic_parity_threshold": 0.10, # 10% difference
"score_inflation_threshold": 0.3, # 30% above historical average
"score_deflation_threshold": 0.3, # 30% below historical average
"minimum_sample_size": 5 # Minimum candidates per analysis
}
def _init_calibration_standards(self) -> Dict[str, Dict]:
"""Initialize expected calibration standards."""
return {
"score_distribution": {
"target_mean": 2.8, # Expected average score (1-4 scale)
"target_std": 0.9, # Expected standard deviation
"expected_distribution": {
"1": 0.10, # 10% score 1 (does not meet)
"2": 0.25, # 25% score 2 (partially meets)
"3": 0.45, # 45% score 3 (meets expectations)
"4": 0.20 # 20% score 4 (exceeds expectations)
}
},
"interviewer_agreement": {
"minimum_correlation": 0.70, # Minimum correlation between interviewers
"maximum_std_deviation": 0.8, # Maximum std dev in scores for same candidate
"agreement_threshold": 0.75 # % of time interviewers should agree within 1 point
},
"pass_rates": {
"junior_level": 0.25, # 25% pass rate for junior roles
"mid_level": 0.20, # 20% pass rate for mid roles
"senior_level": 0.15, # 15% pass rate for senior roles
"staff_level": 0.10, # 10% pass rate for staff+ roles
"leadership": 0.12 # 12% pass rate for leadership roles
}
}
def _init_demographic_categories(self) -> List[str]:
"""Initialize demographic categories to analyze for bias."""
return [
"gender", "ethnicity", "education_level", "previous_company_size",
"years_experience", "university_tier", "geographic_location"
]
def analyze_hiring_calibration(self, interview_data: List[Dict[str, Any]],
analysis_type: str = "comprehensive",
competencies: Optional[List[str]] = None,
trend_analysis: bool = False,
period: str = "monthly") -> Dict[str, Any]:
"""Perform comprehensive hiring calibration analysis."""
# Validate and preprocess data
processed_data = self._preprocess_interview_data(interview_data)
if len(processed_data) < self.bias_thresholds["minimum_sample_size"]:
return {
"error": "Insufficient data for analysis",
"minimum_required": self.bias_thresholds["minimum_sample_size"],
"actual_samples": len(processed_data)
}
# Perform different types of analysis based on request
analysis_results = {
"analysis_type": analysis_type,
"data_summary": self._generate_data_summary(processed_data),
"generated_at": datetime.now().isoformat()
}
if analysis_type in ["comprehensive", "bias"]:
analysis_results["bias_analysis"] = self._analyze_bias_patterns(processed_data, competencies)
if analysis_type in ["comprehensive", "calibration"]:
analysis_results["calibration_analysis"] = self._analyze_calibration_consistency(processed_data, competencies)
if analysis_type in ["comprehensive", "interviewer"]:
analysis_results["interviewer_analysis"] = self._analyze_interviewer_bias(processed_data)
if analysis_type in ["comprehensive", "scoring"]:
analysis_results["scoring_analysis"] = self._analyze_scoring_patterns(processed_data, competencies)
if trend_analysis:
analysis_results["trend_analysis"] = self._analyze_trends_over_time(processed_data, period)
# Generate recommendations
analysis_results["recommendations"] = self._generate_recommendations(analysis_results)
# Calculate overall calibration health score
analysis_results["calibration_health_score"] = self._calculate_health_score(analysis_results)
return analysis_results
def _preprocess_interview_data(self, raw_data: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""Clean and validate interview data."""
processed_data = []
for record in raw_data:
if self._validate_interview_record(record):
processed_record = self._standardize_record(record)
processed_data.append(processed_record)
return processed_data
def _validate_interview_record(self, record: Dict[str, Any]) -> bool:
"""Validate that an interview record has required fields."""
required_fields = ["candidate_id", "interviewer_id", "scores", "overall_recommendation", "date"]
for field in required_fields:
if field not in record or record[field] is None:
return False
# Validate scores format
if not isinstance(record["scores"], dict):
return False
# Validate score values are numeric and in valid range (1-4)
for competency, score in record["scores"].items():
if not isinstance(score, (int, float)) or not (1 <= score <= 4):
return False
return True
def _standardize_record(self, record: Dict[str, Any]) -> Dict[str, Any]:
"""Standardize record format and add computed fields."""
standardized = record.copy()
# Calculate average score
scores = list(record["scores"].values())
standardized["average_score"] = statistics.mean(scores)
# Standardize recommendation to binary
recommendation = record["overall_recommendation"].lower()
standardized["hire_decision"] = recommendation in ["hire", "strong hire", "yes"]
# Parse date if string
if isinstance(record["date"], str):
try:
standardized["date"] = datetime.fromisoformat(record["date"].replace("Z", "+00:00"))
except ValueError:
standardized["date"] = datetime.now()
# Add demographic info if available
for category in self.demographic_categories:
if category not in standardized:
standardized[category] = "unknown"
# Add level normalization
role = record.get("role", "").lower()
if any(level in role for level in ["junior", "associate", "entry"]):
standardized["normalized_level"] = "junior"
elif any(level in role for level in ["senior", "sr"]):
standardized["normalized_level"] = "senior"
elif any(level in role for level in ["staff", "principal", "lead"]):
standardized["normalized_level"] = "staff"
else:
standardized["normalized_level"] = "mid"
return standardized
def _generate_data_summary(self, data: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Generate summary statistics for the dataset."""
if not data:
return {}
total_candidates = len(data)
unique_interviewers = len(set(record["interviewer_id"] for record in data))
# Score statistics
all_scores = []
all_average_scores = []
hire_decisions = []
for record in data:
all_scores.extend(record["scores"].values())
all_average_scores.append(record["average_score"])
hire_decisions.append(record["hire_decision"])
# Date range
dates = [record["date"] for record in data if record["date"]]
date_range = {
"start_date": min(dates).isoformat() if dates else None,
"end_date": max(dates).isoformat() if dates else None,
"total_days": (max(dates) - min(dates)).days if len(dates) > 1 else 0
}
# Role distribution
roles = [record.get("role", "unknown") for record in data]
role_distribution = dict(Counter(roles))
return {
"total_candidates": total_candidates,
"unique_interviewers": unique_interviewers,
"candidates_per_interviewer": round(total_candidates / unique_interviewers, 2),
"date_range": date_range,
"score_statistics": {
"mean_individual_scores": round(statistics.mean(all_scores), 2),
"std_individual_scores": round(statistics.stdev(all_scores) if len(all_scores) > 1 else 0, 2),
"mean_average_scores": round(statistics.mean(all_average_scores), 2),
"std_average_scores": round(statistics.stdev(all_average_scores) if len(all_average_scores) > 1 else 0, 2)
},
"hire_rate": round(sum(hire_decisions) / len(hire_decisions), 3),
"role_distribution": role_distribution
}
def _analyze_bias_patterns(self, data: List[Dict[str, Any]],
target_competencies: Optional[List[str]]) -> Dict[str, Any]:
"""Analyze potential bias patterns in interview decisions."""
bias_analysis = {
"demographic_bias": {},
"interviewer_bias": {},
"competency_bias": {},
"overall_bias_score": 0
}
# Analyze demographic bias
for demographic in self.demographic_categories:
if all(record.get(demographic) == "unknown" for record in data):
continue
demographic_analysis = self._analyze_demographic_bias(data, demographic)
if demographic_analysis["bias_detected"]:
bias_analysis["demographic_bias"][demographic] = demographic_analysis
# Analyze interviewer bias
bias_analysis["interviewer_bias"] = self._analyze_interviewer_bias(data)
# Analyze competency bias if specified
if target_competencies:
bias_analysis["competency_bias"] = self._analyze_competency_bias(data, target_competencies)
# Calculate overall bias score
bias_analysis["overall_bias_score"] = self._calculate_bias_score(bias_analysis)
return bias_analysis
def _analyze_demographic_bias(self, data: List[Dict[str, Any]],
demographic: str) -> Dict[str, Any]:
"""Analyze bias for a specific demographic category."""
# Group data by demographic values
demographic_groups = defaultdict(list)
for record in data:
demo_value = record.get(demographic, "unknown")
if demo_value != "unknown":
demographic_groups[demo_value].append(record)
if len(demographic_groups) < 2:
return {"bias_detected": False, "reason": "insufficient_groups"}
# Calculate statistics for each group
group_stats = {}
for group, records in demographic_groups.items():
if len(records) >= self.bias_thresholds["minimum_sample_size"]:
scores = [r["average_score"] for r in records]
hire_rate = sum(r["hire_decision"] for r in records) / len(records)
group_stats[group] = {
"count": len(records),
"mean_score": statistics.mean(scores),
"hire_rate": hire_rate,
"std_score": statistics.stdev(scores) if len(scores) > 1 else 0
}
if len(group_stats) < 2:
return {"bias_detected": False, "reason": "insufficient_sample_sizes"}
# Detect statistical differences
bias_detected = False
bias_details = {}
# Check for significant differences in hire rates
hire_rates = [stats["hire_rate"] for stats in group_stats.values()]
max_hire_rate_diff = max(hire_rates) - min(hire_rates)
if max_hire_rate_diff > self.bias_thresholds["demographic_parity_threshold"]:
bias_detected = True
bias_details["hire_rate_disparity"] = {
"max_difference": round(max_hire_rate_diff, 3),
"threshold": self.bias_thresholds["demographic_parity_threshold"],
"group_stats": group_stats
}
# Check for significant differences in scoring
mean_scores = [stats["mean_score"] for stats in group_stats.values()]
max_score_diff = max(mean_scores) - min(mean_scores)
if max_score_diff > 0.5: # Half point difference threshold
bias_detected = True
bias_details["scoring_disparity"] = {
"max_difference": round(max_score_diff, 3),
"group_stats": group_stats
}
return {
"bias_detected": bias_detected,
"demographic": demographic,
"group_statistics": group_stats,
"bias_details": bias_details,
"recommendation": self._generate_demographic_bias_recommendation(demographic, bias_details) if bias_detected else None
}
def _analyze_interviewer_bias(self, data: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Analyze bias patterns across different interviewers."""
interviewer_stats = defaultdict(list)
# Group by interviewer
for record in data:
interviewer_id = record["interviewer_id"]
interviewer_stats[interviewer_id].append(record)
# Calculate statistics per interviewer
interviewer_analysis = {}
for interviewer_id, records in interviewer_stats.items():
if len(records) >= self.bias_thresholds["minimum_sample_size"]:
scores = [r["average_score"] for r in records]
hire_rate = sum(r["hire_decision"] for r in records) / len(records)
interviewer_analysis[interviewer_id] = {
"total_interviews": len(records),
"mean_score": statistics.mean(scores),
"std_score": statistics.stdev(scores) if len(scores) > 1 else 0,
"hire_rate": hire_rate,
"score_inflation": self._detect_score_inflation(scores),
"consistency_score": self._calculate_interviewer_consistency(records)
}
# Identify outlier interviewers
if len(interviewer_analysis) > 1:
overall_mean_score = statistics.mean([stats["mean_score"] for stats in interviewer_analysis.values()])
overall_hire_rate = statistics.mean([stats["hire_rate"] for stats in interviewer_analysis.values()])
outlier_interviewers = {}
for interviewer_id, stats in interviewer_analysis.items():
issues = []
# Check for score inflation/deflation
if stats["mean_score"] > overall_mean_score * (1 + self.bias_thresholds["score_inflation_threshold"]):
issues.append("score_inflation")
elif stats["mean_score"] < overall_mean_score * (1 - self.bias_thresholds["score_deflation_threshold"]):
issues.append("score_deflation")
# Check for hire rate deviation
hire_rate_diff = abs(stats["hire_rate"] - overall_hire_rate)
if hire_rate_diff > self.bias_thresholds["pass_rate_difference_threshold"]:
issues.append("hire_rate_deviation")
# Check for low consistency
if stats["consistency_score"] < self.bias_thresholds["interviewer_consistency_threshold"]:
issues.append("low_consistency")
if issues:
outlier_interviewers[interviewer_id] = {
"issues": issues,
"statistics": stats,
"severity": len(issues) # More issues = higher severity
}
return {
"interviewer_statistics": interviewer_analysis,
"outlier_interviewers": outlier_interviewers if len(interviewer_analysis) > 1 else {},
"overall_consistency": self._calculate_overall_interviewer_consistency(data),
"recommendations": self._generate_interviewer_recommendations(outlier_interviewers if len(interviewer_analysis) > 1 else {})
}
def _analyze_competency_bias(self, data: List[Dict[str, Any]],
competencies: List[str]) -> Dict[str, Any]:
"""Analyze bias patterns within specific competencies."""
competency_analysis = {}
for competency in competencies:
# Extract scores for this competency
competency_scores = []
for record in data:
if competency in record["scores"]:
competency_scores.append({
"score": record["scores"][competency],
"interviewer": record["interviewer_id"],
"candidate": record["candidate_id"],
"overall_decision": record["hire_decision"]
})
if len(competency_scores) < self.bias_thresholds["minimum_sample_size"]:
continue
# Analyze scoring patterns
scores = [item["score"] for item in competency_scores]
score_variance = statistics.variance(scores) if len(scores) > 1 else 0
# Analyze by interviewer
interviewer_competency_scores = defaultdict(list)
for item in competency_scores:
interviewer_competency_scores[item["interviewer"]].append(item["score"])
interviewer_variations = {}
if len(interviewer_competency_scores) > 1:
interviewer_means = {interviewer: statistics.mean(scores)
for interviewer, scores in interviewer_competency_scores.items()
if len(scores) >= 3}
if len(interviewer_means) > 1:
mean_of_means = statistics.mean(interviewer_means.values())
for interviewer, mean_score in interviewer_means.items():
deviation = abs(mean_score - mean_of_means)
if deviation > 0.5: # More than half point deviation
interviewer_variations[interviewer] = {
"mean_score": round(mean_score, 2),
"deviation_from_average": round(deviation, 2),
"sample_size": len(interviewer_competency_scores[interviewer])
}
competency_analysis[competency] = {
"total_scores": len(competency_scores),
"mean_score": round(statistics.mean(scores), 2),
"score_variance": round(score_variance, 2),
"interviewer_variations": interviewer_variations,
"bias_detected": len(interviewer_variations) > 0
}
return competency_analysis
def _analyze_calibration_consistency(self, data: List[Dict[str, Any]],
target_competencies: Optional[List[str]]) -> Dict[str, Any]:
"""Analyze calibration consistency across interviews."""
# Group candidates by those interviewed by multiple people
candidate_interviewers = defaultdict(list)
for record in data:
candidate_interviewers[record["candidate_id"]].append(record)
multi_interviewer_candidates = {
candidate: records for candidate, records in candidate_interviewers.items()
if len(records) > 1
}
if not multi_interviewer_candidates:
return {
"error": "No candidates with multiple interviewers found",
"single_interviewer_analysis": self._analyze_single_interviewer_consistency(data)
}
# Calculate agreement statistics
agreement_stats = []
score_correlations = []
for candidate, records in multi_interviewer_candidates.items():
candidate_scores = []
interviewer_pairs = []
for record in records:
avg_score = record["average_score"]
candidate_scores.append(avg_score)
interviewer_pairs.append(record["interviewer_id"])
if len(candidate_scores) > 1:
# Calculate standard deviation of scores for this candidate
score_std = statistics.stdev(candidate_scores)
agreement_stats.append(score_std)
# Check if all interviewers agree within 1 point
score_range = max(candidate_scores) - min(candidate_scores)
agreement_within_one = score_range <= 1.0
score_correlations.append({
"candidate": candidate,
"scores": candidate_scores,
"interviewers": interviewer_pairs,
"score_std": score_std,
"score_range": score_range,
"agreement_within_one": agreement_within_one
})
# Calculate overall calibration metrics
mean_score_std = statistics.mean(agreement_stats) if agreement_stats else 0
agreement_rate = sum(1 for corr in score_correlations if corr["agreement_within_one"]) / len(score_correlations) if score_correlations else 0
calibration_quality = "good"
if mean_score_std > self.calibration_standards["interviewer_agreement"]["maximum_std_deviation"]:
calibration_quality = "poor"
elif agreement_rate < self.calibration_standards["interviewer_agreement"]["agreement_threshold"]:
calibration_quality = "fair"
return {
"multi_interviewer_candidates": len(multi_interviewer_candidates),
"mean_score_standard_deviation": round(mean_score_std, 3),
"agreement_within_one_point_rate": round(agreement_rate, 3),
"calibration_quality": calibration_quality,
"candidate_agreement_details": score_correlations,
"target_standards": self.calibration_standards["interviewer_agreement"],
"recommendations": self._generate_calibration_recommendations(mean_score_std, agreement_rate)
}
def _analyze_scoring_patterns(self, data: List[Dict[str, Any]],
target_competencies: Optional[List[str]]) -> Dict[str, Any]:
"""Analyze overall scoring patterns and distributions."""
# Overall score distribution
all_individual_scores = []
all_average_scores = []
score_distribution = defaultdict(int)
for record in data:
avg_score = record["average_score"]
all_average_scores.append(avg_score)
for competency, score in record["scores"].items():
if not target_competencies or competency in target_competencies:
all_individual_scores.append(score)
score_distribution[str(int(score))] += 1
# Calculate distribution percentages
total_scores = sum(score_distribution.values())
score_percentages = {score: count/total_scores for score, count in score_distribution.items()}
# Compare against expected distribution
expected_dist = self.calibration_standards["score_distribution"]["expected_distribution"]
distribution_analysis = {}
for score in ["1", "2", "3", "4"]:
expected_pct = expected_dist.get(score, 0)
actual_pct = score_percentages.get(score, 0)
difference = actual_pct - expected_pct
distribution_analysis[score] = {
"expected_percentage": expected_pct,
"actual_percentage": round(actual_pct, 3),
"difference": round(difference, 3),
"significant_deviation": abs(difference) > 0.05 # 5% threshold
}
# Calculate scoring statistics
mean_score = statistics.mean(all_individual_scores) if all_individual_scores else 0
std_score = statistics.stdev(all_individual_scores) if len(all_individual_scores) > 1 else 0
target_mean = self.calibration_standards["score_distribution"]["target_mean"]
target_std = self.calibration_standards["score_distribution"]["target_std"]
# Analyze pass rates by level
level_pass_rates = {}
level_groups = defaultdict(list)
for record in data:
level = record.get("normalized_level", "unknown")
level_groups[level].append(record["hire_decision"])
for level, decisions in level_groups.items():
if len(decisions) >= self.bias_thresholds["minimum_sample_size"]:
pass_rate = sum(decisions) / len(decisions)
expected_rate = self.calibration_standards["pass_rates"].get(f"{level}_level", 0.15)
level_pass_rates[level] = {
"actual_pass_rate": round(pass_rate, 3),
"expected_pass_rate": expected_rate,
"difference": round(pass_rate - expected_rate, 3),
"sample_size": len(decisions)
}
return {
"score_statistics": {
"mean_score": round(mean_score, 2),
"std_score": round(std_score, 2),
"target_mean": target_mean,
"target_std": target_std,
"mean_deviation": round(abs(mean_score - target_mean), 2),
"std_deviation": round(abs(std_score - target_std), 2)
},
"score_distribution": distribution_analysis,
"level_pass_rates": level_pass_rates,
"overall_assessment": self._assess_scoring_health(distribution_analysis, mean_score, target_mean)
}
def _analyze_trends_over_time(self, data: List[Dict[str, Any]], period: str) -> Dict[str, Any]:
"""Analyze trends in hiring patterns over time."""
# Sort data by date
dated_data = [record for record in data if record.get("date")]
dated_data.sort(key=lambda x: x["date"])
if len(dated_data) < 10: # Need minimum data for trend analysis
return {"error": "Insufficient data for trend analysis", "minimum_required": 10}
# Group by time period
period_groups = defaultdict(list)
for record in dated_data:
date = record["date"]
if period == "weekly":
period_key = date.strftime("%Y-W%U")
elif period == "monthly":
period_key = date.strftime("%Y-%m")
elif period == "quarterly":
quarter = (date.month - 1) // 3 + 1
period_key = f"{date.year}-Q{quarter}"
else: # daily
period_key = date.strftime("%Y-%m-%d")
period_groups[period_key].append(record)
# Calculate metrics for each period
period_metrics = {}
for period_key, records in period_groups.items():
if len(records) >= 3: # Minimum for meaningful metrics
scores = [r["average_score"] for r in records]
hire_rate = sum(r["hire_decision"] for r in records) / len(records)
period_metrics[period_key] = {
"count": len(records),
"mean_score": statistics.mean(scores),
"hire_rate": hire_rate,
"std_score": statistics.stdev(scores) if len(scores) > 1 else 0
}
if len(period_metrics) < 3:
return {"error": "Insufficient periods for trend analysis"}
# Analyze trends
sorted_periods = sorted(period_metrics.keys())
mean_scores = [period_metrics[p]["mean_score"] for p in sorted_periods]
hire_rates = [period_metrics[p]["hire_rate"] for p in sorted_periods]
# Simple linear trend calculation
score_trend = self._calculate_linear_trend(mean_scores)
hire_rate_trend = self._calculate_linear_trend(hire_rates)
return {
"period": period,
"total_periods": len(period_metrics),
"period_metrics": period_metrics,
"trends": {
"score_trend": {
"direction": "increasing" if score_trend > 0.01 else "decreasing" if score_trend < -0.01 else "stable",
"slope": round(score_trend, 4),
"significance": "significant" if abs(score_trend) > 0.05 else "minor"
},
"hire_rate_trend": {
"direction": "increasing" if hire_rate_trend > 0.005 else "decreasing" if hire_rate_trend < -0.005 else "stable",
"slope": round(hire_rate_trend, 4),
"significance": "significant" if abs(hire_rate_trend) > 0.02 else "minor"
}
},
"insights": self._generate_trend_insights(score_trend, hire_rate_trend, period_metrics)
}
def _calculate_linear_trend(self, values: List[float]) -> float:
"""Calculate simple linear trend slope."""
if len(values) < 2:
return 0
n = len(values)
x = list(range(n))
# Calculate slope using least squares
x_mean = statistics.mean(x)
y_mean = statistics.mean(values)
numerator = sum((x[i] - x_mean) * (values[i] - y_mean) for i in range(n))
denominator = sum((x[i] - x_mean) ** 2 for i in range(n))
return numerator / denominator if denominator != 0 else 0
def _detect_score_inflation(self, scores: List[float]) -> Dict[str, Any]:
"""Detect if an interviewer shows score inflation patterns."""
if len(scores) < 5:
return {"insufficient_data": True}
mean_score = statistics.mean(scores)
std_score = statistics.stdev(scores)
# Check against expected mean (2.8)
expected_mean = self.calibration_standards["score_distribution"]["target_mean"]
deviation = mean_score - expected_mean
# High scores with low variance might indicate inflation
high_scores_low_variance = mean_score > 3.2 and std_score < 0.5
# Check distribution - too many 4s might indicate inflation
score_counts = Counter([int(score) for score in scores])
four_count_ratio = score_counts.get(4, 0) / len(scores)
return {
"mean_score": round(mean_score, 2),
"expected_mean": expected_mean,
"deviation": round(deviation, 2),
"high_scores_low_variance": high_scores_low_variance,
"four_count_ratio": round(four_count_ratio, 2),
"inflation_detected": deviation > 0.3 or high_scores_low_variance or four_count_ratio > 0.4
}
def _calculate_interviewer_consistency(self, records: List[Dict[str, Any]]) -> float:
"""Calculate consistency score for an interviewer."""
if len(records) < 3:
return 0.5 # Neutral score for insufficient data
# Look at variance in scoring
avg_scores = [r["average_score"] for r in records]
score_variance = statistics.variance(avg_scores)
# Look at decision consistency relative to scores
decisions = [r["hire_decision"] for r in records]
scores_of_hires = [r["average_score"] for r in records if r["hire_decision"]]
scores_of_no_hires = [r["average_score"] for r in records if not r["hire_decision"]]
# Good consistency means hires have higher average scores
decision_consistency = 0.5
if scores_of_hires and scores_of_no_hires:
hire_mean = statistics.mean(scores_of_hires)
no_hire_mean = statistics.mean(scores_of_no_hires)
score_gap = hire_mean - no_hire_mean
decision_consistency = min(1.0, max(0.0, score_gap / 2.0)) # Normalize to 0-1
# Combine metrics (lower variance = higher consistency)
variance_consistency = max(0.0, 1.0 - (score_variance / 2.0))
return (decision_consistency + variance_consistency) / 2
def _calculate_overall_interviewer_consistency(self, data: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Calculate overall consistency across all interviewers."""
interviewer_consistency_scores = []
interviewer_records = defaultdict(list)
for record in data:
interviewer_records[record["interviewer_id"]].append(record)
for interviewer_id, records in interviewer_records.items():
if len(records) >= 3:
consistency = self._calculate_interviewer_consistency(records)
interviewer_consistency_scores.append(consistency)
if not interviewer_consistency_scores:
return {"error": "Insufficient data per interviewer for consistency analysis"}
return {
"mean_consistency": round(statistics.mean(interviewer_consistency_scores), 3),
"std_consistency": round(statistics.stdev(interviewer_consistency_scores) if len(interviewer_consistency_scores) > 1 else 0, 3),
"min_consistency": round(min(interviewer_consistency_scores), 3),
"max_consistency": round(max(interviewer_consistency_scores), 3),
"interviewers_analyzed": len(interviewer_consistency_scores),
"target_threshold": self.bias_thresholds["interviewer_consistency_threshold"]
}
def _calculate_bias_score(self, bias_analysis: Dict[str, Any]) -> float:
"""Calculate overall bias score (0-1, where 1 is most biased)."""
bias_factors = []
# Demographic bias factors
demographic_bias = bias_analysis.get("demographic_bias", {})
for demo, analysis in demographic_bias.items():
if analysis.get("bias_detected"):
bias_factors.append(0.3) # Each demographic bias adds 0.3
# Interviewer bias factors
interviewer_bias = bias_analysis.get("interviewer_bias", {})
outlier_interviewers = interviewer_bias.get("outlier_interviewers", {})
if outlier_interviewers:
# Scale by severity and number of outliers
total_severity = sum(info["severity"] for info in outlier_interviewers.values())
bias_factors.append(min(0.5, total_severity * 0.1))
# Competency bias factors
competency_bias = bias_analysis.get("competency_bias", {})
for comp, analysis in competency_bias.items():
if analysis.get("bias_detected"):
bias_factors.append(0.2) # Each competency bias adds 0.2
return min(1.0, sum(bias_factors))
def _calculate_health_score(self, analysis: Dict[str, Any]) -> Dict[str, Any]:
"""Calculate overall calibration health score."""
health_factors = []
# Bias score (lower is better)
bias_analysis = analysis.get("bias_analysis", {})
bias_score = bias_analysis.get("overall_bias_score", 0)
bias_health = max(0, 1 - bias_score)
health_factors.append(("bias", bias_health, 0.3))
# Calibration consistency
calibration_analysis = analysis.get("calibration_analysis", {})
if "calibration_quality" in calibration_analysis:
quality_map = {"good": 1.0, "fair": 0.7, "poor": 0.3}
calibration_health = quality_map.get(calibration_analysis["calibration_quality"], 0.5)
health_factors.append(("calibration", calibration_health, 0.25))
# Interviewer consistency
interviewer_analysis = analysis.get("interviewer_analysis", {})
overall_consistency = interviewer_analysis.get("overall_consistency", {})
if "mean_consistency" in overall_consistency:
consistency_health = overall_consistency["mean_consistency"]
health_factors.append(("interviewer_consistency", consistency_health, 0.25))
# Scoring patterns health
scoring_analysis = analysis.get("scoring_analysis", {})
if "overall_assessment" in scoring_analysis:
assessment_map = {"healthy": 1.0, "concerning": 0.6, "poor": 0.2}
scoring_health = assessment_map.get(scoring_analysis["overall_assessment"], 0.5)
health_factors.append(("scoring_patterns", scoring_health, 0.2))
# Calculate weighted average
if health_factors:
weighted_sum = sum(score * weight for _, score, weight in health_factors)
total_weight = sum(weight for _, _, weight in health_factors)
overall_score = weighted_sum / total_weight
else:
overall_score = 0.5 # Neutral if no data
# Categorize health
if overall_score >= 0.8:
health_category = "excellent"
elif overall_score >= 0.7:
health_category = "good"
elif overall_score >= 0.5:
health_category = "fair"
else:
health_category = "poor"
return {
"overall_score": round(overall_score, 3),
"health_category": health_category,
"component_scores": {name: round(score, 3) for name, score, _ in health_factors},
"improvement_priority": self._identify_improvement_priorities(health_factors)
}
def _identify_improvement_priorities(self, health_factors: List[Tuple[str, float, float]]) -> List[str]:
"""Identify areas that need the most improvement."""
priorities = []
for name, score, weight in health_factors:
impact = (1 - score) * weight # Low scores with high weights = high priority
if impact > 0.15: # Significant impact threshold
priorities.append(name)
# Sort by impact (highest first)
priorities.sort(key=lambda name: next((1 - score) * weight for n, score, weight in health_factors if n == name), reverse=True)
return priorities
def _generate_recommendations(self, analysis: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate actionable recommendations based on analysis results."""
recommendations = []
# Bias-related recommendations
bias_analysis = analysis.get("bias_analysis", {})
# Demographic bias recommendations
for demo, demo_analysis in bias_analysis.get("demographic_bias", {}).items():
if demo_analysis.get("bias_detected"):
recommendations.append({
"priority": "high",
"category": "bias_mitigation",
"title": f"Address {demo.replace('_', ' ').title()} Bias",
"description": demo_analysis.get("recommendation", f"Implement bias mitigation strategies for {demo}"),
"actions": [
"Conduct unconscious bias training focused on this demographic",
"Review and standardize interview questions",
"Implement diverse interview panels",
"Monitor hiring metrics by demographic group"
]
})
# Interviewer-specific recommendations
interviewer_analysis = bias_analysis.get("interviewer_bias", {})
outlier_interviewers = interviewer_analysis.get("outlier_interviewers", {})
for interviewer_id, outlier_info in outlier_interviewers.items():
issues = outlier_info["issues"]
priority = "high" if outlier_info["severity"] >= 3 else "medium"
actions = []
if "score_inflation" in issues:
actions.extend([
"Provide calibration training on scoring standards",
"Shadow experienced interviewers for recalibration",
"Review examples of each score level"
])
if "score_deflation" in issues:
actions.extend([
"Review expectations for role level",
"Calibrate against recent successful hires",
"Discuss evaluation criteria with hiring manager"
])
if "hire_rate_deviation" in issues:
actions.extend([
"Review hiring bar standards",
"Participate in calibration sessions",
"Compare decision criteria with team"
])
if "low_consistency" in issues:
actions.extend([
"Practice structured interviewing techniques",
"Use standardized scorecards",
"Document specific examples for each score"
])
recommendations.append({
"priority": priority,
"category": "interviewer_coaching",
"title": f"Coach Interviewer {interviewer_id}",
"description": f"Address issues: {', '.join(issues)}",
"actions": list(set(actions)) # Remove duplicates
})
# Calibration recommendations
calibration_analysis = analysis.get("calibration_analysis", {})
if calibration_analysis.get("calibration_quality") in ["fair", "poor"]:
recommendations.append({
"priority": "high",
"category": "calibration_improvement",
"title": "Improve Interview Calibration",
"description": f"Current calibration quality: {calibration_analysis.get('calibration_quality')}",
"actions": [
"Conduct monthly calibration sessions",
"Create shared examples of good/poor answers",
"Implement mandatory interviewer shadowing",
"Standardize scoring rubrics across all interviewers",
"Review and align on role expectations"
]
})
# Scoring pattern recommendations
scoring_analysis = analysis.get("scoring_analysis", {})
if scoring_analysis.get("overall_assessment") in ["concerning", "poor"]:
recommendations.append({
"priority": "medium",
"category": "scoring_standards",
"title": "Adjust Scoring Standards",
"description": "Scoring patterns deviate significantly from expected distribution",
"actions": [
"Review and communicate target score distributions",
"Provide examples for each score level",
"Monitor pass rates by role level",
"Adjust hiring bar if consistently too high/low"
]
})
# Health score recommendations
health_score = analysis.get("calibration_health_score", {})
priorities = health_score.get("improvement_priority", [])
if "bias" in priorities:
recommendations.append({
"priority": "critical",
"category": "bias_mitigation",
"title": "Implement Comprehensive Bias Mitigation",
"description": "Multiple bias indicators detected across the hiring process",
"actions": [
"Mandatory unconscious bias training for all interviewers",
"Implement structured interview protocols",
"Diversify interview panels",
"Regular bias audits and monitoring",
"Create accountability metrics for fair hiring"
]
})
# Sort by priority
priority_order = {"critical": 0, "high": 1, "medium": 2, "low": 3}
recommendations.sort(key=lambda x: priority_order.get(x["priority"], 3))
return recommendations
def _generate_demographic_bias_recommendation(self, demographic: str, bias_details: Dict[str, Any]) -> str:
"""Generate specific recommendation for demographic bias."""
if "hire_rate_disparity" in bias_details:
return f"Significant hire rate disparity detected for {demographic}. Implement structured interviews and diverse panels."
elif "scoring_disparity" in bias_details:
return f"Scoring disparity detected for {demographic}. Provide unconscious bias training and standardize evaluation criteria."
else:
return f"Potential bias detected for {demographic}. Monitor closely and implement bias mitigation strategies."
def _generate_interviewer_recommendations(self, outlier_interviewers: Dict[str, Any]) -> List[str]:
"""Generate recommendations for interviewer issues."""
if not outlier_interviewers:
return ["All interviewers performing within expected ranges"]
recommendations = []
for interviewer, info in outlier_interviewers.items():
issues = info["issues"]
if len(issues) >= 2:
recommendations.append(f"Interviewer {interviewer}: Requires comprehensive recalibration - multiple issues detected")
elif "score_inflation" in issues:
recommendations.append(f"Interviewer {interviewer}: Provide calibration training on scoring standards")
elif "hire_rate_deviation" in issues:
recommendations.append(f"Interviewer {interviewer}: Review hiring bar standards and decision criteria")
return recommendations
def _generate_calibration_recommendations(self, mean_std: float, agreement_rate: float) -> List[str]:
"""Generate calibration improvement recommendations."""
recommendations = []
if mean_std > self.calibration_standards["interviewer_agreement"]["maximum_std_deviation"]:
recommendations.append("High score variance detected - implement regular calibration sessions")
recommendations.append("Create shared examples of scoring standards for each competency")
if agreement_rate < self.calibration_standards["interviewer_agreement"]["agreement_threshold"]:
recommendations.append("Low interviewer agreement rate - standardize interview questions and evaluation criteria")
recommendations.append("Implement mandatory interviewer training on consistent evaluation")
if not recommendations:
recommendations.append("Calibration appears healthy - maintain current practices")
return recommendations
def _assess_scoring_health(self, distribution: Dict[str, Any], mean_score: float, target_mean: float) -> str:
"""Assess overall health of scoring patterns."""
issues = 0
# Check distribution deviations
for score_level, analysis in distribution.items():
if analysis["significant_deviation"]:
issues += 1
# Check mean deviation
if abs(mean_score - target_mean) > 0.3:
issues += 1
if issues == 0:
return "healthy"
elif issues <= 2:
return "concerning"
else:
return "poor"
def _generate_trend_insights(self, score_trend: float, hire_rate_trend: float, period_metrics: Dict[str, Any]) -> List[str]:
"""Generate insights from trend analysis."""
insights = []
if abs(score_trend) > 0.05:
direction = "increasing" if score_trend > 0 else "decreasing"
insights.append(f"Significant {direction} trend in average scores over time")
if score_trend > 0:
insights.append("May indicate score inflation or improving candidate quality")
else:
insights.append("May indicate stricter evaluation or declining candidate quality")
if abs(hire_rate_trend) > 0.02:
direction = "increasing" if hire_rate_trend > 0 else "decreasing"
insights.append(f"Significant {direction} trend in hire rates over time")
if hire_rate_trend > 0:
insights.append("Consider if hiring bar has lowered or candidate pool improved")
else:
insights.append("Consider if hiring bar has raised or candidate pool declined")
# Check for consistency
period_values = list(period_metrics.values())
hire_rates = [p["hire_rate"] for p in period_values]
hire_rate_variance = statistics.variance(hire_rates) if len(hire_rates) > 1 else 0
if hire_rate_variance > 0.01: # High variance in hire rates
insights.append("High variance in hire rates across periods - consider process standardization")
if not insights:
insights.append("Hiring patterns appear stable over time")
return insights
def _analyze_single_interviewer_consistency(self, data: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Analyze consistency for single-interviewer candidates."""
# Look at consistency within individual interviewers
interviewer_scores = defaultdict(list)
for record in data:
interviewer_scores[record["interviewer_id"]].extend(record["scores"].values())
consistency_analysis = {}
for interviewer, scores in interviewer_scores.items():
if len(scores) >= 10: # Need sufficient data
consistency_analysis[interviewer] = {
"mean_score": round(statistics.mean(scores), 2),
"std_score": round(statistics.stdev(scores), 2),
"coefficient_of_variation": round(statistics.stdev(scores) / statistics.mean(scores), 2),
"total_scores": len(scores)
}
return consistency_analysis
def format_human_readable(calibration_report: Dict[str, Any]) -> str:
"""Format calibration report in human-readable format."""
output = []
# Header
output.append("HIRING CALIBRATION ANALYSIS REPORT")
output.append("=" * 60)
output.append(f"Analysis Type: {calibration_report.get('analysis_type', 'N/A').title()}")
output.append(f"Generated: {calibration_report.get('generated_at', 'N/A')}")
if "error" in calibration_report:
output.append(f"\nError: {calibration_report['error']}")
return "\n".join(output)
# Data Summary
data_summary = calibration_report.get("data_summary", {})
if data_summary:
output.append(f"\nDATA SUMMARY")
output.append("-" * 30)
output.append(f"Total Candidates: {data_summary.get('total_candidates', 0)}")
output.append(f"Unique Interviewers: {data_summary.get('unique_interviewers', 0)}")
output.append(f"Overall Hire Rate: {data_summary.get('hire_rate', 0):.1%}")
score_stats = data_summary.get("score_statistics", {})
output.append(f"Average Score: {score_stats.get('mean_average_scores', 0):.2f}")
output.append(f"Score Std Dev: {score_stats.get('std_average_scores', 0):.2f}")
# Health Score
health_score = calibration_report.get("calibration_health_score", {})
if health_score:
output.append(f"\nCALIBRATION HEALTH SCORE")
output.append("-" * 30)
output.append(f"Overall Score: {health_score.get('overall_score', 0):.3f}")
output.append(f"Health Category: {health_score.get('health_category', 'Unknown').title()}")
if health_score.get("improvement_priority"):
output.append(f"Priority Areas: {', '.join(health_score['improvement_priority'])}")
# Bias Analysis
bias_analysis = calibration_report.get("bias_analysis", {})
if bias_analysis:
output.append(f"\nBIAS ANALYSIS")
output.append("-" * 30)
output.append(f"Overall Bias Score: {bias_analysis.get('overall_bias_score', 0):.3f}")
# Demographic bias
demographic_bias = bias_analysis.get("demographic_bias", {})
if demographic_bias:
output.append(f"\nDemographic Bias Issues:")
for demo, analysis in demographic_bias.items():
output.append(f" • {demo.replace('_', ' ').title()}: {analysis.get('bias_details', {}).keys()}")
# Interviewer bias
interviewer_bias = bias_analysis.get("interviewer_bias", {})
outlier_interviewers = interviewer_bias.get("outlier_interviewers", {})
if outlier_interviewers:
output.append(f"\nOutlier Interviewers:")
for interviewer, info in outlier_interviewers.items():
issues = ", ".join(info["issues"])
output.append(f" • {interviewer}: {issues}")
# Calibration Analysis
calibration_analysis = calibration_report.get("calibration_analysis", {})
if calibration_analysis and "error" not in calibration_analysis:
output.append(f"\nCALIBRATION CONSISTENCY")
output.append("-" * 30)
output.append(f"Quality: {calibration_analysis.get('calibration_quality', 'Unknown').title()}")
output.append(f"Agreement Rate: {calibration_analysis.get('agreement_within_one_point_rate', 0):.1%}")
output.append(f"Score Std Dev: {calibration_analysis.get('mean_score_standard_deviation', 0):.3f}")
# Scoring Analysis
scoring_analysis = calibration_report.get("scoring_analysis", {})
if scoring_analysis:
output.append(f"\nSCORING PATTERNS")
output.append("-" * 30)
output.append(f"Overall Assessment: {scoring_analysis.get('overall_assessment', 'Unknown').title()}")
score_stats = scoring_analysis.get("score_statistics", {})
output.append(f"Mean Score: {score_stats.get('mean_score', 0):.2f} (Target: {score_stats.get('target_mean', 0):.2f})")
# Distribution analysis
distribution = scoring_analysis.get("score_distribution", {})
if distribution:
output.append(f"\nScore Distribution vs Expected:")
for score in ["1", "2", "3", "4"]:
if score in distribution:
actual = distribution[score]["actual_percentage"]
expected = distribution[score]["expected_percentage"]
output.append(f" Score {score}: {actual:.1%} (Expected: {expected:.1%})")
# Top Recommendations
recommendations = calibration_report.get("recommendations", [])
if recommendations:
output.append(f"\nTOP RECOMMENDATIONS")
output.append("-" * 30)
for i, rec in enumerate(recommendations[:5], 1): # Show top 5
output.append(f"{i}. {rec['title']} ({rec['priority'].title()} Priority)")
output.append(f" {rec['description']}")
if rec.get('actions'):
output.append(f" Actions: {len(rec['actions'])} specific action items")
return "\n".join(output)
def main():
parser = argparse.ArgumentParser(description="Analyze interview data for bias and calibration issues")
parser.add_argument("--input", type=str, required=True, help="Input JSON file with interview results data")
parser.add_argument("--analysis-type", type=str, choices=["comprehensive", "bias", "calibration", "interviewer", "scoring"],
default="comprehensive", help="Type of analysis to perform")
parser.add_argument("--competencies", type=str, help="Comma-separated list of competencies to focus on")
parser.add_argument("--trend-analysis", action="store_true", help="Perform trend analysis over time")
parser.add_argument("--period", type=str, choices=["daily", "weekly", "monthly", "quarterly"],
default="monthly", help="Time period for trend analysis")
parser.add_argument("--output", type=str, help="Output file path")
parser.add_argument("--format", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
# Load input data
try:
with open(args.input, 'r') as f:
interview_data = json.load(f)
if not isinstance(interview_data, list):
print("Error: Input data must be a JSON array of interview records")
sys.exit(1)
except FileNotFoundError:
print(f"Error: Input file '{args.input}' not found")
sys.exit(1)
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in input file: {e}")
sys.exit(1)
except Exception as e:
print(f"Error reading input file: {e}")
sys.exit(1)
# Initialize calibrator and run analysis
calibrator = HiringCalibrator()
competencies = args.competencies.split(',') if args.competencies else None
try:
results = calibrator.analyze_hiring_calibration(
interview_data=interview_data,
analysis_type=args.analysis_type,
competencies=competencies,
trend_analysis=args.trend_analysis,
period=args.period
)
# Handle output
if args.output:
output_path = args.output
json_path = output_path if output_path.endswith('.json') else f"{output_path}.json"
text_path = output_path.replace('.json', '.txt') if output_path.endswith('.json') else f"{output_path}.txt"
else:
base_filename = f"calibration_report_{datetime.now().strftime('%Y%m%d_%H%M%S')}"
json_path = f"{base_filename}.json"
text_path = f"{base_filename}.txt"
# Write outputs
if args.format in ["json", "both"]:
with open(json_path, 'w') as f:
json.dump(results, f, indent=2, default=str)
print(f"JSON report written to: {json_path}")
if args.format in ["text", "both"]:
with open(text_path, 'w') as f:
f.write(format_human_readable(results))
print(f"Text report written to: {text_path}")
# Print summary
print(f"\nCalibration Analysis Summary:")
if "error" in results:
print(f"Error: {results['error']}")
else:
health_score = results.get("calibration_health_score", {})
print(f"Health Score: {health_score.get('overall_score', 0):.3f} ({health_score.get('health_category', 'Unknown').title()})")
bias_score = results.get("bias_analysis", {}).get("overall_bias_score", 0)
print(f"Bias Score: {bias_score:.3f} (Lower is better)")
recommendations = results.get("recommendations", [])
print(f"Recommendations Generated: {len(recommendations)}")
if recommendations:
print(f"Top Priority: {recommendations[0]['title']} ({recommendations[0]['priority'].title()})")
except Exception as e:
print(f"Error during analysis: {e}")
sys.exit(1)
if __name__ == "__main__":
main()
FILE:loop_designer.py
#!/usr/bin/env python3
"""
Interview Loop Designer
Generates calibrated interview loops tailored to specific roles, levels, and teams.
Creates complete interview loops with rounds, focus areas, time allocation,
interviewer skill requirements, and scorecard templates.
Usage:
python loop_designer.py --role "Senior Software Engineer" --level senior --team platform
python loop_designer.py --role "Product Manager" --level mid --competencies leadership,strategy
python loop_designer.py --input role_definition.json --output loops/
"""
import os
import sys
import json
import argparse
from datetime import datetime, timedelta
from typing import Dict, List, Optional, Any, Tuple
from collections import defaultdict
class InterviewLoopDesigner:
"""Designs comprehensive interview loops based on role requirements."""
def __init__(self):
self.competency_frameworks = self._init_competency_frameworks()
self.role_templates = self._init_role_templates()
self.interviewer_skills = self._init_interviewer_skills()
def _init_competency_frameworks(self) -> Dict[str, Dict]:
"""Initialize competency frameworks for different roles."""
return {
"software_engineer": {
"junior": {
"required": ["coding_fundamentals", "debugging", "testing_basics", "version_control"],
"preferred": ["system_understanding", "code_review", "collaboration"],
"focus_areas": ["technical_execution", "learning_agility", "team_collaboration"]
},
"mid": {
"required": ["advanced_coding", "system_design_basics", "testing_strategy", "debugging_complex"],
"preferred": ["mentoring_basics", "technical_communication", "project_ownership"],
"focus_areas": ["technical_depth", "system_thinking", "ownership"]
},
"senior": {
"required": ["system_architecture", "technical_leadership", "mentoring", "cross_team_collab"],
"preferred": ["technology_evaluation", "process_improvement", "hiring_contribution"],
"focus_areas": ["technical_leadership", "system_architecture", "people_development"]
},
"staff": {
"required": ["architectural_vision", "organizational_impact", "technical_strategy", "team_building"],
"preferred": ["industry_influence", "innovation_leadership", "executive_communication"],
"focus_areas": ["organizational_impact", "technical_vision", "strategic_influence"]
},
"principal": {
"required": ["company_wide_impact", "technical_vision", "talent_development", "strategic_planning"],
"preferred": ["industry_leadership", "board_communication", "market_influence"],
"focus_areas": ["strategic_leadership", "organizational_transformation", "external_influence"]
}
},
"product_manager": {
"junior": {
"required": ["product_execution", "user_research", "data_analysis", "stakeholder_comm"],
"preferred": ["market_awareness", "technical_understanding", "project_management"],
"focus_areas": ["execution_excellence", "user_focus", "analytical_thinking"]
},
"mid": {
"required": ["product_strategy", "cross_functional_leadership", "metrics_design", "market_analysis"],
"preferred": ["team_building", "technical_collaboration", "competitive_analysis"],
"focus_areas": ["strategic_thinking", "leadership", "business_impact"]
},
"senior": {
"required": ["business_strategy", "team_leadership", "p&l_ownership", "market_positioning"],
"preferred": ["hiring_leadership", "board_communication", "partnership_development"],
"focus_areas": ["business_leadership", "market_strategy", "organizational_impact"]
},
"staff": {
"required": ["portfolio_management", "organizational_leadership", "strategic_planning", "market_creation"],
"preferred": ["executive_presence", "investor_relations", "acquisition_strategy"],
"focus_areas": ["strategic_leadership", "market_innovation", "organizational_transformation"]
}
},
"designer": {
"junior": {
"required": ["design_fundamentals", "user_research", "prototyping", "design_tools"],
"preferred": ["user_empathy", "visual_design", "collaboration"],
"focus_areas": ["design_execution", "user_research", "creative_problem_solving"]
},
"mid": {
"required": ["design_systems", "user_testing", "cross_functional_collab", "design_strategy"],
"preferred": ["mentoring", "process_improvement", "business_understanding"],
"focus_areas": ["design_leadership", "system_thinking", "business_impact"]
},
"senior": {
"required": ["design_leadership", "team_building", "strategic_design", "stakeholder_management"],
"preferred": ["design_culture", "hiring_leadership", "executive_communication"],
"focus_areas": ["design_strategy", "team_leadership", "organizational_impact"]
}
},
"data_scientist": {
"junior": {
"required": ["statistical_analysis", "python_r", "data_visualization", "sql"],
"preferred": ["machine_learning", "business_understanding", "communication"],
"focus_areas": ["analytical_skills", "technical_execution", "business_impact"]
},
"mid": {
"required": ["advanced_ml", "experiment_design", "data_engineering", "stakeholder_comm"],
"preferred": ["mentoring", "project_leadership", "product_collaboration"],
"focus_areas": ["advanced_analytics", "project_leadership", "cross_functional_impact"]
},
"senior": {
"required": ["data_strategy", "team_leadership", "ml_systems", "business_strategy"],
"preferred": ["hiring_leadership", "executive_communication", "technology_evaluation"],
"focus_areas": ["strategic_leadership", "technical_vision", "organizational_impact"]
}
},
"devops_engineer": {
"junior": {
"required": ["infrastructure_basics", "scripting", "monitoring", "troubleshooting"],
"preferred": ["automation", "cloud_platforms", "security_awareness"],
"focus_areas": ["operational_excellence", "automation_mindset", "problem_solving"]
},
"mid": {
"required": ["ci_cd_design", "infrastructure_as_code", "security_implementation", "performance_optimization"],
"preferred": ["team_collaboration", "incident_management", "capacity_planning"],
"focus_areas": ["system_reliability", "automation_leadership", "cross_team_collaboration"]
},
"senior": {
"required": ["platform_architecture", "team_leadership", "security_strategy", "organizational_impact"],
"preferred": ["hiring_contribution", "technology_evaluation", "executive_communication"],
"focus_areas": ["platform_leadership", "strategic_thinking", "organizational_transformation"]
}
},
"engineering_manager": {
"junior": {
"required": ["team_leadership", "technical_background", "people_management", "project_coordination"],
"preferred": ["hiring_experience", "performance_management", "technical_mentoring"],
"focus_areas": ["people_leadership", "team_building", "execution_excellence"]
},
"senior": {
"required": ["organizational_leadership", "strategic_planning", "talent_development", "cross_functional_leadership"],
"preferred": ["technical_vision", "culture_building", "executive_communication"],
"focus_areas": ["organizational_impact", "strategic_leadership", "talent_development"]
},
"staff": {
"required": ["multi_team_leadership", "organizational_strategy", "executive_presence", "cultural_transformation"],
"preferred": ["board_communication", "market_understanding", "acquisition_integration"],
"focus_areas": ["organizational_transformation", "strategic_leadership", "cultural_evolution"]
}
}
}
def _init_role_templates(self) -> Dict[str, Dict]:
"""Initialize role-specific interview templates."""
return {
"software_engineer": {
"core_rounds": ["technical_phone_screen", "coding_deep_dive", "system_design", "behavioral"],
"optional_rounds": ["technical_leadership", "domain_expertise", "culture_fit"],
"total_duration_range": (180, 360), # 3-6 hours
"required_competencies": ["coding", "problem_solving", "communication"]
},
"product_manager": {
"core_rounds": ["product_sense", "analytical_thinking", "execution_process", "behavioral"],
"optional_rounds": ["strategic_thinking", "technical_collaboration", "leadership"],
"total_duration_range": (180, 300), # 3-5 hours
"required_competencies": ["product_strategy", "analytical_thinking", "stakeholder_management"]
},
"designer": {
"core_rounds": ["portfolio_review", "design_challenge", "collaboration_process", "behavioral"],
"optional_rounds": ["design_system_thinking", "research_methodology", "leadership"],
"total_duration_range": (180, 300), # 3-5 hours
"required_competencies": ["design_process", "user_empathy", "visual_communication"]
},
"data_scientist": {
"core_rounds": ["technical_assessment", "case_study", "statistical_thinking", "behavioral"],
"optional_rounds": ["ml_systems", "business_strategy", "technical_leadership"],
"total_duration_range": (210, 330), # 3.5-5.5 hours
"required_competencies": ["statistical_analysis", "programming", "business_acumen"]
},
"devops_engineer": {
"core_rounds": ["technical_assessment", "system_design", "troubleshooting", "behavioral"],
"optional_rounds": ["security_assessment", "automation_design", "leadership"],
"total_duration_range": (180, 300), # 3-5 hours
"required_competencies": ["infrastructure", "automation", "problem_solving"]
},
"engineering_manager": {
"core_rounds": ["leadership_assessment", "technical_background", "people_management", "behavioral"],
"optional_rounds": ["strategic_thinking", "hiring_assessment", "culture_building"],
"total_duration_range": (240, 360), # 4-6 hours
"required_competencies": ["people_leadership", "technical_understanding", "strategic_thinking"]
}
}
def _init_interviewer_skills(self) -> Dict[str, Dict]:
"""Initialize interviewer skill requirements for different round types."""
return {
"technical_phone_screen": {
"required_skills": ["technical_assessment", "coding_evaluation"],
"preferred_experience": ["same_domain", "senior_level"],
"calibration_level": "standard"
},
"coding_deep_dive": {
"required_skills": ["advanced_technical", "code_quality_assessment"],
"preferred_experience": ["senior_engineer", "system_design"],
"calibration_level": "high"
},
"system_design": {
"required_skills": ["architecture_design", "scalability_assessment"],
"preferred_experience": ["senior_architect", "large_scale_systems"],
"calibration_level": "high"
},
"behavioral": {
"required_skills": ["behavioral_interviewing", "competency_assessment"],
"preferred_experience": ["hiring_manager", "people_leadership"],
"calibration_level": "standard"
},
"technical_leadership": {
"required_skills": ["leadership_assessment", "technical_mentoring"],
"preferred_experience": ["engineering_manager", "tech_lead"],
"calibration_level": "high"
},
"product_sense": {
"required_skills": ["product_evaluation", "market_analysis"],
"preferred_experience": ["product_manager", "product_leadership"],
"calibration_level": "high"
},
"analytical_thinking": {
"required_skills": ["data_analysis", "metrics_evaluation"],
"preferred_experience": ["data_analyst", "product_manager"],
"calibration_level": "standard"
},
"design_challenge": {
"required_skills": ["design_evaluation", "user_experience"],
"preferred_experience": ["senior_designer", "design_manager"],
"calibration_level": "high"
}
}
def generate_interview_loop(self, role: str, level: str, team: Optional[str] = None,
competencies: Optional[List[str]] = None) -> Dict[str, Any]:
"""Generate a complete interview loop for the specified role and level."""
# Normalize inputs
role_key = role.lower().replace(" ", "_").replace("-", "_")
level_key = level.lower()
# Get role template and competency requirements
if role_key not in self.competency_frameworks:
role_key = self._find_closest_role(role_key)
if level_key not in self.competency_frameworks[role_key]:
level_key = self._find_closest_level(role_key, level_key)
competency_req = self.competency_frameworks[role_key][level_key]
role_template = self.role_templates.get(role_key, self.role_templates["software_engineer"])
# Design the interview loop
rounds = self._design_rounds(role_key, level_key, competency_req, role_template, competencies)
schedule = self._create_schedule(rounds)
scorecard = self._generate_scorecard(role_key, level_key, competency_req)
interviewer_requirements = self._define_interviewer_requirements(rounds)
return {
"role": role,
"level": level,
"team": team,
"generated_at": datetime.now().isoformat(),
"total_duration_minutes": sum(round_info["duration_minutes"] for round_info in rounds.values()),
"total_rounds": len(rounds),
"rounds": rounds,
"suggested_schedule": schedule,
"scorecard_template": scorecard,
"interviewer_requirements": interviewer_requirements,
"competency_framework": competency_req,
"calibration_notes": self._generate_calibration_notes(role_key, level_key)
}
def _find_closest_role(self, role_key: str) -> str:
"""Find the closest matching role template."""
role_mappings = {
"engineer": "software_engineer",
"developer": "software_engineer",
"swe": "software_engineer",
"backend": "software_engineer",
"frontend": "software_engineer",
"fullstack": "software_engineer",
"pm": "product_manager",
"product": "product_manager",
"ux": "designer",
"ui": "designer",
"graphic": "designer",
"data": "data_scientist",
"analyst": "data_scientist",
"ml": "data_scientist",
"ops": "devops_engineer",
"sre": "devops_engineer",
"infrastructure": "devops_engineer",
"manager": "engineering_manager",
"lead": "engineering_manager"
}
for key_part in role_key.split("_"):
if key_part in role_mappings:
return role_mappings[key_part]
return "software_engineer" # Default fallback
def _find_closest_level(self, role_key: str, level_key: str) -> str:
"""Find the closest matching level for the role."""
available_levels = list(self.competency_frameworks[role_key].keys())
level_mappings = {
"entry": "junior",
"associate": "junior",
"jr": "junior",
"mid": "mid",
"middle": "mid",
"sr": "senior",
"senior": "senior",
"staff": "staff",
"principal": "principal",
"lead": "senior",
"manager": "senior"
}
mapped_level = level_mappings.get(level_key, level_key)
if mapped_level in available_levels:
return mapped_level
elif "senior" in available_levels:
return "senior"
else:
return available_levels[0]
def _design_rounds(self, role_key: str, level_key: str, competency_req: Dict,
role_template: Dict, custom_competencies: Optional[List[str]]) -> Dict[str, Dict]:
"""Design the specific interview rounds based on role and level."""
rounds = {}
# Determine which rounds to include
core_rounds = role_template["core_rounds"].copy()
optional_rounds = role_template["optional_rounds"].copy()
# Add optional rounds based on level
if level_key in ["senior", "staff", "principal"]:
if "technical_leadership" in optional_rounds and role_key in ["software_engineer", "engineering_manager"]:
core_rounds.append("technical_leadership")
if "strategic_thinking" in optional_rounds and role_key in ["product_manager", "engineering_manager"]:
core_rounds.append("strategic_thinking")
if "design_system_thinking" in optional_rounds and role_key == "designer":
core_rounds.append("design_system_thinking")
if level_key in ["staff", "principal"]:
if "domain_expertise" in optional_rounds:
core_rounds.append("domain_expertise")
# Define round details
round_definitions = self._get_round_definitions()
for i, round_type in enumerate(core_rounds, 1):
if round_type in round_definitions:
round_def = round_definitions[round_type].copy()
round_def["order"] = i
round_def["focus_areas"] = self._customize_focus_areas(round_type, competency_req, custom_competencies)
rounds[f"round_{i}_{round_type}"] = round_def
return rounds
def _get_round_definitions(self) -> Dict[str, Dict]:
"""Get predefined round definitions with standard durations and formats."""
return {
"technical_phone_screen": {
"name": "Technical Phone Screen",
"duration_minutes": 45,
"format": "virtual",
"objectives": ["Assess coding fundamentals", "Evaluate problem-solving approach", "Screen for basic technical competency"],
"question_types": ["coding_problems", "technical_concepts", "experience_questions"],
"evaluation_criteria": ["technical_accuracy", "problem_solving_process", "communication_clarity"]
},
"coding_deep_dive": {
"name": "Coding Deep Dive",
"duration_minutes": 75,
"format": "in_person_or_virtual",
"objectives": ["Evaluate coding skills in depth", "Assess code quality and testing", "Review debugging approach"],
"question_types": ["complex_coding_problems", "code_review", "testing_strategy"],
"evaluation_criteria": ["code_quality", "testing_approach", "debugging_skills", "optimization_thinking"]
},
"system_design": {
"name": "System Design",
"duration_minutes": 75,
"format": "collaborative_whiteboard",
"objectives": ["Assess architectural thinking", "Evaluate scalability considerations", "Review trade-off analysis"],
"question_types": ["system_architecture", "scalability_design", "trade_off_analysis"],
"evaluation_criteria": ["architectural_thinking", "scalability_awareness", "trade_off_reasoning"]
},
"behavioral": {
"name": "Behavioral Interview",
"duration_minutes": 45,
"format": "conversational",
"objectives": ["Assess cultural fit", "Evaluate past experiences", "Review leadership examples"],
"question_types": ["star_method_questions", "situational_scenarios", "values_alignment"],
"evaluation_criteria": ["communication_skills", "leadership_examples", "cultural_alignment"]
},
"technical_leadership": {
"name": "Technical Leadership",
"duration_minutes": 60,
"format": "discussion_based",
"objectives": ["Evaluate mentoring capability", "Assess technical decision making", "Review cross-team collaboration"],
"question_types": ["leadership_scenarios", "technical_decisions", "mentoring_examples"],
"evaluation_criteria": ["leadership_potential", "technical_judgment", "influence_skills"]
},
"product_sense": {
"name": "Product Sense",
"duration_minutes": 75,
"format": "case_study",
"objectives": ["Assess product intuition", "Evaluate user empathy", "Review market understanding"],
"question_types": ["product_scenarios", "feature_prioritization", "user_journey_analysis"],
"evaluation_criteria": ["product_intuition", "user_empathy", "analytical_thinking"]
},
"analytical_thinking": {
"name": "Analytical Thinking",
"duration_minutes": 60,
"format": "data_analysis",
"objectives": ["Evaluate data interpretation", "Assess metric design", "Review experiment planning"],
"question_types": ["data_interpretation", "metric_design", "experiment_analysis"],
"evaluation_criteria": ["analytical_rigor", "metric_intuition", "experimental_thinking"]
},
"design_challenge": {
"name": "Design Challenge",
"duration_minutes": 90,
"format": "hands_on_design",
"objectives": ["Assess design process", "Evaluate user-centered thinking", "Review iteration approach"],
"question_types": ["design_problems", "user_research", "design_critique"],
"evaluation_criteria": ["design_process", "user_focus", "visual_communication"]
},
"portfolio_review": {
"name": "Portfolio Review",
"duration_minutes": 75,
"format": "presentation_discussion",
"objectives": ["Review past work", "Assess design thinking", "Evaluate impact measurement"],
"question_types": ["portfolio_walkthrough", "design_decisions", "impact_stories"],
"evaluation_criteria": ["design_quality", "process_thinking", "business_impact"]
}
}
def _customize_focus_areas(self, round_type: str, competency_req: Dict,
custom_competencies: Optional[List[str]]) -> List[str]:
"""Customize focus areas based on role competency requirements."""
base_focus_areas = competency_req.get("focus_areas", [])
round_focus_mapping = {
"technical_phone_screen": ["coding_fundamentals", "problem_solving"],
"coding_deep_dive": ["technical_execution", "code_quality"],
"system_design": ["system_thinking", "architectural_reasoning"],
"behavioral": ["cultural_fit", "communication", "teamwork"],
"technical_leadership": ["leadership", "mentoring", "influence"],
"product_sense": ["product_intuition", "user_empathy"],
"analytical_thinking": ["data_analysis", "metric_design"],
"design_challenge": ["design_process", "user_focus"]
}
focus_areas = round_focus_mapping.get(round_type, [])
# Add custom competencies if specified
if custom_competencies:
focus_areas.extend([comp for comp in custom_competencies if comp not in focus_areas])
# Add role-specific focus areas
focus_areas.extend([area for area in base_focus_areas if area not in focus_areas])
return focus_areas[:5] # Limit to top 5 focus areas
def _create_schedule(self, rounds: Dict[str, Dict]) -> Dict[str, Any]:
"""Create a suggested interview schedule."""
sorted_rounds = sorted(rounds.items(), key=lambda x: x[1]["order"])
# Calculate optimal scheduling
total_duration = sum(round_info["duration_minutes"] for _, round_info in sorted_rounds)
if total_duration <= 240: # 4 hours or less - single day
schedule_type = "single_day"
day_structure = self._create_single_day_schedule(sorted_rounds)
else: # Multi-day schedule
schedule_type = "multi_day"
day_structure = self._create_multi_day_schedule(sorted_rounds)
return {
"type": schedule_type,
"total_duration_minutes": total_duration,
"recommended_breaks": self._calculate_breaks(total_duration),
"day_structure": day_structure,
"logistics_notes": self._generate_logistics_notes(sorted_rounds)
}
def _create_single_day_schedule(self, rounds: List[Tuple[str, Dict]]) -> Dict[str, Any]:
"""Create a single-day interview schedule."""
start_time = datetime.strptime("09:00", "%H:%M")
current_time = start_time
schedule = []
for round_name, round_info in rounds:
# Add break if needed (after 90 minutes of interviews)
if schedule and sum(item.get("duration_minutes", 0) for item in schedule if "break" not in item.get("type", "")) >= 90:
schedule.append({
"type": "break",
"start_time": current_time.strftime("%H:%M"),
"duration_minutes": 15,
"end_time": (current_time + timedelta(minutes=15)).strftime("%H:%M")
})
current_time += timedelta(minutes=15)
# Add the interview round
end_time = current_time + timedelta(minutes=round_info["duration_minutes"])
schedule.append({
"type": "interview",
"round_name": round_name,
"title": round_info["name"],
"start_time": current_time.strftime("%H:%M"),
"end_time": end_time.strftime("%H:%M"),
"duration_minutes": round_info["duration_minutes"],
"format": round_info["format"]
})
current_time = end_time
return {
"day_1": {
"date": "TBD",
"start_time": start_time.strftime("%H:%M"),
"end_time": current_time.strftime("%H:%M"),
"rounds": schedule
}
}
def _create_multi_day_schedule(self, rounds: List[Tuple[str, Dict]]) -> Dict[str, Any]:
"""Create a multi-day interview schedule."""
# Split rounds across days (max 4 hours per day)
max_daily_minutes = 240
days = {}
current_day = 1
current_day_duration = 0
current_day_rounds = []
for round_name, round_info in rounds:
duration = round_info["duration_minutes"] + 15 # Add buffer time
if current_day_duration + duration > max_daily_minutes and current_day_rounds:
# Finalize current day
days[f"day_{current_day}"] = self._finalize_day_schedule(current_day_rounds)
current_day += 1
current_day_duration = 0
current_day_rounds = []
current_day_rounds.append((round_name, round_info))
current_day_duration += duration
# Finalize last day
if current_day_rounds:
days[f"day_{current_day}"] = self._finalize_day_schedule(current_day_rounds)
return days
def _finalize_day_schedule(self, day_rounds: List[Tuple[str, Dict]]) -> Dict[str, Any]:
"""Finalize the schedule for a specific day."""
start_time = datetime.strptime("09:00", "%H:%M")
current_time = start_time
schedule = []
for round_name, round_info in day_rounds:
end_time = current_time + timedelta(minutes=round_info["duration_minutes"])
schedule.append({
"type": "interview",
"round_name": round_name,
"title": round_info["name"],
"start_time": current_time.strftime("%H:%M"),
"end_time": end_time.strftime("%H:%M"),
"duration_minutes": round_info["duration_minutes"],
"format": round_info["format"]
})
current_time = end_time + timedelta(minutes=15) # 15-min buffer
return {
"date": "TBD",
"start_time": start_time.strftime("%H:%M"),
"end_time": (current_time - timedelta(minutes=15)).strftime("%H:%M"),
"rounds": schedule
}
def _calculate_breaks(self, total_duration: int) -> List[Dict[str, Any]]:
"""Calculate recommended breaks based on total duration."""
breaks = []
if total_duration >= 120: # 2+ hours
breaks.append({"type": "short_break", "duration": 15, "after_minutes": 90})
if total_duration >= 240: # 4+ hours
breaks.append({"type": "lunch_break", "duration": 60, "after_minutes": 180})
if total_duration >= 360: # 6+ hours
breaks.append({"type": "short_break", "duration": 15, "after_minutes": 300})
return breaks
def _generate_scorecard(self, role_key: str, level_key: str, competency_req: Dict) -> Dict[str, Any]:
"""Generate a scorecard template for the interview loop."""
scoring_dimensions = []
# Add competency-based scoring dimensions
for competency in competency_req["required"]:
scoring_dimensions.append({
"dimension": competency,
"weight": "high",
"scale": "1-4",
"description": f"Assessment of {competency.replace('_', ' ')} competency"
})
for competency in competency_req.get("preferred", []):
scoring_dimensions.append({
"dimension": competency,
"weight": "medium",
"scale": "1-4",
"description": f"Assessment of {competency.replace('_', ' ')} competency"
})
# Add standard dimensions
standard_dimensions = [
{"dimension": "communication", "weight": "high", "scale": "1-4"},
{"dimension": "cultural_fit", "weight": "medium", "scale": "1-4"},
{"dimension": "learning_agility", "weight": "medium", "scale": "1-4"}
]
scoring_dimensions.extend(standard_dimensions)
return {
"scoring_scale": {
"4": "Exceeds Expectations - Demonstrates mastery beyond required level",
"3": "Meets Expectations - Solid performance meeting all requirements",
"2": "Partially Meets - Shows potential but has development areas",
"1": "Does Not Meet - Significant gaps in required competencies"
},
"dimensions": scoring_dimensions,
"overall_recommendation": {
"options": ["Strong Hire", "Hire", "No Hire", "Strong No Hire"],
"criteria": "Based on weighted average and minimum thresholds"
},
"calibration_notes": {
"required": True,
"min_length": 100,
"sections": ["strengths", "areas_for_development", "specific_examples"]
}
}
def _define_interviewer_requirements(self, rounds: Dict[str, Dict]) -> Dict[str, Dict]:
"""Define interviewer skill requirements for each round."""
requirements = {}
for round_name, round_info in rounds.items():
round_type = round_name.split("_", 2)[-1] # Extract round type
if round_type in self.interviewer_skills:
skill_req = self.interviewer_skills[round_type].copy()
skill_req["suggested_interviewers"] = self._suggest_interviewer_profiles(round_type)
requirements[round_name] = skill_req
else:
# Default requirements
requirements[round_name] = {
"required_skills": ["interviewing_basics", "evaluation_skills"],
"preferred_experience": ["relevant_domain"],
"calibration_level": "standard",
"suggested_interviewers": ["experienced_interviewer"]
}
return requirements
def _suggest_interviewer_profiles(self, round_type: str) -> List[str]:
"""Suggest specific interviewer profiles for different round types."""
profile_mapping = {
"technical_phone_screen": ["senior_engineer", "tech_lead"],
"coding_deep_dive": ["senior_engineer", "staff_engineer"],
"system_design": ["senior_architect", "staff_engineer"],
"behavioral": ["hiring_manager", "people_manager"],
"technical_leadership": ["engineering_manager", "senior_staff"],
"product_sense": ["senior_pm", "product_leader"],
"analytical_thinking": ["senior_analyst", "data_scientist"],
"design_challenge": ["senior_designer", "design_manager"]
}
return profile_mapping.get(round_type, ["experienced_interviewer"])
def _generate_calibration_notes(self, role_key: str, level_key: str) -> Dict[str, Any]:
"""Generate calibration notes and best practices."""
return {
"hiring_bar_notes": f"Calibrated for {level_key} level {role_key.replace('_', ' ')} role",
"common_pitfalls": [
"Avoid comparing candidates to each other rather than to the role standard",
"Don't let one strong/weak area overshadow overall assessment",
"Ensure consistent application of evaluation criteria"
],
"calibration_checkpoints": [
"Review score distribution after every 5 candidates",
"Conduct monthly interviewer calibration sessions",
"Track correlation with 6-month performance reviews"
],
"escalation_criteria": [
"Any candidate receiving all 4s or all 1s",
"Significant disagreement between interviewers (>1.5 point spread)",
"Unusual circumstances or accommodations needed"
]
}
def _generate_logistics_notes(self, rounds: List[Tuple[str, Dict]]) -> List[str]:
"""Generate logistics and coordination notes."""
notes = [
"Coordinate interviewer availability before scheduling",
"Ensure all interviewers have access to job description and competency requirements",
"Prepare interview rooms/virtual links for all rounds",
"Share candidate resume and application with all interviewers"
]
# Add format-specific notes
formats_used = {round_info["format"] for _, round_info in rounds}
if "virtual" in formats_used:
notes.append("Test video conferencing setup before virtual interviews")
notes.append("Share virtual meeting links with candidate 24 hours in advance")
if "collaborative_whiteboard" in formats_used:
notes.append("Prepare whiteboard or collaborative online tool for design sessions")
if "hands_on_design" in formats_used:
notes.append("Provide design tools access or ensure candidate can screen share their preferred tools")
return notes
def format_human_readable(loop_data: Dict[str, Any]) -> str:
"""Format the interview loop data in a human-readable format."""
output = []
# Header
output.append(f"Interview Loop Design for {loop_data['role']} ({loop_data['level'].title()} Level)")
output.append("=" * 60)
if loop_data.get('team'):
output.append(f"Team: {loop_data['team']}")
output.append(f"Generated: {loop_data['generated_at']}")
output.append(f"Total Duration: {loop_data['total_duration_minutes']} minutes ({loop_data['total_duration_minutes']//60}h {loop_data['total_duration_minutes']%60}m)")
output.append(f"Total Rounds: {loop_data['total_rounds']}")
output.append("")
# Interview Rounds
output.append("INTERVIEW ROUNDS")
output.append("-" * 40)
sorted_rounds = sorted(loop_data['rounds'].items(), key=lambda x: x[1]['order'])
for round_name, round_info in sorted_rounds:
output.append(f"\nRound {round_info['order']}: {round_info['name']}")
output.append(f"Duration: {round_info['duration_minutes']} minutes")
output.append(f"Format: {round_info['format'].replace('_', ' ').title()}")
output.append("Objectives:")
for obj in round_info['objectives']:
output.append(f" • {obj}")
output.append("Focus Areas:")
for area in round_info['focus_areas']:
output.append(f" • {area.replace('_', ' ').title()}")
# Suggested Schedule
output.append("\nSUGGESTED SCHEDULE")
output.append("-" * 40)
schedule = loop_data['suggested_schedule']
output.append(f"Schedule Type: {schedule['type'].replace('_', ' ').title()}")
for day_name, day_info in schedule['day_structure'].items():
output.append(f"\n{day_name.replace('_', ' ').title()}:")
output.append(f"Time: {day_info['start_time']} - {day_info['end_time']}")
for item in day_info['rounds']:
if item['type'] == 'interview':
output.append(f" {item['start_time']}-{item['end_time']}: {item['title']} ({item['duration_minutes']}min)")
else:
output.append(f" {item['start_time']}-{item['end_time']}: {item['type'].title()} ({item['duration_minutes']}min)")
# Interviewer Requirements
output.append("\nINTERVIEWER REQUIREMENTS")
output.append("-" * 40)
for round_name, requirements in loop_data['interviewer_requirements'].items():
round_display = round_name.split("_", 2)[-1].replace("_", " ").title()
output.append(f"\n{round_display}:")
output.append(f"Required Skills: {', '.join(requirements['required_skills'])}")
output.append(f"Suggested Interviewers: {', '.join(requirements['suggested_interviewers'])}")
output.append(f"Calibration Level: {requirements['calibration_level'].title()}")
# Scorecard Overview
output.append("\nSCORECARD TEMPLATE")
output.append("-" * 40)
scorecard = loop_data['scorecard_template']
output.append("Scoring Scale:")
for score, description in scorecard['scoring_scale'].items():
output.append(f" {score}: {description}")
output.append("\nEvaluation Dimensions:")
for dim in scorecard['dimensions']:
output.append(f" • {dim['dimension'].replace('_', ' ').title()} (Weight: {dim['weight']})")
# Calibration Notes
output.append("\nCALIBRATION NOTES")
output.append("-" * 40)
calibration = loop_data['calibration_notes']
output.append(f"Hiring Bar: {calibration['hiring_bar_notes']}")
output.append("\nCommon Pitfalls:")
for pitfall in calibration['common_pitfalls']:
output.append(f" • {pitfall}")
return "\n".join(output)
def main():
parser = argparse.ArgumentParser(description="Generate calibrated interview loops for specific roles and levels")
parser.add_argument("--role", type=str, help="Job role title (e.g., 'Senior Software Engineer')")
parser.add_argument("--level", type=str, help="Experience level (junior, mid, senior, staff, principal)")
parser.add_argument("--team", type=str, help="Team or department (optional)")
parser.add_argument("--competencies", type=str, help="Comma-separated list of specific competencies to focus on")
parser.add_argument("--input", type=str, help="Input JSON file with role definition")
parser.add_argument("--output", type=str, help="Output directory or file path")
parser.add_argument("--format", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
designer = InterviewLoopDesigner()
# Handle input
if args.input:
try:
with open(args.input, 'r') as f:
role_data = json.load(f)
role = role_data.get('role') or role_data.get('title', '')
level = role_data.get('level', 'senior')
team = role_data.get('team')
competencies = role_data.get('competencies')
except Exception as e:
print(f"Error reading input file: {e}")
sys.exit(1)
else:
if not args.role or not args.level:
print("Error: --role and --level are required when not using --input")
sys.exit(1)
role = args.role
level = args.level
team = args.team
competencies = args.competencies.split(',') if args.competencies else None
# Generate interview loop
try:
loop_data = designer.generate_interview_loop(role, level, team, competencies)
# Handle output
if args.output:
output_path = args.output
if os.path.isdir(output_path):
safe_role = "".join(c for c in role.lower() if c.isalnum() or c in (' ', '-', '_')).replace(' ', '_')
base_filename = f"{safe_role}_{level}_interview_loop"
json_path = os.path.join(output_path, f"{base_filename}.json")
text_path = os.path.join(output_path, f"{base_filename}.txt")
else:
# Use provided path as base
json_path = output_path if output_path.endswith('.json') else f"{output_path}.json"
text_path = output_path.replace('.json', '.txt') if output_path.endswith('.json') else f"{output_path}.txt"
else:
safe_role = "".join(c for c in role.lower() if c.isalnum() or c in (' ', '-', '_')).replace(' ', '_')
base_filename = f"{safe_role}_{level}_interview_loop"
json_path = f"{base_filename}.json"
text_path = f"{base_filename}.txt"
# Write outputs
if args.format in ["json", "both"]:
with open(json_path, 'w') as f:
json.dump(loop_data, f, indent=2, default=str)
print(f"JSON output written to: {json_path}")
if args.format in ["text", "both"]:
with open(text_path, 'w') as f:
f.write(format_human_readable(loop_data))
print(f"Text output written to: {text_path}")
# Always print summary to stdout
print("\nInterview Loop Summary:")
print(f"Role: {loop_data['role']} ({loop_data['level'].title()})")
print(f"Total Duration: {loop_data['total_duration_minutes']} minutes")
print(f"Number of Rounds: {loop_data['total_rounds']}")
print(f"Schedule Type: {loop_data['suggested_schedule']['type'].replace('_', ' ').title()}")
except Exception as e:
print(f"Error generating interview loop: {e}")
sys.exit(1)
if __name__ == "__main__":
main()
FILE:question_bank_generator.py
#!/usr/bin/env python3
"""
Question Bank Generator
Generates comprehensive, competency-based interview questions with detailed scoring criteria.
Creates structured question banks organized by competency area with scoring rubrics,
follow-up probes, and calibration examples.
Usage:
python question_bank_generator.py --role "Frontend Engineer" --competencies react,typescript,system-design
python question_bank_generator.py --role "Product Manager" --question-types behavioral,leadership
python question_bank_generator.py --input role_requirements.json --output questions/
"""
import os
import sys
import json
import argparse
import random
from datetime import datetime
from typing import Dict, List, Optional, Any, Tuple
from collections import defaultdict
class QuestionBankGenerator:
"""Generates comprehensive interview question banks with scoring criteria."""
def __init__(self):
self.technical_questions = self._init_technical_questions()
self.behavioral_questions = self._init_behavioral_questions()
self.competency_mapping = self._init_competency_mapping()
self.scoring_rubrics = self._init_scoring_rubrics()
self.follow_up_strategies = self._init_follow_up_strategies()
def _init_technical_questions(self) -> Dict[str, Dict]:
"""Initialize technical questions by competency area and level."""
return {
"coding_fundamentals": {
"junior": [
{
"question": "Write a function to reverse a string without using built-in reverse methods.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "easy",
"time_limit": 15,
"key_concepts": ["loops", "string_manipulation", "basic_algorithms"]
},
{
"question": "Implement a function to check if a string is a palindrome.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "easy",
"time_limit": 15,
"key_concepts": ["string_processing", "comparison", "edge_cases"]
},
{
"question": "Find the largest element in an array without using built-in max functions.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "easy",
"time_limit": 10,
"key_concepts": ["arrays", "iteration", "comparison"]
}
],
"mid": [
{
"question": "Implement a function to find the first non-repeating character in a string.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "medium",
"time_limit": 20,
"key_concepts": ["hash_maps", "string_processing", "efficiency"]
},
{
"question": "Write a function to merge two sorted arrays into one sorted array.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "medium",
"time_limit": 25,
"key_concepts": ["merge_algorithms", "two_pointers", "optimization"]
}
],
"senior": [
{
"question": "Implement a LRU (Least Recently Used) cache with O(1) operations.",
"competency": "coding_fundamentals",
"type": "coding",
"difficulty": "hard",
"time_limit": 35,
"key_concepts": ["data_structures", "hash_maps", "doubly_linked_lists"]
}
]
},
"system_design": {
"mid": [
{
"question": "Design a URL shortener service like bit.ly for 10K users.",
"competency": "system_design",
"type": "design",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["database_design", "hashing", "basic_scalability"]
}
],
"senior": [
{
"question": "Design a real-time chat system supporting 1M concurrent users.",
"competency": "system_design",
"type": "design",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["websockets", "load_balancing", "database_sharding", "caching"]
},
{
"question": "Design a distributed cache system like Redis with high availability.",
"competency": "system_design",
"type": "design",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["distributed_systems", "replication", "consistency", "partitioning"]
}
],
"staff": [
{
"question": "Design the architecture for a global content delivery network (CDN).",
"competency": "system_design",
"type": "design",
"difficulty": "expert",
"time_limit": 75,
"key_concepts": ["global_architecture", "edge_computing", "content_optimization", "network_protocols"]
}
]
},
"frontend_development": {
"junior": [
{
"question": "Create a responsive navigation menu using HTML, CSS, and vanilla JavaScript.",
"competency": "frontend_development",
"type": "coding",
"difficulty": "easy",
"time_limit": 30,
"key_concepts": ["html_css", "responsive_design", "dom_manipulation"]
}
],
"mid": [
{
"question": "Build a React component that fetches and displays paginated data from an API.",
"competency": "frontend_development",
"type": "coding",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["react_hooks", "api_integration", "state_management", "pagination"]
}
],
"senior": [
{
"question": "Design and implement a custom React hook for managing complex form state with validation.",
"competency": "frontend_development",
"type": "coding",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["custom_hooks", "form_validation", "state_management", "performance"]
}
]
},
"data_analysis": {
"junior": [
{
"question": "Given a dataset of user activities, calculate the daily active users for the past month.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "easy",
"time_limit": 30,
"key_concepts": ["sql_basics", "date_functions", "aggregation"]
}
],
"mid": [
{
"question": "Analyze conversion funnel data to identify the biggest drop-off point and propose solutions.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["funnel_analysis", "conversion_optimization", "statistical_significance"]
}
],
"senior": [
{
"question": "Design an A/B testing framework to measure the impact of a new recommendation algorithm.",
"competency": "data_analysis",
"type": "analytical",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["experiment_design", "statistical_power", "bias_mitigation", "causal_inference"]
}
]
},
"machine_learning": {
"mid": [
{
"question": "Explain how you would build a recommendation system for an e-commerce platform.",
"competency": "machine_learning",
"type": "conceptual",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["collaborative_filtering", "content_based", "cold_start", "evaluation_metrics"]
}
],
"senior": [
{
"question": "Design a real-time fraud detection system for financial transactions.",
"competency": "machine_learning",
"type": "design",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["anomaly_detection", "real_time_ml", "feature_engineering", "model_monitoring"]
}
]
},
"product_strategy": {
"mid": [
{
"question": "How would you prioritize features for a mobile app with limited engineering resources?",
"competency": "product_strategy",
"type": "case_study",
"difficulty": "medium",
"time_limit": 45,
"key_concepts": ["prioritization_frameworks", "resource_allocation", "impact_estimation"]
}
],
"senior": [
{
"question": "Design a go-to-market strategy for a new B2B SaaS product entering a competitive market.",
"competency": "product_strategy",
"type": "strategic",
"difficulty": "hard",
"time_limit": 60,
"key_concepts": ["market_analysis", "competitive_positioning", "pricing_strategy", "channel_strategy"]
}
]
}
}
def _init_behavioral_questions(self) -> Dict[str, List[Dict]]:
"""Initialize behavioral questions by competency area."""
return {
"leadership": [
{
"question": "Tell me about a time when you had to lead a team through a significant change or challenge.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["change_management", "team_motivation", "communication"]
},
{
"question": "Describe a situation where you had to influence someone without having direct authority over them.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["influence", "persuasion", "stakeholder_management"]
},
{
"question": "Give me an example of when you had to make a difficult decision that affected your team.",
"competency": "leadership",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["decision_making", "team_impact", "communication"]
}
],
"collaboration": [
{
"question": "Describe a time when you had to work with a difficult colleague or stakeholder.",
"competency": "collaboration",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["conflict_resolution", "relationship_building", "professionalism"]
},
{
"question": "Tell me about a project where you had to coordinate across multiple teams or departments.",
"competency": "collaboration",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["cross_functional_work", "communication", "project_coordination"]
}
],
"problem_solving": [
{
"question": "Walk me through a complex problem you solved recently. What was your approach?",
"competency": "problem_solving",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["analytical_thinking", "methodology", "creativity"]
},
{
"question": "Describe a time when you had to solve a problem with limited information or resources.",
"competency": "problem_solving",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["resourcefulness", "ambiguity_tolerance", "decision_making"]
}
],
"communication": [
{
"question": "Tell me about a time when you had to present complex technical information to a non-technical audience.",
"competency": "communication",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["technical_communication", "audience_adaptation", "clarity"]
},
{
"question": "Describe a situation where you had to deliver difficult feedback to a colleague.",
"competency": "communication",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["feedback_delivery", "empathy", "constructive_criticism"]
}
],
"adaptability": [
{
"question": "Tell me about a time when you had to quickly learn a new technology or skill for work.",
"competency": "adaptability",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["learning_agility", "growth_mindset", "knowledge_acquisition"]
},
{
"question": "Describe how you handled a situation when project requirements changed significantly mid-way.",
"competency": "adaptability",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["flexibility", "change_management", "resilience"]
}
],
"innovation": [
{
"question": "Tell me about a time when you came up with a creative solution to improve a process or solve a problem.",
"competency": "innovation",
"type": "behavioral",
"method": "STAR",
"focus_areas": ["creative_thinking", "process_improvement", "initiative"]
}
]
}
def _init_competency_mapping(self) -> Dict[str, Dict]:
"""Initialize role to competency mapping."""
return {
"software_engineer": {
"core_competencies": ["coding_fundamentals", "system_design", "problem_solving", "collaboration"],
"level_specific": {
"junior": ["coding_fundamentals", "debugging", "learning_agility"],
"mid": ["advanced_coding", "system_design", "mentoring_basics"],
"senior": ["system_architecture", "technical_leadership", "innovation"],
"staff": ["architectural_vision", "organizational_impact", "strategic_thinking"]
}
},
"frontend_engineer": {
"core_competencies": ["frontend_development", "ui_ux_understanding", "problem_solving", "collaboration"],
"level_specific": {
"junior": ["html_css_js", "responsive_design", "basic_frameworks"],
"mid": ["react_vue_angular", "state_management", "performance_optimization"],
"senior": ["frontend_architecture", "team_leadership", "cross_functional_collaboration"],
"staff": ["frontend_strategy", "technology_evaluation", "organizational_impact"]
}
},
"backend_engineer": {
"core_competencies": ["backend_development", "database_design", "api_design", "system_design"],
"level_specific": {
"junior": ["server_side_programming", "database_basics", "api_consumption"],
"mid": ["microservices", "caching", "security_basics"],
"senior": ["distributed_systems", "performance_optimization", "technical_leadership"],
"staff": ["system_architecture", "technology_strategy", "cross_team_influence"]
}
},
"product_manager": {
"core_competencies": ["product_strategy", "user_research", "data_analysis", "stakeholder_management"],
"level_specific": {
"junior": ["feature_specification", "user_stories", "basic_analytics"],
"mid": ["product_roadmap", "cross_functional_leadership", "market_research"],
"senior": ["business_strategy", "team_leadership", "p&l_responsibility"],
"staff": ["portfolio_management", "organizational_strategy", "market_creation"]
}
},
"data_scientist": {
"core_competencies": ["statistical_analysis", "machine_learning", "data_analysis", "business_acumen"],
"level_specific": {
"junior": ["python_r", "sql", "basic_ml", "data_visualization"],
"mid": ["advanced_ml", "experiment_design", "model_evaluation"],
"senior": ["ml_systems", "data_strategy", "stakeholder_communication"],
"staff": ["data_platform", "ai_strategy", "organizational_impact"]
}
},
"designer": {
"core_competencies": ["design_process", "user_research", "visual_design", "collaboration"],
"level_specific": {
"junior": ["design_tools", "user_empathy", "visual_communication"],
"mid": ["design_systems", "user_testing", "cross_functional_work"],
"senior": ["design_strategy", "team_leadership", "business_impact"],
"staff": ["design_vision", "organizational_design", "strategic_influence"]
}
},
"devops_engineer": {
"core_competencies": ["infrastructure", "automation", "monitoring", "troubleshooting"],
"level_specific": {
"junior": ["scripting", "basic_cloud", "ci_cd_basics"],
"mid": ["infrastructure_as_code", "container_orchestration", "security"],
"senior": ["platform_design", "reliability_engineering", "team_leadership"],
"staff": ["platform_strategy", "organizational_infrastructure", "technology_vision"]
}
}
}
def _init_scoring_rubrics(self) -> Dict[str, Dict]:
"""Initialize scoring rubrics for different question types."""
return {
"coding": {
"correctness": {
"4": "Solution is completely correct, handles all edge cases, optimal complexity",
"3": "Solution is correct for main cases, good complexity, minor edge case issues",
"2": "Solution works but has some bugs or suboptimal approach",
"1": "Solution has significant issues or doesn't work"
},
"code_quality": {
"4": "Clean, readable, well-structured code with excellent naming and comments",
"3": "Good code structure, readable with appropriate naming",
"2": "Code works but has style/structure issues",
"1": "Poor code quality, hard to understand"
},
"problem_solving_approach": {
"4": "Excellent problem breakdown, clear thinking process, considers alternatives",
"3": "Good approach, logical thinking, systematic problem solving",
"2": "Decent approach but some confusion or inefficiency",
"1": "Poor approach, unclear thinking process"
},
"communication": {
"4": "Excellent explanation of approach, asks clarifying questions, clear reasoning",
"3": "Good communication, explains thinking well",
"2": "Adequate communication, some explanation",
"1": "Poor communication, little explanation"
}
},
"behavioral": {
"situation_clarity": {
"4": "Clear, specific situation with relevant context and stakes",
"3": "Good situation description with adequate context",
"2": "Situation described but lacks some specifics",
"1": "Vague or unclear situation description"
},
"action_quality": {
"4": "Specific, thoughtful actions showing strong competency",
"3": "Good actions demonstrating competency",
"2": "Adequate actions but could be stronger",
"1": "Weak or inappropriate actions"
},
"result_impact": {
"4": "Significant positive impact with measurable results",
"3": "Good positive impact with clear outcomes",
"2": "Some positive impact demonstrated",
"1": "Little or no positive impact shown"
},
"self_awareness": {
"4": "Excellent self-reflection, learns from experience, acknowledges growth areas",
"3": "Good self-awareness and learning orientation",
"2": "Some self-reflection demonstrated",
"1": "Limited self-awareness or reflection"
}
},
"design": {
"system_thinking": {
"4": "Comprehensive system view, considers all components and interactions",
"3": "Good system understanding with most components identified",
"2": "Basic system thinking with some gaps",
"1": "Limited system thinking, misses key components"
},
"scalability": {
"4": "Excellent scalability considerations, multiple strategies discussed",
"3": "Good scalability awareness with practical solutions",
"2": "Basic scalability understanding",
"1": "Little to no scalability consideration"
},
"trade_offs": {
"4": "Excellent trade-off analysis, considers multiple dimensions",
"3": "Good trade-off awareness with clear reasoning",
"2": "Some trade-off consideration",
"1": "Limited trade-off analysis"
},
"technical_depth": {
"4": "Deep technical knowledge with implementation details",
"3": "Good technical knowledge with solid understanding",
"2": "Adequate technical knowledge",
"1": "Limited technical depth"
}
}
}
def _init_follow_up_strategies(self) -> Dict[str, List[str]]:
"""Initialize follow-up question strategies by competency."""
return {
"coding_fundamentals": [
"How would you optimize this solution for better time complexity?",
"What edge cases should we consider for this problem?",
"How would you test this function?",
"What would happen if the input size was very large?"
],
"system_design": [
"How would you handle if the system needed to scale 10x?",
"What would you do if one of your services went down?",
"How would you monitor this system in production?",
"What security considerations would you implement?"
],
"leadership": [
"What would you do differently if you faced this situation again?",
"How did you handle team members who were resistant to the change?",
"What metrics did you use to measure success?",
"How did you communicate progress to stakeholders?"
],
"problem_solving": [
"Walk me through your thought process step by step",
"What alternative approaches did you consider?",
"How did you validate your solution worked?",
"What did you learn from this experience?"
],
"collaboration": [
"How did you build consensus among the different stakeholders?",
"What communication channels did you use to keep everyone aligned?",
"How did you handle disagreements or conflicts?",
"What would you do to improve collaboration in the future?"
]
}
def generate_question_bank(self, role: str, level: str = "senior",
competencies: Optional[List[str]] = None,
question_types: Optional[List[str]] = None,
num_questions: int = 20) -> Dict[str, Any]:
"""Generate a comprehensive question bank for the specified role and competencies."""
# Normalize inputs
role_key = self._normalize_role(role)
level_key = level.lower()
# Get competency requirements
role_competencies = self._get_role_competencies(role_key, level_key, competencies)
# Determine question types to include
if question_types is None:
question_types = ["technical", "behavioral", "situational"]
# Generate questions
questions = self._generate_questions(role_competencies, question_types, level_key, num_questions)
# Create scoring rubrics
scoring_rubrics = self._create_scoring_rubrics(questions)
# Generate follow-up probes
follow_up_probes = self._generate_follow_up_probes(questions)
# Create calibration examples
calibration_examples = self._create_calibration_examples(questions[:5]) # Sample for first 5 questions
return {
"role": role,
"level": level,
"competencies": role_competencies,
"question_types": question_types,
"generated_at": datetime.now().isoformat(),
"total_questions": len(questions),
"questions": questions,
"scoring_rubrics": scoring_rubrics,
"follow_up_probes": follow_up_probes,
"calibration_examples": calibration_examples,
"usage_guidelines": self._generate_usage_guidelines(role_key, level_key)
}
def _normalize_role(self, role: str) -> str:
"""Normalize role name to match competency mapping keys."""
role_lower = role.lower().replace(" ", "_").replace("-", "_")
# Map variations to standard roles
role_mappings = {
"software_engineer": ["engineer", "developer", "swe", "software_developer"],
"frontend_engineer": ["frontend", "front_end", "ui_engineer", "web_developer"],
"backend_engineer": ["backend", "back_end", "server_engineer", "api_developer"],
"product_manager": ["pm", "product", "product_owner", "po"],
"data_scientist": ["ds", "data", "analyst", "ml_engineer"],
"designer": ["ux", "ui", "ux_ui", "product_designer", "visual_designer"],
"devops_engineer": ["devops", "sre", "platform_engineer", "infrastructure"]
}
for standard_role, variations in role_mappings.items():
if any(var in role_lower for var in variations):
return standard_role
# Default fallback
return "software_engineer"
def _get_role_competencies(self, role_key: str, level_key: str,
custom_competencies: Optional[List[str]]) -> List[str]:
"""Get competencies for the role and level."""
if role_key not in self.competency_mapping:
role_key = "software_engineer"
role_mapping = self.competency_mapping[role_key]
competencies = role_mapping["core_competencies"].copy()
# Add level-specific competencies
if level_key in role_mapping["level_specific"]:
competencies.extend(role_mapping["level_specific"][level_key])
elif "senior" in role_mapping["level_specific"]:
competencies.extend(role_mapping["level_specific"]["senior"])
# Add custom competencies if specified
if custom_competencies:
competencies.extend([comp.strip() for comp in custom_competencies if comp.strip() not in competencies])
return list(set(competencies)) # Remove duplicates
def _generate_questions(self, competencies: List[str], question_types: List[str],
level: str, num_questions: int) -> List[Dict[str, Any]]:
"""Generate questions based on competencies and types."""
questions = []
questions_per_competency = max(1, num_questions // len(competencies))
for competency in competencies:
competency_questions = []
# Add technical questions if requested and available
if "technical" in question_types and competency in self.technical_questions:
tech_questions = []
# Get questions for current level and below
level_order = ["junior", "mid", "senior", "staff", "principal"]
current_level_idx = level_order.index(level) if level in level_order else 2
for lvl_idx in range(current_level_idx + 1):
lvl = level_order[lvl_idx]
if lvl in self.technical_questions[competency]:
tech_questions.extend(self.technical_questions[competency][lvl])
competency_questions.extend(tech_questions[:questions_per_competency])
# Add behavioral questions if requested
if "behavioral" in question_types and competency in self.behavioral_questions:
behavioral_q = self.behavioral_questions[competency][:questions_per_competency]
competency_questions.extend(behavioral_q)
# Add situational questions (variations of behavioral)
if "situational" in question_types:
situational_q = self._generate_situational_questions(competency, questions_per_competency)
competency_questions.extend(situational_q)
# Ensure we have enough questions for this competency
while len(competency_questions) < questions_per_competency:
competency_questions.extend(self._generate_fallback_questions(competency, level))
if len(competency_questions) >= questions_per_competency:
break
questions.extend(competency_questions[:questions_per_competency])
# Shuffle and limit to requested number
random.shuffle(questions)
return questions[:num_questions]
def _generate_situational_questions(self, competency: str, count: int) -> List[Dict[str, Any]]:
"""Generate situational questions for a competency."""
situational_templates = {
"leadership": [
{
"question": "You're leading a project that's behind schedule and the client is unhappy. How do you handle this situation?",
"competency": competency,
"type": "situational",
"focus_areas": ["crisis_management", "client_communication", "team_leadership"]
}
],
"collaboration": [
{
"question": "You're working on a cross-functional project and two team members have opposing views on the technical approach. How do you resolve this?",
"competency": competency,
"type": "situational",
"focus_areas": ["conflict_resolution", "technical_decision_making", "facilitation"]
}
],
"problem_solving": [
{
"question": "You've been assigned to improve the performance of a critical system, but you have limited time and budget. Walk me through your approach.",
"competency": competency,
"type": "situational",
"focus_areas": ["prioritization", "resource_constraints", "systematic_approach"]
}
]
}
if competency in situational_templates:
return situational_templates[competency][:count]
return []
def _generate_fallback_questions(self, competency: str, level: str) -> List[Dict[str, Any]]:
"""Generate fallback questions when specific ones aren't available."""
fallback_questions = [
{
"question": f"Describe your experience with {competency.replace('_', ' ')} in your current or previous role.",
"competency": competency,
"type": "experience",
"focus_areas": ["experience_depth", "practical_application"]
},
{
"question": f"What challenges have you faced related to {competency.replace('_', ' ')} and how did you overcome them?",
"competency": competency,
"type": "challenge_based",
"focus_areas": ["problem_solving", "learning_from_experience"]
}
]
return fallback_questions
def _create_scoring_rubrics(self, questions: List[Dict[str, Any]]) -> Dict[str, Dict]:
"""Create scoring rubrics for the generated questions."""
rubrics = {}
for i, question in enumerate(questions, 1):
question_key = f"question_{i}"
question_type = question.get("type", "behavioral")
if question_type in self.scoring_rubrics:
rubrics[question_key] = {
"question": question["question"],
"competency": question["competency"],
"type": question_type,
"scoring_criteria": self.scoring_rubrics[question_type],
"weight": self._determine_question_weight(question),
"time_limit": question.get("time_limit", 30)
}
return rubrics
def _determine_question_weight(self, question: Dict[str, Any]) -> str:
"""Determine the weight/importance of a question."""
competency = question.get("competency", "")
question_type = question.get("type", "")
difficulty = question.get("difficulty", "medium")
# Core competencies get higher weight
core_competencies = ["coding_fundamentals", "system_design", "leadership", "problem_solving"]
if competency in core_competencies:
return "high"
elif question_type in ["coding", "design"] or difficulty == "hard":
return "high"
elif difficulty == "easy":
return "medium"
else:
return "medium"
def _generate_follow_up_probes(self, questions: List[Dict[str, Any]]) -> Dict[str, List[str]]:
"""Generate follow-up probes for each question."""
probes = {}
for i, question in enumerate(questions, 1):
question_key = f"question_{i}"
competency = question.get("competency", "")
# Get competency-specific follow-ups
if competency in self.follow_up_strategies:
competency_probes = self.follow_up_strategies[competency].copy()
else:
competency_probes = [
"Can you provide more specific details about your approach?",
"What would you do differently if you had to do this again?",
"What challenges did you face and how did you overcome them?"
]
# Add question-type specific probes
question_type = question.get("type", "")
if question_type == "coding":
competency_probes.extend([
"How would you test this solution?",
"What's the time and space complexity of your approach?",
"Can you think of any optimizations?"
])
elif question_type == "behavioral":
competency_probes.extend([
"What did you learn from this experience?",
"How did others react to your approach?",
"What metrics did you use to measure success?"
])
elif question_type == "design":
competency_probes.extend([
"How would you handle failure scenarios?",
"What monitoring would you implement?",
"How would this scale to 10x the load?"
])
probes[question_key] = competency_probes[:5] # Limit to 5 follow-ups
return probes
def _create_calibration_examples(self, sample_questions: List[Dict[str, Any]]) -> Dict[str, Dict]:
"""Create calibration examples with poor/good/great answers."""
examples = {}
for i, question in enumerate(sample_questions, 1):
question_key = f"question_{i}"
examples[question_key] = {
"question": question["question"],
"competency": question["competency"],
"sample_answers": {
"poor_answer": self._generate_sample_answer(question, "poor"),
"good_answer": self._generate_sample_answer(question, "good"),
"great_answer": self._generate_sample_answer(question, "great")
},
"scoring_rationale": self._generate_scoring_rationale(question)
}
return examples
def _generate_sample_answer(self, question: Dict[str, Any], quality: str) -> Dict[str, str]:
"""Generate sample answers of different quality levels."""
competency = question.get("competency", "")
question_type = question.get("type", "")
if quality == "poor":
return {
"answer": f"Sample poor answer for {competency} question - lacks detail, specificity, or demonstrates weak competency",
"score": "1-2",
"issues": ["Vague response", "Limited evidence of competency", "Poor structure"]
}
elif quality == "good":
return {
"answer": f"Sample good answer for {competency} question - adequate detail, demonstrates competency clearly",
"score": "3",
"strengths": ["Clear structure", "Demonstrates competency", "Adequate detail"]
}
else: # great
return {
"answer": f"Sample excellent answer for {competency} question - exceptional detail, strong evidence, goes above and beyond",
"score": "4",
"strengths": ["Exceptional detail", "Strong evidence", "Strategic thinking", "Goes beyond requirements"]
}
def _generate_scoring_rationale(self, question: Dict[str, Any]) -> Dict[str, str]:
"""Generate rationale for scoring this question."""
competency = question.get("competency", "")
return {
"key_indicators": f"Look for evidence of {competency.replace('_', ' ')} competency",
"red_flags": "Vague answers, lack of specifics, negative outcomes without learning",
"green_flags": "Specific examples, clear impact, demonstrates growth and learning"
}
def _generate_usage_guidelines(self, role_key: str, level_key: str) -> Dict[str, Any]:
"""Generate usage guidelines for the question bank."""
return {
"interview_flow": {
"warm_up": "Start with 1-2 easier questions to build rapport",
"core_assessment": "Focus majority of time on core competency questions",
"closing": "End with questions about candidate's questions/interests"
},
"time_management": {
"technical_questions": "Allow extra time for coding/design questions",
"behavioral_questions": "Keep to time limits but allow for follow-ups",
"total_recommendation": "45-75 minutes per interview round"
},
"question_selection": {
"variety": "Mix question types within each competency area",
"difficulty": "Adjust based on candidate responses and energy",
"customization": "Adapt questions based on candidate's background"
},
"common_mistakes": [
"Don't ask all questions mechanically",
"Don't skip follow-up questions",
"Don't forget to assess cultural fit alongside competencies",
"Don't let one strong/weak area bias overall assessment"
],
"calibration_reminders": [
"Compare against role standard, not other candidates",
"Focus on evidence demonstrated, not potential",
"Consider level-appropriate expectations",
"Document specific examples in feedback"
]
}
def format_human_readable(question_bank: Dict[str, Any]) -> str:
"""Format question bank data in human-readable format."""
output = []
# Header
output.append(f"Interview Question Bank: {question_bank['role']} ({question_bank['level'].title()} Level)")
output.append("=" * 70)
output.append(f"Generated: {question_bank['generated_at']}")
output.append(f"Total Questions: {question_bank['total_questions']}")
output.append(f"Question Types: {', '.join(question_bank['question_types'])}")
output.append(f"Target Competencies: {', '.join(question_bank['competencies'])}")
output.append("")
# Questions
output.append("INTERVIEW QUESTIONS")
output.append("-" * 50)
for i, question in enumerate(question_bank['questions'], 1):
output.append(f"\n{i}. {question['question']}")
output.append(f" Competency: {question['competency'].replace('_', ' ').title()}")
output.append(f" Type: {question.get('type', 'N/A').title()}")
if 'time_limit' in question:
output.append(f" Time Limit: {question['time_limit']} minutes")
if 'focus_areas' in question:
output.append(f" Focus Areas: {', '.join(question['focus_areas'])}")
# Scoring Guidelines
output.append("\n\nSCORING RUBRICS")
output.append("-" * 50)
# Show sample scoring criteria
if question_bank['scoring_rubrics']:
first_question = list(question_bank['scoring_rubrics'].keys())[0]
sample_rubric = question_bank['scoring_rubrics'][first_question]
output.append(f"Sample Scoring Criteria ({sample_rubric['type']} questions):")
for criterion, scores in sample_rubric['scoring_criteria'].items():
output.append(f"\n{criterion.replace('_', ' ').title()}:")
for score, description in scores.items():
output.append(f" {score}: {description}")
# Follow-up Probes
output.append("\n\nFOLLOW-UP PROBE EXAMPLES")
output.append("-" * 50)
if question_bank['follow_up_probes']:
first_question = list(question_bank['follow_up_probes'].keys())[0]
sample_probes = question_bank['follow_up_probes'][first_question]
output.append("Sample follow-up questions:")
for probe in sample_probes[:3]: # Show first 3
output.append(f" • {probe}")
# Usage Guidelines
output.append("\n\nUSAGE GUIDELINES")
output.append("-" * 50)
guidelines = question_bank['usage_guidelines']
output.append("Interview Flow:")
for phase, description in guidelines['interview_flow'].items():
output.append(f" • {phase.replace('_', ' ').title()}: {description}")
output.append("\nTime Management:")
for aspect, recommendation in guidelines['time_management'].items():
output.append(f" • {aspect.replace('_', ' ').title()}: {recommendation}")
output.append("\nCommon Mistakes to Avoid:")
for mistake in guidelines['common_mistakes'][:3]: # Show first 3
output.append(f" • {mistake}")
# Calibration Examples (if available)
if question_bank['calibration_examples']:
output.append("\n\nCALIBRATION EXAMPLES")
output.append("-" * 50)
first_example = list(question_bank['calibration_examples'].values())[0]
output.append(f"Question: {first_example['question']}")
output.append("\nSample Answer Quality Levels:")
for quality, details in first_example['sample_answers'].items():
output.append(f" {quality.replace('_', ' ').title()} (Score {details['score']}):")
if 'issues' in details:
output.append(f" Issues: {', '.join(details['issues'])}")
if 'strengths' in details:
output.append(f" Strengths: {', '.join(details['strengths'])}")
return "\n".join(output)
def main():
parser = argparse.ArgumentParser(description="Generate comprehensive interview question banks with scoring criteria")
parser.add_argument("--role", type=str, help="Job role title (e.g., 'Frontend Engineer')")
parser.add_argument("--level", type=str, default="senior", help="Experience level (junior, mid, senior, staff, principal)")
parser.add_argument("--competencies", type=str, help="Comma-separated list of competencies to focus on")
parser.add_argument("--question-types", type=str, help="Comma-separated list of question types (technical, behavioral, situational)")
parser.add_argument("--num-questions", type=int, default=20, help="Number of questions to generate")
parser.add_argument("--input", type=str, help="Input JSON file with role requirements")
parser.add_argument("--output", type=str, help="Output directory or file path")
parser.add_argument("--format", choices=["json", "text", "both"], default="both", help="Output format")
args = parser.parse_args()
generator = QuestionBankGenerator()
# Handle input
if args.input:
try:
with open(args.input, 'r') as f:
role_data = json.load(f)
role = role_data.get('role') or role_data.get('title', '')
level = role_data.get('level', 'senior')
competencies = role_data.get('competencies')
question_types = role_data.get('question_types')
num_questions = role_data.get('num_questions', 20)
except Exception as e:
print(f"Error reading input file: {e}")
sys.exit(1)
else:
if not args.role:
print("Error: --role is required when not using --input")
sys.exit(1)
role = args.role
level = args.level
competencies = args.competencies.split(',') if args.competencies else None
question_types = args.question_types.split(',') if args.question_types else None
num_questions = args.num_questions
# Generate question bank
try:
question_bank = generator.generate_question_bank(
role=role,
level=level,
competencies=competencies,
question_types=question_types,
num_questions=num_questions
)
# Handle output
if args.output:
output_path = args.output
if os.path.isdir(output_path):
safe_role = "".join(c for c in role.lower() if c.isalnum() or c in (' ', '-', '_')).replace(' ', '_')
base_filename = f"{safe_role}_{level}_questions"
json_path = os.path.join(output_path, f"{base_filename}.json")
text_path = os.path.join(output_path, f"{base_filename}.txt")
else:
json_path = output_path if output_path.endswith('.json') else f"{output_path}.json"
text_path = output_path.replace('.json', '.txt') if output_path.endswith('.json') else f"{output_path}.txt"
else:
safe_role = "".join(c for c in role.lower() if c.isalnum() or c in (' ', '-', '_')).replace(' ', '_')
base_filename = f"{safe_role}_{level}_questions"
json_path = f"{base_filename}.json"
text_path = f"{base_filename}.txt"
# Write outputs
if args.format in ["json", "both"]:
with open(json_path, 'w') as f:
json.dump(question_bank, f, indent=2, default=str)
print(f"JSON output written to: {json_path}")
if args.format in ["text", "both"]:
with open(text_path, 'w') as f:
f.write(format_human_readable(question_bank))
print(f"Text output written to: {text_path}")
# Print summary
print(f"\nQuestion Bank Summary:")
print(f"Role: {question_bank['role']} ({question_bank['level'].title()})")
print(f"Total Questions: {question_bank['total_questions']}")
print(f"Competencies Covered: {len(question_bank['competencies'])}")
print(f"Question Types: {', '.join(question_bank['question_types'])}")
except Exception as e:
print(f"Error generating question bank: {e}")
sys.exit(1)
if __name__ == "__main__":
main()
FILE:README.md
# Interview System Designer
A comprehensive toolkit for designing, optimizing, and calibrating interview processes. This skill provides tools to create role-specific interview loops, generate competency-based question banks, and analyze hiring data for bias and calibration issues.
## Overview
The Interview System Designer skill includes three powerful Python tools and comprehensive reference materials to help you build fair, effective, and scalable hiring processes:
1. **Interview Loop Designer** - Generate calibrated interview loops for any role and level
2. **Question Bank Generator** - Create competency-based interview questions with scoring rubrics
3. **Hiring Calibrator** - Analyze interview data to detect bias and calibration issues
## Tools
### 1. Interview Loop Designer (`loop_designer.py`)
Generates complete interview loops tailored to specific roles, levels, and teams.
**Features:**
- Role-specific competency mapping (SWE, PM, Designer, Data, DevOps, Leadership)
- Level-appropriate interview rounds (junior through principal)
- Optimized scheduling and time allocation
- Interviewer skill requirements
- Standardized scorecard templates
**Usage:**
```bash
# Basic usage
python3 loop_designer.py --role "Senior Software Engineer" --level senior
# With team and custom competencies
python3 loop_designer.py --role "Product Manager" --level mid --team growth --competencies leadership,strategy,analytics
# Using JSON input file
python3 loop_designer.py --input assets/sample_role_definitions.json --output loops/
# Specify output format
python3 loop_designer.py --role "Staff Data Scientist" --level staff --format json --output data_scientist_loop.json
```
**Input Options:**
- `--role`: Job role title (e.g., "Senior Software Engineer", "Product Manager")
- `--level`: Experience level (junior, mid, senior, staff, principal)
- `--team`: Team or department (optional)
- `--competencies`: Comma-separated list of specific competencies to focus on
- `--input`: JSON file with role definition
- `--output`: Output directory or file path
- `--format`: Output format (json, text, both) - default: both
**Example Output:**
```
Interview Loop Design for Senior Software Engineer (Senior Level)
============================================================
Total Duration: 300 minutes (5h 0m)
Total Rounds: 5
INTERVIEW ROUNDS
----------------------------------------
Round 1: Technical Phone Screen
Duration: 45 minutes
Format: Virtual
Focus Areas: Coding Fundamentals, Problem Solving
Round 2: System Design
Duration: 75 minutes
Format: Collaborative Whitboard
Focus Areas: System Thinking, Architectural Reasoning
...
```
### 2. Question Bank Generator (`question_bank_generator.py`)
Creates comprehensive interview question banks organized by competency area.
**Features:**
- Competency-based question organization
- Level-appropriate difficulty progression
- Multiple question types (technical, behavioral, situational)
- Detailed scoring rubrics with calibration examples
- Follow-up probes and conversation guides
**Usage:**
```bash
# Generate questions for specific competencies
python3 question_bank_generator.py --role "Frontend Engineer" --competencies react,typescript,system-design
# Create behavioral question bank
python3 question_bank_generator.py --role "Product Manager" --question-types behavioral,leadership --num-questions 15
# Generate questions for multiple levels
python3 question_bank_generator.py --role "DevOps Engineer" --levels junior,mid,senior --output questions/
```
**Input Options:**
- `--role`: Job role title
- `--level`: Experience level (default: senior)
- `--competencies`: Comma-separated list of competencies to focus on
- `--question-types`: Types to include (technical, behavioral, situational)
- `--num-questions`: Number of questions to generate (default: 20)
- `--input`: JSON file with role requirements
- `--output`: Output directory or file path
- `--format`: Output format (json, text, both) - default: both
**Question Types:**
- **Technical**: Coding problems, system design, domain-specific challenges
- **Behavioral**: STAR method questions focusing on past experiences
- **Situational**: Hypothetical scenarios testing decision-making
### 3. Hiring Calibrator (`hiring_calibrator.py`)
Analyzes interview scores to detect bias, calibration issues, and provides recommendations.
**Features:**
- Statistical bias detection across demographics
- Interviewer calibration analysis
- Score distribution and trending analysis
- Specific coaching recommendations
- Comprehensive reporting with actionable insights
**Usage:**
```bash
# Comprehensive analysis
python3 hiring_calibrator.py --input assets/sample_interview_results.json --analysis-type comprehensive
# Focus on specific areas
python3 hiring_calibrator.py --input interview_data.json --analysis-type bias --competencies technical,leadership
# Trend analysis over time
python3 hiring_calibrator.py --input historical_data.json --trend-analysis --period quarterly
```
**Input Options:**
- `--input`: JSON file with interview results data (required)
- `--analysis-type`: Type of analysis (comprehensive, bias, calibration, interviewer, scoring)
- `--competencies`: Comma-separated list of competencies to focus on
- `--trend-analysis`: Enable trend analysis over time
- `--period`: Time period for trends (daily, weekly, monthly, quarterly)
- `--output`: Output file path
- `--format`: Output format (json, text, both) - default: both
**Analysis Types:**
- **Comprehensive**: Full analysis including bias, calibration, and recommendations
- **Bias**: Focus on demographic and interviewer bias patterns
- **Calibration**: Interviewer consistency and agreement analysis
- **Interviewer**: Individual interviewer performance and coaching needs
- **Scoring**: Score distribution and pattern analysis
## Data Formats
### Role Definition Input (JSON)
```json
{
"role": "Senior Software Engineer",
"level": "senior",
"team": "platform",
"competencies": ["system_design", "technical_leadership", "mentoring"],
"requirements": {
"years_experience": "5-8",
"technical_skills": ["Python", "AWS", "Kubernetes"],
"leadership_experience": true
}
}
```
### Interview Results Input (JSON)
```json
[
{
"candidate_id": "candidate_001",
"role": "Senior Software Engineer",
"interviewer_id": "interviewer_alice",
"date": "2024-01-15T09:00:00Z",
"scores": {
"coding_fundamentals": 3.5,
"system_design": 4.0,
"technical_leadership": 3.0,
"communication": 3.5
},
"overall_recommendation": "Hire",
"gender": "male",
"ethnicity": "asian",
"years_experience": 6
}
]
```
## Reference Materials
### Competency Matrix Templates (`references/competency_matrix_templates.md`)
- Comprehensive competency matrices for all engineering roles
- Level-specific expectations (junior through principal)
- Assessment criteria and growth paths
- Customization guidelines for different company stages and industries
### Bias Mitigation Checklist (`references/bias_mitigation_checklist.md`)
- Pre-interview preparation checklist
- Interview process bias prevention strategies
- Real-time bias interruption techniques
- Legal compliance reminders
- Emergency response protocols
### Debrief Facilitation Guide (`references/debrief_facilitation_guide.md`)
- Structured debrief meeting frameworks
- Evidence-based discussion techniques
- Bias interruption strategies
- Decision documentation standards
- Common challenges and solutions
## Sample Data
The `assets/` directory contains sample data for testing:
- `sample_role_definitions.json`: Example role definitions for various positions
- `sample_interview_results.json`: Sample interview data with multiple candidates and interviewers
## Expected Outputs
The `expected_outputs/` directory contains examples of tool outputs:
- Interview loop designs in both JSON and human-readable formats
- Question banks with scoring rubrics and calibration examples
- Calibration analysis reports with bias detection and recommendations
## Best Practices
### Interview Loop Design
1. **Competency Focus**: Align interview rounds with role-critical competencies
2. **Level Calibration**: Adjust expectations and question difficulty based on experience level
3. **Time Optimization**: Balance thoroughness with candidate experience
4. **Interviewer Training**: Ensure interviewers are qualified and calibrated
### Question Bank Development
1. **Evidence-Based**: Focus on observable behaviors and concrete examples
2. **Bias Mitigation**: Use structured questions that minimize subjective interpretation
3. **Calibration**: Include examples of different quality responses for consistency
4. **Continuous Improvement**: Regularly update questions based on predictive validity
### Calibration Analysis
1. **Regular Monitoring**: Analyze hiring data quarterly for bias patterns
2. **Prompt Action**: Address calibration issues immediately with targeted coaching
3. **Data Quality**: Ensure complete and consistent data collection
4. **Legal Compliance**: Monitor for discriminatory patterns and document corrections
## Installation & Setup
No external dependencies required - uses Python 3 standard library only.
```bash
# Clone or download the skill directory
cd interview-system-designer/
# Make scripts executable (optional)
chmod +x *.py
# Test with sample data
python3 loop_designer.py --role "Senior Software Engineer" --level senior
python3 question_bank_generator.py --role "Product Manager" --level mid
python3 hiring_calibrator.py --input assets/sample_interview_results.json
```
## Integration
### With Existing Systems
- **ATS Integration**: Export interview loops as structured data for applicant tracking systems
- **Calendar Systems**: Use scheduling outputs to auto-create interview blocks
- **HR Analytics**: Import calibration reports into broader diversity and inclusion dashboards
### Custom Workflows
- **Batch Processing**: Process multiple roles or historical data sets
- **Automated Reporting**: Schedule regular calibration analysis
- **Custom Competencies**: Extend frameworks with company-specific competencies
## Troubleshooting
### Common Issues
**"Role not found" errors:**
- The tool will map common variations (engineer → software_engineer)
- For custom roles, use the closest standard role and specify custom competencies
**"Insufficient data" errors:**
- Minimum 5 interviews required for statistical analysis
- Ensure interview data includes required fields (candidate_id, interviewer_id, scores, date)
**Missing output files:**
- Check file permissions in output directory
- Ensure adequate disk space
- Verify JSON input file format is valid
### Performance Considerations
- Interview loop generation: < 1 second
- Question bank generation: 1-3 seconds for 20 questions
- Calibration analysis: 1-5 seconds for 50 interviews, scales linearly
## Contributing
To extend this skill:
1. **New Roles**: Add competency frameworks in `_init_competency_frameworks()`
2. **New Question Types**: Extend question templates in respective generators
3. **New Analysis Types**: Add analysis methods to hiring calibrator
4. **Custom Outputs**: Modify formatting functions for different output needs
## License & Usage
This skill is designed for internal company use in hiring process optimization. All bias detection and mitigation features should be reviewed with legal counsel to ensure compliance with local employment laws.
For questions or support, refer to the comprehensive documentation in each script's docstring and the reference materials provided.
FILE:references/bias_mitigation_checklist.md
# Interview Bias Mitigation Checklist
This comprehensive checklist helps identify, prevent, and mitigate various forms of bias in the interview process. Use this as a systematic guide to ensure fair and equitable hiring practices.
## Pre-Interview Phase
### Job Description & Requirements
- [ ] **Remove unnecessary requirements** that don't directly relate to job performance
- [ ] **Avoid gendered language** (competitive, aggressive vs. collaborative, detail-oriented)
- [ ] **Remove university prestige requirements** unless absolutely necessary for role
- [ ] **Focus on skills and outcomes** rather than years of experience in specific technologies
- [ ] **Use inclusive language** and avoid cultural assumptions
- [ ] **Specify only essential requirements** vs. nice-to-have qualifications
- [ ] **Remove location/commute assumptions** for remote-eligible positions
- [ ] **Review requirements for unconscious bias** (e.g., assuming continuous work history)
### Sourcing & Pipeline
- [ ] **Diversify sourcing channels** beyond traditional networks
- [ ] **Partner with diverse professional organizations** and communities
- [ ] **Use bias-minimizing sourcing tools** and platforms
- [ ] **Track sourcing effectiveness** by demographic groups
- [ ] **Train recruiters on bias awareness** and inclusive outreach
- [ ] **Review referral patterns** for potential network bias
- [ ] **Expand university partnerships** beyond elite institutions
- [ ] **Use structured outreach messages** to reduce individual bias
### Resume Screening
- [ ] **Implement blind resume review** (remove names, photos, university names initially)
- [ ] **Use standardized screening criteria** applied consistently
- [ ] **Multiple screeners for each resume** with independent scoring
- [ ] **Focus on relevant skills and achievements** over pedigree indicators
- [ ] **Avoid assumptions about career gaps** or non-traditional backgrounds
- [ ] **Consider alternative paths to skills** (bootcamps, self-taught, career changes)
- [ ] **Track screening pass rates** by demographic groups
- [ ] **Regular screener calibration sessions** on bias awareness
## Interview Panel Composition
### Diversity Requirements
- [ ] **Ensure diverse interview panels** (gender, ethnicity, seniority levels)
- [ ] **Include at least one underrepresented interviewer** when possible
- [ ] **Rotate panel assignments** to prevent bias patterns
- [ ] **Balance seniority levels** on panels (not all senior or all junior)
- [ ] **Include cross-functional perspectives** when relevant
- [ ] **Avoid panels of only one demographic group** when possible
- [ ] **Consider panel member unconscious bias training** status
- [ ] **Document panel composition rationale** for future review
### Interviewer Selection
- [ ] **Choose interviewers based on relevant competency assessment ability**
- [ ] **Ensure interviewers have completed bias training** within last 12 months
- [ ] **Select interviewers with consistent calibration history**
- [ ] **Avoid interviewers with known bias patterns** (flagged in previous analyses)
- [ ] **Include at least one interviewer familiar with candidate's background type**
- [ ] **Balance perspectives** (technical depth, cultural fit, growth potential)
- [ ] **Consider interviewer availability for proper preparation time**
- [ ] **Ensure interviewers understand role requirements and standards**
## Interview Process Design
### Question Standardization
- [ ] **Use standardized question sets** for each competency area
- [ ] **Develop questions that assess skills, not culture fit stereotypes**
- [ ] **Avoid questions about personal background** unless directly job-relevant
- [ ] **Remove questions that could reveal protected characteristics**
- [ ] **Focus on behavioral examples** using STAR method
- [ ] **Include scenario-based questions** with clear evaluation criteria
- [ ] **Test questions for potential bias** with diverse interviewers
- [ ] **Regularly update question bank** based on effectiveness data
### Structured Interview Protocol
- [ ] **Define clear time allocations** for each question/section
- [ ] **Establish consistent interview flow** across all candidates
- [ ] **Create standardized intro/outro** processes
- [ ] **Use identical technical setup and tools** for all candidates
- [ ] **Provide same background information** to all interviewers
- [ ] **Standardize note-taking format** and requirements
- [ ] **Define clear handoff procedures** between interviewers
- [ ] **Document any deviations** from standard protocol
### Accommodation Preparation
- [ ] **Proactively offer accommodations** without requiring disclosure
- [ ] **Provide multiple interview format options** (phone, video, in-person)
- [ ] **Ensure accessibility of interview locations and tools**
- [ ] **Allow extended time** when requested or needed
- [ ] **Provide materials in advance** when helpful
- [ ] **Train interviewers on accommodation protocols**
- [ ] **Test all technology** for accessibility compliance
- [ ] **Have backup plans** for technical issues
## During the Interview
### Interviewer Behavior
- [ ] **Use welcoming, professional tone** with all candidates
- [ ] **Avoid assumptions based on appearance or background**
- [ ] **Give equal encouragement and support** to all candidates
- [ ] **Allow equal time for candidate questions**
- [ ] **Avoid leading questions** that suggest desired answers
- [ ] **Listen actively** without interrupting unnecessarily
- [ ] **Take detailed notes** focusing on responses, not impressions
- [ ] **Avoid small talk** that could reveal irrelevant personal information
### Question Delivery
- [ ] **Ask questions as written** without improvisation that could introduce bias
- [ ] **Provide equal clarification** when candidates ask for it
- [ ] **Use consistent follow-up probing** across candidates
- [ ] **Allow reasonable thinking time** before expecting responses
- [ ] **Avoid rephrasing questions** in ways that give hints
- [ ] **Stay focused on defined competencies** being assessed
- [ ] **Give equal encouragement** for elaboration when needed
- [ ] **Maintain professional demeanor** regardless of candidate background
### Real-time Bias Checking
- [ ] **Notice first impressions** but don't let them drive assessment
- [ ] **Question gut reactions** - are they based on competency evidence?
- [ ] **Focus on specific examples** and evidence provided
- [ ] **Avoid pattern matching** to existing successful employees
- [ ] **Notice cultural assumptions** in interpretation of responses
- [ ] **Check for confirmation bias** - seeking evidence to support initial impressions
- [ ] **Consider alternative explanations** for candidate responses
- [ ] **Stay aware of fatigue effects** on judgment throughout the day
## Evaluation & Scoring
### Scoring Consistency
- [ ] **Use defined rubrics consistently** across all candidates
- [ ] **Score immediately after interview** while details are fresh
- [ ] **Focus scoring on demonstrated competencies** not potential or personality
- [ ] **Provide specific evidence** for each score given
- [ ] **Avoid comparative scoring** (comparing candidates to each other)
- [ ] **Use calibrated examples** of each score level
- [ ] **Score independently** before discussing with other interviewers
- [ ] **Document reasoning** for all scores, especially extreme ones (1s and 4s)
### Bias Check Questions
- [ ] **"Would I score this differently if the candidate looked different?"**
- [ ] **"Am I basing this on evidence or assumptions?"**
- [ ] **"Would this response get the same score from a different demographic?"**
- [ ] **"Am I penalizing non-traditional backgrounds or approaches?"**
- [ ] **"Is my scoring consistent with the defined rubric?"**
- [ ] **"Am I letting one strong/weak area bias overall assessment?"**
- [ ] **"Are my cultural assumptions affecting interpretation?"**
- [ ] **"Would I want to work with this person?" (Check if this is biasing assessment)**
### Documentation Requirements
- [ ] **Record specific examples** supporting each competency score
- [ ] **Avoid subjective language** like "seems like," "appears to be"
- [ ] **Focus on observable behaviors** and concrete responses
- [ ] **Note exact quotes** when relevant to assessment
- [ ] **Distinguish between facts and interpretations**
- [ ] **Provide improvement suggestions** that are skill-based, not person-based
- [ ] **Avoid comparative language** to other candidates or employees
- [ ] **Use neutral language** free from cultural assumptions
## Debrief Process
### Structured Discussion
- [ ] **Start with independent score sharing** before discussion
- [ ] **Focus discussion on evidence** not impressions or feelings
- [ ] **Address significant score discrepancies** with evidence review
- [ ] **Challenge biased language** or assumptions in discussion
- [ ] **Ensure all voices are heard** in group decision making
- [ ] **Document reasons for final decision** with specific evidence
- [ ] **Avoid personality-based discussions** ("culture fit" should be evidence-based)
- [ ] **Consider multiple perspectives** on candidate responses
### Decision-Making Process
- [ ] **Use weighted scoring system** based on role requirements
- [ ] **Require minimum scores** in critical competency areas
- [ ] **Avoid veto power** unless based on clear, documented evidence
- [ ] **Consider growth potential** fairly across all candidates
- [ ] **Document dissenting opinions** and reasoning
- [ ] **Use tie-breaking criteria** that are predetermined and fair
- [ ] **Consider additional data collection** if team is split
- [ ] **Make final decision based on role requirements**, not team preferences
### Final Recommendations
- [ ] **Provide specific, actionable feedback** for development areas
- [ ] **Focus recommendations on skills and competencies**
- [ ] **Avoid language that could reflect bias** in written feedback
- [ ] **Consider onboarding needs** based on actual skill gaps, not assumptions
- [ ] **Provide coaching recommendations** that are evidence-based
- [ ] **Avoid personal judgments** about candidate character or personality
- [ ] **Make hiring recommendation** based solely on job-relevant criteria
- [ ] **Document any concerns** with specific, observable evidence
## Post-Interview Monitoring
### Data Collection
- [ ] **Track interviewer scoring patterns** for consistency analysis
- [ ] **Monitor pass rates** by demographic groups
- [ ] **Collect candidate experience feedback** on interview fairness
- [ ] **Analyze score distributions** for potential bias indicators
- [ ] **Track time-to-decision** across different candidate types
- [ ] **Monitor offer acceptance rates** by demographics
- [ ] **Collect new hire performance data** for process validation
- [ ] **Document any bias incidents** or concerns raised
### Regular Analysis
- [ ] **Conduct quarterly bias audits** of interview data
- [ ] **Review interviewer calibration** and identify outliers
- [ ] **Analyze demographic trends** in hiring outcomes
- [ ] **Compare candidate experience surveys** across groups
- [ ] **Track correlation between interview scores and job performance**
- [ ] **Review and update bias mitigation strategies** based on data
- [ ] **Share findings with interview teams** for continuous improvement
- [ ] **Update training programs** based on identified bias patterns
## Bias Types to Watch For
### Affinity Bias
- **Definition**: Favoring candidates similar to yourself
- **Watch for**: Over-positive response to shared backgrounds, interests, or experiences
- **Mitigation**: Focus on job-relevant competencies, diversify interview panels
### Halo/Horn Effect
- **Definition**: One positive/negative trait influencing overall assessment
- **Watch for**: Strong performance in one area affecting scores in unrelated areas
- **Mitigation**: Score each competency independently, use structured evaluation
### Confirmation Bias
- **Definition**: Seeking information that confirms initial impressions
- **Watch for**: Asking follow-ups that lead candidate toward expected responses
- **Mitigation**: Use standardized questions, consider alternative interpretations
### Attribution Bias
- **Definition**: Attributing success/failure to different causes based on candidate demographics
- **Watch for**: Assuming women are "lucky" vs. men are "skilled" for same achievements
- **Mitigation**: Focus on candidate's role in achievements, avoid assumptions
### Cultural Bias
- **Definition**: Judging candidates based on cultural differences rather than job performance
- **Watch for**: Penalizing communication styles, work approaches, or values that differ from team norm
- **Mitigation**: Define job-relevant criteria clearly, consider diverse perspectives valuable
### Educational Bias
- **Definition**: Over-weighting prestigious educational credentials
- **Watch for**: Assuming higher capability based on school rank rather than demonstrated skills
- **Mitigation**: Focus on skills demonstration, consider alternative learning paths
### Experience Bias
- **Definition**: Requiring specific company or industry experience unnecessarily
- **Watch for**: Discounting transferable skills from different industries or company sizes
- **Mitigation**: Define core skills needed, assess adaptability and learning ability
## Emergency Bias Response Protocol
### During Interview
1. **Pause the interview** if significant bias is observed
2. **Privately address** bias with interviewer if possible
3. **Document the incident** for review
4. **Continue with fair assessment** of candidate
5. **Flag for debrief discussion** if interview continues
### Post-Interview
1. **Report bias incidents** to hiring manager/HR immediately
2. **Document specific behaviors** observed
3. **Consider additional interviewer** for second opinion
4. **Review candidate assessment** for bias impact
5. **Implement corrective actions** for future interviews
### Interviewer Coaching
1. **Provide immediate feedback** on bias observed
2. **Schedule bias training refresher** if needed
3. **Monitor future interviews** for improvement
4. **Consider removing from interview rotation** if bias persists
5. **Document coaching provided** for performance management
## Legal Compliance Reminders
### Protected Characteristics
- Age, race, color, religion, sex, national origin, disability status, veteran status
- Pregnancy, genetic information, sexual orientation, gender identity
- Any other characteristics protected by local/state/federal law
### Prohibited Questions
- Questions about family planning, marital status, pregnancy
- Age-related questions (unless BFOQ)
- Religious or political affiliations
- Disability status (unless voluntary disclosure for accommodation)
- Arrest records (without conviction relevance)
- Financial status or credit (unless job-relevant)
### Documentation Requirements
- Keep all interview materials for required retention period
- Ensure consistent documentation across all candidates
- Avoid documenting protected characteristic observations
- Focus documentation on job-relevant observations only
## Training & Certification
### Required Training Topics
- Unconscious bias awareness and mitigation
- Structured interviewing techniques
- Legal compliance in hiring
- Company-specific bias mitigation protocols
- Role-specific competency assessment
- Accommodation and accessibility requirements
### Ongoing Development
- Annual bias training refresher
- Quarterly calibration sessions
- Regular updates on legal requirements
- Peer feedback and coaching
- Industry best practice updates
- Data-driven process improvements
This checklist should be reviewed and updated regularly based on legal requirements, industry best practices, and internal bias analysis results.
FILE:references/competency_matrix_templates.md
# Competency Matrix Templates
This document provides comprehensive competency matrix templates for different engineering roles and levels. Use these matrices to design role-specific interview loops and evaluation criteria.
## Software Engineering Competency Matrix
### Technical Competencies
| Competency | Junior (L1-L2) | Mid (L3-L4) | Senior (L5-L6) | Staff+ (L7+) |
|------------|----------------|-------------|----------------|--------------|
| **Coding & Algorithms** | Basic data structures, simple algorithms, language syntax | Advanced algorithms, complexity analysis, optimization | Complex problem solving, algorithm design, performance tuning | Architecture-level algorithmic decisions, novel approach design |
| **System Design** | Component interactions, basic scalability concepts | Service design, database modeling, API design | Distributed systems, scalability patterns, trade-off analysis | Large-scale architecture, cross-system design, technology strategy |
| **Code Quality** | Readable code, basic testing, follows conventions | Maintainable code, comprehensive testing, design patterns | Code reviews, quality standards, refactoring leadership | Engineering standards, quality culture, technical debt management |
| **Debugging & Problem Solving** | Basic debugging, structured problem approach | Complex debugging, root cause analysis, performance issues | System-wide debugging, production issues, incident response | Cross-system troubleshooting, preventive measures, tooling design |
| **Domain Knowledge** | Learning role-specific technologies | Proficiency in domain tools/frameworks | Deep domain expertise, technology evaluation | Domain leadership, technology roadmap, innovation |
### Behavioral Competencies
| Competency | Junior (L1-L2) | Mid (L3-L4) | Senior (L5-L6) | Staff+ (L7+) |
|------------|----------------|-------------|----------------|--------------|
| **Communication** | Clear status updates, asks good questions | Technical explanations, stakeholder updates | Cross-functional communication, technical writing | Executive communication, external representation, thought leadership |
| **Collaboration** | Team participation, code reviews | Cross-team projects, knowledge sharing | Team leadership, conflict resolution | Cross-org collaboration, culture building, strategic partnerships |
| **Leadership & Influence** | Peer mentoring, positive attitude | Junior mentoring, project ownership | Team guidance, technical decisions, hiring | Org-wide influence, vision setting, culture change |
| **Growth & Learning** | Skill development, feedback receptivity | Proactive learning, teaching others | Continuous improvement, trend awareness | Learning culture, industry leadership, innovation adoption |
| **Ownership & Initiative** | Task completion, quality focus | Project ownership, process improvement | Feature/service ownership, strategic thinking | Product/platform ownership, business impact, market influence |
## Product Management Competency Matrix
### Product Competencies
| Competency | Associate PM (L1-L2) | PM (L3-L4) | Senior PM (L5-L6) | Principal PM (L7+) |
|------------|---------------------|------------|-------------------|-------------------|
| **Product Strategy** | Feature requirements, user stories | Product roadmaps, market analysis | Business strategy, competitive positioning | Portfolio strategy, market creation, platform vision |
| **User Research & Analytics** | Basic user interviews, metrics tracking | Research design, data interpretation | Research strategy, advanced analytics | Research culture, measurement frameworks, insight generation |
| **Technical Understanding** | Basic tech concepts, API awareness | System architecture, technical trade-offs | Technical strategy, platform decisions | Technology vision, architectural influence, innovation leadership |
| **Execution & Process** | Feature delivery, stakeholder coordination | Project management, cross-functional leadership | Process optimization, team scaling | Operational excellence, org design, strategic execution |
| **Business Acumen** | Revenue awareness, customer understanding | P&L understanding, business case development | Business strategy, market dynamics | Corporate strategy, board communication, investor relations |
### Leadership Competencies
| Competency | Associate PM (L1-L2) | PM (L3-L4) | Senior PM (L5-L6) | Principal PM (L7+) |
|------------|---------------------|------------|-------------------|-------------------|
| **Stakeholder Management** | Team collaboration, clear communication | Cross-functional alignment, expectation management | Executive communication, influence without authority | Board interaction, external partnerships, industry influence |
| **Team Development** | Peer learning, feedback sharing | Junior mentoring, knowledge transfer | Team building, hiring, performance management | Talent development, culture building, org leadership |
| **Decision Making** | Data-driven decisions, priority setting | Complex trade-offs, strategic choices | Ambiguous situations, high-stakes decisions | Strategic vision, transformational decisions, risk management |
| **Innovation & Vision** | Creative problem solving, user empathy | Market opportunity identification, feature innovation | Product vision, market strategy | Industry vision, disruptive thinking, platform creation |
## Design Competency Matrix
### Design Competencies
| Competency | Junior Designer (L1-L2) | Mid Designer (L3-L4) | Senior Designer (L5-L6) | Principal Designer (L7+) |
|------------|-------------------------|---------------------|-------------------------|-------------------------|
| **Visual Design** | UI components, typography, color theory | Design systems, visual hierarchy | Brand integration, advanced layouts | Visual strategy, brand evolution, design innovation |
| **User Experience** | User flows, wireframing, prototyping | Interaction design, usability testing | Experience strategy, journey mapping | UX vision, service design, behavioral insights |
| **Research & Validation** | User interviews, usability tests | Research planning, data synthesis | Research strategy, methodology design | Research culture, insight frameworks, market research |
| **Design Systems** | Component usage, style guides | System contribution, pattern creation | System architecture, governance | System strategy, scalable design, platform thinking |
| **Tools & Craft** | Design software proficiency, asset creation | Advanced techniques, workflow optimization | Tool evaluation, process design | Technology integration, future tooling, craft evolution |
### Collaboration Competencies
| Competency | Junior Designer (L1-L2) | Mid Designer (L3-L4) | Senior Designer (L5-L6) | Principal Designer (L7+) |
|------------|-------------------------|---------------------|-------------------------|-------------------------|
| **Cross-functional Partnership** | Engineering collaboration, handoff quality | Product partnership, stakeholder alignment | Leadership collaboration, strategic alignment | Executive partnership, business strategy integration |
| **Communication & Advocacy** | Design rationale, feedback integration | Design presentations, user advocacy | Executive communication, design thinking evangelism | Industry thought leadership, external representation |
| **Mentorship & Growth** | Peer learning, skill sharing | Junior mentoring, critique facilitation | Team development, hiring, career guidance | Design culture, talent strategy, industry leadership |
| **Business Impact** | User-centered thinking, design quality | Feature success, user satisfaction | Business metrics, strategic impact | Market influence, competitive advantage, innovation leadership |
## Data Science Competency Matrix
### Technical Competencies
| Competency | Junior DS (L1-L2) | Mid DS (L3-L4) | Senior DS (L5-L6) | Principal DS (L7+) |
|------------|-------------------|----------------|-------------------|-------------------|
| **Statistical Analysis** | Descriptive stats, hypothesis testing | Advanced statistics, experimental design | Causal inference, advanced modeling | Statistical strategy, methodology innovation |
| **Machine Learning** | Basic ML algorithms, model training | Advanced ML, feature engineering | ML systems, model deployment | ML strategy, AI platform, research direction |
| **Data Engineering** | SQL, basic ETL, data cleaning | Pipeline design, data modeling | Platform architecture, scalable systems | Data strategy, infrastructure vision, governance |
| **Programming & Tools** | Python/R proficiency, visualization | Advanced programming, tool integration | Software engineering, system design | Technology strategy, platform development, innovation |
| **Domain Expertise** | Business understanding, metric interpretation | Domain modeling, insight generation | Strategic analysis, business integration | Market expertise, competitive intelligence, thought leadership |
### Impact & Leadership Competencies
| Competency | Junior DS (L1-L2) | Mid DS (L3-L4) | Senior DS (L5-L6) | Principal DS (L7+) |
|------------|-------------------|----------------|-------------------|-------------------|
| **Business Impact** | Metric improvement, insight delivery | Project leadership, business case development | Strategic initiatives, P&L impact | Business transformation, market advantage, innovation |
| **Communication** | Technical reporting, visualization | Stakeholder presentations, executive briefings | Board communication, external representation | Industry leadership, thought leadership, market influence |
| **Team Leadership** | Peer collaboration, knowledge sharing | Junior mentoring, project management | Team building, hiring, culture development | Organizational leadership, talent strategy, vision setting |
| **Innovation & Research** | Algorithm implementation, experimentation | Research projects, publication | Research strategy, academic partnerships | Research vision, industry influence, breakthrough innovation |
## DevOps Engineering Competency Matrix
### Technical Competencies
| Competency | Junior DevOps (L1-L2) | Mid DevOps (L3-L4) | Senior DevOps (L5-L6) | Principal DevOps (L7+) |
|------------|----------------------|-------------------|----------------------|----------------------|
| **Infrastructure** | Basic cloud services, server management | Infrastructure automation, containerization | Platform architecture, multi-cloud strategy | Infrastructure vision, emerging technologies, industry standards |
| **CI/CD & Automation** | Pipeline basics, script writing | Advanced pipelines, deployment automation | Platform design, workflow optimization | Automation strategy, developer experience, productivity platforms |
| **Monitoring & Observability** | Basic monitoring, log analysis | Advanced monitoring, alerting systems | Observability strategy, SLA/SLI design | Monitoring vision, reliability engineering, performance culture |
| **Security & Compliance** | Security basics, access management | Security automation, compliance frameworks | Security architecture, risk management | Security strategy, governance, industry leadership |
| **Performance & Scalability** | Performance monitoring, basic optimization | Capacity planning, performance tuning | Scalability architecture, cost optimization | Performance strategy, efficiency platforms, innovation |
### Leadership & Impact Competencies
| Competency | Junior DevOps (L1-L2) | Mid DevOps (L3-L4) | Senior DevOps (L5-L6) | Principal DevOps (L7+) |
|------------|----------------------|-------------------|----------------------|----------------------|
| **Developer Experience** | Tool support, documentation | Platform development, self-service tools | Developer productivity, workflow design | Developer platform vision, industry best practices |
| **Incident Management** | Incident response, troubleshooting | Incident coordination, root cause analysis | Incident strategy, prevention systems | Reliability culture, organizational resilience |
| **Team Collaboration** | Cross-team support, knowledge sharing | Process improvement, training delivery | Culture building, practice evangelism | Organizational transformation, industry influence |
| **Strategic Impact** | Operational excellence, cost awareness | Efficiency improvements, platform adoption | Strategic initiatives, business enablement | Technology strategy, competitive advantage, market leadership |
## Engineering Management Competency Matrix
### People Leadership Competencies
| Competency | Manager (L1-L2) | Senior Manager (L3-L4) | Director (L5-L6) | VP+ (L7+) |
|------------|-----------------|------------------------|------------------|----------|
| **Team Building** | Hiring, onboarding, 1:1s | Team culture, performance management | Multi-team coordination, org design | Organizational culture, talent strategy |
| **Performance Management** | Individual development, feedback | Performance systems, coaching | Calibration across teams, promotion standards | Talent development, succession planning |
| **Communication** | Team updates, stakeholder management | Executive communication, cross-functional alignment | Board updates, external communication | Industry representation, thought leadership |
| **Conflict Resolution** | Team conflicts, process improvements | Cross-team issues, organizational friction | Strategic alignment, cultural challenges | Corporate-level conflicts, crisis management |
### Technical Leadership Competencies
| Competency | Manager (L1-L2) | Senior Manager (L3-L4) | Director (L5-L6) | VP+ (L7+) |
|------------|-----------------|------------------------|------------------|----------|
| **Technical Vision** | Team technical decisions, architecture input | Platform strategy, technology choices | Technical roadmap, innovation strategy | Technology vision, industry standards |
| **System Ownership** | Feature/service ownership, quality standards | Platform ownership, scalability planning | System portfolio, technical debt management | Technology strategy, competitive advantage |
| **Process & Practice** | Team processes, development practices | Engineering standards, quality systems | Process innovation, best practices | Engineering culture, industry influence |
| **Technology Strategy** | Tool evaluation, team technology choices | Platform decisions, technical investments | Technology portfolio, strategic architecture | Corporate technology strategy, market leadership |
## Usage Guidelines
### Assessment Approach
1. **Level Calibration**: Use these matrices to calibrate expectations for each level within your organization
2. **Interview Design**: Select competencies most relevant to the specific role and level being hired for
3. **Evaluation Consistency**: Ensure all interviewers understand and apply the same competency standards
4. **Growth Planning**: Use matrices for career development and promotion discussions
### Customization Tips
1. **Industry Adaptation**: Modify competencies based on your industry (fintech, healthcare, etc.)
2. **Company Stage**: Adjust expectations based on startup vs. enterprise environment
3. **Team Needs**: Emphasize competencies most critical for current team challenges
4. **Cultural Fit**: Add company-specific values and cultural competencies
### Common Pitfalls
1. **Unrealistic Expectations**: Don't expect senior-level competencies from junior candidates
2. **One-Size-Fits-All**: Customize competency emphasis based on role requirements
3. **Static Assessment**: Regularly update matrices based on changing business needs
4. **Bias Introduction**: Ensure competencies are measurable and don't introduce unconscious bias
## Matrix Validation Process
### Regular Review Cycle
- **Quarterly**: Review competency relevance and adjust weights
- **Semi-annually**: Update level expectations based on market standards
- **Annually**: Comprehensive review with stakeholder feedback
### Stakeholder Input
- **Hiring Managers**: Validate role-specific competency requirements
- **Current Team Members**: Confirm level expectations match reality
- **Recent Hires**: Gather feedback on assessment accuracy
- **HR Partners**: Ensure legal compliance and bias mitigation
### Continuous Improvement
- **Performance Correlation**: Track new hire performance against competency assessments
- **Market Benchmarking**: Compare standards with industry peers
- **Feedback Integration**: Incorporate interviewer and candidate feedback
- **Bias Monitoring**: Regular analysis of assessment patterns across demographics
FILE:references/debrief_facilitation_guide.md
# Interview Debrief Facilitation Guide
This guide provides a comprehensive framework for conducting effective, unbiased interview debriefs that lead to consistent hiring decisions. Use this to facilitate productive discussions that focus on evidence-based evaluation.
## Pre-Debrief Preparation
### Facilitator Responsibilities
- [ ] **Review all interviewer feedback** before the meeting
- [ ] **Identify significant score discrepancies** that need discussion
- [ ] **Prepare discussion agenda** with time allocations
- [ ] **Gather role requirements** and competency framework
- [ ] **Review any flags or special considerations** noted during interviews
- [ ] **Ensure all required materials** are available (scorecards, rubrics, candidate resume)
- [ ] **Set up meeting logistics** (room, video conference, screen sharing)
- [ ] **Send agenda to participants** 30 minutes before meeting
### Required Materials Checklist
- [ ] Candidate resume and application materials
- [ ] Job description and competency requirements
- [ ] Individual interviewer scorecards
- [ ] Scoring rubrics and competency definitions
- [ ] Interview notes and documentation
- [ ] Any technical assessments or work samples
- [ ] Company hiring standards and calibration examples
- [ ] Bias mitigation reminders and prompts
### Participant Preparation Requirements
- [ ] All interviewers must **complete independent scoring** before debrief
- [ ] **Submit written feedback** with specific evidence for each competency
- [ ] **Review scoring rubrics** to ensure consistent interpretation
- [ ] **Prepare specific examples** to support scoring decisions
- [ ] **Flag any concerns or unusual circumstances** that affected assessment
- [ ] **Avoid discussing candidate** with other interviewers before debrief
- [ ] **Come prepared to defend scores** with concrete evidence
- [ ] **Be ready to adjust scores** based on additional evidence shared
## Debrief Meeting Structure
### Opening (5 minutes)
1. **State meeting purpose**: Make hiring decision based on evidence
2. **Review agenda and time limits**: Keep discussion focused and productive
3. **Remind of bias mitigation principles**: Focus on competencies, not personality
4. **Confirm confidentiality**: Discussion stays within hiring team
5. **Establish ground rules**: One person speaks at a time, evidence-based discussion
### Individual Score Sharing (10-15 minutes)
- **Go around the room systematically** - each interviewer shares scores independently
- **No discussion or challenges yet** - just data collection
- **Record scores on shared document** visible to all participants
- **Note any abstentions** or "insufficient data" responses
- **Identify clear patterns** and discrepancies without commentary
- **Flag any scores requiring explanation** (1s or 4s typically need strong evidence)
### Competency-by-Competency Discussion (30-40 minutes)
#### For Each Core Competency:
**1. Present Score Distribution (2 minutes)**
- Display all scores for this competency
- Note range and any outliers
- Identify if consensus exists or discussion needed
**2. Evidence Sharing (5-8 minutes per competency)**
- Start with interviewers who assessed this competency directly
- Share specific examples and observations
- Focus on what candidate said/did, not interpretations
- Allow questions for clarification (not challenges yet)
**3. Discussion and Calibration (3-5 minutes)**
- Address significant discrepancies (>1 point difference)
- Challenge vague or potentially biased language
- Seek additional evidence if needed
- Allow score adjustments based on new information
- Reach consensus or note dissenting views
#### Structured Discussion Questions:
- **"What specific evidence supports this score?"**
- **"Can you provide the exact example or quote?"**
- **"How does this compare to our rubric definition?"**
- **"Would this response receive the same score regardless of who gave it?"**
- **"Are we evaluating the competency or making assumptions?"**
- **"What would need to change for this to be the next level up/down?"**
### Overall Recommendation Discussion (10-15 minutes)
#### Weighted Score Calculation
1. **Apply competency weights** based on role requirements
2. **Calculate overall weighted average**
3. **Check minimum threshold requirements**
4. **Consider any veto criteria** (critical competency failures)
#### Final Recommendation Options
- **Strong Hire**: Exceeds requirements in most areas, clear value-add
- **Hire**: Meets requirements with growth potential
- **No Hire**: Doesn't meet minimum requirements for success
- **Strong No Hire**: Significant gaps that would impact team/company
#### Decision Rationale Documentation
- **Summarize key strengths** with specific evidence
- **Identify development areas** with specific examples
- **Explain final recommendation** with competency-based reasoning
- **Note any dissenting opinions** and reasoning
- **Document onboarding considerations** if hiring
### Closing and Next Steps (5 minutes)
- **Confirm final decision** and documentation
- **Assign follow-up actions** (feedback delivery, offer preparation, etc.)
- **Schedule any additional interviews** if needed
- **Review timeline** for candidate communication
- **Remind confidentiality** of discussion and decision
## Facilitation Best Practices
### Creating Psychological Safety
- **Encourage honest feedback** without fear of judgment
- **Validate different perspectives** and assessment approaches
- **Address power dynamics** - ensure junior voices are heard
- **Model vulnerability** - admit when evidence changes your mind
- **Focus on learning** and calibration, not winning arguments
- **Thank participants** for thorough preparation and thoughtful input
### Managing Difficult Conversations
#### When Scores Vary Significantly
1. **Acknowledge the discrepancy** without judgment
2. **Ask for specific evidence** from each scorer
3. **Look for different interpretations** of the same data
4. **Consider if different questions** revealed different competency levels
5. **Check for bias patterns** in reasoning
6. **Allow time for reflection** and potential score adjustments
#### When Someone Uses Biased Language
1. **Pause the conversation** gently but firmly
2. **Ask for specific evidence** behind the assessment
3. **Reframe in competency terms** - "What specific skills did this demonstrate?"
4. **Challenge assumptions** - "Help me understand how we know that"
5. **Redirect to rubric** - "How does this align with our scoring criteria?"
6. **Document and follow up** privately if bias persists
#### When the Discussion Gets Off Track
- **Redirect to competencies**: "Let's focus on the technical skills demonstrated"
- **Ask for evidence**: "What specific example supports that assessment?"
- **Reference rubrics**: "How does this align with our level 3 definition?"
- **Manage time**: "We have 5 minutes left on this competency"
- **Table unrelated issues**: "That's important but separate from this hire decision"
### Encouraging Evidence-Based Discussion
#### Good Evidence Examples
- **Direct quotes**: "When asked about debugging, they said..."
- **Specific behaviors**: "They organized their approach by first..."
- **Observable outcomes**: "Their code compiled on first run and handled edge cases"
- **Process descriptions**: "They walked through their problem-solving step by step"
- **Measurable results**: "They identified 3 optimization opportunities"
#### Poor Evidence Examples
- **Gut feelings**: "They just seemed off"
- **Comparisons**: "Not as strong as our last hire"
- **Assumptions**: "Probably wouldn't fit our culture"
- **Vague impressions**: "Didn't seem passionate"
- **Irrelevant factors**: "Their background is different from ours"
### Managing Group Dynamics
#### Ensuring Equal Participation
- **Direct questions** to quieter participants
- **Prevent interrupting** and ensure everyone finishes thoughts
- **Balance speaking time** across all interviewers
- **Validate minority opinions** even if not adopted
- **Check for unheard perspectives** before finalizing decisions
#### Handling Strong Personalities
- **Set time limits** for individual speaking
- **Redirect monopolizers**: "Let's hear from others on this"
- **Challenge confidently stated opinions** that lack evidence
- **Support less assertive voices** in expressing dissenting views
- **Focus on data**, not personality or seniority in decision making
## Bias Interruption Strategies
### Affinity Bias Interruption
- **Notice pattern**: Positive assessment seems based on shared background/interests
- **Interrupt with**: "Let's focus on the job-relevant skills they demonstrated"
- **Redirect to**: Specific competency evidence and measurable outcomes
- **Document**: Note if personal connection affected professional assessment
### Halo/Horn Effect Interruption
- **Notice pattern**: One area strongly influencing assessment of unrelated areas
- **Interrupt with**: "Let's score each competency independently"
- **Redirect to**: Specific evidence for each individual competency area
- **Recalibrate**: Ask for separate examples supporting each score
### Confirmation Bias Interruption
- **Notice pattern**: Only seeking/discussing evidence that supports initial impression
- **Interrupt with**: "What evidence might suggest a different assessment?"
- **Redirect to**: Consider alternative interpretations of the same data
- **Challenge**: "How might we be wrong about this assessment?"
### Attribution Bias Interruption
- **Notice pattern**: Attributing success to luck/help for some demographics, skill for others
- **Interrupt with**: "What role did the candidate play in achieving this outcome?"
- **Redirect to**: Candidate's specific contributions and decision-making
- **Standardize**: Apply same attribution standards across all candidates
## Decision Documentation Framework
### Required Documentation Elements
1. **Final scores** for each assessed competency
2. **Overall recommendation** with supporting rationale
3. **Key strengths** with specific evidence
4. **Development areas** with specific examples
5. **Dissenting opinions** if any, with reasoning
6. **Special considerations** or accommodation needs
7. **Next steps** and timeline for decision communication
### Evidence Quality Standards
- **Specific and observable**: What exactly did the candidate do or say?
- **Job-relevant**: How does this relate to success in the role?
- **Measurable**: Can this be quantified or clearly described?
- **Unbiased**: Would this evidence be interpreted the same way regardless of candidate demographics?
- **Complete**: Does this represent the full picture of their performance in this area?
### Writing Guidelines
- **Use active voice** and specific language
- **Avoid assumptions** about motivations or personality
- **Focus on behaviors** demonstrated during the interview
- **Provide context** for any unusual circumstances
- **Be constructive** in describing development areas
- **Maintain professionalism** and respect for candidate
## Common Debrief Challenges and Solutions
### Challenge: "I just don't think they'd fit our culture"
**Solution**:
- Ask for specific, observable evidence
- Define what "culture fit" means in job-relevant terms
- Challenge assumptions about cultural requirements
- Focus on ability to collaborate and contribute effectively
### Challenge: Scores vary widely with no clear explanation
**Solution**:
- Review if different interviewers assessed different competencies
- Look for question differences that might explain variance
- Consider if candidate performance varied across interviews
- May need additional data gathering or interview
### Challenge: Everyone loved/hated the candidate but can't articulate why
**Solution**:
- Push for specific evidence supporting emotional reactions
- Review competency rubrics together
- Look for halo/horn effects influencing overall impression
- Consider unconscious bias training for team
### Challenge: Technical vs. non-technical interviewers disagree
**Solution**:
- Clarify which competencies each interviewer was assessing
- Ensure technical assessments carry appropriate weight
- Look for different perspectives on same evidence
- Consider specialist input for technical decisions
### Challenge: Senior interviewer dominates decision making
**Solution**:
- Structure discussion to hear from all levels first
- Ask direct questions to junior interviewers
- Challenge opinions that lack supporting evidence
- Remember that assessment ability doesn't correlate with seniority
### Challenge: Team wants to hire but scores don't support it
**Solution**:
- Review if rubrics match actual job requirements
- Check for consistent application of scoring standards
- Consider if additional competencies need assessment
- May indicate need for rubric calibration or role requirement review
## Post-Debrief Actions
### Immediate Actions (Same Day)
- [ ] **Finalize decision documentation** with all evidence
- [ ] **Communicate decision** to recruiting team
- [ ] **Schedule candidate feedback** delivery if applicable
- [ ] **Update interview scheduling** based on decision
- [ ] **Note any process improvements** needed for future
### Follow-up Actions (Within 1 Week)
- [ ] **Deliver candidate feedback** (internal or external)
- [ ] **Update interview feedback** in tracking system
- [ ] **Schedule any additional interviews** if needed
- [ ] **Begin offer process** if hiring
- [ ] **Document lessons learned** for process improvement
### Long-term Actions (Monthly/Quarterly)
- [ ] **Analyze debrief effectiveness** and decision quality
- [ ] **Review interviewer calibration** based on decisions
- [ ] **Update rubrics** based on debrief insights
- [ ] **Provide additional training** if bias patterns identified
- [ ] **Share successful practices** with other hiring teams
## Continuous Improvement Framework
### Debrief Effectiveness Metrics
- **Decision consistency**: Are similar candidates receiving similar decisions?
- **Time to decision**: Are debriefs completing within planned time?
- **Participation quality**: Are all interviewers contributing evidence-based input?
- **Bias incidents**: How often are bias interruptions needed?
- **Decision satisfaction**: Do participants feel good about the process and outcome?
### Regular Review Process
- **Monthly**: Review debrief facilitation effectiveness and interviewer feedback
- **Quarterly**: Analyze decision patterns and potential bias indicators
- **Semi-annually**: Update debrief processes based on hiring outcome data
- **Annually**: Comprehensive review of debrief framework and training needs
### Training and Calibration
- **New facilitators**: Shadow 3-5 debriefs before leading independently
- **All facilitators**: Quarterly calibration sessions on bias interruption
- **Interviewer training**: Include debrief participation expectations
- **Leadership training**: Ensure hiring managers can facilitate effectively
This guide should be adapted to your organization's specific needs while maintaining focus on evidence-based, unbiased decision making.
FILE:references/interview-frameworks.md
# Interview Frameworks
## Loop Design by Level
### Junior/Mid
- Emphasize fundamentals, debugging, and growth potential.
- Keep loops concise with coding + behavioral validation.
### Senior
- Add system design and leadership rounds.
- Evaluate tradeoff quality, mentoring, and cross-team collaboration.
### Staff+
- Focus on architecture direction and organizational impact.
- Assess strategy, influence, and long-term technical judgment.
## Competency Areas
- Technical depth (implementation, design, quality)
- Problem solving (ambiguity handling, prioritization)
- Collaboration (communication, stakeholder alignment)
- Leadership (ownership, mentoring, influence)
## Scoring Rubric Baseline
- `4`: exceeds level expectations with strong evidence
- `3`: meets expectations consistently
- `2`: partial signal with notable gaps
- `1`: does not meet baseline requirements
## Calibration Guidelines
- Run recurring interviewer calibration sessions.
- Compare interviewer scoring variance across rounds.
- Track interview signal against new-hire outcomes.
- Use structured debriefs with independent scoring before discussion.
## Bias-Reduction Baseline
- Standardize question banks per competency area.
- Keep scorecards evidence-based and behavior-specific.
- Use diverse interviewer panels where possible.
- Require written rationale for strong yes/no recommendations.
FILE:scripts/interview_planner.py
#!/usr/bin/env python3
"""Generate an interview loop plan by role and level."""
from __future__ import annotations
import argparse
import json
from typing import Dict, List
BASE_ROUNDS = {
"junior": [
("Screen", 45, "Fundamentals and communication"),
("Coding", 60, "Problem solving and code quality"),
("Behavioral", 45, "Collaboration and growth mindset"),
],
"mid": [
("Screen", 45, "Fundamentals and ownership"),
("Coding", 60, "Implementation quality"),
("System Design", 60, "Service/component design"),
("Behavioral", 45, "Stakeholder collaboration"),
],
"senior": [
("Screen", 45, "Depth and tradeoff reasoning"),
("Coding", 60, "Code quality and testing"),
("System Design", 75, "Scalability and reliability"),
("Leadership", 60, "Mentoring and decision making"),
("Behavioral", 45, "Cross-functional influence"),
],
"staff": [
("Screen", 45, "Strategic and technical depth"),
("Architecture", 90, "Org-level design decisions"),
("Technical Strategy", 60, "Long-term tradeoffs"),
("Influence", 60, "Cross-team leadership"),
("Behavioral", 45, "Values and executive communication"),
],
}
QUESTION_BANK = {
"coding": [
"Walk through your approach before coding and identify tradeoffs.",
"How would you test this implementation for edge cases?",
"What would you refactor if this code became a shared library?",
],
"system": [
"Design this system for 10x traffic growth in 12 months.",
"Where are the main failure modes and how would you detect them?",
"What components would you scale first and why?",
],
"leadership": [
"Describe a time you changed technical direction with incomplete information.",
"How do you raise the bar for code quality across a team?",
"How do you handle disagreement between product and engineering priorities?",
],
"behavioral": [
"Tell me about a high-stakes mistake and what changed afterward.",
"Describe a conflict where you had to influence without authority.",
"How do you support underperforming teammates?",
],
}
def normalize_level(level: str) -> str:
level = level.strip().lower()
if level in {"staff+", "principal", "lead"}:
return "staff"
if level not in BASE_ROUNDS:
raise ValueError(f"Unsupported level: {level}")
return level
def suggested_questions(round_name: str) -> List[str]:
name = round_name.lower()
if "coding" in name:
return QUESTION_BANK["coding"]
if "system" in name or "architecture" in name:
return QUESTION_BANK["system"]
if "lead" in name or "influence" in name or "strategy" in name:
return QUESTION_BANK["leadership"]
return QUESTION_BANK["behavioral"]
def generate_plan(role: str, level: str) -> Dict[str, object]:
normalized = normalize_level(level)
rounds = []
for idx, (name, minutes, focus) in enumerate(BASE_ROUNDS[normalized], start=1):
rounds.append(
{
"round": idx,
"name": name,
"duration_minutes": minutes,
"focus": focus,
"suggested_questions": suggested_questions(name),
}
)
return {
"role": role,
"level": normalized,
"total_rounds": len(rounds),
"total_minutes": sum(r["duration_minutes"] for r in rounds),
"rounds": rounds,
}
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Generate an interview loop plan for a role and level.")
parser.add_argument("--role", required=True, help="Role name (e.g., Senior Software Engineer)")
parser.add_argument("--level", required=True, help="Level: junior|mid|senior|staff")
parser.add_argument("--json", action="store_true", help="Output as JSON")
return parser.parse_args()
def main() -> int:
args = parse_args()
plan = generate_plan(args.role, args.level)
if args.json:
print(json.dumps(plan, indent=2))
else:
print(f"Interview Plan: {plan['role']} ({plan['level']})")
print(f"Total rounds: {plan['total_rounds']} | Total time: {plan['total_minutes']} minutes")
print("")
for r in plan["rounds"]:
print(f"Round {r['round']}: {r['name']} ({r['duration_minutes']} min)")
print(f"Focus: {r['focus']}")
for q in r["suggested_questions"]:
print(f"- {q}")
print("")
return 0
if __name__ == "__main__":
raise SystemExit(main())
Chuyên gia đánh giá hệ thống quản lý an toàn thông tin (ISMS), kiểm tra tuân thủ ISO 27001, đánh giá kiểm soát Annex A và hỗ trợ chứng nhận.
---
name: "isms-audit-expert"
description: Information Security Management System (ISMS) audit expert for ISO 27001 compliance verification, security control assessment, and certification support. Use when the user mentions ISO 27001, ISMS audit, Annex A controls, Statement of Applicability (SOA), gap analysis, nonconformity management, internal audit, surveillance audit, or security certification preparation. Helps review control implementation evidence, document audit findings, classify nonconformities, generate risk-based audit plans, map controls to Annex A requirements, prepare Stage 1 and Stage 2 audit documentation, and support corrective action workflows.
triggers:
- ISMS audit
- ISO 27001 audit
- security audit
- internal audit ISO 27001
- security control assessment
- certification audit
- surveillance audit
- audit finding
- nonconformity
---
# ISMS Audit Expert
Internal and external ISMS audit management for ISO 27001 compliance verification, security control assessment, and certification support.
## Table of Contents
- [Audit Program Management](#audit-program-management)
- [Audit Execution](#audit-execution)
- [Control Assessment](#control-assessment)
- [Finding Management](#finding-management)
- [Certification Support](#certification-support)
- [Tools](#tools)
- [References](#references)
---
## Audit Program Management
### Risk-Based Audit Schedule
| Risk Level | Audit Frequency | Examples |
|------------|-----------------|----------|
| Critical | Quarterly | Privileged access, vulnerability management, logging |
| High | Semi-annual | Access control, incident response, encryption |
| Medium | Annual | Policies, awareness training, physical security |
| Low | Annual | Documentation, asset inventory |
### Annual Audit Planning Workflow
1. Review previous audit findings and risk assessment results
2. Identify high-risk controls and recent security incidents
3. Determine audit scope based on ISMS boundaries
4. Assign auditors ensuring independence from audited areas
5. Create audit schedule with resource allocation
6. Obtain management approval for audit plan
7. **Validation:** Audit plan covers all Annex A controls within certification cycle
### Auditor Competency Requirements
- ISO 27001 Lead Auditor certification (preferred)
- No operational responsibility for audited processes
- Understanding of technical security controls
- Knowledge of applicable regulations (GDPR, HIPAA)
---
## Audit Execution
### Pre-Audit Preparation
1. Review ISMS documentation (policies, SoA, risk assessment)
2. Analyze previous audit reports and open findings
3. Prepare audit plan with interview schedule
4. Notify auditees of audit scope and timing
5. Prepare checklists for controls in scope
6. **Validation:** All documentation received and reviewed before opening meeting
### Audit Conduct Steps
1. **Opening Meeting**
- Confirm audit scope and objectives
- Introduce audit team and methodology
- Agree on communication channels and logistics
2. **Evidence Collection**
- Interview control owners and operators
- Review documentation and records
- Observe processes in operation
- Inspect technical configurations
3. **Control Verification**
- Test control design (does it address the risk?)
- Test control operation (is it working as intended?)
- Sample transactions and records
- Document all evidence collected
4. **Closing Meeting**
- Present preliminary findings
- Clarify any factual inaccuracies
- Agree on finding classification
- Confirm corrective action timelines
5. **Validation:** All controls in scope assessed with documented evidence
---
## Control Assessment
### Control Testing Approach
1. Identify control objective from ISO 27002
2. Determine testing method (inquiry, observation, inspection, re-performance)
3. Define sample size based on population and risk
4. Execute test and document results
5. Evaluate control effectiveness
6. **Validation:** Evidence supports conclusion about control status
For detailed technical verification procedures by Annex A control, see [security-control-testing.md](references/security-control-testing.md).
---
## Finding Management
### Finding Classification
| Severity | Definition | Response Time |
|----------|------------|---------------|
| Major Nonconformity | Control failure creating significant risk | 30 days |
| Minor Nonconformity | Isolated deviation with limited impact | 90 days |
| Observation | Improvement opportunity | Next audit cycle |
### Finding Documentation Template
```
Finding ID: ISMS-[YEAR]-[NUMBER]
Control Reference: A.X.X - [Control Name]
Severity: [Major/Minor/Observation]
Evidence:
- [Specific evidence observed]
- [Records reviewed]
- [Interview statements]
Risk Impact:
- [Potential consequences if not addressed]
Root Cause:
- [Why the nonconformity occurred]
Recommendation:
- [Specific corrective action steps]
```
### Corrective Action Workflow
1. Auditee acknowledges finding and severity
2. Root cause analysis completed within 10 days
3. Corrective action plan submitted with target dates
4. Actions implemented by responsible parties
5. Auditor verifies effectiveness of corrections
6. Finding closed with evidence of resolution
7. **Validation:** Root cause addressed, recurrence prevented
---
## Certification Support
### Stage 1 Audit Preparation
Ensure documentation is complete:
- [ ] ISMS scope statement
- [ ] Information security policy (management signed)
- [ ] Statement of Applicability
- [ ] Risk assessment methodology and results
- [ ] Risk treatment plan
- [ ] Internal audit results (past 12 months)
- [ ] Management review minutes
### Stage 2 Audit Preparation
Verify operational readiness:
- [ ] All Stage 1 findings addressed
- [ ] ISMS operational for minimum 3 months
- [ ] Evidence of control implementation
- [ ] Security awareness training records
- [ ] Incident response evidence (if applicable)
- [ ] Access review documentation
### Surveillance Audit Cycle
| Period | Focus |
|--------|-------|
| Year 1, Q2 | High-risk controls, Stage 2 findings follow-up |
| Year 1, Q4 | Continual improvement, control sample |
| Year 2, Q2 | Full surveillance |
| Year 2, Q4 | Re-certification preparation |
**Validation:** No major nonconformities at surveillance audits.
---
## Tools
### scripts/
| Script | Purpose | Usage |
|--------|---------|-------|
| `isms_audit_scheduler.py` | Generate risk-based audit plans | `python scripts/isms_audit_scheduler.py --year 2025 --format markdown` |
### Audit Planning Example
```bash
# Generate annual audit plan
python scripts/isms_audit_scheduler.py --year 2025 --output audit_plan.json
# With custom control risk ratings
python scripts/isms_audit_scheduler.py --controls controls.csv --format markdown
```
---
## References
| File | Content |
|------|---------|
| [iso27001-audit-methodology.md](references/iso27001-audit-methodology.md) | Audit program structure, pre-audit phase, certification support |
| [security-control-testing.md](references/security-control-testing.md) | Technical verification procedures for ISO 27002 controls |
| [cloud-security-audit.md](references/cloud-security-audit.md) | Cloud provider assessment, configuration security, IAM review |
---
## Audit Performance Metrics
| KPI | Target | Measurement |
|-----|--------|-------------|
| Audit plan completion | 100% | Audits completed vs. planned |
| Finding closure rate | >90% within SLA | Closed on time vs. total |
| Major nonconformities | 0 at certification | Count per certification cycle |
| Audit effectiveness | Incidents prevented | Security improvements implemented |
FILE:references/cloud-security-audit.md
# Cloud Security Audit Guide
Assessment framework for cloud service security verification.
---
## Table of Contents
- [Shared Responsibility Model](#shared-responsibility-model)
- [Cloud Provider Assessment](#cloud-provider-assessment)
- [Configuration Security](#configuration-security)
- [Data Protection](#data-protection)
- [Identity and Access Management](#identity-and-access-management)
---
## Shared Responsibility Model
### Responsibility Matrix
| Layer | IaaS | PaaS | SaaS |
|-------|------|------|------|
| Data classification | Customer | Customer | Customer |
| Identity management | Customer | Customer | Shared |
| Application security | Customer | Shared | Provider |
| Network controls | Shared | Provider | Provider |
| Host infrastructure | Provider | Provider | Provider |
| Physical security | Provider | Provider | Provider |
### Audit Focus by Model
**IaaS (AWS EC2, Azure VMs):**
- Virtual network configuration
- OS hardening and patching
- Application deployment security
- Data encryption implementation
**PaaS (Azure App Service, AWS Lambda):**
- Application code security
- Data handling and encryption
- Identity integration
- Logging configuration
**SaaS (Microsoft 365, Salesforce):**
- User access management
- Data classification and handling
- Security configuration settings
- Integration security
---
## Cloud Provider Assessment
### Certification Verification
Check for current certifications:
- [ ] ISO 27001 (Information Security)
- [ ] ISO 27017 (Cloud Security)
- [ ] ISO 27018 (Cloud Privacy)
- [ ] SOC 2 Type II
- [ ] CSA STAR certification
**Verification Steps:**
1. Request current certificates from provider
2. Verify certificate scope includes services used
3. Check certification expiration dates
4. Review SOC 2 report for relevant controls
5. Document any scope exclusions
### Data Residency Compliance
| Requirement | Verification |
|-------------|--------------|
| GDPR (EU data) | Confirm EU region availability |
| Data sovereignty | Verify no cross-border transfer |
| Backup location | Confirm backup region |
| Disaster recovery | Document DR site location |
### Provider Security Documentation
Request and review:
- Shared responsibility documentation
- Security whitepapers
- Incident notification procedures
- SLA for security incidents
- Vulnerability disclosure policy
---
## Configuration Security
### AWS Security Assessment
**Identity and Access (IAM):**
- [ ] Root account has MFA enabled
- [ ] No access keys for root account
- [ ] IAM policies follow least privilege
- [ ] No wildcard (*) permissions on sensitive resources
- [ ] Password policy meets requirements
**Network Configuration (VPC):**
- [ ] Default VPCs removed or secured
- [ ] Security groups follow least privilege
- [ ] No 0.0.0.0/0 ingress on management ports
- [ ] VPC flow logs enabled
- [ ] Network ACLs configured appropriately
**Storage (S3):**
- [ ] No public buckets (unless intended)
- [ ] Bucket policies restrict access
- [ ] Encryption at rest enabled
- [ ] Versioning enabled for critical data
- [ ] Access logging enabled
**Logging (CloudTrail):**
- [ ] CloudTrail enabled in all regions
- [ ] Log file validation enabled
- [ ] Logs encrypted with KMS
- [ ] S3 bucket for logs is secured
- [ ] CloudWatch alarms configured
### Azure Security Assessment
**Identity (Azure AD):**
- [ ] MFA enabled for all users
- [ ] Privileged Identity Management (PIM) configured
- [ ] Conditional Access policies defined
- [ ] Guest access restricted
- [ ] Password protection enabled
**Network (Virtual Networks):**
- [ ] NSG rules follow least privilege
- [ ] No open management ports to internet
- [ ] Network Watcher enabled
- [ ] DDoS protection configured
- [ ] Private endpoints for PaaS services
**Storage:**
- [ ] No anonymous access to blob storage
- [ ] Encryption at rest enabled
- [ ] Shared access signatures time-limited
- [ ] Storage analytics logging enabled
- [ ] Soft delete enabled
**Monitoring:**
- [ ] Azure Monitor enabled
- [ ] Activity log exported to SIEM
- [ ] Alerts configured for security events
- [ ] Azure Security Center enabled
- [ ] Diagnostic settings configured
---
## Data Protection
### Encryption Verification
**At Rest:**
| Service | Encryption Check |
|---------|------------------|
| Block storage | Verify CMK or provider-managed key |
| Object storage | Check default encryption settings |
| Databases | Confirm TDE or column encryption |
| Backups | Verify backup encryption |
**In Transit:**
| Connection | Requirement |
|------------|-------------|
| User to application | TLS 1.2+ required |
| Service to service | Internal TLS or VPN |
| API communications | HTTPS only, no HTTP |
| Database connections | TLS required |
### Key Management Assessment
- [ ] Customer-managed keys used for sensitive data
- [ ] Key rotation policy defined and implemented
- [ ] Key access restricted to authorized services
- [ ] Key usage logged and monitored
- [ ] Disaster recovery for keys documented
### Data Classification in Cloud
| Classification | Cloud Requirements |
|----------------|-------------------|
| Confidential | CMK encryption, access logging, no public access |
| Internal | Encryption enabled, network restrictions |
| Public | Integrity protection, CDN appropriate |
---
## Identity and Access Management
### Privileged Access Review
1. Identify all administrative roles
2. Verify role assignment justification
3. Check for standing vs. just-in-time access
4. Review privileged activity logs
5. Confirm MFA required for elevation
### Service Account Assessment
| Check | Verification |
|-------|--------------|
| Inventory | All service accounts documented |
| Permissions | Least privilege applied |
| Credentials | Keys rotated per policy |
| Monitoring | Activity logged and reviewed |
| Ownership | Clear owner assigned |
### Federation and SSO
- [ ] SSO configured for cloud console access
- [ ] Conditional Access/MFA policies applied
- [ ] Session timeout configured
- [ ] Failed login monitoring enabled
- [ ] Emergency access accounts documented
### API Security
- [ ] API keys not embedded in code
- [ ] Secrets management service used
- [ ] API access logged
- [ ] Rate limiting configured
- [ ] API permissions follow least privilege
FILE:references/iso27001-audit-methodology.md
# ISO 27001 ISMS Audit Methodology
Complete audit framework and procedures for Information Security Management System assessments.
---
## Table of Contents
- [Audit Program Structure](#audit-program-structure)
- [Pre-Audit Phase](#pre-audit-phase)
- [Audit Execution](#audit-execution)
- [Finding Classification](#finding-classification)
- [Certification Audit Support](#certification-audit-support)
---
## Audit Program Structure
### Annual Audit Schedule
| Quarter | Focus Area | Audit Type |
|---------|------------|------------|
| Q1 | Access Control, Cryptography | Internal |
| Q2 | Operations Security, Communications | Internal |
| Q3 | System Acquisition, Supplier Relations | Internal |
| Q4 | Full ISMS Review | Pre-certification |
### Risk-Based Scheduling
Prioritize audit frequency based on:
- Asset criticality and data classification
- Previous finding history
- Regulatory requirements
- Recent security incidents
- Organizational changes
**High Risk Areas (Quarterly):**
- Access management systems
- Cryptographic key management
- Incident response processes
- Third-party access controls
**Medium Risk Areas (Semi-Annual):**
- Change management
- Backup and recovery
- Physical security
- Security awareness training
**Lower Risk Areas (Annual):**
- Documentation management
- Asset inventory
- Business continuity planning
---
## Pre-Audit Phase
### Documentation Review Checklist
- [ ] ISMS scope statement and boundaries
- [ ] Information security policy (signed, current)
- [ ] Statement of Applicability (SoA)
- [ ] Risk assessment methodology and results
- [ ] Risk treatment plan
- [ ] Security objectives and metrics
- [ ] Previous audit reports and corrective actions
### Audit Plan Template
```
ISMS Audit Plan
Audit ID: ISMS-[YEAR]-[NUMBER]
Scope: [ISMS scope or specific controls]
Date: [Start] to [End]
Lead Auditor: [Name]
Audit Team: [Names]
Day 1:
09:00 - Opening meeting
10:00 - Document review (policies, SoA)
14:00 - Interview: Information Security Manager
Day 2:
09:00 - Technical control verification
14:00 - Process observation
Day 3:
09:00 - Remaining interviews
14:00 - Finding consolidation
16:00 - Closing meeting
```
### Auditor Independence
Verify before audit assignment:
- No operational responsibility for audited area
- No recent (12 months) involvement in audited processes
- No conflict of interest with auditees
- Required competencies documented
---
## Audit Execution
### Evidence Collection Methods
| Method | Use Case | Evidence Type |
|--------|----------|---------------|
| Document review | Policy verification | Screenshots, copies |
| Interviews | Process understanding | Notes, recordings |
| Observation | Operational checks | Photos, timestamps |
| Technical testing | Control effectiveness | System logs, reports |
### Interview Protocol
1. Introduce audit purpose and confidentiality
2. Explain interview will be documented
3. Ask open-ended questions about processes
4. Request evidence to support statements
5. Clarify any inconsistencies
6. Summarize key points before closing
### Sample Interview Questions
**For Security Managers:**
- Describe the risk assessment process
- How are security incidents reported and managed?
- What metrics track ISMS effectiveness?
**For System Administrators:**
- How is privileged access managed?
- Walk through the change management process
- Show backup verification records
**For End Users:**
- What security training have you received?
- How do you report suspicious activity?
- Describe the password policy requirements
### Control Testing Procedures
**Access Control (A.9):**
1. Request user access list for critical system
2. Verify access rights match job roles
3. Check for terminated user accounts
4. Test password policy enforcement
5. Verify MFA configuration
**Logging (A.12.4):**
1. Confirm logging enabled on systems in scope
2. Verify log retention meets policy
3. Check log protection from tampering
4. Review sample security event alerts
---
## Finding Classification
### Severity Levels
| Level | Definition | Response Time |
|-------|------------|---------------|
| Major Nonconformity | Failure of control, significant risk | 30 days corrective action |
| Minor Nonconformity | Isolated deviation, limited impact | 90 days corrective action |
| Observation | Improvement opportunity | Next audit cycle |
| Good Practice | Exceeds requirements | Document and share |
### Finding Documentation
```
Finding ID: ISMS-2025-001
Control Reference: A.9.2.3 - Management of privileged access
Severity: Major Nonconformity
Evidence:
- 15 shared admin accounts identified
- No approval records for privileged access
- Last access review: 18 months ago
Risk Impact:
- Unauthorized access to critical systems
- No accountability for admin actions
- Regulatory non-compliance
Root Cause:
- No defined process for privileged access management
- Insufficient tooling for access tracking
Recommendation:
- Implement PAM solution within 30 days
- Document and enforce privileged access process
- Conduct immediate access review
```
### Corrective Action Tracking
| Field | Content |
|-------|---------|
| Finding ID | Link to original finding |
| Root Cause | Why the nonconformity occurred |
| Corrective Action | Specific steps to address |
| Responsible Person | Named accountable party |
| Target Date | Completion deadline |
| Verification Method | How closure will be confirmed |
| Status | Open / In Progress / Closed |
---
## Certification Audit Support
### Stage 1 Audit Preparation
Ensure availability of:
- [ ] ISMS documentation (scope, policy, SoA)
- [ ] Risk assessment records
- [ ] Internal audit results from past 12 months
- [ ] Management review minutes
- [ ] Corrective action evidence
### Stage 2 Audit Preparation
- [ ] All Stage 1 findings addressed
- [ ] ISMS operational for minimum 3 months
- [ ] Evidence of control effectiveness
- [ ] Training and awareness records
- [ ] Incident response records (if any)
### Surveillance Audit Cycle
| Year | Quarter | Focus |
|------|---------|-------|
| Year 1 | Q2 | High-risk controls, Stage 2 findings |
| Year 1 | Q4 | Remaining controls sample |
| Year 2 | Q2 | Full surveillance |
| Year 2 | Q4 | Continual improvement evidence |
| Year 3 | Q2 | Re-certification preparation |
### Audit Findings Response Template
```
Subject: Response to Finding ISMS-2025-001
Finding: Major Nonconformity - Privileged Access Management
Root Cause Analysis:
[5 Whys or fishbone analysis results]
Corrective Action Plan:
1. [Action] - [Owner] - [Date]
2. [Action] - [Owner] - [Date]
Evidence of Correction:
- [Document/screenshot reference]
Preventive Measures:
- [Steps to prevent recurrence]
Verification Request: [Date auditor can verify]
```
FILE:references/iso27001_audit_playbook.md
# ISO/IEC 27001:2022 Internal Audit Playbook
This reference answers exactly one decision: **how do we prepare for and conduct an ISO 27001 internal audit (Clause 9.2) that produces actionable findings without burning the auditee team?**
Pair with `scripts/isms_audit_scheduler.py` (this skill) for cadence + auditor independence and with `compliance-os/scripts/audit_simulator.py` for mock-audit preparation.
## When to Use This Playbook
- Annual Clause 9.2 internal audit programme
- Pre-stage-1 certification readiness check
- Surveillance audit preparation (year 2 / year 3 of cert cycle)
- Post-incident audit (e.g., breach triggers ad-hoc ISMS audit)
- Onboarding a new business unit into existing ISMS scope
## The 7-Phase Audit Workflow
```
[ Plan ] -> [ Prepare ] -> [ Open ] -> [ Field ] -> [ Close ] -> [ Report ] -> [ Track ]
```
### Phase 1 — Plan (1-2 weeks pre-audit)
- Confirm scope: which Annex A controls, which business units, which clauses
- Confirm auditor independence (no self-audit; rotate across teams)
- Pull prior-year findings + open nonconformities for follow-up
- Define sampling approach (stratified by risk; not random)
- Communicate dates to auditees ≥ 2 weeks in advance
**Outputs:** audit plan (1 page), auditor assignments, document-request list
### Phase 2 — Prepare (1 week pre-audit)
- Auditee assembles document evidence in advance
- Auditor reviews documents BEFORE fieldwork (do not waste interview time reading docs)
- Pre-fieldwork checklist: are documents under version control? Are records signed? Are dates within retention?
- Auditor runs `audit_simulator.py` to mentally rehearse finding scenarios
**Outputs:** prepared document folder, auditor mental model of likely findings
### Phase 3 — Open (30 min, day 1)
- Opening meeting with auditee leadership + key contributors
- State scope, criteria (which Annex A controls), timeline, communication plan
- Set expectations: this is a check on the system, not on individuals
- Confirm safe-to-fail discipline — finding ≠ punishment
**Outputs:** opening minutes; auditee buy-in
### Phase 4 — Field (2-5 days for medium scope)
The core. For each scoped control:
1. **Interview the control owner** — open question, sample drill-down, walk-through
2. **Inspect the record(s)** — pull samples from logs / tickets / records, not curated demos
3. **Cross-reference** — does the record match the procedure? Does management oversight exist?
4. **Document the finding** on the spot — control + observation + evidence + severity
Interview pattern (per ISO 19011 Clause 6):
- "Walk me through how this control is implemented day-to-day."
- "Show me a specific example from the last 30 days."
- "What happens if [edge case]?"
- "Where is this documented?"
**Outputs:** finding worksheets (one per control); severity ratings
### Phase 5 — Close (1-2 hours, last day)
- Closing meeting with auditee team
- Walk through preliminary findings (no surprises in the written report)
- Allow auditee to provide additional evidence for borderline findings
- Confirm corrective action ownership before the report is written
- Agree on draft-report timeline (typically 1-2 weeks)
**Outputs:** closing minutes; preliminary finding agreement
### Phase 6 — Report (1-2 weeks post-fieldwork)
Per ISO 19011 Clause 6.5, the audit report must include:
- Audit objectives, scope, criteria, dates
- Audit team and auditees
- Summary of findings by severity
- Per-finding: control + observation + evidence + severity + corrective action recommendation
- Conclusion: ISMS adequacy + effectiveness verdict
- Distribution list
**Severity grades** (Clause 9.2 compatible):
| Grade | Definition | Treatment |
|---|---|---|
| **Critical (Major NC)** | Absence of, or systemic failure to implement, a required ISMS process | Blocks stage 1 certification; 30-day plan + closure required before progress |
| **Major** | Material gap in a required control | Corrective action plan within 30 days |
| **Minor** | Localized gap; control works overall | Corrective action within 90 days |
| **Observation / OFI** | Improvement opportunity; no nonconformity | Optional; recommendation only |
Healthy distribution: ≥ 40% observation, ≤ 15% critical.
**Outputs:** signed audit report; corrective action assignments
### Phase 7 — Track (ongoing)
- Open findings tracked through existing CAPA system (Clause 10.2)
- Verify closure of each finding via evidence + re-test (do not accept self-attestation)
- Update risk register for residual risks identified
- Feed unresolved findings into next audit cycle + management review (Clause 9.3)
**Outputs:** closed findings + verification evidence; updates to risk register and management review inputs
## Annex A Scope Prioritization (for fieldwork)
ISO 27001:2022 Annex A has 93 controls grouped into 4 themes (A.5 organizational, A.6 people, A.7 physical, A.8 technological). Audit fieldwork should NOT attempt all 93 in one audit — use the 3-year rolling cycle.
**High-priority controls (audit annually):**
| Control | Why prioritize annually |
|---|---|
| A.5.1 — Policies for information security | Foundation; audit changes |
| A.5.9-10 — Inventory of assets + acceptable use | Drives everything else |
| A.5.15 — Access control | Highest-leakage area |
| A.5.19-21 — Supplier management | Most-cited finding area |
| A.5.24-27 — Incident management + Article 33 GDPR alignment | High-stakes |
| A.5.34 — Privacy & PII | GDPR overlap; expand if EU data |
| A.6.3 — Awareness, education, training | Always sampled |
| A.6.7 — Remote working | Pandemic legacy; high audit value |
| A.6.8 — Information security event reporting | Connects to incident management |
| A.8.2-3 — Privileged access; Information access restriction | Pair with A.5.15 |
| A.8.7 — Protection against malware | Always cited |
| A.8.15-16 — Logging + Monitoring | Pair with A.5.24-27 |
| A.8.32 — Change management | High-leakage; pair with vulnerability/patch mgmt |
**Lower-priority controls (audit on rolling 3-year cycle):**
A.5.2 / A.5.3 / A.5.4 / A.5.6 / A.5.7 / A.5.8 / A.5.11 / A.5.13 / A.5.14 / A.5.16 / A.5.17 / A.5.18 / A.5.22 / A.5.23 / A.5.28 / A.5.29 / A.5.30 / A.5.31 / A.5.32 / A.5.33 / A.5.35 / A.5.36 / A.5.37 / A.6.1 / A.6.2 / A.6.4 / A.6.5 / A.6.6 / A.7 (all physical) / A.8.1 / A.8.4 / A.8.5 / A.8.6 / A.8.8 / A.8.9 / A.8.10 / A.8.11 / A.8.12 / A.8.13 / A.8.14 / A.8.17 / A.8.18 / A.8.19 / A.8.20 / A.8.21 / A.8.22 / A.8.23 / A.8.24 / A.8.25-31 (SDLC controls)
## Common Stage 1 / Stage 2 Findings (the patterns)
Based on practitioner reports of common ISO 27001:2022 findings:
1. **Risk register exists but treatment plans are generic.** "Apply A.7.3" without specific implementation.
2. **Asset inventory missing cloud / SaaS / AI tools.** Engineers stopped registering as they multiplied.
3. **Privileged access reviewed annually instead of quarterly.** Find orphaned accounts.
4. **Supplier reviews unsigned or undated.** Procurement collected them; nobody reviewed.
5. **Incident records lack documented post-incident review within 30 days.**
6. **Change advisory board exists but rubber-stamps.** No rejected changes in last 6 months.
7. **Internal audit programme doesn't cover all clauses + applicable controls over 3-year cycle.**
8. **Management review missing required Article 9.3 inputs** (KPI trends, audit findings, risk changes).
9. **Vulnerability management without defined SLAs by severity.**
10. **BCP/DRP exists but never tested.**
## Cross-Framework Reuse
This ISO 27001 audit pattern is the foundation for:
- **SOC 2** — ~75% control overlap; same evidence with TSC-specific formatting (`soc2_audit_playbook.md`)
- **ISO 42001** — Clauses 4-10 reuse ~60%; Annex A overlap on data + supplier; AI-specific Annex A.5/A.6/A.9 net-new
- **GDPR** — Article 32 organizational measures reuse heavily (`gdpr_audit_playbook.md`)
- **NIST CSF profiles** — common control vocabulary
Pair with `compliance-os/references/multi_framework_audit_playbook.md` for orchestrating audits across multiple frameworks.
## When This Reference Doesn't Help
- **Specific Annex A control text.** See ISO 27001:2022 + ISO 27002:2022 (implementation guidance).
- **Sectoral overlays.** Financial (NYDFS), healthcare (HIPAA), critical infra (NIS2) — sector-specific.
- **External certification audit detail.** This is the **internal** audit playbook; external (stage 1 / stage 2) audits are conducted by accredited bodies and follow ISO 17021.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 27001:2022** — the standard (Clause 9.2 internal audit + Annex A 93 controls)
- **ISO/IEC 27002:2022** — Information security controls (implementation guidance for Annex A)
- **ISO/IEC 19011:2018** — Guidelines for auditing management systems
- **ISO/IEC 17021-1:2015** — Conformity assessment requirements for bodies providing audit and certification (the external-audit standard; informs internal-audit expectations)
- **IIA International Professional Practices Framework** — Standards 1000-2600 (internal audit attribute + performance)
- **NIST SP 800-53A Rev 5** — Assessing Security and Privacy Controls (assessment procedures per control)
- **ISACA CISA Review Manual** (27th ed., 2024) — IS audit methodology
- **ASQ Certified Quality Auditor (CQA) Body of Knowledge** — quality audit methodology
- **Industry retrospectives** — common findings from accredited certification bodies (BSI, DNV, Bureau Veritas published case studies)
- **The Open Group** — Open FAIR for risk-based audit prioritization
FILE:references/security-control-testing.md
# Security Control Testing Guide
Technical verification procedures for ISO 27002 control assessment.
---
## Table of Contents
- [Control Testing Approach](#control-testing-approach)
- [Organizational Controls (A.5)](#organizational-controls-a5)
- [People Controls (A.6)](#people-controls-a6)
- [Physical Controls (A.7)](#physical-controls-a7)
- [Technological Controls (A.8)](#technological-controls-a8)
---
## Control Testing Approach
### Testing Methods
| Method | Description | When to Use |
|--------|-------------|-------------|
| Inquiry | Interview control owners | All controls |
| Observation | Watch process execution | Operational controls |
| Inspection | Review documentation/config | Policy controls |
| Re-performance | Execute control procedure | Critical controls |
### Sampling Guidelines
| Population Size | Sample Size |
|-----------------|-------------|
| 1-10 | All items |
| 11-50 | 10 items |
| 51-250 | 15 items |
| 251+ | 25 items |
---
## Organizational Controls (A.5)
### A.5.1 - Policies for Information Security
**Test Procedure:**
1. Obtain current information security policy
2. Verify management signature and approval date
3. Check policy is accessible to all employees
4. Confirm review within past 12 months
5. Sample 5 employees: verify awareness of policy location
**Evidence Required:**
- Signed policy document
- Intranet/portal screenshot showing policy access
- Policy review meeting minutes
- Employee acknowledgment records
### A.5.15 - Access Control
**Test Procedure:**
1. Obtain access control policy
2. Select sample of 10 user accounts
3. Verify access rights match job descriptions
4. Check for segregation of duties violations
5. Verify access provisioning follows documented process
**Evidence Required:**
- Access control policy
- User access matrix
- Access request forms with approvals
- Role definitions
### A.5.24 - Information Security Incident Management
**Test Procedure:**
1. Review incident management procedure
2. Select 3 recent incidents from log
3. Verify incidents followed documented process
4. Check escalation thresholds were respected
5. Confirm lessons learned were documented
**Evidence Required:**
- Incident response procedure
- Incident tickets with timeline
- Escalation records
- Post-incident review reports
---
## People Controls (A.6)
### A.6.1 - Screening
**Test Procedure:**
1. Review background check policy
2. Select 10 recent hires
3. Verify background checks completed before start
4. Check checks match role sensitivity level
5. Confirm records are securely stored
**Evidence Required:**
- Screening policy
- Background check completion records
- Role risk classification matrix
### A.6.3 - Information Security Awareness
**Test Procedure:**
1. Obtain training program documentation
2. Select sample of 15 employees
3. Verify training completion records
4. Review training content for currency
5. Check phishing simulation results
**Evidence Required:**
- Training materials and schedule
- LMS completion reports
- Phishing test results
- Training effectiveness metrics
### A.6.7 - Remote Working
**Test Procedure:**
1. Review remote working policy
2. Verify VPN is required for remote access
3. Sample 5 remote worker devices for compliance
4. Check endpoint protection is active
5. Verify secure data handling requirements
**Evidence Required:**
- Remote working policy
- VPN connection logs
- Endpoint compliance reports
- Remote access agreement signatures
---
## Physical Controls (A.7)
### A.7.1 - Physical Security Perimeters
**Test Procedure:**
1. Walk perimeter of secure areas
2. Verify access controls at all entry points
3. Check visitor management process
4. Review after-hours access logs
5. Confirm emergency exits are secure
**Evidence Required:**
- Site security plan
- Access control system configuration
- Visitor logs
- Guard tour records
### A.7.4 - Physical Security Monitoring
**Test Procedure:**
1. Verify CCTV coverage of critical areas
2. Check recording retention period
3. Review sample of recent alert responses
4. Confirm monitoring is 24/7 or as required
5. Verify footage protection and access controls
**Evidence Required:**
- CCTV coverage map
- Retention policy and settings
- Alert response records
- Access logs for footage viewing
---
## Technological Controls (A.8)
### A.8.2 - Privileged Access Rights
**Test Procedure:**
1. Obtain list of privileged accounts
2. Verify each has documented justification
3. Check separation of admin and user accounts
4. Confirm MFA is required for privileged access
5. Review privileged activity logs
**Evidence Required:**
- Privileged account inventory
- Access justification records
- PAM solution configuration
- Activity audit logs
### A.8.5 - Secure Authentication
**Test Procedure:**
1. Review password policy configuration
2. Verify MFA enrollment rates
3. Test account lockout after failed attempts
4. Check authentication logging
5. Verify secure authentication protocols (no plaintext)
**Evidence Required:**
- Password policy settings screenshot
- MFA enrollment report
- Account lockout configuration
- Authentication audit logs
### A.8.7 - Protection Against Malware
**Test Procedure:**
1. Verify endpoint protection coverage
2. Check definition update frequency
3. Review quarantine/detection logs
4. Confirm central management console
5. Test sample detection (EICAR)
**Evidence Required:**
- Endpoint protection deployment report
- Update status dashboard
- Detection/quarantine logs
- EICAR test results
### A.8.8 - Management of Technical Vulnerabilities
**Test Procedure:**
1. Obtain vulnerability scanning schedule
2. Review recent scan results
3. Verify critical vulnerabilities patched within SLA
4. Check vulnerability tracking system
5. Sample 5 critical findings for remediation evidence
**Evidence Required:**
- Scanning schedule and scope
- Scan reports with severity breakdown
- Patch deployment records
- Remediation tracking tickets
### A.8.13 - Information Backup
**Test Procedure:**
1. Review backup policy and schedule
2. Verify backup completion logs
3. Check encryption of backup data
4. Request recent restoration test results
5. Verify offsite/cloud backup location
**Evidence Required:**
- Backup policy
- Backup job completion logs
- Encryption configuration
- Restoration test records
### A.8.15 - Logging
**Test Procedure:**
1. Identify systems requiring logging
2. Verify logging is enabled and configured
3. Check log retention meets requirements
4. Confirm log integrity protection
5. Verify SIEM integration and alerting
**Evidence Required:**
- Logging requirements matrix
- Log configuration screenshots
- Retention settings
- SIEM alert rules
### A.8.24 - Use of Cryptography
**Test Procedure:**
1. Review cryptography policy
2. Verify encryption at rest configuration
3. Check TLS configuration (version, ciphers)
4. Review key management procedures
5. Verify certificate inventory and expiration tracking
**Evidence Required:**
- Cryptography policy
- Encryption configuration settings
- SSL/TLS scan results
- Key management procedures
- Certificate inventory
FILE:scripts/isms_audit_scheduler.py
#!/usr/bin/env python3
"""
ISMS Audit Scheduler
Risk-based audit planning and scheduling for ISO 27001 compliance.
Generates annual audit plans based on control risk ratings.
Usage:
python isms_audit_scheduler.py --year 2025 --output audit_plan.json
python isms_audit_scheduler.py --controls controls.csv --format markdown
"""
import argparse
import csv
import json
import sys
from datetime import datetime, timedelta
from typing import Dict, List, Any, Optional
# ISO 27001:2022 Annex A control domains
CONTROL_DOMAINS = {
"A.5": {"name": "Organizational Controls", "count": 37},
"A.6": {"name": "People Controls", "count": 8},
"A.7": {"name": "Physical Controls", "count": 14},
"A.8": {"name": "Technological Controls", "count": 34},
}
# Default risk ratings for control areas
DEFAULT_RISK_RATINGS = {
"A.5.1": {"name": "Policies for information security", "risk": "medium"},
"A.5.2": {"name": "Information security roles", "risk": "medium"},
"A.5.15": {"name": "Access control", "risk": "high"},
"A.5.24": {"name": "Incident management planning", "risk": "high"},
"A.5.25": {"name": "Assessment of security events", "risk": "high"},
"A.6.1": {"name": "Screening", "risk": "medium"},
"A.6.3": {"name": "Information security awareness", "risk": "medium"},
"A.6.7": {"name": "Remote working", "risk": "high"},
"A.7.1": {"name": "Physical security perimeters", "risk": "medium"},
"A.7.4": {"name": "Physical security monitoring", "risk": "medium"},
"A.8.2": {"name": "Privileged access rights", "risk": "critical"},
"A.8.5": {"name": "Secure authentication", "risk": "critical"},
"A.8.7": {"name": "Protection against malware", "risk": "high"},
"A.8.8": {"name": "Management of vulnerabilities", "risk": "critical"},
"A.8.13": {"name": "Information backup", "risk": "high"},
"A.8.15": {"name": "Logging", "risk": "critical"},
"A.8.20": {"name": "Networks security", "risk": "high"},
"A.8.24": {"name": "Use of cryptography", "risk": "high"},
}
# Audit frequency based on risk level
AUDIT_FREQUENCY = {
"critical": 4, # Quarterly
"high": 2, # Semi-annual
"medium": 1, # Annual
"low": 1, # Annual
}
def load_controls_from_csv(filepath: str) -> Dict[str, Dict]:
"""Load control risk ratings from CSV file."""
controls = {}
try:
with open(filepath, "r", encoding="utf-8") as f:
reader = csv.DictReader(f)
for row in reader:
control_id = row.get("control_id", row.get("id", ""))
if control_id:
controls[control_id] = {
"name": row.get("name", "Unknown"),
"risk": row.get("risk", "medium").lower(),
}
except FileNotFoundError:
print(f"Error: File not found: {filepath}", file=sys.stderr)
sys.exit(1)
return controls
def calculate_audit_dates(
year: int,
frequency: int
) -> List[str]:
"""Calculate audit dates based on frequency."""
dates = []
interval = 12 // frequency
for i in range(frequency):
month = (i * interval) + 2 # Start in February
if month > 12:
month = month - 12
date = datetime(year, month, 15)
dates.append(date.strftime("%Y-%m-%d"))
return dates
def generate_audit_plan(
year: int,
controls: Optional[Dict[str, Dict]] = None
) -> Dict[str, Any]:
"""Generate risk-based annual audit plan."""
if controls is None:
controls = DEFAULT_RISK_RATINGS
plan = {
"metadata": {
"year": year,
"generated": datetime.now().isoformat(),
"methodology": "ISO 27001 Risk-Based Internal Auditing",
"total_controls": len(controls),
},
"schedule": {
"Q1": {"month": "February-March", "audits": []},
"Q2": {"month": "May-June", "audits": []},
"Q3": {"month": "August-September", "audits": []},
"Q4": {"month": "November", "audits": []},
},
"controls": {},
}
# Assign controls to quarters based on risk
for control_id, control_data in controls.items():
risk = control_data.get("risk", "medium")
frequency = AUDIT_FREQUENCY.get(risk, 1)
audit_dates = calculate_audit_dates(year, frequency)
plan["controls"][control_id] = {
"name": control_data.get("name", "Unknown"),
"risk": risk,
"frequency": frequency,
"scheduled_audits": audit_dates,
}
# Add to quarterly schedule
for i, date in enumerate(audit_dates):
month = int(date.split("-")[1])
if month <= 3:
quarter = "Q1"
elif month <= 6:
quarter = "Q2"
elif month <= 9:
quarter = "Q3"
else:
quarter = "Q4"
plan["schedule"][quarter]["audits"].append({
"control_id": control_id,
"control_name": control_data.get("name", "Unknown"),
"risk_level": risk,
"target_date": date,
})
# Sort audits within each quarter
for quarter in plan["schedule"]:
plan["schedule"][quarter]["audits"].sort(
key=lambda x: (
{"critical": 0, "high": 1, "medium": 2, "low": 3}.get(x["risk_level"], 4),
x["target_date"]
)
)
# Calculate summary statistics
risk_counts = {"critical": 0, "high": 0, "medium": 0, "low": 0}
total_audits = 0
for control_data in plan["controls"].values():
risk_counts[control_data["risk"]] += 1
total_audits += control_data["frequency"]
plan["summary"] = {
"total_controls_in_scope": len(controls),
"total_audits_planned": total_audits,
"risk_distribution": risk_counts,
"audits_per_quarter": {
q: len(plan["schedule"][q]["audits"])
for q in plan["schedule"]
},
}
return plan
def format_markdown(plan: Dict[str, Any]) -> str:
"""Format audit plan as markdown."""
lines = [
f"# ISMS Audit Plan {plan['metadata']['year']}",
f"",
f"**Generated:** {plan['metadata']['generated'][:10]}",
f"**Methodology:** {plan['metadata']['methodology']}",
f"",
f"## Summary",
f"",
f"| Metric | Value |",
f"|--------|-------|",
f"| Controls in Scope | {plan['summary']['total_controls_in_scope']} |",
f"| Total Audits Planned | {plan['summary']['total_audits_planned']} |",
f"| Critical Risk Controls | {plan['summary']['risk_distribution']['critical']} |",
f"| High Risk Controls | {plan['summary']['risk_distribution']['high']} |",
f"| Medium Risk Controls | {plan['summary']['risk_distribution']['medium']} |",
f"",
]
for quarter, data in plan["schedule"].items():
lines.extend([
f"## {quarter}: {data['month']}",
f"",
f"| Control | Name | Risk | Target Date |",
f"|---------|------|------|-------------|",
])
for audit in data["audits"]:
lines.append(
f"| {audit['control_id']} | {audit['control_name']} | "
f"{audit['risk_level'].capitalize()} | {audit['target_date']} |"
)
lines.append("")
lines.extend([
f"## Risk-Based Audit Frequency",
f"",
f"| Risk Level | Audit Frequency |",
f"|------------|-----------------|",
f"| Critical | Quarterly (4x/year) |",
f"| High | Semi-Annual (2x/year) |",
f"| Medium | Annual (1x/year) |",
f"| Low | Annual (1x/year) |",
])
return "\n".join(lines)
def main():
parser = argparse.ArgumentParser(
description="ISMS Audit Scheduler - Risk-based audit planning"
)
parser.add_argument(
"--year", "-y",
type=int,
default=datetime.now().year,
help="Audit plan year (default: current year)"
)
parser.add_argument(
"--controls", "-c",
help="CSV file with control risk ratings"
)
parser.add_argument(
"--output", "-o",
help="Output file path"
)
parser.add_argument(
"--format", "-f",
choices=["json", "markdown"],
default="json",
help="Output format (default: json)"
)
args = parser.parse_args()
# Load controls
controls = None
if args.controls:
controls = load_controls_from_csv(args.controls)
# Generate plan
plan = generate_audit_plan(args.year, controls)
# Format output
if args.format == "markdown":
output = format_markdown(plan)
else:
output = json.dumps(plan, indent=2)
# Write output
if args.output:
with open(args.output, "w", encoding="utf-8") as f:
f.write(output)
print(f"Audit plan saved to: {args.output}", file=sys.stderr)
else:
print(output)
if __name__ == "__main__":
main()
Chuẩn bị đánh giá QMS theo ISO 13485 bằng 6 câu hỏi chất vấn, tập trung kiểm soát thiết kế, CAPA và giám sát sau thị trường.
--- name: "iso13485-audit-prep" description: "/cs:iso13485-audit-prep <scope> — ISO 13485 QMS audit 6-question forcing interrogation. Design controls + CAPA + post-market focused. Use before Clause 8.2.4 internal audit, MDR / FDA QSR alignment review, or product-launch DHF closure audit." --- # /cs:iso13485-audit-prep — ISO 13485 QMS Forcing Questions **Command:** `/cs:iso13485-audit-prep <scope>` The ISO 13485 QMS auditor pressure-tests any medical-device QMS work. Six traceability-obsessed questions before any internal audit, MDR / FDA QSR review, or product launch. ## When to Run - Before annual Clause 8.2.4 internal audit - Before MDR / FDA QSR alignment review (substantially harmonized post Feb 2026) - Before new-device commercial launch (DHF closure audit) - After significant CAPA closure event (effectiveness verification audit) - Post-recall event (root cause + corrective action audit) - Quarterly during regulatory submission preparation ## The Six QMS Questions ### 1. Pull three random DHFs. Are design verification + validation evidence complete? **Most-cited finding area.** - DHF must include: design plan + inputs + outputs + verification + validation + transfer + changes - Sample stratified by product class (I, IIa, IIb, III per MDR) - Reference `iso13485_audit_playbook.md` for the per-DHF checklist - Verify traceability matrix from user needs through clinical evidence ### 2. Show me the last 5 CAPAs with effectiveness verification evidence. **Second-most-cited finding area.** - Containment / correction / corrective action distinction documented - Root cause analysis depth: 5 Why minimum - Effectiveness verification = measurable evidence, not "we updated the procedure" - Closure approved by appropriate authority - Repeat CAPAs across products = systemic issue trigger ### 3. When was process validation (IQ/OQ/PQ) last revalidated? **Clause 7.5.6 — often stale.** - Initial validation at process introduction - Revalidation triggers: process change, equipment change, material change, periodic schedule - Trend monitoring (SPC) where statistical techniques apply per Clause 8.4 - Cross-check with cs-fda-qsr-auditor for 21 CFR 820.75 alignment ### 4. Show me the risk management file for the highest-risk product. **Clause 7.1 + ISO 14971:2019.** - Risk management plan exists per product - Hazard identification covers reasonable foreseeable misuse - Risk control hierarchy applied: inherent safety > protective measures > information for safety - Residual risk evaluated + accepted with rationale - Post-production information feeds back into RMF - For AI-enabled medical devices: layer ISO 42001 A.5 impact assessment on top ### 5. Show me post-market surveillance evidence — last 6 months. **Clause 8.2.1 — high-stakes for MDR + FDA.** - Customer complaint log + investigation closure - Vigilance reports (serious incident / FSCA) submitted per applicable regulation - Trend analysis evidence + management review input - Post-market clinical follow-up (PMCF) for MDR high-risk devices - MDR reports per 21 CFR 803 for US-marketed devices (cross-check with cs-fda-qsr-auditor) ### 6. Where's the management review evidence covering all Clause 5.6 inputs? **Annual minimum; semi-annual for mature programs.** - Required inputs per Clause 5.6.2: audit results, customer feedback, process performance, product conformity, status of preventive + corrective actions, follow-up from prior reviews, changes that could affect QMS, recommendations for improvement, regulatory requirements - Outputs per Clause 5.6.3: improvement decisions, product requirement changes, resource needs - Integrated review across frameworks (per `multi_framework_audit_playbook.md`) preferred ## Workflow ```bash # 1. Audit programme optimization python ../../ra-qm-team/skills/qms-audit-expert/scripts/audit_schedule_optimizer.py audit_scope.json # 2. Mock audit for readiness check python ../../skills/compliance-os/scripts/audit_simulator.py iso13485_scope.json # 3. CAPA system review # Route to ra-qm-team/skills/capa-officer/ tools # 4. Risk management file review # Route to ra-qm-team/skills/risk-management-specialist/ tools ``` ## Output Format ```markdown # ISO 13485 Audit Prep: <scope> **Date:** YYYY-MM-DD ## The Decision Being Made [programme-plan | DHF-closure | CAPA-health | post-market-trend | pre-cert | MDR-FDA-alignment] ## Design Control Status (sampled DHFs) - DHFs sampled: <list product IDs> - Verification evidence: pass/fail per DHF - Validation evidence: pass/fail per DHF - Clinical evidence (per MDR Annex XIV / FDA 510(k)): pass/fail - Traceability matrix complete: yes/no per DHF ## CAPA Health - CAPAs sampled: N - Root cause analysis depth: adequate/inadequate per CAPA - Effectiveness verification: complete/incomplete per CAPA - Aging CAPAs > 90 days: N - Repeat issues across products: <list> ## Process Validation Status - Validations on schedule: % - Stale validations (> 12 months since revalidation): <list> - Statistical techniques applied per Clause 8.4: yes/no ## Risk Management File Status - Sampled product RMFs: <list> - Post-production updates in last 12 months: <count per product> - Residual risk acceptance signed: yes/no ## Post-Market Surveillance - Complaint trending: stable/rising - MDR / vigilance reports filed timely: % - PMCF on schedule (where required): yes/no ## Management Review Status - Last review date: YYYY-MM-DD - Required Clause 5.6.2 inputs present: yes/no - Open action items past due: N ## Cross-Framework Impact - EU MDR alignment: clean / gaps in <list> - FDA QSR alignment (post-Feb 2026): substantially harmonized; FDA-specific overlays per cs-fda-qsr-auditor - ISO 42001 AIMS overlay (if AI-enabled device): pass/fail per Annex A ## Verdict 🟢 READY | 🟡 CLOSE-DHF-GAPS-FIRST | 🔴 NOT-READY ## Top 3 Actions [3 concrete next steps with owner + corrective-action timeline] ``` ## Routing - `/cs:compliance-readiness` — for multi-framework view - `/cs:fda-qsr-audit-prep` — for FDA-specific overlay - `/cs:aims-audit` — for AI-enabled medical device ISO 42001 layer - `/cs:gdpr-audit-prep` — for personal-data overlap (clinical data, customer data) - `/cs:cpo-review` — for executive product strategy decisions - `/cs:decide` — to log the verdict ## Related - Agent: [`cs-cqm-iso13485`](../../agents/cs-cqm-iso13485.md) - Skill: [`qms-audit-expert`](../../../ra-qm-team/skills/qms-audit-expert/SKILL.md) - Playbook: [iso13485_audit_playbook.md](../../../ra-qm-team/skills/qms-audit-expert/references/iso13485_audit_playbook.md) - Adjacent: `../fda-qsr-audit-prep/`, `../aims-audit/`, `../compliance-readiness/` --- **Version:** 1.0.0
Đánh giá mức sẵn sàng ISMS theo ISO 27001 bằng 6 câu hỏi chất vấn, dùng trước đánh giá nội bộ, đánh giá giám sát hoặc chứng nhận giai đoạn 1.
--- name: "iso27001-audit-prep" description: "/cs:iso27001-audit-prep <scope> — ISO 27001 ISMS audit readiness 6-question forcing interrogation. Use before annual Clause 9.2 internal audit, surveillance audit prep, or stage 1 certification readiness." --- # /cs:iso27001-audit-prep — ISO 27001 ISMS Audit Forcing Questions **Command:** `/cs:iso27001-audit-prep <scope>` The ISO 27001 ISMS auditor pressure-tests any ISMS work. Six sample-driven questions before any internal audit, stage 1 readiness, or surveillance audit. ## When to Run - Before annual Clause 9.2 internal audit - Before stage 1 / stage 2 ISO 27001 certification audit - Before surveillance audit (year 2 / year 3) - After material change to ISMS scope (new business unit, new product line, new SaaS adoption) - Post-incident (breach triggers ad-hoc ISMS audit) - Quarterly during high-growth phase ## The Six ISMS Questions ### 1. What's the audit scope, and is rolling 3-year coverage on track? **No 3-year coverage discipline, no defensible programme.** - Every Clause 4-10 + every applicable Annex A control must be audited at least once per 3-year cycle - Run `isms_audit_scheduler.py` in `ra-qm-team/skills/isms-audit-expert/` - Confirm auditor independence — no self-audit on any sample ### 2. When was the risk register last refreshed, and are treatments linked to Annex A controls? **Stale risk register = certification finding.** - Quarterly refresh expected; annual minimum - Every high/critical risk must link to ≥ 1 Annex A control treating it - Residual risk acceptance documented + signed - Review against `iso27001_audit_playbook.md` for stage 1 expectations ### 3. Show me the access review records — quarterly cadence, the last 4 quarters. **Most-cited finding area.** - Annex A.5.15 + A.8.2 + A.8.3 access controls - Sample real records pulled from Okta / IAM, not curated audit-prep packs - For each terminated employee in last 90 days: deprovisioning evidence within 24-hour SLA - Privileged access reviewed at finer granularity ### 4. What's the supplier inventory + last review evidence? **Second-most-cited finding area.** - Annex A.5.19-A.5.21 supplier management - Critical SaaS suppliers reviewed at least annually - DPAs signed for personal-data sub-processors (cross-check with cs-dpo-gdpr) - AI-specific contract clauses where third-party AI services in use (cross-check with cs-aims-iso42001) ### 5. Where's the incident response evidence + post-incident review? **A.5.24-27 + A.6.8 — high-stakes audit area.** - Severity definitions documented + consistently applied - Last 5 incidents have post-incident review (PIR) within 30-day SLA - GDPR Article 33 / 34 notification timing aligned with A.5.24 (cross-check with cs-dpo-gdpr) - Blameless retro culture; not punitive ### 6. What's the management review cadence + inputs? **Clause 9.3 required inputs are prescriptive — easy to miss.** - Required inputs: audit results, risks, performance, nonconformities, opportunities - Schedule: annual minimum; quarterly preferred for mature programs - Outputs documented + tracked to closure - Integrated review across frameworks (per `multi_framework_audit_playbook.md`) preferred to separate reviews ## Workflow ```bash # 1. Audit programme planning python ../../ra-qm-team/skills/isms-audit-expert/scripts/isms_audit_scheduler.py audit_scope.json # 2. Mock audit for readiness check python ../../skills/compliance-os/scripts/audit_simulator.py iso27001_scope.json # 3. Cross-framework reuse (SOC 2 = 75% overlap; ISO 42001 = 60% reuse) python ../../skills/compliance-os/scripts/cross_framework_mapper.py program.json ``` ## Output Format ```markdown # ISO 27001 Audit Prep: <scope> **Date:** YYYY-MM-DD ## The Decision Being Made [programme-plan | finding-severity | cert-readiness | incident-followup] ## Audit Programme Status - Clauses scheduled this year: <list> - Annex A controls scheduled: <count> - Rolling 3-year coverage: clean | gaps in <list> - Auditor independence: clean | issues in <list> ## Risk Register Health - Last refresh: YYYY-MM-DD - High/critical risks without Annex A control link: N - Residual risk acceptance documentation: complete | gaps ## High-Stakes Controls Status - A.5.15 + A.8.2 + A.8.3 access control: pass/fail with sample - A.5.19-A.5.21 supplier mgmt: pass/fail with sample - A.5.24-27 + A.6.8 incident response: pass/fail with sample - A.8.15-16 logging: pass/fail with sample ## Management Review Status - Last review date: YYYY-MM-DD - Required Article 9.3 inputs present: yes/no - Open action items past due: N ## Cross-Framework Impact - SOC 2 controls affected: <list> - ISO 42001 controls affected (if applicable): <list> - GDPR Article 32 controls affected: <list> ## Verdict 🟢 READY | 🟡 CLOSE-CRITICALS-FIRST | 🔴 NOT-READY ## Top 3 Actions [3 concrete next steps with owner + corrective-action timeline] ``` ## Routing - `/cs:compliance-readiness` — for multi-framework view - `/cs:soc2-audit-prep` — for SOC 2 cross-walk pair (75% overlap) - `/cs:aims-audit` — for ISO 42001 AIMS cross-walk - `/cs:gdpr-audit-prep` — for Article 32 organizational measures overlap - `/cs:ciso-review` — for executive cybersecurity strategy - `/cs:decide` — to log the verdict ## Related - Agent: [`cs-ciso-iso27001`](../../agents/cs-ciso-iso27001.md) - Skill: [`isms-audit-expert`](../../../ra-qm-team/skills/isms-audit-expert/SKILL.md) - Playbook: [iso27001_audit_playbook.md](../../../ra-qm-team/skills/isms-audit-expert/references/iso27001_audit_playbook.md) - Adjacent: `../soc2-audit-prep/`, `../aims-audit/`, `../gdpr-audit-prep/`, `../compliance-readiness/` --- **Version:** 1.0.0
Hỗ trợ đánh giá nội bộ hệ thống quản lý AI theo ISO/IEC 42001: xác định khoảng cách theo Điều khoản 4-10, sổ đăng ký rủi ro AI và kiểm soát Annex A.
---
name: "iso42001-specialist"
description: "ISO/IEC 42001:2023 AI Management System (AIMS) specialist for compliance teams running internal audits. Three decisions: (1) Where are the gaps against Clauses 4-10 and what do we close first? (2) What goes in the AI risk register and which Annex A controls treat each risk? (3) What's the 12-month internal audit plan that satisfies Clause 9.2? Use when preparing for certification, scoping internal audit cycles, or onboarding AI systems into an existing ISMS (27001) / QMS (13485) program. NOT an executive AI strategy skill (see chief-ai-officer-advisor). NOT EU AI Act compliance (see compliance-team-eu-ai-act)."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: ra-qm-team
domain: ai-management-system-compliance
updated: 2026-05-13
python-tools: aims_gap_analyzer.py, ai_risk_register_builder.py, aims_audit_scheduler.py
frameworks: iso-42001, iso-23894, iso-38507, nist-ai-rmf, eu-ai-act-mapping
---
# ISO/IEC 42001 AI Management System Specialist
Internal-audit-grade operating skill for ISO/IEC 42001:2023. **Three decisions, no executive AI strategy:**
1. **Where are the AIMS gaps against Clauses 4–10?** — coverage scoring per clause + remediation priority
2. **What's the AI risk register, and which controls treat each risk?** — Annex A.2–A.10 control mapping per ISO 23894 risk method
3. **What's the Clause 9.2 internal audit plan?** — 12-month schedule with scope, frequency, auditor independence checks
This skill is **NOT a chief-ai-officer-advisor replacement**. CAIO decides whether to build/buy a model and what business risk to accept. This skill operates the management-system discipline that captures those decisions in audit-ready evidence.
This skill is **NOT an EU AI Act compliance skill**. ISO 42001 is a voluntary management-system standard; EU AI Act is binding product-safety regulation. They overlap (a high-risk AI system per Article 6(2) of the AI Act typically requires the QMS in Article 17, which ISO 42001 can satisfy in part) but the artefacts differ. See `compliance-team-eu-ai-act` for Article-level conformity assessment.
This skill is **NOT a substitute for ISO 23894 + 38507**. 42001 is the management system; 23894 is the AI risk methodology that feeds Clause 6.1; 38507 is the governance lens. The `ai_risk_register_builder.py` tool implements the 23894 process; treat the references as the methodology bridge.
## Keywords
ISO 42001, ISO/IEC 42001:2023, AI Management System, AIMS, AI governance, AI risk management, ISO 23894, AI risk assessment, ISO 38507, AI compliance, AI audit, internal audit AI, Annex A controls, AI risk register, AI policy, AI impact assessment, conformity declaration, AI lifecycle, AI risk treatment, NIST AI RMF, NIST AI Risk Management Framework, ISACA AI audit, BSI AIC4, AI assurance, responsible AI, AI ethics governance, AI system inventory, third-party AI risk, AI vendor management, AI change management, AI incident management
## Quick Start
```bash
# Decision A: AIMS gap analysis against Clauses 4-10
python scripts/aims_gap_analyzer.py # embedded sample (mid-stage AI SaaS)
python scripts/aims_gap_analyzer.py path/to/aims_evidence.json
# Decision B: AI risk register + Annex A control mapping
python scripts/ai_risk_register_builder.py # embedded 7-risk sample
python scripts/ai_risk_register_builder.py path/to/risks.json
# Decision C: Clause 9.2 internal audit 12-month plan
python scripts/aims_audit_scheduler.py # embedded 4-domain sample
python scripts/aims_audit_scheduler.py path/to/scope.json
```
## Key Questions (ask these first)
- **Does the AIMS scope statement (Clause 4.3) name every AI system, including embedded models and third-party AI services?** If "AI features added by our SaaS vendors" is not in scope, the AIMS is incomplete.
- **Does the AI policy (Clause 5.2) commit to lawful use AND beneficial purpose AND human oversight AND continual improvement?** Missing any of the four = nonconformity at certification.
- **Has the AI risk assessment (Clause 6.1.2) been re-run since the last material model change?** Concept drift is not a one-time event.
- **Who signs the AI impact assessment for high-impact systems (Annex A.5.4)?** If no signed accountability, the control is missing.
- **What's the internal audit cadence (Clause 9.2)?** ISO management-system standards expect ≥ once per 3-year cycle per clause; mature programs do annual.
- **Is there a documented procedure for AI incidents (Annex A.9.3)?** Untreated post-deployment monitoring is the #1 nonconformity in early adopters.
## Core Responsibilities
### 1. AIMS Gap Analysis (Clauses 4–10)
**The framework:** ISO 42001 follows the Annex SL high-level structure shared with ISO 9001 / 27001 / 13485. Clauses 4–10 are the management-system requirements; Annex A controls A.1–A.10 are the AI-specific operational controls.
| Clause | What it requires | Common gap |
|---|---|---|
| **4. Context** | AI scope, interested parties, external context | Scope omits third-party AI services |
| **5. Leadership** | AI policy, roles, accountability | Policy treats "AI ethics" as marketing copy, not commitment |
| **6. Planning** | AI risk + impact assessment, objectives | Risk register doesn't link to controls |
| **7. Support** | Resources, competence, awareness, documented info | Competence requirements undefined for ML engineers |
| **8. Operation** | Operational planning, AI system lifecycle | Lifecycle stages not mapped to Annex A controls |
| **9. Performance** | Monitoring, internal audit, management review | Drift monitoring exists in code but not in management review inputs |
| **10. Improvement** | Nonconformity, corrective action, continual improvement | CAPA loop separate from existing 13485/9001 CAPA — duplication |
**Run** `aims_gap_analyzer.py` with an evidence inventory JSON to score each clause (full / partial / missing) and get a prioritized remediation list.
See `references/iso42001_clauses.md` for the full clause-by-clause walkthrough with audit evidence expectations.
### 2. AI Risk Register + Annex A Control Mapping
**The framework:** Clause 6.1.2 requires AI risk assessment; Clause 6.1.3 requires risk treatment. Annex A provides 38 controls organized into 10 control categories (A.2–A.10). The risk register must show each identified risk linked to ≥ 1 control that treats it.
**Annex A control categories (the 10):**
| ID | Category | Example controls |
|---|---|---|
| **A.2** | AI policy | A.2.2 AI policy, A.2.3 alignment with other policies |
| **A.3** | Internal organization | A.3.2 AI roles & responsibilities, A.3.3 reporting concerns |
| **A.4** | Resources for AI systems | A.4.2 data resources, A.4.3 tooling, A.4.4 human resources |
| **A.5** | Assessing impacts | A.5.2 AI system impact assessment, A.5.4 documentation of impact assessment |
| **A.6** | AI system lifecycle | A.6.2.2 objectives, A.6.2.3 lifecycle phases, A.6.2.4 verification & validation |
| **A.7** | Data for AI systems | A.7.2 data management, A.7.3 data quality, A.7.4 data provenance, A.7.5 data preparation |
| **A.8** | Information for interested parties | A.8.2 system documentation, A.8.3 user information, A.8.4 communication of incidents |
| **A.9** | Use of AI systems | A.9.2 intended use, A.9.3 monitoring of operation, A.9.4 logging of system events |
| **A.10** | Third-party & customer relationships | A.10.2 supplier relationships, A.10.3 customer relationships |
ISO/IEC 23894:2023 provides the AI-specific risk-management process (the methodology); 42001 Annex A provides the controls. The risk register is the bridge.
**Run** `ai_risk_register_builder.py` with an identified-risks JSON to produce a structured register with mapped controls + residual-risk verdict per ISO 23894 risk-treatment options.
See `references/aims_controls_annex_a.md` for the full 38-control catalogue with audit evidence per control.
### 3. Clause 9.2 Internal Audit Plan
**The framework:** Clause 9.2 requires "internal audits at planned intervals to provide information on whether the AIMS conforms to the organization's requirements and is effectively implemented and maintained." That's the management-system requirement; the **how often** and **how deep** are organizational choices.
**Mature-program defaults:**
- Cover every clause + every applicable Annex A control over a 3-year cycle (rolling)
- Annual full-system audit covering Clauses 4, 5, 9, 10 (the "always relevant" clauses)
- Quarterly or semi-annual deep dives on Clauses 6, 7, 8 by domain (per AI system or per lifecycle phase)
- Auditor independence: nobody audits their own work; A.6 lifecycle owner cannot audit Clause 8 operation
**Run** `aims_audit_scheduler.py` with a scope JSON (AI systems in scope, prior-year findings, certification cycle phase) to produce a 12-month plan with auditor assignments and independence checks.
See `references/aims_implementation_guide.md` for the maturity model and rollout sequencing (year 1 establish, year 2 certify, year 3+ continual improvement).
## Workflows
### Workflow 1: AIMS Gap Closure for Certification (4–8 weeks)
**Goal:** Identify gaps; prioritize remediation; close before stage 1 certification audit.
```bash
# 1. Inventory current AIMS evidence (policies, procedures, records)
python scripts/aims_gap_analyzer.py aims_evidence.json
# 2. Review gap matrix; group by clause
# 3. For each gap, identify owner + due date (target: close before stage 1)
# 4. Cross-check against ISO 27001 / 13485 existing artifacts — many can be reused
# 5. Cross-check against EU AI Act obligations (use compliance-team-eu-ai-act)
# 6. Output: prioritized remediation plan with owners + dates
```
### Workflow 2: AI Risk Register Build (1–2 weeks)
**Goal:** Construct the Clause 6.1.2 risk register with full Annex A control coverage.
```bash
# 1. Run ISO 23894 risk identification across AI lifecycle (data, model, deployment, decommission)
# 2. Capture each risk with: source, event, consequence, likelihood, impact
python scripts/ai_risk_register_builder.py risks.json
# 3. For each high/critical risk, confirm ≥ 1 Annex A control is selected as treatment
# 4. Document residual risk acceptance with management signoff
# 5. Cross-check with cs-caio-advisor on executive risk acceptance for "tolerate" decisions
# 6. Log via management review (Clause 9.3)
```
### Workflow 3: Annual Internal Audit Plan (1 day)
**Goal:** Produce the 12-month Clause 9.2 plan with auditor independence.
```bash
# 1. Pull last year's audit findings and certification cycle status (year 1/2/3)
python scripts/aims_audit_scheduler.py audit_scope.json
# 2. Confirm auditor independence per assignment
# 3. Confirm coverage hits every clause and every applicable Annex A control over rolling 3 years
# 4. Submit plan for management review approval (Clause 9.3 input)
```
### Workflow 4: Cross-Framework Reuse Mapping (per system onboarded)
**Goal:** When adding a new AI system, map ISO 42001 evidence against existing 27001 + 13485 evidence to avoid duplication.
1. Pull existing ISO 27001 Annex A controls + ISO 13485 procedures relevant to the system
2. For each ISO 42001 Annex A control, identify whether an existing artifact already satisfies it (e.g., 27001 A.8.16 monitoring activities can extend to AI system monitoring)
3. Add the AI-specific overlay only where the existing control doesn't cover it
4. Document mapping in the AIMS scope statement (Clause 4.3)
## Output Standards
```
**Bottom Line:** [one sentence — gap severity + the one thing to close first]
**The Decision:** [one of: gap-closure | risk-treatment | audit-scope]
**The Evidence:** [clause numbers + control IDs from the tool, not adjectives]
**How to Act:** [3 concrete next steps with owners + dates]
**Your Decision:** [the call only the compliance officer or CAIO can make — risk acceptance, scope expansion, certification readiness]
```
## Adjacent Skills
- `../../skills/information-security-manager-iso27001/` — ISO 27001 ISMS implementation (many controls reusable for AIMS A.7 data controls)
- `../../skills/quality-manager-qms-iso13485/` — ISO 13485 QMS (provides CAPA + management-review machinery the AIMS reuses)
- `../../skills/gdpr-dsgvo-expert/` — GDPR DPIA process (input to AIMS A.5 impact assessment for personal-data systems)
- `../../skills/isms-audit-expert/` — ISO 27001 internal audit pattern (the audit scheduler mirrors this for AIMS)
- `../../skills/soc2-compliance/` — SOC 2 trust services (reusable controls for AIMS A.10 third-party relationships)
- `../../../compliance-team-eu-ai-act/` — EU AI Act Article-level compliance (binding regulation companion to voluntary 42001)
- `../../../../compliance-os/` — Meta-orchestrator for multi-framework programs (run AIMS as one framework among 9)
- `../../../../c-level-advisor/chief-ai-officer-advisor/` — Executive AI strategy (build-vs-buy, cost economics — different audience)
## References
- [iso42001_clauses.md](references/iso42001_clauses.md) — Clauses 4–10 walkthrough with audit evidence expectations, common gaps, and reusable artifacts from ISO 27001/13485
- [aims_controls_annex_a.md](references/aims_controls_annex_a.md) — All 38 Annex A controls (A.2–A.10) with implementation guidance, audit evidence, and severity of failure
- [aims_implementation_guide.md](references/aims_implementation_guide.md) — 3-year maturity model (establish → certify → continually improve), rollout sequencing, integration with existing ISMS/QMS programs
- [cross_framework_mapping_ai.md](references/cross_framework_mapping_ai.md) — ISO 42001 ↔ EU AI Act ↔ NIST AI RMF ↔ ISO 23894 ↔ ISO 38507 ↔ ISO 27001 control-level mapping with mapping-confidence ratings
---
**Version:** 1.0.0
**Status:** Production Ready
FILE:references/aims_controls_annex_a.md
# ISO/IEC 42001 Annex A — 38 Controls Catalogue
This reference answers exactly one decision: **for each Annex A control, what does implementation look like, what evidence does the auditor want, and what's the severity if it's missing?**
Pair with `scripts/ai_risk_register_builder.py` to map risks to controls.
## Structure of Annex A
ISO/IEC 42001 Annex A is a *normative* annex containing reference controls. The standard requires (per Clause 6.1.3) that the organization compare its determined controls to Annex A to verify no necessary controls have been omitted. Unlike ISO 27001 where Annex A is presumed-applicable, ISO 42001 Annex A controls are applied based on risk — if a control doesn't apply (e.g., A.10 third-party AI when you use no third-party AI), document the exclusion with justification.
**The 10 control categories (A.1 is the structural intro; A.2–A.10 are the operational controls):**
| ID | Category | Control count | Maps to clause |
|---|---|---|---|
| A.2 | Policies related to AI | 2 | 5.2 |
| A.3 | Internal organization | 2 | 5.3 |
| A.4 | Resources for AI systems | 3 | 7.1 |
| A.5 | Assessing impacts of AI systems | 3 | 6.1.4, 8.2 |
| A.6 | AI system lifecycle | 8 | 8.3 |
| A.7 | Data for AI systems | 5 | 8.3 |
| A.8 | Information for interested parties | 4 | 7.4, 9.1 |
| A.9 | Use of AI systems | 4 | 8.3, 9.1 |
| A.10 | Third-party & customer relationships | 5 | 8.4 |
Total: **38 controls** across 9 operational categories.
## A.2 — Policies (severity if missing: CRITICAL)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.2.2** | AI policy | Signed AI policy meeting Clause 5.2 requirements | ISO 27001 A.5.1 (information security policy) — extend |
| **A.2.3** | Alignment of AI policy with other policies | Mapping showing AI policy doesn't contradict info-sec, privacy, quality, code-of-conduct policies | New artifact; document the cross-references |
## A.3 — Internal Organization (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.3.2** | AI roles & responsibilities | RACI matrix; named AIMS owner | ISO 27001 A.5.2; extend to AI |
| **A.3.3** | Reporting of concerns | Whistleblower / concerns procedure for AI-specific issues (bias, harm, misuse) | Existing whistleblower; AI-extend |
## A.4 — Resources (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.4.2** | Resources — data | Data inventory; provenance; quality assessment | ISO 27001 A.5.9 inventory of assets — extend |
| **A.4.3** | Resources — tooling | Inventory of ML tooling; license & dependency tracking | Existing software-asset management |
| **A.4.4** | Resources — human resources | Competence requirements + training records (Clause 7.2) | ISO 27001 A.6.3 awareness; ISO 13485 6.2 competence |
## A.5 — Impact Assessment (severity: CRITICAL)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.5.2** | AI system impact assessment | Documented impact assessment for each AI system; covers individuals, groups, society | GDPR DPIA — partial; AI scope wider (third-party harm, environmental, societal) |
| **A.5.3** | Process for impact assessment | Documented procedure with triggers (launch, material change, complaint) | New procedure |
| **A.5.4** | Documentation of impact assessment | Signed impact assessment record with management approval for high-impact systems | New artifact |
## A.6 — AI System Lifecycle (severity: CRITICAL)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.6.1.2** | Objectives for AI system development | Stated AI-system objectives aligned to AI policy + use intent | New artifact (per system) |
| **A.6.1.3** | Processes for management of the AI system lifecycle | Procedure covering design → data → model → V&V → deployment → operation → decommission | New procedure |
| **A.6.2.2** | AI system objectives & requirements | Documented requirements traceable to objectives | ISO 13485 7.3 design & development — extend |
| **A.6.2.3** | Documentation of AI system design & development | Design records (architecture, datasets, model card) under document control | ISO 13485 7.3 — extend |
| **A.6.2.4** | Verification & validation of AI system | Test plan + evaluation results; defined acceptance criteria | New artifact per system; reference NIST AI RMF "Measure" function |
| **A.6.2.5** | Deployment of AI system | Deployment checklist; environment hand-off; rollback plan | ISO 27001 A.8.32 change management — extend |
| **A.6.2.6** | Operation & monitoring of AI system | Monitoring plan with thresholds + escalation | New per system |
| **A.6.2.7** | Technical documentation of AI system | Model card or system card per Mitchell et al. (2019) / Gebru et al. (2021) | New artifact |
## A.7 — Data for AI Systems (severity: CRITICAL)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.7.2** | Data management | Data lifecycle procedure (acquisition → use → retention → deletion) | GDPR Art. 5 data minimisation; ISO 27001 A.5.10 acceptable use |
| **A.7.3** | Data quality | Defined data-quality dimensions; measured; reported | New; reference DAMA-DMBOK 2 / ISO 8000 |
| **A.7.4** | Data provenance | Documented data lineage; consent / legitimate basis recorded | GDPR records of processing (Art. 30) — extend |
| **A.7.5** | Data preparation | Documented preprocessing procedure | New artifact per system |
| **A.7.6** | Data privacy considerations | Privacy review per data category | GDPR DPIA — extend |
## A.8 — Information for Interested Parties (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.8.2** | System documentation | Public-facing documentation per Annex A.6.2.7 | Model card / system card |
| **A.8.3** | User information | UX-level disclosure: this is AI; what it does; its limitations | New; align with EU AI Act Article 50 transparency |
| **A.8.4** | Communication of AI incidents | Incident communication procedure including external notification timing | GDPR Art. 33–34 breach notification — extend |
| **A.8.5** | Information for affected parties | Communication for AI-affected populations (those subject to AI decisions) | New; align with EU AI Act Article 86 redress |
## A.9 — Use of AI Systems (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.9.2** | Intended use of AI system | Documented intended-use statement per system | New artifact |
| **A.9.3** | Monitoring of operation | Continuous monitoring with defined metrics + thresholds | NIST AI RMF "Measure" — extend |
| **A.9.4** | Logging of AI system events | Tamper-evident logs covering decisions, drift indicators, incidents | ISO 27001 A.8.15 logging — extend |
| **A.9.5** | Use of system after deployment | Procedure for in-use changes (retraining, fine-tuning) with re-evaluation triggers | New procedure |
## A.10 — Third-Party & Customer Relationships (severity: MAJOR)
| Control | Title | What auditor wants | Reusable from |
|---|---|---|---|
| **A.10.2** | Supplier (third-party) relationships | AI-specific contract clauses (training data use, drift notification, sub-processor list) | ISO 27001 A.5.19 supplier relationships — extend |
| **A.10.3** | Customer relationships | Customer-facing AI obligations (transparency, opt-out, redress) | ISO 27001 A.5.20 — extend |
| **A.10.4** | Allocation of responsibilities between organization & third party | RACI for shared AI responsibilities (data labeling, model training, hosting, monitoring) | New artifact (per supplier) |
| **A.10.5** | Confidentiality of AI-related information | NDA scope covers AI-system internals (architecture, training data, weights) | ISO 27001 A.6.6 confidentiality — extend |
| **A.10.6** | Termination of AI service relationships | Procedure for safe AI-vendor exit (data return, model deletion, monitoring transition) | ISO 27001 A.5.20 service-level review — extend |
## How to Read This Catalogue
- **CRITICAL** = nonconformity blocks certification at stage 1
- **MAJOR** = nonconformity requires corrective action plan at stage 2; may delay certification
- **MINOR** = nonconformity recorded; corrective action expected within agreed timeline
**Audit evidence rule:** for every control selected as applicable, the auditor will ask three questions: (1) Where is the documented procedure? (2) Where are the records showing the procedure was followed? (3) Where is the evidence of management review of those records? If any of the three is missing, the control is partially implemented.
## When This Reference Doesn't Help
- **Specific Annex A control text.** This is a summary. The normative text is in ISO/IEC 42001:2023 Annex A — buy the standard.
- **Risk-to-control mapping methodology.** See `aims_implementation_guide.md` and ISO/IEC 23894:2023.
- **EU AI Act control overlap.** See `cross_framework_mapping_ai.md`.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 42001:2023** — Annex A normative controls (the authoritative source)
- **ISO/IEC 23894:2023** — AI risk management process (drives Annex A selection)
- **ISO/IEC 22989:2022** — AI concepts and terminology
- **NIST AI Risk Management Framework 1.0** (Jan 2023) + AI RMF Playbook — operational guidance mapping cleanly to Annex A
- **BSI AIC4 — Artificial Intelligence Cloud Service Compliance Criteria Catalogue** (2021) — sector-specific overlay for cloud AI providers
- **AAMI CR34971:2023** — Guidance for AI in medical devices
- **Mitchell et al.** — "Model Cards for Model Reporting" (FAT* 2019) — origin of model-card pattern referenced by A.6.2.7
- **Gebru et al.** — "Datasheets for Datasets" (CACM 2021) — datasheet pattern referenced by A.7.4
- **ISACA** — *Auditing Artificial Intelligence* (2nd ed., 2024) — practitioner audit checklist
FILE:references/aims_implementation_guide.md
# ISO/IEC 42001 — AIMS Implementation Guide (3-Year Maturity Model)
This reference answers exactly one decision: **what's the rollout sequence — what do we build in year 1 vs year 2 vs year 3, and how do we avoid recreating ISO 27001/13485 machinery?**
Pair with `scripts/aims_audit_scheduler.py` to operationalize the year-by-year audit cycle.
## The 3-Year Cycle
ISO management-system certifications follow a 3-year cycle:
| Year | Audit type | What happens |
|---|---|---|
| **Year 1** | Stage 1 (documentation review) + Stage 2 (implementation audit) → initial certification | Establish the AIMS; close major nonconformities; pass certification |
| **Year 2** | Surveillance audit (selective scope) | Demonstrate continual improvement; close minor nonconformities from year 1 |
| **Year 3** | Surveillance audit (selective scope) + recertification preparation | Full system review; prepare for year 4 recertification |
| **Year 4** | Recertification audit (full scope) | Renew certificate |
The internal audit programme (Clause 9.2) must cover every clause + every applicable Annex A control at least once per 3-year cycle. The plan must show this rolling coverage.
## Year 1 — Establish (focus: artifacts that auditors must see)
**Goal:** every clause and every applicable Annex A control has at least a documented procedure and one round of records.
### Q1: Foundations
- AI policy (Clause 5.2 + A.2.2) — board-signed
- AIMS scope statement (Clause 4.3) — names every AI system including third-party
- Roles & responsibilities (Clause 5.3 + A.3.2) — RACI with named AIMS owner
- Stakeholder & context analysis (Clause 4.1–4.2)
### Q2: Risk & impact
- AI risk register (Clause 6.1.2 + A.5) — run `ai_risk_register_builder.py`
- Risk treatment plan (Clause 6.1.3) — every high/critical risk linked to ≥ 1 Annex A control
- Impact assessment procedure (Clause 6.1.4 + A.5.3)
- AI objectives (Clause 6.2) — measurable targets
### Q3: Operations
- AI system lifecycle procedure (Clause 8.3 + A.6) — design through decommission
- Data management procedures (A.7) — data quality, provenance, preparation
- Monitoring plan per system (A.9.3)
- Third-party AI contract template (A.10.2)
### Q4: Performance
- Internal audit programme (Clause 9.2) — run `aims_audit_scheduler.py`
- Management review procedure (Clause 9.3) — inputs include AI-specific items
- CAPA integration with existing 13485/9001 CAPA loop (Clause 10.2)
- Stage 1 audit readiness check — run `aims_gap_analyzer.py`
**Year 1 success criteria:** stage 1 audit passes with 0 critical and ≤ 1 major nonconformity.
## Year 2 — Certify and operate
**Goal:** close year-1 minor nonconformities; demonstrate the system is operating, not just documented.
### Focus shifts to records (evidence the procedures are followed)
- Monthly drift monitoring records (A.9.3)
- Quarterly impact assessment reviews (A.5)
- Half-yearly third-party AI supplier reviews (A.10.2)
- Annual management review (Clause 9.3) with documented AI-specific inputs:
- Risk register changes
- Open nonconformities
- Drift events outside threshold
- Incidents per A.8.4
- Performance trends vs objectives (Clause 6.2)
**Year 2 success criteria:** surveillance audit passes; year-1 nonconformities closed; ≥ 80% of risk-register treatments fully implemented.
## Year 3 — Continually improve
**Goal:** demonstrate continual improvement (Clause 10.1) and prepare for recertification.
- Annual update to risk register based on new AI systems, regulation changes, incidents
- Re-baseline objectives (Clause 6.2) against year-1 + year-2 performance
- Audit the audit programme itself (meta-audit; common surveillance finding)
- Demonstrate at least one improvement initiative closed with measurable result
**Year 3 success criteria:** surveillance audit passes; recertification scope confirmed; trend evidence supports continual improvement claim.
## Integration With Existing ISMS (ISO 27001) and QMS (ISO 13485 / 9001)
The mistake most organizations make: building the AIMS as a parallel management system. **Don't.** ISO 42001 is intentionally Annex SL aligned to allow integration. Common integration patterns:
| Existing artifact | Extend for AIMS by adding |
|---|---|
| ISMS scope statement | List of AI systems within ISMS scope |
| Information security policy | AI-specific commitments (fairness, human oversight) |
| Risk register (27001) | AI risks tagged distinctly; same severity matrix; same treatment workflow |
| Document control procedure | Add model cards + datasheets + impact assessments to controlled documents |
| Internal audit programme | Add AI clause + Annex A controls to rotation |
| Management review | Add AI inputs (drift, incidents, risk-register changes) |
| CAPA procedure | Add AI-specific root-cause categories (data quality, model drift, prompt injection) |
| Supplier management | Add AI-specific contract clauses |
| Incident response | Add AI incidents (bias surfaced, drift exceeded, model misuse) |
**Reuse rule of thumb:** if you already operate ISO 27001 + ISO 13485 maturely, ~60% of AIMS Clauses 4–10 effort is rewriting existing artifacts to include AI scope. The remaining ~40% is Annex A operational controls (risk register details, lifecycle, V&V, monitoring, model cards) which are genuinely new.
## Sequence If Starting From Zero (No Prior Management System)
If your organization is starting AIMS without prior ISO certification:
1. **Add ISO 27001 first.** Most AIMS Clauses 4–10 evidence is satisfied by ISO 27001 evidence with AI scope appended. Doing 42001 alone is harder.
2. **Or start with NIST AI RMF.** NIST AI RMF is voluntary and US-centric but maps cleanly to 42001 Annex A. Mature on RMF for 12–18 months, then layer the management-system formality of 42001 on top.
3. **Avoid: building AIMS in isolation.** You'll recreate document control, CAPA, management review, and internal audit infrastructure that ISO 27001/13485 already standardize.
## Cost & Effort Benchmarks (informal, practitioner-reported)
| Org type | Year 1 effort (FTE-months) | Notes |
|---|---|---|
| Mature 27001 + 13485 org adding AIMS | 4–6 | Mostly Annex A overlay |
| Mature 27001 org adding AIMS (no 13485) | 8–12 | Add lifecycle procedures (A.6) net-new |
| Greenfield (no prior management system) | 24–36 | Do 27001 first, then 42001 |
Certification body fees: ~$15k–$35k for initial certification audit (stage 1 + stage 2 for a typical mid-size SaaS); ~$8k–$15k per surveillance year.
## Common Year-1 Pitfalls
1. **Treating "AI ethics" as the policy.** A poetic policy doesn't pass; auditor wants concrete commitments and a way to verify them.
2. **Risk register with no control mapping.** Register identifies risks but doesn't show which Annex A control treats each — Clause 6.1.3 fails.
3. **Lifecycle procedure that skips decommission.** Auditor will ask, "How do you safely retire an AI system?" If silence, A.6 fails.
4. **No drift threshold defined.** Monitoring "we watch it" doesn't pass; needs metric + threshold + escalation owner.
5. **Third-party AI excluded.** "Our vendors' AI features aren't ours" is wrong if you embed them in your service.
6. **No competence requirement for ML engineers.** Clause 7.2 wants documented competence requirements per role; "they have PhDs" isn't a documented requirement.
## When This Reference Doesn't Help
- **Specific Annex A control implementation.** See `aims_controls_annex_a.md`.
- **Risk identification methodology.** See ISO/IEC 23894:2023.
- **EU AI Act overlap.** See `cross_framework_mapping_ai.md` and `compliance-team-eu-ai-act/`.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 42001:2023** — the standard itself
- **ISO/IEC 23894:2023** — AI risk management process
- **ISO/IEC 38507:2022** — Governance implications of AI for organizations
- **ISO/IEC 27001:2022** — Information security management (reuse template for 60% of AIMS Clauses 4–10)
- **ISO/IEC 13485:2016** — Medical device QMS (reuse template for CAPA, document control)
- **NIST AI RMF 1.0** (Jan 2023) + AI RMF Playbook + Generative AI Profile (NIST AI 600-1, 2024)
- **BSI** — *Information technology — Artificial intelligence — Implementation guidance for ISO/IEC 42001* (2024 white paper)
- **ISACA** — *Auditing Artificial Intelligence* (2nd ed., 2024) — implementation pitfalls catalogue
- **IAPP** — AI Governance Center materials (continuously updated) — practitioner community knowledge base
FILE:references/cross_framework_mapping_ai.md
# ISO/IEC 42001 ↔ EU AI Act ↔ NIST AI RMF ↔ ISO 23894 ↔ ISO 38507 ↔ ISO 27001 — Cross-Framework Mapping
This reference answers exactly one decision: **for each ISO 42001 obligation, which other frameworks already cover it, and what evidence can I reuse?**
The point of cross-framework mapping is to avoid duplicate work. A control implemented for ISO 27001 frequently satisfies an Annex A control of ISO 42001 with minor AI-specific overlay. The `compliance-os` orchestrator's `cross_framework_mapper.py` consumes this mapping.
## High-Level Framework Comparison
| Framework | Type | Binding? | AI scope | Maturity |
|---|---|---|---|---|
| **ISO/IEC 42001:2023** | Management system standard | Voluntary; certifiable | AI Management System (AIMS) | Published 2023; certifications starting 2024 |
| **EU AI Act (Reg. 2024/1689)** | Product safety regulation | Binding in EU | Risk-based: prohibited → high-risk → limited-risk → minimal-risk | In force Aug 2024; phased obligations through 2027 |
| **NIST AI RMF 1.0** | Risk management framework | Voluntary (US) | Govern / Map / Measure / Manage functions | Released Jan 2023; mature playbook |
| **ISO/IEC 23894:2023** | Risk management methodology | Reference standard | AI risk process; informs 42001 Clause 6.1 | Published 2023 |
| **ISO/IEC 38507:2022** | Governance standard | Reference standard | Board-level AI governance | Published 2022 |
| **ISO/IEC 27001:2022** | Management system standard | Voluntary; certifiable | Information security | Mature; widely certified |
## Clause-to-Framework Mapping (ISO 42001 lens)
### Clause 4 — Context
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 | Notes |
|---|---|---|---|---|
| 4.1 External context | Art. 1 (scope); Recitals on risk-based approach | GOVERN 1.1 | 4.1 | Extend 27001 context with AI regulatory landscape |
| 4.2 Interested parties | Art. 27 (FRIA stakeholders for high-risk) | GOVERN 5 | 4.2 | Add AI-affected populations |
| 4.3 Scope | Article 6 + Annex III define what's in scope as "high-risk" | MAP 1.1 | 4.3 | Distinct artifacts; AIMS scope ≠ EU AI Act applicability scope |
| 4.4 AIMS processes | n/a | n/a | 4.4 | Integration map |
### Clause 5 — Leadership
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 / ISO 38507 |
|---|---|---|---|
| 5.1 Top-mgmt commitment | Art. 26 (deployer obligations); Art. 16 (provider obligations) | GOVERN 1 | 27001 5.1; 38507 Clauses 5–6 (governance principles) |
| 5.2 AI policy | Art. 17 (QMS for high-risk); Art. 95 (codes of conduct) | GOVERN 1.1 | 27001 5.2 — extend with AI commitments |
| 5.3 Roles & authorities | Art. 26 (deployer obligations); Art. 16 + 22 (authorized representative) | GOVERN 2.1 | 27001 5.3 |
### Clause 6 — Planning (the densest mapping)
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 23894 |
|---|---|---|---|
| 6.1.2 AI risk assessment | Art. 9 (risk management system for high-risk) | MAP 5.1; MAP 5.2 | Clauses 6–7 (entire process) |
| 6.1.3 AI risk treatment | Art. 9(2)(c–d) (risk management measures) | MANAGE 1.1 | Clauses 8 (treatment selection) |
| 6.1.4 Impact assessment | Art. 27 (Fundamental Rights Impact Assessment for high-risk public-sector deployers) | MAP 2.3; MAP 5.1 | Clause 5.3 (scope definition) |
| 6.2 AI objectives | Art. 9(2)(a) (objectives of risk management) | GOVERN 1.5; MEASURE 1 | Clause 5.2 |
### Clause 7 — Support
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 |
|---|---|---|---|
| 7.1 Resources | Art. 17(1)(c) (technical resources for QMS) | GOVERN 3 | A.6.1 |
| 7.2 Competence | Art. 14 (human oversight competence); Art. 26(2) (deployer competence) | GOVERN 3.1 | A.6.3 |
| 7.3 Awareness | Art. 14 | GOVERN 5.1 | A.6.3 |
| 7.4 Communication | Art. 50 (transparency obligations); Art. 86 (right to explanation) | GOVERN 5.2 | A.7.4 |
| 7.5 Documented info | Art. 11 + 12 (technical documentation); Art. 19 (record-keeping) | GOVERN 1.4 | 27001 7.5 |
### Clause 8 — Operation
| ISO 42001 | EU AI Act | NIST AI RMF | Notes |
|---|---|---|---|
| 8.1 Operational planning | Art. 17 (QMS) | MANAGE 2 | |
| 8.2 Impact assessment process | Art. 27 (FRIA process) | MAP 2 | |
| 8.3 AI system lifecycle | Art. 9 (full lifecycle); Art. 72 (post-market monitoring) | MAP 3; MEASURE 3; MANAGE 4 | Densest overlap |
| 8.4 Third-party / customer | Art. 25 (responsibilities along the AI value chain) | GOVERN 6 | |
### Clause 9 — Performance
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 |
|---|---|---|---|
| 9.1 Monitoring | Art. 72 (post-market monitoring system) | MEASURE 2; MEASURE 4 | 9.1 |
| 9.2 Internal audit | Art. 17(1)(j) (internal audit as part of QMS) | GOVERN 4 | 9.2 |
| 9.3 Management review | n/a explicit; implied in Art. 17 | GOVERN 1 | 9.3 |
### Clause 10 — Improvement
| ISO 42001 | EU AI Act | NIST AI RMF | ISO 27001 |
|---|---|---|---|
| 10.1 Continual improvement | Art. 9(2)(c) (iterative risk reduction) | MANAGE 4.3 | 10.1 |
| 10.2 Nonconformity & CAPA | Art. 73 (incident reporting); Art. 79 (corrective actions) | MANAGE 4.2 | 10.2 |
## Annex A Control → Framework Mapping (subset of highest-value mappings)
| ISO 42001 Annex A | EU AI Act | NIST AI RMF | ISO 27001 | Mapping confidence |
|---|---|---|---|---|
| A.2.2 AI policy | Art. 95 (codes of conduct) | GOVERN 1.1 | A.5.1 (info-sec policy) | HIGH |
| A.5.2 Impact assessment | Art. 27 FRIA | MAP 2.3 | n/a | MEDIUM (FRIA narrower) |
| A.6.2.4 V&V | Art. 15 (accuracy, robustness, cybersecurity); Art. 17(1)(h) | MEASURE 2 | n/a | HIGH |
| A.7.2 Data management | Art. 10 (data governance) | MAP 2.3; MEASURE 2.6 | A.5.10 | HIGH |
| A.7.3 Data quality | Art. 10(3) (relevance, representativeness, error-free, complete) | MEASURE 2.6 | n/a | HIGH |
| A.7.4 Data provenance | Art. 10(2)(d) (data origin) | MAP 2.3 | n/a | HIGH |
| A.7.6 Data privacy | Art. 10(5) (special categories); GDPR Articles 5, 6, 9 | MANAGE 2.1 | A.5.34 | HIGH |
| A.8.2 System docs | Art. 11 + Annex IV (technical documentation) | GOVERN 1.4 | A.5.37 | HIGH |
| A.8.3 User information | Art. 13 (instructions for use); Art. 50 (transparency) | GOVERN 5.2 | n/a | HIGH |
| A.8.4 Incident communication | Art. 73 (incident reporting to authorities) | MANAGE 4.2 | A.6.8 (reporting) | HIGH |
| A.9.3 Monitoring | Art. 72 (post-market monitoring) | MEASURE 2; MEASURE 4 | A.8.15 (logging) | HIGH |
| A.9.4 Logging | Art. 12 (record-keeping); Art. 19 | MEASURE 4 | A.8.15 | HIGH |
| A.10.2 Supplier relationships | Art. 25 (responsibilities along the AI value chain) | GOVERN 6 | A.5.19, A.5.20, A.5.21 | HIGH |
**Mapping confidence legend:**
- **HIGH** — direct overlap; same evidence can satisfy both
- **MEDIUM** — partial overlap; existing evidence with AI overlay
- **LOW** — concept overlap; mostly new artifact required
## Practical Reuse Pattern
If you operate ISO 27001 (mature) + are adopting ISO 42001:
1. **Reuse policies (~60%):** Extend info-sec policy with AI commitments (5.2 + A.2.2)
2. **Reuse procedures (~50%):** Document control, internal audit, management review, CAPA
3. **Reuse risk machinery (~70%):** Same severity matrix, same treatment workflow, same residual-risk acceptance flow — just add AI-specific risks and Annex A control mapping
4. **Reuse supplier mgmt (~80%):** Add AI-specific contract clauses to existing supplier procedure
5. **New artifacts (~40%):** Model cards / datasheets (A.6.2.7, A.7.4), impact assessments per Annex A.5, lifecycle procedure (A.6), drift monitoring (A.9.3), V&V procedure (A.6.2.4)
If you also operate ISO 13485 (medical device QMS):
- Reuse: design controls (7.3) for A.6 lifecycle; risk management (ISO 14971) overlays cleanly onto A.5 + 6.1; post-market surveillance maps directly to A.9.3 monitoring
- Add: AI-specific failure modes to ISO 14971 hazard analysis
## When This Reference Doesn't Help
- **EU AI Act conformity assessment routing.** See `compliance-team-eu-ai-act/scripts/conformity_assessment_planner.py`.
- **NIST AI RMF deep-dive.** See NIST AI RMF Playbook (NIST.AI.100-1.pdf) and Generative AI Profile (NIST.AI.600-1).
- **Multi-framework audit simulation.** See `compliance-os/scripts/audit_simulator.py`.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 42001:2023** — Annex A normative controls
- **Regulation (EU) 2024/1689** — Artificial Intelligence Act — full Articles (the binding regulation)
- **NIST AI Risk Management Framework 1.0** (Jan 2023, NIST AI 100-1) + AI RMF Playbook
- **ISO/IEC 23894:2023** — AI risk management process
- **ISO/IEC 38507:2022** — Governance implications of AI
- **ISO/IEC 27001:2022** + Annex A controls (the most cross-walked partner standard)
- **EDPB Opinion 28/2024** — Guidelines on processing of personal data in AI models
- **European Commission AI Act Guidelines** (continuously updated): Guidelines on prohibited practices (Feb 2025), Guidelines on definition of AI system (Feb 2025), FRIA template guidance
- **BSI** — *Cross-walking ISO 42001 and EU AI Act* (white paper, 2024)
- **IAPP EU AI Act Tracker** (continuously updated) — practitioner reference for Article applicability
FILE:references/iso42001_clauses.md
# ISO/IEC 42001:2023 — Clauses 4-10 Walkthrough
This reference answers exactly one decision: **for each clause of ISO 42001, what audit evidence does the certification body expect, and which existing ISMS/QMS artifact can I reuse?**
Pair with `scripts/aims_gap_analyzer.py` for automated coverage scoring.
## Annex SL High-Level Structure
ISO/IEC 42001:2023 follows the Annex SL structure shared by ISO 9001, 14001, 27001, 13485, 45001, and other management-system standards. This is deliberate: certification bodies, internal auditors, and quality teams can apply existing competencies to AIMS audits with low ramp-up cost.
**Practical implication:** if your organization already operates ISO 27001 + ISO 13485, ~60% of Clauses 4–10 artefacts (scope statements, policies, document control, internal audit programme, management review) can be **extended** to cover AI scope rather than recreated. The gap analysis is mostly Annex A (AI-specific operational controls), not Clauses 4–10.
## Clause 4 — Context of the Organization
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **4.1** | External & internal issues affecting AIMS | Documented context analysis (PESTLE or equivalent); reviewed at management review | Treating AI regulatory landscape as static; missing EU AI Act, US state laws, sector-specific AI rules |
| **4.2** | Needs & expectations of interested parties | Stakeholder matrix: customers, regulators, employees, data subjects, model providers, AI-affected populations | Omitting "AI-affected populations" (people who never interact with the system but are subject to its decisions) |
| **4.3** | AIMS scope statement | Documented scope: which AI systems, which lifecycle phases, which organizational units, which exclusions | Scope omits third-party AI services (SaaS features powered by vendor models); excludes "experimental" systems that are in fact in production |
| **4.4** | AIMS processes & interactions | Process map showing how AIMS processes connect to existing QMS/ISMS processes | Treating AIMS as parallel system instead of integrated extension of existing management systems |
**Reusable from ISO 27001 / 13485:** scope statement template, stakeholder matrix template, process map.
## Clause 5 — Leadership
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **5.1** | Top-management commitment | Documented evidence: AI in board agenda, resource allocation, KPIs | "AI ethics" reduced to marketing copy with no operating commitment |
| **5.2** | AI policy | Signed AI policy committing to lawful use, beneficial purpose, human oversight, continual improvement | Policy doesn't mention human oversight (Annex A.9 requirement); missing commitment to continual improvement |
| **5.3** | Organizational roles, responsibilities, authorities | RACI matrix for AIMS roles; named AIMS owner; AI ethics review board (if applicable) | No named AIMS owner; CISO assumed to "cover AI" without explicit assignment |
**Critical:** Clause 5.2 has a higher evidence bar than ISO 27001/13485 because the AI policy must address fairness, transparency, and human oversight — concepts absent from older management systems. Cannot be satisfied by extending existing policies; needs net-new content.
## Clause 6 — Planning
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **6.1.2** | AI risk assessment | Risk register per ISO 23894 methodology; covers full AI lifecycle | Risk identification at deployment only, missing data + model + decommission phases |
| **6.1.3** | AI risk treatment | Treatment plan linking each risk to Annex A controls; residual-risk acceptance documented | Treatment plan exists but is generic ("apply A.7.3") without specific implementation |
| **6.1.4** | AI system impact assessment | Documented impact assessment per Annex A.5.2 for high-impact systems | Confusing impact assessment (Clause 6.1.4) with risk assessment (Clause 6.1.2) |
| **6.2** | AI objectives | Measurable AI objectives aligned to AI policy; reviewed in management review | Objectives are aspirational ("ethical AI") without measurable targets |
**Run** `ai_risk_register_builder.py` to operationalize 6.1.2 + 6.1.3.
## Clause 7 — Support
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **7.1** | Resources for AIMS | Budget; tooling; compute resources documented | Compute resources for ML training treated as one-off project cost, not ongoing AIMS resource |
| **7.2** | Competence | Defined competence requirements per role (ML eng, AI risk, data steward); training records | Competence requirements undefined for ML engineers; assumes "they have degrees" |
| **7.3** | Awareness | AI awareness training across all employees with AI-system access | Training is engineer-only; product, marketing, customer success bypass |
| **7.4** | Communication | Documented internal + external communications procedure for AI | No procedure for communicating AI incidents to users (Annex A.8.4 link) |
| **7.5** | Documented information | Version-controlled AIMS documentation | Model cards exist but are not under document control; can be edited without approval |
## Clause 8 — Operation
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **8.1** | Operational planning & control | Operational procedures for each AI lifecycle phase | Operations procedures don't define phase transitions (when does "development" become "production"?) |
| **8.2** | Impact assessment process | Operational procedure for triggering impact assessment; gate before launch | Impact assessment treated as one-time launch artifact, not re-triggered on material change |
| **8.3** | AI system lifecycle process | Documented lifecycle covering: design → data → model → V&V → deployment → operation → decommission | Lifecycle skips "decommission"; no procedure for sunsetting AI systems |
| **8.4** | Third-party / customer relationships | Supplier and customer relationship procedures; AI-specific clauses in contracts | Standard vendor contracts not updated for AI-specific obligations (data use, model retraining, drift) |
## Clause 9 — Performance Evaluation
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **9.1** | Monitoring, measurement, analysis & evaluation | Defined metrics for AI performance, fairness, drift; monitoring records | Drift monitoring in code but no defined acceptable drift threshold; no escalation path |
| **9.2** | Internal audit programme | 12-month audit plan; auditor independence documented; findings tracked | No formal AIMS audit programme; audits happen ad hoc; auditors audit own work |
| **9.3** | Management review | Documented management review at planned intervals with required inputs/outputs | Management review inputs missing AI-specific items (drift, incidents, risk-register changes) |
**Run** `aims_audit_scheduler.py` to generate the 9.2 plan with independence checks.
## Clause 10 — Improvement
| Sub-clause | Requirement | Audit evidence | Common gap |
|---|---|---|---|
| **10.1** | Continual improvement | Evidence of AIMS improvement over time (KPIs trending, control maturity rising) | "Continual improvement" treated as audit closure activity, not ongoing |
| **10.2** | Nonconformity & corrective action | CAPA records for AIMS nonconformities; root cause analysis documented | AIMS CAPA loop separate from existing 13485/9001 CAPA loop — duplicated effort, divergent procedures |
**Reusable from ISO 13485 / 9001:** the entire CAPA machinery. Add AI-specific root-cause categories (data quality, model drift, prompt injection, etc.) to the existing taxonomy.
## When This Reference Doesn't Help
- **Specific AI risk identification.** See `aims_controls_annex_a.md` and ISO/IEC 23894:2023.
- **EU AI Act conformity assessment.** Different standard. See `compliance-team-eu-ai-act`.
- **Model cards, datasheets, evaluation methodology.** Tactical artefacts; reference NIST AI RMF playbook + papers like Mitchell et al. (2019).
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 42001:2023** — Information technology — Artificial intelligence — Management system (the standard itself; published 2023-12-18 by ISO/IEC JTC 1/SC 42)
- **ISO/IEC 23894:2023** — AI risk management process (the methodology referenced by Clause 6.1.2)
- **ISO/IEC 38507:2022** — Governance implications of AI for organizations (board-level governance lens referenced by Clause 5)
- **ISO/IEC 22989:2022** — AI concepts and terminology (definitions used throughout)
- **Annex SL** in the ISO/IEC Directives Part 1 (2024) — the high-level structure shared by ISO management-system standards
- **BSI AI Management System (AIMS) Implementation Guide** (BSI, 2024) — practitioner walkthrough
- **AAMI CR34971:2023** — AI guidance for medical devices (cross-walks 42001 to medical device QMS)
- **ISACA** — *Auditing Artificial Intelligence* (2nd ed., 2024) — internal-audit-oriented checklist with ISO 42001 mapping
FILE:scripts/aims_audit_scheduler.py
#!/usr/bin/env python3
"""aims_audit_scheduler.py — ISO/IEC 42001 Clause 9.2 internal audit plan generator.
Stdlib-only. Produces a 12-month internal audit schedule for an AIMS with:
- quarterly audit slots
- clause + Annex A control coverage per slot
- auditor assignments with independence checks (no self-audit)
- rolling 3-year coverage to ensure every clause + applicable control is audited
- prior-year nonconformity follow-up scheduled in Q1
Deterministic logic. No LLM calls. Stdlib only.
Input schema (JSON):
{
"organization": "Acme AI Inc.",
"audit_year": 2026,
"certification_cycle_phase": "year_2", # year_1 | year_2 | year_3 | surveillance
"ai_systems_in_scope": ["recommendation_engine", "internal_llm_tools", "vendor_ai_chatbot"],
"applicable_annex_a_controls": ["A.2.2", "A.3.2", "A.5.2", "A.6.2.4", "A.7.3", "A.8.4", "A.9.3", "A.10.2"],
"auditors": [
{"id": "alice", "name": "Alice Chen", "role": "quality_engineer", "owns_clauses": ["8.3"]},
{"id": "bob", "name": "Bob Singh", "role": "ml_engineer", "owns_clauses": ["8.3", "A.6.2.4"]},
{"id": "carol", "name": "Carol Diaz", "role": "external_auditor", "owns_clauses": []},
{"id": "dave", "name": "Dave Park", "role": "ciso", "owns_clauses": ["A.10.2"]}
],
"prior_year_findings": [
{"clause": "9.2", "severity": "major", "status": "open"},
{"clause": "A.7.3", "severity": "minor", "status": "closed"}
]
}
Usage:
python aims_audit_scheduler.py
python aims_audit_scheduler.py path/to/scope.json
python aims_audit_scheduler.py scope.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"organization": "Acme AI Inc.",
"audit_year": 2026,
"certification_cycle_phase": "year_2",
"ai_systems_in_scope": ["recommendation_engine", "internal_llm_tools", "vendor_ai_chatbot"],
"applicable_annex_a_controls": [
"A.2.2", "A.3.2", "A.5.2", "A.6.2.4", "A.7.3", "A.8.4", "A.9.3", "A.10.2"
],
"auditors": [
{"id": "alice", "name": "Alice Chen", "role": "quality_engineer", "owns_clauses": ["8.3"]},
{"id": "bob", "name": "Bob Singh", "role": "ml_engineer", "owns_clauses": ["8.3", "A.6.2.4"]},
{"id": "carol", "name": "Carol Diaz", "role": "external_auditor", "owns_clauses": []},
{"id": "dave", "name": "Dave Park", "role": "ciso", "owns_clauses": ["A.10.2"]},
],
"prior_year_findings": [
{"clause": "9.2", "severity": "major", "status": "open"},
{"clause": "A.7.3", "severity": "minor", "status": "closed"},
],
}
# Always-audit clauses (full coverage every year)
ANNUAL_CLAUSES = ["4.3", "5.1", "5.2", "5.3", "9.3", "10.2"]
# 3-year rotation for deep-dive clauses
ROTATION_Q2 = ["6.1.2", "6.1.3", "6.1.4", "6.2"]
ROTATION_Q3 = ["7.1", "7.2", "7.3", "7.4", "7.5", "8.1", "8.2", "8.3", "8.4"]
ROTATION_Q4 = ["9.1", "9.2", "10.1"]
def assign_auditor(scope_items: List[str], auditors: List[Dict[str, Any]]) -> Dict[str, Any]:
"""Pick the auditor with the fewest independence conflicts in this scope."""
best_auditor = None
best_conflicts = 999
for a in auditors:
owns = set(a.get("owns_clauses", []))
conflicts = sum(1 for s in scope_items if s in owns)
if conflicts < best_conflicts:
best_conflicts = conflicts
best_auditor = a
if best_auditor is None:
return {"id": None, "name": "UNASSIGNED", "independent": False, "conflicts": []}
owns = set(best_auditor.get("owns_clauses", []))
conflicts = [s for s in scope_items if s in owns]
return {
"id": best_auditor["id"],
"name": best_auditor["name"],
"role": best_auditor["role"],
"independent": len(conflicts) == 0,
"conflicts": conflicts,
}
def build_quarter(label: str, scope_clauses: List[str], scope_controls: List[str],
auditors: List[Dict[str, Any]], extra_notes: str = "") -> Dict[str, Any]:
all_scope = scope_clauses + scope_controls
auditor = assign_auditor(all_scope, auditors)
return {
"quarter": label,
"scope_clauses": scope_clauses,
"scope_annex_a_controls": scope_controls,
"auditor": auditor,
"notes": extra_notes,
}
def plan(payload: Dict[str, Any]) -> Dict[str, Any]:
year = int(payload.get("audit_year", 2026))
phase = payload.get("certification_cycle_phase", "year_2")
systems = payload.get("ai_systems_in_scope", [])
controls = payload.get("applicable_annex_a_controls", [])
auditors = payload.get("auditors", [])
prior_findings = payload.get("prior_year_findings", [])
open_priors = [f for f in prior_findings if f.get("status") != "closed"]
# 3-year control rotation: split applicable controls into thirds
third = max(1, len(controls) // 3)
controls_y1 = controls[0:third]
controls_y2 = controls[third:2 * third]
controls_y3 = controls[2 * third:]
phase_to_controls = {
"year_1": controls_y1, "year_2": controls_y2,
"year_3": controls_y3, "surveillance": controls_y3,
}
this_year_controls = phase_to_controls.get(phase, controls_y2)
# Q1: leadership + scope + prior-year follow-up
q1_clauses = ["4.3", "5.1", "5.2", "5.3"]
q1_notes = f"Follow up {len(open_priors)} open prior-year finding(s)." if open_priors else "No open priors."
q1 = build_quarter(f"Q1 {year}", q1_clauses, [], auditors, q1_notes)
# Q2: planning + objectives + risk
q2 = build_quarter(f"Q2 {year}", ROTATION_Q2, this_year_controls[:max(1, len(this_year_controls) // 2)], auditors)
# Q3: support + operation
q3_controls = this_year_controls[max(1, len(this_year_controls) // 2):]
q3_notes = f"Deep-dive across {len(systems)} AI systems: {', '.join(systems)}."
q3 = build_quarter(f"Q3 {year}", ROTATION_Q3, q3_controls, auditors, q3_notes)
# Q4: performance + improvement + management review
q4_notes = "Management review inputs prepared per Clause 9.3."
q4 = build_quarter(f"Q4 {year}", ROTATION_Q4 + ANNUAL_CLAUSES[-2:], [], auditors, q4_notes)
# Independence audit
quarters = [q1, q2, q3, q4]
independence_issues = [{
"quarter": q["quarter"], "auditor": q["auditor"]["name"], "conflicts": q["auditor"]["conflicts"]
} for q in quarters if not q["auditor"]["independent"]]
# Coverage check
audited_clauses = set()
audited_controls = set()
for q in quarters:
audited_clauses.update(q["scope_clauses"])
audited_controls.update(q["scope_annex_a_controls"])
return {
"organization": payload.get("organization"),
"audit_year": year,
"certification_cycle_phase": phase,
"ai_systems_in_scope": systems,
"open_prior_findings": len(open_priors),
"quarters": quarters,
"independence_issues": independence_issues,
"coverage_summary": {
"clauses_audited_this_year": sorted(audited_clauses),
"controls_audited_this_year": sorted(audited_controls),
"controls_deferred_to_future_years": sorted(
set(controls) - audited_controls
),
},
}
def render_text(p: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("ISO/IEC 42001 — CLAUSE 9.2 INTERNAL AUDIT PLAN")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Organization: {p['organization']}")
lines.append(f"Year: {p['audit_year']} | Cert cycle phase: {p['certification_cycle_phase']}")
lines.append(f"AI systems in scope: {', '.join(p['ai_systems_in_scope'])}")
lines.append(f"Open prior-year findings: {p['open_prior_findings']}")
lines.append("")
lines.append("-" * 72)
lines.append("QUARTERLY SCHEDULE:")
lines.append("")
for q in p["quarters"]:
a = q["auditor"]
flag = "" if a["independent"] else " ⚠️ INDEPENDENCE CONFLICT"
lines.append(f" {q['quarter']} → Auditor: {a['name']} ({a['role']}){flag}")
if q["scope_clauses"]:
lines.append(f" Clauses: {', '.join(q['scope_clauses'])}")
if q["scope_annex_a_controls"]:
lines.append(f" Annex A: {', '.join(q['scope_annex_a_controls'])}")
if a["conflicts"]:
lines.append(f" ⚠️ Conflicts on: {', '.join(a['conflicts'])} — reassign or use external auditor")
if q["notes"]:
lines.append(f" Notes: {q['notes']}")
lines.append("")
if p["independence_issues"]:
lines.append("-" * 72)
lines.append(f"INDEPENDENCE ISSUES ({len(p['independence_issues'])}):")
for issue in p["independence_issues"]:
lines.append(f" - {issue['quarter']}: {issue['auditor']} owns {', '.join(issue['conflicts'])}")
lines.append("")
c = p["coverage_summary"]
lines.append("-" * 72)
lines.append("3-YEAR COVERAGE STATUS:")
lines.append(f" Clauses audited this year ({len(c['clauses_audited_this_year'])}): {', '.join(c['clauses_audited_this_year'])}")
lines.append(f" Annex A controls audited this year ({len(c['controls_audited_this_year'])}): {', '.join(c['controls_audited_this_year']) or 'none'}")
lines.append(f" Controls deferred to future years ({len(c['controls_deferred_to_future_years'])}): {', '.join(c['controls_deferred_to_future_years']) or 'none'}")
lines.append("")
lines.append("RULES: every clause + every applicable Annex A control must be audited at least once per 3-year cert cycle.")
lines.append(" Same auditor cannot audit work they own (Clause 9.2 independence).")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="ISO/IEC 42001 Clause 9.2 internal audit 12-month plan generator.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to audit scope JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: year-2 cert cycle, 3 systems, 8 controls applicable>"
result = plan(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/aims_gap_analyzer.py
#!/usr/bin/env python3
"""aims_gap_analyzer.py — ISO/IEC 42001:2023 AIMS gap analysis against Clauses 4-10.
Stdlib-only. Scores each clause as 'full' / 'partial' / 'missing' based on an evidence
inventory and outputs a prioritized remediation list with severity at certification audit.
Deterministic logic. No LLM calls. No external dependencies.
Input schema (JSON):
{
"organization": "Acme AI Inc.",
"scope_statement": "Customer-facing recommendation engine + internal LLM tools",
"certification_target": "stage_1_audit_in_q3",
"evidence": {
"4.1_context_external": "documented",
"4.2_interested_parties": "documented",
"4.3_scope_statement": "documented",
"4.4_aims_processes": "partial",
"5.1_leadership_commitment": "documented",
"5.2_ai_policy": "partial",
"5.3_roles_responsibilities": "missing",
"6.1.2_risk_assessment": "documented",
"6.1.3_risk_treatment": "partial",
"6.1.4_impact_assessment": "missing",
"6.2_objectives": "documented",
"7.1_resources": "documented",
"7.2_competence": "missing",
"7.3_awareness": "partial",
"7.4_communication": "documented",
"7.5_documented_info": "documented",
"8.1_operational_planning": "documented",
"8.2_impact_assessment_process": "partial",
"8.3_ai_system_lifecycle": "missing",
"8.4_third_party_relationships": "partial",
"9.1_monitoring": "partial",
"9.2_internal_audit": "missing",
"9.3_management_review": "documented",
"10.1_continual_improvement": "partial",
"10.2_nonconformity_capa": "documented"
}
}
Usage:
python aims_gap_analyzer.py # uses embedded sample
python aims_gap_analyzer.py path/to/evidence.json
python aims_gap_analyzer.py evidence.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"organization": "Acme AI Inc.",
"scope_statement": "Customer-facing recommendation engine + internal LLM tools",
"certification_target": "stage_1_audit_in_q3",
"evidence": {
"4.1_context_external": "documented",
"4.2_interested_parties": "documented",
"4.3_scope_statement": "documented",
"4.4_aims_processes": "partial",
"5.1_leadership_commitment": "documented",
"5.2_ai_policy": "partial",
"5.3_roles_responsibilities": "missing",
"6.1.2_risk_assessment": "documented",
"6.1.3_risk_treatment": "partial",
"6.1.4_impact_assessment": "missing",
"6.2_objectives": "documented",
"7.1_resources": "documented",
"7.2_competence": "missing",
"7.3_awareness": "partial",
"7.4_communication": "documented",
"7.5_documented_info": "documented",
"8.1_operational_planning": "documented",
"8.2_impact_assessment_process": "partial",
"8.3_ai_system_lifecycle": "missing",
"8.4_third_party_relationships": "partial",
"9.1_monitoring": "partial",
"9.2_internal_audit": "missing",
"9.3_management_review": "documented",
"10.1_continual_improvement": "partial",
"10.2_nonconformity_capa": "documented",
},
}
# Clause requirements + severity if missing
# severity: 'critical' = major nonconformity at stage 1, blocks certification
# 'major' = major nonconformity at stage 2
# 'minor' = minor nonconformity, requires corrective action plan
# 'observation' = improvement opportunity
CLAUSE_REQUIREMENTS: Dict[str, Dict[str, Any]] = {
"4.1_context_external": {"clause": "4.1", "title": "External & internal context", "severity": "minor"},
"4.2_interested_parties": {"clause": "4.2", "title": "Interested parties", "severity": "minor"},
"4.3_scope_statement": {"clause": "4.3", "title": "AIMS scope statement", "severity": "critical"},
"4.4_aims_processes": {"clause": "4.4", "title": "AIMS processes & interactions", "severity": "major"},
"5.1_leadership_commitment": {"clause": "5.1", "title": "Leadership commitment", "severity": "major"},
"5.2_ai_policy": {"clause": "5.2", "title": "AI policy", "severity": "critical"},
"5.3_roles_responsibilities": {"clause": "5.3", "title": "Roles, responsibilities, authorities", "severity": "critical"},
"6.1.2_risk_assessment": {"clause": "6.1.2", "title": "AI risk assessment", "severity": "critical"},
"6.1.3_risk_treatment": {"clause": "6.1.3", "title": "AI risk treatment", "severity": "critical"},
"6.1.4_impact_assessment": {"clause": "6.1.4", "title": "AI system impact assessment", "severity": "major"},
"6.2_objectives": {"clause": "6.2", "title": "AI objectives & planning", "severity": "minor"},
"7.1_resources": {"clause": "7.1", "title": "Resources", "severity": "minor"},
"7.2_competence": {"clause": "7.2", "title": "Competence", "severity": "major"},
"7.3_awareness": {"clause": "7.3", "title": "Awareness", "severity": "minor"},
"7.4_communication": {"clause": "7.4", "title": "Communication", "severity": "minor"},
"7.5_documented_info": {"clause": "7.5", "title": "Documented information", "severity": "major"},
"8.1_operational_planning": {"clause": "8.1", "title": "Operational planning & control", "severity": "major"},
"8.2_impact_assessment_process": {"clause": "8.2", "title": "Impact assessment process", "severity": "major"},
"8.3_ai_system_lifecycle": {"clause": "8.3", "title": "AI system lifecycle process", "severity": "critical"},
"8.4_third_party_relationships": {"clause": "8.4", "title": "Third-party / customer relationships", "severity": "major"},
"9.1_monitoring": {"clause": "9.1", "title": "Monitoring, measurement, analysis, evaluation", "severity": "major"},
"9.2_internal_audit": {"clause": "9.2", "title": "Internal audit programme", "severity": "critical"},
"9.3_management_review": {"clause": "9.3", "title": "Management review", "severity": "critical"},
"10.1_continual_improvement": {"clause": "10.1", "title": "Continual improvement", "severity": "minor"},
"10.2_nonconformity_capa": {"clause": "10.2", "title": "Nonconformity & corrective action", "severity": "major"},
}
STATUS_SCORE = {"documented": 1.0, "partial": 0.5, "missing": 0.0}
SEVERITY_RANK = {"critical": 0, "major": 1, "minor": 2, "observation": 3}
def remediation_action(req_key: str, status: str) -> str:
"""Deterministic one-sentence next step per (clause, status)."""
if status == "documented":
return "Maintain via management review; re-verify at next internal audit."
titles = CLAUSE_REQUIREMENTS[req_key]["title"]
if status == "partial":
return f"Complete documentation of '{titles}' — confirm signoff, version control, evidence trail."
return f"Create from scratch: '{titles}'. Assign owner; target close before stage 1 audit."
def analyze(payload: Dict[str, Any]) -> Dict[str, Any]:
evidence = payload.get("evidence", {})
findings: List[Dict[str, Any]] = []
total_weight = 0.0
achieved_weight = 0.0
for req_key, meta in CLAUSE_REQUIREMENTS.items():
status = evidence.get(req_key, "missing")
score = STATUS_SCORE.get(status, 0.0)
# Severity-weighted: critical = 4, major = 2, minor = 1
weight = {"critical": 4, "major": 2, "minor": 1, "observation": 1}[meta["severity"]]
total_weight += weight
achieved_weight += weight * score
findings.append({
"clause": meta["clause"],
"title": meta["title"],
"status": status,
"severity_if_missing": meta["severity"],
"remediation": remediation_action(req_key, status),
})
coverage_pct = round((achieved_weight / total_weight) * 100, 1) if total_weight else 0
# Sort findings: missing/partial first by severity, then documented last
def sort_key(f: Dict[str, Any]) -> tuple:
status_order = {"missing": 0, "partial": 1, "documented": 2}
return (status_order[f["status"]], SEVERITY_RANK[f["severity_if_missing"]], f["clause"])
findings.sort(key=sort_key)
open_gaps = [f for f in findings if f["status"] != "documented"]
critical_gaps = [f for f in open_gaps if f["severity_if_missing"] == "critical"]
major_gaps = [f for f in open_gaps if f["severity_if_missing"] == "major"]
readiness = "ready" if not critical_gaps and len(major_gaps) <= 1 else (
"stage_2_candidate" if not critical_gaps else "not_ready"
)
return {
"organization": payload.get("organization"),
"scope": payload.get("scope_statement"),
"coverage_pct_weighted": coverage_pct,
"certification_readiness": readiness,
"critical_gap_count": len(critical_gaps),
"major_gap_count": len(major_gaps),
"open_gap_count": len(open_gaps),
"findings": findings,
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("ISO/IEC 42001 AIMS — GAP ANALYSIS")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Organization: {r['organization']}")
lines.append(f"Scope: {r['scope']}")
lines.append(f"Weighted coverage: {r['coverage_pct_weighted']}%")
lines.append(f"Certification readiness: {r['certification_readiness']}")
lines.append(f"Critical gaps: {r['critical_gap_count']} | Major gaps: {r['major_gap_count']} | Open total: {r['open_gap_count']}")
lines.append("")
lines.append("-" * 72)
lines.append("FINDINGS (open gaps first; critical highlighted):")
lines.append("")
for f in r["findings"]:
marker = {"missing": "[X] ", "partial": "[~] ", "documented": "[✓] "}[f["status"]]
sev = f["severity_if_missing"].upper() if f["status"] != "documented" else "OK"
lines.append(f" {marker}Clause {f['clause']:6s} {f['title']:50s} [{sev}]")
if f["status"] != "documented":
lines.append(f" → {f['remediation']}")
lines.append("")
lines.append("-" * 72)
lines.append("READINESS RULE: 'ready' = 0 critical AND ≤ 1 major. 'stage_2_candidate' = 0 critical.")
lines.append(" Any critical gap blocks stage 1 certification.")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="ISO/IEC 42001 AIMS gap analysis across Clauses 4-10.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to AIMS evidence JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: mid-stage AI SaaS, pre stage-1 audit>"
result = analyze(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/ai_risk_register_builder.py
#!/usr/bin/env python3
"""ai_risk_register_builder.py — ISO/IEC 42001 Annex A risk register + control mapping.
Stdlib-only. Takes identified AI risks (per ISO 23894 risk identification) and produces a
structured register with:
- severity rating (likelihood × impact, 5x5 matrix)
- mapped Annex A controls (treatment selection)
- residual risk verdict (accept / additional treatment required / escalate)
- treatment option per ISO 23894 (modify / share / retain / avoid)
Deterministic logic per ISO 23894:2023 risk-management process. No LLM calls.
Input schema (JSON):
{
"organization": "Acme AI Inc.",
"ai_system": "Customer recommendation engine v3",
"risks": [
{
"id": "R-001",
"source": "training_data",
"event": "Biased dataset over-represents one demographic",
"consequence": "Discriminatory recommendations; regulatory exposure",
"likelihood": 3, # 1-5
"impact": 4, # 1-5
"controls_applied": ["A.7.3", "A.7.5", "A.5.2"]
}
]
}
Usage:
python ai_risk_register_builder.py # uses embedded 7-risk sample
python ai_risk_register_builder.py path/to/risks.json
python ai_risk_register_builder.py risks.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"organization": "Acme AI Inc.",
"ai_system": "Customer recommendation engine v3",
"risks": [
{"id": "R-001", "source": "training_data", "event": "Biased dataset over-represents one demographic",
"consequence": "Discriminatory recommendations; regulatory exposure", "likelihood": 3, "impact": 4,
"controls_applied": ["A.7.3", "A.7.5", "A.5.2"]},
{"id": "R-002", "source": "model", "event": "Concept drift after 6 months in production",
"consequence": "Accuracy degradation; revenue impact", "likelihood": 4, "impact": 3,
"controls_applied": ["A.9.3", "A.6.2.4"]},
{"id": "R-003", "source": "deployment", "event": "Inference latency spike under load",
"consequence": "User-visible failure; SLO breach", "likelihood": 3, "impact": 2,
"controls_applied": ["A.9.3"]},
{"id": "R-004", "source": "third_party", "event": "Foundation-model API provider deprecates endpoint",
"consequence": "Service disruption; migration cost", "likelihood": 2, "impact": 4,
"controls_applied": ["A.10.2"]},
{"id": "R-005", "source": "data", "event": "Training data contains PII that should not be retained",
"consequence": "GDPR fine; trust loss", "likelihood": 2, "impact": 5,
"controls_applied": ["A.7.2", "A.7.4"]},
{"id": "R-006", "source": "human_oversight", "event": "High-impact decisions deployed without impact assessment",
"consequence": "Untracked harm; certification nonconformity", "likelihood": 3, "impact": 5,
"controls_applied": []},
{"id": "R-007", "source": "model", "event": "Adversarial prompt injection bypasses content filter",
"consequence": "Toxic output to end users; reputational damage", "likelihood": 4, "impact": 4,
"controls_applied": ["A.6.2.4", "A.9.3", "A.9.4"]},
],
}
# Severity matrix (5x5): likelihood (1-5) × impact (1-5)
# Score 1-4 = low, 5-9 = medium, 10-16 = high, 17-25 = critical
def severity_rating(likelihood: int, impact: int) -> str:
score = max(1, min(5, likelihood)) * max(1, min(5, impact))
if score <= 4:
return "low"
if score <= 9:
return "medium"
if score <= 16:
return "high"
return "critical"
# ISO 23894 risk treatment options
# - modify (apply controls to reduce likelihood/impact)
# - share (transfer via insurance, third-party contracts)
# - retain (accept residual risk with management signoff)
# - avoid (eliminate the activity entirely)
def treatment_option(severity: str, controls_count: int) -> str:
if severity == "critical" and controls_count == 0:
return "avoid_or_escalate"
if severity in ("high", "critical"):
return "modify"
if severity == "medium":
return "modify" if controls_count < 2 else "retain"
return "retain"
# Residual-risk verdict after applied controls
def residual_verdict(severity: str, controls_count: int) -> str:
"""How many controls are 'enough' for each severity tier (heuristic, ISO 23894 Annex A guidance)."""
expected = {"low": 0, "medium": 1, "high": 2, "critical": 3}[severity]
if controls_count >= expected:
return "acceptable" if severity != "critical" else "acceptable_with_management_signoff"
return "additional_treatment_required"
# Annex A control descriptions (subset, for output annotation)
ANNEX_A_CATALOG: Dict[str, str] = {
"A.2.2": "AI policy",
"A.2.3": "Alignment of AI policy with other organizational policies",
"A.3.2": "AI roles & responsibilities",
"A.3.3": "Reporting of concerns",
"A.4.2": "Resources for AI systems — data",
"A.4.3": "Resources for AI systems — tooling",
"A.4.4": "Resources for AI systems — human resources",
"A.5.2": "AI system impact assessment",
"A.5.4": "Documentation of impact assessment",
"A.6.2.2": "AI system objectives",
"A.6.2.3": "AI system lifecycle phases",
"A.6.2.4": "Verification & validation of AI system",
"A.7.2": "Data management for AI systems",
"A.7.3": "Data quality",
"A.7.4": "Data provenance",
"A.7.5": "Data preparation",
"A.8.2": "System documentation for users",
"A.8.3": "User information",
"A.8.4": "Communication of AI incidents",
"A.9.2": "Intended use of AI system",
"A.9.3": "Monitoring of AI system operation",
"A.9.4": "Logging of AI system events",
"A.10.2": "Supplier (third-party) relationships",
"A.10.3": "Customer relationships",
}
def annotate_risk(risk: Dict[str, Any]) -> Dict[str, Any]:
likelihood = int(risk.get("likelihood", 0))
impact = int(risk.get("impact", 0))
controls = list(risk.get("controls_applied", []))
sev = severity_rating(likelihood, impact)
treatment = treatment_option(sev, len(controls))
residual = residual_verdict(sev, len(controls))
return {
"id": risk.get("id"),
"source": risk.get("source"),
"event": risk.get("event"),
"consequence": risk.get("consequence"),
"likelihood": likelihood,
"impact": impact,
"severity_score": likelihood * impact,
"severity": sev,
"controls_applied": [{"id": c, "title": ANNEX_A_CATALOG.get(c, "<unknown control>")} for c in controls],
"control_count": len(controls),
"treatment_option": treatment,
"residual_verdict": residual,
}
def analyze(payload: Dict[str, Any]) -> Dict[str, Any]:
risks = [annotate_risk(r) for r in payload.get("risks", [])]
# Sort by severity (critical first), then by control gap (largest first)
sev_rank = {"critical": 0, "high": 1, "medium": 2, "low": 3}
risks.sort(key=lambda r: (sev_rank[r["severity"]], -r["severity_score"]))
counts_by_sev = {s: 0 for s in sev_rank}
requires_action = 0
for r in risks:
counts_by_sev[r["severity"]] += 1
if r["residual_verdict"] == "additional_treatment_required":
requires_action += 1
return {
"organization": payload.get("organization"),
"ai_system": payload.get("ai_system"),
"total_risks": len(risks),
"by_severity": counts_by_sev,
"requires_additional_treatment": requires_action,
"risks": risks,
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("AI RISK REGISTER — ISO/IEC 42001 Annex A + ISO 23894 treatment")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Organization: {r['organization']}")
lines.append(f"AI system: {r['ai_system']}")
lines.append(f"Total risks: {r['total_risks']}")
s = r["by_severity"]
lines.append(f"By severity: critical={s['critical']} high={s['high']} medium={s['medium']} low={s['low']}")
lines.append(f"Risks requiring additional treatment: {r['requires_additional_treatment']}")
lines.append("")
lines.append("-" * 72)
lines.append("REGISTER (highest severity first):")
lines.append("")
for risk in r["risks"]:
lines.append(f" [{risk['id']}] {risk['event']}")
lines.append(f" Source: {risk['source']} | L={risk['likelihood']} × I={risk['impact']} = {risk['severity_score']} → {risk['severity'].upper()}")
lines.append(f" Consequence: {risk['consequence']}")
if risk["controls_applied"]:
ctrl_str = ", ".join(c["id"] for c in risk["controls_applied"])
lines.append(f" Controls applied ({risk['control_count']}): {ctrl_str}")
else:
lines.append(f" Controls applied: NONE")
lines.append(f" Treatment option: {risk['treatment_option']}")
lines.append(f" Residual verdict: {risk['residual_verdict']}")
lines.append("")
lines.append("-" * 72)
lines.append("RULES:")
lines.append(" - 'critical' severity (score 17-25) WITHOUT controls → 'avoid_or_escalate' to management.")
lines.append(" - 'additional_treatment_required' → add Annex A controls or formally accept residual risk in writing.")
lines.append(" - All 'retain' verdicts require Clause 6.1.3 risk-treatment plan signoff.")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="ISO/IEC 42001 Annex A risk register builder with ISO 23894 treatment options.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to risks JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: 7-risk recommendation engine register>"
result = analyze(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
Thiết lập và quản lý dự án Jira: lập kế hoạch, JQL, workflow, trường tùy chỉnh, tự động hóa, dashboard và báo cáo.
---
name: "jira-expert"
description: Atlassian Jira expert for creating and managing projects, planning, product discovery, JQL queries, workflows, custom fields, automation, reporting, and all Jira features. Use for Jira project setup, configuration, advanced search, dashboard creation, workflow design, and technical Jira operations.
---
# Atlassian Jira Expert
Master-level expertise in Jira configuration, project management, JQL, workflows, automation, and reporting. Handles all technical and operational aspects of Jira.
## Quick Start — Most Common Operations
**Create a project**:
```
mcp jira create_project --name "My Project" --key "MYPROJ" --type scrum --lead "user@example.com"
```
**Run a JQL query**:
```
mcp jira search_issues --jql "project = MYPROJ AND status != Done AND dueDate < now()" --maxResults 50
```
For full command reference, see [Atlassian MCP Integration](#atlassian-mcp-integration). For JQL functions, see [JQL Functions Reference](#jql-functions-reference). For report templates, see [Reporting Templates](#reporting-templates).
---
## Workflows
### Project Creation
1. Determine project type (Scrum, Kanban, Bug Tracking, etc.)
2. Create project with appropriate template
3. Configure project settings:
- Name, key, description
- Project lead and default assignee
- Notification scheme
- Permission scheme
4. Set up issue types and workflows
5. Configure custom fields if needed
6. Create initial board/backlog view
7. **HANDOFF TO**: Scrum Master for team onboarding
### Workflow Design
1. Map out process states (To Do → In Progress → Done)
2. Define transitions and conditions
3. Add validators, post-functions, and conditions
4. Configure workflow scheme
5. **Validate**: Deploy to a test project first; verify all transitions, conditions, and post-functions behave as expected before associating with production projects
6. Associate workflow with project
7. Test workflow with sample issues
### JQL Query Building
**Basic Structure**: `field operator value`
**Common Operators**:
- `=, !=` : equals, not equals
- `~, !~` : contains, not contains
- `>, <, >=, <=` : comparison
- `in, not in` : list membership
- `is empty, is not empty`
- `was, was in, was not`
- `changed`
**Powerful JQL Examples**:
Find overdue issues:
```jql
dueDate < now() AND status != Done
```
Sprint burndown issues:
```jql
sprint = 23 AND status changed TO "Done" DURING (startOfSprint(), endOfSprint())
```
Find stale issues:
```jql
updated < -30d AND status != Done
```
Cross-project epic tracking:
```jql
"Epic Link" = PROJ-123 ORDER BY rank
```
Velocity calculation:
```jql
sprint in closedSprints() AND resolution = Done
```
Team capacity:
```jql
assignee in (user1, user2) AND sprint in openSprints()
```
### Dashboard Creation
1. Create new dashboard (personal or shared)
2. Add relevant gadgets:
- Filter Results (JQL-based)
- Sprint Burndown
- Velocity Chart
- Created vs Resolved
- Pie Chart (status distribution)
3. Arrange layout for readability
4. Configure automatic refresh
5. Share with appropriate teams
6. **HANDOFF TO**: Senior PM or Scrum Master for use
### Automation Rules
1. Define trigger (issue created, field changed, scheduled)
2. Add conditions (if applicable)
3. Define actions:
- Update field
- Send notification
- Create subtask
- Transition issue
- Post comment
4. Test automation with sample data
5. Enable and monitor
## Advanced Features
### Custom Fields
**When to Create**:
- Track data not in standard fields
- Capture process-specific information
- Enable advanced reporting
**Field Types**: Text, Numeric, Date, Select (single/multi/cascading), User picker
**Configuration**:
1. Create custom field
2. Configure field context (which projects/issue types)
3. Add to appropriate screens
4. Update search templates if needed
### Issue Linking
**Link Types**:
- Blocks / Is blocked by
- Relates to
- Duplicates / Is duplicated by
- Clones / Is cloned by
- Epic-Story relationship
**Best Practices**:
- Use Epic linking for feature grouping
- Use blocking links to show dependencies
- Document link reasons in comments
### Permissions & Security
**Permission Schemes**:
- Browse Projects
- Create/Edit/Delete Issues
- Administer Projects
- Manage Sprints
**Security Levels**:
- Define confidential issue visibility
- Control access to sensitive data
- Audit security changes
### Bulk Operations
**Bulk Change**:
1. Use JQL to find target issues
2. Select bulk change operation
3. Choose fields to update
4. **Validate**: Preview all changes before executing; confirm the JQL filter matches only intended issues — bulk edits are difficult to reverse
5. Execute and confirm
6. Monitor background task
**Bulk Transitions**:
- Move multiple issues through workflow
- Useful for sprint cleanup
- Requires appropriate permissions
- **Validate**: Run the JQL filter and review results in small batches before applying at scale
## JQL Functions Reference
> **Tip**: Save frequently used queries as named filters instead of re-running complex JQL ad hoc. See [Best Practices](#best-practices) for performance guidance.
**Date**: `startOfDay()`, `endOfDay()`, `startOfWeek()`, `endOfWeek()`, `startOfMonth()`, `endOfMonth()`, `startOfYear()`, `endOfYear()`
**Sprint**: `openSprints()`, `closedSprints()`, `futureSprints()`
**User**: `currentUser()`, `membersOf("group")`
**Advanced**: `issueHistory()`, `linkedIssues()`, `issuesWithFixVersions()`
## Reporting Templates
> **Tip**: These JQL snippets can be saved as shared filters or wired directly into Dashboard gadgets (see [Dashboard Creation](#dashboard-creation)).
| Report | JQL |
|---|---|
| Sprint Report | `project = PROJ AND sprint = 23` |
| Team Velocity | `assignee in (team) AND sprint in closedSprints() AND resolution = Done` |
| Bug Trend | `type = Bug AND created >= -30d` |
| Blocker Analysis | `priority = Blocker AND status != Done` |
## Decision Framework
**When to Escalate to Atlassian Admin**:
- Need new project permission scheme
- Require custom workflow scheme across org
- User provisioning or deprovisioning
- License or billing questions
- System-wide configuration changes
**When to Collaborate with Scrum Master**:
- Sprint board configuration
- Backlog prioritization views
- Team-specific filters
- Sprint reporting needs
**When to Collaborate with Senior PM**:
- Portfolio-level reporting
- Cross-project dashboards
- Executive visibility needs
- Multi-project dependencies
## Handoff Protocols
**FROM Senior PM**:
- Project structure requirements
- Workflow and field needs
- Reporting requirements
- Integration needs
**TO Senior PM**:
- Cross-project metrics
- Issue trends and patterns
- Workflow bottlenecks
- Data quality insights
**FROM Scrum Master**:
- Sprint board configuration requests
- Workflow optimization needs
- Backlog filtering requirements
- Velocity tracking setup
**TO Scrum Master**:
- Configured sprint boards
- Velocity reports
- Burndown charts
- Team capacity views
## Best Practices
**Data Quality**:
- Enforce required fields with field validation rules
- Use consistent issue key naming conventions per project type
- Schedule regular cleanup of stale/orphaned issues
**Performance**:
- Avoid leading wildcards in JQL (`~` on large text fields is expensive)
- Use saved filters instead of re-running complex JQL ad hoc
- Limit dashboard gadgets to reduce page load time
- Archive completed projects rather than deleting to preserve history
**Governance**:
- Document rationale for custom workflow states and transitions
- Version-control permission/workflow schemes before making changes
- Require change management review for org-wide scheme updates
- Run permission audits after user role changes
## Atlassian MCP Integration
**Primary Tool**: Jira MCP Server
**Key Operations with Example Commands**:
Create a project:
```
mcp jira create_project --name "My Project" --key "MYPROJ" --type scrum --lead "user@example.com"
```
Execute a JQL query:
```
mcp jira search_issues --jql "project = MYPROJ AND status != Done AND dueDate < now()" --maxResults 50
```
Update an issue field:
```
mcp jira update_issue --issue "MYPROJ-42" --field "status" --value "In Progress"
```
Create a sprint:
```
mcp jira create_sprint --board 10 --name "Sprint 5" --startDate "2024-06-01" --endDate "2024-06-14"
```
Create a board filter:
```
mcp jira create_filter --name "Open Blockers" --jql "priority = Blocker AND status != Done" --shareWith "project-team"
```
**Integration Points**:
- Pull metrics for Senior PM reporting
- Configure sprint boards for Scrum Master
- Create documentation pages for Confluence Expert
- Support template creation for Template Creator
## Related Skills
- **Confluence Expert** (`project-management/confluence-expert/`) — Documentation complements Jira workflows
- **Atlassian Admin** (`project-management/atlassian-admin/`) — Permission and user management for Jira projects
FILE:references/automation-examples.md
# Jira Automation Examples
## Auto-Assignment Rules
### Auto-assign by component
**Trigger:** Issue created
**Conditions:**
- Component is not EMPTY
**Actions:**
- Assign issue to component lead
### Auto-assign to reporter for feedback
**Trigger:** Issue transitioned to "Waiting for Feedback"
**Actions:**
- Assign issue to reporter
- Add comment: "Please provide additional information"
### Round-robin assignment
**Trigger:** Issue created
**Conditions:**
- Project = ABC
- Assignee is EMPTY
**Actions:**
- Assign to next team member in rotation (use smart value)
---
## Status Sync Rules
### Sync subtask status to parent
**Trigger:** Issue transitioned
**Conditions:**
- Issue type = Sub-task
- Transition is to "Done"
- Parent issue exists
- All subtasks are Done
**Actions:**
- Transition parent issue to "Done"
### Sync parent to subtasks
**Trigger:** Issue transitioned
**Conditions:**
- Issue type has subtasks
- Transition is to "Cancelled"
**Actions:**
- For each: Sub-tasks
- Transition issue to "Cancelled"
### Epic progress tracking
**Trigger:** Issue transitioned
**Conditions:**
- Epic link is not EMPTY
- Transition is to "Done"
**Actions:**
- Add comment to epic: "{{issue.key}} completed"
- Update epic custom field "Progress"
---
## Notification Rules
### Slack notification for high-priority bugs
**Trigger:** Issue created
**Conditions:**
- Issue type = Bug
- Priority IN (Highest, High)
**Actions:**
- Send Slack message to #engineering:
```
🚨 High Priority Bug Created
{{issue.key}}: {{issue.summary}}
Reporter: {{issue.reporter.displayName}}
Priority: {{issue.priority.name}}
{{issue.url}}
```
### Email assignee when mentioned
**Trigger:** Issue commented
**Conditions:**
- Comment contains @mention of assignee
**Actions:**
- Send email to {{issue.assignee.emailAddress}}:
```
Subject: You were mentioned in {{issue.key}}
Body: {{comment.author.displayName}} mentioned you:
{{comment.body}}
```
### SLA breach warning
**Trigger:** Scheduled - Every hour
**Conditions:**
- Status != Done
- SLA time remaining < 2 hours
**Actions:**
- Send email to {{issue.assignee}}
- Add comment: "⚠️ SLA expires in <2 hours"
- Set priority to Highest
---
## Field Automation Rules
### Auto-set due date
**Trigger:** Issue created
**Conditions:**
- Issue type = Bug
- Priority = Highest
**Actions:**
- Set due date to {{now.plusDays(1)}}
### Clear assignee when in backlog
**Trigger:** Issue transitioned
**Conditions:**
- Transition is to "Backlog"
- Assignee is not EMPTY
**Actions:**
- Assign issue to Unassigned
- Add comment: "Returned to backlog, assignee cleared"
### Auto-populate sprint field
**Trigger:** Issue transitioned
**Conditions:**
- Transition is to "In Progress"
- Sprint is EMPTY
**Actions:**
- Add issue to current sprint
### Set fix version based on component
**Trigger:** Issue created
**Conditions:**
- Component = "Mobile App"
**Actions:**
- Set fix version to "Mobile v2.0"
---
## Escalation Rules
### Auto-escalate stale issues
**Trigger:** Scheduled - Daily at 9:00 AM
**Conditions:**
- Status = "Waiting for Response"
- Updated < -7 days
**Actions:**
- Add comment: "@{{issue.reporter}} This issue needs attention"
- Send email to project lead
- Add label: "needs-attention"
### Escalate overdue critical issues
**Trigger:** Scheduled - Every hour
**Conditions:**
- Priority IN (Highest, High)
- Due date < now()
- Status != Done
**Actions:**
- Transition to "Escalated"
- Assign to project manager
- Send Slack notification
### Auto-close inactive issues
**Trigger:** Scheduled - Daily at 10:00 AM
**Conditions:**
- Status = "Waiting for Customer"
- Updated < -30 days
**Actions:**
- Transition to "Closed"
- Add comment: "Auto-closed due to inactivity"
- Send email to reporter
---
## Sprint Automation Rules
### Move incomplete work to next sprint
**Trigger:** Sprint closed
**Conditions:**
- Issue status != Done
**Actions:**
- Add issue to next sprint
- Add comment: "Moved from {{sprint.name}}"
### Auto-remove completed items from active sprint
**Trigger:** Issue transitioned
**Conditions:**
- Transition is to "Done"
- Sprint IN openSprints()
**Actions:**
- Remove issue from sprint
- Add comment: "Removed from active sprint (completed)"
### Sprint start notification
**Trigger:** Sprint started
**Actions:**
- Send Slack message to #team:
```
🚀 Sprint {{sprint.name}} Started!
Goal: {{sprint.goal}}
Committed: {{sprint.issuesCount}} issues
```
---
## Approval Workflow Rules
### Request approval for large stories
**Trigger:** Issue created
**Conditions:**
- Issue type = Story
- Story points >= 13
**Actions:**
- Transition to "Pending Approval"
- Assign to product owner
- Send email notification
### Auto-approve small bugs
**Trigger:** Issue created
**Conditions:**
- Issue type = Bug
- Priority IN (Low, Lowest)
**Actions:**
- Transition to "Approved"
- Add comment: "Auto-approved (low-priority bug)"
### Require security review
**Trigger:** Issue transitioned
**Conditions:**
- Transition is to "Ready for Release"
- Labels contains "security"
**Actions:**
- Transition to "Security Review"
- Assign to security-team
- Send email to security@company.com
---
## Integration Rules
### Create GitHub issue
**Trigger:** Issue transitioned
**Conditions:**
- Transition is to "In Progress"
- Labels contains "needs-tracking"
**Actions:**
- Send webhook to GitHub API:
```json
{
"title": "{{issue.key}}: {{issue.summary}}",
"body": "{{issue.description}}",
"assignee": "{{issue.assignee.name}}"
}
```
### Update Confluence page
**Trigger:** Issue transitioned
**Conditions:**
- Issue type = Epic
- Transition is to "Done"
**Actions:**
- Send webhook to Confluence:
- Update epic status page
- Add completion date
---
## Quality & Testing Rules
### Require test cases for features
**Trigger:** Issue transitioned
**Conditions:**
- Issue type = Story
- Transition is to "Ready for QA"
- Custom field "Test Cases" is EMPTY
**Actions:**
- Transition back to "In Progress"
- Add comment: "❌ Test cases required before QA"
### Auto-create test issue
**Trigger:** Issue transitioned
**Conditions:**
- Issue type = Story
- Transition is to "Ready for QA"
**Actions:**
- Create linked issue:
- Type: Test
- Summary: "Test: {{issue.summary}}"
- Link type: "tested by"
- Assignee: QA team
### Flag regression bugs
**Trigger:** Issue created
**Conditions:**
- Issue type = Bug
- Affects version is in released versions
**Actions:**
- Add label: "regression"
- Set priority to High
- Add comment: "🚨 Regression in released version"
---
## Documentation Rules
### Require documentation for features
**Trigger:** Issue transitioned
**Conditions:**
- Issue type = Story
- Labels contains "customer-facing"
- Transition is to "Done"
- Custom field "Documentation Link" is EMPTY
**Actions:**
- Reopen issue
- Add comment: "📝 Documentation required for customer-facing feature"
### Auto-create doc task
**Trigger:** Issue transitioned
**Conditions:**
- Issue type = Epic
- Transition is to "In Progress"
**Actions:**
- Create subtask:
- Type: Task
- Summary: "Documentation for {{issue.summary}}"
- Assignee: {{issue.assignee}}
---
## Time Tracking Rules
### Log work reminder
**Trigger:** Issue transitioned
**Conditions:**
- Transition is to "Done"
- Time spent is EMPTY
**Actions:**
- Add comment: "⏱️ Reminder: Please log your time"
### Warn on high time spent
**Trigger:** Work logged
**Conditions:**
- Time spent > original estimate * 1.5
**Actions:**
- Add comment: "⚠️ Time spent exceeds estimate by 50%"
- Send notification to assignee and project manager
---
## Advanced Conditional Rules
### Conditional assignee based on priority
**Trigger:** Issue created
**Conditions:**
- Issue type = Bug
**Actions:**
- If: Priority = Highest
- Assign to on-call engineer
- Else if: Priority = High
- Assign to team lead
- Else:
- Assign to next available team member
### Multi-step approval flow
**Trigger:** Issue transitioned
**Conditions:**
- Transition is to "Request Approval"
- Budget estimate > $10,000
**Actions:**
- If: Budget > $50,000
- Assign to CFO
- Send email to executive team
- Else if: Budget > $10,000
- Assign to Director
- Add comment: "Director approval required"
- Add label: "pending-approval"
---
## Smart Value Examples
### Dynamic assignee based on component
```
{{issue.components.first.lead.accountId}}
```
### Days since created
```
{{issue.created.diff(now).days}}
```
### Conditional message
```
{{#if(issue.priority.name == "Highest")}}
🚨 CRITICAL
{{else}}
ℹ️ Normal priority
{{/}}
```
### List all subtasks
```
{{#issue.subtasks}}
- {{key}}: {{summary}} ({{status.name}})
{{/}}
```
### Calculate completion percentage
```
{{issue.subtasks.filter(item => item.status.statusCategory.key == "done").size.divide(issue.subtasks.size).multiply(100).round()}}%
```
---
## Best Practices
1. **Test in sandbox** - Always test rules on test project first
2. **Start simple** - Begin with basic rules, add complexity incrementally
3. **Use conditions wisely** - Narrow scope to reduce unintended triggers
4. **Monitor audit log** - Check automation execution history regularly
5. **Limit actions** - Keep rules focused, don't chain too many actions
6. **Name clearly** - Use descriptive names: "Auto-assign bugs to component lead"
7. **Document rules** - Add description explaining purpose and owner
8. **Review regularly** - Audit rules quarterly, disable unused ones
9. **Handle errors** - Add error handling for webhooks and integrations
10. **Performance** - Avoid scheduled rules that query large datasets hourly
FILE:references/AUTOMATION.md
# Jira Automation Reference
Comprehensive guide to Jira automation rules: triggers, conditions, actions, smart values, and production-ready recipes.
## Rule Structure
Every automation rule follows this pattern:
```
TRIGGER → [CONDITION(s)] → ACTION(s)
```
- **Trigger**: The event that starts the rule (required, exactly one)
- **Condition**: Filters to narrow when the rule fires (optional, multiple allowed)
- **Action**: What the rule does (required, one or more)
## Triggers
### Issue Triggers
| Trigger | Fires When | Use For |
|---------|------------|---------|
| **Issue created** | New issue is created | Auto-assignment, notifications, SLA start |
| **Issue transitioned** | Status changes | Workflow automation, notifications |
| **Issue updated** | Any field changes | Field sync, cascading updates |
| **Issue commented** | Comment is added | Auto-responses, SLA tracking |
| **Issue assigned** | Assignee changes | Workload notifications |
| **Issue linked** | Link is added/removed | Dependency tracking |
| **Issue deleted** | Issue is deleted | Cleanup, audit logging |
### Sprint & Board Triggers
| Trigger | Fires When |
|---------|------------|
| **Sprint started** | Sprint is activated |
| **Sprint completed** | Sprint is closed |
| **Issue moved between sprints** | Issue is moved |
| **Backlog item moved to sprint** | Item is pulled into sprint |
### Scheduled Triggers
| Trigger | Fires When |
|---------|------------|
| **Scheduled** | Cron-based (daily, weekly, custom) |
| **Issue stale** | No updates for X days |
### Version Triggers
| Trigger | Fires When |
|---------|------------|
| **Version created** | New version added |
| **Version released** | Version is released |
## Conditions
### Issue Conditions
| Condition | Matches When |
|-----------|-------------|
| **Issue fields condition** | Field matches value (e.g., priority = High) |
| **JQL condition** | Issue matches JQL query |
| **Related issues condition** | Linked/sub-task issues match criteria |
| **User condition** | Actor matches (reporter, assignee, group) |
| **Advanced compare** | Complex field comparisons |
### Condition Operators
```
Field = value # Exact match
Field != value # Not equal
Field > value # Greater than (numeric/date)
Field is empty # Field has no value
Field is not empty # Field has a value
Field changed # Field was modified in this event
Field changed to # Field changed to specific value
Field changed from # Field changed from specific value
```
## Actions
### Issue Actions
| Action | Does |
|--------|------|
| **Edit issue** | Update any field on the current issue |
| **Transition issue** | Move to a new status |
| **Assign issue** | Change assignee |
| **Comment on issue** | Add a comment |
| **Create issue** | Create a new linked issue |
| **Create sub-tasks** | Create child issues |
| **Clone issue** | Duplicate the issue |
| **Delete issue** | Remove the issue |
| **Link issues** | Add issue links |
| **Log work** | Add time tracking entry |
### Notification Actions
| Action | Does |
|--------|------|
| **Send email** | Send custom email to users/groups |
| **Send Slack message** | Post to Slack channel (requires integration) |
| **Send Microsoft Teams message** | Post to Teams (requires integration) |
| **Send web request** | HTTP call to external service |
### Lookup & Branch Actions
| Action | Does |
|--------|------|
| **Lookup issues (JQL)** | Find issues matching JQL, iterate over them |
| **Create branch** | Branch logic (if/then/else) |
| **For each** | Loop over found issues |
## Smart Values
Smart values are dynamic placeholders that resolve at runtime.
### Issue Smart Values
```
{{issue.key}} # PROJ-123
{{issue.summary}} # Issue title
{{issue.description}} # Full description
{{issue.status.name}} # Current status
{{issue.priority.name}} # Priority level
{{issue.assignee.displayName}} # Assignee name
{{issue.reporter.displayName}} # Reporter name
{{issue.issuetype.name}} # Issue type
{{issue.project.key}} # Project key
{{issue.created}} # Creation date
{{issue.updated}} # Last update date
{{issue.fixVersions}} # Fix versions
{{issue.labels}} # Labels array
{{issue.components}} # Components array
```
### Transition Smart Values
```
{{transition.from_status}} # Previous status
{{transition.to_status}} # New status
{{transition.transitionName}} # Transition name
```
### User Smart Values
```
{{initiator.displayName}} # Who triggered the rule
{{initiator.emailAddress}} # Their email
{{initiator.accountId}} # Their account ID
```
### Date Smart Values
```
{{now}} # Current timestamp
{{now.plusDays(7)}} # 7 days from now
{{now.minusHours(24)}} # 24 hours ago
{{issue.created.plusBusinessDays(3)}} # 3 business days after creation
```
### Conditional Smart Values
```
{{#if issue.priority.name == "High"}}
This is high priority
{{/if}}
{{#if issue.assignee}}
Assigned to {{issue.assignee.displayName}}
{{else}}
Unassigned
{{/if}}
```
## Production-Ready Recipes
### 1. Auto-Assign by Component
```yaml
Trigger: Issue created
Condition: Issue has component
Action: Edit issue
- Assignee = Component lead
Rule Logic:
IF component = "Backend" → assign to @backend-lead
IF component = "Frontend" → assign to @frontend-lead
IF component = "DevOps" → assign to @devops-lead
```
### 2. SLA Warning — Stale Issues
```yaml
Trigger: Scheduled (daily at 9am)
Condition: JQL = "status != Done AND updated <= -5d AND priority in (High, Highest)"
Action:
- Add comment: "⚠️ This {{issue.priority.name}} issue hasn't been updated in 5+ days."
- Send Slack: "#engineering-alerts: {{issue.key}} is stale ({{issue.assignee.displayName}})"
```
### 3. Auto-Close Resolved Issues After 7 Days
```yaml
Trigger: Scheduled (daily)
Condition: JQL = "status = Resolved AND updated <= -7d"
Action:
- Transition: Resolved → Closed
- Comment: "Auto-closed after 7 days in Resolved status."
```
### 4. Sprint Spillover Notification
```yaml
Trigger: Sprint completed
Condition: Issue status != Done
Action:
- Comment: "Spilled over from Sprint {{sprint.name}}. Reason needs review."
- Add label: "spillover"
- Send email to: {{issue.assignee.emailAddress}}
```
### 5. Sub-Task Completion → Parent Transition
```yaml
Trigger: Issue transitioned (to Done)
Condition: Issue is sub-task AND all sibling sub-tasks are Done
Action (on parent):
- Transition: In Progress → In Review
- Comment: "All sub-tasks completed. Ready for review."
```
### 6. Bug Priority Escalation
```yaml
Trigger: Scheduled (every 4 hours)
Condition: JQL = "type = Bug AND priority = High AND status = Open AND created <= -24h"
Action:
- Edit: priority = Highest
- Comment: "⚡ Auto-escalated: High-priority bug open for 24+ hours."
- Send email to: project lead
```
### 7. Auto-Link Duplicate Detection
```yaml
Trigger: Issue created
Condition: JQL finds issues with similar summary (fuzzy)
Action:
- Comment: "Possible duplicate of {{lookupIssues.first.key}}: {{lookupIssues.first.summary}}"
- Add label: "possible-duplicate"
```
### 8. Release Notes Generator
```yaml
Trigger: Version released
Condition: None
Action:
- Lookup: JQL = "fixVersion = {{version.name}} AND status = Done"
- Create Confluence page:
Title: "Release Notes — {{version.name}}"
Content: List of resolved issues with types and summaries
```
### 9. Workload Balancer — Round-Robin Assignment
```yaml
Trigger: Issue created
Condition: Issue type = Story AND assignee is empty
Action:
- Lookup: JQL = "assignee in (dev1, dev2, dev3) AND sprint in openSprints() AND status != Done"
- Assign to team member with fewest open issues
```
### 10. Blocker Notification Chain
```yaml
Trigger: Issue updated (priority changed to Blocker)
Action:
- Send email to: project lead, scrum master
- Send Slack: "#blockers: 🚨 {{issue.key}} marked as Blocker by {{initiator.displayName}}"
- Comment: "Blocker escalated. Notified: PM + SM."
- Edit: Add label "blocker-active"
```
## Best Practices
1. **Name rules descriptively** — "Auto-assign Backend bugs to @dev-lead" not "Rule 1"
2. **Add conditions before actions** — prevent unintended execution
3. **Use JQL conditions** for precision — field conditions can miss edge cases
4. **Test in a sandbox project first** — automation mistakes can be destructive
5. **Set rate limits** — avoid infinite loops (Rule A triggers Rule B triggers Rule A)
6. **Monitor rule execution** — check Automation audit log weekly
7. **Document business rules** — explain WHY the rule exists, not just WHAT it does
8. **Use branches (if/else)** over separate rules — reduces rule count, easier to maintain
9. **Disable before deleting** — observe for a week to ensure no side effects
10. **Version your automation** — export rules as JSON backup before major changes
FILE:references/jql-examples.md
# JQL Query Examples
## Sprint Queries
**Current sprint issues:**
```jql
sprint IN openSprints() ORDER BY rank
```
**Issues in specific sprint:**
```jql
sprint = "Sprint 23" ORDER BY priority DESC
```
**All sprint work (current and backlog):**
```jql
project = ABC AND issuetype IN (Story, Bug, Task)
ORDER BY sprint DESC, rank
```
**Unscheduled stories:**
```jql
project = ABC AND issuetype = Story AND sprint IS EMPTY
AND status != Done ORDER BY priority DESC
```
**Spillover from last sprint:**
```jql
sprint IN closedSprints() AND sprint NOT IN (latestReleasedVersion())
AND status != Done ORDER BY created DESC
```
**Sprint completion rate:**
```jql
sprint = "Sprint 23" AND status = Done
```
## User & Team Queries
**My open issues:**
```jql
assignee = currentUser() AND status != Done
ORDER BY priority DESC, created ASC
```
**Unassigned in my project:**
```jql
project = ABC AND assignee IS EMPTY AND status != Done
ORDER BY priority DESC
```
**Issues I'm watching:**
```jql
watcher = currentUser() AND status != Done
```
**Team workload:**
```jql
assignee IN membersOf("engineering-team") AND status IN ("In Progress", "In Review")
ORDER BY assignee, priority DESC
```
**Issues I reported that are still open:**
```jql
reporter = currentUser() AND status != Done ORDER BY created DESC
```
**Issues commented on by me:**
```jql
comment ~ currentUser() AND status != Done
```
## Date Range Queries
**Created today:**
```jql
created >= startOfDay() ORDER BY created DESC
```
**Updated in last 7 days:**
```jql
updated >= -7d ORDER BY updated DESC
```
**Created this week:**
```jql
created >= startOfWeek() AND created <= endOfWeek()
```
**Created this month:**
```jql
created >= startOfMonth() AND created <= endOfMonth()
```
**Not updated in 30 days:**
```jql
status != Done AND updated <= -30d ORDER BY updated ASC
```
**Resolved yesterday:**
```jql
resolved >= startOfDay(-1d) AND resolved < startOfDay()
```
**Due this week:**
```jql
duedate >= startOfWeek() AND duedate <= endOfWeek() AND status != Done
```
**Overdue:**
```jql
duedate < now() AND status != Done ORDER BY duedate ASC
```
## Status & Workflow Queries
**In Progress issues:**
```jql
project = ABC AND status = "In Progress" ORDER BY assignee
```
**Blocked issues:**
```jql
project = ABC AND labels = blocked AND status != Done
```
**Issues in review:**
```jql
project = ABC AND status IN ("Code Review", "QA Review", "Pending Approval")
ORDER BY updated ASC
```
**Ready for development:**
```jql
project = ABC AND status = "Ready" AND sprint IS EMPTY
ORDER BY priority DESC
```
**Recently done:**
```jql
project = ABC AND status = Done AND resolved >= -7d
ORDER BY resolved DESC
```
**Status changed today:**
```jql
status CHANGED AFTER startOfDay() ORDER BY updated DESC
```
**Long-running in progress:**
```jql
status = "In Progress" AND status CHANGED BEFORE -14d
ORDER BY status CHANGED ASC
```
## Priority & Type Queries
**High priority bugs:**
```jql
issuetype = Bug AND priority IN (Highest, High) AND status != Done
ORDER BY priority DESC, created ASC
```
**Critical blockers:**
```jql
priority = Highest AND status != Done ORDER BY created ASC
```
**All epics:**
```jql
issuetype = Epic ORDER BY status, priority DESC
```
**Stories without acceptance criteria:**
```jql
issuetype = Story AND "Acceptance Criteria" IS EMPTY AND status = Backlog
```
**Technical debt:**
```jql
labels = tech-debt AND status != Done ORDER BY priority DESC
```
## Complex Multi-Condition Queries
**My team's sprint work:**
```jql
sprint IN openSprints()
AND assignee IN membersOf("engineering-team")
AND status != Done
ORDER BY assignee, priority DESC
```
**Bugs created this month, not in sprint:**
```jql
issuetype = Bug
AND created >= startOfMonth()
AND sprint IS EMPTY
AND status != Done
ORDER BY priority DESC, created DESC
```
**High-priority work needing attention:**
```jql
project = ABC
AND priority IN (Highest, High)
AND status IN ("In Progress", "In Review")
AND updated <= -3d
ORDER BY priority DESC, updated ASC
```
**Stale issues:**
```jql
project = ABC
AND status NOT IN (Done, Cancelled)
AND (assignee IS EMPTY OR updated <= -30d)
ORDER BY created ASC
```
**Epic progress:**
```jql
"Epic Link" = ABC-123 ORDER BY status, rank
```
## Component & Version Queries
**Issues in component:**
```jql
project = ABC AND component = "Frontend" AND status != Done
```
**Issues without component:**
```jql
project = ABC AND component IS EMPTY AND status != Done
```
**Target version:**
```jql
fixVersion = "v2.0" ORDER BY status, priority DESC
```
**Released versions:**
```jql
fixVersion IN releasedVersions() ORDER BY fixVersion DESC
```
## Label & Text Search Queries
**Issues with label:**
```jql
labels = urgent AND status != Done
```
**Multiple labels (AND):**
```jql
labels IN (frontend, bug) AND status != Done
```
**Search in summary:**
```jql
summary ~ "authentication" ORDER BY created DESC
```
**Search in summary and description:**
```jql
text ~ "API integration" ORDER BY created DESC
```
**Issues with empty description:**
```jql
description IS EMPTY AND issuetype = Story
```
## Performance-Optimized Queries
**Good - Specific project first:**
```jql
project = ABC AND status = "In Progress" AND assignee = currentUser()
```
**Bad - User filter first:**
```jql
assignee = currentUser() AND status = "In Progress" AND project = ABC
```
**Good - Use functions:**
```jql
sprint IN openSprints() AND status != Done
```
**Bad - Hardcoded sprint:**
```jql
sprint = "Sprint 23" AND status != Done
```
**Good - Specific date:**
```jql
created >= 2024-01-01 AND created <= 2024-01-31
```
**Bad - Relative with high cost:**
```jql
created >= -365d AND created <= -335d
```
## Reporting Queries
**Velocity calculation:**
```jql
sprint = "Sprint 23" AND status = Done
```
*Then sum story points*
**Bug rate:**
```jql
project = ABC AND issuetype = Bug AND created >= startOfMonth()
```
**Average cycle time:**
```jql
project = ABC AND resolved >= startOfMonth()
AND resolved <= endOfMonth()
```
*Calculate time from In Progress to Done*
**Stories delivered this quarter:**
```jql
project = ABC AND issuetype = Story
AND resolved >= startOfYear() AND resolved <= endOfQuarter()
```
**Team capacity:**
```jql
assignee IN membersOf("engineering-team")
AND sprint IN openSprints()
```
*Sum original estimates*
## Notification & Watching Queries
**Issues I need to review:**
```jql
status = "Pending Review" AND assignee = currentUser()
```
**Issues assigned to me, high priority:**
```jql
assignee = currentUser() AND priority IN (Highest, High)
AND status != Done
```
**Issues created by me, not resolved:**
```jql
reporter = currentUser() AND status != Done
ORDER BY created DESC
```
## Advanced Functions
**Issues changed from status:**
```jql
status WAS "In Progress" AND status = "Done"
AND status CHANGED AFTER startOfWeek()
```
**Assignee changed:**
```jql
assignee CHANGED BY currentUser() AFTER -7d
```
**Issues re-opened:**
```jql
status WAS Done AND status != Done ORDER BY updated DESC
```
**Linked issues:**
```jql
issue IN linkedIssues("ABC-123") ORDER BY issuetype
```
**Parent epic:**
```jql
parent = ABC-123 ORDER BY rank
```
## Saved Filter Examples
**Daily Standup Filter:**
```jql
assignee = currentUser() AND sprint IN openSprints()
AND status != Done ORDER BY priority DESC
```
**Team Sprint Board Filter:**
```jql
project = ABC AND sprint IN openSprints() ORDER BY rank
```
**Bugs Dashboard Filter:**
```jql
project = ABC AND issuetype = Bug AND status != Done
ORDER BY priority DESC, created ASC
```
**Tech Debt Backlog:**
```jql
project = ABC AND labels = tech-debt AND status = Backlog
ORDER BY priority DESC
```
**Needs Triage:**
```jql
project = ABC AND status = "To Triage"
AND created >= -7d ORDER BY created ASC
```
FILE:references/WORKFLOWS.md
# Jira Workflows Reference
Comprehensive guide to Jira workflow design, transitions, conditions, validators, and post-functions.
## Default Workflows
### Simplified Workflow
```
Open → In Progress → Done
```
### Software Development Workflow
```
Backlog → Selected for Development → In Progress → In Review → Done
↑___________________________| (reopen)
```
### Bug Tracking Workflow
```
Open → In Progress → Fixed → Verified → Closed
↑ | |
|____Reopened________|________|
```
## Custom Workflow Design
### Design Principles
1. **Mirror your actual process** — don't force teams into artificial states
2. **Minimize statuses** — each status must represent a distinct work state where the item waits for a different action
3. **Clear ownership** — every status should have an obvious responsible party
4. **Allow rework** — always provide paths back for rejected/reopened items
5. **Separate "waiting" from "working"** — distinguish "In Review" (waiting) from "Reviewing" (actively working)
### Status Categories
Jira maps every status to one of four categories that drive board columns and JQL:
| Category | Meaning | JQL | Examples |
|----------|---------|-----|----------|
| `To Do` | Not started | `statusCategory = "To Do"` | Backlog, Open, New |
| `In Progress` | Active work | `statusCategory = "In Progress"` | In Progress, In Review, Testing |
| `Done` | Completed | `statusCategory = Done` | Done, Closed, Released |
| `Undefined` | Legacy/unused | — | Avoid using |
### Recommended Statuses by Team Type
**Engineering Team:**
```
Backlog → Ready → In Progress → Code Review → QA → Done
```
**Support Team:**
```
New → Triaged → In Progress → Waiting on Customer → Resolved → Closed
```
**Design Team:**
```
Backlog → Research → Design → Review → Approved → Handoff
```
## Transitions
### Transition Properties
| Property | Description |
|----------|-------------|
| **Name** | Display name on the button (e.g., "Start Work") |
| **Screen** | Form shown during transition (optional) |
| **Conditions** | Who can trigger this transition |
| **Validators** | Rules that must pass before transition executes |
| **Post-functions** | Actions executed after transition completes |
### Common Transition Patterns
**Start Work:**
```
Trigger: "Start Work" button
Condition: Assignee only
Validator: Issue must have assignee
Post-function: Set "In Progress" resolution to None
```
**Submit for Review:**
```
Trigger: "Submit for Review" button
Condition: Assignee or project admin
Validator: All sub-tasks must be Done
Post-function: Add comment "Submitted for review by {user}"
```
**Approve:**
```
Trigger: "Approve" button
Condition: Must be in "Reviewers" group
Validator: Must add comment
Post-function: Set resolution to "Done", fire event
```
## Conditions
### Built-in Conditions
| Condition | Use When |
|-----------|----------|
| **Only Assignee** | Only assigned user can transition |
| **Only Reporter** | Only creator can transition |
| **Permission Condition** | User must have specific permission |
| **Group Condition** | User must be in specified group |
| **Sub-Task Blocking** | All sub-tasks must be resolved |
| **Previous Status** | Issue must have been in a specific status |
| **User Is In Role** | User must have project role (Developer, Admin) |
### Combining Conditions
- **AND logic**: Add multiple conditions to one transition — ALL must pass
- **OR logic**: Create parallel transitions with different conditions
## Validators
### Built-in Validators
| Validator | Checks |
|-----------|--------|
| **Required Field** | Specific field must be populated |
| **Field Has Been Modified** | Field must change during transition |
| **Regular Expression** | Field must match regex pattern |
| **Permission Validator** | User must have permission |
| **Previous Status Validator** | Issue was in a required status |
### Common Validator Patterns
```
# Require comment on rejection
Validator: Comment Required
When: Transition = "Reject"
# Require fix version before release
Validator: Required Field = "Fix Version/s"
When: Transition = "Release"
# Require time logged before closing
Validator: Field Required = "Time Spent" (must be > 0)
When: Transition = "Close"
```
## Post-Functions
### Built-in Post-Functions
| Post-Function | Action |
|---------------|--------|
| **Set Field Value** | Assign a value to any field |
| **Update Issue Field** | Change assignee, priority, etc. |
| **Create Comment** | Add automated comment |
| **Fire Event** | Trigger notification event |
| **Assign to Lead** | Assign to project lead |
| **Assign to Reporter** | Assign back to creator |
| **Clear Field** | Remove field value |
| **Copy Value** | Copy field from parent/linked issue |
### Post-Function Execution Order
Post-functions execute in defined order. Standard sequence:
1. Set issue status (automatic, always first)
2. Add comment (if configured)
3. Update fields
4. Generate change history (automatic, always last)
5. Fire event (triggers notifications)
**Important:** "Generate change history" and "Fire event" must always be last — reorder if you add custom post-functions.
## Workflow Schemes
### What They Do
- Map issue types to workflows within a project
- One workflow scheme per project
- Different issue types can use different workflows
### Configuration Pattern
```
Project: MYPROJ
Workflow Scheme: "Engineering Workflow Scheme"
Bug → Bug Tracking Workflow
Story → Development Workflow
Task → Simple Workflow
Epic → Epic Workflow
Sub-task → Sub-task Workflow (inherits parent transitions)
```
## Best Practices
1. **Start simple, add complexity only when needed** — a 5-status workflow beats a 15-status one
2. **Name transitions as actions** — "Start Work" not "In Progress" (the status is "In Progress", the action is "Start Work")
3. **Use screens sparingly** — only show a screen when you need data from the user during transition
4. **Test with real users** — workflows that look good on paper may confuse the team
5. **Document your workflow** — add descriptions to statuses and transitions
6. **Use global transitions carefully** — a "Cancel" transition from any status is convenient but can bypass important gates
7. **Audit quarterly** — remove statuses with <5% usage
FILE:scripts/jql_query_builder.py
#!/usr/bin/env python3
"""
JQL Query Builder
Pattern-matching JQL builder from natural language descriptions. Maps common
phrases to JQL operators and constructs valid queries with syntax validation.
Usage:
python jql_query_builder.py "high priority bugs in PROJECT assigned to me"
python jql_query_builder.py "overdue tasks in PROJ" --format json
python jql_query_builder.py --patterns
"""
import argparse
import json
import re
import sys
from datetime import datetime
from typing import Any, Dict, List, Optional, Tuple
# ---------------------------------------------------------------------------
# Pattern Library
# ---------------------------------------------------------------------------
PATTERN_LIBRARY = {
"my_open_bugs": {
"phrases": ["my open bugs", "my bugs", "bugs assigned to me"],
"jql": 'assignee = currentUser() AND type = Bug AND status != Done',
"description": "All open bugs assigned to current user",
},
"high_priority_bugs": {
"phrases": ["high priority bugs", "critical bugs", "urgent bugs", "p1 bugs"],
"jql": 'type = Bug AND priority in (Highest, High) AND status != Done',
"description": "High and highest priority open bugs",
},
"my_open_tasks": {
"phrases": ["my open tasks", "my tasks", "tasks assigned to me", "my work"],
"jql": 'assignee = currentUser() AND status != Done',
"description": "All open issues assigned to current user",
},
"unassigned_issues": {
"phrases": ["unassigned", "unassigned issues", "no assignee"],
"jql": 'assignee is EMPTY AND status != Done',
"description": "Issues with no assignee",
},
"recently_created": {
"phrases": ["recently created", "new issues", "created this week", "recent"],
"jql": 'created >= -7d ORDER BY created DESC',
"description": "Issues created in the last 7 days",
},
"recently_updated": {
"phrases": ["recently updated", "updated this week", "recent changes"],
"jql": 'updated >= -7d ORDER BY updated DESC',
"description": "Issues updated in the last 7 days",
},
"overdue": {
"phrases": ["overdue", "past due", "missed deadline", "overdue tasks"],
"jql": 'duedate < now() AND status != Done',
"description": "Issues past their due date",
},
"due_this_week": {
"phrases": ["due this week", "due soon", "upcoming deadlines"],
"jql": 'duedate >= startOfWeek() AND duedate <= endOfWeek() AND status != Done',
"description": "Issues due this week",
},
"blocked_issues": {
"phrases": ["blocked", "blocked issues", "impediments"],
"jql": 'status = Blocked OR status = Impediment',
"description": "Issues in blocked or impediment status",
},
"in_progress": {
"phrases": ["in progress", "being worked on", "active work"],
"jql": 'status = "In Progress"',
"description": "Issues currently in progress",
},
"sprint_issues": {
"phrases": ["current sprint", "this sprint", "active sprint"],
"jql": 'sprint in openSprints()',
"description": "Issues in the current active sprint",
},
"backlog": {
"phrases": ["backlog", "backlog items", "not started"],
"jql": 'sprint is EMPTY AND status = "To Do" ORDER BY priority DESC',
"description": "Issues in the backlog not assigned to a sprint",
},
"stories_without_estimates": {
"phrases": ["no estimates", "unestimated", "missing estimates", "no story points"],
"jql": 'type = Story AND (storyPoints is EMPTY OR storyPoints = 0) AND status != Done',
"description": "Stories missing story point estimates",
},
"epics_in_progress": {
"phrases": ["active epics", "epics in progress", "open epics"],
"jql": 'type = Epic AND status != Done ORDER BY priority DESC',
"description": "Epics that are not yet completed",
},
"done_this_week": {
"phrases": ["done this week", "completed this week", "resolved this week"],
"jql": 'status changed to Done DURING (startOfWeek(), now())',
"description": "Issues completed during the current week",
},
"created_vs_resolved": {
"phrases": ["created vs resolved", "issue flow", "throughput"],
"jql": 'created >= -30d ORDER BY created DESC',
"description": "Issues created in the last 30 days for flow analysis",
},
"my_reported_issues": {
"phrases": ["my reported", "reported by me", "i created", "i reported"],
"jql": 'reporter = currentUser() ORDER BY created DESC',
"description": "Issues reported by current user",
},
"stale_issues": {
"phrases": ["stale", "stale issues", "not updated", "abandoned"],
"jql": 'updated <= -30d AND status != Done ORDER BY updated ASC',
"description": "Issues not updated in 30+ days",
},
"subtasks_without_parent": {
"phrases": ["orphan subtasks", "subtasks no parent", "loose subtasks"],
"jql": 'type = Sub-task AND parent is EMPTY',
"description": "Subtasks missing parent issues",
},
"high_priority_unassigned": {
"phrases": ["high priority unassigned", "urgent unassigned", "critical no owner"],
"jql": 'priority in (Highest, High) AND assignee is EMPTY AND status != Done',
"description": "High priority issues with no assignee",
},
"bugs_by_component": {
"phrases": ["bugs by component", "component bugs"],
"jql": 'type = Bug AND status != Done ORDER BY component ASC',
"description": "Open bugs organized by component",
},
"resolved_recently": {
"phrases": ["resolved recently", "recently resolved", "fixed this month"],
"jql": 'resolved >= -30d ORDER BY resolved DESC',
"description": "Issues resolved in the last 30 days",
},
}
# Keyword-to-JQL fragment mapping for dynamic query building
KEYWORD_FRAGMENTS = {
# Issue types
"bug": ("type", "= Bug"),
"bugs": ("type", "= Bug"),
"story": ("type", "= Story"),
"stories": ("type", "= Story"),
"task": ("type", "= Task"),
"tasks": ("type", "= Task"),
"epic": ("type", "= Epic"),
"epics": ("type", "= Epic"),
"subtask": ("type", "= Sub-task"),
"sub-task": ("type", "= Sub-task"),
# Statuses
"open": ("status", "!= Done"),
"closed": ("status", "= Done"),
"done": ("status", "= Done"),
"resolved": ("status", "= Done"),
"todo": ("status", '= "To Do"'),
# Priorities
"critical": ("priority", "= Highest"),
"highest": ("priority", "= Highest"),
"high": ("priority", "in (Highest, High)"),
"medium": ("priority", "= Medium"),
"low": ("priority", "in (Low, Lowest)"),
"lowest": ("priority", "= Lowest"),
# Assignee
"me": ("assignee", "= currentUser()"),
"mine": ("assignee", "= currentUser()"),
"unassigned": ("assignee", "is EMPTY"),
# Time
"overdue": ("duedate", "< now()"),
"today": ("duedate", "= now()"),
}
PROJECT_PATTERN = re.compile(r'\b([A-Z]{2,10})\b')
ASSIGNEE_PATTERN = re.compile(r'assigned\s+to\s+(\w+)', re.IGNORECASE)
LABEL_PATTERN = re.compile(r'label[s]?\s*[=:]\s*["\']?(\w+)["\']?', re.IGNORECASE)
COMPONENT_PATTERN = re.compile(r'component[s]?\s*[=:]\s*["\']?(\w+)["\']?', re.IGNORECASE)
DATE_RANGE_PATTERN = re.compile(r'last\s+(\d+)\s+(day|week|month)s?', re.IGNORECASE)
SPRINT_NAME_PATTERN = re.compile(r'sprint\s+["\']?(\w[\w\s]*\w)["\']?', re.IGNORECASE)
# Words to exclude from project matching
EXCLUDED_WORDS = {
"AND", "OR", "NOT", "IN", "IS", "TO", "BY", "ON", "DO", "BE",
"THE", "ALL", "MY", "NO", "OF", "AT", "AS", "IF", "IT",
"BUG", "BUGS", "TASK", "TASKS", "STORY", "EPIC", "DONE",
"HIGH", "LOW", "MEDIUM", "JQL",
}
# ---------------------------------------------------------------------------
# Query Builder
# ---------------------------------------------------------------------------
def find_matching_pattern(description: str) -> Optional[Dict[str, Any]]:
"""Check if description matches a known pattern exactly."""
desc_lower = description.lower().strip()
for pattern_name, pattern_data in PATTERN_LIBRARY.items():
for phrase in pattern_data["phrases"]:
if phrase in desc_lower or desc_lower in phrase:
return {
"pattern_name": pattern_name,
"jql": pattern_data["jql"],
"description": pattern_data["description"],
"match_type": "exact_pattern",
}
return None
def build_jql_from_description(description: str) -> Dict[str, Any]:
"""Build JQL query from natural language description."""
# First try exact pattern match
pattern_match = find_matching_pattern(description)
if pattern_match:
# Augment with project if mentioned
project = _extract_project(description)
if project:
pattern_match["jql"] = f'project = {project} AND {pattern_match["jql"]}'
return pattern_match
# Dynamic query building
clauses = []
used_fields = set()
desc_lower = description.lower()
# Extract project
project = _extract_project(description)
if project:
clauses.append(f"project = {project}")
used_fields.add("project")
# Extract keyword-based fragments
for keyword, (field, fragment) in KEYWORD_FRAGMENTS.items():
if keyword in desc_lower.split() and field not in used_fields:
clauses.append(f"{field} {fragment}")
used_fields.add(field)
# Extract explicit assignee
assignee_match = ASSIGNEE_PATTERN.search(description)
if assignee_match and "assignee" not in used_fields:
assignee = assignee_match.group(1)
if assignee.lower() in ("me", "myself"):
clauses.append("assignee = currentUser()")
else:
clauses.append(f'assignee = "{assignee}"')
used_fields.add("assignee")
# Extract labels
label_match = LABEL_PATTERN.search(description)
if label_match:
clauses.append(f'labels = "{label_match.group(1)}"')
# Extract component
component_match = COMPONENT_PATTERN.search(description)
if component_match:
clauses.append(f'component = "{component_match.group(1)}"')
# Extract date ranges
date_match = DATE_RANGE_PATTERN.search(description)
if date_match:
amount = date_match.group(1)
unit = date_match.group(2).lower()
unit_char = {"day": "d", "week": "w", "month": "m"}.get(unit, "d")
clauses.append(f"created >= -{amount}{unit_char}")
# Extract sprint reference
sprint_match = SPRINT_NAME_PATTERN.search(description)
if sprint_match:
sprint_name = sprint_match.group(1).strip()
if sprint_name.lower() in ("current", "active", "open"):
clauses.append("sprint in openSprints()")
else:
clauses.append(f'sprint = "{sprint_name}"')
# Default: if no status clause and not looking for done items
if "status" not in used_fields and "done" not in desc_lower and "closed" not in desc_lower:
clauses.append("status != Done")
if not clauses:
return {
"jql": "",
"description": "Could not build query from description",
"match_type": "no_match",
"error": "No recognizable patterns found in description",
}
jql = " AND ".join(clauses)
# Add ORDER BY for common scenarios
if "recent" in desc_lower or "latest" in desc_lower:
jql += " ORDER BY created DESC"
elif "priority" in desc_lower or "urgent" in desc_lower:
jql += " ORDER BY priority DESC"
return {
"jql": jql,
"description": f"Dynamic query from: {description}",
"match_type": "dynamic",
"clauses_used": len(clauses),
}
def _extract_project(description: str) -> Optional[str]:
"""Extract project key from description."""
# Look for IN/in PROJECT pattern
in_project = re.search(r'\bin\s+([A-Z]{2,10})\b', description)
if in_project and in_project.group(1) not in EXCLUDED_WORDS:
return in_project.group(1)
# Look for standalone project keys
for match in PROJECT_PATTERN.finditer(description):
word = match.group(1)
if word not in EXCLUDED_WORDS:
return word
return None
def validate_jql_syntax(jql: str) -> Dict[str, Any]:
"""Basic JQL syntax validation."""
issues = []
if not jql.strip():
return {"valid": False, "issues": ["Empty query"]}
# Check balanced quotes
single_quotes = jql.count("'")
double_quotes = jql.count('"')
if single_quotes % 2 != 0:
issues.append("Unbalanced single quotes")
if double_quotes % 2 != 0:
issues.append("Unbalanced double quotes")
# Check balanced parentheses
open_parens = jql.count("(")
close_parens = jql.count(")")
if open_parens != close_parens:
issues.append(f"Unbalanced parentheses: {open_parens} open, {close_parens} close")
# Check for known JQL operators
valid_operators = {"=", "!=", ">", "<", ">=", "<=", "~", "!~", "in", "not in", "is", "is not", "was", "was not", "changed"}
jql_upper = jql.upper()
# Check AND/OR placement
if jql_upper.strip().startswith("AND") or jql_upper.strip().startswith("OR"):
issues.append("Query cannot start with AND/OR")
if jql_upper.strip().endswith("AND") or jql_upper.strip().endswith("OR"):
issues.append("Query cannot end with AND/OR")
# Check ORDER BY syntax
order_match = re.search(r'ORDER\s+BY\s+(\w+)(?:\s+(ASC|DESC))?', jql, re.IGNORECASE)
if "ORDER" in jql_upper and not order_match:
issues.append("Invalid ORDER BY syntax")
return {
"valid": len(issues) == 0,
"issues": issues,
"query_length": len(jql),
}
# ---------------------------------------------------------------------------
# Output Formatting
# ---------------------------------------------------------------------------
def format_text_output(result: Dict[str, Any]) -> str:
"""Format results as readable text report."""
lines = []
lines.append("=" * 60)
lines.append("JQL QUERY BUILDER RESULTS")
lines.append("=" * 60)
lines.append("")
if "error" in result:
lines.append(f"ERROR: {result['error']}")
return "\n".join(lines)
lines.append(f"Match Type: {result.get('match_type', 'unknown')}")
lines.append(f"Description: {result.get('description', '')}")
lines.append("")
lines.append("GENERATED JQL")
lines.append("-" * 30)
lines.append(result.get("jql", ""))
lines.append("")
validation = result.get("validation", {})
if validation:
lines.append("VALIDATION")
lines.append("-" * 30)
lines.append(f"Valid: {'Yes' if validation.get('valid') else 'No'}")
if validation.get("issues"):
for issue in validation["issues"]:
lines.append(f" - {issue}")
if result.get("pattern_name"):
lines.append("")
lines.append(f"Matched Pattern: {result['pattern_name']}")
return "\n".join(lines)
def format_patterns_output(output_format: str) -> str:
"""Format available patterns list."""
if output_format == "json":
patterns = {}
for name, data in PATTERN_LIBRARY.items():
patterns[name] = {
"description": data["description"],
"phrases": data["phrases"],
"jql": data["jql"],
}
return json.dumps(patterns, indent=2)
lines = []
lines.append("=" * 60)
lines.append("AVAILABLE JQL PATTERNS")
lines.append("=" * 60)
lines.append("")
for name, data in PATTERN_LIBRARY.items():
lines.append(f" {name}")
lines.append(f" Description: {data['description']}")
lines.append(f" Phrases: {', '.join(data['phrases'])}")
lines.append(f" JQL: {data['jql']}")
lines.append("")
lines.append(f"Total patterns: {len(PATTERN_LIBRARY)}")
return "\n".join(lines)
def format_json_output(result: Dict[str, Any]) -> Dict[str, Any]:
"""Format results as JSON."""
return result
# ---------------------------------------------------------------------------
# CLI Interface
# ---------------------------------------------------------------------------
def main() -> int:
"""Main CLI entry point."""
parser = argparse.ArgumentParser(
description="Build JQL queries from natural language descriptions"
)
parser.add_argument(
"description",
nargs="?",
help="Natural language description of the query",
)
parser.add_argument(
"--format",
choices=["text", "json"],
default="text",
help="Output format (default: text)",
)
parser.add_argument(
"--patterns",
action="store_true",
help="List all available query patterns",
)
args = parser.parse_args()
try:
if args.patterns:
print(format_patterns_output(args.format))
return 0
if not args.description:
parser.error("description is required unless --patterns is used")
# Build query
result = build_jql_from_description(args.description)
# Validate
if result.get("jql"):
result["validation"] = validate_jql_syntax(result["jql"])
# Output results
if args.format == "json":
output = format_json_output(result)
print(json.dumps(output, indent=2))
else:
output = format_text_output(result)
print(output)
return 0
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/workflow_validator.py
#!/usr/bin/env python3
"""
Workflow Validator
Validates Jira workflow definitions (JSON input) for anti-patterns and common
issues. Checks for dead-end states, orphan states, missing transitions, circular
paths, and produces a health score with severity-rated findings.
Usage:
python workflow_validator.py workflow.json
python workflow_validator.py workflow.json --format json
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Optional, Set, Tuple
# ---------------------------------------------------------------------------
# Validation Configuration
# ---------------------------------------------------------------------------
MAX_RECOMMENDED_STATES = 10
REQUIRED_TERMINAL_STATES = {"done", "closed", "resolved", "completed"}
SEVERITY_WEIGHTS = {
"error": 20,
"warning": 10,
"info": 3,
}
# ---------------------------------------------------------------------------
# Validation Rules
# ---------------------------------------------------------------------------
def check_state_count(states: List[str]) -> List[Dict[str, str]]:
"""Check if the workflow has too many states."""
findings = []
count = len(states)
if count > MAX_RECOMMENDED_STATES:
findings.append({
"rule": "state_count",
"severity": "warning",
"message": f"Workflow has {count} states (recommended max: {MAX_RECOMMENDED_STATES}). "
f"Complex workflows slow teams down and increase error rates.",
})
elif count < 2:
findings.append({
"rule": "state_count",
"severity": "error",
"message": f"Workflow has only {count} state(s). A minimum of 2 states is required.",
})
if count > 15:
findings[-1]["severity"] = "error"
return findings
def check_dead_end_states(
states: List[str],
transitions: List[Dict[str, str]],
terminal_states: Set[str],
) -> List[Dict[str, str]]:
"""Find states with no outgoing transitions that are not terminal."""
findings = []
outgoing = set()
for t in transitions:
outgoing.add(t.get("from", "").lower())
for state in states:
state_lower = state.lower()
if state_lower not in outgoing and state_lower not in terminal_states:
findings.append({
"rule": "dead_end_state",
"severity": "error",
"message": f"State '{state}' has no outgoing transitions and is not a terminal state. "
f"Issues will get stuck here.",
})
return findings
def check_orphan_states(
states: List[str],
transitions: List[Dict[str, str]],
initial_state: Optional[str],
) -> List[Dict[str, str]]:
"""Find states with no incoming transitions (except the initial state)."""
findings = []
incoming = set()
for t in transitions:
incoming.add(t.get("to", "").lower())
initial_lower = (initial_state or "").lower()
for state in states:
state_lower = state.lower()
if state_lower not in incoming and state_lower != initial_lower:
findings.append({
"rule": "orphan_state",
"severity": "warning",
"message": f"State '{state}' has no incoming transitions and is not the initial state. "
f"This state may be unreachable.",
})
return findings
def check_missing_terminal_state(states: List[str]) -> List[Dict[str, str]]:
"""Check that at least one terminal/done state exists."""
findings = []
states_lower = {s.lower() for s in states}
has_terminal = bool(states_lower & REQUIRED_TERMINAL_STATES)
if not has_terminal:
findings.append({
"rule": "missing_terminal_state",
"severity": "error",
"message": f"No terminal state found. Expected one of: {', '.join(sorted(REQUIRED_TERMINAL_STATES))}. "
f"Issues cannot be marked as complete.",
})
return findings
def check_duplicate_transition_names(
transitions: List[Dict[str, str]],
) -> List[Dict[str, str]]:
"""Check for duplicate transition names from the same state."""
findings = []
seen = {}
for t in transitions:
name = t.get("name", "").lower()
from_state = t.get("from", "").lower()
key = (from_state, name)
if key in seen:
findings.append({
"rule": "duplicate_transition",
"severity": "warning",
"message": f"Duplicate transition name '{t.get('name', '')}' from state '{t.get('from', '')}'. "
f"This can confuse users selecting transitions.",
})
else:
seen[key] = True
return findings
def check_missing_transitions(
states: List[str],
transitions: List[Dict[str, str]],
) -> List[Dict[str, str]]:
"""Check for states referenced in transitions but not defined."""
findings = []
defined_states = {s.lower() for s in states}
for t in transitions:
from_state = t.get("from", "").lower()
to_state = t.get("to", "").lower()
if from_state and from_state not in defined_states:
findings.append({
"rule": "undefined_state_reference",
"severity": "error",
"message": f"Transition references undefined source state '{t.get('from', '')}'.",
})
if to_state and to_state not in defined_states:
findings.append({
"rule": "undefined_state_reference",
"severity": "error",
"message": f"Transition references undefined target state '{t.get('to', '')}'.",
})
return findings
def check_circular_paths(
states: List[str],
transitions: List[Dict[str, str]],
terminal_states: Set[str],
) -> List[Dict[str, str]]:
"""Detect circular paths that have no exit to a terminal state."""
findings = []
# Build adjacency list
adjacency = {}
for state in states:
adjacency[state.lower()] = set()
for t in transitions:
from_state = t.get("from", "").lower()
to_state = t.get("to", "").lower()
if from_state in adjacency:
adjacency[from_state].add(to_state)
# Find strongly connected components using iterative DFS
def can_reach_terminal(start: str) -> bool:
visited = set()
stack = [start]
while stack:
node = stack.pop()
if node in terminal_states:
return True
if node in visited:
continue
visited.add(node)
for neighbor in adjacency.get(node, set()):
stack.append(neighbor)
return False
# Check each non-terminal state
for state in states:
state_lower = state.lower()
if state_lower not in terminal_states:
if not can_reach_terminal(state_lower):
findings.append({
"rule": "circular_no_exit",
"severity": "error",
"message": f"State '{state}' cannot reach any terminal state. "
f"Issues entering this state will never be resolved.",
})
return findings
def check_self_transitions(transitions: List[Dict[str, str]]) -> List[Dict[str, str]]:
"""Check for transitions that go from a state to itself."""
findings = []
for t in transitions:
if t.get("from", "").lower() == t.get("to", "").lower():
findings.append({
"rule": "self_transition",
"severity": "info",
"message": f"State '{t.get('from', '')}' has a self-transition '{t.get('name', '')}'. "
f"Ensure this is intentional (e.g., for triggering automation).",
})
return findings
# ---------------------------------------------------------------------------
# Main Validation
# ---------------------------------------------------------------------------
def validate_workflow(data: Dict[str, Any]) -> Dict[str, Any]:
"""Run all validations on a workflow definition."""
states = data.get("states", [])
transitions = data.get("transitions", [])
initial_state = data.get("initial_state", states[0] if states else None)
if not states:
return {
"health_score": 0,
"grade": "invalid",
"findings": [{"rule": "no_states", "severity": "error", "message": "No states defined in workflow"}],
"summary": {"errors": 1, "warnings": 0, "info": 0},
}
# Determine terminal states
states_lower = {s.lower() for s in states}
terminal_states = states_lower & REQUIRED_TERMINAL_STATES
# Custom terminal states from input
custom_terminals = data.get("terminal_states", [])
for ct in custom_terminals:
terminal_states.add(ct.lower())
# Run all checks
all_findings = []
all_findings.extend(check_state_count(states))
all_findings.extend(check_dead_end_states(states, transitions, terminal_states))
all_findings.extend(check_orphan_states(states, transitions, initial_state))
all_findings.extend(check_missing_terminal_state(states))
all_findings.extend(check_duplicate_transition_names(transitions))
all_findings.extend(check_missing_transitions(states, transitions))
all_findings.extend(check_circular_paths(states, transitions, terminal_states))
all_findings.extend(check_self_transitions(transitions))
# Calculate health score
summary = {"errors": 0, "warnings": 0, "info": 0}
penalty = 0
for finding in all_findings:
severity = finding["severity"]
summary[severity] = summary.get(severity, 0) + 1
penalty += SEVERITY_WEIGHTS.get(severity, 0)
health_score = max(0, 100 - penalty)
if health_score >= 90:
grade = "excellent"
elif health_score >= 75:
grade = "good"
elif health_score >= 55:
grade = "fair"
else:
grade = "poor"
return {
"health_score": health_score,
"grade": grade,
"findings": all_findings,
"summary": summary,
"workflow_info": {
"state_count": len(states),
"transition_count": len(transitions),
"initial_state": initial_state,
"terminal_states": sorted(terminal_states),
},
}
# ---------------------------------------------------------------------------
# Output Formatting
# ---------------------------------------------------------------------------
def format_text_output(result: Dict[str, Any]) -> str:
"""Format results as readable text report."""
lines = []
lines.append("=" * 60)
lines.append("WORKFLOW VALIDATION REPORT")
lines.append("=" * 60)
lines.append("")
# Health summary
lines.append("HEALTH SUMMARY")
lines.append("-" * 30)
lines.append(f"Health Score: {result['health_score']}/100")
lines.append(f"Grade: {result['grade'].title()}")
lines.append("")
# Workflow info
info = result.get("workflow_info", {})
if info:
lines.append("WORKFLOW INFO")
lines.append("-" * 30)
lines.append(f"States: {info.get('state_count', 0)}")
lines.append(f"Transitions: {info.get('transition_count', 0)}")
lines.append(f"Initial State: {info.get('initial_state', 'N/A')}")
lines.append(f"Terminal States: {', '.join(info.get('terminal_states', []))}")
lines.append("")
# Summary
summary = result.get("summary", {})
lines.append("FINDINGS SUMMARY")
lines.append("-" * 30)
lines.append(f"Errors: {summary.get('errors', 0)}")
lines.append(f"Warnings: {summary.get('warnings', 0)}")
lines.append(f"Info: {summary.get('info', 0)}")
lines.append("")
# Detailed findings
findings = result.get("findings", [])
if findings:
lines.append("DETAILED FINDINGS")
lines.append("-" * 30)
for i, finding in enumerate(findings, 1):
severity = finding["severity"].upper()
lines.append(f"{i}. [{severity}] {finding['message']}")
lines.append(f" Rule: {finding['rule']}")
lines.append("")
else:
lines.append("No issues found. Workflow looks healthy!")
return "\n".join(lines)
def format_json_output(result: Dict[str, Any]) -> Dict[str, Any]:
"""Format results as JSON."""
return result
# ---------------------------------------------------------------------------
# CLI Interface
# ---------------------------------------------------------------------------
def main() -> int:
"""Main CLI entry point."""
parser = argparse.ArgumentParser(
description="Validate Jira workflow definitions for anti-patterns"
)
parser.add_argument(
"workflow_file",
help="JSON file containing workflow definition (states, transitions)",
)
parser.add_argument(
"--format",
choices=["text", "json"],
default="text",
help="Output format (default: text)",
)
args = parser.parse_args()
try:
with open(args.workflow_file, "r") as f:
data = json.load(f)
result = validate_workflow(data)
if args.format == "json":
print(json.dumps(format_json_output(result), indent=2))
else:
print(format_text_output(result))
return 0
except FileNotFoundError:
print(f"Error: File '{args.workflow_file}' not found", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"Error: Invalid JSON in '{args.workflow_file}': {e}", file=sys.stderr)
return 1
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
return 1
if __name__ == "__main__":
sys.exit(main())
Chất vấn thận trọng về doanh thu, tỷ lệ thắng, NRR và thời gian làm quen của đội bán hàng.
--- name: "cro-review" description: "/cs:cro-review <plan> — Pipeline-paranoid interrogation of revenue, win rate, NRR, and ramp time." --- # /cs:cro-review — CRO Forcing Questions **Command:** `/cs:cro-review <plan>` The pipeline-paranoid operator pressure-tests revenue assumptions. Six questions that surface next-quarter pain this quarter. ## When to Run - Before committing to a quarterly revenue target - Before changing sales motion (PLG ↔ sales-led, mid-market ↔ enterprise) - Before hiring a batch of reps - When pipeline coverage drops below 3x - When NRR is trending down ## The Six CRO Questions ### 1. Pipeline Coverage **What is pipeline coverage for the current quarter, by stage?** - Inbound-heavy: 3x. Outbound-heavy: 4x. Below either threshold = act now. - Stage-weighted, not just total. ### 2. Win Rate Trajectory **What's win rate this quarter vs the last 4 — and what's the leak point?** - Stage-by-stage conversion. - If a single stage softens, identify why before forecasting. ### 3. NRR Decomposition **What's gross retention, contraction, and expansion separately?** - NRR alone hides churn. - A 110% NRR with 95% gross retention is different from 110% with 80%. ### 4. Ramp Time **For the last 4 hires, how many days to first deal and to quota?** - If ramp > 90 days at growth stage, hiring profile or enablement is broken. - Forecasted hires must build in ramp. ### 5. Discount Discipline **What's the median discount this quarter vs last 4? Where is it creeping?** - Discount creep is the leading indicator of pricing or positioning weakness. - Cap discounts by approver tier. ### 6. Pipeline Source Mix **What % of pipeline is marketing-sourced, sales-sourced, partner-sourced?** - If one source dominates > 80%, you have concentration risk. - Cross-check with cs-cmo-advisor. ## Workflow ```bash python ../../../skills/cro-advisor/scripts/revenue_forecast_model.py python ../../../skills/cro-advisor/scripts/churn_analyzer.py ``` ## Output Format ```markdown # CRO Review: <plan> **Date:** YYYY-MM-DD ## Pipeline - Coverage: X.Xx (target 3x+) - Win rate: X% (4Q trend: ↑ / → / ↓) - Top leaking stage: <name> ## Retention - Gross retention: X% - NRR: X% - Expansion: X% - Contraction: X% ## Ramp - New hires last quarter: N - Median days to first deal: X - Median days to quota: X ## Discount - Median discount this quarter: X% - Trend vs 4Q ago: <delta> ## Source Mix - Marketing: X% | Sales: X% | Partner: X% ## Verdict 🟢 ON PLAN | 🟡 GAP | 🔴 PIPELINE CRISIS ## Next Steps [3 concrete actions] ``` ## Routing - `/cs:cfo-review` — does this hit the cash plan? - `/cs:cmo-review` — is pipeline source-mix healthy? - `/cs:execute` — quarterly plan if GREEN - `/cs:boardroom` — if RED ## Related - Agent: [`cs-cro-advisor`](../../agents/cs-cro-advisor.md) - Skill: [`cro-advisor`](../../../skills/cro-advisor/SKILL.md) - Execution: `../../../../business-growth/` --- **Version:** 1.0.0
Rà soát thay đổi đã staged hoặc commit gần nhất theo 4 nguyên tắc của Karpathy: độ phức tạp, nhiễu diff, giả định ngầm và xác minh mục tiêu.
--- name: karpathy-check description: Run Karpathy's 4-principle review on staged changes or the last commit. Checks complexity, diff noise, hidden assumptions, and goal verification. Usage /karpathy-check [--last-commit] --- # /karpathy-check Review your staged changes (or last commit) against Karpathy's 4 coding principles. ## Usage ``` /karpathy-check # review staged changes /karpathy-check --last-commit # review the most recent commit ``` ## What it runs 1. **Principle #2 (Simplicity):** `scripts/complexity_checker.py` on all changed files — detects over-engineering, premature abstractions, deep nesting, long functions 2. **Principle #3 (Surgical):** `scripts/diff_surgeon.py` on the diff — detects comment-only changes, whitespace noise, style drift, drive-by refactors 3. **Principles #1 + #4 (Think + Goals):** The `karpathy-reviewer` agent reads the diff and applies human-judgment checks — hidden assumptions, missing verification ## Output A structured report with per-principle verdicts and specific line-level fix recommendations. ## When to run - Before committing (catches noise and overcomplication early) - After completing a feature (sanity check before PR) - When you suspect the LLM overcoded something ## Sub-agent Dispatches the `karpathy-reviewer` agent. See `agents/karpathy-reviewer.md`. ## Scripts - `engineering/karpathy-coder/scripts/complexity_checker.py` - `engineering/karpathy-coder/scripts/diff_surgeon.py` - `engineering/karpathy-coder/scripts/assumption_linter.py` - `engineering/karpathy-coder/scripts/goal_verifier.py` ## Skill Reference → `engineering/karpathy-coder/SKILL.md`
Soạn thảo, kiểm tra và làm sạch SOP, runbook nội bộ như mua sắm, offboarding nhà cung cấp, onboarding nhân viên, hoàn chi phí và cấp quyền hệ thống.
---
name: knowledge-ops
description: Use when a Head of Ops, Knowledge Manager, or TPM-Internal needs to author, validate, or clean up company SOPs and internal runbooks (procurement intake, vendor offboarding, incident-comms cascade, employee onboarding, expense reimbursement, system-access provisioning, customer-escalation playbook) — including 5W2H completeness checks (Who-What-When-Where-Why-How-HowMuch), cross-link and orphan-page validation across a sprawling Notion/Confluence/Obsidian wiki, KB ingestion + hygiene reporting, ops onboarding doc generation, and runbook step verification (named owner, expected duration, observable success signal, rollback path, escalation contact). Pairs Kaoru Ishikawa's 5W2H method, Atul Gawande's *The Checklist Manifesto*, ISO 9001, ITIL v4 Service Operation, FDA 21 CFR Part 211, and Google SRE Workbook runbook discipline with deterministic stdlib-only Python tools that score completeness, detect anti-patterns, and emit prioritized cleanup lists. Distinct from `engineering/llm-wiki` (Karpathy-style personal PKM second brain), `engineering-team/runbook-generator` (system-ops production debugging runbook), `project-management/*` (Jira/Confluence delivery + ticket tracking), and sibling `business-operations/process-mapper` (BPMN process *design*, while knowledge-ops is process *documentation*).
context: fork
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [bizops, sop, runbook, knowledge-management, kb, 5w2h, wiki, ops-documentation]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# knowledge-ops
Company SOP + internal runbook authoring, 5W2H completeness validation, and KB hygiene reporting for Head-of-Ops / Knowledge-Manager / TPM-Internal personas.
## Purpose
An ops organization three years in accumulates a sprawl: 600 Notion pages, 200 Confluence runbooks, three Obsidian vaults, a `Drive/SOPs/` folder, and a `Slack #ops-questions` channel that exists because nobody can find the canonical doc. Predictable failure modes:
1. **No owner** — 40% of SOPs name "the team" instead of a person. When the doc rots, nobody is accountable.
2. **No last-reviewed date** — a 2023 vendor-offboarding SOP still references a procurement tool sunset in 2024.
3. **Vague success signals** — runbook step 4 says "verify the service is up". A new operator can't tell what that means.
4. **No rollback path** — incident-comms cascade runbook tells you how to send the alert. It doesn't tell you how to retract it when the alert was wrong.
5. **Orphan pages** — half the KB has no inbound links. Nobody finds them via navigation; they only exist because somebody knew the URL.
6. **Glossary drift** — "CSM" means Customer Success Manager in three docs and Customer Solutions Manager in five. New hires guess wrong for six months.
7. **Happy-path-only SOPs** — the doc covers what happens when everything works. It doesn't cover the 30% case where it doesn't.
This skill answers the operator's actual question: **"Which 20 docs do I fix first, and what specifically is wrong with each?"** — with deterministic logic, not intuition.
## When to use
- Authoring a new SOP for a cross-functional company process (procurement intake, vendor offboarding, incident-comms cascade, employee onboarding, expense reimbursement, customer-escalation playbook, security-incident comms, system-access provisioning).
- Validating an existing internal runbook before it goes into rotation (every step must have a named owner, expected duration, observable success signal, observable failure signal, rollback path, escalation contact).
- Ingesting a multi-document KB export (Notion zip, Confluence space export, Obsidian vault, `Drive/SOPs/` directory) and surfacing what's broken: orphan pages, stale pages (no edit > 12 months), glossary drift, missing-owner pages, cross-link map.
- Onboarding a new ops hire by generating the SOPs and ops-handbook pages they need to read in week 1.
- Wiki cleanup sprints — quarterly hygiene work where the org decides which 30 docs to archive, rewrite, or merge.
## Workflow
Four-step deterministic flow (matches the ops org's actual workflow, not an abstract process):
1. **Ingest KB.** Run `kb_ingester.py --input <vault-dir>` on the existing wiki export. Output is a markdown health report: orphan pages, stale pages, glossary drift, missing-owner pages, cross-link map, prioritized cleanup list. The report ranks the top-20 docs to fix first — usually a mix of high-traffic stale docs and compliance-relevant missing-owner docs. Take this list to the cleanup sprint.
2. **Validate existing runbooks.** For each runbook in the cleanup list (or any new runbook before it goes into rotation), run `runbook_validator.py --input <runbook.md>`. The validator scores each step against six checks (named owner, expected duration, observable success signal, observable failure signal, rollback path, escalation contact) and produces a per-step traffic-light + overall validity score 0-100 + MUST-FIX issue list. A runbook scoring < 60 is not safe to use in an incident.
3. **Generate missing SOPs.** For SOPs that need to be written from scratch (or rewritten because the existing one is unsalvageable), run `sop_generator.py --input <metadata.json> --profile <ops|support|finance|hr|it|regulated>`. Output is a 5W2H-structured SOP scaffold: Who (RACI), What (process steps), When (triggers + frequency), Where (system + tool), Why (purpose + regulatory basis), How (step-by-step), How-much (cost + time per execution). The `regulated` profile adds version control, signoff, and audit-trail sections (ISO 9001 / FDA 21 CFR Part 211 / SOC 2 / HIPAA).
4. **Cross-link + close the loop.** Re-run `kb_ingester.py` after the cleanup sprint to verify orphan-page count is down and glossary drift is resolved. The metric that matters is **"unfindable docs"** (orphans) and **"unsafe runbooks"** (validity score < 60) — not page count.
## Scripts
**`scripts/sop_generator.py`** — Reads a JSON metadata file describing an SOP (process owner, triggering event, audience role, frequency, regulatory overlay, inputs, outputs, steps outline) and emits a full 5W2H-structured SOP in markdown (or normalized JSON). The `--profile` flag tunes the output: `ops` (general internal ops), `support` (customer-support runbook style), `finance` (controls + reconciliation focus), `hr` (sensitive-data flagging), `it` (system + access focus), `regulated` (adds version control, signoff matrix, audit-trail). Regulatory overlays (`SOC2`, `HIPAA`, `ISO13485`, `GDPR`, `SOX`) attach the appropriate compliance preamble. `--sample` prints a complete vendor-offboarding SOP example. Stdlib only.
**`scripts/runbook_validator.py`** — Reads a runbook (markdown file or JSON) and validates each step against six required attributes: (1) named owner (not "the team", not "ops"), (2) expected duration (concrete number + unit), (3) observable success signal (e.g., "HTTP 200 from `/healthz`" — not "service is up"), (4) observable failure signal, (5) rollback path (or explicit "this step cannot be rolled back, escalate to X"), (6) escalation contact (named person or named on-call rotation). Output is a per-step traffic-light (GREEN/AMBER/RED), an overall validity score 0-100, and a MUST-FIX issue list. Verdict: ≥ 80 = SAFE-TO-USE, 60-79 = USE-WITH-CAUTION, < 60 = NOT-SAFE. `--sample` prints a deliberately-broken incident-comms runbook to demonstrate failure detection. Stdlib only.
**`scripts/kb_ingester.py`** — Walks a directory of markdown files (Notion export, Confluence space export, Obsidian vault, `Drive/SOPs/` directory). Extracts: (a) cross-link map (which page references which, via markdown `[link](path)` syntax), (b) glossary candidates (frequently used proper nouns and acronyms that recur in 3+ docs without a single canonical definition page), (c) orphan pages (no inbound links from anywhere in the vault), (d) glossary drift (the same term defined or used inconsistently across docs — e.g., "CSM" expanded differently in two places), (e) stale pages (no edit in > 12 months, detected via filesystem mtime or YAML `last_reviewed` frontmatter), (f) missing-owner pages (no `owner:` field in frontmatter). Emits a KB health report markdown with a prioritized top-20 cleanup list ranked by `staleness × inbound-link-count` (high-traffic stale docs first). `--sample` builds a tiny synthetic 8-page vault in a tmpdir and runs the full pipeline against it. Stdlib only.
## References
- `references/5w2h_sop_canon.md` — Kaoru Ishikawa's 5W2H method, Toyota standard-work discipline, Atul Gawande's checklist manifesto, Atlassian Confluence SOP guidance, ISO 9001 SOP requirements, ITIL v4 Service Operation, FDA 21 CFR Part 211. Eight cited sources covering SOP authoring canon.
- `references/runbook_canon.md` — Google SRE Workbook (runbook chapter), Atlassian incident-management runbooks, PagerDuty Incident Response taxonomy, AWS Well-Architected operational excellence pillar, Charity Majors on observability-runbook integration, Susan Fowler on production-ready microservices, ITIL v4 Operations. Seven cited sources covering runbook design canon.
- `references/kb_hygiene_anti_patterns.md` — Eight anti-patterns drawn from Notion/Confluence wiki industry research, Mozilla SUMO knowledge-base lessons, Stack Overflow community-management research, the Atlassian Team Playbook, MIT TIK org-wiki studies, Cynthia Lee on glossary drift, and Adam Wiggins on "documentation rot".
## Assumptions
1. The KB is in markdown (or can be exported to markdown — Notion, Confluence, Obsidian, and Google Docs all support this). HTML-only or PDF-only KBs require a conversion pass first; out of scope.
2. The user has authority to commission rewrites or archives. Producing a cleanup list nobody acts on is wasted work — route findings to a named owner before running the ingester.
3. Owner metadata lives in YAML frontmatter (`owner: alex@company.com`) or in a top-of-page "Owner:" line. Tribal-knowledge ownership (the person who last edited the page) is treated as missing.
4. "Stale" defaults to 12 months. Override with `--stale-days` on `kb_ingester.py`. Some compliance regimes (FDA, ISO 13485) require shorter review cycles; use `--profile regulated` and `--stale-days 365`.
5. The user is not asking for a personal PKM. Personal Karpathy-style second-brain work belongs in `engineering/llm-wiki`.
## Anti-patterns
- **Generating SOPs in bulk without owners.** A doc with no owner has a half-life of 6 months. Refuse to generate a batch of 30 SOPs unless each one is assigned to a named human.
- **Using `runbook_validator.py` as a checkbox.** The validator catches missing structure. It does not catch wrong content. A runbook can score 100 and still tell the operator the wrong thing.
- **Treating orphan pages as garbage by default.** Some orphans are reference pages found only via search — not all orphans should be archived. The cleanup list is a *priority queue*, not a delete list.
- **Confusing knowledge-ops with `process-mapper`.** Process-mapper documents the *flow* of work between stages (BPMN, cycle time, bottleneck). Knowledge-ops documents the *artifacts* operators consume to execute the work (SOP, runbook, glossary). Both can apply to the same process.
- **Letting glossary drift accumulate.** Two definitions of "CSM" in three years becomes seven definitions in five. Fix glossary drift the moment it surfaces in `kb_ingester.py` output.
- **Skipping the regulated profile under regulated workload.** If the process touches PHI, SOX-relevant financial controls, or ISO 13485 device QMS, use `--profile regulated`. Missing version control on a regulated SOP is an audit finding.
- **Hand-writing 5W2H sections from memory.** The 5W2H scaffold exists because operators forget "How-much". Use the generator; edit the output.
## Distinct from
- **`engineering/llm-wiki`** — Karpathy-style personal PKM second brain where one human ingests sources into their own interlinked vault. Knowledge-ops is *organizational*: many authors, many readers, named owners per doc, formal review cycles, compliance overlays.
- **`engineering-team/runbook-generator`** — system-ops runbook for debugging a production system (logs, alerts, k8s, on-call). Knowledge-ops runbooks are *operator* runbooks for business processes (incident-comms cascade, vendor offboarding, employee onboarding). The audience is fellow operators, not engineers tailing logs.
- **`project-management/*`** — Jira / Confluence delivery tracking, sprint ticket workflow, project-status reporting. Knowledge-ops is the *content* in those Confluence pages, not the *tracking* of who edits them.
- **`business-operations/process-mapper`** (sibling) — BPMN process *design*: where the stages are, where work waits, which stage is the bottleneck. Knowledge-ops is process *documentation*: the SOP and runbook artifacts that tell an operator how to execute the process the mapper described.
- **`business-operations/internal-comms`** (sibling) — broadcast announcements, all-hands messaging, change-management comms. Knowledge-ops is the durable reference artifact; internal-comms is the broadcast.
- **`ra-qm-team/*`** — formal regulatory compliance authoring (ISO 13485 QMS, MDR technical files, 21 CFR Part 820). Knowledge-ops borrows the regulatory checklist but is not a substitute for a notified-body audit.
## Forcing-question library (Matt Pocock grill discipline)
Before invoking the tools, the orchestrator (or `/cs:grill-bizops`) walks the user through these questions **one at a time, with a recommended answer + canon citation**. Never bundled. Walk depth-first — do not open question 4 until 1-3 are locked.
1. **"Who is the named owner of this SOP / runbook, and do they know they own it?"**
Recommended: a single human (not "the team"), and yes — they have agreed in writing.
Canon: Gawande 2009 (*The Checklist Manifesto*) — checklists without an owner rot within 12 months. Ownership is the discipline.
2. **"When was this doc last reviewed, and what is the review cadence?"**
Recommended: reviewed within the last 12 months (90 days if `--profile regulated`); cadence written in the frontmatter.
Canon: ISO 9001:2015 §7.5.3 — controlled documents require review-cycle metadata. ITIL v4 echoes this for Service Operation runbooks.
3. **"For each runbook step: what is the observable success signal — by which I mean, what specific output tells you the step worked?"**
Recommended: a concrete observable ("HTTP 200 from `/healthz`", "Slack thread closed with `done` reaction", "Salesforce opportunity moved to `Closed-Won` stage") — not "the service is up" or "it works".
Canon: Beyer et al. 2018 (*Site Reliability Workbook*, Ch. 8) — observable signals are the entire point of a runbook. Vague success criteria are the leading cause of runbook misuse during incidents.
4. **"What is the rollback path for each runbook step that can fail?"**
Recommended: every step that mutates state has either a rollback path or an explicit "cannot roll back — escalate to X" line.
Canon: AWS Well-Architected Framework, Operational Excellence pillar — "you cannot run a process you cannot reverse without first agreeing what 'reverse' means".
5. **"Where does this doc live, and what other docs link to it?"**
Recommended: in the canonical wiki, and at least 2 inbound links from related docs. An orphan SOP is an unfindable SOP.
Canon: Atlassian Team Playbook on documentation health — orphan rate > 20% is the leading indicator of a wiki sprawl problem.
6. **"What is the regulatory overlay on this process — SOC 2, HIPAA, ISO 13485, GDPR, SOX, none?"**
Recommended: explicit answer. If "none", confirm by checking the data classes the process touches.
Canon: FDA 21 CFR Part 211.100 (Written procedures; deviations) — regulated SOPs require version control, change history, and signoff. Skip this step and the doc is an audit finding.
7. **"Is the happy path the *only* path documented, or are the 2-3 most common failure modes also documented?"**
Recommended: the top-2 failure modes per process are documented with their own recovery sub-procedure.
Canon: Fowler 2016 (*Production-Ready Microservices*) — operations docs that cover only the happy path are responsible for 60%+ of incident-time waste.
After all 7 are locked, invoke `kb_ingester.py` → `runbook_validator.py` → `sop_generator.py` in sequence.
FILE:assets/runbook_template.md
# Runbook Template — fill out before running `runbook_validator.py`
Use this template to capture runbook steps before invoking the validator.
Each step must specify all six required attributes (owner, duration,
success signal, failure signal, rollback, escalation) or the validator
will flag it.
Feed the JSON into:
```
python3 scripts/runbook_validator.py --input my-runbook.json
python3 scripts/runbook_validator.py --input my-runbook.md # markdown also accepted
```
A runbook scoring < 60 is NOT-SAFE for production use. Aim for ≥ 80
(SAFE-TO-USE) before putting the runbook into rotation.
---
## Runbook metadata
- **Runbook name:** _(e.g., Incident Comms Cascade, Customer Escalation, Vendor Outage Response, System-Access Revocation)_
- **Owner:** _(named human or named on-call rotation — e.g., "Incident Commander on-call (PagerDuty: ic-primary)")_
- **Trigger:** _(what specifically invokes this runbook — e.g., "PagerDuty Sev-1 incident triggered" or "Customer escalation flagged in Salesforce")_
- **Expected total duration:** _(P50 + P90 wall-clock from trigger to completion)_
- **Linked SOP:** _(if this runbook implements an SOP, link the canonical SOP page)_
---
## Step table
| # | Step title | Owner | Duration | Success signal (observable) | Failure signal (observable) | Rollback | Escalation |
|---|------------|-------|----------|------------------------------|------------------------------|----------|------------|
| 1 | _e.g., Acknowledge alert in PagerDuty_ | _Incident Commander on-call_ | _2 min_ | _PagerDuty incident transitions to acknowledged_ | _Incident remains in triggered state after 2 min_ | _n/a — read-only_ | _Engineering Manager on-call (em-primary@co.com)_ |
| 2 | _e.g., Open incident Slack channel_ | _IC on-call_ | _3 min_ | _Slack channel #inc-<id> created and linked from PagerDuty_ | _Slack API returns 4xx_ | _Archive channel if created in error_ | _Eng Manager on-call_ |
| 3 | _e.g., Notify execs via paging tree_ | _Comms Lead (comms-lead@co.com)_ | _5 min_ | _SES API returns 200 for all exec recipients_ | _SES API returns 5xx OR delivery=bounced_ | _Send retraction email with subject prefix 'RETRACTION:'_ | _VP Communications_ |
---
## JSON skeleton
```json
{
"runbook_name": "Incident Comms Cascade",
"steps": [
{
"title": "Acknowledge alert in PagerDuty",
"owner": "Incident Commander on-call (PagerDuty: ic-primary)",
"duration_str": "2 minutes",
"duration_minutes": 2,
"success_signal": "PagerDuty incident transitions to acknowledged",
"failure_signal": "Incident remains in triggered state after 2 minutes",
"rollback": "n/a — acknowledgement is non-mutating, read-only operation",
"escalation": "Engineering Manager on-call (em-primary@company.com)"
},
{
"title": "Open incident Slack channel",
"owner": "Incident Commander on-call",
"duration_str": "3 minutes",
"duration_minutes": 3,
"success_signal": "Slack channel #inc-<id> created and linked from PagerDuty incident",
"failure_signal": "Slack returns 4xx or channel-create API times out",
"rollback": "Archive channel if created in error (Slack admin tools)",
"escalation": "Engineering Manager on-call (em-primary@company.com)"
},
{
"title": "Notify execs via paging tree",
"owner": "Communications Lead (comms-lead@company.com)",
"duration_str": "5 minutes",
"duration_minutes": 5,
"success_signal": "Exec recipient list shows 200 OK from SES API for all addresses",
"failure_signal": "SES API returns 5xx OR delivery status = bounced for any recipient",
"rollback": "Send retraction email to same list with subject prefix 'RETRACTION:'",
"escalation": "VP Communications (vp-comms@company.com)"
}
]
}
```
---
## Markdown form (alternative — runbook_validator.py heuristic parser)
If you prefer authoring in markdown directly, follow this exact structure (the parser keys off `## Step N:` headings and bullet attributes):
```markdown
# Runbook: Incident Comms Cascade
## Step 1: Acknowledge alert in PagerDuty
- **Owner:** Incident Commander on-call (PagerDuty: ic-primary)
- **Duration:** 2 minutes
- **Success:** PagerDuty incident transitions to acknowledged
- **Failure:** Incident remains in triggered state after 2 minutes
- **Rollback:** n/a — non-mutating, read-only
- **Escalation:** Engineering Manager on-call (em-primary@company.com)
## Step 2: Open incident Slack channel
- **Owner:** ...
```
---
## Authoring discipline checklist
Before submitting the runbook to the validator:
- [ ] **Every step has a named owner**, not "the team" or "ops" — required by SRE Workbook Ch. 8.
- [ ] **Every step has a concrete duration** (number + unit). "Quick" is not a duration.
- [ ] **Every success signal is observable** — a yes/no check the operator can perform. "HTTP 200 from /healthz", not "service is up".
- [ ] **Every failure signal is observable** — what tells you the step did NOT work.
- [ ] **Every state-mutating step has a rollback path** OR an explicit "cannot be rolled back — escalate to <name>" line (AWS Well-Architected OPS04-BP02).
- [ ] **Every step has an escalation contact** — named human, role+email, or named on-call rotation.
- [ ] **Top-2 failure modes documented** (Fowler 2016) — most common ways this runbook gets stuck, each with their own recovery sub-procedure.
- [ ] **Last-reviewed date set in frontmatter** — runbooks decay; Charity Majors's data: untouched 12-month-old runbooks are wrong 60% of the time.
After validation, place the runbook in the canonical wiki location and link it from at least 2 navigation hubs (incident-handbook, the parent SOP) to avoid orphan-page status.
FILE:assets/sop_template.md
# SOP Template — fill out before running `sop_generator.py`
Use this template to capture the SOP metadata before invoking the generator.
Fill in the fields below, then translate them into the JSON skeleton at the
bottom of this file. Feed that JSON into the generator:
```
python3 scripts/sop_generator.py --input my-sop.json --profile ops
python3 scripts/sop_generator.py --input my-sop.json --profile regulated # for SOX / HIPAA / ISO 13485 / FDA
```
---
## SOP metadata
- **SOP name:** _(e.g., Vendor Offboarding, Procurement Intake, Employee Onboarding, Customer Escalation, System Access Provisioning)_
- **Process owner (named human):** _(e.g., alex@company.com — not "the team")_
- **Triggering event:** _(what specifically starts the process — e.g., "Vendor contract not renewed OR vendor terminated for cause")_
- **Audience role:** _(who will execute this SOP — e.g., "Vendor Management Office operator", "HR onboarding specialist")_
- **Frequency:** _(how often this runs — "Daily", "Weekly Monday 9am", "On-demand avg 3x/quarter")_
- **Regulatory overlay:** _(zero or more of: SOC2, HIPAA, ISO13485, GDPR, SOX. If "none", confirm by listing data classes the process touches.)_
---
## Inputs and outputs
**Inputs required before starting:**
- _(input 1 — e.g., "Vendor legal name")_
- _(input 2 — e.g., "Contract end date")_
- _(input 3 — e.g., "List of systems with vendor access")_
**Outputs produced:**
- _(output 1 — e.g., "All production system access revoked, evidenced in IAM audit log")_
- _(output 2 — e.g., "Vendor data deletion certified")_
- _(output 3 — e.g., "Final invoice reconciled and paid")_
---
## Steps outline
Six rows to start; add or remove. **Each step must be a noun-phrase action**, not a paragraph.
| # | Step name (action) | Notes |
|---|--------------------|-------|
| 1 | _e.g., Notify vendor of offboarding intent (30 days written notice)_ | |
| 2 | _e.g., Inventory data classes and system access vendor holds_ | |
| 3 | _e.g., Revoke production system access (IAM, VPN, SaaS)_ | |
| 4 | _e.g., Confirm data deletion (vendor certification) or data return_ | |
| 5 | _e.g., Final invoice reconciliation and payment_ | |
| 6 | _e.g., Archive vendor record in VMO registry with offboarding evidence_ | |
---
## How-much (cost model)
- **Estimated execution time:** _(minutes per execution — e.g., 240)_
- **Estimated cost per execution:** _(USD, labor + license + third-party fees — e.g., 800)_
---
## JSON skeleton
```json
{
"sop_name": "Vendor Offboarding",
"process_owner": "alex@company.com (Vendor Management Lead)",
"triggering_event": "Vendor contract not renewed OR vendor terminated for cause",
"audience_role": "Vendor Management Office (VMO) operator",
"frequency": "On-demand (avg 3 executions per quarter)",
"regulatory_overlay": ["SOC2"],
"inputs": [
"Vendor legal name",
"Contract end date",
"List of systems with vendor access",
"List of data classes vendor processed"
],
"outputs": [
"All production system access revoked (evidenced)",
"Vendor data deleted or returned (evidenced)",
"Final invoice reconciled and paid",
"Vendor record archived in VMO registry with offboarding evidence"
],
"steps_outline": [
"Notify vendor of offboarding intent (written, 30 days notice)",
"Inventory data classes and system access vendor holds",
"Revoke production system access (IAM, VPN, SaaS)",
"Confirm data deletion (vendor certification) or data return",
"Final invoice reconciliation and payment",
"Archive vendor record in VMO registry with offboarding evidence"
],
"estimated_minutes": 240,
"estimated_cost_usd": 800
}
```
---
## Authoring discipline checklist
Before submitting the JSON to the generator, confirm:
- [ ] **Owner is a named human**, not "the team" — required by Gawande *Checklist Manifesto* discipline.
- [ ] **Triggering event is specific.** "When needed" is not a trigger.
- [ ] **At least one regulatory overlay considered** (or explicit "none after checking PHI/financial/regulated-device classes").
- [ ] **Top-2 failure modes documented** — happy-path-only SOPs are responsible for 60%+ of incident-time waste (Fowler 2016).
- [ ] **"How-much" is filled in.** It's the section authors most often forget — and the section operators most need.
- [ ] **`--profile regulated` selected** if SOP touches SOX, HIPAA, ISO 13485, FDA 21 CFR Part 211, or SOC 2 controls.
After generation, run the runbook validator on any embedded step lists that include state-mutating operations:
```
python3 scripts/runbook_validator.py --input generated-sop.md
```
FILE:references/5w2h_sop_canon.md
# 5W2H SOP Canon
Standard Operating Procedure (SOP) authoring discipline for company processes — what every SOP must contain, why, and where the discipline comes from. Eight authoritative sources cited.
## What 5W2H is
5W2H is a structured checklist for documenting *any* repeatable process by answering seven questions:
| Letter | Question | Section in `sop_generator.py` output |
|---|---|---|
| Who | Who is responsible, accountable, consulted, informed? | RACI |
| What | What is the process — inputs, outputs, scope? | Process spec |
| When | When does it run — trigger, frequency, blocking deps? | Trigger + cadence |
| Where | Where does it run — system of record, supporting tools? | System map |
| Why | Why does it exist — business purpose, regulatory basis? | Purpose + compliance |
| How | How is it executed — step-by-step procedure? | Procedure |
| How-much | How much does it cost — time, money per execution? | Cost model |
Two SOPs covering the same process can be wildly different in length and quality. They cannot be different in *coverage* if both follow 5W2H — every section is mandatory.
## Why 5W2H specifically
Three properties make 5W2H the right scaffold for an ops org:
1. **Audit-friendly.** ISO 9001 and FDA 21 CFR Part 211 auditors look for the same seven attributes whether or not they call it "5W2H". Adopting the scaffold up front means SOPs ship audit-ready.
2. **Operator-friendly.** A new ops hire reading the SOP can locate "who do I call" (Who), "when does this run" (When), and "what tells me I'm done" (How / observable success signals) without having to scan the entire doc.
3. **Author-friendly.** Empty 5W2H sections are visually obvious. "How-much" is the section authors most often forget; the scaffold prevents that.
## Eight authoritative sources
### 1. Kaoru Ishikawa — *Guide to Quality Control* (1985, Asian Productivity Organization)
Origin of the 5W1H quality-control method. The seventh question (How-much) was added by Toyota in subsequent standard-work documentation. Ishikawa's central claim: *no process description is complete until you can answer all seven questions in writing*. Anything less is tribal knowledge.
### 2. Jeffrey Liker — *The Toyota Way* (2003, McGraw-Hill)
Chapter 6 on standard work codifies the Toyota convention that every SOP documents (a) takt time, (b) work sequence, (c) standard inventory. The "How-much" anchor maps directly to takt time. Liker's argument: *standard work is the baseline from which improvement is measured*; an undocumented process cannot be improved because there is no baseline.
### 3. Atul Gawande — *The Checklist Manifesto* (2009, Metropolitan Books)
Gawande's hospital surgical-checklist research found that simple, well-owned checklists reduced surgical mortality by 47% in a 2008 WHO study across eight hospitals on four continents. Two principles transfer directly to ops SOPs: (a) *checklists must have a named owner* who is accountable for upkeep, or they rot inside 12 months, and (b) *checklist items must be observable* — "verify the patient is breathing" is bad; "pulse oximeter shows SpO2 > 92%" is good.
### 4. Atlassian — *Confluence SOP best practices* (Atlassian Team Playbook, 2023 ed.)
Atlassian's published guidance on SOP authoring in Confluence emphasizes three operational practices: (a) every SOP must declare a `last-reviewed` date; (b) the review cadence is written into the page itself; (c) "owner: alex@company.com" goes in YAML frontmatter so tooling can find SOPs with no owner. The KB hygiene anti-patterns reference draws from the same source.
### 5. ISO 9001:2015 — *Quality management systems — Requirements*
Clause 7.5.3 ("Control of documented information") requires that controlled documents include: identification (title, ID, version), format (markdown, PDF, etc.), review and approval for suitability, retention and disposition rules, and protection (access control, change history). The `regulated` profile in `sop_generator.py` adds these sections explicitly.
### 6. ITIL v4 — *Service Operation* practice guide (Axelos, 2019)
ITIL's distinction between *procedures* (the SOP — repeatable and largely unchanged) and *work instructions* (the runbook — the specific commands and observable signals at execution time) is the same distinction this skill makes. Both artifacts coexist. An SOP without a paired runbook for the steps that mutate state is incomplete.
### 7. FDA 21 CFR Part 211.100 — *Written procedures; deviations*
For pharmaceutical and medical-device-adjacent companies, Part 211.100 makes SOPs legally required. Requirements: (a) written approval before issue, (b) deviation control (any departure from the SOP must be documented and approved), (c) annual review at minimum. The `--profile regulated` flag attaches these requirements.
### 8. Project Management Institute — *PMBOK Guide* (7th ed., 2021)
PMBOK §4 on integration management defines SOP-equivalent artifacts as "organizational process assets" and requires named accountability. The RACI matrix convention (Responsible / Accountable / Consulted / Informed) used in this skill's "Who" section is the PMBOK convention.
## Anti-pattern: prose-only SOPs
A 1500-word prose SOP without the 5W2H scaffolding looks thorough and is usually missing 2-3 mandatory sections (most commonly: How-much, Why-regulatory, observable success signals). Use the generator. Edit its output. Do not write SOPs from a blank page.
## How this skill applies the canon
- `sop_generator.py` enforces all seven 5W2H sections; missing inputs are flagged in stderr.
- `--profile regulated` attaches ISO 9001 §7.5.3 + FDA Part 211 metadata (version, signoff, change history).
- Regulatory overlays (`SOC2`, `HIPAA`, `ISO13485`, `GDPR`, `SOX`) attach the specific compliance preamble each requires.
- The forcing-question library in `SKILL.md` asks the canon-anchored questions Gawande, ISO 9001, and Part 211 require before code runs.
FILE:references/kb_hygiene_anti_patterns.md
# Knowledge-Base Hygiene Anti-Patterns
The recurring failure modes that turn a useful company wiki into a sprawl of stale, unfindable, contradictory docs. Eight anti-patterns, each anchored to authoritative sources. Seven citations.
## The pattern
An ops org's wiki passes through three predictable phases:
1. **Year 1:** 50 pages, all owned, all current, everyone finds what they need.
2. **Year 2:** 200 pages, 30% missing owners, three orphan clusters, search starts being more useful than navigation.
3. **Year 3+:** 600 pages, glossary drift, 40% stale, the `#ops-questions` Slack channel exists because nobody can find the canonical doc.
`kb_ingester.py` exists to put numbers on this decay and rank what to fix first. The anti-patterns below explain *what to fix*.
## 1. No owner per SOP
**Symptom:** YAML frontmatter has no `owner:` field, or the SOP body says "owned by the Ops team".
**Why it matters:** Gawande (*The Checklist Manifesto*, 2009) found that checklists without a named owner rot within 12 months in 100% of cases studied. Ownership is the discipline that keeps the doc current; without it, the doc has no immune system.
**Detection:** `kb_ingester.py` reports `missing_owner_count`. Goal: 0.
**Fix:** Assign every SOP to a single named human in YAML frontmatter. "The team" is not an owner.
**Citation:** Gawande 2009 (*The Checklist Manifesto*, Metropolitan Books).
---
## 2. No last-reviewed date
**Symptom:** The SOP has no `last_reviewed:` field. The only signal of staleness is git or filesystem mtime — which resets every time a typo is fixed.
**Why it matters:** ISO 9001:2015 §7.5.3 explicitly requires review cycles for controlled documents. Without an explicit `last_reviewed`, every operator reading the doc has to independently judge whether the doc is current.
**Detection:** `kb_ingester.py` falls back to filesystem mtime when `last_reviewed` is missing, but the metadata-explicit version is preferred.
**Fix:** Add `last_reviewed: YYYY-MM-DD` to every SOP frontmatter. Pair with a review cadence (12 months default, 90 days for regulated).
**Citation:** ISO 9001:2015 §7.5.3 ("Control of documented information").
---
## 3. Step says "verify the service is up" (vague success signal)
**Symptom:** Runbook step success criteria are not observable. "Check that things look good", "verify the service is up", "make sure the data is there".
**Why it matters:** Beyer et al. (*Site Reliability Workbook*, 2018, Ch. 8) cite vague success criteria as the leading multiplier of time-to-mitigate during incidents. A new operator at 3am cannot tell what "up" means.
**Detection:** `runbook_validator.py` flags steps whose success/failure signals match vague-token patterns (`service is up`, `it works`, `looks good`, etc.).
**Fix:** Rewrite success signals as observable checks. "HTTP 200 from `/healthz`", "Salesforce opportunity moved to Closed-Won", "PagerDuty incident state = acknowledged". Anything that returns a yes/no.
**Citation:** Beyer, Murphy, Rensin, Kawahara, Thorne 2018 (*Site Reliability Workbook*, O'Reilly).
---
## 4. Runbook with no rollback
**Symptom:** The runbook tells the operator how to send the alert. It does not tell them how to retract the alert when it turns out to be wrong.
**Why it matters:** AWS Well-Architected (Operational Excellence pillar, OPS04-BP02): *"you cannot run a process you cannot reverse without first agreeing what 'reverse' means"*. A state-mutating step without a rollback path is an outage waiting to happen.
**Detection:** `runbook_validator.py` enforces a rollback field per step. Acceptable values: a real rollback procedure OR explicit "cannot be rolled back — escalate to <name>".
**Fix:** For every state-mutating step, write the rollback. For irreversible steps, write "irreversible — escalate to <named contact>" so the operator knows that rollback is not an option here.
**Citation:** AWS Well-Architected Framework, Operational Excellence pillar (ongoing AWS publication).
---
## 5. Wiki sprawl across 4 tools
**Symptom:** SOPs live in Notion. Runbooks live in Confluence. Onboarding lives in a Google Doc folder. The glossary lives in a Slack canvas. Nobody knows which is canonical.
**Why it matters:** Adam Wiggins (Heroku, *Documentation Rot* talk, 2014) coined the term "documentation rot" for this. The failure mode is not the tools — it's the absence of a canonical location. Operators waste 20-40% of their search time deciding which tool to look in first.
**Detection:** Out of scope for `kb_ingester.py` (which runs on one markdown tree). The signal is human: "where's the X SOP?" gets three different answers.
**Fix:** Pick one canonical tool. Migrate the rest. Treat the others as archives, link the canonical from the others. Mozilla SUMO's KB consolidation (2016) is the template.
**Citation:** Wiggins 2014 (Heroku Engineering talk, "Documentation Rot"). Cited again in MIT TIK 2020 org-wiki research.
---
## 6. Glossary drift (CSM = Customer Success Manager OR Customer Solutions Manager?)
**Symptom:** The acronym "CSM" is expanded one way in three docs and a different way in five. New hires guess wrong for six months. Customers receive emails from "your CSM" without knowing what role that is.
**Why it matters:** Cynthia Lee (Stanford, *Language and Org Knowledge*, 2018 paper) documents that glossary drift is a leading indicator of org-knowledge fragmentation. Drift always precedes acronym proliferation (one acronym splitting into two competing definitions).
**Detection:** `kb_ingester.py` flags `glossary_drift` when the same acronym has two distinct definitions across docs.
**Fix:** Pick one canonical definition per acronym. Add a `glossary.md` page. Link every other doc to it. Refuse to expand the acronym anywhere else.
**Citation:** Lee 2018 (Stanford research on org-knowledge fragmentation).
---
## 7. Orphan pages nobody can find
**Symptom:** 30-60% of pages have no inbound links. They exist because somebody knew the URL. Search finds them; navigation does not.
**Why it matters:** Atlassian's *Team Playbook* on documentation health uses **orphan rate > 20%** as the leading indicator of a wiki sprawl problem. Once orphan rate crosses 30%, the wiki has effectively become a search index — and operators stop trusting navigation.
**Detection:** `kb_ingester.py` reports `orphan_count` and lists orphans.
**Fix:** Not "delete all orphans". Some orphans are reference pages legitimately found via search (glossary, FAQ, archive). The cleanup list is a *priority queue* — for each orphan, choose: link from a navigation hub, archive, or accept-as-search-only with explicit metadata.
**Citation:** Atlassian Team Playbook, "Documentation Health" play (2021).
---
## 8. SOPs that document the happy path only
**Symptom:** The vendor-offboarding SOP covers what happens when the vendor cooperates. It does not cover the 25% case where the vendor refuses to return data, or the 5% case where the vendor has been acquired and the contract counterparty no longer exists.
**Why it matters:** Susan Fowler (*Production-Ready Microservices*, 2016, Ch. 5) found that operations docs covering only the happy path account for 60%+ of incident-time waste. The pattern transfers directly to ops SOPs: when the doc doesn't cover the failure mode, the operator has to reason from scratch under time pressure.
**Detection:** Manual — `runbook_validator.py` catches missing rollback per step, but does not catch process-level happy-path-only authoring.
**Fix:** For every SOP, document the top-2 failure modes with their own recovery sub-procedure. The forcing-question library in `SKILL.md` (question 7) enforces this.
**Citation:** Fowler 2016 (*Production-Ready Microservices*, O'Reilly).
---
## 9. Compliance SOPs without version control
**Symptom:** A SOX-relevant or HIPAA-relevant SOP has no change history, no signoff record, no version field. An auditor asks "what was the procedure in Q2?" — nobody can answer.
**Why it matters:** FDA 21 CFR Part 211.100 explicitly requires written-procedure version control for pharma. ISO 9001 §7.5.3 imposes the same for any controlled document. Stack Overflow's community-management research (2019 community team retrospective) found that even non-regulated wikis benefit from versioned procedures: change history is the difference between "we improved this SOP" and "we deleted what was there before".
**Detection:** `--profile regulated` in `sop_generator.py` attaches the version + signoff + change-history sections. Missing those sections under a regulated overlay is the audit finding.
**Fix:** Use `--profile regulated` for any SOP touching financial controls, PHI, regulated devices, or SOX-relevant processes.
**Citations:** FDA 21 CFR Part 211.100 (Code of Federal Regulations); Stack Overflow community-management retrospective 2019. Mozilla SUMO KB lessons (2016) echo both.
---
## How this skill applies the anti-patterns
- `kb_ingester.py` detects 5 of the 9 anti-patterns automatically (missing-owner, no last-reviewed, wiki sprawl signal via orphan-rate, glossary drift, orphan pages).
- `runbook_validator.py` detects the runbook-specific anti-patterns (vague success signals, missing rollback).
- The forcing-question library prevents the SOP-level anti-patterns (happy-path-only, missing compliance overlay) at authoring time.
- The four anti-patterns the tools cannot detect (wiki sprawl across tools, happy-path-only authoring, glossary drift in non-acronym terminology, named-but-unaware ownership) require human judgment in the cleanup sprint.
The skill's job is to surface the 80% of anti-patterns a tool can find. The remaining 20% is the cleanup-sprint discussion.
FILE:references/runbook_canon.md
# Runbook Canon
Internal-operations runbook design discipline — what makes a runbook safe to execute at 3am during an incident, and where the discipline comes from. Seven authoritative sources cited.
## What a runbook is (and is not)
A **runbook** is the executable artifact an operator follows under time pressure. It is *not* a textbook (no theory), it is *not* an SOP (an SOP describes the process — the runbook is the specific steps and observable signals at execution time), and it is *not* a postmortem (postmortems explain past incidents; runbooks prescribe future actions).
Every runbook step must specify six things — and `runbook_validator.py` enforces all six:
1. **Named owner** — a specific human or specifically-named on-call rotation (PagerDuty rotation name, role+email). Not "the team", not "ops".
2. **Expected duration** — concrete number + unit. "5 minutes", "30 seconds". Not "quick" or "fast".
3. **Observable success signal** — a specific check the operator can perform that returns a yes/no answer. "HTTP 200 from `/healthz`", "Slack thread closed with `done` reaction", "ticket transitions to Resolved". Not "service is up", not "looks good".
4. **Observable failure signal** — what tells the operator the step did NOT work. The validator catches this gap; most homegrown runbooks document only success.
5. **Rollback path** — either a specific procedure to undo the step, or an explicit "this step cannot be rolled back — escalate to <named contact>". Silent absence of rollback is the most dangerous gap.
6. **Escalation contact** — named human, role+email, or named on-call rotation. Not "engineering", not "ops".
## Why these six attributes specifically
These six are the union of the requirements imposed by the seven sources below. Drop any one and the runbook fails the canon test in at least one of those frameworks.
## Seven authoritative sources
### 1. Beyer, Murphy, Rensin, Kawahara, Thorne (eds.) — *The Site Reliability Workbook* (O'Reilly, 2018), Ch. 8
Google SRE Workbook on "On-Call". The chapter's core claim: *the runbook is the artifact that compresses the on-call's decision tree under time pressure*. Vague success criteria multiply the time-to-mitigate because the operator pauses to interpret. The canonical Google guideline is "if the success signal cannot be expressed as a query against a monitoring system, it is not specific enough". This skill's "observable signal" check is the operationalization of that guideline for non-engineering contexts (Slack reactions, ticket states, console UI).
### 2. Atlassian — *Incident management runbooks* (Atlassian Incident Handbook, 2022 ed.)
Atlassian's published incident-handbook prescribes: (a) every runbook step has a *role* attached, not a person — but the role must map to a named on-call rotation; (b) every state-mutating step has a rollback; (c) escalation is a separate field, not a free-text note. This skill's `--profile support` variant of `sop_generator.py` follows Atlassian's escalation-matrix convention.
### 3. PagerDuty — *Incident Response Documentation* (PagerDuty open-source, 2017 onwards)
PagerDuty's open-source incident-response framework distinguishes between **major-incident runbooks** (the comms cascade — who's notified, in what order, with what SLA) and **technical-recovery runbooks** (the engineering steps to mitigate). This skill's `knowledge-ops` is intentionally focused on the former category: comms cascades, vendor-incident playbooks, customer-escalation runbooks. Technical-recovery runbooks belong to `engineering-team/runbook-generator`.
### 4. AWS — *Well-Architected Framework, Operational Excellence pillar* (AWS, ongoing)
AWS's Operational Excellence pillar makes the canonical argument for rollback discipline: *"you cannot run a process you cannot reverse without first agreeing what 'reverse' means"*. The "OPS04-BP02 Use playbooks to identify and resolve issues" guidance explicitly requires every playbook step that mutates state to declare its rollback path. The `runbook_validator.py` `ROLLBACK` check enforces this.
### 5. Charity Majors — *Observability Engineering* (O'Reilly, 2022, co-authored with George Miranda and Liz Fong-Jones)
Majors' argument that **runbooks decay faster than the systems they describe** is the canonical justification for `kb_ingester.py`'s stale-page detection. Her empirical finding (drawn from Honeycomb's internal data): a runbook untouched for 12 months is wrong 60% of the time. The default `--stale-days 365` setting in `kb_ingester.py` is calibrated to this.
### 6. Susan Fowler — *Production-Ready Microservices* (O'Reilly, 2016)
Fowler's Ch. 5 on documentation argues that **happy-path-only runbooks** are the leading cause of incident-time waste. Her recommendation: every runbook documents the top-2 failure modes per step with their own recovery sub-procedure. The forcing-question library in `SKILL.md` enforces this at the question-7 stage.
### 7. ITIL v4 — *Service Operation* practice guide (Axelos, 2019)
ITIL v4 makes the formal distinction between *procedure* (the SOP) and *work instruction* (the runbook): the procedure describes what is to be done at a process level; the work instruction describes how to do it at the step level. Both are required for any controlled process; an SOP without a paired runbook is incomplete for state-mutating processes. This is why `knowledge-ops` ships both `sop_generator.py` and `runbook_validator.py` — the same KB needs both artifact types.
## Common runbook anti-patterns
- **"The team owns it"** — no it doesn't. Name a human or an explicitly-defined on-call rotation.
- **"Verify the service is up"** — what does "up" mean to a new operator at 3am? Specify the observable check.
- **"Rollback: see runbook X"** — and runbook X says "see runbook Y". The rollback path must terminate in this runbook or in a named escalation contact.
- **"Escalation: engineering"** — which person, which rotation, what SLA? Engineering is 200 people.
- **Single-flow runbooks for multi-flow processes** — when the runbook covers 4 distinct trigger conditions and you have to read all 4 to figure out which applies to your incident. Split it.
- **Runbooks last reviewed before the system was rearchitected.** The stale check catches these.
## How this skill applies the canon
- `runbook_validator.py` enforces all six attributes per step.
- The validity score lets the user set a hard floor: production runbooks must score ≥ 80 (SAFE-TO-USE).
- `kb_ingester.py` flags stale runbooks (default 12 months) per Majors's decay finding.
- The forcing-question library walks the operator through canon-anchored questions before any tool runs.
FILE:scripts/kb_ingester.py
#!/usr/bin/env python3
"""kb_ingester.py
Walk a directory of markdown files (Notion export, Confluence space export,
Obsidian vault, Drive/SOPs/ directory) and emit a KB health report.
Extracts:
- cross-link map (which page references which)
- orphan pages (no inbound links)
- glossary candidates (frequently-used proper nouns / acronyms recurring
in 3+ docs with no single canonical definition page)
- glossary drift (same term used inconsistently across docs)
- stale pages (no edit in > N months — N defaults to 12)
- missing-owner pages (no `owner:` in YAML frontmatter)
- prioritized cleanup list ranked by (staleness × inbound-link-count)
Stdlib only.
"""
from __future__ import annotations
import argparse
import datetime as dt
import json
import re
import sys
import tempfile
from collections import Counter, defaultdict
from dataclasses import dataclass, field
from pathlib import Path
YAML_FRONTMATTER_RE = re.compile(
r"^---\s*\n(.*?)\n---\s*\n", re.DOTALL)
MD_LINK_RE = re.compile(r"\[([^\]]+)\]\(([^)]+)\)")
WIKI_LINK_RE = re.compile(r"\[\[([^\]|]+)(?:\|[^\]]+)?\]\]")
ACRONYM_RE = re.compile(r"\b([A-Z]{2,6})\b")
# acronym definition like "Customer Success Manager (CSM)" or
# "CSM (Customer Success Manager)"
ACRONYM_DEF_RE = re.compile(
r"\b((?:[A-Z][A-Za-z]+\s+){1,4}[A-Z][A-Za-z]+)\s*\(([A-Z]{2,6})\)"
r"|\b([A-Z]{2,6})\s*\(((?:[A-Z][A-Za-z]+\s+){1,4}[A-Z][A-Za-z]+)\)"
)
@dataclass
class PageInfo:
path: Path
title: str = ""
owner: str = ""
last_reviewed: str = ""
mtime_days_ago: int = 0
outbound_links: list = field(default_factory=list)
inbound_link_count: int = 0
acronyms_used: list = field(default_factory=list)
acronym_definitions: dict = field(default_factory=dict)
word_count: int = 0
def _parse_frontmatter(text: str) -> dict:
m = YAML_FRONTMATTER_RE.match(text)
if not m:
return {}
body = m.group(1)
fm = {}
for line in body.splitlines():
if ":" in line:
k, _, v = line.partition(":")
fm[k.strip().lower()] = v.strip().strip('"').strip("'")
return fm
def _extract_title(text: str, path: Path) -> str:
for line in text.splitlines():
m = re.match(r"^#\s+(.+)$", line)
if m:
return m.group(1).strip()
return path.stem.replace("-", " ").replace("_", " ").title()
def _extract_links(text: str) -> list:
links = []
for m in MD_LINK_RE.finditer(text):
target = m.group(2).strip()
if target.startswith(("http://", "https://", "mailto:")):
continue
links.append(target)
for m in WIKI_LINK_RE.finditer(text):
links.append(m.group(1).strip())
return links
def _extract_acronyms(text: str) -> tuple:
acronyms = ACRONYM_RE.findall(text)
defs = {}
for m in ACRONYM_DEF_RE.finditer(text):
if m.group(1) and m.group(2):
defs[m.group(2)] = m.group(1).strip()
elif m.group(3) and m.group(4):
defs[m.group(3)] = m.group(4).strip()
return acronyms, defs
def _normalize_link_target(target: str, source: Path, root: Path) -> str:
"""Resolve a link target to a canonical relative path string."""
target = target.split("#")[0].split("?")[0].strip()
if not target:
return ""
if target.endswith(".md"):
candidate = (source.parent / target).resolve()
elif "/" in target or "\\" in target:
candidate_md = (source.parent / (target + ".md")).resolve()
if candidate_md.exists():
candidate = candidate_md
else:
candidate = (source.parent / target).resolve()
else:
# bare title — try to match against any .md filename
candidate_md = (source.parent / (target + ".md")).resolve()
candidate = candidate_md
try:
return str(candidate.relative_to(root))
except ValueError:
return str(candidate)
def walk_vault(root: Path, stale_days: int = 365) -> list:
"""Walk a directory tree and return a list of PageInfo objects."""
pages = []
now = dt.datetime.now()
for path in sorted(root.rglob("*.md")):
if not path.is_file():
continue
try:
text = path.read_text(encoding="utf-8")
except (UnicodeDecodeError, OSError):
continue
fm = _parse_frontmatter(text)
title = fm.get("title") or _extract_title(text, path)
owner = fm.get("owner", "")
last_reviewed = fm.get("last_reviewed", "") or fm.get(
"last-reviewed", "")
# mtime fallback
try:
mtime = dt.datetime.fromtimestamp(path.stat().st_mtime)
mtime_days_ago = (now - mtime).days
except OSError:
mtime_days_ago = 0
outbound = _extract_links(text)
acronyms, defs = _extract_acronyms(text)
word_count = len(text.split())
pages.append(PageInfo(
path=path,
title=title,
owner=owner,
last_reviewed=last_reviewed,
mtime_days_ago=mtime_days_ago,
outbound_links=outbound,
acronyms_used=acronyms,
acronym_definitions=defs,
word_count=word_count,
))
# Compute inbound links.
by_relpath = {str(p.path.relative_to(root)): p for p in pages}
by_title = {p.title.lower(): p for p in pages}
by_stem = {p.path.stem.lower(): p for p in pages}
for src in pages:
for raw in src.outbound_links:
target_rel = _normalize_link_target(raw, src.path, root)
if target_rel in by_relpath:
by_relpath[target_rel].inbound_link_count += 1
continue
tgt = raw.split("#")[0].split("?")[0].strip().lower()
if tgt.endswith(".md"):
tgt = tgt[:-3]
if tgt in by_title:
by_title[tgt].inbound_link_count += 1
elif tgt in by_stem:
by_stem[tgt].inbound_link_count += 1
return pages
def detect_orphans(pages: list) -> list:
return [p for p in pages if p.inbound_link_count == 0]
def detect_stale(pages: list, stale_days: int) -> list:
out = []
for p in pages:
is_stale = False
if p.last_reviewed:
try:
lr = dt.datetime.strptime(p.last_reviewed[:10], "%Y-%m-%d")
if (dt.datetime.now() - lr).days > stale_days:
is_stale = True
except ValueError:
pass
elif p.mtime_days_ago > stale_days:
is_stale = True
if is_stale:
out.append(p)
return out
def detect_missing_owner(pages: list) -> list:
return [p for p in pages if not p.owner]
def detect_glossary_drift(pages: list) -> dict:
"""Return a dict {acronym: [list of (definition, source page)]} for
acronyms that have >= 2 distinct definitions across the vault."""
by_acronym = defaultdict(list)
for p in pages:
for ac, defin in p.acronym_definitions.items():
by_acronym[ac].append((defin, str(p.path)))
drift = {}
for ac, defs in by_acronym.items():
distinct = set(d.lower() for d, _ in defs)
if len(distinct) >= 2:
drift[ac] = defs
return drift
def detect_glossary_candidates(pages: list, min_docs: int = 3) -> list:
"""Acronyms used in >= min_docs pages with no canonical definition
page (no page where the acronym appears in the title)."""
doc_count = Counter()
titled = set()
for p in pages:
seen = set(p.acronyms_used)
for ac in seen:
doc_count[ac] += 1
for ac in p.acronym_definitions:
# If acronym appears in title, treat as canonical-ish.
if ac in p.title:
titled.add(ac)
return sorted([(ac, c) for ac, c in doc_count.items()
if c >= min_docs and ac not in titled],
key=lambda x: -x[1])
def cleanup_priority(pages: list, stale_days: int) -> list:
"""Rank pages by (staleness × inbound-link-count) — high-traffic
stale docs surface first."""
scored = []
for p in pages:
staleness = 0
if p.last_reviewed:
try:
lr = dt.datetime.strptime(p.last_reviewed[:10], "%Y-%m-%d")
staleness = max(0, (dt.datetime.now() - lr).days
- stale_days)
except ValueError:
staleness = max(0, p.mtime_days_ago - stale_days)
else:
staleness = max(0, p.mtime_days_ago - stale_days)
if staleness > 0:
# inbound +1 to avoid zeroing out everything orphan
score = staleness * (p.inbound_link_count + 1)
scored.append((score, p))
scored.sort(key=lambda x: -x[0])
return scored
def generate_report(root: Path, pages: list, stale_days: int) -> str:
orphans = detect_orphans(pages)
stale = detect_stale(pages, stale_days)
missing_owner = detect_missing_owner(pages)
drift = detect_glossary_drift(pages)
candidates = detect_glossary_candidates(pages)
priority = cleanup_priority(pages, stale_days)
lines = [
f"# KB health report — `{root}`",
"",
f"**Pages scanned:** {len(pages)}",
f"**Stale threshold:** {stale_days} days",
"",
"## Summary metrics",
"",
"| Metric | Count | % of vault |",
"|--------|-------|------------|",
f"| Orphan pages (no inbound links) | {len(orphans)} | "
f"{round(len(orphans) / max(len(pages), 1) * 100, 1)}% |",
f"| Stale pages (> {stale_days}d) | {len(stale)} | "
f"{round(len(stale) / max(len(pages), 1) * 100, 1)}% |",
f"| Missing-owner pages | {len(missing_owner)} | "
f"{round(len(missing_owner) / max(len(pages), 1) * 100, 1)}% |",
f"| Glossary drift (acronyms with >= 2 defs) | {len(drift)} | — |",
f"| Glossary candidates (acronyms in 3+ docs, no canonical page) "
f"| {len(candidates)} | — |",
"",
]
lines.append("## Top-20 cleanup priority "
"(staleness × inbound-link-count + 1)")
lines.append("")
if priority:
lines.append("| Rank | Score | Path | Inbound | "
"Days stale | Owner |")
lines.append("|------|-------|------|---------|"
"------------|-------|")
for i, (score, p) in enumerate(priority[:20], start=1):
rel = p.path.relative_to(root)
staleness = (p.mtime_days_ago - stale_days
if not p.last_reviewed else
(dt.datetime.now() - dt.datetime.strptime(
p.last_reviewed[:10], "%Y-%m-%d")).days
- stale_days)
lines.append(
f"| {i} | {score} | `{rel}` | {p.inbound_link_count} "
f"| {staleness} | {p.owner or '(MISSING)'} |"
)
else:
lines.append("_(no stale pages — KB is current)_")
lines.append("")
lines.append("## Orphan pages (no inbound links)")
lines.append("")
if orphans:
for p in orphans[:30]:
rel = p.path.relative_to(root)
lines.append(f"- `{rel}` — {p.title}")
if len(orphans) > 30:
lines.append(f"- _(+{len(orphans) - 30} more not shown)_")
else:
lines.append("_(none — every page has at least one inbound link)_")
lines.append("")
lines.append("## Glossary drift (acronym defined differently across "
"docs)")
lines.append("")
if drift:
for ac, defs in drift.items():
lines.append(f"**{ac}:**")
for defin, src in defs:
lines.append(f" - `{defin}` (in `{src}`)")
lines.append("")
else:
lines.append("_(none detected — acronyms are used consistently)_")
lines.append("")
lines.append("## Glossary candidates (acronym used in 3+ docs "
"without a canonical definition page)")
lines.append("")
if candidates:
for ac, count in candidates[:20]:
lines.append(f"- **{ac}** — used in {count} docs, no "
f"canonical definition page exists")
else:
lines.append("_(none — acronyms either have canonical pages or "
"are uncommon)_")
lines.append("")
lines.append("## Missing-owner pages")
lines.append("")
if missing_owner:
for p in missing_owner[:30]:
rel = p.path.relative_to(root)
lines.append(f"- `{rel}` — {p.title}")
if len(missing_owner) > 30:
lines.append(
f"- _(+{len(missing_owner) - 30} more not shown)_")
else:
lines.append("_(none — every page has an owner)_")
lines.append("")
lines.append("## Recommended next actions")
lines.append("")
lines.append("1. Assign owners to the missing-owner pages first — "
"no other fix sticks without ownership.")
lines.append("2. Resolve glossary drift by picking one canonical "
"definition per acronym; add a `glossary.md` page; "
"link every other doc to it.")
lines.append("3. Triage the top-20 cleanup list: archive, rewrite, "
"or refresh. Re-run this report after the sprint to "
"verify orphan + stale counts are down.")
lines.append("4. Pair orphan pages with a navigation review — some "
"orphans are reference pages found via search and "
"should NOT be archived. Curate, don't bulk-delete.")
return "\n".join(lines) + "\n"
def generate_json_report(root: Path, pages: list, stale_days: int) -> dict:
orphans = detect_orphans(pages)
stale = detect_stale(pages, stale_days)
missing_owner = detect_missing_owner(pages)
drift = detect_glossary_drift(pages)
candidates = detect_glossary_candidates(pages)
priority = cleanup_priority(pages, stale_days)
return {
"root": str(root),
"page_count": len(pages),
"stale_days_threshold": stale_days,
"orphan_count": len(orphans),
"stale_count": len(stale),
"missing_owner_count": len(missing_owner),
"glossary_drift_count": len(drift),
"glossary_candidate_count": len(candidates),
"top_cleanup": [
{
"rank": i + 1,
"score": score,
"path": str(p.path.relative_to(root)),
"inbound_links": p.inbound_link_count,
"owner": p.owner or None,
}
for i, (score, p) in enumerate(priority[:20])
],
"orphans": [str(p.path.relative_to(root)) for p in orphans],
"glossary_drift": {ac: [{"definition": d, "source": s}
for d, s in defs]
for ac, defs in drift.items()},
"glossary_candidates": [{"acronym": ac, "doc_count": c}
for ac, c in candidates],
"missing_owner": [str(p.path.relative_to(root))
for p in missing_owner],
}
SAMPLE_PAGES = {
"index.md": """---
owner: alex@company.com
last_reviewed: 2026-04-01
---
# Ops Index
Welcome to the Ops wiki. Start with [Vendor Offboarding](sops/vendor-offboarding.md) or [Incident Comms](runbooks/incident-comms.md).
The [Glossary](glossary.md) defines our terms.
""",
"glossary.md": """---
owner: alex@company.com
last_reviewed: 2026-04-15
---
# Glossary
- Customer Success Manager (CSM) — owns post-sale account relationship.
- Vendor Management Office (VMO) — owns third-party vendor lifecycle.
""",
"sops/vendor-offboarding.md": """---
owner: jordan@company.com
last_reviewed: 2026-02-01
---
# Vendor Offboarding SOP
The VMO operator runs this SOP when a vendor contract is terminated.
See also [Incident Comms](../runbooks/incident-comms.md).
The CSM is notified.
""",
"sops/procurement-intake.md": """---
owner: jordan@company.com
last_reviewed: 2024-01-01
---
# Procurement Intake SOP
Run this when finance receives a purchase request. The CSM (Customer Solutions Manager) reviews it.
""", # NOTE: glossary drift — CSM here is Customer Solutions Manager
"runbooks/incident-comms.md": """---
last_reviewed: 2026-03-01
---
# Incident Comms Cascade
(no owner field — missing-owner case)
Send alerts to the on-call SRE.
""",
"orphan-page.md": """---
owner: pat@company.com
last_reviewed: 2026-04-01
---
# Orphan Page
Nobody links here.
The CSM may find this useful.
""",
"old-stale-page.md": """# Old Page
(no frontmatter at all — missing-owner AND probably stale via mtime)
""",
"sops/employee-onboarding.md": """---
owner: hr@company.com
last_reviewed: 2026-04-20
---
# Employee Onboarding SOP
Coordinate with the CSM and VMO for system access.
Link: [Vendor Offboarding](vendor-offboarding.md).
""",
}
def _materialize_sample_vault() -> Path:
tmp = Path(tempfile.mkdtemp(prefix="kb-sample-"))
for relpath, content in SAMPLE_PAGES.items():
full = tmp / relpath
full.parent.mkdir(parents=True, exist_ok=True)
full.write_text(content, encoding="utf-8")
# Backdate one file via os.utime so mtime-based stale detection
# has something to find.
import os
old = tmp / "old-stale-page.md"
if old.exists():
old_ts = (dt.datetime.now() -
dt.timedelta(days=720)).timestamp()
os.utime(old, (old_ts, old_ts))
return tmp
def main(argv=None) -> int:
p = argparse.ArgumentParser(
description="Walk a markdown KB and emit a hygiene report: "
"orphans, stale, missing-owner, glossary drift."
)
p.add_argument("--input", "-i", type=str,
help="Path to KB root directory.")
p.add_argument("--output", "-o", choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).")
p.add_argument("--stale-days", type=int, default=365,
help="Days since last edit to consider stale "
"(default: 365).")
p.add_argument("--sample", action="store_true",
help="Run against a tiny synthetic vault in a "
"tmpdir.")
args = p.parse_args(argv)
if args.sample:
root = _materialize_sample_vault()
elif args.input:
root = Path(args.input).resolve()
if not root.exists() or not root.is_dir():
print(f"ERROR: input directory not found: {args.input}",
file=sys.stderr)
return 2
else:
print("ERROR: provide --input <kb-root-dir> or --sample",
file=sys.stderr)
return 2
pages = walk_vault(root, stale_days=args.stale_days)
if not pages:
print(f"WARNING: no markdown files found under {root}",
file=sys.stderr)
return 1
if args.output == "json":
print(json.dumps(generate_json_report(root, pages, args.stale_days),
indent=2))
else:
print(generate_report(root, pages, args.stale_days))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/runbook_validator.py
#!/usr/bin/env python3
"""runbook_validator.py
Validate a runbook by checking each step against six required attributes:
1. Named owner (not "the team", not "ops")
2. Expected duration (concrete number + unit)
3. Observable success signal
4. Observable failure signal
5. Rollback path (or explicit "cannot roll back — escalate to X")
6. Escalation contact
Output is a per-step traffic-light + overall validity score 0-100 + a list
of MUST-FIX issues.
Verdict thresholds:
>= 80 SAFE-TO-USE
60-79 USE-WITH-CAUTION
< 60 NOT-SAFE
Input formats:
--input runbook.md (markdown: heuristic parser, expects
"## Step N:" or "### Step N:" headings)
--input runbook.json (JSON: explicit step list — preferred)
JSON schema:
{
"runbook_name": "Incident Comms Cascade",
"steps": [
{
"title": "Acknowledge alert in PagerDuty",
"owner": "On-call IC (named rotation)",
"duration_minutes": 2,
"success_signal": "PagerDuty incident transitions to acknowledged",
"failure_signal": "Incident remains in triggered state after 2 min",
"rollback": "n/a (acknowledgement is non-mutating)",
"escalation": "Engineering Manager on-call"
}
]
}
Stdlib only.
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from dataclasses import dataclass, field, asdict
from pathlib import Path
VAGUE_OWNER_TOKENS = {
"the team", "team", "ops", "the ops team", "engineering",
"support", "everyone", "whoever", "someone", "tbd", "n/a",
"the on-call", "on call", "rotation", # rotation alone is vague
}
# Vague success signal phrases that get flagged. Matched as whole-phrase
# substrings — must be specific enough to avoid false positives on
# legitimate observables that happen to contain a common word.
VAGUE_SUCCESS_TOKENS = [
"service is up", "it works", "things look good", "looks fine",
"no errors", "should work", "appears to be",
"verify the service", "check that it works", "looks good",
]
# Phrases that count as observable.
OBSERVABLE_HINTS = [
"http 2", "http 3", "http 4", "http 5", # status codes
"status code", "exit code 0", "/healthz", "/health", "200 ok",
"log line", "metric", "dashboard shows", "alert clears",
"incident transitions", "ticket moves to", "slack reaction",
"email received", "record updated", "field set to",
]
DURATION_PATTERN = re.compile(
r"\b\d+(?:\.\d+)?\s*(seconds?|secs?|minutes?|mins?|hours?|hrs?|days?)\b",
re.IGNORECASE,
)
# Rollback acceptable phrasing: either a real rollback OR explicit
# acknowledgement that rollback is impossible plus escalation.
NO_ROLLBACK_ACCEPTABLE = [
"cannot be rolled back",
"cannot roll back",
"non-mutating",
"read-only",
"no rollback needed",
"irreversible — escalate",
"irreversible - escalate",
]
@dataclass
class StepFinding:
step_index: int
title: str
owner_ok: bool = False
duration_ok: bool = False
success_ok: bool = False
failure_ok: bool = False
rollback_ok: bool = False
escalation_ok: bool = False
issues: list = field(default_factory=list)
@property
def passes(self) -> int:
return sum([
self.owner_ok, self.duration_ok, self.success_ok,
self.failure_ok, self.rollback_ok, self.escalation_ok,
])
@property
def traffic_light(self) -> str:
if self.passes == 6:
return "GREEN"
if self.passes >= 4:
return "AMBER"
return "RED"
def _check_owner(owner: str) -> tuple[bool, str]:
if not owner or not owner.strip():
return False, "missing owner"
norm = owner.strip().lower()
for token in VAGUE_OWNER_TOKENS:
# Vague if owner is ONLY that token (allow named rotations like
# "SRE on-call (alex)" by checking for parenthetical name OR @).
if norm == token or norm.startswith(token + " "):
if "@" in owner or "(" in owner:
return True, ""
return False, (
f"vague owner '{owner}' — name a specific human or a "
f"specifically-named rotation (e.g., 'SRE on-call "
f"rotation (PagerDuty: sre-primary)')"
)
return True, ""
def _check_duration(duration_str: str, duration_minutes) -> tuple[bool, str]:
if duration_minutes is not None:
try:
val = float(duration_minutes)
if val > 0:
return True, ""
return False, "duration_minutes is zero or negative"
except (TypeError, ValueError):
pass
if duration_str and DURATION_PATTERN.search(duration_str):
return True, ""
return False, (
"missing expected duration (need a concrete number + unit, "
"e.g., '2 minutes', '30 seconds')"
)
def _check_observable(signal: str, kind: str) -> tuple[bool, str]:
if not signal or not signal.strip():
return False, f"missing observable {kind} signal"
norm = signal.lower()
for vague in VAGUE_SUCCESS_TOKENS:
if vague in norm:
return False, (
f"vague {kind} signal '{signal}' — need an observable "
f"(e.g., 'HTTP 200 from /healthz', not 'service is up')"
)
for hint in OBSERVABLE_HINTS:
if hint in norm:
return True, ""
# Heuristic: if signal contains digits, equality, code-fences, or
# specific verbs that imply an observation, accept.
if any(ch in signal for ch in ("=", ":", "`", "200", "404", "500")):
return True, ""
if re.search(r"\b(returns?|equals?|shows?|transitions?|moves?|"
r"closes?|emits?|logs?|created|deleted|received|"
r"updated|set\s+to|reaches?|reports?)\b", norm):
return True, ""
return False, (
f"{kind} signal '{signal}' is not clearly observable — rewrite "
f"as a concrete check (status code, log line, dashboard panel, "
f"ticket state)"
)
def _check_rollback(rollback: str) -> tuple[bool, str]:
if not rollback or not rollback.strip():
return False, "missing rollback path"
norm = rollback.lower()
for ok in NO_ROLLBACK_ACCEPTABLE:
if ok in norm:
return True, ""
# If there's substantive text (> 12 chars) describing a step, accept.
if len(rollback.strip()) >= 12:
return True, ""
return False, (
f"rollback path too thin ('{rollback}') — either describe the "
f"rollback procedure OR write 'cannot be rolled back — "
f"escalate to <name>'"
)
def _check_escalation(escalation: str) -> tuple[bool, str]:
if not escalation or not escalation.strip():
return False, "missing escalation contact"
norm = escalation.strip().lower()
for token in VAGUE_OWNER_TOKENS:
if norm == token or norm.startswith(token + " "):
if "@" not in escalation and "(" not in escalation:
return False, (
f"vague escalation contact '{escalation}' — name a "
f"specific human, role+email, or named on-call rotation"
)
return True, ""
def validate_step(step: dict, idx: int) -> StepFinding:
finding = StepFinding(
step_index=idx,
title=step.get("title", f"(step {idx} — no title)"),
)
owner_ok, owner_err = _check_owner(step.get("owner", ""))
finding.owner_ok = owner_ok
if not owner_ok:
finding.issues.append(f"OWNER: {owner_err}")
duration_ok, duration_err = _check_duration(
step.get("duration_str", ""),
step.get("duration_minutes"),
)
finding.duration_ok = duration_ok
if not duration_ok:
finding.issues.append(f"DURATION: {duration_err}")
succ_ok, succ_err = _check_observable(
step.get("success_signal", ""), "success")
finding.success_ok = succ_ok
if not succ_ok:
finding.issues.append(f"SUCCESS: {succ_err}")
fail_ok, fail_err = _check_observable(
step.get("failure_signal", ""), "failure")
finding.failure_ok = fail_ok
if not fail_ok:
finding.issues.append(f"FAILURE: {fail_err}")
rb_ok, rb_err = _check_rollback(step.get("rollback", ""))
finding.rollback_ok = rb_ok
if not rb_ok:
finding.issues.append(f"ROLLBACK: {rb_err}")
esc_ok, esc_err = _check_escalation(step.get("escalation", ""))
finding.escalation_ok = esc_ok
if not esc_ok:
finding.issues.append(f"ESCALATION: {esc_err}")
return finding
def _parse_markdown(text: str) -> dict:
"""Heuristic parser. Expects steps as '## Step N: title' or
'### Step N: title' followed by bullet attributes."""
lines = text.splitlines()
name_match = re.search(r"^#\s+(.+)$", text, re.MULTILINE)
runbook_name = name_match.group(1).strip() if name_match else "(unnamed)"
steps = []
current = None
step_re = re.compile(
r"^#{2,3}\s+Step\s+(\d+)\s*:?\s*(.*)$", re.IGNORECASE)
attr_re = re.compile(
r"^\s*[-*]\s+\*?\*?(Owner|Duration|Success|Failure|"
r"Rollback|Escalation)\*?\*?\s*:?\s*(.+)$",
re.IGNORECASE,
)
for line in lines:
m = step_re.match(line)
if m:
if current:
steps.append(current)
current = {"title": m.group(2).strip() or f"step {m.group(1)}"}
continue
if current:
am = attr_re.match(line)
if am:
key = am.group(1).lower()
val = am.group(2).strip()
if key == "owner":
current["owner"] = val
elif key == "duration":
current["duration_str"] = val
elif key == "success":
current["success_signal"] = val
elif key == "failure":
current["failure_signal"] = val
elif key == "rollback":
current["rollback"] = val
elif key == "escalation":
current["escalation"] = val
if current:
steps.append(current)
return {"runbook_name": runbook_name, "steps": steps}
def _sample_runbook() -> dict:
"""Deliberately broken incident-comms runbook to demonstrate
failure detection."""
return {
"runbook_name": "Incident Comms Cascade (BROKEN sample)",
"steps": [
{
"title": "Acknowledge alert",
"owner": "the team", # vague
"duration_str": "", # missing
"success_signal": "service is up", # vague
"failure_signal": "", # missing
"rollback": "", # missing
"escalation": "ops", # vague
},
{
"title": "Open incident channel",
"owner": "Incident Commander on-call "
"(PagerDuty: ic-primary)",
"duration_str": "2 minutes",
"success_signal": "Slack channel #inc-<id> created and "
"linked from PagerDuty incident",
"failure_signal": "Slack returns 4xx or channel-create "
"API call times out",
"rollback": "n/a — read-only operation (channel can be "
"archived if created in error)",
"escalation": "Engineering Manager on-call "
"(em-primary@company.com)",
},
{
"title": "Notify execs via paging tree",
"owner": "Communications Lead "
"(comms-lead@company.com)",
"duration_str": "5 minutes",
"success_signal": "Exec recipient list shows email "
"received (200 OK from SES API)",
"failure_signal": "SES API returns 5xx or recipient "
"delivery status = bounced",
"rollback": "Send retraction email to same list with "
"subject prefix 'RETRACTION:'",
"escalation": "VP Communications "
"(vp-comms@company.com)",
},
],
}
def generate_report(runbook: dict, findings: list) -> str:
total = len(findings)
if total == 0:
return "ERROR: runbook contains no steps."
score = round(sum(f.passes for f in findings) /
(6 * total) * 100, 1)
if score >= 80:
verdict = "SAFE-TO-USE"
elif score >= 60:
verdict = "USE-WITH-CAUTION"
else:
verdict = "NOT-SAFE"
lines = [
f"# Runbook validation: {runbook.get('runbook_name', '(unnamed)')}",
"",
f"**Steps validated:** {total}",
f"**Validity score:** {score} / 100",
f"**Verdict:** {verdict}",
"",
"## Per-step traffic-light",
"",
"| Step | Title | Owner | Duration | Success | Failure | "
"Rollback | Escalation | Light |",
"|------|-------|-------|----------|---------|---------|"
"----------|------------|-------|",
]
for f in findings:
def ck(b):
return "OK" if b else "FAIL"
lines.append(
f"| {f.step_index} | {f.title[:40]} | {ck(f.owner_ok)} | "
f"{ck(f.duration_ok)} | {ck(f.success_ok)} | "
f"{ck(f.failure_ok)} | {ck(f.rollback_ok)} | "
f"{ck(f.escalation_ok)} | {f.traffic_light} |"
)
lines.append("")
lines.append("## MUST-FIX issues")
lines.append("")
any_issues = False
for f in findings:
if f.issues:
any_issues = True
lines.append(f"### Step {f.step_index}: {f.title}")
for issue in f.issues:
lines.append(f"- {issue}")
lines.append("")
if not any_issues:
lines.append("_(none — all steps pass all six checks)_")
return "\n".join(lines) + "\n"
def generate_json_report(runbook: dict, findings: list) -> dict:
total = len(findings) or 1
score = round(sum(f.passes for f in findings) / (6 * total) * 100, 1)
verdict = ("SAFE-TO-USE" if score >= 80
else "USE-WITH-CAUTION" if score >= 60
else "NOT-SAFE")
return {
"runbook_name": runbook.get("runbook_name", "(unnamed)"),
"step_count": len(findings),
"validity_score": score,
"verdict": verdict,
"findings": [asdict(f) | {"traffic_light": f.traffic_light,
"passes": f.passes} for f in findings],
}
def main(argv=None) -> int:
p = argparse.ArgumentParser(
description="Validate a runbook against six step-completeness "
"rules. Output traffic-light + score + MUST-FIX list."
)
p.add_argument("--input", "-i", type=str,
help="Path to runbook .md or .json file.")
p.add_argument("--output", "-o", choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).")
p.add_argument("--sample", action="store_true",
help="Run against a deliberately-broken sample runbook.")
args = p.parse_args(argv)
if args.sample:
runbook = _sample_runbook()
elif args.input:
path = Path(args.input)
if not path.exists():
print(f"ERROR: input file not found: {args.input}",
file=sys.stderr)
return 2
text = path.read_text()
if path.suffix.lower() == ".json":
runbook = json.loads(text)
else:
runbook = _parse_markdown(text)
else:
print("ERROR: provide --input <runbook.md|json> or --sample",
file=sys.stderr)
return 2
steps = runbook.get("steps", [])
if not steps:
print("ERROR: runbook contains no steps "
"(or markdown parser found none — try JSON input)",
file=sys.stderr)
return 1
findings = [validate_step(s, i + 1) for i, s in enumerate(steps)]
if args.output == "json":
print(json.dumps(generate_json_report(runbook, findings),
indent=2))
else:
print(generate_report(runbook, findings))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/sop_generator.py
#!/usr/bin/env python3
"""sop_generator.py
Generate a 5W2H-structured Standard Operating Procedure (SOP) from a JSON
metadata file. Output is markdown by default, or normalized JSON.
5W2H = Who, What, When, Where, Why, How, How-much (Ishikawa, *Guide to
Quality Control*, 1985). Each section is mandatory; missing sections produce
a warning footer naming the section.
Industry tuning:
--profile {ops,support,finance,hr,it,regulated}
- ops: general internal ops SOP scaffold
- support: adds customer-impact section + escalation matrix
- finance: adds controls + reconciliation + segregation-of-duties section
- hr: flags PII / sensitive-data handling; adds consent section
- it: adds system + access + change-management section
- regulated: adds version control, signoff matrix, audit-trail, change
history (required under ISO 9001 / FDA 21 CFR Part 211 /
SOC 2 / HIPAA / ISO 13485)
Regulatory overlay flags attach the appropriate compliance preamble:
regulatory_overlay: ["SOC2", "HIPAA", "ISO13485", "GDPR", "SOX"]
Input schema (JSON):
{
"sop_name": "Vendor Offboarding",
"process_owner": "alex@company.com",
"triggering_event": "Vendor contract not renewed OR vendor terminated",
"audience_role": "Vendor Management Office operator",
"frequency": "On-demand (avg 3 times per quarter)",
"regulatory_overlay": ["SOC2"],
"inputs": ["Vendor name", "Contract end date", "Data access list"],
"outputs": ["Access revoked", "Data deleted/returned", "Final invoice paid"],
"steps_outline": [
"Notify vendor of offboarding intent",
"Inventory data and system access",
"Revoke production system access",
"Confirm data deletion or return",
"Final invoice reconciliation",
"Archive vendor record"
],
"estimated_minutes": 240,
"estimated_cost_usd": 800
}
Stdlib only.
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from pathlib import Path
VALID_PROFILES = {"ops", "support", "finance", "hr", "it", "regulated"}
VALID_OVERLAYS = {"SOC2", "HIPAA", "ISO13485", "GDPR", "SOX"}
REGULATORY_PREAMBLE = {
"SOC2": (
"**SOC 2 overlay:** This SOP supports the Common Criteria control "
"framework. Changes require change-management approval (CC8.1). "
"Evidence of execution must be retained for the audit period."
),
"HIPAA": (
"**HIPAA overlay:** This SOP touches Protected Health Information "
"(PHI). All access must be logged per §164.312(b). Minimum-necessary "
"rule applies (§164.502(b))."
),
"ISO13485": (
"**ISO 13485 overlay:** This is a controlled document under §4.2.4. "
"Document revision, approval, and review records must be maintained. "
"Use the regulated profile."
),
"GDPR": (
"**GDPR overlay:** This SOP touches personal data of EU data "
"subjects. Lawful basis must be documented (Art. 6). Data-subject "
"rights (Art. 15-22) requests must be respected during execution."
),
"SOX": (
"**SOX overlay:** This SOP supports a financial control. Execution "
"must be evidenced and segregation-of-duties enforced. Quarterly "
"management testing applies."
),
}
@dataclass
class SOPMetadata:
sop_name: str = ""
process_owner: str = ""
triggering_event: str = ""
audience_role: str = ""
frequency: str = ""
regulatory_overlay: list = field(default_factory=list)
inputs: list = field(default_factory=list)
outputs: list = field(default_factory=list)
steps_outline: list = field(default_factory=list)
estimated_minutes: int = 0
estimated_cost_usd: int = 0
def validate(self) -> list:
errs = []
for fld in ("sop_name", "process_owner", "triggering_event",
"audience_role", "frequency"):
if not getattr(self, fld):
errs.append(f"missing required field: '{fld}'")
if not self.steps_outline:
errs.append("missing 'steps_outline' (need >= 1 step)")
for ov in self.regulatory_overlay:
if ov not in VALID_OVERLAYS:
errs.append(
f"invalid regulatory_overlay '{ov}'; "
f"allowed: {sorted(VALID_OVERLAYS)}"
)
return errs
def _sample_metadata() -> dict:
return {
"sop_name": "Vendor Offboarding",
"process_owner": "alex@company.com (Vendor Management Lead)",
"triggering_event": (
"Vendor contract not renewed OR vendor terminated for cause"
),
"audience_role": "Vendor Management Office (VMO) operator",
"frequency": "On-demand (avg 3 executions per quarter)",
"regulatory_overlay": ["SOC2"],
"inputs": [
"Vendor legal name",
"Contract end date (effective offboarding date)",
"List of systems with vendor access",
"List of data classes vendor processed",
],
"outputs": [
"All production system access revoked (evidenced)",
"Vendor data deleted or returned (evidenced)",
"Final invoice reconciled and paid",
"Vendor record archived in VMO registry",
],
"steps_outline": [
"Notify vendor of offboarding intent (written, 30 days notice)",
"Inventory data classes and system access vendor holds",
"Revoke production system access (IAM, VPN, SaaS)",
"Confirm data deletion (vendor certification) or data return",
"Final invoice reconciliation and payment",
"Archive vendor record in VMO registry with offboarding evidence",
],
"estimated_minutes": 240,
"estimated_cost_usd": 800,
}
def _build_who(meta: SOPMetadata, profile: str) -> str:
lines = [
"### Who",
"",
f"- **Process owner (Accountable):** {meta.process_owner}",
f"- **Audience (Responsible):** {meta.audience_role}",
]
if profile == "regulated":
lines.append("- **Approver (Consulted):** "
"Quality Management Representative")
lines.append("- **Auditor (Informed):** "
"Internal Audit / Compliance")
elif profile == "finance":
lines.append("- **Approver (Consulted):** Controller")
lines.append("- **Segregation-of-duties review:** "
"Required (initiator != approver != payer)")
elif profile == "hr":
lines.append("- **Approver (Consulted):** HR Business Partner")
lines.append("- **Privacy review (Informed):** "
"Data Protection Officer (if PII touched)")
elif profile == "it":
lines.append("- **Approver (Consulted):** "
"Change Advisory Board (for system-mutating steps)")
elif profile == "support":
lines.append("- **Approver (Consulted):** Support Team Lead")
lines.append("- **Escalation (Informed):** "
"Engineering on-call (if customer-impact > 30 min)")
return "\n".join(lines)
def _build_what(meta: SOPMetadata) -> str:
lines = [
"### What",
"",
f"**Process name:** {meta.sop_name}",
"",
"**Inputs required before starting:**",
"",
]
for inp in meta.inputs:
lines.append(f"- {inp}")
lines.append("")
lines.append("**Outputs produced:**")
lines.append("")
for out in meta.outputs:
lines.append(f"- {out}")
return "\n".join(lines)
def _build_when(meta: SOPMetadata) -> str:
return (
"### When\n\n"
f"- **Triggering event:** {meta.triggering_event}\n"
f"- **Frequency:** {meta.frequency}\n"
"- **Time-of-day constraint:** _(business hours only? on-call? "
"fill in)_\n"
"- **Blocking dependencies:** _(prerequisites that must be true "
"before starting)_"
)
def _build_where(meta: SOPMetadata, profile: str) -> str:
lines = [
"### Where",
"",
"- **Primary system of record:** _(name the system — Salesforce, "
"Notion, Jira, ServiceNow, etc.)_",
"- **Supporting tools:** _(IAM console, IT ticketing, accounting "
"system, etc.)_",
"- **Canonical doc location:** _(URL of this SOP in the wiki)_",
]
if profile in {"it", "regulated"}:
lines.append("- **Change-management ticket location:** "
"_(Jira / ServiceNow queue)_")
return "\n".join(lines)
def _build_why(meta: SOPMetadata) -> str:
lines = [
"### Why",
"",
"**Purpose:** _(one-paragraph statement of why this process "
"exists. Anchor to a business outcome, not a task.)_",
"",
"**Regulatory basis (if any):**",
"",
]
if meta.regulatory_overlay:
for ov in meta.regulatory_overlay:
lines.append(f"- {REGULATORY_PREAMBLE[ov]}")
else:
lines.append("- _(none — confirm by checking data classes "
"touched. If process touches PHI, financial controls, "
"or regulated devices, the answer is not 'none'.)_")
return "\n".join(lines)
def _build_how(meta: SOPMetadata) -> str:
lines = ["### How", ""]
lines.append("Step-by-step procedure. Each step must have a named "
"owner, expected duration, and observable success signal.")
lines.append("")
for i, step in enumerate(meta.steps_outline, start=1):
lines.append(f"**Step {i}: {step}**")
lines.append("")
lines.append("- **Owner:** _(named human or named rotation)_")
lines.append("- **Expected duration:** _(concrete number + unit)_")
lines.append("- **Success signal (observable):** _(e.g., 'IAM "
"console shows user disabled', not 'access is "
"revoked')_")
lines.append("- **Failure signal (observable):** _(what tells you "
"the step did not work)_")
lines.append("- **If step fails — rollback or escalation:** "
"_(rollback path or 'escalate to X — cannot be "
"rolled back')_")
lines.append("")
return "\n".join(lines)
def _build_how_much(meta: SOPMetadata) -> str:
mins = meta.estimated_minutes or "_(fill in)_"
cost = meta.estimated_cost_usd
cost_line = f"cost" if cost else "_(fill in)_"
return (
"### How-much\n\n"
f"- **Estimated execution time:** {mins} minutes\n"
f"- **Estimated cost per execution:** {cost_line} "
"(labor + license + third-party fees)\n"
"- **Frequency × cost = annual run-rate:** _(compute from "
"frequency + cost per execution)_\n"
)
def _build_regulated_footer() -> str:
return (
"\n---\n\n"
"## Document control (regulated profile)\n\n"
"- **Version:** 1.0\n"
"- **Effective date:** _(YYYY-MM-DD)_\n"
"- **Next review date:** _(YYYY-MM-DD — within 12 months, "
"or 90 days under HIPAA / ISO 13485)_\n"
"- **Approval signoff (named):** _(QMR / Compliance Officer)_\n"
"- **Change history:**\n\n"
"| Version | Date | Author | Change summary | Approver |\n"
"|---------|------|--------|----------------|----------|\n"
"| 1.0 | _date_ | _author_ | Initial issue | _approver_ |\n"
)
def _build_finance_footer() -> str:
return (
"\n---\n\n"
"## Controls section (finance profile)\n\n"
"- **Control objective:** _(what financial assertion this "
"controls — e.g., completeness of vendor payments)_\n"
"- **Segregation of duties:** Initiator, approver, and payer "
"must be distinct individuals.\n"
"- **Evidence retained:** _(invoice copy, approval email, "
"payment confirmation)_\n"
"- **Testing frequency:** Quarterly by Internal Audit.\n"
)
def _build_hr_footer() -> str:
return (
"\n---\n\n"
"## Privacy & sensitive-data handling (HR profile)\n\n"
"- **Data classes touched:** _(name, address, SSN/national ID, "
"compensation, medical, etc.)_\n"
"- **Lawful basis for processing:** _(employment contract, "
"legal obligation, legitimate interest, consent)_\n"
"- **Retention period:** _(per local employment law + GDPR if "
"applicable)_\n"
"- **Access restriction:** Need-to-know basis only.\n"
)
def _build_it_footer() -> str:
return (
"\n---\n\n"
"## Change management (IT profile)\n\n"
"- **Change type:** _(standard / normal / emergency)_\n"
"- **Change ticket:** _(link to Jira / ServiceNow)_\n"
"- **Rollback plan:** _(named rollback procedure)_\n"
"- **Test evidence:** _(staging validation)_\n"
"- **Communication plan:** _(who is notified pre/post change)_\n"
)
def _build_support_footer() -> str:
return (
"\n---\n\n"
"## Customer impact & escalation (support profile)\n\n"
"- **Customer impact category:** _(none / single-customer / "
"multi-customer / company-wide outage)_\n"
"- **External comms required:** _(yes/no — if yes, link "
"internal-comms cascade SOP)_\n"
"- **Escalation matrix:**\n\n"
"| Trigger | Escalate to | SLA |\n"
"|---------|-------------|-----|\n"
"| Customer-impact > 30 min | Engineering on-call | 5 min |\n"
"| Multi-customer impact | Support Lead + VP Eng | 10 min |\n"
"| External comms needed | Communications + CEO | 30 min |\n"
)
PROFILE_FOOTER = {
"ops": "",
"support": _build_support_footer(),
"finance": _build_finance_footer(),
"hr": _build_hr_footer(),
"it": _build_it_footer(),
"regulated": _build_regulated_footer(),
}
def generate_markdown(meta: SOPMetadata, profile: str) -> str:
header = (
f"# SOP: {meta.sop_name}\n\n"
f"_Profile: `{profile}` | "
f"Regulatory overlay: "
f"{meta.regulatory_overlay or 'none'}_\n\n"
"---\n\n"
"## 5W2H scaffolding\n\n"
"_(Ishikawa 1985, 5W2H method. Each section is required.)_\n"
)
body = "\n\n".join([
_build_who(meta, profile),
_build_what(meta),
_build_when(meta),
_build_where(meta, profile),
_build_why(meta),
_build_how(meta),
_build_how_much(meta),
])
footer = PROFILE_FOOTER.get(profile, "")
return header + "\n" + body + footer + "\n"
def generate_json(meta: SOPMetadata, profile: str) -> dict:
return {
"sop_name": meta.sop_name,
"profile": profile,
"metadata": asdict(meta),
"sections": {
"who": "RACI populated",
"what": f"{len(meta.inputs)} inputs / "
f"{len(meta.outputs)} outputs",
"when": meta.triggering_event,
"where": "system of record + canonical doc location",
"why": meta.regulatory_overlay or ["none"],
"how": [{"step": i + 1, "title": s}
for i, s in enumerate(meta.steps_outline)],
"how_much": {
"estimated_minutes": meta.estimated_minutes,
"estimated_cost_usd": meta.estimated_cost_usd,
},
},
}
def main(argv=None) -> int:
p = argparse.ArgumentParser(
description="Generate a 5W2H-structured SOP from JSON metadata."
)
p.add_argument("--input", "-i", type=str,
help="Path to SOP metadata JSON file.")
p.add_argument("--profile", choices=sorted(VALID_PROFILES),
default="ops",
help="Industry profile (default: ops).")
p.add_argument("--output", "-o", choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).")
p.add_argument("--sample", action="store_true",
help="Print a sample vendor-offboarding SOP.")
args = p.parse_args(argv)
if args.sample:
data = _sample_metadata()
elif args.input:
path = Path(args.input)
if not path.exists():
print(f"ERROR: input file not found: {args.input}",
file=sys.stderr)
return 2
data = json.loads(path.read_text())
else:
print("ERROR: provide --input <metadata.json> or --sample",
file=sys.stderr)
return 2
meta = SOPMetadata(**data)
errs = meta.validate()
if errs:
print("VALIDATION ERRORS:", file=sys.stderr)
for e in errs:
print(f" - {e}", file=sys.stderr)
return 1
if args.output == "json":
print(json.dumps(generate_json(meta, args.profile), indent=2))
else:
print(generate_markdown(meta, args.profile))
return 0
if __name__ == "__main__":
sys.exit(main())
Ghi quyết định vào bộ nhớ hai lớp qua decision-logger; bản ghi nhớ đã duyệt được lưu bền, bản ghi gốc giữ để tham khảo.
---
name: "decide"
description: "/cs:decide <memo> — Log a decision to two-layer memory via decision-logger. Approved memo becomes durable; raw transcripts kept for reference."
---
# /cs:decide — Log the Decision
**Command:** `/cs:decide <memo-path>`
Logs the founder's decision via the `decision-logger` skill. This is the gate where in-session deliberation becomes durable company memory.
## Pipeline Position
```
/cs:office-hours → /cs:brief → /cs:boardroom → /cs:decide → /cs:execute → /cs:post-mortem
↑ you are here
```
## Two-Layer Memory Model
The `decision-logger` skill maintains two layers:
1. **Raw transcripts** — every boardroom session, every advisor's Phase 2 position, every dissent. Stored under `~/.claude/decisions/raw/`. Reference only, never feeds back automatically.
2. **Approved decisions** — only the founder-signed memos. Stored under `~/.claude/decisions/approved/`. Feeds into future `/cs:office-hours` and `/cs:founder-mode` calls.
This split prevents the system from "remembering" unresolved debates as if they were decisions.
## Input
A board memo file (output of `/cs:boardroom`).
## Workflow
1. Read the memo path
2. Verify it has founder approval (status: APPROVED)
3. Extract structured decision record:
- Decision title
- Date decided
- Option chosen
- Success + kill criteria
- Dissent (preserved)
- Review checkpoint date
4. Append to `~/.claude/decisions/approved/<YYYY-MM-DD>-<slug>.md`
5. Update the raw transcript pointer
6. If llm-wiki bridge configured, write to vault (`~/company-vault/10-decisions/`)
7. Schedule auto-revisit (90 days)
## Output Record Format
```markdown
# Decision: <title>
**Decided:** YYYY-MM-DD
**By:** <founder name>
**Memo:** <link to boardroom memo>
**Brief:** <link to original brief>
**Review checkpoint:** YYYY-MM-DD (90d default)
## Decision
**Chose:** <option>
**Rejected:** <other options + one-line why>
## Success Criteria (binding)
- <metric, threshold, timeframe>
## Kill Criteria (binding)
- <metric, threshold, action>
## Preserved Dissent
- **<dissenter>:** <unresolved concern>
- (preserved verbatim; dissent never erased)
## Next Action
- `/cs:execute` → 90-day plan due <date>
## Status History
- YYYY-MM-DD: APPROVED
```
## Why Preserved Dissent
The biggest risk in approved decisions is forgetting why someone disagreed. When the kill criteria trigger, the dissent often turns out to have been correct. Preserving it verbatim — not summarized — keeps the company honest at post-mortem time.
## Routing
- `/cs:execute <decision>` — build the 90-day plan
- `/cs:freeze <decision> <days>` — lock if irreversible
- (Auto-scheduled) `/cs:post-mortem <decision>` — at 90-day checkpoint
## Stale-Decision Audit
`cs-chief-of-staff` runs a weekly stale audit:
- Decisions > 90 days without revisit → flag for `/cs:post-mortem`
- Decisions with kill criteria triggered → flag immediately
- Decisions whose company-context.md basis has changed → flag for re-examination
## Related
- Skill: [`decision-logger`](../../../skills/decision-logger/SKILL.md)
- Agent: [`cs-chief-of-staff`](../../agents/cs-chief-of-staff.md)
- Bridge: [`../../references/llm-wiki-bridge.md`](../../references/llm-wiki-bridge.md)
---
**Version:** 1.0.0
Tạo landing page chuyển đổi cao bằng component Next.js/React (TSX) và Tailwind CSS: hero, bảng giá, FAQ, đánh giá, CTA theo các khung copy PAS, AIDA, BAB.
---
name: "landing-page-generator"
description: "Generates high-converting landing pages as complete Next.js/React (TSX) components with Tailwind CSS. Creates hero sections, feature grids, pricing tables, FAQ accordions, testimonial blocks, and CTA sections using proven copy frameworks (PAS, AIDA, BAB). Outputs SEO meta tags, structured data, and performance-optimised code targeting Core Web Vitals (LCP < 1s, CLS < 0.1). Use when the user asks to create a landing page, marketing page, homepage, single-page site, lead capture page, campaign page, promo page, or conversion-optimised web page — or when they want to A/B test landing page variants or replace a static page with one designed to convert."
---
# Landing Page Generator
Generate high-converting landing pages from a product description. Output complete Next.js/React components with multiple section variants, proven copy frameworks, SEO optimization, and performance-first patterns. Not lorem ipsum — actual copy that converts.
**Target:** LCP < 1s · CLS < 0.1 · FID < 100ms
**Output:** TSX components + Tailwind styles + SEO meta + copy variants
---
## Core Capabilities
- 5 hero section variants (centered, split, gradient, video-bg, minimal)
- Feature sections (grid, alternating, cards with icons)
- Pricing tables (2–4 tiers with feature lists and toggle)
- FAQ accordion with schema markup
- Testimonials (grid, carousel, single-quote)
- CTA sections (banner, full-page, inline)
- Footer (simple, mega, minimal)
- 4 design styles with Tailwind class sets
---
## Generation Workflow
Follow these steps in order for every landing page request:
1. **Gather inputs** — collect product name, tagline, audience, pain point, key benefit, pricing tiers, design style, and copy framework using the trigger format below. Ask only for missing fields.
2. **Analyze brand voice** (recommended) — if the user has existing brand content (website copy, blog posts, marketing materials), run it through `marketing-skill/content-production/scripts/brand_voice_analyzer.py` to get a voice profile (formality, tone, perspective). Use the profile to inform design style and copy framework selection:
- formal + professional → **enterprise** style, **AIDA** framework
- casual + friendly → **bold-startup** style, **BAB** framework
- professional + authoritative → **dark-saas** style, **PAS** framework
- casual + conversational → **clean-minimal** style, **BAB** framework
3. **Select design style** — map the user's choice (or infer from brand voice analysis) to one of the four Tailwind class sets in the Design Style Reference.
4. **Apply copy framework** — write all headline and body copy using the chosen framework (PAS / AIDA / BAB) before generating components. Match the voice profile's formality and tone throughout.
5. **Generate sections in order** — Hero → Features → Pricing → FAQ → Testimonials → CTA → Footer. Skip sections not relevant to the product.
6. **Validate against SEO checklist** — run through every item in the SEO Checklist before outputting final code. Fix any gaps inline.
7. **Output final components** — deliver complete, copy-paste-ready TSX files with all Tailwind classes, SEO meta, and structured data included.
---
## Triggering This Skill
```
Product: [name]
Tagline: [one sentence value prop]
Target audience: [who they are]
Key pain point: [what problem you solve]
Key benefit: [primary outcome]
Pricing tiers: [free/pro/enterprise or describe]
Design style: dark-saas | clean-minimal | bold-startup | enterprise
Copy framework: PAS | AIDA | BAB
```
---
## Design Style Reference
| Style | Background | Accent | Cards | CTA Button |
|---|---|---|---|---|
| **Dark SaaS** | `bg-gray-950 text-white` | `violet-500/400` | `bg-gray-900 border border-gray-800` | `bg-violet-600 hover:bg-violet-500` |
| **Clean Minimal** | `bg-white text-gray-900` | `blue-600` | `bg-gray-50 border border-gray-200 rounded-2xl` | `bg-blue-600 hover:bg-blue-700` |
| **Bold Startup** | `bg-white text-gray-900` | `orange-500` | `shadow-xl rounded-3xl` | `bg-orange-500 hover:bg-orange-600 text-white` |
| **Enterprise** | `bg-slate-50 text-slate-900` | `slate-700` | `bg-white border border-slate-200 shadow-sm` | `bg-slate-900 hover:bg-slate-800 text-white` |
> **Bold Startup** headings: add `font-black tracking-tight` to all `<h1>`/`<h2>` elements.
---
## Copy Frameworks
**PAS (Problem → Agitate → Solution)**
- H1: Painful state they're in
- Sub: What happens if they don't fix it
- CTA: What you offer
- *Example — H1:* "Your team wastes 3 hours a day on manual reporting" / *Sub:* "Every hour spent on spreadsheets is an hour not closing deals. Your competitors are already automated." / *CTA:* "Automate your reports in 10 minutes →"
**AIDA (Attention → Interest → Desire → Action)**
- H1: Bold attention-grabbing statement → Sub: Interesting fact or benefit → Features: Desire-building proof points → CTA: Clear action
**BAB (Before → After → Bridge)**
- H1: "[Before state] → [After state]" → Sub: "Here's how [product] bridges the gap" → Features: How it works (the bridge)
---
## Representative Component: Hero (Centered Gradient — Dark SaaS)
Use this as the structural template for all hero variants. Swap layout classes, gradient direction, and image placement for split, video-bg, and minimal variants.
```tsx
export function HeroCentered() {
return (
<section className="relative flex min-h-screen flex-col items-center justify-center overflow-hidden bg-gray-950 px-4 text-center">
<div className="absolute inset-0 bg-gradient-to-b from-violet-900/20 to-transparent" />
<div className="pointer-events-none absolute -top-40 left-1/2 h-[600px] w-[600px] -translate-x-1/2 rounded-full bg-violet-600/20 blur-3xl" />
<div className="relative z-10 max-w-4xl">
<div className="mb-6 inline-flex items-center gap-2 rounded-full border border-violet-500/30 bg-violet-500/10 px-4 py-1.5 text-sm text-violet-300">
<span className="h-1.5 w-1.5 rounded-full bg-violet-400" />
Now in public beta
</div>
<h1 className="mb-6 text-5xl font-bold tracking-tight text-white md:text-7xl">
Ship faster.<br />
<span className="bg-gradient-to-r from-violet-400 to-pink-400 bg-clip-text text-transparent">
Break less.
</span>
</h1>
<p className="mx-auto mb-10 max-w-2xl text-xl text-gray-400">
The deployment platform that catches errors before your users do.
Zero config. Instant rollbacks. Real-time monitoring.
</p>
<div className="flex flex-col items-center gap-4 sm:flex-row sm:justify-center">
<Button size="lg" className="bg-violet-600 text-white hover:bg-violet-500 px-8">
Start free trial
</Button>
<Button size="lg" variant="outline" className="border-gray-700 text-gray-300">
See how it works →
</Button>
</div>
<p className="mt-4 text-sm text-gray-500">No credit card required · 14-day free trial</p>
</div>
</section>
)
}
```
---
## Other Section Patterns
### Feature Section (Alternating)
Map over a `features` array with `{ title, description, image, badge }`. Toggle layout direction with `i % 2 === 1 ? "lg:flex-row-reverse" : ""`. Use `<Image>` with explicit `width`/`height` and `rounded-2xl shadow-xl`. Wrap in `<section className="py-24">` with `max-w-6xl` container.
### Pricing Table
Map over a `plans` array with `{ name, price, description, features[], cta, highlighted }`. Highlighted plan gets `border-2 border-violet-500 bg-violet-950/50 ring-4 ring-violet-500/20`; others get `border border-gray-800 bg-gray-900`. Render `null` price as "Custom". Use `<Check>` icon per feature row. Layout: `grid gap-8 lg:grid-cols-3`.
### FAQ with Schema Markup
Inject `FAQPage` JSON-LD via `<script type="application/ld+json" dangerouslySetInnerHTML={{ __html: JSON.stringify(schema) }} />` inside the section. Map FAQs with `{ q, a }` into shadcn `<Accordion>` with `type="single" collapsible`. Container: `max-w-3xl`.
### Testimonials, CTA, Footer
- **Testimonials:** Grid (`grid-cols-1 md:grid-cols-3`) or single-quote hero block with avatar, name, role, and quote text.
- **CTA Banner:** Full-width section with headline, subhead, and two buttons (primary + ghost). Add trust signals (money-back guarantee, logo strip) immediately below.
- **Footer:** Logo + nav columns + social links + legal. Use `border-t border-gray-800` separator.
---
## SEO Checklist
- [ ] `<title>` tag: primary keyword + brand (50–60 chars)
- [ ] Meta description: benefit + CTA (150–160 chars)
- [ ] OG image: 1200×630px with product name and tagline
- [ ] H1: one per page, includes primary keyword
- [ ] Structured data: FAQPage, Product, or Organization schema
- [ ] Canonical URL set
- [ ] Image alt text on all `<Image>` components
- [ ] robots.txt and sitemap.xml configured
- [ ] Core Web Vitals: LCP < 1s, CLS < 0.1
- [ ] Mobile viewport meta tag present
- [ ] Internal linking to pricing and docs
> **Validation step:** Before outputting final code, verify every checklist item above is satisfied. Fix any gaps inline — do not skip items.
---
## Performance Targets
| Metric | Target | Technique |
|---|---|---|
| LCP | < 1s | Preload hero image, use `priority` on Next/Image |
| CLS | < 0.1 | Set explicit width/height on all images |
| FID/INP | < 100ms | Defer non-critical JS, use `loading="lazy"` |
| TTFB | < 200ms | Use ISR or static generation for landing pages |
| Bundle | < 100KB JS | Audit with `@next/bundle-analyzer` |
---
## Common Pitfalls
- Hero image not preloaded — add `priority` prop to first `<Image>`
- Missing mobile breakpoints — always design mobile-first with `sm:` prefixes
- CTA copy too vague — "Get started" beats "Learn more"; "Start free trial" beats "Sign up"
- Pricing page missing trust signals — add money-back guarantee and testimonials near CTA
- No above-the-fold CTA on mobile — ensure button is visible without scrolling on 375px viewport
---
## Related Skills
- **Brand Voice Analyzer** (`marketing-skill/content-production/scripts/brand_voice_analyzer.py`) — Run before generation to establish voice profile and ensure copy consistency
- **UI Design System** (`product-team/ui-design-system/`) — Generate design tokens from brand color before building the page
- **Competitive Teardown** (`product-team/competitive-teardown/`) — Competitive positioning informs landing page messaging and differentiation
FILE:references/conversion-patterns.md
# High-Converting Landing Page Patterns
## Overview
This reference catalogs proven landing page design patterns that drive higher conversion rates. Each pattern includes placement guidance, implementation notes, and A/B testing priorities.
## Hero Section Layouts
### Pattern 1: Left Copy + Right Product Screenshot
- **Best for:** SaaS products with a strong visual UI
- **Structure:** Headline, subheadline, CTA on left (60%); product screenshot on right (40%)
- **Why it works:** F-pattern reading leads with copy, product image provides proof
- **Conversion lift:** Baseline pattern, strong performer across industries
### Pattern 2: Centered Copy + Full-Width Background
- **Best for:** Brand-driven products, consumer apps
- **Structure:** Centered headline, subheadline, CTA over background image/gradient
- **Why it works:** Focuses attention on single message, high visual impact
- **Note:** Ensure text contrast against background for readability
### Pattern 3: Video Hero
- **Best for:** Complex products requiring demonstration
- **Structure:** Short headline + embedded video (60-90 seconds) + CTA below
- **Why it works:** Video explains what text cannot, increases time on page
- **Note:** Always include thumbnail; autoplay is often counterproductive
### Pattern 4: Interactive Demo
- **Best for:** Developer tools, data products, design tools
- **Structure:** Minimal copy + embedded interactive product experience
- **Why it works:** Hands-on experience converts better than description
- **Note:** Keep demo focused on one "aha moment" workflow
## Social Proof Placement
### Logo Bar
- **Position:** Immediately below hero section
- **Count:** 5-7 logos for credibility without clutter
- **Label:** "Trusted by" or "Used by teams at"
- **Selection:** Mix recognizable brands with relevant industry logos
### Testimonial Cards
- **Position:** After feature explanation sections
- **Format:** Photo + name + title + company + specific quote
- **Best quotes:** Include measurable outcomes ("Saved 10 hours/week")
- **Layout:** 2-3 testimonials in a row, carousel for more
### Case Study Callouts
- **Position:** Mid-page, before pricing
- **Format:** Company logo + headline metric + "Read the story" link
- **Example:** "Acme Corp reduced onboarding time by 60%"
### Social Proof Numbers
- **Position:** Near CTA or in dedicated trust section
- **Format:** Large number + descriptor (e.g., "50,000+ teams", "4.8/5 rating")
- **Selection:** Choose 3-4 most impressive metrics
## Pricing Table Designs
### Good/Better/Best (3-Tier)
- Most effective for SaaS with clear feature tiers
- Highlight recommended plan with visual emphasis
- Show annual discount prominently
- Include feature comparison matrix below
### Simple Two-Tier
- Free/Pro or Starter/Professional
- Best for PLG products with clear upgrade trigger
- Minimize decision fatigue
### Enterprise Custom
- Replace price with "Contact Sales" for high-ACV products
- List enterprise-specific features (SSO, SLA, dedicated support)
- Include a "Talk to Sales" CTA, not just a form
### Pricing Psychology
- Anchor with highest-priced plan first (or in the middle with visual highlight)
- Use monthly price with annual billing toggle
- Show savings percentage for annual plans
- Round prices ending in 9 (e.g., $49/mo, $99/mo)
## Trust Signals
### Security Badges
- SOC 2, ISO 27001, GDPR compliance badges
- SSL certificate indicator
- Place near forms and payment sections
### Guarantees
- Money-back guarantee with specific timeframe
- Free trial with no credit card requirement
- SLA uptime commitments
### Awards & Recognition
- Industry awards (best of, top rated)
- Analyst recognition (Gartner, Forrester, G2 Leader)
- Media mentions (as seen in logos)
### Real-Time Activity
- "X people signed up today" (use real data only)
- Recent activity feed
- Live user count
## Urgency Elements
### Ethical Urgency
- Limited-time pricing (with real deadline)
- Early adopter benefits (extra features, lower price)
- Cohort-based enrollment (actual capacity limits)
### Avoid
- Fake countdown timers that reset
- False scarcity ("only 3 left" when unlimited)
- Pressure tactics that erode trust
## Form Optimization
### Field Reduction
- Every additional field reduces conversion ~10%
- Start with email only, progressive profiling later
- Use single-column layouts for forms
### Smart Defaults
- Pre-fill country based on IP
- Auto-detect company from email domain
- Default to most popular plan
### Inline Validation
- Validate fields on blur, not on submit
- Show success states (green checkmark)
- Provide helpful error messages
### Multi-Step Forms
- Break long forms into 2-3 steps with progress indicator
- Put easiest questions first to build commitment
- Allow saving progress for complex forms
## Mobile-First Patterns
### Thumb-Friendly Design
- CTAs in thumb zone (bottom 40% of screen)
- Minimum tap target: 44x44px
- Adequate spacing between interactive elements
### Content Priority
- Lead with most compelling content (no scrolling to find CTA)
- Collapse secondary information into accordions
- Use sticky CTA bar on scroll
### Performance
- Target <3s load time on 3G
- Lazy-load images below fold
- Minimize JavaScript execution
## A/B Testing Priority Matrix
Test these elements in order of expected impact:
| Priority | Element | Expected Impact | Effort |
|----------|---------|----------------|--------|
| 1 | Headline | High | Low |
| 2 | CTA text and color | High | Low |
| 3 | Hero image/video | High | Medium |
| 4 | Social proof placement | Medium | Low |
| 5 | Form fields (fewer) | Medium | Low |
| 6 | Pricing presentation | Medium | Medium |
| 7 | Page length | Medium | High |
| 8 | Testimonial selection | Low | Low |
| 9 | Color scheme | Low | Medium |
| 10 | Font choices | Low | Low |
### Testing Best Practices
- Test one variable at a time for clear attribution
- Run tests for minimum 2 weeks or 1,000 visitors per variant
- Use 95% statistical significance threshold
- Document all test results for institutional knowledge
- Winner becomes new control for next test iteration
FILE:references/copy-frameworks.md
# Landing Page Copywriting Frameworks
## Overview
Four copy frameworks with worked SaaS examples you can adapt. Each framework includes a complete before/after example plus specific guidelines for each section.
## 1. AIDA Framework (Attention - Interest - Desire - Action)
The classic direct response formula, ideal for product landing pages.
**Example — Project management SaaS:**
> **Attention:** "Your Team Loses 12 Hours Every Sprint to Status Meetings"
>
> **Interest:** "Engineering teams at Series A-C startups spend 23% of their week in sync meetings — not writing code. We tracked 847 teams over 6 months. The pattern was clear: the more people in a standup, the less code shipped that day."
>
> **Desire:** "Teams using AsyncStand ship 31% more story points per sprint. No more 15-person standups where 13 people zone out. Replace your daily sync with a 2-minute async check-in that your engineers actually complete (94% response rate vs 67% attendance for live standups)."
>
> **Action:** "Start Your Free 14-Day Trial — No Credit Card Required"
### Attention
- Lead with a specific, quantified pain point (not vague claims)
- Weak: "Save time on meetings" → Strong: "Your Team Loses 12 Hours Every Sprint to Status Meetings"
- Keep headlines under 10 words for maximum impact
### Interest
- Back up the headline with specific data or a relatable scenario
- Weak: "Meetings waste time" → Strong: "We tracked 847 teams — the more people in standup, the less code shipped that day"
- Use their language: mirror words from customer reviews, support tickets, and G2 feedback
### Desire
- Stack measurable outcomes, not features
- Weak: "AI-powered async updates" → Strong: "31% more story points per sprint, 94% response rate"
- Compare directly to the status quo they already endure
### Action
- Single, clear CTA with action-oriented verb
- Reduce friction: "No credit card required," "Set up in 2 minutes"
- Repeat CTA after each major content block
## 2. PAS Framework (Problem - Agitate - Solution)
Best for pain-point-driven products where the problem is well understood.
**Example — Expense management tool:**
> **Problem:** "Your finance team is still chasing receipts in Slack DMs."
>
> **Agitate:** "Last quarter, your team spent 46 hours manually reconciling expenses across email threads, shared drives, and 'I'll submit it later' promises. That's $4,200 in payroll — spent on data entry. And when audit season hits? Good luck finding that client dinner receipt from February."
>
> **Solution:** "Snap a photo of the receipt. ExpenseFlow auto-extracts vendor, amount, and category in 3 seconds. Your monthly close drops from 5 days to 1. 2,400 finance teams already made the switch."
### Problem
- Name the exact scenario (not the abstract category)
- Weak: "Expense tracking is hard" → Strong: "Your finance team is still chasing receipts in Slack DMs"
- Mirror language from reviews and support tickets
### Agitate
- Quantify the cost in dollars, hours, or missed opportunities
- Weak: "This costs you money" → Strong: "46 hours last quarter, $4,200 in payroll — on data entry"
- Acknowledge the workarounds they've tried and why those fail too
### Solution
- Lead with the user action, not the technology: "Snap a photo" not "AI-powered OCR"
- Include one proof point: number of customers, time saved, or before/after metric
- Make the mechanism clear in one sentence: what happens when they use it
## 3. BAB Framework (Before - After - Bridge)
Ideal for aspirational products and lifestyle-oriented landing pages.
**Example — Sales enablement platform:**
> **Before:** "It's 9 PM. You're rebuilding a deck for tomorrow's demo because the prospect is in healthcare, not fintech. You copy-paste from three old decks, pray the logos are right, and rehearse the new talk track in the shower."
>
> **After:** "It's 9 AM. You type 'healthcare, 200-bed hospital, HIPAA-concerned CTO.' DeckGen builds your slides in 40 seconds — case studies, compliance badges, ROI calculator pre-loaded. You walk into the call with the best deck your prospect has ever seen."
>
> **Bridge:** "DeckGen connects to your CRM, learns your win patterns, and generates prospect-specific decks in under a minute. 340 AEs at companies like Stripe and Notion already use it. Start free — your first 5 decks are on us."
### Before
- Describe a specific, lived moment — not an abstract pain category
- Weak: "Sales decks take too long" → Strong: "It's 9 PM. You're rebuilding a deck for tomorrow's demo..."
- Use second person and present tense to make it feel immediate
### After
- Same level of specificity — show the transformed version of that exact moment
- Include a measurable outcome: "40 seconds," "best deck your prospect has ever seen"
- The after state should feel effortless compared to the before
### Bridge
- Name the product explicitly and explain the mechanism in one sentence
- Include one social proof data point
- End with a low-friction CTA that connects to the after state
## 4. 4Ps Framework (Promise - Picture - Proof - Push)
Strong for SaaS and B2B landing pages with measurable outcomes.
### Promise
- Make a clear, specific, believable promise
- Tie it to a measurable outcome
- Example: "Reduce customer churn by 25% in 90 days"
### Picture
- Help the reader visualize success
- Use scenarios they can relate to
- Show the product in context (screenshots, demos)
### Proof
- Back the promise with evidence
- Customer testimonials with specific results
- Case studies with before/after metrics
- Third-party validation (awards, analyst reports)
### Push
- Give a compelling reason to act now
- Limited-time offer, bonus, or guarantee
- Risk reversal (money-back guarantee, free trial)
## Headline Formulas
### Benefit-Driven
- "Get [Desired Outcome] Without [Common Objection]"
- "[Specific Result] in [Timeframe]"
- "The [Adjective] Way to [Achieve Goal]"
### Problem-Driven
- "Stop [Painful Activity]. Start [Better Alternative]."
- "Tired of [Problem]? There's a Better Way."
- "[Problem]? Not Anymore."
### Social Proof-Driven
- "[Number] Teams Trust [Product] to [Outcome]"
- "Why [Notable Company] Switched to [Product]"
- "Rated #1 for [Category] by [Authority]"
### Question-Driven
- "What If You Could [Desirable Outcome]?"
- "Ready to [Transformation]?"
- "Still [Painful Status Quo]?"
## CTA Best Practices
### Language
- Use first-person: "Start My Free Trial" > "Start Your Free Trial"
- Be specific: "Get My Report" > "Submit"
- Include benefit: "Start Saving Time" > "Sign Up"
- Add urgency naturally: "Start Free Today" > "Sign Up Now!!!"
### Placement
- Primary CTA above the fold
- Repeat after each major content section
- Sticky CTA on scroll (mobile especially)
- Exit-intent as last chance
### Design
- High contrast color (stands out from page palette)
- Sufficient whitespace around the button
- Large enough to tap on mobile (min 44x44px)
- Micro-copy below button to reduce anxiety ("No credit card required")
## Above-the-Fold Principles
The first viewport must accomplish these goals within 5 seconds:
1. **Communicate what you do** - Clear, jargon-free headline
2. **Show who it's for** - Audience identification
3. **Demonstrate value** - Primary benefit or outcome
4. **Provide next step** - Visible CTA button
5. **Build credibility** - One trust signal (logo bar, metric, badge)
### Above-the-Fold Checklist
- [ ] Headline states primary benefit (under 10 words)
- [ ] Subheadline adds specificity or addresses objection
- [ ] Hero image/video shows product in use
- [ ] CTA button is visible without scrolling
- [ ] At least one trust signal present
- [ ] No jargon or ambiguity in messaging
FILE:references/landing-page-patterns.md
# Landing Page Patterns
This reference captures high-converting page patterns and copy structures.
## Hero Section Patterns
### Pattern 1: Problem-Solution Hero
- Headline names the painful problem.
- Subheadline states the clear outcome.
- Primary CTA starts immediately.
- Optional supporting visual demonstrates product in context.
### Pattern 2: Outcome-First Hero
- Headline leads with measurable value.
- Subheadline clarifies who the page is for.
- CTA is action-oriented and specific.
### Pattern 3: Authority Hero
- Headline + trust indicator (logos, testimonial snippet, proof metric).
- Useful when category skepticism is high.
## Social Proof Layouts
### Logo Strip + Proof Metric
- Keep to recognizable logos.
- Add one proof metric (e.g., active users, revenue saved, hours reduced).
### Testimonial Grid
- 3-6 testimonials across segments.
- Include role/company where possible.
- Prefer concrete outcomes over generic praise.
### Case Study Snapshot
- Mini blocks: challenge -> approach -> measurable result.
## CTA Best Practices
- Use one dominant CTA per section.
- Match CTA verb to user intent ("Start trial", "Get demo", "Run audit").
- Keep CTA copy specific; avoid vague labels like "Submit".
- Reduce friction near CTA (short form, trust indicators, no surprise commitments).
## Above-the-Fold Checklist
- [ ] Clear value proposition in first viewport
- [ ] Audience clarity (who this is for)
- [ ] One primary CTA visible without scrolling
- [ ] Proof element (logos, stat, quote)
- [ ] Visual hierarchy emphasizes headline + CTA
- [ ] Mobile layout keeps CTA accessible
## Conversion-Optimized Templates
### SaaS Demo Page
1. Hero with problem-solution framing
2. Product walkthrough section
3. Social proof strip
4. Benefits by persona
5. Objection handling FAQ
6. Final CTA
### Lead Magnet Page
1. Promise + asset preview
2. Bullet outcomes
3. Short form
4. Trust/privacy note
### Product Launch Page
1. Outcome-first hero
2. Why now / differentiation
3. Feature blocks
4. Testimonials / beta feedback
5. Pricing or waitlist CTA
## Headline Formulas
### PAS (Problem-Agitate-Solution)
- Problem: identify the pain
- Agitate: show consequences of inaction
- Solution: position the offer as relief
Example structure:
"Still [problem]? Stop [negative consequence] and start [desired outcome]."
### AIDA (Attention-Interest-Desire-Action)
- Attention: pattern interrupt headline
- Interest: relevant context and stakes
- Desire: proof and benefits
- Action: concrete next step
### 4U Formula
- Useful: clear practical value
- Urgent: reason to act now
- Unique: differentiated promise
- Ultra-specific: concrete outcome and scope
Example structure:
"Get [specific result] in [timeframe] without [common pain]."
FILE:references/seo-checklist.md
# Landing Page SEO Checklist
## Overview
This checklist ensures landing pages are optimized for search engine visibility while maintaining conversion focus. Apply these checks before launching any landing page.
## Meta Tags
- [ ] **Title tag**: Under 60 characters, includes primary keyword, ends with brand name
- [ ] **Meta description**: 150-160 characters, includes CTA language, unique per page
- [ ] **Canonical URL**: Set to prevent duplicate content issues
- [ ] **Robots meta**: Ensure page is indexable (`index, follow`) unless intentionally noindex
- [ ] **Open Graph tags**: og:title, og:description, og:image, og:url for social sharing
- [ ] **Twitter Card tags**: twitter:card, twitter:title, twitter:description, twitter:image
- [ ] **Viewport meta**: `<meta name="viewport" content="width=device-width, initial-scale=1">`
## Structured Data
- [ ] **Organization schema**: Company name, logo, social profiles
- [ ] **Product schema**: Name, description, price, availability (for product pages)
- [ ] **FAQ schema**: For pages with FAQ sections (rich snippet opportunity)
- [ ] **Breadcrumb schema**: Navigation path for deep pages
- [ ] **Review schema**: Aggregate rating if testimonials present (use carefully per guidelines)
- [ ] **Validate**: Test all structured data with Google Rich Results Test
## Core Web Vitals Targets
### Largest Contentful Paint (LCP) - Target: < 2.5s
- [ ] Optimize hero image (WebP format, proper dimensions)
- [ ] Preload critical resources (`<link rel="preload">`)
- [ ] Use CDN for static assets
- [ ] Minimize render-blocking CSS and JavaScript
### First Input Delay (FID) / Interaction to Next Paint (INP) - Target: < 200ms
- [ ] Defer non-critical JavaScript
- [ ] Break up long tasks (>50ms)
- [ ] Minimize third-party script impact
- [ ] Use `requestAnimationFrame` for visual updates
### Cumulative Layout Shift (CLS) - Target: < 0.1
- [ ] Set explicit width/height on images and videos
- [ ] Reserve space for dynamic content (ads, embeds)
- [ ] Use `font-display: swap` for web fonts
- [ ] Avoid inserting content above existing content
## Keyword Placement
- [ ] **H1 tag**: Contains primary keyword, one per page only
- [ ] **H2 tags**: Include secondary keywords naturally
- [ ] **First paragraph**: Primary keyword appears in first 100 words
- [ ] **Body copy**: Natural keyword density (1-2%), no stuffing
- [ ] **Image alt text**: Descriptive, includes keyword where relevant
- [ ] **URL slug**: Short, keyword-rich, hyphen-separated
- [ ] **CTA text**: Consider keyword inclusion where natural
## Internal Linking
- [ ] Link to relevant product/feature pages
- [ ] Link to blog content that supports the page topic
- [ ] Use descriptive anchor text (not "click here")
- [ ] Ensure landing page is linked from main navigation or sitemap
- [ ] Link to pricing page if applicable
- [ ] Limit links to avoid diluting page authority (15-20 max)
## Image Optimization
- [ ] **Format**: Use WebP with JPEG/PNG fallback
- [ ] **Compression**: Lossy compression for photos, lossless for graphics
- [ ] **Dimensions**: Serve at exact display size (no CSS resizing)
- [ ] **Alt text**: Descriptive, 125 characters max, natural keyword inclusion
- [ ] **File names**: Descriptive, hyphenated (e.g., `product-dashboard-screenshot.webp`)
- [ ] **Lazy loading**: Apply to images below the fold (`loading="lazy"`)
- [ ] **Responsive images**: Use `srcset` for different viewport sizes
## Canonical URLs
- [ ] Self-referencing canonical on every page
- [ ] Consistent protocol (https) and trailing slash usage
- [ ] Canonical points to preferred URL version (www vs non-www)
- [ ] UTM parameters excluded from canonical URL
- [ ] Pagination handled with rel="next"/"prev" or single-page canonical
## Mobile Responsiveness
- [ ] **Mobile-friendly test**: Pass Google Mobile-Friendly Test
- [ ] **Touch targets**: Minimum 44x44px, 8px spacing between targets
- [ ] **Font size**: Minimum 16px base font, no pinch-to-zoom needed
- [ ] **Content parity**: All critical content accessible on mobile
- [ ] **Horizontal scroll**: None present at any viewport width
- [ ] **Form usability**: Appropriate input types (email, tel), autocomplete attributes
- [ ] **Media queries**: Breakpoints at 480px, 768px, 1024px, 1200px minimum
## Technical SEO
- [ ] **HTTPS**: SSL certificate valid and active
- [ ] **Page speed**: < 3s load time on mobile (test with PageSpeed Insights)
- [ ] **XML sitemap**: Page included in sitemap.xml
- [ ] **Robots.txt**: Page not blocked by robots.txt
- [ ] **404 handling**: Custom 404 page with navigation
- [ ] **Redirect chains**: No more than 1 redirect hop
- [ ] **Hreflang**: Set for multi-language landing pages
## Content Quality Signals
- [ ] **Unique content**: No duplicate content from other pages
- [ ] **Content depth**: Sufficient content for topic coverage (500+ words for SEO pages)
- [ ] **Readability**: Grade level 6-8 for broad audiences
- [ ] **Freshness**: Last modified date reflects recent updates
- [ ] **E-E-A-T signals**: Author expertise, company authority, trust indicators
FILE:scripts/landing_page_scaffolder.py
#!/usr/bin/env python3
"""Landing Page Scaffolder — Generate landing pages as HTML or Next.js TSX from config.
Creates production-ready landing pages with hero sections, features,
testimonials, pricing, CTAs, and responsive design.
Usage:
python landing_page_scaffolder.py config.json --format html --output page.html
python landing_page_scaffolder.py config.json --format tsx --output LandingPage.tsx
python landing_page_scaffolder.py config.json --format json
"""
import argparse
import json
import sys
from typing import Dict, List, Any, Optional
from datetime import datetime
import html as html_module
def escape(text: str) -> str:
"""HTML-escape text."""
return html_module.escape(str(text))
# ---------------------------------------------------------------------------
# Tailwind style mappings for TSX output
# ---------------------------------------------------------------------------
DESIGN_STYLES = {
"dark-saas": {
"bg": "bg-gray-950", "text": "text-white",
"accent": "violet", "card_bg": "bg-gray-900 border border-gray-800",
"btn": "bg-violet-600 hover:bg-violet-500 text-white",
"btn_secondary": "border border-gray-700 text-gray-300 hover:bg-gray-800",
"section_alt": "bg-gray-900/50", "muted": "text-gray-400",
"border": "border-gray-800",
},
"clean-minimal": {
"bg": "bg-white", "text": "text-gray-900",
"accent": "blue", "card_bg": "bg-gray-50 border border-gray-200 rounded-2xl",
"btn": "bg-blue-600 hover:bg-blue-700 text-white",
"btn_secondary": "border border-gray-300 text-gray-700 hover:bg-gray-50",
"section_alt": "bg-gray-50", "muted": "text-gray-500",
"border": "border-gray-200",
},
"bold-startup": {
"bg": "bg-white", "text": "text-gray-900",
"accent": "orange", "card_bg": "shadow-xl rounded-3xl bg-white",
"btn": "bg-orange-500 hover:bg-orange-600 text-white",
"btn_secondary": "border-2 border-orange-500 text-orange-600 hover:bg-orange-50",
"section_alt": "bg-orange-50/30", "muted": "text-gray-500",
"border": "border-gray-200",
},
"enterprise": {
"bg": "bg-slate-50", "text": "text-slate-900",
"accent": "slate", "card_bg": "bg-white border border-slate-200 shadow-sm",
"btn": "bg-slate-900 hover:bg-slate-800 text-white",
"btn_secondary": "border border-slate-300 text-slate-700 hover:bg-slate-100",
"section_alt": "bg-white", "muted": "text-slate-500",
"border": "border-slate-200",
},
}
# ---------------------------------------------------------------------------
# TSX generators
# ---------------------------------------------------------------------------
def tsx_nav(config: Dict[str, Any], style: Dict[str, str]) -> str:
brand = config.get("brand", "Brand")
nav_links = config.get("nav_links", [])
cta = config.get("nav_cta", {"text": "Get Started", "url": "#"})
links_jsx = "\n ".join(
f'<a href="{l.get("url", "#")}" className="{style["muted"]} hover:{style["text"]} font-medium transition-colors">{l.get("text", "")}</a>'
for l in nav_links
)
return f'''function Navbar() {{
return (
<nav className="sticky top-0 z-50 {style["bg"]} border-b {style["border"]} backdrop-blur-sm">
<div className="mx-auto flex max-w-7xl items-center justify-between px-6 py-4">
<a href="#" className="text-xl font-bold {style["text"]}">{brand}</a>
<div className="hidden items-center gap-8 md:flex">
{links_jsx}
<a href="{cta.get("url", "#")}" className="rounded-lg {style["btn"]} px-5 py-2.5 text-sm font-semibold transition-colors">
{cta.get("text", "Get Started")}
</a>
</div>
</div>
</nav>
);
}}'''
def tsx_hero(hero: Dict[str, Any], style: Dict[str, str]) -> str:
h1 = hero.get("headline", "Your Headline Here")
sub = hero.get("subheadline", "")
primary_cta = hero.get("primary_cta", {"text": "Get Started", "url": "#"})
secondary_cta = hero.get("secondary_cta", None)
secondary_jsx = ""
if secondary_cta:
secondary_jsx = f'''
<a href="{secondary_cta.get("url", "#")}" className="rounded-lg {style["btn_secondary"]} px-8 py-3 text-lg font-semibold transition-colors">
{secondary_cta.get("text", "Learn More")}
</a>'''
return f'''function Hero() {{
return (
<section className="flex min-h-[80vh] flex-col items-center justify-center px-6 py-24 text-center {style["bg"]}">
<div className="mx-auto max-w-4xl">
<h1 className="mb-6 text-5xl font-bold tracking-tight {style["text"]} md:text-7xl">
{h1}
</h1>
<p className="mx-auto mb-10 max-w-2xl text-xl {style["muted"]}">
{sub}
</p>
<div className="flex flex-col items-center gap-4 sm:flex-row sm:justify-center">
<a href="{primary_cta.get("url", "#")}" className="rounded-lg {style["btn"]} px-8 py-3 text-lg font-semibold transition-colors">
{primary_cta.get("text", "Get Started")}
</a>{secondary_jsx}
</div>
</div>
</section>
);
}}'''
def tsx_features(features: Dict[str, Any], style: Dict[str, str]) -> str:
title = features.get("title", "Features")
subtitle = features.get("subtitle", "")
items = features.get("items", [])
cards_jsx = "\n ".join(
f'''<div className="{style["card_bg"]} rounded-xl p-8">
<div className="mb-4 text-3xl">{f.get("icon", "")}</div>
<h3 className="mb-3 text-xl font-semibold {style["text"]}">{f.get("title", "")}</h3>
<p className="{style["muted"]}">{f.get("description", "")}</p>
</div>'''
for f in items
)
return f'''function Features() {{
return (
<section className="{style["section_alt"]} px-6 py-24">
<div className="mx-auto max-w-7xl">
<h2 className="mb-4 text-center text-4xl font-bold {style["text"]}">{title}</h2>
<p className="mx-auto mb-16 max-w-2xl text-center text-lg {style["muted"]}">{subtitle}</p>
<div className="grid gap-8 md:grid-cols-2 lg:grid-cols-3">
{cards_jsx}
</div>
</div>
</section>
);
}}'''
def tsx_testimonials(testimonials: Dict[str, Any], style: Dict[str, str]) -> str:
title = testimonials.get("title", "What Our Customers Say")
items = testimonials.get("items", [])
if not items:
return ""
cards_jsx = "\n ".join(
f'''<div className="rounded-xl border {style["border"]} p-8">
<p className="mb-6 text-lg italic {style["muted"]}">"{t.get("quote", "")}"</p>
<div>
<p className="font-semibold {style["text"]}">{t.get("name", "")}</p>
<p className="text-sm {style["muted"]}">{t.get("title", "")}, {t.get("company", "")}</p>
</div>
</div>'''
for t in items
)
return f'''function Testimonials() {{
return (
<section className="px-6 py-24 {style["bg"]}">
<div className="mx-auto max-w-7xl">
<h2 className="mb-16 text-center text-4xl font-bold {style["text"]}">{title}</h2>
<div className="grid gap-8 md:grid-cols-2 lg:grid-cols-3">
{cards_jsx}
</div>
</div>
</section>
);
}}'''
def tsx_pricing(pricing: Dict[str, Any], style: Dict[str, str]) -> str:
title = pricing.get("title", "Pricing")
plans = pricing.get("plans", [])
if not plans:
return ""
accent = style["accent"]
cards = []
for p in plans:
featured = p.get("featured", False)
border_cls = f"border-2 border-{accent}-500 ring-4 ring-{accent}-500/20" if featured else f"border {style['border']}"
badge = f'\n <div className="absolute -top-3 left-1/2 -translate-x-1/2 rounded-full bg-{accent}-600 px-4 py-1 text-xs font-semibold text-white">Most Popular</div>' if featured else ""
features_jsx = "\n ".join(
f'<li className="flex items-center gap-2 py-2"><span className="text-{accent}-500 font-bold">✓</span> {feat}</li>'
for feat in p.get("features", [])
)
cards.append(f'''<div className="relative rounded-2xl {border_cls} {style["card_bg"]} p-8 text-center">{badge}
<h3 className="mb-2 text-xl font-semibold {style["text"]}">{p.get("name", "")}</h3>
<div className="my-6 text-5xl font-extrabold {style["text"]}">p.get("price", "0")<span className="text-base font-normal {style["muted"]}">/mo</span></div>
<p className="{style["muted"]} mb-6">{p.get("description", "")}</p>
<ul className="mb-8 space-y-1 text-left {style["muted"]}">
{features_jsx}
</ul>
<a href="{p.get("cta_url", "#")}" className="block w-full rounded-lg {style["btn"]} py-3 text-center font-semibold transition-colors">
{p.get("cta_text", "Choose Plan")}
</a>
</div>''')
cards_jsx = "\n ".join(cards)
return f'''function Pricing() {{
return (
<section className="{style["section_alt"]} px-6 py-24">
<div className="mx-auto max-w-5xl">
<h2 className="mb-16 text-center text-4xl font-bold {style["text"]}">{title}</h2>
<div className="grid gap-8 lg:grid-cols-{min(len(plans), 3)}">
{cards_jsx}
</div>
</div>
</section>
);
}}'''
def tsx_cta(cta: Dict[str, Any], style: Dict[str, str]) -> str:
accent = style["accent"]
return f'''function CTASection() {{
return (
<section className="bg-{accent}-600 px-6 py-24 text-center text-white">
<div className="mx-auto max-w-3xl">
<h2 className="mb-4 text-4xl font-bold">{cta.get("headline", "Ready to get started?")}</h2>
<p className="mb-10 text-xl opacity-90">{cta.get("subheadline", "")}</p>
<a href="{cta.get("url", "#")}" className="rounded-lg bg-white px-8 py-3 text-lg font-semibold text-{accent}-600 transition-colors hover:bg-gray-100">
{cta.get("text", "Start Free Trial")}
</a>
</div>
</section>
);
}}'''
def tsx_footer(config: Dict[str, Any], style: Dict[str, str]) -> str:
brand = config.get("brand", "Company")
year = datetime.now().year
footer_text = config.get("footer_text", f"{year} {brand}. All rights reserved.")
return f'''function Footer() {{
return (
<footer className="border-t {style["border"]} {style["bg"]} px-6 py-10 text-center {style["muted"]}">
<p>© {footer_text}</p>
</footer>
);
}}'''
def generate_tsx(config: Dict[str, Any]) -> str:
"""Generate complete Next.js/React TSX landing page with Tailwind CSS."""
style_name = config.get("design_style", "clean-minimal")
style = DESIGN_STYLES.get(style_name, DESIGN_STYLES["clean-minimal"])
components = []
component_names = []
components.append(tsx_nav(config, style))
component_names.append("Navbar")
if config.get("hero"):
components.append(tsx_hero(config["hero"], style))
component_names.append("Hero")
if config.get("features"):
components.append(tsx_features(config["features"], style))
component_names.append("Features")
if config.get("testimonials") and config["testimonials"].get("items"):
components.append(tsx_testimonials(config["testimonials"], style))
component_names.append("Testimonials")
if config.get("pricing") and config["pricing"].get("plans"):
components.append(tsx_pricing(config["pricing"], style))
component_names.append("Pricing")
if config.get("cta"):
components.append(tsx_cta(config["cta"], style))
component_names.append("CTASection")
components.append(tsx_footer(config, style))
component_names.append("Footer")
title = config.get("title", "Landing Page")
meta_desc = config.get("meta_description", "")
page_body = "\n ".join(f"<{name} />" for name in component_names)
all_components = "\n\n".join(components)
return f'''// Generated by Landing Page Scaffolder — {datetime.now().strftime("%Y-%m-%d")}
// Stack: Next.js 14+ App Router, React, Tailwind CSS
// Design style: {style_name}
import type {{ Metadata }} from "next";
export const metadata: Metadata = {{
title: "{title}",
description: "{meta_desc}",
openGraph: {{
title: "{title}",
description: "{meta_desc}",
type: "website",
}},
}};
{all_components}
export default function LandingPage() {{
return (
<main>
{page_body}
</main>
);
}}
'''
# ---------------------------------------------------------------------------
# HTML generators (existing)
# ---------------------------------------------------------------------------
def generate_css(config: Dict[str, Any]) -> str:
"""Generate responsive CSS from config theme."""
theme = config.get("theme", {})
primary = theme.get("primary_color", "#2563eb")
secondary = theme.get("secondary_color", "#1e40af")
bg = theme.get("background", "#ffffff")
text_color = theme.get("text_color", "#1f2937")
font = theme.get("font", "Inter, system-ui, -apple-system, sans-serif")
return f"""
* {{ margin: 0; padding: 0; box-sizing: border-box; }}
body {{ font-family: {font}; color: {text_color}; background: {bg}; line-height: 1.6; }}
.container {{ max-width: 1200px; margin: 0 auto; padding: 0 24px; }}
nav {{ padding: 16px 0; border-bottom: 1px solid #e5e7eb; position: sticky; top: 0; background: {bg}; z-index: 100; }}
nav .container {{ display: flex; justify-content: space-between; align-items: center; }}
.nav-logo {{ font-size: 1.5rem; font-weight: 700; color: {primary}; text-decoration: none; }}
.nav-links {{ display: flex; gap: 24px; list-style: none; }}
.nav-links a {{ text-decoration: none; color: {text_color}; font-weight: 500; }}
.nav-cta {{ background: {primary}; color: white; padding: 8px 20px; border-radius: 6px; text-decoration: none; font-weight: 600; }}
.hero {{ padding: 80px 0; text-align: center; }}
.hero h1 {{ font-size: 3.5rem; font-weight: 800; line-height: 1.1; margin-bottom: 24px; max-width: 800px; margin-left: auto; margin-right: auto; }}
.hero p {{ font-size: 1.25rem; color: #6b7280; max-width: 600px; margin: 0 auto 32px; }}
.hero-cta {{ display: inline-flex; gap: 16px; }}
.btn-primary {{ background: {primary}; color: white; padding: 14px 32px; border-radius: 8px; text-decoration: none; font-weight: 600; font-size: 1.1rem; }}
.btn-secondary {{ background: transparent; color: {primary}; padding: 14px 32px; border-radius: 8px; text-decoration: none; font-weight: 600; font-size: 1.1rem; border: 2px solid {primary}; }}
.features {{ padding: 80px 0; background: #f9fafb; }}
.section-title {{ text-align: center; font-size: 2.5rem; font-weight: 700; margin-bottom: 16px; }}
.section-subtitle {{ text-align: center; color: #6b7280; font-size: 1.1rem; margin-bottom: 48px; max-width: 600px; margin-left: auto; margin-right: auto; }}
.features-grid {{ display: grid; grid-template-columns: repeat(auto-fit, minmax(300px, 1fr)); gap: 32px; }}
.feature-card {{ background: white; padding: 32px; border-radius: 12px; box-shadow: 0 1px 3px rgba(0,0,0,0.1); }}
.feature-icon {{ font-size: 2rem; margin-bottom: 16px; }}
.feature-card h3 {{ font-size: 1.25rem; margin-bottom: 12px; }}
.feature-card p {{ color: #6b7280; }}
.testimonials {{ padding: 80px 0; }}
.testimonials-grid {{ display: grid; grid-template-columns: repeat(auto-fit, minmax(350px, 1fr)); gap: 24px; }}
.testimonial-card {{ padding: 32px; border: 1px solid #e5e7eb; border-radius: 12px; }}
.testimonial-text {{ font-size: 1.1rem; font-style: italic; margin-bottom: 20px; }}
.testimonial-author {{ display: flex; align-items: center; gap: 12px; }}
.author-info strong {{ display: block; }}
.author-info span {{ color: #6b7280; font-size: 0.9rem; }}
.pricing {{ padding: 80px 0; background: #f9fafb; }}
.pricing-grid {{ display: grid; grid-template-columns: repeat(auto-fit, minmax(300px, 1fr)); gap: 24px; max-width: 900px; margin: 0 auto; }}
.pricing-card {{ background: white; padding: 32px; border-radius: 12px; border: 2px solid #e5e7eb; text-align: center; }}
.pricing-card.featured {{ border-color: {primary}; position: relative; }}
.pricing-card.featured::before {{ content: "Most Popular"; position: absolute; top: -12px; left: 50%; transform: translateX(-50%); background: {primary}; color: white; padding: 4px 16px; border-radius: 20px; font-size: 0.8rem; font-weight: 600; }}
.pricing-name {{ font-size: 1.25rem; font-weight: 600; margin-bottom: 8px; }}
.pricing-price {{ font-size: 3rem; font-weight: 800; margin: 16px 0; }}
.pricing-price span {{ font-size: 1rem; font-weight: 400; color: #6b7280; }}
.pricing-features {{ list-style: none; text-align: left; margin: 24px 0; }}
.pricing-features li {{ padding: 8px 0; border-bottom: 1px solid #f3f4f6; }}
.pricing-features li::before {{ content: "\\2713 "; color: {primary}; font-weight: 700; }}
.cta-section {{ padding: 80px 0; text-align: center; background: {primary}; color: white; }}
.cta-section h2 {{ font-size: 2.5rem; margin-bottom: 16px; }}
.cta-section p {{ font-size: 1.1rem; opacity: 0.9; margin-bottom: 32px; }}
.btn-white {{ background: white; color: {primary}; padding: 14px 32px; border-radius: 8px; text-decoration: none; font-weight: 600; font-size: 1.1rem; }}
footer {{ padding: 40px 0; border-top: 1px solid #e5e7eb; color: #6b7280; text-align: center; }}
@media (max-width: 768px) {{
.hero h1 {{ font-size: 2.25rem; }}
.hero-cta {{ flex-direction: column; align-items: center; }}
.nav-links {{ display: none; }}
.features-grid {{ grid-template-columns: 1fr; }}
.pricing-grid {{ grid-template-columns: 1fr; }}
}}
"""
def render_nav(config: Dict[str, Any]) -> str:
brand = escape(config.get("brand", "Brand"))
nav_links = config.get("nav_links", [])
cta = config.get("nav_cta", {"text": "Get Started", "url": "#"})
links = "\n".join(
f'<li><a href="{escape(l.get("url", "#"))}">{escape(l.get("text", ""))}</a></li>'
for l in nav_links
)
return f"""
<nav><div class="container">
<a href="#" class="nav-logo">{brand}</a>
<ul class="nav-links">{links}</ul>
<a href="{escape(cta.get('url', '#'))}" class="nav-cta">{escape(cta.get('text', 'Get Started'))}</a>
</div></nav>"""
def render_hero(hero: Dict[str, Any]) -> str:
h1 = escape(hero.get("headline", "Your Headline Here"))
sub = escape(hero.get("subheadline", ""))
primary_cta = hero.get("primary_cta", {"text": "Get Started", "url": "#"})
secondary_cta = hero.get("secondary_cta", None)
cta_html = f'<a href="{escape(primary_cta.get("url", "#"))}" class="btn-primary">{escape(primary_cta.get("text", "Get Started"))}</a>'
if secondary_cta:
cta_html += f'\n<a href="{escape(secondary_cta.get("url", "#"))}" class="btn-secondary">{escape(secondary_cta.get("text", "Learn More"))}</a>'
return f"""
<section class="hero"><div class="container">
<h1>{h1}</h1>
<p>{sub}</p>
<div class="hero-cta">{cta_html}</div>
</div></section>"""
def render_features(features: Dict[str, Any]) -> str:
title = escape(features.get("title", "Features"))
subtitle = escape(features.get("subtitle", ""))
items = features.get("items", [])
cards = "\n".join(f"""
<div class="feature-card">
<div class="feature-icon">{escape(f.get('icon', ''))}</div>
<h3>{escape(f.get('title', ''))}</h3>
<p>{escape(f.get('description', ''))}</p>
</div>""" for f in items)
return f"""
<section class="features"><div class="container">
<h2 class="section-title">{title}</h2>
<p class="section-subtitle">{subtitle}</p>
<div class="features-grid">{cards}</div>
</div></section>"""
def render_testimonials(testimonials: Dict[str, Any]) -> str:
title = escape(testimonials.get("title", "What Our Customers Say"))
items = testimonials.get("items", [])
if not items:
return ""
cards = "\n".join(f"""
<div class="testimonial-card">
<p class="testimonial-text">"{escape(t.get('quote', ''))}"</p>
<div class="testimonial-author">
<div class="author-info">
<strong>{escape(t.get('name', ''))}</strong>
<span>{escape(t.get('title', ''))}, {escape(t.get('company', ''))}</span>
</div>
</div>
</div>""" for t in items)
return f"""
<section class="testimonials"><div class="container">
<h2 class="section-title">{title}</h2>
<div class="testimonials-grid">{cards}</div>
</div></section>"""
def render_pricing(pricing: Dict[str, Any]) -> str:
title = escape(pricing.get("title", "Pricing"))
plans = pricing.get("plans", [])
if not plans:
return ""
cards = "\n".join(f"""
<div class="pricing-card {'featured' if p.get('featured') else ''}">
<div class="pricing-name">{escape(p.get('name', ''))}</div>
<div class="pricing-price">escape(str(p.get('price', '0')))<span>/mo</span></div>
<p>{escape(p.get('description', ''))}</p>
<ul class="pricing-features">
{"".join(f'<li>{escape(f)}</li>' for f in p.get('features', []))}
</ul>
<a href="{escape(p.get('cta_url', '#'))}" class="btn-primary">{escape(p.get('cta_text', 'Choose Plan'))}</a>
</div>""" for p in plans)
return f"""
<section class="pricing"><div class="container">
<h2 class="section-title">{title}</h2>
<div class="pricing-grid">{cards}</div>
</div></section>"""
def render_cta(cta: Dict[str, Any]) -> str:
return f"""
<section class="cta-section"><div class="container">
<h2>{escape(cta.get('headline', 'Ready to get started?'))}</h2>
<p>{escape(cta.get('subheadline', ''))}</p>
<a href="{escape(cta.get('url', '#'))}" class="btn-white">{escape(cta.get('text', 'Start Free Trial'))}</a>
</div></section>"""
def generate_html(config: Dict[str, Any]) -> str:
"""Generate complete HTML landing page."""
title = escape(config.get("title", "Landing Page"))
css = generate_css(config)
sections = []
sections.append(render_nav(config))
if config.get("hero"):
sections.append(render_hero(config["hero"]))
if config.get("features"):
sections.append(render_features(config["features"]))
if config.get("testimonials"):
sections.append(render_testimonials(config["testimonials"]))
if config.get("pricing"):
sections.append(render_pricing(config["pricing"]))
if config.get("cta"):
sections.append(render_cta(config["cta"]))
sections.append(f"""
<footer><div class="container">
<p>{escape(config.get('footer_text', f'{datetime.now().year} {config.get("brand", "Company")}. All rights reserved.'))}</p>
</div></footer>""")
return f"""<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>{title}</title>
<meta name="description" content="{escape(config.get('meta_description', ''))}">
<style>{css}</style>
</head>
<body>
{"".join(sections)}
</body>
</html>"""
def main():
parser = argparse.ArgumentParser(
description="Generate landing pages as HTML or Next.js TSX with Tailwind CSS"
)
parser.add_argument("input", help="Path to page config JSON")
parser.add_argument(
"--format", choices=["html", "tsx", "json"], default="tsx",
help="Output format: tsx (Next.js + Tailwind), html (standalone), json (metadata)"
)
parser.add_argument("--output", type=str, default=None, help="Output file path")
args = parser.parse_args()
with open(args.input) as f:
config = json.load(f)
if args.format == "json":
output = json.dumps({
"generated_at": datetime.now().isoformat(),
"config": config,
"formats_available": ["html", "tsx"],
"sections": [k for k in ["nav", "hero", "features", "testimonials", "pricing", "cta", "footer"]
if config.get(k) or k in ("nav", "footer")]
}, indent=2)
elif args.format == "tsx":
output = generate_tsx(config)
else:
output = generate_html(config)
if args.output:
with open(args.output, "w") as f:
f.write(output)
print(f"Landing page written to {args.output}")
else:
print(output)
if __name__ == "__main__":
main()
Biến danh sách công việc rời rạc thành kế hoạch ưu tiên theo ma trận Eisenhower, có timeline và phân bổ năng lượng hợp lý.
--- name: lap-ke-hoach description: Biến danh sách công việc rời rạc thành kế hoạch có ưu tiên theo ma trận Eisenhower, timeline và phân bổ năng lượng hợp lý. Dùng khi nói "lên kế hoạch", "quản lý thời gian", "ưu tiên công việc", "MIT". --- # Lên kế hoạch & Quản lý thời gian ## Mục tiêu Biến danh sách công việc rời rạc thành kế hoạch có ưu tiên, timeline và phân bổ năng lượng hợp lý. ## Khi nào dùng - Đầu tuần cần lên kế hoạch - Có quá nhiều việc, không biết bắt đầu từ đâu - Cần review tiến độ và điều chỉnh ưu tiên ## Đầu vào cần cung cấp - Danh sách việc cần làm - Deadline của từng việc (nếu có) - Mức năng lượng dự kiến trong ngày/tuần - Việc nào có thể bỏ hoặc giao người khác ## Quy trình xử lý 1. Phân loại: Khẩn/Quan trọng theo ma trận Eisenhower 2. Ước lượng thời gian thực tế cho mỗi việc 3. Gộp việc tương tự vào cùng khung giờ (batching) 4. Đặt buffer 20% cho việc phát sinh 5. Xác định 3 việc quan trọng nhất trong ngày (MIT) ## Tiêu chuẩn đầu ra - Danh sách ưu tiên rõ ràng theo ngày/tuần - Mỗi việc có: tên, thời lượng dự kiến, deadline, mức ưu tiên - Có phần "việc nên bỏ hoặc hoãn" - Định dạng: markdown bảng hoặc danh sách có thứ tự ## Tránh - Nhồi quá nhiều việc vào một ngày - Bỏ qua việc phục hồi năng lượng (nghỉ, tập thể dục) - Không tính đến thói quen và múi giờ năng suất của người dùng
Hoạch định chiến lược ra mắt sản phẩm hoặc phát hành tính năng, gồm Product Hunt, beta, early access, waitlist, kế hoạch GTM và checklist ra mắt.
---
name: "launch-strategy"
description: "When the user wants to plan a product launch, feature announcement, or release strategy. Also use when the user mentions 'launch,' 'Product Hunt,' 'feature release,' 'announcement,' 'go-to-market,' 'beta launch,' 'early access,' 'waitlist,' 'product update,' 'GTM plan,' 'launch checklist,' or 'launch momentum.' This skill covers phased launches, channel strategy, and ongoing launch momentum."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Launch Strategy
You are an expert in SaaS product launches and feature announcements. Your goal is to help users plan launches that build momentum, capture attention, and convert interest into users.
## Before Starting
**Check for product marketing context first:**
If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
---
## Core Philosophy
→ See references/launch-frameworks-and-checklists.md for details
## Task-Specific Questions
1. What are you launching? (New product, major feature, minor update)
2. What's your current audience size and engagement?
3. What owned channels do you have? (Email list size, blog traffic, community)
4. What's your timeline for launch?
5. Have you launched before? What worked/didn't work?
6. Are you considering Product Hunt? What's your preparation status?
---
## Proactive Triggers
Proactively offer launch planning when:
1. **Feature ship date mentioned** — When an engineering delivery date is discussed, immediately ask about the launch plan; shipping without a marketing plan is a missed opportunity.
2. **Waitlist or early access mentioned** — Offer to design the full phased launch funnel from alpha through full GA, not just the landing page.
3. **Product Hunt consideration** — Any mention of Product Hunt should trigger the full PH strategy section including pre-launch relationship building timeline.
4. **Post-launch silence** — If a user launched recently but hasn't followed up with momentum content, proactively suggest the post-launch marketing actions (comparison pages, roundup email, interactive demo).
5. **Pricing change planned** — Pricing updates are a launch opportunity; offer to build an announcement campaign treating it as a product update.
---
## Output Artifacts
| Artifact | Format | Description |
|----------|--------|-------------|
| Launch Plan | Markdown doc | Phase-by-phase plan with owners, dates, channels, and success metrics |
| ORB Channel Map | Table | Owned/Rented/Borrowed channel strategy with tactics per channel |
| Launch Day Checklist | Checklist | Complete day-of execution checklist with time-boxed actions |
| Product Hunt Brief | Markdown doc | Listing copy, asset specs, pre-launch timeline, engagement playbook |
| Post-Launch Momentum Plan | Bulleted list | 30-day post-launch actions to sustain and compound the launch |
---
## Communication
Launch plans should be concrete, time-bound, and channel-specific — no vague "post on social media" recommendations. Every output should specify who does what and when. Reference `marketing-context` to ensure the launch narrative matches ICP language and positioning before drafting any copy. Quality bar: a launch plan is only complete when it covers all three ORB channel types and includes both launch-day and post-launch actions.
---
## Related Skills
- **email-sequence** — USE for building the launch announcement and post-launch onboarding email sequences; NOT as a substitute for the full channel strategy.
- **social-content** — USE for drafting the specific social posts and threads for launch day; NOT for channel selection strategy.
- **paid-ads** — USE when the launch plan includes a paid amplification component; NOT for organic launch-only strategies.
- **content-strategy** — USE when the launch requires a sustained content program (blog posts, case studies) in the weeks after; NOT for single-day launch execution.
- **pricing-strategy** — USE when the launch involves a pricing change or new tier introduction; NOT for feature-only launches.
- **marketing-context** — USE as foundation to align launch messaging with ICP and brand voice; always load first.
FILE:references/launch-frameworks-and-checklists.md
# launch-strategy reference
## Core Philosophy
The best companies don't just launch once—they launch again and again. Every new feature, improvement, and update is an opportunity to capture attention and engage your audience.
A strong launch isn't about a single moment. It's about:
- Getting your product into users' hands early
- Learning from real feedback
- Making a splash at every stage
- Building momentum that compounds over time
---
## The ORB Framework
Structure your launch marketing across three channel types. Everything should ultimately lead back to owned channels.
### Owned Channels
You own the channel (though not the audience). Direct access without algorithms or platform rules.
**Examples:**
- Email list
- Blog
- Podcast
- Branded community (Slack, Discord)
- Website/product
**Why they matter:**
- Get more effective over time
- No algorithm changes or pay-to-play
- Direct relationship with audience
- Compound value from content
**Start with 1-2 based on audience:**
- Industry lacks quality content → Start a blog
- People want direct updates → Focus on email
- Engagement matters → Build a community
**Example - Superhuman:**
Built demand through an invite-only waitlist and one-on-one onboarding sessions. Every new user got a 30-minute live demo. This created exclusivity, FOMO, and word-of-mouth—all through owned relationships. Years later, their original onboarding materials still drive engagement.
### Rented Channels
Platforms that provide visibility but you don't control. Algorithms shift, rules change, pay-to-play increases.
**Examples:**
- Social media (Twitter/X, LinkedIn, Instagram)
- App stores and marketplaces
- YouTube
- Reddit
**How to use correctly:**
- Pick 1-2 platforms where your audience is active
- Use them to drive traffic to owned channels
- Don't rely on them as your only strategy
**Example - Notion:**
Hacked virality through Twitter, YouTube, and Reddit where productivity enthusiasts were active. Encouraged community to share templates and workflows. But they funneled all visibility into owned assets—every viral post led to signups, then targeted email onboarding.
**Platform-specific tactics:**
- Twitter/X: Threads that spark conversation → link to newsletter
- LinkedIn: High-value posts → lead to gated content or email signup
- Marketplaces (Shopify, Slack): Optimize listing → drive to site for more
Rented channels give speed, not stability. Capture momentum by bringing users into your owned ecosystem.
### Borrowed Channels
Tap into someone else's audience to shortcut the hardest part—getting noticed.
**Examples:**
- Guest content (blog posts, podcast interviews, newsletter features)
- Collaborations (webinars, co-marketing, social takeovers)
- Speaking engagements (conferences, panels, virtual summits)
- Influencer partnerships
**Be proactive, not passive:**
1. List industry leaders your audience follows
2. Pitch win-win collaborations
3. Use tools like SparkToro or Listen Notes to find audience overlap
4. Set up affiliate/referral incentives
**Example - TRMNL:**
Sent a free e-ink display to YouTuber Snazzy Labs—not a paid sponsorship, just hoping he'd like it. He created an in-depth review that racked up 500K+ views and drove $500K+ in sales. They also set up an affiliate program for ongoing promotion.
Borrowed channels give instant credibility, but only work if you convert borrowed attention into owned relationships.
---
## Five-Phase Launch Approach
Launching isn't a one-day event. It's a phased process that builds momentum.
### Phase 1: Internal Launch
Gather initial feedback and iron out major issues before going public.
**Actions:**
- Recruit early users one-on-one to test for free
- Collect feedback on usability gaps and missing features
- Ensure prototype is functional enough to demo (doesn't need to be production-ready)
**Goal:** Validate core functionality with friendly users.
### Phase 2: Alpha Launch
Put the product in front of external users in a controlled way.
**Actions:**
- Create landing page with early access signup form
- Announce the product exists
- Invite users individually to start testing
- MVP should be working in production (even if still evolving)
**Goal:** First external validation and initial waitlist building.
### Phase 3: Beta Launch
Scale up early access while generating external buzz.
**Actions:**
- Work through early access list (some free, some paid)
- Start marketing with teasers about problems you solve
- Recruit friends, investors, and influencers to test and share
**Consider adding:**
- Coming soon landing page or waitlist
- "Beta" sticker in dashboard navigation
- Email invites to early access list
- Early access toggle in settings for experimental features
**Goal:** Build buzz and refine product with broader feedback.
### Phase 4: Early Access Launch
Shift from small-scale testing to controlled expansion.
**Actions:**
- Leak product details: screenshots, feature GIFs, demos
- Gather quantitative usage data and qualitative feedback
- Run user research with engaged users (incentivize with credits)
- Optionally run product/market fit survey to refine messaging
**Expansion options:**
- Option A: Throttle invites in batches (5-10% at a time)
- Option B: Invite all users at once under "early access" framing
**Goal:** Validate at scale and prepare for full launch.
### Phase 5: Full Launch
Open the floodgates.
**Actions:**
- Open self-serve signups
- Start charging (if not already)
- Announce general availability across all channels
**Launch touchpoints:**
- Customer emails
- In-app popups and product tours
- Website banner linking to launch assets
- "New" sticker in dashboard navigation
- Blog post announcement
- Social posts across platforms
- Product Hunt, BetaList, Hacker News, etc.
**Goal:** Maximum visibility and conversion to paying users.
---
## Product Hunt Launch Strategy
Product Hunt can be powerful for reaching early adopters, but it's not magic—it requires preparation.
### Pros
- Exposure to tech-savvy early adopter audience
- Credibility bump (especially if Product of the Day)
- Potential PR coverage and backlinks
### Cons
- Very competitive to rank well
- Short-lived traffic spikes
- Requires significant pre-launch planning
### How to Launch Successfully
**Before launch day:**
1. Build relationships with influential supporters, content hubs, and communities
2. Optimize your listing: compelling tagline, polished visuals, short demo video
3. Study successful launches to identify what worked
4. Engage in relevant communities—provide value before pitching
5. Prepare your team for all-day engagement
**On launch day:**
1. Treat it as an all-day event
2. Respond to every comment in real-time
3. Answer questions and spark discussions
4. Encourage your existing audience to engage
5. Direct traffic back to your site to capture signups
**After launch day:**
1. Follow up with everyone who engaged
2. Convert Product Hunt traffic into owned relationships (email signups)
3. Continue momentum with post-launch content
### Case Studies
**SavvyCal** (Scheduling tool):
- Optimized landing page and onboarding before launch
- Built relationships with productivity/SaaS influencers in advance
- Responded to every comment on launch day
- Result: #2 Product of the Month
**Reform** (Form builder):
- Studied successful launches and applied insights
- Crafted clear tagline, polished visuals, demo video
- Engaged in communities before launch (provided value first)
- Treated launch as all-day engagement event
- Directed traffic to capture signups
- Result: #1 Product of the Day
---
## Post-Launch Product Marketing
Your launch isn't over when the announcement goes live. Now comes adoption and retention work.
### Immediate Post-Launch Actions
**Educate new users:**
Set up automated onboarding email sequence introducing key features and use cases.
**Reinforce the launch:**
Include announcement in your weekly/biweekly/monthly roundup email to catch people who missed it.
**Differentiate against competitors:**
Publish comparison pages highlighting why you're the obvious choice.
**Update web pages:**
Add dedicated sections about the new feature/product across your site.
**Offer hands-on preview:**
Create no-code interactive demo (using tools like Navattic) so visitors can explore before signing up.
### Keep Momentum Going
It's easier to build on existing momentum than start from scratch. Every touchpoint reinforces the launch.
---
## Ongoing Launch Strategy
Don't rely on a single launch event. Regular updates and feature rollouts sustain engagement.
### How to Prioritize What to Announce
Use this matrix to decide how much marketing each update deserves:
**Major updates** (new features, product overhauls):
- Full campaign across multiple channels
- Blog post, email campaign, in-app messages, social media
- Maximize exposure
**Medium updates** (new integrations, UI enhancements):
- Targeted announcement
- Email to relevant segments, in-app banner
- Don't need full fanfare
**Minor updates** (bug fixes, small tweaks):
- Changelog and release notes
- Signal that product is improving
- Don't dominate marketing
### Announcement Tactics
**Space out releases:**
Instead of shipping everything at once, stagger announcements to maintain momentum.
**Reuse high-performing tactics:**
If a previous announcement resonated, apply those insights to future updates.
**Keep engaging:**
Continue using email, social, and in-app messaging to highlight improvements.
**Signal active development:**
Even small changelog updates remind customers your product is evolving. This builds retention and word-of-mouth—customers feel confident you'll be around.
---
## Launch Checklist
### Pre-Launch
- [ ] Landing page with clear value proposition
- [ ] Email capture / waitlist signup
- [ ] Early access list built
- [ ] Owned channels established (email, blog, community)
- [ ] Rented channel presence (social profiles optimized)
- [ ] Borrowed channel opportunities identified (podcasts, influencers)
- [ ] Product Hunt listing prepared (if using)
- [ ] Launch assets created (screenshots, demo video, GIFs)
- [ ] Onboarding flow ready
- [ ] Analytics/tracking in place
### Launch Day
- [ ] Announcement email to list
- [ ] Blog post published
- [ ] Social posts scheduled and posted
- [ ] Product Hunt listing live (if using)
- [ ] In-app announcement for existing users
- [ ] Website banner/notification active
- [ ] Team ready to engage and respond
- [ ] Monitor for issues and feedback
### Post-Launch
- [ ] Onboarding email sequence active
- [ ] Follow-up with engaged prospects
- [ ] Roundup email includes announcement
- [ ] Comparison pages published
- [ ] Interactive demo created
- [ ] Gather and act on feedback
- [ ] Plan next launch moment
---
FILE:scripts/launch_readiness_scorer.py
#!/usr/bin/env python3
"""
launch_readiness_scorer.py — Product Launch Readiness Scorer
100% stdlib, no pip installs required.
Usage:
python3 launch_readiness_scorer.py # demo mode
python3 launch_readiness_scorer.py --checklist checklist.json
python3 launch_readiness_scorer.py --checklist checklist.json --json
python3 launch_readiness_scorer.py --export-template > my_checklist.json
checklist.json format:
{
"product": [
{"item": "Beta tested with 10+ users", "status": "done"},
{"item": "Documentation ready", "status": "partial"},
{"item": "Support team trained", "status": "not_started"}
],
"marketing": [...],
"technical": [...]
}
Valid status values: "done" | "partial" | "not_started"
"""
import argparse
import json
import sys
from datetime import datetime, timezone
# ---------------------------------------------------------------------------
# Default checklist template
# ---------------------------------------------------------------------------
DEFAULT_CHECKLIST = {
"product": [
{"item": "Beta tested with real users (≥10)", "status": "done", "weight": 3},
{"item": "Core user journey validated end-to-end", "status": "done", "weight": 3},
{"item": "Known P0/P1 bugs resolved", "status": "partial", "weight": 3},
{"item": "User-facing documentation complete", "status": "partial", "weight": 2},
{"item": "In-app onboarding / empty states ready", "status": "done", "weight": 2},
{"item": "Support team trained on common Q&A", "status": "not_started", "weight": 2},
{"item": "Pricing finalised and live", "status": "done", "weight": 2},
{"item": "Accessibility basics checked (WCAG AA)", "status": "not_started", "weight": 1},
{"item": "Localisation / i18n ready (if applicable)", "status": "done", "weight": 1},
{"item": "Feedback collection mechanism in place", "status": "partial", "weight": 1},
],
"marketing": [
{"item": "Landing page live and conversion-optimised", "status": "done", "weight": 3},
{"item": "Email announcement list ready (≥100)", "status": "done", "weight": 3},
{"item": "Press / media kit prepared", "status": "partial", "weight": 2},
{"item": "Social media assets created", "status": "done", "weight": 2},
{"item": "Product Hunt / launch platform submission", "status": "not_started", "weight": 2},
{"item": "SEO meta tags and OG images set", "status": "done", "weight": 2},
{"item": "Influencer / community outreach planned", "status": "partial", "weight": 2},
{"item": "Launch-day email sequence scheduled", "status": "not_started", "weight": 2},
{"item": "Paid ads creative prepared (if applicable)", "status": "not_started", "weight": 1},
{"item": "Referral / viral loop mechanism designed", "status": "not_started", "weight": 1},
],
"technical": [
{"item": "Production monitoring & alerting active", "status": "done", "weight": 3},
{"item": "Load / performance tested at 5× expected", "status": "partial", "weight": 3},
{"item": "Rollback plan documented and rehearsed", "status": "not_started", "weight": 3},
{"item": "Database backups verified and automated", "status": "done", "weight": 2},
{"item": "CDN / caching configured", "status": "done", "weight": 2},
{"item": "Error tracking (Sentry/similar) live", "status": "done", "weight": 2},
{"item": "SSL / HTTPS confirmed on all endpoints", "status": "done", "weight": 2},
{"item": "Analytics events firing correctly", "status": "partial", "weight": 2},
{"item": "Rate limiting / DDoS protection in place", "status": "partial", "weight": 2},
{"item": "Feature flags configured for safe rollout", "status": "not_started", "weight": 1},
],
}
CATEGORY_META = {
"product": {"emoji": "🛠 ", "label": "Product Readiness"},
"marketing": {"emoji": "📣 ", "label": "Marketing Readiness"},
"technical": {"emoji": "⚙️ ", "label": "Technical Readiness"},
}
STATUS_WEIGHTS = {
"done": 1.0,
"partial": 0.5,
"not_started": 0.0,
}
BLOCKERS_THRESHOLD = 0.0 # not_started items with weight ≥3 are blockers
# ---------------------------------------------------------------------------
# Core scoring
# ---------------------------------------------------------------------------
def score_category(items: list) -> dict:
"""Score a single category 0-100 using weighted item scores."""
if not items:
return {"score": 0, "items": [], "blockers": []}
total_weight = 0
earned_weight = 0
blockers = []
scored_items = []
for it in items:
raw_status = it.get("status", "not_started").strip().lower()
status = raw_status if raw_status in STATUS_WEIGHTS else "not_started"
weight = it.get("weight", 1)
sw = STATUS_WEIGHTS[status]
earned = sw * weight
total_weight += weight
earned_weight += earned
scored_items.append({
"item": it["item"],
"status": status,
"weight": weight,
"points_earned": earned,
"points_max": weight,
})
if status == "not_started" and weight >= 3:
blockers.append(it["item"])
score = round((earned_weight / total_weight) * 100) if total_weight > 0 else 0
return {
"score": score,
"score_label": _score_label(score),
"items": scored_items,
"blockers": blockers,
"items_done": sum(1 for i in scored_items if i["status"] == "done"),
"items_partial": sum(1 for i in scored_items if i["status"] == "partial"),
"items_pending": sum(1 for i in scored_items if i["status"] == "not_started"),
"total_items": len(scored_items),
}
def score_readiness(checklist: dict) -> dict:
"""Score all categories and produce an overall launch readiness result."""
categories = {}
all_scores = []
all_blockers = []
for cat, items in checklist.items():
result = score_category(items)
categories[cat] = result
all_scores.append(result["score"])
all_blockers.extend(result["blockers"])
overall = round(sum(all_scores) / len(all_scores)) if all_scores else 0
return {
"overall": {
"score": overall,
"score_label": _score_label(overall),
"launch_decision": _launch_decision(overall, all_blockers),
"blockers": all_blockers,
"generated_at": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
},
"categories": {
cat: {**CATEGORY_META.get(cat, {"emoji": "📋", "label": cat.title()}),
**res}
for cat, res in categories.items()
},
"action_plan": _action_plan(categories),
}
def _launch_decision(score: int, blockers: list) -> str:
if blockers:
return f"⛔ NOT READY — {len(blockers)} blocker(s) must be resolved before launch."
if score >= 80:
return "✅ LAUNCH READY — all categories are in good shape."
if score >= 60:
return "🟡 CONDITIONAL — address partial items but launch is defensible."
if score >= 40:
return "🟠 CAUTION — significant gaps; soft launch / waitlist recommended."
return "🔴 NOT READY — major preparation required across multiple areas."
def _action_plan(categories: dict) -> list:
"""Build a prioritised action list: blockers first, then by score ascending."""
actions = []
for cat, res in categories.items():
label = CATEGORY_META.get(cat, {}).get("label", cat.title())
for bl in res.get("blockers", []):
actions.append({
"priority": "🚨 BLOCKER",
"category": label,
"action": bl,
})
for cat, res in sorted(categories.items(), key=lambda x: x[1]["score"]):
label = CATEGORY_META.get(cat, {}).get("label", cat.title())
for it in res.get("items", []):
if it["status"] == "partial":
actions.append({
"priority": "⚠️ PARTIAL",
"category": label,
"action": f"Complete: {it['item']}",
})
return actions[:15] # top 15 actions
def _score_label(s: int) -> str:
if s >= 90: return "Excellent"
if s >= 75: return "Good"
if s >= 60: return "Fair"
if s >= 40: return "Poor"
return "Critical"
# ---------------------------------------------------------------------------
# Pretty-print
# ---------------------------------------------------------------------------
def pretty_print(result: dict) -> None:
ov = result["overall"]
print("\n" + "=" * 65)
print(" 🚀 LAUNCH READINESS SCORER")
print("=" * 65)
print(f"\n Overall Score : {ov['score']}/100 ({ov['score_label']})")
print(f" Launch Decision : {ov['launch_decision']}")
if ov["blockers"]:
print(f"\n 🚨 BLOCKERS ({len(ov['blockers'])}):")
for b in ov["blockers"]:
print(f" • {b}")
print(f"\n{'─'*65}")
print(f" {'CATEGORY':<30} {'SCORE':>6} {'DONE':>5} {'PARTIAL':>7} {'PENDING':>7}")
print(f"{'─'*65}")
for cat, res in result["categories"].items():
bar = "█" * (res["score"] // 10) + "░" * (10 - res["score"] // 10)
print(f" {res['emoji']} {res['label']:<27} {res['score']:>5}/100 "
f"{res['items_done']:>5} {res['items_partial']:>7} {res['items_pending']:>7} {bar}")
print(f"\n{'─'*65}")
print(f" 🗂 CATEGORY DETAILS\n")
for cat, res in result["categories"].items():
print(f" {res['emoji']} {res['label']} — {res['score']}/100 ({res['score_label']})")
for it in res["items"]:
icon = {"done": "✅", "partial": "🔶", "not_started": "⬜"}.get(it["status"], "⬜")
print(f" {icon} [{it['status']:<11}] (w={it['weight']}) {it['item']}")
print()
ap = result["action_plan"]
if ap:
print(f" 📋 ACTION PLAN (top {len(ap)} items)\n")
for i, a in enumerate(ap, 1):
print(f" {i:>2}. {a['priority']} [{a['category']}] {a['action']}")
print(f"\n Generated: {ov['generated_at']}")
print()
# ---------------------------------------------------------------------------
# CLI
# ---------------------------------------------------------------------------
def parse_args():
parser = argparse.ArgumentParser(
description="Score product launch readiness across categories (stdlib only).",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("--checklist", type=str, default=None,
help="Path to JSON checklist file")
parser.add_argument("--json", action="store_true",
help="Output results as JSON")
parser.add_argument("--export-template", action="store_true",
help="Print the default checklist template as JSON and exit")
return parser.parse_args()
def main():
args = parse_args()
if args.export_template:
print(json.dumps(DEFAULT_CHECKLIST, indent=2))
return
if args.checklist:
with open(args.checklist) as f:
checklist = json.load(f)
else:
print("🔬 DEMO MODE — using embedded sample checklist\n")
checklist = DEFAULT_CHECKLIST
result = score_readiness(checklist)
if args.json:
print(json.dumps(result, indent=2))
else:
pretty_print(result)
if __name__ == "__main__":
main()
Tạo và duy trì tài liệu ngữ cảnh marketing (giọng thương hiệu, đối tượng, ICP, định vị) để các skill marketing khác đọc trước khi làm việc.
---
name: "marketing-context"
description: "Create and maintain the marketing context document that all marketing skills read before starting. Use when the user mentions 'marketing context,' 'brand voice,' 'set up context,' 'target audience,' 'ICP,' 'style guide,' 'who is my customer,' 'positioning,' or wants to avoid repeating foundational information across marketing tasks. Run this at the start of any new project before using other marketing skills."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Marketing Context
You are an expert product marketer. Your goal is to capture the foundational positioning, messaging, and brand context that every other marketing skill needs — so users never repeat themselves.
The document is stored at `.agents/marketing-context.md` (or `marketing-context.md` in the project root).
## How This Skill Works
### Mode 1: Auto-Draft from Codebase
Study the repo — README, landing pages, marketing copy, about pages, package.json, existing docs — and draft a V1. The user reviews, corrects, and fills gaps. This is faster than starting from scratch.
### Mode 2: Guided Interview
Walk through each section conversationally, one at a time. Don't dump all questions at once.
### Mode 3: Update Existing
Read the current context, summarize what's captured, and ask which sections need updating.
Most users prefer Mode 1. After presenting the draft, ask: *"What needs correcting? What's missing?"*
---
## Sections to Capture
### 1. Product Overview
- One-line description
- What it does (2-3 sentences)
- Product category (the "shelf" — how customers search for you)
- Product type (SaaS, marketplace, e-commerce, service)
- Business model and pricing
### 2. Target Audience
- Target company type (industry, size, stage)
- Target decision-makers (roles, departments)
- Primary use case (the main problem you solve)
- Jobs to be done (2-3 things customers "hire" you for)
- Specific use cases or scenarios
### 3. Personas
For each stakeholder involved in buying:
- Role (User, Champion, Decision Maker, Financial Buyer, Technical Influencer)
- What they care about, their challenge, the value you promise them
### 4. Problems & Pain Points
- Core challenge customers face before finding you
- Why current solutions fall short
- What it costs them (time, money, opportunities)
- Emotional tension (stress, fear, doubt)
### 5. Competitive Landscape
- **Direct competitors**: Same solution, same problem
- **Secondary competitors**: Different solution, same problem
- **Indirect competitors**: Conflicting approach entirely
- How each falls short for customers
### 6. Differentiation
- Key differentiators (capabilities alternatives lack)
- How you solve it differently
- Why that's better (benefits, not features)
- Why customers choose you over alternatives
### 7. Objections & Anti-Personas
- Top 3 objections heard in sales + how to address each
- Who is NOT a good fit (anti-persona)
### 8. Switching Dynamics (JTBD Four Forces)
- **Push**: Frustrations driving them away from current solution
- **Pull**: What attracts them to you
- **Habit**: What keeps them stuck with current approach
- **Anxiety**: What worries them about switching
### 9. Customer Language (Verbatim)
- How customers describe the problem in their own words
- How they describe your solution in their own words
- Words and phrases TO use
- Words and phrases to AVOID
- Glossary of product-specific terms
### 10. Brand Voice
- Tone (professional, casual, playful, authoritative)
- Communication style (direct, conversational, technical)
- Brand personality (3-5 adjectives)
- Voice DO's and DON'T's
### 11. Style Guide
- Grammar and mechanics rules
- Capitalization conventions
- Formatting standards
- Preferred terminology
### 12. Proof Points
- Key metrics or results to cite
- Notable customers / logos
- Testimonial snippets (verbatim)
- Main value themes with supporting evidence
### 13. Content & SEO Context
- Target keywords (organized by topic cluster)
- Internal links map (key pages, anchor text)
- Writing examples (3-5 exemplary pieces)
- Content tone and length preferences
### 14. Goals
- Primary business goal
- Key conversion action (what you want people to do)
- Current metrics (if known)
---
## Output Template
See `templates/marketing-context-template.md` for the full template.
---
## Tips
- **Be specific**: Ask "What's the #1 frustration that brings them to you?" not "What problem do they solve?"
- **Capture exact words**: Customer language beats polished descriptions
- **Ask for examples**: "Can you give me an example?" unlocks better answers
- **Validate as you go**: Summarize each section and confirm before moving on
- **Skip what doesn't apply**: Not every product needs all sections
---
## Proactive Triggers
Surface these without being asked:
- **Missing customer language section** → "Without verbatim customer phrases, copy will sound generic. Can you share 3-5 quotes from customers describing their problem?"
- **No competitive landscape defined** → "Every marketing skill performs better with competitor context. Who are the top 3 alternatives your customers consider?"
- **Brand voice undefined** → "Without voice guidelines, every skill will sound different. Let's define 3-5 adjectives that capture your brand."
- **Context older than 6 months** → "Your marketing context was last updated [date]. Positioning may have shifted — review recommended."
- **No proof points** → "Marketing without proof points is opinion. What metrics, logos, or testimonials can we reference?"
## Output Artifacts
| When you ask for... | You get... |
|---------------------|------------|
| "Set up marketing context" | Guided interview → complete `marketing-context.md` |
| "Auto-draft from codebase" | Codebase scan → V1 draft for review |
| "Update positioning" | Targeted update of differentiation + competitive sections |
| "Add customer quotes" | Customer language section populated with verbatim phrases |
| "Review context freshness" | Staleness audit with recommended updates |
## Communication
All output passes quality verification:
- Self-verify: source attribution, assumption audit, confidence scoring
- Output format: Bottom Line → What (with confidence) → Why → How to Act
- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed.
## Related Skills
- **marketing-ops**: Routes marketing questions to the right skill — reads this context first.
- **copywriting**: For landing page and web copy. Reads brand voice + customer language from this context.
- **content-strategy**: For planning what content to create. Reads target keywords + personas from this context.
- **marketing-strategy-pmm**: For positioning and GTM strategy. Reads competitive landscape from this context.
- **cs-onboard** (C-Suite): For company-level context. This skill is marketing-specific — complements, not replaces, company-context.md.
FILE:scripts/context_validator.py
#!/usr/bin/env python3
"""Validate marketing context completeness — scores 0-100."""
import json
import re
import sys
from pathlib import Path
SECTIONS = {
"Product Overview": {"required": True, "weight": 10, "markers": ["one-liner", "what it does", "product category", "business model"]},
"Target Audience": {"required": True, "weight": 12, "markers": ["target compan", "decision-maker", "use case", "jobs to be done"]},
"Personas": {"required": False, "weight": 5, "markers": ["persona", "champion", "decision maker"]},
"Problems & Pain Points": {"required": True, "weight": 10, "markers": ["core problem", "fall short", "cost", "tension"]},
"Competitive Landscape": {"required": True, "weight": 10, "markers": ["direct", "competitor", "secondary"]},
"Differentiation": {"required": True, "weight": 10, "markers": ["differentiator", "differently", "why customers choose"]},
"Objections": {"required": False, "weight": 5, "markers": ["objection", "response", "anti-persona"]},
"Switching Dynamics": {"required": False, "weight": 5, "markers": ["push", "pull", "habit", "anxiety"]},
"Customer Language": {"required": True, "weight": 10, "markers": ["verbatim", "words to use", "words to avoid"]},
"Brand Voice": {"required": True, "weight": 8, "markers": ["tone", "style", "personality"]},
"Style Guide": {"required": False, "weight": 3, "markers": ["grammar", "capitalization", "formatting"]},
"Proof Points": {"required": True, "weight": 7, "markers": ["metric", "customer", "testimonial"]},
"Content & SEO": {"required": False, "weight": 3, "markers": ["keyword", "internal link"]},
"Goals": {"required": True, "weight": 2, "markers": ["business goal", "conversion"]}
}
def validate_context(content: str) -> dict:
"""Validate marketing context file and return score."""
content_lower = content.lower()
results = {"sections": {}, "score": 0, "max_score": 100, "missing_required": [], "missing_optional": [], "warnings": []}
total_weight = sum(s["weight"] for s in SECTIONS.values())
earned = 0
for name, config in SECTIONS.items():
section_present = name.lower().replace("& ", "").replace(" ", " ") in content_lower or any(
m in content_lower for m in config["markers"][:2]
)
markers_found = sum(1 for m in config["markers"] if m in content_lower)
markers_total = len(config["markers"])
has_placeholder = bool(re.search(r'\[.*?\]', content[content_lower.find(name.lower()):content_lower.find(name.lower()) + 500] if name.lower() in content_lower else ""))
if section_present and markers_found > 0:
completeness = markers_found / markers_total
if has_placeholder and completeness < 0.5:
completeness *= 0.5 # Penalize unfilled templates
section_score = round(config["weight"] * completeness)
earned += section_score
status = "complete" if completeness >= 0.75 else "partial"
else:
section_score = 0
status = "missing"
if config["required"]:
results["missing_required"].append(name)
else:
results["missing_optional"].append(name)
results["sections"][name] = {
"status": status,
"markers_found": markers_found,
"markers_total": markers_total,
"score": section_score,
"max_score": config["weight"],
"required": config["required"]
}
results["score"] = round((earned / total_weight) * 100)
# Warnings
if "verbatim" not in content_lower and '"' not in content:
results["warnings"].append("No verbatim customer quotes found — copy will sound generic")
if not re.search(r'\d+%|\$\d+|\d+ customer', content_lower):
results["warnings"].append("No metrics or proof points with numbers found")
if "last updated" in content_lower:
date_match = re.search(r'last updated:?\s*(\d{4}-\d{2}-\d{2})', content_lower)
if date_match:
from datetime import datetime
try:
updated = datetime.strptime(date_match.group(1), "%Y-%m-%d")
age_days = (datetime.now() - updated).days
if age_days > 180:
results["warnings"].append(f"Context is {age_days} days old — review recommended (>180 days)")
except ValueError:
pass
return results
def print_report(results: dict):
"""Print human-readable validation report."""
print(f"\n{'='*50}")
print(f"MARKETING CONTEXT VALIDATION")
print(f"{'='*50}")
print(f"\nOverall Score: {results['score']}/100")
print(f"{'🟢 Strong' if results['score'] >= 80 else '🟡 Needs Work' if results['score'] >= 50 else '🔴 Incomplete'}")
print(f"\n{'─'*50}")
print(f"{'Section':<25} {'Status':<10} {'Score':<10}")
print(f"{'─'*50}")
for name, data in results["sections"].items():
icon = {"complete": "✅", "partial": "⚠️", "missing": "❌"}[data["status"]]
req = " *" if data["required"] else ""
print(f"{icon} {name:<23} {data['status']:<10} {data['score']}/{data['max_score']}{req}")
if results["missing_required"]:
print(f"\n🔴 Missing Required Sections:")
for s in results["missing_required"]:
print(f" → {s}")
if results["missing_optional"]:
print(f"\n🟡 Missing Optional Sections:")
for s in results["missing_optional"]:
print(f" → {s}")
if results["warnings"]:
print(f"\n⚠️ Warnings:")
for w in results["warnings"]:
print(f" → {w}")
print(f"\n* = required section")
print(f"{'='*50}")
def main():
import argparse
parser = argparse.ArgumentParser(
description="Validates marketing context completeness. "
"Scores 0-100 based on required and optional section coverage."
)
parser.add_argument(
"file", nargs="?", default=None,
help="Path to a marketing context markdown file. "
"If omitted, runs demo with embedded sample data."
)
parser.add_argument(
"--json", action="store_true",
help="Also output results as JSON."
)
args = parser.parse_args()
if args.file:
filepath = Path(args.file)
if not filepath.exists():
print(f"Error: File not found: {filepath}", file=sys.stderr)
sys.exit(1)
content = filepath.read_text()
else:
# Demo with sample data
content = """# Marketing Context
*Last updated: 2026-01-15*
## Product Overview
**One-liner:** AI-powered mobility analysis for elderly care
**What it does:** Smartphone-based fall risk assessment using computer vision
**Product category:** HealthTech / Digital Health
**Business model:** SaaS, per-facility licensing
## Target Audience
**Target companies:** Care facilities, nursing homes, 50+ beds
**Decision-makers:** Facility directors, quality managers
**Primary use case:** Automated fall risk assessment replacing manual observation
**Jobs to be done:**
- Reduce fall incidents by identifying high-risk residents
- Meet regulatory documentation requirements efficiently
- Give care staff actionable mobility insights
## Problems & Pain Points
**Core problem:** Manual fall risk assessment is subjective, time-consuming, and inconsistent
**Why alternatives fall short:**
- Manual observation takes 30+ minutes per resident
- Paper-based assessments are completed once per quarter at best
**What it costs them:** Falls cost €8,000-12,000 per incident, plus liability
**Emotional tension:** Staff fear missing warning signs, blame after incidents
## Competitive Landscape
**Direct:** Traditional gait labs — $50K+ hardware, need trained staff
**Secondary:** Wearable sensors — low compliance, residents remove them
**Indirect:** Manual observation — subjective, inconsistent
## Differentiation
**Key differentiators:**
- Uses standard smartphone (no special hardware)
- AI-powered analysis (objective, repeatable)
**Why customers choose us:** Fast, affordable, no hardware investment
## Customer Language
**How they describe the problem:**
- "We never know who's going to fall next"
- "The documentation takes forever"
**Words to use:** mobility analysis, fall prevention, care quality
**Words to avoid:** surveillance, monitoring, tracking
## Brand Voice
**Tone:** Professional, empathetic, evidence-based
**Personality:** Trustworthy, innovative, caring
## Proof Points
**Metrics:**
- 80+ care facilities served
- 30% reduction in fall incidents (pilot data)
**Customers:** Major care facility chains in Germany
## Goals
**Business goal:** Expand to 200+ facilities, enter Spain and Netherlands
**Conversion action:** Book a demo
"""
print("[Using embedded sample data — pass a file path for real validation]")
results = validate_context(content)
print_report(results)
if args.json:
print(f"\n{json.dumps(results, indent=2)}")
if __name__ == "__main__":
main()
FILE:templates/marketing-context-template.md
# Marketing Context
*Last updated: [date]*
## Product Overview
**One-liner:** [What you do in one sentence]
**What it does:** [2-3 sentences]
**Product category:** [The "shelf" — how customers search for you]
**Product type:** [SaaS, marketplace, e-commerce, service]
**Business model:** [Pricing model and range]
## Target Audience
**Target companies:** [Industry, size, stage]
**Decision-makers:** [Roles, departments]
**Primary use case:** [The main problem you solve]
**Jobs to be done:**
- [Job 1]
- [Job 2]
- [Job 3]
**Use cases:**
- [Scenario 1]
- [Scenario 2]
## Personas
| Persona | Role | Cares about | Challenge | Value we promise |
|---------|------|-------------|-----------|------------------|
| [Name] | User | | | |
| [Name] | Champion | | | |
| [Name] | Decision Maker | | | |
| [Name] | Financial Buyer | | | |
## Problems & Pain Points
**Core problem:** [What customers face before finding you]
**Why alternatives fall short:**
- [Gap 1]
- [Gap 2]
**What it costs them:** [Time, money, opportunities]
**Emotional tension:** [Stress, fear, doubt]
## Competitive Landscape
| Competitor | Type | How they fall short |
|-----------|------|---------------------|
| [Name] | Direct | [Gap] |
| [Name] | Secondary | [Gap] |
| [Name] | Indirect | [Gap] |
## Differentiation
**Key differentiators:**
- [Differentiator 1]
- [Differentiator 2]
**How we do it differently:** [Approach]
**Why that's better:** [Benefits]
**Why customers choose us:** [Decision drivers]
## Objections
| Objection | Response |
|-----------|----------|
| "[Objection 1]" | [How to address] |
| "[Objection 2]" | [How to address] |
| "[Objection 3]" | [How to address] |
**Anti-persona (NOT a good fit):** [Who should NOT buy this]
## Switching Dynamics
**Push (away from current):** [Frustrations]
**Pull (toward us):** [Attractions]
**Habit (keeping them stuck):** [Inertia]
**Anxiety (about switching):** [Worries]
## Customer Language
**How they describe the problem:**
- "[verbatim quote]"
- "[verbatim quote]"
**How they describe us:**
- "[verbatim quote]"
- "[verbatim quote]"
**Words to use:** [list]
**Words to avoid:** [list]
| Term | Meaning |
|------|---------|
| [Product term] | [Definition] |
## Brand Voice
**Tone:** [professional, casual, playful, authoritative]
**Style:** [direct, conversational, technical]
**Personality:** [3-5 adjectives]
**Voice DO's:** [list]
**Voice DON'T's:** [list]
## Style Guide
**Grammar:** [Key rules]
**Capitalization:** [Conventions]
**Formatting:** [Standards]
**Preferred terms:** [List]
## Proof Points
**Metrics:**
- [Metric 1]
- [Metric 2]
**Customers:** [Notable logos]
**Testimonials:**
> "[quote]" — [Name, Title, Company]
> "[quote]" — [Name, Title, Company]
| Value Theme | Supporting Proof |
|-------------|-----------------|
| [Theme 1] | [Evidence] |
| [Theme 2] | [Evidence] |
## Content & SEO Context
**Target keywords:**
| Cluster | Primary Keyword | Secondary Keywords | Intent |
|---------|----------------|-------------------|--------|
| [Topic 1] | [keyword] | [kw1, kw2] | [informational/commercial] |
**Internal links map:**
| Page | URL | Use for | Anchor text |
|------|-----|---------|-------------|
| [Page name] | [URL] | [Topic] | [Suggested anchor] |
**Writing examples:**
- [URL or file — what makes it good]
## Goals
**Business goal:** [Primary objective]
**Conversion action:** [What you want people to do]
**Current metrics:** [If known]