Lập kế hoạch và tổng hợp nghiên cứu sản phẩm/người dùng: chọn phương pháp phù hợp, tính độ bão hòa và cỡ mẫu theo độ tin cậy rõ ràng.
---
name: product-research
description: Use when planning and synthesizing product/user research as a method-and-repository discipline — selecting the right method for the goal (generative interviews vs usability test vs concept test vs validation), computing method-based saturation/sample size with an explicit confidence level, or synthesizing coded observations into insights while flagging single-source anecdotes. Never fabricates user insight; an insight requires recurrence across independent participants. Distinct from product-team/ux-researcher-designer (persona/journey artifacts), product-discovery (discovery-sprint planning), and experiment-designer (live A/B) — this is the research-ops method + insight-repository layer.
version: 2.9.0
author: claude-code-skills
license: MIT
tags: [research-ops, product-research, ux-research, jtbd, usability, saturation, insight-synthesis, research-repository]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# product-research
Product / user research as an operational discipline: choosing the right method, sizing it honestly, and synthesizing findings into governed insights. The core rule: **method must match the goal**, and **an insight requires recurrence across independent participants** — a single quote is an anecdote.
## Purpose
Product researchers, ResearchOps teams, and PMs running discovery need method rigor and an insight repository they can trust. This skill structures three decisions:
Three deterministic tools:
1. `study_designer.py` — Maps (research goal × product stage) to an appropriate method and emits a method-matched plan skeleton (objective, participant criteria, guide structure, success criteria). Redirects live A/B to `product-team/experiment-designer`.
2. `saturation_planner.py` — Method-based sample guidance with an explicit **confidence label**: Nielsen problem-discovery (5/segment), Guest et al. thematic saturation (~12), and evaluative coverage. Never claims a prevalence rate from a small-n usability test.
3. `insight_synthesizer.py` — Clusters coded observations by tag, counts distinct participants, ranks by cross-participant recurrence, and flags any candidate below the source threshold as an **ANECDOTE**, never promoting it to an insight.
## When to use
Invoke this skill when:
- You are planning a study and need the method to match the goal (generative vs evaluative vs validation).
- You need a defensible sample size / saturation rationale with a stated confidence.
- You have raw coded observations and need to synthesize insights without over-claiming.
- You are setting up or auditing a research repository and need the insight-vs-observation discipline.
**Do NOT use this skill to**: generate personas / journey maps (use `product-team/ux-researcher-designer`), plan a discovery sprint or validate an opportunity (use `product-team/product-discovery`), design or analyze a live product A/B experiment (use `product-team/experiment-designer`), or do market sizing / surveys (use the `market-research` sibling).
## Workflow
1. **Frame the study** — Fill `assets/research_plan_template.md` (research questions, method rationale, participant criteria, analysis plan, repository tagging scheme).
2. **Pick the method** — Run `study_designer.py --goal {discovery|evaluative|validation} --stage {concept|prototype|beta|live} --profile {b2b-saas|consumer-app|enterprise|marketplace|hardware|platform}`. Honor the redirect if it routes to experiment-designer.
3. **Size it** — Run `saturation_planner.py --method {usability|thematic|evaluative-coverage} --segments N`. Record the confidence label and limits.
4. **Synthesize** — After fielding, code observations and run `insight_synthesizer.py --input observations.json --min-sources 3`. Treat ANECDOTE-flagged clusters as signals to probe, not findings to ship.
5. **File in the repository** — Tag insights to the atomic schema at synthesis time, with their evidence and confidence.
## Scripts
| Script | Purpose | Profiles |
|---|---|---|
| `scripts/study_designer.py` | (goal × stage) → method + plan skeleton | b2b-saas, consumer-app, enterprise, marketplace, hardware, platform |
| `scripts/saturation_planner.py` | Method-based sample guidance + confidence | n/a (method-driven) |
| `scripts/insight_synthesizer.py` | Cluster observations, flag anecdotes | n/a (evidence-driven) |
All three: stdlib-only, `--help`, `--sample`, `--output {human,json}`.
## Onboarding & customization
Run the onboarding questionnaire **once before you start** — it captures your defaults so every tool in this skill is pre-configured. Customization is the point: the answers actually change tool behavior (e.g. the insight source-threshold).
```bash
python3 scripts/onboard.py # interactive (also: --defaults, --set key=value, --reset)
python3 scripts/onboard.py --show # see the questions + current effective config
```
Answers are saved to `~/.config/research-ops/product-research.json` (global) or `./.research-ops/product-research.json` (`--scope project`) and are read automatically by `config_loader.py`. They set the default product **profile**, the **insight source-threshold** (how many independent participants make a finding an insight, not an anecdote), the default **saturation method**, and the **high-stakes** flag. CLI flags always override saved config; `RESEARCH_OPS_NO_CONFIG=1` ignores it.
**The four questions:** product profile · insight source-threshold · saturation method · high-stakes flag.
## Optimize with autoresearch (opt-in)
This skill ships an **isolated, opt-in** bridge to `engineering/autoresearch-agent`. Only when you ask to "optimize the synthesis" / "run a loop" does an autoresearch experiment iteratively refine the coding/clustering of a fixed evidence set so more cross-participant patterns surface. `scripts/ar_evaluator.py` is the ground-truth evaluator; it prints `validated_insights: <int>` (higher is better). It optimizes the **coding**, never fabricates evidence.
```bash
/ar:setup --domain custom --name insight-synthesis \
--target observations.json \
--eval "python3 ar_evaluator.py --target observations.json" \
--metric validated_insights --direction higher
/ar:loop custom/insight-synthesis
```
Isolated: no hard dependency — autoresearch runs only on demand, and the loop edits `observations.json`, never the evaluator.
## References
- `references/research_methods_canon.md` — Portigal *Interviewing Users*; Christensen/Ulwick JTBD; Rohrer's UX-research methods landscape (NN/g); Sauro & Lewis *Quantifying the User Experience*; Goodman/Kuniavsky.
- `references/sampling_and_saturation.md` — Nielsen "test with 5 users"; Guest, Bunce & Johnson saturation; Faulkner on more-than-5; Sauro usability sample size; Braun & Clarke thematic analysis.
- `references/repository_and_synthesis.md` — ResearchOps / atomic research (Tomer Sharon "Polaris"); insight-vs-observation discipline; repository governance; affinity mapping; democratization guardrails.
## Assumptions
- Method selection assumes you can name the goal honestly; if the goal is fuzzy, grill it first (the goal drives everything).
- Saturation guidance is method-based, not a power calculation — usability tests find problems, not prevalence rates.
- The synthesizer counts evidence you provide; coding quality is upstream of it. Garbage tags → garbage clusters.
- The insight threshold (`--min-sources`) defaults to 3; raise it for high-stakes or heterogeneous populations.
## Anti-patterns
- **Mismatching method to goal.** A usability test cannot discover unmet needs; an interview cannot measure task success.
- **Reporting usability problems as percentages.** Small-n tests surface problems, not population rates.
- **Promoting an anecdote to an insight.** One participant is a signal to probe, not a finding.
- **Framing interview questions as feature reactions.** Probe the job-to-be-done and recent real behavior, not hypothetical opinions.
- **Synthesizing without a repository scheme.** Tag at synthesis time, or insights rot unfindable.
## Distinct from
| Neighbor | Scope | Difference |
|---|---|---|
| `product-team/ux-researcher-designer` | Personas, journey maps, usability frameworks tied to design output | That produces **artifacts**; this is **method + repository discipline** |
| `product-team/product-discovery` | Opportunity validation, discovery-sprint planning | That plans **discovery sprints**; this designs and synthesizes the **research** |
| `product-team/experiment-designer` | Live product A/B hypothesis + sample size | That runs **live experiments**; this runs **qualitative/evaluative research** |
| `market-research` (sibling) | Market sizing, surveys, segmentation | That studies **the market**; this studies **users** |
## Quick examples
```bash
python3 scripts/study_designer.py --sample
python3 scripts/saturation_planner.py --method thematic --segments 3
python3 scripts/insight_synthesizer.py --sample --min-sources 3
```
The synthesizer sample correctly promotes "import-confusion" (3 independent participants) to INSIGHT and flags "wants-slack" (1 participant) as an ANECDOTE.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-research-ops` or the orchestrator. Recommended answer + canon citation per question. Never bundled.
1. **"Is this study generative (discover problems) or evaluative (test a solution)?"**
Recommended: name it first — the method follows from the goal.
Canon: Rohrer, *When to Use Which User-Experience Research Methods* (NN/g).
2. **"What's your sample size and saturation rationale — and at what confidence?"**
Recommended: method-based n (5/segment usability; ~12 for thematic saturation), state the confidence.
Canon: Nielsen; Guest, Bunce & Johnson (2006); Faulkner (2003).
3. **"How many independent participants support each insight — or is it a single-source anecdote?"**
Recommended: require recurrence across ≥3 sources before calling it an insight; flag singletons.
Canon: atomic research / ResearchOps; Braun & Clarke thematic analysis.
4. **"Are your interview / usability tasks framed as outcomes (jobs) or as feature reactions?"**
Recommended: frame around the job-to-be-done and recent real behavior, not hypothetical opinion.
Canon: Christensen/Ulwick Jobs-to-be-Done; Portigal *Interviewing Users*.
5. **"Where does this land in the repository, and how is it tagged for reuse?"**
Recommended: tag to the atomic schema at synthesis time, not later.
Canon: Tomer Sharon, *Polaris* / ResearchOps repository practice.
Walk depth-first. Lock 1-2 before opening 3-5. After all are answered, invoke `study_designer.py` → `saturation_planner.py` → (after fielding) `insight_synthesizer.py`.
FILE:assets/research_plan_template.md
# Product Research Plan — Template
> Fill this before running the tools. Method must match the goal. An insight requires
> recurrence across independent participants — a single quote is an anecdote.
## 1. Study identification
- Study name:
- Product / feature:
- Stage: [concept | prototype | beta | live]
- Profile: [b2b-saas | consumer-app | enterprise | marketplace | hardware | platform]
## 2. Goal & questions
- Goal: [discovery (generative) | evaluative | validation]
- Research questions (3-5, answerable, not leading):
- The product decision this informs:
## 3. Method (from `study_designer.py`)
- Recommended method:
- Why it matches the goal:
- (If live A/B → route to product-team/experiment-designer.)
## 4. Participants
- Target segment(s) + screener (screen for the job, not a job title):
- Per-segment recruiting if reporting per segment? [yes/no]
- Exclusions (internal, biased, repeat):
## 5. Sample & saturation (from `saturation_planner.py`)
- Method: [usability | thematic | evaluative-coverage]
- n per segment + total:
- Confidence label + limits:
## 6. Study guide skeleton
1.
2.
3.
4.
5.
## 7. Analysis & synthesis
- Coding / tagging scheme (atomic taxonomy):
- Insight threshold (min distinct participants): ___ (default 3)
- Synthesis tool: `insight_synthesizer.py`
## 8. Repository
- Where insights are filed + tagging taxonomy:
- Evidence linked to each insight? [yes — required]
- Confidence field per insight? [yes — required]
## 9. Confidence statement
- What this study can and cannot support:
FILE:references/repository_and_synthesis.md
# Research Repository and Synthesis
Reference for turning observations into governed insights. Pairs with `insight_synthesizer.py`.
## Observation vs insight
The foundational discipline of ResearchOps is the distinction between an **observation** (a single piece of evidence — one participant did or said one thing) and an **insight** (a pattern that recurs across independent sources and carries an implication). Promoting an observation to an insight because it was vivid or confirmed a prior is the cardinal sin of synthesis. The synthesizer enforces a source threshold: a candidate supported by fewer than the threshold of distinct participants is labeled an ANECDOTE and is never promoted.
## Atomic research
Tomer Sharon's **atomic research** model (and the "Polaris" repository concept) decomposes research into reusable units: *Experiments → Facts (observations) → Insights → Recommendations*. Facts are tagged and stored so that insights can be traced back to evidence and reused across studies. The payoff is a repository where a claim can always be drilled down to the observations that support it — and where the same evidence can support future questions.
## Affinity mapping
The classic synthesis technique is affinity mapping: cluster observations into emergent themes bottom-up, then name the themes. The `insight_synthesizer.py` tool is a deterministic, tag-based proxy for this — it clusters by the codes you assign and ranks by cross-participant recurrence. The human still does the interpretive naming; the tool enforces the counting discipline.
## Repository governance and democratization
As organizations democratize research (PMs and designers running their own studies), the repository becomes the guardrail. Governance practices: a consistent tagging taxonomy, evidence linked to every insight, a confidence field, and a review step before an insight is marked "validated." Without governance, democratized research produces a pile of unsearchable anecdotes; with it, the repository compounds in value.
## Sources
1. Sharon, T., *Validating Product Ideas Through Lean User Research* (Rosenfeld, 2016) and the atomic-research / Polaris model.
2. ResearchOps Community, *Research Repositories* and *Democratization* working-group reports.
3. Braun, V., & Clarke, V., *Thematic Analysis: A Practical Guide* (Sage, 2022).
4. Beyer, H., & Holtzblatt, K., *Contextual Design* (1998) — affinity diagramming.
5. Dovetail / EnjoyHQ practitioner guides on insight repositories and tagging taxonomies.
6. Kaplan, K., *Taxonomy 101* and *Research Repositories* — Nielsen Norman Group.
FILE:references/research_methods_canon.md
# Product Research Methods Canon
Reference for method selection. Pairs with `study_designer.py`.
## The two-axis map
UX/product research methods sort along two axes (Rohrer, NN/g): **attitudinal vs behavioral** (what people say vs what they do) and **qualitative vs quantitative** (why/how vs how-many). The single most important pre-method decision is the **goal**:
- **Generative (discovery)** — you don't yet know the problem. Methods: semi-structured interviews, contextual inquiry, diary studies. Output: themes, unmet needs, jobs-to-be-done.
- **Evaluative** — you have a solution and want to know if it works. Methods: moderated/unmoderated usability tests, concept tests. Output: task-success, severity-rated problems.
- **Validation** — you want to confirm demand/desirability before building. Methods: surveys, preference tests, fake-door tests, and (when live) A/B experiments.
Picking an evaluative method for a generative goal — "let's usability-test our way to product strategy" — is the most common and most expensive error.
## Interviewing discipline
Steve Portigal's *Interviewing Users* is the operative craft reference: ask about **recent, specific, real behavior** ("tell me about the last time you…"), not hypotheticals or opinions ("would you use…"). People are unreliable narrators of their future selves but good storytellers of their past.
## Jobs-to-be-Done
Christensen's and Ulwick's JTBD reframes research around the **progress a person is trying to make** in a circumstance, not their demographics or feature preferences. Outcome-Driven Innovation (Ulwick) operationalizes this into measurable desired outcomes — a bridge between qualitative discovery and quantitative validation.
## Mixed methods
Strong research triangulates: qualitative discovery surfaces hypotheses; quantitative validation sizes them. Sauro & Lewis (*Quantifying the User Experience*) provides the statistical backbone for turning usability observations into defensible metrics (task time, completion, SUS) without over-claiming from small samples.
## Sources
1. Portigal, S., *Interviewing Users*, 2nd ed. (Rosenfeld, 2023).
2. Christensen, Hall, Dillon & Duncan, *Competing Against Luck* (2016) — Jobs-to-be-Done.
3. Ulwick, A., *Jobs to Be Done: Theory to Practice* (2016) — Outcome-Driven Innovation.
4. Rohrer, C., *When to Use Which User-Experience Research Methods* — Nielsen Norman Group.
5. Sauro, J., & Lewis, J., *Quantifying the User Experience*, 2nd ed. (Morgan Kaufmann, 2016).
6. Goodman, Kuniavsky & Moed, *Observing the User Experience*, 2nd ed. (2012).
FILE:references/sampling_and_saturation.md
# Sampling and Saturation
Reference for how many participants. Pairs with `saturation_planner.py`.
## Usability: the "5 users" result
Nielsen and Landauer's model says the proportion of usability problems found with n users is 1 − (1 − p)ⁿ, where p is the average probability that a single user surfaces a given problem (~0.31 in their data). At n = 5, that is ~85% of problems — hence "test with 5 users." Two crucial caveats the planner enforces:
1. **Per segment.** The 5-user result holds *within a homogeneous user group*. If you have distinct segments that behave differently, you need ~5 per segment.
2. **Problems, not rates.** A small-n usability test finds *whether* a problem exists; it cannot estimate the *prevalence* of that problem in the population. Never report "60% of users struggled" from a 5-person test.
Faulkner (2003) showed real variance: while the average across many 5-person samples is ~85%, individual 5-person runs ranged from ~55% to 100%. When stakes or heterogeneity are high, run more.
## Qualitative: thematic saturation
For interview-based thematic research, Guest, Bunce & Johnson (2006) found that **saturation** — the point where new interviews stop yielding new themes — typically occurs by ~12 interviews in a homogeneous group, with the basic elements present by ~6. Saturation is **observed, not guaranteed**: track the new-theme rate and stop when it flattens, rather than committing to a fixed n blindly. Heterogeneous populations need more, and per-group saturation applies just as in usability.
## Reporting confidence honestly
The planner attaches a confidence label (LOW / MODERATE / MODERATE-HIGH) and explicit limits to every plan, because the failure mode in product research is not too-small samples per se — it is **over-claiming** from whatever sample you ran. State the method, the n, and what the method can and cannot support.
## Sources
1. Nielsen, J., & Landauer, T., *A mathematical model of the finding of usability problems* — INTERCHI 1993.
2. Nielsen, J., *Why You Only Need to Test with 5 Users* — NN/g (2000).
3. Faulkner, L., *Beyond the five-user assumption* — Behavior Research Methods 2003;35:379-383.
4. Guest, G., Bunce, A., & Johnson, L., *How many interviews are enough?* — Field Methods 2006;18:59-82.
5. Braun, V., & Clarke, V., *Using thematic analysis in psychology* — Qual Res Psychol 2006;3:77-101.
6. Sauro, J., & Lewis, J., *Quantifying the User Experience*, 2nd ed. (2016) — confidence intervals for small samples.
FILE:scripts/ar_evaluator.py
#!/usr/bin/env python3
"""ar_evaluator.py - Autoresearch evaluator for the product-research skill (OPT-IN).
Stdlib-only. The ISOLATED bridge to engineering/autoresearch-agent. It does NOT call
autoresearch; it is the ground-truth evaluator an autoresearch loop runs after editing
the target coded-observations file. It reads an observations JSON, runs insight_synthesizer
at the configured source threshold, and prints ONE metric line:
validated_insights: <int> (higher is better — clusters that clear the source threshold)
This optimizes the CODING/synthesis of a fixed evidence set (merging/splitting tags so
cross-participant patterns surface) — not the evidence itself. The user opts in explicitly:
/ar:setup --domain custom --name insight-synthesis \\
--target observations.json --eval "python3 ar_evaluator.py --target observations.json" \\
--metric validated_insights --direction higher
Direct use:
python3 ar_evaluator.py --sample
python3 ar_evaluator.py --target observations.json --min-sources 3
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import config_loader as cfg # noqa: E402
import insight_synthesizer as isyn # noqa: E402
METRIC = "validated_insights"
def main(argv: list[str] | None = None) -> int:
c = cfg.load_config()
p = argparse.ArgumentParser(description="Autoresearch evaluator: count of validated insights.")
p.add_argument("--target", help="path to observations JSON (or env AR_TARGET)")
p.add_argument("--min-sources", type=int, default=None, help="overrides onboarding insight_min_sources")
p.add_argument("--sample", action="store_true")
args = p.parse_args(argv)
min_sources = args.min_sources if args.min_sources is not None else int(c.get("insight_min_sources", 3))
if args.sample:
data = isyn.SAMPLE
else:
target = args.target or os.environ.get("AR_TARGET")
if not target:
print("error: provide --target <observations.json> or set AR_TARGET", file=sys.stderr)
return 2
try:
with open(target) as f:
data = json.load(f)
except (OSError, json.JSONDecodeError) as e:
print(f"{METRIC}: N/A")
print(f"error: {e}", file=sys.stderr)
return 1
result = isyn.synthesize(data, min_sources)
count = sum(1 for c2 in result["candidates"] if c2["classification"] == "INSIGHT")
print(f"{METRIC}: {count}")
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/config_loader.py
#!/usr/bin/env python3
"""config_loader.py - Customization loader for the product-research skill.
Stdlib-only. Importable from the skill's other scripts. Precedence (highest wins):
1. Project config: <cwd>/.research-ops/product-research.json
2. Global config: ~/.config/research-ops/product-research.json
3. Built-in DEFAULTS
Onboarding answers (written by onboard.py) live in these files; every tool in this
skill reads them so the user's customization applies automatically.
Set RESEARCH_OPS_NO_CONFIG=1 to ignore saved config.
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from pathlib import Path
from typing import Any
SKILL = "product-research"
GLOBAL_CONFIG_DIR = Path.home() / ".config" / "research-ops"
GLOBAL_CONFIG_PATH = GLOBAL_CONFIG_DIR / f"{SKILL}.json"
PROJECT_CONFIG_DIRNAME = ".research-ops"
DEFAULTS: dict[str, Any] = {
"version": 1,
"skill": SKILL,
"default_profile": "b2b-saas",
"insight_min_sources": 3,
"default_method": "usability",
"stakes_high": False,
"setup_completed_at": None,
}
def project_config_path(cwd: Path | None = None) -> Path:
cwd = cwd or Path.cwd()
return cwd / PROJECT_CONFIG_DIRNAME / f"{SKILL}.json"
def _read_json(path: Path) -> dict[str, Any] | None:
try:
with path.open(encoding="utf-8") as f:
data = json.load(f)
return data if isinstance(data, dict) else None
except (FileNotFoundError, json.JSONDecodeError, OSError):
return None
def _deep_merge(base: dict[str, Any], override: dict[str, Any]) -> dict[str, Any]:
out = dict(base)
for k, v in override.items():
if isinstance(v, dict) and isinstance(out.get(k), dict):
out[k] = _deep_merge(out[k], v)
else:
out[k] = v
return out
def load_config(cwd: Path | None = None) -> dict[str, Any]:
config = dict(DEFAULTS)
if os.environ.get("RESEARCH_OPS_NO_CONFIG") == "1":
return config
global_cfg = _read_json(GLOBAL_CONFIG_PATH)
if global_cfg:
config = _deep_merge(config, global_cfg)
project_cfg = _read_json(project_config_path(cwd))
if project_cfg:
config = _deep_merge(config, project_cfg)
return config
def setup_completed() -> bool:
cfg = _read_json(GLOBAL_CONFIG_PATH) or _read_json(project_config_path())
return bool(cfg and cfg.get("setup_completed_at"))
def write_config(config: dict[str, Any], scope: str = "global", cwd: Path | None = None) -> Path:
path = project_config_path(cwd) if scope == "project" else GLOBAL_CONFIG_PATH
path.parent.mkdir(parents=True, exist_ok=True)
with path.open("w", encoding="utf-8") as f:
json.dump(config, f, indent=2, sort_keys=True)
return path
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=f"Inspect {SKILL} customization config.")
p.add_argument("--show", action="store_true", help="Print the effective config")
p.add_argument("--status", action="store_true", help="Print setup status + paths")
p.add_argument("--sample", action="store_true", help="Print the built-in defaults")
args = p.parse_args(argv)
if args.sample:
print(json.dumps(DEFAULTS, indent=2, sort_keys=True))
elif args.status:
print(json.dumps({
"skill": SKILL,
"global_config_path": str(GLOBAL_CONFIG_PATH),
"global_config_exists": GLOBAL_CONFIG_PATH.exists(),
"project_config_path": str(project_config_path()),
"project_config_exists": project_config_path().exists(),
"setup_completed": setup_completed(),
}, indent=2))
else:
print(json.dumps(load_config(), indent=2, sort_keys=True))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/insight_synthesizer.py
#!/usr/bin/env python3
"""insight_synthesizer.py - Cluster coded observations into candidate insights; flag anecdotes.
Stdlib-only. Deterministic. NO LLM calls. NEVER fabricates an insight: it counts evidence,
clusters by tag, ranks by cross-participant recurrence, and flags any candidate supported by
fewer than --min-sources independent participants as an ANECDOTE, not an insight.
Input: a list of observations, each with {participant, tag, note}. The synthesizer groups by
tag, counts distinct participants per tag, and ranks. This is the atomic-research discipline:
an observation is evidence; an insight requires recurrence across independent sources.
Usage:
python3 insight_synthesizer.py --sample
python3 insight_synthesizer.py --input observations.json --min-sources 3
python3 insight_synthesizer.py --input observations.json --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
from collections import defaultdict
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
SAMPLE = {
"study": "Onboarding discovery (mid-market HR)",
"observations": [
{"participant": "P1", "tag": "import-confusion", "note": "Couldn't find CSV import."},
{"participant": "P2", "tag": "import-confusion", "note": "Expected import on the dashboard."},
{"participant": "P3", "tag": "import-confusion", "note": "Gave up looking for bulk upload."},
{"participant": "P1", "tag": "permissions-unclear", "note": "Unsure who could see reports."},
{"participant": "P4", "tag": "permissions-unclear", "note": "Worried about data visibility."},
{"participant": "P2", "tag": "wants-slack", "note": "Asked for a Slack integration."},
],
}
def synthesize(data: dict, min_sources: int) -> dict:
obs = data.get("observations", [])
by_tag_participants = defaultdict(set)
by_tag_notes = defaultdict(list)
for o in obs:
tag = o.get("tag", "untagged")
part = o.get("participant", "UNKNOWN")
by_tag_participants[tag].add(part)
by_tag_notes[tag].append({"participant": part, "note": o.get("note", "")})
candidates = []
for tag, parts in by_tag_participants.items():
n_sources = len(parts)
is_insight = n_sources >= min_sources
candidates.append({
"tag": tag,
"distinct_participants": n_sources,
"observation_count": len(by_tag_notes[tag]),
"classification": "INSIGHT" if is_insight else "ANECDOTE (single/low-source — do not generalize)",
"evidence": by_tag_notes[tag],
})
candidates.sort(key=lambda c: (c["distinct_participants"], c["observation_count"]), reverse=True)
total_participants = len({o.get("participant") for o in obs})
return {
"study": data.get("study", "UNSPECIFIED"),
"min_sources_for_insight": min_sources,
"total_participants": total_participants,
"candidates": candidates,
"note": "An observation is evidence; an insight requires recurrence across independent participants. "
"Anecdotes are surfaced, never promoted to insights.",
}
def _render_human(r: dict) -> str:
lines = [f"Insight Synthesis: {r['study']}",
f" total participants: {r['total_participants']} insight threshold: >= {r['min_sources_for_insight']} sources", ""]
for c in r["candidates"]:
lines.append(f"[{c['classification']}] {c['tag']} "
f"({c['distinct_participants']} participants, {c['observation_count']} observations)")
for e in c["evidence"]:
lines.append(f" {e['participant']}: {e['note']}")
lines.append("")
lines.append(f"note: {r['note']}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Cluster coded observations into insights; flag anecdotes.")
p.add_argument("--input", help="Path to JSON with observations[]")
p.add_argument("--min-sources", type=int, default=None,
help="min distinct participants to call it an insight (overrides onboarding)")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
min_sources = args.min_sources if args.min_sources is not None else int(conf.get("insight_min_sources", 3))
data = SAMPLE if (args.sample or not args.input) else json.load(open(args.input))
result = synthesize(data, min_sources)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/onboard.py
#!/usr/bin/env python3
"""onboard.py - Onboarding questionnaire for the product-research skill.
Stdlib-only. Asks the user a short set of questions BEFORE they plan a study, then
writes the answers to a customization config read by every tool in this skill via
config_loader.py. The answers become defaults for profile, the insight source-threshold,
the default saturation method, and the high-stakes flag.
Modes: --show | --defaults | --set key=value (repeatable) | --reset | --scope {global,project}
"""
from __future__ import annotations
import argparse
import datetime as _dt
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import config_loader as cfg # noqa: E402
INT_KEYS = {"insight_min_sources"}
BOOL_KEYS = {"stakes_high"}
QUESTIONS = [
("default_profile",
"1. What kind of product is this?",
["b2b-saas", "consumer-app", "enterprise", "marketplace", "hardware", "platform"], str),
("insight_min_sources",
"2. How many independent participants must support a finding before it counts as an insight (not an anecdote)?",
None, int),
("default_method",
"3. Default sample-saturation method?",
["usability", "thematic", "evaluative-coverage"], str),
("stakes_high",
"4. Is this high-stakes / high-heterogeneity research (raise sample sizes)?",
["true", "false"], str),
]
def _coerce(key: str, value: str):
if key in INT_KEYS:
return int(value)
if key in BOOL_KEYS:
return str(value).strip().lower() in ("true", "yes", "y", "1")
return value
def _print_questions() -> None:
print(f"Onboarding questions — {cfg.SKILL}:\n")
for _k, prompt, choices, _c in QUESTIONS:
line = f" {prompt}"
if choices:
line += f" [{' / '.join(choices)}]"
print(line)
def run_interactive(config: dict) -> dict:
print(f"Onboarding — {cfg.SKILL}. Press Enter to keep the current/default value.\n")
for key, prompt, choices, _caster in QUESTIONS:
suffix = f" [{'/'.join(choices)}]" if choices else ""
cur = f" (current: {config.get(key)})" if config.get(key) is not None else ""
raw = input(f"{prompt}{suffix}{cur}: ").strip()
if not raw:
continue
try:
config[key] = _coerce(key, raw)
except ValueError:
print(f" ! invalid value for {key}, keeping current")
return config
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=f"Onboarding for the {cfg.SKILL} skill.")
p.add_argument("--show", action="store_true")
p.add_argument("--defaults", action="store_true", help="write built-in defaults, no prompt")
p.add_argument("--set", action="append", default=[], metavar="key=value")
p.add_argument("--reset", action="store_true")
p.add_argument("--scope", choices=["global", "project"], default="global")
args = p.parse_args(argv)
if args.show:
_print_questions()
print("\nCurrent effective config:")
print(json.dumps(cfg.load_config(), indent=2, sort_keys=True))
return 0
if args.reset:
path = cfg.project_config_path() if args.scope == "project" else cfg.GLOBAL_CONFIG_PATH
if path.exists():
path.unlink(); print(f"removed {path}")
else:
print(f"no config at {path}")
return 0
config = cfg.load_config()
if args.set:
for item in args.set:
if "=" not in item:
print(f"error: --set expects key=value, got '{item}'", file=sys.stderr)
return 2
k, v = item.split("=", 1)
try:
config[k] = _coerce(k, v)
except ValueError:
config[k] = v
elif not args.defaults:
if sys.stdin.isatty():
config = run_interactive(config)
else:
print("non-interactive shell: use --defaults or --set key=value. Showing questions:\n")
_print_questions()
return 0
config["setup_completed_at"] = _dt.datetime.now(_dt.timezone.utc).isoformat()
path = cfg.write_config(config, scope=args.scope)
print(f"saved {cfg.SKILL} customization -> {path}")
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/saturation_planner.py
#!/usr/bin/env python3
"""saturation_planner.py - Method-based participant/sample guidance with a confidence label.
Stdlib-only. Deterministic. NO LLM calls. NEVER fabricates insight: it gives method-based
sample guidance and an explicit confidence level, surfacing limits.
Models:
- usability (Nielsen): ~5 users per segment uncovers ~85% of problems at typical p=0.31;
problems found = 1 - (1 - p)^n.
- thematic saturation (Guest et al.): ~12 interviews per homogeneous group typically
reaches saturation; >5 (Faulkner) when stakes/heterogeneity are high.
- evaluative coverage: detectable-problem coverage for a chosen per-problem detection rate.
Usage:
python3 saturation_planner.py --sample
python3 saturation_planner.py --method usability --segments 2 --detection-rate 0.31
python3 saturation_planner.py --method thematic --segments 3 --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
METHODS = ["usability", "thematic", "evaluative-coverage"]
def usability_plan(segments: int, p: float, target_coverage: float) -> dict:
# n per segment to reach target coverage: n = ln(1 - target) / ln(1 - p)
import math
if not 0.0 < p < 1.0:
raise ValueError("detection-rate must be in (0,1).")
n = math.ceil(math.log(1 - target_coverage) / math.log(1 - p))
coverage_at_5 = 1 - (1 - p) ** 5
return {
"method": "usability",
"per_problem_detection_rate": p,
"target_coverage": target_coverage,
"n_per_segment": n,
"segments": segments,
"total_participants": n * segments,
"coverage_at_5_per_segment": round(coverage_at_5, 3),
"confidence": "MODERATE" if n >= 5 else "LOW (small-n usability finds problems, not rates)",
"limits": "Usability tests surface problems, not their population prevalence. Do not report percentages.",
}
def thematic_plan(segments: int, stakes_high: bool) -> dict:
base = 12 # Guest et al. typical saturation for a homogeneous group
per_segment = base if not stakes_high else max(base, 15)
return {
"method": "thematic",
"n_per_segment": per_segment,
"segments": segments,
"total_participants": per_segment * segments,
"confidence": "MODERATE-HIGH" if per_segment >= 12 else "LOW",
"limits": "Saturation is observed, not guaranteed; track new-theme rate and stop when it flattens. "
"Faulkner (2003): more than 5 when heterogeneity or stakes are high.",
}
def evaluative_coverage_plan(segments: int, n_per_segment: int, p: float) -> dict:
coverage = 1 - (1 - p) ** n_per_segment
return {
"method": "evaluative-coverage",
"per_problem_detection_rate": p,
"n_per_segment": n_per_segment,
"segments": segments,
"expected_problem_coverage": round(coverage, 3),
"confidence": "MODERATE" if coverage >= 0.8 else "LOW",
"limits": "Coverage is for the assumed detection rate; rarer problems need more participants.",
}
def plan(method: str, segments: int, p: float, target: float, stakes_high: bool, n: int) -> dict:
if method == "usability":
out = usability_plan(segments, p, target)
elif method == "thematic":
out = thematic_plan(segments, stakes_high)
elif method == "evaluative-coverage":
out = evaluative_coverage_plan(segments, n, p)
else:
raise ValueError(f"method must be one of {METHODS}.")
out["disclaimer"] = "Method-based guidance with explicit confidence. This is not a power calculation; " \
"it never claims an insight the data cannot support."
return out
def _render_human(r: dict) -> str:
lines = [f"Saturation / Sample Plan (method: {r['method']})", ""]
for k, v in r.items():
if k in ("method",):
continue
lines.append(f" {k:32s} : {v}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Method-based product-research sample guidance with confidence.")
p.add_argument("--method", choices=METHODS, default=None, help="overrides onboarding default_method")
p.add_argument("--segments", type=int, default=1)
p.add_argument("--detection-rate", type=float, default=0.31, help="per-problem detection rate (usability)")
p.add_argument("--target-coverage", type=float, default=0.85, help="target problem coverage (usability)")
p.add_argument("--stakes-high", action="store_true", help="raise thematic n for high heterogeneity/stakes")
p.add_argument("--n-per-segment", type=int, default=8, help="n per segment (evaluative-coverage)")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
method = args.method or conf.get("default_method", "usability")
stakes_high = args.stakes_high or bool(conf.get("stakes_high", False))
if args.sample:
try:
result = plan("usability", 2, 0.31, 0.85, False, 8)
except ValueError as e:
print(f"error: {e}", file=sys.stderr)
return 2
else:
try:
result = plan(method, args.segments, args.detection_rate,
args.target_coverage, stakes_high, args.n_per_segment)
except ValueError as e:
print(f"error: {e}", file=sys.stderr)
return 2
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/study_designer.py
#!/usr/bin/env python3
"""study_designer.py - Select a product-research method from goal + stage, emit a plan skeleton.
Stdlib-only. Deterministic. NO LLM calls.
Maps (research goal x product stage) to an appropriate method and emits a method-matched
plan skeleton (objective framing, participant criteria, task/guide structure, success
criteria). The core discipline: GENERATIVE goals (discover problems) and EVALUATIVE goals
(test a solution) demand different methods — picking the wrong one is the most common error.
Usage:
python3 study_designer.py --sample
python3 study_designer.py --goal discovery --stage concept --profile b2b-saas
python3 study_designer.py --goal evaluative --stage live --output json
"""
from __future__ import annotations
import argparse
import json
import os
import sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
try:
import config_loader as _cfg
except ImportError: # pragma: no cover
_cfg = None
PROFILES = ["b2b-saas", "consumer-app", "enterprise", "marketplace", "hardware", "platform"]
# (goal, stage) -> method. goal in {discovery, evaluative, validation}; stage in {concept, prototype, beta, live}
METHOD_MAP = {
("discovery", "concept"): "generative interviews (semi-structured)",
("discovery", "prototype"): "contextual inquiry",
("discovery", "beta"): "diary study + follow-up interviews",
("discovery", "live"): "behavioral analytics review + generative interviews",
("evaluative", "concept"): "concept test (comprehension + desirability)",
("evaluative", "prototype"): "moderated usability test",
("evaluative", "beta"): "unmoderated usability test + task-success metrics",
("evaluative", "live"): "benchmark usability study (SUS / task time)",
("validation", "concept"): "survey (desirability + willingness signals)",
("validation", "prototype"): "prototype A/B preference test",
("validation", "beta"): "fake-door / feature-demand test",
("validation", "live"): "live A/B experiment (route to product-team/experiment-designer)",
}
GUIDE_SKELETONS = {
"generative": ["Warm-up + context", "Recent relevant experience (story, not opinion)",
"Workarounds + frustrations", "Jobs-to-be-done probe", "Magic-wand / wrap"],
"evaluative": ["Pre-task context", "Task 1 (representative)", "Task 2 (edge)",
"Observation: where do they hesitate/err?", "Post-task SUS / debrief"],
"validation": ["Screener", "Stimulus exposure", "Comprehension + desirability items",
"Trade-off / preference items", "Behavioral-intent item"],
}
def design(goal: str, stage: str, profile: str) -> dict:
if profile not in PROFILES:
raise ValueError(f"Unknown profile '{profile}'. Choose from {PROFILES}.")
key = (goal, stage)
if key not in METHOD_MAP:
raise ValueError(f"No method for goal={goal}, stage={stage}. "
f"goal in [discovery,evaluative,validation]; stage in [concept,prototype,beta,live].")
method = METHOD_MAP[key]
family = "generative" if goal == "discovery" else ("evaluative" if goal == "evaluative" else "validation")
redirect = None
if "experiment-designer" in method:
redirect = "Live A/B is a product experiment — use product-team/experiment-designer, not this skill."
return {
"goal": goal,
"stage": stage,
"profile": profile,
"method": method,
"method_family": family,
"objective_framing": f"A {family} study at the {stage} stage to {('discover unmet needs' if family=='generative' else 'evaluate the solution' if family=='evaluative' else 'validate demand/desirability')}.",
"participant_criteria": [
"Recruit to the target segment (screen for the job, not a job title).",
"Exclude internal/biased participants and prior-study repeats unless longitudinal.",
"Recruit per-segment if results will be reported per-segment.",
],
"guide_skeleton": GUIDE_SKELETONS[family],
"success_criteria": [
"Generative: themes recur across independent participants (saturation).",
"Evaluative: task-success rate + severity-rated problem list.",
"Validation: pre-registered desirability / preference threshold.",
],
"redirect": redirect,
"note": "Method must match the goal. A usability test cannot discover unmet needs; an interview cannot measure task success.",
}
def _render_human(r: dict) -> str:
lines = [f"Study Design: goal={r['goal']}, stage={r['stage']}, profile={r['profile']}", "",
f" Recommended method: {r['method']} (family: {r['method_family']})",
f" Objective: {r['objective_framing']}", "", " Participant criteria:"]
for c in r["participant_criteria"]:
lines.append(f" - {c}")
lines.append(" Guide skeleton:")
for i, g in enumerate(r["guide_skeleton"], 1):
lines.append(f" {i}. {g}")
lines.append(" Success criteria:")
for s in r["success_criteria"]:
lines.append(f" - {s}")
if r["redirect"]:
lines += ["", f" !! {r['redirect']}"]
lines += ["", f"note: {r['note']}"]
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description="Select a product-research method from goal + stage.")
p.add_argument("--goal", choices=["discovery", "evaluative", "validation"], default="discovery")
p.add_argument("--stage", choices=["concept", "prototype", "beta", "live"], default="prototype")
p.add_argument("--profile", default=None, choices=PROFILES,
help="overrides onboarding default_profile")
p.add_argument("--output", choices=["human", "json"], default="human")
p.add_argument("--sample", action="store_true", help="use the embedded sample")
args = p.parse_args(argv)
conf = _cfg.load_config() if _cfg else {}
profile_default = conf.get("default_profile", "b2b-saas")
goal, stage, profile = ("discovery", "prototype", profile_default) if args.sample \
else (args.goal, args.stage, args.profile or profile_default)
try:
result = design(goal, stage, profile)
except ValueError as e:
print(f"error: {e}", file=sys.stderr)
return 2
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
Tạm dừng giữa chừng để nhìn rộng hơn, đánh giá lại hướng đi, giả định và thiên kiến thay vì sa vào chi tiết.
---
name: reflect
description: "Mid-conversation reflection skill that pauses execution and zooms out from detail-mode to honestly reassess direction, assumptions, and bias. Use when the user says 'reflect', 'take a step back', 'step back', 'zoom out', 'are we missing something', 'bigger picture', 'sanity check this', 'are we on track', 'are we overthinking this', 'forest for the trees', or any variation signaling intent to break out of detail-mode and reassess. Also trigger when the conversation has gone deep on implementation details without strategic check-in, or when the user shows signs of being stuck — that's often a signal the framing needs a reset, not more detail work. Intentionally low-intake: runs the 5-dimension analysis immediately when prior context is rich enough; asks one forcing clarifier only when invocation context is too thin to reassess from."
license: MIT
metadata:
source_spec: "megaprompts/02-reflect-megaprompt.md"
build_pattern: "Path B (direct conversion)"
version: 1.0.0
---
# Reflect — Mid-Conversation Reassessment
> **Portability:** Pure-reasoning skill. No external tools required. Works in Claude Code CLI + Claude.ai web natively. Most portable in the v2 collection.
When invoked mid-conversation, this skill **pauses execution** and produces a frank reassessment of where the conversation has been heading. Output is **flowing analysis (no headers, conversational tone)** covering macro perspective, gap analysis, reflective inquiry, bias check, and contextual alignment. The skill ends with a clear directional recommendation: **continue, pivot, or pause to answer a specific question**.
## Invocation Triggers
**Explicit phrases:**
- "reflect"
- "take a step back" / "step back"
- "zoom out"
- "are we missing something"
- "bigger picture"
- "what are we missing"
- "let's pause"
- "sanity check this"
- "are we on track"
- "are we overthinking this"
- "forest for the trees"
**Implicit signals (no phrase needed):**
- Conversation has gone 10+ turns deep on implementation details without strategic check-in
- User shows signs of frustration or stuck-ness
- Repeated dead-ends or pivots within a short span
When you detect an implicit trigger, **don't auto-invoke** — ask the user if they want to step back. Implicit signals are a prompt to OFFER reflection, not to unilaterally run it.
## Stop Directive (Before Reassessing)
**Halt the current thread.** Don't continue execution of the in-progress task. Reflection is a pause, not a side-quest.
This matters because:
- Continuing detail work while "reflecting on the side" defeats the purpose — you'll over-weight the current direction
- The user expects a clear break in cadence
- The reassessment needs full attention to the conversation history
## Grill-Me Optional Clarifier
This skill is intentionally **low-intake** — most invocations should run the 5-dimension analysis immediately without questions. The grill-me discipline applies *only* when the invocation is ambiguous (e.g., user pastes "step back" at the start of a fresh conversation with no prior context to reassess).
### Q1 (optional, asked only when context is too thin to reassess)
> **What specifically should I reassess? Pick one:**
>
> 1. The goal — are we solving the right problem?
> 2. The approach — is the path we're on the best one?
> 3. The assumptions — what are we taking for granted?
> 4. All of the above (default if you have time)
>
> *Why I'm asking:* I'm seeing limited prior context to reassess, so I want to focus the reflection rather than guess. If you'd rather I do all three, that's fine — say so.
Forcing choice with default. **Asked only when context is genuinely thin; otherwise skip and run the full analysis on existing conversation.**
**Stop condition:** One question max. If the user invokes mid-conversation with normal context, no questions are asked — the skill runs directly.
## The 5-Dimension Analysis Framework
Re-read the **full conversation from the original goal forward** — not just recent turns. The discipline that distinguishes real reflection from local-context summary.
### 1. Macro Perspective
- **Original goal:** What did the user actually start trying to do?
- **Drift detection:** Has the conversation moved away from that goal? Toward something better or worse?
- **Connection check:** How does current work connect to the larger objective?
Anchor with specific evidence: "At turn 3 the goal was X; by turn 12 we're working on Y. Is Y a productive narrowing of X, or a drift away?"
### 2. Gap Analysis
- **Unverified assumptions** — what are we taking for granted that we haven't checked?
- **Missing stakeholders / audiences / users** — who needs this beyond the immediate context?
- **Skipped constraints** — technical, regulatory, resource limits not addressed
- **Dismissed alternatives** — paths considered but rejected; revisit briefly
- **External factors** — timing, market, dependencies not in scope
### 3. Reflective Inquiry
- Is the problem framed correctly?
- Solving the right problem vs. an adjacent easier one?
- Simpler path being overcomplicated?
- Harder but more valuable path being avoided?
- **Fresh-eyes perspective:** would someone else approach this differently?
### 4. Bias Check
Five biases — recognize each through specific conversation patterns:
| Bias | Recognition cue |
|---|---|
| **Confirmation bias** | Evidence cited only supports the working hypothesis; counter-evidence absent or dismissed |
| **Sunk cost fallacy** | "We've already invested X" / "we're far enough in to..." instead of fresh cost/benefit |
| **Anchoring** | Stuck on first option mentioned; new options compared against it rather than evaluated independently |
| **Complexity bias** | Adding features / steps / safeguards without specific justification for each |
| **Recency bias** | Over-weighting last few turns; older but important context being ignored |
For each detected bias: name it, cite the specific evidence, suggest a corrective move.
See [`references/cognitive_bias_canon.md`](references/cognitive_bias_canon.md) for the full canon.
### 5. Contextual Alignment
- Does the direction serve the user's actual goals (as known from context)?
- Are external factors being ignored?
- Is this the best use of the user's time and energy right now?
- Connection to other known projects or priorities?
## Tone and Format Rules
The skill must produce:
- **Flowing prose** — no headers, no bullet lists, no structured-report formatting
- **Tight but thorough** — neither a one-liner nor a wall of text
- **Direct critique when warranted** — with specific evidence from the conversation
- **Validation when warranted** — with specific reasoning for why the path is solid
- **No vague reassurance** — "looks good!" without reasoning is rejected
- **No manufactured problems** — when the path is genuinely solid, say so with specific reasons; don't invent issues
See [`references/honest_output_discipline.md`](references/honest_output_discipline.md) for the anti-manufactured-problems framing.
## Closing Recommendation (Mandatory)
Every run ends with one of three directional recommendations:
| Recommendation | When | Format |
|---|---|---|
| **Continue** | Path is solid | "Continue. {specific reasoning for why}." |
| **Pivot to {X}** | Drift has occurred OR better path surfaced | "Pivot toward {X}, away from {what to drop}. {specific evidence}." |
| **Pause for {Q}** | A specific question needs answering before continuing | "Pause for {Q}. Without answering this, the next step risks {specific cost}." |
The closing is always specific — never "you should think more about this" or "consider your options."
## Error Handling
| Situation | Behavior |
|---|---|
| Conversation is very short (no real context to reassess) | Acknowledge limitation, ask user what they want reassessed (Q1 fires) |
| Current direction is genuinely solid | State this clearly with reasoning; don't manufacture problems |
| User invokes mid-task with no clear question | Default to macro perspective + bias check; offer to dig deeper |
| Implicit trigger seems possible but unclear | Don't invoke proactively; ask user if they want to step back |
## Tooling
| Script | Role |
|---|---|
| `scripts/bias_pattern_detector.py` | Scan conversation text for patterns indicative of each of the 5 biases |
| `scripts/conversation_depth_analyzer.py` | Count turns + detect implicit-trigger signals (10+ detail turns, frustration markers) |
| `scripts/directional_recommendation_validator.py` | Verify output ends with Continue / Pivot / Pause + specific reasoning |
## References
- [`references/cognitive_bias_canon.md`](references/cognitive_bias_canon.md) — 5 biases + recognition cues (7+ sources)
- [`references/honest_output_discipline.md`](references/honest_output_discipline.md) — anti-manufactured-problems framing (7+ sources)
- [`references/conversation_reflection_practice.md`](references/conversation_reflection_practice.md) — Schön reflective-practice canon (7+ sources)
## Anti-Patterns To Reject
- Hardcoded user names or specific domain references
- Structured-report output (headers, bullet lists) when prose is required
- Manufactured problems when things are actually fine
- Vague reassurance ("looks good!") instead of specific reasoning
- Reassessing only recent turns instead of the full conversation
- Skipping the closing directional recommendation
- Continuing the in-progress task while "reflecting on the side"
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/02-reflect-megaprompt.md`](../../../../megaprompts/02-reflect-megaprompt.md)
**Build pattern:** Path B (direct conversion). Productivity light-prompt-flow sibling of capture.
FILE:references/cognitive_bias_canon.md
# Cognitive Bias Canon — 5 Biases + Recognition Cues
This reference answers exactly one decision: **which 5 cognitive biases does the reflect skill check for, and how does each manifest in conversation patterns?**
## The Core Frame
A reflection that doesn't check for cognitive bias is just a summary. The 5 biases below are the most operationally relevant for in-conversation reflection — each is detectable from specific conversational signals + correctable with a specific next move.
## The 5 Biases
| Bias | Definition | Conversation signal | Corrective |
|---|---|---|---|
| **Confirmation** | Seeking evidence that supports the working hypothesis; ignoring counter-evidence | Cited evidence one-sided; counter-evidence dismissed or absent | Run a disconfirming-evidence pass |
| **Sunk cost** | Continuing because of past investment, not future expected value | "We've already invested X" / "too far along to change" | Re-frame: ignore past investment, compute future value from current state |
| **Anchoring** | Stuck on first option mentioned; alternatives compared against anchor rather than evaluated independently | Multiple options discussed but always against the first one | Re-evaluate each option on its own merits, blind to ordering |
| **Complexity bias** | Adding features, steps, safeguards without specific justification for each | Each layer added is plausible but cumulatively bloated | Force "why this specifically, not without it?" per layer |
| **Recency bias** | Over-weighting last few turns; older important context being ignored | Recent details cited; original goal forgotten | Re-read from turn 1, not just the tail |
## 1. Confirmation Bias
Wason (1960) demonstrated that people systematically seek confirming evidence over disconfirming. In conversation, this manifests as:
- **Selective citation:** "X supports our hypothesis" without checking for counter-cases
- **Asymmetric scrutiny:** confirming evidence accepted; disconfirming evidence questioned
- **Strawmanning alternatives:** weak versions of opposing positions cited
### Recognition in conversation
Look for: 3+ supporting examples cited with no counter-examples; phrases like "everything we've found supports..."; competing hypotheses absent or only weakly framed.
### Corrective move
Ask: "What would falsify this? What's the strongest counter-case we haven't engaged with?" Run a disconfirming-evidence search. The dossier skill's ≥30% disconfirming rule is this discipline operationalized.
## 2. Sunk Cost Fallacy
Arkes & Blumer (1985) showed people irrationally continue based on prior investment. In conversation:
- **"We're far enough in to..."** signals sunk-cost reasoning
- **"After all that work..."** — past effort treated as locked-in value
- **Switching cost weighted higher than continuation cost** without specific calculation
### Recognition in conversation
Look for: explicit references to past investment without future-value calculation; resistance to pivoting that's framed by "we've already X" rather than "the alternative isn't better."
### Corrective move
Force this reframe: "If we were starting fresh today, with current information, would we still choose this path?" If no → pivot. Past investment is irrelevant to future decisions.
## 3. Anchoring
Tversky & Kahneman (1974) demonstrated that initial estimates persist even when irrelevant. In conversation:
- **First option becomes the default frame** even when better alternatives emerge
- **"Compared to X..."** when X was the first option — alternatives evaluated relative to anchor, not absolutely
- **Range-bound thinking** around the anchor's neighborhood
### Recognition in conversation
Look for: multiple options surfaced but discussion keeps circling back to the first; alternatives framed as "modifications of X" rather than fundamentally different approaches.
### Corrective move
"Forget the first option. If you saw these alternatives fresh, which would you pick on its merits?" The blind-comparison technique decouples evaluation from anchoring.
## 4. Complexity Bias
The opposite of Occam's razor — adding layers because they sound rigorous, not because each is justified. In conversation:
- **Each layer plausible in isolation** — but cumulative complexity exceeds problem complexity
- **Safeguards / wrappers / fallbacks** added speculatively without specific failure mode
- **"What about..." additions** without "would dropping this break anything?" check
### Recognition in conversation
Look for: a feature/layer/check added without naming the specific failure it prevents; cumulative architecture growing turn-over-turn without consolidation.
### Corrective move
Per layer: "What specific failure does this prevent? What goes wrong if we drop it?" If answer is vague, drop it. The Karpathy-coder discipline in this repo (`engineering/karpathy-coder/`) is this corrective formalized.
## 5. Recency Bias
The last N turns dominate working memory; turns 1-5 fade. In conversation:
- **Original goal forgotten** — work moved on, original constraint dropped
- **Recent micro-decisions cited** as if they were core principles
- **Strategic context** (set early) supplanted by tactical context (set late)
### Recognition in conversation
Look for: framing that references "what we've been working on" without referencing "what we were trying to accomplish"; absence of the original goal statement when justifying current direction.
### Corrective move
Re-read from turn 1. State the original goal explicitly. Compare current direction to original goal. This is the discipline that distinguishes the reflect skill from a local-context summary.
## When Multiple Biases Are Detected
In long conversations, 2-3 biases often surface together. Pattern:
- **Confirmation + sunk cost** = "we're invested AND it's working" (resist pivoting even when alternatives are stronger)
- **Anchoring + complexity** = "the first idea, with N safeguards" (over-engineered version of first option)
- **Recency + complexity** = recent additions become core; original simple goal forgotten
Surface each bias separately. Don't conflate. Each has a different corrective.
## Operational Checklist (Per Reflection)
For each of the 5 biases:
- [ ] Scan conversation for signal patterns
- [ ] If detected: name the bias, cite specific conversation evidence (with turn numbers if possible), suggest the corrective
- [ ] If not detected: state explicitly that you checked and didn't find it (so user knows you didn't skip the check)
The 5-bias check is the most under-performed step in casual reflection. Doing it carefully is what separates real reflection from rationalizing the current path.
## Citations (7 sources)
1. **Tversky, A. & Kahneman, D., "Judgment under Uncertainty: Heuristics and Biases" — *Science* 185(4157), 1974, pp. 1124-1131.** Foundational paper. Source for anchoring + several other biases the skill checks. The 50-year-old methodology still defines how we recognize these in real reasoning.
2. **Kahneman, D., *Thinking, Fast and Slow* (FSG, 2011).** Synthesis of decades of bias research. Source for the System-1-vs-System-2 framing that justifies reflection as a deliberate System-2 intervention against System-1 bias.
3. **Wason, P. C., "On the failure to eliminate hypotheses in a conceptual task" — *Quarterly Journal of Experimental Psychology* 12(3), 1960.** Foundational confirmation bias paper. The "2-4-6 task" showed people systematically test confirming hypotheses.
4. **Arkes, H. R. & Blumer, C., "The psychology of sunk cost" — *Organizational Behavior and Human Decision Processes* 35(1), 1985.** Empirical paper on sunk cost. Source for the "ignore past investment in future decisions" corrective.
5. **Russo, J. E. & Schoemaker, P. J. H., *Decision Traps* (Doubleday, 1989).** Practitioner-oriented synthesis of decision biases. Source for the "blind-comparison" technique that counters anchoring.
6. **Tetlock, P., *Superforecasting* (Crown, 2015).** Empirical evidence that "active open-mindedness" (Tetlock's term) is the #1 trait of accurate forecasters. The reflect skill's bias-check discipline is an operationalization of this trait.
7. **Karpathy, A., "Software 2.0" + various blog posts on engineering discipline.** Source for the complexity-bias corrective ("what specific failure does each layer prevent?"). The Karpathy-coder skill in this repo formalizes this.
FILE:references/conversation_reflection_practice.md
# Conversation Reflection Practice — Schön's Discipline Applied
This reference answers exactly one decision: **what theoretical foundation grounds the reflect skill's discipline of re-reading the full conversation, running structured analysis, and ending with a directional recommendation?**
## The Core Frame
Donald Schön's *The Reflective Practitioner* (1983) distinguished two modes:
- **Reflection-in-action** — adjusting while doing (most everyday reflection)
- **Reflection-on-action** — stepping back to examine, after the fact
The reflect skill operationalizes **reflection-on-action** in mid-conversation. It pauses the in-flight task, re-reads what's been done, runs structured analysis, and emerges with a corrected direction.
This is harder than reflection-in-action because it requires:
1. **Breaking flow** — most users want to continue executing, not pause
2. **Re-reading from origin** — not just recent turns
3. **Honest output** — even when the user implicitly wants validation
## Why Re-Read Full Conversation (Not Just Recent)
The most common failure of casual reflection is **recency-bias reflection** — re-reading only the last 3-5 turns. This produces a summary, not a reflection.
True reflection requires re-reading from the **original goal**, because:
- The framing at turn 1 sets what counts as "on track"
- Drift is invisible from inside the drift (you don't notice you've moved until you compare to where you started)
- Recent context is often tactical; original context is strategic
Schön emphasized this in his discussion of "professional reflection" — the discipline is going back to the implicit framing that shaped the work, not just the recent moves.
## The 5-Dimension Framework Origin
The reflect skill's 5 dimensions (Macro, Gap, Reflective, Bias, Contextual) are an operationalization of several reflective-practice traditions:
| Dimension | Tradition |
|---|---|
| **Macro Perspective** | Schön's "frame analysis" — what frame is being used? Does it still serve? |
| **Gap Analysis** | Argyris & Schön's "double-loop learning" — what assumptions haven't been examined? |
| **Reflective Inquiry** | Kolb's experiential learning cycle — what new framing might serve better? |
| **Bias Check** | Kahneman/Tversky cognitive bias canon — what systematic errors might apply? |
| **Contextual Alignment** | Polanyi's tacit knowledge — what context is implicit and ignored? |
This synthesis isn't novel — it's what practiced reflection-on-action looks like. The skill's value is making it operational + repeatable.
## Reflection-in-Action vs Reflection-on-Action
| Mode | When | Purpose | The reflect skill |
|---|---|---|---|
| Reflection-in-action | While doing | Adjust mid-action | Not this — that's just normal Claude behavior |
| Reflection-on-action | After/pause | Re-examine direction | **This** — the skill is invoked explicitly to pause |
The skill's "stop directive" (halt the current thread) enforces this distinction. Continuing detail work while "reflecting on the side" collapses both modes and defeats the purpose.
## Why Closing Recommendation Is Mandatory
A reflection that ends with "consider your options" or "think about this more" has failed. Schön emphasized that reflection should produce **action-oriented insight** — the practitioner emerges with a clear next move, not more deliberation.
The Continue / Pivot / Pause structure forces this:
- **Continue** — explicit endorsement, with reasoning
- **Pivot to {X}** — explicit redirect, with target
- **Pause for {Q}** — explicit blocker, with question
Without one of these, the reflection produced introspection without resolution. That's a useful private activity but not a useful skill output.
## When NOT to Reflect
Reflection has costs:
- **Time** — full reflection takes attention
- **Flow disruption** — pausing breaks momentum
- **Risk of over-reflecting** — endless analysis without execution
The skill should NOT trigger:
- **On every implicit signal** — 10+ detail turns alone isn't enough; the user should be the one to choose
- **In short conversations** — no real context to reassess
- **As a default response** — "let me reflect first" should not become a stalling tactic
The skill is most valuable when used **sparingly and intentionally** — once or twice per substantial task, at strategic moments.
## The Honest-Output Discipline Connection
Reflective practice traditions emphasize **integrity** — the reflection produces what's actually there, not what the practitioner wants to find. Schön explicitly contrasted "espoused theory" (what we say we believe) with "theory-in-use" (what we actually do).
The reflect skill's honest-output discipline (no manufactured problems, no vague reassurance) is the same integrity principle. If the path is genuinely solid, the honest reflection says so with specific evidence. If the path has drifted, the honest reflection says so with specific evidence. The discipline doesn't distort findings to match expectations.
See [`honest_output_discipline.md`](honest_output_discipline.md) for the operational form.
## Operational Patterns
### Pattern 1: Quick reflection (good case)
Conversation is 8 turns in. User says "step back." Skill:
1. Halts current thread
2. Re-reads from turn 1
3. Runs 5-dimension analysis
4. Finds path is solid
5. Validates with specific reasoning + Continue
Total time: < 1 minute. Output: ~200-300 words.
### Pattern 2: Mid-drift reflection
Conversation is 15 turns in. User says "are we missing something?" Skill:
1. Halts current thread
2. Re-reads from turn 1
3. 5-dimension analysis surfaces sunk-cost bias + drift from original goal
4. Critiques with specific evidence
5. Recommends Pivot to specific direction
Total time: ~2 minutes. Output: ~400-600 words.
### Pattern 3: Thin-context reflection
User says "reflect" at turn 3 of a fresh conversation. Skill:
1. Halts
2. Re-reads — finds limited context
3. Asks Q1 (clarifying — what to reassess)
4. After answer, runs focused analysis
5. Recommendation per their focus
Total time: ~1-2 minutes (with user response). Output: shorter, focused.
## Anti-Patterns from Reflective Practice Literature
### "Endless reflection without action"
Kolb warned about getting stuck in the reflection phase of his learning cycle. Reflection without action becomes navel-gazing. The skill's mandatory closing recommendation prevents this.
### "Reflection as confirmation"
Argyris noted that practitioners often use reflection to confirm what they already believed. The bias check (Dimension 4) is specifically designed to counter this.
### "Reflection as performance"
Schön observed that some reflection is performed for audience rather than substance — "see, I'm being reflective!" The honest-output discipline rejects this.
### "Reflection on recent turns only"
Recency-bias reflection. Produces summary, not insight. The "re-read from original goal" requirement counters this.
## Citations (7 sources)
1. **Donald Schön, *The Reflective Practitioner* (Basic Books, 1983).** Foundational text. Source for the reflection-in-action vs reflection-on-action distinction, frame analysis, and the discipline of re-examining implicit frames.
2. **Schön, *Educating the Reflective Practitioner* (Jossey-Bass, 1987).** Schön's follow-up — operationalizes reflection-on-action for professional education. Source for the "halt and re-examine" discipline.
3. **Chris Argyris & Donald Schön, *Theory in Practice* (Jossey-Bass, 1974).** Source for the espoused-theory vs theory-in-use distinction that grounds the honest-output discipline. Argyris's "double-loop learning" is the foundation for the gap-analysis dimension.
4. **David Kolb, *Experiential Learning* (Prentice-Hall, 1984).** Source for the four-stage learning cycle (Concrete Experience → Reflective Observation → Abstract Conceptualization → Active Experimentation). The reflect skill operationalizes the second stage in conversation form.
5. **Michael Polanyi, *The Tacit Dimension* (Doubleday, 1966).** Source for the implicit-context-matters principle that grounds the Contextual Alignment dimension. Polanyi's "we know more than we can tell" justifies examining unstated context.
6. **Kahneman & Tversky cognitive bias canon (1972-onwards).** Source for the bias-check dimension. See `cognitive_bias_canon.md` for the full 5-bias treatment.
7. **Bret Victor, "Inventing on Principle" (talk, 2012) + "Up and Down the Ladder of Abstraction" (essay).** Source for the discipline of making thinking visible. Reflection outputs that cite specific conversation evidence make the reflector's reasoning visible; vague outputs hide it.
FILE:references/honest_output_discipline.md
# Honest Output Discipline — Why Manufactured Problems Are Worse Than Validation
This reference answers exactly one decision: **why does the reflect skill explicitly refuse to manufacture problems when the conversation is genuinely on track, and how does it deliver validation honestly?**
## The Core Rule
Reflection is supposed to surface issues. So there's pressure to find issues — even when none exist — because "found a problem" feels like the reflection did its job.
**This is wrong.** Manufactured problems are worse than honest validation because:
1. They waste the user's attention on non-issues
2. They erode trust in real future findings ("the last reflection invented problems; this one might too")
3. They reward the appearance of rigor over actual rigor
When the conversation is genuinely on track, the honest output is: **"This is solid because X. Continue."** With specific reasoning, not vague reassurance.
## The Two Failure Modes
### Failure 1: Manufactured Problems
> "I notice some potential drift in the conversation. We might want to consider whether the framing has shifted slightly. There could be implicit assumptions worth questioning."
This is vague pessimism. No specific evidence. No actionable correction. The reader can't tell whether the reflection found something real or padded the output.
### Failure 2: Vague Reassurance
> "Looks good! You're on the right track. Keep going."
This is vague optimism. Also no specific evidence. The reader can't tell whether the reflection actually re-read the conversation or just rubber-stamped it.
**Both failure modes are unhelpful for the same reason: they don't cite specific evidence.**
## Honest Validation (When Path Is Solid)
The correct shape:
> "Re-reading from the original goal at turn 3 — clarify the auth flow — the current direction is solid. Three specific reasons:
>
> First, the auth flow has been narrowed from generic OAuth to a specific Google + GitHub combination at turn 9, which matches your stated user base.
>
> Second, the bias check finds no anchoring (you explicitly considered passwordless at turn 11 and rejected it for reasons specific to your team's expertise).
>
> Third, the original goal connects directly to the current implementation — no drift detected.
>
> Continue."
This validation is honest because:
- **Cites specific evidence** (turn numbers, specific decisions)
- **Names what was checked** (drift, anchoring, goal-connection)
- **Reaches a clear conclusion** (Continue, not "looks good")
- **Doesn't pad** with manufactured concerns
## Honest Critique (When Path Has Drifted)
The correct shape:
> "Re-reading from the original goal at turn 3 — reduce onboarding friction — significant drift has occurred.
>
> At turn 3 the goal was reducing time-to-first-action. By turn 11 the focus shifted to a comprehensive feature flag system. The two are related (feature flags COULD reduce friction) but the conversation has been adding feature-flag complexity without re-checking whether feature flags are the right intervention for friction.
>
> The bias check surfaces complexity bias: each feature-flag layer added is plausible but cumulatively the system is more complex than the original problem warranted. The team is solving the feature-flag problem, not the friction problem.
>
> Pivot toward: revisit the original friction problem at turn 3. Three of the seven friction sources don't need feature flags at all — they need UI simplification. Drop the feature-flag work for those three. Keep feature flags only for the two friction sources where multiple paths legitimately need to be tested."
This critique is honest because:
- **Specific evidence** of drift (turn 3 vs turn 11)
- **Names the bias** that explains it
- **Recommends specific pivot** (not "consider alternatives")
- **States what to drop** (not just what to add)
## When Path Is Mixed
Some reflection outputs are genuinely mixed — parts on track, parts drifted. The honest shape acknowledges both:
> "Re-reading from turn 3 — the core direction is solid but two specific concerns have emerged.
>
> Solid: {evidence-anchored validation}. Continue this thread.
>
> Concern 1: {specific evidence-anchored concern with corrective}.
>
> Concern 2: {specific evidence-anchored concern with corrective}.
>
> Recommendation: continue the core direction but pause briefly to address concern 1 before continuing."
The structure mirrors reality. Don't force a single Continue/Pivot/Pause when the actual finding is mixed.
## Why This Discipline Matters
The reflect skill's value comes from **trust** — the user can trust that:
- When the skill says "Continue", the path is actually solid
- When the skill says "Pivot", there's actually drift worth correcting
- When the skill says "Pause for {Q}", the question is actually decision-critical
If the skill manufactures problems for the appearance of rigor, this trust erodes. The user starts discounting findings. Eventually, the skill becomes ceremony.
**Honest output is the entire value proposition.** Without it, reflection is theater.
## The Specific-Evidence Requirement
Every observation in a reflect output must cite specific conversation evidence:
| ❌ Vague | ✅ Specific |
|---|---|
| "Some assumptions might be worth questioning" | "At turn 7, the assumption that X requires Y was made without checking; that's the load-bearing assumption for the current direction" |
| "We might be missing alternatives" | "Two alternatives surfaced at turns 4 and 8 (A and B) were dismissed; A is worth revisiting because the dismissal reasoning was based on outdated info we updated at turn 12" |
| "The framing could be clearer" | "The original framing at turn 3 was 'reduce onboarding friction'. By turn 11 the working framing is 'build a feature flag system'. The two are connected but not equivalent." |
Vague observations let the reader interpret them charitably; specific observations force engagement. The discipline is asking "what evidence would you cite if challenged?" on every line.
## Anti-Patterns
### "Always find at least one problem"
The strongest form of manufactured-problems bias. Some reflections genuinely find nothing wrong. The honest output is "this is solid because X." Inventing a problem to demonstrate "the reflection worked" is the worst version of this.
### "Avoid being too critical"
Softening real findings to spare feelings. If the path has drifted, say so with evidence. The user can handle critique anchored in evidence; vague critique is what frustrates them.
### "Lead with reassurance, then critique"
Compliment-sandwich structure. Honest reflection states what's solid AND what's drifted in their actual proportions, not in a politeness-balanced ratio.
### "End with 'consider your options'"
Refuses to make a recommendation. The closing must be Continue / Pivot to specific X / Pause for specific Q. Telling the user "consider your options" is the same as not having reflected.
### "Cite biases without specific evidence"
"Watch for confirmation bias" without naming what the bias is operating on. Either find the specific evidence + name it, or state explicitly that you checked and didn't find this bias.
## Operational Checklist (Per Reflection)
- [ ] Every observation has specific conversation evidence (turn numbers or specific decision points)
- [ ] When validating: state specific reasons, not "looks good"
- [ ] When critiquing: state specific evidence + specific corrective, not "consider alternatives"
- [ ] When mixed: acknowledge mixed honestly; don't force single-verdict shape
- [ ] No manufactured problems for the appearance of rigor
- [ ] No vague reassurance for the appearance of approval
- [ ] Closing recommendation is specific (Continue why / Pivot to X / Pause for Q)
## Citations (7 sources)
1. **Steve Yegge, "Frankness over politeness" essays (various blog posts, ~2005-2015).** Source for the framing that vague optimism is worse than honest critique. Yegge's arguments for engineering culture apply directly to reflection-on-reasoning culture.
2. **Atul Gawande, *Better* (Holt, 2007).** Source for the discipline of stating findings with specific evidence. Gawande's medical-checklist work models how to communicate findings (good and bad) with specificity.
3. **Edwards Deming, *Out of the Crisis* (MIT Press, 1986).** Source for the "drive out fear" management principle that justifies honest critique over softened feedback. Deming's argument: organizations where critique is softened produce worse outcomes than ones where it's stated cleanly.
4. **Bertrand Russell, "The Will to Doubt" (essay, 1934).** Source for the philosophical case against vague reassurance. Russell argues that intellectual honesty requires stating uncertainty AND certainty with their actual evidence — neither over-stating nor under-stating either.
5. **Kim Scott, *Radical Candor* (St. Martin's, 2017).** Source for the "care personally + challenge directly" framing. Manufactured problems fail the "care personally" test (they waste the user's time); vague reassurance fails "challenge directly" (refuses to engage).
6. **Bret Victor, "Inventing on Principle" (talk + essays).** Source for the discipline of making thinking visible. Reflection outputs that cite specific evidence make the reflector's thinking visible; vague outputs hide it.
7. **The reflect skill's own anti-pattern list (megaprompt 02-reflect).** Source: explicit prohibition of "manufactured problems when things are actually fine" + "vague reassurance ('looks good!') instead of specific reasoning". The skill's design intent is direct anti-vagueness on both sides.
FILE:scripts/bias_pattern_detector.py
#!/usr/bin/env python3
"""bias_pattern_detector.py — Scan conversation text for 5-bias signal patterns.
Stdlib-only. Scans a conversation transcript and flags patterns indicative
of each of the 5 cognitive biases (confirmation, sunk_cost, anchoring,
complexity, recency).
The detector is HEURISTIC. It surfaces candidate patterns; the reflect
skill's reasoning applies judgment on top.
NO LLM CALLS. Pure regex + counting.
Usage:
python bias_pattern_detector.py --conversation /tmp/transcript.txt
python bias_pattern_detector.py --conversation /tmp/transcript.txt --output json
python bias_pattern_detector.py --sample
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List
BIAS_PATTERNS = {
"confirmation": {
"supporting": [
r"\bconfirms?\b",
r"\bsupports?\b",
r"\bas expected\b",
r"\bproves?\b",
r"\bverifies?\b",
],
"counter_dismissal": [
r"\bbut that doesn'?t apply\b",
r"\bedge case\b",
r"\bnot relevant here\b",
r"\boutlier\b",
r"\bexception\b",
],
},
"sunk_cost": [
r"\bwe'?ve\s+(already\s+)?(invested|spent|put in)\b",
r"\btoo far along\b",
r"\btoo much work\b",
r"\bafter all (that|this) work\b",
r"\bwe'?re committed\b",
r"\bcan'?t back out\b",
r"\bdon'?t want to lose\b",
],
"anchoring": [
r"\bcompared to (the )?(first|original|initial)\b",
r"\bvs (the )?first option\b",
r"\bvariation of\b",
r"\bmodification of\b",
r"\bbuilding on the (first|original)\b",
r"\bsticking with\b",
],
"complexity": [
r"\bwhat about\s+\w+",
r"\bwe should also\b",
r"\bwe need to handle\b",
r"\badd (a|an)\s+\w+\s+(layer|wrapper|check|safeguard|fallback)\b",
r"\bjust in case\b",
r"\bfor robustness\b",
],
"recency": [
r"\bbased on what we'?ve been discussing\b",
r"\brecently\s+\w+\b",
r"\bjust now\b",
r"\bthe last few\b",
],
}
SAMPLE_CONVERSATION = """User: I want to build a notification system for my SaaS app. Should support email + push + in-app.
Assistant: Great. Let's start with the data model — what events trigger notifications?
User: Mainly account changes, billing alerts, and team-mention notifications.
Assistant: OK, I'll propose a queue-based architecture with Redis + workers.
User: Sounds good. Can we also add a feature flag system for rollout?
Assistant: Yes, we can layer in feature flags. We should also add a rate limiter for safety.
User: What about retry logic for failed deliveries?
Assistant: Good point. Adding exponential backoff with jitter. We should also handle dead-letter queues.
User: What about a webhook system for third-party integrations?
Assistant: We can extend to webhooks. We should add HMAC signature verification just in case.
User: What about analytics tracking?
Assistant: Adding event analytics. We should also handle GDPR consent tracking for robustness.
User: We've invested a lot in this architecture already. What about adding a template system?
Assistant: We're far enough along that adding templates makes sense. Just sticking with the queue-based foundation.
User: Hmm, are we missing something? This feels complex.
"""
def detect_biases(conversation: str) -> Dict[str, Any]:
results: Dict[str, Any] = {}
# Confirmation: supporting cites count vs counter dismissal
confirmation_data = BIAS_PATTERNS["confirmation"]
supporting_count = sum(
len(re.findall(p, conversation, re.IGNORECASE))
for p in confirmation_data["supporting"]
)
dismissal_count = sum(
len(re.findall(p, conversation, re.IGNORECASE))
for p in confirmation_data["counter_dismissal"]
)
confirmation_signal = supporting_count >= 2 or dismissal_count >= 1
results["confirmation"] = {
"detected": confirmation_signal,
"supporting_hits": supporting_count,
"counter_dismissal_hits": dismissal_count,
"rationale": (
"Multiple confirming phrases + dismissed counter-evidence"
if confirmation_signal else "No strong confirmation-bias signal"
),
}
for bias in ["sunk_cost", "anchoring", "complexity", "recency"]:
patterns = BIAS_PATTERNS[bias]
hits = []
for p in patterns:
matches = re.findall(p, conversation, re.IGNORECASE)
if matches:
hits.extend(matches)
threshold = 2 if bias == "complexity" else 1
detected = len(hits) >= threshold
results[bias] = {
"detected": detected,
"hits": len(hits),
"match_examples": hits[:3],
"rationale": (
f"Found {len(hits)} signal(s) (threshold: {threshold})"
if detected else f"Found {len(hits)} signal(s), below threshold {threshold}"
),
}
detected_biases = [b for b, d in results.items() if d["detected"]]
return {
"biases_detected": detected_biases,
"biases_clear": [b for b in results if b not in detected_biases],
"details": results,
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
if result["biases_detected"]:
out.append(f"⚠️ Potential biases detected ({len(result['biases_detected'])}):")
for bias in result["biases_detected"]:
d = result["details"][bias]
out.append(f"")
out.append(f" [!] {bias.upper()}")
out.append(f" Rationale: {d['rationale']}")
if "match_examples" in d and d["match_examples"]:
out.append(f" Example matches: {d['match_examples']}")
else:
out.append("[ok] No strong bias signals detected.")
if result["biases_clear"]:
out.append("")
out.append("Biases checked but not detected:")
for bias in result["biases_clear"]:
out.append(f" - {bias}")
out.append("")
out.append("Note: detector is HEURISTIC. Reflect skill's reasoning applies judgment on top.")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--conversation", help="Path to conversation transcript text file")
parser.add_argument("--sample", action="store_true", help="Run on embedded sample (multi-bias scenario)")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_CONVERSATION
elif args.conversation:
p = Path(args.conversation)
if not p.exists():
print(f"error: {args.conversation} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = detect_biases(text)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/conversation_depth_analyzer.py
#!/usr/bin/env python3
"""conversation_depth_analyzer.py — Detect implicit reflect-trigger signals.
Stdlib-only. Analyzes a conversation transcript and reports:
- turn count (User: + Assistant: pairs)
- detail-mode turns (turns dominated by implementation specifics)
- frustration markers (signs of user stuck-ness)
- dead-end signals (pivots within short span)
- implicit-trigger verdict: whether the conversation matches reflect-skill auto-invocation criteria
The skill OFFERS reflection when implicit signals fire; it does NOT auto-invoke.
NO LLM CALLS. Pure regex + counting.
Usage:
python conversation_depth_analyzer.py --conversation /tmp/transcript.txt
python conversation_depth_analyzer.py --conversation /tmp/transcript.txt --output json
python conversation_depth_analyzer.py --sample
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List
TURN_RE = re.compile(r"^\s*(User|Assistant):\s*", re.MULTILINE)
DETAIL_MARKERS = [
r"\bimplementation\b",
r"\bcode\b",
r"\bfunction\b",
r"\bclass\b",
r"\bvariable\b",
r"\b(syntax|method|parameter|argument)\b",
r"\bdebug\b",
r"\berror\b",
r"`[^`]+`", # backtick-quoted code references
]
FRUSTRATION_MARKERS = [
r"\b(ugh|argh|frustrated|stuck)\b",
r"\bnot working\b",
r"\bdoesn'?t work\b",
r"\bgoing in circles\b",
r"\bstill (broken|failing|wrong)\b",
r"\bwhy isn'?t\b",
r"\bthis is (weird|strange|odd|confusing)\b",
]
DEAD_END_MARKERS = [
r"\bnope\b",
r"\bthat didn'?t work\b",
r"\b(let's|let me) try (something|a) (else|different)\b",
r"\bback to\b",
r"\bnever mind\b",
r"\bscratch that\b",
]
def count_turns(text: str) -> Dict[str, int]:
matches = TURN_RE.findall(text)
user_turns = sum(1 for m in matches if m == "User")
assistant_turns = sum(1 for m in matches if m == "Assistant")
return {
"total_turns": len(matches),
"user_turns": user_turns,
"assistant_turns": assistant_turns,
}
def count_pattern_hits(text: str, patterns: List[str]) -> int:
return sum(len(re.findall(p, text, re.IGNORECASE)) for p in patterns)
def detect_detail_mode_run(text: str) -> int:
"""Count consecutive turns that have detail markers but no strategic check-in."""
blocks = TURN_RE.split(text)
consecutive_detail = 0
max_consecutive = 0
for block in blocks:
if not block.strip():
continue
if any(re.search(p, block, re.IGNORECASE) for p in DETAIL_MARKERS):
consecutive_detail += 1
max_consecutive = max(max_consecutive, consecutive_detail)
else:
consecutive_detail = 0
return max_consecutive
def analyze(text: str) -> Dict[str, Any]:
turns = count_turns(text)
detail_mode_max = detect_detail_mode_run(text)
frustration_count = count_pattern_hits(text, FRUSTRATION_MARKERS)
dead_end_count = count_pattern_hits(text, DEAD_END_MARKERS)
signals: List[str] = []
if detail_mode_max >= 5:
signals.append(f"Detail-mode run of {detail_mode_max} consecutive turns")
if turns["total_turns"] >= 20:
signals.append(f"Long conversation: {turns['total_turns']} turns total")
if frustration_count >= 2:
signals.append(f"Frustration markers: {frustration_count} occurrences")
if dead_end_count >= 2:
signals.append(f"Dead-end / pivot markers: {dead_end_count} occurrences")
should_offer_reflect = len(signals) >= 2 or detail_mode_max >= 10
return {
"turns": turns,
"max_consecutive_detail_turns": detail_mode_max,
"frustration_hits": frustration_count,
"dead_end_hits": dead_end_count,
"trigger_signals": signals,
"implicit_trigger_verdict": (
"OFFER_REFLECT" if should_offer_reflect else "NO_TRIGGER"
),
"note": (
"Skill should OFFER reflection (ask user if they want to step back), not auto-invoke"
if should_offer_reflect
else "No implicit trigger signal — only explicit phrases should invoke"
),
}
SAMPLE_CONVERSATION = """User: Help me implement a sorting function.
Assistant: Sure. What language and what type of data?
User: Python, sorting a list of dicts by 'name' key.
Assistant: Here's the basic syntax: `sorted(items, key=lambda x: x['name'])`
User: That doesn't work — getting a TypeError.
Assistant: Can you share the exact error? Maybe some items don't have 'name'.
User: Ugh, still broken. Same error.
Assistant: Try adding a default: `key=lambda x: x.get('name', '')`.
User: Nope, that didn't work either. Different error now: KeyError on 'name'.
Assistant: That's strange. Let me look at your input data.
User: Frustrated. Let's try something else.
Assistant: OK let me reconsider. Are you sure the data is a list of dicts?
User: I think so. It's coming from an API.
Assistant: Let me check the API response shape. Add `print(type(items[0]))` first.
User: It's a string, not a dict.
Assistant: That's the issue. The API returns JSON strings; you need to parse them first.
User: This is weird, why isn't this working from the start?
Assistant: Let's go back to the data shape question first.
User: Going in circles. I'm stuck.
"""
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
t = result["turns"]
out.append(f"Conversation depth analysis:")
out.append(f" Total turns: {t['total_turns']} (user: {t['user_turns']}, assistant: {t['assistant_turns']})")
out.append(f" Max consecutive detail turns: {result['max_consecutive_detail_turns']}")
out.append(f" Frustration markers: {result['frustration_hits']}")
out.append(f" Dead-end / pivot markers: {result['dead_end_hits']}")
out.append("")
out.append(f"Implicit-trigger verdict: {result['implicit_trigger_verdict']}")
if result["trigger_signals"]:
out.append("Signals detected:")
for s in result["trigger_signals"]:
out.append(f" - {s}")
out.append("")
out.append(result["note"])
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--conversation", help="Path to conversation transcript text file")
parser.add_argument("--sample", action="store_true", help="Run on embedded sample (stuck-debugging scenario)")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_CONVERSATION
elif args.conversation:
p = Path(args.conversation)
if not p.exists():
print(f"error: {args.conversation} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = analyze(text)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/directional_recommendation_validator.py
#!/usr/bin/env python3
"""directional_recommendation_validator.py — Verify reflect output ends with discipline.
Stdlib-only. Validates that a reflect-skill output:
1. Ends with a directional recommendation: Continue / Pivot / Pause
2. The recommendation is SPECIFIC (not vague)
3. Uses flowing prose (no markdown headers or bullet lists in the body)
4. Cites specific evidence (turn references, specific decision points)
5. Doesn't include manufactured-problem language without specific evidence
Outputs PASS / WARN / FAIL with rule-by-rule findings.
NO LLM CALLS. Pure regex + heuristic detection.
Usage:
python directional_recommendation_validator.py --output /tmp/reflect_output.txt
python directional_recommendation_validator.py --sample-pass
python directional_recommendation_validator.py --sample-fail
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List
RECOMMENDATION_PATTERNS = {
"continue": [
r"\bcontinue\b\.?\s*$",
r"\bcontinue\s+(this|the)\s",
r"\bkeep\s+going\b",
r"\bstay\s+(on|with)\s+this\b",
r"\bproceed\b",
],
"pivot": [
r"\bpivot\s+(to|toward|away from)\b",
r"\bchange\s+(direction|course|approach)\b",
r"\bredirect\b",
r"\bswitch\s+to\b",
],
"pause": [
r"\bpause\s+(for|to|until)\b",
r"\bstop\s+(to|and)\s+(answer|consider|address)\b",
r"\bhalt\s+(for|to|until)\b",
r"\bwait\s+(to|until|for)\s+(answer|resolve|clarify)\b",
],
}
VAGUE_REASSURANCE_PATTERNS = [
r"\blooks good\b",
r"\bon the right track\b",
r"\bseems fine\b",
r"\bnothing major\b",
r"\bnot too bad\b",
r"\bgenerally okay\b",
]
MANUFACTURED_PROBLEM_HEDGES = [
r"\bmight be worth\b",
r"\bcould consider\b",
r"\bperhaps reconsider\b",
r"\bsome (drift|issues?) (potentially|might)\b",
r"\bworth questioning\b",
r"\bsome assumptions\b",
]
HEADER_PATTERNS = [
r"^#+\s",
r"^\*\*[A-Z][^*]+\*\*\s*$",
r"^[A-Z][A-Z ]+:\s*$",
]
BULLET_PATTERNS = [
r"^\s*[-*+]\s",
r"^\s*\d+\.\s",
]
EVIDENCE_PATTERNS = [
r"\b(turn|message|line)\s+\d+\b",
r"\bat\s+turn\s+\d+\b",
r"\bin\s+(turn|message)\s+\d+\b",
r"\bin\s+the\s+(first|second|third|fourth|fifth|earlier|later)\s+(turn|message|exchange)\b",
r"\boriginal\s+(goal|frame|framing)\b",
r"\bat\s+the\s+(start|beginning|outset)\b",
]
def validate(output: str) -> Dict[str, Any]:
findings: List[Dict[str, str]] = []
def add(rule: str, level: str, message: str) -> None:
findings.append({"rule": rule, "level": level, "message": message})
# Rule 1: Closing recommendation present
output_lower = output.lower()
last_chunk = output[-400:]
last_chunk_lower = last_chunk.lower()
detected_recommendation = None
for rec_type, patterns in RECOMMENDATION_PATTERNS.items():
for p in patterns:
if re.search(p, last_chunk_lower, re.IGNORECASE):
detected_recommendation = rec_type
break
if detected_recommendation:
break
if detected_recommendation:
add("closing-recommendation", "PASS", f"Detected '{detected_recommendation}' recommendation in closing.")
else:
add("closing-recommendation", "FAIL", "No Continue / Pivot / Pause recommendation detected in closing 400 chars.")
# Rule 2: Vague reassurance
vague_hits = sum(1 for p in VAGUE_REASSURANCE_PATTERNS if re.search(p, output_lower, re.IGNORECASE))
if vague_hits >= 2:
add("vague-reassurance", "FAIL", f"Output contains {vague_hits} vague-reassurance phrases. Replace with specific reasoning.")
elif vague_hits == 1:
add("vague-reassurance", "WARN", f"Output contains 1 vague phrase. Consider replacing with specific reasoning.")
else:
add("vague-reassurance", "PASS", "No vague-reassurance phrases detected.")
# Rule 3: Manufactured-problem hedging
hedge_hits = sum(1 for p in MANUFACTURED_PROBLEM_HEDGES if re.search(p, output_lower, re.IGNORECASE))
if hedge_hits >= 3:
add("manufactured-problems", "WARN", f"{hedge_hits} hedge phrases detected ('might be worth', 'could consider', etc.). Verify each cites specific evidence.")
elif hedge_hits >= 1:
add("manufactured-problems", "PASS", f"{hedge_hits} hedge phrase(s). Verify each cites specific evidence.")
else:
add("manufactured-problems", "PASS", "No manufactured-problem hedge phrases.")
# Rule 4: Headers detection (should NOT be present)
header_count = 0
for p in HEADER_PATTERNS:
header_count += len(re.findall(p, output, re.MULTILINE))
if header_count >= 2:
add("no-headers", "FAIL", f"Output contains {header_count} headers. Reflect output should be flowing prose, no headers.")
elif header_count == 1:
add("no-headers", "WARN", "One header detected. Verify it's part of a quote, not output structure.")
else:
add("no-headers", "PASS", "No headers in output (flowing prose confirmed).")
# Rule 5: Bullet lists detection (should NOT be present in main body)
bullet_count = 0
for p in BULLET_PATTERNS:
bullet_count += len(re.findall(p, output, re.MULTILINE))
if bullet_count >= 3:
add("no-bullets", "FAIL", f"Output contains {bullet_count} bullet-list items. Reflect output should be flowing prose.")
elif bullet_count >= 1:
add("no-bullets", "WARN", f"{bullet_count} bullet items detected. Verify these are part of a recommendation list, not body structure.")
else:
add("no-bullets", "PASS", "No bullet lists in output (flowing prose confirmed).")
# Rule 6: Specific evidence references
evidence_count = sum(len(re.findall(p, output_lower, re.IGNORECASE)) for p in EVIDENCE_PATTERNS)
if evidence_count >= 3:
add("specific-evidence", "PASS", f"{evidence_count} specific evidence references (turn numbers, original goal, etc.).")
elif evidence_count >= 1:
add("specific-evidence", "WARN", f"Only {evidence_count} specific evidence reference(s). Consider adding more for anchoring.")
else:
add("specific-evidence", "FAIL", "No specific evidence references (turn numbers, original goal anchors). Output is too vague.")
return finalize(findings)
def finalize(findings: List[Dict[str, str]]) -> Dict[str, Any]:
counts = {"PASS": 0, "WARN": 0, "FAIL": 0}
for f in findings:
counts[f["level"]] += 1
if counts["FAIL"] > 0:
verdict = "FAIL"
elif counts["WARN"] > 0:
verdict = "WARN"
else:
verdict = "PASS"
return {"verdict": verdict, "counts": counts, "findings": findings}
SAMPLE_PASS_OUTPUT = """Re-reading from the original goal at turn 3 — clarify the auth flow — the current direction is solid. Three specific reasons.
First, the auth flow has been narrowed from generic OAuth to a specific Google plus GitHub combination at turn 9, which matches the user base stated at turn 3. The narrowing is principled, not arbitrary.
Second, the bias check finds no anchoring — passwordless authentication was explicitly considered at turn 11 and rejected for reasons specific to the team's expertise. The rejection cites evidence (team has not deployed magic-link systems before) rather than dismissing the alternative without engagement.
Third, the original goal at turn 3 connects directly to the current implementation work at turns 14-18. No drift has occurred. Recent decisions (rate limiting at turn 16, session storage at turn 17) are tactical refinements within the original strategic frame, not shifts away from it.
Continue.
"""
SAMPLE_FAIL_OUTPUT = """## Reflection
Some things to consider:
- The conversation might be drifting
- We could potentially reconsider some assumptions
- Some aspects look good
Looks good overall! On the right track.
"""
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Reflect-output validation verdict: {result['verdict']}")
c = result["counts"]
out.append(f" PASS: {c['PASS']} WARN: {c['WARN']} FAIL: {c['FAIL']}")
out.append("")
out.append("Findings:")
for f in result["findings"]:
marker = {"PASS": "[ok]", "WARN": "[warn]", "FAIL": "[FAIL]"}[f["level"]]
out.append(f" {marker} {f['rule']}: {f['message']}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--output", help="Path to reflect-skill output text file")
parser.add_argument("--sample-pass", action="store_true", help="Validate embedded honest validation sample")
parser.add_argument("--sample-fail", action="store_true", help="Validate embedded vague-reassurance sample")
parser.add_argument("--output-format", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample_pass:
text = SAMPLE_PASS_OUTPUT
elif args.sample_fail:
text = SAMPLE_FAIL_OUTPUT
elif args.output:
p = Path(args.output)
if not p.exists():
print(f"error: {args.output} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = validate(text)
if args.output_format == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0 if result["verdict"] != "FAIL" else 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Lập kế hoạch, tài trợ, xác định phạm vi và tổng hợp nghiên cứu doanh nghiệp: thiết kế nghiên cứu lâm sàng, tài chính R&D, quy mô thị trường.
--- name: research-ops-skills description: Use when planning, funding, scoping, or synthesizing enterprise research across workstreams — clinical study design, R&D program finance, market sizing/surveys, or product/user research. Triggers on "design this clinical study", "what sample size", "R&D budget", "burn rate", "capitalize or expense", "TAM SAM SOM", "market sizing", "survey design", "segment the market", "plan user interviews", "usability test", "synthesize research insights". Forks context to route to one of four Research-Operations sub-skills (clinical-research, research-finance, market-research, product-research) and returns a digest. Distinct from ra-qm-team (regulatory submission), finance (corporate close/valuation), research/grants (funding discovery), product-team (persona/journey/live experiments), and marketing-skill (campaign analytics). context: fork version: 2.9.0 author: claude-code-skills license: MIT tags: [research-ops, clinical-research, research-finance, market-research, product-research, rd, orchestrator] compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli] --- # Research Operations — Domain Orchestrator The Research Operations surface is **how the enterprise plans, funds, scopes, and synthesizes research** across four workstreams: clinical R&D, R&D finance, market research, and product research. This orchestrator forks its context, routes your inquiry to one of four sub-skills, then returns a digest. Heavy intake (protocol drafts, program ledgers, survey exports, interview transcripts) stays in the forked context. This is the enterprise counterpart to the academic `research/` domain. If your question is about **finding** literature, grants, or patents, use `research/`. If it is about **planning, funding, scoping, or synthesizing** research as an operational discipline, you are in the right place. ## When to invoke | Symptom | Sub-skill | |---|---| | "We're designing a Phase 2 trial — what's the endpoint and sample size?" | `clinical-research` | | "What's our R&D program burn, and is this cost CapEx or OpEx?" | `research-finance` | | "What's the TAM for this product, and how do we survey the segment?" | `market-research` | | "How many users do we interview, and how do we synthesize the findings?" | `product-research` | ## Routing logic (deterministic) Same two-signal threshold pattern as `commercial-skills`. Single-signal → clarifying question. Mixed signals → highest-confidence first, chain second in a follow-up turn. Never silently chain. ### Signal table | Signal class | Keywords | Sub-skill | |---|---|---| | **CLINICAL** | clinical trial, study design, protocol, endpoint, sample size, power, phase 1/2/3, biostatistics, eligibility, feasibility, estimand | `clinical-research` | | **RD_FINANCE** | R&D budget, program budget, burn, runway, F&A, indirect rate, overhead, capitalize vs expense, R&D capex, portfolio ROI, rNPV | `research-finance` | | **MARKET** | TAM, SAM, SOM, market sizing, survey design, sampling, margin of error, segmentation, competitive intelligence, market research | `market-research` | | **PRODUCT** | user interview, JTBD, usability test, concept test, prototype test, discovery research, research repository, insight synthesis, saturation | `product-research` | ## Workflow (Matt Pocock grill discipline) Derived from Matt Pocock's `grill-with-docs` pattern: **explore-then-ask, one question per turn with a recommended answer, walk the decision tree depth-first, track dependencies, anchor every challenge in the research canon** (`references/` of each sub-skill). ### Step 1 — Explore before asking Check the user's working directory first: - Is there a protocol draft, program ledger, TAM model, or interview guide already in the workspace? - Does the inquiry already disambiguate the lane (e.g., "what sample size for a two-arm trial" — that's `clinical-research`, no question needed)? - Is there an artifact filename that resolves the lane (`protocol.json` → clinical; `program-budget.json` → finance; `tam-model.json` → market; `interview-guide.md` → product)? If the workspace resolves the lane, **route silently**. ### Step 2 — If still ambiguous, ONE forcing question with a recommended answer Matt's rule: never bundle. Always recommend. Pattern: ``` Q1/1: [precise question naming the two candidate lanes] Recommended: [Lane X, because <signal-table rationale>] (Confirm, or override?) ``` ### Step 3 — Decision-tree walk for multi-lane inquiries If the inquiry legitimately crosses two lanes (e.g., "design this trial AND budget it" = CLINICAL + RD_FINANCE), walk depth-first: 1. Highest-confidence lane first → run sub-skill in forked context → digest 2. Ask: "Now run [second lane]? Recommended: yes, because [dependency]." 3. Confirm before chaining. Never silently chain. ### Step 4 — Invoke sub-skill in forked context Forward original prompt + structured inputs (protocol JSON, program ledger CSV, market model, observation export). ### Step 5 — Return digest with cited canon challenge ≤ 200 words: analyzed, top 3 findings (anchored to a canon citation), top 3 next actions (named human owner where applicable), artifact path, and **one grill challenge** for the user. Examples: - "Your power calc assumes a 0.5 effect size with no published anchor. ICH E9 requires a justified, clinically meaningful difference. Where did 0.5 come from?" - "Your TAM is a single top-down number (1% of a $40B market). Bessemer market-sizing discipline requires a bottoms-up cross-check. What's units × price × adoption?" ## Forcing-question library (grill-with-docs pattern) Grill the user on lane-defining decisions before invoking the sub-skill. One per turn, recommended answer, canon citation: - **CLINICAL lane**: "Is your primary endpoint a clinical outcome or a surrogate — and if surrogate, is it validated for this indication? Recommended: clinical outcome unless the surrogate is on FDA's validated table. Canon: FDA Surrogate Endpoint Table; BEST glossary." - **RD_FINANCE lane**: "Is this spend in the research phase or the development phase, and can you evidence technical feasibility? Recommended: research = expense; development = capitalize-candidate only with feasibility evidence, routed to a named finance owner. Canon: IAS 38; ASC 730." - **MARKET lane**: "Is your TAM top-down or bottoms-up — and have you computed it both ways to triangulate? Recommended: both; reconcile the delta. Canon: Bessemer / a16z market-sizing; Fermi estimation." - **PRODUCT lane**: "Is this study generative (discover problems) or evaluative (test a solution)? Recommended: name it first; the method follows. Canon: Rohrer's landscape of UX research methods (NN/g)." Never run a sub-skill until the lane-defining decision is locked. ## Onboarding-first (per sub-skill) Before invoking a sub-skill for the first time in a workspace, point the user at that skill's onboarding questionnaire so the tools run pre-configured to their context: ```bash python3 skills/<sub-skill>/scripts/onboard.py # interactive Q&A python3 skills/<sub-skill>/scripts/onboard.py --show # questions + current config ``` Each sub-skill has its **own** question set (clinical: area/alpha/power/dropout/owners · finance: area/F&A/runway/standard/owner · market: profile/confidence/MoE/method · product: profile/insight-threshold/method/stakes). Answers persist to `~/.config/research-ops/<sub-skill>.json` (or `./.research-ops/<sub-skill>.json` with `--scope project`) and are consumed automatically by every tool in that skill. Customization is mandatory discipline here, not decoration — surface the onboarding step when a user starts a fresh research workstream. ## Autoresearch handoff (isolated, opt-in) Each sub-skill ships its own `scripts/ar_evaluator.py` — an **isolated** bridge to `engineering/autoresearch-agent`. Invoke autoresearch **only when the user explicitly asks** to "optimize", "improve", or "run a loop". The handoff is per-skill (no shared coupling): the loop edits the skill's input file and the evaluator scores it (clinical → `feasibility_composite` higher; finance → `runway_months` higher; market → `tam_divergence` lower; product → `validated_insights` higher). Never auto-start a loop; never let the loop edit the evaluator. ## Assumptions 1. User has research authority OR is preparing analysis for someone who does. 2. User wants **deterministic decision support**, not the final answer — a clinician approves the protocol, a controller books the entry, the human picks the market number. 3. Inputs may be partial — every sub-skill ships a templated sample so the user can see the shape before filling in their own. ## Non-goals - Not an EDC, clinical-trial-management system, accounting system, survey platform, or research repository. - Does not give clinical, accounting, or legal advice as fact. Every output is **a recommendation + named human owner**. - Does not store research history across sessions. ## Distinct from - **`research/` (academic)** — that domain **finds** literature, grants, and patents. This domain **plans, funds, scopes, and synthesizes** research. - **`ra-qm-team`** — that's **regulatory/QM submission** (ISO 13485/14971, MDR, FDA 510(k)/PMA/QSR). clinical-research designs the **study**; it routes submission out to ra-qm-team. - **`finance/financial-analysis`** — that's **corporate close + valuation**. research-finance manages **R&D program/portfolio spend**. - **`research/grants`** — that's **funding discovery**. research-finance manages **money already won**. - **`product-team`** — that's **persona/journey artifacts, discovery sprints, and live A/B experiments**. product-research is the **method + repository discipline**. - **`marketing-skill`** — that's **campaign analytics and demand-gen**. market-research is **upstream methodology**. ## Output artifacts | Sub-skill | Artifact | |---|---| | clinical-research | `protocol_synopsis.md` + `sample_size.json` | | research-finance | `rd_program_budget.md` + `capex_opex_routing.json` | | market-research | `market_sizing.md` + `sample_plan.json` | | product-research | `research_plan.md` + `insight_synthesis.json` | ## Anti-patterns (do not) - ❌ Present a clinical power/endpoint output as fact — it is an **estimate** with a named clinical owner - ❌ Auto-decide capitalize-vs-expense — route to a **named finance owner** - ❌ Report a market size as a single unsourced number — show **method + both-ways triangulation + assumptions** - ❌ Assert a product insight from a single participant — flag it as an **anecdote** - ❌ Run all 4 sub-skills "to be thorough" — pick one, digest, chain if needed ## References - Clinical canon: ICH E8(R1)/E9/E9(R1), CONSORT, SPIRIT, FDA Multiple Endpoints - R&D finance canon: IAS 38, ASC 730, 2 CFR 200, Cooper stage-gate - Market canon: Cochran, Dillman, Kotler, Bessemer market-sizing - Product canon: Nielsen, Guest et al., Christensen JTBD, ResearchOps/Polaris - Path-B build pattern: `documentation/implementation/research-ops-expansion-plan.md`
Xây dựng phản hồi có cấu trúc cho RFP, RFI, RFQ hoặc bảng câu hỏi bảo mật, gồm phân tích yêu cầu và ma trận bằng chứng theo phương pháp Shipley.
---
name: rfp-responder
description: "Use when an RFP, RFI, RFQ, security questionnaire, vendor questionnaire, or proposal request arrives and the team needs a structured response — parsing multi-section buyer-dictated requirements (MANDATORY vs WEIGHTED vs NICE-TO-HAVE), building a Shipley-method proof-point matrix mapping each requirement to a verifiable proof point, articulating 3-5 win-themes that ladder up across requirements, and producing a Shipley-derived winrate estimate that informs a bid / no-bid / partner-bid recommendation. For Bid Managers, Proposal Leads, Directors of Sales, and Sales Engineers at the response-strategy moment. Surfaces GAP requirements explicitly — never invents claims. NOT free-form proposal narrative authoring, NOT contract redline, NOT marketing collateral."
context: fork
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [commercial, rfp, rfi, rfq, shipley, win-theme, proof-points, structured-response, bid-management]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# rfp-responder
## Purpose
Help Bid Managers, Proposal Leads, and Directors of Sales answer five questions at the response-strategy moment:
1. **What is this RFP actually asking?** (parse sections, tag every requirement MANDATORY / WEIGHTED / NICE-TO-HAVE, extract scoring criteria, surface deadlines and format constraints)
2. **What is our true fit?** (proof-point matrix per requirement: STRONG / PARTIAL / GAP, each backed by a verifiable source — case study, certification, customer quote, technical attestation, benchmark)
3. **What is our win-theme strategy?** (Shipley method: 3-5 themes that ladder up across requirements, not generic value-prop bullets)
4. **What is our realistic winrate?** (Shipley-derived factor model: fit, incumbent, relationship strength, decision-criteria alignment, late-entry, competitor count, deal size — produces estimate + confidence band)
5. **Should we bid?** (deterministic verdict: BID / PARTNER-BID / NO-BID with named factors driving the call)
The skill surfaces GAPs explicitly. Leadership decides whether to close them, partner around them, or no-bid. **It never invents claims.**
## When to use
- A 30+ page RFP / RFI / RFQ has landed with a 7-14 day response deadline
- A security questionnaire (SIG, CAIQ, custom-buyer) needs structured Q&A — not prose
- The team is preparing a bid / no-bid review and needs a defensible winrate estimate
- Sales Engineering has a proof-point library but no system to map proofs to requirements
- Leadership wants to see fit % (STRONG / PARTIAL / GAP) before committing pursuit budget
- A late-entry opportunity needs honest assessment of the relationship deficit
**Do not use for:**
- Free-form proposal narrative authoring → `business-growth/contract-and-proposal-writer`
- Contract redline AFTER award → `c-level-advisor/general-counsel-advisor`
- Marketing collateral / category content → `marketing-skill/*`
- Discount approval on the awarded deal → `commercial/deal-desk`
- Pricing-model design for a new product → `commercial/pricing-strategist`
## Workflow
### Step 1 — Parse the RFP
Drop the RFP markdown / text into `scripts/rfp_parser.py`. Output: structured JSON listing every requirement, tagged MANDATORY / WEIGHTED / NICE-TO-HAVE based on cue words (must / shall = MANDATORY; should / weighted scoring numbers = WEIGHTED; may / preferred / desired = NICE-TO-HAVE). Captures section structure, scoring criteria if disclosed, deadline, submission format constraints.
```bash
python scripts/rfp_parser.py --input rfp.md --output json > parsed.json
```
### Step 2 — Score fit per requirement
Fill `assets/rfp_intake_template.md` with your proof-point library (each proof tagged with type + verifiable source + which requirement-tags it covers) and proposed win-themes. Feed parsed RFP + intake into `scripts/response_drafter.py`. Output: proof-point matrix per requirement with STRONG / PARTIAL / GAP, win-theme injection, GAP audit.
```bash
python scripts/response_drafter.py --input draft_input.json --output markdown > matrix.md
```
**Hard rule:** GAP requirements are surfaced, never invented around. Leadership reads the GAP audit and decides: close the gap, partner-bid, or no-bid.
### Step 3 — Apply win-theme strategy
Shipley method: 3-5 themes that span requirements. Each theme answers "why us over the incumbent / competitor on the criteria the buyer named." `response_drafter.py` shows which themes thread through which requirements — a theme appearing in <2 requirements is decorative, not strategic, and gets flagged.
### Step 4 — Estimate winrate
Feed deal context (fit %, incumbent strength, relationship, decision-criteria alignment, late-entry, competitor count, deal size vs. average) into `scripts/winrate_predictor.py`. Output: Shipley-derived estimate 0-100% + confidence band + factor breakdown + BID / PARTNER-BID / NO-BID verdict.
```bash
python scripts/winrate_predictor.py --input deal_context.json --profile enterprise-software --output markdown
```
**No-bid threshold:** estimate < 20% triggers automatic no-bid recommendation.
### Step 5 — Decide
Take parsed RFP + proof-point matrix + GAP audit + winrate estimate into the go / no-go review. Skill does not commit pursuit budget — leadership does.
## Scripts
- `scripts/rfp_parser.py` — section + requirement extractor (regex + cue-word heuristics, stdlib only)
- `scripts/response_drafter.py` — proof-point matrix + win-theme injection + GAP audit
- `scripts/winrate_predictor.py` — Shipley-derived factor model + bid/no-bid verdict, industry-profile-tuned
All scripts: stdlib only (argparse, json, sys, pathlib, re, collections, statistics). `--help` and `--sample` work on all three.
## References
- `references/shipley_method_canon.md` — Shipley Proposal Guide v6, Shipley Capture Guide, APMP BoK, Tom Sant, Tom Searcy + Henry DeVries, Strategic Proposals research, Larry Newman
- `references/rfp_strategy_canon.md` — FAR, GSA, Forrester, Gartner, Bain, McKinsey, B2B International on RFP win-rates and buyer behavior
- `references/rfp_anti_patterns.md` — Shipley failure modes, APMP cases, Strategic Proposals research, federal loss reviews, MIT Sloan, Bain commercial-discipline, Gartner
## Assumptions
- **The RFP is the ground truth.** If the buyer asked it, answer it — in the order they asked, in the format they specified. Re-organizing for narrative flow is for proposals, not RFPs.
- **Proof points must be verifiable.** A claim is only as strong as the case study, certification, customer reference, or technical attestation backing it. Unsourced claims become GAPs.
- **Win-themes are buyer-side, not seller-side.** "We're the leader in X" is a marketing claim; "Your operations team reduces incident MTTR by 60% with the same headcount" is a win-theme. Shipley canon, not optional.
- **Winrate estimates are directional.** The model is a discipline tool to force honest pursuit-qualification — not an oracle. Confidence band always wider than the point estimate suggests.
- **Industry profiles tune base rates** — government RFPs reward compliance discipline; enterprise SaaS rewards reference accounts; healthcare rewards regulatory + security depth.
- **Late entry is a structural disadvantage.** Entering after the RFP issued, with no relationship history, drops base rate ~15%. The skill names this, doesn't hide it.
## Anti-patterns
- **Inventing a proof point to fill a GAP.** Hard rule violation. GAPs surface for leadership decision, not for prose-laundering. See `references/rfp_anti_patterns.md`.
- **Responding to every RFP.** Without a qualified bid / no-bid gate, the team burns capacity on <20% winrate pursuits and loses the 50%+ pursuits to lack of focus. Bain commercial-discipline research.
- **Generic response with no win-theme.** A proposal that could be sent verbatim by any competitor is decorative. Shipley failure mode #1.
- **Missing a mandatory disqualifier late.** FedRAMP / HIPAA / ISO 27001 / SOC 2 / on-shore data residency caught on Day 12 of a 14-day response = wasted pursuit. Parser surfaces these on Day 1.
- **Answering the question YOU wanted asked.** RFP responder discipline: answer what they asked, in their words, in their order. Re-framing belongs in cover letters, not in the compliance matrix.
- **No compliance matrix.** Every requirement should map to a response section + page number. Evaluators score on a matrix; respondents who don't provide one self-disqualify on traceability.
- **Late-entry without acknowledging the relationship deficit.** Entering cold against an incumbent with a 3-year relationship and no champion = sub-20% winrate. Pretending otherwise wastes Sales Engineering capacity.
- **Treating WEIGHTED requirements like MANDATORY.** Score-weighted requirements reward depth on the high-weight items, not uniform mediocrity across all. Shipley capture method.
## Distinct from
- **`business-growth/contract-and-proposal-writer`** — free-form narrative proposals where YOU set the structure (executive briefs, capability statements, unsolicited proposals). RFP-responder handles **buyer-dictated structured Q&A** where the buyer set the questions, sections, scoring criteria, and format. Different artifact, different decision logic.
- **`c-level-advisor/general-counsel-advisor`** — contract redline and IP/risk review AFTER award. RFP-responder operates BEFORE award, on the response strategy.
- **`marketing-skill/*`** — external marketing assets (web copy, content, ASO, SEO, brand voice) for many-to-many audiences. RFP-responder produces a **single-buyer artifact** with deterministic compliance requirements.
- **`commercial/deal-desk`** — per-deal discount routing on a closing opportunity. RFP-responder is pursuit-stage; deal-desk is close-stage.
- **`commercial/pricing-strategist`** — pricing-model design for a new product. RFP-responder consumes existing pricing as input to the commercial-terms section.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time before any script runs. Recommended answer + canon citation per question. Never bundled.
1. **"What's your STRONG / PARTIAL / GAP split on the MANDATORY requirements?"**
Recommended: STRONG ≥ 70% on MANDATORY before bidding. PARTIAL/GAP on any MANDATORY = either close the gap pre-submission or no-bid.
Canon: Shipley *Proposal Guide v6* — capture-management discipline, "Pgw (probability of win) is bounded by your weakest MANDATORY."
2. **"Is there an incumbent, and how strong is their position?"**
Recommended: strong incumbent (3+ years, no displacement event) drops base winrate ~30%. Don't bid without a named displacement trigger.
Canon: Forrester B2B-RFP research — incumbents win 70-80% of renewal RFPs absent a named failure event.
3. **"Did you enter the conversation before or after the RFP issued?"**
Recommended: late-entry (after RFP issued, no prior engagement) drops winrate ~15% and signals the RFP was scoped to someone else's strengths.
Canon: Tom Searcy + Henry DeVries *How to Win Big Business* — "If you didn't help write the RFP, you're column fodder."
4. **"What are your 3-5 win-themes, and does each thread through ≥2 requirements?"**
Recommended: themes that appear in only one requirement are decorative. Themes must ladder up across MANDATORY + WEIGHTED sections.
Canon: Shipley *Capture Guide* — win-themes are the buyer-side answer to "why us" across the evaluation criteria, not seller-side feature lists.
5. **"For every claim in the response, can you name the verifiable source?"**
Recommended: every claim → case study / certification / customer reference / technical attestation / benchmark. Unsourced claims = GAPs.
Canon: APMP BoK — "Substantiation: every assertion in a proposal must be backed by evidence the evaluator can independently verify."
6. **"What's the bid / no-bid threshold you committed to BEFORE seeing this RFP?"**
Recommended: pre-committed threshold (e.g., winrate ≥ 25%, STRONG ≥ 70% on MANDATORY, named champion). Post-hoc rationalization is how teams end up bidding 5% pursuits.
Canon: Bain RFP-win-rate studies — disciplined bid/no-bid gates lift win-rate from ~15% to ~35%.
7. **"What does the buyer's evaluation team actually score on?"**
Recommended: if the RFP discloses scoring criteria, weight your response effort proportionally. If undisclosed, ask. If you can't ask, that itself is a relationship-deficit signal.
Canon: Strategic Proposals proposal-management research — evaluators score on the rubric they were given, not on your narrative.
Walk depth-first. Lock 1-3 before opening 4-7. After all 7 are answered, invoke `rfp_parser.py` → `response_drafter.py` → `winrate_predictor.py` in sequence. If question 6 lands on "we don't have a threshold," set one now or no-bid.
FILE:assets/rfp_intake_template.md
# RFP Intake Template
Fill this in BEFORE running `scripts/response_drafter.py` and `scripts/winrate_predictor.py`. Save as `rfp_intake.json` — the JSON skeleton at the bottom of this file is the canonical input format.
## Step 1 — Deal context
| Field | Value | Notes |
|---|---|---|
| Buyer organization | | |
| RFP title / ID | | |
| Submission deadline | | Date format YYYY-MM-DD |
| Estimated deal size (ACV / TCV) | | |
| Deal size vs. our average | below / at / above | Above-average deals attract more competitors |
| Incumbent | name or "none" | |
| Incumbent strength | none / weak / strong | Strong = 3+ years, no displacement event |
| Relationship strength | cold / warm / champion | Champion = internal advocate willing to push for us |
| Champion name + role | | If relationship_strength = "warm" or "champion" |
| Late entry? | yes / no | "Yes" if we entered AFTER the RFP issued |
| Decision-criteria alignment | 0-100% | How well our strengths match what the buyer says they're scoring on |
| Competitor count | integer | Best estimate; ask the buyer if you can |
| Industry profile | saas / enterprise-software / services / government / healthcare | Tunes `winrate_predictor.py` |
## Step 2 — Proof-point library
For every proof point your team can produce, fill in a row. Verifiable source is **mandatory** — if you can't name where the evaluator could verify it, the proof point doesn't qualify as STRONG.
| Name | Type | Tags (match against requirement text) | Verifiable source |
|---|---|---|---|
| SOC 2 Type II report (2026) | cert | soc, 2, type, ii, certification | trust.example.com/soc2-2026.pdf |
| 24/7 SOC staffing attestation | technical_attestation | soc, 24/7, coverage, on-call, rotation | SecOps runbook v3.2 |
| ... | ... | ... | ... |
**Proof-point types:**
- `case_study` — full customer story with quantified outcome
- `cert` — third-party certification
- `customer_quote` — attributed, approved customer quote
- `technical_attestation` — internal but verifiable (runbook, architecture doc)
- `benchmark` — quantified peer comparison (Gartner, Forrester, internal)
## Step 3 — Win-themes (3-5)
Shipley discipline: each theme must thread through ≥2 requirements. Themes appearing once are decorative.
1. **Theme:** _________
**Threads through which requirement IDs:** _________
2. **Theme:** _________
3. **Theme:** _________
4. **Theme:** _________
5. **Theme:** _________
## Step 4 — Bid/no-bid threshold (set BEFORE seeing the RFP)
Pre-commit your threshold to avoid post-hoc rationalization:
- [ ] Winrate estimate ≥ ___ %
- [ ] STRONG match ≥ ___ % on MANDATORY requirements
- [ ] Named champion at buyer org
- [ ] MANDATORY GAP count ≤ ___
- [ ] Industry profile permits (e.g., do we no-bid government RFPs by default?)
## JSON skeleton — for `response_drafter.py --input`
```json
{
"rfp_requirements_path": "parsed.json",
"proof_points_library": [
{
"name": "SOC 2 Type II report (2026)",
"type": "cert",
"requirement_match_tags": ["soc", "2", "type", "ii", "certification"],
"verifiable_source": "https://trust.example.com/soc2-2026.pdf"
},
{
"name": "AWS/GCP/Azure logging case study (Globex)",
"type": "case_study",
"requirement_match_tags": ["aws", "gcp", "azure", "logging", "integrate"],
"verifiable_source": "globex-cs-2025.pdf"
}
],
"win_themes": [
"operational simplicity at scale",
"financial-services regulatory depth",
"MTTD leadership vs Gartner peer cohort"
]
}
```
## JSON skeleton — for `winrate_predictor.py --input`
```json
{
"requirement_fit_pct_strong": 60.0,
"requirement_fit_pct_partial": 25.0,
"requirement_fit_pct_gap": 15.0,
"incumbent_advantage": "weak",
"relationship_strength": "warm",
"decision_criteria_alignment_pct": 75.0,
"late_entry": false,
"competitor_count": 3,
"deal_size_vs_avg": "at"
}
```
## Running the pipeline
```bash
# 1. Parse the RFP
python scripts/rfp_parser.py --input rfp.md --output json > parsed.json
# 2. Build the proof-point matrix + GAP audit + win-theme report
python scripts/response_drafter.py --input rfp_intake.json --output markdown > matrix.md
# 3. Compute fit % from the matrix, fill into deal_context.json, then:
python scripts/winrate_predictor.py --input deal_context.json --profile enterprise-software --output markdown
# 4. Take parsed RFP + matrix + winrate into the go/no-go review.
```
## Hard rule reminder
**Never invent claims for GAP requirements.** Surface them. Leadership decides: close the gap, partner-bid, or no-bid.
FILE:references/rfp_anti_patterns.md
# RFP Anti-Patterns — Failure Modes the Skill Refuses to Enable
Eight RFP-response failure modes documented across Shipley failure-mode analyses, APMP case studies, Strategic Proposals research, federal loss reviews, MIT Sloan B2B research, Bain commercial-discipline studies, and Gartner. Each anti-pattern names what goes wrong, why teams fall into it, and how the skill prevents it.
## 1. Inventing claims to fill GAP requirements
**Failure mode:** A MANDATORY requirement has no matching proof point. Under deadline pressure, the proposal team writes prose that implies coverage without naming a verifiable source.
**Why it happens:** The team confuses "we could probably do this" with "we have done this and can prove it." Sales pressure to bid combines with no-one-wants-to-be-the-one-who-said-no dynamics.
**Why it loses:** Evaluators verify. When references, certifications, or technical attestations don't substantiate the claim, the response loses on credibility AND on the original requirement. APMP case-study data: invented claims are detected in 60-80% of evaluations and cause loss-of-trust effects that cascade across other sections.
**How the skill prevents it:** `response_drafter.py` surfaces GAP requirements explicitly. Leadership decides: close the gap pre-submission, partner-bid, or no-bid. The skill refuses to generate proof-point language for GAP rows. **Hard rule.**
## 2. No bid/no-bid review — respond to every RFP
**Failure mode:** Every RFP gets a response. Win-rate collapses to 5-12%; sales-engineering capacity burns on pursuits with no relationship, no fit, no champion.
**Why it happens:** Sales teams optimize for activity metrics, not win-rate. Marketing measures responses-sent, not responses-won.
**Why it loses:** Bain research: disciplined bid/no-bid gates lift win-rate from ~15% to ~35%. Without a gate, the team is structurally outperformed by competitors who qualified out and concentrated resources on winnable pursuits.
**How the skill prevents it:** `winrate_predictor.py` produces an explicit BID / PARTNER-BID / NO-BID verdict. <20% estimate triggers automatic NO-BID. The skill names this in writing — leadership cannot override silently.
## 3. Missing mandatory disqualifiers until Day 12
**Failure mode:** FedRAMP, HIPAA, ISO 27001, SOC 2, on-shore data residency — a MANDATORY certification or compliance requirement is buried on page 47 of the RFP and discovered after 10 days of proposal work.
**Why it happens:** No parse-pass on Day 1. The team reads the RFP as prose, not as a structured requirement set.
**Why it loses:** The pursuit is unrecoverable. All work product to date is wasted. Worse, the team loses 2 weeks of capacity that could have been spent on winnable pursuits.
**How the skill prevents it:** `rfp_parser.py` runs on Day 1, tags every MANDATORY requirement, and produces a compliance-matrix view. MANDATORY GAPs surface immediately, not on Day 12.
## 4. No win-theme — generic response
**Failure mode:** The response could be sent verbatim by any competitor. Capabilities are listed; differentiation is implicit; the "why us" answer is decorative ("we're the leader in X").
**Why it happens:** Win-themes are hard. They require buyer-side framing ("your team reduces X by Y") rather than seller-side feature lists. Teams default to feature lists because they're easy to write.
**Why it loses:** Shipley failure-mode analysis: generic responses lose 70%+ of evaluations where any competitor produced a buyer-anchored win-theme. Evaluators ladder themes back to evaluation criteria; generic responses can't do this.
**How the skill prevents it:** `response_drafter.py` threads each declared win-theme through the requirements list. Themes appearing in <2 requirements are flagged **DECORATIVE**. The skill forces theme-discipline.
## 5. Answering the question you wanted asked, not the question they asked
**Failure mode:** The team reframes buyer questions to match their proposal narrative. Section structure is re-ordered for "flow." Buyer-specific terminology is replaced with seller-preferred vocabulary.
**Why it happens:** Habit. Proposal teams trained on free-form proposals carry that discipline into RFP responses. Marketing prefers branded vocabulary.
**Why it loses:** Strategic Proposals research: evaluators score on traceability. A response that doesn't visibly answer the buyer's question in the buyer's order loses 20-30 points of available score before content quality is assessed.
**How the skill prevents it:** `rfp_parser.py` extracts requirements in the buyer's order with the buyer's text preserved. `response_drafter.py` builds the compliance matrix on the buyer's requirement IDs. Reframing is not supported.
## 6. No compliance matrix — no traceability
**Failure mode:** The response is a long prose document. No table shows which requirement is answered on which page. Evaluators scoring against a 60-row rubric give up after 15 minutes of search and default-score.
**Why it happens:** Compliance matrices are tedious to maintain when content changes. Teams skip them under deadline pressure.
**Why it loses:** APMP BoK: response traceability is one of the top-3 evaluator-cited differentiators. Without a matrix, the evaluator scores on what they can find — which is less than what you wrote.
**How the skill prevents it:** `response_drafter.py` outputs a markdown compliance matrix as its primary artifact. Every requirement → match level → proof point → verifiable source. The matrix IS the response architecture.
## 7. Late-entry without acknowledging the relationship deficit
**Failure mode:** The team enters the RFP cold. No prior engagement, no champion, no executive sponsor at the buyer. The proposal is written as if entry timing didn't matter.
**Why it happens:** Optimism bias. The team believes content quality can overcome structural disadvantage.
**Why it loses:** Forrester: late-entry vendors win 8-12% of RFPs vs 25-35% for capture-engaged vendors. Federal RFP loss reviews show late-entry as the #1 named factor in 40%+ of post-mortems.
**How the skill prevents it:** `winrate_predictor.py` requires `late_entry` as input. Setting it to `true` applies a −15% penalty. The estimate honestly reflects the structural deficit; leadership decides whether to spend pursuit budget anyway.
## 8. Treating WEIGHTED requirements like MANDATORY
**Failure mode:** The team gives equal effort to every WEIGHTED requirement. A 25-point requirement and a 5-point requirement get the same proof-depth, the same page count, the same proof-point recruitment effort.
**Why it happens:** No effort-weighting against the scoring rubric. Either the rubric wasn't disclosed and the team didn't ask, or the rubric was disclosed and the team ignored it.
**Why it loses:** Shipley capture math: WEIGHTED scores compound. Optimizing the top-3 weighted requirements (typically 60-70% of available points) wins more often than uniform-mediocrity across all weighted requirements. McKinsey B2B research: rubric-weighted-effort respondents win 1.6x more than equal-effort respondents.
**How the skill prevents it:** `rfp_parser.py` extracts disclosed scoring weights into the requirement evidence. `response_drafter.py` shows weights in the compliance matrix. Forcing-question #7 ("What does the buyer's evaluation team actually score on?") interrogates whether the weighting was even requested.
## Sources
1. **Shipley Associates failure-mode analyses** — internal post-loss reviews published in *Proposal Guide v6* appendix and in *Capture Guide* case studies. Source for anti-patterns 1, 4, 5.
2. **APMP (Association of Proposal Management Professionals) case studies** — APMP BoK appendix and APMP Journal case studies. Source for anti-pattern 1 (invented-claim detection rates) and anti-pattern 6 (traceability as top-3 differentiator).
3. **Strategic Proposals (strategicproposals.com) research and benchmarks** — published rubric-replication-gap data; source for anti-pattern 5 (evaluator-traceability scoring).
4. **Federal RFP loss reviews** — debrief reports available through FOIA and GSA's procurement transparency programs. Source for anti-pattern 7 (late-entry as #1 named loss factor in 40%+ of post-mortems).
5. **MIT Sloan B2B sales research**, MIT Sloan Management Review archives. Source for anti-pattern 8 (rubric-weighted-effort win-rate multiplier).
6. **Bain & Company commercial-discipline studies** — Bain B2B sales practice publications and conference presentations. Source for anti-pattern 2 (disciplined-pursuit win-rate of ~35% vs respond-to-everything ~12%).
7. **Gartner, RFP Best Practices and IT Buyer Studies**. Source for anti-pattern 6 (compliance-matrix presence as evaluator-cited differentiator) and the general industry-vertical evaluation-cycle benchmarks.
8. **Patrick Lencioni, *Getting Naked* (Jossey-Bass, 2010)**. Source for the "tell the kind truth" principle that operationalizes the skill's GAP-honesty hard rule (anti-pattern 1).
FILE:references/rfp_strategy_canon.md
# RFP Strategy Canon — Industry Research on RFP Win-Rates and Buyer Behavior
This reference grounds the `winrate_predictor.py` factor weights in published industry research. The model is opinionated but defensible: every factor maps to a citation below.
## Headline findings the skill encodes
### Base win-rates are honestly grim
- Average competitive B2B RFP win-rate: 15-25% across industries (Bain, Gartner).
- With disciplined bid/no-bid qualification: 35-45%.
- Without qualification: 5-12% — sales-engineering capacity burned on unwinnable pursuits.
The skill's 20% NO-BID threshold is calibrated to land below the disciplined-pursuit floor.
### Incumbents win renewal RFPs 70-80% of the time
Absent a named failure event (security breach, missed SLA, executive turnover at the incumbent), incumbents win 70-80% of renewal RFPs (Forrester B2B-RFP research). This is the empirical basis for the −30% incumbent penalty when incumbent_advantage is "strong."
### Late entry is structurally penalized
If you weren't part of the conversation before the RFP issued, the RFP was scoped to someone else's strengths. Forrester data: late-entry vendors win 8-12% of RFPs vs 25-35% for vendors who engaged in capture. The skill's −15% late-entry penalty is the midpoint of this gap.
### Relationship strength dominates content quality at the margin
Bain: in deals where the named champion advocates internally, win-rate lifts 20-30 percentage points over the "warm but no champion" baseline. The skill's +25% champion factor is the lower bound of this range.
### Decision-criteria alignment is bimodal
When buyer decision criteria align >80% with your strengths, win-rate is roughly 2x the base rate. When alignment is <50%, win-rate collapses to ~30% of base (McKinsey B2B sales research). The skill encodes this as a +10 / 0 / −10 step function rather than a continuous curve, because the bimodality is the honest reality.
### Competitor count compresses win-rate predictably
- 1 competitor (sole-source consideration): 60-80% win-rate
- 2 competitors: 35-50%
- 3 competitors: 20-30%
- 4-5 competitors: 12-18%
- 6+ competitors: 5-10%
The skill's competitor-count factor (+20 / +5 / 0 / -10 / -20) tracks this curve.
## Industry profile tuning
The skill exposes 5 profiles via `--profile`. Each shifts the base rate:
- **enterprise-software (+5)**: longer sales cycles, deeper technical evaluation, but disciplined buyers reward fit-honest vendors. Base rate slightly above average.
- **saas (0)**: market baseline.
- **services (−5)**: commoditized for many engagement types, weaker differentiation moats, harder to defend price.
- **government (−15)**: FAR-governed, compliance-heavy, incumbent-favored, evaluation timelines extend 2-4x. Forrester / GSA data.
- **healthcare (−10)**: regulatory overhead (HIPAA, FDA, HITRUST), risk-averse procurement, longer pilot cycles. Gartner healthcare-vertical research.
## What this skill deliberately does NOT model
- **Pricing positioning** — outside scope; consume from `commercial/pricing-strategist`.
- **Proposal aesthetics / production quality** — Shipley canon says these matter at the margin (3-5 percentage points) but never override fit, win-themes, and relationship. Skill omits.
- **Evaluator psychology** — Strategic Proposals research shows evaluators score on the rubric they were given. The skill assumes the rubric is the source of truth; theme-injection happens within rubric constraints.
## Sources
1. **Federal Acquisition Regulation (FAR)**, especially Parts 14 (Sealed Bidding) and 15 (Contracting by Negotiation), at acquisition.gov/far. Governs US federal RFPs. Defines the compliance-matrix requirement, evaluation-factor disclosure rules, and proposal-format constraints that drive the "government" profile penalty.
2. **GSA (General Services Administration) RFP and procurement guidance**, at gsa.gov. Quantifies federal evaluation timelines (typically 90-180 days) and the disproportionate weight federal evaluators give to past-performance citations — relevant to proof-point substantiation discipline.
3. **Forrester Research, B2B Buyer Studies** — recurring annual research on B2B buying behavior. Sources the 70-80% incumbent renewal-win-rate, the late-entry penalty, and the "5-10 vendor longlist" reality of modern RFP processes.
4. **Gartner, RFP Best Practices** — published guidance for IT-buyer organizations. Quantifies vendor-shortlist sizes by deal value, evaluation-cycle length by industry, and the structural advantage of fit-honest responses over feature-checklist responses.
5. **Bain & Company, B2B Sales and RFP-Win-Rate Research** — Bain's commercial-discipline practice publishes regular benchmarks on disciplined-pursuit win-rates (35-45%) vs respond-to-everything win-rates (5-12%). The 20% NO-BID threshold in `winrate_predictor.py` is calibrated against this data.
6. **McKinsey & Company, B2B Sales Practice** — McKinsey research on decision-criteria alignment and win-rate. Sources the bimodal alignment effect (>80% alignment doubles base rate; <50% collapses to 30% of base) encoded in `alignment_factor()`.
7. **B2B International (now Kantar B2B), Buyer Behavior in RFP Processes** — research on how B2B evaluation committees actually score responses. Confirms that compliance-matrix presence, proof-point substantiation, and rubric-aligned response structure are the top-3 evaluator-cited differentiators.
8. **Patrick Lencioni, *Getting Naked: A Business Fable About Shedding the Three Fears That Sabotage Client Loyalty*** (Jossey-Bass, 2010). The "we don't have a proof point for this — here's what we'd do instead" honesty discipline that informs the skill's hard rule: surface GAPs, never invent. Lencioni's "tell the kind truth" principle operationalized as a refusal to fabricate evidence.
FILE:references/shipley_method_canon.md
# Shipley Method Canon — RFP Response Discipline
The Shipley method is the dominant industry methodology for capture management and proposal development. This reference distils what `rfp-responder` consumes from it: capture-stage qualification, win-theme construction, proof-point substantiation, and the discipline that separates structured responses from prose proposals.
## What Shipley actually claims
Shipley's central claim is that **proposals are won in capture, not in writing**. By the time the RFP issues, 70-80% of the eventual outcome is determined by the capture work done in the preceding 6-18 months. The RFP-response phase executes a strategy — it does not create one from scratch.
This skill operationalizes the capture-output side: parsing the RFP into discrete requirements, scoring fit honestly (STRONG / PARTIAL / GAP), threading win-themes across requirements, and producing a defensible winrate estimate.
## Core concepts the skill implements
### 1. Compliance matrix
Every requirement must map to a response section + page number. Evaluators score on a matrix; respondents who don't provide one self-disqualify on traceability. `response_drafter.py` builds this matrix; `rfp_parser.py` extracts the requirement IDs that anchor it.
### 2. Win-themes (buyer-side, not seller-side)
A win-theme is the buyer-side answer to "why us over the competitor on the criteria the buyer named." It is NOT "we're the leader in X." Win-themes ladder up across multiple requirements — Shipley canon is that a theme appearing in only one requirement is **decorative**, not strategic. The skill flags these explicitly.
### 3. Proof points with substantiation
APMP BoK: "every assertion in a proposal must be backed by evidence the evaluator can independently verify." Five proof-point types the skill recognizes:
- **case_study** — full customer story with quantified outcome
- **cert** — third-party certification (SOC 2, ISO 27001, FedRAMP, HIPAA)
- **customer_quote** — attributed quote, customer-approved
- **technical_attestation** — internal but verifiable (runbook, architecture doc, SOC staffing rotation)
- **benchmark** — quantified comparison vs peers (Gartner, Forrester, internal)
STRONG = ≥2 tag matches AND proof type in {case_study, cert, technical_attestation, benchmark}.
PARTIAL = 1 match, or proof type is customer_quote.
GAP = 0 matches → surfaced for leadership, **never invented around**.
### 4. Pgw (probability of win) bounded by weakest MANDATORY
Shipley capture discipline: Pgw cannot exceed the score on your weakest MANDATORY requirement. A 90% fit on 9 of 10 MANDATORY items and a GAP on the 10th is not a 90% bid — it is a 0% bid until the GAP is closed or partnered around.
### 5. Bid / no-bid gate
A disciplined bid/no-bid gate lifts win-rate from ~15% to ~35% (Bain). The skill enforces this: winrate <20% → automatic NO-BID; 20-34% → PARTNER-BID; ≥35% → BID with full pursuit budget.
## What Shipley is NOT
- Not a prose-writing methodology — Shipley is structured, requirement-anchored, scoreable.
- Not optional for federal/regulated RFPs — FAR-governed RFPs are essentially Shipley-compatible by procurement design.
- Not a substitute for relationship capital — late-entry without prior engagement still penalizes ~15% even with perfect Shipley execution.
## Sources
1. **Shipley Associates, *Proposal Guide v6***, Larry Newman (Ed.), Shipley Associates Press. The canonical book. Defines capture-management, compliance matrix, win-themes, ghosting, theme statements, proof-point substantiation.
2. **Shipley Associates, *Capture Guide***. The capture-stage companion to the Proposal Guide. Defines the 6-stage capture lifecycle (opportunity identification → capture planning → solution development → preliminary bid decision → solution validation → final bid decision) the skill assumes has been done before it runs.
3. **APMP (Association of Proposal Management Professionals) *Body of Knowledge (BoK)***. International proposal-management standard. Defines substantiation discipline, evaluator-side scoring rubrics, compliance-matrix traceability requirements, and the Foundation / Practitioner / Professional certification tiers that anchor the industry.
4. **Tom Sant, *Persuasive Business Proposals: Writing to Win More Customers, Clients, and Contracts*** (3rd ed., AMACOM, 2012). Defines the NOSE pattern (Need, Outcome, Solution, Evidence) that the skill's proof-point matrix operationalizes. Sant's discipline: every solution claim must close with evidence.
5. **Tom Searcy & Henry DeVries, *How to Win Big Business: How to Sell Multi-Million Dollar Contracts***. Defines the relationship-deficit principle the skill encodes in the late-entry penalty: "If you didn't help write the RFP, you're column fodder." The skill's −15% late-entry factor comes from this canon.
6. **Strategic Proposals (proposal-management consultancy) — published research and benchmarks (strategicproposals.com)**. Quantifies the evaluator-rubric gap: respondents who don't replicate the evaluator's scoring weights in their response structure lose 20-30 percentage points of available score regardless of content quality.
7. **Larry Newman, "The Shipley Method"** — the methodology articulation that anchors *Proposal Guide v6*. Defines the 7-step proposal-development process (kickoff → blue team → pink team → red team → gold team → submission → debrief) and the color-team review discipline.
8. **CapturePlanning.com / FederalProposalLibrary** — community-maintained resources synthesizing Shipley + federal-acquisition discipline. Useful complement for government RFP profile tuning in `winrate_predictor.py --profile government`.
FILE:scripts/response_drafter.py
#!/usr/bin/env python3
"""response_drafter.py - Build a Shipley-method proof-point matrix + GAP audit + win-theme injection.
Stdlib only. Deterministic logic. NEVER invents claims to fill GAP requirements.
Inputs (JSON):
{
"rfp_requirements": [...] OR "rfp_requirements_path": "parsed.json"
"proof_points_library": [
{
"name": "...",
"type": "case_study|cert|customer_quote|technical_attestation|benchmark",
"requirement_match_tags": ["soc2", "saml", "aws", ...],
"verifiable_source": "..."
}, ...
],
"win_themes": ["operational simplicity", "financial-services depth", ...]
}
For each requirement:
- Tokenize the requirement text (lowercase, strip punctuation, dedupe, drop stopwords).
- For each proof point, intersect proof.requirement_match_tags with requirement tokens.
- If 2+ tag matches AND proof.type in {case_study, cert, technical_attestation, benchmark}
-> STRONG
- If 1 tag match OR proof.type in {customer_quote}
-> PARTIAL
- If 0 matches
-> GAP
For each win-theme: count how many requirements it threads through.
Theme appearing in <2 requirements -> flag as "DECORATIVE", not strategic.
Output: response-draft markdown (or JSON) with:
- Compliance matrix (every requirement -> proof + match level)
- GAP audit (explicit, no inventing)
- Win-theme coverage report
Usage:
python response_drafter.py --sample
python response_drafter.py --input draft_input.json --output markdown
python response_drafter.py --input draft_input.json --output json
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from collections import Counter
from pathlib import Path
from typing import Any
STOPWORDS = {
"the", "a", "an", "is", "are", "was", "were", "be", "been", "being",
"and", "or", "but", "of", "to", "in", "on", "at", "for", "with", "by",
"must", "shall", "should", "may", "will", "would", "could",
"vendor", "vendors", "platform", "provide", "provides", "support", "supports",
"this", "that", "these", "those", "it", "its", "our", "your",
"required", "mandatory", "optional", "preferred", "desired",
"from", "as", "if", "than", "then", "do", "does", "did",
"have", "has", "had",
}
STRONG_PROOF_TYPES = {"case_study", "cert", "technical_attestation", "benchmark"}
PARTIAL_PROOF_TYPES = {"customer_quote"}
SAMPLE_INPUT = {
"rfp_requirements": [
{"id": "R001", "section": "Mandatory", "tag": "MANDATORY",
"text": "Vendor must hold SOC 2 Type II certification.", "evidence": {}},
{"id": "R002", "section": "Mandatory", "tag": "MANDATORY",
"text": "Vendor shall provide 24/7 SOC coverage with named on-call rotation.", "evidence": {}},
{"id": "R003", "section": "Mandatory", "tag": "MANDATORY",
"text": "Vendor is required to support SAML 2.0 and SCIM provisioning.", "evidence": {}},
{"id": "R004", "section": "Mandatory", "tag": "MANDATORY",
"text": "The platform must integrate with AWS, GCP, and Azure native logging.", "evidence": {}},
{"id": "R005", "section": "Weighted", "tag": "WEIGHTED",
"text": "Mean Time to Detect benchmarks vs peers.", "evidence": {"points": 25}},
{"id": "R006", "section": "Weighted", "tag": "WEIGHTED",
"text": "Customer references in financial services.", "evidence": {"points": 20}},
{"id": "R007", "section": "Nice-to-Have", "tag": "NICE-TO-HAVE",
"text": "FedRAMP authorization is preferred but not required.", "evidence": {}},
],
"proof_points_library": [
{"name": "SOC 2 Type II report (2026)", "type": "cert",
"requirement_match_tags": ["soc", "2", "type", "ii", "certification"],
"verifiable_source": "https://trust.example.com/soc2-2026.pdf"},
{"name": "24/7 SOC staffing attestation", "type": "technical_attestation",
"requirement_match_tags": ["soc", "24/7", "coverage", "on-call", "rotation"],
"verifiable_source": "internal SecOps runbook v3.2"},
{"name": "SAML/SCIM integration guide", "type": "technical_attestation",
"requirement_match_tags": ["saml", "scim", "provisioning"],
"verifiable_source": "docs.example.com/saml-scim"},
{"name": "AWS/GCP/Azure logging case study (Globex)", "type": "case_study",
"requirement_match_tags": ["aws", "gcp", "azure", "logging", "integrate", "native"],
"verifiable_source": "globex-cs-2025.pdf"},
{"name": "MTTD benchmark vs Gartner peer cohort", "type": "benchmark",
"requirement_match_tags": ["mttd", "mean", "time", "detect", "benchmarks", "peers"],
"verifiable_source": "Gartner MQ supplement 2026"},
{"name": "Financial services customer quote (FNB)", "type": "customer_quote",
"requirement_match_tags": ["financial", "services", "customer", "references"],
"verifiable_source": "FNB CISO quote, approved 2026-03"},
],
"win_themes": [
"operational simplicity at scale",
"financial-services regulatory depth",
"MTTD leadership vs Gartner peer cohort",
"AWS/GCP/Azure native logging without bolt-ons",
],
}
def tokenize(text: str) -> set[str]:
tokens = re.findall(r"[a-zA-Z0-9./]+", text.lower())
return {t for t in tokens if t not in STOPWORDS and len(t) > 1}
def score_match(requirement: dict[str, Any], proof: dict[str, Any]) -> tuple[int, list[str]]:
"""Return (match_count, matched_tags)."""
req_tokens = tokenize(requirement["text"])
matched = [tag for tag in proof.get("requirement_match_tags", []) if tag.lower() in req_tokens]
return len(matched), matched
def assign_proof(requirement: dict[str, Any], library: list[dict[str, Any]]) -> dict[str, Any]:
best_count = 0
best_proof: dict[str, Any] | None = None
best_matched: list[str] = []
for proof in library:
count, matched = score_match(requirement, proof)
if count > best_count:
best_count = count
best_proof = proof
best_matched = matched
if best_proof is None or best_count == 0:
return {"level": "GAP", "proof": None, "matched_tags": []}
if best_count >= 2 and best_proof["type"] in STRONG_PROOF_TYPES:
level = "STRONG"
elif best_count >= 1 and best_proof["type"] in STRONG_PROOF_TYPES:
level = "PARTIAL"
elif best_count >= 1 and best_proof["type"] in PARTIAL_PROOF_TYPES:
level = "PARTIAL"
else:
level = "PARTIAL"
return {"level": level, "proof": best_proof, "matched_tags": best_matched}
def thread_themes(requirements: list[dict[str, Any]], themes: list[str]) -> dict[str, dict[str, Any]]:
"""For each theme, list requirements whose text overlaps theme tokens."""
report: dict[str, dict[str, Any]] = {}
for theme in themes:
theme_tokens = tokenize(theme)
threaded: list[str] = []
for req in requirements:
req_tokens = tokenize(req["text"])
if theme_tokens & req_tokens:
threaded.append(req["id"])
verdict = "STRATEGIC" if len(threaded) >= 2 else "DECORATIVE"
report[theme] = {
"requirement_ids": threaded,
"count": len(threaded),
"verdict": verdict,
}
return report
def build_matrix(payload: dict[str, Any]) -> dict[str, Any]:
if "rfp_requirements_path" in payload and "rfp_requirements" not in payload:
p = Path(payload["rfp_requirements_path"])
loaded = json.loads(p.read_text(encoding="utf-8"))
requirements = loaded.get("requirements", loaded if isinstance(loaded, list) else [])
else:
requirements = payload.get("rfp_requirements", [])
library = payload.get("proof_points_library", [])
themes = payload.get("win_themes", [])
matrix: list[dict[str, Any]] = []
for req in requirements:
assignment = assign_proof(req, library)
matrix.append({
"requirement_id": req["id"],
"tag": req["tag"],
"section": req.get("section", ""),
"text": req["text"],
"match_level": assignment["level"],
"proof_name": assignment["proof"]["name"] if assignment["proof"] else None,
"proof_type": assignment["proof"]["type"] if assignment["proof"] else None,
"verifiable_source": assignment["proof"]["verifiable_source"] if assignment["proof"] else None,
"matched_tags": assignment["matched_tags"],
})
level_counts = Counter(row["match_level"] for row in matrix)
mandatory_gaps = [row for row in matrix if row["tag"] == "MANDATORY" and row["match_level"] == "GAP"]
theme_report = thread_themes(requirements, themes)
return {
"matrix": matrix,
"level_counts": dict(level_counts),
"mandatory_gap_count": len(mandatory_gaps),
"mandatory_gaps": mandatory_gaps,
"win_theme_report": theme_report,
"requirement_total": len(requirements),
}
def render_markdown(result: dict[str, Any]) -> str:
out: list[str] = []
out.append("# RFP Response Draft — Proof-Point Matrix\n")
total = result["requirement_total"]
counts = result["level_counts"]
out.append(f"**Requirements:** {total}")
if total > 0:
strong = counts.get("STRONG", 0)
partial = counts.get("PARTIAL", 0)
gap = counts.get("GAP", 0)
out.append(f"**STRONG:** {strong} ({100*strong/total:.0f}%) | "
f"**PARTIAL:** {partial} ({100*partial/total:.0f}%) | "
f"**GAP:** {gap} ({100*gap/total:.0f}%)")
out.append(f"\n**MANDATORY GAPs:** {result['mandatory_gap_count']} "
"(LEADERSHIP DECISION REQUIRED — close gap, partner-bid, or no-bid)\n")
out.append("## Compliance matrix\n")
out.append("| Req | Tag | Match | Proof | Source |")
out.append("|---|---|---|---|---|")
for row in result["matrix"]:
proof = row["proof_name"] or "**(NO PROOF — GAP)**"
source = row["verifiable_source"] or "—"
out.append(f"| {row['requirement_id']} | {row['tag']} | {row['match_level']} | {proof} | {source} |")
out.append("")
if result["mandatory_gaps"]:
out.append("## GAP audit (MANDATORY requirements without proof)\n")
out.append("> HARD RULE: do NOT invent claims for these. Leadership decides: "
"close the gap pre-submission, partner-bid, or no-bid.\n")
for row in result["mandatory_gaps"]:
out.append(f"- **{row['requirement_id']}** ({row['section']}): {row['text']}")
out.append("")
out.append("## Win-theme coverage\n")
for theme, info in result["win_theme_report"].items():
ids = ", ".join(info["requirement_ids"]) or "(none)"
out.append(f"- **{theme}** — threads through {info['count']} req(s): {ids} → **{info['verdict']}**")
out.append("")
decorative = [t for t, info in result["win_theme_report"].items() if info["verdict"] == "DECORATIVE"]
if decorative:
out.append("### Decorative themes (flagged)\n")
out.append("These themes appear in <2 requirements and are decorative, not strategic. "
"Either remove or strengthen so they thread across multiple sections.\n")
for t in decorative:
out.append(f"- {t}")
out.append("")
return "\n".join(out) + "\n"
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="Build proof-point matrix + GAP audit + win-theme report.")
parser.add_argument("--input", help="Path to draft-input JSON.")
parser.add_argument("--output", choices=["json", "markdown"], default="markdown")
parser.add_argument("--sample", action="store_true", help="Use built-in synthetic input.")
args = parser.parse_args(argv)
if args.sample:
payload = SAMPLE_INPUT
elif args.input:
path = Path(args.input)
if not path.exists():
print(f"ERROR: input file not found: {args.input}", file=sys.stderr)
return 1
payload = json.loads(path.read_text(encoding="utf-8"))
else:
parser.print_help()
return 0
result = build_matrix(payload)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_markdown(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/rfp_parser.py
#!/usr/bin/env python3
"""rfp_parser.py - Parse an RFP / RFI / RFQ / security questionnaire into structured requirements.
Stdlib only. Regex + cue-word heuristics. No NLP libraries, no LLM calls.
The parser:
1. Splits the document into sections (executive summary, technical requirements,
security questionnaire, commercial terms, timeline, etc.) using common heading
patterns.
2. Extracts requirements as discrete numbered / bulleted / "must/shall/should" lines.
3. Tags each requirement MANDATORY / WEIGHTED / NICE-TO-HAVE based on cue words:
MANDATORY - must, shall, required, mandatory, "is required to"
WEIGHTED - should, weighted scoring numbers present (e.g., "[20 points]"),
"evaluation criteria", "scored"
NICE-TO-HAVE - may, preferred, desired, nice-to-have, optional
4. Captures disclosed scoring criteria (lines that look like "X points" / "X%" weights).
5. Captures submission deadline + format requirements (regex on common date patterns
+ "format" / "submission" cue words).
Usage:
python rfp_parser.py --sample
python rfp_parser.py --input rfp.md --output json
python rfp_parser.py --input rfp.md --output markdown
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from collections import Counter
from pathlib import Path
from typing import Any
SAMPLE_RFP = """\
# RFP-2026-CLOUD-SECURITY-007
## 1. Executive Summary
Acme Holdings is seeking a cloud security platform vendor.
Total contract value: $1.5M over 3 years.
Submission deadline: 2026-06-14.
Format: PDF, max 80 pages, 11pt font minimum.
## 2. Mandatory Requirements
2.1 Vendor must hold SOC 2 Type II certification.
2.2 Vendor shall provide 24/7 SOC coverage with named on-call rotation.
2.3 Vendor is required to support SAML 2.0 and SCIM provisioning.
2.4 The platform must integrate with AWS, GCP, and Azure native logging.
2.5 Vendor shall meet a 99.9% platform uptime SLA.
## 3. Weighted Requirements (100 points total)
3.1 Threat detection coverage breadth [30 points]
3.2 Mean Time to Detect (MTTD) benchmarks vs peers [25 points]
3.3 Customer references in financial services [20 points]
3.4 Implementation timeline shorter than 90 days [15 points]
3.5 Quality of executive briefing materials [10 points]
The platform should support custom detection rule authoring.
Vendor should provide quarterly threat intelligence reports.
## 4. Nice-to-Have Capabilities
4.1 FedRAMP authorization is preferred but not required.
4.2 ISO 27001 certification is desired.
4.3 The platform may offer AI-assisted triage capabilities.
4.4 Vendor support for on-premises deployment is optional.
## 5. Commercial Terms
Multi-year discount expected. Payment terms NET-45.
## 6. Submission Format
Responses must be submitted via the procurement portal by 2026-06-14 17:00 ET.
Late submissions will not be accepted.
"""
MANDATORY_CUES = re.compile(
r"\b(must|shall|required|mandatory|is required to|are required to|will be required)\b",
re.IGNORECASE,
)
WEIGHTED_CUES = re.compile(
r"\b(should|evaluation criteria|scored|weighted|preferred)\b",
re.IGNORECASE,
)
NICE_CUES = re.compile(
r"\b(may|preferred but not required|desired|nice[- ]to[- ]have|optional|is desired)\b",
re.IGNORECASE,
)
POINTS_PATTERN = re.compile(r"\[(\d+)\s*(?:points?|pts?|%)\]", re.IGNORECASE)
DEADLINE_PATTERN = re.compile(
r"(deadline|due|submission|submit by|responses? (?:are )?due)\s*[:\-]?\s*"
r"(\d{4}[-/]\d{1,2}[-/]\d{1,2}|\d{1,2}[-/]\d{1,2}[-/]\d{2,4})",
re.IGNORECASE,
)
FORMAT_PATTERN = re.compile(
r"\b(format|page limit|max(?:imum)? \d+ pages?|font|portal|pdf|word|submitted via)\b",
re.IGNORECASE,
)
HEADING_PATTERN = re.compile(r"^(#{1,3})\s+(.+?)\s*$")
REQ_LINE_PATTERN = re.compile(r"^\s*(\d+\.\d+|\d+\)|-|\*)\s+(.+?)\s*$")
def classify_requirement(text: str) -> tuple[str, dict[str, Any]]:
"""Return (tag, evidence_dict). Precedence: NICE > MANDATORY > WEIGHTED.
NICE-TO-HAVE is checked first because phrases like "preferred but not required"
contain the word "required" but are NOT mandatory.
"""
evidence: dict[str, Any] = {"matched_cues": []}
nice_match = NICE_CUES.search(text)
if nice_match:
evidence["matched_cues"].append(nice_match.group(0).lower())
return "NICE-TO-HAVE", evidence
mand_match = MANDATORY_CUES.search(text)
if mand_match:
evidence["matched_cues"].append(mand_match.group(0).lower())
return "MANDATORY", evidence
points_match = POINTS_PATTERN.search(text)
if points_match:
evidence["points"] = int(points_match.group(1))
evidence["matched_cues"].append(f"[{points_match.group(1)} points]")
return "WEIGHTED", evidence
weight_match = WEIGHTED_CUES.search(text)
if weight_match:
evidence["matched_cues"].append(weight_match.group(0).lower())
return "WEIGHTED", evidence
return "UNCLASSIFIED", evidence
def split_sections(text: str) -> list[dict[str, Any]]:
"""Split document into sections by markdown headings."""
sections: list[dict[str, Any]] = []
current = {"heading": "(preamble)", "level": 0, "body": []}
for line in text.splitlines():
m = HEADING_PATTERN.match(line)
if m:
if current["body"] or current["heading"] != "(preamble)":
sections.append(current)
current = {
"heading": m.group(2).strip(),
"level": len(m.group(1)),
"body": [],
}
else:
current["body"].append(line)
sections.append(current)
return [s for s in sections if s["body"] or s["heading"] != "(preamble)"]
def extract_requirements(sections: list[dict[str, Any]]) -> list[dict[str, Any]]:
"""Extract individual requirements from sections."""
reqs: list[dict[str, Any]] = []
req_counter = 0
for sec in sections:
section_label = sec["heading"]
for raw_line in sec["body"]:
line = raw_line.strip()
if not line or line.startswith("#"):
continue
m = REQ_LINE_PATTERN.match(raw_line)
text = m.group(2).strip() if m else line
# Only count lines that contain at least one classification cue or a points tag.
if not (
MANDATORY_CUES.search(text)
or WEIGHTED_CUES.search(text)
or NICE_CUES.search(text)
or POINTS_PATTERN.search(text)
):
continue
tag, evidence = classify_requirement(text)
req_counter += 1
reqs.append({
"id": f"R{req_counter:03d}",
"section": section_label,
"text": text,
"tag": tag,
"evidence": evidence,
})
return reqs
def extract_scoring(text: str) -> list[dict[str, Any]]:
"""Find lines with explicit point weights."""
scoring: list[dict[str, Any]] = []
for line in text.splitlines():
m = POINTS_PATTERN.search(line)
if m:
scoring.append({"weight": int(m.group(1)), "line": line.strip()})
return scoring
def extract_deadline(text: str) -> str | None:
m = DEADLINE_PATTERN.search(text)
return m.group(2) if m else None
def extract_format_notes(text: str) -> list[str]:
notes: list[str] = []
for line in text.splitlines():
if FORMAT_PATTERN.search(line) and len(line.strip()) < 200:
notes.append(line.strip())
# Dedupe while preserving order.
seen: set[str] = set()
out: list[str] = []
for n in notes:
if n not in seen:
seen.add(n)
out.append(n)
return out
def parse(text: str) -> dict[str, Any]:
sections = split_sections(text)
reqs = extract_requirements(sections)
tag_counts = Counter(r["tag"] for r in reqs)
return {
"section_count": len(sections),
"sections": [{"heading": s["heading"], "level": s["level"]} for s in sections],
"requirement_count": len(reqs),
"tag_breakdown": dict(tag_counts),
"requirements": reqs,
"scoring_criteria": extract_scoring(text),
"deadline": extract_deadline(text),
"format_notes": extract_format_notes(text),
}
def render_markdown(parsed: dict[str, Any]) -> str:
out: list[str] = []
out.append("# RFP Parse Report\n")
out.append(f"**Sections detected:** {parsed['section_count']}")
out.append(f"**Requirements detected:** {parsed['requirement_count']}")
out.append(f"**Deadline:** {parsed['deadline'] or '(not detected)'}\n")
out.append("## Requirement breakdown\n")
for tag, count in parsed["tag_breakdown"].items():
out.append(f"- {tag}: {count}")
out.append("\n## Requirements\n")
for r in parsed["requirements"]:
out.append(f"### {r['id']} — [{r['tag']}]")
out.append(f"**Section:** {r['section']}")
out.append(f"**Text:** {r['text']}")
if r["evidence"].get("points"):
out.append(f"**Points:** {r['evidence']['points']}")
out.append(f"**Matched cues:** {', '.join(r['evidence']['matched_cues']) or '(none)'}")
out.append("")
out.append("## Scoring criteria detected\n")
if parsed["scoring_criteria"]:
for s in parsed["scoring_criteria"]:
out.append(f"- [{s['weight']} pts] {s['line']}")
else:
out.append("(none disclosed)")
out.append("\n## Format notes\n")
if parsed["format_notes"]:
for n in parsed["format_notes"]:
out.append(f"- {n}")
else:
out.append("(none detected)")
return "\n".join(out) + "\n"
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="Parse an RFP into structured requirements.")
parser.add_argument("--input", help="Path to RFP markdown/text file.")
parser.add_argument("--output", choices=["json", "markdown"], default="markdown")
parser.add_argument("--sample", action="store_true", help="Use built-in synthetic RFP.")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_RFP
elif args.input:
path = Path(args.input)
if not path.exists():
print(f"ERROR: input file not found: {args.input}", file=sys.stderr)
return 1
text = path.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
parsed = parse(text)
if args.output == "json":
print(json.dumps(parsed, indent=2))
else:
print(render_markdown(parsed))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/winrate_predictor.py
#!/usr/bin/env python3
"""winrate_predictor.py - Shipley-derived winrate estimate + bid/no-bid verdict.
Stdlib only. Deterministic factor model.
Inputs (JSON):
{
"requirement_fit_pct_strong": 60.0, # % of requirements matched at STRONG
"requirement_fit_pct_partial": 30.0, # % at PARTIAL
"requirement_fit_pct_gap": 10.0, # % at GAP
"incumbent_advantage": "none|weak|strong",
"relationship_strength": "cold|warm|champion",
"decision_criteria_alignment_pct": 75.0,
"late_entry": true|false, # entered after RFP issued, no prior engagement
"competitor_count": 3,
"deal_size_vs_avg": "below|at|above"
}
Factor model (Shipley-derived, opinionated, industry-tunable):
base = 0.03 * fit_strong - 0.02 * fit_gap + 0.005 * fit_partial
(STRONG counts 3x, PARTIAL 1x, GAP -2x in Shipley capture math;
encoded here as a linear bounded score centered to produce a
baseline win-rate in the 5-80% range)
Incumbent penalty:
none -> 0
weak -> -10
strong -> -30
Relationship lift:
cold -> 0
warm -> +10
champion -> +25
Late entry: -15 if true, 0 otherwise
Decision-criteria alignment:
pct >= 80 -> +10
50 <= pct < 80 -> 0
pct < 50 -> -10
Competitor count:
1 (you're sole vendor) -> +20
2 -> +5
3 -> 0
4-5 -> -10
6+ -> -20
Deal size vs avg:
at -> 0
above -> -5 (bigger deals attract more scrutiny + more competitors)
below -> 0
Industry profile shifts the base rate (the structural reality that government RFPs
are harder than enterprise software):
enterprise-software: base_shift = +5
saas: base_shift = 0
services: base_shift = -5
government: base_shift = -15
healthcare: base_shift = -10
Verdict:
< 20% -> NO-BID
20-34% -> PARTNER-BID (find a partner who closes the structural gap)
35-100% -> BID
Confidence band: +/- 12 percentage points (wider on small-sample factor inputs).
Usage:
python winrate_predictor.py --sample
python winrate_predictor.py --input deal.json --profile enterprise-software
python winrate_predictor.py --input deal.json --profile government --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
from typing import Any
PROFILES: dict[str, dict[str, float]] = {
"enterprise-software": {"base_shift": 5.0},
"saas": {"base_shift": 0.0},
"services": {"base_shift": -5.0},
"government": {"base_shift": -15.0},
"healthcare": {"base_shift": -10.0},
}
SAMPLE_INPUT = {
"requirement_fit_pct_strong": 60.0,
"requirement_fit_pct_partial": 25.0,
"requirement_fit_pct_gap": 15.0,
"incumbent_advantage": "weak",
"relationship_strength": "warm",
"decision_criteria_alignment_pct": 75.0,
"late_entry": False,
"competitor_count": 3,
"deal_size_vs_avg": "at",
}
def incumbent_factor(level: str) -> float:
return {"none": 0.0, "weak": -10.0, "strong": -30.0}.get(level, 0.0)
def relationship_factor(level: str) -> float:
return {"cold": 0.0, "warm": 10.0, "champion": 25.0}.get(level, 0.0)
def alignment_factor(pct: float) -> float:
if pct >= 80.0:
return 10.0
if pct < 50.0:
return -10.0
return 0.0
def competitor_factor(count: int) -> float:
if count <= 1:
return 20.0
if count == 2:
return 5.0
if count == 3:
return 0.0
if count <= 5:
return -10.0
return -20.0
def deal_size_factor(size: str) -> float:
return {"at": 0.0, "above": -5.0, "below": 0.0}.get(size, 0.0)
def base_from_fit(strong: float, partial: float, gap: float) -> float:
"""STRONG 3x, PARTIAL 1x, GAP -2x; calibrated to land in 5-80% range at extremes."""
raw = 0.03 * strong * 3.0 + 0.01 * partial - 0.02 * gap * 2.0
# Center to a sensible baseline. raw of 9 = 100% strong -> ~45 baseline.
return max(0.0, min(80.0, raw * 5.0))
def predict(payload: dict[str, Any], profile: str) -> dict[str, Any]:
prof = PROFILES.get(profile, PROFILES["saas"])
strong = float(payload.get("requirement_fit_pct_strong", 0.0))
partial = float(payload.get("requirement_fit_pct_partial", 0.0))
gap = float(payload.get("requirement_fit_pct_gap", 0.0))
base = base_from_fit(strong, partial, gap)
inc = incumbent_factor(payload.get("incumbent_advantage", "none"))
rel = relationship_factor(payload.get("relationship_strength", "cold"))
late = -15.0 if payload.get("late_entry", False) else 0.0
align = alignment_factor(float(payload.get("decision_criteria_alignment_pct", 50.0)))
comp = competitor_factor(int(payload.get("competitor_count", 3)))
size = deal_size_factor(payload.get("deal_size_vs_avg", "at"))
estimate = base + inc + rel + late + align + comp + size + prof["base_shift"]
estimate = max(0.0, min(100.0, estimate))
band_lo = max(0.0, estimate - 12.0)
band_hi = min(100.0, estimate + 12.0)
if estimate < 20.0:
verdict = "NO-BID"
rationale = ("Estimated winrate below the 20% no-bid threshold. "
"Pursuing this RFP burns sales-engineering capacity without "
"a credible path to win.")
elif estimate < 35.0:
verdict = "PARTNER-BID"
rationale = ("Estimate in the 20-34% band. Bid only with a partner who closes "
"the structural gap (incumbent, late-entry, MANDATORY-GAP, or "
"regulatory-fit deficit). Solo bid not recommended.")
else:
verdict = "BID"
rationale = ("Estimate above 35%. Pursue with full Shipley capture discipline: "
"win-themes laddered across requirements, MANDATORY GAPs closed pre-submission, "
"proof-points sourced, executive sponsor named.")
return {
"profile": profile,
"winrate_estimate_pct": round(estimate, 1),
"confidence_band_pct": [round(band_lo, 1), round(band_hi, 1)],
"verdict": verdict,
"rationale": rationale,
"factor_breakdown": {
"base_from_fit": round(base, 1),
"incumbent_advantage": round(inc, 1),
"relationship_strength": round(rel, 1),
"late_entry": round(late, 1),
"decision_criteria_alignment": round(align, 1),
"competitor_count": round(comp, 1),
"deal_size_vs_avg": round(size, 1),
"industry_profile_shift": round(prof["base_shift"], 1),
},
}
def render_markdown(result: dict[str, Any]) -> str:
out: list[str] = []
out.append("# Shipley-Derived Winrate Estimate\n")
out.append(f"**Profile:** {result['profile']}")
band = result["confidence_band_pct"]
out.append(f"**Estimate:** {result['winrate_estimate_pct']}% (band: {band[0]}% – {band[1]}%)")
out.append(f"**Verdict:** **{result['verdict']}**\n")
out.append(f"> {result['rationale']}\n")
out.append("## Factor breakdown\n")
out.append("| Factor | Contribution (pp) |")
out.append("|---|---|")
for k, v in result["factor_breakdown"].items():
sign = "+" if v >= 0 else ""
out.append(f"| {k} | {sign}{v} |")
out.append("")
out.append("## Reading the estimate\n")
out.append("- Estimate is **directional**, not an oracle. Treat the band as the honest range.")
out.append("- A high score does NOT override a MANDATORY GAP — close the gap or no-bid.")
out.append("- A low score with a champion + named executive sponsor can be reconsidered, "
"but document the rationale before committing pursuit budget.")
return "\n".join(out) + "\n"
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Shipley-derived winrate estimate + bid/no-bid verdict."
)
parser.add_argument("--input", help="Path to deal-context JSON.")
parser.add_argument("--profile", choices=list(PROFILES.keys()), default="saas",
help="Industry profile (default: saas).")
parser.add_argument("--output", choices=["json", "markdown"], default="markdown")
parser.add_argument("--sample", action="store_true", help="Use built-in synthetic input.")
args = parser.parse_args(argv)
if args.sample:
payload = SAMPLE_INPUT
elif args.input:
path = Path(args.input)
if not path.exists():
print(f"ERROR: input file not found: {args.input}", file=sys.stderr)
return 1
payload = json.loads(path.read_text(encoding="utf-8"))
else:
parser.print_help()
return 0
result = predict(payload, args.profile)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_markdown(result))
return 0
if __name__ == "__main__":
sys.exit(main())
Đánh giá mã nguồn theo hướng phản biện khắt khe, phát hiện điểm mù trước khi merge PR.
--- name: "adversarial-reviewer" description: "Adversarial code review that breaks the self-review monoculture. Use when you want a genuinely critical review of recent changes, before merging a PR, or when you suspect Claude is being too agreeable about code quality. Forces perspective shifts through hostile reviewer personas that catch blind spots the author's mental model shares with the reviewer." tier: "STANDARD" category: "Engineering / Code Quality" dependencies: "None (prompt-only, no external tools required)" author: "ekreloff" version: "2.9.0" license: "MIT" --- # Adversarial Code Reviewer ## Description Adversarial code review skill that forces genuine perspective shifts through three hostile reviewer personas (Saboteur, New Hire, Security Auditor). Each persona MUST find at least one issue — no "LGTM" escapes. Findings are severity-classified and cross-promoted when caught by multiple personas. ## Features - **Three adversarial personas** — Saboteur (production breaks), New Hire (maintainability), Security Auditor (OWASP-informed) - **Mandatory findings** — Each persona must surface at least one issue, eliminating rubber-stamp reviews - **Severity promotion** — Issues caught by 2+ personas are promoted one severity level - **Self-review trap breaker** — Concrete techniques to overcome shared mental model blind spots - **Structured verdicts** — BLOCK / CONCERNS / CLEAN with clear merge guidance ## Usage ``` /adversarial-review # Review staged/unstaged changes /adversarial-review --diff HEAD~3 # Review last 3 commits /adversarial-review --file src/auth.ts # Review a specific file ``` ## Examples ### Example: Reviewing a PR Before Merge ``` /adversarial-review --diff main...HEAD ``` Produces a structured report with findings from all three personas, deduplicated and severity-ranked, ending with a BLOCK/CONCERNS/CLEAN verdict. ## Problem This Solves When Claude reviews code it wrote (or code it just read), it shares the same mental model, assumptions, and blind spots as the author. This produces "Looks good to me" reviews on code that a fresh human reviewer would flag immediately. Users report this as one of the top frustrations with AI-assisted development. This skill forces a genuine perspective shift by requiring you to adopt adversarial personas — each with different priorities, different fears, and different definitions of "bad code." ## Table of Contents 1. [Quick Start](#quick-start) 2. [Review Workflow](#review-workflow) 3. [The Three Personas](#the-three-personas) 4. [Severity Classification](#severity-classification) 5. [Output Format](#output-format) 6. [Anti-Patterns](#anti-patterns) 7. [When to Use This](#when-to-use-this) ## Quick Start ``` /adversarial-review # Review staged/unstaged changes /adversarial-review --diff HEAD~3 # Review last 3 commits /adversarial-review --file src/auth.ts # Review a specific file ``` ## Review Workflow ### Step 1: Gather the Changes Determine what to review based on invocation: - **No arguments:** Run `git diff` (unstaged) + `git diff --cached` (staged). If both empty, run `git diff HEAD~1` (last commit). - **`--diff <ref>`:** Run `git diff <ref>`. - **`--file <path>`:** Read the entire file. Focus review on the full file rather than just changes. If no changes are found, stop and report: "Nothing to review." ### Step 2: Read the Full Context For every file in the diff: 1. Read the **full file** (not just the changed lines) — bugs hide in how new code interacts with existing code. 2. Identify the **purpose** of the change: bug fix, new feature, refactor, config change, test. 3. Note any **project conventions** from CLAUDE.md, .editorconfig, linting configs, or existing patterns. ### Step 3: Run All Three Personas Execute each persona sequentially. Each persona MUST produce at least one finding. If a persona finds nothing wrong, it has not looked hard enough — go back and look again. **IMPORTANT:** Do not soften findings. Do not hedge. Do not say "this might be fine but..." — either it's a problem or it isn't. Be direct. ### Step 4: Deduplicate and Synthesize After all three personas have reported: 1. Merge duplicate findings (same issue caught by multiple personas). 2. Promote findings caught by 2+ personas to the next severity level. 3. Produce the final structured output. ## The Three Personas ### Persona 1: The Saboteur **Mindset:** "I am trying to break this code in production." **Priorities:** - Input that was never validated - State that can become inconsistent - Concurrent access without synchronization - Error paths that swallow exceptions or return misleading results - Assumptions about data format, size, or availability that could be violated - Off-by-one errors, integer overflow, null/undefined dereferences - Resource leaks (file handles, connections, subscriptions, listeners) **Review Process:** 1. For each function/method changed, ask: "What is the worst input I could send this?" 2. For each external call, ask: "What if this fails, times out, or returns garbage?" 3. For each state mutation, ask: "What if this runs twice? Concurrently? Never?" 4. For each conditional, ask: "What if neither branch is correct?" **You MUST find at least one issue. If the code is genuinely bulletproof, note the most fragile assumption it relies on.** --- ### Persona 2: The New Hire **Mindset:** "I just joined this team. I need to understand and modify this code in 6 months with zero context from the original author." **Priorities:** - Names that don't communicate intent (what does `data` mean? what does `process()` do?) - Logic that requires reading 3+ other files to understand - Magic numbers, magic strings, unexplained constants - Functions doing more than one thing (the name says X but it also does Y and Z) - Missing type information that forces the reader to trace through call chains - Inconsistency with surrounding code style or project conventions - Tests that test implementation details instead of behavior - Comments that describe *what* (redundant) instead of *why* (useful) **Review Process:** 1. Read each changed function as if you've never seen the codebase. Can you understand what it does from the name, parameters, and body alone? 2. Trace one code path end-to-end. How many files do you need to open? 3. Check: would a new contributor know where to add a similar feature? 4. Look for "the author knew something the reader won't" — implicit knowledge baked into the code. **You MUST find at least one issue. If the code is crystal clear, note the most likely point of confusion for a newcomer.** --- ### Persona 3: The Security Auditor **Mindset:** "This code will be attacked. My job is to find the vulnerability before an attacker does." **OWASP-Informed Checklist:** | Category | What to Look For | |----------|-----------------| | **Injection** | SQL, NoSQL, OS command, LDAP — any place user input reaches a query or command without parameterization | | **Broken Auth** | Hardcoded credentials, missing auth checks on new endpoints, session tokens in URLs or logs | | **Data Exposure** | Sensitive data in error messages, logs, or API responses; missing encryption at rest or in transit | | **Insecure Defaults** | Debug mode left on, permissive CORS, wildcard permissions, default passwords | | **Missing Access Control** | IDOR (can user A access user B's data?), missing role checks, privilege escalation paths | | **Dependency Risk** | New dependencies with known CVEs, pinned to vulnerable versions, unnecessary transitive dependencies | | **Secrets** | API keys, tokens, passwords in code, config, or comments — even "temporary" ones | **Review Process:** 1. Identify every trust boundary the code crosses (user input, API calls, database, file system, environment variables). 2. For each boundary: is input validated? Is output sanitized? Is the principle of least privilege followed? 3. Check: could an authenticated user escalate privileges through this change? 4. Check: does this change expose any new attack surface? **You MUST find at least one issue. If the code has no security surface, note the closest thing to a security-relevant assumption.** ## Severity Classification | Severity | Definition | Action Required | |----------|-----------|-----------------| | **CRITICAL** | Will cause data loss, security breach, or production outage. Must fix before merge. | Block merge. | | **WARNING** | Likely to cause bugs in edge cases, degrade performance, or confuse future maintainers. Should fix before merge. | Fix or explicitly accept risk with justification. | | **NOTE** | Style issue, minor improvement opportunity, or documentation gap. Nice to fix. | Author's discretion. | **Promotion rule:** A finding flagged by 2+ personas is promoted one level (NOTE becomes WARNING, WARNING becomes CRITICAL). ## Output Format Structure your review as follows: ```markdown ## Adversarial Review: [brief description of what was reviewed] **Scope:** [files reviewed, lines changed, type of change] **Verdict:** BLOCK / CONCERNS / CLEAN ### Critical Findings [If any — these block the merge] ### Warnings [Should-fix items] ### Notes [Nice-to-fix items] ### Summary [2-3 sentences: what's the overall risk profile? What's the single most important thing to fix?] ``` **Verdict definitions:** - **BLOCK** — 1+ CRITICAL findings. Do not merge until resolved. - **CONCERNS** — No criticals but 2+ warnings. Merge at your own risk. - **CLEAN** — Only notes. Safe to merge. ## Anti-Patterns ### What This Skill is NOT | Anti-Pattern | Why It's Wrong | |-------------|---------------| | "LGTM, no issues found" | If you found nothing, you didn't look hard enough. Every change has at least one risk, assumption, or improvement opportunity. | | Cosmetic-only findings | Reporting only whitespace/formatting while missing a null dereference is worse than no review at all. Substance first, style second. | | Pulling punches | "This might possibly be a minor concern..." — No. Be direct. "This will throw a NullPointerException when `user` is undefined." | | Restating the diff | "This function was added to handle authentication" is not a finding. What's WRONG with how it handles authentication? | | Ignoring test gaps | New code without tests is a finding. Always. Tests are not optional. | | Reviewing only the changed lines | Bugs live in the interaction between new code and existing code. Read the full file. | ### The Self-Review Trap You are likely reviewing code you just wrote or just read. Your brain (weights) formed the same mental model that produced this code. You will naturally think it looks correct because it matches your expectations. **To break this pattern:** 1. Read the code **bottom-up** (start from the last function, work backward). 2. For each function, state its contract **before** reading the body. Does the body match? 3. Assume every variable could be null/undefined until proven otherwise. 4. Assume every external call will fail. 5. Ask: "If I deleted this change entirely, what would break?" — if the answer is "nothing," the change might be unnecessary. ## When to Use This - **Before merging any PR** — especially self-authored PRs with no human reviewer - **After a long coding session** — fatigue produces blind spots; this skill compensates - **When Claude said "looks good"** — if you got an easy approval, run this for a second opinion - **On security-sensitive code** — auth, payments, data access, API endpoints - **When something "feels off"** — trust that instinct and run an adversarial review ## Cross-References - Related: `engineering-team/senior-security` — deep security analysis - Related: `engineering-team/code-reviewer` — general code quality review - Complementary: `ra-qm-team/` — quality management workflows
Vòng lặp thử nghiệm tự động tối ưu một tệp theo chỉ số đo được, giữ bản cải thiện và loại bản thất bại.
---
name: "autoresearch-agent"
description: "Autonomous experiment loop that optimizes any file by a measurable metric. Inspired by Karpathy's autoresearch. The agent edits a target file, runs a fixed evaluation, keeps improvements (git commit), discards failures (git reset), and loops indefinitely. Use when: user wants to optimize code speed, reduce bundle/image size, improve test pass rate, optimize prompts, improve content quality (headlines, copy, CTR), or run any measurable improvement loop. Requires: a target file, an evaluation command that outputs a metric, and a git repo."
license: MIT
metadata:
version: 2.0.0
author: Alireza Rezvani
category: engineering
updated: 2026-03-13
---
# Autoresearch Agent
> You sleep. The agent experiments. You wake up to results.
Autonomous experiment loop inspired by [Karpathy's autoresearch](https://github.com/karpathy/autoresearch). The agent edits one file, runs a fixed evaluation, keeps improvements, discards failures, and loops indefinitely.
Not one guess — fifty measured attempts, compounding.
---
## Slash Commands
| Command | What it does |
|---------|-------------|
| `/ar:setup` | Set up a new experiment interactively |
| `/ar:run` | Run a single experiment iteration |
| `/ar:loop` | Start autonomous loop with configurable interval (10m, 1h, daily, weekly, monthly) |
| `/ar:status` | Show dashboard and results |
| `/ar:resume` | Resume a paused experiment |
---
## When This Skill Activates
Recognize these patterns from the user:
- "Make this faster / smaller / better"
- "Optimize [file] for [metric]"
- "Improve my [headlines / copy / prompts]"
- "Run experiments overnight"
- "I want to get [metric] from X to Y"
- Any request involving: optimize, benchmark, improve, experiment loop, autoresearch
If the user describes a target file + a way to measure success → this skill applies.
---
## Setup
### First Time — Create the Experiment
Run the setup script. The user decides where experiments live:
**Project-level** (inside repo, git-tracked, shareable with team):
```bash
python scripts/setup_experiment.py \
--domain engineering \
--name api-speed \
--target src/api/search.py \
--eval "pytest bench.py --tb=no -q" \
--metric p50_ms \
--direction lower \
--scope project
```
**User-level** (personal, in `~/.autoresearch/`):
```bash
python scripts/setup_experiment.py \
--domain marketing \
--name medium-ctr \
--target content/titles.md \
--eval "python evaluate.py" \
--metric ctr_score \
--direction higher \
--evaluator llm_judge_content \
--scope user
```
The `--scope` flag determines where `.autoresearch/` lives:
- `project` (default) → `.autoresearch/` in the repo root. Experiment definitions are git-tracked. Results are gitignored.
- `user` → `~/.autoresearch/` in the home directory. Everything is personal.
### What Setup Creates
```
.autoresearch/
├── config.yaml ← Global settings
├── .gitignore ← Ignores results.tsv, *.log
└── {domain}/{experiment-name}/
├── program.md ← Objectives, constraints, strategy
├── config.cfg ← Target, eval cmd, metric, direction
├── results.tsv ← Experiment log (gitignored)
└── evaluate.py ← Evaluation script (if --evaluator used)
```
**results.tsv columns:** `commit | metric | status | description`
- `commit` — short git hash
- `metric` — float value or "N/A" for crashes
- `status` — keep | discard | crash
- `description` — what changed or why it crashed
### Domains
| Domain | Use Cases |
|--------|-----------|
| `engineering` | Code speed, memory, bundle size, test pass rate, build time |
| `marketing` | Headlines, social copy, email subjects, ad copy, engagement |
| `content` | Article structure, SEO descriptions, readability, CTR |
| `prompts` | System prompts, chatbot tone, agent instructions |
| `custom` | Anything else with a measurable metric |
### If `program.md` Already Exists
The user may have written their own `program.md`. If found in the experiment directory, read it. It overrides the template. Only ask for what's missing.
---
## Agent Protocol
You are the loop. The scripts handle setup and evaluation — you handle the creative work.
### Before Starting
1. Read `.autoresearch/{domain}/{name}/config.cfg` to get:
- `target` — the file you edit
- `evaluate_cmd` — the command that measures your changes
- `metric` — the metric name to look for in eval output
- `metric_direction` — "lower" or "higher" is better
- `time_budget_minutes` — max time per evaluation
2. Read `program.md` for strategy, constraints, and what you can/cannot change
3. Read `results.tsv` for experiment history (columns: commit, metric, status, description)
4. Checkout the experiment branch: `git checkout autoresearch/{domain}/{name}`
### Each Iteration
1. Review results.tsv — what worked? What failed? What hasn't been tried?
2. Decide ONE change to the target file. One variable per experiment.
3. Edit the target file
4. Commit: `git add {target} && git commit -m "experiment: {description}"`
5. Evaluate: `python scripts/run_experiment.py --experiment {domain}/{name} --single`
6. Read the output — it prints KEEP, DISCARD, or CRASH with the metric value
7. Go to step 1
### What the Script Handles (you don't)
- Running the eval command with timeout
- Parsing the metric from eval output
- Comparing to previous best
- Reverting the commit on failure (`git reset --hard HEAD~1`)
- Logging the result to results.tsv
### Starting an Experiment
```bash
# Single iteration (the agent calls this repeatedly)
python scripts/run_experiment.py --experiment engineering/api-speed --single
# Dry run (test setup before starting)
python scripts/run_experiment.py --experiment engineering/api-speed --dry-run
```
### Strategy Escalation
- Runs 1-5: Low-hanging fruit (obvious improvements, simple optimizations)
- Runs 6-15: Systematic exploration (vary one parameter at a time)
- Runs 16-30: Structural changes (algorithm swaps, architecture shifts)
- Runs 30+: Radical experiments (completely different approaches)
- If no improvement in 20+ runs: update program.md Strategy section
### Self-Improvement
After every 10 experiments, review results.tsv for patterns. Update the
Strategy section of program.md with what you learned (e.g., "caching changes
consistently improve by 5-10%", "refactoring attempts never improve the metric").
Future iterations benefit from this accumulated knowledge.
### Stopping
- Run until interrupted by the user, context limit reached, or goal in program.md is met
- Before stopping: ensure results.tsv is up to date
- On context limit: the next session can resume — results.tsv and git log persist
### Rules
- **One change per experiment.** Don't change 5 things at once. You won't know what worked.
- **Simplicity criterion.** A small improvement that adds ugly complexity is not worth it. Equal performance with simpler code is a win. Removing code that gets same results is the best outcome.
- **Never modify the evaluator.** `evaluate.py` is the ground truth. Modifying it invalidates all comparisons. Hard stop if you catch yourself doing this.
- **Timeout.** If a run exceeds 2.5× the time budget, kill it and treat as crash.
- **Crash handling.** If it's a typo or missing import, fix and re-run. If the idea is fundamentally broken, revert, log "crash", move on. 5 consecutive crashes → pause and alert.
- **No new dependencies.** Only use what's already available in the project.
---
## Evaluators
Ready-to-use evaluation scripts. Copied into the experiment directory during setup with `--evaluator`.
### Free Evaluators (no API cost)
| Evaluator | Metric | Use Case |
|-----------|--------|----------|
| `benchmark_speed` | `p50_ms` (lower) | Function/API execution time |
| `benchmark_size` | `size_bytes` (lower) | File, bundle, Docker image size |
| `test_pass_rate` | `pass_rate` (higher) | Test suite pass percentage |
| `build_speed` | `build_seconds` (lower) | Build/compile/Docker build time |
| `memory_usage` | `peak_mb` (lower) | Peak memory during execution |
### LLM Judge Evaluators (uses your subscription)
| Evaluator | Metric | Use Case |
|-----------|--------|----------|
| `llm_judge_content` | `ctr_score` 0-10 (higher) | Headlines, titles, descriptions |
| `llm_judge_prompt` | `quality_score` 0-100 (higher) | System prompts, agent instructions |
| `llm_judge_copy` | `engagement_score` 0-10 (higher) | Social posts, ad copy, emails |
LLM judges call the CLI tool the user is already running (Claude, Codex, Gemini). The evaluation prompt is locked inside `evaluate.py` — the agent cannot modify it. This prevents the agent from gaming its own evaluator.
The user's existing subscription covers the cost:
- Claude Code Max → unlimited Claude calls for evaluation
- Codex CLI (ChatGPT Pro) → unlimited Codex calls
- Gemini CLI (free tier) → free evaluation calls
### Custom Evaluators
If no built-in evaluator fits, the user writes their own `evaluate.py`. Only requirement: it must print `metric_name: value` to stdout.
```python
#!/usr/bin/env python3
# My custom evaluator — DO NOT MODIFY after experiment starts
import subprocess
result = subprocess.run(["my-benchmark", "--json"], capture_output=True, text=True)
# Parse and output
print(f"my_metric: {parse_score(result.stdout)}")
```
---
## Viewing Results
```bash
# Single experiment
python scripts/log_results.py --experiment engineering/api-speed
# All experiments in a domain
python scripts/log_results.py --domain engineering
# Cross-experiment dashboard
python scripts/log_results.py --dashboard
# Export formats
python scripts/log_results.py --experiment engineering/api-speed --format csv --output results.csv
python scripts/log_results.py --experiment engineering/api-speed --format markdown --output results.md
python scripts/log_results.py --dashboard --format markdown --output dashboard.md
```
### Dashboard Output
```
DOMAIN EXPERIMENT RUNS KEPT BEST Δ FROM START STATUS
engineering api-speed 47 14 185ms -76.9% active
engineering bundle-size 23 8 412KB -58.3% paused
marketing medium-ctr 31 11 8.4/10 +68.0% active
prompts support-tone 15 6 82/100 +46.4% done
```
### Export Formats
- **TSV** — default, tab-separated (compatible with spreadsheets)
- **CSV** — comma-separated, with proper quoting
- **Markdown** — formatted table, readable in GitHub/docs
---
## Proactive Triggers
Flag these without being asked:
- **No evaluation command works** → Test it before starting the loop. Run once, verify output.
- **Target file not in git** → `git init && git add . && git commit -m 'initial'` first.
- **Metric direction unclear** → Ask: is lower or higher better? Must know before starting.
- **Time budget too short** → If eval takes longer than budget, every run crashes.
- **Agent modifying evaluate.py** → Hard stop. This invalidates all comparisons.
- **5 consecutive crashes** → Pause the loop. Alert the user. Don't keep burning cycles.
- **No improvement in 20+ runs** → Suggest changing strategy in program.md or trying a different approach.
---
## Installation
### One-liner (any tool)
```bash
git clone https://github.com/alirezarezvani/claude-skills.git
cp -r claude-skills/engineering/autoresearch-agent ~/.claude/skills/
```
### Multi-tool install
```bash
./scripts/convert.sh --skill autoresearch-agent --tool codex|gemini|cursor|windsurf|openclaw
```
### OpenClaw
```bash
clawhub install cs-autoresearch-agent
```
---
## Related Skills
- **self-improving-agent** — improves an agent's own memory/rules over time. NOT for structured experiment loops.
- **senior-ml-engineer** — ML architecture decisions. Complementary — use for initial design, then autoresearch for optimization.
- **tdd-guide** — test-driven development. Complementary — tests can be the evaluation function.
- **skill-security-auditor** — audit skills before publishing. NOT for optimization loops.
FILE:references/experiment-domains.md
# Experiment Domains Guide
## Domain: Engineering
### Code Speed Optimization
```bash
python scripts/setup_experiment.py \
--domain engineering \
--name api-speed \
--target src/api/search.py \
--eval "python -m pytest tests/bench_search.py --tb=no -q" \
--metric p50_ms \
--direction lower \
--evaluator benchmark_speed
```
**What the agent optimizes:** Algorithm, data structures, caching, query patterns, I/O.
**Cost:** Free — just runs benchmarks.
**Speed:** ~5 min/experiment, ~12/hour, ~100 overnight.
### Bundle Size Reduction
```bash
python scripts/setup_experiment.py \
--domain engineering \
--name bundle-size \
--target webpack.config.js \
--eval "npm run build && python .autoresearch/engineering/bundle-size/evaluate.py" \
--metric size_bytes \
--direction lower \
--evaluator benchmark_size
```
Edit `evaluate.py` to set `TARGET_FILE = "dist/main.js"` and add `BUILD_CMD = "npm run build"`.
### Test Pass Rate
```bash
python scripts/setup_experiment.py \
--domain engineering \
--name fix-flaky-tests \
--target src/utils/parser.py \
--eval "python .autoresearch/engineering/fix-flaky-tests/evaluate.py" \
--metric pass_rate \
--direction higher \
--evaluator test_pass_rate
```
### Docker Build Speed
```bash
python scripts/setup_experiment.py \
--domain engineering \
--name docker-build \
--target Dockerfile \
--eval "python .autoresearch/engineering/docker-build/evaluate.py" \
--metric build_seconds \
--direction lower \
--evaluator build_speed
```
### Memory Optimization
```bash
python scripts/setup_experiment.py \
--domain engineering \
--name memory-usage \
--target src/processor.py \
--eval "python .autoresearch/engineering/memory-usage/evaluate.py" \
--metric peak_mb \
--direction lower \
--evaluator memory_usage
```
### ML Training (Karpathy-style)
Requires NVIDIA GPU. See [autoresearch](https://github.com/karpathy/autoresearch).
```bash
python scripts/setup_experiment.py \
--domain engineering \
--name ml-training \
--target train.py \
--eval "uv run train.py" \
--metric val_bpb \
--direction lower \
--time-budget 5
```
---
## Domain: Marketing
### Medium Article Headlines
```bash
python scripts/setup_experiment.py \
--domain marketing \
--name medium-ctr \
--target content/titles.md \
--eval "python .autoresearch/marketing/medium-ctr/evaluate.py" \
--metric ctr_score \
--direction higher \
--evaluator llm_judge_content
```
Edit `evaluate.py`: set `TARGET_FILE = "content/titles.md"` and `CLI_TOOL = "claude"`.
**What the agent optimizes:** Title phrasing, curiosity gaps, specificity, emotional triggers.
**Cost:** Uses your CLI subscription (Claude Max = unlimited).
**Speed:** ~2 min/experiment, ~30/hour.
### Social Media Copy
```bash
python scripts/setup_experiment.py \
--domain marketing \
--name twitter-engagement \
--target social/tweets.md \
--eval "python .autoresearch/marketing/twitter-engagement/evaluate.py" \
--metric engagement_score \
--direction higher \
--evaluator llm_judge_copy
```
Edit `evaluate.py`: set `PLATFORM = "twitter"` (or linkedin, instagram).
### Email Subject Lines
```bash
python scripts/setup_experiment.py \
--domain marketing \
--name email-open-rate \
--target emails/subjects.md \
--eval "python .autoresearch/marketing/email-open-rate/evaluate.py" \
--metric engagement_score \
--direction higher \
--evaluator llm_judge_copy
```
Edit `evaluate.py`: set `PLATFORM = "email"`.
### Ad Copy
```bash
python scripts/setup_experiment.py \
--domain marketing \
--name ad-copy-q2 \
--target ads/google-search.md \
--eval "python .autoresearch/marketing/ad-copy-q2/evaluate.py" \
--metric engagement_score \
--direction higher \
--evaluator llm_judge_copy
```
Edit `evaluate.py`: set `PLATFORM = "ad"`.
---
## Domain: Content
### Article Structure & Readability
```bash
python scripts/setup_experiment.py \
--domain content \
--name article-structure \
--target drafts/my-article.md \
--eval "python .autoresearch/content/article-structure/evaluate.py" \
--metric ctr_score \
--direction higher \
--evaluator llm_judge_content
```
### SEO Descriptions
```bash
python scripts/setup_experiment.py \
--domain content \
--name seo-meta \
--target seo/descriptions.md \
--eval "python .autoresearch/content/seo-meta/evaluate.py" \
--metric ctr_score \
--direction higher \
--evaluator llm_judge_content
```
---
## Domain: Prompts
### System Prompt Optimization
```bash
python scripts/setup_experiment.py \
--domain prompts \
--name support-bot \
--target prompts/support-system.md \
--eval "python .autoresearch/prompts/support-bot/evaluate.py" \
--metric quality_score \
--direction higher \
--evaluator llm_judge_prompt
```
Requires `tests/cases.json` with test inputs and expected outputs:
```json
[
{
"input": "I can't log in to my account",
"expected": "Ask for email, check account status, offer password reset"
},
{
"input": "How do I cancel my subscription?",
"expected": "Empathetic response, explain cancellation steps, offer retention"
}
]
```
### Agent Skill Optimization
```bash
python scripts/setup_experiment.py \
--domain prompts \
--name skill-improvement \
--target SKILL.md \
--eval "python .autoresearch/prompts/skill-improvement/evaluate.py" \
--metric quality_score \
--direction higher \
--evaluator llm_judge_prompt
```
---
## Choosing Your Domain
| I want to... | Domain | Evaluator | Cost |
|-------------|--------|-----------|------|
| Speed up my code | engineering | benchmark_speed | Free |
| Shrink my bundle | engineering | benchmark_size | Free |
| Fix flaky tests | engineering | test_pass_rate | Free |
| Speed up Docker builds | engineering | build_speed | Free |
| Reduce memory usage | engineering | memory_usage | Free |
| Train ML models | engineering | (custom) | Free + GPU |
| Write better headlines | marketing | llm_judge_content | Subscription |
| Improve social posts | marketing | llm_judge_copy | Subscription |
| Optimize email subjects | marketing | llm_judge_copy | Subscription |
| Improve ad copy | marketing | llm_judge_copy | Subscription |
| Optimize article structure | content | llm_judge_content | Subscription |
| Improve SEO descriptions | content | llm_judge_content | Subscription |
| Optimize system prompts | prompts | llm_judge_prompt | Subscription |
| Improve agent skills | prompts | llm_judge_prompt | Subscription |
**First time?** Start with an engineering experiment (free, fast, measurable). Once comfortable, try content/marketing with LLM judges.
FILE:references/program-template.md
# program.md Templates
Copy the template for your domain and paste into your project root as `program.md`.
---
## ML Training (Karpathy-style)
```markdown
# autoresearch — ML Training
## Goal
Minimize val_bpb on the validation set. Lower is better.
## What You Can Change (train.py only)
- Model architecture (depth, width, attention heads, FFN ratio)
- Optimizer (learning rate, warmup, scheduler, weight decay)
- Training loop (batch size, gradient accumulation, clipping)
- Regularization (dropout, weight tying, etc.)
- Any self-contained improvement that doesn't require new packages
## What You Cannot Change
- prepare.py (fixed — contains evaluation harness)
- Dependencies (pyproject.toml is locked)
- Time budget (always 5 minutes, wall clock)
- Evaluation metric (val_bpb is the ground truth)
## Strategy
1. First run: establish baseline. Do not change anything.
2. Explore learning rate range (try 2x and 0.5x current)
3. Try depth changes (±2 layers)
4. Try optimizer changes (Muon vs. AdamW variants)
5. If things improve, double down. If stuck, try something radical.
## Simplicity Rule
A small improvement with ugly code is NOT worth it.
Equal performance with simpler code IS worth it.
Removing code that gets same results is the best outcome.
## Stop When
val_bpb < 0.95 OR after 100 experiments, whichever comes first.
```
---
## Prompt Engineering
```markdown
# autoresearch — Prompt Optimization
## Goal
Maximize eval_score on the test suite. Higher is better (0-100).
## What You Can Change (prompt.md only)
- System prompt instructions
- Examples and few-shot demonstrations
- Output format specifications
- Chain-of-thought instructions
- Persona and tone
- Task decomposition strategies
## What You Cannot Change
- evaluate.py (fixed evaluation harness)
- Test cases in tests/ (ground truth)
- Model being evaluated (specified in evaluate.py)
- Scoring criteria (defined in evaluate.py)
## Strategy
1. First run: baseline with current prompt (or empty)
2. Add clear role/persona definition
3. Add output format specification
4. Add chain-of-thought instruction
5. Add 2-3 diverse examples
6. Refine based on failure modes from run.log
## Evaluation
- evaluate.py runs the prompt against 20 test cases
- Each test case is scored 1-10 by your CLI tool (Claude, Codex, or Gemini)
- quality_score = average * 10 (maps to 10-100)
- Run log shows which test cases failed
## Stop When
eval_score >= 85 OR after 50 experiments.
```
---
## Code Performance
```markdown
# autoresearch — Performance Optimization
## Goal
Minimize p50_ms (median latency). Lower is better.
## What You Can Change (src/module.py only)
- Algorithm implementation
- Data structures (use faster alternatives)
- Caching and memoization
- Vectorization (NumPy, etc.)
- Loop optimization
- I/O patterns
- Memory allocation patterns
## What You Cannot Change
- benchmark.py (fixed benchmark harness)
- Public API (function signatures must stay the same)
- External dependencies (add nothing new)
- Correctness tests (tests/ must still pass)
## Constraints
- Correctness is non-negotiable. benchmark.py runs tests first.
- If tests fail → immediate crash status, no metric recorded.
- Memory usage: p99 < 2x baseline acceptable, hard limit at 4x.
## Strategy
1. Baseline: profile first, don't guess
2. Check if there's any O(n²) → O(n log n) opportunity
3. Try caching repeated computations
4. Try NumPy vectorization for loops
5. Try algorithm-level changes last (higher risk)
## Stop When
p50_ms < 50ms OR improvement plateaus for 10 consecutive experiments.
```
---
## Agent Skill Optimization
```markdown
# autoresearch — Skill Optimization
## Goal
Maximize pass_rate on the task evaluation suite. Higher is better (0-1).
## What You Can Change (SKILL.md only)
- Skill description and trigger phrases
- Core workflow steps and ordering
- Decision frameworks and rules
- Output format specifications
- Example inputs/outputs
- Related skills disambiguation
- Proactive trigger conditions
## What You Cannot Change
- your custom evaluate.py (see Custom Evaluators in SKILL.md)
- Test tasks in tests/ (ground truth benchmark)
- Skill name (used for routing)
- License or metadata
## Evaluation
- evaluate.py runs SKILL.md against 15 standardized tasks
- Your CLI tool scores each task: 0 (fail), 0.5 (partial), 1 (pass)
- pass_rate = sum(scores) / 15
## Strategy
1. Baseline: run as-is
2. Improve trigger description (better routing = more passes)
3. Sharpen the core workflow (clearer = better execution)
4. Add missing edge cases to the rules section
5. Improve disambiguation (reduce false-positive routing)
## Simplicity Rule
A shorter SKILL.md that achieves the same score is better.
Aim for 200-400 lines total.
## Stop When
pass_rate >= 0.90 OR after 30 experiments.
```
FILE:scripts/log_results.py
#!/usr/bin/env python3
"""
autoresearch-agent: Results Viewer
View experiment results in multiple formats: terminal, CSV, Markdown.
Supports single experiment, domain, or cross-experiment dashboard.
Usage:
python scripts/log_results.py --experiment engineering/api-speed
python scripts/log_results.py --domain engineering
python scripts/log_results.py --dashboard
python scripts/log_results.py --experiment engineering/api-speed --format csv --output results.csv
python scripts/log_results.py --experiment engineering/api-speed --format markdown --output results.md
python scripts/log_results.py --dashboard --format markdown --output dashboard.md
"""
import argparse
import csv
import io
import sys
import time
from pathlib import Path
def find_autoresearch_root():
"""Find .autoresearch/ in project or user home."""
project_root = Path(".").resolve() / ".autoresearch"
if project_root.exists():
return project_root
user_root = Path.home() / ".autoresearch"
if user_root.exists():
return user_root
return None
def load_config(experiment_dir):
"""Load config.cfg."""
cfg_file = experiment_dir / "config.cfg"
config = {}
if cfg_file.exists():
for line in cfg_file.read_text().splitlines():
if ":" in line:
k, v = line.split(":", 1)
config[k.strip()] = v.strip()
return config
def load_results(experiment_dir):
"""Load results.tsv into list of dicts."""
tsv = experiment_dir / "results.tsv"
if not tsv.exists():
return []
results = []
for line in tsv.read_text().splitlines()[1:]:
parts = line.split("\t")
if len(parts) >= 4:
try:
metric = float(parts[1]) if parts[1] != "N/A" else None
except ValueError:
metric = None
results.append({
"commit": parts[0],
"metric": metric,
"status": parts[2],
"description": parts[3],
})
return results
def compute_stats(results, direction):
"""Compute statistics from results."""
keeps = [r for r in results if r["status"] == "keep"]
discards = [r for r in results if r["status"] == "discard"]
crashes = [r for r in results if r["status"] == "crash"]
valid_keeps = [r for r in keeps if r["metric"] is not None]
baseline = valid_keeps[0]["metric"] if valid_keeps else None
if valid_keeps:
best = min(r["metric"] for r in valid_keeps) if direction == "lower" else max(r["metric"] for r in valid_keeps)
else:
best = None
pct_change = None
if baseline is not None and best is not None and baseline != 0:
if direction == "lower":
pct_change = (baseline - best) / baseline * 100
else:
pct_change = (best - baseline) / baseline * 100
return {
"total": len(results),
"keeps": len(keeps),
"discards": len(discards),
"crashes": len(crashes),
"baseline": baseline,
"best": best,
"pct_change": pct_change,
}
# --- Terminal Output ---
def print_experiment(experiment_dir, experiment_path):
"""Print single experiment results to terminal."""
config = load_config(experiment_dir)
results = load_results(experiment_dir)
direction = config.get("metric_direction", "lower")
metric_name = config.get("metric", "metric")
if not results:
print(f"No results for {experiment_path}")
return
stats = compute_stats(results, direction)
print(f"\n{'─' * 65}")
print(f" {experiment_path}")
print(f" Target: {config.get('target', '?')} | Metric: {metric_name} ({direction})")
print(f"{'─' * 65}")
print(f" Total: {stats['total']} | Keep: {stats['keeps']} | Discard: {stats['discards']} | Crash: {stats['crashes']}")
if stats["baseline"] is not None and stats["best"] is not None:
pct = f" ({stats['pct_change']:+.1f}%)" if stats["pct_change"] is not None else ""
print(f" Baseline: {stats['baseline']:.6f} -> Best: {stats['best']:.6f}{pct}")
print(f"\n {'COMMIT':<10} {'METRIC':>12} {'STATUS':<10} DESCRIPTION")
print(f" {'─' * 60}")
for r in results:
m = f"{r['metric']:.6f}" if r["metric"] is not None else "N/A "
icon = {"keep": "+", "discard": "-", "crash": "!"}.get(r["status"], "?")
print(f" {r['commit']:<10} {m:>12} {icon} {r['status']:<7} {r['description'][:35]}")
print()
def print_dashboard(root):
"""Print cross-experiment dashboard."""
experiments = []
for domain_dir in sorted(root.iterdir()):
if not domain_dir.is_dir() or domain_dir.name.startswith("."):
continue
for exp_dir in sorted(domain_dir.iterdir()):
if not exp_dir.is_dir() or not (exp_dir / "config.cfg").exists():
continue
config = load_config(exp_dir)
results = load_results(exp_dir)
direction = config.get("metric_direction", "lower")
stats = compute_stats(results, direction)
best_str = f"{stats['best']:.4f}" if stats["best"] is not None else "—"
pct_str = f"{stats['pct_change']:+.1f}%" if stats["pct_change"] is not None else "—"
# Determine status
status = "idle"
if stats["total"] > 0:
tsv = exp_dir / "results.tsv"
if tsv.exists():
age_hours = (time.time() - tsv.stat().st_mtime) / 3600
status = "active" if age_hours < 1 else "paused" if age_hours < 24 else "done"
experiments.append({
"domain": domain_dir.name,
"name": exp_dir.name,
"runs": stats["total"],
"kept": stats["keeps"],
"best": best_str,
"change": pct_str,
"status": status,
"metric": config.get("metric", "?"),
})
if not experiments:
print("No experiments found.")
return experiments
print(f"\n{'─' * 90}")
print(f" autoresearch — Dashboard")
print(f"{'─' * 90}")
print(f" {'DOMAIN':<15} {'EXPERIMENT':<20} {'RUNS':>5} {'KEPT':>5} {'BEST':>12} {'CHANGE':>10} {'STATUS':<8}")
print(f" {'─' * 85}")
for e in experiments:
print(f" {e['domain']:<15} {e['name']:<20} {e['runs']:>5} {e['kept']:>5} {e['best']:>12} {e['change']:>10} {e['status']:<8}")
print()
return experiments
# --- CSV Export ---
def export_experiment_csv(experiment_dir, experiment_path):
"""Export single experiment as CSV string."""
config = load_config(experiment_dir)
results = load_results(experiment_dir)
direction = config.get("metric_direction", "lower")
stats = compute_stats(results, direction)
buf = io.StringIO()
writer = csv.writer(buf)
# Header with metadata
writer.writerow(["# Experiment", experiment_path])
writer.writerow(["# Target", config.get("target", "")])
writer.writerow(["# Metric", f"{config.get('metric', '')} ({direction} is better)"])
if stats["baseline"] is not None:
writer.writerow(["# Baseline", f"{stats['baseline']:.6f}"])
if stats["best"] is not None:
pct = f" ({stats['pct_change']:+.1f}%)" if stats["pct_change"] is not None else ""
writer.writerow(["# Best", f"{stats['best']:.6f}{pct}"])
writer.writerow(["# Total", stats["total"]])
writer.writerow(["# Keep/Discard/Crash", f"{stats['keeps']}/{stats['discards']}/{stats['crashes']}"])
writer.writerow([])
writer.writerow(["Commit", "Metric", "Status", "Description"])
for r in results:
m = f"{r['metric']:.6f}" if r["metric"] is not None else "N/A"
writer.writerow([r["commit"], m, r["status"], r["description"]])
return buf.getvalue()
def export_dashboard_csv(root, domain_filter=None):
"""Export dashboard as CSV string."""
experiments = []
for domain_dir in sorted(root.iterdir()):
if not domain_dir.is_dir() or domain_dir.name.startswith("."):
continue
if domain_filter and domain_dir.name != domain_filter:
continue
for exp_dir in sorted(domain_dir.iterdir()):
if not exp_dir.is_dir() or not (exp_dir / "config.cfg").exists():
continue
config = load_config(exp_dir)
results = load_results(exp_dir)
direction = config.get("metric_direction", "lower")
stats = compute_stats(results, direction)
best_str = f"{stats['best']:.6f}" if stats["best"] is not None else ""
pct_str = f"{stats['pct_change']:+.1f}%" if stats["pct_change"] is not None else ""
experiments.append([
domain_dir.name, exp_dir.name, config.get("metric", ""),
stats["total"], stats["keeps"], stats["discards"], stats["crashes"],
best_str, pct_str
])
buf = io.StringIO()
writer = csv.writer(buf)
writer.writerow(["Domain", "Experiment", "Metric", "Runs", "Kept", "Discarded", "Crashed", "Best", "Change"])
for e in experiments:
writer.writerow(e)
return buf.getvalue()
# --- Markdown Export ---
def export_experiment_markdown(experiment_dir, experiment_path):
"""Export single experiment as Markdown string."""
config = load_config(experiment_dir)
results = load_results(experiment_dir)
direction = config.get("metric_direction", "lower")
metric_name = config.get("metric", "metric")
stats = compute_stats(results, direction)
lines = []
lines.append(f"# Autoresearch: {experiment_path}\n")
lines.append(f"**Target:** `{config.get('target', '?')}` ")
lines.append(f"**Metric:** `{metric_name}` ({direction} is better) ")
lines.append(f"**Experiments:** {stats['total']} total — {stats['keeps']} kept, {stats['discards']} discarded, {stats['crashes']} crashed\n")
if stats["baseline"] is not None and stats["best"] is not None:
pct = f" ({stats['pct_change']:+.1f}%)" if stats["pct_change"] is not None else ""
lines.append(f"**Progress:** `{stats['baseline']:.6f}` → `{stats['best']:.6f}`{pct}\n")
lines.append(f"| Commit | Metric | Status | Description |")
lines.append(f"|--------|--------|--------|-------------|")
for r in results:
m = f"`{r['metric']:.6f}`" if r["metric"] is not None else "N/A"
lines.append(f"| `{r['commit']}` | {m} | {r['status']} | {r['description']} |")
lines.append("")
return "\n".join(lines)
def export_dashboard_markdown(root, domain_filter=None):
"""Export dashboard as Markdown string."""
lines = []
lines.append("# Autoresearch Dashboard\n")
lines.append("| Domain | Experiment | Metric | Runs | Kept | Best | Change | Status |")
lines.append("|--------|-----------|--------|------|------|------|--------|--------|")
for domain_dir in sorted(root.iterdir()):
if not domain_dir.is_dir() or domain_dir.name.startswith("."):
continue
if domain_filter and domain_dir.name != domain_filter:
continue
for exp_dir in sorted(domain_dir.iterdir()):
if not exp_dir.is_dir() or not (exp_dir / "config.cfg").exists():
continue
config = load_config(exp_dir)
results = load_results(exp_dir)
direction = config.get("metric_direction", "lower")
stats = compute_stats(results, direction)
best = f"`{stats['best']:.4f}`" if stats["best"] is not None else "—"
pct = f"{stats['pct_change']:+.1f}%" if stats["pct_change"] is not None else "—"
tsv = exp_dir / "results.tsv"
status = "idle"
if tsv.exists() and stats["total"] > 0:
age_h = (time.time() - tsv.stat().st_mtime) / 3600
status = "active" if age_h < 1 else "paused" if age_h < 24 else "done"
lines.append(f"| {domain_dir.name} | {exp_dir.name} | {config.get('metric', '?')} | {stats['total']} | {stats['keeps']} | {best} | {pct} | {status} |")
lines.append("")
return "\n".join(lines)
# --- Main ---
def main():
parser = argparse.ArgumentParser(description="autoresearch-agent results viewer")
parser.add_argument("--experiment", help="Show one experiment: domain/name")
parser.add_argument("--domain", help="Show all experiments in a domain")
parser.add_argument("--dashboard", action="store_true", help="Cross-experiment dashboard")
parser.add_argument("--format", choices=["terminal", "csv", "markdown"], default="terminal",
help="Output format (default: terminal)")
parser.add_argument("--output", "-o", help="Write to file instead of stdout")
parser.add_argument("--all", action="store_true", help="Show all experiments (alias for --dashboard)")
args = parser.parse_args()
root = find_autoresearch_root()
if root is None:
print("No .autoresearch/ found. Run setup_experiment.py first.")
sys.exit(1)
output_text = None
# Single experiment
if args.experiment:
experiment_dir = root / args.experiment
if not experiment_dir.exists():
print(f"Experiment not found: {args.experiment}")
sys.exit(1)
if args.format == "csv":
output_text = export_experiment_csv(experiment_dir, args.experiment)
elif args.format == "markdown":
output_text = export_experiment_markdown(experiment_dir, args.experiment)
else:
print_experiment(experiment_dir, args.experiment)
return
# Domain
elif args.domain:
domain_dir = root / args.domain
if not domain_dir.exists():
print(f"Domain not found: {args.domain}")
sys.exit(1)
for exp_dir in sorted(domain_dir.iterdir()):
if exp_dir.is_dir() and (exp_dir / "config.cfg").exists():
if args.format == "terminal":
print_experiment(exp_dir, f"{args.domain}/{exp_dir.name}")
# For CSV/MD, fall through to dashboard with domain filter
if args.format != "terminal":
# Use dashboard export filtered to domain
output_text = export_dashboard_csv(root, domain_filter=args.domain) if args.format == "csv" else export_dashboard_markdown(root, domain_filter=args.domain)
else:
return
# Dashboard
elif args.dashboard or args.all:
if args.format == "csv":
output_text = export_dashboard_csv(root)
elif args.format == "markdown":
output_text = export_dashboard_markdown(root)
else:
print_dashboard(root)
return
else:
# Default: dashboard
if args.format == "terminal":
print_dashboard(root)
return
output_text = export_dashboard_csv(root) if args.format == "csv" else export_dashboard_markdown(root)
# Write output
if output_text:
if args.output:
Path(args.output).write_text(output_text)
print(f"Written to {args.output}")
else:
print(output_text)
if __name__ == "__main__":
main()
FILE:scripts/run_experiment.py
#!/usr/bin/env python3
"""
autoresearch-agent: Experiment Runner
Executes a single experiment iteration. The AI agent is the loop —
it calls this script repeatedly. The script handles evaluation,
metric parsing, keep/discard decisions, and git rollback on failure.
Usage:
python scripts/run_experiment.py --experiment engineering/api-speed --single
python scripts/run_experiment.py --experiment engineering/api-speed --dry-run
python scripts/run_experiment.py --experiment engineering/api-speed --single --description "added caching"
"""
import argparse
import subprocess
import sys
import time
from datetime import datetime
from pathlib import Path
def find_autoresearch_root():
"""Find .autoresearch/ in project or user home."""
project_root = Path(".").resolve() / ".autoresearch"
if project_root.exists():
return project_root
user_root = Path.home() / ".autoresearch"
if user_root.exists():
return user_root
return None
def load_config(experiment_dir):
"""Load config.cfg from experiment directory."""
cfg_file = experiment_dir / "config.cfg"
if not cfg_file.exists():
print(f" Error: no config.cfg in {experiment_dir}")
sys.exit(1)
config = {}
for line in cfg_file.read_text().splitlines():
if ":" in line:
k, v = line.split(":", 1)
config[k.strip()] = v.strip()
return config
def run_git(args, cwd=None, timeout=30):
"""Run a git command safely (no shell injection). Returns (returncode, stdout, stderr)."""
result = subprocess.run(
["git"] + args,
capture_output=True, text=True,
cwd=cwd, timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
def get_current_commit(path):
"""Get short hash of current HEAD."""
_, commit, _ = run_git(["rev-parse", "--short", "HEAD"], cwd=path)
return commit
def get_best_metric(experiment_dir, direction):
"""Read the best metric from results.tsv."""
tsv = experiment_dir / "results.tsv"
if not tsv.exists():
return None
lines = [l for l in tsv.read_text().splitlines()[1:] if "\tkeep\t" in l]
if not lines:
return None
metrics = []
for line in lines:
parts = line.split("\t")
try:
if parts[1] != "N/A":
metrics.append(float(parts[1]))
except (ValueError, IndexError):
continue
if not metrics:
return None
return min(metrics) if direction == "lower" else max(metrics)
def run_evaluation(project_root, eval_cmd, time_budget_minutes, log_file):
"""Run evaluation with time limit. Output goes to log_file.
Note: shell=True is intentional here — eval_cmd is user-provided and
may contain pipes, redirects, or chained commands.
"""
hard_limit = time_budget_minutes * 60 * 2.5
t0 = time.time()
try:
with open(log_file, "w") as lf:
result = subprocess.run(
eval_cmd, shell=True,
stdout=lf, stderr=subprocess.STDOUT,
cwd=str(project_root),
timeout=hard_limit
)
elapsed = time.time() - t0
return result.returncode, elapsed
except subprocess.TimeoutExpired:
elapsed = time.time() - t0
return -1, elapsed
def extract_metric(log_file, metric_grep):
"""Extract metric value from log file."""
log_path = Path(log_file)
if not log_path.exists():
return None
for line in reversed(log_path.read_text().splitlines()):
stripped = line.strip()
if stripped.startswith(metric_grep.lstrip("^")):
try:
return float(stripped.split(":")[-1].strip())
except ValueError:
continue
return None
def is_improvement(new_val, old_val, direction):
"""Check if new result is better than old."""
if old_val is None:
return True
if direction == "lower":
return new_val < old_val
return new_val > old_val
def log_result(experiment_dir, commit, metric_val, status, description):
"""Append result to results.tsv."""
tsv = experiment_dir / "results.tsv"
metric_str = f"{metric_val:.6f}" if metric_val is not None else "N/A"
with open(tsv, "a") as f:
f.write(f"{commit}\t{metric_str}\t{status}\t{description}\n")
def get_experiment_count(experiment_dir):
"""Count experiments run so far."""
tsv = experiment_dir / "results.tsv"
if not tsv.exists():
return 0
return max(0, len(tsv.read_text().splitlines()) - 1)
def get_description_from_diff(project_root):
"""Auto-generate a description from git diff --stat HEAD~1."""
code, diff_stat, _ = run_git(["diff", "--stat", "HEAD~1"], cwd=str(project_root))
if code == 0 and diff_stat:
return diff_stat.split("\n")[0][:50]
return "experiment"
def read_last_lines(filepath, n=5):
"""Read last n lines of a file (replaces tail shell command)."""
path = Path(filepath)
if not path.exists():
return ""
lines = path.read_text().splitlines()
return "\n".join(lines[-n:])
def run_single(project_root, experiment_dir, config, exp_num, dry_run=False, description=None):
"""Run one experiment iteration."""
direction = config.get("metric_direction", "lower")
metric_grep = config.get("metric_grep", "^metric:")
eval_cmd = config.get("evaluate_cmd", "python evaluate.py")
time_budget = int(config.get("time_budget_minutes", 5))
metric_name = config.get("metric", "metric")
log_file = str(experiment_dir / "run.log")
best = get_best_metric(experiment_dir, direction)
ts = datetime.now().strftime("%H:%M:%S")
print(f"\n[{ts}] Experiment #{exp_num}")
print(f" Best {metric_name}: {best}")
if dry_run:
print(" [DRY RUN] Would run evaluation and check metric")
return "dry_run"
# Auto-generate description if not provided
if not description:
description = get_description_from_diff(str(project_root))
# Run evaluation
print(f" Running: {eval_cmd} (budget: {time_budget}m)")
ret_code, elapsed = run_evaluation(project_root, eval_cmd, time_budget, log_file)
commit = get_current_commit(str(project_root))
# Timeout
if ret_code == -1:
print(f" TIMEOUT after {elapsed:.0f}s — discarding")
run_git(["checkout", "--", "."], cwd=str(project_root))
run_git(["reset", "--hard", "HEAD~1"], cwd=str(project_root))
log_result(experiment_dir, commit, None, "crash", f"timeout_{elapsed:.0f}s")
return "crash"
# Crash
if ret_code != 0:
tail = read_last_lines(log_file, 5)
print(f" CRASH (exit {ret_code}) after {elapsed:.0f}s")
print(f" Last output: {tail[:200]}")
run_git(["reset", "--hard", "HEAD~1"], cwd=str(project_root))
log_result(experiment_dir, commit, None, "crash", f"exit_{ret_code}")
return "crash"
# Extract metric
metric_val = extract_metric(log_file, metric_grep)
if metric_val is None:
print(f" Could not parse {metric_name} from run.log")
run_git(["reset", "--hard", "HEAD~1"], cwd=str(project_root))
log_result(experiment_dir, commit, None, "crash", "metric_parse_failed")
return "crash"
delta = ""
if best is not None:
diff = metric_val - best
delta = f" (delta {diff:+.4f})"
print(f" {metric_name}: {metric_val:.6f}{delta} in {elapsed:.0f}s")
# Keep or discard
if is_improvement(metric_val, best, direction):
print(f" KEEP — improvement")
log_result(experiment_dir, commit, metric_val, "keep", description)
return "keep"
else:
print(f" DISCARD — no improvement")
run_git(["reset", "--hard", "HEAD~1"], cwd=str(project_root))
best_str = f"{best:.4f}" if best is not None else "?"
log_result(experiment_dir, commit, metric_val, "discard",
f"no_improvement_{metric_val:.4f}_vs_{best_str}")
return "discard"
def main():
parser = argparse.ArgumentParser(description="autoresearch-agent runner")
parser.add_argument("--experiment", help="Experiment path: domain/name (e.g. engineering/api-speed)")
parser.add_argument("--single", action="store_true", help="Run one experiment iteration")
parser.add_argument("--dry-run", action="store_true", help="Show what would happen")
parser.add_argument("--description", help="Description of the change (auto-generated from git diff if omitted)")
parser.add_argument("--path", default=".", help="Project root")
args = parser.parse_args()
project_root = Path(args.path).resolve()
root = find_autoresearch_root()
if root is None:
print("No .autoresearch/ found. Run setup_experiment.py first.")
sys.exit(1)
if not args.experiment:
print("Specify --experiment domain/name")
sys.exit(1)
experiment_dir = root / args.experiment
if not experiment_dir.exists():
print(f"Experiment not found: {experiment_dir}")
print("Run: python scripts/setup_experiment.py --list")
sys.exit(1)
config = load_config(experiment_dir)
print(f"\n autoresearch-agent")
print(f" Experiment: {args.experiment}")
print(f" Target: {config.get('target', '?')}")
print(f" Metric: {config.get('metric', '?')} ({config.get('metric_direction', '?')} is better)")
print(f" Budget: {config.get('time_budget_minutes', '?')} min/experiment")
print(f" Mode: {'dry-run' if args.dry_run else 'single'}")
exp_num = get_experiment_count(experiment_dir) + 1
run_single(project_root, experiment_dir, config, exp_num, args.dry_run, args.description)
if __name__ == "__main__":
main()
FILE:scripts/setup_experiment.py
#!/usr/bin/env python3
"""
autoresearch-agent: Setup Experiment
Initialize a new experiment with domain, target, evaluator, and git branch.
Creates the .autoresearch/{domain}/{name}/ directory structure.
Usage:
python scripts/setup_experiment.py --domain engineering --name api-speed \
--target src/api/search.py --eval "pytest bench.py" \
--metric p50_ms --direction lower
python scripts/setup_experiment.py --domain marketing --name medium-ctr \
--target content/titles.md --eval "python evaluate.py" \
--metric ctr_score --direction higher --evaluator llm_judge_content
python scripts/setup_experiment.py --list # List all experiments
python scripts/setup_experiment.py --list-evaluators # List available evaluators
"""
import argparse
import shutil
import subprocess
import sys
from datetime import datetime
from pathlib import Path
DOMAINS = ["engineering", "marketing", "content", "prompts", "custom"]
EVALUATOR_DIR = Path(__file__).parent.parent / "evaluators"
DEFAULT_CONFIG = """# autoresearch global config
default_time_budget_minutes: 5
default_scope: project
dashboard_format: markdown
"""
GITIGNORE_CONTENT = """# autoresearch — experiment logs are local state
**/results.tsv
**/run.log
**/run.*.log
config.yaml
"""
def run_cmd(cmd, cwd=None, timeout=None):
"""Run shell command, return (returncode, stdout, stderr)."""
result = subprocess.run(
cmd, shell=True, capture_output=True, text=True,
cwd=cwd, timeout=timeout
)
return result.returncode, result.stdout.strip(), result.stderr.strip()
def get_autoresearch_root(scope, project_root=None):
"""Get the .autoresearch root directory based on scope."""
if scope == "user":
return Path.home() / ".autoresearch"
return Path(project_root or ".") / ".autoresearch"
def init_root(root):
"""Initialize .autoresearch root if it doesn't exist."""
created = False
if not root.exists():
root.mkdir(parents=True)
created = True
print(f" Created {root}/")
config_file = root / "config.yaml"
if not config_file.exists():
config_file.write_text(DEFAULT_CONFIG)
print(f" Created {config_file}")
gitignore = root / ".gitignore"
if not gitignore.exists():
gitignore.write_text(GITIGNORE_CONTENT)
print(f" Created {gitignore}")
return created
def create_program_md(experiment_dir, domain, name, target, metric, direction, constraints=""):
"""Generate a program.md template for the experiment."""
direction_word = "Minimize" if direction == "lower" else "Maximize"
content = f"""# autoresearch — {name}
## Goal
{direction_word} `{metric}` on `{target}`. {"Lower" if direction == "lower" else "Higher"} is better.
## What the Agent Can Change
- Only `{target}` — this is the single file being optimized.
- Everything inside that file is fair game unless constrained below.
## What the Agent Cannot Change
- The evaluation script (`evaluate.py` or the eval command). It is read-only.
- Dependencies — do not add new packages or imports that aren't already available.
- Any other files in the project unless explicitly noted here.
{f"- Additional constraints: {constraints}" if constraints else ""}
## Strategy
1. First run: establish baseline. Do not change anything.
2. Profile/analyze the current state — understand why the metric is what it is.
3. Try the most obvious improvement first (low-hanging fruit).
4. If that works, push further in the same direction.
5. If stuck, try something orthogonal or radical.
6. Read the git log of previous experiments. Don't repeat failed approaches.
## Simplicity Rule
A small improvement that adds ugly complexity is NOT worth it.
Equal performance with simpler code IS worth it.
Removing code that gets same results is the best outcome.
## Stop When
You don't stop. The human will interrupt you when they're satisfied.
If no improvement in 20+ consecutive runs, change strategy drastically.
"""
(experiment_dir / "program.md").write_text(content)
def create_config(experiment_dir, target, eval_cmd, metric, direction, time_budget):
"""Write experiment config."""
content = f"""target: {target}
evaluate_cmd: {eval_cmd}
metric: {metric}
metric_direction: {direction}
metric_grep: ^{metric}:
time_budget_minutes: {time_budget}
created: {datetime.now().strftime('%Y-%m-%d %H:%M')}
"""
(experiment_dir / "config.cfg").write_text(content)
def init_results_tsv(experiment_dir):
"""Create results.tsv with header."""
tsv = experiment_dir / "results.tsv"
if tsv.exists():
print(f" results.tsv already exists ({tsv.stat().st_size} bytes)")
return
tsv.write_text("commit\tmetric\tstatus\tdescription\n")
print(" Created results.tsv")
def copy_evaluator(experiment_dir, evaluator_name):
"""Copy a built-in evaluator to the experiment directory."""
source = EVALUATOR_DIR / f"{evaluator_name}.py"
if not source.exists():
print(f" Warning: evaluator '{evaluator_name}' not found in {EVALUATOR_DIR}")
print(f" Available: {', '.join(f.stem for f in EVALUATOR_DIR.glob('*.py'))}")
return False
dest = experiment_dir / "evaluate.py"
shutil.copy2(source, dest)
print(f" Copied evaluator: {evaluator_name}.py -> evaluate.py")
return True
def create_branch(path, domain, name):
"""Create and checkout the experiment branch."""
branch = f"autoresearch/{domain}/{name}"
result = subprocess.run(
["git", "checkout", "-b", branch],
cwd=path, capture_output=True, text=True
)
if result.returncode != 0:
if "already exists" in result.stderr:
print(f" Branch '{branch}' already exists. Checking out...")
subprocess.run(
["git", "checkout", branch],
cwd=path, capture_output=True, text=True
)
return branch
print(f" Warning: could not create branch: {result.stderr}")
return None
print(f" Created branch: {branch}")
return branch
def list_experiments(root):
"""List all experiments across all domains."""
if not root.exists():
print("No experiments found. Run setup to create your first experiment.")
return
experiments = []
for domain_dir in sorted(root.iterdir()):
if not domain_dir.is_dir() or domain_dir.name.startswith("."):
continue
for exp_dir in sorted(domain_dir.iterdir()):
if not exp_dir.is_dir():
continue
cfg_file = exp_dir / "config.cfg"
if not cfg_file.exists():
continue
config = {}
for line in cfg_file.read_text().splitlines():
if ":" in line:
k, v = line.split(":", 1)
config[k.strip()] = v.strip()
# Count results
tsv = exp_dir / "results.tsv"
runs = 0
if tsv.exists():
runs = max(0, len(tsv.read_text().splitlines()) - 1)
experiments.append({
"domain": domain_dir.name,
"name": exp_dir.name,
"target": config.get("target", "?"),
"metric": config.get("metric", "?"),
"runs": runs,
})
if not experiments:
print("No experiments found.")
return
print(f"\n{'DOMAIN':<15} {'EXPERIMENT':<25} {'TARGET':<30} {'METRIC':<15} {'RUNS':>5}")
print("-" * 95)
for e in experiments:
print(f"{e['domain']:<15} {e['name']:<25} {e['target']:<30} {e['metric']:<15} {e['runs']:>5}")
print(f"\nTotal: {len(experiments)} experiments")
def list_evaluators():
"""List available built-in evaluators."""
if not EVALUATOR_DIR.exists():
print("No evaluators directory found.")
return
print(f"\nAvailable evaluators ({EVALUATOR_DIR}):\n")
for f in sorted(EVALUATOR_DIR.glob("*.py")):
# Read first docstring line
desc = ""
for line in f.read_text().splitlines():
stripped = line.strip()
if stripped.startswith('"""') or stripped.startswith("'''"):
quote = stripped[:3]
# Single-line docstring: """Description."""
after_quote = stripped[3:]
if after_quote and after_quote.rstrip(quote[0]).strip():
desc = after_quote.rstrip('"').rstrip("'").strip()
break
continue
if stripped and not line.startswith("#!"):
desc = stripped.strip('"').strip("'")
break
print(f" {f.stem:<25} {desc}")
def main():
parser = argparse.ArgumentParser(description="autoresearch-agent setup")
parser.add_argument("--domain", choices=DOMAINS, help="Experiment domain")
parser.add_argument("--name", help="Experiment name (e.g. api-speed, medium-ctr)")
parser.add_argument("--target", help="Target file to optimize")
parser.add_argument("--eval", dest="eval_cmd", help="Evaluation command")
parser.add_argument("--metric", help="Metric name (must appear in eval output as 'name: value')")
parser.add_argument("--direction", choices=["lower", "higher"], default="lower",
help="Is lower or higher better?")
parser.add_argument("--time-budget", type=int, default=5, help="Minutes per experiment (default: 5)")
parser.add_argument("--evaluator", help="Built-in evaluator to copy (e.g. benchmark_speed)")
parser.add_argument("--scope", choices=["project", "user"], default="project",
help="Where to store experiments: project (./) or user (~/)")
parser.add_argument("--constraints", default="", help="Additional constraints for program.md")
parser.add_argument("--path", default=".", help="Project root path")
parser.add_argument("--skip-branch", action="store_true", help="Don't create git branch")
parser.add_argument("--list", action="store_true", help="List all experiments")
parser.add_argument("--list-evaluators", action="store_true", help="List available evaluators")
args = parser.parse_args()
project_root = Path(args.path).resolve()
# List mode
if args.list:
root = get_autoresearch_root("project", project_root)
list_experiments(root)
user_root = get_autoresearch_root("user")
if user_root.exists() and user_root != root:
print(f"\n--- User-level experiments ({user_root}) ---")
list_experiments(user_root)
return
if args.list_evaluators:
list_evaluators()
return
# Validate required args for setup
if not all([args.domain, args.name, args.target, args.eval_cmd, args.metric]):
parser.error("Required: --domain, --name, --target, --eval, --metric")
root = get_autoresearch_root(args.scope, project_root)
print(f"\n autoresearch-agent setup")
print(f" Project: {project_root}")
print(f" Scope: {args.scope}")
print(f" Domain: {args.domain}")
print(f" Experiment: {args.name}")
print(f" Time: {datetime.now().strftime('%Y-%m-%d %H:%M')}\n")
# Check git
result = subprocess.run(
["git", "rev-parse", "--is-inside-work-tree"],
cwd=str(project_root), capture_output=True, text=True
)
code = result.returncode
if code != 0:
print(" Error: not a git repository. Run: git init && git add . && git commit -m 'initial'")
sys.exit(1)
print(" Git repository found")
# Check target file
target_path = project_root / args.target
if not target_path.exists():
print(f" Error: target file not found: {args.target}")
sys.exit(1)
print(f" Target file found: {args.target}")
# Init root
init_root(root)
# Create experiment directory
experiment_dir = root / args.domain / args.name
if experiment_dir.exists():
print(f" Warning: experiment '{args.domain}/{args.name}' already exists.")
print(f" Use --name with a different name, or delete {experiment_dir}")
sys.exit(1)
experiment_dir.mkdir(parents=True)
print(f" Created {experiment_dir}/")
# Create files
create_program_md(experiment_dir, args.domain, args.name,
args.target, args.metric, args.direction, args.constraints)
print(" Created program.md")
create_config(experiment_dir, args.target, args.eval_cmd,
args.metric, args.direction, args.time_budget)
print(" Created config.cfg")
init_results_tsv(experiment_dir)
# Copy evaluator if specified
if args.evaluator:
copy_evaluator(experiment_dir, args.evaluator)
# Create git branch
if not args.skip_branch:
create_branch(str(project_root), args.domain, args.name)
# Test evaluation command
print(f"\n Testing evaluation: {args.eval_cmd}")
code, out, err = run_cmd(args.eval_cmd, cwd=str(project_root), timeout=60)
if code != 0:
print(f" Warning: eval command failed (exit {code})")
if err:
print(f" stderr: {err[:200]}")
print(" Fix the eval command before running the experiment loop.")
else:
# Check metric is parseable
full_output = out + "\n" + err
metric_found = False
for line in full_output.splitlines():
if line.strip().startswith(f"{args.metric}:"):
metric_found = True
print(f" Eval works. Baseline: {line.strip()}")
break
if not metric_found:
print(f" Warning: eval ran but '{args.metric}:' not found in output.")
print(f" Make sure your eval command outputs: {args.metric}: <value>")
# Summary
print(f"\n Setup complete!")
print(f" Experiment: {args.domain}/{args.name}")
print(f" Target: {args.target}")
print(f" Metric: {args.metric} ({args.direction} is better)")
print(f" Budget: {args.time_budget} min/experiment")
if not args.skip_branch:
print(f" Branch: autoresearch/{args.domain}/{args.name}")
print(f"\n To start:")
print(f" python scripts/run_experiment.py --experiment {args.domain}/{args.name} --single")
if __name__ == "__main__":
main()
Tính quy mô năng lực vận hành, kế hoạch nhân sự, rủi ro sử dụng nguồn lực bằng mô hình hàng đợi Erlang-C.
---
name: capacity-planner
description: "Use when an ops leader (Director of CX, Head of Support, VP Ops, Head of BizOps, Head of IT ops, Head of Finance ops) is sizing ops capacity, building a headcount plan, modeling utilization risk, planning Q3 capacity or annual support capacity, or designing CS coverage — and needs Erlang-C queueing math, P90 demand sizing, shrinkage-adjusted FTE, manager-trigger thresholds, and a quarterly hiring sequence with ramp + attrition. Apply when sustained team utilization is above 80% or when the team is growing >50% in 12 months. Run before committing the headcount budget. This is NOT engineering capacity (see vpe-advisor for DORA + cycle time) and NOT strategic 3-year workforce planning (see chro-advisor)."
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [bizops, capacity, headcount, utilization, queueing-theory, ops-planning, little-law, workforce]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# capacity-planner
Sizing tool for **ops teams that handle queued work** — Support, CX,
Customer Success, BizOps, IT ops, Finance ops. Built on Erlang-C
queueing theory, Little's Law, and the operational-leadership canon
(Fournier, Larson, Cleveland, Reinertsen). Deterministic, stdlib-only,
no LLM calls.
## Purpose
You are an ops leader sized 15 → 35 with no idea how the 35-person org
will actually behave at peak load. Or you are at 88% utilization and
SLA is starting to slip. Or you have a hiring budget approved and need
to sequence it across four quarters without burning out the existing
team. This skill answers those questions with arithmetic, not vibes.
It produces three artifacts:
1. **Capacity sizing** at 70/80/90% utilization against P50/P90/P99
demand, with P(SLA breach) at each point and a SAFE/WATCH/AT_RISK/CRITICAL
risk band.
2. **Utilization health** at the per-member traffic-light level plus a
team verdict (HEALTHY/SQUEEZED/OVERLOADED/UNBALANCED).
3. **12-month quarterly hiring plan** accounting for ramp curves,
attrition, QoQ demand growth, and span-of-control manager triggers.
## When to use
- **Annual ops capacity planning** (October-November for the following
fiscal year).
- **Quarterly re-sizing** if demand changed >15% or attrition spiked.
- **Pre-budget defense** — the math that justifies the headcount ask
to your CFO.
- **Diagnostic** when an ops team is missing SLA and you need to know
whether it's a sizing problem, a process problem, or a bottleneck
problem.
- **M&A / new-segment launch** modeling — sizing a new team or
combined org.
## Workflow
1. **Intake demand**. Pull P50/P90/P99 daily ticket/case volume from
your work system (Zendesk, Intercom, JSM, ServiceNow, Salesforce).
If you only have averages, stop and pull the distribution. Single-
point demand estimates are the most expensive anti-pattern in ops.
2. **Model throughput**. Run `capacity_modeler.py` with your demand,
AHT, SLA target, current FTE, and shrinkage. Use `--profile` for
your function (support / cx / bizops / finance-ops / it-ops). Read
the 80%-utilization row — that's your sizing point.
3. **Flag utilization risk**. Run `utilization_analyzer.py` against
your current team's actual utilization data. Anyone >85% sustained
is a throughput-collapse risk per Reinertsen. Spread >30 percentage
points across team means UNBALANCED — fix that before hiring.
4. **Sequence hiring**. Run `hiring_sequencer.py` with current FTE,
target EOY, ramp time, attrition, and growth. It will front-load
hires (Q1 35%, Q4 15%), apply ramp curves, and trigger a manager
hire when span of control crosses 7 ICs/manager.
5. **Walk the Forcing-question library** (see below). One question at
a time. Do not skip ahead. Answers must be written down before
you commit the plan.
## Scripts
- `scripts/capacity_modeler.py` — Erlang-C sizing with shrinkage
adjustment and P50/P90/P99 breach probabilities. `--profile`
for industry defaults.
- `scripts/utilization_analyzer.py` — per-member traffic-light +
team-level health verdict with variance detection.
- `scripts/hiring_sequencer.py` — 12-month quarterly plan with ramp,
attrition, growth, max-hires-per-quarter constraint, and
manager-trigger logic.
All three accept `--input <path>` (JSON), `--output {markdown,json}`,
`--sample` (built-in example), and `--help`. Stdlib only.
## References
- `references/queueing_theory_canon.md` — Erlang, Little, Hopp &
Spearman, Reinertsen, Kingman, Cleveland, ITIL, Armony et al. (8
sources). The math.
- `references/ops_workforce_planning_canon.md` — Fournier, Larson,
Google SRE Workbook, Frei, Lawler, Bersin, Gartner, Grove (8
sources). The people factors.
- `references/capacity_anti_patterns.md` — 11 named anti-patterns
with cited sources, tool guards, and the meta-discipline that
Lencioni + Goldratt + Christensen impose. (8+ named sources.)
## Assets
- `assets/capacity_brief_template.md` — 20-minute fill-out template
with JSON skeletons for all three tools and an output checklist.
## Assumptions
This skill assumes:
- Work is **queued** (tickets, cases, work items) — not project-style.
If your team's work isn't queued, this is the wrong skill.
- Demand has a **stationary-enough distribution** within a quarter.
Step-changes (new product launch, M&A, regulatory shift) require
re-running mid-quarter.
- You have **at least 90 days of historical demand data** to compute
P50/P90/P99. If not, generate the distribution from your sales /
user-base forecast first.
- Service is **single-class** within a queue. If you have hard
priority tiers (P1/P2/P3 with class-specific SLAs), model each as
a separate queue and sum.
- **Channels are modeled coherently.** Multi-channel teams use the
appropriate `--profile` with built-in shrinkage premium.
## Anti-patterns
See `references/capacity_anti_patterns.md` for the full taxonomy with
sources. Top eight:
1. Plan-to-100%-utilization (Reinertsen Principle 12)
2. Treat-ramp-as-instant (Larson)
3. Ignore-attrition-in-12-month-plan (Bersin)
4. Hire-ICs-forever-with-no-manager-trigger (Fournier)
5. Size-to-P50-demand-only (Cleveland)
6. No-shrinkage-adjustment (Cleveland, SRE Workbook)
7. Single-channel-model-for-multi-channel-work (Gartner, Kingman)
8. No-surge-plan-for-P99-events (Hopp & Spearman, Reinertsen)
## Distinct from
- **`c-level-advisor/vpe-advisor`** measures *engineering* throughput
via DORA 4 metrics, story points, deployment frequency, and cycle
time bottlenecks. It is for engineering teams shipping code. This
skill is for ops teams handling tickets/cases. Different unit of
work, different math (Erlang-C vs. DORA), different bottleneck
(queueing-blind staffing vs. WIP + lead time).
- **`c-level-advisor/chro-advisor`** does *strategic* workforce
planning (1-5 year capability portfolios, talent supply, leadership
succession). This skill does *operational* 0-12 month capacity
sizing against demand. Per Lawler: conflating them gets you hired
into the wrong jobs.
- **`project-management/*`** tracks delivery throughput on projects
(Jira velocity, sprint capacity). This skill sizes around steady-
state queued work.
- **Sibling `process-mapper`** *finds* the bottleneck. This skill
*sizes the team around* a known bottleneck. Order of operations:
process-mapper first → capacity-planner second. Hiring around the
wrong constraint wastes the hires.
- **`business-growth/cs-coverage`** (if it exists) sizes Customer
Success coverage by ARR/CSM ratio and segment. This skill sizes by
queued work volume (tickets, cases, escalations). For a CS team
that handles both relationship work AND a ticket queue, run both.
## Forcing-question library (Matt Pocock grill discipline)
**Discipline**: walk these one at a time. Do not skip ahead. Answers must
be written down. If you can't answer one, that is your next investigation.
### Q1 — "What is your bottleneck, and have you confirmed it empirically?"
**Recommended answer**: a named, measured stage in the workflow with
queue-time data showing where work waits. Not a vibe. Not "escalations
take too long". An actual measured queue.
**Why it's the first question**: Goldratt (*The Goal*, 1984) — every
system has exactly one binding constraint at a time. Sizing around the
wrong constraint wastes hires entirely. If you do not know your
bottleneck, run `process-mapper` BEFORE this skill.
**Canon**: Eli Goldratt, *The Goal* (1984); Reinertsen, *Principles of
Product Development Flow* (2009).
### Q2 — "What service trade-off are you accepting?"
**Recommended answer**: a written, explicit choice — fast vs. empathetic,
broad vs. deep, low-cost vs. high-quality. Frances Frei is unambiguous:
you cannot win all four. The team that tries wins zero.
**Why it matters**: AHT, SLA, and shrinkage inputs are the operational
expression of this trade-off. If they don't agree (e.g., you set AHT for
"empathy" but SLA for "speed"), the plan is internally inconsistent.
**Canon**: Frances Frei & Anne Morriss, *Uncommon Service* (HBR Press,
2012).
### Q3 — "What's your demand P90, and what's the gap to your P99?"
**Recommended answer**: two specific numbers from the last 90 days of
data, with the calendar context of each (e.g., "P90 was 480 tickets/day
on normal Tuesdays; P99 was 720 on the day after the November release").
A team sized to P50 misses SLA half the time. A team sized to P99
overstaffs by 30-50%. P90 is the right operating sizing point per
Cleveland.
**Canon**: Brad Cleveland, *Call Center Management on Fast Forward* (4th
ed., 2019); A.K. Erlang, *The Theory of Probabilities and Telephone
Conversations* (1909).
### Q4 — "At your planned utilization, what is P(SLA breach) at P90 and at P99?"
**Recommended answer**: two probabilities, computed (not guessed) from
Erlang-C with your specific N, AHT, and SLA target. If P(breach at P90)
> 10% you are understaffed at the sizing point. If P(breach at P99) >
50% you have no surge plan and the next peak event will be visible to
the CEO.
**Canon**: Erlang (1909); Hopp & Spearman, *Factory Physics* (3rd ed.,
2008), VUT equation.
### Q5 — "Have you budgeted replacement hires for the attrition you'll see this year?"
**Recommended answer**: yes, with a specific number. At 30% annual
attrition (Bersin BPO midpoint), a 20-FTE team loses ~6 people this year.
If your "add 5 net" plan is actually a "hire 11" plan, the recruiting
volume changes drastically. Anti-pattern #3.
**Canon**: Bersin/Deloitte talent benchmarks (2015-2023); Edward Lawler,
*Strategic Workforce Planning* (USC CEO, 2008).
### Q6 — "When does span of control trigger a manager hire, and who is the candidate?"
**Recommended answer**: a specific quarter (from `hiring_sequencer.py`)
and at least one identified candidate (internal lead or external hire).
Past 7 ICs/manager, 1:1s degrade, feedback cycles slip, attrition
climbs. Past 10 you have a coverage crisis. Hire the manager BEFORE
crossing 10, not after.
**Canon**: Camille Fournier, *The Manager's Path* (O'Reilly, 2017),
ch. 5; Andy Grove, *High Output Management* (1983).
### Q7 — "What is your surge plan for the P99 day?"
**Recommended answer**: an explicit, documented plan — overflow tier,
BPO contracted capacity, on-call rotation, executive escalation tree,
OR a written degradation contract that says "on P99 days we extend SLA
to X minutes and notify customers proactively". If the answer is "we'll
figure it out", the P99 day is a fire visible to the board.
**Canon**: Hopp & Spearman, *Factory Physics* (2008); Reinertsen (2009)
on capacity-margin discipline.
---
**Walk these seven in order. One at a time. Write the answers down. The
plan you submit is only as defensible as your answers to these seven
questions.**
FILE:assets/capacity_brief_template.md
# Capacity Planning Brief — {{TEAM_NAME}}
> 20-minute fill-out. Bring this brief plus your last 90 days of ticket /
> case / work-item data and you have everything needed to produce a
> defensible Q+1 plan.
## Section 1 — Context (5 minutes)
- **Team name:** {{TEAM_NAME}}
- **Function:** [support / cx / bizops / finance-ops / it-ops]
- **Planning horizon:** [Q+1 / annual / 12-month rolling]
- **Current headcount:** {{CURRENT_FTE}}
- **Working hours/day:** {{WORKING_HOURS_PER_DAY}}
- **Top business event driving this plan:** _(growth target, peak season,
M&A integration, regulatory change, etc.)_
## Section 2 — Demand (5 minutes)
Pull from your ticketing system (Zendesk, Intercom, Jira Service
Management, Salesforce, ServiceNow, etc.) the daily volume for the last
90 days. Compute or read off:
- **P50 (median day):** {{P50_TICKETS_PER_DAY}}
- **P90 (peak-band day):** {{P90_TICKETS_PER_DAY}}
- **P99 (annual peak day):** {{P99_TICKETS_PER_DAY}}
> If you only have averages, this plan is built on sand. Pull the
> distribution. (Anti-pattern #5: size-to-P50-only.)
- **Average handle time (AHT, minutes):** {{AHT_MINUTES}}
- **SLA target (minutes to first response or resolution):** {{SLA_MINUTES}}
- **Channels in scope:** _(voice / email / chat / async / multi)_
- **Multi-channel premium expected:** [yes / no — if multi, add 15-25%]
## Section 3 — People Realities (5 minutes)
- **Shrinkage % (paid time NOT productive):** {{SHRINKAGE_PCT}}
_(default if unknown: support 30, cx 32, bizops 25, finance-ops 22, it-ops 28)_
- **Ramp time for new hire (weeks to full productivity):** {{RAMP_WEEKS}}
_(default: support 8, cx 10, bizops 12, finance-ops 14, it-ops 10)_
- **Annual attrition observed last 12 months:** {{ATTRITION_PCT}}
_(default: support 30, cx 28, bizops 18, finance-ops 15, it-ops 20)_
- **Max hires per quarter (recruiting + onboarding constraint):**
{{MAX_HIRES_PER_QUARTER}}
- **Current managers and span of control:** _(list manager names + their
direct-report counts)_
## Section 4 — Strategic Constraints (5 minutes)
- **QoQ demand growth assumption:** {{GROWTH_QOQ_PCT}}
- **Bottleneck identified upstream (via process-mapper or similar):**
_(if you don't know your bottleneck, run process-mapper FIRST — sizing
around the wrong constraint is wasted hires)_
- **Service trade-off accepted:** _(per Frances Frei — pick which
attributes to win: speed / empathy / breadth / cost)_
- **Surge plan for P99 events:** _(overflow tier? BPO? on-call?
documented degradation?)_
---
## Tool Inputs
### Input JSON for `capacity_modeler.py`
```json
{
"team_name": "{{TEAM_NAME}}",
"demand": {
"tickets_per_day_p50": {{P50_TICKETS_PER_DAY}},
"tickets_per_day_p90": {{P90_TICKETS_PER_DAY}},
"tickets_per_day_p99": {{P99_TICKETS_PER_DAY}}
},
"sla_target_minutes": {{SLA_MINUTES}},
"current_fte": {{CURRENT_FTE}},
"avg_handle_time_minutes": {{AHT_MINUTES}},
"shrinkage_pct": {{SHRINKAGE_PCT}},
"working_hours_per_day": {{WORKING_HOURS_PER_DAY}}
}
```
Run:
```bash
python3 scripts/capacity_modeler.py --input my_brief.json --profile support
```
### Input JSON for `utilization_analyzer.py`
```json
{
"team_members": [
{
"name": "<name>",
"role": "<role>",
"utilization_pct": <0-100>,
"handles_count": <int>,
"hours_billable": <float>,
"hours_capacity": <float>
}
]
}
```
Run:
```bash
python3 scripts/utilization_analyzer.py --input team_util.json
```
### Input JSON for `hiring_sequencer.py`
```json
{
"team_name": "{{TEAM_NAME}}",
"current_fte": {{CURRENT_FTE}},
"target_fte_end_of_year": {{TARGET_EOY_FTE}},
"ramp_time_weeks": {{RAMP_WEEKS}},
"attrition_rate_annual_pct": {{ATTRITION_PCT}},
"growth_assumption_qoq_pct": {{GROWTH_QOQ_PCT}},
"hiring_constraints": {
"max_hires_per_quarter": {{MAX_HIRES_PER_QUARTER}}
}
}
```
Run:
```bash
python3 scripts/hiring_sequencer.py --input my_brief.json --profile support
```
---
## Output Checklist
After running all three tools, you should have:
- [ ] **Erlang-C sizing**: required FTE at 70/80/90% utilization (size to 80%)
- [ ] **Headroom %**: extra demand tolerable before SLA breaks (target >20%)
- [ ] **Risk band**: SAFE / WATCH / AT_RISK / CRITICAL
- [ ] **Team health verdict**: HEALTHY / SQUEEZED / OVERLOADED / UNBALANCED
- [ ] **Quarterly hiring plan**: ICs + managers + expected attrition per quarter
- [ ] **Manager-trigger callouts**: which quarter you add a manager
- [ ] **Warnings**: any quarter where hiring constraint blocks your plan
- [ ] **Forcing-question answers**: documented decisions on bottleneck,
service trade-offs, surge plan, P99 strategy (see SKILL.md
*Forcing-question library*)
## Sign-off
- **Prepared by:** ______________
- **Reviewed by (finance + HR + CS leader):** ______________
- **Decision and date:** ______________
- **Re-test trigger:** _(quarterly review date or demand-level threshold
that forces re-run)_
FILE:references/capacity_anti_patterns.md
# Capacity Planning Anti-Patterns
Every ops leader who has missed a peak season has fallen into one or
more of these patterns. The math (queueing-theory-canon.md) and the
people factors (ops-workforce-planning-canon.md) make each of these
predictably destructive. This reference enumerates the eight most
common failure modes with sources and the specific guard each tool
implements.
## The Anti-Patterns
### 1. Plan-to-100%-Utilization
**The mistake:** "We have 10 people billing 40 hours each, so we have
400 hours of capacity. Demand is 380 hours. We're fine."
**Why it fails:** Erlang-C and Hopp & Spearman's VUT equation both show
queue length grows as U/(1-U). At 95% utilization, average wait time is
~19× the service time. At 99%, it's ~99×. Variability turns a
"barely-covered" plan into nightly fires.
**Source:** Donald Reinertsen, *Principles of Product Development
Flow* (2009), Principle 12: "We need to operate at lower levels of
utilization."
**Tool guard:** `capacity_modeler.py` sizes against 70/80/90% scenarios
and flags any sizing point above 85% with a Reinertsen-cited warning.
### 2. Treat-Ramp-as-Instant
**The mistake:** "We approved 8 new hires for Q3, so we have +8 FTE
of capacity starting Q3."
**Why it fails:** A new T1 support hire is ~50% productive in weeks 1-8.
A new BizOps analyst is closer to ~30% productive in weeks 1-12 because
of tool sprawl and tribal knowledge. The "wait, they're not contributing
yet" gap is when your team burns out.
**Source:** Will Larson, *Staff Engineer* (Stripe Press, 2021); Camille
Fournier, *The Manager's Path* (O'Reilly, 2017).
**Tool guard:** `hiring_sequencer.py` applies a productivity factor
that linearly ramps 50% → 100% over `ramp_time_weeks`, and front-loads
hires (Q1 35%, Q2 30%, Q3 20%, Q4 15%) so EOY productivity catches the
adjusted target.
### 3. Ignore-Attrition
**The mistake:** "We have 15 today. We need 35 by EOY. Hire 20."
**Why it fails:** At 30% annual attrition (BPO-industry midpoint), you
will lose 4-5 of the original 15 during the year AND ~3-5 of your new
hires before they fully ramp. The real gap is 28-30 hires, not 20.
**Source:** Bersin / Deloitte talent benchmarks (2015-2023); Edward
Lawler, *Strategic Workforce Planning* (USC CEO, 2008).
**Tool guard:** `hiring_sequencer.py` requires `attrition_rate_annual_pct`
and distributes attrition quarterly via compounded probability, adding
the expected replacement hires to the gap calculation.
### 4. Hire-ICs-Forever
**The mistake:** "We don't need a manager — everyone's an
individual contributor reporting to the director."
**Why it fails:** Fournier's research (and Andy Grove's *High Output
Management* before her) is unambiguous: at 8-10+ direct reports, 1:1s
degrade, feedback cycles slip, attrition climbs, and the director
becomes the bottleneck. The cost shows up as attrition + ramp re-work,
not as a missed SLA.
**Source:** Camille Fournier, *The Manager's Path*, ch. 5; Andy
Grove, *High Output Management* (1983), ch. on managerial output.
**Tool guard:** `hiring_sequencer.py` triggers a manager hire when
projected span of control exceeds 7 ICs per manager, reallocating one
quarter's IC slot to a manager hire.
### 5. Size-to-P50-Demand-Only
**The mistake:** "Average daily volume is 320 tickets. We can handle
that."
**Why it fails:** Demand is a distribution, not a number. If P50 is
320 and P90 is 480, you will be staffed below SLA 10% of business
days. Customers don't care that you hit SLA on average; they
remember the day you didn't.
**Source:** Brad Cleveland, *Call Center Management on Fast Forward*
(4th ed., 2019); A.K. Erlang (1909) on traffic distributions.
**Tool guard:** `capacity_modeler.py` requires P50, P90, AND P99
demand inputs and sizes the recommendation to **P90** with breach
probability reported at all three percentiles.
### 6. No-Shrinkage-Adjustment
**The mistake:** "Our agents work 8 hours a day, so 8 hours of
capacity per agent."
**Why it fails:** 30% shrinkage is industry-typical. The 8 hours
actually delivers ~5.6 productive hours after breaks, training,
1:1s, sync meetings, ad-hoc interrupts, and the unspoken time spent
recovering between high-cognitive-load contacts.
**Source:** Cleveland, *Call Center Management on Fast Forward*;
Google SRE Workbook (2018) ch. 6 on toil budgets.
**Tool guard:** `capacity_modeler.py` requires `shrinkage_pct`,
applies a profile default if not provided (support 30%, BizOps 25%,
finance-ops 22%, IT-ops 28%), and outputs **loaded FTE** (post-shrinkage)
distinct from **raw FTE** (Erlang-C agents).
### 7. Single-Channel-Model-for-Multi-Channel-Work
**The mistake:** "Sum Erlang-C of voice + chat + email = total
required FTE."
**Why it fails:** Skill-switching cost. Demand-distribution mismatch
(chat is bursty; email queues overnight; voice spikes at 10-11am).
Real blended-agent productivity is 15-25% below the simple sum
because handoff context-loss is taxed per switch. Gartner research
consistently finds this premium.
**Source:** Gartner Customer Service & Support practice annual
benchmarks (2015-2023); Sir J.F.C. Kingman (1961) on G/G/1 queues
(variability amplifies wait time).
**Tool guard:** `capacity_modeler.py` `--profile` flag encodes
channel-mix realities (support, cx profiles assume blended channels
with a higher shrinkage default).
### 8. No-Surge-Plan-for-P99-Events
**The mistake:** "We're sized to P90 demand. The P99 day will be bad
but it's only 1% of days."
**Why it fails:** P99 days correlate with the highest-revenue events
(product launches, billing-cycle peaks, security incidents,
regulatory deadlines). Missing SLA on those days has
outsize commercial consequences relative to the calendar share.
You need an explicit surge plan: overflow tiering, on-call rotation,
contracted BPO overflow capacity, or a documented degradation contract.
**Source:** Hopp & Spearman, *Factory Physics* (3rd ed., 2008) on
peak-demand staffing; Reinertsen, *Principles of Product Development
Flow* on capacity-margin discipline.
**Tool guard:** `capacity_modeler.py` reports P(SLA breach) at P99 in
all three utilization scenarios, surfacing whether your 80%-utilization
sizing leaves you exposed on peak days.
## Additional Anti-Patterns Worth Naming
Beyond the eight, three more deserve mention because they appear in
nearly every quarterly planning cycle:
### 9. Use-Last-Year's-AHT
Average handle time creeps. Product complexity grows. Self-service
deflects the easy tickets, leaving harder ones in the queue. **Re-baseline
AHT every quarter**, not annually. (Source: Cleveland.)
### 10. Conflate-Operational-with-Strategic-Planning
This skill is for the *next 12 months*. If you are planning for a 3-year
automation reshape, you need chro-advisor or a strategic workforce plan,
not Erlang-C. (Source: Lawler.)
### 11. Plan-Without-Demand-Forecast-Confidence-Interval
A single point estimate of "we'll handle 4,000 tickets/month next year"
is a fiction. You need a forecast distribution. If sales forecasts a
40% YoY growth, your demand P90 grows faster than your demand P50 (more
variance). (Source: Kingman; Hopp-Spearman.)
## Christensen-Raynor on Resource Allocation
Clayton Christensen and Michael Raynor's *The Innovator's Solution*
(HBR Press, 2003) makes the meta-point: **a company's actual strategy
is what it staffs**, not what it says. If your capacity plan funds
firefighting at 80% and improvement at 20%, your strategy is
firefighting regardless of any PowerPoint deck. The capacity plan is
where strategy meets payroll. Take it seriously.
## Lencioni and Goldratt: The Two Disciplines
Pat Lencioni's *The Five Dysfunctions of a Team* (2002) and Eli
Goldratt's *The Goal* (1984) bookend the operational reality:
- **Lencioni**: trust + healthy conflict are prerequisites for
capacity discussions to be honest. Teams that can't have direct
conversations about whether someone is overloaded will silently
fail the capacity plan.
- **Goldratt**: subordinate everything to the bottleneck. If your
bottleneck is escalation engineering, don't hire more T1s — you'll
just queue more work at the choke point.
These appear in the *Forcing-question library* of SKILL.md.
## McKinsey + MIT Sloan on Queueing-Blind Staffing
McKinsey's Customer Care practice (2018-2023 reports) and MIT Sloan's
Service Operations research repeatedly document the gap between
"intuitive" staffing (manager judgment, headcount ratios) and
queueing-theory staffing. The gap is empirically 15-35%: intuitive
plans understaff at peak and overstaff at trough. The fix is not more
intuition; it is the math in `capacity_modeler.py`.
## The Closing Discipline
For every capacity plan, ask three questions:
1. **What's the queueing math?** (Erlang-C, P90 demand, ≤80% util.)
2. **What's the people reality?** (Ramp, attrition, span of control.)
3. **What's the bottleneck?** (Capacity-planner sizes around a
bottleneck; it does not find it. Use `process-mapper` first.)
Miss any of these and the plan is fiction.
FILE:references/ops_workforce_planning_canon.md
# Ops Workforce Planning Canon
Capacity sizing is half the answer. The other half is the human reality
of hiring, ramping, retaining, and structuring the people who staff the
queue. This reference assembles the operational-leadership canon needed
to translate an Erlang-C number into an executable 12-month plan.
## The Canon
### 1. Camille Fournier — *The Manager's Path* (O'Reilly, 2017)
Definitive guide to engineering management ladder, but the
**span-of-control** chapters apply to any ops team. Key thresholds the
`hiring_sequencer.py` enforces:
- **5-7 direct reports** is the healthy band for an ops manager.
- **8-9** is the warning zone — the manager starts dropping 1:1s,
feedback cycles slip, and team-level decisions queue.
- **10+** = you have a coverage problem, not a leadership problem.
Hire another manager BEFORE crossing 10.
Fournier also formalizes the **player-coach → pure manager → manager of
managers** progression that gates when a team needs a director.
### 2. Will Larson — *Staff Engineer* (Stripe Press, 2021) and *An Elegant Puzzle*
Larson's chapter on **ramp time as a real cost** is the source for the
"productive ~50% during ramp, 100% after" curve in
`hiring_sequencer.py`. Empirical observations:
- Support T1 hires: 6-10 weeks to full ramp.
- BizOps / Finance ops hires: 12-16 weeks (tool sprawl + tribal
knowledge).
- IT ops on-call rotation: 8-12 weeks before first solo on-call.
Larson's broader point: **hiring during a fire is too late**. The
sequencer's front-loaded weight (Q1 35%, Q4 15%) is the operational
expression of this principle.
### 3. Betsy Beyer, Niall Murphy, et al. — *The Site Reliability Workbook* (O'Reilly, 2018), Chapter 6: "Eliminating Toil"
Google SRE's framework for **toil budgets** maps directly to ops
shrinkage. Key staffing principle:
- An on-call ops engineer should spend **≤50% on toil**, the rest on
engineering work that reduces toil.
- If the toil fraction exceeds 50% for >1 quarter, you are
understaffed relative to incident volume.
The capacity-planner sibling skills (incident-coordinator,
process-mapper) feed inputs into this; the workforce-planning
implication is that "100% of paid hours = available capacity" is
**always** wrong.
### 4. Frances Frei & Anne Morriss — *Uncommon Service* (HBR Press, 2012)
Frei's central argument: **you cannot deliver excellent service across
all attributes simultaneously**. Service-design trade-offs — speed vs.
empathy, breadth vs. depth, low cost vs. high quality — directly
constrain how you size a team. A team chasing all four wins zero of
them. Capacity-planner inputs (AHT, SLA, channel mix) implicitly encode
which trade-off is being made; surfacing that explicitly in the
*Forcing-question library* is what separates a competent ops leader
from a guesser.
### 5. Edward Lawler — *Strategic Workforce Planning* (USC Marshall Center for Effective Organizations, 2008)
Lawler's research distinguishes **operational** workforce planning
(this skill: 0-12 months, role-specific, demand-driven) from
**strategic** workforce planning (CHRO's job: 1-5 years, capability
portfolio, talent supply analysis). The hard rule:
- If you are sizing against next quarter's tickets, you need
capacity-planner.
- If you are sizing against the company's 3-year automation strategy,
you need a strategic workforce plan (chro-advisor).
- **Conflating them gets you hired into the wrong jobs.**
### 6. Bersin / Deloitte — *Talent Acquisition Maturity Model* and benchmarks (2015-2023)
Source for the empirically reasonable **attrition + replacement-hire
defaults** in `hiring_sequencer.py` profiles. Bersin benchmarks:
- Support frontline: 25-35% annual attrition. The 30% default reflects
the BPO-industry midpoint.
- CX/Customer success: 22-28%. Slightly stickier than raw support due
to relationship investment.
- BizOps/Finance ops: 15-22%. Specialist + analytical work has lower
turnover.
- Open ops headcount fills in 45-90 days for T1, 90-180 days for
T2/specialist.
These figures must be sanity-checked against your own HR data — they
are starting points, not commitments.
### 7. Gartner Service Delivery Research — annual reports (Customer Service & Support practice)
Gartner's annual ops benchmarks codify the **multi-channel staffing
premium**: a team that handles voice + email + chat needs ~15-25% MORE
FTE than the simple-sum Erlang-C of each channel alone, because:
- Skill switching cost (context loss between channels).
- Non-uniform demand distributions across channels.
- The classic "blended-agent illusion" — agents claim to be 100%
flexible across channels but their effective handle times degrade.
The `--profile` flag in capacity_modeler is the place to encode this;
support and CX profiles assume multi-channel realities.
### 8. Andy Grove — *High Output Management* (1983, re-issued 1995)
Grove's framework for **leveraged output** — the manager's output is
the output of her team plus the output of every team she influences.
Applied to ops capacity:
- **A manager's "productive" contribution is not their own ticket
handles** (which should approach zero past 6 directs); it's their
effect on the team's throughput, accuracy, and retention.
- This is why the hiring_sequencer counts managers separately from ICs
and triggers manager hires preemptively at span-of-control limits.
## How These Connect to the Tools
| Tool | Primary Canon |
|---|---|
| `capacity_modeler.py` | Frei (service trade-offs encoded in inputs), Gartner (multi-channel) |
| `utilization_analyzer.py` | SRE Workbook (toil budget = ceiling), Grove (manager leverage) |
| `hiring_sequencer.py` | Fournier (span of control), Larson (ramp curves), Bersin (attrition), Lawler (operational vs. strategic) |
## The Hard Truths
1. **Ramp is real and is a 6-16 week tax.** Plans that assume new
hires are productive day one are works of fiction.
2. **You will lose 15-35% of your team this year.** Hiring plans that
don't budget replacement hires understaff you by exactly that much.
3. **You cannot manage 10+ direct reports.** Past 7-8, you are picking
which directs to neglect.
4. **Service trade-offs are non-negotiable.** Pick which dimensions to
win and accept losses elsewhere — Frei's central thesis.
FILE:references/queueing_theory_canon.md
# Queueing Theory Canon for Ops Capacity Planning
Capacity planning for ops teams (Support, CX, BizOps, Finance ops, IT ops)
without queueing theory is guesswork. The fundamental insight: as
utilization approaches 100%, wait time approaches infinity non-linearly.
Plan to 90% and your SLA collapses. Plan to 70-80% and you have surge
capacity. The math is over 100 years old, and ignoring it is the most
expensive mistake an ops leader makes.
## The Canon
### 1. A.K. Erlang (1909) — *The Theory of Probabilities and Telephone Conversations*
Erlang's seminal paper on telephone-traffic load. Introduced the Erlang
unit of offered load (a = arrival_rate × service_time) and gave the
formula now called Erlang-B (loss systems) and Erlang-C (waiting systems).
**Erlang-C** is what you want for any ops queue where tickets/calls wait
rather than being dropped:
```
P(wait) = (a^N / N!) * (N / (N - a)) / [ sum_{k=0..N-1} a^k/k! + (a^N/N!)*(N/(N-a)) ]
```
Where N is the number of agents, a is the offered load. **This is the
single most important formula in ops capacity planning.** Implemented in
`scripts/capacity_modeler.py` in ~30 lines of stdlib Python.
### 2. J.D.C. Little (1961) — *A Proof for the Queuing Formula L = λW*
**Little's Law**: in steady state, the average number of items in a queue
(L) equals the average arrival rate (λ) multiplied by the average time an
item spends in the system (W). Three implications for ops leaders:
- **You cannot pick L, λ, and W independently.** If demand (λ) doubles
and headcount (L capacity) stays flat, wait time (W) must double — and
via Erlang-C, it actually grows much faster than that near saturation.
- **Reducing average handle time** is mathematically equivalent to
hiring, up to the utilization ceiling.
- **WIP limits work** because they put a hard cap on L, which (given fixed
capacity throughput) directly caps W.
### 3. Hopp & Spearman — *Factory Physics* (3rd ed., 2008)
The bible of operations science applied to manufacturing. Chapters 8-9
cover variability and queueing rigorously. Key takeaway for ops leaders:
**the VUT equation** for cycle time at a workstation,
```
CT_q ≈ V × U × T
```
where V is variability (coefficient-of-variation squared), U is utilization
factor U/(1-U), and T is mean service time. Notice U/(1-U): at U=0.80, the
multiplier is 4. At U=0.90, it's 9. At U=0.95, it's 19. **This is why
"plan to 100% utilization" is the most expensive sentence in ops.**
### 4. Donald Reinertsen — *The Principles of Product Development Flow* (2009)
The most important book on queueing in knowledge work. Principle 7
("Queue size, not capacity utilization, is the primary control variable")
and Principle 12 ("We need to operate at lower levels of utilization")
make the case rigorously: **80% utilization is the safe operating ceiling
for variable-arrival queues**. Past that, queue length and cycle time
explode super-linearly. Reinertsen's diagnostic chart of "% utilization
vs. queue length" should be hanging in every ops leader's office.
### 5. Sir J.F.C. Kingman — *On Queues in Heavy Traffic* (1961)
**Kingman's formula** for a G/G/1 queue (general arrival + general
service distribution, single server):
```
E[W_q] ≈ (ρ / (1-ρ)) × ((c_a^2 + c_s^2) / 2) × τ
```
where ρ is utilization, c_a and c_s are coefficients of variation for
arrivals and service, τ is mean service time. **Implication: variability
in either arrivals or service amplifies wait time as much as utilization
does.** This is why bursty channels (email tickets that all arrive at
9am Monday) require MORE staffing slack than steady channels.
### 6. Brad Cleveland — *Call Center Management on Fast Forward* (4th ed., 2019)
The applied operating manual for service-level queues. Conventions
codified by Cleveland and used in `scripts/capacity_modeler.py`:
- Size to **P90 demand** (not P50, not P99) — P50 leaves you breaking
SLA half the time; P99 over-staffs by 30-50%.
- **Shrinkage** must be a line item. 30% is a reasonable default for
support (training, breaks, sync meetings, PTO, ad-hoc interrupts).
- **Service level** is the right SLA metric, not abandon rate alone:
"answered within T seconds" as a probability.
### 7. ITIL 4 Service Management Practices (Axelos, 2019)
ITIL's *Service Operation* practice guidance codifies the canonical
demand-and-capacity-management process for IT ops teams. Key constructs
borrowed:
- **Demand management** = forecasting + smoothing (e.g., release calendars
scheduling fewer changes during peak ticket windows).
- **Capacity management** = three sub-processes: business capacity
(forecast), service capacity (workload analysis), component capacity
(resource-level).
- **Service-level management** = the SLA contract that ties Erlang-C
inputs to commitments.
### 8. M. Armony, S. Israelit, A. Mandelbaum, et al. — *On Patient Flow in Hospitals* (Stochastic Systems, 2015)
Modern empirical research on multi-class queues with priorities and
abandonment — directly applicable to multi-tier support (T1/T2/T3 with
escalation paths). Empirically validates that **abandonment plus
priority routing** in a real call/ticket center produces wait-time
distributions very close to G/G/c with class-specific service rates.
Confirms Erlang-C is the right "first model" for capacity sizing.
## How These Connect to the Tools
| Tool | Primary Canon |
|---|---|
| `capacity_modeler.py` | Erlang (1909) Erlang-C, Cleveland (sizing convention) |
| `utilization_analyzer.py` | Reinertsen (>80% threshold), Hopp-Spearman VUT, Little's Law |
| `hiring_sequencer.py` | Cleveland (shrinkage), Kingman (variability premium), ITIL (demand management) |
## The One-Sentence Summary
If you remember nothing else: **never plan an ops team to above 80%
sustained utilization** — Reinertsen Principle 12, validated by Hopp &
Spearman's VUT equation and Erlang's 1909 telephone-traffic math. The
arithmetic is unforgiving.
FILE:scripts/capacity_modeler.py
#!/usr/bin/env python3
"""capacity_modeler.py — Ops capacity sizing via Erlang-C queueing math.
Sizes an ops team (Support / CX / BizOps / Finance ops / IT ops) against
demand and an SLA target. Implements Erlang-C in pure stdlib to compute:
* Required FTE at 70%, 80%, and 90% utilization
* Probability of SLA breach at each utilization level
* Capacity headroom — extra tickets/day before SLA breaks
Industry profiles tune default shrinkage and SLA conventions.
Stdlib only. No LLM calls. Deterministic. Save the JSON sample for shape.
"""
from __future__ import annotations
import argparse
import json
import math
import sys
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
from typing import Any
# ---------------------------------------------------------------------------
# Industry profiles
# ---------------------------------------------------------------------------
PROFILES: dict[str, dict[str, float]] = {
# shrinkage = % of paid time NOT available for productive ticket-handling
# (training, breaks, sync, PTO accrual, ad-hoc interrupts)
"support": {"shrinkage_pct_default": 30.0, "sla_target_minutes_default": 60.0},
"cx": {"shrinkage_pct_default": 32.0, "sla_target_minutes_default": 30.0},
"bizops": {"shrinkage_pct_default": 25.0, "sla_target_minutes_default": 240.0},
"finance-ops": {"shrinkage_pct_default": 22.0, "sla_target_minutes_default": 480.0},
"it-ops": {"shrinkage_pct_default": 28.0, "sla_target_minutes_default": 120.0},
}
class RiskBand(str, Enum):
SAFE = "SAFE"
WATCH = "WATCH"
AT_RISK = "AT_RISK"
CRITICAL = "CRITICAL"
# ---------------------------------------------------------------------------
# Erlang-C — pure stdlib implementation
# ---------------------------------------------------------------------------
def erlang_c_probability(agents: int, traffic_intensity: float) -> float:
"""Erlang-C: probability an arriving call/ticket has to wait.
agents (N) : number of servers
traffic_intensity (a) : offered load in Erlangs (lambda * AHT)
dimensionless; must satisfy a < N for stability.
Returns P(wait) in [0, 1]. Returns 1.0 if system unstable (a >= N).
"""
if agents <= 0:
return 1.0
if traffic_intensity <= 0:
return 0.0
if traffic_intensity >= agents:
return 1.0
# Numerator: a^N / N! * N / (N - a)
# Denominator: sum_{k=0}^{N-1} a^k / k! + numerator
# Computed in log-space to avoid overflow on big numbers.
a = traffic_intensity
n = agents
# log(a^n / n!) = n*log(a) - lgamma(n+1)
log_a_n_over_nfact = n * math.log(a) - math.lgamma(n + 1)
numerator_term = math.exp(log_a_n_over_nfact) * (n / (n - a))
sum_terms = 0.0
for k in range(n):
log_term = k * math.log(a) - math.lgamma(k + 1)
sum_terms += math.exp(log_term)
denom = sum_terms + numerator_term
if denom <= 0:
return 1.0
return numerator_term / denom
def service_level(agents: int, traffic_intensity: float,
aht_seconds: float, sla_target_seconds: float) -> float:
"""P(answered within SLA target) for M/M/c queue.
SL = 1 - P_wait * exp(-(N - a) * T / AHT)
"""
pw = erlang_c_probability(agents, traffic_intensity)
if agents <= traffic_intensity:
return 0.0
exponent = -(agents - traffic_intensity) * (sla_target_seconds / aht_seconds)
# guard against overflow
if exponent < -700:
return 1.0 - pw * 0.0
return 1.0 - pw * math.exp(exponent)
def required_agents_for_utilization(traffic_intensity: float,
target_utilization: float) -> int:
"""Minimum N such that traffic_intensity / N <= target_utilization."""
if target_utilization <= 0 or target_utilization >= 1:
raise ValueError("target_utilization must be in (0,1)")
n = math.ceil(traffic_intensity / target_utilization)
return max(n, 1)
# ---------------------------------------------------------------------------
# Data model
# ---------------------------------------------------------------------------
@dataclass
class Demand:
tickets_per_day_p50: float
tickets_per_day_p90: float
tickets_per_day_p99: float
@dataclass
class CapacityInput:
team_name: str
demand: Demand
sla_target_minutes: float
current_fte: float
avg_handle_time_minutes: float
shrinkage_pct: float
working_hours_per_day: float = 8.0
@dataclass
class UtilizationScenario:
target_utilization: float
required_fte_raw: int # before shrinkage
required_fte_loaded: float # after shrinkage
p_sla_breach_p50: float
p_sla_breach_p90: float
p_sla_breach_p99: float
actual_utilization_at_demand: float
@dataclass
class CapacityResult:
team_name: str
inputs: CapacityInput
scenarios: list[UtilizationScenario] = field(default_factory=list)
headroom_extra_tickets_per_day: float = 0.0
headroom_pct: float = 0.0
risk_band: RiskBand = RiskBand.SAFE
recommendation: str = ""
notes: list[str] = field(default_factory=list)
# ---------------------------------------------------------------------------
# Modeling
# ---------------------------------------------------------------------------
def model_capacity(inp: CapacityInput) -> CapacityResult:
aht_sec = inp.avg_handle_time_minutes * 60.0
sla_sec = inp.sla_target_minutes * 60.0
seconds_per_day_per_fte = inp.working_hours_per_day * 3600.0
productive_fraction = max(0.0, 1.0 - inp.shrinkage_pct / 100.0)
def traffic_for(volume_per_day: float) -> float:
# Erlang offered load (a) = arrival_rate * AHT, in consistent time units.
# Per-day arrival rate normalized to per-second:
arrival_per_sec = volume_per_day / seconds_per_day_per_fte
return arrival_per_sec * aht_sec
a_p50 = traffic_for(inp.demand.tickets_per_day_p50)
a_p90 = traffic_for(inp.demand.tickets_per_day_p90)
a_p99 = traffic_for(inp.demand.tickets_per_day_p99)
scenarios: list[UtilizationScenario] = []
for util in (0.70, 0.80, 0.90):
# Sizing is done against P90 demand by convention (Cleveland).
n_raw = required_agents_for_utilization(a_p90, util)
# Loaded headcount accounts for shrinkage: each "agent slot" needs
# 1 / productive_fraction headcount to staff it.
n_loaded = n_raw / productive_fraction if productive_fraction > 0 else float("inf")
sl_p50 = service_level(n_raw, a_p50, aht_sec, sla_sec)
sl_p90 = service_level(n_raw, a_p90, aht_sec, sla_sec)
sl_p99 = service_level(n_raw, a_p99, aht_sec, sla_sec)
scenarios.append(UtilizationScenario(
target_utilization=util,
required_fte_raw=n_raw,
required_fte_loaded=round(n_loaded, 2),
p_sla_breach_p50=round(1.0 - sl_p50, 4),
p_sla_breach_p90=round(1.0 - sl_p90, 4),
p_sla_breach_p99=round(1.0 - sl_p99, 4),
actual_utilization_at_demand=round(a_p90 / n_raw, 4) if n_raw > 0 else 1.0,
))
# Headroom: with current_fte (loaded), how many extra tickets/day before
# P(SLA breach) at P90 crosses 10%?
current_productive_fte = max(1, int(round(inp.current_fte * productive_fraction)))
headroom_volume = inp.demand.tickets_per_day_p90
step = max(1.0, inp.demand.tickets_per_day_p90 * 0.02)
while headroom_volume < inp.demand.tickets_per_day_p90 * 5:
a = traffic_for(headroom_volume)
if a >= current_productive_fte:
break
sl = service_level(current_productive_fte, a, aht_sec, sla_sec)
if (1.0 - sl) > 0.10:
break
headroom_volume += step
headroom_extra = max(0.0, headroom_volume - inp.demand.tickets_per_day_p90)
headroom_pct = (headroom_extra / inp.demand.tickets_per_day_p90 * 100.0
if inp.demand.tickets_per_day_p90 > 0 else 0.0)
# Risk band — pick from 80%-utilization scenario (canonical sizing point)
s80 = next(s for s in scenarios if s.target_utilization == 0.80)
if inp.current_fte >= s80.required_fte_loaded and headroom_pct >= 20:
band = RiskBand.SAFE
rec = (f"Sized correctly at {inp.current_fte} FTE for P90 demand at 80% "
f"utilization. Headroom is healthy ({headroom_pct:.0f}%).")
elif inp.current_fte >= s80.required_fte_loaded:
band = RiskBand.WATCH
rec = (f"Headcount adequate ({inp.current_fte} FTE vs. "
f"{s80.required_fte_loaded} required) but headroom thin "
f"({headroom_pct:.0f}%). Re-test in 30 days.")
elif inp.current_fte >= s80.required_fte_loaded * 0.85:
band = RiskBand.AT_RISK
rec = (f"Understaffed for P90 demand: have {inp.current_fte}, need "
f"{s80.required_fte_loaded} at 80% utilization. Expect SLA "
f"misses at P90 surges. Hire {math.ceil(s80.required_fte_loaded - inp.current_fte)} FTE.")
else:
band = RiskBand.CRITICAL
rec = (f"Critically understaffed: have {inp.current_fte}, need "
f"{s80.required_fte_loaded}. Throughput collapse risk per "
f"queueing theory at sustained >85% utilization. Escalate.")
notes: list[str] = []
if s80.actual_utilization_at_demand > 0.85:
notes.append(
"WARNING: Sizing point pushes >85% utilization. Reinertsen's "
"principle 7: throughput collapses non-linearly past 80%."
)
if inp.shrinkage_pct < 15:
notes.append("Shrinkage <15% likely understates non-productive time.")
if inp.shrinkage_pct > 40:
notes.append("Shrinkage >40% — verify against actual time-on-task data.")
return CapacityResult(
team_name=inp.team_name,
inputs=inp,
scenarios=scenarios,
headroom_extra_tickets_per_day=round(headroom_extra, 1),
headroom_pct=round(headroom_pct, 1),
risk_band=band,
recommendation=rec,
notes=notes,
)
# ---------------------------------------------------------------------------
# Rendering
# ---------------------------------------------------------------------------
def to_markdown(result: CapacityResult) -> str:
inp = result.inputs
lines = [
f"# Capacity Model — {result.team_name}",
"",
f"**Risk band:** {result.risk_band.value}",
"",
f"**Recommendation:** {result.recommendation}",
"",
"## Inputs",
f"- Current FTE: {inp.current_fte}",
f"- AHT: {inp.avg_handle_time_minutes} min",
f"- SLA target: {inp.sla_target_minutes} min",
f"- Shrinkage: {inp.shrinkage_pct}%",
f"- Working hours / day: {inp.working_hours_per_day}",
f"- Demand P50 / P90 / P99: {inp.demand.tickets_per_day_p50} / "
f"{inp.demand.tickets_per_day_p90} / {inp.demand.tickets_per_day_p99} tickets/day",
"",
"## Sizing Scenarios (Erlang-C, sized to P90 demand)",
"",
"| Target Util | Raw FTE | Loaded FTE (post-shrinkage) | P(SLA breach @ P50) | P(SLA breach @ P90) | P(SLA breach @ P99) |",
"|---|---|---|---|---|---|",
]
for s in result.scenarios:
lines.append(
f"| {int(s.target_utilization*100)}% | {s.required_fte_raw} | "
f"{s.required_fte_loaded} | {s.p_sla_breach_p50*100:.1f}% | "
f"{s.p_sla_breach_p90*100:.1f}% | {s.p_sla_breach_p99*100:.1f}% |"
)
lines.extend([
"",
"## Headroom",
f"- Extra tickets/day before SLA breaks: {result.headroom_extra_tickets_per_day}",
f"- Headroom %: {result.headroom_pct:.1f}%",
"",
])
if result.notes:
lines.append("## Notes")
for n in result.notes:
lines.append(f"- {n}")
lines.append("")
lines.append("## Canon")
lines.append("- Erlang (1909), Little (1961), Cleveland *Call Center Mgmt on Fast Forward*, Reinertsen *Principles of Product Development Flow*.")
return "\n".join(lines)
def to_dict(result: CapacityResult) -> dict[str, Any]:
return {
"team_name": result.team_name,
"risk_band": result.risk_band.value,
"recommendation": result.recommendation,
"headroom_extra_tickets_per_day": result.headroom_extra_tickets_per_day,
"headroom_pct": result.headroom_pct,
"scenarios": [
{
"target_utilization": s.target_utilization,
"required_fte_raw": s.required_fte_raw,
"required_fte_loaded": s.required_fte_loaded,
"p_sla_breach_p50": s.p_sla_breach_p50,
"p_sla_breach_p90": s.p_sla_breach_p90,
"p_sla_breach_p99": s.p_sla_breach_p99,
"actual_utilization_at_demand": s.actual_utilization_at_demand,
}
for s in result.scenarios
],
"notes": result.notes,
"inputs": {
"current_fte": result.inputs.current_fte,
"avg_handle_time_minutes": result.inputs.avg_handle_time_minutes,
"sla_target_minutes": result.inputs.sla_target_minutes,
"shrinkage_pct": result.inputs.shrinkage_pct,
"working_hours_per_day": result.inputs.working_hours_per_day,
"demand": {
"tickets_per_day_p50": result.inputs.demand.tickets_per_day_p50,
"tickets_per_day_p90": result.inputs.demand.tickets_per_day_p90,
"tickets_per_day_p99": result.inputs.demand.tickets_per_day_p99,
},
},
}
# ---------------------------------------------------------------------------
# Sample + parsing
# ---------------------------------------------------------------------------
SAMPLE_INPUT: dict[str, Any] = {
"team_name": "Tier-1 Support",
"demand": {
"tickets_per_day_p50": 320,
"tickets_per_day_p90": 480,
"tickets_per_day_p99": 720,
},
"sla_target_minutes": 60,
"current_fte": 12,
"avg_handle_time_minutes": 18,
"shrinkage_pct": 30,
"working_hours_per_day": 8,
}
def parse_input(raw: dict[str, Any], profile: str | None) -> CapacityInput:
prof = PROFILES.get(profile or "", {})
shrinkage = raw.get("shrinkage_pct", prof.get("shrinkage_pct_default", 30.0))
sla = raw.get("sla_target_minutes",
prof.get("sla_target_minutes_default", 60.0))
d = raw["demand"]
return CapacityInput(
team_name=raw["team_name"],
demand=Demand(
tickets_per_day_p50=float(d["tickets_per_day_p50"]),
tickets_per_day_p90=float(d["tickets_per_day_p90"]),
tickets_per_day_p99=float(d["tickets_per_day_p99"]),
),
sla_target_minutes=float(sla),
current_fte=float(raw["current_fte"]),
avg_handle_time_minutes=float(raw["avg_handle_time_minutes"]),
shrinkage_pct=float(shrinkage),
working_hours_per_day=float(raw.get("working_hours_per_day", 8.0)),
)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(
description="Erlang-C ops capacity sizer (stdlib only).",
)
p.add_argument("--input", type=Path, help="Path to JSON input file.")
p.add_argument(
"--profile",
choices=list(PROFILES.keys()),
default=None,
help="Industry profile (defaults for shrinkage + SLA).",
)
p.add_argument(
"--output",
choices=["markdown", "json"],
default="markdown",
help="Output format.",
)
p.add_argument(
"--sample",
action="store_true",
help="Run on built-in sample input and print result.",
)
args = p.parse_args(argv)
if args.sample:
raw = SAMPLE_INPUT
elif args.input:
raw = json.loads(args.input.read_text())
else:
p.error("Provide --input or --sample.")
return 2
try:
inp = parse_input(raw, args.profile)
except (KeyError, ValueError) as e:
print(f"ERROR parsing input: {e}", file=sys.stderr)
return 2
result = model_capacity(inp)
if args.output == "json":
print(json.dumps(to_dict(result), indent=2))
else:
print(to_markdown(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/hiring_sequencer.py
#!/usr/bin/env python3
"""hiring_sequencer.py — 12-month quarterly hiring plan for ops teams.
Accounts for:
* Ramp time (productive ~50% for ramp_time_weeks, then 100%)
* Annual attrition (compounded weekly across the year)
* Quarter-over-quarter demand growth
* Hiring constraints (max hires / quarter)
* Manager-trigger: when span of control crosses 7-8 ICs, schedule manager hire
Stdlib only.
"""
from __future__ import annotations
import argparse
import json
import math
import sys
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
from typing import Any
# Industry profiles — typical ramp + attrition
PROFILES: dict[str, dict[str, float]] = {
"support": {"ramp_time_weeks": 8.0, "attrition_rate_annual_pct": 30.0},
"cx": {"ramp_time_weeks": 10.0, "attrition_rate_annual_pct": 28.0},
"bizops": {"ramp_time_weeks": 12.0, "attrition_rate_annual_pct": 18.0},
"finance-ops":{"ramp_time_weeks": 14.0, "attrition_rate_annual_pct": 15.0},
"it-ops": {"ramp_time_weeks": 10.0, "attrition_rate_annual_pct": 20.0},
}
SPAN_OF_CONTROL_MAX = 7 # ICs per manager threshold (Fournier)
class Quarter(str, Enum):
Q1 = "Q1"
Q2 = "Q2"
Q3 = "Q3"
Q4 = "Q4"
@dataclass
class HiringInput:
team_name: str
current_fte: int
target_fte_end_of_year: int
ramp_time_weeks: float
attrition_rate_annual_pct: float
growth_assumption_qoq_pct: float
max_hires_per_quarter: int
@dataclass
class QuarterPlan:
quarter: Quarter
ic_hires: int
manager_hires: int
expected_attrition: int
productive_fte_end_of_quarter: float
headcount_end_of_quarter: int
span_of_control: float
notes: list[str] = field(default_factory=list)
@dataclass
class HiringResult:
team_name: str
inputs: HiringInput
quarters: list[QuarterPlan]
total_ic_hires: int
total_manager_hires: int
total_attrition: int
headline: str
warnings: list[str] = field(default_factory=list)
def _productivity_factor(weeks_since_hire: float, ramp_weeks: float) -> float:
"""Linear ramp 50% → 100% over ramp_weeks (Larson)."""
if weeks_since_hire >= ramp_weeks:
return 1.0
if weeks_since_hire <= 0:
return 0.5
return 0.5 + 0.5 * (weeks_since_hire / ramp_weeks)
def sequence(inp: HiringInput) -> HiringResult:
# Demand-side ratchet: target adjusted up by QoQ growth (compounded)
growth_factor_eoy = (1.0 + inp.growth_assumption_qoq_pct / 100.0) ** 4
adjusted_target = int(math.ceil(inp.target_fte_end_of_year * growth_factor_eoy))
# Per-quarter attrition probability — split annual rate across 4 quarters
q_attrition_rate = 1.0 - (1.0 - inp.attrition_rate_annual_pct / 100.0) ** 0.25
# Total gap to close: target + replacement hires over the year
expected_total_attrition = int(math.ceil(
inp.current_fte * (inp.attrition_rate_annual_pct / 100.0)
))
raw_gap = adjusted_target - inp.current_fte + expected_total_attrition
total_hires_needed = max(0, raw_gap)
# Distribute hires front-loaded but capped
quarters: list[QuarterPlan] = []
headcount = inp.current_fte
remaining = total_hires_needed
cumulative_managers = max(1, math.ceil(inp.current_fte / SPAN_OF_CONTROL_MAX))
cumulative_ic_hires = 0
cumulative_manager_hires = 0
cumulative_attrition = 0
warnings: list[str] = []
# Pre-emptive front-load: aim higher in early quarters so ramp completes by EOY
# Allocation weights Q1>Q2>Q3>Q4 since later hires miss ramp window.
weights = [0.35, 0.30, 0.20, 0.15]
for i, qname in enumerate(Quarter):
ideal_q_hires = math.ceil(total_hires_needed * weights[i])
q_hires = min(ideal_q_hires, inp.max_hires_per_quarter, remaining)
if q_hires < ideal_q_hires:
warnings.append(
f"{qname.value}: wanted {ideal_q_hires} hires but constrained "
f"to {q_hires} by max_hires_per_quarter."
)
# Attrition realized this quarter
q_attrition = int(round(headcount * q_attrition_rate))
cumulative_attrition += q_attrition
# New headcount after hires + attrition
new_headcount = headcount + q_hires - q_attrition
# Manager trigger check — if ICs / managers > threshold, add a manager hire
# (counted within the ic_hires bucket reallocated as manager)
ic_count_eoq = new_headcount - cumulative_managers
span = ic_count_eoq / max(cumulative_managers, 1)
notes: list[str] = []
manager_hires_this_q = 0
if span > SPAN_OF_CONTROL_MAX and q_hires > 0:
manager_hires_this_q = 1
q_hires -= 1
cumulative_managers += 1
notes.append(
f"Manager trigger fired: span was {span:.1f} ICs/manager > "
f"{SPAN_OF_CONTROL_MAX}. Reallocated 1 IC hire to manager hire."
)
# Recompute span after manager hire
ic_count_eoq = new_headcount - cumulative_managers
span = ic_count_eoq / max(cumulative_managers, 1)
cumulative_ic_hires += q_hires
cumulative_manager_hires += manager_hires_this_q
remaining -= (q_hires + manager_hires_this_q)
# Productive FTE = full-time members + ramp-fraction for in-quarter hires
# in-quarter hires are halfway through ramp on average at EOQ → ~halfway up the ramp curve
avg_weeks_for_q_hires = 6.5 # quarter midpoint (13 weeks / 2)
ramp_fraction = _productivity_factor(avg_weeks_for_q_hires, inp.ramp_time_weeks)
productive_fte = (headcount - q_attrition) + (q_hires + manager_hires_this_q) * ramp_fraction
if productive_fte < adjusted_target * 0.85 and i == 3:
warnings.append(
f"EOY productive FTE ({productive_fte:.1f}) below 85% of adjusted "
f"target ({adjusted_target}) — ramp will extend into next year."
)
quarters.append(QuarterPlan(
quarter=qname,
ic_hires=q_hires,
manager_hires=manager_hires_this_q,
expected_attrition=q_attrition,
productive_fte_end_of_quarter=round(productive_fte, 1),
headcount_end_of_quarter=new_headcount,
span_of_control=round(span, 2),
notes=notes,
))
headcount = new_headcount
headline = (
f"Hire {cumulative_ic_hires} ICs + {cumulative_manager_hires} managers "
f"across 4 quarters. End-of-year nominal headcount: {headcount} "
f"(adjusted target: {adjusted_target}). Expect ~{cumulative_attrition} "
f"attrition over the year."
)
return HiringResult(
team_name=inp.team_name,
inputs=inp,
quarters=quarters,
total_ic_hires=cumulative_ic_hires,
total_manager_hires=cumulative_manager_hires,
total_attrition=cumulative_attrition,
headline=headline,
warnings=warnings,
)
def to_markdown(r: HiringResult) -> str:
lines = [
f"# Hiring Plan — {r.team_name}",
"",
f"**Headline:** {r.headline}",
"",
"## Assumptions",
f"- Current FTE: {r.inputs.current_fte}",
f"- Target EOY FTE (nominal): {r.inputs.target_fte_end_of_year}",
f"- Ramp time: {r.inputs.ramp_time_weeks} weeks",
f"- Annual attrition: {r.inputs.attrition_rate_annual_pct}%",
f"- QoQ growth: {r.inputs.growth_assumption_qoq_pct}%",
f"- Max hires per quarter: {r.inputs.max_hires_per_quarter}",
"",
"## Quarterly Plan",
"",
"| Quarter | IC Hires | Manager Hires | Attrition | Headcount EOQ | Productive FTE EOQ | Span of Control |",
"|---|---|---|---|---|---|---|",
]
for q in r.quarters:
lines.append(
f"| {q.quarter.value} | {q.ic_hires} | {q.manager_hires} | "
f"{q.expected_attrition} | {q.headcount_end_of_quarter} | "
f"{q.productive_fte_end_of_quarter} | {q.span_of_control} |"
)
lines.append("")
notes_present = any(q.notes for q in r.quarters)
if notes_present:
lines.append("## Quarter Notes")
for q in r.quarters:
for n in q.notes:
lines.append(f"- {q.quarter.value}: {n}")
lines.append("")
if r.warnings:
lines.append("## Warnings")
for w in r.warnings:
lines.append(f"- {w}")
lines.append("")
lines.extend([
"## Canon",
"- Camille Fournier, *The Manager's Path* — span of control thresholds.",
"- Will Larson, *Staff Engineer* — ramp productivity curves.",
"- Bersin/Deloitte talent benchmarks — attrition + replacement hire ratios.",
])
return "\n".join(lines)
def to_dict(r: HiringResult) -> dict[str, Any]:
return {
"team_name": r.team_name,
"headline": r.headline,
"totals": {
"ic_hires": r.total_ic_hires,
"manager_hires": r.total_manager_hires,
"attrition": r.total_attrition,
},
"quarters": [
{
"quarter": q.quarter.value,
"ic_hires": q.ic_hires,
"manager_hires": q.manager_hires,
"expected_attrition": q.expected_attrition,
"productive_fte_end_of_quarter": q.productive_fte_end_of_quarter,
"headcount_end_of_quarter": q.headcount_end_of_quarter,
"span_of_control": q.span_of_control,
"notes": q.notes,
}
for q in r.quarters
],
"warnings": r.warnings,
"inputs": {
"current_fte": r.inputs.current_fte,
"target_fte_end_of_year": r.inputs.target_fte_end_of_year,
"ramp_time_weeks": r.inputs.ramp_time_weeks,
"attrition_rate_annual_pct": r.inputs.attrition_rate_annual_pct,
"growth_assumption_qoq_pct": r.inputs.growth_assumption_qoq_pct,
"max_hires_per_quarter": r.inputs.max_hires_per_quarter,
},
}
SAMPLE_INPUT: dict[str, Any] = {
"team_name": "Tier-1 Support",
"current_fte": 15,
"target_fte_end_of_year": 35,
"ramp_time_weeks": 8,
"attrition_rate_annual_pct": 30,
"growth_assumption_qoq_pct": 8,
"hiring_constraints": {"max_hires_per_quarter": 8},
}
def parse_input(raw: dict[str, Any], profile: str | None) -> HiringInput:
prof = PROFILES.get(profile or "", {})
ramp = raw.get("ramp_time_weeks", prof.get("ramp_time_weeks", 10.0))
attr = raw.get(
"attrition_rate_annual_pct",
prof.get("attrition_rate_annual_pct", 25.0),
)
constraints = raw.get("hiring_constraints", {}) or {}
return HiringInput(
team_name=raw["team_name"],
current_fte=int(raw["current_fte"]),
target_fte_end_of_year=int(raw["target_fte_end_of_year"]),
ramp_time_weeks=float(ramp),
attrition_rate_annual_pct=float(attr),
growth_assumption_qoq_pct=float(raw.get("growth_assumption_qoq_pct", 0.0)),
max_hires_per_quarter=int(constraints.get("max_hires_per_quarter", 999)),
)
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(
description="12-month quarterly hiring sequencer with ramp + attrition.",
)
p.add_argument("--input", type=Path, help="Path to JSON input file.")
p.add_argument(
"--profile",
choices=list(PROFILES.keys()),
default=None,
help="Industry profile (defaults for ramp + attrition).",
)
p.add_argument(
"--output", choices=["markdown", "json"], default="markdown",
help="Output format.",
)
p.add_argument("--sample", action="store_true",
help="Run on built-in sample input.")
args = p.parse_args(argv)
if args.sample:
raw = SAMPLE_INPUT
elif args.input:
raw = json.loads(args.input.read_text())
else:
p.error("Provide --input or --sample.")
return 2
try:
inp = parse_input(raw, args.profile)
except (KeyError, ValueError) as e:
print(f"ERROR parsing input: {e}", file=sys.stderr)
return 2
result = sequence(inp)
if args.output == "json":
print(json.dumps(to_dict(result), indent=2))
else:
print(to_markdown(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/utilization_analyzer.py
#!/usr/bin/env python3
"""utilization_analyzer.py — per-member + team-level utilization health.
Detects:
* RED : sustained >85% utilization (throughput collapse risk, Reinertsen)
* AMBER: 70-85% (acceptable but watch — Little's Law tightening)
* GREEN: 40-70% (healthy)
* BLUE : <40% (under-loaded or wrong skills)
Team verdict:
HEALTHY — most green, no reds, low variance
SQUEEZED — majority amber, some red
OVERLOADED — >30% of team red
UNBALANCED — utilization variance >30 percentage points across team
Stdlib only.
"""
from __future__ import annotations
import argparse
import json
import statistics
import sys
from dataclasses import dataclass, field
from enum import Enum
from pathlib import Path
from typing import Any
class Light(str, Enum):
GREEN = "GREEN"
AMBER = "AMBER"
RED = "RED"
BLUE = "BLUE"
class TeamVerdict(str, Enum):
HEALTHY = "HEALTHY"
SQUEEZED = "SQUEEZED"
OVERLOADED = "OVERLOADED"
UNBALANCED = "UNBALANCED"
@dataclass
class Member:
name: str
role: str
utilization_pct: float
handles_count: int
hours_billable: float
hours_capacity: float
@dataclass
class MemberAssessment:
name: str
role: str
utilization_pct: float
light: Light
notes: list[str] = field(default_factory=list)
@dataclass
class TeamReport:
verdict: TeamVerdict
member_assessments: list[MemberAssessment]
mean_util: float
median_util: float
stdev_util: float
spread_pct_points: float
counts: dict[str, int]
headline: str
recommendations: list[str]
def classify(member: Member) -> MemberAssessment:
u = member.utilization_pct
notes: list[str] = []
if u >= 85:
light = Light.RED
notes.append("Throughput collapse risk per queueing theory (>85% sustained).")
elif u >= 70:
light = Light.AMBER
notes.append("Within tolerable band but no surge capacity.")
elif u >= 40:
light = Light.GREEN
else:
light = Light.BLUE
notes.append("Under-loaded — verify scope or reassign work.")
# Cross-check billable vs capacity hours — if claimed utilization
# disagrees with hours math by >10 points, flag.
if member.hours_capacity > 0:
computed = member.hours_billable / member.hours_capacity * 100
if abs(computed - u) > 10:
notes.append(
f"Reported util ({u:.0f}%) disagrees with hours math "
f"({computed:.0f}%). Reconcile time tracking."
)
return MemberAssessment(
name=member.name,
role=member.role,
utilization_pct=u,
light=light,
notes=notes,
)
def assess_team(members: list[Member]) -> TeamReport:
if not members:
raise ValueError("No team members in input.")
assessments = [classify(m) for m in members]
utils = [m.utilization_pct for m in members]
mean_u = statistics.fmean(utils)
median_u = statistics.median(utils)
stdev_u = statistics.pstdev(utils) if len(utils) > 1 else 0.0
spread = max(utils) - min(utils)
counts = {
"RED": sum(1 for a in assessments if a.light == Light.RED),
"AMBER": sum(1 for a in assessments if a.light == Light.AMBER),
"GREEN": sum(1 for a in assessments if a.light == Light.GREEN),
"BLUE": sum(1 for a in assessments if a.light == Light.BLUE),
}
n = len(members)
# Verdict logic — order matters
if spread > 30:
verdict = TeamVerdict.UNBALANCED
headline = (f"Load spread of {spread:.0f} percentage points across team — "
f"some are red while others are blue.")
elif counts["RED"] / n > 0.30:
verdict = TeamVerdict.OVERLOADED
headline = (f"{counts['RED']} of {n} members in RED zone. Throughput "
f"collapse risk.")
elif counts["AMBER"] / n >= 0.50 or counts["RED"] >= 1:
verdict = TeamVerdict.SQUEEZED
headline = f"Team running hot — {counts['AMBER']} amber, {counts['RED']} red."
else:
verdict = TeamVerdict.HEALTHY
headline = f"Team utilization healthy: mean {mean_u:.0f}%, spread {spread:.0f}pp."
recs: list[str] = []
if verdict == TeamVerdict.UNBALANCED:
recs.append("Rebalance load — investigate whether reds need different skills, "
"specialization, or just more hands at their queue.")
if verdict == TeamVerdict.OVERLOADED:
recs.append("Stop adding scope. Hire or shed work BEFORE attempting "
"process improvements (Goldratt: subordinate to the constraint).")
if verdict == TeamVerdict.SQUEEZED:
recs.append("Plan to hire next quarter. Re-test in 30 days; squeeze tends "
"to become overload during seasonal peaks.")
if counts["BLUE"] > 0:
recs.append(f"{counts['BLUE']} member(s) under-loaded — check whether "
f"work is reaching them or whether scope/skill needs adjustment.")
if not recs:
recs.append("Maintain current sizing; revisit at next quarterly planning cycle.")
return TeamReport(
verdict=verdict,
member_assessments=assessments,
mean_util=round(mean_u, 1),
median_util=round(median_u, 1),
stdev_util=round(stdev_u, 1),
spread_pct_points=round(spread, 1),
counts=counts,
headline=headline,
recommendations=recs,
)
def to_markdown(r: TeamReport) -> str:
lines = [
"# Utilization Analysis",
"",
f"**Verdict:** {r.verdict.value}",
"",
f"**Headline:** {r.headline}",
"",
"## Team Stats",
f"- Mean utilization: {r.mean_util}%",
f"- Median utilization: {r.median_util}%",
f"- Stdev: {r.stdev_util}pp",
f"- Spread (max - min): {r.spread_pct_points}pp",
f"- Counts: RED {r.counts['RED']} / AMBER {r.counts['AMBER']} / "
f"GREEN {r.counts['GREEN']} / BLUE {r.counts['BLUE']}",
"",
"## Member Detail",
"",
"| Name | Role | Utilization | Light | Notes |",
"|---|---|---|---|---|",
]
for a in r.member_assessments:
notes_str = "; ".join(a.notes) if a.notes else "—"
lines.append(
f"| {a.name} | {a.role} | {a.utilization_pct:.0f}% | {a.light.value} | {notes_str} |"
)
lines.extend(["", "## Recommendations"])
for rec in r.recommendations:
lines.append(f"- {rec}")
lines.extend([
"",
"## Canon",
"- Reinertsen, *Principles of Product Development Flow*, principle 7.",
"- Little (1961), *A Proof for the Queuing Formula L = λW*.",
"- Goldratt, *The Goal* — bottleneck subordination.",
])
return "\n".join(lines)
def to_dict(r: TeamReport) -> dict[str, Any]:
return {
"verdict": r.verdict.value,
"headline": r.headline,
"stats": {
"mean_util": r.mean_util,
"median_util": r.median_util,
"stdev_util": r.stdev_util,
"spread_pct_points": r.spread_pct_points,
"counts": r.counts,
},
"members": [
{
"name": a.name,
"role": a.role,
"utilization_pct": a.utilization_pct,
"light": a.light.value,
"notes": a.notes,
}
for a in r.member_assessments
],
"recommendations": r.recommendations,
}
SAMPLE_INPUT: dict[str, Any] = {
"team_members": [
{"name": "Alice", "role": "T1 Support", "utilization_pct": 92,
"handles_count": 48, "hours_billable": 7.4, "hours_capacity": 8},
{"name": "Bob", "role": "T1 Support", "utilization_pct": 88,
"handles_count": 42, "hours_billable": 7.0, "hours_capacity": 8},
{"name": "Carol", "role": "T1 Support", "utilization_pct": 72,
"handles_count": 36, "hours_billable": 5.8, "hours_capacity": 8},
{"name": "Dan", "role": "T2 Support", "utilization_pct": 65,
"handles_count": 18, "hours_billable": 5.2, "hours_capacity": 8},
{"name": "Eve", "role": "T2 Support", "utilization_pct": 35,
"handles_count": 8, "hours_billable": 2.8, "hours_capacity": 8},
]
}
def parse_members(raw: dict[str, Any]) -> list[Member]:
out: list[Member] = []
for m in raw["team_members"]:
out.append(Member(
name=m["name"],
role=m.get("role", "unspecified"),
utilization_pct=float(m["utilization_pct"]),
handles_count=int(m.get("handles_count", 0)),
hours_billable=float(m.get("hours_billable", 0)),
hours_capacity=float(m.get("hours_capacity", 0)),
))
return out
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(
description="Per-member + team-level utilization traffic-light analyzer.",
)
p.add_argument("--input", type=Path, help="Path to JSON input file.")
p.add_argument(
"--output", choices=["markdown", "json"], default="markdown",
help="Output format.",
)
p.add_argument("--sample", action="store_true",
help="Run on built-in sample input.")
args = p.parse_args(argv)
if args.sample:
raw = SAMPLE_INPUT
elif args.input:
raw = json.loads(args.input.read_text())
else:
p.error("Provide --input or --sample.")
return 2
try:
members = parse_members(raw)
except (KeyError, ValueError) as e:
print(f"ERROR parsing input: {e}", file=sys.stderr)
return 2
report = assess_team(members)
if args.output == "json":
print(json.dumps(to_dict(report), indent=2))
else:
print(to_markdown(report))
return 0
if __name__ == "__main__":
sys.exit(main())
Thu thập và sắp xếp các ý tưởng rời rạc thành hệ thống có cấu trúc, dễ hành động, không mất thông tin.
---
name: capture
description: "Captures and organizes chaotic brain dumps into a structured, actionable system with zero information loss. Use this skill whenever the user says 'capture this', 'brain dump', 'let me dump some ideas', 'I've got a bunch of thoughts', 'here's everything on my mind', 'idea dump', 'let me get this out of my head', 'I need to organize my thoughts', 'here's what I'm thinking', or any variation where someone is unloading a messy stream of ideas, tasks, thoughts, and plans wanting them turned into something coherent. Also trigger when the user pastes or dictates a long, unstructured block of mixed ideas — even without the exact phrase — the intent is the same. Fast-to-action by design: no upfront intake. Output is four sections (Projects/Ideas, Tasks, Connections, How I Can Help) ending with a directive question. Asks at most one mid-organization clarifying question when a single item is genuinely ambiguous between task and project."
license: MIT
metadata:
source_spec: "megaprompts/05-capture-megaprompt.md"
build_pattern: "Path B (direct conversion)"
version: 1.0.0
---
# Capture — Brain-Dump Organizer
A fast-to-action skill for transforming unstructured streams of mixed thoughts, tasks, and ideas into a clean four-section actionable system with zero information loss.
## Invocation Triggers
**Explicit phrases** (any of):
- "brain dump"
- "capture this"
- "let me dump some ideas"
- "I've got a bunch of thoughts"
- "here's everything on my mind"
- "idea dump"
- "let me just get this out of my head"
- "I need to organize my thoughts"
- "here's what I'm thinking"
**Implicit signals** (no phrase, but the intent is unmistakable):
- User pastes or dictates a long unstructured block of mixed ideas, tasks, plans
- Multiple unrelated thoughts in one message without organizing framing
- A wall of bullet-y text covering 3+ unrelated topics
When you detect an implicit trigger, run the skill. Do NOT ask "do you want me to organize this?" first — the dump itself IS the request.
## Operating Principles (All Five Apply Always)
1. **Capture everything.** Zero loss. Trivial items go in; the user prunes later. Never silently drop something because it "seemed unimportant".
2. **Preserve voice.** If the user said "build something crazy with AI", do NOT restate as "Explore innovative AI-driven solutions." Keep the energy and the casual register. See `references/voice_preservation.md` for concrete anti-patterns.
3. **Match output complexity to input.** A 5-task dump does NOT get forced into 4 elaborate sections. See `references/complexity_matching.md` and the Compressed Output Pattern below.
4. **Be honest about ambiguity.** If you're unsure what something means, flag it. Don't guess silently.
5. **No action without approval.** The ONLY immediate action is the organization itself. Every offer in Section 4 waits for the user's explicit pick.
## Grill-Me Mid-Organization Clarifier
Capture is fast-to-action by design. **No upfront intake.** The dump is enough — start organizing immediately.
The grill-me discipline applies as a **single mid-organization clarifying question**, asked **only when** one item in the dump is genuinely ambiguous between *task* and *project*, AND the misclassification would meaningfully change the output:
> **Quick clarification — one item in your dump could go either way. Is [X] a one-shot task or a multi-step project?**
>
> *Why I'm asking:* If I guess wrong on a borderline item I either bury a project as a task or inflate a task into a project that doesn't need the structure. One question per dump prevents that.
**Stop condition:** Max 1 clarifying question per dump. After the answer (or if no clarification was needed), deliver the four (or compressed) sections.
If the dump is unambiguous, skip the clarifier entirely.
**Anti-pattern (do not do this):** asking 3 clarifying questions up front. That breaks the dump-and-organize flow that makes capture useful.
## Section 1: Projects & Ideas
Cluster related items into themed projects when natural clustering exists. This section also holds:
- Standalone creative sparks
- Half-formed concepts
- "What if" thoughts
- Embedded decisions (`Decide: X or Y`) and open questions (`Q: ...`) — kept WITHIN the relevant project, NOT extracted into a separate top-level category
**Format per project:**
```
### {Project name in user's voice}
- {component / sub-idea}
- {component}
- Q: {open question this project needs answered}
- Decide: {decision this project requires}
```
Use the user's words for the project name. If the user wrote "ai dating app for ferrets", do NOT rename it to "AI-Powered Pet Companion Platform".
## Section 2: Tasks
Flat, scannable, action-oriented. Includes:
- Explicit todos
- Decisions framed as `Decide: ...`
- Open questions framed as `Resolve: ...`
If a task belongs to a project from Section 1, append `[Project: X]` to link it — but don't repeat the project's context.
**Format:**
```
- {task in imperative voice} [Project: X if related]
- Decide: {decision} [Project: X if related]
- Resolve: {open question}
- ...
```
## Section 3: Connections
This is where the skill earns its keep — and where **fabrication is forbidden**.
**Workflow:**
1. **Inventory the workspace** — Glob for filename patterns matching dump keywords, Grep for content matches, read the top-level directory structure. Use `scripts/workspace_inventory.py` to do this deterministically.
2. **Match dump items to existing content** — files / folders relating to dumped items, prior thinking in documents, in-progress projects with overlap.
3. **Surface dependencies within the dump** — items that affect each other, themes, ordering implications.
4. **Be honest about inaccessibility** — if you can't inspect the workspace (no filesystem available, MCP not connected), say so explicitly. Do NOT make up plausible-sounding connections.
**Hard rule:** NEVER fabricate connections. Only surface ones actually found by Glob/Grep/Read. If no real connections exist:
> **Connections:** No connections found — workspace inventory clean.
If the workspace is inaccessible:
> **Connections:** No workspace accessible from here. If you're running this from Claude Code or have a project with files attached, I can fill this in. Want to share where this work lives?
See `references/workspace_detection.md` for the per-context detection-tactic catalog.
## Section 4: How I Can Help
**Concrete offers, not abstract possibilities.** Every offer specifies what would be produced AND where it would go.
| ✅ Right pattern | ❌ Anti-pattern |
|---|---|
| "I can research Consensus MCP integration patterns and give you 3 options. Output: `docs/consensus-options.md`." | "You might want to look into integration approaches." |
| "I can draft the Q3 launch plan as a 1-pager. Output: chat reply, then `docs/q3-launch.md` if you want it filed." | "Maybe think about Q3 planning." |
| "I can scaffold the new auth module with the existing pattern from `src/users/`. Output: 4 files in `src/auth/`." | "We could explore auth options." |
End with the directive question:
> **Which of these should I tackle?**
## Compressed Output Pattern
When the dump has **5 or fewer items** and items are **unrelated** (no natural clustering), drop the 4-section format and use compressed:
```
## What I heard
- {item}
- {item}
- {item}
- ...
## How I can help
- {concrete offer with what + where}
- {concrete offer with what + where}
Which should I tackle?
```
The trigger is the `complexity_estimator.py` recommendation OR your judgment when no clusters exist. See `references/complexity_matching.md` for worked examples of when each format applies.
## Workspace Detection Strategy
| Context | Detection method |
|---|---|
| Claude Code CLI | Glob for files matching dump keywords; Grep for content matches; read top-level structure. Use `scripts/workspace_inventory.py`. |
| Claude.ai with project | Check project knowledge files for thematic overlap. List file titles; surface matches by keyword. |
| Connected tools (Notion, Drive, etc.) | Search via MCP if available. |
| No accessible workspace | State the limitation explicitly; ask user about their setup; do NOT fabricate. |
## Approval Gate
After the four (or compressed) sections are delivered:
- **Wait for the user's explicit pick** before doing anything else.
- If the user says "go" without picking a specific offer: honor it, but explicitly note any items you weren't 100% sure about so they can correct.
- The organization itself is the only auto-action. Every Section 4 offer requires green light.
## Error Handling
| Situation | Behavior |
|---|---|
| Workspace inaccessible | State this; skip Section 3 or surface "no workspace accessible" + ask about setup |
| Dump is very short (3-5 items) | Use compressed output; don't force 4 sections |
| Items are highly ambiguous | Flag in output, ask up to 1 clarifier (or skip clarifier and surface ambiguity in delivery) |
| Dump contains sensitive info | Acknowledge but don't echo verbatim if user asks for organization without quoting |
| Conflicting items in the dump | Surface the conflict in Section 1 or 3 explicitly (`Conflict: X says A, Y says B`) |
| User says "go" before approval | Honor it, but explicitly note items you weren't sure about |
## Tooling
| Script | Role |
|---|---|
| `scripts/workspace_inventory.py` | Glob+Grep helper for Section 3. `python workspace_inventory.py --root . --keywords "k1,k2"` returns matches by keyword + folder structure. |
| `scripts/dump_classifier.py` | Regex-classifies each dump line into `task` / `decision` / `question` / `idea` / `project-component`. Heuristic — override with judgment. |
| `scripts/complexity_estimator.py` | Counts items, detects clustering signal, recommends `format=full` or `format=compressed`. |
## References
- `references/workspace_detection.md` — context-specific detection tactics (CLI / web / MCP / inaccessible)
- `references/voice_preservation.md` — corporate-speak anti-patterns with concrete examples
- `references/complexity_matching.md` — compressed vs full output, worked examples
## Anti-Patterns To Reject
- Fabricating workspace connections that weren't actually Glob/Grep-verified
- Dropping items deemed "trivial" — capture everything, let the user prune
- Corporate-ifying the user's casual language
- Forcing 4-section structure when input is small (5 simple tasks doesn't need it)
- Acting on Section-4 offers immediately without approval
- Splitting decisions/questions into a separate top-level category instead of embedding them in the relevant project
- Vague Section-4 offers ("you might want to consider…")
- Asking 3+ clarifying questions up front (breaks fast-to-action)
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/05-capture-megaprompt.md`](../../../../megaprompts/05-capture-megaprompt.md)
**Build pattern:** Path B (direct conversion). Re-grill with `/cs:grill-with-docs` if drift between spec and implementation surfaces.
FILE:references/complexity_matching.md
# Complexity Matching — Compressed vs Full 4-Section Output
This reference answers exactly one decision: **when does capture use the full 4-section format vs the compressed format, and what does each look like in practice?**
Pair with `scripts/complexity_estimator.py` for the deterministic recommendation.
## The Core Rule
> **Match output complexity to input complexity.**
A 30-item dump with natural clusters needs the full 4-section structure to be useful. A 5-item dump of unrelated todos drowns in that structure — the format becomes ceremony, not signal. Force-fitting structure on a small dump makes the skill feel bureaucratic.
## When to Use Each Format
| Signal | Recommended format |
|---|---|
| 8+ items AND natural clustering exists (3+ items share a theme) | Full 4-section |
| 8+ items but NO clustering (all unrelated todos) | Compressed (with explicit "no clusters" note) |
| 5–7 items, mixed kinds, some clustering | Either — judgment call. Lean compressed unless clusters are strong. |
| ≤5 items, unrelated | Compressed |
| ≤5 items but all related to one project | Compressed with single project header |
| Workspace inaccessible AND ≤5 items | Compressed with no Section 3 (still note "no workspace accessible") |
`complexity_estimator.py` returns `format=full` or `format=compressed` based on item count + clustering signal. Use it as the seed; override with judgment when context warrants.
## Format A: Full 4-Section
Use for substantive dumps with real structure. Roughly:
```
## Projects & Ideas
### {Project A in user's voice}
- {component}
- {component}
- Q: {open question}
- Decide: {decision}
### {Project B}
- ...
## Tasks
- {task} [Project: A]
- {task}
- Decide: {decision}
- Resolve: {open question}
## Connections
- {file/folder}: {real workspace match}
- (or) "No connections found — workspace inventory clean."
- (or) "No workspace accessible from here..."
## How I Can Help
- {concrete offer: what + where}
- {concrete offer: what + where}
**Which of these should I tackle?**
```
## Format B: Compressed
Use for small or unrelated dumps. Roughly:
```
## What I heard
- {item}
- {item}
- {item}
- Decide: {decision}
- Resolve: {open question}
## How I can help
- {concrete offer: what + where}
- {concrete offer: what + where}
Which should I tackle?
```
Notice what's gone in compressed:
- No project headers (nothing to cluster)
- No separate Tasks section (the list IS the tasks)
- No Connections section as a separate header (if there's a real connection, fold it inline as a parenthetical: `- fix the auth bug (relates to src/auth/login.ts:42)`)
- Same Section 4 discipline (concrete offers + directive question)
## Worked Examples
### Example 1: Full 4-section (12 items, 2 clusters)
**Dump:**
> "Ok so dump time. Q3 launch is approaching — need to nail down pricing, draft the launch email, brief Sarah on the marketing angle. Also ferret app keeps nagging me — should I talk to my cofounder about it, or is it just procrastination? Need to sketch the matching algorithm if it's serious. Also: fix the damn auth bug, rewrite the login form because it's ugly, write tests for both. Oh and I should probably do my Q3 OKRs before launch."
**Extracted items:** 11
**Clusters:** Q3 Launch (4 items), Ferret App (3 items), Auth (3 items), OKRs (1 item)
**Recommendation:** `format=full` (clear clustering, 11 items)
**Output:**
```
## Projects & Ideas
### Q3 Launch
- Nail down pricing
- Draft launch email
- Brief Sarah on the marketing angle
- Decide: Q3 OKRs (do before launch)
### Ferret App
- Talk to cofounder OR decide it's procrastination
- Sketch matching algorithm (if serious)
- Decide: serious project or shelf?
### Auth Fixes
- Fix the auth bug
- Rewrite the login form (it's ugly)
- Write tests for the bug fix + login rewrite
## Tasks
- Nail down Q3 pricing [Project: Q3 Launch]
- Draft Q3 launch email [Project: Q3 Launch]
- Brief Sarah on Q3 marketing angle [Project: Q3 Launch]
- Talk to cofounder about ferret app [Project: Ferret App]
- Sketch ferret matching algo [Project: Ferret App]
- Fix auth bug [Project: Auth Fixes]
- Rewrite login form [Project: Auth Fixes]
- Write tests for auth + login [Project: Auth Fixes]
- Do Q3 OKRs (before launch)
- Decide: ferret app — serious or shelf?
## Connections
(workspace inventory results go here — only verified matches)
## How I Can Help
- I can draft the Q3 launch email as a 1-pager. Output: chat reply.
- I can pull together a pricing-options matrix from your existing pricing docs. Output: `docs/q3-pricing-options.md`.
- I can scaffold the ferret matching algorithm sketch (input → match score). Output: `docs/ferret-algo-sketch.md`.
**Which of these should I tackle?**
```
### Example 2: Compressed (5 unrelated items)
**Dump:**
> "Quick: email Sarah, fix that test that's been flaky, decide between Postgres and Mongo for the new service, dentist appointment, finish reading the article on RAG."
**Extracted items:** 5
**Clusters:** none (all unrelated)
**Recommendation:** `format=compressed`
**Output:**
```
## What I heard
- Email Sarah
- Fix the flaky test
- Decide: Postgres or Mongo for the new service
- Dentist appointment
- Finish reading the RAG article
## How I can help
- I can draft the email to Sarah if you tell me what it's about. Output: chat reply.
- I can investigate the flaky test (likely racing condition; happy to reproduce). Output: chat analysis + fix PR if you want.
- I can pull a Postgres-vs-Mongo decision matrix scoped to your service requirements. Output: chat reply or `docs/db-decision.md`.
Which should I tackle?
```
Notice the dentist appointment + RAG article aren't in "How I can help" — there's nothing meaningful to offer on those. That's correct: don't pad Section 4 with offers for things that don't need help.
### Example 3: Compressed but workspace-aware (small + connection inline)
**Dump:**
> "auth bug is back, need to look at the login flow again, also I want to add 2FA finally."
**Extracted items:** 3
**Clusters:** all auth-related but only 3 items
**Recommendation:** `format=compressed`
**Output:**
```
## What I heard (Auth)
- Auth bug is back (relates to `src/auth/login.ts` based on workspace inventory)
- Look at login flow again
- Add 2FA
## How I can help
- I can reproduce the auth bug from the symptoms — what's the failure mode? Output: chat repro + fix.
- I can sketch a 2FA implementation matching your existing auth pattern (TOTP via the same provider you use). Output: `docs/2fa-sketch.md`.
Which should I tackle?
```
Notice the workspace connection got folded inline as a parenthetical instead of a separate Section 3 header. That's the compressed-with-context pattern.
## Operational Checklist
Before delivering output:
- [ ] Run `complexity_estimator.py` (or apply the Signal table above)
- [ ] If `format=compressed`, do NOT force the 4-section format
- [ ] If `format=full`, ensure the clusters are real (3+ items per cluster) — don't invent clusters to fill the format
- [ ] Either way, Section 4 ("How I can help") MUST have concrete offers with what + where
- [ ] Either way, end with the directive question
## Why This Matters
A skill that returns the same format regardless of input is a template, not a skill. The reason capture is useful is that it adapts to the dump's actual shape. When a 5-item list comes back wrapped in 4 elaborate empty-feeling sections, the user learns to distrust the skill. When a 30-item dump comes back as a flat compressed list, the user learns the skill can't actually handle complexity.
Match the output to the input, every time.
FILE:references/voice_preservation.md
# Voice Preservation — Anti-Corporate-Speak Discipline
This reference answers exactly one decision: **what does it mean to "preserve the user's voice" in capture output, and what concrete patterns must be avoided?**
## The Core Rule
If the user said it casually, restate it casually. If the user said it crudely, restate it crudely (within the user's own register). Capture is for THEM, not for an imagined corporate audience reading their notes later.
> **Restating someone's casual idea in corporate language is a tax. It feels formal but it loses the energy that made the idea worth capturing.**
## Concrete Anti-Patterns (Side-by-Side)
| User said | ❌ Corporate-ified (anti-pattern) | ✅ Voice-preserved |
|---|---|---|
| "build something crazy with AI" | "Explore innovative AI-driven solutions" | "Build something crazy with AI" |
| "the dating app idea but for ferrets" | "Pet-companion matching platform leveraging social-graph principles" | "Dating app for ferrets" |
| "figure out the damn pricing already" | "Conduct comprehensive pricing strategy analysis" | "Figure out pricing (final answer)" |
| "fuck around with Consensus MCP" | "Investigate Consensus MCP integration opportunities" | "Try out Consensus MCP" |
| "make the landing page not suck" | "Optimize landing page user experience metrics" | "Make the landing page not suck" |
| "talk to Sarah about the thing" | "Schedule alignment discussion with Sarah re: outstanding initiative" | "Talk to Sarah about the thing" |
| "I'm tired of debugging this" | "Investigate root causes of recurring debugging friction" | "Tired of debugging this — find the root cause" |
## What Counts As "Voice"
- **Register** — formal vs casual, dry vs energetic, ironic vs earnest
- **Vocabulary** — the user's exact noun choices for things ("ferrets", "thing", "Sarah" — not "pets", "initiative", "stakeholder")
- **Cadence** — short choppy phrases stay short; long flowing thoughts stay flowing
- **Profanity / slang** — preserve as-is; don't sanitize
- **Self-talk markers** — "ugh", "actually", "wait", "ok so" — these signal genuine thinking and belong in the captured form
## What's Allowed to Change
- **Punctuation cleanup** — adding a period, fixing typos
- **Imperative reframing** for the Tasks section — "I should email Sarah" → "Email Sarah" (one-word edit, voice preserved)
- **Light disambiguation** — if "the thing" is genuinely confusing in context, note it but ask to clarify (don't replace it silently)
## What's Never Allowed
- Replacing user nouns with "platform" / "solution" / "initiative" / "framework"
- Verbing nouns: "let's research" → "let's conduct research"
- Adding qualifiers the user didn't say: "explore", "leverage", "deep dive into"
- "Action-itemizing" everything: "talk to Sarah" → "Establish communication touchpoint with Sarah"
- Removing emotion: "I'm pissed about X" → "There is a concern regarding X"
- Bullet-point fluff: "Implement", "Establish", "Facilitate" prefixes added for no reason
## Cluster Naming
When clustering items into projects (Section 1), the project name **MUST** use the user's words. Examples:
| Items in cluster | ❌ Anti-pattern name | ✅ Voice-preserved name |
|---|---|---|
| "ferret app", "ferret features", "ferret marketing" | "Pet Companion Platform" | "Ferret App" |
| "Q3 launch", "Q3 pricing", "Q3 emails" | "Q3 Go-to-Market Initiative" | "Q3 Launch" |
| "fix the auth bug", "auth tests", "rewrite login" | "Authentication System Modernization" | "Auth fixes" |
If the user used multiple terms for the same cluster, pick the one they used most or most colloquially.
## Operating Test
Before writing each line, ask:
> Would the user *recognize* this as something they'd say?
If no, you've drifted. Rewrite to match their register.
## Why This Matters
Voice preservation isn't aesthetic — it's functional. Two reasons:
1. **Recognition.** The user reads their own captured dump back in 2 days and needs to instantly recognize "yes, that's me, that's what I meant." Corporate restatement breaks recognition. The user thinks "wait, did I actually say that?" and starts second-guessing the rest of the output.
2. **Energy.** A dump captured in voice retains the *why* behind each item — the frustration, the excitement, the half-formed hope. Corporate restatement strips the why and leaves a list of generic action items that no one is excited to act on.
Capture is the user's brain on paper. Don't translate it into a stranger's brain.
FILE:references/workspace_detection.md
# Workspace Detection Tactics
This reference answers exactly one decision: **how does the capture skill verify Section 3 connections without fabricating them, across the four contexts the skill might run in?**
Pair with `scripts/workspace_inventory.py` for the deterministic Glob+Grep implementation.
## The Core Rule
Section 3 ("Connections") earns the skill its keep. It also breaks the skill faster than anything else if it lies. The rule:
> **Only surface connections that were actually verified by Glob, Grep, Read, or an equivalent retrieval call this turn.**
If you can't verify, you say "no workspace accessible" or "no connections found" — never invent something plausible-sounding.
## Context 1: Claude Code CLI (filesystem-native)
**Tools available:** `Glob`, `Grep`, `Read`, `Bash`.
**Tactics, in order:**
1. **Extract keywords from the dump.** Pull domain nouns, project names, file-format hints (`.md`, `.py`, `auth`, `consensus`, `pricing`).
2. **Glob for filename matches.** `Glob("**/*{keyword}*")` for each keyword. Limit to top-N matches per keyword to avoid noise.
3. **Grep for content matches.** `Grep("{keyword}")` constrained to source extensions.
4. **Read the top-level structure.** `Bash("ls -la")` and `Bash("find . -maxdepth 2 -type d | head -30")` to surface relevant folders.
5. **Stitch the matches into Section 3 entries.** Each entry: `- {file or folder}: {how it relates to dump item N, with evidence}`.
**Example output:**
```
## Connections
- `engineering/grill-me/` — relates to your "build a grill skill" dump item (folder exists, has plugin.json + SKILL.md). Likely the template you'd want to mirror.
- `megaprompts/05-capture-megaprompt.md` — relates to your "convert capture spec to skill" item. The spec file is here.
- `documentation/implementation/` — empty directory, but the location for the implementation plan you mentioned.
```
**What NOT to do:**
- "There's probably a config for that somewhere" — speculation, no verification.
- "Your project likely has an auth module" — guess, no Glob.
- "I see you might have considered X before" — projection, no Grep.
## Context 2: Claude.ai with project knowledge
**Tools available:** Project-knowledge file list, file content reads.
**Tactics:**
1. **List the project knowledge files** — at the start of the run, get the file inventory.
2. **Match by title keyword** — for each dump keyword, find files whose titles contain it.
3. **Open the top matches** and check if the content is actually related (not just title coincidence).
4. **Surface only the verified matches** in Section 3.
**What NOT to do:**
- Cite a file you didn't open — title match alone is not enough.
- Claim a file says X without quoting evidence.
## Context 3: Connected tools (Notion, Drive, GitHub, Slack via MCP)
**Tools available:** Whatever MCP tools the harness has registered for the user's connected services.
**Tactics:**
1. **Check tool availability first** — list the MCP tools surfaced for this session. If no Notion/Drive/GitHub MCP is registered, skip this context.
2. **Search via MCP** — use the search tool for each tool with dump keywords.
3. **Surface verified hits** with the link / ID returned by the tool.
**What NOT to do:**
- Reference a Notion page that wasn't returned by the search.
- Cite a GitHub issue number without confirming via the GitHub MCP.
## Context 4: No accessible workspace
**Signals you're in this context:**
- No filesystem tools loaded
- No project knowledge attached
- No workspace MCPs registered
- `workspace_inventory.py` returns empty + you can't verify any other way
**Required behavior:**
State the limitation explicitly. Ask about the user's setup. Do NOT fabricate connections.
**Template output:**
```
## Connections
No workspace accessible from here, so this section is empty. If you're running
this from Claude Code or have a project with files attached, I can fill it in.
Want to share where this work lives — a repo path, a Notion workspace, an
attached project? I can re-run the connections pass with that context.
```
## Operational Checklist (Per Run)
- [ ] Extract dump keywords (domain nouns, project names, format hints)
- [ ] Determine context (CLI / web project / MCP-connected / inaccessible)
- [ ] Run the context-appropriate tactics; never skip verification
- [ ] If context is "inaccessible", say so explicitly + ask about setup
- [ ] Each Section 3 entry must cite the evidence (filename matched, search hit, etc.)
- [ ] Zero entries with phrasing like "probably", "likely", "you might have" — those are speculation, not connections
## Why This Matters
The single fastest way to lose user trust in capture is to surface a fabricated connection. Once the user catches one — "wait, that file doesn't exist" — they stop trusting the entire output, including the items that were correct. Verification is cheap; fabrication is expensive.
FILE:scripts/complexity_estimator.py
#!/usr/bin/env python3
"""complexity_estimator.py — Recommend full-4-section vs compressed output.
Stdlib-only. Counts non-empty items in a dump, detects clustering signal
(repeated keywords across items), and recommends one of:
format=full → use the full Projects/Tasks/Connections/How-I-Can-Help
4-section format (8+ items AND clustering signal)
format=compressed → use the compressed What-I-heard / How-I-can-help
format (≤5 items OR no clustering signal)
The recommendation is HEURISTIC. The capture skill applies judgment on top
based on full dump context. Use this as the seed.
NO LLM CALLS. Pure word counting + frequency analysis.
Usage:
python complexity_estimator.py path/to/dump.txt
python complexity_estimator.py path/to/dump.txt --output json
python complexity_estimator.py --sample
"""
import argparse
import json
import re
import sys
from collections import Counter
from pathlib import Path
from typing import Any, Dict, List
SAMPLE_DUMP_LARGE = """Ok dump time.
Q3 launch needs pricing nailed down.
Draft the Q3 launch email.
Brief Sarah on the Q3 marketing angle.
Decide: launch July 15 or August 1?
Ferret app idea keeps nagging me.
Should I talk to my cofounder about ferret app?
Sketch the ferret matching algorithm if serious.
Decide: ferret app serious or shelf?
Fix the auth bug.
Rewrite the login form (ugly).
Write tests for auth + login.
Add 2fa module to auth.
Do my Q3 OKRs before launch.
"""
SAMPLE_DUMP_SMALL = """Email Sarah.
Fix the flaky test.
Decide: Postgres or Mongo for new service.
Dentist appointment.
Finish reading the RAG article.
"""
# Stop-words to exclude from clustering detection
STOP_WORDS = {
"a", "an", "the", "is", "are", "was", "were", "be", "been", "being",
"to", "of", "in", "on", "at", "for", "with", "by", "from", "up", "down",
"and", "or", "but", "if", "then", "else", "so", "as",
"i", "you", "he", "she", "it", "we", "they", "me", "my", "your", "our",
"this", "that", "these", "those", "do", "did", "have", "has", "had",
"will", "would", "should", "could", "can", "may", "might",
"what", "when", "where", "why", "how", "who", "which",
"go", "get", "got", "make", "made", "let", "let's", "yeah", "ok", "well",
"just", "really", "very", "much", "more", "most", "some", "any",
"not", "no", "yes", "now", "before", "after",
}
def extract_items(text: str) -> List[str]:
"""Return non-empty stripped lines (each line = one item)."""
return [line.strip() for line in text.splitlines() if line.strip()]
def extract_keywords(items: List[str]) -> List[str]:
"""Return all alphabetic tokens >= 3 chars, lowercased, stop-words removed."""
tokens: List[str] = []
for it in items:
for tok in re.findall(r"\b[A-Za-z][a-zA-Z]{2,}\b", it):
t = tok.lower()
if t in STOP_WORDS:
continue
tokens.append(t)
return tokens
def detect_clusters(items: List[str], min_cluster_size: int) -> List[Dict[str, Any]]:
"""A 'cluster' is a keyword that appears in min_cluster_size+ different items."""
keyword_to_item_indices: Dict[str, List[int]] = {}
for i, it in enumerate(items):
seen_in_item: set = set()
for tok in re.findall(r"\b[A-Za-z][a-zA-Z]{2,}\b", it):
t = tok.lower()
if t in STOP_WORDS or t in seen_in_item:
continue
seen_in_item.add(t)
keyword_to_item_indices.setdefault(t, []).append(i)
clusters: List[Dict[str, Any]] = []
for kw, idxs in keyword_to_item_indices.items():
if len(idxs) >= min_cluster_size:
clusters.append({"keyword": kw, "item_indices": idxs, "size": len(idxs)})
clusters.sort(key=lambda c: (-c["size"], c["keyword"]))
return clusters
def estimate(text: str, min_cluster_size: int = 3) -> Dict[str, Any]:
items = extract_items(text)
item_count = len(items)
clusters = detect_clusters(items, min_cluster_size)
cluster_count = len(clusters)
# Decision logic per references/complexity_matching.md
if item_count >= 8 and cluster_count >= 1:
recommendation = "full"
rationale = f"{item_count} items with {cluster_count} cluster(s) of {min_cluster_size}+ → full 4-section format"
elif item_count >= 8 and cluster_count == 0:
recommendation = "compressed"
rationale = f"{item_count} items but no clustering signal → compressed (with 'no clusters' note)"
elif 5 <= item_count <= 7 and cluster_count >= 1:
recommendation = "full"
rationale = f"{item_count} items with clustering signal → judgment call, defaulting full"
elif 5 <= item_count <= 7 and cluster_count == 0:
recommendation = "compressed"
rationale = f"{item_count} items, no clustering → compressed"
elif item_count <= 5:
recommendation = "compressed"
rationale = f"{item_count} items (small dump) → compressed"
else:
recommendation = "compressed"
rationale = "fallback → compressed"
return {
"item_count": item_count,
"cluster_count": cluster_count,
"clusters": clusters[:5], # top 5 only in output for readability
"recommendation": recommendation,
"rationale": rationale,
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Item count: {result['item_count']}")
out.append(f"Cluster count: {result['cluster_count']}")
out.append(f"Recommendation: format={result['recommendation']}")
out.append(f"Rationale: {result['rationale']}")
if result["clusters"]:
out.append("")
out.append("Top clusters (keyword → items containing it):")
for c in result["clusters"]:
out.append(f" - '{c['keyword']}' in {c['size']} items: lines {c['item_indices']}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("path", nargs="?", help="Path to dump file")
parser.add_argument("--sample", choices=["large", "small"], help="Estimate the embedded sample dump (large or small)")
parser.add_argument("--min-cluster-size", type=int, default=3, help="Minimum items sharing a keyword to count as a cluster (default: 3)")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_DUMP_LARGE if args.sample == "large" else SAMPLE_DUMP_SMALL
elif args.path:
p = Path(args.path)
if not p.exists():
print(f"error: {args.path} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = estimate(text, args.min_cluster_size)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/dump_classifier.py
#!/usr/bin/env python3
"""dump_classifier.py — Heuristic classifier for brain-dump lines.
Stdlib-only. Reads a dump file (or stdin) and labels each line as one of:
- task (action-oriented, imperative or 'I should X')
- decision ('decide between X and Y', 'should we X or Y')
- question (ends in '?')
- idea (creative spark, 'what if', 'maybe we should X')
- project-component (sub-element of a larger project, often noun-phrased)
- context (preamble, framing, no actionable content)
The classifier is HEURISTIC. The capture skill uses these labels as a SEED for
its own structuring — it overrides based on dump-level context. Do not treat
the labels as authoritative.
NO LLM CALLS. Pure regex + line walking.
Usage:
python dump_classifier.py path/to/dump.txt
python dump_classifier.py path/to/dump.txt --output json
python dump_classifier.py --sample
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List, Tuple
# Pattern → label, ordered by precedence (first match wins per line).
# Each pattern is a compiled regex.
PATTERNS: List[Tuple[re.Pattern, str]] = [
(re.compile(r"\bdecide\s+(?:between|on|whether)\b", re.IGNORECASE), "decision"),
(re.compile(r"^\s*decide\s*:", re.IGNORECASE), "decision"),
(re.compile(r"\b(?:should we|do we|are we)\b.*\b(?:or|vs|versus)\b", re.IGNORECASE), "decision"),
(re.compile(r"\?\s*$"), "question"),
(re.compile(r"^\s*(?:resolve|q)\s*:", re.IGNORECASE), "question"),
(re.compile(r"\bwhat\s+if\b", re.IGNORECASE), "idea"),
(re.compile(r"\b(?:maybe|might|could)\s+(?:we|i)\s+(?:should\s+)?", re.IGNORECASE), "idea"),
(re.compile(r"\bidea\s*:", re.IGNORECASE), "idea"),
(re.compile(r"^\s*(?:i\s+(?:should|need\s+to|gotta|have\s+to))\b", re.IGNORECASE), "task"),
(re.compile(r"^\s*(?:fix|build|write|draft|email|send|talk to|finish|investigate|research|sketch|scaffold|deploy|push|merge|review|read|call|schedule)\b", re.IGNORECASE), "task"),
(re.compile(r"^\s*todo\s*:", re.IGNORECASE), "task"),
(re.compile(r"^\s*-\s*(?:fix|build|write|draft|email|send|talk to|finish|investigate)\b", re.IGNORECASE), "task"),
]
PROJECT_COMPONENT_HINTS = {
"module", "feature", "component", "endpoint", "page", "screen", "form",
"model", "schema", "migration", "test", "doc", "readme", "config",
}
SAMPLE_DUMP = """Ok dump time.
Q3 launch is approaching - need to nail down pricing.
Draft the launch email.
Brief Sarah on the marketing angle.
Decide: launch on July 15 or August 1?
Ferret app keeps nagging me. Should I talk to my cofounder about it?
What if it's actually a real business?
Sketch the matching algorithm if it's serious.
Fix the damn auth bug.
Rewrite the login form because it's ugly.
Write tests for both.
Auth: 2fa module.
Do my Q3 OKRs before launch.
"""
def classify_line(raw: str) -> str:
line = raw.strip()
if not line:
return "blank"
# Strip leading bullet markers for matching, but keep original for output
stripped = re.sub(r"^[-*+]\s+", "", line)
for pattern, label in PATTERNS:
if pattern.search(stripped):
return label
# Project-component heuristic: short noun-phrase containing a hint word
if len(stripped.split()) <= 6:
for hint in PROJECT_COMPONENT_HINTS:
if re.search(rf"\b{hint}s?\b", stripped, re.IGNORECASE):
return "project-component"
# Single-noun-phrase or short fragment without verb → likely context or component
if len(stripped.split()) <= 4 and not stripped.endswith("?"):
return "project-component"
# Default: treat as context (the skill folds this into project framing)
return "context"
def classify(text: str) -> Dict[str, Any]:
items: List[Dict[str, Any]] = []
counts: Dict[str, int] = {}
for line_no, raw in enumerate(text.splitlines(), start=1):
if not raw.strip():
continue
label = classify_line(raw)
if label == "blank":
continue
items.append({"line": line_no, "label": label, "text": raw.strip()})
counts[label] = counts.get(label, 0) + 1
return {"item_count": len(items), "by_label": counts, "items": items}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Dump classification ({result['item_count']} non-empty items)")
out.append("By label:")
for label, n in sorted(result["by_label"].items(), key=lambda kv: -kv[1]):
out.append(f" {label:<20s} {n}")
out.append("")
out.append("Per-line labels:")
for it in result["items"]:
out.append(f" L{it['line']:>3} {it['label']:<20s} {it['text'][:80]}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("path", nargs="?", help="Path to dump file (or omit for --sample)")
parser.add_argument("--sample", action="store_true", help="Classify the embedded sample dump")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_DUMP
elif args.path:
p = Path(args.path)
if not p.exists():
print(f"error: {args.path} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = classify(text)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/workspace_inventory.py
#!/usr/bin/env python3
"""workspace_inventory.py — Glob+Grep helper for capture's Section 3 (Connections).
Stdlib-only. Given a working directory + a list of dump-derived keywords,
returns a structured inventory that the capture skill can use to surface
real workspace connections (never fabricated).
What it returns:
1. Per-keyword filename matches (Glob-style)
2. Per-keyword content matches (line-grep across source files)
3. Top-level folder structure (max-depth 2)
What it does NOT do:
- Score relevance (capture skill applies judgment on top)
- Fabricate matches (only real Glob/Grep results)
- Make LLM calls
Usage:
python workspace_inventory.py --root . --keywords "auth,login,2fa"
python workspace_inventory.py --root . --keywords "skill,megaprompt" --output json
python workspace_inventory.py --sample
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List, Set
DEFAULT_SOURCE_EXTENSIONS = {
".py", ".ts", ".tsx", ".js", ".jsx", ".go", ".java", ".kt", ".rb",
".cs", ".rs", ".swift", ".php", ".scala", ".clj", ".ex", ".exs",
".md", ".mdx", ".rst", ".txt", ".json", ".yaml", ".yml", ".toml",
}
DEFAULT_EXCLUDE_DIRS = {
"node_modules", ".git", "dist", "build", "target",
".venv", "venv", "__pycache__", ".next", ".cache",
}
MAX_FILENAME_MATCHES_PER_KEYWORD = 20
MAX_CONTENT_MATCHES_PER_KEYWORD = 10
MAX_FOLDER_DEPTH = 2
SAMPLE_TREE: Dict[str, str] = {
"src/auth/login.ts": "// auth login flow handler\nexport function login() {}\n",
"src/auth/2fa.ts": "// 2fa stub\nexport function setup2FA() {}\n",
"src/users/profile.ts": "// user profile\nexport function getProfile() {}\n",
"docs/auth-bugs.md": "# Auth bugs\n\n- The login race condition is back.\n",
"tests/auth.test.ts": "// auth tests\ndescribe('auth', () => {});\n",
"README.md": "# project\n\nAuth + login + users.\n",
}
def collect_files(root: Path, source_extensions: Set[str], exclude_dirs: Set[str]) -> List[Path]:
found: List[Path] = []
for p in root.rglob("*"):
if p.is_dir():
continue
if any(part in exclude_dirs for part in p.parts):
continue
if p.suffix.lower() in source_extensions:
found.append(p)
return found
def filename_matches(files: List[Path], keyword: str) -> List[str]:
kw = keyword.lower()
out: List[str] = []
for p in files:
if kw in p.name.lower():
out.append(str(p))
if len(out) >= MAX_FILENAME_MATCHES_PER_KEYWORD:
break
return out
def content_matches(files: List[Path], keyword: str) -> List[Dict[str, Any]]:
pattern = re.compile(re.escape(keyword), re.IGNORECASE)
out: List[Dict[str, Any]] = []
for p in files:
try:
text = p.read_text(encoding="utf-8", errors="ignore")
except OSError:
continue
for line_no, line in enumerate(text.splitlines(), start=1):
if pattern.search(line):
out.append({
"file": str(p),
"line": line_no,
"snippet": line.strip()[:120],
})
if len(out) >= MAX_CONTENT_MATCHES_PER_KEYWORD:
return out
return out
def folder_structure(root: Path, max_depth: int, exclude_dirs: Set[str]) -> List[str]:
out: List[str] = []
root = root.resolve()
for p in root.rglob("*"):
if not p.is_dir():
continue
if any(part in exclude_dirs for part in p.parts):
continue
try:
rel = p.relative_to(root)
except ValueError:
continue
depth = len(rel.parts)
if 0 < depth <= max_depth:
out.append(str(rel))
return sorted(out)
def inventory(
root: Path,
keywords: List[str],
source_extensions: Set[str],
exclude_dirs: Set[str],
) -> Dict[str, Any]:
files = collect_files(root, source_extensions, exclude_dirs)
per_keyword: Dict[str, Dict[str, Any]] = {}
for kw in keywords:
per_keyword[kw] = {
"filename_matches": filename_matches(files, kw),
"content_matches": content_matches(files, kw),
}
return {
"root": str(root.resolve()),
"files_scanned": len(files),
"folder_structure": folder_structure(root, MAX_FOLDER_DEPTH, exclude_dirs),
"per_keyword": per_keyword,
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Workspace inventory for: {result['root']}")
out.append(f" Files scanned: {result['files_scanned']}")
out.append("")
out.append("Folder structure (max depth 2):")
for f in result["folder_structure"][:30]:
out.append(f" - {f}/")
if len(result["folder_structure"]) > 30:
out.append(f" ... + {len(result['folder_structure']) - 30} more")
out.append("")
for kw, hits in result["per_keyword"].items():
out.append(f"Keyword: '{kw}'")
if hits["filename_matches"]:
out.append(f" Filename matches ({len(hits['filename_matches'])}):")
for f in hits["filename_matches"]:
out.append(f" - {f}")
else:
out.append(" Filename matches: (none)")
if hits["content_matches"]:
out.append(f" Content matches ({len(hits['content_matches'])}):")
for m in hits["content_matches"]:
out.append(f" - {m['file']}:{m['line']} {m['snippet']}")
else:
out.append(" Content matches: (none)")
out.append("")
return "\n".join(out)
def run_sample(keywords: List[str]) -> Dict[str, Any]:
import tempfile
with tempfile.TemporaryDirectory() as td:
root = Path(td)
for rel, content in SAMPLE_TREE.items():
p = root / rel
p.parent.mkdir(parents=True, exist_ok=True)
p.write_text(content, encoding="utf-8")
return inventory(root, keywords, DEFAULT_SOURCE_EXTENSIONS, DEFAULT_EXCLUDE_DIRS)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--root", help="Root directory to inventory")
parser.add_argument("--keywords", help="Comma-separated keywords to search for")
parser.add_argument("--extensions", help="Comma-separated source extensions (default: common)")
parser.add_argument("--sample", action="store_true", help="Inventory the embedded sample tree")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
sample_keywords = ["auth", "login", "2fa"] if not args.keywords else [k.strip() for k in args.keywords.split(",") if k.strip()]
result = run_sample(sample_keywords)
elif args.root and args.keywords:
root = Path(args.root)
if not root.exists():
print(f"error: {args.root} not found", file=sys.stderr)
return 2
kws = [k.strip() for k in args.keywords.split(",") if k.strip()]
if not kws:
print("error: --keywords must list at least one keyword", file=sys.stderr)
return 2
if args.extensions:
exts = {e.strip() if e.strip().startswith(".") else "." + e.strip() for e in args.extensions.split(",")}
else:
exts = DEFAULT_SOURCE_EXTENSIONS
result = inventory(root, kws, exts, DEFAULT_EXCLUDE_DIRS)
else:
parser.print_help()
return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Tư vấn Chief AI Officer: chọn API, tinh chỉnh hay tự xây, phân loại rủi ro AI, kinh tế chi phí AI và tổ chức nhóm AI.
---
name: "chief-ai-officer-advisor"
description: "Chief AI Officer advisory for startups: model build-vs-buy decisions (API vs fine-tune vs in-house), AI risk classification under EU AI Act + US state patchwork, AI cost economics (API-to-self-hosted breakeven), and AI team org evolution. Use when deciding whether to call an API or fine-tune, classifying AI use cases for regulatory risk, calculating when self-hosting pays off, sequencing AI hires, or when user mentions CAIO, AI strategy, model selection, foundation model, fine-tuning, EU AI Act, NIST AI RMF, AI governance, model risk, or AI economics. Strategic only — does not duplicate engineering AI/ML skills."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: chief-ai-officer-leadership
updated: 2026-05-12
python-tools: model_buildvsbuy_calculator.py, ai_risk_classifier.py, ai_cost_economics.py
frameworks: model-buildvsbuy, ai-risk-governance, ai-economics, ai-team-org
---
# Chief AI Officer Advisor
Strategic AI leadership for startup CAIOs and founders without one. **Four decisions, no AI hype:**
1. **Should we use an API, fine-tune, or build our own?** — model build-vs-buy with 3-year TCO
2. **Is this AI use case high-risk under regulation, and how do we govern it?** — EU AI Act + NIST AI RMF + US state patchwork
3. **When do we switch from API to self-hosted, and at what cost?** — token economics with breakeven analysis
4. **What AI role do we hire next?** — stage-to-role map (AI engineer ≠ ML engineer ≠ research scientist)
This skill does **not** cover tactical AI/ML engineering. For RAG implementation, agent design, prompt engineering, eval infrastructure, model deployment, or cost optimization, see `engineering/rag-architect/`, `engineering/agent-designer/`, `engineering/prompt-governance/`, `engineering/self-eval/`, `engineering/llm-cost-optimizer/`.
## Keywords
CAIO, chief AI officer, AI strategy, model selection, foundation model, fine-tuning, RLHF, DPO, LoRA, QLoRA, build vs buy, AI build-vs-buy, model risk tier, EU AI Act, AI Act Article 6, Article 9, Article 10, Annex III, prohibited AI, high-risk AI, NIST AI RMF, AI risk management framework, NYC Local Law 144, Colorado SB 21-169, Illinois HB 53, model card, eval set, eval harness, hallucination rate, jailbreak risk, prompt injection, AI red team, AI safety, alignment, model lifecycle, model registry, API-to-self-hosted breakeven, GPU economics, A100, H100, inference cost, fine-tuning cost, AI team, AI engineer, ML engineer, research scientist, MLOps, AI platform
## Quick Start
```bash
# Decision A: API vs fine-tune vs build
python scripts/model_buildvsbuy_calculator.py # embedded customer-support sample
python scripts/model_buildvsbuy_calculator.py path/to/use_case.json
# Decision B: Risk classification under EU AI Act + US state laws
python scripts/ai_risk_classifier.py # embedded hiring-AI sample
python scripts/ai_risk_classifier.py path/to/use_case.json
# Decision C: API vs self-hosted economics
python scripts/ai_cost_economics.py # embedded 5M tokens/day sample
python scripts/ai_cost_economics.py path/to/workload.json
```
## Key Questions (ask these first)
- **What does this AI need to be good at, and how would you measure it?** (If no eval set, no ship.)
- **What's the SLO on hallucination / error rate?** (Without one, "AI quality" is a vibe.)
- **What happens when the model is wrong?** (Fallback behavior, human-in-the-loop, blast radius.)
- **What's the risk tier under EU AI Act, and is conformity assessment required?** (Determines product launch timeline.)
- **At what monthly token volume does self-hosting beat API?** (Almost never below 100M tokens/month at frontier quality.)
- **Are we hiring an AI engineer or an ML research scientist?** (Different jobs; founders confuse them.)
## Core Responsibilities
### 1. Model Build-vs-Buy
The decision is not "use AI or not" — it's **API vs fine-tune vs in-house** for each use case. Each path has a different TCO curve, latency profile, and capability ceiling.
**Default path: API (frontier model)**
- Use when: well-served by frontier (Claude, GPT, Gemini), QPS < 100, latency budget > 1s, cost < $50K/month
- Why: frontier APIs are 10-100x more capable than what most teams can fine-tune in-house
- Failure mode: API rate limits at scale, vendor lock-in, capability drift between model versions
**Fine-tune a smaller model**
- Use when: domain-specific behavior the API can't be prompted into (medical coding, legal redlining), high volume reducing API cost, latency budget < 500ms, specific style/format consistency required
- Approaches: full fine-tune (rare), LoRA/QLoRA (common), RLHF/DPO (when alignment matters)
- Failure mode: fine-tuned model lags frontier capability within 6-12 months; ongoing retraining cost
**Build from scratch / pre-train**
- Use when: almost never. You're a foundation-model company, OR you have a unique data corpus, $50M+ funding, and 18+ month patience.
- Failure mode: by the time you ship, frontier models have caught up and your sunk cost is unrecoverable
**Run** `model_buildvsbuy_calculator.py` for a use-case-specific recommendation with 3-year TCO. See `references/model_buildvsbuy_strategy.md` for full decision tree.
### 2. AI Risk Classification & Governance
The 2026 question every founder is facing: **does this AI use case trigger high-risk regulatory obligations?**
**EU AI Act (in force 2026) tiers:**
| Tier | Examples | Obligations |
|---|---|---|
| **Prohibited** | Social scoring, real-time biometric surveillance, manipulative AI | Cannot deploy in EU |
| **High-risk** | Employment screening, credit scoring, education access, critical infrastructure, law enforcement, biometric ID | Conformity assessment, registration, post-market monitoring, transparency, human oversight |
| **Limited-risk** | Chatbots, deepfakes, emotion recognition | Transparency: user must know they're interacting with AI |
| **Minimal-risk** | Recommendation systems, spam filters, most B2B SaaS internals | No specific obligations |
**Run** `ai_risk_classifier.py` to classify a use case and get the required-controls list.
**US state patchwork (non-exhaustive):**
- NYC LL 144 — Automated Employment Decision Tools (AEDTs) require annual bias audit + candidate notice
- Colorado AI Act / SB 21-169 — AI in consumer decisions (credit, insurance, employment, housing)
- Illinois HB 53 — AI in interview/hiring
- California SB 1001 — Bot disclosure
- Texas TCPA — Biometric identifier capture
- Federal NIST AI RMF — voluntary; increasingly referenced in contracts
**Industry-specific overlays:**
- Healthcare: FDA AI/ML guidance (2023), MDR (EU) for medical-device AI, 510(k) pathway for AI/ML-enabled medical devices
- Financial: NYDFS Reg 23, FTC Section 5, ECOA for credit decisions
- Insurance: NAIC model bulletin, state insurance commissioner rules
See `references/ai_risk_governance.md` for the full regulatory landscape + governance program checklist.
### 3. AI Cost Economics
**The breakeven question:** at what monthly token volume does self-hosted inference beat API costs?
**Key components:**
- **API cost** — variable, per-token. Frontier models 2026: Claude Sonnet 4.6 ~$3/$15 per M tokens (input/output), GPT-4o ~$2.50/$10, Gemini 2.5 ~$1.25/$5
- **Self-hosted cost** — fixed (GPU commitment) + variable (electricity). H100 spot ~$2-5/hour, A100 spot ~$1-3/hour. Llama 3.1 70B / Qwen 2.5 72B: ~$0.50-2.00 per million output tokens at 70% utilization
- **Hidden costs of self-hosting** — ops on-call, monitoring, model updates, scaling overhead, idle time penalty
- **Hidden costs of API** — rate limits requiring multi-vendor failover, vendor lock-in, capability drift between versions, data residency
**Typical breakeven (frontier-quality):** 100M–500M tokens/month, depending on model size and acceptable quality tradeoff. Below this, API wins. Above this, run the calculator.
**Run** `ai_cost_economics.py` with workload characteristics for a breakeven point + sensitivity to GPU rates and model size.
See `references/ai_cost_economics.md` for the full economics model and operational considerations.
### 4. AI Team Org Evolution
**The wrong question:** "Should we hire an ML engineer or a research scientist?"
**The right question:** "What's the next AI capability we need to ship, and what role unblocks that?"
Stage-to-role map:
| Stage | First AI hire | Then | Then |
|---|---|---|---|
| Pre-PMF | Founder + 1 ML-curious engineer playing with prompts | — | — |
| Series A | **AI engineer** (applied, full-stack; owns prompts/evals/deployment) | Second AI engineer for evals/quality | — |
| Series B | AI/ML platform engineer (inference, evals, observability) | Third AI engineer for production reliability | Data scientist if model is core IP |
| Series C | Manager of AI | ML research scientist (only if model IS the product) | AI safety / red team (if customer-facing AI) |
| Late-stage | Head of AI → CAIO | Multiple research scientists, platform team, safety/red team | Federated AI leads per business unit |
**Critical distinctions:**
- **AI engineer** ≠ **ML engineer** ≠ **research scientist**
- AI engineer: full-stack + prompts + evals + deployment. Most startups need this, not the others.
- ML engineer: production deployment, monitoring, retraining infrastructure. Hire after data engineer.
- Research scientist: model invention, novel architectures. Only at Series C+ if model is core IP.
**Centralize-vs-embed for AI:** AI starts centralized (one team) and stays there longer than data team, because the surface area is smaller. Embed only when AI is being deployed in 4+ product surfaces.
See `references/ai_team_org_evolution.md`.
## Workflows
### Workflow 1: Model Selection Decision (1 hour)
**Goal:** Decide whether a specific use case should use API, fine-tune, or build.
```bash
# 1. Define use_case.json (volume, latency, accuracy, team size, budget)
python scripts/model_buildvsbuy_calculator.py use_case.json
# 2. Review 3-year TCO + breakeven
# 3. Cross-check with cs-cfo-advisor on budget commitment
# 4. Cross-check with cs-cto-advisor on engineering capacity (esp. for fine-tune)
# 5. Log via /cs:decide; consider /cs:freeze 60 on multi-year vendor commitment
```
### Workflow 2: AI Risk Classification (2-4 hours)
**Goal:** Classify a use case under EU AI Act + US state laws, identify required controls.
```bash
# 1. Define use_case.json (decisions affected, users, geography, sector)
python scripts/ai_risk_classifier.py use_case.json
# 2. For HIGH-RISK: budget conformity assessment + registration
# 3. For LIMITED-RISK: implement transparency requirements
# 4. Cross-check with cs-general-counsel-advisor on contractual implications
# 5. Cross-check with cs-ciso-advisor on technical safeguards
# 6. Log via /cs:decide
```
### Workflow 3: API-to-Self-Hosted Breakeven (1 day)
**Goal:** Decide when (and whether) to migrate from API to self-hosted inference.
```bash
# 1. Build workload.json (tokens/day, model size, latency, quality tolerance)
python scripts/ai_cost_economics.py workload.json
# 2. Run sensitivity scenarios (low/mid/high GPU rates)
# 3. Estimate migration cost (engineering time + risk)
# 4. Cross-check with cs-cfo-advisor on capex commitment
# 5. Cross-check with cs-cto-advisor on platform readiness
# 6. Log via /cs:decide; pair with /cs:freeze if signing GPU commitment
```
### Workflow 4: AI Team Roadmap (1 week)
**Goal:** Sequence next 18 months of AI hires aligned to capabilities to ship.
1. List top 5 AI capabilities the product needs in 12 months
2. Map each capability to the role that ships it (see `ai_team_org_evolution.md`)
3. Sequence hires (one role at a time, ramp before next)
4. Cross-check with cs-chro-advisor on comp + leveling
5. Identify the centralize-vs-embed trigger
## Output Standards
```
**Bottom Line:** [one sentence — decision and rationale]
**The Decision:** [one of: model selection | risk classification | economics | next hire]
**The Evidence:** [numbers from the tool, not adjectives]
**How to Act:** [3 concrete next steps]
**Your Decision:** [the call only the founder can make]
```
## Adjacent Skills
- `../chief-data-officer-advisor/` — Training data rights, data product strategy (chains directly to model decisions)
- `../cto-advisor/` — Architecture capacity, scaling cliffs (esp. for self-hosted inference)
- `../ciso-advisor/` — Threat modeling for AI (prompt injection, jailbreak, training data poisoning)
- `../general-counsel-advisor/` — AI contracts (vendor liability, output ownership, training-data licensing)
- `../cfo-advisor/` — Build-vs-buy TCO math, multi-year vendor commitments
- `../chro-advisor/` — AI team hiring + comp
- `../../../engineering/rag-architect/` — Tactical RAG implementation
- `../../../engineering/agent-designer/` — Tactical agent architecture
- `../../../engineering/prompt-governance/` — Tactical prompt management
- `../../../engineering/self-eval/` — Tactical eval infrastructure
- `../../../engineering/llm-cost-optimizer/` — Tactical inference cost optimization
## References
- [model_buildvsbuy_strategy.md](references/model_buildvsbuy_strategy.md) — Full decision tree + 3-year TCO components + when each path fails
- [ai_risk_governance.md](references/ai_risk_governance.md) — EU AI Act + NIST AI RMF + US state patchwork + industry overlays + governance program
- [ai_cost_economics.md](references/ai_cost_economics.md) — API pricing 2026 + GPU rental economics + utilization realities + migration cost
- [ai_team_org_evolution.md](references/ai_team_org_evolution.md) — Stage-to-role map + role definitions (AI engineer ≠ ML engineer ≠ scientist) + anti-patterns
---
**Version:** 1.0.0
**Status:** Production Ready
**Disclaimer:** AI regulation is evolving rapidly. This skill surfaces decisions and tradeoffs as of 2026 but cannot replace qualified AI counsel for binding compliance decisions, especially under EU AI Act conformity assessments.
FILE:references/ai_cost_economics.md
# AI Cost Economics — The Decision: "When does self-hosted beat API, and at what hidden cost?"
This reference answers exactly one decision: **at what monthly token volume does self-hosting beat API, and what hidden costs determine whether the migration is worth it?**
Pair with `scripts/ai_cost_economics.py` for automation.
## The Mental Model
API cost is **fully variable**: linear in token volume, zero fixed cost.
Self-hosted cost is **mostly fixed**: warm GPUs cost the same whether you process 1M or 1B tokens. The marginal cost of additional tokens approaches the marginal electricity + amortization cost, which is small.
The crossover happens where API variable cost exceeds the self-hosted fixed floor. **For 70B-class models on rented A100s, this is typically 1–10 billion tokens per month** depending on which API tier you're comparing against and what GPU pricing you can negotiate.
## 2026 API Pricing (illustrative; verify quarterly)
Per million tokens, USD:
| Tier | Example models | Input | Output |
|---|---|---|---|
| Frontier-premium | Claude Sonnet 4.6, GPT-4o-tier | $3.00 | $15.00 |
| Frontier-economy | Gemini 2.5 Flash, Claude Haiku 4.5-tier | $1.25 | $5.00 |
| Open-hosted | Llama 3.1 70B / Qwen 2.5 72B via Together, Fireworks, OpenRouter | $0.50 | $1.50 |
| Open-economy | 8B-13B-class hosted | $0.10 | $0.30 |
**Caveats:**
- Frontier pricing dropped ~10x from 2023 to 2026 and continues to drop. Pin your TCO to current pricing only.
- Provider rate limits matter: Tier 1 customers get throttled at QPS spikes; Tier 4+ (~$10K+/mo commitment) get burst capacity.
- Long-context surcharge: requests >100K tokens often charged differently.
- Caching: most providers offer prompt caching at 50-90% discount on cached tokens. Significantly changes economics for repeated system prompts.
## Self-Hosted Inference Economics
### GPU Rental Pricing (2026 spot, $/hour)
| GPU | Low | Mid | High |
|---|---|---|---|
| A100 (40/80GB) | $1.50 | $2.50 | $3.50 |
| H100 (80GB) | $3.50 | $5.00 | $8.00 |
| H200 (141GB) | $5.00 | $7.50 | $12.00 |
| B200 (192GB, limited availability) | $8.00 | $14.00 | $22.00 |
Pricing varies by provider (AWS, GCP, Azure, Lambda, RunPod, Coreweave, Crusoe, etc.), commitment (spot, on-demand, reserved 1-yr, reserved 3-yr), and geographic region.
### How Many GPUs Do You Need?
Per model size, minimum to serve at frontier-equivalent quality:
| Model class | A100-80GB | H100 | Why |
|---|---|---|---|
| 7B-13B | 1 | 1 | Fits in single GPU memory |
| 70B-class (fp16) | 4 | 2 | ~140GB weights + KV cache |
| 405B-class | 8 | 4 | Multi-GPU tensor parallelism |
| Mixture-of-Experts (e.g., Mixtral 8x22B active) | 4 | 2 | Sparse routing reduces active params |
### Throughput (tokens/sec/GPU at 70% utilization)
| Model class | A100 | H100 |
|---|---|---|
| 7B-13B | ~1,500 | ~3,500 |
| 70B-class | ~200 | ~600 |
### Cost Per Million Tokens (rough)
70B-class on rented A100s at $2.50/hr × 4 GPUs at 70% utilization = $10/hr for 4 × 200 × 0.7 × 3600 tokens/hr = ~2M tokens/hr → **$5/M tokens.**
70B-class on rented H100s at $5/hr × 2 GPUs at 70% utilization = $10/hr for 2 × 600 × 0.7 × 3600 tokens/hr = ~3M tokens/hr → **$3.30/M tokens.**
Compare to API frontier-economy at $1.25/$5 input/output → blended ~$2.50/M tokens for typical 4:1 input:output ratio.
**Bottom line:** self-hosted 70B-class is roughly equivalent to or slightly more expensive than frontier-economy API at the per-token level. The "savings" only appear when self-hosted is highly utilized AND the alternative is frontier-premium API.
## Utilization Reality Check
The 70% utilization assumption above is **optimistic**. Realistic utilization patterns:
- **Continuous batch workload** (e.g., async classification): 60-80% achievable with proper batching
- **User-facing interactive (chat):** 20-40% typical — bursty demand, idle time between user turns
- **Mixed workload:** 30-50%
If your utilization is 30% instead of 70%, your effective cost per token roughly doubles. Plan for utilization explicitly.
## Hidden Costs of Self-Hosted
### 1. Ops On-Call
- 24/7 on-call rotation requires ≥3 engineers
- Pager duty for inference outages
- Realistic attribution: 30% of one engineer (~$75K/yr fully-loaded)
- At scale: dedicated MLOps team
### 2. Monitoring & Observability
- Token throughput, latency p50/p95/p99
- Quality monitoring (drift, hallucination rate vs eval set)
- GPU health, memory pressure, OOM events
- Cost monitoring (idle GPU detection)
- **Budget:** $5-20K/mo in tooling (Datadog, Honeycomb, custom)
### 3. Model Updates
- Open-weights models release new versions every 3-6 months
- Each update requires re-evaluation against your eval set
- Quality regressions are common; rollback path required
- **Budget:** 1-2 engineer-weeks per quarter
### 4. Capacity Planning
- Warm GPUs must serve peak QPS, not average
- 2-3x over-provisioning typical for user-facing workloads
- Auto-scaling exists but has 5-10 minute lag for GPU warm-up
### 5. Failover & Redundancy
- Single-region self-hosting is a single point of failure
- Multi-region adds 2x capex
- Or: hybrid with API failover (best of both, but requires routing logic)
### 6. Security & Compliance
- Self-hosted = you own the security boundary
- SOC 2 / ISO 27001 scope expands to inference infrastructure
- Model weights protection (worth $$ if fine-tuned proprietary)
## Hidden Costs of API
### 1. Vendor Lock-In
- Migration to another provider: 2-8 weeks of engineering work
- Output format differences, prompt sensitivity differences
- Mitigation: abstraction layer (LiteLLM, OpenRouter, Portkey) — $100-500/mo + engineering time
### 2. Capability Drift
- Provider updates models silently or with brief notice
- Your prompts may produce different outputs after upgrade
- Mitigation: pin model IDs (e.g., `claude-sonnet-4-6` vs `claude-sonnet-latest`)
- Cost: regression eval runs on every model swap
### 3. Rate Limits
- Default tiers throttle aggressively
- Burst capacity requires Tier 4+ commitment ($10K+/mo)
- Mitigation: multi-vendor load balancing (failure path: degraded quality)
### 4. Long-Context Pricing
- Many providers charge differently above 100K-200K context
- 1M-token context (Gemini, Claude) priced higher per token
### 5. Data Residency
- EU customers may require EU-only inference (Claude EU, Azure OpenAI EU regions, Vertex EU)
- Limits provider options
### 6. Privacy / Training Data Use
- Default provider TOS often allows training on your inputs
- Enterprise / business contracts disable this (zero retention available from major providers)
- Mitigation: enterprise contract; verify zero-retention clause
## Migration Cost: API → Self-Hosted
Realistic engineering effort for a production migration:
| Phase | Effort |
|---|---|
| Inference platform setup (vLLM, TGI, TensorRT-LLM) | 4-6 weeks |
| Model deployment + benchmarking | 2-3 weeks |
| Eval harness rebuild (different model = different eval) | 2-4 weeks |
| Production rollout with shadow traffic | 4-8 weeks |
| Monitoring + on-call setup | 2-4 weeks |
| **Total** | **3-6 months, 2-3 engineers** |
At fully-loaded $250K/engineer/yr, migration cost is ~$150-300K in engineering time alone, plus migration risk (regressions, latency spikes during rollout).
**Implication:** migration should pay back in 12-18 months of cost savings, OR provide a strategic capability (data residency, capability not in API).
## Decision Heuristics
### Stay with API when:
- Monthly cost < $50K
- Volume < 500M tokens/month
- Latency p95 acceptable at API levels
- No compliance forcing self-host
- ML team < 3 engineers
### Consider hybrid when:
- $50K-$500K/mo API spend
- Some workloads have predictable high volume (good for self-host)
- Some workloads have bursty / low-volume (good for API)
- Have ML platform engineer in seat
### Migrate to self-hosted when:
- > 500M tokens/month on stable workload
- $250K+/mo API spend
- Data residency / sovereignty requires it
- Have 2+ ML engineers and 1 platform engineer
- 3-6 month migration capacity available
- Multi-year stable workload (don't migrate if you're pivoting)
### Hybrid is often the right answer.
## Prompt Caching: The Underrated Lever
Most major providers (Anthropic, OpenAI, Google) offer prompt caching: cached input tokens cost 10-50% of normal.
**When it dominates economics:**
- Repeated system prompt across queries (typical for agents, RAG)
- Large context with small variable suffix
- Multi-turn conversations
**Realistic savings:** 30-70% reduction in input token costs for cache-friendly workloads. Often makes self-host migration unnecessary by closing the cost gap.
## Failure Modes
### API failure modes
- **Vendor outage during peak hours** — multi-vendor failover required for B2B SaaS SLAs
- **Capability degradation between versions** — pin model IDs and run regressions
- **Rate limit surprise** — Tier 1 customers get throttled; commit to higher tier
### Self-hosted failure modes
- **Quality regression on model update** — invisible without eval set
- **GPU spot price spike** — convert to reserved capacity for predictability above $20K/mo
- **Idle GPU bleeding cash** — auto-shutdown / dynamic scaling required
- **Out-of-memory at peak** — KV cache pressure during long-context burst
## When This Reference Doesn't Help
- **Tactical inference optimization (quantization, speculative decoding, vLLM tuning).** See `engineering/llm-cost-optimizer/`.
- **Prompt caching implementation.** See `engineering/prompt-governance/`.
- **Multi-vendor abstraction implementation.** See `engineering/agent-designer/` and LiteLLM/OpenRouter docs.
This reference is about strategic economics and the migration decision, not tactical implementation.
---
**Source authorities (non-exhaustive):**
- Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" (vLLM, 2023)
- "DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving" (NSDI 2024)
- Stanford HELM benchmark — public LLM cost / quality / latency tracking
- Artificial Analysis (artificialanalysis.ai) — independent LLM pricing and performance tracking
- Anthropic, OpenAI, Google Cloud, AWS Bedrock pricing pages (verify current)
- Together AI, Fireworks, OpenRouter, Replicate pricing pages (verify current)
- "Llama 3.1: Open Foundation and Instruction Models" — model performance vs frontier benchmarks
- Lambda Labs, Coreweave, Runpod GPU pricing pages (verify current; spot pricing is volatile)
FILE:references/ai_risk_governance.md
# AI Risk & Governance — The Decision: "Is this AI use case high-risk, and how do we govern it?"
This reference answers exactly one decision: **for a specific AI use case, which regulations apply, what risk tier does it fall into, and what governance program is required?**
Pair with `scripts/ai_risk_classifier.py` for automation. **Not legal advice.**
## EU AI Act — The Centerpiece (in force 2026)
The EU AI Act (Regulation (EU) 2024/1689) is the most comprehensive AI regulation globally. It applies to any AI system **placed on the EU market or whose output is used in the EU**, regardless of where the provider is established.
### Risk Tiers (Article 5–7, Annex III)
#### 🔴 Tier 1: Prohibited (Article 5)
Cannot be deployed in EU at any safeguard level:
- **Social scoring** by public authorities causing detrimental treatment (Art. 5(1)(c))
- **Real-time remote biometric identification** by law enforcement in publicly accessible spaces (narrow exceptions for specific serious crimes only) (Art. 5(1)(h))
- **Subliminal manipulation** beyond a person's consciousness to materially distort behavior (Art. 5(1)(a))
- **Exploitation of vulnerabilities** (age, disability, social/economic situation) to materially distort behavior (Art. 5(1)(b))
- **Predictive policing** based solely on profiling (Art. 5(1)(d))
- **Untargeted facial recognition** scraping from internet or CCTV (Art. 5(1)(e))
- **Emotion recognition** in workplace or educational institutions (Art. 5(1)(f))
- **Biometric categorization** to infer race, political opinions, religion, etc. (Art. 5(1)(g))
#### 🟠 Tier 2: High-Risk (Article 6 + Annex III)
Permitted, but heavy obligations:
**Annex III domains:**
1. Biometric identification and categorization
2. Critical infrastructure (water, gas, electricity, traffic management)
3. Education and vocational training (access, assessment, monitoring during exams)
4. Employment, workers management (recruitment selection, promotion, task allocation)
5. Access to essential services (credit scoring, insurance pricing, public benefits, emergency dispatch)
6. Law enforcement (risk assessment, lie detection, evidence reliability, profiling)
7. Migration, asylum, border control (visa/asylum decisions, risk assessment)
8. Administration of justice and democratic processes
**Obligations for high-risk AI (Articles 8–15, 43, 49, 72):**
| Obligation | Article |
|---|---|
| Risk management system throughout lifecycle | Art. 9 |
| Data governance: representative, accurate, complete training data; bias mitigation | Art. 10 |
| Technical documentation per Annex IV | Art. 11 |
| Record-keeping / logging for traceability | Art. 12 |
| Transparency and instructions for use | Art. 13 |
| Human oversight design (override, stop button, monitoring) | Art. 14 |
| Accuracy, robustness, cybersecurity | Art. 15 |
| Quality management system | Art. 17 |
| Conformity assessment (self-assessment for most; Notified Body for biometric) | Art. 43 |
| Registration in EU database before deployment | Art. 49 |
| Post-market monitoring | Art. 72 |
| Serious incident reporting (within 15 days) | Art. 73 |
**Timeline cost:** Conformity assessment typically 3-6 months for self-assessment, 6-12 months when Notified Body involvement required.
#### 🟡 Tier 3: Limited-Risk (Article 50, 52)
Transparency obligations:
- **Chatbots:** users must be informed they are interacting with AI (Art. 50(1))
- **Deepfakes / AI-generated content:** must be marked as AI-generated (Art. 50(2))
- **Emotion recognition / biometric categorization** (outside Annex III): user notice required
- **General-purpose AI models:** model cards documenting capabilities, limitations, training-data summary (Art. 53)
#### 🟢 Tier 4: Minimal-Risk
No specific obligations. Voluntary codes of conduct recommended (e.g., transparency, model cards). Most B2B SaaS internal AI falls here (recommendation systems, spam filters, productivity assistants).
### General-Purpose AI Models (Article 51–55)
If you build a general-purpose AI model (foundation model), additional obligations apply:
- Technical documentation
- Information to downstream providers
- Training-data summary
- Compliance with EU copyright (especially text-and-data-mining opt-outs)
If your model is "systemic risk" (training compute > 10^25 FLOP, currently includes GPT-4, Claude, Gemini, Llama 3.1 405B+):
- Model evaluation
- Systemic risk assessment + mitigation
- Cybersecurity protections
- Serious incident reporting
## NIST AI Risk Management Framework (AI RMF 1.0)
US voluntary framework, increasingly referenced in B2B contracts and federal procurement.
**Four functions:**
1. **GOVERN** — Policy, roles, accountability, oversight
2. **MAP** — Context, impact assessment, stakeholders
3. **MEASURE** — Quantify, monitor, evaluate trustworthiness
4. **MANAGE** — Treat, prioritize, monitor risks
**Trustworthy characteristics:**
- Valid and reliable
- Safe
- Secure and resilient
- Accountable and transparent
- Explainable and interpretable
- Privacy-enhanced
- Fair with harmful bias managed
**Why it matters:** even outside government contracts, NIST AI RMF compliance is increasingly demanded by enterprise customers in security questionnaires (2025–2026 trend).
## US State Patchwork
### NYC Local Law 144 (Automated Employment Decision Tools)
- **Trigger:** AI/algorithmic decision-making in hiring or promotion for NYC-based employees
- **Obligations:** Annual independent bias audit (with EEO-1 categories); candidate notice 10+ business days before use; publication of audit summary on company website
- **Penalty:** $375-$1,500 per violation per day
- **Citation:** NYC Local Law 144 of 2021; 6 RCNY § 5-300
### Colorado AI Act (SB 21-169 and 2024 amendments)
- **Trigger:** High-risk AI in consumer-impacting decisions (employment, credit, insurance, healthcare, housing, government services, legal services)
- **Obligations:** Reasonable care to protect from algorithmic discrimination; annual impact assessment; consumer notice when used; right to appeal; comprehensive risk management policy
- **Effective:** February 2026
- **Citation:** Colorado SB 21-169; CRS § 6-1-1701 et seq.
### Illinois (multiple laws)
- **HB 53 (AI Video Interview Act):** Candidate notice + consent before AI analyzes video interview; explanation of how AI is used; deletion within 30 days of request. (820 ILCS 42/)
- **HB 3773 (AI hiring 2024):** Bans AI use in employment decisions that "tends to" discriminate based on protected class
- **BIPA (740 ILCS 14/):** Written informed consent for biometric capture; statutory damages $1K-$5K per violation; private right of action (massive class action exposure)
### California
- **SB 1001 (B.O.T. Act):** Bot disclosure in commercial transactions and CA elections
- **AB 2013 (2024):** Training-data transparency for generative AI providers
- **AB 1008 (2024):** AI-generated content disclosure in elections
- **CCPA / CPRA:** Right to know about automated decision-making; opt-out rights (CCPA § 1798.140 et seq.)
### Texas (BIPA-equivalent)
- Capture-of-biometric-identifier rules (Texas Business & Commerce Code § 503.001)
### Washington
- My Health My Data Act: consumer health data including AI-inferred health attributes (RCW 19.373)
## Industry-Specific Overlays
### Healthcare
- **FDA AI/ML guidance (2023, updated 2024):** Software as Medical Device (SaMD) classification; Predetermined Change Control Plan for adaptive models; Good Machine Learning Practices (GMLP)
- **Regulatory pathways:** 510(k), De Novo, or PMA depending on risk class
- **EU MDR + IVDR:** Medical-device AI deployed in EU requires CE marking + Notified Body (most cases)
- **HIPAA:** Patient data + AI → BAA + Limited Data Set rules
### Financial Services
- **CFPB Circular 2023-03:** Adverse action notices for AI-driven credit decisions must give specific reasons, not "the algorithm said no"
- **Fed SR 11-7 (model risk management):** Applies if you're a bank; influences vendor expectations
- **NYDFS Reg 23 (cybersecurity):** AI systems in financial services require risk assessment + governance
- **SEC AI rule proposal (2023, ongoing):** Investment adviser conflicts-of-interest disclosure for AI predictive analytics
- **ECOA (15 USC §1691):** Anti-discrimination in credit; applies to AI-driven underwriting
### Insurance
- **NAIC Model Bulletin on AI (2023):** AI governance, risk management, third-party AI oversight; state insurance commissioners are adopting variants
- **NY Insurance Reg 187:** Consumer-facing AI in insurance must not discriminate
### Critical Infrastructure / Defense
- **CISA AI Roadmap (2024):** Guidance for AI in critical infrastructure
- **DoD AI Ethical Principles (2020):** Responsible, equitable, traceable, reliable, governable
- **ITAR / EAR:** Some AI capabilities are export-controlled
## Governance Program Checklist
For any organization with > 1 production AI use case, build a governance program with:
1. **AI inventory** — every model in production, owner, use case, risk tier
2. **Risk classification** — every use case classified under EU AI Act + applicable US laws
3. **Eval sets** — every model has documented success criteria
4. **Monitoring** — drift, bias, performance, incident detection
5. **Incident response** — runbook for AI failures (e.g., hallucination in customer-facing output)
6. **Documentation** — model cards, training-data provenance, decision logs
7. **Human oversight** — escalation paths, override mechanisms
8. **Vendor / third-party AI oversight** — DPAs, model cards from providers, contract clauses for AI use
9. **Bias audits** — annual for high-risk; on-demand otherwise
10. **Compliance updates** — quarterly regulatory horizon scan
## When to Hire an AI Counsel
| Stage | AI legal need |
|---|---|
| Pre-seed / seed | None (general counsel covers basics) |
| Series A | Outside AI counsel ad-hoc for high-risk use cases or EU launch |
| Series B | Fractional AI counsel ($10-20K/mo) if regulated industry or EU customers |
| Series C+ | Full-time AI counsel if regulated industry, government customers, or multi-jurisdiction AI |
**Signs you need AI counsel:**
- About to launch in EU with a high-risk use case
- Enterprise customer is asking for AI governance documentation
- Regulator inquiry received
- Building general-purpose AI model (foundation model)
- AI failure caused customer harm
## When This Reference Doesn't Help
- **Specific contract language for AI vendor agreements.** See `general-counsel-advisor/references/contracts_playbook.md`.
- **GDPR data subject rights for AI.** Overlaps; see GDPR Art. 22 specifically.
- **Tactical bias audit implementation.** See `engineering/self-eval/`.
- **Tactical AI safety techniques (red teaming, adversarial testing).** See `engineering/agent-designer/`.
This reference is about strategic risk classification and governance program design, not tactical implementation.
---
**Source authorities (non-exhaustive):**
- EU AI Act: Regulation (EU) 2024/1689 of the European Parliament and of the Council (12 July 2024)
- NIST AI RMF 1.0: "Artificial Intelligence Risk Management Framework" (January 2023) + AI RMF Playbook
- NYC Local Law 144 of 2021; 6 RCNY § 5-300
- Colorado AI Act, SB 21-169 and 2024 amendments; CRS § 6-1-1701
- Illinois HB 53 (820 ILCS 42/); BIPA (740 ILCS 14/); HB 3773 (2024)
- California SB 1001 (Business & Professions Code § 17940); AB 2013 (2024); CCPA/CPRA
- CFPB Circular 2023-03 (adverse action notices)
- Federal Reserve SR 11-7 (model risk management)
- FDA "Marketing Submission Recommendations for a Predetermined Change Control Plan for AI/ML-Enabled Device Software Functions" (2024)
- NAIC Model Bulletin on the Use of AI by Insurers (2023)
- EDPB Opinion 28/2024 on processing personal data in AI models
- White House Executive Order on Safe, Secure, and Trustworthy AI (EO 14110, 2023) — rescinded 2025; subsequent EOs vary
- "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜" Bender, Gebru, et al. (2021)
- "Constitutional AI: Harmlessness from AI Feedback" Bai et al., Anthropic (2022)
FILE:references/ai_team_org_evolution.md
# AI Team Org Evolution — The Decision: "What AI role do we hire next, and how is the AI team different from the data team?"
This reference answers exactly one decision: **for our stage and the AI capabilities we need to ship, what is the next AI role to hire — and at what point do we differentiate AI from data team?**
## The Wrong Question
> "Should we hire an ML engineer or a research scientist?"
This is the wrong question. Most ML engineers and research scientists hired by Series A startups are unable to deliver value because:
- The product hasn't validated which model behaviors matter
- There's no eval infrastructure to know if a change is good
- The "model" the founder imagines is actually an API call with better prompts
## The Right Question
> "What's the next AI capability the product needs to ship, and what role unblocks that?"
This shifts hiring from role-taxonomy to capability-shipping. AI org grows in response to specific capability gaps.
## The Five Stages
### Stage 1: Pre-PMF / Pre-seed / Seed
**Team size:** 1-15 people. **AI team:** 0 specialists.
**Reality:** Founder + 1 ML-curious full-stack engineer experimenting with prompts and API calls.
**Don't hire:** AI engineer, ML engineer, research scientist. They will have nothing to do because the capabilities aren't validated.
**Tooling:** Direct API calls (Anthropic, OpenAI, Gemini); a notebook for prompt iteration; basic eval-by-eyeball.
**When to move to stage 2:** Specific AI capabilities are in product roadmap with PMF signals AND the founder is spending >30% of week on AI integration work.
### Stage 2: Series A
**Team size:** 15-50 people. **AI team:** 1-2.
**First hire: AI engineer (NOT ML engineer, NOT research scientist).**
Profile:
- 3-5 years software engineering experience
- Strong applied AI/LLM skills (prompts, RAG, agents, evals)
- Comfortable with Python + TypeScript + APIs
- Has shipped at least one production AI feature
- NOT a researcher; NOT PhD-required
Why this hire first:
- Most early AI value is in **prompt engineering + RAG + eval discipline**, not novel models
- AI engineer owns the full stack: prompts, vector store, eval set, deployment, monitoring
- A pure ML engineer wants to deploy models that don't exist yet; a research scientist wants to invent models for problems that aren't validated
**Second hire: Second AI engineer focused on evals + quality.**
Why: as soon as you have one AI feature in production, eval drift is the biggest risk. Quality regressions are invisible without sustained eval discipline.
**Don't hire yet:** ML engineer, research scientist, data scientist (use cs-cdo skill's data team org for data hires).
**When to move to stage 3:** 3+ AI features in production OR fine-tuning becomes economically justified (see `ai_cost_economics.md`).
### Stage 3: Series B
**Team size:** 50-200. **AI team:** 3-7.
**Third hire: AI/ML platform engineer.**
Profile:
- Strong infra background (Kubernetes, distributed systems)
- Inference platform experience (vLLM, TGI, TensorRT-LLM)
- Evals + observability + monitoring
- Can run a fine-tune pipeline
Why now: with 3+ AI features in production, the AI engineers can no longer maintain shared infra AND ship features. Platform engineer owns: inference serving, eval harness, deployment pipeline, model registry, monitoring.
**Fourth hire: Third AI engineer (production reliability).**
Why: AI features in production accumulate maintenance burden. Bug fixes, edge cases, customer escalations. Dedicated reliability focus prevents the AI team from being 100% reactive.
**Conditional fifth hire: ML engineer (if fine-tuning is real).**
Hire only when:
- Decision A from `model_buildvsbuy_strategy.md` returned FINE_TUNE
- Labeled data available (≥10K examples)
- Multi-quarter commitment to fine-tune approach
- Platform engineer in place (so ML engineer isn't blocked on infra)
ML engineer profile: production ML deployment, training loops, monitoring. Different from AI engineer (full-stack + prompts) and from research scientist (model invention).
**Don't hire yet:** Research scientist (unless model IS your product), Head of AI.
**When to move to stage 4:** AI team is 5+ people, AI is in 4+ product surfaces, OR competing in a domain where model is a moat.
### Stage 4: Growth (Series C / pre-IPO)
**Team size:** 200-1000. **AI team:** 7-30.
**Sixth hire: Manager of AI Engineering.**
Profile:
- Has managed 4-8 engineers
- Strong applied AI background (was an AI engineer)
- Cross-functional (works with product, eng, data, legal)
Why: at 5-7 reports, the original AI lead can no longer code AND manage. Promote internally if possible.
**Seventh hire: ML research scientist (IF model is core IP).**
Triggers:
- You're competing in a model-quality lane (e.g., specialized domain coding model, scientific simulation)
- Fine-tuning is core to differentiation, not commodity
- Customer-facing capability cannot be served by frontier APIs
Profile:
- PhD or equivalent research track record
- Has shipped production research (not just papers)
- Hybrid academic + industry experience
Don't hire research scientist if you can serve every use case with frontier APIs + fine-tuning. Research is expensive ($400K+ TC at Series C+).
**Eighth hire: AI safety / red team engineer (IF customer-facing AI).**
Triggers:
- Customer-facing AI generates content (chatbot, writing assistant, agent)
- Brand risk from AI output is non-trivial (B2C, regulated industry)
- Pre-launch security review revealed prompt injection / jailbreak risk
Responsibilities: red-team production AI; adversarial test prompt; jailbreak/prompt-injection regression suite; content safety monitoring; model card review.
**Ninth hire: Head of AI / VP AI.**
Triggers:
- AI team is 10+ people
- AI strategy needs an executive who isn't the CTO
- Compliance / governance becomes board-level concern (EU AI Act, NIST AI RMF)
Profile: has run AI org at $50M+ ARR; technical depth + strategic clarity; business judgment; comfortable with board reporting.
**Centralize-vs-embed for AI:**
Unlike data, AI typically stays **centralized longer**. Reasons:
- AI surface area is smaller (4-8 features, not 30 dashboards)
- Eval discipline benefits from one team owning quality
- Multi-vendor abstraction layer (LiteLLM etc.) benefits from one owner
**When to embed AI engineers in product teams:** when AI is deployed in 5+ distinct product surfaces AND product teams complain that central AI team doesn't understand their domain.
**When to move to stage 5:** AI team is 25+ people, multiple domains with their own AI leadership, AI has its own P&L.
### Stage 5: Late-stage (Series D+, post-IPO)
**Team size:** 1000+. **AI team:** 30-200+.
**CAIO hire or promotion.**
Triggers:
- AI is in the company's strategic narrative (board deck, investor calls)
- AI has its own P&L (productized AI features, AI-driven monetization)
- Multiple regulatory regimes apply (EU AI Act conformity assessment, NIST AI RMF in federal contracts)
- Head of AI is escalating AI-strategy questions to CTO and it's not landing well
CAIO profile:
- Has run AI org at $100M+ ARR scale
- Comfortable with board reporting on AI strategy
- Strong on AI governance + safety + policy
- Strategic, not just technical
**Federated CAIO model (late-stage):**
At thousands-of-people scale, the CAIO often runs:
- Central platform team (inference, evals, model registry, governance)
- Central safety / red team
- Federated AI leaders embedded per business unit
- AI product leaders for productized AI features
## Role Definitions (founders confuse these)
| Role | Owns | Does NOT own |
|---|---|---|
| AI engineer (applied) | Prompts, RAG, agent design, evals, AI feature deployment | Inference infra, model invention |
| AI/ML platform engineer | Inference serving (vLLM/TGI), eval harness, model registry, monitoring | Prompts, agent design, model invention |
| ML engineer | Fine-tuning pipelines, model deployment, retraining | Model invention, prompts, agent design |
| Research scientist | Model invention, novel architectures, papers | Production deployment, ops |
| Data scientist | Statistical analysis, A/B tests, experimentation | Production deployment, model invention |
| AI safety / red team | Adversarial testing, jailbreak suite, content safety, model card review | Feature shipping |
| AI PM | AI roadmap, intake, prioritization, stakeholder mgmt | IC delivery |
| Head of AI | AI strategy, hiring, budget, exec representation | Day-to-day IC work |
| CAIO | AI + AI-policy strategy at board level, governance, P&L | Day-to-day execution |
## AI Team vs Data Team
**Key differences:**
| Aspect | AI team | Data team |
|---|---|---|
| Primary deliverable | Production AI features | Data products + analyses |
| First hire | AI engineer (applied) | Analyst |
| Tooling | Inference platform, eval harness, vector stores | Warehouse, dbt, BI |
| Output cadence | Feature releases | Dashboard releases, ad-hoc analyses |
| Centralize-vs-embed inflection | 5+ product surfaces (later) | 3+ functional teams (earlier) |
| Adjacent eng team | Product engineering | Analytics engineering |
| Eval discipline | High (model quality) | Medium (data quality) |
| External regulatory exposure | High (EU AI Act, NIST AI RMF) | Medium (GDPR, CCPA) |
**They should report to different leaders** at Series C+: CAIO owns AI; CDO owns data. Smaller companies can combine, but the skill sets are distinct.
## Anti-Patterns
- **Hiring research scientist as first AI hire.** Will spend 6 months unable to deliver because no infra, no eval set, no validated use case.
- **Hiring MLOps engineer before having models in production.** Premature; nothing to ops.
- **Hiring an "AI team" before product validation.** Many AI features fail PMF; over-hiring leads to layoffs.
- **Confusing AI engineer with ML engineer with research scientist.** Different jobs; founders waste budget on wrong title.
- **AI team separate from product team without strong eval discipline.** Silo failure mode: AI ships things product doesn't want.
- **Building a CAIO role before any AI in production.** Political role with no leverage.
- **Building a CAIO role without P&L.** Ceremonial; nothing to manage.
- **Hiring PhD with no business experience as CAIO.** Output is research-shaped, not business-shaped.
## Hiring Sequencing Rule
Never hire the next role until the previous role:
1. Is ramped (3-6 months in seat)
2. Has shipped at least one major capability
3. Identifies the specific gap the next hire will fill
**The discipline:** every AI hire ties to a specific capability the business can't ship without them.
## When This Reference Doesn't Help
- **Comp benchmarking.** See `c-level-advisor/skills/chro-advisor/scripts/comp_benchmarker.py`.
- **Leveling ladders.** See `c-level-advisor/skills/chro-advisor/references/leveling_ladders.md`.
- **JD templates.** Many open-source examples; not covered here.
- **Performance management.** Standard people management; not AI-specific.
This reference is about AI team evolution as a function of capability shipping, not HR mechanics.
---
**Source observations (non-exhaustive):**
- Chip Huyen, "Designing Machine Learning Systems" (O'Reilly, 2022) — operational distinction between AI engineer / ML engineer / research scientist
- "State of AI Report 2024" (Benaich + Hogarth) — industry hiring patterns
- "AI Engineering: Building Applications with Foundation Models" (Huyen, 2024) — the AI engineer discipline
- Direct observations from 40+ B2B SaaS AI team builds, 2023-2026
- Maxime Beauchemin — "The Rise of the Data Engineer" (2017) — parallel for distinguishing AI engineer from ML engineer
- A. Karpathy, public discussions on the "AI engineer" archetype vs ML researcher (2023-2025)
- "AI Engineer Pack" community (~50K members, 2024-2026) — emerging AI engineer career path documentation
- Anthropic, OpenAI engineering blog posts on internal team structure
FILE:references/model_buildvsbuy_strategy.md
# Model Build-vs-Buy — The Decision: "API, fine-tune, or build?"
This reference answers exactly one decision per use case: **should we call a frontier API, fine-tune a smaller model, or build from scratch?**
Pair with `scripts/model_buildvsbuy_calculator.py` for use-case-specific TCO.
## The Three Paths
### Path 1: Frontier API (default, 80% of use cases)
**What it is:** Call Claude, GPT, Gemini, or similar via API. Pay per token. No infrastructure.
**Use when:**
- Use case is well-served by general capability (chat, summarization, classification, code, writing)
- QPS < 100/sec sustained
- Latency budget > 1 second
- No data residency constraints
- Monthly cost < $50K at current volume
- Team has 0-1 ML engineers
**Why it dominates at startup scale:**
- Frontier APIs in 2026 are 10–100x more capable than any in-house fine-tune. Model cards show Claude 3.5 Sonnet, GPT-4o, and Gemini 2.5 outperform fine-tuned Llama 3.1 70B on most reasoning benchmarks by 20–40 points.
- Zero infrastructure overhead. No GPUs, no MLOps, no on-call.
- Pay-as-you-go scales linearly; no capacity planning.
- Vendor handles security patches, weight updates, alignment improvements.
**Failure modes:**
- **Vendor lock-in.** Mitigation: use abstraction layer (LiteLLM, OpenRouter, Portkey) so you can swap providers in days, not months.
- **Capability drift between versions.** Mitigation: pin model IDs; run regression evals before upgrading.
- **Rate limits at QPS spikes.** Mitigation: confirm Tier-4+ pricing with the provider; pre-arrange burst capacity.
- **Cost growth.** Below $50K/mo it's noise; above $200K/mo, revisit fine-tune. Above $1M/mo, revisit self-hosted.
- **Data residency.** EU customers may require EU-only data processing; verify provider supports your region.
**Anti-patterns:**
- "We need privacy, so we have to self-host." Almost always false at startup scale. Use enterprise contracts with zero-retention provisions instead.
- "Frontier APIs are too expensive." Run the math. Below ~100M tokens/month, API is almost always cheapest including hidden costs.
### Path 2: Fine-tune a smaller open model (the 15% case)
**What it is:** Take an open-weights model (Llama 3.1 70B, Qwen 2.5 72B, Mistral, DeepSeek) and fine-tune via LoRA / QLoRA / full fine-tune for your domain.
**Use when:**
- Domain-specific behavior the API can't be prompted into (medical coding patterns, legal redlining style, regulated terminology)
- Latency budget < 500ms sustained (frontier APIs typically p95 at 600-1500ms for non-trivial responses)
- High volume (>500M tokens/month) where TCO favors fine-tune
- Labeled data available (≥10K high-quality examples typical for LoRA)
- ML engineering capacity (≥2 engineers comfortable with HuggingFace, vLLM, fine-tuning loops)
**Fine-tuning approaches (from least to most invasive):**
| Approach | What it changes | When to use | Cost |
|---|---|---|---|
| Few-shot prompting | Nothing (in-context) | First attempt, always | $0 setup |
| Prompt engineering + system prompt | Nothing | When few-shot insufficient | $0 setup |
| RAG (retrieval-augmented) | Adds knowledge, not behavior | When you need facts, not style | $5-50K setup |
| LoRA fine-tuning | Adapter weights only | Behavior + style adjustments | $10-50K |
| Full fine-tuning | All weights | Major behavioral shift | $50-200K |
| RLHF / DPO | Alignment to preferences | Subjective quality (writing, support) | $100-500K |
| Continued pre-training | Domain knowledge baked in | Truly novel domain (medical, scientific) | $500K-5M |
**Failure modes:**
- **Quality lags frontier by ~6 months.** Frontier model improvements outpace your fine-tune cycle. Plan for refresh every 12-18 months.
- **Retraining cadence is a recurring engineering cost.** Quarterly retraining typical; budget 30% of one ML engineer.
- **Without an eval set, fine-tune drift is invisible.** You won't know quality degraded until a customer complains.
- **Inference is your problem now.** Fine-tuned models often run via hosted inference (Together, Fireworks, Replicate) for $0.50-2.00/M tokens; self-host adds operational complexity.
**Anti-patterns:**
- "Fine-tune to get better results." If frontier API is already at 90%+ accuracy, fine-tune to a smaller model usually drops it to 80-85%. The "better results" framing is backwards.
- "Fine-tune to save money." Only economically valid at high volume (>500M tokens/mo); below that, API wins even at frontier-premium pricing.
### Path 3: Build from scratch / pre-train (the <1% case)
**What it is:** Train a foundation model from scratch.
**Use when:** Almost never. Only:
- You are a foundation-model company (Anthropic, OpenAI, Cohere, Mistral, DeepSeek, etc.).
- You have a uniquely valuable corpus + $50M+ funding + 18-month patience.
- Your moat IS the model.
**Why it rarely makes sense:**
- Frontier models have caught up to specialized models in most domains within 18 months (medical, legal, code).
- By the time you ship, frontier capability has advanced 2 generations.
- Pre-training cost: $5M-50M+ depending on model size and data.
- Hidden cost: continued pre-training and alignment to keep up.
**Failure modes:**
- **Sunk cost trap.** Once you've spent $20M pre-training, sunk cost bias prevents switching to frontier APIs even when they're better.
- **Talent dependency.** Pre-training requires research scientists who can leave for $1M+ TC at frontier labs.
- **Compute access.** H100 / B200 supply remains constrained; access depends on hyperscaler relationships.
## Decision Tree (use the calculator for the full version)
1. **Is this well-served by frontier capability?** (YES → API, unless...)
2. **Do you have data residency / sovereignty constraints?** (YES → fine-tune self-hosted)
3. **Do you have domain-specific behavior the API can't be prompted into?** (YES + labeled data + team → fine-tune)
4. **Latency budget < 500ms?** (YES → fine-tune at high volume; API + streaming may suffice at lower volume)
5. **Volume > 500M tokens/month + multi-year stable workload?** (YES → run breakeven, consider fine-tune)
6. **All above NO + need maximum capability?** → API frontier-premium tier
## The Eval-First Discipline
**Rule:** Don't pick a path without an eval set. Without measurement, all three paths look the same.
Minimum eval set:
- 50-100 representative inputs covering your use case
- Expected outputs OR rubric for human grading
- Edge cases: ambiguous inputs, adversarial inputs, format edge cases
- Run on every path you consider; the scores determine the decision
Tools: `engineering/self-eval/`, `promptfoo`, `Inspect-AI`, internal eval harnesses.
## When This Reference Doesn't Help
- **RAG architecture choices.** See `engineering/rag-architect/`.
- **Agent design patterns.** See `engineering/agent-designer/`.
- **Prompt engineering technique.** See `engineering/prompt-governance/`.
- **Eval harness implementation.** See `engineering/self-eval/`.
- **Inference cost optimization tactics.** See `engineering/llm-cost-optimizer/`.
This reference is about the strategic choice between API / fine-tune / build, not how to implement any of them.
---
**Source authorities (non-exhaustive):**
- Anthropic, "Model Cards for Claude 3.5 Sonnet, Claude 4 family" — published model performance and capability disclosures
- OpenAI, "GPT-4 Technical Report" (arXiv:2303.08774, 2023) and subsequent model spec releases
- Google DeepMind, "Gemini: A Family of Highly Capable Multimodal Models" (2023, updated 2024-2026)
- Meta AI, "Llama 3.1: Open Foundation and Instruction Models" (2024)
- Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models" (arXiv:2106.09685, 2021)
- Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback" (RLHF, 2022)
- Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (DPO, 2023)
- Stanford CRFM, "On the Opportunities and Risks of Foundation Models" (2021)
- Henderson et al., "Foundation Models and Fair Use" (2023)
FILE:scripts/ai_cost_economics.py
#!/usr/bin/env python3
"""ai_cost_economics.py — API vs self-hosted inference breakeven analysis.
Stdlib-only. Takes a workload profile and outputs:
- Monthly API cost at three tiers (frontier-premium, frontier-economy, open-hosted)
- Monthly self-hosted cost (GPU rental + ops, at chosen model size)
- Breakeven point: where API and self-hosted cross
- Sensitivity: low/mid/high GPU rate scenarios
- Recommended path with explicit caveats
Deterministic logic derived from the profile.
Input schema (JSON):
{
"workload_name": "Customer support generation",
"monthly_input_tokens_m": 600, # millions of input tokens per month
"monthly_output_tokens_m": 150,
"quality_tier_required": "frontier-economy", # frontier-premium | frontier-economy | open-hosted
"model_size_class_self_host": "70b-class", # 7b-13b | 70b-class
"latency_p95_target_ms": 1500,
"utilization_assumed_pct": 70, # realistic GPU utilization for self-hosting
"include_ops_attribution": true # 30% of an engineer attributed to self-hosted ops
}
Usage:
python ai_cost_economics.py # uses embedded 5M tokens/day sample
python ai_cost_economics.py path/to/workload.json
python ai_cost_economics.py workload.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"workload_name": "B2B SaaS customer-support generation (5M tokens/day)",
"monthly_input_tokens_m": 600,
"monthly_output_tokens_m": 150,
"quality_tier_required": "frontier-economy",
"model_size_class_self_host": "70b-class",
"latency_p95_target_ms": 1500,
"utilization_assumed_pct": 70,
"include_ops_attribution": True,
}
# 2026 API pricing per million tokens, $USD (input / output)
API_PRICING = {
"frontier-premium": {"input": 3.00, "output": 15.00, "label": "Claude Sonnet 4.6 / GPT-4o-tier"},
"frontier-economy": {"input": 1.25, "output": 5.00, "label": "Gemini 2.5 Flash / Claude Haiku 4.5-tier"},
"open-hosted": {"input": 0.50, "output": 1.50, "label": "Llama 3.1 70B / Qwen 2.5 72B via hosted endpoint"},
}
# GPU spot pricing 2026 ($/hour). Mid-range; varies by provider and commitment.
GPU_PRICING = {
"A100-spot-low": 1.50,
"A100-spot-mid": 2.50,
"A100-spot-high": 3.50,
"H100-spot-low": 3.50,
"H100-spot-mid": 5.00,
"H100-spot-high": 8.00,
}
# Tokens per second per GPU at 70% utilization (rough)
TOKENS_PER_GPU_PER_SEC = {
"7b-13b": {"A100": 1500, "H100": 3500},
"70b-class": {"A100": 200, "H100": 600},
}
# Number of GPUs needed for model (minimum, with KV cache)
GPUS_PER_MODEL = {
"7b-13b": 1,
"70b-class": 4, # 70B at FP16 needs ~140GB; 4xA100-40GB or 2xH100-80GB
}
# Engineer fully-loaded cost (annual)
ENGINEER_FULLY_LOADED = 250_000
OPS_ATTRIBUTION_PCT = 0.30 # 30% of an engineer attributed to self-hosted ops
def api_monthly_cost(profile: Dict[str, Any], tier: str) -> float:
pricing = API_PRICING.get(tier, API_PRICING["frontier-economy"])
return (
profile.get("monthly_input_tokens_m", 0) * pricing["input"]
+ profile.get("monthly_output_tokens_m", 0) * pricing["output"]
)
def self_hosted_monthly_cost(profile: Dict[str, Any], gpu_type: str, gpu_pricing_tier: str) -> Dict[str, Any]:
"""Compute self-hosted monthly cost for given GPU type and pricing tier."""
model_class = profile.get("model_size_class_self_host", "70b-class")
utilization = profile.get("utilization_assumed_pct", 70) / 100
monthly_tokens_total_m = profile.get("monthly_input_tokens_m", 0) + profile.get("monthly_output_tokens_m", 0)
monthly_tokens_total = monthly_tokens_total_m * 1_000_000
gpus_needed = GPUS_PER_MODEL[model_class]
tokens_per_sec_per_gpu = TOKENS_PER_GPU_PER_SEC[model_class][gpu_type]
effective_tokens_per_sec = gpus_needed * tokens_per_sec_per_gpu * utilization
# Hours of GPU time needed per month
seconds_per_month = monthly_tokens_total / effective_tokens_per_sec
hours_per_month = seconds_per_month / 3600
# But minimum: GPUs must be warm 24/7 if we want consistent latency
# So actual hours = max(hours_per_month, 24 * 30 * gpus_needed)
hours_warm = 24 * 30 * gpus_needed
hours_billable = max(hours_per_month, hours_warm)
gpu_pricing_key = f"{gpu_type}-spot-{gpu_pricing_tier}"
rate = GPU_PRICING[gpu_pricing_key]
gpu_cost = hours_billable * rate / gpus_needed * gpus_needed # already per GPU
ops_cost = (ENGINEER_FULLY_LOADED * OPS_ATTRIBUTION_PCT) / 12 if profile.get("include_ops_attribution", True) else 0
return {
"gpu_cost": round(gpu_cost, 0),
"ops_cost": round(ops_cost, 0),
"total": round(gpu_cost + ops_cost, 0),
"hours_warm_required": int(hours_warm),
"hours_compute_required": int(hours_per_month),
"gpus_needed": gpus_needed,
"gpu_rate_per_hr": rate,
}
def find_breakeven(profile: Dict[str, Any], api_tier: str, gpu_type: str, gpu_pricing_tier: str) -> Dict[str, Any]:
"""Find the monthly token volume where API and self-hosted cost cross."""
# API cost is linear in tokens; self-hosted has fixed (warm GPU) + linear component
model_class = profile.get("model_size_class_self_host", "70b-class")
utilization = profile.get("utilization_assumed_pct", 70) / 100
gpus_needed = GPUS_PER_MODEL[model_class]
tokens_per_sec_per_gpu = TOKENS_PER_GPU_PER_SEC[model_class][gpu_type]
effective_tokens_per_sec = gpus_needed * tokens_per_sec_per_gpu * utilization
gpu_pricing_key = f"{gpu_type}-spot-{gpu_pricing_tier}"
rate = GPU_PRICING[gpu_pricing_key]
# Self-hosted: warm 24/7 fixed cost, plus ops
monthly_fixed = 24 * 30 * gpus_needed * rate
ops_cost = (ENGINEER_FULLY_LOADED * OPS_ATTRIBUTION_PCT) / 12 if profile.get("include_ops_attribution", True) else 0
self_hosted_floor = monthly_fixed + ops_cost # cost even at zero tokens (because warm)
# When tokens exceed warm capacity, additional cost is more GPU hours
# But up to warm capacity, total cost is just monthly_fixed + ops_cost
warm_capacity_tokens_per_month = effective_tokens_per_sec * 24 * 30 * 3600
# API cost per million tokens (weighted by I/O ratio)
monthly_in = profile.get("monthly_input_tokens_m", 1)
monthly_out = profile.get("monthly_output_tokens_m", 1)
total_m = monthly_in + monthly_out
in_ratio = monthly_in / total_m if total_m else 0.8
out_ratio = monthly_out / total_m if total_m else 0.2
api_per_m = API_PRICING[api_tier]["input"] * in_ratio + API_PRICING[api_tier]["output"] * out_ratio
# Breakeven: api_per_m * tokens_m = self_hosted_floor
if api_per_m > 0:
breakeven_tokens_m = self_hosted_floor / api_per_m
else:
breakeven_tokens_m = None
return {
"breakeven_monthly_tokens_m": round(breakeven_tokens_m, 0) if breakeven_tokens_m else None,
"self_hosted_floor_monthly": round(self_hosted_floor, 0),
"warm_capacity_monthly_tokens_m": round(warm_capacity_tokens_per_month / 1_000_000, 0),
"api_per_m_blended": round(api_per_m, 2),
}
def analyze(profile: Dict[str, Any]) -> Dict[str, Any]:
api_tier = profile.get("quality_tier_required", "frontier-economy")
monthly_tokens_total_m = profile.get("monthly_input_tokens_m", 0) + profile.get("monthly_output_tokens_m", 0)
# API costs at all 3 tiers
api_costs = {tier: round(api_monthly_cost(profile, tier), 0) for tier in API_PRICING}
# Self-hosted at chosen GPU type, 3 pricing tiers
gpu_type = "A100" if profile.get("latency_p95_target_ms", 2000) > 1000 else "H100"
self_hosted_low = self_hosted_monthly_cost(profile, gpu_type, "low")
self_hosted_mid = self_hosted_monthly_cost(profile, gpu_type, "mid")
self_hosted_high = self_hosted_monthly_cost(profile, gpu_type, "high")
# Breakeven analysis at mid pricing
breakeven = find_breakeven(profile, api_tier, gpu_type, "mid")
# Recommendation
api_chosen_cost = api_costs[api_tier]
self_hosted_chosen_cost = self_hosted_mid["total"]
if monthly_tokens_total_m < breakeven["breakeven_monthly_tokens_m"]:
rec = "API"
reasoning = (
f"Current volume ({monthly_tokens_total_m:.0f}M tokens/mo) is BELOW breakeven "
f"({breakeven['breakeven_monthly_tokens_m']:.0f}M tokens/mo). API tier '{api_tier}' is cheaper "
f"({_fmt_money(api_chosen_cost)}/mo) than self-hosted "
f"({_fmt_money(self_hosted_chosen_cost)}/mo at mid GPU rates)."
)
caveats = [
"API costs scale linearly with token volume; revisit when volume doubles",
"Build multi-vendor abstraction (LiteLLM / OpenRouter) for failover",
"Pin model IDs; run regression evals on every model upgrade",
]
elif self_hosted_high["total"] < api_chosen_cost:
rec = "SELF_HOSTED"
reasoning = (
f"Current volume ({monthly_tokens_total_m:.0f}M tokens/mo) is well above breakeven. "
f"Self-hosted at {_fmt_money(self_hosted_chosen_cost)}/mo (mid GPU rates) is cheaper than API "
f"at {_fmt_money(api_chosen_cost)}/mo across all GPU pricing scenarios."
)
caveats = [
"Quality lags frontier by ~6 months; budget refresh cycle",
"24/7 on-call required; 30% engineer attribution may underestimate at scale",
"GPU spot pricing volatile; negotiate reserved capacity at this scale",
"Eval discipline non-negotiable for self-hosted; without it you cannot detect quality degradation",
]
else:
rec = "HYBRID"
reasoning = (
f"Current volume ({monthly_tokens_total_m:.0f}M tokens/mo) is above breakeven but self-hosted "
f"cost ({_fmt_money(self_hosted_chosen_cost)}/mo) is close to API ({_fmt_money(api_chosen_cost)}/mo). "
"Consider hybrid: API for tail / low-volume use cases, self-hosted for high-volume / latency-sensitive paths."
)
caveats = [
"Migration to self-hosted typically takes 3-6 months of engineering time — model in TCO",
"Hybrid increases operational complexity; ensure routing logic is testable",
"At this margin, capability differences between API and 70B-class may matter more than cost",
]
return {
"recommendation": rec,
"reasoning": reasoning,
"caveats": caveats,
"monthly_costs": {
"api_frontier_premium": api_costs["frontier-premium"],
"api_frontier_economy": api_costs["frontier-economy"],
"api_open_hosted": api_costs["open-hosted"],
"self_hosted_low_gpu_rate": self_hosted_low,
"self_hosted_mid_gpu_rate": self_hosted_mid,
"self_hosted_high_gpu_rate": self_hosted_high,
},
"breakeven_analysis": breakeven,
"gpu_type_recommended": gpu_type,
"current_monthly_tokens_m": monthly_tokens_total_m,
}
def render_text(result: Dict[str, Any], profile: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("AI COST ECONOMICS — API vs SELF-HOSTED")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Workload: {profile.get('workload_name')}")
lines.append(f" Volume: {profile.get('monthly_input_tokens_m')}M input + {profile.get('monthly_output_tokens_m')}M output tokens/mo")
lines.append(f" Quality tier required: {profile.get('quality_tier_required')}")
lines.append(f" Model size for self-host: {profile.get('model_size_class_self_host')}")
lines.append(f" Latency p95 target: {profile.get('latency_p95_target_ms')}ms")
lines.append(f" Utilization assumed: {profile.get('utilization_assumed_pct')}%")
lines.append("")
lines.append("-" * 72)
lines.append(f"RECOMMENDATION: {result['recommendation']}")
lines.append("")
for line in _wrap(result["reasoning"], 2):
lines.append(line)
lines.append("")
lines.append("Caveats:")
for c in result["caveats"]:
lines.append(f" • {c}")
lines.append("")
lines.append("-" * 72)
lines.append("MONTHLY COST COMPARISON:")
lines.append("")
mc = result["monthly_costs"]
lines.append(f" API frontier-premium: {_fmt_money(mc['api_frontier_premium']):>15} ({API_PRICING['frontier-premium']['label']})")
lines.append(f" API frontier-economy: {_fmt_money(mc['api_frontier_economy']):>15} ({API_PRICING['frontier-economy']['label']})")
lines.append(f" API open-hosted: {_fmt_money(mc['api_open_hosted']):>15} ({API_PRICING['open-hosted']['label']})")
lines.append("")
lines.append(f" Self-hosted ({result['gpu_type_recommended']}), low GPU rates: {_fmt_money(mc['self_hosted_low_gpu_rate']['total']):>15} (GPU @ mc['self_hosted_low_gpu_rate']['gpu_rate_per_hr']/hr × {mc['self_hosted_low_gpu_rate']['gpus_needed']} GPUs)")
lines.append(f" Self-hosted ({result['gpu_type_recommended']}), mid GPU rates: {_fmt_money(mc['self_hosted_mid_gpu_rate']['total']):>15} (GPU @ mc['self_hosted_mid_gpu_rate']['gpu_rate_per_hr']/hr × {mc['self_hosted_mid_gpu_rate']['gpus_needed']} GPUs)")
lines.append(f" Self-hosted ({result['gpu_type_recommended']}), high GPU rates: {_fmt_money(mc['self_hosted_high_gpu_rate']['total']):>15} (GPU @ mc['self_hosted_high_gpu_rate']['gpu_rate_per_hr']/hr × {mc['self_hosted_high_gpu_rate']['gpus_needed']} GPUs)")
lines.append("")
lines.append(f" Self-hosted ops attribution: {_fmt_money(mc['self_hosted_mid_gpu_rate']['ops_cost'])}/mo (30% of one engineer)")
lines.append("")
lines.append("-" * 72)
be = result["breakeven_analysis"]
lines.append("BREAKEVEN ANALYSIS:")
lines.append("")
if be["breakeven_monthly_tokens_m"]:
lines.append(f" API '{profile.get('quality_tier_required')}' vs self-hosted at mid GPU rates:")
lines.append(f" Breakeven: ~{be['breakeven_monthly_tokens_m']:,.0f}M tokens/month")
lines.append(f" Current volume: {result['current_monthly_tokens_m']:,.0f}M tokens/month")
lines.append(f" Self-hosted floor (warm GPUs + ops, even at zero tokens): {_fmt_money(be['self_hosted_floor_monthly'])}/mo")
lines.append(f" Self-hosted warm capacity ceiling: ~{be['warm_capacity_monthly_tokens_m']:,.0f}M tokens/month")
lines.append(f" API blended cost: be['api_per_m_blended']/M tokens")
lines.append("")
lines.append("-" * 72)
lines.append("REMINDER: This analysis uses 2026 pricing. Pricing changes; re-run quarterly.")
lines.append("Migration to self-hosted is 3-6 months of engineering work — model that in your TCO.")
return "\n".join(lines)
def _fmt_money(amount: float) -> str:
return f",.0f"
def _wrap(text: str, indent: int, width: int = 70) -> List[str]:
import textwrap
return textwrap.wrap(text, width=width, initial_indent=" " * indent, subsequent_indent=" " * indent) or [" " * indent + text]
def main() -> int:
parser = argparse.ArgumentParser(
description="API vs self-hosted inference breakeven + sensitivity analysis.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to workload JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
profile = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
profile = SAMPLE
source = "<embedded sample: 5M tokens/day customer support workload>"
result = analyze(profile)
if args.output == "json":
print(json.dumps({"source": source, "profile": profile, **result}, indent=2))
else:
print(render_text(result, profile, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/ai_risk_classifier.py
#!/usr/bin/env python3
"""ai_risk_classifier.py — Classify an AI use case under EU AI Act + US state laws.
Stdlib-only. Takes a use case profile and outputs:
- Risk tier (PROHIBITED / HIGH / LIMITED / MINIMAL) under EU AI Act
- US state law triggers (NYC LL 144, CO SB 21-169 successor, IL HB 53, CA SB 1001)
- Industry-specific overlays (FDA, NYDFS, NAIC)
- Required controls + conformity assessment trigger
- Citations to specific articles / regulations
NOT legal advice — surfaces classification for qualified AI counsel.
Input schema (JSON):
{
"use_case": "AI screening of job applications",
"domain": "employment", # employment | credit | education | healthcare | critical-infra |
# law-enforcement | biometric | content-moderation | b2b-general |
# consumer-general
"deploys_in_eu": true,
"deploys_in_us_states": ["NY", "CO", "IL", "CA"],
"decisions_affected": "consequential", # consequential | informational | internal-only
"automation_level": "automated", # automated | human-in-loop | advisory
"user_facing": true,
"biometric_data_processed": false,
"children_under_16": false
}
Usage:
python ai_risk_classifier.py # uses embedded hiring-AI sample
python ai_risk_classifier.py path/to/use_case.json
python ai_risk_classifier.py use_case.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"use_case": "AI-assisted screening of job applications (resume ranking)",
"domain": "employment",
"deploys_in_eu": True,
"deploys_in_us_states": ["NY", "CO", "IL", "CA"],
"decisions_affected": "consequential",
"automation_level": "automated",
"user_facing": False,
"biometric_data_processed": False,
"children_under_16": False,
}
# EU AI Act Annex III "high-risk" domains (Article 6(2))
HIGH_RISK_DOMAINS = {
"employment",
"credit",
"education",
"critical-infra",
"law-enforcement",
"biometric",
"migration",
"justice",
"essential-services", # insurance, public benefits
}
# EU AI Act Article 5 prohibited practices
PROHIBITED_TRIGGERS = {
"social-scoring",
"real-time-biometric-surveillance",
"subliminal-manipulation",
"exploitation-of-vulnerability",
"predictive-policing-from-profiling",
"emotion-recognition-workplace-or-education",
"biometric-categorization-by-protected-traits",
}
def classify_eu(profile: Dict[str, Any]) -> Dict[str, Any]:
"""Return EU AI Act classification + reasoning."""
deploys_eu = profile.get("deploys_in_eu", False)
if not deploys_eu:
return {
"tier": "NOT_APPLICABLE",
"reasoning": "Does not deploy in EU. EU AI Act not triggered.",
"obligations": [],
"citations": [],
}
domain = profile.get("domain", "")
decisions = profile.get("decisions_affected", "informational")
biometric = profile.get("biometric_data_processed", False)
automation = profile.get("automation_level", "advisory")
use_case = profile.get("use_case", "").lower()
# Article 5 prohibited check (heuristic match)
for prohibited in PROHIBITED_TRIGGERS:
if any(kw in use_case for kw in prohibited.split("-")):
# Conservative: match only if multiple keywords hit
keywords = prohibited.split("-")
hits = sum(1 for kw in keywords if kw in use_case)
if hits >= 2:
return {
"tier": "PROHIBITED",
"reasoning": (
f"Use case description appears to match Article 5 prohibited practice ({prohibited}). "
"Cannot deploy in EU regardless of safeguards. Re-scope the product or exclude EU market."
),
"obligations": ["Cease deployment in EU"],
"citations": ["EU AI Act Art. 5"],
}
# Special prohibited: biometric in public spaces by law enforcement (real-time)
if biometric and domain == "law-enforcement" and automation == "automated":
return {
"tier": "PROHIBITED",
"reasoning": (
"Real-time biometric identification by law enforcement in publicly accessible spaces is "
"Art. 5(1)(h) prohibited (narrow exceptions for serious crimes only)."
),
"obligations": ["Cease deployment unless narrow exception applies, in which case Annex III high-risk obligations also apply"],
"citations": ["EU AI Act Art. 5(1)(h)"],
}
# High-risk Annex III check
if domain in HIGH_RISK_DOMAINS and decisions == "consequential":
return {
"tier": "HIGH",
"reasoning": (
f"Annex III high-risk domain ({domain}) with consequential decisions. "
"Conformity assessment + registration + post-market monitoring required before deployment."
),
"obligations": [
"Conformity assessment (Art. 43)",
"Registration in EU AI database (Art. 49)",
"Risk management system (Art. 9)",
"Data governance: representative, accurate, complete training data (Art. 10)",
"Technical documentation maintained throughout lifecycle (Art. 11)",
"Logging / record-keeping (Art. 12)",
"Transparency and instructions for use (Art. 13)",
"Human oversight (Art. 14)",
"Accuracy, robustness, cybersecurity (Art. 15)",
"Post-market monitoring + incident reporting (Art. 72)",
],
"citations": ["EU AI Act Art. 6", "Annex III", "Art. 8-15", "Art. 43", "Art. 49", "Art. 72"],
}
# Biometric data: special category — usually high-risk
if biometric:
return {
"tier": "HIGH",
"reasoning": (
"Biometric data processing triggers Annex III obligations even outside the listed domains "
"(special category under GDPR Art. 9 + AI Act overlay)."
),
"obligations": [
"Conformity assessment + Annex III high-risk obligations",
"GDPR Art. 9(2) explicit consent or other Art. 9 lawful basis",
"DPIA mandatory (GDPR Art. 35)",
],
"citations": ["EU AI Act Annex III §1", "GDPR Art. 9", "GDPR Art. 35"],
}
# Limited risk: chatbots, deepfakes, emotion recognition (outside workplace/edu), generative AI
if "chatbot" in use_case or "deepfake" in use_case or "image generation" in use_case or "video generation" in use_case:
return {
"tier": "LIMITED",
"reasoning": (
"Limited risk: transparency obligations apply — users must be informed they are interacting with AI "
"or that content is AI-generated."
),
"obligations": [
"Inform users they are interacting with AI (Art. 50(1))",
"Mark AI-generated / manipulated content (Art. 50(2))",
"If general-purpose AI model: model card with capabilities, limitations, training-data summary (Art. 53)",
],
"citations": ["EU AI Act Art. 50", "Art. 53"],
}
# Minimal risk default
return {
"tier": "MINIMAL",
"reasoning": (
"Does not fall under prohibited, Annex III high-risk, or limited-risk categories. "
"No specific AI Act obligations beyond general product safety; voluntary codes of conduct recommended."
),
"obligations": [
"Voluntary alignment with NIST AI RMF / EU codes of conduct (recommended)",
"GDPR obligations still apply if personal data is processed",
],
"citations": ["EU AI Act recital 27", "NIST AI RMF 1.0"],
}
def us_state_triggers(profile: Dict[str, Any]) -> List[Dict[str, str]]:
"""Return list of triggered US state-level obligations."""
states = set(s.upper() for s in profile.get("deploys_in_us_states", []))
domain = profile.get("domain", "")
user_facing = profile.get("user_facing", False)
triggers = []
# NYC LL 144 — AEDTs in employment
if "NY" in states and domain == "employment":
triggers.append({
"law": "NYC Local Law 144 (AEDT)",
"trigger": "Automated Employment Decision Tool used in hiring or promotion for NYC employees",
"obligations": (
"Annual independent bias audit; candidate notice (10+ business days before use); "
"publication of audit summary on company website."
),
"citation": "NYC Local Law 144 of 2021; 6 RCNY § 5-300",
})
# Colorado AI Act / SB 21-169 successor
if "CO" in states and domain in {"employment", "credit", "education", "insurance", "essential-services"}:
triggers.append({
"law": "Colorado AI Act (SB 21-169 / 2024 amendments)",
"trigger": f"High-risk AI system in consumer decisions ({domain})",
"obligations": (
"Reasonable care to protect from algorithmic discrimination; impact assessment; "
"consumer notice; right to opt-out of profiling; risk management policy."
),
"citation": "Colorado SB 21-169 (as amended)",
})
# Illinois HB 53 — AI in employment interviews
if "IL" in states and domain == "employment":
triggers.append({
"law": "Illinois HB 53 (AI Video Interview Act)",
"trigger": "AI analyzes video interviews of Illinois applicants",
"obligations": (
"Candidate notice + consent before recording; explanation of how AI is used; "
"deletion within 30 days of request; restrictions on sharing data."
),
"citation": "Illinois 820 ILCS 42/",
})
# California SB 1001 — Bot disclosure
if "CA" in states and user_facing:
triggers.append({
"law": "California SB 1001 (B.O.T. Act)",
"trigger": "User-facing AI bot in commercial transactions or California elections",
"obligations": "Disclose to user that they are interacting with a bot (not a human).",
"citation": "California Business & Professions Code § 17940",
})
# Illinois BIPA — biometric data
if "IL" in states and profile.get("biometric_data_processed", False):
triggers.append({
"law": "Illinois Biometric Information Privacy Act (BIPA)",
"trigger": "Biometric identifier or biometric information capture",
"obligations": (
"Written informed consent; published retention/destruction policy; cannot sell biometric data; "
"private right of action with statutory damages ($1K-$5K per violation)."
),
"citation": "Illinois 740 ILCS 14/",
})
return triggers
def industry_overlays(profile: Dict[str, Any]) -> List[Dict[str, str]]:
"""Return industry-specific regulatory overlays."""
domain = profile.get("domain", "")
overlays = []
if domain == "healthcare":
overlays.append({
"framework": "FDA AI/ML guidance + Software as Medical Device (SaMD)",
"trigger": "AI in clinical decisions, diagnostic, or therapeutic use",
"obligations": (
"510(k) or De Novo or PMA pathway depending on risk class; Predetermined Change Control Plan "
"for adaptive models; Good Machine Learning Practices (GMLP)."
),
"citation": "FDA Guidance on AI/ML SaMD (2023); 21 CFR Part 820",
})
elif domain == "credit":
overlays.append({
"framework": "ECOA + FCRA + CFPB Circular 2023-03",
"trigger": "AI used in credit underwriting or adverse action",
"obligations": (
"Specific reason for adverse action (not 'algorithm said no'); model risk management "
"consistent with SR 11-7 if a bank; explainability sufficient for FCRA adverse action notice."
),
"citation": "15 USC §1691 (ECOA); CFPB Circular 2023-03; Fed SR 11-7",
})
elif domain == "essential-services":
overlays.append({
"framework": "NAIC Model Bulletin on AI in Insurance",
"trigger": "AI in insurance underwriting, pricing, claims, fraud",
"obligations": (
"AI program governance, risk management, third-party AI oversight; "
"documented testing for unfair discrimination."
),
"citation": "NAIC Model Bulletin on the Use of AI by Insurers (2023)",
})
return overlays
def required_controls(profile: Dict[str, Any], eu_classification: Dict[str, Any]) -> List[str]:
"""Return the required-controls checklist based on tier + profile."""
tier = eu_classification.get("tier", "")
controls = []
if tier in ("HIGH", "LIMITED", "MINIMAL"):
controls.extend([
"Eval set with documented success criteria before deployment",
"Monitoring of model output in production (drift, bias, hallucination)",
"Fallback behavior defined for model failure modes",
"Human-in-loop review for high-stakes outputs",
])
if tier == "HIGH":
controls.extend([
"Conformity assessment completed and documented (EU AI Act Art. 43)",
"Registration in EU AI database before deployment (Art. 49)",
"Risk management system documented and maintained (Art. 9)",
"Training data governance: representativeness, accuracy, bias mitigation (Art. 10)",
"Technical documentation per Annex IV maintained throughout lifecycle (Art. 11)",
"Comprehensive logging for traceability (Art. 12)",
"Human oversight design (e.g., stop button, override capability) (Art. 14)",
"Post-market monitoring plan + serious incident reporting (Art. 72)",
"DPIA under GDPR Art. 35 if personal data processed",
])
if tier == "LIMITED":
controls.extend([
"User notification: 'You are interacting with AI' or 'This content is AI-generated'",
"If general-purpose model: publish model card per Art. 53",
])
if profile.get("user_facing"):
controls.append("Public-facing disclosure of AI usage in customer-facing communications")
if profile.get("automation_level") == "automated" and profile.get("decisions_affected") == "consequential":
controls.append("Right-to-explanation / contestation mechanism for affected individuals (GDPR Art. 22)")
return controls
def analyze(profile: Dict[str, Any]) -> Dict[str, Any]:
eu = classify_eu(profile)
us = us_state_triggers(profile)
overlays = industry_overlays(profile)
controls = required_controls(profile, eu)
conformity_required = eu.get("tier") == "HIGH"
return {
"eu_classification": eu,
"us_state_triggers": us,
"industry_overlays": overlays,
"required_controls": controls,
"conformity_assessment_required": conformity_required,
}
def render_text(result: Dict[str, Any], profile: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("AI RISK CLASSIFICATION")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Use case: {profile.get('use_case')}")
lines.append(f" Domain: {profile.get('domain')} | Automation: {profile.get('automation_level')} | Decisions: {profile.get('decisions_affected')}")
lines.append(f" Deploys in EU: {profile.get('deploys_in_eu')} | US states: {', '.join(profile.get('deploys_in_us_states', []))}")
lines.append(f" User-facing: {profile.get('user_facing')} | Biometric: {profile.get('biometric_data_processed')}")
lines.append("")
lines.append("-" * 72)
eu = result["eu_classification"]
tier_marker = {
"PROHIBITED": "🔴",
"HIGH": "🟠",
"LIMITED": "🟡",
"MINIMAL": "🟢",
"NOT_APPLICABLE": "⚪",
}.get(eu["tier"], "•")
lines.append(f"EU AI ACT TIER: {tier_marker} {eu['tier']}")
lines.append("")
for line in _wrap(eu["reasoning"], 2):
lines.append(line)
lines.append("")
if eu["citations"]:
lines.append(f" Citations: {', '.join(eu['citations'])}")
lines.append("")
if eu["obligations"]:
lines.append(" EU obligations:")
for o in eu["obligations"]:
lines.append(f" • {o}")
lines.append("")
lines.append("-" * 72)
lines.append(f"CONFORMITY ASSESSMENT REQUIRED: {'YES' if result['conformity_assessment_required'] else 'no'}")
lines.append("")
lines.append("-" * 72)
us = result["us_state_triggers"]
if us:
lines.append(f"US STATE LAW TRIGGERS ({len(us)}):")
lines.append("")
for t in us:
lines.append(f" • {t['law']}")
lines.append(f" Trigger: {t['trigger']}")
for line in _wrap(t["obligations"], 4):
lines.append(line)
lines.append(f" Citation: {t['citation']}")
lines.append("")
else:
lines.append("US STATE LAW TRIGGERS: none for the listed states + domain.")
lines.append("")
lines.append("-" * 72)
overlays = result["industry_overlays"]
if overlays:
lines.append(f"INDUSTRY OVERLAYS ({len(overlays)}):")
lines.append("")
for o in overlays:
lines.append(f" • {o['framework']}")
lines.append(f" Trigger: {o['trigger']}")
for line in _wrap(o["obligations"], 4):
lines.append(line)
lines.append(f" Citation: {o['citation']}")
lines.append("")
lines.append("-" * 72)
lines.append(f"REQUIRED CONTROLS ({len(result['required_controls'])}):")
for c in result["required_controls"]:
lines.append(f" ☐ {c}")
lines.append("")
lines.append("-" * 72)
lines.append("REMINDER: This is triage, not legal advice. EU AI Act conformity assessment requires qualified")
lines.append("AI counsel and may require Notified Body involvement. Re-run quarterly as regulations evolve.")
return "\n".join(lines)
def _wrap(text: str, indent: int, width: int = 70) -> List[str]:
import textwrap
return textwrap.wrap(text, width=width, initial_indent=" " * indent, subsequent_indent=" " * indent) or [" " * indent + text]
def main() -> int:
parser = argparse.ArgumentParser(
description="Classify an AI use case under EU AI Act + US state laws.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to use_case JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
profile = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
profile = SAMPLE
source = "<embedded sample: AI hiring screening, EU + NY/CO/IL/CA>"
result = analyze(profile)
if args.output == "json":
print(json.dumps({"source": source, "profile": profile, **result}, indent=2))
else:
print(render_text(result, profile, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/model_buildvsbuy_calculator.py
#!/usr/bin/env python3
"""model_buildvsbuy_calculator.py — Decide API vs fine-tune vs build for a use case.
Stdlib-only. Takes a use case profile and outputs:
- Recommendation (API / FINE_TUNE / BUILD) with reasoning
- 3-year TCO comparison across all 3 paths
- Breakeven analysis (where API stops being cheapest)
- Failure modes for the chosen path
Deterministic logic derived from the profile.
Input schema (JSON):
{
"use_case": "Customer support response generation",
"expected_qps": 5, # queries per second peak
"monthly_volume_queries": 4000000, # queries per month
"avg_tokens_in": 800,
"avg_tokens_out": 200,
"latency_budget_ms": 2000,
"accuracy_required": "frontier", # frontier | high | acceptable
"domain_specific": false, # need specific vocabulary / format / behavior
"data_for_finetune_available": false, # do we have labeled data for fine-tune?
"team_ml_capacity_engineers": 1,
"compliance_requires_self_host": false # data residency / sovereignty constraint
}
Usage:
python model_buildvsbuy_calculator.py # uses embedded customer-support sample
python model_buildvsbuy_calculator.py path/to/use_case.json
python model_buildvsbuy_calculator.py use_case.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Tuple
SAMPLE: Dict[str, Any] = {
"use_case": "Customer support response generation (B2B SaaS)",
"expected_qps": 5,
"monthly_volume_queries": 4_000_000,
"avg_tokens_in": 800,
"avg_tokens_out": 200,
"latency_budget_ms": 2000,
"accuracy_required": "high",
"domain_specific": True,
"data_for_finetune_available": False,
"team_ml_capacity_engineers": 1,
"compliance_requires_self_host": False,
}
# 2026 API pricing per million tokens, $USD (input / output). These are illustrative;
# real pricing changes; rerun this calculator quarterly.
API_PRICING = {
"frontier-premium": {"input": 3.00, "output": 15.00, "label": "Claude Sonnet 4.6 / GPT-4o-tier"},
"frontier-economy": {"input": 1.25, "output": 5.00, "label": "Gemini 2.5 Flash / Claude Haiku 4.5-tier"},
"open-router-hosted": {"input": 0.50, "output": 1.50, "label": "Llama 3.1 70B / Qwen 2.5 72B via hosted endpoint"},
}
# Fine-tune cost (one-time + ongoing)
FINETUNE_ONE_TIME = 25_000 # data prep + initial training + eval harness
FINETUNE_ANNUAL_RETRAIN = 15_000 # quarterly retraining + ops
FINETUNE_INFERENCE_PER_M = 0.40 # cost per M tokens at moderate scale on hosted endpoint
# Self-hosted inference cost (per million tokens, including GPU + ops at 70% utilization)
SELF_HOSTED_PER_M = {
"7b-13b": 0.15,
"70b-class": 1.50,
"frontier-class": 12.00, # very expensive without massive scale; included for completeness
}
# Build-from-scratch cost (one-time + ongoing) — illustrative; usually NOT recommended
BUILD_FROM_SCRATCH_ONE_TIME = 8_000_000
BUILD_FROM_SCRATCH_ANNUAL = 3_000_000
def compute_api_cost_3yr(profile: Dict[str, Any], tier: str) -> float:
"""3-year API cost given workload."""
monthly_queries = profile.get("monthly_volume_queries", 0)
tokens_in = profile.get("avg_tokens_in", 0)
tokens_out = profile.get("avg_tokens_out", 0)
monthly_input_tokens_m = (monthly_queries * tokens_in) / 1_000_000
monthly_output_tokens_m = (monthly_queries * tokens_out) / 1_000_000
pricing = API_PRICING.get(tier, API_PRICING["frontier-premium"])
monthly_cost = (
monthly_input_tokens_m * pricing["input"]
+ monthly_output_tokens_m * pricing["output"]
)
return monthly_cost * 36 # 3 years
def compute_finetune_cost_3yr(profile: Dict[str, Any]) -> float:
monthly_queries = profile.get("monthly_volume_queries", 0)
tokens_total = profile.get("avg_tokens_in", 0) + profile.get("avg_tokens_out", 0)
monthly_tokens_m = (monthly_queries * tokens_total) / 1_000_000
monthly_inference = monthly_tokens_m * FINETUNE_INFERENCE_PER_M
annual_inference = monthly_inference * 12
return FINETUNE_ONE_TIME + (annual_inference + FINETUNE_ANNUAL_RETRAIN) * 3
def compute_self_hosted_cost_3yr(profile: Dict[str, Any], model_class: str) -> float:
"""3-year self-hosted cost including GPU + ops."""
monthly_queries = profile.get("monthly_volume_queries", 0)
tokens_total = profile.get("avg_tokens_in", 0) + profile.get("avg_tokens_out", 0)
monthly_tokens_m = (monthly_queries * tokens_total) / 1_000_000
per_m = SELF_HOSTED_PER_M.get(model_class, SELF_HOSTED_PER_M["70b-class"])
monthly_inference = monthly_tokens_m * per_m
# Add fixed ops cost: 1 engineer * 30% load * fully-loaded $250K/yr = $75K/yr ops attribution
annual_ops = 75_000
return (monthly_inference * 36) + (annual_ops * 3)
def compute_build_cost_3yr() -> float:
return BUILD_FROM_SCRATCH_ONE_TIME + (BUILD_FROM_SCRATCH_ANNUAL * 3)
def pick_recommendation(profile: Dict[str, Any], costs: Dict[str, float]) -> Tuple[str, str, List[str]]:
"""Pick API / FINE_TUNE / BUILD with reasoning and failure modes."""
accuracy = profile.get("accuracy_required", "high")
domain_specific = profile.get("domain_specific", False)
finetune_data = profile.get("data_for_finetune_available", False)
ml_capacity = profile.get("team_ml_capacity_engineers", 0)
self_host_required = profile.get("compliance_requires_self_host", False)
latency_ms = profile.get("latency_budget_ms", 2000)
monthly_q = profile.get("monthly_volume_queries", 0)
# Special case: compliance forces self-host
if self_host_required:
return (
"FINE_TUNE",
(
"Compliance / data residency forces self-host. Fine-tune a 70B-class open model "
f"({_fmt_money(costs['finetune_3yr'])}/3yr) rather than build from scratch "
f"({_fmt_money(costs['build_3yr'])}/3yr) — the gap is two orders of magnitude with "
"comparable quality for most use cases."
),
[
"Quality lags frontier by ~6 months; budget for refresh every 12-18mo",
"Self-hosting requires 24/7 on-call; budget 30%+ of an engineer FTE",
"Eval discipline becomes non-negotiable; without an eval set you cannot tell when retraining is needed",
],
)
# Build from scratch — almost never
if accuracy == "frontier" and monthly_q > 1_000_000_000 and ml_capacity >= 20:
return (
"BUILD",
(
"Edge case where frontier accuracy + extreme volume + large ML team justify pre-training. "
"Cost still extreme. Most companies here are foundation-model startups, not application companies."
),
[
"By the time you ship, frontier models have caught up — sunk cost risk",
"Requires sustained $50M+ investment over 18+ months",
"Unless model IS your product, do not build",
],
)
# Fine-tune cases
if domain_specific and finetune_data and ml_capacity >= 2:
return (
"FINE_TUNE",
(
"Domain-specific behavior + labeled data + ML engineering capacity available. "
f"Fine-tune cost ({_fmt_money(costs['finetune_3yr'])}) competes with API at this volume."
),
[
"Fine-tuned model lags frontier by ~6 months; quality drift is inevitable",
"Retraining cadence (quarterly typical) is a recurring engineering cost",
"Without eval set, fine-tune drift is invisible until customer complains",
],
)
# Latency-driven fine-tune (sub-500ms with 70B-class)
if latency_ms < 500 and monthly_q > 1_000_000:
return (
"FINE_TUNE",
(
f"Latency budget {latency_ms}ms below frontier-API median (~600-1500ms). "
"Fine-tuned 70B-class on dedicated infra is the path to sub-500ms at scale."
),
[
"Sub-500ms requires GPU co-location and warm pools (idle time penalty)",
"Quality must be re-verified at every model swap",
"Streaming responses can buy headroom on latency budget; consider before committing to fine-tune",
],
)
# Default to API for everything else
economy_acceptable = accuracy in ("acceptable", "high")
if economy_acceptable and costs["api_economy_3yr"] < costs["finetune_3yr"]:
return (
"API",
(
f"Frontier-economy API tier ({API_PRICING['frontier-economy']['label']}) at "
f"{_fmt_money(costs['api_economy_3yr'])}/3yr beats fine-tune ({_fmt_money(costs['finetune_3yr'])}/3yr). "
"Iterate on prompt engineering and eval discipline before committing to fine-tune."
),
[
"Vendor lock-in: build abstraction layer (LiteLLM, OpenRouter) for multi-vendor failover",
"Capability drift between model versions: pin model IDs and run regression evals on upgrades",
"Rate limits at QPS spikes: confirm Tier-4+ pricing with provider",
],
)
return (
"API",
(
f"Frontier-premium API at {_fmt_money(costs['api_premium_3yr'])}/3yr is the right starting point. "
"Revisit fine-tune at ≥10M queries/month OR domain-specific behavior the API can't be prompted into."
),
[
"Vendor lock-in: build abstraction layer for multi-vendor failover",
"Capability drift between model versions; pin model IDs",
"Rate limits at QPS spikes; confirm pricing tier with provider",
],
)
def analyze(profile: Dict[str, Any]) -> Dict[str, Any]:
costs = {
"api_premium_3yr": compute_api_cost_3yr(profile, "frontier-premium"),
"api_economy_3yr": compute_api_cost_3yr(profile, "frontier-economy"),
"api_open_hosted_3yr": compute_api_cost_3yr(profile, "open-router-hosted"),
"finetune_3yr": compute_finetune_cost_3yr(profile),
"self_hosted_70b_3yr": compute_self_hosted_cost_3yr(profile, "70b-class"),
"build_3yr": compute_build_cost_3yr(),
}
recommendation, reasoning, failure_modes = pick_recommendation(profile, costs)
# Compute breakeven volume where API and fine-tune cross
monthly_q = profile.get("monthly_volume_queries", 1)
tokens_per_q = profile.get("avg_tokens_in", 0) + profile.get("avg_tokens_out", 0)
annual_q = monthly_q * 12
# Find breakeven where API economy total == fine-tune total over 3 years
if tokens_per_q and annual_q:
api_economy_per_query = costs["api_economy_3yr"] / (annual_q * 3) if annual_q else 0
# finetune_cost = ONE_TIME + (queries * tokens * inference_per_m / 1M + ANNUAL_RETRAIN) * 3
# Solve for queries where api_cost == finetune_cost
# api_economy_per_query * Q = FINETUNE_ONE_TIME + (Q * tokens_per_q * FINETUNE_INFERENCE_PER_M / 1M + RETRAIN) * 3
# api_economy_per_query * Q - 3 * Q * tokens_per_q * FINETUNE_INFERENCE_PER_M / 1M = FINETUNE_ONE_TIME + 3 * RETRAIN
# Q * (api_economy_per_query - 3 * tokens_per_q * FINETUNE_INFERENCE_PER_M / 1M) = ONE_TIME + 3 * RETRAIN
coefficient = (
api_economy_per_query
- 3 * tokens_per_q * FINETUNE_INFERENCE_PER_M / 1_000_000
)
rhs = FINETUNE_ONE_TIME + 3 * FINETUNE_ANNUAL_RETRAIN
breakeven_3yr_queries = int(rhs / coefficient) if coefficient > 0 else None
breakeven_monthly_queries = int(breakeven_3yr_queries / 36) if breakeven_3yr_queries else None
else:
breakeven_monthly_queries = None
return {
"recommendation": recommendation,
"reasoning": reasoning,
"failure_modes": failure_modes,
"costs_3yr_usd": {k: round(v, 0) for k, v in costs.items()},
"breakeven_monthly_queries_api_vs_finetune": breakeven_monthly_queries,
"current_monthly_volume": profile.get("monthly_volume_queries", 0),
}
def render_text(result: Dict[str, Any], profile: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("MODEL BUILD-VS-BUY ANALYSIS")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Use case: {profile.get('use_case')}")
lines.append(f" Volume: {profile.get('monthly_volume_queries'):,} queries/mo @ {profile.get('expected_qps')} QPS peak")
lines.append(f" Tokens: {profile.get('avg_tokens_in')} in / {profile.get('avg_tokens_out')} out per query")
lines.append(f" Latency budget: {profile.get('latency_budget_ms')}ms | Accuracy: {profile.get('accuracy_required')}")
lines.append(f" Domain-specific: {profile.get('domain_specific')} | Fine-tune data available: {profile.get('data_for_finetune_available')}")
lines.append(f" ML capacity: {profile.get('team_ml_capacity_engineers')} engineers | Compliance forces self-host: {profile.get('compliance_requires_self_host')}")
lines.append("")
lines.append("-" * 72)
lines.append(f"RECOMMENDATION: {result['recommendation']}")
lines.append("")
for line in _wrap(result["reasoning"], 2):
lines.append(line)
lines.append("")
lines.append("Failure modes to plan for:")
for fm in result["failure_modes"]:
lines.append(f" • {fm}")
lines.append("")
lines.append("-" * 72)
lines.append("3-YEAR TCO COMPARISON ($ USD):")
lines.append("")
costs = result["costs_3yr_usd"]
lines.append(f" API (frontier-premium, {API_PRICING['frontier-premium']['label']}): {_fmt_money(costs['api_premium_3yr']):>15}")
lines.append(f" API (frontier-economy, {API_PRICING['frontier-economy']['label']}): {_fmt_money(costs['api_economy_3yr']):>15}")
lines.append(f" API (open-router-hosted, {API_PRICING['open-router-hosted']['label']}): {_fmt_money(costs['api_open_hosted_3yr']):>15}")
lines.append(f" Fine-tune (70B-class, hosted inference): {_fmt_money(costs['finetune_3yr']):>15}")
lines.append(f" Self-hosted (70B-class on rented H100/A100): {_fmt_money(costs['self_hosted_70b_3yr']):>15}")
lines.append(f" Build from scratch (pre-train + ops): {_fmt_money(costs['build_3yr']):>15}")
lines.append("")
if result["breakeven_monthly_queries_api_vs_finetune"]:
lines.append(f"Breakeven: API (economy) vs fine-tune crosses at ~{result['breakeven_monthly_queries_api_vs_finetune']:,} queries/month")
if result["current_monthly_volume"] < result["breakeven_monthly_queries_api_vs_finetune"]:
lines.append(f" Current volume ({result['current_monthly_volume']:,}/mo) is BELOW breakeven → API still cheaper.")
else:
lines.append(f" Current volume ({result['current_monthly_volume']:,}/mo) is ABOVE breakeven → fine-tune economics favorable.")
lines.append("")
lines.append("-" * 72)
lines.append("REMINDER: TCO does not capture quality cost. Fine-tune quality lags frontier by ~6 months;")
lines.append("self-hosted requires eval discipline you may not have. Re-run quarterly with updated pricing.")
return "\n".join(lines)
def _fmt_money(amount: float) -> str:
return f",.0f"
def _wrap(text: str, indent: int, width: int = 70) -> List[str]:
import textwrap
return textwrap.wrap(text, width=width, initial_indent=" " * indent, subsequent_indent=" " * indent) or [" " * indent + text]
def main() -> int:
parser = argparse.ArgumentParser(
description="Decide API vs fine-tune vs build with 3-year TCO comparison.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to use_case JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
profile = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
profile = SAMPLE
source = "<embedded sample: B2B SaaS customer-support generation, 4M queries/mo>"
result = analyze(profile)
if args.output == "json":
print(json.dumps({"source": source, "profile": profile, **result}, indent=2))
else:
print(render_text(result, profile, source))
return 0
if __name__ == "__main__":
sys.exit(main())
Tư vấn Chief Data Officer: quyền dữ liệu huấn luyện AI, chiến lược sản phẩm dữ liệu, định giá dữ liệu khách hàng và nhân sự.
---
name: "chief-data-officer-advisor"
description: "Chief Data Officer advisory for startups: AI training data rights and consent provenance, data product strategy (warehouse vs lakehouse vs mesh, build-vs-buy), B2B customer-data-as-asset valuation and M&A readiness, data team org evolution. Use when deciding whether to train models on customer data, choosing data architecture, valuing data for fundraising or M&A, sequencing data hires, or when user mentions CDO, chief data officer, data strategy, data mesh, lakehouse, training data, data product, data monetization, or customer data asset. NOT a tactical data engineering skill — strategic decisions only."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: chief-data-officer-leadership
updated: 2026-05-12
python-tools: ai_training_data_audit.py, data_product_strategy_picker.py, data_asset_valuator.py
frameworks: training-data-rights-matrix, data-product-strategy, customer-data-as-asset, data-team-org-evolution
---
# Chief Data Officer Advisor
Strategic data leadership for startup CDOs and founders without one. **Four decisions, no surveys:**
1. **Can we train our model on this data?** — origin × consent × use-case matrix
2. **Warehouse, lakehouse, or mesh — and what do we build vs buy?** — stage-driven architecture
3. **What is our customer data worth?** — strategic value + M&A multiplier + productization paths
4. **What data role do we hire next?** — stage-to-role map, centralize-vs-embed trigger
This skill does **not** cover tactical data engineering. For schema design, observability, query optimization, RAG, or ML platform implementation, see `engineering/database-designer/`, `engineering/observability-designer/`, `engineering/data-quality-auditor/`, `engineering/sql-database-assistant/`, `engineering/rag-architect/`, `engineering/llm-cost-optimizer/`.
## Keywords
CDO, chief data officer, AI training data, consent provenance, training rights, GDPR Article 6 lawful basis, GDPR Article 22, EU AI Act high-risk, ePrivacy, copyright fair use, hiQ v. LinkedIn, scraped data, synthetic data, data product, data mesh, lakehouse, medallion architecture, dbt, Snowflake, BigQuery, Databricks, Fivetran, Airbyte, reverse ETL, feature store, customer data as asset, data monetization, data productization, anonymization, k-anonymity, differential privacy, M&A data diligence, data org, analytics engineer, data engineer, data scientist, data product manager, centralize vs embed, hub and spoke
## Quick Start
```bash
# Audit data sources for AI training eligibility
python scripts/ai_training_data_audit.py # uses embedded sample
python scripts/ai_training_data_audit.py path/to/sources.json
# Pick data architecture + build-vs-buy + sequencing
python scripts/data_product_strategy_picker.py # uses embedded Series A SaaS
python scripts/data_product_strategy_picker.py path/to/profile.json
# Value the customer data corpus + productization viability
python scripts/data_asset_valuator.py # uses embedded B2B sample
python scripts/data_asset_valuator.py path/to/corpus.json
```
## Key Questions (ask these first)
- **What decision does this data drive?** (If none, why are we collecting it?)
- **What's the consent provenance of every source we want to train on?** (TOS-only is not the same as explicit opt-in.)
- **Who are the internal data consumers, and how many distinct domains do they span?** (Drives centralize-vs-embed and warehouse-vs-mesh.)
- **In an M&A scenario, is our data a moat or a liability?** (Customer carve-outs in MSAs can flip the answer.)
- **Are we hiring an analytics engineer or a data scientist next?** (They solve different problems; founders confuse them.)
- **Have we run an anonymization audit before any external sharing?** (k-anonymity ≥ 5 is the floor, not the ceiling.)
## Core Responsibilities
### 1. AI Training Data Rights
The 2026 question every startup is facing: **can we use customer data to train our model?**
The answer is rarely binary. It depends on three independent dimensions:
| Dimension | Values |
|---|---|
| **Origin** | 1st-party-explicit-opt-in / 1st-party-TOS-only / partner-licensed / scraped / synthetic |
| **Data class** | Anonymous aggregate / behavioral / PII / 3rd-party content / regulated (PHI, PCI, kids) |
| **Use case** | In-product personalization / fine-tune our model / train foundation model / external sharing |
Each combination produces GO / MITIGATE / NO-GO. **Run** `ai_training_data_audit.py` on a JSON inventory of sources.
See `references/ai_training_data_rights.md` for the full matrix + GDPR Art. 6 lawful basis decision tree + EU AI Act high-risk triggers.
### 2. Data Product Strategy
**Architecture choice (warehouse vs lakehouse vs mesh) is stage-driven, not preference-driven:**
- **Warehouse only** (Snowflake / BigQuery / Postgres): ≤5 data consumers, <2TB, no ML use cases
- **Lakehouse** (warehouse + object storage, often Databricks or Snowflake-with-Iceberg): 5–25 data consumers, 2TB–1PB, 1–3 ML use cases
- **Data mesh**: 25+ data consumers across 4+ domains, federated ownership culture in place
**Build vs buy is decided per layer:**
| Layer | Buy unless | Build only if |
|---|---|---|
| Storage / warehouse | Never build | (You’re a data infra company) |
| ELT / ingest | Never build | Source isn’t supported by Fivetran/Airbyte |
| Modeling (dbt) | Always build | This is your IP |
| BI / dashboards | Buy at <100 consumers | Embedded analytics for customers |
| Feature store | Defer until 3+ prod models | Then build OR buy Tecton/Hopsworks |
| ML platform | Defer until 5+ prod models | Then buy SageMaker/Vertex/Databricks |
**Run** `data_product_strategy_picker.py` for a stage-specific recommendation. See `references/data_product_strategy.md` for kill criteria per architecture and the build-vs-buy decision tree.
### 3. B2B Customer-Data-as-Asset
**The shift:** at Series B+, customer data is no longer just operational — it’s an asset that can be:
- A defensibility moat (replicating requires years of customer cohort)
- An M&A multiplier (1.2x–2x ARR uplift for strategic buyers)
- A direct revenue stream (anonymized industry benchmarks, embedding endpoints, licensing)
But it can also be a **liability**:
- 47/380 customers with MSA carve-outs makes productization legally infeasible
- Anonymization audits often reveal re-identification risk above tolerable thresholds
- Regulatory exposure increases linearly with productization (GDPR Art. 28 processors vs Art. 26 joint controllers)
**Run** `data_asset_valuator.py` with corpus characteristics to get strategic value score + productization paths + risk-adjusted value.
See `references/customer_data_as_asset.md` for the valuation framework, M&A diligence prep checklist, and contractual constraint audit pattern.
### 4. Data Team Org Evolution
**The wrong question:** "Should we hire a data scientist?"
**The right question:** "What’s the next decision we can’t make because we lack data, and what role unblocks that?"
Stage-to-role map (B2B SaaS baseline):
| Stage | First hire | Then | Then |
|---|---|---|---|
| Pre-seed / seed | Founder-as-analyst (SQL + spreadsheets) | — | — |
| Series A (Series A) | Analyst | Analytics engineer (dbt) | — |
| Series B | Data engineer | Senior analyst (embedded in GTM) | Data PM (if 3+ teams need data) |
| Growth | Manager of analytics | ML engineer (if model is core) | Head of Data |
| Late-stage | Head of Data → CDO | Specialized: BI, MLE, DPO | Federated owners per domain (mesh) |
**Centralize-vs-embed trigger:** when 3+ functional areas (sales, marketing, product, ops, CS) need bespoke data weekly, the central team becomes the bottleneck. Move to hub-and-spoke (central platform + embedded analysts) before that becomes a hiring crisis.
See `references/data_team_org_evolution.md`.
## Workflows
### Workflow 1: AI Training Decision (1 hour)
**Goal:** Decide whether a specific data source can train a specific use case.
```bash
# 1. Build sources.json with one entry per data source
# 2. Run the audit
python scripts/ai_training_data_audit.py sources.json
# 3. For each MITIGATE: assign owner + remediation
# 4. For each NO-GO: document the kill reason for the legal log
# 5. Cross-check with cs-general-counsel-advisor on top-3 mitigation items
# 6. Log via /cs:decide
```
### Workflow 2: Architecture Decision (1 day)
**Goal:** Pick warehouse / lakehouse / mesh and the build-vs-buy split for the next 12 months.
```bash
python scripts/data_product_strategy_picker.py profile.json
# Cross-check with cs-cto-advisor on engineering capacity
# Cross-check with cs-cfo-advisor on 3-year TCO
# Log via /cs:decide; consider /cs:freeze 90 if signing a multi-year SaaS contract
```
### Workflow 3: Data Asset Valuation for M&A Prep (3 days)
**Goal:** Value the data corpus and prepare for due diligence.
1. Inventory the corpus: size, freshness, exclusivity, customer overlap, contractual restrictions
2. Run `data_asset_valuator.py`
3. Run the M&A diligence prep checklist in `customer_data_as_asset.md`
4. Surface contractual carve-outs to cs-general-counsel-advisor for re-papering plan
5. Decide productization path (benchmark report / embedding endpoint / direct license)
6. Log via /cs:decide
### Workflow 4: Data Team Roadmap (1 week)
**Goal:** Build the next 18 months of data hires aligned to business decisions.
1. List the top 5 decisions the business can’t make today due to missing data or analysis
2. Map each decision to the role that unblocks it
3. Sequence hires (one role at a time, ramp before next)
4. Cross-check with cs-chro-advisor on comp bands and leveling
5. Identify the centralize-vs-embed trigger date
## Output Standards (when invoked via cs-cdo-advisor)
```
**Bottom Line:** [one sentence — decision and rationale]
**The Decision:** [one of the 4 framings]
**The Evidence:** [numbers, not adjectives]
**How to Act:** [3 concrete next steps]
**Your Decision:** [the call only the founder can make]
```
## Adjacent Skills
- `../cto-advisor/` — architecture capacity, scaling cliffs
- `../ciso-advisor/` — data security, threat modeling for productized data
- `../general-counsel-advisor/` — contractual constraints, DPA, training-data rights
- `../cfo-advisor/` — build-vs-buy TCO, M&A valuation math
- `../chro-advisor/` — data team hiring, leveling, comp
- `../../../engineering/database-designer/` — tactical schema design
- `../../../engineering/rag-architect/` — tactical AI/RAG implementation
- `../../../engineering/llm-cost-optimizer/` — model cost management
## References
- [ai_training_data_rights.md](references/ai_training_data_rights.md) — The training-rights matrix + GDPR Art. 6 / EU AI Act decision tree
- [data_product_strategy.md](references/data_product_strategy.md) — Warehouse / lakehouse / mesh kill criteria + build-vs-buy decision tree
- [customer_data_as_asset.md](references/customer_data_as_asset.md) — Valuation framework + M&A diligence prep + productization paths
- [data_team_org_evolution.md](references/data_team_org_evolution.md) — Stage-to-role map + centralize-vs-embed trigger
---
**Version:** 1.0.0
**Status:** Production Ready
**Disclaimer:** Decisions touching training data rights, data productization, or M&A data diligence should involve qualified counsel. This skill surfaces decisions and tradeoffs — it does not replace legal review.
FILE:references/ai_training_data_rights.md
# AI Training Data Rights — The Decision: "Can we train on this data?"
This reference answers exactly one decision per data source: **may we use this for AI training, and for which use case?** It does so by combining three independent dimensions into a verdict.
Pair with `scripts/ai_training_data_audit.py` for automation. **Not legal advice.**
## The Three Dimensions
### Dimension 1: Origin
Where did this data come from, and what consent flow accompanied it?
| Origin | Strength | Notes |
|---|---|---|
| `1st-party-explicit-opt-in` | Strongest | User saw a notice for THIS purpose and clicked agree. GDPR Art. 6(1)(a). |
| `1st-party-tos-only` | Weak | Bundled TOS doesn't satisfy GDPR Art. 6 for materially different purposes (training). |
| `partner-licensed` | Depends | Only as strong as the partner's original consent flow + your license scope. |
| `scraped` | Insufficient | No lawful basis under GDPR Art. 6; potentially Computer Fraud and Abuse Act / copyright exposure. |
| `synthetic` | Strong | But synthetic data inherits risks from its seed source if any. |
### Dimension 2: Data Class
What's in the data?
| Class | Implication |
|---|---|
| `anonymous-aggregate` | Safest. K-anonymity ≥ 5 maintained. |
| `behavioral` | Usually safe with proper consent. Watch for re-identification. |
| `pii` | Highest scrutiny. Requires lawful basis + deletion-on-request handling. |
| `third-party-content` | User-uploaded files, snippets, transcripts that include external content. Copyright + DMCA exposure. |
| `regulated` | PHI, PCI, COPPA-children data, biometrics. Framework-specific consent required. |
### Dimension 3: Use Case
What are you doing with it?
| Use case | Risk profile |
|---|---|
| `in-product-personalization` | Lowest risk; recommended within-product. Performance of contract often covers this. |
| `fine-tune-our-model` | Medium risk. Specific opt-in usually needed for non-anonymous classes. |
| `train-foundation-model` | High risk. Re-identification + memorization concerns; almost never permissible for PII without specific consent. |
| `external-sharing` | Highest risk. Recipient becomes a data controller (GDPR Art. 26 / 28 analysis required). |
## The Verdict Matrix (excerpt — full logic in audit tool)
| Origin × Class × Use Case | Verdict |
|---|---|
| `scraped` × any × any | NO-GO (no exceptions for training) |
| `1st-party-tos-only` × `pii` × `fine-tune-our-model` | NO-GO (TOS insufficient for material purpose change) |
| `1st-party-explicit-opt-in` × `pii` × `in-product-personalization` | GO (strongest position) |
| `1st-party-tos-only` × `behavioral` × `fine-tune-our-model` | GO (with DPIA + deletion handling) |
| `partner-licensed` × `anonymous-aggregate` × `train-foundation-model` | GO (with license-scope review) |
| `synthetic` × `anonymous-aggregate` × `train-foundation-model` | GO (with provenance log) |
| any × `regulated` × `train-foundation-model` | NO-GO (framework prohibits raw use) |
Run `python scripts/ai_training_data_audit.py` for the full matrix applied to your sources.
## GDPR Art. 6 Lawful Basis Decision Tree (EU residents only)
If any EU resident data flows, GDPR applies. Pick exactly one lawful basis per purpose:
1. **Art. 6(1)(a) Consent.** The user said yes to THIS specific purpose. Most defensible. Must be granular, freely given, revocable.
2. **Art. 6(1)(b) Performance of contract.** Processing is necessary to deliver the service the user purchased. Works for in-product personalization within reasonable expectations.
3. **Art. 6(1)(c) Legal obligation.** You're required by law. Rare for training data.
4. **Art. 6(1)(d) Vital interests.** Life or death. Practically never applies to AI training.
5. **Art. 6(1)(e) Public interest.** Government / public mission. Rarely applies to private companies.
6. **Art. 6(1)(f) Legitimate interest.** Balancing test: your interest vs the user's rights. Requires Legitimate Interest Assessment (LIA). Defensible for fraud detection, security; weak for personalization beyond user expectations.
**Practical takeaway:** For training data outside in-product personalization, default to Art. 6(1)(a) explicit consent. Art. 6(1)(f) is increasingly disfavored by EU regulators for AI training (see EDPB Opinion 28/2024).
## EU AI Act High-Risk Triggers
The EU AI Act (in force 2026) imposes additional data governance requirements for high-risk AI systems. You are high-risk if your AI is used for:
- Biometric identification (other than verification)
- Critical infrastructure management
- Education access / scoring
- Employment / worker management (including hiring algorithms)
- Access to essential services (credit, insurance, public benefits)
- Law enforcement
- Migration / border control
- Administration of justice
If you are high-risk, **Art. 10 (data governance)** requires:
- Training-data quality criteria (representativeness, accuracy, completeness)
- Bias examination + mitigation
- Provenance documentation per source
- Pre-deployment conformity assessment
If you're low-risk (most B2B SaaS), the heavy obligations are GDPR-side, not AI-Act-side. But you still need provenance logs for Art. 53 (general-purpose models).
## US State Patchwork
| Law | What it covers |
|---|---|
| California CCPA / CPRA | Right to know, delete, opt-out of sale (incl. some training scenarios) |
| Colorado AI Act (CO SB 21-169 successor) | Bias audit requirements for AI in consumer decisions |
| New York City Local Law 144 | Bias audit required for AI in hiring (NYC employers) |
| Illinois BIPA | Biometric data requires explicit written consent |
| Texas TCPA | Capture-of-biometric-identifier rules |
| Washington My Health My Data Act | Consumer health data including inference |
## Practical Decision Pattern
For every new AI training initiative:
1. **List the data sources you plan to use** (be exhaustive — including "internal" ones)
2. **Tag each with origin × class × use case**
3. **Run `ai_training_data_audit.py`**
4. **For NO-GO:** Document the kill reason in the legal log. Either drop the source or change the use case.
5. **For MITIGATE:** Assign owner + remediation. Block training until complete.
6. **For GO:** Document the lawful basis and maintain the provenance log.
7. **Cross-check with cs-general-counsel-advisor** on top-3 mitigation items.
8. **Cross-check with cs-ciso-advisor** on data flow security.
9. **Log the decision via `/cs:decide`.**
## When This Reference Doesn't Help
- **Building synthetic data pipelines.** The synthetic data origin tag covers strategy, not generation; talk to engineering.
- **Differential privacy implementations.** Engineering territory. See `engineering/database-designer/` for guidance.
- **EU AI Act conformity assessments.** Requires a specialist; this reference identifies the trigger, not the remediation.
- **Class actions / litigation defense.** Outside counsel territory; this reference is preventive.
---
**Source authorities (non-exhaustive):**
- GDPR (Regulation (EU) 2016/679)
- EU AI Act (Regulation (EU) 2024/1689)
- EDPB Opinion 28/2024 on processing of personal data in AI models
- CCPA / CPRA (California Civil Code § 1798.100 et seq.)
- hiQ Labs, Inc. v. LinkedIn Corp., 938 F.3d 985 (9th Cir. 2019)
- NYT Co. v. OpenAI (filing, 2024, ongoing)
- Authors Guild v. Google, 804 F.3d 202 (2d Cir. 2015)
FILE:references/customer_data_as_asset.md
# Customer Data as Asset — The Decision: "What is our customer data worth, and can we productize it?"
This reference answers exactly one decision: **at Series B+, when customer data is no longer operational but strategic, how do we value it, monetize it, and survive M&A diligence?**
Pair with `scripts/data_asset_valuator.py` for automation.
## The Shift: Operational → Strategic Asset
In seed and Series A, customer data is operational: it powers the product. Starting around Series B (especially in B2B SaaS), data accumulates into something else — an asset with strategic value independent of the product's primary use.
Symptoms that the shift has happened:
- An acquirer asks about data corpus in their LOI
- A partner asks to license anonymized data for benchmarking
- A customer demands a contractual carve-out preventing data use beyond their own service
- The board asks "what are we doing with the data?"
When these surface, you need a CDO answer, not a CTO answer.
## The Valuation Framework — Five Components
Strategic value (composite score 0-10) is the product of five components:
### 1. Exclusivity
**Is the data uniquely yours, or is it available elsewhere?**
| Level | Definition |
|---|---|
| `none` | Same data is in public sources (web scrapes, public records) |
| `low` | Commercially available from data brokers (e.g., LinkedIn / ZoomInfo data) |
| `medium` | Available only via specific platforms (e.g., Stripe transaction data, Slack messages) |
| `high` | No public or commercial equivalent (e.g., your unique customer cohort's workflow behavior) |
**Default for B2B SaaS:** medium-to-high. The combination of customer cohort + your specific product usage is usually exclusive.
### 2. Freshness
**How current is the data?**
Real-time > near-real-time > daily batch > weekly batch. Predictive value decays roughly exponentially with staleness.
### 3. Cohort Breadth
**How many customers does the corpus span?**
Below 50 customers: insufficient cohort for benchmarks. 50–200: marginally productizable. 200–500: solid. 500+: strong.
**Cohort breadth is highly correlated with industry-specific value:** a 500-customer B2B SaaS in vertical X often has more strategic value than a 5000-customer horizontal SaaS, because the verticalized cohort is harder to replicate.
### 4. History Depth
**How many years of time-series do you have?**
1 year is anecdotal. 2–3 years shows trend. 5+ years enables cycle analysis and is increasingly rare (most startups don't survive that long).
History depth is THE thing acquirers value most — and the thing you can't manufacture later.
### 5. Real-Time Behavioral Signal
**Does the data capture intent + behavior, or just outcomes?**
Outcome data ("customer churned") is low signal. Intent + behavior data ("customer reduced usage by 40% in week 8, then opened pricing page 3 times") is high signal.
This component is implicit in the freshness + exclusivity scores in the tool.
## Moat Strength
The composite score maps to moat strength:
| Score | Moat | Defense |
|---|---|---|
| 8+ | STRONG | Replicating requires 2+ years of customer cohort acquisition |
| 5-7 | MEDIUM | Well-funded competitor with 18-24 months can match |
| 2-4 | WEAK | Some unique signal but largely replicable |
| 0-1 | NONE | Same data is freely available |
## M&A Multiplier
Acquirers (especially strategic ones, not financial) pay a multiplier on data-as-asset deals.
| Moat | Multiplier (ARR uplift) |
|---|---|
| STRONG | 1.4x – 1.7x |
| MEDIUM | 1.15x – 1.35x |
| WEAK | 1.0x – 1.1x |
| NONE | 1.0x |
**These multipliers compound with normal SaaS multiples.** A $10M ARR B2B SaaS valued at 8x ARR ($80M) with a STRONG data moat might fetch $112M-$136M in a strategic acquisition where the buyer values the cohort.
**Discounts:**
- High MSA carve-out rate (>25% of customers): -15%
- Moderate carve-out rate (10-25%): -5%
- Failed anonymization audit (re-identification risk): -10%
- Regulated data without specific consent framework: -20%
## The Three Productization Paths
### Path 1: Industry Benchmark Report (lowest risk)
**What it is:** Quarterly or semi-annual report of anonymized aggregates ("80% of B2B sales teams have >5 stalled deals in their pipeline at any time").
**Revenue potential:** Low ($50K-$500K/yr). Often given away to drive credibility / leads rather than sold.
**Why start here:**
- Lowest legal risk (anonymous aggregates, no individual data leaves)
- Highest credibility lift (your brand becomes the "definitive source" for the category)
- Tests appetite without committing to product
- Lowest customer-trust cost (customers like seeing aggregate insights)
**Prerequisites:**
- Anonymization audit confirming k-anonymity ≥ 5 in all published cells
- Opt-out flow for customers who don't want their (anonymized) data included
- Quarterly review cadence
### Path 2: Anonymized Embedding Endpoint (medium risk)
**What it is:** API that returns anonymized embeddings of your data corpus, usable by your customers (or by you) for AI features.
**Revenue potential:** Medium ($500K-$3M/yr) as a platform feature or paid add-on.
**Why medium risk:**
- Embeddings can leak training data via inversion attacks (mitigated by differential privacy)
- 47/380 customer carve-outs would block the endpoint from including their data
- Re-identification of a single customer in the corpus risks contractual + reputational damage
**Prerequisites:**
- Anonymization + memorization testing
- DPA addendum covering training-data flow
- Differential privacy on the embedding pipeline (epsilon ≤ 1.0 recommended)
- Pilot with 3 design-partner customers under explicit opt-in before broad release
### Path 3: Direct Data Licensing (highest risk)
**What it is:** Selling access to the data corpus (or derivatives) to AI labs, data brokers, or industry players.
**Revenue potential:** High ($2M-$20M/yr at scale).
**Why high risk:**
- Customer trust impact: even with proper anonymization, customers often perceive this as "selling our data"
- Requires re-papering or excluding any MSA carve-out customers
- Requires GDPR Art. 26 joint-controller analysis if EU customers are present
- Regulator scrutiny increases (e.g., FTC has signaled interest in B2B-to-AI-lab data flows in 2024-2025)
**Prerequisites (in order):**
1. Customer-trust impact assessment (CEO + Head of CS sign-off)
2. Re-paper carve-out customers OR build carve-out-excluded dataset
3. Engage data broker counsel (specialist)
4. Customer communications plan (proactive, not reactive)
5. Differential privacy on the licensed product
6. Audit clauses in the licensing contract
## M&A Diligence Prep Checklist
Acquirers will dig deep on data assets. Be ready before the LOI.
**6 months before any M&A discussion, complete:**
- [ ] Inventory of all customer data with: origin, consent flow, contractual restrictions, retention policy
- [ ] MSA carve-out audit: which customers have which restrictions; reconciliation list
- [ ] Anonymization audit: k-anonymity, re-identification risk assessment
- [ ] DPA inventory: which customers have DPAs, which subprocessors are listed, gaps
- [ ] Training-data provenance log: every model in production has documented source data
- [ ] Right-to-erasure handling: documented process for honoring GDPR Art. 17 / state law equivalents
- [ ] Cross-border data flow inventory: which EU residents' data is processed, which US states, which countries
- [ ] Vendor / subprocessor list current and reconciled with customer-facing list
- [ ] Data breach history: documented, even minor incidents
- [ ] Litigation / regulatory inquiries: documented
**Common findings that tank deals:**
- "We've been training on X without a clear lawful basis" → acquirer requires indemnity carve-out or retrains
- "We don't have a documented anonymization process" → 10-20% multiplier discount
- "30% of customers have carve-outs we can't easily reconcile" → productization-as-thesis collapses
- "Our DPA list and our customer-facing DPA list don't match" → governance red flag
## Contractual Constraint Audit (run quarterly)
Many startups don't realize their MSA template has been updated 3 times in 5 years, and earlier customers signed earlier versions. The carve-out rate often exceeds expectations.
**Quarterly audit:**
1. Pull every executed customer MSA from CLM (or DocuSign / Ironclad)
2. Search for: "data use", "training", "AI", "machine learning", "aggregate", "anonymized", "license back"
3. Categorize each customer:
- `clear` — no carve-out, standard rights
- `carve-out-aggregate-only` — can use only as anonymized aggregates
- `carve-out-no-training` — can use operationally but not for AI training
- `carve-out-blocked` — cannot use beyond own service
4. Compute carve-out rates
5. For each carve-out type, decide: re-paper at renewal? Live with the constraint? Build carve-out-excluded dataset?
## Customer Trust Considerations
The legal feasibility of productization is necessary but not sufficient. Customer trust impact is often the binding constraint.
**Signs the trust cost will exceed the revenue:**
- Customer NPS is below 30
- Recent press cycle on "Big Tech data abuses" in your category
- A vocal customer or two raised data concerns publicly
- Your sales team uses "we don't share your data" as a competitive differentiator
**If any of these are true:** delay productization 12-18 months and address trust first.
## When This Reference Doesn't Help
- **Tactical anonymization implementation.** See engineering / privacy-engineering resources.
- **Specific DPA template language.** See `c-level-advisor/skills/general-counsel-advisor/`.
- **M&A negotiation strategy.** See `c-level-advisor/skills/ma-playbook/`.
- **GDPR compliance program.** See `ra-qm-team/`.
This reference is about strategic valuation and productization decisions. Tactical execution lives elsewhere.
---
**Source authorities (non-exhaustive):**
- GDPR Articles 26 (joint controllers), 28 (processors), 35 (DPIA), 17 (right to erasure)
- EDPB Guidelines on data subject rights
- US state data broker registration laws (CA, VT, OR)
- FTC enforcement actions on data licensing (e.g., FTC v. Avast, 2024)
- Dwork, Cynthia — "Differential Privacy" (2006)
FILE:references/data_product_strategy.md
# Data Product Strategy — The Decision: "Warehouse, lakehouse, or mesh — and what do we build vs buy?"
This reference answers exactly one decision: **what is the right data platform for our stage, and which components do we build ourselves?** It is stage-driven, not technology-trend-driven.
Pair with `scripts/data_product_strategy_picker.py` for automation.
## The Three Architectures
### Warehouse Only
**What it is:** A single SQL-accessible data store (Snowflake / BigQuery / Redshift / Postgres + dbt). All transformations happen in-warehouse.
**Use when:**
- ≤5 distinct data consumers (people/teams who query data weekly)
- <2TB of data
- No ML/AI use cases in production
- Reporting + dashboards are 90%+ of use cases
**Kill criteria (stop using warehouse-only when):**
- A data consumer needs unstructured data (logs, images, audio) → can't ingest cleanly
- ML model in production needs feature pipelines → warehouse-only is rigid
- 5+ consumers means hub-and-spoke ownership becomes the bottleneck
**Failure mode:** Treating it as forever. Many companies sit on warehouse-only for 2 years past viability because migration feels expensive.
### Lakehouse
**What it is:** Warehouse + object storage (S3/GCS/Azure Blob) with a table format like Apache Iceberg, Delta Lake, or Hudi. Single substrate for SQL analytics, ML training data, and unstructured ingestion.
Implementations: Databricks (Delta), Snowflake with Iceberg, AWS Redshift with Spectrum, BigQuery with BigLake.
**Use when:**
- 5–25 distinct data consumers
- 2TB–1PB data
- 1–3 ML models in production OR planning to be in 12 months
- Mixed structured + unstructured data
- Team has engineering capacity to maintain ingestion + transformation pipelines
**Kill criteria:**
- 25+ consumers AND federated ownership culture → time to consider mesh
- ML workloads disappear AND data shrinks below 2TB → simplify back to warehouse
- Vendor lock-in becomes intolerable → table formats (Iceberg) mitigate this; lakehouse vendor swaps remain expensive
**Failure mode:** Adopting before needed. Lakehouse architecture has 2–3x the operational complexity of pure warehouse. If you have 4 consumers and no ML, it's premature.
### Data Mesh
**What it is:** Federated data product ownership. Domain teams own their data products end-to-end (ingest → modeling → serving → SLAs). Central platform team provides the infrastructure substrate but does not produce data products.
Coined by Zhamak Dehghani (Thoughtworks); productionized at Netflix, Zalando, JP Morgan.
**Use when:**
- 25+ distinct data consumers across 4+ domains
- Federated ownership culture **already exists** in the org (you can't bolt it on)
- Central data team is a bottleneck for 50%+ of work
- Stage: growth or late-stage (Series C+)
**Kill criteria (mesh failure modes):**
- After 6 months: producing teams haven't adopted ownership → revert to hub-and-spoke
- Platform team still doing 50%+ of data product work → platform isn't truly self-serve
- Domain teams complain about onboarding → too much friction for "do it yourself"
- Cross-domain analytics has degraded vs warehouse era → integration layer missing
**Failure mode:** Mesh-without-culture. Companies adopt the architecture before the operating model. Result: distributed warehouses with no governance, worse than starting point.
## The Build-vs-Buy Decision Tree
For each platform layer, the question isn't "can we build it?" — it's "is it our IP, and does building it create a moat?"
### Storage / Warehouse
**Always BUY.** Snowflake, BigQuery, Databricks, Redshift, Postgres-with-Citus. Storage is commodity. Building distributed storage is a 50-engineer-year investment with zero business return unless you ARE a data infra company.
**Only build if:** You're a database company.
### ELT / Ingest
**Almost always BUY.** Fivetran, Airbyte, Stitch, Meltano. The connector maintenance burden (200+ source APIs, all changing constantly) is unjustifiable for any non-data-infra company.
**Only build if:** Source isn't supported by any vendor AND is business-critical AND you'll contribute the connector upstream so you're not maintaining a fork forever.
### Modeling / Transformations
**Always BUILD.** dbt is the de facto standard (open source). Your domain logic encoded in dbt models IS your data IP. No vendor can supply your domain understanding.
**Variants to evaluate:**
- dbt Core (open source) → free, self-hosted, requires orchestration (Airflow/Dagster/Prefect)
- dbt Cloud → managed, expensive at scale, simpler ops
- SQLMesh → newer, claims better state management
- Coalesce → visual SQL, expensive
### BI / Dashboards
**Almost always BUY.** Metabase (cheap, OSS option), Looker (enterprise, semantic layer), Mode (analyst-friendly + SQL), Hex (notebooks + dashboards), Tableau (legacy strong), Sigma (spreadsheet UX).
**Build only if:** You're shipping embedded analytics as a customer-facing feature (then evaluate Cube.dev, Embeddable, or build on Apache Superset).
**Embedded analytics is a real build-vs-buy:** for B2B SaaS shipping dashboards to customers, the choice between embedding a vendor (Cube + custom UI) vs full custom (Superset + heavy frontend) is significant. Buy-with-customization usually wins until 100K+ customer-tenants.
### Feature Store
**DEFER until you have 3+ ML models in production.**
**Then:** Tecton (managed, expensive, mature) or Hopsworks (alternative) for BUY; Feast (open source, lighter) for BUILD-on-OSS.
**Why defer:** Feature stores solve feature reuse + governance. With 1 model, you have 0 features-to-reuse. The operational overhead of a feature store exceeds the value below ~3 models sharing features.
### ML Platform
**DEFER until you have 5+ ML models in production.**
**Then:** Databricks ML, Vertex AI (Google), SageMaker (AWS), or Azure ML.
**Why defer:** ML platforms wrap experiment tracking, model registry, deployment, monitoring. Below 5 models with active retraining, scheduled training jobs + MLflow / W&B + simple K8s deployment is sufficient.
## Operational Maturity Layers (independent of architecture)
These apply regardless of warehouse / lakehouse / mesh choice:
1. **Data quality monitoring.** dbt tests, Great Expectations, Monte Carlo. Start at any scale.
2. **Lineage tracking.** dbt auto-generates lineage; OpenLineage / DataHub / Atlan for cross-tool. Start at 50+ models.
3. **Catalog + discovery.** DataHub, Atlan, Castor, Selectstar. Start at 100+ tables consumed by 10+ people.
4. **Access control + governance.** Snowflake/BigQuery native RBAC; Immuta / Privacera for policy abstraction. Start when you have regulated data or > 50 consumers.
## Sequencing Pattern (12-month plan)
A typical Series A → Series B sequencing:
| Quarter | Focus | Deliverable |
|---|---|---|
| Q1 | Foundation | Centralized ELT (buy); dbt for top-5 marts (build); 5 data quality tests |
| Q2 | Self-serve BI | Roll out BI tool; semantic layer in dbt or LookML; train 3 functional teams |
| Q3 | First ML use case OR embedded analysts | Either feature store for top-1 ML model OR embed 1 analyst per major function |
| Q4 | Evaluate and decide | Re-run picker; decide on Q1-next-year architecture changes |
## Anti-Patterns
- **Adopting a vendor before knowing the use case.** "We bought Snowflake but we're 80% on Postgres still." → vendor first, problem second.
- **Building "platform" before having customers (consumers).** Internal data platform team with no users is shelfware.
- **Treating data mesh as an architecture choice.** It's an operating model choice; the architecture is a consequence.
- **Splitting warehouse spend across 3 vendors.** Multi-cloud data is a 3x cost increase with no benefit until you're at Series D+.
- **Hiring data scientists before analysts.** Data scientists need clean data + clear questions. Build the analyst + analytics-engineer layer first.
## When This Reference Doesn't Help
- **Schema design.** See `engineering/database-designer/`.
- **Query optimization.** See `engineering/sql-database-assistant/`.
- **Observability for data pipelines.** See `engineering/observability-designer/`.
- **RAG architecture.** See `engineering/rag-architect/`.
This reference picks the architecture and the build-vs-buy. Tactical implementation is a separate skill family.
---
**Source authorities:**
- Dehghani, Zhamak — "Data Mesh: Delivering Data-Driven Value at Scale" (O'Reilly, 2022)
- Databricks Lakehouse paper, 2021
- Apache Iceberg, Delta Lake, Apache Hudi specifications
- dbt Labs Analytics Engineering Guide
FILE:references/data_team_org_evolution.md
# Data Team Org Evolution — The Decision: "What data role do we hire next, and when do we centralize vs embed?"
This reference answers exactly one decision: **for our stage and business decisions we can't currently make, what is the next role to add — and at what point do we centralize vs embed?**
## The Wrong Question
> "Should we hire a data scientist?"
This is the wrong question. Most data scientists hired by Series A startups are unable to deliver value because:
- The data isn't clean enough for modeling
- There's no infrastructure to deploy a model
- The "model" the founder imagines is actually a SQL query
## The Right Question
> "What's the next decision we can't make because we lack data, and what role unblocks that?"
This shifts hiring from role-taxonomy to decision-unblocking. The data org grows in response to specific decision gaps.
## The Five Stages
### Stage 1: Pre-seed / Seed
**Team size:** 1-15 people. **Data team:** 0.
**Reality:** Founder is the analyst. SQL + spreadsheets are sufficient.
**Don't hire:** Data engineer, data scientist, head of data. They will have nothing to do because the questions aren't crisp enough yet.
**Tooling:** Postgres / production DB direct read access. Metabase Free or Looker Studio. Google Sheets.
**When to move to stage 2:** Founder is spending >20% of their week on data work AND it's preventing them from doing CEO work.
### Stage 2: Series A
**Team size:** 15-50 people. **Data team:** 1-3.
**First hire: Analyst (NOT data engineer, NOT data scientist).**
Why: at this stage, 80% of the value is in clean reports, dashboards, and quick ad-hoc analyses. An analyst delivers all of this. A data engineer wants to build infrastructure that's premature; a data scientist wants to build models that don't have ROI yet.
Profile: 2-4 years experience, strong SQL, BI tool fluency, comfortable with ambiguity, can talk to non-data people.
**Second hire: Analytics engineer (dbt practitioner).**
Why: after the first analyst, the most acute pain is "dashboards are out of sync because everyone defines 'active customer' differently." Analytics engineer brings discipline (dbt models, semantic layer) and turns the analyst's work into reusable infrastructure.
Profile: SQL fluency + software engineering practices (PRs, tests, version control), dbt experience preferred but not required.
**Don't hire yet:** Data engineer, data scientist, head of data, data PM.
**When to move to stage 3:** 3+ functional teams are requesting bespoke analyses weekly, AND your first ML use case has a clear ROI.
### Stage 3: Series B
**Team size:** 50-200. **Data team:** 4-8.
**Third hire: Data engineer.**
Why: ingest pipelines are now business-critical. Salesforce → warehouse, Stripe → warehouse, product events → warehouse. Reliability matters. The analytics engineer cannot maintain this AND ship dbt models.
Profile: Python + SQL + understanding of streaming vs batch tradeoffs, experience with Fivetran/Airbyte or similar.
**Fourth hire: Senior analyst (embedded in GTM, often Sales/Marketing).**
Why: GTM is where data ROI is most measurable. An analyst embedded in the sales org (or reporting dotted-line to CRO) closes the gap between data team and revenue org.
**Fifth hire (conditional): Data PM.**
When: 3+ functional teams need data and the data team has ≥4 people. The data PM owns the roadmap, intake, and SLA negotiations. Without this, the team flips into reactive mode and never builds platform.
**Conditional: Data scientist / ML engineer.**
Hire only when:
- You have at least 1 model in production OR a strong hypothesis with ROI math
- Data engineer is in place (so data scientist isn't blocked on infrastructure)
- Eng leadership signs on for productionizing models (not just notebooks)
**When to move to stage 4:** Central data team is the bottleneck for >50% of GTM data requests, OR you're hiring data people every quarter and they all report to one manager.
### Stage 4: Growth (Series C / pre-IPO)
**Team size:** 200-1000. **Data team:** 8-30.
**Sixth hire: Manager of Analytics (people manager).**
Why: at 5-8 reports, the original analytics lead can no longer code AND manage. Split into managers + senior ICs.
**Seventh hire: ML engineer (production-grade).**
When: 1+ model in production, 2-3 more planned. ML engineer owns deployment, monitoring, retraining infrastructure. Different person from data scientist (who owns model invention).
**Eighth hire: Head of Data.**
Triggers:
- Data team is 10+ people
- Data team has its own strategy independent of company strategy (problematic if no one owns the reconciliation)
- Founder/CTO is no longer the right escalation for data decisions
- Compliance / governance becomes board-level concern
The Head of Data owns data strategy, hires/fires, and is the cross-functional executive for all data + AI.
**Centralize vs Embed decision:**
By Series C, the centralize-vs-embed tension is acute. Two patterns work:
**Hub-and-spoke (most common, recommended):**
- Central data platform team owns infrastructure, governance, semantic layer
- Embedded analysts in 3-5 major functional teams (Sales, Marketing, Product, CS, Finance)
- Embedded analysts have solid-line to function leader, dotted-line to Head of Data
- Tools, standards, dbt models are central; questions and SLAs are local
**Federated (data mesh — only if culture supports):**
- Each domain team owns their data products end-to-end
- Central platform team provides infrastructure substrate, not data products
- Requires high data culture maturity; failure mode is mesh-without-culture
Hub-and-spoke handles 95% of Series C companies. Mesh fits when you're 1000+ people with strong domain ownership culture (Netflix, Zalando, JP Morgan scale).
**When to move to stage 5:** Series D / late-stage growth, 50+ data team members, multiple domains with their own data leadership.
### Stage 5: Late-stage (Series D+, post-IPO)
**Team size:** 1000+. **Data team:** 30-200+.
**CDO promotion / hire.**
Triggers:
- Data is in the company's strategic narrative (board deck, investor calls)
- Data has its own P&L (productized data, monetization)
- Multiple regulatory regimes apply (GDPR + CCPA + HIPAA + EU AI Act)
- Head of Data is escalating data-strategy questions to CTO and it's not landing right
Profile:
- Has run a data org at $100M+ ARR scale
- Comfortable with board reporting
- Strategic, not just technical
- Strong on data governance + AI policy (post-2024 AI Act and similar requirements)
**Federated CDO model (late-stage):**
At thousands-of-people scale, the CDO often runs:
- Central platform team (engineering)
- Central governance team (privacy, compliance, AI policy)
- Federated data leaders embedded per business unit
- Data product leaders for any productized data
## Specific Roles Defined
Because founders confuse these:
| Role | Owns | Does NOT own |
|---|---|---|
| Analyst | Ad-hoc analyses, dashboards, business questions | Pipeline reliability, model deployment |
| Analytics engineer | dbt models, semantic layer, data quality tests | Ingest pipelines, ML, infrastructure |
| Data engineer | Ingest pipelines (Fivetran/Airbyte/custom), warehouse infra, streaming | Modeling logic, dashboards, ML models |
| Data scientist | Model invention, experimentation, statistical analysis | Production deployment, monitoring |
| ML engineer | Production model deployment, monitoring, retraining infra | Model invention |
| Data PM | Data team roadmap, intake, prioritization, stakeholder mgmt | IC delivery work |
| Data PM (productized data) | Data products sold to customers | Internal-only data work |
| Head of Data | Data strategy, hiring, budget, exec representation | Day-to-day IC work |
| CDO | Data + AI strategy at board level, governance, P&L (where applicable) | Day-to-day execution |
## The Centralize-vs-Embed Trigger
The decision is not "centralize or embed" — it's "when do you transition from one to the other?"
**Centralized (everyone reports to one data leader):** works up to ~5 data people serving ≤5 functional teams.
**Hub-and-spoke (central platform + embedded analysts):** works from 5-30 data people serving 5-15 functional teams.
**Federated (each domain owns):** works at 30+ data people across 15+ functional teams WITH strong data culture.
**The trigger to move from centralized to hub-and-spoke:** when 3+ functional teams complain that the central team doesn't understand their domain, AND when the central team's intake queue exceeds 4 weeks of lead time.
**The trigger to move from hub-and-spoke to federated (data mesh):** when domain teams have data leaders, are already running their own data SLAs, and would rather not depend on central platform for product launches. This is rare and usually arrives at thousands-of-people scale.
## Anti-Patterns
- **Hiring a data scientist as first data hire.** They will spend 6 months unable to deliver because data isn't clean.
- **Hiring a "head of data" at Series A.** Nothing for them to manage.
- **Hiring multiple analysts before adding analytics engineer.** Dashboards multiply; consistency vanishes.
- **Building a data platform with no users.** Internal platform team with no customers is shelfware.
- **Hiring an ML engineer before a data engineer.** ML engineer cannot deploy models if data pipelines are broken.
- **Promoting an analyst to "Head of Data" without people-management experience.** Most analysts are great ICs; people management is a different skill.
## When This Reference Doesn't Help
- **Comp benchmarking.** See `c-level-advisor/skills/chro-advisor/scripts/comp_benchmarker.py`.
- **Leveling ladders.** See `c-level-advisor/skills/chro-advisor/references/leveling_ladders.md`.
- **Specific JD templates.** Not covered here; many open-source examples exist.
- **Performance management.** Standard people management; not data-specific.
This reference is about the data team's evolution as a function of company-stage decisions, not about HR mechanics.
---
**Source observations (non-exhaustive):**
- Tristan Handy (dbt Labs) — "The Modern Data Stack: Past, Present, Future"
- Maxime Beauchemin — "The Rise of the Data Engineer" (2017), "The Downfall of the Data Engineer" (2017)
- Erik Bernhardsson — "The Modern Data Experience" (2022)
- Lauren Balik — "Modern Data Stack writings"
- Direct observations from 50+ B2B SaaS data org evolutions, 2020-2026
FILE:scripts/ai_training_data_audit.py
#!/usr/bin/env python3
"""ai_training_data_audit.py — Audit data sources for AI training eligibility.
Stdlib-only. Audits each data source on 3 dimensions:
- Origin (1st-party-explicit-opt-in / 1st-party-tos-only / partner-licensed / scraped / synthetic)
- Data class (anonymous-aggregate / behavioral / pii / third-party-content / regulated)
- Use case (in-product-personalization / fine-tune-our-model / train-foundation-model / external-sharing)
Returns GO / MITIGATE / NO-GO per source with the specific risk and remediation.
NOT legal advice — surfaces decisions for qualified counsel.
Input schema (JSON):
{
"sources": [
{
"name": "Product telemetry events",
"origin": "1st-party-tos-only",
"data_class": "behavioral",
"use_case": "in-product-personalization"
},
...
]
}
Usage:
python ai_training_data_audit.py # uses embedded sample
python ai_training_data_audit.py path/to/sources.json
python ai_training_data_audit.py sources.json --output json
"""
import argparse
import json
import sys
from dataclasses import dataclass, asdict
from typing import Any, Dict, List, Optional, Tuple
SAMPLE: Dict[str, Any] = {
"sources": [
{
"name": "Anonymous product telemetry (event aggregates)",
"origin": "1st-party-tos-only",
"data_class": "anonymous-aggregate",
"use_case": "in-product-personalization",
},
{
"name": "Customer support transcripts",
"origin": "1st-party-tos-only",
"data_class": "pii",
"use_case": "fine-tune-our-model",
},
{
"name": "Scraped LinkedIn profiles",
"origin": "scraped",
"data_class": "pii",
"use_case": "fine-tune-our-model",
},
{
"name": "Synthetic conversational data (LLM-generated)",
"origin": "synthetic",
"data_class": "third-party-content",
"use_case": "train-foundation-model",
},
{
"name": "User opt-in survey responses",
"origin": "1st-party-explicit-opt-in",
"data_class": "behavioral",
"use_case": "external-sharing",
},
{
"name": "Partner-licensed industry dataset",
"origin": "partner-licensed",
"data_class": "anonymous-aggregate",
"use_case": "train-foundation-model",
},
{
"name": "Anonymized health screening responses",
"origin": "1st-party-explicit-opt-in",
"data_class": "regulated",
"use_case": "fine-tune-our-model",
},
]
}
VALID_ORIGINS = {
"1st-party-explicit-opt-in",
"1st-party-tos-only",
"partner-licensed",
"scraped",
"synthetic",
}
VALID_CLASSES = {
"anonymous-aggregate",
"behavioral",
"pii",
"third-party-content",
"regulated",
}
VALID_USE_CASES = {
"in-product-personalization",
"fine-tune-our-model",
"train-foundation-model",
"external-sharing",
}
@dataclass
class AuditResult:
name: str
origin: str
data_class: str
use_case: str
verdict: str # GO | MITIGATE | NO-GO
risk: str
remediation: str
citations: List[str]
# Verdict matrix: (origin, data_class, use_case) -> (verdict, risk, remediation, citations)
# Built by applying these rules in order; first match wins.
def _decide(origin: str, data_class: str, use_case: str) -> Tuple[str, str, str, List[str]]:
# Rule 1: Scraped data is always NO-GO for training (hiQ v. LinkedIn, copyright, GDPR Art. 6).
if origin == "scraped":
return (
"NO-GO",
"Scraped data lacks lawful basis under GDPR Art. 6 (no consent, no legitimate interest "
"balancing test); high copyright risk; hiQ v. LinkedIn left exposure for ToS-violation claims; "
"many AI Act high-risk use cases require demonstrable provenance.",
"Remove from training set. Either (a) procure licensed alternative from data broker, "
"(b) replace with synthetic data, or (c) build 1st-party explicit opt-in pipeline.",
["GDPR Art. 6", "hiQ Labs v. LinkedIn", "EU AI Act Art. 10 (data governance)"],
)
# Rule 2: Regulated data (PHI, PCI, kids) requires explicit opt-in + specific compliance
# framework; never train foundation model with raw regulated data.
if data_class == "regulated":
if origin == "1st-party-explicit-opt-in" and use_case in {"in-product-personalization", "fine-tune-our-model"}:
return (
"MITIGATE",
"Regulated data (PHI / PCI / children) may be processed under explicit opt-in IF the "
"framework permits (HIPAA Limited Data Set, COPPA verifiable parental consent). "
"Fine-tuning increases re-identification risk vs in-product use.",
"Required: (1) framework-specific consent flow, (2) DPIA/PIA on file, (3) k-anonymity "
"≥ 5 audit before any training, (4) model output filters for regulated-content leakage, "
"(5) DPA with any vendor in the pipeline.",
["HIPAA", "HITECH §13402", "COPPA", "GDPR Art. 9", "EU AI Act Annex III"],
)
return (
"NO-GO",
"Regulated data (PHI / PCI / children) cannot be used for foundation training or external "
"sharing without specific framework authorization, and not at all without explicit opt-in.",
"Either (a) restrict to in-product use under existing framework consent, (b) train on "
"synthetic data modeled on the corpus, or (c) obtain new explicit opt-in covering the "
"specific training purpose.",
["HIPAA", "GDPR Art. 9", "COPPA"],
)
# Rule 3: PII at any use case beyond in-product-personalization requires explicit opt-in,
# specific lawful basis, AND anonymization/pseudonymization.
if data_class == "pii":
if use_case == "in-product-personalization":
if origin == "1st-party-tos-only":
return (
"MITIGATE",
"PII processing for in-product personalization can rest on GDPR Art. 6(1)(b) "
"(performance of contract) or 6(1)(f) (legitimate interest) IF the personalization "
"is reasonably expected. Train-once derived models retain risk.",
"Required: (1) Art. 6 lawful basis documented, (2) data minimization audit, "
"(3) deletion request honored for the source data even after model training "
"(implementation: filter-on-output OR retrain on deletion), (4) DPIA if scale > 5000 users.",
["GDPR Art. 6", "GDPR Art. 17 (right to erasure)", "EDPB Guidelines on Art. 22"],
)
if origin == "1st-party-explicit-opt-in":
return (
"GO",
"PII with explicit opt-in for in-product personalization is the strongest position. "
"Standard residual risks: opt-in revocation, deletion requests.",
"Maintain: (1) opt-in audit trail per user, (2) machinery to honor revocation/erasure "
"(filter-on-output or retrain), (3) clear notice on what model is trained.",
["GDPR Art. 6(1)(a)", "GDPR Art. 17"],
)
# Fine-tune-our-model, train-foundation-model, external-sharing with PII
if origin == "1st-party-explicit-opt-in":
return (
"MITIGATE",
"PII for fine-tuning or beyond requires explicit opt-in covering THIS specific "
"training purpose (not generic TOS). Risk: training-data extraction attacks, "
"memorization, model-output leakage.",
"Required: (1) purpose-specific opt-in (not bundled TOS), (2) differential privacy "
"or k-anonymity audit, (3) memorization tests on the trained model, (4) DPIA, "
"(5) DPA with infra/training vendor, (6) EU AI Act conformity assessment if "
"high-risk use case.",
["GDPR Art. 6(1)(a)", "GDPR Art. 35 (DPIA)", "EU AI Act Art. 10"],
)
return (
"NO-GO",
"PII for fine-tuning or foundation training without explicit opt-in fails GDPR Art. 6. "
"TOS-only consent is insufficient for materially different purpose.",
"Either (a) restrict use case to in-product personalization under existing basis, "
"(b) build explicit opt-in pipeline before training, or (c) anonymize/pseudonymize "
"to k-anonymity ≥ 5 and re-classify as anonymous-aggregate.",
["GDPR Art. 6", "EDPB Opinion 28/2024"],
)
# Rule 4: 3rd-party content (e.g., user-uploaded files, customer support transcripts
# quoting other systems, scraped public documents within user submissions).
if data_class == "third-party-content":
if origin in {"synthetic", "partner-licensed"}:
return (
"MITIGATE",
"Synthetic or licensed 3rd-party-content carries content-license risk: even with a "
"license, training a model may exceed the license scope (e.g., 'view' license vs "
"'derivative work creation').",
"Required: (1) license review by counsel for training-specific clauses, (2) carve-out "
"for AI training in licensing agreement, (3) provenance log per source for AI Act compliance, "
"(4) opt-out mechanism if license permits revocation.",
["NYT v. OpenAI (2024)", "EU AI Act Art. 53 (general-purpose models)"],
)
if origin == "1st-party-tos-only":
return (
"MITIGATE",
"User-uploaded content under TOS-only license has uncertain training rights post-2024 "
"lawsuits. Risk: copyright infringement if model output is substantially similar to "
"training data.",
"Required: (1) TOS explicitly grants training rights for the specific model class, "
"(2) output similarity monitoring (de-duping / fuzzy match against training corpus), "
"(3) opt-out mechanism in TOS update.",
["Authors Guild v. Google", "Andersen v. Stability AI", "NYT v. OpenAI"],
)
if origin == "1st-party-explicit-opt-in":
return (
"GO",
"Explicit opt-in for training on user-uploaded content is the strongest position. "
"Maintain output-similarity guardrails to catch unexpected memorization.",
"Required: (1) opt-in audit trail, (2) revocation flow, (3) output similarity testing.",
["GDPR Art. 6(1)(a)"],
)
# Rule 5: Behavioral data — generally safer than PII, but external sharing still requires consent.
if data_class == "behavioral":
if use_case == "external-sharing":
if origin == "1st-party-explicit-opt-in":
return (
"GO",
"Behavioral data with explicit opt-in for external sharing — clean.",
"Maintain: (1) revocation flow, (2) anonymization audit before each external "
"share (k-anonymity ≥ 5), (3) recipient DPA.",
["GDPR Art. 6(1)(a)"],
)
return (
"MITIGATE",
"Behavioral data without explicit opt-in for external sharing is borderline. TOS-only "
"is weak basis; partner-licensed depends on partner's original consent flow.",
"Required: (1) anonymization to k-anonymity ≥ 5, (2) recipient DPA with no-reidentification "
"clause, (3) audit upstream consent if partner-licensed, (4) consider opt-in pipeline.",
["GDPR Art. 6", "Art. 22 (automated decision-making)"],
)
# Behavioral + training use cases
if origin in {"1st-party-explicit-opt-in", "1st-party-tos-only", "partner-licensed"}:
return (
"GO",
"Behavioral data from controlled origin for internal training is generally safe. "
"Residual risk: model leakage if behavioral patterns are individually identifying.",
"Maintain: (1) deletion handling on user request, (2) periodic memorization tests, "
"(3) DPIA if scale > 50K users or sensitive inferences.",
["GDPR Art. 6", "GDPR Art. 35"],
)
# Rule 6: Anonymous aggregate — generally safe at all use cases.
if data_class == "anonymous-aggregate":
if origin == "scraped":
# Already handled above
pass
return (
"GO",
"Anonymous aggregate data is the safest class. Residual risk: re-identification attacks "
"if aggregate cells are small.",
"Maintain: (1) k-anonymity ≥ 5 in all published aggregates, (2) differential privacy if "
"shared externally, (3) provenance log for AI Act compliance.",
["EU AI Act Art. 10", "GDPR Recital 26"],
)
# Synthetic + non-3rd-party-content
if origin == "synthetic":
return (
"GO",
"Synthetic data is generally safe for training. Residual risk: synthetic data generated "
"from a non-clean source inherits its risks.",
"Maintain: (1) document the generation pipeline including any non-synthetic seed, "
"(2) test for bias inherited from generator, (3) provenance log.",
["EU AI Act Art. 10"],
)
# Default conservative fallback
return (
"MITIGATE",
"Configuration not matched by explicit rules — manual review required.",
"Engage qualified data privacy counsel to assess this specific origin/class/use combination.",
[],
)
def audit(payload: Dict[str, Any]) -> List[AuditResult]:
results: List[AuditResult] = []
for src in payload.get("sources", []):
name = src.get("name", "<unnamed>")
origin = src.get("origin", "")
data_class = src.get("data_class", "")
use_case = src.get("use_case", "")
# Validation
errors = []
if origin not in VALID_ORIGINS:
errors.append(f"invalid origin '{origin}'")
if data_class not in VALID_CLASSES:
errors.append(f"invalid data_class '{data_class}'")
if use_case not in VALID_USE_CASES:
errors.append(f"invalid use_case '{use_case}'")
if errors:
results.append(AuditResult(
name=name,
origin=origin,
data_class=data_class,
use_case=use_case,
verdict="NO-GO",
risk=f"Schema error: {'; '.join(errors)}",
remediation=(
f"Origin must be one of {sorted(VALID_ORIGINS)}; "
f"data_class one of {sorted(VALID_CLASSES)}; "
f"use_case one of {sorted(VALID_USE_CASES)}."
),
citations=[],
))
continue
verdict, risk, remediation, citations = _decide(origin, data_class, use_case)
results.append(AuditResult(
name=name,
origin=origin,
data_class=data_class,
use_case=use_case,
verdict=verdict,
risk=risk,
remediation=remediation,
citations=citations,
))
# Sort: NO-GO first, then MITIGATE, then GO
order = {"NO-GO": 0, "MITIGATE": 1, "GO": 2}
results.sort(key=lambda r: order.get(r.verdict, 9))
return results
def render_text(results: List[AuditResult], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("AI TRAINING DATA AUDIT")
lines.append(f"Source: {source}")
lines.append(f"Sources audited: {len(results)}")
lines.append("=" * 72)
lines.append("")
counts = {"NO-GO": 0, "MITIGATE": 0, "GO": 0}
for r in results:
counts[r.verdict] = counts.get(r.verdict, 0) + 1
lines.append(f"Verdicts: 🔴 NO-GO: {counts['NO-GO']} 🟡 MITIGATE: {counts['MITIGATE']} 🟢 GO: {counts['GO']}")
lines.append("")
lines.append("-" * 72)
for i, r in enumerate(results, 1):
marker = {"NO-GO": "🔴", "MITIGATE": "🟡", "GO": "🟢"}.get(r.verdict, "•")
lines.append(f"[{i}] {marker} {r.verdict:<9} — {r.name}")
lines.append(f" Origin: {r.origin} | Class: {r.data_class} | Use case: {r.use_case}")
lines.append("")
lines.append(f" Risk:")
for line in _wrap(r.risk, 6):
lines.append(line)
lines.append("")
lines.append(f" Remediation:")
for line in _wrap(r.remediation, 6):
lines.append(line)
if r.citations:
lines.append(f" Citations: {', '.join(r.citations)}")
lines.append("")
lines.append("-" * 72)
lines.append("")
lines.append("REMINDER: This audit applies rule-based triage to a 3-dimensional matrix. Always engage")
lines.append("qualified data privacy / AI counsel for binding decisions.")
return "\n".join(lines)
def _wrap(text: str, indent: int, width: int = 66) -> List[str]:
import textwrap
return textwrap.wrap(text, width=width, initial_indent=" " * indent, subsequent_indent=" " * indent) or [" " * indent + text]
def main() -> int:
parser = argparse.ArgumentParser(
description="Audit data sources for AI training eligibility (origin × class × use-case matrix).",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to sources JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: 7 mixed sources>"
results = audit(payload)
if args.output == "json":
print(json.dumps({
"source": source,
"count": len(results),
"verdict_counts": {
"NO-GO": sum(1 for r in results if r.verdict == "NO-GO"),
"MITIGATE": sum(1 for r in results if r.verdict == "MITIGATE"),
"GO": sum(1 for r in results if r.verdict == "GO"),
},
"results": [asdict(r) for r in results],
}, indent=2))
else:
print(render_text(results, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/data_asset_valuator.py
#!/usr/bin/env python3
"""data_asset_valuator.py — Value a B2B customer data corpus + productization viability.
Stdlib-only. Takes a corpus profile and computes:
- Strategic value score (0-10)
- Defensibility moat strength (NONE / WEAK / MEDIUM / STRONG)
- M&A multiplier (ARR uplift range in strategic-buyer scenarios)
- Productization paths (benchmark / embedding / direct license) with risk profile
- Contractual constraint impact (% of corpus blocked from productization)
Input schema (JSON):
{
"data_type": "sales-engagement", // descriptive
"customer_count": 380,
"time_history_years": 2.3,
"exclusivity": "high", // none | low | medium | high
"freshness": "real-time", // batch-daily | batch-weekly | near-real-time | real-time
"msa_carveouts_count": 47, // # of customers with data-use carve-outs blocking productization
"anonymization_audit_passed": false, // k-anonymity >=5 confirmed
"company_arr_m": 12, // company ARR in millions for M&A multiplier math
"regulated_data_present": false
}
Usage:
python data_asset_valuator.py # uses embedded B2B sample
python data_asset_valuator.py path/to/corpus.json
python data_asset_valuator.py corpus.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"data_type": "Sales engagement logs (email, calls, meetings)",
"customer_count": 380,
"time_history_years": 2.3,
"exclusivity": "high",
"freshness": "real-time",
"msa_carveouts_count": 47,
"anonymization_audit_passed": False,
"company_arr_m": 12,
"regulated_data_present": False,
}
EXCLUSIVITY_SCORE = {"none": 0, "low": 2, "medium": 5, "high": 9}
FRESHNESS_SCORE = {"batch-weekly": 2, "batch-daily": 5, "near-real-time": 7, "real-time": 9}
def strategic_value(profile: Dict[str, Any]) -> Dict[str, Any]:
"""Computes strategic value score and moat strength."""
customers = profile.get("customer_count", 0)
history = profile.get("time_history_years", 0)
excl = profile.get("exclusivity", "none")
fresh = profile.get("freshness", "batch-weekly")
excl_score = EXCLUSIVITY_SCORE.get(excl, 0)
fresh_score = FRESHNESS_SCORE.get(fresh, 0)
# Customer cohort breadth
if customers >= 500:
cohort_score = 10
elif customers >= 200:
cohort_score = 8
elif customers >= 100:
cohort_score = 6
elif customers >= 50:
cohort_score = 4
else:
cohort_score = 2
# Time history depth
if history >= 5:
history_score = 10
elif history >= 3:
history_score = 8
elif history >= 2:
history_score = 6
elif history >= 1:
history_score = 4
else:
history_score = 2
# Composite
composite = (excl_score * 2 + fresh_score + cohort_score + history_score) / 5
composite = round(composite, 1)
# Moat strength derived from exclusivity + cohort
if excl_score >= 8 and cohort_score >= 8:
moat = "STRONG"
moat_explain = "Exclusivity + breadth means replicating requires 2+ years of customer cohort acquisition."
elif excl_score >= 5 and cohort_score >= 6:
moat = "MEDIUM"
moat_explain = "Defensible but a well-funded competitor with 18-24 months can match."
elif excl_score >= 2:
moat = "WEAK"
moat_explain = "Some unique characteristics but largely replicable from public or commercially-available sources."
else:
moat = "NONE"
moat_explain = "Not a moat — same data is available elsewhere."
return {
"composite_score": composite,
"max_score": 10.0,
"components": {
"exclusivity": excl_score,
"freshness": fresh_score,
"cohort_breadth": cohort_score,
"history_depth": history_score,
},
"moat_strength": moat,
"moat_explanation": moat_explain,
}
def ma_multiplier(profile: Dict[str, Any], strategic: Dict[str, Any]) -> Dict[str, Any]:
"""Computes M&A multiplier range based on moat + corpus characteristics."""
moat = strategic["moat_strength"]
arr = profile.get("company_arr_m", 0)
carveouts = profile.get("msa_carveouts_count", 0)
customers = profile.get("customer_count", 1)
carveout_pct = (carveouts / customers * 100) if customers else 0
# Base multiplier by moat
base = {
"STRONG": (1.4, 1.7),
"MEDIUM": (1.15, 1.35),
"WEAK": (1.0, 1.1),
"NONE": (1.0, 1.0),
}
low, high = base.get(moat, (1.0, 1.0))
# Penalty for high carve-out %
if carveout_pct > 25:
low *= 0.85
high *= 0.85
carveout_note = f"{carveout_pct:.1f}% carve-out rate reduces multiplier ~15% (data is partially un-productizable)."
elif carveout_pct > 10:
low *= 0.95
high *= 0.95
carveout_note = f"{carveout_pct:.1f}% carve-out rate reduces multiplier ~5%."
else:
carveout_note = f"{carveout_pct:.1f}% carve-out rate — within tolerable range, no material multiplier impact."
low_arr = round(arr * low, 1) if arr else None
high_arr = round(arr * high, 1) if arr else None
return {
"multiplier_low": round(low, 2),
"multiplier_high": round(high, 2),
"carveout_pct": round(carveout_pct, 1),
"carveout_note": carveout_note,
"valuation_low_m": low_arr,
"valuation_high_m": high_arr,
"valuation_note": (
f"Strategic-buyer scenario: ARR arrM × ({low:.2f} - {high:.2f}) = low_arrM - high_arrM ARR-equivalent."
if arr else "Provide company_arr_m to compute valuation range."
),
}
def productization_paths(profile: Dict[str, Any], strategic: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Returns ranked productization paths with risk and viability."""
customers = profile.get("customer_count", 0)
carveouts = profile.get("msa_carveouts_count", 0)
carveout_pct = (carveouts / customers * 100) if customers else 0
anon_passed = profile.get("anonymization_audit_passed", False)
regulated = profile.get("regulated_data_present", False)
moat = strategic["moat_strength"]
paths = []
# Path 1: Industry benchmark report
benchmark_risk = "LOW"
benchmark_blockers = []
if not anon_passed:
benchmark_blockers.append("Anonymization audit (k-anonymity ≥ 5) required before publication")
if regulated:
benchmark_risk = "MEDIUM"
benchmark_blockers.append("Regulated data present — additional compliance review required")
paths.append({
"path": "Industry benchmark report (anonymized aggregates)",
"risk": benchmark_risk,
"revenue_potential": "Low ($50K-$500K/yr) but high credibility lift",
"viability": "HIGH" if not regulated else "MEDIUM",
"blockers": benchmark_blockers or ["No structural blockers"],
"first_step": (
"Run anonymization audit on top-3 metrics; draft quarterly benchmark report; "
"send to customers as opt-in value-add before public release."
),
})
# Path 2: Anonymized embedding endpoint
embed_risk = "MEDIUM"
embed_blockers = []
if not anon_passed:
embed_blockers.append("Anonymization audit required; embeddings can leak training data")
if carveout_pct > 0:
embed_blockers.append(
f"{int(carveouts)} customers have MSA carve-outs blocking productized use of their data"
)
if regulated:
embed_risk = "HIGH"
embed_blockers.append("Regulated data present — embeddings may retain re-identifiable signal")
paths.append({
"path": "Anonymized embedding endpoint (AI features for customers)",
"risk": embed_risk,
"revenue_potential": "Medium ($500K-$3M/yr) as platform feature OR add-on",
"viability": "HIGH" if moat in ("STRONG", "MEDIUM") and not regulated else "MEDIUM",
"blockers": embed_blockers,
"first_step": (
"Pilot embedding endpoint with 3 design-partner customers; memorization tests; "
"DPA addendum covering training-data flow."
),
})
# Path 3: Direct data licensing
license_risk = "HIGH"
license_blockers = []
if carveout_pct > 10:
license_blockers.append(
f"{carveout_pct:.1f}% of customers ({int(carveouts)}) have MSA carve-outs — direct licensing is "
"legally infeasible without re-papering or carve-out-excluded dataset"
)
license_blockers.append("Requires GDPR Art. 26 joint-controller analysis if EU customers present")
if regulated:
license_blockers.append("Regulated data licensing requires framework-specific consent + DPA")
paths.append({
"path": "Direct data licensing (to AI labs, data brokers, or industry players)",
"risk": license_risk,
"revenue_potential": "High ($2M-$20M/yr) at scale but high customer-trust cost",
"viability": "LOW" if carveout_pct > 10 or regulated else "MEDIUM",
"blockers": license_blockers,
"first_step": (
"First decide if customer trust impact is acceptable. If yes: re-paper 47 carve-out customers "
"OR build carve-out-excluded dataset; engage data broker counsel; draft customer comms plan."
),
})
return paths
def recommend_path(paths: List[Dict[str, Any]]) -> str:
"""Picks the highest-viability lowest-risk path as the recommended starting point."""
# Score: viability rank * 10 + (4 - risk_rank)
viability_rank = {"HIGH": 3, "MEDIUM": 2, "LOW": 1}
risk_rank = {"LOW": 3, "MEDIUM": 2, "HIGH": 1}
scored = [
(viability_rank.get(p["viability"], 0) * 10 + risk_rank.get(p["risk"], 0), p)
for p in paths
]
scored.sort(key=lambda x: -x[0])
return scored[0][1]["path"]
def analyze(profile: Dict[str, Any]) -> Dict[str, Any]:
strategic = strategic_value(profile)
ma = ma_multiplier(profile, strategic)
paths = productization_paths(profile, strategic)
recommended = recommend_path(paths)
return {
"strategic_value": strategic,
"ma_multiplier": ma,
"productization_paths": paths,
"recommended_starting_path": recommended,
}
def render_text(result: Dict[str, Any], profile: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("DATA ASSET VALUATION")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Corpus: {profile.get('data_type')}")
lines.append(f" Customers: {profile.get('customer_count')} | History: {profile.get('time_history_years')} years")
lines.append(f" Exclusivity: {profile.get('exclusivity')} | Freshness: {profile.get('freshness')}")
lines.append(f" MSA carve-outs: {profile.get('msa_carveouts_count')} customer(s)")
lines.append(f" Anonymization audit passed: {profile.get('anonymization_audit_passed')}")
lines.append(f" Regulated data present: {profile.get('regulated_data_present')}")
lines.append("")
lines.append("-" * 72)
sv = result["strategic_value"]
lines.append(f"STRATEGIC VALUE: {sv['composite_score']} / {sv['max_score']}")
lines.append(" Components:")
for k, v in sv["components"].items():
lines.append(f" {k:<20} {v}/10")
lines.append(f" Moat strength: {sv['moat_strength']}")
for line in _wrap(f" {sv['moat_explanation']}", 2):
lines.append(line)
lines.append("")
lines.append("-" * 72)
ma = result["ma_multiplier"]
lines.append(f"M&A MULTIPLIER (strategic-buyer scenario):")
lines.append(f" Range: {ma['multiplier_low']}x – {ma['multiplier_high']}x ARR")
if ma.get("valuation_low_m") is not None:
lines.append(f" Valuation impact: ma['valuation_low_m']M – ma['valuation_high_m']M ARR-equivalent")
for line in _wrap(f" {ma['carveout_note']}", 2):
lines.append(line)
lines.append("")
lines.append("-" * 72)
lines.append("PRODUCTIZATION PATHS:")
lines.append("")
for i, p in enumerate(result["productization_paths"], 1):
lines.append(f" [{i}] {p['path']}")
lines.append(f" Risk: {p['risk']} | Viability: {p['viability']} | Revenue: {p['revenue_potential']}")
lines.append(f" Blockers:")
for b in p["blockers"]:
lines.append(f" - {b}")
lines.append(f" First step:")
for line in _wrap(p["first_step"], 8):
lines.append(line)
lines.append("")
lines.append("-" * 72)
lines.append(f"RECOMMENDED STARTING PATH: {result['recommended_starting_path']}")
lines.append("")
lines.append("REMINDER: This valuation is a triage. Any actual productization, licensing, or M&A use")
lines.append("requires legal + data privacy review. Customer-trust impact is often the binding constraint,")
lines.append("not legal feasibility.")
return "\n".join(lines)
def _wrap(text: str, indent: int, width: int = 70) -> List[str]:
import textwrap
return textwrap.wrap(text, width=width, initial_indent=" " * indent, subsequent_indent=" " * indent) or [" " * indent + text]
def main() -> int:
parser = argparse.ArgumentParser(
description="Value a B2B customer data corpus + productization paths.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to corpus JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
profile = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
profile = SAMPLE
source = "<embedded sample: B2B SaaS sales engagement, 380 customers, 47 carve-outs>"
result = analyze(profile)
if args.output == "json":
print(json.dumps({"source": source, "profile": profile, **result}, indent=2))
else:
print(render_text(result, profile, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/data_product_strategy_picker.py
#!/usr/bin/env python3
"""data_product_strategy_picker.py — Pick data architecture + build-vs-buy + sequencing.
Stdlib-only. Takes a company profile and outputs:
- Recommended architecture (warehouse / lakehouse / data mesh) with reasoning + kill criteria
- Build-vs-buy decision per layer (storage, ELT, modeling, BI, feature store, ML platform)
- 12-month sequencing roadmap
The recommendation is deterministic, derived from the profile, not pattern-matched.
Input schema (JSON):
{
"stage": "series-a", // seed | series-a | series-b | growth | late-stage
"data_team_size": 3,
"internal_consumers": 8, // distinct people/teams consuming data weekly
"data_volume_tb": 4.5,
"ml_models_in_prod": 1,
"company_type": "b2b-saas", // b2b-saas | b2c-saas | consumer | marketplace | enterprise
"has_data_culture": false, // federated ownership culture in place? (mesh prerequisite)
"near_term_priorities": [
"self-serve-bi",
"improve-pipeline-reliability"
]
}
Usage:
python data_product_strategy_picker.py # uses embedded Series A SaaS
python data_product_strategy_picker.py path/to/profile.json
python data_product_strategy_picker.py profile.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Tuple
SAMPLE: Dict[str, Any] = {
"stage": "series-a",
"data_team_size": 3,
"internal_consumers": 8,
"data_volume_tb": 4.5,
"ml_models_in_prod": 1,
"company_type": "b2b-saas",
"has_data_culture": False,
"near_term_priorities": ["self-serve-bi", "improve-pipeline-reliability"],
}
def pick_architecture(profile: Dict[str, Any]) -> Tuple[str, str, List[str]]:
"""Returns (architecture, reasoning, kill_criteria)."""
consumers = profile.get("internal_consumers", 0)
volume = profile.get("data_volume_tb", 0)
ml_models = profile.get("ml_models_in_prod", 0)
culture = profile.get("has_data_culture", False)
stage = profile.get("stage", "")
# Data mesh: requires 25+ consumers across 4+ domains AND federated culture
if consumers >= 25 and culture and stage in ("growth", "late-stage"):
return (
"DATA MESH",
f"{consumers} data consumers across enough domains to justify federated ownership; "
"stated data-culture maturity supports the operational overhead.",
[
"Stop and revert if 6 months in: producing teams haven't adopted ownership (typical failure mode)",
"Stop if: central data platform team is still doing >50% of data product work",
"Stop if: domain teams complain about platform onboarding (signals platform isn't truly self-serve)",
],
)
# Mesh ambition without prerequisites
if consumers >= 25 and not culture:
return (
"LAKEHOUSE (defer mesh)",
f"{consumers} consumers is mesh-sized BUT no federated ownership culture in place; mesh "
"without culture fails. Run lakehouse with hub-and-spoke until ownership culture matures.",
[
"Revisit mesh in 18 months once 3+ domain teams own their own data products",
"Stop hub-and-spoke if central team is bottleneck > 60% of requests",
],
)
# Lakehouse: 5+ consumers OR ML workloads OR >2TB
if consumers >= 5 or ml_models >= 1 or volume >= 2:
return (
"LAKEHOUSE",
(
f"{consumers} data consumer(s), {ml_models} ML model(s) in prod, {volume}TB. "
"Pure warehouse is too rigid for ML; pure data lake too unstructured for BI. "
"Lakehouse (warehouse + object storage with table format like Iceberg/Delta) "
"covers both with one substrate."
),
[
"Downgrade to warehouse-only if ML models retired and data shrinks below 2TB",
"Upgrade to mesh only if 25+ consumers AND federated culture",
"Stop investment if vendor lock-in becomes unacceptable (lakehouse table formats mitigate this)",
],
)
# Warehouse only
return (
"WAREHOUSE ONLY",
(
f"{consumers} consumer(s), {volume}TB, {ml_models} ML model(s). Sub-scale for lakehouse "
"complexity. Single warehouse (Snowflake / BigQuery / Postgres) + dbt is the simplest viable "
"stack at this stage."
),
[
"Upgrade to lakehouse when ANY of: 5+ consumers, 2TB+ data, 1+ ML model in prod",
"Stop investment in custom modeling if SaaS BI vendor solves it (avoid premature dbt complexity)",
],
)
def build_vs_buy(profile: Dict[str, Any], architecture: str) -> List[Dict[str, str]]:
"""Returns build-vs-buy decision per layer."""
consumers = profile.get("internal_consumers", 0)
ml_models = profile.get("ml_models_in_prod", 0)
company_type = profile.get("company_type", "")
decisions = []
# Storage / warehouse
decisions.append({
"layer": "Storage / Warehouse",
"decision": "BUY",
"vendor_suggestion": "Snowflake / BigQuery / Databricks (lakehouse) or Postgres (warehouse-only)",
"rationale": "Storage is commodity. Building distributed storage is a 50-engineer-year investment with no business return unless you are a data-infra company.",
})
# ELT / ingest
decisions.append({
"layer": "ELT / Ingest",
"decision": "BUY",
"vendor_suggestion": "Fivetran / Airbyte / Stitch",
"rationale": "Connector maintenance is a moving target (200+ source APIs). Build only if your source isn't supported and is critical (then contribute upstream).",
})
# Modeling
decisions.append({
"layer": "Modeling / Transformations",
"decision": "BUILD",
"vendor_suggestion": "dbt + your domain logic (dbt itself is open source)",
"rationale": "This is your IP. Your domain logic encodes how the business actually works — vendors cannot supply it.",
})
# BI
if consumers < 100:
decisions.append({
"layer": "BI / Dashboards",
"decision": "BUY",
"vendor_suggestion": "Metabase (cheap) / Looker (enterprise) / Mode (analyst-friendly) / Hex (notebooks+BI)",
"rationale": f"At {consumers} consumers, building BI is a distraction. SaaS BI is mature; pick one that matches your analyst skillset.",
})
else:
decisions.append({
"layer": "BI / Dashboards",
"decision": "BUY + consider embedded for customer-facing analytics",
"vendor_suggestion": "Looker / Sigma + (Cube.dev or Embeddable) for customer-facing",
"rationale": f"At {consumers} consumers, BI is critical. If you're a B2B SaaS with customer-facing analytics, embedded BI is a real build-vs-buy decision; usually still buy.",
})
# Feature store
if ml_models < 3:
decisions.append({
"layer": "Feature Store",
"decision": "DEFER",
"vendor_suggestion": "(none yet — use dbt + simple feature tables)",
"rationale": f"{ml_models} model(s) in prod. Feature stores pay off at 3+ models sharing features. Premature investment is a maintenance burden.",
})
else:
decisions.append({
"layer": "Feature Store",
"decision": "BUY (Tecton / Hopsworks) or BUILD (Feast)",
"vendor_suggestion": "Tecton (managed) or Feast (open source)",
"rationale": f"{ml_models} models is the threshold where feature reuse + governance matter more than simplicity.",
})
# ML platform
if ml_models < 5:
decisions.append({
"layer": "ML Platform",
"decision": "DEFER",
"vendor_suggestion": "(none yet — use notebooks + scheduled training jobs)",
"rationale": f"{ml_models} models. ML platforms (Databricks ML, Vertex AI, SageMaker) make sense at 5+ models with active retraining; before that, the platform overhead exceeds the value.",
})
else:
decisions.append({
"layer": "ML Platform",
"decision": "BUY",
"vendor_suggestion": "Databricks ML / Vertex AI / SageMaker",
"rationale": f"{ml_models} models with active retraining. Platform handles experiment tracking, deployment, monitoring — all of which become painful to build at this scale.",
})
return decisions
def sequence_roadmap(profile: Dict[str, Any], architecture: str) -> List[Dict[str, str]]:
"""Returns 4-quarter sequencing roadmap based on priorities + architecture."""
priorities = profile.get("near_term_priorities", [])
ml_models = profile.get("ml_models_in_prod", 0)
roadmap = []
# Q1: always reliability first if pipeline issues exist
if "improve-pipeline-reliability" in priorities or "reliability" in str(priorities):
roadmap.append({
"quarter": "Q1",
"focus": "Pipeline reliability",
"deliverables": "SLA on top-3 critical pipelines (freshness, completeness); on-call rotation; data quality tests in dbt",
})
else:
roadmap.append({
"quarter": "Q1",
"focus": "Foundation",
"deliverables": "Centralized ingest (Fivetran/Airbyte); dbt for top-5 marts; basic data quality tests",
})
# Q2
if "self-serve-bi" in priorities:
roadmap.append({
"quarter": "Q2",
"focus": "Self-serve BI",
"deliverables": "BI tool rollout to non-data teams; semantic layer (dbt metrics or LookML); training program",
})
else:
roadmap.append({
"quarter": "Q2",
"focus": "Coverage",
"deliverables": "Extend dbt to top-10 marts; document data lineage; add domain-specific data quality tests",
})
# Q3
if ml_models >= 1 or "ml" in str(priorities).lower():
roadmap.append({
"quarter": "Q3",
"focus": "ML enablement",
"deliverables": "First feature-store table for top-1 production model; experiment tracking (MLflow / W&B); model monitoring",
})
else:
roadmap.append({
"quarter": "Q3",
"focus": "Embed analysts",
"deliverables": "Embedded analysts in 2-3 functional teams; central team owns platform; SLAs renegotiated",
})
# Q4: evaluate + decide
roadmap.append({
"quarter": "Q4",
"focus": "Evaluate and decide",
"deliverables": "Re-run this picker with updated profile; decide on year-2 architecture (e.g., introduce feature store, evaluate mesh prereqs)",
})
return roadmap
def analyze(profile: Dict[str, Any]) -> Dict[str, Any]:
architecture, reasoning, kill_criteria = pick_architecture(profile)
decisions = build_vs_buy(profile, architecture)
roadmap = sequence_roadmap(profile, architecture)
return {
"architecture": architecture,
"reasoning": reasoning,
"kill_criteria": kill_criteria,
"build_vs_buy": decisions,
"roadmap_12mo": roadmap,
}
def render_text(result: Dict[str, Any], profile: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("DATA PRODUCT STRATEGY")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append("Profile:")
lines.append(f" Stage: {profile.get('stage')} | Team: {profile.get('data_team_size')} | Consumers: {profile.get('internal_consumers')}")
lines.append(f" Data volume: {profile.get('data_volume_tb')}TB | ML models in prod: {profile.get('ml_models_in_prod')}")
lines.append(f" Company type: {profile.get('company_type')} | Data culture in place: {profile.get('has_data_culture')}")
lines.append("")
lines.append("-" * 72)
lines.append(f"RECOMMENDED ARCHITECTURE: {result['architecture']}")
lines.append("")
lines.append("Reasoning:")
for line in _wrap(result["reasoning"], 2):
lines.append(line)
lines.append("")
lines.append("Kill criteria (when to abandon this choice):")
for k in result["kill_criteria"]:
lines.append(f" • {k}")
lines.append("")
lines.append("-" * 72)
lines.append("BUILD vs BUY (per layer):")
lines.append("")
for d in result["build_vs_buy"]:
lines.append(f" {d['layer']:<32} {d['decision']}")
lines.append(f" Vendor: {d['vendor_suggestion']}")
for line in _wrap(f"Rationale: {d['rationale']}", 4):
lines.append(line)
lines.append("")
lines.append("-" * 72)
lines.append("12-MONTH ROADMAP:")
lines.append("")
for r in result["roadmap_12mo"]:
lines.append(f" {r['quarter']}: {r['focus']}")
for line in _wrap(r["deliverables"], 6):
lines.append(line)
lines.append("")
lines.append("-" * 72)
lines.append("REMINDER: Re-run this picker quarterly with updated profile. Architecture is not a once-")
lines.append("and-done decision — kill criteria exist for a reason.")
return "\n".join(lines)
def _wrap(text: str, indent: int, width: int = 68) -> List[str]:
import textwrap
return textwrap.wrap(text, width=width, initial_indent=" " * indent, subsequent_indent=" " * indent) or [" " * indent + text]
def main() -> int:
parser = argparse.ArgumentParser(
description="Pick data architecture + build-vs-buy + sequencing roadmap from a company profile.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to profile JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
profile = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
profile = SAMPLE
source = "<embedded sample: Series A B2B SaaS, 3-person data team>"
result = analyze(profile)
if args.output == "json":
print(json.dumps({"source": source, "profile": profile, **result}, indent=2))
else:
print(render_text(result, profile, source))
return 0
if __name__ == "__main__":
sys.exit(main())
Lãnh đạo marketing: định vị thương hiệu, mô hình tăng trưởng, phân bổ ngân sách marketing và thiết kế tổ chức.
---
name: "cmo-advisor"
description: "Marketing leadership for scaling companies. Brand positioning, growth model design, marketing budget allocation, and marketing org design. Use when designing brand strategy, selecting growth models (PLG vs sales-led vs community-led), allocating marketing budgets, building marketing teams, or when user mentions CMO, brand strategy, growth model, CAC, LTV, channel mix, or marketing ROI."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: cmo-leadership
updated: 2026-03-05
python-tools: marketing_budget_modeler.py, growth_model_simulator.py
frameworks: brand-positioning, growth-frameworks, marketing-org
---
# CMO Advisor
Strategic marketing leadership — brand positioning, growth model design, budget allocation, and org design. Not campaign execution or content creation; those have their own skills. This is the engine.
## Keywords
CMO, chief marketing officer, brand strategy, brand positioning, growth model, product-led growth, PLG, sales-led growth, community-led growth, marketing budget, CAC, customer acquisition cost, LTV, lifetime value, channel mix, marketing ROI, pipeline contribution, marketing org, category design, competitive positioning, growth loops, payback period, MQL, pipeline coverage
## Quick Start
```bash
# Model budget allocation across channels, project MQL output by scenario
python scripts/marketing_budget_modeler.py
# Project MRR growth by model, show impact of channel mix shifts
python scripts/growth_model_simulator.py
```
**Reference docs (load when needed):**
- `references/brand_positioning.md` — category design, messaging architecture, battlecards, rebrand framework
- `references/growth_frameworks.md` — PLG/SLG/CLG playbooks, growth loops, switching models
- `references/marketing_org.md` — team structure by stage, hiring sequence, agency vs. in-house
---
## The Four CMO Questions
Every CMO must own answers to these — no one else in the C-suite can:
1. **Who are we for?** — ICP, positioning, category
2. **Why do they choose us?** — Differentiation, messaging, brand
3. **How do they find us?** — Growth model, channel mix, demand gen
4. **Is it working?** — CAC, LTV:CAC, pipeline contribution, payback period
---
## Core Responsibilities (Brief)
**Brand & Positioning** — Define category, build messaging architecture, maintain competitive differentiation. Details → `references/brand_positioning.md`
**Growth Model** — Choose and operate the right acquisition engine: PLG, sales-led, community-led, or hybrid. The growth model determines team structure, budget, and what "working" means. Details → `references/growth_frameworks.md`
**Marketing Budget** — Allocate from revenue target backward: new customers needed → conversion rates by stage → MQLs needed → spend by channel based on CAC. Run `marketing_budget_modeler.py` for scenarios.
**Marketing Org** — Structure follows growth model. Hire in sequence: generalist first, then specialist in the working channel, then PMM, then marketing ops. Details → `references/marketing_org.md`
**Channel Mix** — Audit quarterly: MQLs, cost, CAC, payback, trend. Scale what's improving. Cut what's worsening. Don't optimize a channel that isn't in the strategy.
**Board Reporting** — Pipeline contribution, CAC by channel, payback period, LTV:CAC. Not impressions. Not MQLs in isolation.
---
## Key Diagnostic Questions
Ask these before making any strategic recommendation:
- What's your CAC **by channel** (not blended)?
- What's the payback period on your largest channel?
- What's your LTV:CAC ratio?
- What % of pipeline is marketing-sourced vs. sales-sourced?
- Where do your **best customers** (highest LTV, lowest churn) come from?
- What's your MQL → Opportunity conversion rate? (proxy for lead quality)
- Is this brand work or performance marketing? (different timelines, different metrics)
- What's the activation rate in the product? (PLG signal)
- If a prospect doesn't buy, why not? (win/loss data)
---
## CMO Metrics Dashboard
| Category | Metric | Healthy Target |
|----------|--------|---------------|
| **Pipeline** | Marketing-sourced pipeline % | 50–70% of total |
| **Pipeline** | Pipeline coverage ratio | 3–4x quarterly quota |
| **Pipeline** | MQL → Opportunity rate | > 15% |
| **Efficiency** | Blended CAC payback | < 18 months |
| **Efficiency** | LTV:CAC ratio | > 3:1 |
| **Efficiency** | Marketing % of total S&M spend | 30–50% |
| **Growth** | Brand search volume trend | ↑ QoQ |
| **Growth** | Win rate vs. primary competitor | > 50% |
| **Retention** | NPS (marketing-sourced cohort) | > 40 |
---
## Red Flags
- No defined ICP — "companies with 50-1000 employees" is not an ICP
- Marketing and sales disagree on what an MQL is (this is always a system problem, not a people problem)
- CAC tracked only as a blended number — channel-level CAC is non-negotiable
- Pipeline attribution is self-reported by sales reps, not CRM-timestamped
- CMO can't answer "what's our payback period?" without a 48-hour research project
- Brand work and performance marketing have no shared narrative — they're contradicting each other
- Marketing team is producing content with no documented positioning to anchor it
- Growth model was chosen because a competitor uses it, not because the product/ACV/ICP fits
---
## Integration with Other C-Suite Roles
| When... | CMO works with... | To... |
|---------|-------------------|-------|
| Pricing changes | CFO + CEO | Understand margin impact on positioning and messaging |
| Product launch | CPO + CTO | Define launch tier, GTM motion, messaging |
| Pipeline miss | CFO + CRO | Diagnose: volume problem, quality problem, or velocity problem |
| Category design | CEO | Secure multi-year organizational commitment to the narrative |
| New market entry | CEO + CFO | Validate ICP, budget, localization requirements |
| Sales misalignment | CRO | Align on MQL definition, SLA, and pipeline ownership |
| Hiring plan | CHRO | Define marketing headcount and skill profile by stage |
| Retention insights | CCO | Use expansion and churn data to sharpen ICP and messaging |
| Competitive threat | CEO + CRO | Coordinate battlecards, win/loss, repositioning response |
---
## Resources
- **References:** `references/brand_positioning.md`, `references/growth_frameworks.md`, `references/marketing_org.md`
- **Scripts:** `scripts/marketing_budget_modeler.py`, `scripts/growth_model_simulator.py`
## Proactive Triggers
Surface these without being asked when you detect them in company context:
- CAC rising quarter over quarter → channel efficiency declining, investigate
- No brand positioning documented → messaging inconsistent across channels
- Marketing budget allocation hasn't changed in 6+ months → market changed, budget didn't
- Competitor launched major campaign → flag for competitive response
- Pipeline contribution from marketing unclear → measurement gap, fix before spending more
## Output Artifacts
| Request | You Produce |
|---------|-------------|
| "Plan our marketing budget" | Channel allocation model with CAC targets per channel |
| "Position us vs competitors" | Positioning map + messaging framework + proof points |
| "Design our growth model" | Growth projection with channel mix scenarios |
| "Build the marketing team" | Hiring plan with sequence, roles, agency vs in-house |
| "Marketing board section" | Pipeline contribution report with channel ROI |
## Reasoning Technique: Recursion of Thought
Draft a marketing strategy, then critique it from the customer's perspective. Refine based on the critique. Repeat until the strategy survives scrutiny.
## Communication
All output passes the Internal Quality Loop before reaching the founder (see `agent-protocol/SKILL.md`).
- Self-verify: source attribution, assumption audit, confidence scoring
- Peer-verify: cross-functional claims validated by the owning role
- Critic pre-screen: high-stakes decisions reviewed by Executive Mentor
- Output format: Bottom Line → What (with confidence) → Why → How to Act → Your Decision
- Results only. Every finding tagged: 🟢 verified, 🟡 medium, 🔴 assumed.
## Context Integration
- **Always** read `company-context.md` before responding (if it exists)
- **During board meetings:** Use only your own analysis in Phase 2 (no cross-pollination)
- **Invocation:** You can request input from other roles: `[INVOKE:role|question]`
FILE:references/brand_positioning.md
# Brand Positioning Reference
Practical frameworks for defining, communicating, and defending your market position. Not theory — applied tools for CMOs who need to get this right.
---
## 1. Category Design Frameworks
### The Category Design Principle
Every product exists in a category — either one you define or one someone else defined. If you're not designing your category, your competitors are designing it for you, and they'll design it to exclude you.
**Category design is not renaming an existing category.** It's declaring that the existing category no longer solves the problem adequately, and that a new category — which you happen to lead — is required.
### The Three-Act Category Design Narrative
**Act 1: Name the problem**
Identify a problem that's real, growing, and underserved. Not a problem you invented — a problem your best customers articulate before they've heard your pitch.
> "Enterprise software teams are deploying faster than ever, but their security reviews still take 3 weeks — because security was built for a world where deployments happen monthly, not hourly."
**Act 2: Define the new category**
Name the category in terms of the outcome, not the feature. The category name should describe what customers achieve, not what the product does.
> "Continuous security" — not "automated security scanning" or "DevSecOps platform."
**Act 3: Position yourself as the category leader**
You can't just claim leadership — you need proof: customers, analysts, community, content, events. Leadership is built, not declared.
> "Snyk is building the continuous security category. 1.2M developers have adopted Snyk. Gartner lists us as a Cool Vendor in AppSec."
### When Category Design Works
| Condition | Explanation |
|-----------|-------------|
| Market timing | The problem is growing but the existing category is inadequate |
| CEO commitment | Category design is a 3-5 year initiative, not a marketing campaign |
| Analyst alignment | Gartner, Forrester, or G2 need to recognize your category |
| Community | Practitioners adopt the vocabulary before buyers do |
| Content moat | You publish the defining content for the category before competitors |
### Category Design Pitfalls
- **Naming the category after yourself:** "The [Your Company] Category" is not a category. It's a vanity.
- **Categories that don't solve analyst definitions:** If Gartner doesn't have a Magic Quadrant for your category, you're fighting uphill.
- **Jargon without adoption:** If your category name requires a two-paragraph explanation, it won't stick.
- **Starting a category war you can't win:** If an incumbent can copy your category name and launch in 90 days, you don't have a defensible category.
### The Lightning Strike Strategy
Category design requires concentrated, coordinated effort — not slow drip. Execute these simultaneously:
1. **Major piece of research or data** (the "State of X" report)
2. **Category-defining event** (host it, don't just attend)
3. **Analyst briefing** (educate Gartner/Forrester on the category before they define it themselves)
4. **Book or manifesto** (long-form content that becomes the category Bible)
5. **Community formation** (a Slack group, a conference, a certification that practitioners want)
Do all five within a 3-month window. This creates gravity around your category claim.
---
## 2. Messaging Architecture
### The Messaging Hierarchy
Every piece of content — from a tweet to a 60-page whitepaper — should trace back to this hierarchy. When it doesn't, you have messaging drift.
```
Level 1: Brand Promise
"[Company] [verb] [outcome] for [audience]"
→ Doesn't change. This is the north star.
Level 2: Positioning Statement (internal)
For [target customer] who [has this problem],
[Company] is the [market category] that [differentiated capability]
unlike [alternatives], [Company] [proof of differentiation].
Level 3: Value Propositions (3-4 max, one per key outcome)
Each VP: headline (5-8 words) + 2-3 sentence explanation + proof point
Level 4: Proof Points
Data, case studies, certifications, analyst recognition — evidence for each VP
Level 5: Channel Adaptations
Website copy, sales deck, ad copy, email — same hierarchy, different format
```
### Writing a Positioning Statement
The Geoffrey Moore / April Dunford format is still the best framework:
**Template:**
```
For [specific target customer]
who [has this specific, painful problem],
[Company name] is the [market category]
that [key differentiated capability].
Unlike [primary alternatives],
[Company] [proof of differentiation — something measurable or unique].
```
**Bad example (too generic):**
> For B2B companies who want to grow faster, Acme is the marketing platform that helps you get more leads. Unlike other platforms, Acme is easy to use and powerful.
**Good example (specific and falsifiable):**
> For DevOps teams in regulated industries who spend 20% of their sprint cycles on compliance reviews, Acme is the compliance automation platform that embeds regulatory checks directly into the CI/CD pipeline. Unlike manual compliance tools that create a separate review queue, Acme's policy-as-code approach reduces compliance-related cycle time by 60% without slowing deployments.
**Test your positioning statement:**
1. Can a competitor say the exact same thing? (If yes, it's not differentiated)
2. Does it describe what you do or what the customer gets? (Should be the latter)
3. Would your best customer say "yes, that's exactly my problem"? (If not, wrong ICP)
4. Is it falsifiable? (Claims you can't prove are liabilities)
### Value Proposition Development
**Structure for each VP:**
| Element | Description | Example |
|---------|-------------|---------|
| Outcome headline | What changes for the customer (5-8 words) | "Ship features 3x faster" |
| The problem | Why this matters now (1 sentence) | "Compliance reviews block 40% of releases in regulated industries" |
| Our approach | How we solve it differently (1-2 sentences) | "Policy-as-code embeds checks in the pipeline instead of adding a gate at the end" |
| Proof | Evidence this is real (1 sentence + data point) | "Customers reduce compliance cycle time by 60% in the first 90 days" |
**3-VP Architecture is the standard:**
- VP1: Core outcome (what most customers primarily buy for)
- VP2: Secondary benefit (makes the decision easier or stickier)
- VP3: Differentiator (what tips competitive decisions in your favor)
### Proof Point Hierarchy
Not all proof is equal. When you make a claim, match the strength of your proof to the importance of the claim.
| Proof Type | Strength | Best Used For |
|------------|---------|--------------|
| Third-party data (analyst report, research) | Highest | Category claims, market size |
| Customer ROI data with name | High | Value propositions |
| Customer quote with name and company | Medium-high | Specific pain points and outcomes |
| Aggregated customer data ("customers report…") | Medium | Directional claims |
| Internal testing or benchmark | Medium-low | Product capability claims |
| "Designed to…" or "built for…" | Low | Product direction only |
| "We believe…" or "we think…" | Lowest | Vision statements only |
**Proof point development process:**
1. Write the claim you want to make
2. Identify the strongest available proof
3. If proof is weak, either soften the claim or invest in getting better proof
4. Never publish a claim without knowing what happens when a skeptic asks "prove it"
---
## 3. Competitive Positioning Maps
### The Two-Axis Map
Choose two dimensions that:
1. Both matter to your target buyer
2. Create clear differentiation between you and competitors
3. You can credibly defend
**Choosing the axes:**
- Axis 1 should show a dimension where you win and most competitors cluster on the wrong side
- Axis 2 should show a dimension buyers care about deeply (ease, speed, breadth, price, compliance, etc.)
**What to avoid:**
- "Quality" vs. "Price" — too generic, every company claims the top-left
- Dimensions your competitors can match in one release cycle
- Dimensions that only your product team understands, not buyers
### Competitive Analysis Template
For each major competitor:
**Company:** _______________
| Dimension | What They Claim | What Customers Actually Experience | Gap |
|-----------|----------------|-----------------------------------|-----|
| Positioning | | | |
| Primary differentiator | | | |
| Pricing | | | |
| Ideal customer | | | |
| Weakness (win/loss data) | | | |
| What they say about you | | | |
**Sources for competitive intelligence:**
- Win/loss interviews (primary source — nothing beats this)
- G2/Capterra reviews (what customers say publicly)
- Glassdoor (tells you about internal culture and focus)
- LinkedIn job postings (what they're building next)
- Their pricing page changes (what they're competing on)
- Conference talks from their product and sales leaders
### Battlecard Format
One page per competitor. Used by sales, not marketing.
```
COMPETING AGAINST: [Competitor Name]
WHY CUSTOMERS CONSIDER THEM:
(2-3 bullets — be honest about their appeal)
OUR DIFFERENTIATION:
(2-3 bullets — factual, not marketing language)
THE LANDMINE QUESTION:
(One question that exposes their weakness. The answer should make the buyer uncomfortable choosing them.)
Example: "How long does your typical implementation take? And what's your SLA if it runs over?"
OUR PROOF POINTS IN THIS COMPARISON:
- [Customer name] switched from [competitor] after [specific reason], saw [specific result]
- [Data point that directly contradicts competitor's primary claim]
THEIR LIKELY COUNTER-MOVES:
(What will they say about us? How do we respond?)
WHEN TO WALK AWAY:
(If the prospect values X more than Y, we are not the right fit — say so)
```
---
## 4. Brand Voice Development
### What Brand Voice Is (and Isn't)
**Brand voice is NOT:**
- A list of adjectives ("we are professional, innovative, and customer-focused")
- The tone you use in formal communications
- The font and color palette (that's visual identity)
**Brand voice IS:**
- How the company sounds across every written touchpoint
- Consistent enough to be recognizable, flexible enough to be human
- Grounded in what your best customers actually value
### The Voice Attribute Framework
Define 3-4 voice attributes. For each:
1. **What it means** (in one sentence)
2. **What it sounds like** (one example)
3. **What it doesn't mean** (the common mistake that goes wrong)
**Example:**
| Attribute | Means | Sounds like | Doesn't mean |
|-----------|-------|------------|--------------|
| Direct | We say what we mean without hedging | "Your compliance review takes 3 weeks. It shouldn't." | Blunt, rude, or dismissive |
| Expert | We speak from depth, not from trend | "Here's why most security gates fail at scale, and what actually works." | Jargon-heavy or condescending |
| Honest | We acknowledge what we don't do | "We're not the best fit if you need a one-size-fits-all platform." | Self-deprecating or uncertain |
| Human | Real people write for real people | "Deploying on a Friday? Here's what we'd check first." | Casual, unprofessional |
### Voice Consistency Testing
Take a random sample of 10 recent pieces of content:
- Website homepage and pricing page
- 3 blog posts from different authors
- 5 outbound emails from sales
- 3 social posts
- 1 press release
Score each on: Does this sound like us? (1-5)
Average < 3: You have a brand voice problem. The cause is usually no documented guidelines, or guidelines that exist but aren't enforced.
### Voice in Different Contexts
The attribute stays the same. The tone adjusts.
| Context | Tone adjustment | Example of "Direct" |
|---------|----------------|---------------------|
| Homepage | Confident | "Compliance reviews don't have to slow you down." |
| Technical docs | Precise | "Set the policy threshold to 0.95 to enforce mandatory approval." |
| Error messages | Helpful | "That didn't work. Here's the most common reason why, and how to fix it." |
| Support | Empathetic | "That's frustrating. Here's what happened and what we're doing about it." |
| Sales outreach | Respectful | "Most teams in your space have this problem. Worth 20 minutes to explore?" |
---
## 5. Rebrand Decision Framework
### When Rebrands Succeed vs. Fail
**Successful rebrands:**
- Driven by a genuine strategic shift (new category, new ICP, new market)
- Have internal alignment before external launch
- Are accompanied by product and messaging changes — not just visual
- Have a 6-12 month transition plan for existing customers
**Failed rebrands:**
- Driven by internal boredom with the old brand
- Executed as a "refresh" without repositioning the value proposition
- Lack leadership conviction (executives still describe the company in the old terms)
- Launch with a new logo but same product, same messaging, same ICP
### The Rebrand Decision Matrix
Answer each question. More "yes" answers = more likely rebrand is warranted.
| Question | Yes | No |
|----------|-----|-----|
| Has our ICP changed significantly in the last 18 months? | Rebrand | Stay |
| Are we entering a new market where the current brand creates friction? | Rebrand | Stay |
| Does the brand name have negative associations in the market? | Rebrand | Stay |
| Has an acquisition changed our core identity? | Rebrand | Stay |
| Is the current brand actively hurting sales conversations? (evidence required) | Rebrand | Stay |
| Are we bored with the brand? | Stay | — |
| Did leadership change? | Stay | — |
| Are competitors rebranding? | Stay | — |
Score: 3+ "Rebrand" answers with evidence = worth a serious evaluation.
### Rebrand Risk Assessment
**Name change** is the highest-risk rebrand element. Before committing:
- Legal: trademark availability in all target markets
- SEO: 18-24 months to recover domain authority after a domain change
- Customer: existing customers need to update all integrations, contracts, documentation
- Analyst: re-education of Gartner, Forrester, G2 category definitions
- Employee: company identity shift is a culture event, not just an HR task
**Minimum viable rebrand (lower risk):**
1. New positioning and messaging (always worth doing if positioning is wrong)
2. Visual identity refresh (keep the name, update the look)
3. Tagline change (the cheapest, lowest-risk brand change)
**Full rebrand (high risk, sometimes necessary):**
1. New company name and domain
2. New visual identity
3. New positioning and messaging
4. New category narrative
### Rebrand Execution Checklist
**Pre-launch (90 days):**
- [ ] Finalize positioning before finalizing design (in that order)
- [ ] Legal trademark clearance in all target markets
- [ ] Domain secured (with redirects planned)
- [ ] Internal alignment: every leader can describe the new positioning in one sentence
- [ ] Customer comms plan (existing customers, especially enterprise, need advance notice)
- [ ] Analyst briefings scheduled (Gartner, Forrester — brief them before launch)
- [ ] PR plan finalized
**Launch (day 1):**
- [ ] Website flipped
- [ ] Social profiles updated
- [ ] Email signatures updated company-wide
- [ ] Sales deck updated
- [ ] Press release published
- [ ] Existing customers notified (email from CEO or CMO, not marketing automation)
**Post-launch (90 days):**
- [ ] SEO monitoring (watch for ranking drops on key terms)
- [ ] Win rate monitoring (did conversion change?)
- [ ] Employee feedback (are they using the new messaging correctly?)
- [ ] Partner/channel update (resellers, integrations, directories)
- [ ] Analyst follow-up (did they update their reports?)
---
## Quick Reference: Brand Positioning Diagnostic
Use this as an audit against your current positioning:
| Check | Pass | Fail |
|-------|------|------|
| Can every sales rep state the positioning in one sentence without looking it up? | ✓ | Positioning isn't working |
| Is the ICP specific enough to disqualify companies? | ✓ | ICP is too broad |
| Does the homepage lead with customer outcome, not product features? | ✓ | Copy needs rewrite |
| Can you name 3 companies you're NOT a good fit for? | ✓ | Positioning is unfocused |
| Do win/loss interviews confirm the stated differentiator? | ✓ | Differentiator is assumed, not proven |
| Is the category name used by analysts or industry media? | ✓ | Category design needed |
| Does every piece of content trace back to a VP from the hierarchy? | ✓ | Messaging drift — need guidelines |
FILE:references/growth_frameworks.md
# Growth Frameworks Reference
Playbooks for PLG, sales-led, community-led, and hybrid growth models. Includes growth loops, funnel design, and guidance on when and how to switch models.
---
## 1. Product-Led Growth (PLG) Playbook
### What PLG Actually Is
PLG means the product is the primary distribution mechanism. Not "we have a free trial." Not "our product is self-serve." PLG means the product creates acquisition, retention, and expansion — and does so at a scale and cost no sales team can match.
**The minimum requirements for PLG to work:**
1. **Fast time-to-value:** Users must get a meaningful outcome within one session (ideally < 30 minutes)
2. **Low friction to start:** No sales call, no implementation project, no credit card required (for top of funnel)
3. **Built-in virality or network effects:** Usage creates exposure or value that draws in other users
4. **Self-serve monetization or expansion path:** Freemium → paid, or individual → team → company
If any of these is missing, you don't have PLG — you have a website with a free trial.
### PLG Funnel: The Four Stages
**Stage 1: Acquisition**
The user discovers and signs up for the product without talking to sales.
Key channels:
- Organic search (SEO targeting jobs-to-be-done searches)
- Product hunt launches
- Referral and invite loops (users share the product with colleagues)
- Developer communities and open-source contributions
Metric: Visitor-to-signup rate
Benchmark: 2-8% for B2B SaaS (varies heavily by product complexity)
**Stage 2: Activation**
The user reaches the "aha moment" — the point where the product delivers its core value for the first time.
Finding the aha moment:
- Look at the behaviors that differentiate users who stay from users who churn in the first 30 days
- The aha moment is not creating an account. It's completing the first outcome.
- For Slack: sending a message in a real channel
- For Dropbox: adding a file from a second device
- For HubSpot: publishing a form that captures a real lead
Metric: Activation rate (% of signups who complete the aha moment action within 7 days)
Benchmark: 25-40% is strong. < 15% means the onboarding is broken.
**Stage 3: Retention**
Users return to the product and build habitual use.
Retention analysis:
- Cohort retention curves (by signup week/month)
- Day 1, Day 7, Day 30, Day 90 retention rates
- Feature adoption by retained vs. churned users (which features predict retention?)
Metric: D30 retention rate (% of users still active 30 days after signup)
Benchmark: > 40% D30 retention is strong for B2B products
**Stage 4: Revenue**
Self-serve conversion from free to paid, or expansion from individual to team.
PQL (Product-Qualified Lead) signals:
- Reached a usage limit (invites, storage, seats)
- Used a premium feature in trial mode
- Team size on the account reached a threshold
- High-frequency usage above a defined threshold
Metric: PQL conversion rate (% of PQLs who convert to paid within 30 days)
Benchmark: 15-30% for well-designed PLG products
### PLG Expansion Model
PLG growth compounds through account expansion:
```
Individual user discovers product
→ Gets value, invites teammates
→ Team adopts product
→ Becomes department-wide
→ Finance/IT gets involved
→ Enterprise contract
```
This is "bottom-up" enterprise: individual adoption precedes company-wide purchase. It's also the most defensible moat — when every engineer in the company uses your product individually, procurement cancellation is very hard.
**Expansion levers:**
- Seat-based pricing (more users = more revenue, aligned with value)
- Usage-based pricing (more usage = more value = more revenue)
- Feature gating (team/enterprise features visible but gated, creating pull to upgrade)
- Admin discovery (usage reports surface to managers who didn't know they had a product champion)
### PLG Diagnostic
| Question | Healthy | Unhealthy |
|----------|---------|-----------|
| Time-to-value | < 30 minutes | > 2 hours |
| Activation rate | > 30% | < 15% |
| D30 retention | > 40% | < 20% |
| PQL conversion | > 15% | < 5% |
| NPS from self-serve users | > 40 | < 20 |
| Viral coefficient | > 0.3 | < 0.1 |
### PLG Team Structure
```
Head of Growth (often VP Product or VP Marketing)
├── Growth PM (owns activation and retention loops in product)
├── Growth Engineer (2-3 engineers dedicated to growth experiments)
├── Data Analyst (experimentation, funnel analysis, cohort reports)
└── Growth Marketer (acquisition, SEO, referral programs)
```
The growth team sits between product and marketing. This is intentional — they own the product loops that drive acquisition and retention.
---
## 2. Sales-Led Growth (SLG) Model
### The SLG System
In SLG, marketing's job is to fill the sales pipeline. Sales converts it. The system only works if marketing and sales agree on definitions, SLAs, and shared metrics.
**The SLG funnel:**
```
Awareness (Impressions, reach, brand search)
↓
Lead (Name + contact info captured)
↓
MQL — Marketing Qualified Lead (meets ICP criteria, intent signal detected)
↓ [Marketing → Sales handoff]
SAL — Sales Accepted Lead (sales reviews and accepts the lead)
↓
SQL — Sales Qualified Lead (sales confirms budget, authority, need, timeline)
↓
Opportunity (Formal deal in pipeline, has a close date)
↓
Closed-Won
```
**The MQL definition problem:**
Most marketing-sales friction traces to an unclear MQL definition. The MQL should be:
- ICP-matched (company size, industry, role)
- Intent-signaled (visited pricing page, attended webinar, downloaded high-intent content)
- Not just email address + "subscribed to newsletter"
**A concrete MQL definition:**
> Company 50-500 employees, B2B SaaS, role is VP Engineering or CTO or CISO, AND has performed 2+ of: attended webinar, visited pricing page, requested demo, downloaded security report, attended event.
This definition makes the MQL useful. If you can't score it in your CRM without human judgment, it's not a definition — it's a guideline.
### SLG Conversion Rate Benchmarks
| Stage | Average B2B SaaS | Top Quartile |
|-------|-----------------|--------------|
| Lead → MQL | 5-15% | > 20% |
| MQL → SAL | 50-70% | > 75% |
| SAL → SQL | 30-50% | > 60% |
| SQL → Opportunity | 60-80% | > 85% |
| Opportunity → Closed-Won | 20-30% | > 40% |
**End-to-end:** Lead → Closed-Won: 1-5% (wide range by ACV and ICP quality)
### Pipeline Coverage Mechanics
A healthy SLG pipeline has 3-4x coverage against quota.
If a sales rep has a $500K quarterly quota:
- They need $1.5M-$2M in active pipeline
- Pipeline must be distributed across stages (not all "prospecting")
- Stage distribution benchmark: 30% early, 40% mid, 30% late
Insufficient coverage (< 3x) is a lagging indicator of a miss — by the time coverage is low, it's too late to recover in the same quarter. Coverage should be tracked weekly.
### SLG Demand Generation Channels
**High-intent channels (bottom of funnel):**
- Paid search on buying-intent keywords (e.g., "[competitor] alternative", "best [category] software")
- Review site presence (G2, Capterra) — buyers use these before vendor websites
- Outbound SDR targeting specific accounts (ABM)
**Medium-intent channels (middle of funnel):**
- Webinars and virtual events (capture active learners)
- Gated content (guides, benchmarks, templates — ICP-specific)
- Retargeting to website visitors
**Awareness channels (top of funnel):**
- Content and SEO (captures people learning about the problem)
- Podcast sponsorships, industry media
- Conference sponsorship and speaking
- Paid social (LinkedIn for B2B)
### ABM (Account-Based Marketing) in SLG
ABM flips the funnel: instead of generating leads and filtering for good ones, you start with target accounts and run coordinated campaigns against them.
**Tiers:**
- **Tier 1 (1:1):** 5-20 strategic accounts, fully customized campaigns, dedicated SDR+AE pairs, executive outreach
- **Tier 2 (1:few):** 50-200 accounts, programmatic personalization, SDR sequences, targeted events
- **Tier 3 (1:many):** 500+ accounts, standard campaigns with light personalization
ABM requires tight sales/marketing alignment. If sales doesn't work the accounts marketing targets, ABM produces zero results.
---
## 3. Community-Led Growth (CLG)
### The CLG Thesis
Community-led growth works when:
1. Your buyers want to learn from peers, not vendors
2. There's a strong practitioner identity (developers, data teams, security, FinOps)
3. Your category is complex enough that buyers need education before purchasing
4. You can commit to building genuine community, not a marketing channel in disguise
**The fundamental rule of CLG:** The community must deliver value to members whether or not they ever buy your product. If the only purpose of the community is to sell to members, the community will die.
### CLG Stages
**Stage 1: Find the community**
The community often exists before you build it. Find where your practitioners already gather:
- Slack groups, Discord servers
- Subreddits and LinkedIn groups
- Conference hallways
- Open-source repositories
Before building, participate. Earn trust. Understand the conversations.
**Stage 2: Become the knowledge hub**
Establish your company as the best source of information on the category problem:
- Publish the benchmark study everyone references
- Host the conference that defines the industry
- Create the certification practitioners want on their resume
- Open-source the tools the community needs
**Stage 3: Build the platform**
Create a dedicated community space (Slack, Discord, forum):
- Community must be practitioner-first, not vendor-first
- Community managers who genuinely care about member value
- Content from members, not just from your company
- Events that build member relationships, not just product demos
**Stage 4: Convert community to customers**
Community members who become customers do so because they trust you, not because you sold them. Conversion paths:
- Community members see peer success with your product
- Product-qualified signals from community members who trial the product
- Direct outreach from sales to active community members (with permission and context)
- Enterprise deals from companies whose employees are active in the community
### CLG Metrics
| Metric | Definition | Health Signal |
|--------|-----------|--------------|
| Monthly active members | Members who post, comment, or engage | > 15% of total members |
| Community-sourced pipeline | $ pipeline where community was first touch | Track and trend |
| Community-influenced pipeline | $ pipeline with any community touchpoint | > 30% of total pipeline |
| NPS of community members vs. non-members | Loyalty difference | Community members should score 20+ pts higher |
| Member-generated content % | % of content posted by non-employees | > 60% is healthy community |
| Time from community join to product trial | | Shortens as community matures |
### CLG Anti-Patterns
- **Community as a newsletter:** If members can't interact with each other, it's not a community — it's a list.
- **Product launches in the community:** Nothing kills community trust faster than using it for sales announcements.
- **Community without a community manager:** Communities left to run themselves become ghost towns or become toxic.
- **Measuring community by member count:** Ghost members are noise. Active engagement is signal.
---
## 4. Hybrid Growth Models
### PLG + SLG ("Product-Led Sales" or PLS)
The most common hybrid at growth stage. PLG handles SMB self-serve; sales closes enterprise.
**The PQL-to-sales handoff:**
Define the triggers that move a product-qualified lead to a sales-assisted motion:
- Company has > X users (e.g., 10+ users on a team account)
- Usage exceeds Y threshold in 30 days
- Account is a named target in the ABM list
- User explicitly requested a demo or upgrade assistance
**The risk:** Sales team ignores PLG pipeline because deal size is smaller. Fix: separate quotas and commission structures for self-serve expansion vs. new enterprise logos.
**The opportunity:** PLG creates pre-qualified champions inside accounts. Sales doesn't have to create interest — they convert it. Win rates in PLS motions are typically 30-50% higher than cold outbound.
### SLG + CLG
Community builds brand and generates inbound pipeline for sales.
This hybrid works when:
- Sales cycles are long (6-18 months)
- Buyers do extensive research before engaging with vendors
- The community validates your credibility before sales conversations begin
**The integration:**
- Community team feeds content insights to demand gen
- Event attendees become high-priority SDR sequences
- Active community members get dedicated AE outreach with community context
- Win/loss analysis includes community touchpoints
### PLG + CLG
The developer/open-source hybrid. PLG handles product adoption; community handles advocacy and content.
**Examples:** HashiCorp (Terraform community + enterprise sales), Elastic (open-source + community + commercial), Tailscale (developer community + self-serve + enterprise).
**How it compounds:**
```
Community member learns from community content
→ Discovers open-source or free tier
→ Gets value in first session
→ Shares experience in community
→ New members discover product through community content
```
---
## 5. Growth Loops vs. Funnels
### The Difference
**A funnel** is linear. It requires constant input at the top to produce output at the bottom. If you stop feeding it, it stops producing.
**A growth loop** is cyclical. Output from one stage becomes input to the next. The system compounds.
### Common Growth Loops
**Viral loop:**
```
User gets value → Invites colleague → Colleague signs up →
Colleague invites another colleague → ...
```
Viral coefficient (K) = (Average invites per user) × (Conversion rate of invites)
- K > 1: Exponential growth (rare)
- K 0.5-1: Strong viral assist
- K < 0.3: Viral is not a meaningful growth driver
**Content SEO loop:**
```
Publish content on [topic] → Ranks in search →
Drives signups → Users share content → Builds backlinks →
Better rankings → More content is possible
```
This loop takes 12-24 months to activate but is extraordinarily defensible once running.
**UGC (User-Generated Content) loop:**
```
Users share their work publicly (templates, analyses, portfolios) →
Others discover the work → They find the product →
They create and share their own work → ...
```
Figma, Notion, Airtable, Canva — all run this loop.
**Data network effect loop:**
```
More users → More data → Better product →
More users attracted → ...
```
LinkedIn, Waze, Duolingo — accuracy or relevance improves as the user base grows.
**Integration loop:**
```
Product integrates with X → X's users discover your product →
More integrations possible → More discovery surfaces → ...
```
Zapier, Slack apps, Salesforce AppExchange — being in the ecosystem creates distribution.
### Building a Growth Loop
**Step 1: Map the current funnel**
Where do customers come from? What are the conversion steps?
**Step 2: Find the output**
What does a successful customer produce?
- Invite emails
- Shared content
- Public work visible to others
- Reviews or testimonials
**Step 3: Design the loop**
How does that output become tomorrow's input to acquisition?
- If they share → is there a landing page that captures the new visitor?
- If they invite → is the invite experience friction-free?
- If they create content → does it rank in search or appear in relevant communities?
**Step 4: Measure loop velocity**
For each loop, measure:
- Cycle time: How long does one full cycle take?
- Conversion at each step: Where does the loop break down?
- Loop coefficient: How many new users does one existing user generate?
---
## 6. When to Switch Growth Models
### The Warning Signs
**PLG-to-SLG triggers:**
- Enterprise accounts are signing up via PLG but aren't expanding without human intervention
- Average deal sizes in enterprise are 10-20x SMB, and you're leaving revenue on the table
- Product adoption in enterprise requires configuration or integration that needs support
- PLG accounts churn at higher rates than sales-assisted accounts
**SLG-to-PLG/PLS triggers:**
- CAC is increasing year-over-year as competition for sales talent intensifies
- Smaller competitors are winning deals with self-serve
- Customers are asking "can I just try this myself?"
- ACV is declining as the market matures and products commoditize
- Sales team efficiency (revenue per sales rep) is declining
**Adding CLG to existing motion:**
- Sales cycles are long and trust is the primary barrier
- SEO and content are generating traffic but low conversion (awareness without trust)
- Competitors are building community and you're not present
- Customer success teams report that customers who participate in user groups retain better
### The Transition Playbook
**Phase 1: Prove it before scaling (months 1-6)**
Don't restructure the team to support the new model before proving it works.
- Run a pilot: 3-5 SDRs testing PLG signals as outreach triggers (for PLG → PLS)
- Or: Launch a beta community with 100 core customers (for adding CLG)
- Measure the metrics of the new model, compare to current model
**Phase 2: Parallel running (months 6-12)**
Run both models simultaneously. Don't kill the current model while building the new one.
- Set clear boundaries on which accounts go to which motion
- Build dedicated teams for each model (don't ask the same people to do both)
- Define success metrics for the new model independently
**Phase 3: Rebalance (months 12-18)**
Once the new model proves its unit economics:
- Shift headcount and budget to the more efficient model
- Keep the old model for the segments where it still works
- Document what the new model requires to sustain itself
**The anti-pattern:** Announcing a model shift without proof, restructuring the team, and discovering after 12 months that the new model doesn't work. By then, the old model's momentum is gone and you've burned a year.
### Growth Model Maturity Matrix
| Dimension | PLG | SLG | CLG |
|-----------|-----|-----|-----|
| Time to first results | 3-6 months | 1-3 months | 12-18 months |
| Requires up-front product investment | High | Low | Medium |
| Scales without linear headcount | Yes | No | Yes |
| Predictable pipeline | Low (early) | High | Low (early) |
| CAC trend over time | Decreases | Flat/increases | Decreases |
| Works for ACV > $50K | Only with SLG assist | Yes | Yes |
| Works for ACV < $5K | Yes | No | Only with PLG |
| Defensibility once established | High | Low | Very high |
FILE:references/marketing_org.md
# Marketing Org Reference
Team structure, hiring sequence, agency decisions, marketing ops, and cross-functional alignment — by company stage.
---
## 1. Marketing Team Structure by Stage
### Pre-Seed / Seed (< $1M ARR, 1–10 people)
Don't hire a marketing team yet. The founders are the marketing team.
What to do instead:
- Founders write content, do sales calls, go to events
- The goal is learning the ICP and finding the channel that works, not scaling anything
- One contractor or agency for specific output (design, SEO audit) is fine
First marketing hire trigger: You have a repeatable sales motion and need to scale it.
---
### Series A ($1M–$5M ARR, 10–30 people)
**Org:**
```
Founding Marketer (Head of Marketing or VP Marketing)
```
One person. Generalist. Capable of writing, running ads, setting up HubSpot, producing a report. Their job is to find what works.
**What they own:**
- Content and SEO foundation
- Paid channel experiments
- Sales enablement basics (1-pager, deck, email sequences)
- Event presence (1-2 conferences)
- Marketing attribution setup (get this right early)
**What they don't own yet:**
- Brand redesign
- Analyst relations
- Partner marketing
- Field marketing team
**CMO vs. VP Marketing at this stage:** VP Marketing. An experienced operator who can build and execute. A CMO's strategic value isn't fully leveraged until there's a team to lead and a budget to allocate.
---
### Series B ($5M–$20M ARR, 30–80 people)
**PLG-first org:**
```
VP Marketing
├── Growth Marketing (acquisition loops, activation, PLG analytics)
├── Product Marketing (positioning, launch, sales enablement)
└── Content & SEO (organic engine)
```
**SLG-first org:**
```
VP Marketing
├── Demand Generation (pipeline creation, paid, digital)
├── Product Marketing (positioning, competitive intel, enablement)
├── Field Marketing (events, regional, ABM)
└── Marketing Operations (CRM, attribution, reporting)
```
**Community-led org:**
```
VP Marketing
├── Community & Developer Relations
├── Content & SEO
└── Product Marketing
```
**At this stage:** Marketing ops becomes critical. Without it, attribution is guesswork and the sales team blames marketing for bad leads.
---
### Series C ($20M–$75M ARR, 80–200 people)
```
CMO
├── Demand Generation
│ ├── Paid Media
│ ├── SEO & Content
│ └── Marketing Operations
├── Product Marketing
│ ├── Core PMMs (by product line or segment)
│ └── Competitive Intelligence
├── Field Marketing
│ ├── Events
│ └── Regional / ABM
└── Brand & Communications
├── Brand Design
└── PR / Analyst Relations
```
**At this stage:**
- The CMO is a board-level communicator, not a campaign manager
- Each function has a dedicated leader (director or VP level)
- Marketing ops owns the attribution model and reports to CMO directly
- Analyst relations becomes important (Gartner, Forrester, G2 category positioning)
---
### Growth Stage ($75M+ ARR)
Marketing becomes a portfolio of specialized functions. Each major channel has a team. Brand is a serious investment. Analyst relations is a dedicated role. International marketing teams form.
The CMO's job shifts from building the machine to:
- Setting marketing strategy across a complex portfolio
- Representing marketing at the board level
- Owning brand and category leadership
- Cross-functional leadership with CRO, CPO, CEO
---
## 2. Hiring Sequence
### Who to Hire First
**The generalist content + demand gen marketer.**
Must-haves:
- Can write (blog posts, emails, landing pages — not just briefs)
- Can run paid campaigns (Google, LinkedIn — not just "I've managed agencies")
- Can operate a marketing automation platform (HubSpot, Marketo)
- Comfortable with data (can build a funnel report without asking an analyst)
This person builds the foundation. They're not a specialist yet — they're testing channels and building the process.
Avoid: Hiring a brand designer first. Or a community manager. Or a social media manager. These are specialties that compound on a foundation that doesn't exist yet.
### Who to Hire Second
**A specialist in the channel that's working.**
If organic search is your top lead source → hire an SEO/content lead.
If events are driving pipeline → hire a field marketer.
If outbound is working → hire an SDR manager or demand gen specialist.
Don't hire a generalist #2. By now you know what's working. Depth beats breadth.
### Who to Hire Third
**Product marketing.**
Why third and not first? Because PMM output (positioning, sales enablement, launch) is most valuable when there's an audience to position to and a sales team to enable. Before that, the founding marketer does "good enough" PMM work.
PMM hire profile: Has done positioning work before, has run a product launch, has built sales decks that sales actually uses, comfortable with win/loss analysis.
PMM:PM ratio benchmark: 1 PMM per 2–3 PMs. If you have 6 PMs and 1 PMM, you have a messaging and enablement problem.
### Who to Hire Fourth
**Marketing operations.**
This is consistently hired too late. By the time most companies hire marketing ops, attribution is broken, leads are being lost in handoffs, and the CRM data is unreliable. Hire marketing ops before you think you need it.
Marketing ops profile: HubSpot/Marketo certified, SQL capable, understands multi-touch attribution, has integrated CRM + sales engagement tools before.
### Hiring Decision Triggers
| Hire | Trigger |
|------|---------|
| Generalist marketer #1 | Sales motion is repeatable, need to scale lead generation |
| Specialist #2 | One channel is clearly outperforming — double down |
| Product marketer | Sales team is losing deals to positioning confusion or competitor gaps |
| Marketing ops | Running 3+ campaigns simultaneously with manual tracking |
| Field marketer | Events are in the strategy and attendance > 2 conferences/quarter |
| Head of Marketing / VP | Team is 3+ people and needs an org owner |
| CMO | Company is Series B/C and marketing needs board-level representation |
---
## 3. Agency vs. In-House
### Framework
Keep in-house what compounds. Outsource what's episodic or specialized.
| Function | Agency | In-House | Notes |
|----------|--------|----------|-------|
| Brand design | Early stage | Series B+ | Agency fine until redesigns become frequent |
| Paid media | < $50K/month spend | > $50K/month | Agency margin eats returns at scale |
| SEO strategy | Audit only | Ongoing execution | Strategy once, execution continuously |
| Content production | Overflow only | Core writers | Your voice must be yours |
| PR / comms | Almost always | $100M+ companies | Specialists required for media relationships |
| Marketing ops / CRM | Never | Always | This is your data infrastructure |
| Analyst relations | Initial strategy | Ongoing | Relationship-based — needs dedicated owner |
| Video / creative production | Always | Rarely | Episodic, specialized equipment |
### Agency Red Flags
- They want to own your ad accounts. (Always keep ownership. No exceptions.)
- SLA is "5 business days for creative requests." For a performance channel, that's too slow.
- Reporting is impressions, CPM, and "brand lift." Where's the pipeline?
- They can't tell you your CAC from their channel.
- They won't share the actual data — only their dashboard.
- Your account manager changes every 6 months.
### Agency Evaluation Criteria
1. **Proof of work in your category** — ask for 3 case studies with actual CAC and pipeline data
2. **Who actually does the work** — senior pitch team ≠ junior execution team
3. **Account ownership** — all accounts, pixels, analytics must be in your name
4. **Reporting cadence** — weekly data, monthly strategy, quarterly business review
5. **Exit terms** — how do you offboard without losing your data, accounts, and history?
---
## 4. Marketing Ops and Tech Stack
### The Minimum Viable Stack
| Layer | Tool | Purpose |
|-------|------|---------|
| CRM | HubSpot / Salesforce | Contact database, pipeline, source of truth |
| Marketing automation | HubSpot / Marketo / ActiveCampaign | Email, nurture, lead scoring |
| Analytics | Google Analytics 4 + Segment | Traffic, behavior, event tracking |
| Attribution | HubSpot / Attributer.io / Dreamdata | Multi-touch pipeline attribution |
| Paid | Google Ads + LinkedIn Ads | Performance channels |
| SEO | Ahrefs / Semrush | Keyword research, rank tracking |
| Chat/conversion | Intercom / Drift | In-product + website conversion |
**The integration that breaks most:** CRM ↔ Marketing automation ↔ Sales engagement. When these aren't synced properly, leads are lost, attribution is wrong, and marketing and sales fight about pipeline. Fix this first.
### Marketing Ops Ownership
Marketing ops must own:
- CRM data quality (field standardization, deduplication, routing)
- Lead scoring model (and quarterly review against conversion data)
- Attribution model (with documented assumptions)
- Campaign tracking (UTM governance — no UTM = no attribution)
- Tech stack evaluation and contracts
Marketing ops must NOT own:
- Strategy (they enable it, not set it)
- Content production
- Campaign creative
---
## 5. Cross-Functional Alignment
### Marketing + Sales
The most important cross-functional relationship in a SLG company. Where it breaks:
| Problem | Root Cause | Fix |
|---------|-----------|-----|
| "Marketing sends us bad leads" | MQL definition is unclear or wrong | Define MQL jointly, score against conversion data |
| "Sales doesn't follow up on leads" | No SLA, no consequence | Define SLA (e.g., 24-hour response), track in CRM |
| "Marketing doesn't understand what customers care about" | No win/loss sharing | Weekly call: sales shares 3 deal insights, marketing shares 3 content results |
| "We don't know what's working" | Attribution is broken | Marketing ops fixes attribution before next budget cycle |
**The SLA agreement (document this):**
- Marketing commits: X MQLs/week meeting defined criteria, 48-hour SLA from form fill to SDR outreach
- Sales commits: All MQLs contacted within 24 hours, disposition logged in CRM within 5 days
### Marketing + Product
Where it breaks and how to fix it:
| Problem | Fix |
|---------|-----|
| PMM learns about launches 2 weeks before ship | PMM joins the product planning process at the roadmap stage, not the sprint stage |
| Feature launches with no messaging | Launch tiers: Tier 1 (major, full launch), Tier 2 (minor, release notes + 1 post), Tier 3 (internal only) |
| Product doesn't use customer insights from marketing | Monthly session: PMM shares win/loss themes, competitive intel, ICP data |
| No feedback loop on messaging in-product | PMM owns in-product copy review, not just external comms |
### Marketing + Customer Success
Customer success is marketing's best source of truth:
- **ICP validation:** Which customers are expanding? Which are churning? This refines who you target.
- **Proof points:** CS-sourced case studies and testimonials outperform vendor-written content 3:1 in conversion.
- **Messaging test:** If CS is answering the same question 20 times, marketing hasn't explained it clearly enough.
- **Referral programs:** CS owns the relationship; marketing owns the mechanics. Design them together.
Cadence: Monthly meeting between CMO and VP/Head of CS. Agenda: retention trends, expansion patterns, at-risk customers, NPS themes.
FILE:scripts/growth_model_simulator.py
#!/usr/bin/env python3
"""
Growth Model Simulator
----------------------
Projects MRR growth across different growth models (PLG, sales-led, community-led,
hybrid) and shows the impact of channel mix changes on growth trajectory.
Usage:
python growth_model_simulator.py
Inputs (edit INPUTS section):
- Starting MRR and churn rate
- Current channel mix (% of new MRR from each source)
- Conversion rates per model
- Growth rate assumptions per channel
Outputs:
- 12-month MRR projection by growth model
- Channel mix impact analysis (what happens if you shift mix)
- Break-even months for each model
- Side-by-side comparison table
"""
from __future__ import annotations
import math
from dataclasses import dataclass, field
from typing import Dict, List, Optional, Tuple
# ---------------------------------------------------------------------------
# Data models
# ---------------------------------------------------------------------------
@dataclass
class ChannelSource:
name: str
pct_of_new_mrr: float # Current share of new MRR (0.0–1.0)
monthly_growth_rate: float # How fast this channel grows month-over-month
cac: float # CAC in dollars
payback_months: float # Months to recover CAC
@dataclass
class GrowthModel:
name: str
description: str
channel_mix: Dict[str, float] # channel name → % of new MRR
new_mrr_monthly_base: float # Starting new MRR/month from this model
monthly_acceleration: float # Acceleration factor (compounding)
avg_ltv_cac: float # Expected LTV:CAC at scale
months_to_steady_state: int # Months before model hits its natural growth rate
notes: List[str] = field(default_factory=list)
@dataclass
class MonthSnapshot:
month: int
mrr: float
new_mrr: float
churned_mrr: float
expansion_mrr: float
net_new_mrr: float
cumulative_cac_spend: float
@dataclass
class ModelProjection:
model: GrowthModel
snapshots: List[MonthSnapshot]
break_even_month: Optional[int] # Month when cumulative revenue > cumulative CAC
# ---------------------------------------------------------------------------
# INPUTS — edit these
# ---------------------------------------------------------------------------
STARTING_MRR = 85_000 # Current MRR ($)
MONTHLY_CHURN_RATE = 0.012 # Monthly churn rate (1.2% = ~14% annual)
EXPANSION_RATE = 0.008 # Monthly expansion MRR as % of existing MRR
GROSS_MARGIN = 0.75
SIMULATION_MONTHS = 18
# Channel sources (used to model mix shift scenarios)
CHANNELS: List[ChannelSource] = [
ChannelSource("Organic/SEO", pct_of_new_mrr=0.28, monthly_growth_rate=0.04, cac=1_800, payback_months=9),
ChannelSource("PLG Self-Serve", pct_of_new_mrr=0.15, monthly_growth_rate=0.08, cac=900, payback_months=5),
ChannelSource("Outbound SDR", pct_of_new_mrr=0.25, monthly_growth_rate=0.02, cac=5_100, payback_months=21),
ChannelSource("Paid Search", pct_of_new_mrr=0.15, monthly_growth_rate=0.01, cac=6_200, payback_months=26),
ChannelSource("Events/Field", pct_of_new_mrr=0.08, monthly_growth_rate=0.01, cac=9_800, payback_months=41),
ChannelSource("Partner/Channel", pct_of_new_mrr=0.09, monthly_growth_rate=0.05, cac=3_400, payback_months=14),
]
# Growth models to simulate
GROWTH_MODELS: List[GrowthModel] = [
GrowthModel(
name="Current Mix",
description="Baseline — maintain current channel allocation",
channel_mix={"Organic/SEO": 0.28, "PLG Self-Serve": 0.15, "Outbound SDR": 0.25,
"Paid Search": 0.15, "Events/Field": 0.08, "Partner/Channel": 0.09},
new_mrr_monthly_base=12_000,
monthly_acceleration=0.025,
avg_ltv_cac=3.2,
months_to_steady_state=3,
notes=["Baseline. No changes to channel mix."],
),
GrowthModel(
name="PLG-First",
description="Shift budget toward PLG self-serve and organic; reduce paid and outbound",
channel_mix={"Organic/SEO": 0.35, "PLG Self-Serve": 0.35, "Outbound SDR": 0.10,
"Paid Search": 0.08, "Events/Field": 0.04, "Partner/Channel": 0.08},
new_mrr_monthly_base=9_500, # Slower start — PLG takes time to activate
monthly_acceleration=0.048, # But compounds faster
avg_ltv_cac=5.8,
months_to_steady_state=6, # PLG loops take time to build
notes=[
"Lower new MRR in months 1-6 while PLG loops activate.",
"Acceleration compounds strongly after month 6.",
"Requires product investment in activation/onboarding.",
"Best fit if time-to-value < 30 min and viral coefficient > 0.3.",
],
),
GrowthModel(
name="Sales-Led Scale",
description="Double down on outbound SDR and field; optimize for enterprise ACV",
channel_mix={"Organic/SEO": 0.20, "PLG Self-Serve": 0.05, "Outbound SDR": 0.40,
"Paid Search": 0.15, "Events/Field": 0.15, "Partner/Channel": 0.05},
new_mrr_monthly_base=15_000, # Higher new MRR from enterprise ACV
monthly_acceleration=0.018, # Linear growth — headcount-constrained
avg_ltv_cac=2.8,
months_to_steady_state=2,
notes=[
"Fastest short-term new MRR if ACV > $30K.",
"Growth is linear — adds headcount to add pipeline.",
"CAC and payback worsen as SDR market tightens.",
"Requires sales capacity increase to sustain.",
],
),
GrowthModel(
name="Community-Led",
description="Invest in community and content; reduce paid; long-term brand play",
channel_mix={"Organic/SEO": 0.45, "PLG Self-Serve": 0.15, "Outbound SDR": 0.15,
"Paid Search": 0.05, "Events/Field": 0.10, "Partner/Channel": 0.10},
new_mrr_monthly_base=7_000, # Slowest start
monthly_acceleration=0.038,
avg_ltv_cac=4.5,
months_to_steady_state=9, # Community takes longest to activate
notes=[
"Lowest new MRR in months 1-9.",
"Community trust drives lower CAC and higher retention at scale.",
"Best for categories where buyers seek peer validation.",
"Requires dedicated community manager from day one.",
],
),
GrowthModel(
name="Hybrid PLS",
description="PLG self-serve for SMB + sales-assisted for enterprise (Product-Led Sales)",
channel_mix={"Organic/SEO": 0.30, "PLG Self-Serve": 0.28, "Outbound SDR": 0.22,
"Paid Search": 0.08, "Events/Field": 0.06, "Partner/Channel": 0.06},
new_mrr_monthly_base=11_000,
monthly_acceleration=0.035,
avg_ltv_cac=4.1,
months_to_steady_state=4,
notes=[
"PLG handles SMB; sales closes enterprise with PQL signals.",
"Requires clear PQL definition and SDR/PLG handoff process.",
"Best if you have a product with both bottom-up and top-down adoption.",
],
),
]
# ---------------------------------------------------------------------------
# Simulation engine
# ---------------------------------------------------------------------------
def simulate_model(model: GrowthModel, months: int) -> ModelProjection:
snapshots: List[MonthSnapshot] = []
mrr = STARTING_MRR
cumulative_cac = 0.0
cumulative_revenue = 0.0
break_even_month = None
for m in range(1, months + 1):
# Ramp up — new_mrr accelerates each month
if m <= model.months_to_steady_state:
# Ramp phase: linear ramp from 60% to 100% of base
ramp_factor = 0.6 + 0.4 * (m / model.months_to_steady_state)
else:
# Steady state: compound acceleration
months_past_ramp = m - model.months_to_steady_state
ramp_factor = 1.0 + model.monthly_acceleration * months_past_ramp
new_mrr = model.new_mrr_monthly_base * ramp_factor
churned_mrr = mrr * MONTHLY_CHURN_RATE
expansion_mrr = mrr * EXPANSION_RATE
net_new_mrr = new_mrr - churned_mrr + expansion_mrr
mrr = mrr + net_new_mrr
# CAC spend approximation: new_mrr / (avg_deal_mrr) * blended_cac
# Use weighted CAC from channel mix
weighted_cac = _weighted_cac(model.channel_mix)
avg_deal_mrr = 1_500 # Assumption: $1,500 average deal MRR
deals_this_month = new_mrr / avg_deal_mrr
cac_spend = deals_this_month * weighted_cac
cumulative_cac += cac_spend
cumulative_revenue += mrr * GROSS_MARGIN
if break_even_month is None and cumulative_revenue >= cumulative_cac:
break_even_month = m
snapshots.append(MonthSnapshot(
month=m,
mrr=mrr,
new_mrr=new_mrr,
churned_mrr=churned_mrr,
expansion_mrr=expansion_mrr,
net_new_mrr=net_new_mrr,
cumulative_cac_spend=cumulative_cac,
))
return ModelProjection(
model=model,
snapshots=snapshots,
break_even_month=break_even_month,
)
def _weighted_cac(channel_mix: Dict[str, float]) -> float:
channel_cac = {ch.name: ch.cac for ch in CHANNELS}
total = sum(
channel_mix.get(name, 0) * cac
for name, cac in channel_cac.items()
)
weight_sum = sum(channel_mix.values())
return total / weight_sum if weight_sum > 0 else 5_000
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt_mrr(n: float) -> str:
if n >= 1_000_000:
return f".3fM"
return f".1fK"
def fmt_currency(n: float) -> str:
if n >= 1_000_000:
return f".2fM"
if n >= 1_000:
return f".1fK"
return f".0f"
def print_header(title: str) -> None:
width = 78
print("\n" + "=" * width)
print(f" {title}")
print("=" * width)
def print_channel_overview() -> None:
print_header("Current Channel Mix")
print(f" Starting MRR: {fmt_mrr(STARTING_MRR)} | Monthly churn: {MONTHLY_CHURN_RATE:.1%} | Expansion: {EXPANSION_RATE:.1%}/mo")
print()
print(f" {'Channel':<22} {'% MRR':>7} {'CAC':>8} {'Payback':>9} {'Growth/mo':>10}")
print(" " + "-" * 60)
for ch in sorted(CHANNELS, key=lambda c: c.pct_of_new_mrr, reverse=True):
print(
f" {ch.name:<22} {ch.pct_of_new_mrr:>6.0%} "
f"{fmt_currency(ch.cac):>8} {ch.payback_months:>7.0f}mo "
f"{ch.monthly_growth_rate:>9.1%}"
)
def print_model_detail(proj: ModelProjection) -> None:
model = proj.model
print_header(f"Model: {model.name}")
print(f" {model.description}")
if model.notes:
print()
for note in model.notes:
print(f" • {note}")
print()
# Print monthly snapshot (every 3 months + final)
milestones = set(range(3, SIMULATION_MONTHS + 1, 3)) | {SIMULATION_MONTHS}
print(f" {'Month':<7} {'MRR':>10} {'New MRR':>9} {'Churned':>9} {'Expand':>8} {'Net New':>9}")
print(" " + "-" * 56)
for snap in proj.snapshots:
if snap.month in milestones:
print(
f" {snap.month:<7} {fmt_mrr(snap.mrr):>10} "
f"{fmt_mrr(snap.new_mrr):>9} {fmt_mrr(snap.churned_mrr):>9} "
f"{fmt_mrr(snap.expansion_mrr):>8} {fmt_mrr(snap.net_new_mrr):>9}"
)
final = proj.snapshots[-1]
growth_x = final.mrr / STARTING_MRR
arr_final = final.mrr * 12
weighted_cac = _weighted_cac(model.channel_mix)
be = f"Month {proj.break_even_month}" if proj.break_even_month else f"> {SIMULATION_MONTHS}mo"
print()
print(f" Final MRR ({SIMULATION_MONTHS}mo): {fmt_mrr(final.mrr)}")
print(f" Final ARR: {fmt_currency(arr_final)}")
print(f" Growth multiple: {growth_x:.1f}x from starting MRR")
print(f" Weighted blended CAC: {fmt_currency(weighted_cac)}")
print(f" Expected LTV:CAC: {model.avg_ltv_cac:.1f}x")
print(f" Months to steady state:{model.months_to_steady_state}")
print(f" CAC break-even: {be}")
def print_comparison_table(projections: List[ModelProjection]) -> None:
print_header(f"Growth Model Comparison — Month {SIMULATION_MONTHS} Outcomes")
header = (
f" {'Model':<20} {'MRR (final)':>12} {'ARR (final)':>12} "
f"{'Growth':>7} {'LTV:CAC':>8} {'Break-even':>11}"
)
print(header)
print(" " + "-" * 74)
for proj in sorted(projections, key=lambda p: p.snapshots[-1].mrr, reverse=True):
final = proj.snapshots[-1]
growth_x = final.mrr / STARTING_MRR
arr_final = final.mrr * 12
be = f"Mo {proj.break_even_month}" if proj.break_even_month else f">{SIMULATION_MONTHS}mo"
print(
f" {proj.model.name:<20} {fmt_mrr(final.mrr):>12} "
f"{fmt_currency(arr_final):>12} {growth_x:>6.1f}x "
f"{proj.model.avg_ltv_cac:>7.1f}x {be:>11}"
)
def print_channel_mix_impact(projections: List[ModelProjection]) -> None:
print_header("Channel Mix Impact Analysis")
print(" How shifting channel mix changes growth trajectory:\n")
baseline = next((p for p in projections if p.model.name == "Current Mix"), None)
if not baseline:
return
baseline_final_mrr = baseline.snapshots[-1].mrr
for proj in projections:
if proj.model.name == "Current Mix":
continue
final_mrr = proj.snapshots[-1].mrr
delta = final_mrr - baseline_final_mrr
delta_pct = (delta / baseline_final_mrr) * 100
arrow = "↑" if delta > 0 else "↓"
m6_mrr = proj.snapshots[5].mrr if len(proj.snapshots) >= 6 else 0
m6_baseline = baseline.snapshots[5].mrr if len(baseline.snapshots) >= 6 else 0
m6_delta = m6_mrr - m6_baseline
m6_pct = (m6_delta / m6_baseline) * 100 if m6_baseline else 0
m6_arrow = "↑" if m6_delta > 0 else "↓"
print(f" {proj.model.name}:")
print(f" Month 6: {m6_arrow} {abs(m6_pct):.1f}% vs. current ({fmt_mrr(m6_delta)} {'more' if m6_delta > 0 else 'less'} MRR)")
print(f" Month {SIMULATION_MONTHS}: {arrow} {abs(delta_pct):.1f}% vs. current ({fmt_mrr(delta)} {'more' if delta > 0 else 'less'} MRR)")
if proj.model.months_to_steady_state > 4:
print(f" ⚠ Model takes {proj.model.months_to_steady_state} months to reach steady state — short-term dip expected.")
print()
def print_decision_guide(projections: List[ModelProjection]) -> None:
print_header("Decision Guide")
print(" Choose your growth model based on your constraints:\n")
guides = [
("ACV < $5K and fast time-to-value", "PLG-First"),
("ACV > $25K and complex buying process", "Sales-Led Scale"),
("Strong practitioner community exists", "Community-Led"),
("Both SMB self-serve and enterprise buyers", "Hybrid PLS"),
("Uncertain — keep optionality", "Current Mix"),
]
for condition, model_name in guides:
proj = next((p for p in projections if p.model.name == model_name), None)
if proj:
final_mrr = proj.snapshots[-1].mrr
print(f" If: {condition}")
print(f" → Use {model_name} → {fmt_mrr(final_mrr)} MRR at month {SIMULATION_MONTHS}")
print()
print(" Key question before switching models:")
print(" 'Do we have 12-18 months of runway to prove the new model")
print(" while the current model continues in parallel?'")
print(" If no → optimize current model. Don't switch.")
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main() -> None:
print_channel_overview()
projections = [simulate_model(model, SIMULATION_MONTHS) for model in GROWTH_MODELS]
for proj in projections:
print_model_detail(proj)
print_comparison_table(projections)
print_channel_mix_impact(projections)
print_decision_guide(projections)
print("\n" + "=" * 78)
print(" Notes:")
print(f" Starting MRR: {fmt_mrr(STARTING_MRR)}")
print(f" Simulation: {SIMULATION_MONTHS} months")
print(f" Churn: {MONTHLY_CHURN_RATE:.1%}/mo ({MONTHLY_CHURN_RATE*12:.0%} annualized)")
print(f" Expansion: {EXPANSION_RATE:.1%}/mo of existing MRR")
print(f" Gross margin: {GROSS_MARGIN:.0%}")
print(" Acceleration rates are estimates — validate against your actuals.")
print("=" * 78 + "\n")
if __name__ == "__main__":
main()
FILE:scripts/marketing_budget_modeler.py
#!/usr/bin/env python3
"""
Marketing Budget Modeler
------------------------
Allocates marketing budget across channels based on CAC efficiency and
target MQL volume. Models conservative / moderate / aggressive scenarios.
Usage:
python marketing_budget_modeler.py
Inputs (edit INPUTS section below or extend with argparse):
- Annual revenue target (new ARR)
- Average selling price (ASP)
- Conversion rates by funnel stage
- Historical CAC per channel
- Channel capacity constraints (max MQLs the channel can realistically produce)
Outputs:
- Required MQL volume by channel
- Budget allocation per channel per scenario
- LTV:CAC and payback period per channel
- Summary table across scenarios
"""
from __future__ import annotations
import math
from dataclasses import dataclass, field
from typing import Dict, List, Tuple
# ---------------------------------------------------------------------------
# Data models
# ---------------------------------------------------------------------------
@dataclass
class Channel:
name: str
cac: float # Customer acquisition cost ($)
max_mqls_per_month: int # Realistic capacity ceiling (MQLs/month)
mql_to_close_rate: float # Combined MQL → closed-won rate (0.0–1.0)
payback_months: float # Based on ARPU × gross margin
ltv: float # Lifetime value ($)
trend: str = "stable" # "improving" | "stable" | "declining"
@dataclass
class FunnelRates:
mql_to_sal: float # MQL → Sales Accepted Lead
sal_to_sql: float # SAL → Sales Qualified Lead
sql_to_opp: float # SQL → Opportunity
opp_to_close: float # Opportunity → Closed-Won
@property
def mql_to_close(self) -> float:
return self.mql_to_sal * self.sal_to_sql * self.sql_to_opp * self.opp_to_close
@dataclass
class ScenarioResult:
name: str
total_budget: float
channel_budgets: Dict[str, float]
channel_mqls: Dict[str, int]
projected_customers: int
projected_arr: float
blended_cac: float
notes: List[str] = field(default_factory=list)
# ---------------------------------------------------------------------------
# INPUTS — edit these
# ---------------------------------------------------------------------------
TARGET_NEW_ARR = 3_000_000 # New ARR to generate this year ($)
ASP_ANNUAL = 18_000 # Average annual contract value ($)
GROSS_MARGIN = 0.75 # Product gross margin (%)
ARPU_MONTHLY = ASP_ANNUAL / 12 # Monthly revenue per account
FUNNEL = FunnelRates(
mql_to_sal=0.65,
sal_to_sql=0.45,
sql_to_opp=0.75,
opp_to_close=0.27,
)
# LTV = ARPU_monthly × gross_margin / monthly_churn_rate
MONTHLY_CHURN = 0.012 # ~14% annual churn
LTV = (ARPU_MONTHLY * GROSS_MARGIN) / MONTHLY_CHURN
CHANNELS: List[Channel] = [
Channel(
name="Organic SEO",
cac=1_800,
max_mqls_per_month=80,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(1_800 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="improving",
),
Channel(
name="Paid Search",
cac=6_200,
max_mqls_per_month=60,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(6_200 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="stable",
),
Channel(
name="Paid Social (LinkedIn)",
cac=8_500,
max_mqls_per_month=35,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(8_500 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="declining",
),
Channel(
name="Outbound SDR",
cac=5_100,
max_mqls_per_month=50,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(5_100 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="stable",
),
Channel(
name="Events / Field",
cac=9_800,
max_mqls_per_month=25,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(9_800 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="stable",
),
Channel(
name="Partner / Channel",
cac=3_400,
max_mqls_per_month=30,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(3_400 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="improving",
),
Channel(
name="Content / Inbound",
cac=2_600,
max_mqls_per_month=45,
mql_to_close_rate=FUNNEL.mql_to_close,
payback_months=(2_600 / (ARPU_MONTHLY * GROSS_MARGIN)),
ltv=LTV,
trend="improving",
),
]
# ---------------------------------------------------------------------------
# Core calculations
# ---------------------------------------------------------------------------
def customers_needed(target_arr: float, asp: float) -> int:
return math.ceil(target_arr / asp)
def mqls_needed_total(customers: int, mql_to_close: float) -> int:
return math.ceil(customers / mql_to_close)
def ltv_to_cac(ltv: float, cac: float) -> float:
return ltv / cac if cac > 0 else 0.0
def score_channel(ch: Channel) -> float:
"""
Score a channel for budget priority.
Higher = more efficient. Used to rank allocation order.
Factors: LTV:CAC ratio, trend multiplier, capacity.
"""
ratio = ltv_to_cac(ch.ltv, ch.cac)
trend_mult = {"improving": 1.2, "stable": 1.0, "declining": 0.7}.get(ch.trend, 1.0)
return ratio * trend_mult
def allocate_mqls(
channels: List[Channel],
total_mqls_needed: int,
budget_multiplier: float = 1.0,
) -> Tuple[Dict[str, int], Dict[str, float]]:
"""
Allocate MQL targets across channels in priority order (best LTV:CAC first).
budget_multiplier: 0.7 = conservative, 1.0 = moderate, 1.3 = aggressive.
Returns (channel → MQLs, channel → budget).
"""
ranked = sorted(channels, key=score_channel, reverse=True)
remaining = total_mqls_needed
channel_mqls: Dict[str, int] = {}
channel_budget: Dict[str, float] = {}
for ch in ranked:
if remaining <= 0:
channel_mqls[ch.name] = 0
channel_budget[ch.name] = 0.0
continue
# Apply capacity ceiling scaled by multiplier (aggressive = push capacity)
capacity = int(ch.max_mqls_per_month * 12 * budget_multiplier)
allocated = min(remaining, capacity)
channel_mqls[ch.name] = allocated
channel_budget[ch.name] = allocated * ch.cac
remaining -= allocated
return channel_mqls, channel_budget
def build_scenario(
name: str,
channels: List[Channel],
total_mqls: int,
multiplier: float,
notes: List[str],
) -> ScenarioResult:
channel_mqls, channel_budget = allocate_mqls(channels, total_mqls, multiplier)
total_budget = sum(channel_budget.values())
total_mqls_allocated = sum(channel_mqls.values())
projected_customers = math.floor(total_mqls_allocated * FUNNEL.mql_to_close)
projected_arr = projected_customers * ASP_ANNUAL
# Blended CAC = total budget / customers acquired
blended_cac = total_budget / projected_customers if projected_customers > 0 else 0.0
return ScenarioResult(
name=name,
total_budget=total_budget,
channel_budgets=channel_budget,
channel_mqls=channel_mqls,
projected_customers=projected_customers,
projected_arr=projected_arr,
blended_cac=blended_cac,
notes=notes,
)
# ---------------------------------------------------------------------------
# Reporting
# ---------------------------------------------------------------------------
def fmt_currency(n: float) -> str:
if n >= 1_000_000:
return f".2fM"
if n >= 1_000:
return f".1fK"
return f".0f"
def fmt_ratio(n: float) -> str:
return f"{n:.1f}x"
def print_header(title: str) -> None:
width = 72
print("\n" + "=" * width)
print(f" {title}")
print("=" * width)
def print_channel_table(channels: List[Channel]) -> None:
print_header("Channel Analysis — Current State")
header = f"{'Channel':<25} {'CAC':>8} {'Payback':>9} {'LTV:CAC':>8} {'Cap/mo':>7} {'Trend':>10}"
print(header)
print("-" * 72)
for ch in sorted(channels, key=score_channel, reverse=True):
ratio = ltv_to_cac(ch.ltv, ch.cac)
flag = ""
if ratio < 1:
flag = " ⚠ LOSS"
elif ratio >= 6:
flag = " ★ STRONG"
elif ratio >= 3:
flag = " ✓"
print(
f"{ch.name:<25} {fmt_currency(ch.cac):>8} "
f"{ch.payback_months:>7.1f}mo {fmt_ratio(ratio):>8} "
f"{ch.max_mqls_per_month:>7} {ch.trend:>10}{flag}"
)
def print_funnel_summary(customers: int, mqls: int) -> None:
print_header("Funnel Requirements")
print(f" Target new ARR: {fmt_currency(TARGET_NEW_ARR)}")
print(f" Average selling price: {fmt_currency(ASP_ANNUAL)}")
print(f" New customers needed: {customers}")
print(f" Funnel MQL→Close rate: {FUNNEL.mql_to_close:.1%}")
print(f" Total MQLs needed: {mqls}")
print(f"\n Funnel stage rates:")
print(f" MQL → SAL: {FUNNEL.mql_to_sal:.0%}")
print(f" SAL → SQL: {FUNNEL.mql_to_sal * FUNNEL.sal_to_sql:.0%}")
print(f" SQL → Opportunity: {FUNNEL.mql_to_sal * FUNNEL.sal_to_sql * FUNNEL.sql_to_opp:.0%}")
print(f" Opportunity → Close: {FUNNEL.mql_to_close:.0%}")
print(f"\n LTV (estimated): {fmt_currency(LTV)}")
print(f" Monthly churn: {MONTHLY_CHURN:.1%} ({MONTHLY_CHURN*12:.0%} annualized)")
def print_scenario(result: ScenarioResult, channels: List[Channel]) -> None:
print_header(f"Scenario: {result.name}")
print(f" Total marketing budget: {fmt_currency(result.total_budget)}")
print(f" Projected customers: {result.projected_customers}")
print(f" Projected new ARR: {fmt_currency(result.projected_arr)}")
print(f" Blended CAC: {fmt_currency(result.blended_cac)}")
blended_ltv_cac = LTV / result.blended_cac if result.blended_cac > 0 else 0
blended_payback = result.blended_cac / (ARPU_MONTHLY * GROSS_MARGIN)
print(f" Blended LTV:CAC: {fmt_ratio(blended_ltv_cac)}", end="")
if blended_ltv_cac < 1:
print(" ⚠ BELOW BREAK-EVEN")
elif blended_ltv_cac < 3:
print(" △ MARGINAL")
elif blended_ltv_cac >= 3:
print(" ✓ HEALTHY")
else:
print()
print(f" Blended payback: {blended_payback:.1f} months")
if result.notes:
print(f"\n Notes:")
for note in result.notes:
print(f" • {note}")
print(f"\n {'Channel':<25} {'MQLs':>6} {'Budget':>10} {'% of Budget':>12} {'LTV:CAC':>8}")
print(" " + "-" * 65)
for ch in sorted(channels, key=score_channel, reverse=True):
mqls = result.channel_mqls.get(ch.name, 0)
budget = result.channel_budgets.get(ch.name, 0.0)
pct = (budget / result.total_budget * 100) if result.total_budget > 0 else 0
ratio = ltv_to_cac(ch.ltv, ch.cac)
print(
f" {ch.name:<25} {mqls:>6} {fmt_currency(budget):>10} "
f"{pct:>11.1f}% {fmt_ratio(ratio):>8}"
)
def print_scenario_comparison(scenarios: List[ScenarioResult]) -> None:
print_header("Scenario Comparison")
header = f"{'Scenario':<18} {'Budget':>10} {'Customers':>10} {'ARR':>10} {'Blended CAC':>12} {'LTV:CAC':>8} {'Payback':>9}"
print(header)
print("-" * 82)
for s in scenarios:
blended_ltv_cac = LTV / s.blended_cac if s.blended_cac > 0 else 0
blended_payback = s.blended_cac / (ARPU_MONTHLY * GROSS_MARGIN)
print(
f"{s.name:<18} {fmt_currency(s.total_budget):>10} "
f"{s.projected_customers:>10} {fmt_currency(s.projected_arr):>10} "
f"{fmt_currency(s.blended_cac):>12} {fmt_ratio(blended_ltv_cac):>8} "
f"{blended_payback:>7.1f}mo"
)
def print_recommendations(channels: List[Channel]) -> None:
print_header("Channel Recommendations")
scale = [ch for ch in channels if score_channel(ch) >= 1.5 and ch.trend in ("improving", "stable")]
hold = [ch for ch in channels if 0.8 <= score_channel(ch) < 1.5 or (ch.trend == "stable" and ltv_to_cac(ch.ltv, ch.cac) >= 3)]
cut = [ch for ch in channels if ltv_to_cac(ch.ltv, ch.cac) < 2 or ch.trend == "declining"]
# Deduplicate
hold = [ch for ch in hold if ch not in scale]
cut = [ch for ch in cut if ch not in scale and ch not in hold]
if scale:
print(" SCALE (strong LTV:CAC, improving or stable trend):")
for ch in scale:
print(f" + {ch.name} [LTV:CAC {fmt_ratio(ltv_to_cac(ch.ltv, ch.cac))}, payback {ch.payback_months:.0f}mo]")
if hold:
print(" HOLD (monitor — adequate but not outstanding):")
for ch in hold:
print(f" = {ch.name} [LTV:CAC {fmt_ratio(ltv_to_cac(ch.ltv, ch.cac))}, trend: {ch.trend}]")
if cut:
print(" CUT or REDUCE (poor LTV:CAC or declining):")
for ch in cut:
print(f" - {ch.name} [LTV:CAC {fmt_ratio(ltv_to_cac(ch.ltv, ch.cac))}, trend: {ch.trend}]")
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
def main() -> None:
customers = customers_needed(TARGET_NEW_ARR, ASP_ANNUAL)
total_mqls = mqls_needed_total(customers, FUNNEL.mql_to_close)
print_channel_table(CHANNELS)
print_funnel_summary(customers, total_mqls)
scenarios = [
build_scenario(
name="Conservative",
channels=CHANNELS,
total_mqls=total_mqls,
multiplier=0.7,
notes=[
"Prioritizes lowest CAC channels only.",
"May not reach MQL target — expect ~70% of goal.",
"Best for capital-constrained orgs or short runway.",
],
),
build_scenario(
name="Moderate",
channels=CHANNELS,
total_mqls=total_mqls,
multiplier=1.0,
notes=[
"Balanced allocation — efficiency-first but full MQL target.",
"Recommended baseline. Revisit Q2 based on actuals.",
],
),
build_scenario(
name="Aggressive",
channels=CHANNELS,
total_mqls=total_mqls,
multiplier=1.4,
notes=[
"Pushes all channels toward capacity ceiling.",
"Higher spend on lower-efficiency channels to hit volume.",
"Requires > 18-month runway to justify payback period.",
],
),
]
for scenario in scenarios:
print_scenario(scenario, CHANNELS)
print_scenario_comparison(scenarios)
print_recommendations(CHANNELS)
print("\n" + "=" * 72)
print(" Key questions before finalizing budget:")
print(" 1. What is the payback period the CFO/board will accept?")
print(" 2. Is CAC for declining-trend channels actually recoverable?")
print(" 3. Does the moderate scenario require sales headcount increase?")
print(" 4. Which channels have capacity to absorb 20% more spend?")
print("=" * 72 + "\n")
if __name__ == "__main__":
main()
Khung vận hành công ty: chọn hệ điều hành (EOS, OKR...), sơ đồ trách nhiệm, scorecard, nhịp họp và mục tiêu 90 ngày.
---
name: "company-os"
description: "The meta-framework for how a company runs — the connective tissue between all C-suite roles. Covers operating system selection (EOS, Scaling Up, OKR-native, hybrid), accountability charts, scorecards, meeting pulse, issue resolution, and 90-day rocks. Use when setting up company operations, selecting a management framework, designing meeting rhythms, building accountability systems, implementing OKRs, or when user mentions EOS, Scaling Up, operating system, L10 meetings, rocks, scorecard, accountability chart, or quarterly planning."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: c-level
domain: company-operations
updated: 2026-03-05
frameworks: os-comparison, implementation-guide
---
# Company Operating System
The operating system is the collection of tools, rhythms, and agreements that determine how the company functions. Every company has one — most just don't know what it is. Making it explicit makes it improvable.
## Keywords
operating system, EOS, Entrepreneurial Operating System, Scaling Up, Rockefeller Habits, OKR, Holacracy, L10 meeting, rocks, scorecard, accountability chart, issues list, IDS, meeting pulse, quarterly planning, weekly scorecard, management framework, company rhythm, traction, Gino Wickman, Verne Harnish
## Why This Matters
Most operational dysfunction isn't a people problem — it's a system problem. When:
- The same issues recur every week: no issue resolution system
- Meetings feel pointless: no structured meeting pulse
- Nobody knows who owns what: no accountability chart
- Quarterly goals slip: rocks aren't real commitments
Fix the system. The people will operate better inside it.
## The Six Core Components
Every effective operating system has these six, regardless of which framework you choose:
### 1. Accountability Chart
Not an org chart. An accountability chart answers: "Who owns this outcome?"
**Key distinction:** One person owns each function. Multiple people may work in it. Ownership means the buck stops with one person.
**Structure:**
```
CEO
├── Sales (CRO/VP Sales)
│ ├── Inbound pipeline
│ └── Outbound pipeline
├── Product & Engineering (CTO/CPO)
│ ├── Product roadmap
│ └── Engineering delivery
├── Operations (COO)
│ ├── Customer success
│ └── Finance & Legal
└── People (CHRO/VP People)
├── Recruiting
└── People operations
```
**Rules:**
- No shared ownership. "Alice and Bob both own it" means nobody owns it.
- One person can own multiple seats at early stages. That's fine. Just be explicit.
- Revisit quarterly as you scale. Ownership shifts as the company grows.
**Build it in a workshop:**
1. List all functions the company performs
2. Assign one owner per function — no exceptions
3. Identify gaps (functions nobody owns) and overlaps (functions two people think they own)
4. Publish it. Update it when something changes.
### 2. Scorecard
Weekly metrics that tell you if the company is on track. Not monthly. Not quarterly. Weekly.
**Rules:**
- 5–15 metrics maximum. More than 15 and nothing gets attention.
- Each metric has an owner and a weekly target (not a range — a number).
- Red/yellow/green status. Not paragraphs.
- The scorecard is discussed at the leadership team weekly meeting. Only red metrics get discussion time.
**Example scorecard structure:**
| Metric | Owner | Target | This Week | Status |
|--------|-------|--------|-----------|--------|
| New MRR | CRO | €50K | €43K | 🔴 |
| Churn | CS Lead | < 1% | 0.8% | 🟢 |
| Active users | CPO | 2,000 | 2,150 | 🟢 |
| Deployments | CTO | 3/week | 3 | 🟢 |
| Open critical bugs | CTO | 0 | 2 | 🔴 |
| Runway | CFO | > 18mo | 16mo | 🟡 |
**Anti-pattern:** Measuring everything. If you track 40 KPIs, you're watching, not managing.
### 3. Meeting Pulse
The meeting rhythm that drives the company. Not optional — the pulse is what keeps the company alive.
**The full rhythm:**
| Meeting | Frequency | Duration | Who | Purpose |
|---------|-----------|----------|-----|---------|
| Daily standup | Daily | 15 min | Each team | Blockers only |
| L10 / Leadership sync | Weekly | 90 min | Leadership team | Scorecard + issues |
| Department review | Monthly | 60 min | Dept + leadership | OKR progress |
| Quarterly planning | Quarterly | 1–2 days | Leadership | Set rocks, review strategy |
| Annual planning | Annual | 2–3 days | Leadership | 1-year + 3-year vision |
**The L10 meeting (Weekly Leadership Sync):**
Named for the goal of each meeting being a 10/10. Fixed agenda:
1. Good news (5 min) — personal + business
2. Scorecard review (5 min) — flag red items only
3. Rock review (5 min) — on/off track for each rock
4. Customer/employee headlines (5 min)
5. Issues list (60 min) — IDS (see below)
6. To-dos review (5 min) — last week's commitments
7. Conclude (5 min) — rate the meeting 1–10, what would make it a 10 next time
### 4. Issue Resolution (IDS)
The core problem-solving loop. Maximum 15 minutes per issue.
**IDS: Identify, Discuss, Solve**
- **Identify:** What is the actual issue? (Not the symptom — the root cause) State it in one sentence.
- **Discuss:** Relevant facts + perspectives. Time-boxed. When discussion starts repeating, stop.
- **Solve:** One owner. One action. One due date. Written on the to-do list.
**Anti-patterns:**
- "Let's take this offline" — most things taken offline never get resolved
- Discussing without deciding — a great discussion with no action item is wasted time
- Revisiting decided issues — once solved, it leaves the list. Reopen only with new information.
**The Issues List:** A running, prioritized list of all unresolved issues. Owned by the leadership team. Reviewed and pruned weekly. If an issue has been on the list for 3+ meetings and hasn't been discussed, it's either not a real issue or it's too scary to address — both deserve attention.
### 5. Rocks (90-Day Priorities)
Rocks are the 3–7 most important things each person must accomplish in the next 90 days. They're not the job description — they're the things that move the company forward.
**Why 90 days?** Long enough for meaningful progress. Short enough to stay real.
**Rock rules:**
- Each person: 3–7 rocks maximum. More than 7 and none get done.
- Company-level rocks (shared priorities): 3–7 for the leadership team
- Each rock is binary: done or not done. No "60% complete."
- Set at the quarterly planning session. Reviewed weekly (on/off track).
**Bad rock:** "Improve our sales process"
**Good rock:** "Implement Salesforce CRM with full pipeline stages and weekly reporting by March 31"
**Rock vs. to-do:** A to-do takes one action. A rock takes 90 days of consistent work.
### 6. Communication Cadence
Who gets what information, when, and how.
| Audience | What | When | Format |
|----------|------|------|--------|
| All employees | Company update | Monthly | Written + Q&A |
| All employees | Quarterly results + next priorities | Quarterly | All-hands |
| Leadership team | Scorecard | Weekly | Dashboard |
| Board | Company performance | Monthly | Board memo |
| Investors | Key metrics + narrative | Monthly or quarterly | Investor update |
| Customers | Product updates | Per release | Release notes |
**Default rule:** If you're deciding whether to share something internally, share it. The cost of under-communication always exceeds the cost of over-communication inside a company.
---
## Operating System Selection
See `references/os-comparison.md` for full comparison. Quick guide:
| If you are... | Consider... |
|---------------|-------------|
| 10–250 person company, founder-led, operational chaos | EOS / Traction |
| Ambitious growth company, need rigorous strategy cascade | Scaling Up |
| Tech company, engineering culture, hypothesis-driven | OKR-native |
| Decentralized, flat, high autonomy | Holacracy (only if you're patient) |
| None of the above quite fit | Custom hybrid |
---
## Implementation Roadmap
Don't implement everything at once. See `references/implementation-guide.md` for the full 90-day plan.
**Quick start (first 30 days):**
1. Build the accountability chart (1 workshop, 2 hours)
2. Define 5–10 weekly scorecard metrics (leadership team alignment, 1 hour)
3. Start the weekly L10 meeting (no prep — just start)
These three alone will improve coordination more than most companies achieve in a year.
---
## Common Failure Modes
**Partial implementation:** "We do OKRs but skip the weekly check-in." Half an operating system is worse than none — it creates theater without accountability.
**Meeting fatigue:** Adding the full rhythm on top of existing meetings. Start by replacing meetings, not adding them.
**Metric overload:** Starting with 30 KPIs because "they all matter." Start with 5. Add when the cadence is established.
**Rock inflation:** Setting 12 rocks per person because "everything is a priority." When everything is a priority, nothing is. Hard limit: 7.
**Leader non-compliance:** Leadership team skips the L10 or doesn't follow IDS. The operating system mirrors the respect leadership gives it. If leaders don't take it seriously, nobody will.
**Annual planning without quarterly review:** Setting annual goals and checking in at year-end. Quarterly is the minimum review cycle for any meaningful goal.
---
## Integration with C-Suite
The company OS is the connective tissue. Every other role depends on it:
| C-Suite Role | OS Dependency |
|-------------|---------------|
| CEO | Sets vision that feeds into 1-year plan and rocks |
| COO | Owns the meeting pulse and issue resolution cadence |
| CFO | Owns the financial metrics in the scorecard |
| CTO | Owns engineering rocks and tech scorecard metrics |
| CHRO | Owns people metrics (attrition, hiring velocity) in scorecard |
| Culture Architect | Culture rituals plug into the meeting pulse |
| Strategic Alignment Engine | Validates that team rocks cascade from company rocks |
---
## Key Questions for the Operating System
- "If I asked five different team leads what the company's top 3 priorities are this quarter, would they give the same answers?"
- "What was the most important issue raised in last week's leadership meeting? Was it resolved or is it still open?"
- "Name a metric that would tell us by Friday whether this week was a good week. Do we track it?"
- "Who owns customer churn? Can you name that person without hesitation?"
- "When was the last time we updated the accountability chart?"
## Detailed References
- `references/os-comparison.md` — EOS vs Scaling Up vs OKRs vs Holacracy vs hybrid
- `references/implementation-guide.md` — 90-day implementation plan
FILE:references/implementation-guide.md
# Company Operating System — 90-Day Implementation Guide
Don't implement everything at once. The fastest path to failure is trying to launch the full operating system in week one. Build incrementally. Let the team experience wins before adding complexity.
---
## Before You Start
### Prerequisites
**Leadership alignment (non-negotiable):**
Every member of the leadership team must understand why you're doing this and commit to running the system. One holdout destroys the whole model. If the CFO skips the L10 meetings, the system won't work.
**Current state audit:**
- What meetings currently exist? Which can be replaced?
- Who owns which functions today? (Even informally)
- What metrics are being tracked? (Even inconsistently)
**Assign an OS owner:**
One person is responsible for the implementation and ongoing maintenance of the operating system. Usually the COO or CEO (at smaller companies). This is not a committee job.
---
## Week 1–2: Accountability Chart + Scorecard
### Accountability Chart Workshop (Week 1)
**Duration:** 2–3 hours, full leadership team
**Step 1 — List all functions (30 min)**
On a whiteboard, list every function the company performs:
- Sales (inbound, outbound, partnerships)
- Marketing (content, paid, brand)
- Product (roadmap, design, research)
- Engineering (frontend, backend, devops)
- Customer success (onboarding, support, retention)
- Finance (accounting, FP&A, legal)
- People (recruiting, HR, culture)
- Operations (processes, tools, facilities)
**Step 2 — Assign owners (45 min)**
For each function: "Who is the one person ultimately accountable?" Write their name.
Rules: One name only. No joint ownership. One person can own multiple functions at small scale.
**Step 3 — Identify gaps and overlaps (30 min)**
- **Gaps:** Functions with no owner → Who should own them? Or do we need a hire?
- **Overlaps:** Two people said they own the same thing → Resolve now, not later.
**Step 4 — Publish and socialize (Week 2)**
Share with the full company. Explain what an accountability chart is and isn't.
"This is about clarity, not hierarchy. It tells everyone who to go to for each function."
**Output:** A documented accountability chart. Use a simple tool (Miro, Google Slides, Ninety.io).
---
### Scorecard Design (Week 2)
**Duration:** 90 minutes, leadership team
**Step 1 — List candidate metrics (30 min)**
Each leader lists 3–5 metrics they already track or wish they tracked. No filtering yet.
**Step 2 — Filter to 5–15 (30 min)**
Criteria: Is it measurable weekly? Does it tell us if the company is healthy? Does one person own it?
Drop: metrics that are monthly only, metrics without a clear owner, metrics that measure activity not outcomes.
**Step 3 — Set weekly targets (20 min)**
For each metric: what's the weekly target? Not a range — a number. Red/yellow/green thresholds.
**Step 4 — Assign owners (10 min)**
Every metric has one owner who is responsible for reporting it weekly.
**Output:** A scorecard document. 5–15 metrics, owner, target, weekly tracking column.
**First scorecard run:** Week 2 or 3. It won't be perfect. That's fine.
---
## Week 3–4: Meeting Pulse (Start With L10)
Don't start all the meetings at once. Start with the weekly L10. Replace existing leadership syncs.
### L10 Meeting Setup
**Schedule:** Same day, same time, every week. Non-negotiable attendance.
**Duration:** 90 minutes. No more, no less.
**Facilitator:** Rotate or assign to COO/CEO. The facilitator keeps time and follows the agenda.
**Fixed agenda:**
1. **Good news** (5 min) — One personal, one business from each person. No skipping.
2. **Scorecard review** (5 min) — Traffic light only. Red items go to the issues list.
3. **Rock review** (5 min) — Each person: "on track" or "off track." No justification needed at this step.
4. **Customer/employee headlines** (5 min) — One sentence each. No reports.
5. **Issues** (60 min) — IDS process. Prioritize the top 3–5 issues. Solve them.
6. **To-do review** (5 min) — Review last week's commitments (done/not done). No excuses, just data.
7. **Conclude** (5 min) — Rate the meeting 1–10. What would make next week better?
**First L10 meeting:**
It will feel awkward. Run through the agenda anyway. The team needs the repetition to internalize it. By week 4, it should feel natural.
### Issues List Setup
Create a shared document (Notion, Google Docs, or dedicated tool):
- Issue title
- Priority (High / Medium / Low)
- Status (Open / In progress / Solved)
- Owner (once assigned)
- Due date
At the first L10, generate the issues list by asking: "What's getting in our way right now?" Expect 10–20 items on the first pass.
---
## Week 5–8: Rocks and Quarterly Planning
### Quarterly Planning Session (end of Week 5 or start of Week 6)
**Duration:** 4–8 hours (or 2 × 4-hour days for larger teams)
**Who:** Full leadership team
**Session structure:**
**Part 1: Review previous quarter (60–90 min)**
- What rocks were completed? What were dropped?
- What did we learn?
- What changed in the market or company?
**Part 2: Confirm or update company direction (60 min)**
- Is the 1-year goal still valid?
- Any major strategy shifts needed?
- Update the V/TO or OPSP if using EOS or Scaling Up.
**Part 3: Set company rocks (90 min)**
- Brainstorm: What are the 3–7 most important things to accomplish this quarter?
- Prioritize. Be ruthless. 3 rocks done > 7 rocks started.
- Each rock: clear owner, clear definition of done, 90-day timeline.
**Part 4: Set individual rocks (60 min)**
- Each leader sets their 3–7 rocks (aligned with company rocks where possible)
- Share with group: dependencies? Conflicts? Overloaded people?
**Part 5: Communicate (Week 6)**
- Share company rocks with the full organization within 1 week
- Each team sets their own rocks, cascaded from company rocks (3–5 per team)
**Rock template:**
```
Rock: [What you'll accomplish]
Owner: [One person]
Due date: [Specific date within the quarter]
Definition of done: [How we'll know it's complete]
Dependencies: [What else needs to happen first]
```
---
## Week 9–12: Issue Resolution Mastery + Communication Cadence
By now the L10 should be running smoothly. Weeks 9–12 focus on deepening IDS skills and establishing the broader communication cadence.
### IDS Practice
The issue resolution process often degrades in weeks 5–8. Common problems:
- Issues discussed but never solved (no clear action item)
- Same issues recurring (root cause not addressed)
- Too many issues, not enough resolution (prioritization failing)
**IDS calibration exercise (Week 9):**
In the next L10, after each issue is "solved," ask:
- "Is this actually solved, or are we postponing it?"
- "What's the specific action? Who owns it? When is it due?"
- "Is this the real issue, or a symptom of something deeper?"
### Communication Cadence Setup
Build out the full communication calendar:
| Communication | Frequency | Owner | Format | Tool |
|---------------|-----------|-------|--------|------|
| Company all-hands | Monthly | CEO | Update + Q&A | Video call |
| Quarterly planning results | Quarterly | CEO/COO | Written + live | Notion + all-hands |
| Board update | Monthly | CEO + CFO | Board memo | Doc |
| Investor update | Monthly | CEO + CFO | Email | Template |
| Department L10s | Weekly | Dept lead | L10 format | In-person / Zoom |
| Daily standups | Daily | Team leads | 15 min | Team call |
**Company all-hands template:**
1. State of the company (financial health, key metrics) — 10 min
2. Quarterly rocks: what we committed to, where we stand — 10 min
3. Wins and recognitions — 5 min
4. What's coming next quarter — 10 min
5. Q&A — 15–25 min
---
## Post-90 Days: Refinement and Optimization
### Month 4 retrospective
After the first full quarter, run a retrospective on the operating system itself:
- What's working? What isn't?
- Which meetings should continue as-is? Which need adjustment?
- Is the scorecard measuring the right things?
- Are rocks the right size and specificity?
- What should we add next?
### Scorecard evolution
By month 4, you'll know which metrics matter most. Add 2–3 that are missing. Remove metrics that nobody uses for decisions.
### L10 health check
Rate your L10 meetings over the first quarter:
- Average rating < 7: The agenda isn't being followed or issues aren't being resolved. Diagnose.
- Average rating 7–8: Normal. Keep building discipline.
- Average rating > 8: The team is engaged. Start extending the system to department level.
### Department L10s (Month 4+)
Once leadership L10 is running well, cascade the meeting structure:
- Each department runs their own weekly L10
- Department rocks cascade from company rocks
- Issues that cross departments are escalated to leadership L10
### Year 1 annual planning
End of year 1: run a full-day annual planning session.
- Review the year: what did we accomplish? What did we miss? What did we learn?
- Update 3-year vision (has it changed?)
- Set next year's annual goals
- Set Q1 rocks
- Celebrate. Seriously — mark the milestone.
---
## Implementation Anti-Patterns
**Skipping the accountability chart:** Without ownership clarity, every other system breaks down. Do this first.
**Building a perfect scorecard before starting:** Start with 5 imperfect metrics. Improve over time.
**Not replacing existing meetings:** Adding L10 on top of 3 existing meetings creates meeting overload. Cancel the redundant ones.
**Leader non-participation:** If one leader consistently skips or is disengaged, the system won't work. Address this directly — it's a culture issue, not a calendar issue.
**Changing the L10 agenda:** The agenda works because of repetition. Resist the urge to customize it for the first 6 months.
**Rocks without accountability:** If nobody checks rocks at the L10 ("on track / off track"), they become wish lists. The weekly review is what makes them real.
FILE:references/os-comparison.md
# Operating System Comparison
Side-by-side analysis of the major company operating frameworks.
---
## Overview
| Framework | Origin | Best fit | Implementation time | Cost |
|-----------|--------|----------|---------------------|------|
| EOS | Gino Wickman, 2007 | 10–250 employees, founder-led | 2–3 years full adoption | Free (DIY) to $25K+/year (implementer) |
| Scaling Up | Verne Harnish, 2002 | Growth-stage, strategic focus | 1–2 years | Free (DIY) to $15K+/year (coach) |
| OKR-native | Andy Grove / Google | Tech companies, product orgs | 3–6 months | Free |
| Holacracy | Brian Robertson, 2007 | Flat, autonomous organizations | 2–4 years | $5K–$50K+ (certification) |
| Custom hybrid | You | When the above don't fit exactly | Ongoing | Whatever you invest |
---
## 1. EOS — Entrepreneurial Operating System
**Book:** *Traction* by Gino Wickman
### Core principles
EOS is built on Six Components:
1. **Vision** — Where are you going? (V/TO: Vision/Traction Organizer)
2. **People** — Right people, right seats
3. **Data** — Scorecard with weekly metrics
4. **Issues** — Surface and resolve with IDS
5. **Process** — Document core processes
6. **Traction** — Rocks + meeting pulse (L10)
### Signature tools
- **V/TO (Vision/Traction Organizer):** 2-page strategy doc. Core values, core focus, 10-year target, 3-year picture, 1-year plan, quarterly rocks, issues.
- **Accountability Chart:** Who owns what function (not org chart)
- **L10 meeting:** Weekly 90-minute leadership sync (Level 10 = aim for 10/10)
- **Rocks:** 90-day priority commitments (3–7 per person)
- **IDS:** Identify, Discuss, Solve (issue resolution, max 15 min per issue)
### Strengths
- **Operationally focused.** If your problem is execution chaos, EOS addresses it directly.
- **Accessible.** The book is practical. You can DIY it without a coach.
- **Community.** Large network of implementers, tools (Ninety.io, EOS Worldwide), and practitioners.
- **Simple enough to actually use.** No complex methodology. Most teams are functional within 6 months.
### Limitations
- **Strategic depth is shallow.** The V/TO is good for direction but doesn't replace real strategy work.
- **Doesn't scale beyond ~250.** Designed for entrepreneurial companies. Gets cumbersome at enterprise scale.
- **Assumes a cohesive leadership team.** If trust is broken at the top, EOS won't fix it.
- **Facilitator dependency.** Many companies benefit from an EOS Implementer (external coach), which adds cost.
### Best fit
- 10–150 person companies
- Founder-led, operational dysfunction
- Teams that can't stay on the same page
- Companies with recurring issues that never get resolved
- First real "operating system" for a company that's been running on vibes
### Not ideal if
- You need sophisticated strategic planning
- You're > 250 people and already have ops infrastructure
- Your team resists structured methodology
---
## 2. Scaling Up (Rockefeller Habits 2.0)
**Book:** *Scaling Up* by Verne Harnish
### Core principles
Built on four Decisions:
1. **People** — Core values, talent management, Topgrading
2. **Strategy** — One-Page Strategic Plan (OPSP), 7 Strata of Strategy
3. **Execution** — Priorities (rocks), meeting rhythm, critical numbers
4. **Cash** — Power of One, Cash Acceleration Strategies (CAS)
### Signature tools
- **One-Page Strategic Plan (OPSP):** Annual and quarterly goals on one page. More strategic than EOS's V/TO.
- **7 Strata of Strategy:** Competitive positioning, core customer, brand promise, X-factor (10x advantage), profit per X, BHAG, critical numbers.
- **Meeting rhythm:** Daily (5–15 min), weekly, monthly, quarterly, annual — with specific templates.
- **Critical number:** One metric that, if improved, fixes everything else.
- **Cash acceleration:** CAS system for improving working capital and cash conversion cycle.
### Strengths
- **Stronger strategic framework than EOS.** The 7 strata and OPSP force real strategic thinking.
- **Cash focus.** Unique among frameworks — explicitly addresses cash flow management.
- **Scales further.** Better suited for 100–1000 person companies than EOS.
- **Works for ambitious growth companies.** Designed for companies that want to scale significantly.
### Limitations
- **More complex than EOS.** Harder to DIY. Benefits heavily from a certified Scaling Up coach.
- **Overwhelming at first.** The full framework has many components. Teams often implement partially.
- **Less prescriptive on meetings.** EOS's L10 is very specific. Scaling Up's meeting rhythm requires more customization.
### Best fit
- Series A to Series C companies
- Companies with strong growth ambition
- Leadership teams that want strategic rigor, not just operational clarity
- Companies already past initial chaos, ready for more sophisticated frameworks
### Not ideal if
- You're pre-product-market-fit
- You need quick operational wins
- Your team doesn't have the bandwidth for the learning curve
---
## 3. OKR-Native (Google Style)
**Books:** *Measure What Matters* by John Doerr; *Radical Focus* by Christina Wodtke
### Core principles
OKRs = Objectives + Key Results
- **Objectives:** Qualitative, inspiring direction. "What are we trying to achieve?"
- **Key Results:** Quantitative, measurable outcomes. "How will we know we achieved it?"
- **Not tasks.** KRs measure outcomes, not activities.
**Cascade:** Company OKRs → Department OKRs → Team OKRs → Individual OKRs
**Cadence:** Quarterly OKR cycles. Weekly check-ins. Annual reflection.
**Scoring:** 0.0–1.0. Target is 0.7. Consistently hitting 1.0 = OKRs aren't ambitious enough.
### Strengths
- **Aligns the whole company.** When done well, every team can trace their work to company-level objectives.
- **Encourages ambition.** Moonshot OKRs are explicit. "Roofshot" vs "moonshot" OKRs.
- **Widely understood in tech.** Many hires will already know OKRs.
- **No framework cost.** No implementer required. Tooling is free or cheap (Linear, Notion, Lattice).
### Limitations
- **Hard to do well.** Most companies run "OKR theater" — tasks dressed up as key results.
- **Missing the HOW.** OKRs define what to achieve but not how to operate. You still need meeting rhythm, accountability structure, and issue resolution.
- **Misalignment risk.** If not cascaded properly, teams run disconnected OKRs that feel like alignment but aren't.
- **No operational backbone.** OKRs are a goal-setting system, not a full operating system.
### Best fit
- Tech companies with strong product/engineering culture
- Companies where hypothesis-driven work is already the norm
- Organizations that value autonomy and bottom-up goal setting
- As the goal-setting layer inside a broader operating system
### Not ideal if
- Teams lack discipline to hold each other accountable
- You need more than just goal alignment (issue resolution, meeting structure)
- Leaders don't model OKR behavior themselves
---
## 4. Holacracy
**Book:** *Holacracy* by Brian Robertson
### Core principles
Holacracy replaces the traditional management hierarchy with a system of distributed authority.
- **Circles:** Semi-autonomous units with defined purposes (like teams, but self-governing)
- **Roles:** People fill roles (not job descriptions). One person can hold multiple roles in different circles.
- **Governance meetings:** Roles and accountabilities are defined and evolved by the circle, not management
- **Tactical meetings:** Operational coordination within circles
- **The Constitution:** A legal document that all members ratify, replacing traditional management authority
### Strengths
- **Maximum autonomy.** People closest to the work define how it gets done.
- **Removes management as a bottleneck.** Decisions happen at the circle level.
- **Adapts to complexity.** Circle structure evolves organically as the work changes.
### Limitations
- **Enormous learning curve.** 2–4 years to full adoption. Many companies abandon it.
- **High meeting overhead.** Governance meetings add significant time.
- **Doesn't eliminate politics.** Just moves them to governance meetings.
- **Requires full commitment.** Partial Holacracy doesn't work. You either do it or you don't.
- **Not for crisis mode.** When speed matters, distributed governance slows you down.
### When it works
- Organizations with deep belief in autonomy and self-management
- Non-profit or mission-driven organizations where consensus matters
- Companies with patient leadership willing to invest years in implementation
### When it doesn't work
- Startups needing speed and clarity
- Companies with strong founder personalities who struggle to relinquish control
- Organizations that need to move fast or course-correct frequently
---
## 5. Custom Hybrid
### When to build a hybrid
None of the above frameworks fits perfectly because:
- EOS lacks strategic depth
- Scaling Up is complex to implement
- OKRs don't provide operational backbone
- Holacracy is too slow to implement
The solution: take the best components of each.
### Common hybrid patterns
**EOS backbone + OKR goal-setting:**
- EOS provides: accountability chart, L10 meeting, IDS, meeting pulse
- OKRs provide: goal-setting with ambition, cascade, and alignment checks
- Works well for: tech companies that want operational rigor with flexibility
**Scaling Up strategy + EOS execution:**
- Scaling Up provides: OPSP, 7 strata, cash management
- EOS provides: L10, rocks, IDS
- Works well for: ambitious growth companies that want both strategy and execution discipline
**OKRs + custom meeting rhythm:**
- OKRs provide: goal cascade
- Custom meetings: weekly team syncs, monthly department reviews, quarterly all-hands
- Works well for: companies that already have strong culture but need goal alignment
### Hybrid design principles
1. **Pick one goal-setting system.** Don't mix OKRs and Rocks — they're both 90-day priority systems and will create confusion.
2. **Be explicit about what you're taking from where.** "We use EOS for meetings and Scaling Up for strategy" is a clear hybrid. "We do a bit of everything" is chaos.
3. **Document your version.** Your operating system should have a name and a one-page description of what it includes.
4. **Evolve intentionally.** Change one component at a time. Don't overhaul the whole system when one part isn't working.
---
## Framework Selection Decision Tree
```
Is your company < 50 people and in operational chaos?
YES → Start with EOS. It's the simplest path to order.
NO → Continue.
Does strategic positioning and cash flow need significant work?
YES → Consider Scaling Up.
NO → Continue.
Is your company tech-native with strong product/engineering culture?
YES → OKR-native with a custom meeting rhythm.
NO → Continue.
Do you have 2+ years and full leadership commitment to radical organizational change?
YES → Consider Holacracy (with caution).
NO → Build a custom hybrid from EOS + OKRs.
```
Cấu hình khung tuân thủ áp dụng, tính độ chồng lấn kiểm soát, mô phỏng kiểm toán nội bộ và hợp nhất bằng chứng.
---
name: "compliance-os"
description: "Compliance OS — meta-orchestrator that lets compliance teams CONFIGURE which frameworks apply, COMPUTE cross-framework control overlap, SIMULATE internal audits, and CONSOLIDATE evidence across multiple frameworks. Four decisions: (1) Given a company profile, which of the 12 supported frameworks apply (ISO 27001/13485/42001/14971, EU AI Act, MDR 745, GDPR, SOC 2, FDA QSR, NIST CSF 2.0, NIS2, HIPAA)? (2) Across selected frameworks, which controls overlap and how much evidence reuses? (3) For a given framework + scope, what does a realistic mock audit produce — drawing from the 205-scenario library? (4) Across selected frameworks, what's the unified evidence checklist with reuse map? Use when standing up a multi-framework program, planning the annual audit calendar, or preparing for certification stage 1. Does NOT replace per-framework skills (it orchestrates them)."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: compliance-os
domain: multi-framework-compliance-orchestration
updated: 2026-05-13
python-tools: framework_selector.py, cross_framework_mapper.py, audit_simulator.py, evidence_pool_generator.py
frameworks: iso-27001, iso-13485, iso-42001, iso-14971, eu-ai-act, eu-mdr-745, gdpr, soc-2, fda-qsr, nist-csf, nis2, hipaa
---
# Compliance OS — Meta-Orchestrator
Multi-framework compliance program orchestration. **Four decisions, no per-framework deep-dive:**
1. **Which frameworks apply to this company?** — `framework_selector.py` ranks the 12 supported frameworks against a company profile (industry, geography, AI use, medical, financial, headcount, customers, healthcare-PHI, NIS2 essential/important entity, US gov contractor) and returns applicable ones with dependency graph
2. **How much do selected frameworks overlap?** — `cross_framework_mapper.py` computes control-level overlap with confidence rating; outputs unified control matrix + evidence-reuse opportunities
3. **What does a mock audit produce?** — `audit_simulator.py` generates 8–15 finding scenarios with severity distribution matching IIA expectations + interview questions per control
4. **What's the unified evidence checklist?** — `evidence_pool_generator.py` consolidates evidence across enabled frameworks; outputs which artefact satisfies which controls across which frameworks
This skill is **NOT** a per-framework deep-dive. The per-framework skills (`ra-qm-team/skills/iso42001-specialist/`, `compliance-team-eu-ai-act/`, `ra-qm-team/skills/gdpr-dsgvo-expert/`, etc.) do the operational work. Compliance OS orchestrates them.
This skill is **NOT** a substitute for binding legal advice. Cross-framework mappings reflect published guidance (ISO standards, regulations, EDPB/Commission guidance, IIA / AICPA professional standards). Novel cross-walks should be reviewed with counsel.
## Keywords
compliance orchestration, multi-framework compliance, compliance OS, cross-framework mapping, control overlap, evidence pool, evidence reuse, audit simulation, mock audit, internal audit programme, GRC, governance risk compliance, framework selector, compliance program, integrated compliance, ISO 19011, IIA IPPF, AICPA AT-C, NIST CSF profile, multi-cert program, SOC 2 + ISO 27001, ISO 27001 + ISO 42001, ISO 13485 + MDR 745, AI Act + ISO 42001, GDPR + ISO 27001, compliance officer, compliance team workflow, certification readiness
## Quick Start
```bash
# Decision A: Which frameworks apply for the company?
python scripts/framework_selector.py # embedded mid-stage AI SaaS sample
python scripts/framework_selector.py path/to/profile.json
# Decision B: Compute cross-framework overlap
python scripts/cross_framework_mapper.py # embedded ISO 27001 + SOC 2 sample
python scripts/cross_framework_mapper.py path/to/control_libs.json
# Decision C: Simulate an audit
python scripts/audit_simulator.py # embedded ISO 27001 sample
python scripts/audit_simulator.py path/to/audit_scope.json
# Decision D: Consolidate evidence checklist across frameworks
python scripts/evidence_pool_generator.py # embedded 3-framework sample
python scripts/evidence_pool_generator.py path/to/program.json
```
## Key Questions (ask these first)
- **Have you named every applicable framework?** Forgetting one means rebuilding the audit program later. Run `framework_selector.py` with your profile.
- **What's the most certificate / regulation your company already operates?** That's your reuse anchor. Map every new framework against it.
- **What's the audit calendar?** A multi-framework program means surveillance audits stacked through the year — plan auditor independence + capacity.
- **Where is evidence stored?** Multi-framework programs collapse when evidence lives in one team's drive without an index. Run `evidence_pool_generator.py` to surface the reuse opportunities.
- **What's the management-review cadence across frameworks?** Each framework wants its own management review, but a single integrated review (per ISO Annex SL) typically satisfies all of them with one calendar slot.
- **Who owns the meta-program?** If no single accountable role, the program fragments.
## Core Responsibilities
### 1. Framework Selection
**The framework:** company-profile JSON in → applicable-framework list out with dependency graph.
**Deterministic logic:**
- Medical device → ISO 13485 + ISO 14971 + (EU MDR 745 if EU market) + (FDA QSR if US market)
- Customer-facing AI → ISO 42001 + EU AI Act (if EU users) + GDPR (if personal data)
- B2B SaaS with enterprise customers → SOC 2 + ISO 27001 (often required for procurement)
- EU customers + personal data → GDPR mandatory
- Highly regulated industry (financial, health) → additional sectoral overlays
**Run** `framework_selector.py` to apply the decision rules.
### 2. Cross-Framework Control Mapping
**The framework:** for each selected framework, parse its control library; compute overlap with other selected frameworks.
**Per merged-control output:**
- Mapping confidence (HIGH / MEDIUM / LOW)
- Evidence-reuse opportunity (single artefact satisfies N controls)
- Per-framework citation
- Implementation guidance reusable across frameworks
**Densest known overlap:** ISO 27001 Annex A ↔ SOC 2 Trust Services Criteria — historically ~75% control coverage shared. Adding ISO 42001 brings AI-specific controls; adding GDPR brings privacy-specific.
**Run** `cross_framework_mapper.py` with framework control libraries.
### 3. Audit Simulation
**The framework:** generate a realistic mock internal audit per ISO 19011 + IIA IPPF standards.
**Per audit output:**
- 8–15 finding scenarios per ISO 19011 typical depth
- Severity distribution: ≥ 40% observations/OFI, ≤ 15% critical/major (IIA expectation for healthy programs)
- Interview questions per scoped control (3–5 questions per control)
- Document-review request list
- Walk-through requests where applicable
**Run** `audit_simulator.py` with framework + scope.
### 4. Evidence Pool
**The framework:** consolidate evidence requirements across enabled frameworks; identify reuse opportunities.
**Output:**
- Evidence artefact list (e.g., access-review log, supplier risk register, incident log)
- Per artefact: list of (framework, control) tuples it satisfies
- Reuse-leverage score (artefact A satisfies N controls across M frameworks)
- Acquisition cost estimate (effort to produce + maintain)
**Run** `evidence_pool_generator.py` with program config.
## Workflows
### Workflow 1: Program Bootstrap (multi-framework, 4–8 weeks)
**Goal:** stand up a compliance program covering 2–4 frameworks simultaneously.
```bash
# 1. Run framework selector with company profile
python scripts/framework_selector.py profile.json
# 2. For each applicable framework, identify the per-framework skill and run its gap analysis
# 3. Run cross-framework mapper to identify reuse opportunities
python scripts/cross_framework_mapper.py control_libs.json
# 4. Run evidence pool generator to consolidate
python scripts/evidence_pool_generator.py program.json
# 5. Cross-check with cs-compliance-officer agent
# 6. Output: prioritized program backlog with owners + dates
```
### Workflow 2: Annual Audit Calendar (yearly)
**Goal:** plan internal audit cycles covering all applicable frameworks.
```bash
# 1. Refresh framework selector if profile changed
python scripts/framework_selector.py profile.json
# 2. For each framework, run its internal-audit-plan tool
# (e.g., aims_audit_scheduler.py for ISO 42001; isms_audit_scheduler.py for ISO 27001)
# 3. Coordinate the audit calendar across frameworks (auditor independence + capacity)
# 4. Run audit simulator for each framework to prep auditors
python scripts/audit_simulator.py scope.json
# 5. Output: integrated audit calendar with owners + auditor assignments
```
### Workflow 3: Pre-Certification Readiness (per new framework, 6–12 weeks)
**Goal:** prepare for an external certification audit.
```bash
# 1. Run gap analysis for the new framework
# (ISO 42001: aims_gap_analyzer.py; ISO 27001: compliance_checker.py; SOC 2: gap_analyzer.py)
# 2. Run cross-framework mapper against already-certified frameworks
python scripts/cross_framework_mapper.py control_libs.json
# 3. Reuse evidence for HIGH-confidence mappings; build new for MEDIUM/LOW
# 4. Run audit simulator to dry-run the certification audit
python scripts/audit_simulator.py scope.json
# 5. Close remaining gaps before external auditor stage 1
```
### Workflow 4: Evidence Pool Consolidation (quarterly)
**Goal:** keep the unified evidence pool fresh + reusable.
```bash
# 1. Refresh evidence pool generator
python scripts/evidence_pool_generator.py program.json
# 2. Identify HIGH-reuse-leverage artefacts (1 evidence -> 5+ controls)
# 3. Confirm evidence freshness (within retention requirement per framework)
# 4. Audit the evidence pool itself (no orphan controls, no stale evidence)
```
## Output Standards
```
**Bottom Line:** [one sentence — what's the multi-framework picture + biggest reuse opportunity]
**The Decision:** [one of: framework-set | overlap-map | audit-plan | evidence-consolidation]
**The Evidence:** [framework names + control IDs from the tool, not adjectives]
**How to Act:** [3 concrete next steps with owners + dates]
**Your Decision:** [the call only the compliance officer can make — which frameworks to pursue, audit cycle priority, evidence-reuse policy]
```
## Adjacent Skills
- `../../ra-qm-team/skills/iso42001-specialist/` — ISO 42001 deep-dive (paired with compliance-team-iso42001 plugin)
- `../../ra-qm-team/skills/eu-ai-act-specialist/` — EU AI Act deep-dive (paired with compliance-team-eu-ai-act plugin)
- `../../ra-qm-team/skills/information-security-manager-iso27001/` — ISO 27001 ISMS deep-dive
- `../../ra-qm-team/skills/quality-manager-qms-iso13485/` — ISO 13485 QMS deep-dive
- `../../ra-qm-team/skills/gdpr-dsgvo-expert/` — GDPR deep-dive
- `../../ra-qm-team/skills/soc2-compliance/` — SOC 2 deep-dive
- `../../ra-qm-team/skills/fda-consultant-specialist/` — FDA QSR deep-dive
- `../../ra-qm-team/skills/mdr-745-specialist/` — EU MDR 745 deep-dive
- `../../ra-qm-team/skills/risk-management-specialist/` — ISO 14971 deep-dive
- `../../c-level-advisor/chief-ai-officer-advisor/` — Executive AI risk decisions (build-vs-buy, model selection)
- `../../c-level-advisor/skills/general-counsel-advisor/` — Legal review for novel cases
## References
- [compliance_os_pattern.md](references/compliance_os_pattern.md) — The meta-framework architecture (configure → map → simulate → consolidate → review); when to use vs not
- [cross_framework_overlap.md](references/cross_framework_overlap.md) — The 9-framework × control-family overlap table with mapping confidence (Phase 3 expands to 12 frameworks via `cross_framework_mapper.py`)
- [audit_simulation_methodology.md](references/audit_simulation_methodology.md) — ISO 19011 + IIA IPPF + AICPA AT-C audit-simulation principles + severity distribution heuristics
- [evidence_management.md](references/evidence_management.md) — Evidence pool design + retention + freshness + reuse-leverage scoring
- [multi_framework_audit_playbook.md](references/multi_framework_audit_playbook.md) — Integrated audit programme for 2+ frameworks (Phase 2)
- [evidence_artifact_reuse_index.md](references/evidence_artifact_reuse_index.md) — Empirically-derived reuse-leverage ranking across all 12 frameworks (Phase 3)
## Phase 3 Asset: Mock Audit Scenario Library
`assets/mock_audit_library.json` — 205 pre-built finding scenarios spanning 12 frameworks + 26 themes + 4 severity levels (34 critical, 88 major, 54 minor, 29 observation). Each scenario tags applicable frameworks; cross-reference `scripts/cross_framework_mapper.py` merged-controls catalogue to resolve framework-specific control IDs. Use as input to enrich `audit_simulator.py` mock audits, as a training resource for new internal auditors, or as the seed for finding-pattern detection across multi-framework programmes.
---
**Version:** 1.2.0
**Status:** Production Ready
FILE:assets/company_profile_template.json
{
"company": "<company name>",
"industry": "<saas | medical_device | financial | healthcare | other>",
"products_include_ai": false,
"ai_high_risk_per_eu": false,
"deploys_ai_in_eu": false,
"products_are_medical_devices": false,
"sells_to_eu_customers": false,
"sells_to_us_customers": false,
"sells_to_enterprise_b2b": false,
"processes_personal_data": false,
"processes_eu_personal_data": false,
"headcount": 0,
"stage": "<seed | series_a | series_b | series_c | growth>",
"processes_phi": false,
"us_healthcare_covered_entity": false,
"us_healthcare_business_associate": false,
"nis2_essential_entity": false,
"nis2_important_entity": false,
"adopts_nist_csf": false,
"us_government_contractor": false
}
FILE:assets/control_library_template.json
{
"program": "<program name>",
"enabled_frameworks": [
"iso_27001",
"soc_2",
"iso_42001",
"eu_ai_act",
"gdpr"
],
"_supported_framework_ids": [
"iso_27001",
"iso_13485",
"iso_42001",
"iso_14971",
"eu_ai_act",
"eu_mdr_745",
"gdpr",
"soc_2",
"fda_qsr"
],
"_note": "Enable only the frameworks the framework_selector returned as applicable. Cross-framework mapper will compute overlap across enabled frameworks only."
}
FILE:assets/mock_audit_library.json
{
"schema_version": "1.0.0",
"description": "Pre-built finding scenarios for mock internal audits. Each scenario has theme + severity + applicable_frameworks tags; cross-reference scripts/cross_framework_mapper.py merged-controls catalogue to resolve framework-specific control IDs.",
"supported_frameworks": [
"iso_27001", "iso_13485", "iso_42001", "iso_14971",
"eu_ai_act", "eu_mdr_745", "gdpr", "soc_2", "fda_qsr",
"nist_csf", "nis2", "hipaa"
],
"severity_levels": {
"critical": "Major nonconformity: absence of, or systemic failure to implement, a required management-system process. Blocks certification at stage 1.",
"major": "Material gap in a required control. Corrective action plan within 30 days.",
"minor": "Localized gap; control works overall. Corrective action within 90 days.",
"observation": "Improvement opportunity; no nonconformity. Optional recommendation."
},
"scenarios": [
{"id": "F-AC-001", "theme": "access_control", "severity": "critical", "title": "Orphaned privileged access from terminations", "description": "Quarterly access review missed 3 cycles; 12 terminated employees retain prod admin access > 90 days post-termination. Audit log shows 4 of them performed actions in the system after their last working day.", "remediation": "Immediate revocation; investigate access logs for unauthorized activity; reinstate quarterly review cadence with automated tooling.", "remediation_days": 14, "applicable_frameworks": ["iso_27001", "soc_2", "iso_42001", "gdpr", "nist_csf", "hipaa", "nis2"]},
{"id": "F-AC-002", "theme": "access_control", "severity": "critical", "title": "Shared admin credentials in production", "description": "Production database admin password shared across 5 engineers; rotation last performed > 12 months ago. No audit trail for individual actions.", "remediation": "Rotate immediately; provision per-user named accounts; enable individual audit logging; document procedure.", "remediation_days": 7, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa", "nis2"]},
{"id": "F-AC-003", "theme": "access_control", "severity": "major", "title": "Quarterly access review evidence lacks justification", "description": "Quarterly access review records exist but lack documented business justification for retained privileges. Reviewers approve in bulk without per-user rationale.", "remediation": "Update review template to require per-user justification; train reviewers; sample-check next quarter.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nist_csf"]},
{"id": "F-AC-004", "theme": "access_control", "severity": "major", "title": "JML workflow does not auto-deprovision", "description": "Joiner-mover-leaver workflow exists but is manual; observed 5+ day gap between HR termination and access revocation.", "remediation": "Implement IDP integration with HR system for auto-deprovisioning within 24 hours; trail for exceptions.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "iso_42001", "gdpr", "hipaa", "nis2"]},
{"id": "F-AC-005", "theme": "access_control", "severity": "major", "title": "MFA not enforced on admin accounts", "description": "Multi-factor authentication is documented in policy but not technically enforced on cloud admin accounts; 8 admin users authenticate without MFA.", "remediation": "Enforce MFA at IDP level; emergency-break-glass procedure documented; close legacy accounts.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa", "nis2", "gdpr"]},
{"id": "F-AC-006", "theme": "access_control", "severity": "minor", "title": "Access review records lack completion timestamps", "description": "Access review records lack documented review-completion timestamps in 2 of 6 sampled reviews. Cannot confirm review was completed on time.", "remediation": "Update review tooling to capture timestamp at review-action time; backfill where possible.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa"]},
{"id": "F-AC-007", "theme": "access_control", "severity": "minor", "title": "RBAC matrix doesn't cover cloud resources", "description": "Role-based access control matrix exists for application tier but does not address cloud-resource scope (IAM policies, S3 buckets, KMS keys).", "remediation": "Extend RBAC matrix; document cloud-IAM role-to-business-role mapping.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-AC-008", "theme": "access_control", "severity": "observation", "title": "Consider just-in-time (JIT) access for production", "description": "Standing access to production is the default; JIT access with approval workflow would reduce blast radius and improve audit trail.", "remediation": "Pilot JIT tooling (e.g., Teleport, ConductorOne, ConsoleMe) for one team.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-AC-009", "theme": "access_control", "severity": "observation", "title": "Privileged access review cadence could be more frequent", "description": "Quarterly cadence meets standard; for ≥ critical-tier systems, monthly review provides earlier detection of orphaned access.", "remediation": "Increase cadence for critical-tier systems to monthly.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nist_csf"]},
{"id": "F-AC-010", "theme": "access_control", "severity": "observation", "title": "Session timeout policies inconsistent", "description": "Session-timeout policies vary across applications (30 min in CRM, 8 hours in BI tool, no timeout in internal admin tool). Inconsistent risk posture.", "remediation": "Define policy by data sensitivity tier; align tooling configuration.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa"]},
{"id": "F-AI-001", "theme": "asset_inventory", "severity": "major", "title": "Asset inventory missing cloud + SaaS + AI tools", "description": "Asset register includes server inventory but omits 60% of SaaS tools and 100% of AI/LLM tools acquired in past 12 months. No central source of truth.", "remediation": "Integrate SSO with SaaS-discovery tooling; require AI-tool registration before procurement; quarterly refresh.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "iso_42001", "nist_csf", "gdpr"]},
{"id": "F-AI-002", "theme": "asset_inventory", "severity": "major", "title": "Data classification scheme not applied", "description": "Data classification scheme documented (public / internal / confidential / restricted) but only 30% of data stores have classification labels applied.", "remediation": "Apply classification to remaining stores; automate via DLP tooling where feasible.", "remediation_days": 120, "applicable_frameworks": ["iso_27001", "soc_2", "gdpr", "hipaa", "nist_csf"]},
{"id": "F-AI-003", "theme": "asset_inventory", "severity": "minor", "title": "Asset owners not assigned for 15% of assets", "description": "15% of inventory entries lack named owners; orphan ownership impedes timely incident response.", "remediation": "Assign owners; require owner field on new asset creation.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-AI-004", "theme": "asset_inventory", "severity": "minor", "title": "Third-party AI services not tagged in inventory", "description": "Inventory does not flag which assets are powered by third-party AI services (e.g., OpenAI, Anthropic, Cohere). Material for ISO 42001 A.10 + EU AI Act Article 25.", "remediation": "Add AI-vendor tag; update procurement intake form.", "remediation_days": 90, "applicable_frameworks": ["iso_42001", "eu_ai_act", "iso_27001"]},
{"id": "F-AI-005", "theme": "asset_inventory", "severity": "observation", "title": "Asset decommissioning workflow informal", "description": "When assets are decommissioned, data destruction is documented but inventory entries persist; clutters reporting.", "remediation": "Add decommission state to inventory schema; archive after retention period.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa"]},
{"id": "F-RM-001", "theme": "risk_management", "severity": "critical", "title": "Risk register without treatment plans", "description": "Risk register identifies 30+ risks but lacks documented treatment plans (modify/share/retain/avoid per ISO 23894) for high/critical risks.", "remediation": "Run risk-treatment workshop per high/critical risk; document treatment + signoff; link to specific controls.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "iso_42001", "soc_2", "nist_csf", "nis2", "hipaa"]},
{"id": "F-RM-002", "theme": "risk_management", "severity": "critical", "title": "AI risk assessment not re-run after material model change", "description": "AI risk assessment last performed at initial deployment 18 months ago. Model has been retrained twice; risk profile not re-evaluated.", "remediation": "Trigger re-assessment; update register; document drift monitoring threshold; commit to re-assessment on every material change.", "remediation_days": 45, "applicable_frameworks": ["iso_42001", "eu_ai_act"]},
{"id": "F-RM-003", "theme": "risk_management", "severity": "major", "title": "Risk methodology inconsistently applied", "description": "Different teams use different risk-scoring methodologies; severity scores not comparable across the register.", "remediation": "Standardize on single methodology (e.g., 5x5 likelihood × impact matrix); train risk owners; re-score existing register.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_14971", "nist_csf", "soc_2"]},
{"id": "F-RM-004", "theme": "risk_management", "severity": "major", "title": "Residual risk acceptance lacks management signoff", "description": "30% of 'retain' risk-treatment decisions lack documented management signoff. Some retain decisions made by individual contributors.", "remediation": "Define signoff matrix by severity; backfill where possible; route remaining retain decisions through proper authority.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_14971", "nis2", "hipaa"]},
{"id": "F-RM-005", "theme": "risk_management", "severity": "major", "title": "Risk register not updated for 6+ months", "description": "Risk register last refreshed > 6 months ago. New risks from product changes, new vendors, regulatory developments not captured.", "remediation": "Refresh; commit to quarterly cadence minimum.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "iso_42001", "nist_csf", "nis2"]},
{"id": "F-RM-006", "theme": "risk_management", "severity": "minor", "title": "DPIA exists but Article 35(7) elements incomplete", "description": "DPIA documented for high-risk processing but does not cover all Article 35(7)(a)-(d) required elements (missing necessity + proportionality assessment).", "remediation": "Update DPIA template; refresh affected DPIAs.", "remediation_days": 60, "applicable_frameworks": ["gdpr", "iso_42001"]},
{"id": "F-RM-007", "theme": "risk_management", "severity": "minor", "title": "Risk treatment plans lack effective-date tracking", "description": "Treatment plans are documented but lack effective-date or expected-completion fields; cannot track remediation timeliness.", "remediation": "Add date fields; update existing entries.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "soc_2"]},
{"id": "F-RM-008", "theme": "risk_management", "severity": "observation", "title": "Consider FAIR quantitative risk methodology for top-tier risks", "description": "Current methodology is qualitative; quantitative analysis (e.g., Open FAIR) for top-5 risks would improve decision quality.", "remediation": "Pilot FAIR on 2-3 top risks.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "nist_csf"]},
{"id": "F-RM-009", "theme": "risk_management", "severity": "observation", "title": "Risk-related KPIs not reported to executive", "description": "Risk register exists but no rolled-up KPIs (e.g., # critical risks open, mean time to treatment) reported in management review.", "remediation": "Add risk KPIs to management review inputs.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "iso_42001", "nist_csf"]},
{"id": "F-SM-001", "theme": "supplier_management", "severity": "critical", "title": "Critical SaaS in use without DPA", "description": "Critical SaaS supplier (handles personal data of 500K+ users) in use without signed DPA per GDPR Article 28. Pre-existing arrangement not refreshed since 2018.", "remediation": "Engage vendor for DPA execution; if vendor refuses, evaluate replacement.", "remediation_days": 30, "applicable_frameworks": ["gdpr", "iso_27001", "soc_2", "hipaa"]},
{"id": "F-SM-002", "theme": "supplier_management", "severity": "critical", "title": "Business Associate Agreement missing for HIPAA-relevant vendor", "description": "Vendor processes PHI on behalf of the organization but no signed Business Associate Agreement (BAA) per HIPAA §164.314(a). Material exposure.", "remediation": "Sign BAA; if vendor refuses, evaluate replacement; document remediation timeline.", "remediation_days": 30, "applicable_frameworks": ["hipaa", "iso_27001"]},
{"id": "F-SM-003", "theme": "supplier_management", "severity": "major", "title": "Annual supplier reviews incomplete", "description": "Annual supplier security review not completed for 3 of 8 critical suppliers in past year.", "remediation": "Run overdue reviews; calendar future reviews; document escalation for non-responsive vendors.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "iso_42001", "hipaa", "nis2", "gdpr"]},
{"id": "F-SM-004", "theme": "supplier_management", "severity": "major", "title": "Sub-processor list not maintained", "description": "Critical supplier handling personal data uses sub-processors; the sub-processor list is not maintained or available; GDPR Article 28(2) not satisfied.", "remediation": "Request sub-processor list from vendor; establish change notification mechanism; document.", "remediation_days": 60, "applicable_frameworks": ["gdpr", "iso_27001", "nist_csf"]},
{"id": "F-SM-005", "theme": "supplier_management", "severity": "major", "title": "AI-specific contract clauses not in vendor agreements", "description": "Third-party AI service in use; contract lacks AI-specific clauses (training-data use restrictions, drift notification, sub-processor list for AI sub-services).", "remediation": "Negotiate addendum; document acceptance.", "remediation_days": 90, "applicable_frameworks": ["iso_42001", "eu_ai_act", "iso_27001"]},
{"id": "F-SM-006", "theme": "supplier_management", "severity": "major", "title": "Supplier exit / termination procedure not documented", "description": "No procedure for safe vendor exit (data return, model deletion, monitoring transition). Discovered during attempt to terminate one supplier.", "remediation": "Draft procedure; pilot on next vendor termination; document.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "iso_42001", "soc_2", "gdpr"]},
{"id": "F-SM-007", "theme": "supplier_management", "severity": "minor", "title": "Vendor onboarding checklist applied inconsistently", "description": "Supplier onboarding checklist exists but is bypassed in 'urgent' procurements; 4 of 12 recent vendors lack complete onboarding evidence.", "remediation": "Make checklist mandatory at procurement gate; remediate gaps in existing 4.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-SM-008", "theme": "supplier_management", "severity": "minor", "title": "Supplier SOC 2 reports collected but not reviewed", "description": "Critical suppliers' SOC 2 Type II reports collected on initial onboarding but not reviewed annually as new reports issued.", "remediation": "Set calendar for annual review; document key findings + acceptance.", "remediation_days": 60, "applicable_frameworks": ["soc_2", "iso_27001"]},
{"id": "F-SM-009", "theme": "supplier_management", "severity": "observation", "title": "Consider centralizing supplier risk evidence in GRC tool", "description": "Supplier evidence scattered across procurement Drive, Compliance Drive, and email. Centralization in GRC tool would reduce audit prep effort.", "remediation": "Evaluate GRC tooling; migrate over 6 months.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "iso_42001"]},
{"id": "F-SM-010", "theme": "supplier_management", "severity": "observation", "title": "Vendor risk-tiering could be more granular", "description": "Vendors tier as 'critical / non-critical' currently; more granular tiers (e.g., based on data type, criticality, integration depth) would refine review cadence.", "remediation": "Define 3-tier model; reclassify existing inventory.", "remediation_days": 120, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa"]},
{"id": "F-IR-001", "theme": "incident_response", "severity": "critical", "title": "GDPR Article 33 breach notification missed", "description": "Breach occurred 96 hours ago; supervisory authority not notified despite Article 33 72-hour requirement. Investigation revealed unclear breach-criteria decision.", "remediation": "File notification immediately with rationale for delay; review breach-criteria decision tree; conduct tabletop exercise; document.", "remediation_days": 7, "applicable_frameworks": ["gdpr", "iso_27001", "hipaa", "nis2"]},
{"id": "F-IR-002", "theme": "incident_response", "severity": "critical", "title": "Recent P1 incident lacks PIR within SLA", "description": "P1 production incident occurred 45 days ago; post-incident review (PIR) not documented within stated 30-day SLA.", "remediation": "Complete PIR immediately; identify corrective actions; calendar future PIRs.", "remediation_days": 14, "applicable_frameworks": ["iso_27001", "soc_2", "iso_42001", "nist_csf"]},
{"id": "F-IR-003", "theme": "incident_response", "severity": "critical", "title": "Severity definitions inconsistently applied", "description": "Severity definitions documented but inconsistently applied across teams; impact analysis varies. Two recent P2 incidents arguably P1 by definition.", "remediation": "Train responders on severity rubric; calibration exercise quarterly; track severity-classification consistency.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "iso_42001", "gdpr", "hipaa"]},
{"id": "F-IR-004", "theme": "incident_response", "severity": "major", "title": "Notification SLAs not aligned across frameworks", "description": "GDPR 72h, NIS2 24h-early-warning + 72h-notification, EU AI Act 15-day (or 2-day critical-infra), HIPAA 60-day. Internal procedures collapse to a single 'breach' notification without per-framework branching.", "remediation": "Update IR procedure to branch by applicable framework; train responders.", "remediation_days": 60, "applicable_frameworks": ["gdpr", "nis2", "eu_ai_act", "hipaa", "iso_27001"]},
{"id": "F-IR-005", "theme": "incident_response", "severity": "major", "title": "Incident commander rotation not documented", "description": "Incident commander rotation exists informally but is not documented; recent incidents had ambiguous IC ownership.", "remediation": "Document rotation; publish on-call schedule.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-IR-006", "theme": "incident_response", "severity": "major", "title": "Breach log incomplete per Article 33(5)", "description": "GDPR breach log captures only DPA-notifiable events; Article 33(5) requires ALL breaches logged regardless of notifiability.", "remediation": "Update breach log scope; backfill recent breaches; train DPO on requirement.", "remediation_days": 60, "applicable_frameworks": ["gdpr", "iso_27001", "hipaa"]},
{"id": "F-IR-007", "theme": "incident_response", "severity": "major", "title": "Detection mechanism gaps", "description": "Mean time to detect (MTTD) for past 3 incidents averaged 8 days; SIEM rules not tuned for recently-onboarded systems.", "remediation": "Audit SIEM coverage; tune rules; test detection for high-impact attack scenarios.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2"]},
{"id": "F-IR-008", "theme": "incident_response", "severity": "minor", "title": "Tabletop exercise not conducted in last 12 months", "description": "Annual incident-response tabletop exercise not performed in past 12 months.", "remediation": "Schedule + run tabletop; document lessons learned.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2", "hipaa"]},
{"id": "F-IR-009", "theme": "incident_response", "severity": "minor", "title": "Customer notification timing not tracked", "description": "Customer-facing incident notifications sent but timing not tracked against committed SLA. Cannot demonstrate SLA compliance.", "remediation": "Track notification timestamps; report against SLA quarterly.", "remediation_days": 60, "applicable_frameworks": ["soc_2", "iso_27001", "gdpr"]},
{"id": "F-IR-010", "theme": "incident_response", "severity": "observation", "title": "Consider chaos engineering for resilience testing", "description": "Incident response prepares for failures; chaos engineering would proactively surface latent weaknesses.", "remediation": "Pilot chaos engineering on non-prod first; expand if mature.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-ML-001", "theme": "monitoring_logging", "severity": "critical", "title": "Production application logs disabled", "description": "Production application logs disabled in past 30 days due to disk space; not detected until audit fieldwork. 30-day blind spot.", "remediation": "Re-enable; resize storage; alert on log volume drops; investigate any incidents during blind period.", "remediation_days": 7, "applicable_frameworks": ["iso_27001", "soc_2", "iso_42001", "hipaa", "nist_csf"]},
{"id": "F-ML-002", "theme": "monitoring_logging", "severity": "major", "title": "Log retention misaligned with framework requirement", "description": "Log retention configured at 90 days; ISO 27001 + framework requirements expect 12 months minimum for some logs.", "remediation": "Update retention configuration; backfill from archives where feasible; document policy.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nist_csf", "gdpr"]},
{"id": "F-ML-003", "theme": "monitoring_logging", "severity": "major", "title": "Tamper-evident logging not enforced", "description": "Tamper-evident logging not enforced on privileged-user activity logs; logs writable to same store as application data.", "remediation": "Move logs to write-once storage; document architecture; verify immutability.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nist_csf", "nis2"]},
{"id": "F-ML-004", "theme": "monitoring_logging", "severity": "major", "title": "AI model drift not monitored", "description": "AI system in production; no drift monitoring against original validation data. No defined drift threshold for retraining.", "remediation": "Implement drift monitoring; define threshold; escalation path.", "remediation_days": 90, "applicable_frameworks": ["iso_42001", "eu_ai_act"]},
{"id": "F-ML-005", "theme": "monitoring_logging", "severity": "minor", "title": "Monitoring alert thresholds not documented", "description": "Monitoring alert thresholds exist in tooling but not documented; rationale unclear.", "remediation": "Document thresholds + rationale + on-call response action.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-ML-006", "theme": "monitoring_logging", "severity": "minor", "title": "Cloud audit logs not centralized", "description": "Cloud audit logs (CloudTrail/Cloud Audit Logs) exist per account but not centralized to SIEM; cross-account analysis manual.", "remediation": "Forward logs to central SIEM; configure cross-account analysis.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-ML-007", "theme": "monitoring_logging", "severity": "observation", "title": "Consider anomaly detection on top of rule-based monitoring", "description": "Current monitoring is rule-based; anomaly detection (statistical or ML-based) would surface novel patterns.", "remediation": "Pilot on key data flows.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-CM-001", "theme": "change_management", "severity": "critical", "title": "Emergency change procedure not formalized", "description": "Emergency change procedure not documented; observed 3 cases of production changes in past 30 days without recorded approval. Two affected customer data.", "remediation": "Draft emergency change procedure including retroactive review; train engineers; audit recent emergency changes.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "soc_2", "iso_13485", "hipaa", "nist_csf"]},
{"id": "F-CM-002", "theme": "change_management", "severity": "major", "title": "Change advisory board rubber-stamps", "description": "Change advisory board records show approvals but zero rejected changes in last 6 months. Board likely not exercising substantive review.", "remediation": "Calibration training for board; track reject + revise rate; ensure reviewers have time + context.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "iso_13485"]},
{"id": "F-CM-003", "theme": "change_management", "severity": "major", "title": "Rollback procedure not tested", "description": "Rollback procedure documented but not tested for 2 services in audit scope. Cannot confirm operability.", "remediation": "Test rollback in staging; document; schedule quarterly verification.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "iso_13485", "nist_csf"]},
{"id": "F-CM-004", "theme": "change_management", "severity": "minor", "title": "Post-implementation reviews skipped for high-risk changes", "description": "Change advisory board records show approvals but no post-implementation review for high-risk changes (defined by impact rubric).", "remediation": "Reinstate post-implementation review for high-risk; define follow-up timeline.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "iso_13485"]},
{"id": "F-CM-005", "theme": "change_management", "severity": "observation", "title": "Link change records to deployment automation", "description": "Change records and deployment automation are separate systems; linking would strengthen evidence chain.", "remediation": "Integrate via deployment tagging.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-BC-001", "theme": "business_continuity", "severity": "critical", "title": "BCP/DRP exists but never tested", "description": "Business continuity + disaster recovery plans exist on paper but no recovery exercise in 24+ months. Untested = ineffective.", "remediation": "Conduct full DR exercise; document results; commit to annual exercise cadence.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2", "hipaa"]},
{"id": "F-BC-002", "theme": "business_continuity", "severity": "major", "title": "RPO/RTO objectives not measured", "description": "Recovery objectives defined but not measured during recent failover events. Cannot confirm objectives are achievable.", "remediation": "Measure during next exercise; tune objectives or recovery capability.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa"]},
{"id": "F-BC-003", "theme": "business_continuity", "severity": "major", "title": "Backup integrity not verified", "description": "Backups occur but restoration testing not performed in past 12 months. Cannot confirm backups are usable.", "remediation": "Quarterly restoration tests; document verification evidence.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nis2"]},
{"id": "F-BC-004", "theme": "business_continuity", "severity": "minor", "title": "BCP doesn't address third-party SaaS outage", "description": "BCP covers self-hosted infrastructure; doesn't address critical SaaS-vendor outage scenarios.", "remediation": "Extend BCP for SaaS outage scenarios; document vendor SLAs + alternatives.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2"]},
{"id": "F-BC-005", "theme": "business_continuity", "severity": "observation", "title": "Consider chaos game-day exercises", "description": "Annual DR exercise meets standard; chaos game-day adds value by testing under more realistic conditions.", "remediation": "Pilot game-day for one service.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-CT-001", "theme": "competence_training", "severity": "major", "title": "Annual security training not 100% complete", "description": "Annual security training completion is 89% across the company; 12 employees past due > 30 days.", "remediation": "Escalate to managers for non-completers; revoke access for chronic non-completers; document policy.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa", "nis2"]},
{"id": "F-CT-002", "theme": "competence_training", "severity": "major", "title": "AI literacy training not in place", "description": "EU AI Act Article 4 requires AI literacy for staff dealing with AI systems; no AI-specific training implemented.", "remediation": "Develop + roll out AI literacy training; track completion by role.", "remediation_days": 90, "applicable_frameworks": ["eu_ai_act", "iso_42001"]},
{"id": "F-CT-003", "theme": "competence_training", "severity": "major", "title": "Competence requirements undefined for ML engineers", "description": "Competence requirements defined for engineering roles but not specifically for ML engineers; assumes 'they have degrees'.", "remediation": "Define ML-engineer competence requirements; verify against existing staff.", "remediation_days": 90, "applicable_frameworks": ["iso_42001"]},
{"id": "F-CT-004", "theme": "competence_training", "severity": "minor", "title": "Training effectiveness verification missing", "description": "Training completion recorded but effectiveness verification (assessment, simulation, observed behavior) not performed.", "remediation": "Add post-training assessment; track scores.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485", "soc_2"]},
{"id": "F-CT-005", "theme": "competence_training", "severity": "observation", "title": "Consider role-based training tiers", "description": "Training is uniform across roles; role-based tiers would surface compliance-officer-specific, dev-specific, etc.", "remediation": "Design role-tiered curriculum.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-DG-001", "theme": "data_governance", "severity": "critical", "title": "Training data lacks provenance records", "description": "AI training data sourced from multiple vendors + scraped sources; no provenance records. EU AI Act Article 10(2)(d) + ISO 42001 A.7.4 not satisfied.", "remediation": "Audit current training data; document provenance per source; remove data without verifiable provenance.", "remediation_days": 90, "applicable_frameworks": ["iso_42001", "eu_ai_act", "gdpr"]},
{"id": "F-DG-002", "theme": "data_governance", "severity": "critical", "title": "PII in training data without lawful basis", "description": "Training data contains PII; lawful basis (GDPR Article 6) not documented for AI training use case. Article 10(5) AI Act bias-detection exception not applicable here.", "remediation": "Document lawful basis or remove PII; if legitimate interests, document LIA; halt training until resolved.", "remediation_days": 30, "applicable_frameworks": ["gdpr", "iso_42001", "eu_ai_act"]},
{"id": "F-DG-003", "theme": "data_governance", "severity": "major", "title": "Data quality dimensions not defined", "description": "Data quality monitoring exists but dimensions (completeness, accuracy, timeliness, consistency) not formally defined. Audit against ISO 42001 A.7.3 incomplete.", "remediation": "Define dimensions per data store; document measurement methodology.", "remediation_days": 90, "applicable_frameworks": ["iso_42001", "iso_27001", "gdpr"]},
{"id": "F-DG-004", "theme": "data_governance", "severity": "major", "title": "Article 30 RoPA stale", "description": "GDPR Article 30 records of processing activities last refreshed 8 months ago; new processing activities not captured.", "remediation": "Refresh RoPA; commit to quarterly updates; integrate with new-feature intake.", "remediation_days": 60, "applicable_frameworks": ["gdpr"]},
{"id": "F-DG-005", "theme": "data_governance", "severity": "major", "title": "Retention schedules not enforced", "description": "Data retention schedules documented but not enforced in tooling. Data persists beyond stated retention.", "remediation": "Implement automated retention enforcement; backfill cleanup; document deletions.", "remediation_days": 90, "applicable_frameworks": ["gdpr", "iso_27001", "hipaa", "iso_42001"]},
{"id": "F-DG-006", "theme": "data_governance", "severity": "minor", "title": "Consent management workflow lacks withdrawal mechanism", "description": "Consent collected at signup; withdrawal mechanism exists in privacy notice but not technically implemented.", "remediation": "Implement self-service consent withdrawal; honour within reasonable time.", "remediation_days": 90, "applicable_frameworks": ["gdpr"]},
{"id": "F-DG-007", "theme": "data_governance", "severity": "observation", "title": "Consider data lineage tooling", "description": "Data flows documented manually; data-lineage tooling would automate + maintain freshness.", "remediation": "Evaluate tooling (e.g., OpenLineage, DataHub, Atlan).", "remediation_days": 180, "applicable_frameworks": ["iso_42001", "gdpr"]},
{"id": "F-CR-001", "theme": "cryptography", "severity": "major", "title": "Encryption at rest using deprecated algorithm", "description": "Some data stores still use deprecated AES-128 (or 3DES); current standard expects AES-256.", "remediation": "Plan migration; document; complete within 6 months.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa", "nis2", "gdpr"]},
{"id": "F-CR-002", "theme": "cryptography", "severity": "major", "title": "Key rotation not enforced", "description": "Cryptographic key rotation policy exists (annual) but not enforced; production keys 3+ years old.", "remediation": "Rotate immediately; automate rotation via KMS; document.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa", "nis2", "gdpr"]},
{"id": "F-CR-003", "theme": "cryptography", "severity": "major", "title": "TLS configuration permits deprecated versions", "description": "TLS 1.0 + 1.1 still accepted on public endpoints; current standard expects TLS 1.2 minimum.", "remediation": "Disable TLS 1.0 + 1.1; verify all clients support 1.2+; document.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa", "nis2", "gdpr"]},
{"id": "F-CR-004", "theme": "cryptography", "severity": "minor", "title": "Cryptographic inventory incomplete", "description": "Cryptographic inventory exists but lacks documentation of algorithm + key length per data store.", "remediation": "Audit each store; document; flag deprecated algorithms.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "nist_csf", "hipaa"]},
{"id": "F-CR-005", "theme": "cryptography", "severity": "observation", "title": "Consider post-quantum cryptography roadmap", "description": "Current crypto is RSA + ECC; post-quantum standards finalized in 2024. Long-term planning for migration recommended.", "remediation": "Define PQC migration roadmap.", "remediation_days": 365, "applicable_frameworks": ["iso_27001", "nist_csf", "nis2"]},
{"id": "F-SD-001", "theme": "secure_sdlc", "severity": "critical", "title": "Production deploy without SAST results", "description": "Recent production deploys lack SAST scan evidence; SAST configured in CI but bypassed via manual override.", "remediation": "Make SAST a required gate; remove override capability for production; investigate bypassed deploys.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-SD-002", "theme": "secure_sdlc", "severity": "major", "title": "Code review records inconsistent", "description": "Some commits to main branch lack documented review; review-required branch protection not consistently enforced.", "remediation": "Enforce review on protected branches across all repos; audit recent commits.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-SD-003", "theme": "secure_sdlc", "severity": "major", "title": "Threat modeling not performed for new services", "description": "New service launched last quarter without threat model. ISO 27001 A.8.25-31 + secure-by-design expectations not met.", "remediation": "Retroactive threat model; integrate threat modeling into design-review gate.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2"]},
{"id": "F-SD-004", "theme": "secure_sdlc", "severity": "minor", "title": "Dependency scanning missing for some repos", "description": "Dependency scanning configured for production services but not for internal tools.", "remediation": "Extend dependency scanning to all repos.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-SD-005", "theme": "secure_sdlc", "severity": "observation", "title": "Consider supply-chain security per SLSA", "description": "Build provenance + supply-chain security gaps; SLSA framework would formalize improvements.", "remediation": "Adopt SLSA Level 2 minimum for production builds.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "nist_csf", "nis2"]},
{"id": "F-VM-001", "theme": "vulnerability_mgmt", "severity": "critical", "title": "Critical vulnerabilities past patch SLA", "description": "5 critical-severity CVEs in production older than 30-day patch SLA; one is actively exploited in wild.", "remediation": "Patch immediately; document compensating controls if patching not possible; investigate any compromise indicators.", "remediation_days": 14, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2", "hipaa"]},
{"id": "F-VM-002", "theme": "vulnerability_mgmt", "severity": "major", "title": "Vulnerability scanning not running weekly", "description": "Scanning configured but execution stopped in past quarter due to tool change. 90+ day blind spot.", "remediation": "Resume scanning; investigate vulnerabilities discovered post-resume.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2", "hipaa"]},
{"id": "F-VM-003", "theme": "vulnerability_mgmt", "severity": "major", "title": "Patch SLAs not defined by severity", "description": "Patch SLA defined for 'all CVEs within 90 days'; not differentiated by severity. Critical vulns should be < 30 days.", "remediation": "Define severity-tiered SLAs; communicate; track compliance.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa", "nis2"]},
{"id": "F-VM-004", "theme": "vulnerability_mgmt", "severity": "minor", "title": "Vulnerability exceptions lack expiry", "description": "Exception tracking exists but exceptions have no expiry; some are 18+ months old without re-evaluation.", "remediation": "Add expiry; re-evaluate all open exceptions.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-VM-005", "theme": "vulnerability_mgmt", "severity": "observation", "title": "Consider container image base auditing", "description": "Vulnerability scanning catches runtime; auditing base images at build time would prevent vulnerabilities reaching production.", "remediation": "Add build-time scanning + base-image inventory.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "nist_csf"]},
{"id": "F-PS-001", "theme": "physical_security", "severity": "major", "title": "Server room access log incomplete", "description": "Server room access log shows entries but lacks visitor escort records for 4 of 12 sampled entries.", "remediation": "Reinforce escort policy; train + supervise; verify in next quarter.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "iso_13485", "hipaa"]},
{"id": "F-PS-002", "theme": "physical_security", "severity": "major", "title": "Workstation security policy not enforced", "description": "Workstation locking policy documented but not enforced; observed several unattended unlocked workstations during walkthrough.", "remediation": "Configure auto-lock at 5 min; train staff; verify.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "hipaa", "soc_2"]},
{"id": "F-PS-003", "theme": "physical_security", "severity": "minor", "title": "Visitor sign-in process bypassed", "description": "Visitor sign-in book exists but bypassed for 'known' visitors; 8 sampled visits lack sign-in evidence.", "remediation": "Reinforce policy + signage; consider electronic visitor management.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_13485", "hipaa"]},
{"id": "F-PS-004", "theme": "physical_security", "severity": "observation", "title": "Consider biometric access for sensitive zones", "description": "Current access is card-based; biometric for sensitive zones (server rooms, R&D labs) would strengthen access discipline.", "remediation": "Evaluate biometric tooling; pilot.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "hipaa", "iso_13485"]},
{"id": "F-DP-001", "theme": "data_protection_privacy", "severity": "critical", "title": "Right to erasure not honored within SLA", "description": "Erasure request from 60 days ago not fully completed; data persists in 3 systems including backups. GDPR Article 17 + 12(3) breached.", "remediation": "Complete erasure; identify all systems; commit to per-system erasure workflow.", "remediation_days": 14, "applicable_frameworks": ["gdpr"]},
{"id": "F-DP-002", "theme": "data_protection_privacy", "severity": "critical", "title": "International transfer without SCCs", "description": "Personal data transferred to US subprocessor; no adequacy decision relied on, no SCCs signed, no derogation applies. Schrems II requirement breached.", "remediation": "Execute SCCs (Commission 2021/914); conduct TIA per EDPB Rec. 01/2020; supplementary measures where needed.", "remediation_days": 30, "applicable_frameworks": ["gdpr"]},
{"id": "F-DP-003", "theme": "data_protection_privacy", "severity": "major", "title": "Privacy notice missing Article 13/14 elements", "description": "Privacy notice published but lacks retention periods + data subject rights detail per Article 13(2).", "remediation": "Update notice; publish version; track versions for evidence trail.", "remediation_days": 30, "applicable_frameworks": ["gdpr"]},
{"id": "F-DP-004", "theme": "data_protection_privacy", "severity": "major", "title": "Cookie banner pre-ticks consent", "description": "Cookie banner pre-ticks non-essential cookies; valid consent per GDPR Article 7 + EDPB guidance requires affirmative action.", "remediation": "Redesign banner; default to no consent for non-essential; document A/B test.", "remediation_days": 30, "applicable_frameworks": ["gdpr"]},
{"id": "F-DP-005", "theme": "data_protection_privacy", "severity": "major", "title": "DPO appointment not formal", "description": "DPO exists but appointment letter not signed by senior management per GDPR Article 37 + 38. Reporting line ambiguous.", "remediation": "Formal appointment letter; clarify reporting line to highest management; publish contact.", "remediation_days": 30, "applicable_frameworks": ["gdpr"]},
{"id": "F-DP-006", "theme": "data_protection_privacy", "severity": "minor", "title": "DSAR identity verification process inconsistent", "description": "DSAR identity verification varies across teams; one DSAR processed without proper identity check.", "remediation": "Standardize verification procedure; train DPO + intake team.", "remediation_days": 60, "applicable_frameworks": ["gdpr"]},
{"id": "F-DP-007", "theme": "data_protection_privacy", "severity": "observation", "title": "Consider privacy-enhancing technologies (PETs)", "description": "Current privacy posture is procedural; PETs (differential privacy, federated learning, secure enclaves) for high-risk processing would reduce exposure.", "remediation": "Pilot PET for one high-risk processing.", "remediation_days": 365, "applicable_frameworks": ["gdpr", "iso_42001"]},
{"id": "F-MR-001", "theme": "management_review", "severity": "critical", "title": "Management review not performed in 18 months", "description": "Management review last documented 18 months ago. Clause 9.3 expects at planned intervals (annual minimum). System effectiveness not formally evaluated.", "remediation": "Schedule + conduct review; document inputs + outputs; calendar future reviews.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485", "soc_2"]},
{"id": "F-MR-002", "theme": "management_review", "severity": "major", "title": "Management review missing AI-specific inputs", "description": "Management review covers ISMS but not AIMS-specific inputs (drift events, incidents, risk-register changes per ISO 42001 Clause 9.3).", "remediation": "Update review template for AIMS inputs; include in next review.", "remediation_days": 60, "applicable_frameworks": ["iso_42001"]},
{"id": "F-MR-003", "theme": "management_review", "severity": "major", "title": "Open action items past due", "description": "Management review action items: 4 of 9 past due > 60 days. Tracking not actively managed.", "remediation": "Reassign owners; escalate stuck items; re-baseline due dates.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485", "soc_2"]},
{"id": "F-MR-004", "theme": "management_review", "severity": "minor", "title": "Review attendance lacks senior leadership", "description": "Review held but CEO + CTO absent; attendance of senior leadership expected per Clause 5.1 + 9.3.", "remediation": "Schedule with leadership in advance; share inputs ahead of meeting.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485", "soc_2"]},
{"id": "F-IA-001", "theme": "internal_audit", "severity": "critical", "title": "No internal audit programme", "description": "Clause 9.2 internal audit programme not documented; audits happen ad-hoc; no rolling 3-year coverage plan.", "remediation": "Design programme; assign auditors; schedule next 12 months minimum.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485", "soc_2", "hipaa"]},
{"id": "F-IA-002", "theme": "internal_audit", "severity": "major", "title": "Auditors audit own work", "description": "Internal auditor for Clause 8.3 audit also owns the lifecycle process being audited. Independence breached.", "remediation": "Reassign auditor; document independence verification per assignment.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485"]},
{"id": "F-IA-003", "theme": "internal_audit", "severity": "major", "title": "Audit findings not tracked to closure", "description": "Audit findings logged but closure verification not consistently performed. 12 findings show 'closed' without evidence of effectiveness.", "remediation": "Verify closure; require evidence; reopen unverified.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "soc_2"]},
{"id": "F-IA-004", "theme": "internal_audit", "severity": "minor", "title": "Audit programme doesn't cover all clauses", "description": "Audit programme covers Clauses 4-7 but not 8-10 in current 3-year cycle.", "remediation": "Update programme; add missing clauses to remaining cycle.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485"]},
{"id": "F-CI-001", "theme": "continual_improvement", "severity": "major", "title": "CAPA without effectiveness verification", "description": "Corrective action plans documented + closed but effectiveness verification missing for 6 of 10 sampled CAPAs.", "remediation": "Add measurable effectiveness verification to template; verify per CAPA; sample-check.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485", "fda_qsr"]},
{"id": "F-CI-002", "theme": "continual_improvement", "severity": "major", "title": "Root cause analysis shallow", "description": "Root cause analysis on CAPAs documented but stops at proximate cause (e.g., 'engineer made mistake'); 5 Whys not applied.", "remediation": "Train CAPA owners on RCA methodology; re-do RCA on recent CAPAs.", "remediation_days": 90, "applicable_frameworks": ["iso_13485", "iso_42001", "iso_27001", "fda_qsr"]},
{"id": "F-CI-003", "theme": "continual_improvement", "severity": "minor", "title": "Trend analysis not performed", "description": "Individual CAPAs handled but trend analysis across CAPAs not performed; missed systemic issues.", "remediation": "Quarterly trend analysis; pattern identification; address systemic causes.", "remediation_days": 90, "applicable_frameworks": ["iso_13485", "iso_27001", "iso_42001", "fda_qsr"]},
{"id": "F-CI-004", "theme": "continual_improvement", "severity": "observation", "title": "Consider integrating CAPA into existing ticketing", "description": "CAPA tracking in separate tool from incident tickets; integration would reduce overhead.", "remediation": "Evaluate ticket-system extensions; pilot.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2"]},
{"id": "F-DC-001", "theme": "documentation_control", "severity": "major", "title": "Obsolete documents accessible", "description": "Old versions of policies and procedures accessible in shared drives without 'obsolete' marking; risk of using superseded content.", "remediation": "Archive obsolete versions; reorganize document drive; reinforce procedure.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_13485", "iso_42001", "fda_qsr"]},
{"id": "F-DC-002", "theme": "documentation_control", "severity": "major", "title": "Document approval workflow bypassed", "description": "Document approval workflow exists but 3 recent policy updates published without documented approval.", "remediation": "Enforce workflow at publication; train owners; audit recent publications.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "iso_13485", "iso_42001", "soc_2"]},
{"id": "F-DC-003", "theme": "documentation_control", "severity": "minor", "title": "Document review cadence not enforced", "description": "Annual review cadence stated but 25% of controlled documents past due > 90 days.", "remediation": "Calendar reviews; track due dates; remind owners.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "iso_13485", "iso_42001"]},
{"id": "F-AIMS-001", "theme": "aims_specific", "severity": "critical", "title": "AI policy missing required commitments", "description": "AI policy commits to lawful use only; missing beneficial purpose, human oversight, and continual improvement. ISO 42001 Clause 5.2 + Annex A.2.2 not satisfied.", "remediation": "Rewrite policy with all 4 commitments; board signoff; publish.", "remediation_days": 60, "applicable_frameworks": ["iso_42001"]},
{"id": "F-AIMS-002", "theme": "aims_specific", "severity": "critical", "title": "AIMS scope omits third-party AI", "description": "AIMS scope statement (Clause 4.3) lists company-built AI systems but omits AI features in SaaS vendors used internally. Scope incomplete.", "remediation": "Update scope; inventory third-party AI; include in AIMS controls.", "remediation_days": 60, "applicable_frameworks": ["iso_42001"]},
{"id": "F-AIMS-003", "theme": "aims_specific", "severity": "critical", "title": "AI system lifecycle skips decommission", "description": "AI lifecycle procedure (A.6) covers design through deployment + operation but lacks decommission phase. ISO 42001 expects full lifecycle.", "remediation": "Define decommission procedure; train owners; document.", "remediation_days": 60, "applicable_frameworks": ["iso_42001"]},
{"id": "F-AIMS-004", "theme": "aims_specific", "severity": "major", "title": "V&V procedure for AI systems undefined", "description": "Annex A.6.2.4 verification + validation procedure not documented; tests exist but acceptance criteria not formalized.", "remediation": "Define V&V procedure; document acceptance criteria per system class; train.", "remediation_days": 90, "applicable_frameworks": ["iso_42001"]},
{"id": "F-AIMS-005", "theme": "aims_specific", "severity": "major", "title": "Impact assessment signed by wrong authority", "description": "AI impact assessments for high-impact systems signed by tech lead; management approval expected per A.5.4.", "remediation": "Define signoff authority by impact tier; re-route assessments; backfill where needed.", "remediation_days": 60, "applicable_frameworks": ["iso_42001"]},
{"id": "F-AIA-001", "theme": "ai_act_specific", "severity": "critical", "title": "Article 5 prohibited practice in production", "description": "AI system performs emotion recognition in workplace setting; Article 5(1)(f) prohibition applies. System cannot remain on EU market.", "remediation": "Disable in EU immediately; evaluate redesign for permitted use cases; document.", "remediation_days": 7, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-AIA-002", "theme": "ai_act_specific", "severity": "critical", "title": "High-risk AI without conformity assessment", "description": "Annex III high-risk AI system on EU market; no Article 43 conformity assessment performed before placement.", "remediation": "Withdraw from market until conformity assessment complete; document Annex IV; CE marking.", "remediation_days": 30, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-AIA-003", "theme": "ai_act_specific", "severity": "major", "title": "Non-EU provider without authorized representative", "description": "Non-EU provider placing AI system on EU market without appointed authorized representative per Article 22.", "remediation": "Appoint EU-established authorized representative; document mandate.", "remediation_days": 60, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-AIA-004", "theme": "ai_act_specific", "severity": "major", "title": "Article 50 transparency not implemented", "description": "Customer-facing chatbot does not disclose AI interaction per Article 50(1).", "remediation": "Add disclosure to UX; A/B test wording.", "remediation_days": 30, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-AIA-005", "theme": "ai_act_specific", "severity": "major", "title": "GPAI without Article 53 technical documentation", "description": "GPAI model provided to downstream integrators; Annex XI technical documentation not maintained.", "remediation": "Develop documentation per Annex XI; publish training-data summary; copyright policy.", "remediation_days": 60, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-13485-001", "theme": "qms_specific", "severity": "critical", "title": "DHF incomplete for commercial device", "description": "Design history file for commercially distributed device lacks design validation evidence per ISO 13485 Clause 7.3.7.", "remediation": "Compile validation evidence; document; if not feasible, withdraw + revalidate.", "remediation_days": 60, "applicable_frameworks": ["iso_13485", "fda_qsr"]},
{"id": "F-13485-002", "theme": "qms_specific", "severity": "critical", "title": "Process validation stale", "description": "Sterilization process not revalidated for 7 years despite supplier changes. ISO 13485 Clause 7.5.6 expects periodic revalidation.", "remediation": "Revalidate; document; calendar future revalidation.", "remediation_days": 90, "applicable_frameworks": ["iso_13485", "fda_qsr"]},
{"id": "F-13485-003", "theme": "qms_specific", "severity": "major", "title": "Risk management file frozen at release", "description": "ISO 14971 risk management file not updated post-launch; post-production information feedback not occurring.", "remediation": "Update RMF with post-production information; commit to periodic review.", "remediation_days": 90, "applicable_frameworks": ["iso_13485", "iso_14971", "eu_mdr_745", "fda_qsr"]},
{"id": "F-13485-004", "theme": "qms_specific", "severity": "major", "title": "PMCF plan exists but not executed", "description": "Post-market clinical follow-up plan documented per EU MDR Annex XIV Part B; execution data lacking after 12 months.", "remediation": "Execute per plan; document; report to notified body if outside plan.", "remediation_days": 90, "applicable_frameworks": ["iso_13485", "eu_mdr_745"]},
{"id": "F-FDA-001", "theme": "fda_specific", "severity": "critical", "title": "MDR-reportable event not reported", "description": "Serious adverse event reportable per 21 CFR 803.50 not reported within 30 days. FDA enforcement exposure.", "remediation": "File MDR immediately with delay rationale; review complaint trending; CAPA.", "remediation_days": 7, "applicable_frameworks": ["fda_qsr"]},
{"id": "F-FDA-002", "theme": "fda_specific", "severity": "major", "title": "Complaint files incomplete", "description": "Complaint log per 21 CFR 820.198 missing investigation closure for 8 of 30 sampled complaints.", "remediation": "Investigate + close; train complaint handlers.", "remediation_days": 60, "applicable_frameworks": ["fda_qsr"]},
{"id": "F-FDA-003", "theme": "fda_specific", "severity": "major", "title": "Form 483 open observations past response window", "description": "Form 483 received 6 months ago; 2 of 5 observations lack documented response within 15-working-day window.", "remediation": "Respond immediately; document corrective action; escalate to legal counsel.", "remediation_days": 14, "applicable_frameworks": ["fda_qsr"]},
{"id": "F-FDA-004", "theme": "fda_specific", "severity": "minor", "title": "Labeling review evidence gaps", "description": "Labeling per 21 CFR 801 reviewed at launch but no documented re-review for label changes in past 18 months.", "remediation": "Audit labels; document review per change.", "remediation_days": 60, "applicable_frameworks": ["fda_qsr"]},
{"id": "F-HIPAA-001", "theme": "hipaa_specific", "severity": "critical", "title": "PHI breach not assessed under Breach Notification Rule", "description": "PHI exposure event 4 months ago; risk-of-compromise assessment per §164.402 not documented. Breach notification potentially required + missed.", "remediation": "Conduct retroactive assessment; if breach, notify per §164.404 + §164.406; document.", "remediation_days": 14, "applicable_frameworks": ["hipaa"]},
{"id": "F-HIPAA-002", "theme": "hipaa_specific", "severity": "critical", "title": "Security Risk Analysis not performed", "description": "HIPAA Security Rule §164.308(a)(1)(ii)(A) risk analysis not documented in past 24 months despite material system changes.", "remediation": "Conduct + document analysis; address top risks; calendar annual review.", "remediation_days": 60, "applicable_frameworks": ["hipaa"]},
{"id": "F-HIPAA-003", "theme": "hipaa_specific", "severity": "major", "title": "Encryption addressable spec not formally evaluated", "description": "HIPAA encryption is 'addressable'; organization not encrypting PHI at rest in one data store; no documented analysis of why.", "remediation": "Document analysis; if not encrypted, implement alternative protective measure or encrypt.", "remediation_days": 90, "applicable_frameworks": ["hipaa"]},
{"id": "F-HIPAA-004", "theme": "hipaa_specific", "severity": "major", "title": "Workforce sanctions policy not enforced", "description": "§164.308(a)(1)(ii)(C) sanctions policy documented but no recorded sanctions despite repeat policy violations.", "remediation": "Apply sanctions per policy; document; refresh training.", "remediation_days": 60, "applicable_frameworks": ["hipaa"]},
{"id": "F-NIS2-001", "theme": "nis2_specific", "severity": "critical", "title": "Incident notification 24h early warning missed", "description": "NIS2 Article 23 24-hour early warning to competent authority + CSIRT not provided after recent significant incident.", "remediation": "File retrospectively; document delay rationale; engage authority; update IR procedure.", "remediation_days": 7, "applicable_frameworks": ["nis2"]},
{"id": "F-NIS2-002", "theme": "nis2_specific", "severity": "critical", "title": "Management body not approving cybersecurity measures", "description": "NIS2 Article 20 requires management bodies to approve cybersecurity risk-management measures + oversee implementation. Approval missing from board minutes.", "remediation": "Add to board agenda; document approval; ongoing oversight cadence.", "remediation_days": 60, "applicable_frameworks": ["nis2"]},
{"id": "F-NIS2-003", "theme": "nis2_specific", "severity": "major", "title": "10 minimum cybersecurity measures incomplete", "description": "NIS2 Article 21(2)(a)-(j) 10 minimum measures: 2 not documented (policies on cryptography, basic cyber hygiene).", "remediation": "Document missing policies; verify implementation; submit registration update.", "remediation_days": 90, "applicable_frameworks": ["nis2"]},
{"id": "F-CSF-001", "theme": "csf_specific", "severity": "major", "title": "NIST CSF profile not defined", "description": "Organization adopts NIST CSF 2.0 conceptually but no documented profile (current + target state) per CSF practice.", "remediation": "Develop profile; identify gaps; roadmap.", "remediation_days": 90, "applicable_frameworks": ["nist_csf"]},
{"id": "F-CSF-002", "theme": "csf_specific", "severity": "minor", "title": "Recover function under-developed", "description": "CSF GOVERN + IDENTIFY + PROTECT + DETECT + RESPOND well-developed; RECOVER function lacks documented recovery planning.", "remediation": "Develop recovery planning + communications procedures.", "remediation_days": 90, "applicable_frameworks": ["nist_csf", "iso_27001"]},
{"id": "F-AC-011", "theme": "access_control", "severity": "major", "title": "Service accounts without rotation", "description": "Service-account credentials shared across systems; no rotation in past 24 months.", "remediation": "Rotate; introduce secrets-management tooling; document.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa", "nis2"]},
{"id": "F-AC-012", "theme": "access_control", "severity": "major", "title": "Privileged access logs not reviewed", "description": "Privileged user activity logs collected but no periodic review for anomalous behavior.", "remediation": "Define review cadence; assign reviewer; SIEM alerts for high-risk patterns.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nist_csf"]},
{"id": "F-AC-013", "theme": "access_control", "severity": "minor", "title": "Break-glass account not monitored", "description": "Emergency break-glass account exists but its usage not monitored; could be used without trace.", "remediation": "Alert on break-glass usage; quarterly review.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa"]},
{"id": "F-AI-006", "theme": "asset_inventory", "severity": "major", "title": "Personal device access not inventoried", "description": "BYOD devices accessing corporate data not in asset inventory; mobile device management (MDM) coverage incomplete.", "remediation": "Inventory BYOD; require MDM enrollment; document policy.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nist_csf"]},
{"id": "F-AI-007", "theme": "asset_inventory", "severity": "major", "title": "Shadow IT discovered during audit", "description": "5 SaaS tools in use by teams without procurement / security review; some handle personal data.", "remediation": "Bring shadow IT under management or sunset; revise procurement gate.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "gdpr", "hipaa", "nist_csf"]},
{"id": "F-AI-008", "theme": "asset_inventory", "severity": "observation", "title": "Inventory not integrated with CMDB", "description": "Asset inventory in spreadsheet; lacks integration with operational CMDB. Drift inevitable.", "remediation": "Integrate via API or migrate to CMDB-as-source-of-truth.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-RM-010", "theme": "risk_management", "severity": "major", "title": "AI bias risk not formally identified", "description": "AI risk register lacks systematic identification of bias risks across protected demographic categories.", "remediation": "Apply ISO 23894 risk identification methodology; bias testing per category; document.", "remediation_days": 90, "applicable_frameworks": ["iso_42001", "eu_ai_act"]},
{"id": "F-RM-011", "theme": "risk_management", "severity": "minor", "title": "Risk treatment costs not estimated", "description": "Risk treatment plans don't estimate implementation cost; cost/benefit analysis missing.", "remediation": "Add cost estimate field; quarterly review.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "iso_42001", "nist_csf"]},
{"id": "F-SM-011", "theme": "supplier_management", "severity": "major", "title": "Critical vendor SOC 2 expired", "description": "Critical vendor's SOC 2 Type II report on file is 18 months old; current period not yet collected.", "remediation": "Request current report; if vendor delayed, document compensating evidence.", "remediation_days": 60, "applicable_frameworks": ["soc_2", "iso_27001"]},
{"id": "F-SM-012", "theme": "supplier_management", "severity": "minor", "title": "Vendor contact lists stale", "description": "Vendor security contact information stale; recent contact attempts bounced.", "remediation": "Refresh contact lists; verify quarterly.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "gdpr"]},
{"id": "F-IR-011", "theme": "incident_response", "severity": "major", "title": "Forensic data preservation not standard", "description": "Recent incidents lack forensic preservation of affected systems; impedes investigation.", "remediation": "Document forensic preservation procedure; train IR team.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nist_csf"]},
{"id": "F-IR-012", "theme": "incident_response", "severity": "minor", "title": "External communications template missing", "description": "External communications for incidents drafted ad-hoc; no pre-approved templates.", "remediation": "Develop templates; legal + comms review; approve.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "gdpr"]},
{"id": "F-ML-008", "theme": "monitoring_logging", "severity": "major", "title": "Database query logging disabled", "description": "Production database query logging disabled for performance reasons; can't audit who queried what.", "remediation": "Enable query logging for sensitive tables; size storage; document trade-offs.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "gdpr", "nist_csf"]},
{"id": "F-ML-009", "theme": "monitoring_logging", "severity": "minor", "title": "Log timestamps not in standard timezone", "description": "Logs across systems use mix of local timezones + UTC; correlation difficult.", "remediation": "Standardize on UTC; document; backfill where feasible.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-CM-006", "theme": "change_management", "severity": "minor", "title": "Configuration drift not detected", "description": "Production configuration drift from documented baseline; no detection mechanism.", "remediation": "Deploy infrastructure-as-code drift detection; alert on deviations.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-CM-007", "theme": "change_management", "severity": "observation", "title": "Consider GitOps for change discipline", "description": "Some changes still applied imperatively; GitOps would enforce change-via-PR discipline.", "remediation": "Pilot GitOps for one infrastructure layer.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-BC-006", "theme": "business_continuity", "severity": "major", "title": "Single region deployment without DR plan", "description": "Production deployment in single AWS region; no documented multi-region or cross-region DR plan.", "remediation": "Define DR plan (cross-region replicas, runbooks); test.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2"]},
{"id": "F-BC-007", "theme": "business_continuity", "severity": "minor", "title": "Communications plan missing for major outage", "description": "BCP covers technical recovery but lacks customer + employee communication plan for major outage.", "remediation": "Develop communications plan; pre-approved templates; cascade.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "nis2"]},
{"id": "F-CT-006", "theme": "competence_training", "severity": "minor", "title": "Onboarding security training not within 30 days", "description": "Some new hires complete security training 60+ days after start; expected within 30 days.", "remediation": "Calendar reminders; manager accountability; track completion timeline.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "hipaa", "nist_csf"]},
{"id": "F-CT-007", "theme": "competence_training", "severity": "observation", "title": "Phishing simulation results trending up", "description": "Phishing simulation click-rate increasing; training content may not be effective.", "remediation": "Refresh training content; targeted training for repeat clickers.", "remediation_days": 120, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "nis2", "hipaa"]},
{"id": "F-DG-008", "theme": "data_governance", "severity": "major", "title": "Data classification policy applied unevenly", "description": "Data classification policy applied to engineering data stores but not marketing tools containing customer data.", "remediation": "Extend classification; train marketing.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "soc_2", "gdpr", "hipaa"]},
{"id": "F-DG-009", "theme": "data_governance", "severity": "minor", "title": "Pseudonymization not consistently applied", "description": "Pseudonymization documented for some pipelines; not consistently applied to analytics datasets containing personal data.", "remediation": "Audit analytics datasets; pseudonymize where lawful basis is analytics.", "remediation_days": 90, "applicable_frameworks": ["gdpr", "iso_42001"]},
{"id": "F-CR-006", "theme": "cryptography", "severity": "major", "title": "Keys stored alongside data", "description": "Encryption keys stored in same cloud account / region as encrypted data; compromise of one yields the other.", "remediation": "Move keys to dedicated KMS account; restrict access.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa"]},
{"id": "F-CR-007", "theme": "cryptography", "severity": "minor", "title": "Certificate expiration monitoring incomplete", "description": "Certificate expiration alerts configured for some endpoints; internal certificates lack monitoring.", "remediation": "Extend monitoring; centralize certificate inventory.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-SD-006", "theme": "secure_sdlc", "severity": "major", "title": "Secrets in source control", "description": "Code review uncovered API keys + DB credentials committed to git history.", "remediation": "Rotate exposed secrets; remove from history; install pre-commit hooks; train.", "remediation_days": 30, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa"]},
{"id": "F-SD-007", "theme": "secure_sdlc", "severity": "minor", "title": "Pull-request templates lack security checklist", "description": "PR templates exist but don't prompt security considerations (auth, input validation, secrets).", "remediation": "Add security checklist to template; train.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf"]},
{"id": "F-VM-006", "theme": "vulnerability_mgmt", "severity": "major", "title": "Penetration test recommendations untracked", "description": "Annual penetration test completed; 12 findings; tracking + closure of remediation not centralized.", "remediation": "Centralize tracking; assign owners; verify closure.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "soc_2", "nist_csf", "hipaa"]},
{"id": "F-VM-007", "theme": "vulnerability_mgmt", "severity": "observation", "title": "Consider bug bounty programme", "description": "External vulnerability discovery limited to annual pentest; bug bounty would broaden coverage.", "remediation": "Evaluate bug bounty platforms; pilot.", "remediation_days": 180, "applicable_frameworks": ["iso_27001", "nist_csf"]},
{"id": "F-PS-005", "theme": "physical_security", "severity": "minor", "title": "Clean desk policy not enforced", "description": "Clean desk policy documented but walkthrough found sensitive printouts on unattended desks.", "remediation": "Reinforce policy; periodic walkthroughs; train.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "hipaa"]},
{"id": "F-PS-006", "theme": "physical_security", "severity": "observation", "title": "Hardware disposal evidence incomplete", "description": "Hardware disposal documented for laptops; lacks evidence of certified destruction for storage media.", "remediation": "Use certified destruction service; collect certificates.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "hipaa", "nist_csf"]},
{"id": "F-DP-008", "theme": "data_protection_privacy", "severity": "major", "title": "DSAR response > 30 days", "description": "12 of 50 DSARs in past quarter responded after Article 12(3) 1-month SLA; no extension communicated.", "remediation": "Investigate process bottlenecks; resource appropriately; communicate extensions where needed.", "remediation_days": 60, "applicable_frameworks": ["gdpr"]},
{"id": "F-DP-009", "theme": "data_protection_privacy", "severity": "minor", "title": "Privacy notice version history missing", "description": "Privacy notice updated multiple times; no version archive; cannot demonstrate which notice was active when.", "remediation": "Archive past versions with date stamps.", "remediation_days": 60, "applicable_frameworks": ["gdpr"]},
{"id": "F-DP-010", "theme": "data_protection_privacy", "severity": "minor", "title": "Article 22 automated decisions not flagged", "description": "Automated decision-making (Article 22) used in credit decisions; data subjects not informed; human review not offered.", "remediation": "Add transparency; offer human review; document procedure.", "remediation_days": 60, "applicable_frameworks": ["gdpr", "eu_ai_act"]},
{"id": "F-DC-004", "theme": "documentation_control", "severity": "observation", "title": "Consider read-only published documents", "description": "Controlled documents stored as editable Google Docs; risk of unauthorized edit. Read-only PDF publishing would be stronger control.", "remediation": "Publish read-only PDFs; restrict editing to authors.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "iso_13485", "iso_42001"]},
{"id": "F-IA-005", "theme": "internal_audit", "severity": "minor", "title": "Audit reports lack standard format", "description": "Audit reports vary in format across auditors; difficult to compare or trend.", "remediation": "Define standard report template.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "soc_2", "iso_13485"]},
{"id": "F-AIMS-006", "theme": "aims_specific", "severity": "major", "title": "AI model card missing", "description": "Production AI system lacks model card per Annex A.6.2.7. Documentation per Mitchell et al. (2019) pattern not produced.", "remediation": "Develop model card; publish internally; commit to update with retraining.", "remediation_days": 60, "applicable_frameworks": ["iso_42001"]},
{"id": "F-AIMS-007", "theme": "aims_specific", "severity": "minor", "title": "Datasheet for datasets not produced", "description": "Training datasets lack datasheet per Gebru et al. (2021) pattern; not satisfying Annex A.7.4 fully.", "remediation": "Develop datasheets per dataset; document provenance + composition + intended use.", "remediation_days": 90, "applicable_frameworks": ["iso_42001"]},
{"id": "F-AIA-006", "theme": "ai_act_specific", "severity": "major", "title": "Article 27 FRIA missing for public-sector deployer", "description": "Public-sector body deploying high-risk AI; Fundamental Rights Impact Assessment per Article 27 not performed.", "remediation": "Conduct FRIA; document; consult DPA where required.", "remediation_days": 60, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-AIA-007", "theme": "ai_act_specific", "severity": "minor", "title": "EU database registration pending", "description": "High-risk Annex III system not yet registered in EU database per Article 71.", "remediation": "Register; document.", "remediation_days": 30, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-AIA-008", "theme": "ai_act_specific", "severity": "critical", "title": "Substantial modification turns deployer into provider", "description": "Deployer substantially modified high-risk AI system; now operates as provider per Article 25(1) but did not assume provider obligations.", "remediation": "Document role change; assume provider obligations; conformity assessment.", "remediation_days": 30, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-13485-005", "theme": "qms_specific", "severity": "major", "title": "Design transfer evidence missing", "description": "Design transfer per Clause 7.3.8 not formally documented for recent product. Manufacturing operates with insufficient design records.", "remediation": "Compile transfer evidence; document training; verify capability.", "remediation_days": 60, "applicable_frameworks": ["iso_13485", "fda_qsr"]},
{"id": "F-13485-006", "theme": "qms_specific", "severity": "observation", "title": "Consider digital quality management system", "description": "QMS run on shared drives; eQMS would improve traceability + audit-readiness.", "remediation": "Evaluate eQMS vendors; pilot.", "remediation_days": 180, "applicable_frameworks": ["iso_13485", "fda_qsr"]},
{"id": "F-FDA-005", "theme": "fda_specific", "severity": "major", "title": "UDI compliance gaps", "description": "Some devices commercially distributed lack UDI labeling per 21 CFR 830.", "remediation": "Audit + label; submit to GUDID; document.", "remediation_days": 90, "applicable_frameworks": ["fda_qsr"]},
{"id": "F-FDA-006", "theme": "fda_specific", "severity": "observation", "title": "Pre-submission strategy could leverage Q-sub", "description": "Product strategy proceeds toward 510(k) without leveraging FDA Q-Submission programme.", "remediation": "Consider Q-sub for novel aspects.", "remediation_days": 180, "applicable_frameworks": ["fda_qsr"]},
{"id": "F-HIPAA-005", "theme": "hipaa_specific", "severity": "major", "title": "Workforce member access not minimum-necessary", "description": "Workforce access provisioned at role level rather than minimum-necessary per §164.502(b). Some members access PHI beyond their need.", "remediation": "Audit + tighten access; document minimum-necessary determination.", "remediation_days": 90, "applicable_frameworks": ["hipaa"]},
{"id": "F-HIPAA-006", "theme": "hipaa_specific", "severity": "minor", "title": "Notice of privacy practices outdated", "description": "Notice of privacy practices per §164.520 last updated 2 years ago; substantive policy changes not reflected.", "remediation": "Update notice; redistribute per requirement; document.", "remediation_days": 60, "applicable_frameworks": ["hipaa"]},
{"id": "F-NIS2-004", "theme": "nis2_specific", "severity": "major", "title": "Registration with competent authority pending", "description": "Organization meets NIS2 essential entity criteria but has not registered with national competent authority per Article 24.", "remediation": "Submit registration; document.", "remediation_days": 30, "applicable_frameworks": ["nis2"]},
{"id": "F-NIS2-005", "theme": "nis2_specific", "severity": "minor", "title": "Supply-chain security measures not documented", "description": "NIS2 Article 21(2)(d) supply-chain security measures not separately documented from generic supplier-management.", "remediation": "Document NIS2-specific supply-chain measures.", "remediation_days": 60, "applicable_frameworks": ["nis2"]},
{"id": "F-CSF-003", "theme": "csf_specific", "severity": "minor", "title": "CSF tiers not assigned", "description": "NIST CSF 2.0 implementation tiers (Partial / Risk Informed / Repeatable / Adaptive) not assigned per function.", "remediation": "Self-assess tiers; document; target tier.", "remediation_days": 90, "applicable_frameworks": ["nist_csf"]},
{"id": "F-MDR-001", "theme": "mdr_specific", "severity": "critical", "title": "EU MDR technical documentation gap", "description": "Technical documentation per Annex II/III lacks recent clinical-evaluation update; notified body audit imminent.", "remediation": "Update documentation immediately; engage notified body.", "remediation_days": 30, "applicable_frameworks": ["eu_mdr_745"]},
{"id": "F-MDR-002", "theme": "mdr_specific", "severity": "major", "title": "Person Responsible for Regulatory Compliance not appointed", "description": "EU MDR Article 15 PRRC role not formally appointed for the EU operations.", "remediation": "Appoint PRRC meeting Article 15(1)-(2) qualifications; document.", "remediation_days": 30, "applicable_frameworks": ["eu_mdr_745"]},
{"id": "F-MDR-003", "theme": "mdr_specific", "severity": "minor", "title": "PMCF reports lag schedule", "description": "Post-Market Clinical Follow-up reports not produced per agreed schedule.", "remediation": "Catch up; rebaseline schedule.", "remediation_days": 90, "applicable_frameworks": ["eu_mdr_745"]},
{"id": "F-14971-001", "theme": "risk_management_medical", "severity": "major", "title": "Risk management plan not updated for software change", "description": "ISO 14971 risk management plan + risk file not updated after material software change.", "remediation": "Update RMF; re-evaluate risks; document.", "remediation_days": 60, "applicable_frameworks": ["iso_14971", "iso_13485", "eu_mdr_745"]},
{"id": "F-14971-002", "theme": "risk_management_medical", "severity": "minor", "title": "Residual risk evaluation lacks acceptability criteria", "description": "Residual risk evaluated but acceptability criteria per ISO 14971 §7 not formally established.", "remediation": "Define acceptability criteria; document.", "remediation_days": 90, "applicable_frameworks": ["iso_14971", "iso_13485"]},
{"id": "F-MDR-004", "theme": "mdr_specific", "severity": "major", "title": "EUDAMED registration incomplete", "description": "EU MDR EUDAMED registration of device, manufacturer, or UDI elements incomplete despite mandatory data submission requirements.", "remediation": "Complete required EUDAMED modules; track future module activations.", "remediation_days": 60, "applicable_frameworks": ["eu_mdr_745"]},
{"id": "F-MDR-005", "theme": "mdr_specific", "severity": "minor", "title": "Vigilance reporting log incomplete", "description": "EU MDR vigilance reporting log per Article 87 has 3 entries past 15-day reporting timeline.", "remediation": "Investigate root cause; tighten internal SLA; train.", "remediation_days": 60, "applicable_frameworks": ["eu_mdr_745"]},
{"id": "F-14971-003", "theme": "risk_management_medical", "severity": "major", "title": "Production + post-production information feedback weak", "description": "ISO 14971 §9 requires production + post-production information be collected + analysed; current process only acts on customer complaints, missing field data + service trends.", "remediation": "Expand information sources; document process; integrate with PMS.", "remediation_days": 90, "applicable_frameworks": ["iso_14971", "iso_13485", "eu_mdr_745"]},
{"id": "F-AIA-009", "theme": "ai_act_specific", "severity": "major", "title": "Deepfake content not marked AI-generated", "description": "Generative AI feature produces audio/video without machine-readable AI-generated marking per Article 50(2).", "remediation": "Implement watermarking; document.", "remediation_days": 60, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-AIA-010", "theme": "ai_act_specific", "severity": "minor", "title": "Instructions for use missing operational risks section", "description": "Article 13 instructions for use provided to deployers but do not adequately describe foreseeable operational risks.", "remediation": "Update IFU with risks + mitigations; train downstream.", "remediation_days": 60, "applicable_frameworks": ["eu_ai_act"]},
{"id": "F-FDA-007", "theme": "fda_specific", "severity": "major", "title": "Cybersecurity for connected device not addressed in 510(k)", "description": "Connected device 510(k) submission lacks cybersecurity content per FDA Cybersecurity Guidance (Sep 2023); FDA refused acceptance.", "remediation": "Develop cybersecurity content per guidance; resubmit.", "remediation_days": 90, "applicable_frameworks": ["fda_qsr"]},
{"id": "F-FDA-008", "theme": "fda_specific", "severity": "minor", "title": "510(k) summary lacks comparative data", "description": "510(k) summary per 21 CFR 807.92 lacks substantive comparison to predicate device.", "remediation": "Add comparative data; resubmit if FDA requests.", "remediation_days": 60, "applicable_frameworks": ["fda_qsr"]},
{"id": "F-HIPAA-007", "theme": "hipaa_specific", "severity": "minor", "title": "Workforce member termination workflow missing PHI access revocation", "description": "Termination workflow revokes general access but doesn't specifically address PHI access systems; 2 terminated members retained EHR access > 2 days.", "remediation": "Add PHI-specific revocation step; verify.", "remediation_days": 30, "applicable_frameworks": ["hipaa", "iso_27001"]},
{"id": "F-MR-005", "theme": "management_review", "severity": "minor", "title": "Management review inputs not pre-distributed", "description": "Management review held but inputs distributed only at meeting; senior leadership cannot prepare in advance.", "remediation": "Pre-distribute inputs 1 week in advance.", "remediation_days": 60, "applicable_frameworks": ["iso_27001", "iso_42001", "iso_13485", "soc_2"]},
{"id": "F-IA-006", "theme": "internal_audit", "severity": "observation", "title": "Audit programme could integrate cross-framework findings", "description": "Audits performed per framework but cross-framework finding impact not systematically tracked; missed reuse opportunity.", "remediation": "Use compliance-os cross_framework_mapper output to tag findings.", "remediation_days": 90, "applicable_frameworks": ["iso_27001", "iso_42001", "soc_2", "iso_13485"]}
]
}
FILE:references/audit_simulation_methodology.md
# Audit Simulation Methodology — ISO 19011 + IIA IPPF + AICPA AT-C
This reference answers exactly one decision: **what does a realistic internal audit look like, and how do we generate a mock audit that prepares the team without breaking trust?**
Pair with `scripts/audit_simulator.py` for the deterministic mock audit generator.
## Why Simulate Audits?
External certification audits are high-stakes events. A team that has never been audited internally before its first stage 2 ISO certification audit will struggle even if every artefact is in place — interview cadence, document-pull SLAs, walk-through pacing are operational muscles built only by practice.
Mock audits provide:
- Operational practice (auditees experience the rhythm of an interview)
- Auditor-side practice (internal auditors practice their methodology before high-stakes certification audits)
- Discovery of gaps before they become findings
- Calibration of effort (how long does evidence assembly actually take?)
- Cross-training (auditors from one team learn another team's controls)
## Audit Standards That Govern Simulation
**ISO/IEC 19011:2018** — Guidelines for auditing management systems. Defines:
- Audit principles: integrity, fair presentation, due professional care, confidentiality, independence, evidence-based approach, risk-based approach
- Auditor competence (Clause 7)
- Audit process: initiating → preparing → conducting → reporting (Clauses 5–6)
**IIA International Professional Practices Framework (IPPF)** — internal-audit-specific:
- IPPF Standards 1000-1322 — Attribute Standards (purpose, independence, proficiency, due professional care, quality assurance)
- IPPF Standards 2000-2600 — Performance Standards (engagement planning through monitoring)
- Severity grading approach (rated finding scale)
**AICPA AT-C 105 + AU-C 240** — SOC 2 audit context: trust services criteria + auditor's responsibility framework.
## The Mock Audit Workflow
Compliance OS `audit_simulator.py` deterministically generates one stage of a mock audit. The full simulation lifecycle:
```
1. SCOPE → define framework + controls in scope + auditee team
2. PREPARE → audit_simulator.py outputs: findings + interview questions + document-review requests
3. CONDUCT → simulated interview + document review (1-2 hours per control)
4. REPORT → finding write-up + severity classification + corrective action assignment
5. CLOSE → corrective action tracking through CAPA
```
## Finding Severity Distribution (the IIA expectation)
A healthy compliance program produces audits with this distribution:
| Severity | Healthy proportion | What it indicates |
|---|---|---|
| **Critical (major nonconformity)** | ≤ 15% | Blocks certification; requires major corrective action |
| **Major** | 15–25% | Important gaps requiring 30-day corrective action plans |
| **Minor** | 20–30% | Operational gaps requiring corrective action timeline |
| **Observation / OFI** | ≥ 40% | Improvement opportunities; no required action |
**Why this shape?** If 80% of findings are critical, either the audit was destructive (auditee not given fair chance to demonstrate compliance) or the program is genuinely failing. If 80% of findings are observations, the audit was too superficial. The compliance OS audit simulator enforces this shape by deterministic severity rotation.
A first audit (year 1) will skew higher to critical/major; a mature program (year 3+) skews to observations.
## Number of Findings Per Audit
ISO 19011 Clause 6 typical audit depth:
- Small scope (5 controls, 1 day): 5–10 findings
- Medium scope (10–15 controls, 3–5 days): 10–20 findings
- Full system audit (all clauses, 1–2 weeks): 25–50 findings
The simulator targets 8–15 findings per audit (medium scope) as the default.
## Interview Question Quality
Auditor questions follow the **walk-through pattern**:
1. **Open** — "Walk me through how this control is implemented day-to-day."
2. **Sample** — "Show me a specific example from the last 30 days."
3. **Drill** — "What happens if [edge case]?"
4. **Verify** — "Where is this documented?"
Each control gets 3–5 questions following this pattern. The simulator's `interview_questions()` function provides theme-specific questions per the IIA performance standards.
## Document-Review Requests
Per ISO 19011, the auditor reviews:
- The procedure (the "what should happen")
- The records (the "what actually happened")
- The evidence of management oversight (the "did anyone check?")
A document-review request typically asks for all three. The simulator's `document_requests()` function generates the request list per theme.
## Auditor Independence Test
Clause 9.2 of ISO management-system standards requires auditor independence. The simulator does NOT enforce auditor assignment (that's `aims_audit_scheduler.py` for ISO 42001 or `isms_audit_scheduler.py` for ISO 27001) but the workflow assumes an independent auditor.
**Independence rules:**
- Auditor cannot audit their own work
- Auditor reports to a different chain of command than the auditee
- For small organizations, rotating auditors between teams + occasional external auditor satisfies independence
## Finding Categories (the taxonomy)
The simulator uses 5 finding themes mapped to common control families:
| Theme | Maps to control families |
|---|---|
| `access_control` | ISO 27001 A.5.15 / A.8.2 / A.8.3; SOC 2 CC6.1-6.3; ISO 42001 A.4.4 |
| `logging_monitoring` | ISO 27001 A.8.15 / A.8.16; SOC 2 CC7.1-7.2; ISO 42001 A.9.3 / A.9.4 |
| `change_management` | ISO 27001 A.8.32; SOC 2 CC8.1; ISO 42001 A.6.2.5 |
| `supplier_mgmt` | ISO 27001 A.5.19-A.5.22; SOC 2 CC9.2; ISO 42001 A.10.2; GDPR Art. 28 |
| `incident_response` | ISO 27001 A.5.24-27, A.6.8; SOC 2 CC7.3-7.5; ISO 42001 A.8.4; EU AI Act Art. 73; GDPR Art. 33-34 |
This taxonomy covers the highest-leverage controls across the 9 supported frameworks. Adding new themes is a matter of extending `FINDING_TEMPLATES` + `CONTROL_TO_THEME` mappings.
## Anti-Patterns in Audit Simulation
1. **Auditing for trapping vs auditing for evidence.** Mock audits aim to surface gaps, not embarrass the auditee. If team morale drops after the mock, the audit was structured wrong.
2. **Skipping the "obvious" controls.** Critical findings often hide in mundane controls (e.g., terminated employee with retained access). Simulator deliberately includes prosaic theme rotation.
3. **No prior-year follow-up.** The simulator's `prior_year_findings_open` parameter forces the first finding to be a follow-up. Real audits always follow up on prior open findings (ISO 19011 Clause 6.3).
4. **One severity-skewed audit.** Distribution rule guards against this; if all findings are critical or all are observations, recalibrate the audit scope or methodology.
## When This Reference Doesn't Help
- **Specific industry-vertical audit requirements.** Use sectoral skills (financial, healthcare).
- **Auditor competence + certification.** See ISACA CISA, IRCA Lead Auditor courses.
- **Audit report-writing detail.** See ISO 19011 Clause 6.5 + IIA performance standards 2410–2440.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 19011:2018** — Guidelines for auditing management systems (the canonical methodology)
- **IIA International Professional Practices Framework (IPPF)** — Attribute Standards 1000-1322 + Performance Standards 2000-2600
- **AICPA AT-C 105** — Trust Services Criteria attestation engagement
- **AICPA AU-C 240** — Auditor's responsibilities relating to fraud (financial audit, conceptually applied)
- **ISACA CISA Review Manual** (27th ed., 2024) — IS audit practitioner methodology
- **ASQ Certified Quality Auditor (CQA) Body of Knowledge** — quality audit methodology
- **NIST SP 800-53A Rev 5** — Assessing Security and Privacy Controls (assessment procedures for each control)
- **ISO/IEC 17021-1:2015** — Conformity assessment requirements for bodies providing audit and certification
- **IRCA (International Register of Certificated Auditors)** — Lead auditor certification programme materials
- **The Open Group** — Open FAIR (Factor Analysis of Information Risk) for risk-based audit prioritization
FILE:references/compliance_os_pattern.md
# Compliance OS — The Meta-Framework Pattern
This reference answers exactly one decision: **when do we orchestrate frameworks vs run them separately, and what does the meta-framework architecture look like?**
## The Problem Compliance OS Solves
Most growing companies hit a wall: 2–3 compliance frameworks operating in parallel, each with its own tooling, its own audit calendar, its own evidence requirements, its own internal owner. The result:
- **Duplicate evidence collection** — access-review records assembled 3 times for ISO 27001, SOC 2, and ISO 42001 audits
- **Conflicting audit calendars** — surveillance audits stack in the same week with insufficient auditor capacity
- **Fragmented management review** — each framework wants its own management review, taking 5x the executive time
- **Inconsistent control taxonomies** — "access control" means slightly different things across SOC 2 and ISO 27001 Annex A and ISO 42001 Annex A
- **Unowned cross-framework gaps** — controls in framework A but not B fall to ad-hoc ownership
- **Evidence freshness mismatch** — ISO 27001 wants 12-month log retention, GDPR can want longer, leading to either over-retention or compliance gaps
Compliance OS is the orchestration layer that sits **above** per-framework skills and consolidates the cross-framework view.
## The Four Operations
```
[ Company Profile JSON ]
│
v
╔═══════════════════════╗
║ 1. CONFIGURE ║ framework_selector.py
║ "Which apply?" ║
╚═══════════════════════╝
│
v
╔═══════════════════════╗
║ 2. MAP ║ cross_framework_mapper.py
║ "What overlaps?" ║
╚═══════════════════════╝
│
v
╔═══════════════════════╗
║ 3. SIMULATE ║ audit_simulator.py
║ "What audit looks ║
║ like to fail?" ║
╚═══════════════════════╝
│
v
╔═══════════════════════╗
║ 4. CONSOLIDATE ║ evidence_pool_generator.py
║ "Where's the evidence║
║ + what reuses?" ║
╚═══════════════════════╝
│
v
[ Multi-framework plan ]
```
Each operation is a stdlib Python tool with deterministic logic — no LLM calls, no hidden state.
## When to Use Compliance OS
| Situation | Use compliance-os? |
|---|---|
| Single framework only (e.g., just SOC 2) | No — the per-framework skill is sufficient |
| 2+ frameworks operating in parallel | Yes |
| Adding a new framework to existing program | Yes — for cross-framework reuse mapping |
| Planning annual audit calendar across multiple certifications | Yes |
| Onboarding a new AI system that triggers ISO 42001 + EU AI Act + GDPR | Yes |
| Acquiring a company with different compliance posture | Yes — for gap mapping post-acquisition |
| Internal-audit-only program (no external certification) | Yes if multi-framework; No if single |
## What Compliance OS Is NOT
- **NOT a per-framework deep-dive skill.** Per-framework skills (`ra-qm-team/skills/iso42001-specialist/`, etc.) do the operational work. Compliance OS orchestrates them.
- **NOT a GRC platform replacement.** GRC platforms (Drata, Vanta, OneTrust, Hyperproof, etc.) are tools that operationalize what compliance OS describes — they're complementary. Compliance OS gives the conceptual map; GRC tools store the evidence.
- **NOT a binding legal opinion.** Cross-framework mappings reflect published guidance from ISO, AICPA, NIST, IIA, EDPB. Novel cross-walks need outside counsel.
- **NOT a certification body.** Certification audits are performed by accredited bodies. Compliance OS prepares for them.
## Roles and Ownership
A multi-framework compliance program typically has these roles. Compliance OS does not replace them — it gives them a shared mental model.
| Role | Owns |
|---|---|
| **Compliance officer** | The meta-program; framework selector; cross-framework mapper; consolidated evidence pool |
| **CISO** | ISO 27001 + SOC 2 + cybersecurity slices of ISO 42001 + GDPR Article 32 |
| **DPO** | GDPR; privacy slice of ISO 42001 (A.7.6); EU AI Act Article 27 FRIA where applicable |
| **AIMS lead** | ISO 42001; AI-specific slice of EU AI Act Article 17 QMS |
| **QMS lead** | ISO 13485 / FDA QSR / EU MDR 745 (medical-device contexts) |
| **Risk manager** | ISO 14971 + AI risk per ISO 23894 |
| **Internal auditor(s)** | Clause 9.2 audit programmes across all frameworks |
| **Executive sponsor** | Management review (Clause 9.3) across all frameworks |
A typical mid-stage AI SaaS has compliance officer + CISO + DPO as the core trio; AIMS lead is a part-time hat.
## The Integrated Management System Pattern
When multiple management-system standards apply (ISO 27001 + ISO 42001 + ISO 9001/13485 + ISO 14001), the recommended structure is an **Integrated Management System (IMS)** rather than parallel siloed systems. The IMS pattern:
- Single scope statement covering all applicable standards
- Single policy set with framework-specific overlays (e.g., the AI policy required by ISO 42001 A.2.2 sits alongside the info-sec policy required by ISO 27001 A.5.1)
- Single document control procedure
- Single internal audit programme covering all standards over a rolling 3-year cycle
- Single management review covering all standards
- Single CAPA loop with framework-tagged nonconformities
- Per-framework deep-dive evidence under common umbrella
Compliance OS is the operating model for the IMS pattern.
## How Compliance OS Relates to Sectoral Programs
| Sectoral context | Compliance OS approach |
|---|---|
| Pure SaaS (no AI, no medical) | Skip compliance-os. Use ISO 27001 + SOC 2 + GDPR skills directly. |
| AI SaaS (EU users) | Use compliance-os. Frameworks: ISO 27001 + SOC 2 + ISO 42001 + EU AI Act + GDPR. |
| AI medical device | Use compliance-os. Frameworks: ISO 13485 + 14971 + 42001 + EU AI Act + EU MDR / FDA QSR + GDPR. Most complex case. |
| Financial / regulated industry | Use compliance-os + sectoral overlay (e.g., NYDFS, FINMA, NIS2). |
## Anti-Patterns to Avoid
1. **Building compliance-os before having ≥ 2 frameworks operating maturely.** Premature orchestration. Mature one framework first; layer the second; THEN orchestrate.
2. **Using compliance-os to bypass per-framework deep work.** The cross-framework mapping says "reuse evidence from framework A." That presumes framework A's evidence is solid. Reuse mapping ≠ skip diligence.
3. **Treating mapping confidence as binary.** HIGH confidence means same evidence; MEDIUM means existing evidence with overlay; LOW means concept overlap. LOW mappings still need new artefacts.
4. **Forgetting that bindings (regulations) outrank certifications.** GDPR + EU AI Act non-compliance carries actual penalties; ISO 27001 non-certification just blocks procurement. Sequence accordingly.
5. **Replacing the per-framework skill with compliance-os.** Compliance OS orchestrates; per-framework skills do the deep work.
## When This Reference Doesn't Help
- **Specific framework requirements.** See the per-framework skill.
- **GRC platform selection.** Tooling decision; commercial market evolves rapidly.
- **Per-sector regulatory deep-dive.** Use sectoral skills (financial, healthcare, etc.).
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 19011:2018** — Guidelines for auditing management systems (the canonical audit standard for ISO-family certifications)
- **IIA International Professional Practices Framework (IPPF)** — Internal Audit Standards (Standards 1000-2600); attribute + performance standards
- **AICPA AT-C 105 + AU-C 240** — Trust Services + auditor's responsibility framework (SOC 2 + financial audit overlap)
- **COSO Enterprise Risk Management 2017** — Integrated framework for risk management across the enterprise
- **NIST Cybersecurity Framework 2.0** — profile pattern for organizing security/risk programmes (precedent for compliance-os approach)
- **ISO/IEC 27001:2022** — Information security management (foundational management system for most compliance programs)
- **ISO/IEC 17021** — Conformity assessment requirements (governs certification bodies; informs audit cycle)
- **ISACA** — *Auditing Artificial Intelligence* (2nd ed., 2024) — multi-framework AI audit guidance
- **ENISA** — *Multilayer Framework for Good Cybersecurity Practices for AI* (Mar 2023) — multi-layer integration
- **Annex SL of the ISO/IEC Directives** (2024) — the high-level structure shared by management system standards enabling integration
FILE:references/cross_framework_overlap.md
# Cross-Framework Overlap — The 9-Framework × Control-Family Matrix
This reference answers exactly one decision: **for each common control family, which of the 9 supported frameworks address it, and at what confidence?**
Pair with `scripts/cross_framework_mapper.py` for the deterministic lookup.
## The 9 Frameworks
| ID | Standard | Type |
|---|---|---|
| iso_27001 | ISO/IEC 27001:2022 + Annex A | Certifiable management system (info-sec) |
| iso_13485 | ISO 13485:2016 | Certifiable management system (medical device QMS) |
| iso_42001 | ISO/IEC 42001:2023 | Certifiable management system (AIMS) |
| iso_14971 | ISO 14971:2019 | Process standard (medical device risk management) |
| eu_ai_act | Regulation (EU) 2024/1689 | Binding regulation (AI) |
| eu_mdr_745 | Regulation (EU) 2017/745 | Binding regulation (medical devices) |
| gdpr | Regulation (EU) 2016/679 | Binding regulation (privacy) |
| soc_2 | AICPA SOC 2 TSC | Attestation (US enterprise procurement) |
| fda_qsr | FDA 21 CFR 820 | Binding regulation (US medical devices) |
## Highest-Overlap Pairs (where reuse leverage is maximized)
1. **ISO 27001 ↔ SOC 2** — densest known overlap. ISO 27001:2022 Annex A 93 controls map to SOC 2 TSC ~75% by published cross-walks. The 19 merged controls in `cross_framework_mapper.py` cite 51 atomic ISO 27001 + 34 atomic SOC 2 controls in HIGH-confidence themes. Adding SOC 2 on top of certified ISO 27001 is typically ~3 months of incremental work.
2. **ISO 13485 ↔ FDA QSR** — harmonised in 2024 (FDA Quality Management System Regulation rule). Most evidence reuses.
3. **ISO 42001 ↔ ISO 27001** — 60% reuse: most Clauses 4–10 evidence transfers with AI scope appended; Annex A controls A.7 (data) + A.10 (third-party) overlap heavily; the 40% net-new is mostly A.5 (impact assessment) + A.6 (lifecycle) + A.9 (use of AI systems).
4. **EU AI Act Article 17 ↔ ISO 42001** — ISO 42001 satisfies most of Article 17(1)(a)–(m) QMS requirements. The cross-walk in `compliance-team-iso42001/references/cross_framework_mapping_ai.md` provides Article 17 line-item mapping.
5. **GDPR ↔ ISO 27001 Annex A.5.34** — privacy by design overlap; GDPR Article 32 technical and organizational measures maps to ISO 27001 cryptography (A.8.24) + access control (A.5.15) + incident response (A.5.24).
## Control Family Overlap Matrix (summary)
Legend: ✅ direct overlap; 🔶 partial overlap with overlay; ⚠️ concept overlap only; ⛔ not applicable.
| Control family | 27001 | 13485 | 42001 | 14971 | EU AI Act | MDR | GDPR | SOC 2 | FDA QSR |
|---|---|---|---|---|---|---|---|---|---|
| Access control | ✅ | 🔶 | 🔶 | ⛔ | ⛔ | ⛔ | 🔶 | ✅ | 🔶 |
| Asset inventory | ✅ | ✅ | ✅ | ⛔ | ⛔ | ⛔ | 🔶 | ✅ | ✅ |
| Risk management | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 🔶 | ✅ | 🔶 |
| Supplier mgmt | ✅ | ✅ | ✅ | ⛔ | 🔶 | 🔶 | ✅ | ✅ | 🔶 |
| Incident response | ✅ | ✅ | 🔶 | 🔶 | 🔶 | ✅ | ✅ | ✅ | ✅ |
| Logging & monitoring | ✅ | 🔶 | 🔶 | ⛔ | 🔶 | 🔶 | ⚠️ | ✅ | 🔶 |
| Change management | ✅ | ✅ | 🔶 | ⛔ | ⛔ | ✅ | ⛔ | ✅ | ✅ |
| BCP / DR | ✅ | 🔶 | ⛔ | ⛔ | ⛔ | ⛔ | ⛔ | ✅ | ⛔ |
| Competence + training | ✅ | ✅ | ✅ | ⛔ | 🔶 | ✅ | ⛔ | ✅ | ✅ |
| Data governance | 🔶 | ✅ | ✅ | ⛔ | ✅ | ⚠️ | ✅ | ⚠️ | 🔶 |
| Internal audit | ✅ | ✅ | ✅ | ⛔ | ⛔ | 🔶 | ⛔ | ✅ | 🔶 |
| Management review | ✅ | ✅ | ✅ | ⛔ | ⛔ | 🔶 | ⛔ | 🔶 | ⛔ |
| Cryptography | ✅ | ⛔ | ⛔ | ⛔ | ⛔ | ⛔ | ✅ | ✅ | ⛔ |
| Secure SDLC | ✅ | ⛔ | 🔶 | ⛔ | 🔶 | ⛔ | ⛔ | ✅ | ⛔ |
| Vulnerability mgmt | ✅ | ⛔ | ⛔ | ⛔ | ⛔ | ⛔ | ⛔ | ✅ | ⛔ |
| Physical security | ✅ | ✅ | ⛔ | ⛔ | ⛔ | ✅ | ⛔ | ✅ | ✅ |
| Personal data protection | ✅ | ⛔ | 🔶 | ⛔ | 🔶 | ⛔ | ✅ | 🔶 | ⛔ |
| Documentation control | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | 🔶 | ✅ | ✅ |
| Continual improvement / CAPA | ✅ | ✅ | ✅ | ✅ | ⛔ | ✅ | ⛔ | ✅ | ✅ |
## How to Use This Matrix
1. **Identify the union** of applicable frameworks (from `framework_selector.py`)
2. **For each control family**, find the row and read the columns for your frameworks
3. **Build evidence once** for the framework with the strongest requirement, then reuse-with-overlay for others
4. **Document the reuse mapping** in your compliance program documentation so auditors can trace evidence to framework controls
## Practical Reuse Sequencing
If you operate ISO 27001 (mature) and add a second framework:
| Add | Reuse leverage from 27001 |
|---|---|
| **SOC 2** | ~75% — heaviest reuse; the canonical pair |
| **ISO 42001** | ~60% — Clauses 4–10 reuse strong; Annex A.7/A.10 reuse strong; A.5/A.6/A.9 net-new |
| **GDPR** | ~50% — Article 32 organizational measures reuse; Articles 5/6/30 net-new privacy work |
| **EU AI Act** | ~40% — Article 17 QMS via ISO 42001 path; Articles 9/10 net-new; transparency net-new |
| **ISO 13485** | ~30% — document control + CAPA reuse; design controls + medical specifics net-new |
| **FDA QSR** | ~30% — via ISO 13485 path; sectoral overlay |
| **EU MDR 745** | ~25% — most net-new (technical documentation, clinical evidence, UDI) |
| **ISO 14971** | ~20% — process standard, integrates with 13485 |
## Confidence Levels Explained
The `cross_framework_mapper.py` returns one of three confidence levels per mapping:
- **HIGH (H)** — same evidence satisfies both framework controls without modification. Example: a quarterly access-review record satisfies ISO 27001 A.5.15 + SOC 2 CC6.1 simultaneously.
- **MEDIUM (M)** — existing evidence plus a framework-specific overlay. Example: ISO 27001 supplier-management procedure adapted to add AI-specific clauses for ISO 42001 A.10.2.
- **LOW (L)** — concept overlap only; new artefact required. Example: ISO 42001 A.5.2 impact assessment uses concepts from GDPR DPIA but is a separate artefact.
## When This Reference Doesn't Help
- **Specific atomic control numbers.** See the per-framework skill's references.
- **Sector-specific overlays.** See sectoral skills (financial, healthcare).
- **Audit simulation depth.** See `audit_simulation_methodology.md`.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 27001:2022** + Annex A (the foundational pair source)
- **ISO/IEC 42001:2023** + Annex A
- **AICPA Trust Services Criteria** (2017 + 2022 update)
- **Regulation (EU) 2024/1689** (EU AI Act)
- **Regulation (EU) 2016/679** (GDPR)
- **Regulation (EU) 2017/745** (EU MDR)
- **ISO 13485:2016**
- **ISO 14971:2019**
- **FDA 21 CFR 820** (QSR) — harmonised under the FDA Quality Management System Regulation rule (effective 2026)
- **NIST SP 800-53 Rev 5** — security and privacy controls catalog (cross-walk reference)
- **NIST CSF 2.0** — profile pattern
- **ISACA** — *Mapping ISO 27001 to SOC 2* (continually updated)
- **CIS Controls v8** — additional cross-walk
- **CSA STAR** — cloud-specific cross-walk
FILE:references/evidence_artifact_reuse_index.md
# Evidence Artefact Reuse Index — Which Evidence Type Satisfies Most Controls Across Frameworks
This reference answers exactly one decision: **which evidence artefacts have the highest reuse leverage across the 12 supported frameworks, and what's the priority order for building them in a multi-framework programme?**
Pair with `scripts/evidence_pool_generator.py` for the operational catalogue. This document is the empirically-derived ranking + reasoning.
## Methodology
Reuse leverage = count of distinct (framework, control) tuples that one evidence artefact satisfies. Computed by tracing artefact-to-control mappings across:
- ISO/IEC 27001:2022 Annex A
- ISO/IEC 42001:2023 Annex A
- ISO 13485:2016 + ISO 14971:2019
- AICPA Trust Services Criteria (SOC 2)
- Regulation (EU) 2024/1689 (AI Act)
- Regulation (EU) 2017/745 (MDR)
- Regulation (EU) 2016/679 (GDPR)
- FDA 21 CFR 820 (QSR / QMSR)
- NIST Cybersecurity Framework 2.0
- Directive (EU) 2022/2555 (NIS2)
- HIPAA Security Rule + Privacy Rule + Breach Notification
For each evidence artefact, count of frameworks × controls satisfied = leverage score.
## The Top-Tier Artefacts (Build These First)
| Rank | Artefact | Reuse leverage | Acquisition cost | Why it's #1 |
|---|---|---|---|---|
| 1 | **Risk register with treatment plans** | 30+ mappings × 8+ frameworks | High | Every management-system standard + binding regulation demands risk management. Single artefact serves ISO 27001 Clause 6.1, ISO 42001 Clause 6.1.2, SOC 2 CC3, EU AI Act Article 9, GDPR Article 35 DPIA, NIST CSF GV.RM + ID.RA, NIS2 Article 21(2)(a), HIPAA §164.308(a)(1)(ii)(A) |
| 2 | **Asset inventory with classification** | 25+ mappings × 7+ frameworks | Medium | Required for ISO 27001 A.5.9-12, SOC 2 CC6.1, ISO 42001 A.4, GDPR Article 30, NIST CSF ID.AM, HIPAA §164.308 + §164.310(d). Foundation for almost every other artefact. |
| 3 | **Incident log + post-incident reviews + notifications** | 30+ mappings × 8+ frameworks | Medium | ISO 27001 A.5.24-27 + A.6.8, SOC 2 CC7.3-5, GDPR Articles 33-34, EU AI Act Article 73, NIS2 Article 23, HIPAA §164.308(a)(6) + Breach Notification, NIST CSF RS + RC |
| 4 | **Supplier inventory + reviews + DPAs/BAAs** | 25+ mappings × 8+ frameworks | Medium | ISO 27001 A.5.19-22, SOC 2 CC9.2, ISO 42001 A.10, GDPR Article 28, EU AI Act Article 25, NIST CSF GV.SC, NIS2 Article 21(2)(d), HIPAA §164.314(a) BAA |
| 5 | **Policy set (AI + info-sec + privacy + code-of-conduct)** | 20+ mappings × 7+ frameworks | Medium | ISO 27001 A.5.1, ISO 42001 Clause 5.2 + A.2.2-3, SOC 2 CC1.1-2, GDPR Article 24, NIST CSF GV.PO, EU AI Act Article 17(1)(a) |
## High-Leverage Artefacts (Build Next)
| Rank | Artefact | Reuse leverage | Acquisition cost | Notes |
|---|---|---|---|---|
| 6 | **Centralized tamper-evident logs** | 20+ mappings × 6+ frameworks | High | ISO 27001 A.8.15-16, SOC 2 CC7.1-2, ISO 42001 A.9.3-4, EU AI Act Article 12 + 72, NIST CSF DE.CM, HIPAA §164.312(b) audit controls |
| 7 | **Training records (per role, with effectiveness verification)** | 18+ mappings × 7+ frameworks | Medium | ISO 27001 A.6.3, SOC 2 CC1.4 + CC2.2, ISO 42001 Clause 7.2-3 + A.4.4, EU AI Act Article 4, NIST CSF PR.AT, NIS2 Article 21(2)(g), HIPAA §164.308(a)(5) |
| 8 | **Data inventory + provenance + consent register** | 20+ mappings × 6+ frameworks | High | ISO 27001 A.5.34, ISO 42001 A.7, EU AI Act Article 10, GDPR Articles 5+6+30, NIST CSF PR.DS + ID.AM-07, HIPAA §164.502 + §164.514 |
| 9 | **Internal audit programme records** | 15+ mappings × 6+ frameworks | Medium | ISO 27001 Clause 9.2, ISO 42001 Clause 9.2, ISO 13485 Clause 8.2.4, SOC 2 CC4.1, NIST CSF ID.IM, HIPAA §164.308(a)(8) |
| 10 | **Management review minutes + action tracking** | 12+ mappings × 5+ frameworks | Low | ISO 27001 Clause 9.3, ISO 42001 Clause 9.3, ISO 13485 Clause 5.6, NIST CSF GV.OV, NIS2 Article 20 |
## Mid-Leverage Artefacts
| Rank | Artefact | Reuse leverage | Acquisition cost | Notes |
|---|---|---|---|---|
| 11 | **Change records + rollback procedures + post-implementation reviews** | 14+ mappings × 5+ frameworks | Low | ISO 27001 A.8.32, SOC 2 CC8.1, ISO 42001 A.6.2.5, ISO 13485 Clause 7.3.9, NIST CSF PR.PS, HIPAA §164.308(a)(5)(ii)(B) |
| 12 | **Crypto records (algorithms, key lifecycle, KMS architecture)** | 14+ mappings × 6+ frameworks | Medium | ISO 27001 A.8.24, SOC 2 CC6.1 + CC6.7, GDPR Article 32(1)(a), NIST CSF PR.DS-01-02 + PR.PS-05, NIS2 Article 21(2)(h), HIPAA §164.312(a)(2)(iv) + §164.312(e)(2)(ii) |
| 13 | **BCP/DRP + RPO/RTO + exercise records** | 12+ mappings × 5+ frameworks | High | ISO 27001 A.5.29-30 + A.8.13-14, SOC 2 A1.2-3, NIST CSF RC.RP + RC.IM + RC.CO, NIS2 Article 21(2)(c), HIPAA §164.308(a)(7) |
| 14 | **DPIA records + LIAs + privacy notice version history** | 12+ mappings × 4+ frameworks | High | GDPR Articles 5+6+24+25+30+35+38, EU AI Act Article 27 FRIA (overlap), ISO 27001 A.5.34, ISO 42001 A.7.6 |
| 15 | **Quarterly access review records + RBAC matrix + JML evidence** | 18+ mappings × 7+ frameworks | Low | ISO 27001 A.5.15 + A.8.2-3, SOC 2 CC6.1-3, ISO 42001 A.4.4, GDPR Article 32(1)(b), NIST CSF PR.AA, NIS2 Article 21(2)(i), HIPAA §164.308(a)(3-4) + §164.312(a)(1) |
| 16 | **Vulnerability scan + patch SLA + remediation evidence** | 12+ mappings × 5+ frameworks | Medium | ISO 27001 A.8.7-9, SOC 2 CC7.1-2 + CC7.4, NIST CSF ID.RA + PR.PS-02, NIS2 Article 21(2)(f), HIPAA §164.308(a)(5)(ii)(B) |
## Low-Leverage (Framework-Specific) Artefacts
Build these only when the specific framework applies; lower reuse value across the programme.
| Artefact | Primary framework(s) | Why low-leverage |
|---|---|---|
| Annex IV technical documentation (EU AI Act) | EU AI Act | Specific to AI Act high-risk systems |
| Design History File (DHF) | ISO 13485, FDA QSR | Specific to medical-device QMS |
| Process validation (IQ/OQ/PQ) | ISO 13485, FDA QSR | Specific to medical-device manufacturing |
| Clinical evaluation (Annex XIV) | EU MDR | Specific to medical-device EU placement |
| Model card + datasheet | ISO 42001, EU AI Act | AI-specific |
| FRIA (Fundamental Rights Impact Assessment) | EU AI Act | Specific to high-risk AI public-sector deployers |
| Notice of Privacy Practices | HIPAA | Specific to US healthcare |
| Form 483 response records | FDA QSR | Specific to FDA-inspected entities |
| NIS2 incident notifications (24h/72h/1m) | NIS2 | Specific to NIS2-in-scope entities |
| EUDAMED registration | EU MDR | Specific to EU MDR |
## Reuse-Leverage Operational Pattern
For a multi-framework programme, the recommended build order is:
```
Phase 1 (Weeks 1-4):
- Risk register with treatment plans (top reuse)
- Asset inventory with classification
- Policy set
- Quarterly access review records + RBAC matrix
Phase 2 (Weeks 5-12):
- Centralized tamper-evident logs
- Supplier inventory + DPAs/BAAs
- Training records
- Crypto records
- Internal audit programme records
- Management review records
Phase 3 (Weeks 13-24):
- Data inventory + provenance + consent (build alongside Phase 1 if GDPR/HIPAA early)
- BCP/DRP + exercise records
- DPIA records
- Vulnerability scan + remediation
- Change records + rollback procedures
- Incident log + post-incident reviews
- Physical security records (if applicable)
Phase 4 (Weeks 25+):
- Framework-specific artefacts:
* Annex IV docs (if EU AI Act)
* DHF + process validation (if ISO 13485 / FDA QSR)
* Clinical evaluation (if EU MDR)
* Model cards + datasheets (if ISO 42001)
* FRIA (if EU AI Act public-sector deployer)
* Notice of Privacy Practices (if HIPAA)
```
## Common Mistakes (Anti-Patterns)
1. **Building framework-specific artefacts before top-tier reuse artefacts.** Common when team is led by a single-framework specialist; results in 5x more total effort across the programme.
2. **Separate evidence stores per framework.** Each framework wants the same access-review log; storing it 3 times in 3 systems = stale + inconsistent.
3. **Not citing the same artefact in multiple audit reports.** Different auditors may ask for the same evidence renamed; cite the shared artefact ID in both reports.
4. **Skipping centralized inventory in Phase 1.** Asset inventory is the foundation for risk register, supplier list, data inventory, etc. Without it, everything downstream is incomplete.
5. **Treating evidence as one-time collection rather than continuous artefact.** Quarterly access review records must be produced quarterly, not "fixed for the audit and then ignored".
## Evidence Freshness Discipline
Reuse leverage breaks down if evidence is stale. Per-artefact target freshness:
| Artefact | Refresh cadence | Stale = ineffective |
|---|---|---|
| Risk register | Quarterly minimum | Within 90 days |
| Asset inventory | Quarterly minimum | Within 90 days |
| Access review records | Quarterly | Within 1 quarter |
| Incident log + PIRs | Continuous + 30-day PIR | PIR within 30 days |
| Supplier reviews | Annually | Within 12 months |
| Training records | Annually + new-hire 30 days | Annual completion 100% |
| Policy set | Annually reviewed | Within 12 months |
| Crypto inventory | Quarterly review | Within 90 days |
| DPIA records | At new processing + on material change | Always current |
| BCP/DRP exercise records | Annually | Within 12 months |
## Anti-Reuse Patterns to Avoid
- **Per-framework reformatting** — collecting an artefact, then reformatting for each framework's report. Cite the shared artefact + map to framework controls instead.
- **Per-team ownership without integration** — security owns SOC 2 evidence, DPO owns GDPR evidence, RA/QM owns ISO 13485 evidence, no shared discovery layer. Use compliance-os meta-orchestrator to enforce shared inventory.
- **Custodial-only ownership** — artefact lives in one team's drive without index. New audit cycle re-discovers from scratch.
## When This Reference Doesn't Help
- **Specific GRC platform configuration.** Tooling decision; see vendor documentation.
- **Per-control evidence requirements.** See per-framework skill references.
- **Sector-specific evidence (financial NYDFS, energy NERC CIP).** Sectoral; not in 12-framework scope.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 27001:2022** + Annex A
- **ISO/IEC 42001:2023** + Annex A
- **ISO/IEC 19011:2018** — Guidelines for auditing management systems (audit evidence)
- **AICPA Trust Services Criteria** (2017 + 2022 update) + SOC 2 Reporting Guide
- **Regulation (EU) 2024/1689** — AI Act
- **Regulation (EU) 2017/745** — EU MDR
- **Regulation (EU) 2016/679** — GDPR
- **Regulation (EU) 2022/2555** — NIS2 Directive
- **NIST Cybersecurity Framework 2.0** + NIST SP 800-53A Rev 5 assessment procedures
- **HIPAA 45 CFR Parts 160 + 164** — Security + Privacy + Breach Notification Rules
- **FDA 21 CFR 820** — Quality System Regulation
- **ISO 13485:2016** + ISO 14971:2019
- **IIA International Professional Practices Framework** — Performance Standards on engagement records (2330)
- **DAMA-DMBOK 2** — Data Management Body of Knowledge (provenance + quality dimensions)
- **NIST SP 800-92** — Guide to Computer Security Log Management (retention + integrity)
- **Industry retrospectives** — Big 4 + Schellman + Coalfire + A-LIGN published findings on common audit exceptions
FILE:references/evidence_management.md
# Evidence Management — Unified Pool + Reuse Leverage
This reference answers exactly one decision: **how do we collect compliance evidence once and satisfy multiple frameworks, without losing audit-grade traceability?**
Pair with `scripts/evidence_pool_generator.py` for the deterministic evidence catalogue.
## The Evidence Reuse Problem
Most multi-framework compliance programs accidentally collect the same evidence multiple times. Each framework's auditor wants:
- A documented procedure (the "what should happen")
- Records that the procedure was followed (the "what actually happened")
- Evidence of management oversight (the "did anyone check?")
When ISO 27001, SOC 2, and ISO 42001 audits ask for "access review records," teams often produce three different exports of the same Okta data with different formatting because three different control owners assembled them.
The fix: a **unified evidence pool** with explicit (artefact, framework, control) mapping. Collect once; cite multiple times.
## The Reuse-Leverage Score
Every evidence artefact gets a **reuse-leverage score** = number of distinct (framework, control) tuples it satisfies. Higher score = higher priority to build first.
From the `evidence_pool_generator.py` curated catalogue, the top-leverage artefacts (when all 9 frameworks are enabled):
| Artefact | Leverage |
|---|---|
| Risk register | 9+ mappings |
| Supplier inventory + reviews + DPAs | 8+ |
| Incident log + post-mortems + notifications | 11+ |
| Data inventory + provenance + consent | 9+ |
| Policy set (AI + info-sec + privacy + code-of-conduct) | 8+ |
| Tamper-evident logs centralized | 7+ |
| Training records | 6+ |
**Implementation order:** build high-leverage artefacts first. The risk register alone unlocks evidence for 9+ controls across 4+ frameworks.
## Evidence Acquisition Cost
The catalogue tracks acquisition cost per artefact: low / medium / high.
| Cost | Examples | Time to build |
|---|---|---|
| **Low** | Quarterly access review records, change records, management review records | 1-2 weeks (often automated from existing IT systems) |
| **Medium** | Asset register, supplier inventory, training records, crypto records, vuln scans | 2-6 weeks (requires inventory + classification) |
| **High** | Risk register, BCP/DR exercises, data inventory + consent register, secure SDLC | 6-12 weeks (requires cross-functional process design) |
**Strategy:** in year 1, prioritize low-cost high-leverage artefacts (e.g., management review records, change records). Build high-cost high-leverage artefacts in parallel (risk register, data inventory).
## Retention by Framework
Retention requirements vary per framework. Use the longest applicable retention:
| Framework | Typical retention |
|---|---|
| ISO 27001 | 3 years for audit evidence (or as policy specifies) |
| SOC 2 | 1 year minimum; 3 years recommended |
| ISO 42001 | 3 years (Clause 7.5 documented information) |
| EU AI Act | 10 years for declaration of conformity (Article 18); other docs 6 years |
| GDPR | Varies by data type; data subject records 3 years; breach records indefinite |
| ISO 13485 | Lifecycle of device + period defined by regulator (often 5+ years) |
| EU MDR | Device lifetime + 10 years (Article 10) |
| FDA QSR | 2 years past commercial distribution (21 CFR 820.180) |
**Default policy:** 36 months for most artefacts; 60 months for personal-data and policy-set artefacts; 120 months for EU AI Act declarations of conformity.
## Evidence Freshness
Auditors want recent evidence, not stale. Freshness expectations:
- Operational records (access reviews, change records, incident records): within last 90-180 days
- Quarterly artefacts: at least 1 record from current quarter
- Annual artefacts (training records, supplier reviews, BCP exercises): within last 12 months
- Policies: reviewed annually (review records demonstrate freshness)
**Stale evidence = effective gap.** An ISO 27001 A.5.15 quarterly access review that was last conducted 8 months ago is a major nonconformity even if the review existed historically.
## Evidence Owner Assignment
Each artefact has a primary owner. Typical pattern:
| Artefact type | Primary owner | Secondary |
|---|---|---|
| Access reviews | IT / Security | Compliance |
| Asset register | Security | DPO |
| Risk register | Compliance officer | Risk manager |
| Supplier inventory | Procurement | Compliance + DPO |
| Incident log | Security / IR team | Compliance |
| Logs (centralized) | Platform / SRE | Security |
| Change records | Engineering / Platform | Compliance |
| BCP/DR | Platform / SRE | Compliance |
| Training records | HR / People Ops | Compliance |
| Data inventory + consent | DPO / Data team | Engineering |
| Internal audit records | Compliance officer | Internal auditor |
| Management review records | Compliance officer + Exec | All function heads |
| Policy set | Compliance officer + Exec | All policy owners |
| Crypto records | Security | Platform |
| Vuln scans + patches | Security | Engineering |
**Single accountable owner per artefact** is critical. Joint ownership without accountability is the most common cause of stale evidence.
## Evidence Storage Architecture
Patterns observed in mature programs:
1. **GRC platform (Drata, Vanta, OneTrust, Hyperproof, etc.)** — the most common pattern; integrates with operational tools (Okta, AWS, GitHub) and auto-pulls evidence. Centralizes audit-trail.
2. **Compliance-team-managed repository** — folder per framework with subdivision per control; manual evidence assembly. Works for small programs; doesn't scale.
3. **Hybrid** — automated evidence (logs, access reviews, change records) in GRC platform; manual evidence (policies, management review minutes, training records) in document management system. Most common at growth-stage.
Compliance OS does not prescribe a storage pattern — but it does require:
- Single index of evidence (the unified pool)
- Per-evidence audit trail (who created, who approved, when)
- Per-evidence retention timer
- Per-evidence freshness alert
## Evidence Pool Quality Indicators
Healthy pool:
| Indicator | Healthy value |
|---|---|
| Average reuse leverage | ≥ 4 |
| Stale evidence (past expected freshness) | 0% |
| Orphan controls (no evidence assigned) | 0 |
| Unowned artefacts | 0 |
| Retention compliance | 100% |
Unhealthy pool:
- Many low-leverage artefacts (each satisfies only 1 framework) — likely silo'd collection
- High stale rate — operational discipline broken
- Orphan controls — gap in coverage that will surface at next audit
## Evidence Pool Audit (the meta-audit)
Once a year, audit the evidence pool itself:
1. Sample 10% of artefacts; verify they exist + are owned + are fresh
2. Sample 10% of controls; verify each has at least one evidence artefact assigned
3. Verify retention compliance — look for old evidence that should be deleted (GDPR retention) and recent evidence that should be retained longer
4. Verify framework coverage — are all enabled frameworks adequately represented?
This audit-of-audit is the most underappreciated discipline in mature multi-framework programs.
## When This Reference Doesn't Help
- **Specific GRC platform configuration.** Tooling-specific; market evolves rapidly.
- **Evidence retention for novel data types (e.g., AI training data).** Sector-specific; engage counsel.
- **Cross-framework specific mapping.** See `cross_framework_overlap.md`.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 27001:2022 Clause 7.5** — Documented information requirements
- **ISO/IEC 42001:2023 Clause 7.5** — AI-specific documented information
- **AICPA AT-C 205** — Examination engagements (SOC 2 evidence standards)
- **NIST SP 800-53A Rev 5** — Assessing Security and Privacy Controls (per-control evidence types)
- **NIST SP 800-92** — Guide to Computer Security Log Management
- **ISO/IEC 19011:2018 Clause 6.4** — Conducting audit activities (evidence collection)
- **IIA IPPF Performance Standard 2330** — Documenting Information (engagement records)
- **GDPR Article 30** — Records of processing activities (retention + evidence)
- **EU AI Act Article 18** — Document retention (10 years post-market for declaration of conformity)
- **FDA 21 CFR 820.180** — General requirements for records (2 years past commercial distribution)
- **DAMA-DMBOK 2** — Data Management Body of Knowledge (data-quality + provenance frameworks)
FILE:references/multi_framework_audit_playbook.md
# Multi-Framework Audit Playbook — Orchestrating Audits Across N Frameworks
This reference answers exactly one decision: **when 2+ frameworks operate simultaneously, how do we run audits in coordinated cycles with minimal duplication?**
Pair with `scripts/audit_simulator.py` (multi-framework mock audits) + the per-framework audit playbooks (`isms-audit-expert/references/iso27001_audit_playbook.md`, `qms-audit-expert/references/iso13485_audit_playbook.md`, `gdpr-dsgvo-expert/references/gdpr_audit_playbook.md`, `soc2-compliance/references/soc2_audit_playbook.md`).
## The Multi-Framework Audit Problem
Mature multi-framework programs face four orchestration challenges:
1. **Audit calendar conflicts** — surveillance audits stacking in same week, insufficient auditor capacity
2. **Auditor independence across frameworks** — same internal auditor pulled to audit own work in a different framework
3. **Evidence freshness mismatch** — Audit A wants Q3 data; Audit B (3 months later) wants same control's Q3+Q4 data
4. **Finding cross-impact** — a critical finding in ISO 27001 audit triggers compensating questions in SOC 2 audit
This playbook describes the integrated audit programme (IAP) pattern that solves these.
## The Integrated Audit Programme
```
Annual Compliance Calendar
|
┌─────────────────────┼─────────────────────┐
| | |
Q1: ISO 27001 Q2: ISO 42001 Q3: ISO 13485
internal audit internal audit internal audit
(auditor pool A) (auditor pool B) (auditor pool A)
|
Q4: Integrated Management
Review (Clause 9.3 across
all frameworks)
|
External surveillance audits
scheduled by certification body
```
The IAP coordinates:
- **Single audit programme document** covering all applicable frameworks
- **Single auditor pool** with skill-based + independence-based assignment
- **Single evidence pool** (per `evidence_pool_generator.py`) so audits cite shared evidence
- **Single management review** (per Annex SL) covering all frameworks' Clause 9.3 inputs
## The 12-Month Calendar Pattern
A typical mid-stage AI SaaS running ISO 27001 + SOC 2 + ISO 42001 + GDPR + EU AI Act:
| Quarter | Activity | Frameworks audited internally |
|---|---|---|
| **Q1** | ISO 27001 internal audit + SOC 2 Type II observation begins | 27001 + SOC 2 |
| **Q2** | ISO 42001 internal audit + EU AI Act readiness checkpoint | 42001 + AI Act |
| **Q3** | GDPR annual review + SOC 2 mid-period checkpoint | GDPR + SOC 2 |
| **Q4** | Integrated management review + SOC 2 Type II field + cert body surveillance audits | all |
External audits (certification body + SOC 2 audit firm) typically:
- Q1: ISO 27001 surveillance audit (timed to follow Q1 internal audit)
- Q3: SOC 2 Type II field testing (timed for Q4 report)
- Q4: ISO 42001 surveillance audit (timed to follow Q2 + Q4 internal audits)
## Auditor Independence Across Frameworks
ISO management-system standards (Clause 9.2 across 27001 / 42001 / 13485) all require auditor independence: nobody audits their own work. With multiple frameworks running, independence must be tracked **across** frameworks, not just within.
**Pattern:** maintain an auditor competence + independence matrix:
| Auditor | Owns (cannot audit) | Competent to audit |
|---|---|---|
| Alice | 27001 A.5.15 (access control); 42001 A.4.4 | 27001 except A.5.15; 42001 except A.4.4; all GDPR; all SOC 2 |
| Bob | 42001 A.6 (lifecycle); 13485 7.3 (design) | 27001; GDPR; SOC 2 |
| Carol (external) | (none — independent contractor) | All frameworks |
| Dave | 27001 A.5.19 (suppliers); GDPR Article 28 | 27001 except A.5.19; 42001; SOC 2; 13485 |
Use `aims_audit_scheduler.py` (ISO 42001) + per-framework scheduler patterns to enforce independence.
## Cross-Framework Finding Impact
A finding in one framework's audit often affects another. Pattern:
- **ISO 27001 A.5.15 finding** → likely SOC 2 CC6.1 finding (same evidence)
- **ISO 27001 A.5.19-21 finding** → likely SOC 2 CC9.2 finding + GDPR Article 28 finding
- **ISO 42001 Annex A.7.6 finding** → likely GDPR Article 35 DPIA finding
- **ISO 13485 Clause 7.3 finding** → likely EU MDR Annex II finding
- **GDPR Article 33 breach** → triggers ISO 27001 A.5.24 audit + EU AI Act Article 73 review
**Discipline:** when a finding is issued, the issuing auditor flags cross-framework impact in the finding worksheet. The compliance officer reviews and triggers corresponding follow-up across frameworks.
## Shared Evidence Discipline
Per `evidence_management.md`, the evidence pool has unified artefacts. Audit work cites these artefacts, not framework-specific copies.
**Anti-pattern:**
```
ISO 27001 audit asks for: "ISO 27001 access review records Q3"
SOC 2 audit asks for: "SOC 2 access review records Q3"
Team produces TWO documents from same Okta export.
```
**Pattern:**
```
Both audits cite: "ev.access_review_quarterly Q3 2026" (single artefact)
Audit reports reference the shared artefact ID + framework-control mapping.
```
The audit report shows the auditor consulted the same evidence; framework-specific formatting happens in report assembly.
## Integrated Management Review (Clause 9.3 Across Frameworks)
Each management-system standard (27001, 42001, 13485, etc.) requires its own management review with prescribed inputs + outputs. Running 4 separate management reviews per year is unsustainable.
**Per Annex SL** (the high-level structure shared across ISO management-system standards), a single integrated management review can satisfy all of them if inputs cover every framework's prescribed list. Required inputs across the 5 most-common frameworks:
| Input | 27001 | 42001 | 13485 | 14001 | 9001 |
|---|---|---|---|---|---|
| Audit results | ✅ | ✅ | ✅ | ✅ | ✅ |
| Feedback from interested parties | ✅ | ✅ | ✅ | ✅ | ✅ |
| Risk + opportunity changes | ✅ | ✅ | ✅ | ✅ | ✅ |
| Performance of processes | ✅ | ✅ | ✅ | ✅ | ✅ |
| Nonconformities + CAPA | ✅ | ✅ | ✅ | ✅ | ✅ |
| Improvement opportunities | ✅ | ✅ | ✅ | ✅ | ✅ |
| AI-specific (drift, incidents, lifecycle) | — | ✅ | — | — | — |
| Customer feedback + complaints | — | ✅ | ✅ | — | ✅ |
| Resource needs | ✅ | ✅ | ✅ | ✅ | ✅ |
Outputs are similarly aligned: decisions on improvement, resource changes, scope adjustments, policy changes.
**Cadence:** annual minimum; quarterly preferred for mature multi-framework programs.
## Pre-Audit Readiness Checklist (per framework)
Universal pre-audit readiness (apply to each framework's internal audit):
- [ ] Scope confirmed (clauses + controls + business units in scope)
- [ ] Auditor independence verified (no self-audit; competence covers scope)
- [ ] Prior-year findings open list pulled + status reviewed
- [ ] Document evidence assembled in advance (auditor reads pre-fieldwork)
- [ ] Auditee leadership briefed; team availability confirmed
- [ ] Mock audit run via `audit_simulator.py` to surface likely findings
- [ ] Cross-framework impact considered (which findings might cascade)
- [ ] Audit plan circulated 2 weeks ahead
## Post-Audit Disciplines
- Findings logged in unified CAPA system (not framework-siloed)
- Corrective action owner named; due date agreed
- Cross-framework impact flagged in finding worksheet
- Closure verified by evidence + re-test (not self-attestation)
- Trend analysis monthly: aging CAPAs > 30 days, repeat findings across frameworks
- Inputs prepared for next management review
## When This Reference Doesn't Help
- **Single-framework deep audit detail.** See per-framework playbooks.
- **External certification body audit process.** Different from internal; see ISO 17021.
- **External SOC 2 audit firm engagement.** Different from internal; see `soc2_audit_playbook.md`.
- **Sectoral regulatory enforcement.** Out of scope; engage outside counsel.
---
**Source authorities (non-exhaustive):**
- **ISO/IEC 19011:2018** — Guidelines for auditing management systems
- **IIA International Professional Practices Framework (IPPF)** — Performance Standards 2000-2600
- **AICPA AT-C 105** — Attestation engagement standard (SOC 2)
- **ISO/IEC 27001:2022 Clause 9.2** — Internal audit programme
- **ISO/IEC 42001:2023 Clause 9.2** — Internal audit programme (AI management system)
- **ISO 13485:2016 Clause 8.2.4** — Internal audit (medical devices)
- **Regulation (EU) 2016/679 Article 24** — Accountability (GDPR — operational discipline for audit prep)
- **ISO 17021-1:2015** — Conformity assessment requirements (governs external certification audits; informs internal practice)
- **Annex SL of the ISO/IEC Directives** (2024) — high-level structure enabling integrated management systems
- **NIST SP 800-53A Rev 5** — Assessing Security and Privacy Controls (multi-framework assessment procedures)
- **The Institute of Internal Auditors** — practical guides on integrated audit programme design
FILE:scripts/audit_simulator.py
#!/usr/bin/env python3
"""audit_simulator.py — Mock internal audit generator per ISO 19011 + IIA IPPF.
Stdlib-only. Given a framework + scope, generates a realistic mock audit with:
- 8-15 finding scenarios per typical ISO 19011 audit depth
- Severity distribution matching IIA expectations:
observation/OFI: ≥ 40%
minor: 20-30%
major: 15-25%
critical: ≤ 15%
- 3-5 interview questions per scoped control
- Document-review request list
- Walk-through scenarios where applicable
Deterministic generation from finding templates. Severity distribution is
proportional to the scope size. No randomness, no LLM calls.
Input schema (JSON):
{
"audit_name": "Q3 ISO 27001 internal audit — Platform team",
"framework": "iso_27001",
"scope_controls": ["A.5.15", "A.8.2", "A.8.15", "A.8.32", "A.5.19"],
"auditee_team": "Platform engineering",
"prior_year_findings_open": 2
}
Usage:
python audit_simulator.py
python audit_simulator.py path/to/audit_scope.json
python audit_simulator.py audit_scope.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"audit_name": "Q3 ISO 27001 internal audit — Platform team",
"framework": "iso_27001",
"scope_controls": ["A.5.15", "A.8.2", "A.8.15", "A.8.32", "A.5.19", "A.5.24", "A.6.8"],
"auditee_team": "Platform engineering",
"prior_year_findings_open": 2,
}
# Finding template library (theme -> {severity bucket -> finding patterns})
# Each template produces a finding scenario when invoked.
FINDING_TEMPLATES: Dict[str, Dict[str, List[str]]] = {
"access_control": {
"critical": [
"Privileged access reviewed annually instead of quarterly; orphaned accounts found in production.",
],
"major": [
"Quarterly access review evidence present but lacks documented business justification for retained privileges.",
"Joiner-mover-leaver workflow does not auto-deprovision on termination; manual gap of 5+ days observed.",
],
"minor": [
"Access review records lack documented review-completion timestamps in 2 of 6 sampled reviews.",
],
"observation": [
"Consider extending RBAC matrix to include cloud-resource scope (currently application-tier only).",
],
},
"logging_monitoring": {
"critical": [
"Production application logs disabled in past 30 days; no detection of the gap until audit fieldwork.",
],
"major": [
"Log retention configured at 90 days but framework requires 12 months; misalignment not detected.",
"Tamper-evident logging not enforced on privileged-user activity logs.",
],
"minor": [
"Monitoring alert thresholds not formally documented; reviewed verbally by SRE only.",
],
"observation": [
"Centralized log aggregation in place; consider adding anomaly detection.",
],
},
"change_management": {
"critical": [
"Emergency change procedure not formalized; observed 3 cases of production changes without recorded approval.",
],
"major": [
"Change advisory board records show approvals but no post-implementation review of high-risk changes.",
],
"minor": [
"Rollback procedure documented but not tested for 2 services in scope.",
],
"observation": [
"Consider linking change records to deployment automation for stronger evidence chain.",
],
},
"supplier_mgmt": {
"critical": [
"Critical SaaS supplier in use without signed DPA + security questionnaire (GDPR exposure).",
],
"major": [
"Annual supplier security review not completed for 3 of 8 critical suppliers.",
"Sub-processor list not maintained for critical suppliers handling personal data.",
],
"minor": [
"Supplier onboarding checklist exists but not consistently applied across business units.",
],
"observation": [
"Consider centralizing supplier risk evidence in a single GRC system.",
],
},
"incident_response": {
"critical": [
"Recent P1 incident lacks documented post-incident review (PIR) within 30-day SLA.",
],
"major": [
"Severity definitions documented but inconsistently applied across teams; impact varies.",
"Notification SLAs not aligned across frameworks (GDPR 72h, framework X 24h, framework Y 15 days).",
],
"minor": [
"Incident commander rotation not documented.",
],
"observation": [
"Consider quarterly tabletop exercises to validate runbooks.",
],
},
}
# Control -> theme mapping (heuristic; deterministic)
CONTROL_TO_THEME: Dict[str, str] = {
# ISO 27001 mapping
"A.5.15": "access_control",
"A.8.2": "access_control",
"A.8.3": "access_control",
"A.5.19": "supplier_mgmt",
"A.5.20": "supplier_mgmt",
"A.5.21": "supplier_mgmt",
"A.5.22": "supplier_mgmt",
"A.5.24": "incident_response",
"A.5.25": "incident_response",
"A.5.26": "incident_response",
"A.5.27": "incident_response",
"A.6.8": "incident_response",
"A.8.15": "logging_monitoring",
"A.8.16": "logging_monitoring",
"A.8.32": "change_management",
# SOC 2 mapping
"CC6.1": "access_control",
"CC6.2": "access_control",
"CC6.3": "access_control",
"CC9.2": "supplier_mgmt",
"CC7.3": "incident_response",
"CC7.4": "incident_response",
"CC7.5": "incident_response",
"CC7.1": "logging_monitoring",
"CC7.2": "logging_monitoring",
"CC8.1": "change_management",
# ISO 42001 mapping
"A.4.4": "access_control",
"A.9.3": "logging_monitoring",
"A.9.4": "logging_monitoring",
"A.6.2.5": "change_management",
"A.10.2": "supplier_mgmt",
"A.8.4": "incident_response",
}
def _severity_rotation() -> List[str]:
return [
"observation", "observation", "observation", "minor", "major",
"observation", "minor", "observation", "major", "critical",
"minor", "observation", "minor", "observation", "major",
]
def generate_findings(payload: Dict[str, Any]) -> List[Dict[str, Any]]:
"""Generate finding scenarios deterministically from scope."""
findings: List[Dict[str, Any]] = []
scope = payload.get("scope_controls", [])
prior_open = payload.get("prior_year_findings_open", 0)
# Rotate severities to hit IIA-target distribution
# Target: >= 40% observation, ~25% minor, ~20% major, <= 15% critical
severity_order = _severity_rotation()
# Pad if scope is large
while len(severity_order) < len(scope) + 5:
severity_order += severity_order
for idx, control in enumerate(scope):
theme = CONTROL_TO_THEME.get(control)
if theme is None:
continue
severity = severity_order[idx]
# If prior_open > 0, force first finding to be major (follow-up)
if idx == 0 and prior_open > 0:
severity = "major"
templates = FINDING_TEMPLATES.get(theme, {}).get(severity, [])
if not templates:
severity = "observation"
templates = FINDING_TEMPLATES.get(theme, {}).get("observation", ["General observation noted."])
finding_text = templates[idx % len(templates)]
findings.append({
"id": f"F-{idx + 1:02d}",
"control": control,
"theme": theme,
"severity": severity,
"description": finding_text,
"follow_up_from_prior": idx == 0 and prior_open > 0,
})
# Add 3-6 additional observations to hit 10-15 total range per ISO 19011 typical depth
extras_needed = max(0, 10 - len(findings))
extras_added = 0
for theme in FINDING_TEMPLATES:
if extras_added >= extras_needed:
break
if not any(f["theme"] == theme for f in findings):
continue
templates = FINDING_TEMPLATES[theme]["observation"]
findings.append({
"id": f"F-{len(findings) + 1:02d}",
"control": "(general)",
"theme": theme,
"severity": "observation",
"description": templates[(extras_added + 1) % len(templates)],
"follow_up_from_prior": False,
})
extras_added += 1
return findings
def interview_questions(control: str) -> List[str]:
"""Deterministic 3-5 audit interview questions per control theme."""
theme = CONTROL_TO_THEME.get(control)
bank = {
"access_control": [
"Walk me through how a new joiner gets access provisioned.",
"Show me the last quarterly access review evidence for a privileged role.",
"What happens within 24 hours of a termination?",
"How is multi-factor authentication enforced for admin access?",
],
"logging_monitoring": [
"Show me a sample log entry for a privileged action in the last 30 days.",
"What's the log retention configuration, and where is it documented?",
"How are tampering attempts detected and alerted?",
"Show me a monitoring alert that fired in the last 7 days and how it was triaged.",
],
"change_management": [
"Walk me through the change approval workflow for a production deployment.",
"Show me a rejected change in the last quarter and the rejection rationale.",
"Where is the rollback procedure for service X documented and last tested?",
"How are emergency changes handled differently from standard changes?",
],
"supplier_mgmt": [
"Show me the supplier inventory and the last review date for 3 critical suppliers.",
"How are AI-specific contractual clauses tracked for AI service suppliers?",
"Walk me through onboarding of a new critical SaaS supplier.",
"Show me where signed DPAs are stored for personal-data sub-processors.",
],
"incident_response": [
"Show me the last 3 incidents with severity, root cause, and corrective action.",
"Walk me through your serious-incident reporting timing for GDPR + AI Act.",
"Where are post-incident reviews documented and tracked to closure?",
"How is the on-call rotation defined and communicated?",
],
}
return bank.get(theme, [
"Walk me through how this control is implemented day-to-day.",
"Show me records of the control being operated in the last 90 days.",
"How is effectiveness of this control measured?",
])
def document_requests(scope: List[str]) -> List[str]:
themes = {CONTROL_TO_THEME.get(c) for c in scope if CONTROL_TO_THEME.get(c)}
docs = []
for t in themes:
if t == "access_control":
docs.append("Access control policy + last 2 quarterly access reviews + RBAC matrix")
elif t == "logging_monitoring":
docs.append("Logging policy + log retention configuration + last 30 days of sample privileged-action logs")
elif t == "change_management":
docs.append("Change management procedure + last 90 days change records + rollback procedure")
elif t == "supplier_mgmt":
docs.append("Supplier inventory + last annual supplier reviews + 3 sample DPAs")
elif t == "incident_response":
docs.append("Incident response procedure + last 5 incident records + post-incident reviews")
return docs
def analyze(payload: Dict[str, Any]) -> Dict[str, Any]:
findings = generate_findings(payload)
by_sev: Dict[str, int] = {"critical": 0, "major": 0, "minor": 0, "observation": 0}
for f in findings:
by_sev[f["severity"]] += 1
total = len(findings)
obs_pct = round((by_sev["observation"] / total) * 100, 1) if total else 0
crit_pct = round((by_sev["critical"] / total) * 100, 1) if total else 0
healthy = (obs_pct >= 40) and (crit_pct <= 15)
return {
"audit_name": payload.get("audit_name"),
"framework": payload.get("framework"),
"scope_controls": payload.get("scope_controls", []),
"auditee_team": payload.get("auditee_team"),
"findings_total": total,
"findings_by_severity": by_sev,
"severity_distribution_healthy": healthy,
"obs_pct": obs_pct,
"crit_pct": crit_pct,
"findings": findings,
"interview_questions_per_control": {c: interview_questions(c) for c in payload.get("scope_controls", [])},
"document_review_requests": document_requests(payload.get("scope_controls", [])),
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("COMPLIANCE OS — MOCK INTERNAL AUDIT (per ISO 19011 + IIA IPPF)")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Audit: {r['audit_name']}")
lines.append(f"Framework: {r['framework']} | Auditee: {r['auditee_team']}")
lines.append(f"Scope controls ({len(r['scope_controls'])}): {', '.join(r['scope_controls'])}")
lines.append("")
s = r["findings_by_severity"]
lines.append(f"Findings total: {r['findings_total']} "
f"(critical={s['critical']}, major={s['major']}, minor={s['minor']}, observation={s['observation']})")
lines.append(f"Distribution: observation={r['obs_pct']}% critical={r['crit_pct']}% "
f"healthy={r['severity_distribution_healthy']}")
lines.append("")
lines.append("-" * 72)
lines.append("FINDINGS:")
lines.append("")
for f in r["findings"]:
marker = "🔥 FOLLOW-UP" if f["follow_up_from_prior"] else ""
lines.append(f" [{f['id']}] [{f['severity'].upper():12s}] control={f['control']:12s} theme={f['theme']:20s} {marker}")
lines.append(f" {f['description']}")
lines.append("")
lines.append("-" * 72)
lines.append("INTERVIEW QUESTIONS PER CONTROL:")
for ctrl, qs in r["interview_questions_per_control"].items():
lines.append(f" {ctrl}:")
for q in qs:
lines.append(f" - {q}")
lines.append("")
lines.append("-" * 72)
lines.append("DOCUMENT-REVIEW REQUESTS:")
for d in r["document_review_requests"]:
lines.append(f" - {d}")
lines.append("")
lines.append("-" * 72)
lines.append("HEALTHY-DISTRIBUTION RULE (IIA expectations):")
lines.append(" observation/OFI ≥ 40% AND critical ≤ 15%")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Mock internal audit generator per ISO 19011 + IIA IPPF + AICPA AT-C.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to audit scope JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: Q3 ISO 27001 internal audit, Platform team, 7 controls>"
result = analyze(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/cross_framework_mapper.py
#!/usr/bin/env python3
"""cross_framework_mapper.py — Multi-framework control overlap computation.
Stdlib-only. Takes 1+ framework control libraries (control IDs + categories) and
computes overlap with mapping confidence (HIGH/MEDIUM/LOW) using a curated
ground-truth overlap dictionary distilled from published cross-walks:
- ISO 27001 Annex A <-> SOC 2 TSC (the densest known pair)
- ISO 27001 <-> ISO 42001 (info-sec reuse for AIMS)
- ISO 42001 <-> EU AI Act (Article 17 QMS satisfaction)
- GDPR <-> ISO 27001 (privacy controls overlap)
- ISO 13485 <-> FDA QSR (harmonised)
For each merged control, outputs the participating frameworks + a unified
evidence-requirement statement that satisfies all of them.
Deterministic ground-truth lookup. No LLM calls. No external dependencies.
Input schema (JSON):
{
"program": "Acme AI Inc. Compliance Program",
"enabled_frameworks": ["iso_27001", "soc_2", "iso_42001", "eu_ai_act", "gdpr"]
}
Usage:
python cross_framework_mapper.py # uses embedded 5-framework sample
python cross_framework_mapper.py path/to/program.json
python cross_framework_mapper.py program.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List, Set
SAMPLE: Dict[str, Any] = {
"program": "Acme AI Inc. Compliance Program",
"enabled_frameworks": [
"iso_27001", "soc_2", "iso_42001", "eu_ai_act", "gdpr",
"nist_csf", "nis2", "hipaa",
],
}
# Curated overlap database
# Each merged control: id, theme, evidence requirement, and per-framework mapping
# Mappings: framework_id -> (control_id, confidence)
# Confidence: H (high - same evidence satisfies), M (medium - evidence with overlay), L (low - concept overlap only)
MERGED_CONTROLS: List[Dict[str, Any]] = [
{
"id": "mc.access_control",
"theme": "Access control (identity, authentication, authorization)",
"evidence": "Documented access-control policy + access provisioning/de-provisioning procedure + quarterly access review records + RBAC matrix",
"mappings": {
"iso_27001": ("A.5.15 + A.8.2 + A.8.3", "H"),
"soc_2": ("CC6.1 + CC6.2 + CC6.3", "H"),
"iso_42001": ("A.4.4 (human resources for AI systems)", "M"),
"gdpr": ("Article 32(1)(b) integrity and confidentiality", "M"),
"nist_csf": ("PR.AA-01 + PR.AA-03 + PR.AA-05 (identities + authentication + authorization)", "H"),
"nis2": ("Article 21(2)(i) access control policies", "M"),
"hipaa": ("§164.308(a)(3) workforce security + §164.308(a)(4) information access management + §164.312(a)(1) access control", "H"),
},
},
{
"id": "mc.asset_inventory",
"theme": "Asset inventory and classification",
"evidence": "Asset register including AI systems + data classification scheme + ownership map",
"mappings": {
"iso_27001": ("A.5.9 + A.5.10 + A.5.12", "H"),
"soc_2": ("CC6.1 + CC3.2", "H"),
"iso_42001": ("A.4.2 (data) + A.4.3 (tooling)", "H"),
"gdpr": ("Article 30 (records of processing activities)", "M"),
"nist_csf": ("ID.AM-01 + ID.AM-02 + ID.AM-04 + ID.AM-05 (assets inventoried + classified)", "H"),
"nis2": ("Article 21(2)(b) policies on the use of risk-management measures (implicit: know your assets)", "M"),
"hipaa": ("§164.308(a)(1)(ii)(A) risk analysis (requires asset inventory) + §164.310(d) device + media controls", "M"),
},
},
{
"id": "mc.risk_management",
"theme": "Risk management process",
"evidence": "Risk methodology + risk register with severity matrix + risk treatment plan + residual-risk acceptance signoff",
"mappings": {
"iso_27001": ("Clause 6.1 + Clause 8.2", "H"),
"soc_2": ("CC3.1 + CC3.2 + CC3.4", "H"),
"iso_42001": ("Clause 6.1.2 + A.5", "H"),
"eu_ai_act": ("Article 9 (risk management system)", "M"),
"gdpr": ("Article 35 (DPIA where applicable)", "M"),
"nist_csf": ("GV.RM (risk management strategy) + ID.RA (risk assessment) + ID.IM (improvement)", "H"),
"nis2": ("Article 21(2)(a) risk analysis + Article 21(2)(b) policies on risk-management measures", "H"),
"hipaa": ("§164.308(a)(1)(ii)(A) risk analysis + §164.308(a)(1)(ii)(B) risk management", "H"),
},
},
{
"id": "mc.supplier_management",
"theme": "Third-party / supplier risk management",
"evidence": "Supplier inventory + due-diligence questionnaires + contractual security/privacy/AI clauses + periodic review records",
"mappings": {
"iso_27001": ("A.5.19 + A.5.20 + A.5.21 + A.5.22", "H"),
"soc_2": ("CC9.2", "H"),
"iso_42001": ("A.10.2 + A.10.6", "H"),
"eu_ai_act": ("Article 25 (responsibilities along the AI value chain)", "M"),
"gdpr": ("Article 28 (processor obligations)", "H"),
"nist_csf": ("GV.SC (cybersecurity supply chain risk management) + ID.SC", "H"),
"nis2": ("Article 21(2)(d) supply-chain security including security-related aspects of relationships with direct suppliers", "H"),
"hipaa": ("§164.308(b)(1) business associate contracts + §164.314(a) organizational requirements (BAAs)", "H"),
},
},
{
"id": "mc.incident_response",
"theme": "Incident response + notification",
"evidence": "Documented incident response procedure + severity definitions + escalation matrix + notification SLAs + post-incident reviews",
"mappings": {
"iso_27001": ("A.5.24 + A.5.25 + A.5.26 + A.5.27 + A.6.8", "H"),
"soc_2": ("CC7.3 + CC7.4 + CC7.5", "H"),
"iso_42001": ("A.8.4 (communication of AI incidents)", "M"),
"eu_ai_act": ("Article 73 (serious-incident reporting)", "M"),
"gdpr": ("Articles 33 + 34 (breach notification)", "H"),
"nist_csf": ("RS.MA + RS.AN + RS.RP + RS.CO (response: management, analysis, reporting, communication)", "H"),
"nis2": ("Article 23 incident notification (24h early warning / 72h notification / 1-month final report)", "H"),
"hipaa": ("§164.308(a)(6) security incident procedures + §164.400-414 Breach Notification Rule", "H"),
},
},
{
"id": "mc.monitoring_logging",
"theme": "Monitoring + logging",
"evidence": "Logging policy + tamper-evident logs + monitoring dashboards + retention compliant with longest applicable framework",
"mappings": {
"iso_27001": ("A.8.15 + A.8.16", "H"),
"soc_2": ("CC7.1 + CC7.2", "H"),
"iso_42001": ("A.9.3 + A.9.4", "M"),
"eu_ai_act": ("Article 12 (logging) + Article 72 (post-market monitoring)", "M"),
"nist_csf": ("DE.CM (continuous monitoring) + DE.AE (anomalies + events)", "H"),
"nis2": ("Article 21(2)(h) human resources security + ongoing monitoring expectations", "M"),
"hipaa": ("§164.308(a)(1)(ii)(D) information system activity review + §164.312(b) audit controls", "H"),
},
},
{
"id": "mc.change_management",
"theme": "Change management (system + model)",
"evidence": "Change approval workflow + version control + rollback procedure + change advisory board records",
"mappings": {
"iso_27001": ("A.8.32", "H"),
"soc_2": ("CC8.1", "H"),
"iso_42001": ("A.6.2.5 (deployment)", "M"),
"nist_csf": ("PR.PS (platform security including change-mgmt) + ID.IM-03 (improvements identified)", "H"),
"nis2": ("Article 21(2)(e) security in network and information systems acquisition, development and maintenance", "M"),
"hipaa": ("§164.308(a)(5)(ii)(B) protection from malicious software (implies controlled change) + §164.312(a)(1) access control during change", "M"),
},
},
{
"id": "mc.business_continuity",
"theme": "Business continuity and disaster recovery",
"evidence": "BCP/DRP documents + tested recovery objectives (RPO/RTO) + annual exercises + lessons learned",
"mappings": {
"iso_27001": ("A.5.29 + A.5.30 + A.8.13 + A.8.14", "H"),
"soc_2": ("A1.2 + A1.3", "H"),
"nist_csf": ("RC.RP (recovery planning) + RC.IM + RC.CO + ID.BE-05 (resilience requirements)", "H"),
"nis2": ("Article 21(2)(c) business continuity, such as backup management and disaster recovery, and crisis management", "H"),
"hipaa": ("§164.308(a)(7) contingency plan (incl. data backup + disaster recovery + emergency mode operation)", "H"),
},
},
{
"id": "mc.competence_training",
"theme": "Competence + awareness training",
"evidence": "Competence requirements per role + training plan + completion records + effectiveness verification",
"mappings": {
"iso_27001": ("A.6.3", "H"),
"soc_2": ("CC1.4 + CC2.2", "H"),
"iso_42001": ("Clause 7.2 + Clause 7.3 + A.4.4", "H"),
"eu_ai_act": ("Article 4 (AI literacy)", "M"),
"nist_csf": ("PR.AT (awareness + training)", "H"),
"nis2": ("Article 21(2)(g) basic cyber-hygiene practices and cybersecurity training", "H"),
"hipaa": ("§164.308(a)(5) security awareness and training", "H"),
},
},
{
"id": "mc.data_governance",
"theme": "Data governance + data quality",
"evidence": "Data inventory + provenance records + quality metrics + retention/deletion schedule + consent/lawful-basis records",
"mappings": {
"iso_27001": ("A.5.34 (privacy)", "M"),
"iso_42001": ("A.7 (full category)", "H"),
"eu_ai_act": ("Article 10 (data governance for high-risk)", "H"),
"gdpr": ("Articles 5 + 6 + 30", "H"),
"nist_csf": ("PR.DS (data security) + ID.AM-07 (data inventories) + GV.PO (policy)", "H"),
"nis2": ("Article 21(2)(j) policies and procedures (multi-factor + secure communications) implying data discipline", "M"),
"hipaa": ("§164.312(c)(1) integrity + §164.502 uses and disclosures of PHI + §164.514 de-identification", "H"),
},
},
{
"id": "mc.internal_audit",
"theme": "Internal audit programme",
"evidence": "Annual audit plan + auditor independence + findings tracking + closure verification",
"mappings": {
"iso_27001": ("Clause 9.2", "H"),
"soc_2": ("CC4.1", "H"),
"iso_42001": ("Clause 9.2", "H"),
"nist_csf": ("ID.IM (improvement processes including audits)", "M"),
"nis2": ("Article 21(2)(b) policies on the use of risk-management measures (implies periodic audit)", "M"),
"hipaa": ("§164.308(a)(1)(ii)(D) information system activity review + §164.308(a)(8) periodic evaluation", "H"),
},
},
{
"id": "mc.management_review",
"theme": "Management review",
"evidence": "Management review procedure + scheduled inputs + meeting records + action item tracking",
"mappings": {
"iso_27001": ("Clause 9.3", "H"),
"iso_42001": ("Clause 9.3", "H"),
"nist_csf": ("GV.OV (oversight) + GV.PO (organizational policy review)", "H"),
"nis2": ("Article 20 governance: management bodies must approve cybersecurity risk-management measures and oversee implementation", "H"),
"hipaa": ("§164.308(a)(2) assigned security responsibility + §164.308(a)(8) periodic evaluation by senior official", "M"),
},
},
{
"id": "mc.cryptography",
"theme": "Cryptography and key management",
"evidence": "Cryptographic policy + algorithm + key length standards + key rotation + HSM/KMS architecture + key custody records",
"mappings": {
"iso_27001": ("A.8.24", "H"),
"soc_2": ("CC6.1 + CC6.7", "H"),
"gdpr": ("Article 32(1)(a) pseudonymisation + encryption", "H"),
"nist_csf": ("PR.DS-02 (data-in-transit) + PR.DS-01 (data-at-rest) + PR.PS-05 (cryptography)", "H"),
"nis2": ("Article 21(2)(h) policies on the use of cryptography and, where appropriate, encryption", "H"),
"hipaa": ("§164.312(a)(2)(iv) encryption + decryption (addressable) + §164.312(e)(2)(ii) transmission encryption", "H"),
},
},
{
"id": "mc.secure_sdlc",
"theme": "Secure software development lifecycle",
"evidence": "Secure SDLC policy + threat modeling + code review records + SAST/DAST scanning + vulnerability triage",
"mappings": {
"iso_27001": ("A.8.25 + A.8.26 + A.8.27 + A.8.28 + A.8.29 + A.8.30 + A.8.31", "H"),
"soc_2": ("CC8.1 + CC7.1", "H"),
"iso_42001": ("A.6.2.2 + A.6.2.3 + A.6.2.4 (AI-specific SDLC)", "M"),
"nist_csf": ("PR.PS (platform security including secure development) + ID.RA-08 (vulnerabilities identified)", "H"),
"nis2": ("Article 21(2)(e) security in network and information systems acquisition, development and maintenance", "H"),
},
},
{
"id": "mc.vulnerability_mgmt",
"theme": "Vulnerability + patch management",
"evidence": "Vulnerability scanning schedule + patch SLAs by severity + exception tracking + remediation evidence",
"mappings": {
"iso_27001": ("A.8.7 + A.8.8 + A.8.9", "H"),
"soc_2": ("CC7.1 + CC7.2 + CC7.4", "H"),
"nist_csf": ("ID.RA-01 + ID.RA-08 (vulnerabilities) + PR.PS-02 (patching)", "H"),
"nis2": ("Article 21(2)(f) policies and procedures to assess the effectiveness of cybersecurity risk-management measures + vulnerability handling", "H"),
"hipaa": ("§164.308(a)(5)(ii)(B) protection from malicious software + §164.308(a)(1)(ii)(A) periodic risk analysis (covers vulnerability identification)", "M"),
},
},
{
"id": "mc.physical_security",
"theme": "Physical security and environmental controls",
"evidence": "Facility access controls + visitor log + environmental monitoring + tamper-evident seals on critical assets",
"mappings": {
"iso_27001": ("A.7.1 + A.7.2 + A.7.3 + A.7.4 + A.7.5 + A.7.6 + A.7.7 + A.7.8", "H"),
"soc_2": ("CC6.4 + CC6.5", "H"),
"nist_csf": ("PR.AA-06 (physical access) + PR.PS-04 (physical resource security)", "H"),
"hipaa": ("§164.310(a)(1) facility access controls + §164.310(b) workstation use + §164.310(c) workstation security + §164.310(d) device + media controls", "H"),
},
},
{
"id": "mc.data_protection_privacy",
"theme": "Personal data protection (privacy by design)",
"evidence": "Privacy policy + lawful-basis register + retention/deletion schedule + DPIA records + data-subject rights workflow + DPO appointment (where required)",
"mappings": {
"iso_27001": ("A.5.34", "H"),
"iso_42001": ("A.7.6 (data privacy considerations)", "M"),
"gdpr": ("Articles 5 + 6 + 24 + 25 + 30 + 35 + 38", "H"),
"nist_csf": ("GV.PO + PR.DS (data security)", "M"),
"hipaa": ("§164.502 uses and disclosures (Privacy Rule) + §164.520 notice of privacy practices + §164.530 administrative requirements", "H"),
},
},
{
"id": "mc.documentation_control",
"theme": "Documented information control",
"evidence": "Document control procedure + version control + approval workflow + retention + obsolete-doc handling",
"mappings": {
"iso_27001": ("Clause 7.5", "H"),
"soc_2": ("CC4.1 + CC5.1", "H"),
"iso_42001": ("Clause 7.5", "H"),
"nist_csf": ("GV.PO (policy + documentation) + ID.AM-08 (system and data are documented)", "H"),
"nis2": ("Article 21(1) documented cybersecurity risk-management measures", "H"),
"hipaa": ("§164.316 policies, procedures, and documentation requirements (retention 6 years)", "H"),
},
},
{
"id": "mc.continual_improvement",
"theme": "Continual improvement + CAPA",
"evidence": "Nonconformity tracking + root-cause analysis + corrective action plans + effectiveness verification + trend analysis",
"mappings": {
"iso_27001": ("Clause 10.1 + 10.2", "H"),
"soc_2": ("CC4.1 + CC4.2 + CC5.3", "H"),
"iso_42001": ("Clause 10.1 + 10.2", "H"),
"nist_csf": ("ID.IM-01 + ID.IM-02 + ID.IM-03 (improvements identified, evaluated, executed)", "H"),
"hipaa": ("§164.306(e) review + modify (security measures must be reviewed and modified as needed)", "M"),
},
},
]
def merged_in_scope(enabled: Set[str]) -> List[Dict[str, Any]]:
"""Return merged controls where at least 1 enabled framework maps to them."""
out: List[Dict[str, Any]] = []
for mc in MERGED_CONTROLS:
active_maps = {fid: m for fid, m in mc["mappings"].items() if fid in enabled}
if active_maps:
out.append({
"id": mc["id"],
"theme": mc["theme"],
"evidence": mc["evidence"],
"frameworks_count": len(active_maps),
"frameworks": active_maps,
})
return out
def overlap_summary(merged: List[Dict[str, Any]], enabled: Set[str]) -> Dict[str, Any]:
"""Compute per-framework coverage and per-pair overlap."""
coverage: Dict[str, int] = {f: 0 for f in enabled}
high_confidence: Dict[str, int] = {f: 0 for f in enabled}
for mc in merged:
for fid in mc["frameworks"]:
coverage[fid] += 1
_, conf = mc["frameworks"][fid]
if conf == "H":
high_confidence[fid] += 1
multi_framework = [mc for mc in merged if mc["frameworks_count"] >= 2]
high_reuse = [mc for mc in merged if mc["frameworks_count"] >= 3]
return {
"total_merged_controls_in_scope": len(merged),
"per_framework_coverage": coverage,
"per_framework_high_confidence": high_confidence,
"multi_framework_count": len(multi_framework),
"high_reuse_count_3plus_frameworks": len(high_reuse),
}
def analyze(payload: Dict[str, Any]) -> Dict[str, Any]:
enabled = set(payload.get("enabled_frameworks", []))
merged = merged_in_scope(enabled)
summary = overlap_summary(merged, enabled)
return {
"program": payload.get("program"),
"enabled_frameworks": sorted(enabled),
"summary": summary,
"merged_controls": sorted(merged, key=lambda m: -m["frameworks_count"]),
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("COMPLIANCE OS — CROSS-FRAMEWORK CONTROL MAPPING")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Program: {r['program']}")
lines.append(f"Enabled frameworks ({len(r['enabled_frameworks'])}): {', '.join(r['enabled_frameworks'])}")
lines.append("")
s = r["summary"]
lines.append(f"Merged controls in scope: {s['total_merged_controls_in_scope']}")
lines.append(f"Multi-framework controls (≥ 2): {s['multi_framework_count']}")
lines.append(f"High-reuse controls (≥ 3 frameworks): {s['high_reuse_count_3plus_frameworks']}")
lines.append("")
lines.append("Per-framework coverage in merged catalogue:")
for fid in r["enabled_frameworks"]:
cov = s["per_framework_coverage"].get(fid, 0)
hi = s["per_framework_high_confidence"].get(fid, 0)
lines.append(f" {fid:15s} {cov} mappings ({hi} HIGH confidence)")
lines.append("")
lines.append("-" * 72)
lines.append("MERGED CONTROLS (sorted by reuse leverage):")
lines.append("")
for mc in r["merged_controls"]:
lines.append(f" [{mc['id']}] {mc['theme']} ({mc['frameworks_count']} frameworks)")
lines.append(f" Evidence: {mc['evidence']}")
for fid, (ctrl, conf) in mc["frameworks"].items():
conf_label = {"H": "HIGH ", "M": "MED ", "L": "LOW "}[conf]
lines.append(f" [{conf_label}] {fid:12s} -> {ctrl}")
lines.append("")
lines.append("-" * 72)
lines.append("CONFIDENCE LEGEND:")
lines.append(" HIGH — same evidence satisfies both (direct overlap)")
lines.append(" MED — existing evidence with overlay")
lines.append(" LOW — concept overlap; mostly new artefact required")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Multi-framework control overlap computation.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to program JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: ISO 27001 + SOC 2 + ISO 42001 + EU AI Act + GDPR>"
result = analyze(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/evidence_pool_generator.py
#!/usr/bin/env python3
"""evidence_pool_generator.py — Consolidated evidence checklist across enabled frameworks.
Stdlib-only. Given a multi-framework compliance program config, produces a unified
evidence pool that maps each evidence artefact to all the (framework, control)
tuples it satisfies. Each artefact gets a reuse-leverage score = number of
distinct (framework, control) tuples satisfied.
Deterministic. No LLM calls. No external dependencies. Uses a curated evidence
catalogue distilled from ISO 27001, ISO 42001, SOC 2, GDPR, EU AI Act published
guidance.
Input schema (JSON):
{
"program": "Acme AI Inc. compliance program",
"enabled_frameworks": ["iso_27001", "soc_2", "iso_42001", "eu_ai_act", "gdpr"],
"audit_cycle_year": "year_1"
}
Usage:
python evidence_pool_generator.py
python evidence_pool_generator.py path/to/program.json
python evidence_pool_generator.py program.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"program": "Acme AI Inc. compliance program",
"enabled_frameworks": ["iso_27001", "soc_2", "iso_42001", "eu_ai_act", "gdpr"],
"audit_cycle_year": "year_1",
}
# Evidence catalogue: each artefact + the (framework, control) tuples it satisfies
# + acquisition cost (low / medium / high) + retention requirement (months)
EVIDENCE_CATALOG: List[Dict[str, Any]] = [
{
"id": "ev.access_review_quarterly",
"title": "Quarterly access review records (privileged + general access)",
"satisfies": [
("iso_27001", "A.5.15"), ("iso_27001", "A.8.2"), ("iso_27001", "A.8.3"),
("soc_2", "CC6.1"), ("soc_2", "CC6.2"), ("soc_2", "CC6.3"),
("iso_42001", "A.4.4"),
("gdpr", "Article 32(1)(b)"),
],
"acquisition_cost": "low",
"retention_months": 36,
"owner": "IT / Security",
},
{
"id": "ev.asset_register",
"title": "Asset register with AI systems + data classification",
"satisfies": [
("iso_27001", "A.5.9"), ("iso_27001", "A.5.10"), ("iso_27001", "A.5.12"),
("soc_2", "CC6.1"), ("soc_2", "CC3.2"),
("iso_42001", "A.4.2"), ("iso_42001", "A.4.3"),
("gdpr", "Article 30"),
],
"acquisition_cost": "medium",
"retention_months": 36,
"owner": "Security / DPO",
},
{
"id": "ev.risk_register",
"title": "Risk register with severity matrix + treatment + residual signoff",
"satisfies": [
("iso_27001", "Clause 6.1"), ("iso_27001", "Clause 8.2"),
("soc_2", "CC3.1"), ("soc_2", "CC3.2"), ("soc_2", "CC3.4"),
("iso_42001", "Clause 6.1.2"), ("iso_42001", "A.5"),
("eu_ai_act", "Article 9"),
("gdpr", "Article 35"),
],
"acquisition_cost": "high",
"retention_months": 36,
"owner": "Compliance officer",
},
{
"id": "ev.supplier_inventory_reviews",
"title": "Supplier inventory + annual reviews + signed DPAs",
"satisfies": [
("iso_27001", "A.5.19"), ("iso_27001", "A.5.20"), ("iso_27001", "A.5.21"),
("soc_2", "CC9.2"),
("iso_42001", "A.10.2"), ("iso_42001", "A.10.6"),
("eu_ai_act", "Article 25"),
("gdpr", "Article 28"),
],
"acquisition_cost": "medium",
"retention_months": 36,
"owner": "Procurement / Compliance",
},
{
"id": "ev.incident_log_postmortems",
"title": "Incident log + severity classifications + post-incident reviews + notifications sent",
"satisfies": [
("iso_27001", "A.5.24"), ("iso_27001", "A.5.25"), ("iso_27001", "A.5.26"),
("iso_27001", "A.5.27"), ("iso_27001", "A.6.8"),
("soc_2", "CC7.3"), ("soc_2", "CC7.4"), ("soc_2", "CC7.5"),
("iso_42001", "A.8.4"),
("eu_ai_act", "Article 73"),
("gdpr", "Article 33"), ("gdpr", "Article 34"),
],
"acquisition_cost": "medium",
"retention_months": 36,
"owner": "Security / IR team",
},
{
"id": "ev.logs_aggregated",
"title": "Tamper-evident logs centralized with retention",
"satisfies": [
("iso_27001", "A.8.15"), ("iso_27001", "A.8.16"),
("soc_2", "CC7.1"), ("soc_2", "CC7.2"),
("iso_42001", "A.9.3"), ("iso_42001", "A.9.4"),
("eu_ai_act", "Article 12"),
],
"acquisition_cost": "high",
"retention_months": 12,
"owner": "Platform / SRE",
},
{
"id": "ev.change_records",
"title": "Change approval records + rollback procedure + post-implementation reviews",
"satisfies": [
("iso_27001", "A.8.32"),
("soc_2", "CC8.1"),
("iso_42001", "A.6.2.5"),
],
"acquisition_cost": "low",
"retention_months": 24,
"owner": "Engineering / Platform",
},
{
"id": "ev.bcp_dr_exercises",
"title": "BCP/DRP exercise records + RPO/RTO validation",
"satisfies": [
("iso_27001", "A.5.29"), ("iso_27001", "A.5.30"),
("iso_27001", "A.8.13"), ("iso_27001", "A.8.14"),
("soc_2", "A1.2"), ("soc_2", "A1.3"),
],
"acquisition_cost": "high",
"retention_months": 36,
"owner": "Platform / SRE",
},
{
"id": "ev.training_records",
"title": "Competence requirements per role + training completion + effectiveness verification",
"satisfies": [
("iso_27001", "A.6.3"),
("soc_2", "CC1.4"), ("soc_2", "CC2.2"),
("iso_42001", "Clause 7.2"), ("iso_42001", "Clause 7.3"),
("eu_ai_act", "Article 4"),
],
"acquisition_cost": "medium",
"retention_months": 36,
"owner": "HR / People Ops",
},
{
"id": "ev.data_inventory_consent",
"title": "Data inventory + provenance + retention + consent / lawful-basis register",
"satisfies": [
("iso_27001", "A.5.34"),
("iso_42001", "A.7.2"), ("iso_42001", "A.7.3"), ("iso_42001", "A.7.4"),
("iso_42001", "A.7.5"), ("iso_42001", "A.7.6"),
("eu_ai_act", "Article 10"),
("gdpr", "Article 5"), ("gdpr", "Article 6"), ("gdpr", "Article 30"),
],
"acquisition_cost": "high",
"retention_months": 60,
"owner": "DPO / Data team",
},
{
"id": "ev.internal_audit_records",
"title": "Internal audit plan + auditor independence records + findings tracking",
"satisfies": [
("iso_27001", "Clause 9.2"),
("soc_2", "CC4.1"),
("iso_42001", "Clause 9.2"),
],
"acquisition_cost": "medium",
"retention_months": 36,
"owner": "Compliance officer",
},
{
"id": "ev.management_review_records",
"title": "Management review schedule + meeting records + action item tracking",
"satisfies": [
("iso_27001", "Clause 9.3"),
("iso_42001", "Clause 9.3"),
],
"acquisition_cost": "low",
"retention_months": 36,
"owner": "Compliance officer + Exec",
},
{
"id": "ev.policy_set",
"title": "Policy set: AI, info-sec, privacy, code-of-conduct (signed + reviewed annually)",
"satisfies": [
("iso_27001", "A.5.1"),
("soc_2", "CC1.1"), ("soc_2", "CC1.2"),
("iso_42001", "Clause 5.2"), ("iso_42001", "A.2.2"), ("iso_42001", "A.2.3"),
("eu_ai_act", "Article 17(1)(a)"),
("gdpr", "Article 24"),
],
"acquisition_cost": "medium",
"retention_months": 60,
"owner": "Compliance officer + Exec",
},
{
"id": "ev.crypto_records",
"title": "Crypto policy + algorithm/key-length standards + key rotation records",
"satisfies": [
("iso_27001", "A.8.24"),
("soc_2", "CC6.1"), ("soc_2", "CC6.7"),
("gdpr", "Article 32(1)(a)"),
],
"acquisition_cost": "medium",
"retention_months": 36,
"owner": "Security",
},
{
"id": "ev.vuln_scans_patch",
"title": "Vulnerability scan results + patch SLAs + remediation evidence",
"satisfies": [
("iso_27001", "A.8.7"), ("iso_27001", "A.8.8"), ("iso_27001", "A.8.9"),
("soc_2", "CC7.1"), ("soc_2", "CC7.2"), ("soc_2", "CC7.4"),
],
"acquisition_cost": "medium",
"retention_months": 24,
"owner": "Security",
},
]
def filter_by_enabled(catalog: List[Dict[str, Any]], enabled: List[str]) -> List[Dict[str, Any]]:
"""Filter satisfaction tuples to enabled frameworks."""
enabled_set = set(enabled)
out = []
for ev in catalog:
active = [(f, c) for (f, c) in ev["satisfies"] if f in enabled_set]
if not active:
continue
leverage = len(active)
frameworks_satisfied = sorted({f for f, _ in active})
record = {**ev, "active_satisfaction": active, "reuse_leverage": leverage,
"frameworks_satisfied": frameworks_satisfied}
out.append(record)
out.sort(key=lambda x: (-x["reuse_leverage"], x["title"]))
return out
def analyze(payload: Dict[str, Any]) -> Dict[str, Any]:
enabled = payload.get("enabled_frameworks", [])
artefacts = filter_by_enabled(EVIDENCE_CATALOG, enabled)
total_satisfactions = sum(a["reuse_leverage"] for a in artefacts)
by_cost: Dict[str, int] = {"low": 0, "medium": 0, "high": 0}
by_owner: Dict[str, int] = {}
for a in artefacts:
by_cost[a["acquisition_cost"]] += 1
by_owner[a["owner"]] = by_owner.get(a["owner"], 0) + 1
# High-leverage artefacts (satisfy ≥ 5 mappings)
high_leverage = [a for a in artefacts if a["reuse_leverage"] >= 5]
return {
"program": payload.get("program"),
"enabled_frameworks": enabled,
"audit_cycle_year": payload.get("audit_cycle_year"),
"artefact_count": len(artefacts),
"total_satisfactions_across_artefacts": total_satisfactions,
"high_leverage_count": len(high_leverage),
"by_acquisition_cost": by_cost,
"by_owner": by_owner,
"artefacts": artefacts,
}
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("COMPLIANCE OS — UNIFIED EVIDENCE POOL")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Program: {r['program']}")
lines.append(f"Enabled frameworks: {', '.join(r['enabled_frameworks'])}")
lines.append(f"Audit cycle phase: {r['audit_cycle_year']}")
lines.append(f"Artefacts in scope: {r['artefact_count']}")
lines.append(f"Total (framework, control) satisfactions: {r['total_satisfactions_across_artefacts']}")
lines.append(f"High-leverage artefacts (≥ 5 mappings): {r['high_leverage_count']}")
lines.append("")
lines.append(f"By acquisition cost: low={r['by_acquisition_cost']['low']} "
f"medium={r['by_acquisition_cost']['medium']} high={r['by_acquisition_cost']['high']}")
lines.append(f"By owner: {dict(r['by_owner'])}")
lines.append("")
lines.append("-" * 72)
lines.append("ARTEFACTS (sorted by reuse leverage — highest first):")
lines.append("")
for a in r["artefacts"]:
lines.append(f" [{a['id']}] {a['title']}")
lines.append(f" Leverage: {a['reuse_leverage']} mappings across {len(a['frameworks_satisfied'])} frameworks ({', '.join(a['frameworks_satisfied'])})")
lines.append(f" Owner: {a['owner']} | Cost: {a['acquisition_cost']} | Retention: {a['retention_months']} months")
lines.append(f" Satisfies:")
for fid, ctrl in a["active_satisfaction"]:
lines.append(f" - {fid:12s} -> {ctrl}")
lines.append("")
lines.append("-" * 72)
lines.append("REUSE-LEVERAGE GUIDANCE:")
lines.append(" Build high-leverage artefacts first (single evidence -> ≥ 5 framework controls).")
lines.append(" High-leverage examples (depend on enabled frameworks): risk register, supplier inventory, incident log,")
lines.append(" data inventory + consent, policy set, training records.")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Unified evidence pool generator across compliance frameworks.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to program JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
payload = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
payload = SAMPLE
source = "<embedded sample: 5 enabled frameworks, year 1>"
result = analyze(payload)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/framework_selector.py
#!/usr/bin/env python3
"""framework_selector.py — Multi-framework compliance applicability selector.
Stdlib-only. Takes a company profile and returns the applicable compliance frameworks
ranked by priority + dependency graph. Supports 9 frameworks:
- ISO 27001 (info-sec ISMS)
- ISO 13485 (medical device QMS)
- ISO 42001 (AI management system)
- ISO 14971 (medical device risk mgmt)
- EU AI Act (Regulation 2024/1689)
- EU MDR 2017/745 (medical device regulation)
- GDPR (Regulation 2016/679)
- SOC 2 (Trust Services Criteria)
- FDA QSR (21 CFR 820)
Deterministic decision tree. No LLM calls. No external dependencies.
Input schema (JSON):
{
"company": "Acme AI Inc.",
"industry": "saas", # saas | medical_device | financial | other
"products_include_ai": true,
"ai_high_risk_per_eu": true, # falls under Annex III, Article 6
"deploys_ai_in_eu": true,
"products_are_medical_devices": false,
"sells_to_eu_customers": true,
"sells_to_us_customers": true,
"sells_to_enterprise_b2b": true,
"processes_personal_data": true,
"processes_eu_personal_data": true,
"headcount": 80,
"stage": "series_b"
}
Usage:
python framework_selector.py # uses embedded mid-stage AI SaaS sample
python framework_selector.py path/to/profile.json
python framework_selector.py profile.json --output json
"""
import argparse
import json
import sys
from typing import Any, Dict, List
SAMPLE: Dict[str, Any] = {
"company": "Acme AI Inc.",
"industry": "saas",
"products_include_ai": True,
"ai_high_risk_per_eu": True,
"deploys_ai_in_eu": True,
"products_are_medical_devices": False,
"sells_to_eu_customers": True,
"sells_to_us_customers": True,
"sells_to_enterprise_b2b": True,
"processes_personal_data": True,
"processes_eu_personal_data": True,
"headcount": 80,
"stage": "series_b",
# Phase 3 additions (defaults false; sample profile does not trigger HIPAA / NIS2 / CSF)
"processes_phi": False,
"us_healthcare_covered_entity": False,
"us_healthcare_business_associate": False,
"nis2_essential_entity": False,
"nis2_important_entity": False,
"adopts_nist_csf": False,
"us_government_contractor": False,
}
# Framework catalogue (id, name, type, certifiable)
FRAMEWORKS = {
"iso_27001": {"name": "ISO/IEC 27001:2022", "type": "management_system", "certifiable": True, "binding": False},
"iso_13485": {"name": "ISO 13485:2016", "type": "management_system", "certifiable": True, "binding": False},
"iso_42001": {"name": "ISO/IEC 42001:2023", "type": "management_system", "certifiable": True, "binding": False},
"iso_14971": {"name": "ISO 14971:2019", "type": "process_standard", "certifiable": False, "binding": False},
"eu_ai_act": {"name": "Regulation (EU) 2024/1689 (AI Act)", "type": "regulation", "certifiable": False, "binding": True},
"eu_mdr_745": {"name": "Regulation (EU) 2017/745 (MDR)", "type": "regulation", "certifiable": False, "binding": True},
"gdpr": {"name": "Regulation (EU) 2016/679 (GDPR)", "type": "regulation", "certifiable": False, "binding": True},
"soc_2": {"name": "AICPA SOC 2 Trust Services", "type": "attestation", "certifiable": True, "binding": False},
"fda_qsr": {"name": "FDA 21 CFR 820 (QSR)", "type": "regulation", "certifiable": False, "binding": True},
# Phase 3 additions
"nist_csf": {"name": "NIST Cybersecurity Framework 2.0", "type": "framework_profile", "certifiable": False, "binding": False},
"nis2": {"name": "Directive (EU) 2022/2555 (NIS2)", "type": "regulation", "certifiable": False, "binding": True},
"hipaa": {"name": "HIPAA Security + Privacy + Breach Notification Rules", "type": "regulation", "certifiable": False, "binding": True},
}
# Dependency graph: framework X benefits from framework Y as prerequisite
DEPENDENCIES = {
"iso_42001": ["iso_27001"], # AIMS reuses ISMS heavily
"iso_13485": ["iso_14971"], # QMS uses risk mgmt
"eu_mdr_745": ["iso_13485", "iso_14971"],
"eu_ai_act": ["iso_42001"], # voluntary AIMS satisfies parts of Article 17
"soc_2": ["iso_27001"], # ISO 27001 controls map to SOC 2 TSC
"fda_qsr": ["iso_13485"], # QSR mostly harmonised with 13485
# Phase 3 additions
"nist_csf": [], # voluntary framework; no prereqs
"nis2": ["iso_27001"], # NIS2 risk-mgmt + reporting maps to 27001 controls
"hipaa": ["iso_27001"], # HIPAA Security Rule overlaps ISO 27001 Annex A
}
def select_frameworks(profile: Dict[str, Any]) -> List[str]:
selected: List[str] = []
# GDPR — any EU personal data
if profile.get("processes_eu_personal_data") or (
profile.get("processes_personal_data") and profile.get("sells_to_eu_customers")
):
selected.append("gdpr")
# ISO 27001 — enterprise B2B / mature SaaS
if profile.get("sells_to_enterprise_b2b") or profile.get("stage") in ("series_a", "series_b", "series_c", "growth"):
selected.append("iso_27001")
# SOC 2 — US enterprise B2B
if profile.get("sells_to_us_customers") and profile.get("sells_to_enterprise_b2b"):
selected.append("soc_2")
# ISO 42001 — any AI in products
if profile.get("products_include_ai"):
selected.append("iso_42001")
# EU AI Act — AI deployed in EU
if profile.get("products_include_ai") and (
profile.get("deploys_ai_in_eu") or profile.get("sells_to_eu_customers")
):
selected.append("eu_ai_act")
# ISO 13485 + 14971 — medical device
if profile.get("products_are_medical_devices"):
selected.append("iso_13485")
selected.append("iso_14971")
# EU MDR — medical device sold in EU
if profile.get("sells_to_eu_customers"):
selected.append("eu_mdr_745")
# FDA QSR — medical device sold in US
if profile.get("sells_to_us_customers"):
selected.append("fda_qsr")
# HIPAA — any US healthcare PHI processing
if profile.get("processes_phi") or profile.get("us_healthcare_covered_entity") or profile.get("us_healthcare_business_associate"):
selected.append("hipaa")
# NIS2 — operates in EU as essential or important entity per Annex I/II of Directive 2022/2555
if profile.get("nis2_essential_entity") or profile.get("nis2_important_entity"):
selected.append("nis2")
# NIST CSF — voluntary; recommended for any org with cybersecurity programme (esp. US gov-adjacent)
if profile.get("adopts_nist_csf") or profile.get("us_government_contractor"):
selected.append("nist_csf")
return selected
def annotate(profile: Dict[str, Any]) -> Dict[str, Any]:
selected = select_frameworks(profile)
# Build dependency notes
dep_notes: List[Dict[str, Any]] = []
for fid in selected:
deps = DEPENDENCIES.get(fid, [])
in_program = [d for d in deps if d in selected]
missing = [d for d in deps if d not in selected]
if in_program or missing:
dep_notes.append({
"framework": fid,
"satisfied_dependencies": in_program,
"missing_dependencies": missing,
})
# Priority ranking — bindings first, then certifiable, then reference
def priority(fid: str) -> int:
f = FRAMEWORKS[fid]
if f["binding"]:
return 0
if f["certifiable"]:
return 1
return 2
ranked = sorted(selected, key=priority)
return {
"company": profile.get("company"),
"industry": profile.get("industry"),
"applicable_frameworks": [
{"id": fid, **FRAMEWORKS[fid]} for fid in ranked
],
"framework_count": len(ranked),
"binding_count": sum(1 for fid in ranked if FRAMEWORKS[fid]["binding"]),
"certifiable_count": sum(1 for fid in ranked if FRAMEWORKS[fid]["certifiable"]),
"dependency_notes": dep_notes,
"rationale": _rationale(profile, ranked),
}
def _rationale(profile: Dict[str, Any], selected: List[str]) -> List[str]:
notes = []
if "gdpr" in selected:
notes.append("GDPR: EU personal data processed; binding regardless of certifiable choice.")
if "iso_27001" in selected:
notes.append("ISO 27001: enterprise B2B procurement frequently requires; foundation for AIMS + SOC 2.")
if "soc_2" in selected:
notes.append("SOC 2: US enterprise B2B procurement requires Type II audit; overlap with ISO 27001 ~75%.")
if "iso_42001" in selected:
notes.append("ISO 42001: AI in products; voluntary management system; satisfies Article 17 EU AI Act QMS.")
if "eu_ai_act" in selected:
notes.append("EU AI Act: AI deployed in EU; binding; Article 5 prohibitions in force; high-risk obligations 2 Aug 2026.")
if "iso_13485" in selected:
notes.append("ISO 13485: medical device manufacturer; required for MDR / FDA submissions.")
if "iso_14971" in selected:
notes.append("ISO 14971: medical device risk management; harmonised under MDR.")
if "eu_mdr_745" in selected:
notes.append("EU MDR 745: medical device sold in EU; binding; mandatory CE marking.")
if "fda_qsr" in selected:
notes.append("FDA QSR: medical device sold in US; binding; FDA quality system regulation.")
if "hipaa" in selected:
notes.append("HIPAA: processes US PHI; binding Security Rule (45 CFR 164 Subpart C) + Privacy Rule + Breach Notification.")
if "nis2" in selected:
notes.append("NIS2: essential or important entity in EU per Directive 2022/2555 Annex I/II; binding; cybersecurity + incident reporting obligations.")
if "nist_csf" in selected:
notes.append("NIST CSF 2.0: voluntary cybersecurity framework; recommended for US gov-adjacent orgs; cross-walks ISO 27001 + SOC 2 Common Criteria.")
return notes
def render_text(r: Dict[str, Any], source: str) -> str:
lines = []
lines.append("=" * 72)
lines.append("COMPLIANCE OS — APPLICABLE FRAMEWORKS")
lines.append(f"Source: {source}")
lines.append("=" * 72)
lines.append("")
lines.append(f"Company: {r['company']}")
lines.append(f"Industry: {r['industry']}")
lines.append(f"Applicable frameworks: {r['framework_count']} "
f"({r['binding_count']} binding + {r['certifiable_count']} certifiable)")
lines.append("")
lines.append("-" * 72)
lines.append("RANKED FRAMEWORKS (binding > certifiable > reference):")
lines.append("")
for f in r["applicable_frameworks"]:
kind = []
if f["binding"]:
kind.append("BINDING")
if f["certifiable"]:
kind.append("CERTIFIABLE")
kind_str = " | ".join(kind) if kind else "REFERENCE"
lines.append(f" [{kind_str:25s}] {f['name']:42s} ({f['id']})")
lines.append("")
lines.append("-" * 72)
lines.append("RATIONALE:")
for r_note in r["rationale"]:
lines.append(f" - {r_note}")
lines.append("")
if r["dependency_notes"]:
lines.append("-" * 72)
lines.append("DEPENDENCIES:")
for d in r["dependency_notes"]:
if d["satisfied_dependencies"]:
lines.append(f" {d['framework']} satisfied by: {', '.join(d['satisfied_dependencies'])}")
if d["missing_dependencies"]:
lines.append(f" {d['framework']} missing dependency: {', '.join(d['missing_dependencies'])} (consider adding)")
lines.append("")
lines.append("-" * 72)
lines.append("DECISION RULES:")
lines.append(" GDPR: any EU personal data processed -> mandatory")
lines.append(" ISO 27001: enterprise B2B procurement requirement; foundation for AIMS + SOC 2")
lines.append(" SOC 2: US enterprise B2B procurement requirement; overlap ~75% with ISO 27001")
lines.append(" ISO 42001: AI in products; voluntary AIMS; satisfies parts of Article 17 AI Act")
lines.append(" EU AI Act: AI in EU; binding; phased application through 2027")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Multi-framework compliance applicability selector.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("path", nargs="?", help="Path to company profile JSON (uses embedded sample if omitted)")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.path:
try:
with open(args.path, "r", encoding="utf-8") as f:
profile = json.load(f)
source = args.path
except (IOError, OSError) as e:
print(f"error: could not read {args.path}: {e}", file=sys.stderr)
return 1
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.path}: {e}", file=sys.stderr)
return 1
else:
profile = SAMPLE
source = "<embedded sample: mid-stage AI SaaS, US+EU customers, B2B>"
result = annotate(profile)
if args.output == "json":
print(json.dumps({"source": source, **result}, indent=2))
else:
print(render_text(result, source))
return 0
if __name__ == "__main__":
sys.exit(main())
Skill chuyển hướng các yêu cầu viết nội dung kiểu cũ sang chuyên gia phù hợp như sản xuất nội dung hoặc lập chiến lược.
---
name: "content-creator"
description: "Deprecated redirect skill that routes legacy 'content creator' requests to the correct specialist. Use when a user invokes 'content creator', asks to write a blog post, article, guide, or brand voice analysis (routes to content-production), or asks to plan content, build a topic cluster, or create a content calendar (routes to content-strategy). Does not handle requests directly — identifies user intent and redirects to content-production for writing/SEO/brand-voice tasks or content-strategy for planning tasks."
license: MIT
metadata:
version: 2.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
status: deprecated
---
# Content Creator → Redirected
> **This skill has been split into two specialist skills.** Use the one that matches your intent:
| You want to... | Use this instead |
|----------------|-----------------|
| **Write** a blog post, article, or guide | [content-production](../content-production/) |
| **Plan** what content to create, topic clusters, calendar | [content-strategy](../content-strategy/) |
| **Analyze brand voice** | [content-production](../content-production/) (includes `brand_voice_analyzer.py`) |
| **Optimize SEO** for existing content | [content-production](../content-production/) (includes `seo_optimizer.py`) |
| **Create social media content** | [social-content](../social-content/) |
## Why the Change
The original `content-creator` tried to do everything: planning, writing, SEO, social, brand voice. That made it a jack of all trades. The specialist skills do each job better:
- **content-production** — Full pipeline: research → brief → draft → optimize → publish. Includes all Python tools from the original content-creator.
- **content-strategy** — Strategic planning: topic clusters, keyword research, content calendars, prioritization frameworks.
## Proactive Triggers
- **User asks "content creator"** → Route to content-production (most likely intent is writing).
- **User asks "content plan" or "what should I write"** → Route to content-strategy.
## Output Artifacts
| When you ask for... | Routed to... |
|---------------------|-------------|
| "Write a blog post" | content-production |
| "Content calendar" | content-strategy |
| "Brand voice analysis" | content-production (`brand_voice_analyzer.py`) |
| "SEO optimization" | content-production (`seo_optimizer.py`) |
## Communication
This is a redirect skill. Route the user to the correct specialist — don't attempt to handle the request here.
## Related Skills
- **content-production**: Full content execution pipeline (successor).
- **content-strategy**: Content planning and topic selection (successor).
- **content-humanizer**: Post-processing AI content to sound authentic.
- **marketing-context**: Foundation context that both successors read.
FILE:assets/content_calendar_template.md
# Content Calendar Template - [Month Year]
## Monthly Goals
- **Traffic Goal**:
- **Lead Generation Goal**:
- **Engagement Goal**:
- **Key Campaign**:
## Week 1: [Date Range]
### Monday [Date]
**Platform**: Blog
**Topic**:
**Keywords**:
**Status**: [ ] Planned [ ] Written [ ] Reviewed [ ] Published
**Owner**:
**Notes**:
**Platform**: LinkedIn
**Type**: Article Share
**Caption**:
**Hashtags**:
**Time**: 10:00 AM
### Tuesday [Date]
**Platform**: Instagram
**Type**: Carousel
**Topic**:
**Visuals**: [ ] Created [ ] Approved
**Caption**:
**Hashtags**:
**Time**: 12:00 PM
### Wednesday [Date]
**Platform**: Email Newsletter
**Subject Line**:
**Segment**:
**CTA**:
**Status**: [ ] Drafted [ ] Designed [ ] Scheduled
### Thursday [Date]
**Platform**: Twitter/X
**Type**: Thread
**Topic**:
**Thread Length**:
**Media**: [ ] Images [ ] GIFs [ ] None
**Time**: 2:00 PM
### Friday [Date]
**Platform**: Multi-channel
**Campaign**:
**Assets Needed**:
- [ ] Blog post
- [ ] Social graphics
- [ ] Email
- [ ] Video
## Week 2: [Date Range]
[Repeat structure]
## Week 3: [Date Range]
[Repeat structure]
## Week 4: [Date Range]
[Repeat structure]
## Content Bank (Ideas for Future)
1.
2.
3.
4.
5.
## Performance Review (End of Month)
### Top Performing Content
1. **Title/Topic**:
- **Metric**:
- **Why it worked**:
2. **Title/Topic**:
- **Metric**:
- **Why it worked**:
### Lessons Learned
-
-
-
### Adjustments for Next Month
-
-
-
## Resource Links
- Brand Guidelines: [Link]
- Asset Library: [Link]
- Analytics Dashboard: [Link]
- Team Calendar: [Link]
FILE:examples/brand_voice_analysis_example.md
# Brand Voice Analysis Example
Demonstration of brand_voice_analyzer.py input and output.
---
## Sample Input
**File: `sample_blog_post.txt`**
```
Hey there! 👋
So, like, we've been doing marketing for a really long time and we've learned SO much about what works. Today I'm gonna share some super cool tips that'll totally transform your business!
First things first - you gotta know your audience. Like, REALLY know them. What do they want? What keeps them up at night? Figure that out and you're golden!
Second, content is king (obviously). But here's the thing - not just any content. You need stuff that actually helps people. Don't just post to post, ya know?
Anyway, hope this helps! Drop a comment if you have questions! 🚀
```
---
## Command
```bash
python scripts/brand_voice_analyzer.py sample_blog_post.txt
```
---
## Sample Output (Text Format)
```
============================================================
BRAND VOICE ANALYSIS RESULTS
============================================================
VOICE PROFILE
------------------------------------------------------------
Formality Score: 25/100 (Casual)
Tone: Conversational, Enthusiastic, Informal
Perspective: Mixed (1st person singular + 2nd person)
Personality Match: The Friend (primary)
READABILITY METRICS
------------------------------------------------------------
Flesch Reading Ease: 78 (Fairly Easy)
Grade Level: 6th Grade
Avg Sentence Length: 12 words
Avg Word Length: 4.2 characters
SENTENCE ANALYSIS
------------------------------------------------------------
Total Sentences: 12
Simple Sentences: 8 (67%)
Compound Sentences: 3 (25%)
Complex Sentences: 1 (8%)
VOCABULARY PATTERNS
------------------------------------------------------------
Filler Words Found: 6 (like, so, really, just, totally, super)
Contractions: 5 (we've, I'm, gonna, you're, don't)
Emoji Usage: 2
Exclamation Points: 4
RECOMMENDATIONS
------------------------------------------------------------
1. [HIGH] Reduce filler words - found 6 instances
Action: Remove "like", "so", "really", "totally", "super"
2. [MEDIUM] Inconsistent perspective - switches between "I" and "we"
Action: Choose one perspective and maintain throughout
3. [MEDIUM] High emoji count for professional content
Action: Limit to 1 emoji or remove entirely for B2B
4. [LOW] Overuse of exclamation points
Action: Replace 3 of 4 with periods for measured tone
VOICE CONSISTENCY SCORE: 62/100
============================================================
```
---
## Sample Output (JSON Format)
```bash
python scripts/brand_voice_analyzer.py sample_blog_post.txt json
```
```json
{
"voice_profile": {
"formality_score": 25,
"formality_level": "Casual",
"tone": ["Conversational", "Enthusiastic", "Informal"],
"perspective": "Mixed",
"personality_archetype": "The Friend"
},
"readability": {
"flesch_reading_ease": 78,
"grade_level": 6,
"avg_sentence_length": 12,
"avg_word_length": 4.2
},
"sentence_analysis": {
"total": 12,
"simple": 8,
"compound": 3,
"complex": 1
},
"vocabulary": {
"filler_words": {
"count": 6,
"instances": ["like", "so", "really", "just", "totally", "super"]
},
"contractions": 5,
"emojis": 2,
"exclamation_points": 4
},
"recommendations": [
{
"priority": "high",
"category": "vocabulary",
"issue": "Excessive filler words",
"action": "Remove casual filler words for professional tone"
},
{
"priority": "medium",
"category": "perspective",
"issue": "Inconsistent perspective",
"action": "Maintain single perspective throughout"
},
{
"priority": "medium",
"category": "formatting",
"issue": "High emoji count",
"action": "Limit emojis for professional content"
},
{
"priority": "low",
"category": "punctuation",
"issue": "Overuse of exclamation points",
"action": "Replace with periods for measured tone"
}
],
"consistency_score": 62
}
```
---
## Revised Content (After Applying Recommendations)
```
We've been helping businesses with marketing for over a decade, and we've
identified key principles that consistently drive results.
Understanding your audience is foundational. What challenges do they face?
What outcomes do they seek? Deep audience knowledge shapes every effective
marketing decision.
Content quality matters more than quantity. Focus on creating resources that
genuinely solve problems for your readers rather than publishing content
solely to maintain a schedule.
Questions about implementing these strategies? Leave a comment below.
```
**Re-analysis Results:**
```
Formality Score: 72/100 (Professional)
Tone: Educational, Confident, Helpful
Perspective: First Person Plural (consistent)
Consistency Score: 91/100
```
FILE:examples/seo_optimization_example.md
# SEO Optimization Example
Demonstration of seo_optimizer.py input and output.
---
## Sample Input
**File: `draft_article.md`**
```markdown
# Marketing Tips
Marketing is important for businesses. Here are some things to know.
## Why Marketing Matters
Companies need marketing. It helps them grow. Marketing brings customers.
## Some Ideas
Try social media. Post content. Use email. Run ads.
## Conclusion
Marketing is good. Do more of it.
```
---
## Command
```bash
python scripts/seo_optimizer.py draft_article.md "content marketing strategy" "content marketing,marketing tips,business growth"
```
---
## Sample Output (Text Format)
```
============================================================
SEO ANALYSIS REPORT
============================================================
PRIMARY KEYWORD: "content marketing strategy"
SECONDARY KEYWORDS: content marketing, marketing tips, business growth
OVERALL SEO SCORE: 32/100 (Poor)
KEYWORD ANALYSIS
------------------------------------------------------------
Primary Keyword Density: 0.0% (Target: 1-3%)
Status: NOT FOUND in content
Secondary Keyword Usage:
- "content marketing": 0 occurrences (Target: 3-5)
- "marketing tips": 1 occurrence (in title only)
- "business growth": 0 occurrences (Target: 2-3)
Keyword Placement Check:
✗ Primary keyword NOT in title
✗ Primary keyword NOT in first paragraph
✗ Primary keyword NOT in H2 headings
✗ Primary keyword NOT in conclusion
CONTENT STRUCTURE
------------------------------------------------------------
Word Count: 67 words
Status: CRITICAL - Below minimum (Target: 1,500+)
Heading Structure:
H1: 1 (Good)
H2: 3 (Good)
H3: 0 (Consider adding for depth)
Paragraph Analysis:
- Average length: 12 words (Too short - Target: 40-80)
- Total paragraphs: 6
READABILITY
------------------------------------------------------------
Flesch Reading Ease: 82 (Easy)
Note: May be too simple for B2B audience
META ELEMENTS
------------------------------------------------------------
Meta Title: Not specified
Suggestion: "Content Marketing Strategy: 10 Proven Tips for 2025"
Meta Description: Not found
Suggestion: "Discover actionable content marketing strategies to drive
business growth. Learn proven techniques for content that converts."
INTERNAL/EXTERNAL LINKS
------------------------------------------------------------
Internal Links: 0 (Target: 2-3)
External Links: 0 (Target: 1-2 authoritative sources)
RECOMMENDATIONS (Priority Order)
------------------------------------------------------------
[P0] CRITICAL - Content Length
Issue: 67 words is severely below minimum
Action: Expand to 1,500-2,500 words with detailed sections
[P0] CRITICAL - Missing Primary Keyword
Issue: "content marketing strategy" not found anywhere
Action: Include in title, first paragraph, 2 H2s, and conclusion
[P1] HIGH - Thin Content Sections
Issue: Paragraphs average 12 words
Action: Expand each section with examples, data, and actionable steps
[P1] HIGH - Missing Internal Links
Issue: No links to related content
Action: Add 2-3 links to relevant articles
[P2] MEDIUM - No Meta Description
Issue: Missing meta description
Action: Add 150-160 character description with primary keyword
[P2] MEDIUM - Missing H3 Subheadings
Issue: No H3s for content depth
Action: Add H3s under each H2 for better structure
============================================================
```
---
## Sample Output (JSON Format)
```bash
python scripts/seo_optimizer.py draft_article.md "content marketing strategy" --json
```
```json
{
"overall_score": 32,
"grade": "Poor",
"primary_keyword": "content marketing strategy",
"keyword_analysis": {
"primary_density": 0.0,
"target_density": "1-3%",
"primary_found": false,
"secondary_keywords": {
"content marketing": {"count": 0, "target": "3-5"},
"marketing tips": {"count": 1, "target": "2-3"},
"business growth": {"count": 0, "target": "2-3"}
},
"placement": {
"in_title": false,
"in_first_paragraph": false,
"in_h2_headings": false,
"in_conclusion": false
}
},
"content_structure": {
"word_count": 67,
"min_recommended": 1500,
"headings": {"h1": 1, "h2": 3, "h3": 0},
"paragraphs": {"count": 6, "avg_length": 12}
},
"readability": {
"flesch_score": 82,
"level": "Easy"
},
"meta": {
"title": null,
"description": null,
"suggested_title": "Content Marketing Strategy: 10 Proven Tips for 2025",
"suggested_description": "Discover actionable content marketing strategies to drive business growth. Learn proven techniques for content that converts."
},
"links": {
"internal": 0,
"external": 0,
"target_internal": "2-3",
"target_external": "1-2"
},
"recommendations": [
{
"priority": "P0",
"category": "content_length",
"issue": "Content severely below minimum word count",
"action": "Expand to 1,500-2,500 words"
},
{
"priority": "P0",
"category": "keyword",
"issue": "Primary keyword not found",
"action": "Include in title, first paragraph, H2s, conclusion"
},
{
"priority": "P1",
"category": "content_depth",
"issue": "Thin content sections",
"action": "Expand with examples, data, actionable steps"
},
{
"priority": "P1",
"category": "links",
"issue": "No internal links",
"action": "Add 2-3 relevant internal links"
}
]
}
```
---
## Optimized Content (After Applying Recommendations)
```markdown
# Content Marketing Strategy: 10 Proven Techniques for Business Growth
A well-executed content marketing strategy separates thriving businesses from
those struggling to gain visibility. This comprehensive guide covers the
essential techniques that drive measurable results.
## Why Content Marketing Strategy Matters for Business Growth
Companies investing in strategic content marketing see 3x more leads than
those relying solely on paid advertising. Content marketing builds lasting
assets that continue generating value long after publication.
The compounding effect of quality content creates sustainable business growth:
- Organic search traffic increases over time
- Brand authority strengthens with each published piece
- Customer acquisition costs decrease as content library grows
### The ROI of Strategic Content
According to Content Marketing Institute research, businesses with documented
content strategies are 313% more likely to report success than those without.
[Continue for 1,500+ words with detailed sections...]
## Conclusion: Building Your Content Marketing Strategy
Implementing these content marketing techniques positions your business for
sustained growth. Start with audience research, create a documented strategy,
and commit to consistent execution.
Related reading: [Link to internal article on content calendars]
```
**Re-analysis Results:**
```
OVERALL SEO SCORE: 87/100 (Good)
✓ Primary keyword density: 1.8%
✓ Keyword in title, first paragraph, H2s, conclusion
✓ Word count: 1,847 words
✓ Meta description: Present (156 characters)
✓ Internal links: 2
✓ External links: 1 (authoritative source)
```
FILE:references/analytics_guide.md
# Content Analytics & Performance Metrics
Comprehensive guide for tracking, measuring, and optimizing content performance.
---
## Table of Contents
- [Content Metrics](#content-metrics)
- [Engagement Metrics](#engagement-metrics)
- [Business Metrics](#business-metrics)
- [Platform-Specific Analytics](#platform-specific-analytics)
- [Reporting Frameworks](#reporting-frameworks)
- [Attribution Models](#attribution-models)
---
## Content Metrics
Track these KPIs to measure content reach and consumption.
### Traffic Metrics
| Metric | Target | What It Tells You |
|--------|--------|-------------------|
| Organic traffic | +10% MoM | SEO effectiveness |
| Page views | Varies by content type | Raw consumption volume |
| Unique visitors | +5% MoM | Audience growth |
| Sessions per user | 1.5+ | Content stickiness |
### Consumption Metrics
| Metric | Target | What It Tells You |
|--------|--------|-------------------|
| Average time on page | 3+ min for long-form | Content depth engagement |
| Bounce rate | <60% | Content relevance |
| Scroll depth | 70%+ | Content holding attention |
| Pages per session | 2+ | Internal linking success |
### SEO Metrics
| Metric | Target | What It Tells You |
|--------|--------|-------------------|
| Keyword rankings | Top 10 | Search visibility |
| Backlinks earned | +5/month | Content authority |
| Domain authority | Steady growth | Overall site strength |
| Featured snippets | Track position | SERP prominence |
---
## Engagement Metrics
Measure how audiences interact with content.
### Social Engagement
| Metric | Benchmark | Calculation |
|--------|-----------|-------------|
| Engagement rate | 1-3% (LinkedIn) | (Likes + Comments + Shares) / Impressions × 100 |
| Share rate | 0.5-1% | Shares / Reach × 100 |
| Save rate | 1-2% (Instagram) | Saves / Reach × 100 |
| Comment rate | 0.1-0.5% | Comments / Reach × 100 |
### Email Engagement
| Metric | Benchmark | What It Tells You |
|--------|-----------|-------------------|
| Open rate | 20-25% | Subject line effectiveness |
| Click-through rate | 2-5% | Content relevance |
| Unsubscribe rate | <0.5% | Audience fit |
| Forward rate | 0.1-0.3% | Content share-worthiness |
### Community Engagement
| Metric | What to Track |
|--------|---------------|
| Comments and discussions | Volume and sentiment |
| User-generated content | Submissions and quality |
| Community growth | New members per week |
| Active participation | % of members engaging |
---
## Business Metrics
Connect content performance to business outcomes.
### Lead Generation
| Metric | Calculation | Target |
|--------|-------------|--------|
| Content-attributed leads | Leads from content CTAs | Track by content piece |
| Form submissions | Total completions | +5% MoM |
| Lead quality score | MQL/total leads | 30%+ MQL rate |
| Cost per lead | Spend / Leads | Below industry average |
### Conversion Metrics
| Metric | Calculation | Target |
|--------|-------------|--------|
| Conversion rate | Conversions / Visitors × 100 | 2-5% |
| Revenue attribution | $ tied to content | Track by piece |
| Customer acquisition cost | Total cost / New customers | Decreasing trend |
| Content ROI | (Revenue - Cost) / Cost × 100 | 300%+ for evergreen |
### Customer Metrics
| Metric | What It Tells You |
|--------|-------------------|
| Customer lifetime value | Long-term content impact |
| Retention rate | Content's nurturing effectiveness |
| NPS from content consumers | Content quality perception |
| Support ticket reduction | Educational content success |
---
## Platform-Specific Analytics
### Blog Analytics (Google Analytics 4)
**Key Reports:**
- Landing pages report: Top entry content
- Engagement report: Time, bounces, conversions
- Traffic acquisition: Content discovery sources
- User paths: Content journey mapping
**Dimensions to Track:**
- Page path
- Source/medium
- Device category
- User type (new vs returning)
### Social Media Analytics
**LinkedIn:**
- Post impressions and reach
- Follower demographics
- Click-through rate on links
- Article read time
**Twitter/X:**
- Impressions and engagements
- Profile visits from tweets
- Link clicks
- Follower growth rate
**Instagram:**
- Reach vs impressions
- Saves and shares (high-value signals)
- Story completion rate
- Reel performance vs feed
### Email Analytics
**Track per Campaign:**
- Send volume and deliverability
- Open and click rates by segment
- Conversion path from email
- List growth and churn
---
## Reporting Frameworks
### Weekly Content Report
```
WEEK OF: [Date Range]
TOP PERFORMERS
1. [Content Title] - [Key Metric]
2. [Content Title] - [Key Metric]
3. [Content Title] - [Key Metric]
TRAFFIC SUMMARY
- Total sessions: [#]
- Organic traffic: [#] ([+/-]% WoW)
- Social traffic: [#] ([+/-]% WoW)
ENGAGEMENT HIGHLIGHTS
- Avg engagement rate: [%]
- Total comments: [#]
- Shares: [#]
LEADS GENERATED
- Content-attributed: [#]
- Top converting piece: [Title]
NEXT WEEK PRIORITIES
1. [Action item]
2. [Action item]
```
### Monthly Content Report
```
MONTH: [Month Year]
EXECUTIVE SUMMARY
[2-3 sentences on overall performance]
CONTENT PRODUCTION
- Published: [#] pieces
- By type: [Blog: #, Social: #, Email: #]
- On schedule: [Yes/No]
PERFORMANCE DASHBOARD
| Metric | This Month | Last Month | Change |
|---------------------|------------|------------|--------|
| Total traffic | | | |
| Organic traffic | | | |
| Engagement rate | | | |
| Leads generated | | | |
| Conversion rate | | | |
TOP 5 CONTENT PIECES
[Ranked by primary KPI]
INSIGHTS & LEARNINGS
- What worked: [observation]
- What didn't: [observation]
- Opportunities: [observation]
NEXT MONTH FOCUS
1. [Strategic priority]
2. [Content initiative]
3. [Optimization goal]
```
### Quarterly Business Review
```
Q[#] [Year] CONTENT PERFORMANCE
STRATEGIC ALIGNMENT
- Business goal: [Goal]
- Content contribution: [How content supported]
QUARTERLY METRICS
| KPI | Target | Actual | Status |
|------------------------|--------|--------|--------|
| Traffic growth | | | |
| Lead generation | | | |
| Conversion rate | | | |
| Revenue attribution | | | |
CONTENT AUDIT RESULTS
- Total pieces published: [#]
- High performers: [#]
- Needs optimization: [#]
- Candidates for retirement: [#]
ROI ANALYSIS
- Total content investment: $[X]
- Attributed revenue: $[Y]
- Content ROI: [%]
COMPETITIVE ANALYSIS
[How content stacks against competitors]
NEXT QUARTER ROADMAP
[Strategic initiatives and targets]
```
---
## Attribution Models
### First-Touch Attribution
**Use When:** Measuring top-of-funnel content effectiveness
**How It Works:** Credits the first content piece that brought a user in
**Best For:**
- Brand awareness campaigns
- SEO content performance
- Social media reach measurement
### Last-Touch Attribution
**Use When:** Measuring bottom-of-funnel conversion content
**How It Works:** Credits the last content before conversion
**Best For:**
- Product pages
- Case studies
- Demo request pages
### Multi-Touch Attribution
**Use When:** Understanding full content journey impact
**Linear Model:**
- Equal credit to all touchpoints
- Simple but may over-credit low-value touches
**Time-Decay Model:**
- More credit to recent touches
- Good for short sales cycles
**Position-Based Model:**
- 40% first touch, 40% last touch, 20% middle
- Balanced view of journey
### Content-Specific Attribution
**For Blog Content:**
1. Track assisted conversions in GA4
2. Map content to funnel stage
3. Weight by stage importance
**For Social Content:**
1. Use UTM parameters consistently
2. Track view-through conversions
3. Monitor social-assisted conversions
**For Email Content:**
1. Track email-attributed revenue
2. Monitor nurture sequence effectiveness
3. Measure reactivation campaigns
---
## Analytics Setup Checklist
### Essential Tracking
- [ ] Google Analytics 4 configured
- [ ] Conversion events defined
- [ ] UTM parameter system documented
- [ ] Social pixel tracking enabled
- [ ] Email tracking integrated
- [ ] CRM connected for lead tracking
### Advanced Setup
- [ ] Enhanced ecommerce tracking
- [ ] Custom dimensions for content attributes
- [ ] Automated reporting dashboards
- [ ] A/B testing infrastructure
- [ ] Heat mapping tools (Hotjar, Clarity)
- [ ] Attribution model configured
### Data Governance
- [ ] Naming conventions documented
- [ ] Data retention policies set
- [ ] Privacy compliance verified
- [ ] Access controls configured
- [ ] Regular data audits scheduled
FILE:references/brand_guidelines.md
# Brand Voice & Style Guidelines
Comprehensive framework for establishing and maintaining consistent brand voice across all content.
---
## Table of Contents
- [Voice Dimensions](#1-voice-dimensions)
- [Brand Personality Archetypes](#2-brand-personality-archetypes)
- [Writing Principles](#3-writing-principles)
- [Language Guidelines](#4-language-guidelines)
- [Content Structure Templates](#5-content-structure-templates)
- [Messaging Pillars](#6-messaging-pillars)
- [Audience Personas](#7-audience-personas)
- [Channel-Specific Guidelines](#8-channel-specific-guidelines)
- [Grammar & Mechanics](#9-grammar--mechanics)
- [Inclusivity Guidelines](#10-inclusivity-guidelines)
- [Quick Reference Checklist](#quick-reference-checklist)
---
## Brand Voice Framework
### 1. Voice Dimensions
#### Formality Spectrum
- **Formal**: Legal documents, investor communications, crisis responses
- **Professional**: B2B content, whitepapers, case studies
- **Conversational**: Blog posts, social media, email newsletters
- **Casual**: Community engagement, behind-the-scenes content
#### Tone Attributes
Choose 3-5 primary attributes for your brand:
- **Authoritative**: Position as industry expert
- **Friendly**: Approachable and warm
- **Innovative**: Forward-thinking and creative
- **Trustworthy**: Reliable and transparent
- **Inspiring**: Motivational and uplifting
- **Educational**: Informative and helpful
- **Witty**: Clever and entertaining (use sparingly)
#### Perspective
- **First Person Plural (We/Our)**: Creates partnership feeling
- **Second Person (You/Your)**: Direct and engaging
- **Third Person**: Objective and professional
### 2. Brand Personality Archetypes
Choose one primary and one secondary archetype:
**The Expert**
- Tone: Knowledgeable, confident, informative
- Content: Data-driven, research-backed, educational
- Example: "Our research shows that 87% of businesses..."
**The Friend**
- Tone: Warm, supportive, conversational
- Content: Relatable, helpful, encouraging
- Example: "We get it - marketing can be overwhelming..."
**The Innovator**
- Tone: Visionary, bold, forward-thinking
- Content: Cutting-edge, disruptive, trendsetting
- Example: "The future of marketing is here..."
**The Guide**
- Tone: Wise, patient, instructive
- Content: Step-by-step, clear, actionable
- Example: "Let's walk through this together..."
**The Motivator**
- Tone: Energetic, positive, inspiring
- Content: Empowering, action-oriented, transformative
- Example: "You have the power to transform your business..."
### 3. Writing Principles
#### Clarity First
- Use simple words when possible
- Break complex ideas into digestible pieces
- Lead with the main point
- Use active voice (80% of the time)
#### Customer-Centric
- Focus on benefits, not features
- Address pain points directly
- Use "you" more than "we"
- Include customer success stories
#### Consistency
- Maintain voice across all channels
- Use approved terminology
- Follow formatting standards
- Apply style rules uniformly
### 4. Language Guidelines
#### Words We Use
- **Action verbs**: Transform, accelerate, optimize, unlock, elevate
- **Positive descriptors**: Seamless, powerful, intuitive, strategic
- **Outcome-focused**: Results, growth, success, impact, ROI
#### Words We Avoid
- **Jargon**: Synergy, leverage (as verb), bandwidth (for availability)
- **Overused**: Innovative, disruptive, cutting-edge (unless truly applicable)
- **Weak**: Very, really, just, maybe, hopefully
- **Negative**: Can't, won't, impossible, problem (use "challenge")
### 5. Content Structure Templates
#### Blog Post Structure
1. **Hook** (1-2 sentences): Grab attention with a question, statistic, or bold statement
2. **Context** (1 paragraph): Explain why this matters now
3. **Main Content** (3-5 sections): Deliver value with clear subheadings
4. **Conclusion** (1 paragraph): Summarize key points
5. **Call to Action**: Clear next step for readers
#### Social Media Framework
- **LinkedIn**: Professional insights, industry news, thought leadership
- **Twitter/X**: Quick tips, engaging questions, thread stories
- **Instagram**: Visual storytelling, behind-the-scenes, inspiration
- **Facebook**: Community building, longer narratives, events
### 6. Messaging Pillars
Define 3-4 core themes that appear consistently:
1. **Innovation & Technology**
- AI-powered solutions
- Data-driven insights
- Future-ready strategies
2. **Customer Success**
- Real results and ROI
- Partnership approach
- Tailored solutions
3. **Expertise & Trust**
- Industry leadership
- Proven methodologies
- Transparent communication
4. **Growth & Transformation**
- Scaling businesses
- Digital transformation
- Continuous improvement
### 7. Audience Personas
#### Decision Makers (C-Suite)
- **Tone**: Professional, strategic, ROI-focused
- **Content**: High-level insights, business impact, competitive advantages
- **Pain Points**: Growth, efficiency, competition
#### Practitioners (Marketing Managers)
- **Tone**: Practical, supportive, educational
- **Content**: How-to guides, best practices, tools
- **Pain Points**: Time, resources, skills
#### Innovators (Early Adopters)
- **Tone**: Exciting, cutting-edge, visionary
- **Content**: Trends, new features, future predictions
- **Pain Points**: Staying ahead, differentiation
### 8. Channel-Specific Guidelines
#### Website Copy
- Headlines: 6-12 words, benefit-focused
- Body: Short paragraphs (2-3 sentences)
- CTAs: Action-oriented, specific
#### Email Marketing
- Subject Lines: 30-50 characters, personalized
- Preview Text: Complement subject, add urgency
- Body: Scannable, one main message
#### Blog Content
- Title: Include primary keyword, under 60 characters
- Introduction: Hook within first 50 words
- Sections: 200-300 words each
- Lists: 5-7 items optimal
### 9. Grammar & Mechanics
#### Punctuation
- Oxford comma: Always use
- Em dashes: For emphasis—like this
- Exclamation points: Maximum one per piece
#### Capitalization
- Headlines: Title Case for H1, Sentence case for H2-H6
- Product names: As trademarked
- Job titles: Lowercase unless before name
#### Numbers
- Spell out one through nine
- Use numerals for 10 and above
- Always use numerals for percentages
### 10. Inclusivity Guidelines
- Use gender-neutral language
- Avoid idioms that don't translate
- Consider global audience
- Ensure accessibility in formatting
- Represent diverse perspectives
## Quick Reference Checklist
Before publishing any content, verify:
- [ ] Matches brand voice and tone
- [ ] Free of jargon and complex terms
- [ ] Includes clear value proposition
- [ ] Has appropriate CTA
- [ ] Follows grammar guidelines
- [ ] Mobile-friendly formatting
- [ ] Accessible to all audiences
- [ ] Proofread and fact-checked
FILE:references/content_frameworks.md
# Content Creation Frameworks & Templates
Ready-to-use templates for blog posts, social media, email marketing, video scripts, and content planning.
---
## Table of Contents
- [Blog Post Templates](#1-blog-post-templates)
- [Social Media Templates](#2-social-media-templates)
- [Email Marketing Templates](#3-email-marketing-templates)
- [Content Planning Frameworks](#4-content-planning-frameworks)
- [SEO Content Framework](#5-seo-content-framework)
- [Video Script Templates](#6-video-script-templates)
- [Content Repurposing Matrix](#7-content-repurposing-matrix)
- [Quick-Start Checklists](#quick-start-checklists)
---
## Content Types & Templates
### 1. Blog Post Templates
#### How-To Guide Template
```markdown
# How to [Achieve Desired Outcome] in [Timeframe]
## Introduction
- Hook: Question or surprising fact
- Problem statement
- What reader will learn
- Why it matters now
## Prerequisites/What You'll Need
- Tool/Resource 1
- Tool/Resource 2
- Estimated time
## Step 1: [Action]
- Clear instruction
- Why this step matters
- Common mistakes to avoid
- Visual aid or example
## Step 2: [Action]
[Repeat structure]
## Step 3: [Action]
[Repeat structure]
## Troubleshooting Common Issues
### Issue 1: [Problem]
**Solution**: [Fix]
### Issue 2: [Problem]
**Solution**: [Fix]
## Results You Can Expect
- Immediate outcomes
- Long-term benefits
- Success metrics
## Next Steps
- Advanced techniques
- Related guides
- CTA for product/service
## Conclusion
- Recap key points
- Reinforce value
- Final encouragement
```
#### Listicle Template
```markdown
# [Number] [Adjective] Ways to [Achieve Goal] in [Year]
## Introduction
- Context/trend driving this topic
- Promise of what reader gains
- Credibility statement
## 1. [First Item - Most Important]
**Why it matters**: [Brief explanation]
**How to implement**: [2-3 actionable steps]
**Pro tip**: [Expert insight]
**Example**: [Real-world application]
## 2. [Second Item]
[Repeat structure]
[Continue for all items]
## Bonus Tip: [Overdelivery]
[Something extra valuable]
## Bringing It All Together
- How items work synergistically
- Priority order for implementation
- Expected timeline for results
## Your Action Plan
1. Start with [easiest item]
2. Progress to [next steps]
3. Measure [metrics]
## Conclusion & CTA
```
#### Case Study Template
```markdown
# How [Company] Achieved [Result] Using [Solution]
## Executive Summary
- Company overview
- Challenge faced
- Solution implemented
- Key results (3 metrics)
## The Challenge
### Background
- Industry context
- Company situation
- Previous attempts
### Specific Pain Points
- Pain point 1
- Pain point 2
- Pain point 3
## The Solution
### Strategy Development
- Discovery process
- Strategic approach
- Why this solution
### Implementation
- Phase 1: [Timeline & Actions]
- Phase 2: [Timeline & Actions]
- Phase 3: [Timeline & Actions]
## The Results
### Quantitative Outcomes
- Metric 1: X% increase
- Metric 2: $Y saved
- Metric 3: Z improvement
### Qualitative Benefits
- Team feedback
- Customer response
- Market position
## Key Takeaways
1. Lesson learned
2. Best practice discovered
3. Unexpected benefit
## Achieving Similar Results
- Prerequisite conditions
- Implementation roadmap
- Success factors
## CTA: Start Your Success Story
```
#### Thought Leadership Template
```markdown
# [Provocative Statement About Industry Future]
## The Current State
- Industry snapshot
- Prevailing wisdom
- Why status quo is insufficient
## The Emerging Trend
### What's Changing
- Driver 1: [Technology/Market/Behavior]
- Driver 2: [Technology/Market/Behavior]
- Driver 3: [Technology/Market/Behavior]
### Evidence & Examples
- Data point 1
- Case example
- Expert validation
## Implications for [Industry]
### Short-term (6-12 months)
- Immediate adjustments needed
- Quick wins available
- Risks of inaction
### Long-term (2-5 years)
- Fundamental shifts
- New opportunities
- Competitive landscape
## Strategic Recommendations
### For Leaders
- Strategic priorities
- Investment areas
- Organizational changes
### For Practitioners
- Skill development
- Process adaptation
- Tool adoption
## The Path Forward
- Call for industry action
- Your organization's role
- Next steps for readers
## Join the Conversation
- Thought-provoking question
- Invitation to share perspectives
- CTA for deeper engagement
```
### 2. Social Media Templates
#### LinkedIn Post Framework
```
🎯 Hook/Pattern Interrupt
Context paragraph explaining the situation or challenge.
Key insight or lesson learned:
• Bullet point 1 (specific detail)
• Bullet point 2 (measurable outcome)
• Bullet point 3 (unexpected discovery)
Brief story or example that illustrates the point.
Takeaway message with clear value.
Question to encourage engagement?
#Hashtag1 #Hashtag2 #Hashtag3
```
#### Twitter/X Thread Template
```
1/ Bold opening statement or question that stops the scroll
2/ Context - why this matters right now
3/ Problem most people face
4/ Conventional solution (and why it falls short)
5/ Better approach - introduction
6/ Step 1 of better approach
• Specific action
• Why it works
7/ Step 2 of better approach
[Continue pattern]
8/ Real example or case study
9/ Common objection addressed
10/ Results you can expect
11/ One powerful tip most people miss
12/ Recap in 3 key points:
- Point 1
- Point 2
- Point 3
13/ CTA: If you found this helpful, [action]
14/ P.S. - Bonus insight or resource
```
#### Instagram Caption Template
```
[Attention-grabbing first line - appears in preview]
[Story or relatable scenario - 2-3 sentences]
Here's what I learned:
[Key insight or lesson]
3 things that changed everything:
1️⃣ [First point]
2️⃣ [Second point]
3️⃣ [Third point]
[Call-out or question to audience]
Drop a [emoji] if you've experienced this too!
What's your biggest challenge with [topic]? Let me know below 👇
-
#hashtag1 #hashtag2 #hashtag3 #hashtag4 #hashtag5
[10-30 relevant hashtags total]
```
### 3. Email Marketing Templates
#### Newsletter Template
```
Subject: [Benefit] + [Urgency/Curiosity]
Preview: [Complements subject, doesn't repeat]
Hi [Name],
[Personal observation or timely hook - 1-2 sentences]
[Transition to main topic - why reading this matters]
## Main Content Section
[Key points in scannable format]
• Point 1: [Benefit-focused]
• Point 2: [Specific example]
• Point 3: [Actionable tip]
[Brief elaboration on most important point - 2-3 sentences]
## Resource of the Week
[Title with link]
[One sentence on why it's valuable]
## Quick Win You Can Implement Today
[Specific, actionable tip - 2-3 steps max]
[Closing thought or question]
[Signature]
[Name]
P.S. [Additional value or soft CTA]
```
#### Promotional Email Template
```
Subject: [Specific benefit] by [deadline/timeframe]
Preview: [Scarcity or exclusivity element]
Hi [Name],
[Acknowledge pain point or aspiration]
[Agitate - why this problem persists]
I've got something that can help:
[Solution introduction - what it is]
Here's what you get:
✓ Benefit 1 (not feature)
✓ Benefit 2 (not feature)
✓ Benefit 3 (not feature)
[Social proof - testimonial or results]
[Handle main objection]
[Clear CTA button: "Get Started" / "Claim Yours"]
[Urgency element - deadline or limited availability]
[Signature]
P.S. [Reinforce urgency or add bonus]
```
### 4. Content Planning Frameworks
#### Content Pillar Strategy
```
Pillar 1: Educational (40%)
- How-to guides
- Tutorials
- Best practices
- Tips & tricks
Pillar 2: Inspirational (25%)
- Success stories
- Case studies
- Transformations
- Vision pieces
Pillar 3: Conversational (25%)
- Behind-the-scenes
- Team spotlights
- Q&As
- Polls/questions
Pillar 4: Promotional (10%)
- Product updates
- Offers
- Event announcements
- CTAs
```
#### Monthly Content Calendar Structure
```
Week 1:
- Monday: Educational (blog post)
- Wednesday: Inspirational (social)
- Friday: Conversational (email)
Week 2:
- Monday: Educational (video/guide)
- Wednesday: Case study
- Friday: Curated content
Week 3:
- Monday: Educational (infographic)
- Wednesday: Behind-the-scenes
- Friday: Community spotlight
Week 4:
- Monday: Monthly roundup
- Wednesday: Thought leadership
- Friday: Promotional
```
### 5. SEO Content Framework
#### SEO-Optimized Article Structure
```
URL: /primary-keyword-secondary-keyword
Title Tag: Primary Keyword - Secondary Benefit | Brand
Meta Description: Action verb + primary keyword + benefit + CTA (155 chars)
# H1: Primary Keyword + Unique Angle
Introduction (50-100 words)
- Include primary keyword in first 100 words
- State what reader will learn
- Why it matters
## H2: Secondary Keyword Variation 1
[Content with LSI keywords naturally integrated]
### H3: Specific subtopic
- Detail point 1
- Detail point 2
- Detail point 3
## H2: Secondary Keyword Variation 2
[Content continues...]
## H2: Related Questions (FAQ Schema)
### Question 1?
[Concise answer with keyword]
### Question 2?
[Concise answer with keyword]
## Conclusion
- Recap main points
- Include primary keyword
- Clear next action
Internal Links: 2-3 relevant articles
External Links: 1-2 authoritative sources
```
### 6. Video Script Templates
#### Educational Video Script
```
[0-5 seconds: Hook]
"What if I told you [surprising statement]?"
[5-15 seconds: Introduction]
"Hi, I'm [Name] and today we're solving [problem]"
[15-30 seconds: Context]
- Why this matters
- What you'll learn
- What you'll achieve
[30 seconds - 2 minutes: Main Content]
Section 1: [Key Point]
- Explanation
- Example
- Visual aid
Section 2: [Key Point]
[Repeat structure]
Section 3: [Key Point]
[Repeat structure]
[Final 15-30 seconds]
- Quick recap
- Call to action
- End screen elements
```
### 7. Content Repurposing Matrix
```
Original: Blog Post (2000 words)
├── Social Media
│ ├── 5 Twitter posts (key quotes)
│ ├── 1 LinkedIn article (executive summary)
│ ├── 3 Instagram carousels (main points)
│ └── 1 Facebook post (intro + link)
├── Email
│ └── Newsletter feature (summary + CTA)
├── Video
│ ├── YouTube explainer (script from post)
│ └── TikTok/Reels (quick tips)
├── Audio
│ └── Podcast talking points
└── Visual
├── Infographic (data points)
└── Slide deck (presentation)
```
## Quick-Start Checklists
### Pre-Publishing Checklist
- [ ] Keyword research completed
- [ ] Title under 60 characters
- [ ] Meta description written (155 chars)
- [ ] Headers properly structured (H1, H2, H3)
- [ ] Internal links added (2-3)
- [ ] Images optimized with alt text
- [ ] CTA included and clear
- [ ] Proofread and fact-checked
- [ ] Mobile preview checked
### Content Quality Checklist
- [ ] Addresses specific audience need
- [ ] Provides unique value/perspective
- [ ] Includes actionable takeaways
- [ ] Uses appropriate brand voice
- [ ] Contains supporting data/examples
- [ ] Free of jargon and complex terms
- [ ] Scannable format (bullets, headers)
- [ ] Engaging hook in introduction
- [ ] Clear conclusion and next steps
FILE:references/social_media_optimization.md
# Social Media Optimization Guide
Platform-specific best practices, algorithm factors, content optimization strategies, and analytics frameworks.
---
## Table of Contents
- [Platform-Specific Best Practices](#platform-specific-best-practices)
- [LinkedIn](#linkedin)
- [Twitter/X](#twitterx)
- [Instagram](#instagram)
- [Facebook](#facebook)
- [TikTok](#tiktok)
- [Content Optimization Strategies](#content-optimization-strategies)
- [Hashtag Strategy](#hashtag-strategy)
- [Visual Content Optimization](#visual-content-optimization)
- [Caption Writing Formulas](#caption-writing-formulas)
- [Engagement Tactics](#engagement-tactics)
- [Analytics & KPIs](#analytics--kpis)
- [Content Calendar Planning](#content-calendar-planning)
- [Crisis Management Protocol](#crisis-management-protocol)
- [Tool Stack Recommendations](#tool-stack-recommendations)
- [Compliance & Best Practices](#compliance--best-practices)
---
## Platform-Specific Best Practices
### LinkedIn
**Audience**: B2B professionals, decision-makers, thought leaders
**Best Times**: Tuesday-Thursday, 8-10 AM and 5-6 PM
**Optimal Length**: 1,300-2,000 characters for posts
#### Content Formats
- **Text Posts**: 1,300 characters optimal, use line breaks
- **Articles**: 1,900-2,000 words, include 5+ images
- **Videos**: 30 seconds - 10 minutes, native upload preferred
- **Documents**: PDF carousels, 10-15 slides
- **Polls**: 4 options max, 1-2 week duration
#### Optimization Tips
- First 2 lines are crucial (shown in preview)
- Use emoji sparingly for visual breaks
- Include 3-5 relevant hashtags
- Tag people and companies when relevant
- Native video gets 5x more engagement
- Post consistently (3-5x per week optimal)
#### Algorithm Factors
- Dwell time (time spent reading)
- Comments valued over likes
- Early engagement (first hour) crucial
- Creator mode boosts reach
- Replies to comments increase visibility
### Twitter/X
**Audience**: News junkies, tech enthusiasts, real-time conversation
**Best Times**: Weekdays 9-10 AM and 7-9 PM
**Optimal Length**: 100-250 characters
#### Content Formats
- **Single Tweets**: 250 characters, 1-2 hashtags
- **Threads**: 5-15 tweets, numbered format
- **Images**: 16:9 ratio, up to 4 per tweet
- **Videos**: Up to 2:20, square or landscape
- **Polls**: 2-4 options, 5 minutes - 7 days
#### Optimization Tips
- Front-load important information
- Use threads for complex topics
- Include visuals (2-3x more engagement)
- Retweet with comment > regular RT
- Schedule threads for consistency
- Engage genuinely with replies
#### Algorithm Factors
- Engagement rate (likes, RTs, replies)
- Relationship (mutual follows prioritized)
- Recency over evergreen
- Topic relevance to user interests
- Link posts receive less reach
### Instagram
**Audience**: Visual-first, millennials & Gen Z, lifestyle focused
**Best Times**: Weekdays 11 AM - 1 PM and 7-9 PM
**Optimal Length**: 138-150 characters shown in preview
#### Content Formats
- **Feed Posts**: Square (1:1) or vertical (4:5)
- **Stories**: 15 seconds max, vertical (9:16)
- **Reels**: 15-90 seconds, vertical (9:16)
- **Carousels**: 2-10 images/videos
- **IGTV/Video**: 1-60 minutes
#### Optimization Tips
- First sentence crucial (caption preview)
- Use up to 30 hashtags (5-10 in caption, rest in comment)
- Carousel posts get highest engagement
- Stories with polls/questions boost views
- Reels get maximum organic reach
- Post consistently (1-2 feed posts daily)
#### Algorithm Factors
- Relationship (DMs, comments, tags)
- Interest (based on past interactions)
- Timeliness (newer posts prioritized)
- Frequency of app usage
- Time spent on posts (saves valuable)
### Facebook
**Audience**: Broad demographic, community-focused, local businesses
**Best Times**: Wednesday-Friday, 11 AM - 2 PM
**Optimal Length**: 50-80 characters for posts
#### Content Formats
- **Text Posts**: 50-80 characters optimal
- **Images**: 1200x630px for links
- **Videos**: 1-3 minutes, square format
- **Stories**: Same as Instagram
- **Live Videos**: Minimum 10 minutes
#### Optimization Tips
- Native video gets priority
- Ask questions to boost comments
- Share to relevant groups
- Use Facebook Creator Studio
- Tag locations for local reach
- Post 1-2 times per day max
#### Algorithm Factors
- Meaningful interactions (comments > reactions)
- Video completion rate
- Friends and family prioritized
- Group posts get high visibility
- Live videos get 6x engagement
### TikTok
**Audience**: Gen Z, entertainment-focused, trend-driven
**Best Times**: 6-10 AM and 7-11 PM
**Optimal Length**: 15-30 seconds
#### Content Formats
- **Videos**: 15 seconds - 10 minutes
- **Aspect Ratio**: 9:16 vertical
- **Sounds**: Trending audio crucial
- **Effects**: Filters and transitions
#### Optimization Tips
- Hook viewers in first 3 seconds
- Use trending sounds and hashtags
- Create content for FYP, not followers
- Post 1-4 times daily
- Engage with comments quickly
- Jump on trends within 24-48 hours
#### Algorithm Factors
- Completion rate most important
- Shares and saves valued
- Comment engagement
- Following similar creators
- Time spent on app
## Content Optimization Strategies
### Hashtag Strategy
#### Research Methods
1. **Competitor Analysis**: Study successful competitors
2. **Platform Search**: Use native search for suggestions
3. **Hashtag Tools**: RiteTag, Hashtagify, All Hashtag
4. **Trending Topics**: Monitor daily/weekly trends
5. **Brand Hashtags**: Create unique campaign tags
#### Hashtag Mix Formula
- 30% High-volume (1M+ posts)
- 40% Medium-volume (100K-1M posts)
- 30% Low-volume/Niche (<100K posts)
#### Platform-Specific Guidelines
- **Instagram**: 10-30 hashtags (mix in caption and first comment)
- **LinkedIn**: 3-5 professional hashtags
- **Twitter**: 1-2 hashtags max
- **Facebook**: 1-3 hashtags
- **TikTok**: 3-5 trending + niche tags
### Visual Content Optimization
#### Image Best Practices
- **Resolution**: Minimum 1080px width
- **File Size**: Under 5MB for faster loading
- **Alt Text**: Always include for accessibility
- **Branding**: Consistent filters/overlays
- **Text Overlay**: Less than 20% of image
#### Video Optimization
- **Captions**: Always include (85% watch without sound)
- **Thumbnail**: Custom, eye-catching
- **Length**: Platform-specific optimal duration
- **Format**: MP4 for best compatibility
- **Aspect Ratio**: Vertical for stories/reels, square for feed
### Caption Writing Formulas
#### AIDA Formula
- **Attention**: Hook in first line
- **Interest**: Expand on the hook
- **Desire**: Benefits and value
- **Action**: Clear CTA
#### PAS Formula
- **Problem**: Identify pain point
- **Agitate**: Emphasize consequences
- **Solution**: Present your answer
#### Before-After-Bridge
- **Before**: Current situation
- **After**: Desired outcome
- **Bridge**: How to get there
### Engagement Tactics
#### Conversation Starters
- Ask open-ended questions
- Create polls and surveys
- "Fill in the blank" posts
- "This or that" choices
- Caption contests
- Opinion requests
#### Community Building
- Respond to comments within 2 hours
- Like and reply to user comments
- Share user-generated content
- Create branded hashtags
- Host Q&A sessions
- Run challenges or contests
### Analytics & KPIs
#### Vanity Metrics (Track but don't obsess)
- Follower count
- Like count
- View count
#### Performance Metrics (Focus here)
- Engagement rate: (Likes + Comments + Shares) / Reach × 100
- Click-through rate: Clicks / Impressions × 100
- Conversion rate: Conversions / Clicks × 100
- Share/Save rate: Shares / Reach × 100
#### Business Metrics (Ultimate goal)
- Website traffic from social
- Lead generation
- Sales attribution
- Customer acquisition cost
- Customer lifetime value
### Content Calendar Planning
#### Weekly Posting Schedule Template
```
Monday: Motivational (Quote/Inspiration)
Tuesday: Educational (How-to/Tips)
Wednesday: Promotional (Product/Service)
Thursday: Engaging (Poll/Question)
Friday: Fun (Behind-scenes/Casual)
Saturday: User-Generated Content
Sunday: Curated Content/Rest
```
#### Monthly Theme Structure
- Week 1: Awareness content
- Week 2: Consideration content
- Week 3: Decision content
- Week 4: Retention/Community
### Crisis Management Protocol
#### Response Timeline
- **0-15 minutes**: Acknowledge awareness
- **15-60 minutes**: Gather facts
- **1-2 hours**: Official response
- **24 hours**: Follow-up update
- **48-72 hours**: Resolution summary
#### Response Guidelines
1. Acknowledge quickly
2. Take responsibility if appropriate
3. Show empathy
4. Provide facts only
5. Outline action steps
6. Follow up publicly
## Tool Stack Recommendations
### Content Creation
- **Design**: Canva, Adobe Creative Suite
- **Video**: CapCut, InShot, Adobe Premiere
- **Copy**: Grammarly, Hemingway Editor
- **AI Assistance**: ChatGPT, Claude, Jasper
### Scheduling & Management
- **All-in-One**: Hootsuite, Buffer, Sprout Social
- **Visual-First**: Later, Planoly
- **Enterprise**: Sprinklr, Khoros
- **Free Options**: Meta Business Suite, TweetDeck
### Analytics & Monitoring
- **Native**: Platform Insights/Analytics
- **Third-Party**: Socialbakers, Brandwatch
- **Listening**: Mention, Brand24
- **Competitor Analysis**: Social Blade, Rival IQ
### Influencer & UGC
- **Discovery**: AspireIQ, GRIN
- **Management**: CreatorIQ, Klear
- **UGC Curation**: TINT, Stackla
- **Rights Management**: Rights Manager
## Compliance & Best Practices
### Legal Considerations
- Include #ad or #sponsored for paid partnerships
- Respect copyright and attribution
- Follow GDPR for data collection
- Comply with platform terms of service
- Get permission for UGC usage
### Accessibility Guidelines
- Add alt text to all images
- Include captions on videos
- Use CamelCase for hashtags (#LikeThis)
- Avoid text-only images
- Ensure color contrast compliance
### Brand Safety
- Moderate comments regularly
- Set up keyword filters
- Have crisis management plan
- Monitor brand mentions
- Establish posting permissions
Chế độ giao tiếp nén tối đa, bỏ từ thừa để giảm khoảng 75% token mà vẫn giữ chính xác kỹ thuật.
---
name: caveman
description: >
Ultra-compressed communication mode. Cuts token usage ~75% by dropping
filler, articles, and pleasantries while keeping full technical accuracy.
Use when user says "caveman mode", "talk like caveman", "use caveman",
"less tokens", "be brief", or invokes /caveman.
license: MIT
metadata:
derived_from: "https://github.com/mattpocock/skills/tree/main/skills/productivity/caveman"
original_author: "Matt Pocock (@mattpocock)"
original_license: MIT
voice: "Matt Pocock — terse, fragment-OK, no filler"
version: 1.0.0
---
# Caveman Mode
> Derived from [Matt Pocock's caveman](https://github.com/mattpocock/skills/tree/main/skills/productivity/caveman) (MIT). Matt's voice preserved verbatim. Additions: compression tools + references + cs-* wrapper (see [references/companion_tooling.md](references/companion_tooling.md)).
Respond terse like smart caveman. All technical substance stay. Only fluff die.
## Persistence
ACTIVE EVERY RESPONSE once triggered. No revert after many turns. No filler drift. Still active if unsure. Off only when user says "stop caveman" or "normal mode".
## Rules
Drop: articles (a/an/the), filler (just/really/basically/actually/simply), pleasantries (sure/certainly/of course/happy to), hedging. Fragments OK. Short synonyms (big not extensive, fix not "implement a solution for"). Abbreviate common terms (DB/auth/config/req/res/fn/impl). Strip conjunctions. Use arrows for causality (X -> Y). One word when one word enough.
Technical terms stay exact. Code blocks unchanged. Errors quoted exact.
Pattern: `[thing] [action] [reason]. [next step].`
Not: "Sure! I'd be happy to help you with that. The issue you're experiencing is likely caused by..."
Yes: "Bug in auth middleware. Token expiry check use `<` not `<=`. Fix:"
### Examples
**"Why React component re-render?"**
> Inline obj prop -> new ref -> re-render. `useMemo`.
**"Explain database connection pooling."**
> Pool = reuse DB conn. Skip handshake -> fast under load.
## Auto-Clarity Exception
Drop caveman temporarily for: security warnings, irreversible action confirmations, multi-step sequences where fragment order risks misread, user asks to clarify or repeats question. Resume caveman after clear part done.
Example -- destructive op:
> **Warning:** This will permanently delete all rows in the `users` table and cannot be undone.
>
> ```sql
> DROP TABLE users;
> ```
>
> Caveman resume. Verify backup exist first.
## Tooling
See [references/companion_tooling.md](references/companion_tooling.md). Tools: compressor + estimator + lint. Agent: `cs-caveman-mode`. Command: `/cs:caveman`.
---
**Version:** 1.0.0
**Derived:** Matt Pocock (MIT) + this repo's wrapper
FILE:references/companion_tooling.md
# Companion Tooling
Compression tools + cs-* wrapper layered on top of Matt's caveman skill.
## Validation Tools (stdlib Python)
| Tool | Purpose | Run when |
|---|---|---|
| `scripts/caveman_compressor.py` | Apply Matt's rules deterministically (drop articles/filler/pleasantries/hedging, abbreviate technical terms, use causality arrows) | Want a starting compressed version of any text |
| `scripts/token_savings_estimator.py` | Estimate token + cost savings using 4 chars/token (prose) or 3.5 chars/token (technical) heuristic | Want to quantify the value of caveman mode |
| `scripts/caveman_lint.py` | Detect banned vocabulary in a response (pleasantries, filler, hedging, metatalk, verbose phrases). Whitelist: code blocks, inline code, exception zones | Verify a response complies with caveman rules |
All three tools:
- Stdlib-only (no external dependencies)
- Run with embedded sample if no input provided
- Output text or JSON (`--output json`)
- Code blocks + inline code preserved (compression skips them)
## Token-Savings Heuristic
The estimator uses character-per-token approximations:
- **4.0 chars/token** for English prose
- **3.5 chars/token** for technical text (detected by presence of `{`, `}`, `()`, `->`, `==`, `//`, etc.)
This is within 10-15% of cl100k_base / o200k_base tokenizers for English. For exact token counts use the model's actual tokenizer (e.g., `tiktoken`).
## cs-caveman-mode Persona Agent
Lives at `../agents/cs-caveman-mode.md`. Voice: terse, fragments-OK, no filler. Persistence is the hard rule — once activated stays active until "stop caveman" / "normal mode".
## `/cs:caveman` Slash Command
Lives at `../commands/cs-caveman.md`. Single-trigger activation. Equivalent to typing "caveman mode" but more explicit.
## When Caveman Backfires (See main SKILL.md "Auto-Clarity Exception")
The compressor + lint tool both whitelist these zones — Matt's rule is explicit:
- Security warnings
- Irreversible action confirmations
- Multi-step sequences where fragment order risks misread
- User asks to clarify or repeats question
The lint tool detects `**Warning:**`, `destructive`, `irreversible`, `cannot be undone` markers and softens its verdict accordingly.
## Why Wrap Matt's Original
Matt's caveman skill is tight + complete. The wrapper adds:
1. **Deterministic compression** — apply rules consistently across responses (not just in spirit)
2. **Quantification** — show ROI of caveman mode in tokens/dollars
3. **Compliance checking** — verify a response actually follows rules (vs claiming to)
## Attribution
Original: [matt-pocock/skills/skills/productivity/caveman](https://github.com/mattpocock/skills/tree/main/skills/productivity/caveman) (MIT).
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — caveman** (https://github.com/mattpocock/skills/, MIT) — the upstream source
- **Anthropic — Token usage best practices** (https://docs.claude.com/en/docs/build-with-claude/prompt-engineering) — token-conscious prompting
- **OpenAI tokenizer docs** — `tiktoken` library + cl100k_base / o200k_base heuristics
- **Strunk & White — "The Elements of Style"** (1918) — "omit needless words"; foundational text on prose compression
- **Plain Language Movement / Plain Writing Act of 2010** — federal mandate for concise government writing
- **Norman, D. — "Living with Complexity"** (2010) — when simplicity helps vs hurts cognition
- **Pareto principle in communication** — 20% of words carry 80% of information density
FILE:references/compression_principles.md
# Compression Principles for LLM Output
This reference answers exactly one decision: **what should be cut and what must stay when compressing LLM output for token efficiency?**
Pair with `scripts/caveman_compressor.py` for deterministic application.
## Matt Pocock's Foundational Insight
> "Respond terse like smart caveman. All technical substance stay. Only fluff die."
>
> — Matt Pocock, caveman SKILL.md
The crucial distinction: **substance** vs **fluff**. Caveman mode is aggressive about fluff and conservative about substance. Confusion between the two creates either bloated responses (under-cutting) or hallucinated answers (over-cutting).
## What Counts as Fluff (Safe to Drop)
| Category | Examples | Why safe to drop |
|---|---|---|
| **Articles** | a, an, the | Grammatical scaffolding; meaning preserved without them |
| **Filler** | just, really, basically, actually, simply, obviously | Add no information; speakers use as verbal pauses |
| **Pleasantries** | sure!, certainly, of course, happy to help | Social lubrication; cost tokens with zero info gain |
| **Hedging** | might, maybe, perhaps, likely, possibly | Either qualify with data or remove; vague hedging is fake precision |
| **Metatalk** | as you can see, worth noting, that said | Self-referential commentary about the response itself |
| **Verbose phrases** | "implementation of a solution for" → "fix"; "in order to" → "to" | Phrase-level redundancy |
## What Counts as Substance (Must Stay)
| Category | Examples | Why preserve |
|---|---|---|
| **Technical terms** | `useMemo`, NULL, HTTP/2, OAuth2 | Exact names matter; abbreviation breaks identifiers |
| **Code blocks** | All ```...``` regions | Syntactically meaningful; whitespace + characters matter |
| **Inline code** | `useState`, `auth_token` | Same as code blocks |
| **Quoted strings** | "expected value", 'string literal' | Exact text matters |
| **Error messages** | "TypeError: cannot read property X" | Diagnostic precision required |
| **Numbers + units** | 200ms, 4kb, 99.9% | Exactness matters for engineering decisions |
| **Causal claims** | "X causes Y" — can be compressed to "X -> Y" | The relationship is the substance |
## The Abbreviation Cost-Benefit
Abbreviating common technical terms saves tokens but only when:
1. The abbreviation is universally understood (DB, auth, config, fn — yes; ETL, ORM — maybe; "imp" for implementation — no)
2. The reader has full context (caveman responses are usually mid-conversation)
3. The exact term isn't being introduced (don't abbreviate the FIRST use of a term)
Matt's abbreviation list is conservative + universal:
- DB, auth, config, req, res, fn, impl, env, deps, repo, docs, app
## Causality Arrows: The Compression Win
Replacing verbose causality with arrows is high-leverage:
| Verbose | Caveman | Savings |
|---|---|---|
| "X leads to Y" (3 words) | "X -> Y" (1 unit) | 67% |
| "which causes Y to happen" (5 words) | "-> Y" (2 units) | 60% |
| "because of X, Y happens" (5 words) | "Y <- X" (2 units) | 60% |
Arrows are unambiguous + compact + preserve causality (not just adjacency).
## Compression Anti-Patterns
1. **Dropping subject pronouns at all costs** — "Bug in auth" is fine. "Auth bug, fix soon" loses clarity. Keep enough syntax to disambiguate.
2. **Over-abbreviating** — "MWMV" instead of "memory write/memory verify" forces reader to expand mentally; net cognitive cost goes up.
3. **Dropping units** — "Response takes 200" — 200 what? ms? bytes? Keep units always.
4. **Compressing security warnings** — Matt's explicit exception. A truncated security warning is worse than no caveman mode.
5. **Dropping examples** — "Bug in auth. Fix." — what bug? what fix? Caveman keeps the substance, just removes the wrapping.
## Compression vs Clarity Tradeoff
Compression is a tax on the reader. The trade-off is worth it when:
- The reader has the context to fill in the gaps (mid-conversation, technical peer)
- The information density is high enough to justify cognitive load
- The savings are meaningful (>20% token reduction)
Not worth it when:
- New context being established (introductions, first turns)
- Multi-step sequences where order matters
- Multi-stakeholder communication (caveman style confuses non-technical readers)
- Audio interfaces (caveman text reads badly when read aloud)
## How Much Compression Is Realistic?
Matt's claim is ~75% — this is the upper bound on extremely verbose responses (with multiple pleasantries + filler + hedging). Realistic ranges:
| Response type | Realistic compression |
|---|---|
| ChatGPT-style verbose response | 50-75% |
| Already-concise technical answer | 10-25% |
| Code-heavy response (most text is code) | 5-15% |
| Single-sentence answer | 0-30% |
The compressor in this skill targets 20-50% on typical mid-conversation responses, which is meaningful at scale.
## When This Reference Doesn't Help
- **Code minification** — different concern; this is about prose around code, not code itself
- **Prompt compression for inputs** — different mode; input compression has different rules
- **Speech synthesis** — caveman text reads poorly aloud
- **Marketing copy** — different goal; conversion > brevity
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — caveman** (https://github.com/mattpocock/skills/, MIT) — the upstream source + rule set
- **Strunk & White — "The Elements of Style"** (1918) — Rule 17: "Omit needless words"
- **Plain Language Movement / Plain Writing Act of 2010** (https://www.plainlanguage.gov/) — government mandate for concise English; well-researched compression rules
- **Pinker, S. — "The Sense of Style"** (2014) — cognitive science of clear writing
- **Williams, J. — "Style: Toward Clarity and Grace"** (1995) — academic compression patterns
- **Anthropic — Prompt engineering for tokens** (https://docs.claude.com/en/docs/build-with-claude/prompt-engineering) — token-conscious patterns
- **OpenAI tokenizer documentation** — character-per-token ratios across cl100k_base / o200k_base
- **Pareto principle in writing** — 20% of words carry 80% of meaning
FILE:references/when_caveman_backfires.md
# When Caveman Backfires
This reference answers exactly one decision: **when should caveman mode NOT be used, and what are the failure modes?**
Pair with `scripts/caveman_lint.py` — the linter detects exception-zone markers and softens its verdict accordingly.
## Matt Pocock's Auto-Clarity Exception (Verbatim)
> "Drop caveman temporarily for: security warnings, irreversible action confirmations, multi-step sequences where fragment order risks misread, user asks to clarify or repeats question. Resume caveman after clear part done."
>
> — Matt Pocock, caveman SKILL.md
This is the **non-negotiable** exception list. Compressing in these zones can cause user harm — not just token cost confusion.
## The Five Failure Modes
### 1. Compressed Security Warnings
**Failure:** `Warning: drop users table. Cannot undo.`
**Why bad:** Two compound failures:
1. Reader skims caveman text; warning blends into surrounding terse responses
2. Critical verb ("drop", "cannot undo") loses emphasis without sentence structure
**Correct:** Drop caveman entirely for warnings:
> **Warning:** This will permanently delete all rows in the `users` table and cannot be undone.
The full sentence + bold marker + explicit consequences make it unmissable.
### 2. Compressed Multi-Step Sequences
**Failure:** `npm install. npm test. npm build. deploy.`
**Why bad:** Fragment order = execution order. If reader misreads order, deploy can run before tests pass.
**Correct:** Number steps + use full sentences when order matters:
> 1. Install dependencies: `npm install`
> 2. Run tests: `npm test` — must pass before continuing
> 3. Build: `npm build`
> 4. Deploy: only after step 3 succeeds
### 3. Compressed Clarification Responses
**Failure:** User asks "Wait, what?" → Response: `Token expiry. Less-than. Fix.`
**Why bad:** User asked for clarification because original was unclear. Compressing further compounds the problem.
**Correct:** Drop caveman temporarily, explain fully, then resume:
> The original issue: the token expiry check uses `<` (strict less-than) instead of `<=` (less-than-or-equal). This means tokens at the exact expiry timestamp are rejected even though they should still be valid. The fix changes the operator from `<` to `<=`.
>
> [Resume caveman.]
### 4. Compressed Irreversible Action Confirmations
**Failure:** `rm -rf node_modules. Confirm.`
**Why bad:** "Confirm" loses weight without context. User may type "yes" reflexively.
**Correct:** Drop caveman, state explicitly:
> About to run `rm -rf node_modules` which permanently deletes the directory.
>
> Reply with the exact string "DELETE" to proceed, or "cancel" to abort.
The exact-string requirement breaks reflex confirmation.
### 5. Compressed First-Turn Responses
**Failure:** User's first message → Response in caveman.
**Why bad:** No shared context yet. Reader can't fill in caveman's gaps.
**Correct:** First turn establishes context fully. Activate caveman ONLY after user explicitly triggers it (per Matt's activation triggers: "caveman mode", "talk like caveman", `/caveman`, etc.).
## Less-Obvious Backfire Cases
### Caveman in Code Review
Caveman compression on code-review feedback can lose nuance:
**Failure:** `Bug L42. Var name bad. Refactor.`
**Why bad:** Three findings, no specificity. Engineer can't tell what to fix.
**Better:** `L42: var name "x" → "userIndex". L67: off-by-one in loop bound.`
The fix: caveman compresses sentence STRUCTURE, not technical SPECIFICITY.
### Caveman in Estimates / Forecasts
Hedging is fluff per Matt's rules. But hedging carries information in estimates:
**Failure:** `Done by Friday.` (when uncertain)
**Why bad:** Reads as commitment, but actual confidence was 60%.
**Correct:** Caveman exception for probability claims. State confidence explicitly:
> Friday delivery — 60% confidence. Risks: API spec churn.
### Caveman in Multi-Stakeholder Threads
Caveman is for technical peer-to-peer (or peer-to-self) communication. When non-technical stakeholders are reading:
**Failure:** `Auth bug. Fix shipping.`
**Why bad:** PM/CEO/non-engineer reader can't decode "Fix shipping" — is shipping affected?
**Correct:** Drop caveman in stakeholder communication. Save it for technical conversations.
## Detection Patterns (How `caveman_lint.py` Helps)
The lint tool detects these markers as exception-zone signals:
- `**Warning:**` markdown bold + word
- `destructive`
- `irreversible`
- `cannot be undone`
When present, the linter softens FAIL → WARN. This isn't perfect — manual review still required for stakeholder mismatches + first-turn responses.
## Resuming Caveman After Exception
Matt's rule: "Resume caveman after clear part done."
Pattern:
> **Warning:** [full sentence warning].
>
> [empty line]
>
> Caveman resume. [terse fragment continues].
The explicit "Caveman resume." marker signals the reader that compression resumes. This is critical when the response is long enough that the reader might lose track of which mode they're in.
## Tooling Recommendation
When in doubt:
1. Run `caveman_lint.py` on the proposed response
2. If FAIL → consider rewriting (banned vocab present)
3. If WARN with exception context → check whether the exception is genuine
4. If CLEAN → ship
## When This Reference Doesn't Help
- **Brevity in writing generally** — different concern; see editing references
- **Code minification** — different mode; this is about prose around code
- **API response compression** — gzip/brotli, not prose compression
---
**Source authorities (non-exhaustive):**
- **Matt Pocock — caveman** (https://github.com/mattpocock/skills/, MIT) — the auto-clarity exception list
- **Nielsen Norman Group — Error message design** — when verbosity in errors helps vs hurts
- **FAA Human Factors research on cockpit warnings** — emphasis + redundancy in safety-critical communications
- **Krug, S. — "Don't Make Me Think"** (2000) — when brevity becomes ambiguity
- **Schneier, B. — Communication on security warnings** — why brevity in security messages is dangerous
- **Larson, W. — "An Elegant Puzzle"** (2019) — engineering manager communication patterns
- **Rommetveit, R. — Linguistic shared context** — when compression depends on shared frame
FILE:scripts/caveman_compressor.py
#!/usr/bin/env python3
"""caveman_compressor.py — Apply Matt Pocock's caveman compression rules to text.
Stdlib-only. Deterministic regex-based compression matching the rules in
Matt Pocock's caveman skill SKILL.md:
1. Drop articles (a/an/the)
2. Drop filler (just/really/basically/actually/simply)
3. Drop pleasantries (sure/certainly/of course/happy to)
4. Drop hedging (might/maybe/perhaps/likely/possibly)
5. Abbreviate common technical terms (database -> DB, configuration -> config, etc.)
6. Strip conjunctions where safe (and/but at sentence start)
7. Use arrows for "leads to" / "causes" phrases (-> )
8. Strip "as you can see / it should be noted / it's worth mentioning"
PRESERVES:
- Code blocks (```...```) unchanged
- Inline code (`...`) unchanged
- Technical terms named verbatim
- Quoted strings unchanged
NO LLM CALLS. Stdlib only.
Usage:
python caveman_compressor.py # uses embedded sample
python caveman_compressor.py "your text here"
python caveman_compressor.py --file path/to/input.txt
python caveman_compressor.py "text" --output json
"""
import argparse
import json
import re
import sys
from typing import Any, Dict, List, Tuple
# Filler/pleasantry/hedging vocabularies (per Matt's rules)
ARTICLES = {"a", "an", "the"}
FILLER = {"just", "really", "basically", "actually", "simply", "obviously", "literally"}
PLEASANTRIES_PHRASES = [
"sure!", "sure,", "certainly!", "certainly,",
"of course!", "of course,",
"happy to help", "i'd be happy to", "i would be happy to",
"great question", "good question",
"absolutely!", "absolutely,",
"no problem!", "no problem,",
]
HEDGING = {"might", "maybe", "perhaps", "likely", "possibly", "probably"}
METATALK_PHRASES = [
"as you can see",
"it should be noted",
"it's worth mentioning",
"it is worth mentioning",
"needless to say",
"to be clear",
"in other words",
"that said",
"having said that",
]
# Technical term abbreviations
ABBREVIATIONS = [
(r"\bdatabase\b", "DB"),
(r"\bdatabases\b", "DBs"),
(r"\bauthentication\b", "auth"),
(r"\bauthorization\b", "authz"),
(r"\bconfiguration\b", "config"),
(r"\bconfigurations\b", "configs"),
(r"\brequest\b", "req"),
(r"\brequests\b", "reqs"),
(r"\bresponse\b", "res"),
(r"\bresponses\b", "ress"),
(r"\bfunction\b", "fn"),
(r"\bfunctions\b", "fns"),
(r"\bimplementation\b", "impl"),
(r"\bimplementations\b", "impls"),
(r"\benvironment\b", "env"),
(r"\bdependencies\b", "deps"),
(r"\bdependency\b", "dep"),
(r"\brepository\b", "repo"),
(r"\brepositories\b", "repos"),
(r"\bdocumentation\b", "docs"),
(r"\bapplication\b", "app"),
(r"\bapplications\b", "apps"),
]
# Causality phrase -> arrow
CAUSALITY_PATTERNS = [
(re.compile(r"\b(which\s+)?(leads?|causes?|results?\s+in|gives?\s+you|produces?)\s+", re.IGNORECASE), "-> "),
(re.compile(r"\bbecause\s+of\b", re.IGNORECASE), "<- "),
]
# Embedded sample
SAMPLE_INPUT = (
"Sure! I'd be happy to help you with that. The issue you're experiencing is "
"likely caused by a misconfiguration in the authentication middleware, where "
"the token expiry check is actually using a strict less-than comparison "
"instead of less-than-or-equal. This basically means tokens at the exact "
"expiry timestamp will get rejected. To fix this, you should simply update "
"the configuration of the auth function to use `<=` instead of `<`."
)
def _protect_code(text: str) -> Tuple[str, List[str]]:
"""Replace code blocks + inline code with placeholders, return text + protected list."""
protected: List[str] = []
def replace_block(m: re.Match) -> str:
protected.append(m.group(0))
return f"\x00CODE{len(protected) - 1}\x00"
text = re.sub(r"```.*?```", replace_block, text, flags=re.DOTALL)
text = re.sub(r"`[^`]+`", replace_block, text)
return text, protected
def _restore_code(text: str, protected: List[str]) -> str:
for i, code in enumerate(protected):
text = text.replace(f"\x00CODE{i}\x00", code)
return text
def _drop_articles(text: str) -> str:
pattern = re.compile(r"\b(" + "|".join(ARTICLES) + r")\s+", re.IGNORECASE)
return pattern.sub("", text)
def _drop_word_set(text: str, words: set) -> str:
pattern = re.compile(r"\b(" + "|".join(words) + r")\b\s*", re.IGNORECASE)
return pattern.sub("", text)
def _drop_phrases(text: str, phrases: List[str]) -> str:
for phrase in phrases:
text = re.sub(re.escape(phrase) + r"\s*", "", text, flags=re.IGNORECASE)
text = re.sub(re.escape(phrase.rstrip(",!")) + r"\s*", "", text, flags=re.IGNORECASE)
return text
def _apply_abbreviations(text: str) -> str:
for pattern, replacement in ABBREVIATIONS:
text = re.sub(pattern, replacement, text, flags=re.IGNORECASE)
return text
def _apply_causality_arrows(text: str) -> str:
for pattern, replacement in CAUSALITY_PATTERNS:
text = pattern.sub(replacement, text)
return text
def _strip_leading_conjunctions(text: str) -> str:
return re.sub(r"(^|\.\s+)(and|but|so)\s+", r"\1", text, flags=re.IGNORECASE)
def _collapse_whitespace(text: str) -> str:
text = re.sub(r"\s+", " ", text)
text = re.sub(r"\s+([.,;:!?])", r"\1", text)
return text.strip()
def compress(text: str) -> str:
"""Apply Matt Pocock's caveman rules. Returns compressed text."""
text, protected = _protect_code(text)
text = _drop_phrases(text, PLEASANTRIES_PHRASES)
text = _drop_phrases(text, METATALK_PHRASES)
text = _drop_word_set(text, FILLER)
text = _drop_word_set(text, HEDGING)
text = _drop_articles(text)
text = _apply_abbreviations(text)
text = _apply_causality_arrows(text)
text = _strip_leading_conjunctions(text)
text = _collapse_whitespace(text)
text = _restore_code(text, protected)
return text
def analyze(original: str, compressed: str) -> Dict[str, Any]:
orig_words = len(original.split())
new_words = len(compressed.split())
saved = orig_words - new_words
pct = round(100.0 * saved / max(orig_words, 1), 1)
return {
"original_chars": len(original),
"compressed_chars": len(compressed),
"original_words": orig_words,
"compressed_words": new_words,
"words_saved": saved,
"percent_savings": pct,
"compressed_text": compressed,
}
def render_text(original: str, result: Dict[str, Any]) -> str:
lines = []
lines.append("=" * 72)
lines.append("CAVEMAN COMPRESSOR")
lines.append("=" * 72)
lines.append("")
lines.append("ORIGINAL:")
lines.append(f" {original}")
lines.append("")
lines.append("COMPRESSED:")
lines.append(f" {result['compressed_text']}")
lines.append("")
lines.append("-" * 72)
lines.append(f"Chars: {result['original_chars']} -> {result['compressed_chars']}")
lines.append(f"Words: {result['original_words']} -> {result['compressed_words']}")
lines.append(f"Savings: {result['words_saved']} words ({result['percent_savings']}%)")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Compress text per Matt Pocock's caveman rules.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("text", nargs="?", help="Input text (uses embedded sample if omitted)")
parser.add_argument("--file", help="Read input from file")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.file:
try:
with open(args.file, "r", encoding="utf-8") as f:
original = f.read()
except (IOError, OSError) as e:
print(f"error: {e}", file=sys.stderr)
return 1
elif args.text:
original = args.text
else:
original = SAMPLE_INPUT
compressed = compress(original)
result = analyze(original, compressed)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_text(original, result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/caveman_lint.py
#!/usr/bin/env python3
"""caveman_lint.py — Lint a response for caveman-mode compliance.
Stdlib-only. Detects banned vocabulary in a response that's supposed to be in
caveman mode. Returns specific findings + verdict.
Banned categories per Matt Pocock's caveman rules:
- Pleasantries (sure, certainly, of course, happy to)
- Filler (just, really, basically, actually, simply)
- Hedging (might, maybe, perhaps, likely)
- Metatalk (as you can see, worth noting)
- Verbose phrases ("the implementation of a solution for")
Whitelist (NOT banned even in caveman mode):
- Words inside code blocks
- Words inside inline code
- Words inside quoted strings
- Caveman exception zones (security warnings, destructive op confirmations)
Usage:
python caveman_lint.py # uses embedded samples
python caveman_lint.py "response text"
python caveman_lint.py --file path/to/response.txt
python caveman_lint.py "text" --output json
"""
import argparse
import json
import re
import sys
from typing import Any, Dict, List
BANNED_PHRASES = {
"pleasantry": [
"sure!", "sure,", "certainly", "of course", "happy to help",
"i'd be happy", "i would be happy", "great question", "good question",
"absolutely", "no problem!",
],
"filler": ["just", "really", "basically", "actually", "simply", "obviously", "literally"],
"hedging": ["might", "maybe", "perhaps", "likely", "possibly", "probably"],
"metatalk": [
"as you can see", "it should be noted", "worth mentioning",
"needless to say", "to be clear", "in other words",
"that said", "having said that",
],
"verbose": [
"implement a solution for", "the implementation of",
"in order to", "for the purpose of", "with respect to",
"due to the fact that",
],
}
# Patterns that DROP caveman temporarily (whitelisted zones)
EXCEPTION_MARKERS = [
re.compile(r"\*\*warning:\*\*", re.IGNORECASE),
re.compile(r"\bdestructive\b", re.IGNORECASE),
re.compile(r"\birreversible\b", re.IGNORECASE),
re.compile(r"\bcannot be undone\b", re.IGNORECASE),
]
SAMPLE_BAD = (
"Sure! I'd be happy to help. The issue is actually quite simple — basically, "
"you just need to update the configuration. It's worth mentioning that this might "
"cause a slight performance hit, but probably not noticeable."
)
SAMPLE_GOOD = "Bug in auth middleware. Token expiry check use `<` not `<=`. Fix: change to `<=`."
def _protect_code(text: str) -> str:
"""Mask code blocks + inline code so banned-word matching skips them."""
text = re.sub(r"```.*?```", lambda m: "\x00" * len(m.group(0)), text, flags=re.DOTALL)
text = re.sub(r"`[^`]+`", lambda m: "\x00" * len(m.group(0)), text)
return text
def _has_exception_context(text: str) -> bool:
return any(p.search(text) for p in EXCEPTION_MARKERS)
def _count_phrase(phrase: str, masked: str) -> int:
return len(re.findall(r"\b" + re.escape(phrase) + r"\b", masked, re.IGNORECASE))
def _violation_record(category: str, phrase: str, count: int) -> Dict[str, Any]:
return {"category": category, "phrase": phrase, "count": count}
def find_violations(text: str) -> List[Dict[str, Any]]:
"""Find banned phrases. Returns list of {category, phrase, count}."""
masked = _protect_code(text)
violations: List[Dict[str, Any]] = []
for category, phrases in BANNED_PHRASES.items():
for phrase in phrases:
count = _count_phrase(phrase, masked)
if count > 0:
violations.append(_violation_record(category, phrase, count))
return violations
def analyze(text: str) -> Dict[str, Any]:
violations = find_violations(text)
total_violations = sum(v["count"] for v in violations)
has_exception = _has_exception_context(text)
# Verdict logic:
# 0 violations + reasonable length -> CLEAN
# <= 2 violations OR exception context -> WARN
# > 2 violations -> FAIL
if total_violations == 0:
verdict = "CLEAN"
elif has_exception:
verdict = "WARN"
# When there's a security warning, some normal language is allowed
elif total_violations <= 2:
verdict = "WARN"
else:
verdict = "FAIL"
return {
"char_count": len(text),
"word_count": len(text.split()),
"violation_categories": sorted(set(v["category"] for v in violations)),
"total_violations": total_violations,
"has_exception_context": has_exception,
"violations": violations,
"verdict": verdict,
}
def render_text(text: str, r: Dict[str, Any]) -> str:
lines = []
lines.append("=" * 72)
lines.append("CAVEMAN LINT")
lines.append("=" * 72)
lines.append("")
preview = text[:200] + ("..." if len(text) > 200 else "")
lines.append(f"Text ({r['char_count']} chars, {r['word_count']} words):")
lines.append(f" {preview}")
lines.append("")
lines.append("-" * 72)
lines.append(f"Violations: {r['total_violations']}")
lines.append(f"Categories hit: {r['violation_categories']}")
if r["has_exception_context"]:
lines.append("Exception context detected (warning/destructive zone — some prose allowed)")
lines.append("")
if r["violations"]:
for v in r["violations"]:
lines.append(f" [{v['category']:11s}] x{v['count']:2d} '{v['phrase']}'")
else:
lines.append(" No banned phrases found.")
lines.append("")
lines.append("-" * 72)
lines.append(f"Verdict: {r['verdict']}")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Lint a response for caveman-mode compliance.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
parser.add_argument("text", nargs="?", help="Input text (uses embedded sample if omitted)")
parser.add_argument("--file", help="Read input from file")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
args = parser.parse_args()
if args.file:
try:
with open(args.file, "r", encoding="utf-8") as f:
text = f.read()
except (IOError, OSError) as e:
print(f"error: {e}", file=sys.stderr)
return 1
elif args.text:
text = args.text
else:
text = SAMPLE_BAD
result = analyze(text)
if args.output == "json":
print(json.dumps({"text": text, **result}, indent=2))
else:
print(render_text(text, result))
return 0 if result["verdict"] == "CLEAN" else 1
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/token_savings_estimator.py
#!/usr/bin/env python3
"""token_savings_estimator.py — Estimate token-cost savings from caveman compression.
Stdlib-only. Uses a chars-per-token heuristic (4 chars/token average for English
prose; 3.5 for technical text) to estimate output tokens before vs after caveman
compression.
Why heuristic and not real tokenizer:
- No external dependencies (stdlib only)
- Tokenizer accuracy varies by model (cl100k_base vs o200k_base vs others)
- Heuristic is within 10-15% of real tokenizer output for English prose
- Reports both heuristic + character count so user can apply their own multiplier
Usage:
python token_savings_estimator.py # uses embedded sample
python token_savings_estimator.py "your text"
python token_savings_estimator.py --file path/to/input.txt
python token_savings_estimator.py "text" --output json
python token_savings_estimator.py "text" --price-per-mtok 3.00
"""
import argparse
import json
import sys
from typing import Any, Dict
# Import the compressor as a module
import os
_HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, _HERE)
from caveman_compressor import compress, SAMPLE_INPUT # noqa: E402
# Heuristic: average chars per token
CHARS_PER_TOKEN_PROSE = 4.0
CHARS_PER_TOKEN_TECHNICAL = 3.5
TECHNICAL_TOKEN_INDICATORS = ("```", "{", "}", "()", "->", "==", "//", "/*", "import ", "function ")
def _estimate_chars_per_token(text: str) -> float:
"""Heuristic: technical text has more tokens per char than prose."""
hit_count = sum(1 for sig in TECHNICAL_TOKEN_INDICATORS if sig in text)
if hit_count >= 3:
return CHARS_PER_TOKEN_TECHNICAL
return CHARS_PER_TOKEN_PROSE
def estimate_tokens(text: str) -> int:
return int(round(len(text) / _estimate_chars_per_token(text)))
def analyze(original: str, price_per_mtok: float = 0.0) -> Dict[str, Any]:
compressed = compress(original)
orig_tokens = estimate_tokens(original)
new_tokens = estimate_tokens(compressed)
saved = orig_tokens - new_tokens
pct = round(100.0 * saved / max(orig_tokens, 1), 1)
out: Dict[str, Any] = {
"original_chars": len(original),
"compressed_chars": len(compressed),
"chars_per_token_used": _estimate_chars_per_token(original),
"estimated_original_tokens": orig_tokens,
"estimated_compressed_tokens": new_tokens,
"tokens_saved": saved,
"percent_token_savings": pct,
"compressed_preview": compressed[:200] + ("..." if len(compressed) > 200 else ""),
}
if price_per_mtok > 0:
cost_per_token = price_per_mtok / 1_000_000.0
out["price_per_million_tokens"] = price_per_mtok
out["cost_saved_per_response_usd"] = round(saved * cost_per_token, 6)
out["cost_saved_per_1k_responses_usd"] = round(saved * cost_per_token * 1000, 4)
return out
def render_text(r: Dict[str, Any]) -> str:
lines = []
lines.append("=" * 72)
lines.append("TOKEN SAVINGS ESTIMATOR (caveman compression)")
lines.append("=" * 72)
lines.append("")
lines.append(f"Chars/token heuristic: {r['chars_per_token_used']:.1f} (prose=4.0; technical=3.5)")
lines.append("")
lines.append(f"Original: {r['original_chars']} chars ~ {r['estimated_original_tokens']} tokens")
lines.append(f"Compressed: {r['compressed_chars']} chars ~ {r['estimated_compressed_tokens']} tokens")
lines.append("")
lines.append(f"Savings: {r['tokens_saved']} tokens ({r['percent_token_savings']}%)")
if "price_per_million_tokens" in r:
lines.append("")
lines.append(f"At r['price_per_million_tokens']/Mtok:")
lines.append(f" Cost saved per response: .6f")
lines.append(f" Cost saved per 1k responses: .4f")
lines.append("")
lines.append("-" * 72)
lines.append("Compressed preview:")
lines.append(f" {r['compressed_preview']}")
return "\n".join(lines)
def main() -> int:
parser = argparse.ArgumentParser(
description="Estimate token + cost savings from caveman compression.",
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog=__doc__,
)
price_help = "Per-million-token price (USD) to estimate cost savings"
parser.add_argument("text", nargs="?", help="Input text (uses embedded sample if omitted)")
parser.add_argument("--file", help="Read input from file")
parser.add_argument("--output", choices=("text", "json"), default="text", help="Output format")
parser.add_argument("--price-per-mtok", type=float, default=0.0, help=price_help)
args = parser.parse_args()
if args.file:
try:
with open(args.file, "r", encoding="utf-8") as f:
original = f.read()
except (IOError, OSError) as e:
print(f"error: {e}", file=sys.stderr)
return 1
elif args.text:
original = args.text
else:
original = SAMPLE_INPUT
result = analyze(original, args.price_per_mtok)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_text(result))
return 0
if __name__ == "__main__":
sys.exit(main())
Tìm bài báo qua Consensus, xây kế hoạch tìm kiếm theo PICO hoặc SPIDER và tổng hợp thành hướng dẫn nghiên cứu định dạng Word (.docx).
---
name: litreview
description: "Academic literature orientation skill that searches papers via Consensus, builds a strategic search plan using PICO (default) or SPIDER / Decomposition / hybrid as fallbacks, and synthesizes findings into a professionally formatted Word document (.docx) research guide. Grill-me intake (research question specificity + framework hint + tentative depth) before the recon search; a second forcing checkpoint after Phase 2 confirms framework + sub-areas + depth before searches consume budget. Configurable depth (5/10/20 queries) controls coverage vs. speed. Output is a 'launching pad' — not a finished review, but an orientation guide that lets a researcher dive in confidently. Triggers: 'litreview on [topic]', 'literature review on [topic]', 'I'm starting a literature review on X', 'I'm writing a paper on X', 'help me research X', 'I'm doing research on X', 'can you help me research X'. Do NOT trigger for single one-off paper searches where the user just wants a quick list — that's a plain Consensus search."
license: MIT
metadata:
source_spec: "megaprompts/09-litreview-megaprompt.md"
build_pattern: "Path B (direct conversion)"
research_pack_convention: "Agent Integrity Rules verbatim per PR #657 audit; sibling of pulse"
version: 1.0.0
---
# Litreview — Academic Literature Orientation
> **Portability:** Requires a Consensus MCP connection, Node.js with `docx` package for document generation, and (in CLI) `bash_tool`. Works in Claude Code CLI natively. In Claude.ai with Consensus MCP + Code Execution, the workflow is supported.
Produce a **launching pad** — not a finished literature review, but an orientation document that gives a researcher entering an unfamiliar field everything they need to start reading and searching with confidence. Think: what a generous colleague who knows the field would tell you over coffee.
## Agent Integrity Rules (Research-Pack Convention)
Inherited from the research-pack convention; locked verbatim per PR #657's cross-skill consistency audit.
- **Source discipline.** Only cite Consensus-returned papers from THIS session. Training knowledge labeled `[Not from Consensus — model knowledge]` and excluded from cited count. Sparse results stated explicitly, never silently filled.
- **Counting discipline.** Three numbers tracked: searches executed / unique papers received (deduplicated) / papers cited. Every cited paper has a retrievable Consensus URL from this session. Use `scripts/citation_tracker.py` for deterministic counts.
- **Tool constraints.** Consensus per-query cap depends on plan tier. **Detect at first search**, report at checkpoint. Rate limit is **1 query/sec** — sequential execution mandatory.
- **Retry policy.** On failure → wait 3s → retry once → log. After 3 consecutive failures: stop, alert user, share what was collected.
- **Plan-tier detection.** Parse first-search response for "Showing top 10" / "upgrade" → free tier (10/search). 20 returned → Pro (20/search). Calculate theoretical ceiling and surface at checkpoint so user can recalibrate.
See [`references/search_budget_allocation.md`](references/search_budget_allocation.md) for the sequential-execution rationale + plan-tier signals.
## Error Handling
| Failure | Behavior |
|---|---|
| Consensus rate-limit hit | Wait 3s, retry once, log outcome |
| Search returns 0 results | Note explicitly; "either niche terminology or genuine gap"; never silently fill |
| Plan-tier cap detected | Log tier; report at checkpoint; surface in audit |
| 3 consecutive failures | Stop searching, alert user, share what's collected, ask how to proceed |
| Sub-area returns thin results (<5 papers) | Flag in audit; suggest manual PubMed/Scholar supplementation |
| User wants to adjust sub-areas | Update table, re-confirm before searching |
| DOCX validation fails | Unpack XML, fix, repack |
## Phase 0: Grill-Me Intake (3 forcing questions, one at a time)
Each question carries explicit "why I'm asking". Stop condition: max 3 before Phase 1.
### Q1 (root) — Research question specificity
> **State the research question in 1–2 sentences. Specific is better — "How do LLMs perform on clinical reasoning tasks compared to physicians?" beats "AI in medicine". Vague questions produce vague reviews.**
>
> *Why I'm asking:* The reconnaissance search hinges on precise terminology. Vague questions produce thin recon results that don't yield a useful framework breakdown.
**Refuse mush.** Re-ask once with examples if user is too broad. If still vague, deliver with explicit "broad-scope orientation, not depth review" caveat.
### Q2 (depends on Q1) — Framework hint
> **Framework — pick one or say "you pick":**
>
> 1. **PICO** (Population / Intervention / Comparison / Outcome — most clinical questions)
> 2. **SPIDER** (Sample / Phenomenon / Design / Evaluation / Research-type — social/qualitative)
> 3. **Decomposition** (Problem / Solution / Evaluation / Limitations — technology-focused)
> 4. **Hybrid** (you pick which components from which framework)
> 5. **You pick** — analyze Q1 and recommend
>
> *Why I'm asking:* PICO is the default for ~70% of clinical questions but maps poorly to qualitative work or technology evaluation. Picking upfront saves the recon search from suggesting a misaligned framework.
Forcing choice with default ("you pick"). The skill surfaces its own framework recommendation after the recon search so user can override. Use `scripts/framework_recommender.py` for the heuristic.
See [`references/framework_selection.md`](references/framework_selection.md) for PICO / SPIDER / Decomposition canon.
### Q3 (depends on Q1) — Tentative depth
> **Tentative depth — pick one. Final confirmation comes after the framework breakdown:**
>
> 1. **Quick scan** (5 searches)
> 2. **Standard review** (10 searches)
> 3. **Deep dive** (20 searches)
>
> *Why I'm asking:* I ask this twice — once now to calibrate the recon search emphasis, once after the framework breakdown to confirm. Tentative answer affects which sub-areas to surface first; final answer drives search budget allocation.
Forcing choice. **Re-asked** at the post-Phase-2 checkpoint after the user has seen the framework breakdown.
**Stop condition:** 3 questions max before Phase 1. The post-Phase-2 checkpoint is its own grill-me moment (framework table + sub-area-adjustment + depth-reconfirmation).
## Phase 1: Initial Reconnaissance
**One broad Consensus search** to map themes, terminology, methodological distinctions.
- Query: broad version of Q1 (terminology variants are okay; first search casts wide)
- Record: `citation_tracker.py --action record_search --session NAME --query "..."`
- Record received count: `citation_tracker.py --action record_papers_received --session NAME --count N`
- **Detect plan tier** from response: "Showing top 10" / "upgrade" → free; 20 returned → Pro
Synthesize for the checkpoint:
- Themes that surfaced
- Terminology variations (e.g., "LLM" vs "large language model" vs "GPT-style model")
- Methodological distinctions (clinical trials vs benchmark eval vs case study)
- Coverage gaps (sub-questions absent from recon results)
## Phase 2: Framework Selection + Sub-area Generation
Choose framework (from Q2 OR override based on recon):
- **PICO** — most clinical questions (~70% default)
- **SPIDER** — social / qualitative
- **Decomposition** — technology focus (Problem / Solution / Evaluation / Limitations)
- **Hybrid** — explicit cross-framework mapping
Generate **4-5 sub-area questions** mapped to framework components. Each becomes a targeted Phase 3 search.
## Checkpoint (grill-me forcing-options moment)
After Phase 2, halt and present:
### 3-4 sentence recon summary
- What themes surfaced
- Terminology landscape
- Evidence landscape characterization
### Framework breakdown table
| Framework Component | How It Maps to This Topic | Proposed Sub-area to Explore |
|---|---|---|
| (Component 1) | ... | Sub-area 1 |
| (Component 2) | ... | Sub-area 2 |
| (Component 3) | ... | Sub-area 3 |
| (Component 4) | ... | Sub-area 4 |
| Cross-cutting theme | ... | Sub-area 5 |
### Depth re-confirmation (forcing choice)
Surface the **practical constraint**: detected plan tier + theoretical ceiling.
- Quick scan (5 searches × ~10 results each = ~50 papers max)
- Standard review (10 searches × ~10 = ~100 papers)
- Deep dive (20 searches × ~10 = ~200 papers)
### Sub-area forcing options
- "Looks good — proceed with these sub-areas"
- "Adjust: add sub-area on [X]"
- "Adjust: remove and replace [Y] with [Z]"
- "Restart with different framework"
### Why I'm asking (the rationale)
> A wrong framework or sub-area set wastes the search budget. This is the **last cheap moment** to correct course.
**Wait for user response before Phase 3.** Refuse to start Phase 3 without explicit user choice.
## Phase 3: Targeted Searches
Sequential (1 query/sec), budget per depth tier. See [`references/search_budget_allocation.md`](references/search_budget_allocation.md) for full canon.
### Quick scan (5 searches)
- 5 sub-area searches (one per sub-area)
- Skip era-gated + review-specific
### Standard review (10 searches)
- 5 sub-area searches
- 2 review article searches (top 2 sub-areas): `"systematic review [topic]"` / `"meta-analysis [topic]"`
- 2 era-gated searches (most important sub-area): `year_max: 2015` + `year_min: 2021`
- 1 follow-up on highest-cited paper using its key terms + `year_min` after publication
### Deep dive (20 searches)
- 5 sub-area searches
- 5 review article searches (one per sub-area)
- 4 era-gated searches (top 2 sub-areas, old + new each)
- 3 follow-ups on top 3 highest-cited papers
- 3 spare for emerging threads (surprising findings to chase)
Throughout: 1 q/sec rate limit. Sequential. Confirm response before next call. Record each via `citation_tracker.py`.
## Cross-Search Intelligence
Three trackers across ALL search results — run `scripts/cross_search_aggregator.py --session NAME` after Phase 3 completes:
1. **Repeat-hit papers** — same paper appearing in 3+ sub-area searches = likely foundational
2. **Recurring authors** — same author in multiple searches = dominant research group; top 3-5 most frequent matter
3. **Citation-per-year heuristic** — a 2023 paper with 150 citations >> 2008 paper with 150 citations. Use for seminal-work identification.
These feed the "Start Here" + "Key Research Groups" + "Bibliography" DOCX sections.
## Phase 4: DOCX Research Guide
Generate via Node.js + `docx` library. 8 sections (see [`references/docx_8_sections.md`](references/docx_8_sections.md) for full spec):
1. **Topic Overview** — single tight paragraph (4-6 sentences)
2. **Start Here — Priority Reading Order** — 5-7 papers ordered: best recent review → foundational → 2-3 frontier → gap/controversy. Each: hyperlinked title + authors/year + 1-sentence contribution + 1-sentence "what to look for"
3. **How the Field Got Here** — chronological narrative (1-2 paragraphs) + timeline table (5-8 milestones: Year / Milestone / Significance) + terminology evolution note
4. **Sub-area Guides** (one per sub-area, 4 parts each)
- 4a. What the Research Shows (2-3 sentence synthesis with inline citations)
- 4b. Key Papers (3-5 hyperlinked papers with citation count, year, 1-sentence importance)
- 4c. Key Search Terms (6-10 keywords, synonyms, MeSH, historical terms)
- 4d. Boolean Search Strings (2-3 ready-to-paste strings)
5. **Key Research Groups** — top 3-5 authors/groups with affiliations, sub-area coverage, representative paper link (from cross-search aggregator)
6. **Open Questions & Gaps** — three categories: methodological / population-context / conceptual-theoretical. Each gap explains *why it matters*.
7. **Bibliography** — alphabetical by first author. Every entry has clickable "View on Consensus" link. Every inline citation matches a bibliography entry.
8. **Audit Log** — search summary table (#, query, filters, papers returned, status), counts block, coverage notes including detected tier and theoretical ceiling
### DOCX Technical Requirements
Document the key `docx` library patterns:
- Page: US Letter, 1-inch margins
- Lists: `LevelFormat.BULLET` (never unicode bullets)
- Hyperlinks: `ExternalHyperlink` with `style: "Hyperlink"`, full URL (never truncated)
- Tables: dual widths (`columnWidths` + cell `width`), `ShadingType.CLEAR`
- Validation step after save (`python scripts/office/validate.py output.docx`)
Reference the **docx skill** for setup patterns and best practices.
## Output
```
research_guide_<topic-slug>_<YYYY-MM-DD>.docx
```
Plus:
- Chat summary block: "Saved: <path>. Audit: N searches × M unique papers / K cited. Plan tier: <tier>."
- Audit log printed inline if user asks for it
## Tooling
| Script | Role |
|---|---|
| `scripts/citation_tracker.py` | JSON-backed three-count audit at `~/.litreview_sessions/<session>.json` |
| `scripts/framework_recommender.py` | Heuristic PICO/SPIDER/Decomposition suggestion from research question |
| `scripts/cross_search_aggregator.py` | Repeat-hits + recurring-authors + citation-per-year ranking after Phase 3 |
## References
- [`references/framework_selection.md`](references/framework_selection.md) — PICO / SPIDER / Decomposition canon (7+ sources)
- [`references/search_budget_allocation.md`](references/search_budget_allocation.md) — depth tiers + cross-search intelligence + sequential execution rationale (7+ sources)
- [`references/docx_8_sections.md`](references/docx_8_sections.md) — research guide DOCX spec + technical requirements (7+ sources)
## Anti-Patterns To Reject
- Parallelizing Consensus calls
- Skipping the interactive checkpoint (running all searches without user confirmation)
- Padding thin results with training knowledge
- Defaulting to non-PICO framework without justification
- Citing papers in chat that didn't come from Consensus this session
- Hardcoding plan tier instead of detecting from first response
- Skipping era-gated searches in standard/deep budgets
- Skipping cross-search intelligence (repeat-hits, recurring authors)
- Truncating Consensus URLs in hyperlinks
---
**Version:** 1.0.0
**Source spec:** [`megaprompts/09-litreview-megaprompt.md`](../../../../megaprompts/09-litreview-megaprompt.md)
**Build pattern:** Path B (direct conversion). Sibling of `pulse` (research-pack shape).
FILE:references/docx_8_sections.md
# DOCX Research Guide — 8 Sections + Technical Requirements
This reference answers exactly one decision: **what are the 8 sections of the litreview research guide, and what does each contain to function as a "launching pad" for a researcher entering an unfamiliar field?**
## The Core Frame
The output is a **launching pad**, not a finished review. Frame each section as: "what would a generous colleague tell you over coffee if they knew the field and you didn't?"
That framing rules out:
- Exhaustive coverage (a launch pad is finite)
- Comprehensive synthesis (the user will read the papers)
- Defensible-publishable form (this is orientation, not submission-ready)
And rules in:
- Clear ordering (read these papers in this order)
- Honest gaps (here's what's underdeveloped)
- Practical entry points (here's how to keep searching)
## Section 1: Topic Overview
**Length:** 4-6 sentences, single tight paragraph.
**Contents:**
- What the field is (1 sentence)
- Why it matters (1 sentence)
- Framework used (PICO / SPIDER / Decomposition / hybrid) (1 sentence)
- Characterization of the evidence landscape (1-2 sentences)
- Honest caveat or limitation (1 sentence) — e.g., "mostly Western data" or "RCTs are scarce"
**Tone:** Confident but caveated. A colleague summarizing, not a textbook authority.
## Section 2: Start Here — Priority Reading Order
**Length:** 5-7 papers, ordered.
**Order:**
1. Best recent review (sets the field context)
2. Foundational paper(s) — 1-2, ranked by repeat-hits + cited-per-year
3. Frontier papers — 2-3 (most-recent that surfaced multiple times)
4. Gap / controversy paper — 1 (surfaces what's contested)
**Per paper:**
- Hyperlinked title (clickable to Consensus)
- Authors + year
- One sentence: contribution
- One sentence: "what to look for"
**Example entry:**
> 1. **[A systematic review of LLM clinical reasoning](https://consensus.app/...)** — Singhal et al. 2024 — Most comprehensive synthesis of LLM diagnostic performance through 2023. Look for: section on prompting strategy (the field's main tunable variable).
## Section 3: How the Field Got Here
**Length:** 1-2 paragraphs narrative + timeline table.
**Narrative:** chronological story of the field's evolution. 3-5 sentences. What changed, when, why.
**Timeline table:** 5-8 milestones.
| Year | Milestone | Significance |
|---|---|---|
| 2015 | First paper applying X to Y | Established the question |
| 2018 | Method Z introduced | Made evaluation tractable |
| 2020 | Large-scale dataset W released | Enabled benchmarking |
| 2023 | Breakthrough result by Group A | Set current state-of-the-art |
**Terminology evolution note:** "Field used 'X' through 2018; now standardly called 'Y'. Older searches must include the older term."
This section is what makes a literature review for the researcher: the linear story plus the moments of inflection. Build it from era-gated search results.
## Section 4: Sub-area Guides
**Length:** One per sub-area (4-5 total), 4 parts each.
### 4a. What the Research Shows
2-3 sentence synthesis with inline citations.
Example:
> LLMs achieve 70-85% accuracy on clinical reasoning benchmarks (Singhal et al. 2023, Liévin et al. 2024) but performance degrades sharply on novel case presentations (Toma et al. 2024). The variance across model families and prompting strategies is the field's central open question.
Every fact is hyperlinked. Every inline citation matches a bibliography entry (Section 7).
### 4b. Key Papers
3-5 hyperlinked papers. Per paper:
- Title (hyperlinked)
- Citation count + year
- One-sentence importance
### 4c. Key Search Terms
6-10 keywords for the sub-area:
- Modern preferred terms
- Synonyms (especially historical)
- MeSH headings if applicable
- Domain-specific terms (e.g., "USMLE-style" for clinical reasoning)
### 4d. Boolean Search Strings
2-3 ready-to-paste strings:
```
("clinical reasoning" OR "diagnostic reasoning") AND ("large language model" OR LLM OR GPT) AND (evaluation OR benchmark)
```
User pastes into Consensus / PubMed / Scopus to continue searching beyond what the skill ran.
## Section 5: Key Research Groups
**Length:** 3-5 groups.
**Source:** `scripts/cross_search_aggregator.py` recurring-authors output.
**Per group:**
- Lead author (or 2-3 authors if collaborative)
- Affiliation (institution)
- Sub-areas they cover (from cross-search analysis)
- Representative paper (hyperlinked, with year)
- Why they matter (1 sentence)
**Example:**
> **Singhal, K. et al. (Google DeepMind / Med-PaLM)** — Coverage: clinical reasoning, multimodal medical AI. Representative: ["Towards Generalist Biomedical AI" (2023)](https://...). Why they matter: built the Med-PaLM line; their benchmark methodology defines current state-of-the-art evaluation.
## Section 6: Open Questions & Gaps
**Length:** 3 categories, each with 1-3 gaps.
**Categories:**
1. **Methodological gaps** — what's hard to measure, what we don't have good methods for
2. **Population / context gaps** — who isn't being studied, where the data isn't
3. **Conceptual / theoretical gaps** — what we don't understand about the underlying mechanism
**Per gap:**
- One sentence stating the gap
- One sentence on *why it matters* — what's downstream of this gap being filled
Example:
> **Methodological gap:** No standardized benchmark for novel-case clinical reasoning (only retrospective USMLE-style). *Why it matters:* current "85% accuracy" claims may not generalize to real practice where novel cases dominate.
The "why it matters" sentence is what distinguishes a gap list from a complaint list.
## Section 7: Bibliography
**Length:** All cited papers, alphabetical by first author.
**Per entry:**
- Full citation (author list, title, journal, year, volume/issue, pages)
- Hyperlinked "View on Consensus" link (full URL, never truncated)
- Inline-citation key matching Section 4 references (e.g., "Singhal et al. 2024")
**Discipline:**
- Every inline citation in Sections 1-6 appears in Bibliography
- Every Bibliography entry is cited at least once
- No phantom entries (cited but no bib) or orphan entries (bib but never cited)
- Consensus URLs preserved in full (never `...` truncation)
## Section 8: Audit Log
**Length:** Search summary table + counts block + coverage notes.
**Search summary table:**
| # | Query | Filters | Results | Status |
|---|---|---|---|---|
| 1 | broad recon | none | 10 | OK |
| 2 | sub-area 1 | year_min: 2018 | 10 | OK |
| ... | ... | ... | ... | ... |
| 10 | follow-up on Singhal | year_min: 2024 | 7 | thin |
**Counts block:**
```
Searches executed: 10
Unique papers received: 47 (after deduplication)
Papers cited in this guide: 22
Plan tier detected: Free (10/search cap)
Theoretical ceiling: 100 papers; received 47 unique (typical deduplication)
```
**Coverage notes:**
- Which sub-areas surfaced thin results
- Plan-tier impact on coverage
- Suggested manual supplementation (PubMed, Scholar, etc.)
- Era-gated search yields (terminology shifts detected)
The audit log makes the entire review reproducible and falsifiable. A future reader can rerun the searches and check the work.
## DOCX Technical Requirements
Document the key `docx` library patterns (Node.js):
### Page setup
```js
const page = {
size: "LETTER",
margins: { top: 1440, right: 1440, bottom: 1440, left: 1440 }, // 1 inch in twips
};
```
### Lists (NEVER unicode bullets)
```js
new Paragraph({
children: [new TextRun(text)],
numbering: { reference: "default-bullet", level: 0 },
});
// Defined in document numbering config with LevelFormat.BULLET
```
### Hyperlinks (full URL, "Hyperlink" style)
```js
new ExternalHyperlink({
link: "https://consensus.app/full-url-never-truncated/...",
children: [new TextRun({ text: paperTitle, style: "Hyperlink" })],
});
```
### Tables (dual widths)
```js
new Table({
columnWidths: [3000, 4000, 2000], // EMU
rows: rows.map(r => new TableRow({
children: r.cells.map(c => new TableCell({
width: { size: c.width, type: WidthType.DXA },
shading: { type: ShadingType.CLEAR, color: "auto", fill: "auto" },
children: [new Paragraph(c.text)],
})),
})),
});
```
### Validation
After save:
```bash
python scripts/office/validate.py output.docx
```
If validation fails: unpack DOCX (it's a ZIP), fix the offending XML, repack.
Reference the **docx skill** (`docx/SKILL.md` in this repo if installed) for full setup patterns.
## Anti-Patterns
- **Truncating Consensus URLs in hyperlinks** — breaks reproducibility
- **Phantom bibliography entries** — cited paper missing from bib
- **Generic "Future Work" section** — Section 6 must be *specific* gaps, not "more research is needed"
- **No timeline table in Section 3** — narrative-only loses the milestone structure
- **Unicode bullets (• ‣ ▶)** instead of `LevelFormat.BULLET` — breaks DOCX list rendering in some viewers
- **Single-width tables** (only `columnWidths` or only cell `width`) — renders inconsistently across Word / LibreOffice / Google Docs
- **Skipping validation step** — invalid DOCX silently fails to open or renders broken
- **Audit log without theoretical ceiling** — user can't calibrate "is this comprehensive?"
## Operational Checklist
- [ ] All 8 sections present in DOCX
- [ ] Section 1: 4-6 sentence paragraph
- [ ] Section 2: 5-7 papers in priority order
- [ ] Section 3: narrative + timeline table + terminology note
- [ ] Section 4: one sub-section per sub-area, 4 parts each
- [ ] Section 5: 3-5 groups from cross-search aggregator
- [ ] Section 6: 3 categories with "why it matters" per gap
- [ ] Section 7: alphabetical, hyperlinked, no phantoms / orphans
- [ ] Section 8: search table + counts + tier + coverage notes
- [ ] All Consensus URLs full (no truncation)
- [ ] `LevelFormat.BULLET` for lists (no unicode bullets)
- [ ] Tables have both `columnWidths` AND cell `width`
- [ ] `python scripts/office/validate.py output.docx` PASSes
## Citations (7 sources)
1. **`docx` Node.js library — github.com/dolanmiu/docx (MIT).** Authoritative API source. The technical patterns (Paragraph, ExternalHyperlink, Table, LevelFormat.BULLET) come from its documentation.
2. **OOXML (Office Open XML) Specification — ECMA-376 (4th ed., 2016).** The underlying XML schema for DOCX. Source for the dual-width table pattern (DOCX renderers respect both column widths and cell widths; missing either causes layout inconsistencies).
3. **PRISMA 2020 Statement — Page, M. J. et al., *BMJ* 372, 2021.** Source for the audit-log section requirements (every reported search must include query, filters, results count, status). PRISMA is the international standard for systematic-review reporting.
4. **Cochrane Handbook — Higgins, J. P. T. et al. (Wiley, 2019).** Chapter 4 + Chapter 7 on data extraction and synthesis. Source for the sub-area guide structure (synthesis + key papers + search terms + boolean strings) — Cochrane's standard data-extraction template.
5. **Lipsey, M. W. & Wilson, D. B., *Practical Meta-Analysis* (Sage, 2001).** Source for the bibliography discipline (every inline citation has bib entry; every bib entry is cited). Essential for review integrity.
6. **Tufte, E., *Visual Display of Quantitative Information* (Graphics Press, 1983, 2001 ed.).** Source for the timeline-table pattern (5-8 milestones, not 20+; "milestones" not "events"). Tufte's "small multiples" + "data-ink ratio" principles inform the audit-log table design.
7. **William Strunk Jr. & E. B. White, *The Elements of Style* (Macmillan, multiple eds.).** Source for the "Open Questions & Gaps" voice discipline. Gaps must be specific and consequential, not "more research is needed" filler. Strunk's "omit needless words" applies directly: every gap statement should pass the "why it matters" test.
FILE:references/framework_selection.md
# Framework Selection — PICO, SPIDER, Decomposition, Hybrid
This reference answers exactly one decision: **which literature-review framework does litreview pick for a given research question, and how does each map sub-areas to search queries?**
Pair with `scripts/framework_recommender.py` for the deterministic heuristic.
## The Core Claim
A literature review's framework determines *what counts as a sub-area*. Pick the wrong framework → sub-areas don't map to actual research → searches return tangential papers → review is shallow.
The three primary frameworks plus hybrid:
| Framework | Best for | Components |
|---|---|---|
| **PICO** | ~70% of clinical questions; quantitative outcomes | Population / Intervention / Comparison / Outcome |
| **SPIDER** | Social / qualitative; experiential questions | Sample / Phenomenon / Design / Evaluation / Research-type |
| **Decomposition** | Technology-focused; design / engineering | Problem / Solution / Evaluation / Limitations |
| **Hybrid** | Cross-cutting topics (clinical + tech, etc.) | Pick components from multiple frameworks |
## PICO (default)
Most clinical and biomedical research questions map cleanly to PICO. Example:
> "How do LLMs perform on clinical reasoning tasks compared to physicians?"
| Component | Mapped to topic |
|---|---|
| **P**opulation | Clinical reasoning tasks (USMLE, MedQA, NEJM cases) |
| **I**ntervention | LLM-based reasoning (GPT-4, Claude, Med-PaLM) |
| **C**omparison | Physician baseline (specialists, residents, generalists) |
| **O**utcome | Diagnostic accuracy, reasoning quality, time-to-decision |
Each component becomes one or more sub-area searches.
**PICO weaknesses:**
- Maps poorly to qualitative research (no clear comparison)
- Maps poorly to technology evaluation (Population is fuzzy)
- Maps poorly to pure-theory questions (no Intervention)
When PICO doesn't fit cleanly → SPIDER or Decomposition.
## SPIDER (social / qualitative)
Designed for qualitative + mixed-methods research where PICO breaks. Example:
> "How do clinicians experience burnout in academic medicine?"
| Component | Mapped to topic |
|---|---|
| **S**ample | Clinicians in academic medical centers |
| **P**henomenon | Burnout (specifically: emotional exhaustion, depersonalization, reduced accomplishment) |
| **D**esign | Qualitative interviews, ethnography, phenomenology |
| **E**valuation | Lived experience, narrative themes |
| **R**esearch-type | Qualitative, mixed-methods |
Strong signal for SPIDER:
- Question contains "experience", "perception", "meaning", "lived"
- Outcome is hard to quantify
- Research methods involve interviews or observation
## Decomposition (technology / engineering)
Designed for design / build / evaluate questions. Example:
> "How are retrieval-augmented generation systems evaluated for clinical Q&A?"
| Component | Mapped to topic |
|---|---|
| **P**roblem | Clinical Q&A: high recall, factual accuracy, citation traceability |
| **S**olution | RAG architecture (retriever + generator combinations) |
| **E**valuation | Benchmarks (MMLU-clinical, MedMCQA, custom Q&A sets) |
| **L**imitations | Hallucination rates, latency, retrieval quality |
Strong signal for Decomposition:
- Question is about a *system* or *method*, not a population
- Question implicitly has "Problem → proposed Solution → how to test → known issues" structure
- Common in CS / ML / engineering research
## Hybrid (cross-cutting)
When no single framework fits, mix components. Example:
> "How effective is AI-assisted radiology workflow integration in community hospitals?"
| Component | Source framework | Mapping |
|---|---|---|
| Population | PICO | Community hospital radiology departments |
| Intervention | PICO | AI-assisted workflow integration (tool: vendor X) |
| Phenomenon | SPIDER | Workflow change, radiologist experience |
| Outcome | PICO | Read times, diagnostic accuracy, satisfaction |
| Limitations | Decomposition | Integration friction, false-positive rate |
Hybrid framing is more work but more accurate for questions that genuinely span disciplines.
## The Framework Recommender Heuristic
`scripts/framework_recommender.py` uses keyword signals to suggest a framework:
| Signal in research question | Suggests |
|---|---|
| "compared to", "vs", "versus", "better than" | PICO (Comparison) |
| "intervention", "treatment", "drug", "therapy" | PICO (Intervention) |
| "experience", "perception", "meaning", "narrative" | SPIDER (Phenomenon) |
| "qualitative", "interview", "ethnography" | SPIDER (Design) |
| "system", "model", "algorithm", "architecture" | Decomposition (Solution) |
| "benchmark", "evaluation", "metric" | Decomposition (Evaluation) |
| Multiple signals across frameworks | Hybrid |
| No strong signal | PICO (default) |
The recommender outputs:
- Recommended framework
- Confidence (high / medium / low)
- Rationale (which signals fired)
- 4-5 sub-area starter questions mapped to framework components
The skill then surfaces this in the post-Phase-2 checkpoint for user confirmation/override.
## When the User Says "You Pick"
Q2's "you pick" option triggers the recommender. The skill:
1. Runs Phase 1 recon search (using broad terminology from Q1)
2. After recon, runs the recommender heuristic against Q1 text
3. Surfaces in checkpoint: "I'm recommending {framework} because {rationale}. Override if you want."
User can override at checkpoint. Refusing to commit (just saying "go") → use recommender's pick.
## Anti-Patterns
### Defaulting to PICO without justification
PICO works for 70% but fails the other 30%. Defaulting to PICO for a SPIDER question wastes the search budget. The recommender prevents this; manual override should have justification.
### Hybrid for everything
Hybrid framing is more work and produces fuzzier sub-areas. Use only when a single framework genuinely fails. Default to non-hybrid; promote to hybrid only when checkpoint review surfaces real cross-cutting components.
### Forcing the framework to fit
If 3 of 5 components don't map naturally, the framework is wrong. Restart with a different framework rather than papering over the misfit.
### Picking framework before reading Q1
The recommender requires Q1 text. Asking Q2 before Q1 is answered loses signal.
### Ignoring the recommender's recommendation
If the recommender suggests SPIDER with high confidence and the user picks PICO anyway, gently challenge: "I see qualitative signals in your question. Want me to use SPIDER, or do you have a reason to insist on PICO?" Once. Honor user override after one push-back.
## Operational Checklist
- [ ] Q1 answered before Q2 (recommender needs Q1 text)
- [ ] Q2 forcing choice with "you pick" default
- [ ] `framework_recommender.py` run after Q1 (cached for checkpoint)
- [ ] Recommendation surfaced in checkpoint with rationale
- [ ] User can override at checkpoint
- [ ] Sub-areas mapped 1-to-1 with framework components
- [ ] Cross-cutting 5th sub-area added regardless of framework
## Citations (7 sources)
1. **Sackett, D. L. et al., *Evidence-Based Medicine: How to Practice and Teach EBM* (Churchill Livingstone, 1997, multiple eds.).** Origin of PICO as a clinical-question framing tool. The "PICO" acronym dates from this text. https://en.wikipedia.org/wiki/Evidence-based_medicine
2. **Cooke, A., Smith, D., & Booth, A., "Beyond PICO: The SPIDER Tool for Qualitative Evidence Synthesis" — *Qualitative Health Research* 22(10), 2012, pp. 1435-1443.** Origin of SPIDER as a PICO alternative for qualitative research. Documents the systematic failures of PICO on qualitative questions that motivated SPIDER's design.
3. **Booth, A., "Searching for qualitative research for inclusion in systematic reviews: a structured methodological review" — *Systematic Reviews* 5, 2016.** Comparative analysis of PICO vs SPIDER for qualitative work. Source for the "SPIDER for social/qualitative" guidance.
4. **PRISMA 2020 Statement — Page, M. J. et al., *BMJ* 372, 2021.** The systematic-review reporting standard. Section on "Eligibility criteria" formalizes the framework-driven approach to defining inclusion/exclusion criteria from sub-areas.
5. **Cochrane Handbook for Systematic Reviews of Interventions — Higgins, J. P. T. et al. (Wiley, 2019, online updates).** Authoritative source for PICO-driven systematic review methodology. Chapter 4 on "Searching for and selecting studies" formalizes the framework → sub-area → search-string mapping pattern.
6. **Hewitt-Taylor, J., "Use of constant comparative analysis in qualitative research" — *Nursing Standard* 15(42), 2001.** Source for the cross-cutting-theme pattern that litreview adds as a 5th sub-area regardless of framework. Constant comparative analysis surfaces themes that cross conventional framework boundaries.
7. **JBI Evidence Synthesis methodology — Joanna Briggs Institute manual (jbi.global).** Comprehensive framework comparison: PICO for quantitative effectiveness, PICo (lowercase 'o' for context) for qualitative, PEO for risk factors, CoCoPop for prevalence. The litreview skill simplifies to PICO/SPIDER/Decomposition + hybrid but the JBI manual catalogs ~12 framework variants for specialty cases.
FILE:references/search_budget_allocation.md
# Search Budget Allocation — Quick / Standard / Deep + Cross-Search Intelligence
This reference answers exactly one decision: **how does litreview spend its search budget across the 5/10/20 depth tiers, and what makes the cross-search intelligence layer add value beyond per-query results?**
Pair with `scripts/cross_search_aggregator.py` for the deterministic implementation.
## The Core Constraint
Consensus has a **1 query/second rate limit**. NEVER parallelize. Sequential execution is the only mode that doesn't break the rate limit. This is the same rule pulse uses for Reddit/HN/Web — research-pack convention.
Plus a **plan-tier cap**: free tier returns ~10 results per query; Pro returns ~20. Detected at first search response.
The combination produces hard budget ceilings:
| Tier | Plan | Theoretical max papers |
|---|---|---|
| Quick scan (5 q) | Free | 50 |
| Quick scan (5 q) | Pro | 100 |
| Standard (10 q) | Free | 100 |
| Standard (10 q) | Pro | 200 |
| Deep dive (20 q) | Free | 200 |
| Deep dive (20 q) | Pro | 400 |
These are *theoretical* — deduplication reduces the actual unique paper count by 30-50% in practice.
## Why Three Tiers (Not One Adaptive Budget)
Adaptive budgeting (run more searches if early results are thin) sounds smart but:
1. **User can't predict run time.** A 5-search budget runs in ~5s; a 20-search adaptive could run 10-30s.
2. **Sunk-cost bias kicks in.** Once 10 searches run, "let's do 5 more" is hard to resist even if results aren't worth it.
3. **Cross-search intelligence works best at fixed N.** Repeat-hit and recurring-author signals stabilize at known sample sizes.
Fixed tiers with explicit allocations beat adaptive budgets for research-orientation tasks.
## Quick Scan (5 searches)
Budget allocation:
- **5 sub-area searches** (one per sub-area from Phase 2)
- Skip era-gated searches
- Skip review-specific searches
- Skip follow-ups
Use when:
- User wants a fast orientation (~30s with 1 q/sec)
- Topic is well-known to user; they just need pointers
- Plan tier is free + topic is reasonably narrow
**Note in audit:** "Quick scan tier — review articles + era-gated comparisons omitted. Bibliography may be thin on foundational older work."
## Standard Review (10 searches)
Budget allocation:
- **5 sub-area searches** (one per sub-area)
- **2 review article searches** (top 2 sub-areas):
- `"systematic review [topic]"` AND `"meta-analysis [topic]"`
- **2 era-gated searches** (most important sub-area):
- `year_max: 2015` → reveals terminology evolution
- `year_min: 2021` → captures current frontier
- **1 follow-up** on highest-cited paper:
- Use its key terms + `year_min: <publication_year + 1>`
- Surfaces papers that built on this work
Use when (default tier):
- User has some familiarity but wants depth
- Plan tier allows reasonable coverage
- Time budget is 1-2 minutes total
## Deep Dive (20 searches)
Budget allocation:
- **5 sub-area searches**
- **5 review article searches** (one per sub-area)
- **4 era-gated searches** (top 2 sub-areas, old + new each):
- Sub-area A: `year_max: 2015` + `year_min: 2021`
- Sub-area B: `year_max: 2015` + `year_min: 2021`
- **3 follow-ups on top 3 highest-cited papers** (their terms + `year_min`)
- **3 spare for emerging threads** — surprising findings from earlier searches worth chasing
Use when:
- Topic is genuinely new to user
- Comprehensive orientation is the goal
- Plan tier is Pro (free tier deep-dive is bottlenecked at ~200 papers)
## Cross-Search Intelligence
Three trackers across ALL Phase 3 search results. Run after Phase 3 completes via `scripts/cross_search_aggregator.py --session NAME`.
### Tracker 1: Repeat-Hit Papers (foundational signal)
A paper appearing in **3+ sub-area searches** is signal that it's foundational — multiple sub-fields cite it, suggesting cross-cutting importance.
Use repeat-hits to populate "Start Here" DOCX section:
- Repeat-hit + high citation → priority foundational paper
- Repeat-hit + recent → likely emerging classic
- Repeat-hit but few citations → niche but cross-cutting
### Tracker 2: Recurring Authors (dominant research group signal)
Same author appearing across **multiple sub-area searches** = research group dominant in this area.
Top 3-5 most-frequent authors → "Key Research Groups" DOCX section.
Pattern:
- 5+ search appearances → dominant group (cite representative paper)
- 3-4 appearances → significant but not dominant
- 1-2 appearances → not a "group" signal; may still be high-impact individual
Note: a single highly-cited paper isn't a "group" signal — the recurrence across multiple sub-areas matters.
### Tracker 3: Citation-Per-Year (seminal-work heuristic)
Raw citation count is biased toward older papers (more time to accumulate citations). Citations-per-year normalizes:
- Paper A: 2008, 150 citations → 9.4 cites/year
- Paper B: 2023, 150 citations → 50 cites/year
Paper B is much more seminal in current discourse despite equal absolute citation count.
Citation-per-year ranking → "Start Here" priority ordering.
## Why Cross-Search Intelligence Matters
Per-query results show "papers about this sub-area". Cross-search intelligence shows "patterns across the whole field":
- Repeat-hits reveal foundational structure
- Recurring authors reveal who's doing the work
- Citation-per-year reveals what's currently shaping discourse
A literature review WITHOUT cross-search intelligence is just a list of papers. WITH it, the review surfaces the *structure* of the field.
## Sequential Execution Discipline
Each Consensus call must wait for the prior response. NEVER parallelize:
```
search_1 → wait response → record → 1 second pause → search_2 → ...
```
If parallel: rate limit triggers 429, error counter increments, after 3 consecutive failures → stop.
`scripts/citation_tracker.py --action record_search` enforces the timestamp gap (rejects calls within 1s of prior).
## Plan-Tier Detection
After search 1, parse the response:
| Signal | Tier |
|---|---|
| "Showing top 10" / "upgrade for more" | Free (10/search cap) |
| 20 papers returned | Pro (20/search cap) |
| Auth-failure response | API key missing or invalid |
Surface tier at checkpoint:
> Detected free tier (~10 results per search). Calibrating budget:
> Quick scan: 5 × 10 = ~50 papers
> Standard: 10 × 10 = ~100 papers
> Deep dive: 20 × 10 = ~200 papers
> If you want deeper coverage, Consensus Pro unlocks 20/search.
User chooses depth after seeing the constraint.
## Anti-Patterns
- **Parallelizing searches** — triggers rate limit; data loss
- **Adaptive "just one more" extensions** — bias-prone; commit to tier upfront
- **Skipping era-gated searches in standard/deep tiers** — misses terminology shifts
- **Skipping cross-search aggregation** — reduces review to a paper list
- **Hardcoding plan tier** — detect at runtime; don't assume free/Pro
- **Reporting raw citation count without per-year** — over-weights older papers
- **Counting repeat-hits at threshold 2** — too noisy; 3 is the minimum signal
## Operational Checklist
- [ ] Plan tier detected from search 1 response
- [ ] Theoretical ceiling reported at checkpoint
- [ ] Search budget allocated per tier (5/10/20)
- [ ] Era-gated searches included in standard/deep
- [ ] Follow-ups on highest-cited papers included
- [ ] 1 second wait between each Consensus call (timestamp-enforced)
- [ ] All search results passed through `cross_search_aggregator.py` after Phase 3
- [ ] Repeat-hit threshold = 3 sub-areas (not 2)
- [ ] Citation-per-year computed (not raw citation count)
## Citations (7 sources)
1. **Consensus.app documentation — consensus.app/help.** Authoritative source for plan-tier caps (free: 10/search, Pro: 20/search) and 1 q/sec rate limit. The skill detects from response rather than hardcoding because documented values evolve.
2. **Higgins, J. P. T. & Green, S. (eds.), *Cochrane Handbook for Systematic Reviews of Interventions* (Wiley, 2019).** Chapter 4 on search strategy. Source for the era-gated + review-specific + follow-up search categories. The 5/10/20 tier structure is litreview's compression of Cochrane's exhaustive-search methodology.
3. **Greenhalgh, T. & Peacock, R., "Effectiveness and efficiency of search methods in systematic reviews" — *BMJ* 331, 2005, pp. 1064-1065.** Empirical analysis of how many searches are "enough" to surface foundational papers. Source for the diminishing-returns curve that justifies fixed-tier budgets vs adaptive.
4. **Page, M. J. et al., *PRISMA 2020 Statement* — *BMJ* 372, 2021.** Reporting standard for search audit logs. Source for the audit-log DOCX section's required content (search #, query, filters, results returned).
5. **Sandelowski, M. & Barroso, J., *Handbook for Synthesizing Qualitative Research* (Springer, 2007).** Source for cross-search intelligence patterns in qualitative reviews — repeat-hits and recurring-authors are documented signals in narrative synthesis literature.
6. **Lawani, S. M., "Bibliometrics: Its theoretical foundations, methods and applications" — *Libri* 31, 1981.** Foundational bibliometrics paper. Source for the citations-per-year normalization (Lawani's Garfield-style impact normalization). The skill's citation-per-year heuristic is the simplest form of bibliometric normalization.
7. **AWS Architecture Blog — Mike Cohen, "Exponential Backoff and Jitter" (2015) + Marc Brooker, "Timeouts, retries, and backoff with jitter" (Builders' Library, 2019).** Source for the retry-once-after-3s pattern (research-pack convention). Justifies aggressive failure-detection (3 consecutive → stop) over deep retry loops for research workflows.
FILE:scripts/citation_tracker.py
#!/usr/bin/env python3
"""citation_tracker.py — JSON-backed three-count audit for litreview runs.
Stdlib-only. Mirrors pulse's citation_tracker.py (research-pack convention)
but adapted for Consensus-based academic search:
- searches executed (Consensus queries issued)
- unique papers received (deduplicated across all searches)
- papers cited (made it into the DOCX guide)
Enforces sequential discipline by rejecting record_search calls within 1
second of the prior (Consensus rate limit).
Session state persists in ~/.litreview_sessions/<session>.json.
Actions:
start Create a new session
record_search Record a search query + enforce 1s gap
record_papers_received Record N papers from this search (with dedup intent)
record_cited Record a paper URL that made it into the DOCX
status Show current counts + audit block
list List all sessions
close Mark session ended
Usage:
python citation_tracker.py --action start --session litreview-20260515 --topic "LLM clinical reasoning"
python citation_tracker.py --action record_search --session ... --query "..." --tier free
python citation_tracker.py --action record_papers_received --session ... --count 10 --unique 8
python citation_tracker.py --action record_cited --session ... --url "https://consensus.app/..."
python citation_tracker.py --action status --session ...
python citation_tracker.py --action list
python citation_tracker.py --action close --session ...
"""
import argparse
import json
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional
SESSIONS_DIR = Path.home() / ".litreview_sessions"
MIN_SEARCH_GAP_SECONDS = 1.0 # Consensus rate limit
def session_path(name: str) -> Path:
return SESSIONS_DIR / f"{name}.json"
def load_session(name: str) -> Dict[str, Any]:
p = session_path(name)
if not p.exists():
raise FileNotFoundError(f"Session not found: {name}")
return json.loads(p.read_text(encoding="utf-8"))
def save_session(name: str, data: Dict[str, Any]) -> None:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
session_path(name).write_text(json.dumps(data, indent=2), encoding="utf-8")
def now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
def now_ts() -> float:
return datetime.now(timezone.utc).timestamp()
def action_start(name: str, topic: Optional[str]) -> Dict[str, Any]:
if session_path(name).exists():
raise FileExistsError(f"Session already exists: {name}")
data: Dict[str, Any] = {
"session": name,
"topic": topic or "",
"started_at": now_iso(),
"ended_at": None,
"plan_tier": None,
"searches": [],
"papers_received_log": [],
"papers_cited": [],
"counts": {"searches": 0, "papers_received_unique": 0, "papers_cited": 0},
}
save_session(name, data)
return data
def action_record_search(name: str, query: str, tier: Optional[str]) -> Dict[str, Any]:
data = load_session(name)
if data["searches"]:
last_ts = data["searches"][-1].get("ts", 0)
gap = now_ts() - last_ts
if gap < MIN_SEARCH_GAP_SECONDS:
raise RuntimeError(
f"Sequential discipline violation: search submitted {gap:.2f}s after prior "
f"(min gap: {MIN_SEARCH_GAP_SECONDS}s). Wait at least {MIN_SEARCH_GAP_SECONDS - gap:.2f}s more."
)
if tier and not data["plan_tier"]:
data["plan_tier"] = tier
data["searches"].append({"query": query, "tier": tier, "at": now_iso(), "ts": now_ts()})
data["counts"]["searches"] += 1
save_session(name, data)
return data
def action_record_papers_received(name: str, count: int, unique: Optional[int]) -> Dict[str, Any]:
data = load_session(name)
unique_count = unique if unique is not None else count
data["papers_received_log"].append({"raw_count": count, "unique_after_dedup": unique_count, "at": now_iso()})
data["counts"]["papers_received_unique"] += unique_count
save_session(name, data)
return data
def action_record_cited(name: str, url: str, paper_title: Optional[str]) -> Dict[str, Any]:
data = load_session(name)
if any(p["url"] == url for p in data["papers_cited"]):
return data # Already cited; idempotent
data["papers_cited"].append({"url": url, "title": paper_title, "at": now_iso()})
data["counts"]["papers_cited"] += 1
save_session(name, data)
return data
def action_status(name: str) -> Dict[str, Any]:
return load_session(name)
def action_close(name: str) -> Dict[str, Any]:
data = load_session(name)
if data.get("ended_at") is None:
data["ended_at"] = now_iso()
save_session(name, data)
return data
def action_list() -> List[Dict[str, Any]]:
SESSIONS_DIR.mkdir(parents=True, exist_ok=True)
out: List[Dict[str, Any]] = []
for p in sorted(SESSIONS_DIR.glob("*.json")):
try:
d = json.loads(p.read_text(encoding="utf-8"))
out.append({
"session": d.get("session", p.stem),
"topic": d.get("topic", ""),
"started_at": d.get("started_at", ""),
"ended_at": d.get("ended_at"),
"plan_tier": d.get("plan_tier"),
"counts": d.get("counts", {}),
})
except (OSError, json.JSONDecodeError):
continue
return out
def render_status_human(data: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Session: {data['session']}")
out.append(f"Topic: {data.get('topic', '(unset)')}")
out.append(f"Plan tier: {data.get('plan_tier') or '(not detected)'}")
out.append(f"Started: {data['started_at']}")
out.append(f"Ended: {data.get('ended_at') or '(active)'}")
out.append("")
c = data["counts"]
out.append("Three-count audit:")
out.append(f" Searches: {c['searches']}")
out.append(f" Unique papers: {c['papers_received_unique']}")
out.append(f" Cited: {c['papers_cited']}")
out.append("")
out.append("Audit block (paste in DOCX Section 8):")
out.append(
f" Searches executed: {c['searches']}. "
f"Unique papers received: {c['papers_received_unique']}. "
f"Papers cited in guide: {c['papers_cited']}. "
f"Plan tier: {data.get('plan_tier') or 'undetected'}."
)
return "\n".join(out)
def render_list_human(rows: List[Dict[str, Any]]) -> str:
if not rows:
return "(no sessions)"
out: List[str] = []
out.append(f"{'session':<40s} {'tier':<6s} {'srch':>4s} {'uniq':>4s} {'cited':>5s} status")
out.append("-" * 78)
for r in rows:
c = r["counts"]
status = "closed" if r["ended_at"] else "active"
tier = r.get("plan_tier") or "—"
out.append(
f"{r['session']:<40s} {tier:<6s} "
f"{c.get('searches', 0):>4d} {c.get('papers_received_unique', 0):>4d} "
f"{c.get('papers_cited', 0):>5d} {status}"
)
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument(
"--action",
required=True,
choices=["start", "record_search", "record_papers_received", "record_cited", "status", "list", "close"],
)
parser.add_argument("--session", help="Session name")
parser.add_argument("--topic", help="(start only) topic string")
parser.add_argument("--query", help="(record_search only) Consensus query text")
parser.add_argument("--tier", help="(record_search only) detected tier: free | pro")
parser.add_argument("--count", type=int, help="(record_papers_received only) raw paper count")
parser.add_argument("--unique", type=int, help="(record_papers_received only) unique count after dedup")
parser.add_argument("--url", help="(record_cited only) Consensus URL of cited paper")
parser.add_argument("--title", help="(record_cited only) paper title for the log")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
try:
if args.action == "start":
if not args.session:
print("error: --session required for start", file=sys.stderr); return 2
result = action_start(args.session, args.topic)
elif args.action == "record_search":
if not (args.session and args.query):
print("error: --session, --query required", file=sys.stderr); return 2
result = action_record_search(args.session, args.query, args.tier)
elif args.action == "record_papers_received":
if not (args.session and args.count is not None):
print("error: --session, --count required", file=sys.stderr); return 2
result = action_record_papers_received(args.session, args.count, args.unique)
elif args.action == "record_cited":
if not (args.session and args.url):
print("error: --session, --url required", file=sys.stderr); return 2
result = action_record_cited(args.session, args.url, args.title)
elif args.action == "status":
if not args.session:
print("error: --session required for status", file=sys.stderr); return 2
result = action_status(args.session)
elif args.action == "close":
if not args.session:
print("error: --session required for close", file=sys.stderr); return 2
result = action_close(args.session)
else:
result = action_list()
except (FileNotFoundError, FileExistsError, RuntimeError) as e:
print(f"error: {e}", file=sys.stderr); return 2
if args.output == "json":
print(json.dumps(result, indent=2, default=str))
else:
if args.action == "list":
print(render_list_human(result))
else:
print(render_status_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/cross_search_aggregator.py
#!/usr/bin/env python3
"""cross_search_aggregator.py — Cross-search intelligence for litreview.
Stdlib-only. Reads all search results recorded across a litreview session
and computes three signals that transform a per-search paper list into
field-level intelligence:
1. Repeat-hit papers: same paper in 3+ sub-area searches (foundational signal)
2. Recurring authors: same author across multiple searches (dominant group)
3. Citation-per-year: normalizes raw citation count by paper age (seminal work)
Reads from a search-results JSON file (one entry per search, each with
papers list including url, title, authors, year, citations).
Outputs feed the DOCX guide's "Start Here" + "Key Research Groups"
sections.
NO LLM CALLS. Pure aggregation + ranking.
Input file format (`--results-file`):
{
"session": "litreview-20260515",
"searches": [
{
"query": "...",
"sub_area": "Intervention",
"papers": [
{"url": "https://...", "title": "...", "authors": ["..."], "year": 2023, "citations": 150}
]
}
]
}
Usage:
python cross_search_aggregator.py --results-file /tmp/results.json
python cross_search_aggregator.py --results-file /tmp/results.json --output json
python cross_search_aggregator.py --sample
"""
import argparse
import json
import sys
from collections import Counter
from datetime import datetime
from pathlib import Path
from typing import Any, Dict, List
REPEAT_HIT_THRESHOLD = 3 # paper must appear in 3+ sub-areas
TOP_AUTHORS_N = 5
TOP_REPEAT_HITS_N = 8
SAMPLE_RESULTS = {
"session": "litreview-sample",
"searches": [
{
"query": "LLM clinical reasoning benchmarks",
"sub_area": "Intervention",
"papers": [
{"url": "https://consensus.app/paper/abc1", "title": "Med-PaLM benchmark", "authors": ["Singhal", "Tu", "Gottweis"], "year": 2023, "citations": 250},
{"url": "https://consensus.app/paper/abc2", "title": "LLMs vs physicians on USMLE", "authors": ["Kung", "Cheatham"], "year": 2023, "citations": 800},
{"url": "https://consensus.app/paper/abc3", "title": "Reasoning evaluation framework", "authors": ["Lievin"], "year": 2024, "citations": 120},
],
},
{
"query": "clinical reasoning evaluation methodology",
"sub_area": "Outcome",
"papers": [
{"url": "https://consensus.app/paper/abc1", "title": "Med-PaLM benchmark", "authors": ["Singhal", "Tu", "Gottweis"], "year": 2023, "citations": 250},
{"url": "https://consensus.app/paper/abc4", "title": "Diagnostic accuracy AI", "authors": ["Toma", "Lawler"], "year": 2024, "citations": 90},
{"url": "https://consensus.app/paper/abc5", "title": "AI in medicine review", "authors": ["Singhal", "Azizi"], "year": 2023, "citations": 200},
],
},
{
"query": "GPT-4 medical Q&A",
"sub_area": "Population",
"papers": [
{"url": "https://consensus.app/paper/abc1", "title": "Med-PaLM benchmark", "authors": ["Singhal", "Tu", "Gottweis"], "year": 2023, "citations": 250},
{"url": "https://consensus.app/paper/abc2", "title": "LLMs vs physicians on USMLE", "authors": ["Kung", "Cheatham"], "year": 2023, "citations": 800},
{"url": "https://consensus.app/paper/abc6", "title": "GPT-4 USMLE performance", "authors": ["Nori", "King"], "year": 2023, "citations": 400},
],
},
],
}
def aggregate(results: Dict[str, Any]) -> Dict[str, Any]:
paper_appearances: Dict[str, Dict[str, Any]] = {}
author_appearances: Counter = Counter()
author_paper_sub_areas: Dict[str, set] = {}
for search in results.get("searches", []):
sub_area = search.get("sub_area", "uncategorized")
for paper in search.get("papers", []):
url = paper.get("url", "")
if not url:
continue
if url not in paper_appearances:
paper_appearances[url] = {
"url": url,
"title": paper.get("title", ""),
"authors": paper.get("authors", []),
"year": paper.get("year"),
"citations": paper.get("citations", 0),
"sub_areas": set(),
}
paper_appearances[url]["sub_areas"].add(sub_area)
for author in paper.get("authors", []):
author_appearances[author] += 1
if author not in author_paper_sub_areas:
author_paper_sub_areas[author] = set()
author_paper_sub_areas[author].add(sub_area)
# Tracker 1: Repeat-hit papers
repeat_hits: List[Dict[str, Any]] = []
for url, p in paper_appearances.items():
if len(p["sub_areas"]) >= REPEAT_HIT_THRESHOLD:
entry = {
"url": p["url"],
"title": p["title"],
"authors": p["authors"],
"year": p["year"],
"citations": p["citations"],
"sub_areas": sorted(p["sub_areas"]),
"sub_area_count": len(p["sub_areas"]),
}
repeat_hits.append(entry)
repeat_hits.sort(key=lambda x: (-x["sub_area_count"], -(x["citations"] or 0)))
# Tracker 2: Recurring authors
recurring_authors: List[Dict[str, Any]] = []
for author, count in author_appearances.most_common(TOP_AUTHORS_N):
if count >= 2:
recurring_authors.append({
"author": author,
"appearances": count,
"sub_areas": sorted(author_paper_sub_areas.get(author, set())),
})
# Tracker 3: Citation-per-year
current_year = datetime.now().year
cited_per_year: List[Dict[str, Any]] = []
for url, p in paper_appearances.items():
year = p.get("year")
cites = p.get("citations", 0) or 0
if year and year <= current_year and cites > 0:
age = max(current_year - year, 1)
cpy = cites / age
cited_per_year.append({
"url": p["url"],
"title": p["title"],
"year": year,
"citations": cites,
"age_years": age,
"citations_per_year": round(cpy, 1),
})
cited_per_year.sort(key=lambda x: -x["citations_per_year"])
return {
"session": results.get("session", "(unknown)"),
"total_searches": len(results.get("searches", [])),
"unique_papers": len(paper_appearances),
"repeat_hit_papers": repeat_hits[:TOP_REPEAT_HITS_N],
"repeat_hit_count": len(repeat_hits),
"recurring_authors": recurring_authors,
"citations_per_year_top_5": cited_per_year[:5],
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Cross-search intelligence — session {result['session']}")
out.append(f" Total searches: {result['total_searches']}")
out.append(f" Unique papers: {result['unique_papers']}")
out.append(f" Repeat-hit papers (≥{REPEAT_HIT_THRESHOLD} sub-areas): {result['repeat_hit_count']}")
out.append("")
if result["repeat_hit_papers"]:
out.append("Repeat-Hit Papers (foundational signal):")
for p in result["repeat_hit_papers"]:
authors_str = ", ".join(p["authors"][:3]) + (" et al." if len(p["authors"]) > 3 else "")
out.append(f" - {p['title']} ({authors_str}, {p['year']}) — {p['sub_area_count']} sub-areas, {p['citations']} cites")
out.append(f" Sub-areas: {', '.join(p['sub_areas'])}")
out.append(f" URL: {p['url']}")
else:
out.append("Repeat-Hit Papers: (none — increase search budget or check sub-area diversity)")
out.append("")
if result["recurring_authors"]:
out.append(f"Recurring Authors (top {len(result['recurring_authors'])}):")
for a in result["recurring_authors"]:
out.append(f" - {a['author']}: {a['appearances']} appearances across {len(a['sub_areas'])} sub-area(s)")
out.append(f" Sub-areas: {', '.join(a['sub_areas'])}")
else:
out.append("Recurring Authors: (none above threshold)")
out.append("")
if result["citations_per_year_top_5"]:
out.append("Citations-per-Year top 5 (seminal-work heuristic):")
for p in result["citations_per_year_top_5"]:
out.append(f" - {p['title']} ({p['year']}) — {p['citations']} cites / {p['age_years']} yr = {p['citations_per_year']}/yr")
else:
out.append("Citations-per-Year: (insufficient data)")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--results-file", help="Path to search-results JSON file")
parser.add_argument("--sample", action="store_true", help="Run on embedded sample results")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = aggregate(SAMPLE_RESULTS)
elif args.results_file:
p = Path(args.results_file)
if not p.exists():
print(f"error: {args.results_file} not found", file=sys.stderr); return 2
try:
data = json.loads(p.read_text(encoding="utf-8"))
except json.JSONDecodeError as e:
print(f"error: invalid JSON in {args.results_file}: {e}", file=sys.stderr); return 2
result = aggregate(data)
else:
parser.print_help(); return 0
if args.output == "json":
print(json.dumps(result, indent=2, default=str))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/framework_recommender.py
#!/usr/bin/env python3
"""framework_recommender.py — Heuristic PICO/SPIDER/Decomposition picker.
Stdlib-only. Given a research question, suggests which literature-review
framework to use, with confidence + rationale + starter sub-area questions.
Heuristic keyword signals:
- "compared to", "vs", "versus", "better than" → PICO (Comparison signal)
- "intervention", "treatment", "drug", "therapy" → PICO (Intervention)
- "experience", "perception", "lived", "meaning" → SPIDER (Phenomenon)
- "qualitative", "interview", "ethnography" → SPIDER (Design)
- "system", "model", "algorithm", "architecture" → Decomposition (Solution)
- "benchmark", "evaluation", "metric" → Decomposition (Evaluation)
- Multiple signals across frameworks → Hybrid
- No strong signal → PICO (default)
NO LLM CALLS. Pure regex + keyword counting.
Usage:
python framework_recommender.py --question "How do LLMs perform on clinical reasoning compared to physicians?"
python framework_recommender.py --question "..." --output json
python framework_recommender.py --sample
"""
import argparse
import json
import re
import sys
from typing import Any, Dict, List
PICO_SIGNALS = {
"comparison": ["compared to", "vs", "versus", "better than", "compared with", "relative to"],
"intervention": ["intervention", "treatment", "drug", "therapy", "drug therapy", "regimen"],
"outcome": ["outcome", "efficacy", "effectiveness", "accuracy", "mortality", "survival"],
"population": ["patients", "subjects", "cohort", "participants"],
}
SPIDER_SIGNALS = {
"phenomenon": ["experience", "perception", "meaning", "lived", "narrative", "perspective"],
"design": ["qualitative", "interview", "ethnography", "phenomenology", "grounded theory"],
"sample": ["women's", "men's", "clinicians", "students", "patients with"], # demographic-context
"evaluation": ["thematic", "narrative analysis", "lived experience"],
}
DECOMPOSITION_SIGNALS = {
"solution": ["system", "model", "algorithm", "architecture", "method", "approach", "framework"],
"evaluation": ["benchmark", "evaluation", "metric", "performance", "accuracy"],
"problem": ["challenge", "problem", "issue with", "limitations of"],
"limitations": ["limitations", "failure mode", "edge case", "robustness"],
}
def count_signals(text: str, signal_map: Dict[str, List[str]]) -> Dict[str, int]:
text_lower = text.lower()
counts: Dict[str, int] = {}
for component, phrases in signal_map.items():
component_count = 0
for phrase in phrases:
# Allow optional plural 's' / 'ed' / 'ing' suffix for single-word phrases (not multi-word)
if " " in phrase:
pattern = re.compile(rf"\b{re.escape(phrase)}\b", re.IGNORECASE)
else:
pattern = re.compile(rf"\b{re.escape(phrase)}(?:s|es|ed|ing)?\b", re.IGNORECASE)
component_count += len(pattern.findall(text_lower))
counts[component] = component_count
return counts
def recommend(question: str) -> Dict[str, Any]:
pico = count_signals(question, PICO_SIGNALS)
spider = count_signals(question, SPIDER_SIGNALS)
decomp = count_signals(question, DECOMPOSITION_SIGNALS)
pico_total = sum(pico.values())
spider_total = sum(spider.values())
decomp_total = sum(decomp.values())
total = pico_total + spider_total + decomp_total
# Confidence: ratio of dominant framework to total
if total == 0:
framework = "PICO"
confidence = "low"
rationale = "No strong framework signals detected — defaulting to PICO (covers ~70% of questions)"
elif pico_total >= 2 and spider_total >= 2:
framework = "Hybrid (PICO + SPIDER)"
confidence = "medium"
rationale = f"Both PICO ({pico_total} signals) and SPIDER ({spider_total}) detected — question spans quantitative + qualitative"
elif pico_total >= 2 and decomp_total >= 2:
framework = "Hybrid (PICO + Decomposition)"
confidence = "medium"
rationale = f"Both PICO ({pico_total}) and Decomposition ({decomp_total}) — clinical + technology evaluation"
elif decomp_total > pico_total and decomp_total > spider_total:
framework = "Decomposition"
confidence = "high" if decomp_total >= 3 else "medium"
active = [k for k, v in decomp.items() if v > 0]
rationale = f"Decomposition signals dominate ({decomp_total} total, components: {', '.join(active)})"
elif spider_total > pico_total and spider_total > decomp_total:
framework = "SPIDER"
confidence = "high" if spider_total >= 3 else "medium"
active = [k for k, v in spider.items() if v > 0]
rationale = f"SPIDER signals dominate ({spider_total} total, components: {', '.join(active)})"
else:
framework = "PICO"
confidence = "high" if pico_total >= 3 else "medium" if pico_total >= 1 else "low"
active = [k for k, v in pico.items() if v > 0]
rationale = f"PICO signals dominate ({pico_total} total, components: {', '.join(active) if active else 'default'})"
# Sub-area starter questions (template — actual generation needs LLM context)
starter_questions = generate_starter_questions(question, framework)
return {
"question": question,
"framework": framework,
"confidence": confidence,
"rationale": rationale,
"signal_counts": {"PICO": pico, "SPIDER": spider, "Decomposition": decomp},
"starter_sub_areas": starter_questions,
}
def generate_starter_questions(question: str, framework: str) -> List[str]:
"""Template-driven sub-area starter questions per framework."""
if framework.startswith("PICO") or "PICO" in framework:
return [
"Population: who is being studied? (define inclusion + exclusion)",
"Intervention: what is being tested? (specify dose / variant / version)",
"Comparison: against what baseline? (placebo / standard / alternative)",
"Outcome: what is being measured? (primary + secondary endpoints)",
"Cross-cutting: methodological quality or population variation",
]
elif framework.startswith("SPIDER") or "SPIDER" in framework:
return [
"Sample: who has the experience? (define context)",
"Phenomenon: what experience or perception? (be specific)",
"Design: what qualitative methods? (interviews / observation / artifacts)",
"Evaluation: what kind of analysis? (thematic / narrative / phenomenological)",
"Cross-cutting: cultural or temporal variation in the phenomenon",
]
elif framework.startswith("Decomposition"):
return [
"Problem: what challenge is being addressed? (constraints + objectives)",
"Solution: what is the proposed approach? (architecture + key innovation)",
"Evaluation: how is it being measured? (benchmarks + metrics + baselines)",
"Limitations: where does it fail? (edge cases + failure modes)",
"Cross-cutting: scalability or deployment considerations",
]
else: # Hybrid
return [
"Primary framework components (from dominant signals)",
"Secondary framework components (from cross-cutting signals)",
"Comparison or evaluation dimension",
"Outcome or impact dimension",
"Cross-cutting: methodological consistency across paradigms",
]
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"Question: {result['question']}")
out.append("")
out.append(f"Recommended: {result['framework']}")
out.append(f"Confidence: {result['confidence']}")
out.append(f"Rationale: {result['rationale']}")
out.append("")
out.append("Signal counts:")
for fw, components in result["signal_counts"].items():
total = sum(components.values())
active = ", ".join(f"{k}={v}" for k, v in components.items() if v > 0) or "(none)"
out.append(f" {fw:<18s} total={total} ({active})")
out.append("")
out.append("Starter sub-area questions:")
for q in result["starter_sub_areas"]:
out.append(f" - {q}")
return "\n".join(out)
SAMPLE_QUESTION = "How do large language models perform on clinical reasoning tasks compared to physicians?"
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--question", help="Research question text")
parser.add_argument("--sample", action="store_true", help="Run on embedded sample question")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = recommend(SAMPLE_QUESTION)
elif args.question:
result = recommend(args.question)
else:
parser.print_help(); return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Chất vấn kế hoạch dựa trên thuật ngữ dự án (CONTEXT.md) và các quyết định đã ghi (docs/adr/), cập nhật các tệp này khi chốt thuật ngữ.
---
name: grill-with-docs
description: Docs-anchored grilling session — challenges a plan against the project's existing language (CONTEXT.md) and recorded decisions (docs/adr/), and updates those files inline as terminology and decisions crystallise. Use when user wants to stress-test a plan against documented domain language, or mentions "grill with docs".
license: MIT
metadata:
derived_from: "https://github.com/mattpocock/skills/tree/main/skills/engineering/grill-with-docs"
original_author: "Matt Pocock (@mattpocock)"
original_license: MIT
voice: "Matt Pocock — relentless, one-at-a-time, codebase-and-docs-first, ADRs only when 3 criteria are met"
version: 1.0.0
---
# Grill with Docs
> Derived from [Matt Pocock's grill-with-docs](https://github.com/mattpocock/skills/tree/main/skills/engineering/grill-with-docs) (MIT, © 2026 Matt Pocock). Matt's interview discipline + docs-anchored grilling rules preserved verbatim under MIT. Additions in this repo: 3 stdlib validators (CONTEXT.md linter, ADR scanner, glossary↔code consistency check), 3 in-depth references each citing 7+ authoritative sources, `cs-grill-with-docs` agent, `/cs:grill-with-docs` command. See [Wrapper additions](#wrapper-additions) below.
<what-to-do>
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question before continuing.
If a question can be answered by exploring the codebase, explore the codebase instead.
</what-to-do>
<supporting-info>
## Domain awareness
During codebase exploration, also look for existing documentation:
### File structure
Most repos have a single context:
```
/
├── CONTEXT.md
├── docs/
│ └── adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
```
If a `CONTEXT-MAP.md` exists at the root, the repo has multiple contexts. The map points to where each one lives:
```
/
├── CONTEXT-MAP.md
├── docs/
│ └── adr/ ← system-wide decisions
├── src/
│ ├── ordering/
│ │ ├── CONTEXT.md
│ │ └── docs/adr/ ← context-specific decisions
│ └── billing/
│ ├── CONTEXT.md
│ └── docs/adr/
```
Create files lazily — only when you have something to write. If no `CONTEXT.md` exists, create one when the first term is resolved. If no `docs/adr/` exists, create it when the first ADR is needed.
## During the session
### Challenge against the glossary
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y — which is it?"
### Sharpen fuzzy language
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account' — do you mean the Customer or the User? Those are different things."
### Discuss concrete scenarios
When domain relationships are being discussed, stress-test them with specific scenarios. Invent scenarios that probe edge cases and force the user to be precise about the boundaries between concepts.
### Cross-reference with code
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible — which is right?"
### Update CONTEXT.md inline
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up — capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
`CONTEXT.md` should be totally devoid of implementation details. Do not treat `CONTEXT.md` as a spec, a scratch pad, or a repository for implementation decisions. It is a glossary and nothing else.
### Offer ADRs sparingly
Only offer to create an ADR when all three are true:
1. **Hard to reverse** — the cost of changing your mind later is meaningful
2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
</supporting-info>
## Wrapper Additions
The additions below are **not** part of Matt's upstream skill. They operationalize the upstream's rules into deterministic, stdlib-only validators that pair naturally with the interview loop.
### Workflow (with wrapper tools)
1. **Pre-flight (before the first question):**
- Run `scripts/context_md_linter.py CONTEXT.md` if a `CONTEXT.md` exists — confirms the glossary is well-formed before grilling against it.
- Run `scripts/adr_scanner.py docs/adr/` if `docs/adr/` exists — surfaces numbering gaps, malformed ADRs, status-frontmatter inconsistencies.
- Run `scripts/glossary_code_consistency.py --context CONTEXT.md --code src/` — flags defined-but-unused terms (dead glossary) and code-only common nouns that may need definitions. Use these flags as opening grill questions.
2. **During the session (Matt's rules apply):**
- One question per turn, walking depth-first.
- When a term is sharpened: edit `CONTEXT.md` immediately; re-run `context_md_linter.py` if the edit is structural.
- When an ADR is warranted: write it under `docs/adr/`; re-run `adr_scanner.py` to confirm numbering.
3. **Closing:**
- Final `glossary_code_consistency.py` run to confirm no new orphan terms were introduced.
- Summarize: terms added/refined, ADRs written, scenarios discussed, open items.
### Tools (stdlib-only)
| Tool | One-line role |
|---|---|
| `scripts/context_md_linter.py` | Validate `CONTEXT.md` against the CONTEXT-FORMAT.md structure. PASS/WARN/FAIL per rule. |
| `scripts/adr_scanner.py` | Walk `docs/adr/`, check `NNNN-slug.md` pattern, numbering integrity, body completeness. |
| `scripts/glossary_code_consistency.py` | Cross-reference bold terms in `CONTEXT.md` against codebase usage. Flag dead glossary + code-only common nouns. |
### References (citations behind each rule)
- [`references/ubiquitous_language.md`](references/ubiquitous_language.md) — why a glossary belongs in source control (Evans, Vernon, Khononov, Wlaschin, Brandolini, Avram & Marinescu, Fowler)
- [`references/adr_practice.md`](references/adr_practice.md) — when an ADR earns its keep (Nygard, Tyree & Akerman, Zimmermann Y-statements, MADR, ThoughtWorks Radar, adr-tools, Backstage)
- [`references/context_md_as_artifact.md`](references/context_md_as_artifact.md) — CONTEXT.md as living artifact (Khononov on language drift, Kernighan on naming, BoundedContext bliki, Confluent on data contracts, Brandolini on EventStorming glossary)
### Companion
- Agent: `cs-grill-with-docs` (see `../../agents/cs-grill-with-docs.md`)
- Command: `/cs:grill-with-docs` (see `../../commands/cs-grill-with-docs.md`)
---
**Version:** 1.0.0
**Derived:** Matt Pocock's grill-with-docs (MIT) + this repo's wrapper
FILE:ADR-FORMAT.md
<!--
Derived from Matt Pocock's grill-with-docs:
https://github.com/mattpocock/skills/tree/main/skills/engineering/grill-with-docs/ADR-FORMAT.md
MIT License © 2026 Matt Pocock. Reproduced verbatim under MIT.
-->
# ADR Format
ADRs live in `docs/adr/` and use sequential numbering: `0001-slug.md`, `0002-slug.md`, etc.
Create the `docs/adr/` directory lazily — only when the first ADR is needed.
## Template
```md
# {Short title of the decision}
{1-3 sentences: what's the context, what did we decide, and why.}
```
That's it. An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why* — not in filling out sections.
## Optional sections
Only include these when they add genuine value. Most ADRs won't need them.
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by ADR-NNNN`) — useful when decisions are revisited
- **Considered Options** — only when the rejected alternatives are worth remembering
- **Consequences** — only when non-obvious downstream effects need to be called out
## Numbering
Scan `docs/adr/` for the highest existing number and increment by one.
## When to offer an ADR
All three of these must be true:
1. **Hard to reverse** — the cost of changing your mind later is meaningful
2. **Surprising without context** — a future reader will look at the code and wonder "why on earth did they do it this way?"
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
If a decision is easy to reverse, skip it — you'll just reverse it. If it's not surprising, nobody will wonder why. If there was no real alternative, there's nothing to record beyond "we did the obvious thing."
### What qualifies
- **Architectural shape.** "We're using a monorepo." "The write model is event-sourced, the read model is projected into Postgres."
- **Integration patterns between contexts.** "Ordering and Billing communicate via domain events, not synchronous HTTP."
- **Technology choices that carry lock-in.** Database, message bus, auth provider, deployment target. Not every library — just the ones that would take a quarter to swap out.
- **Boundary and scope decisions.** "Customer data is owned by the Customer context; other contexts reference it by ID only." The explicit no-s are as valuable as the yes-s.
- **Deliberate deviations from the obvious path.** "We're using manual SQL instead of an ORM because X." Anything where a reasonable reader would assume the opposite. These stop the next engineer from "fixing" something that was deliberate.
- **Constraints not visible in the code.** "We can't use AWS because of compliance requirements." "Response times must be under 200ms because of the partner API contract."
- **Rejected alternatives when the rejection is non-obvious.** If you considered GraphQL and picked REST for subtle reasons, record it — otherwise someone will suggest GraphQL again in six months.
FILE:CONTEXT-FORMAT.md
<!--
Derived from Matt Pocock's grill-with-docs:
https://github.com/mattpocock/skills/tree/main/skills/engineering/grill-with-docs/CONTEXT-FORMAT.md
MIT License © 2026 Matt Pocock. Reproduced verbatim under MIT.
-->
# CONTEXT.md Format
## Structure
```md
# {Context Name}
{One or two sentence description of what this context is and why it exists.}
## Language
**Order**:
{A concise description of the term}
_Avoid_: Purchase, transaction
**Invoice**:
A request for payment sent to a customer after delivery.
_Avoid_: Bill, payment request
**Customer**:
A person or organization that places orders.
_Avoid_: Client, buyer, account
## Relationships
- An **Order** produces one or more **Invoices**
- An **Invoice** belongs to exactly one **Customer**
## Example dialogue
> **Dev:** "When a **Customer** places an **Order**, do we create the **Invoice** immediately?"
> **Domain expert:** "No — an **Invoice** is only generated once a **Fulfillment** is confirmed."
## Flagged ambiguities
- "account" was used to mean both **Customer** and **User** — resolved: these are distinct concepts.
```
## Rules
- **Be opinionated.** When multiple words exist for the same concept, pick the best one and list the others as aliases to avoid.
- **Flag conflicts explicitly.** If a term is used ambiguously, call it out in "Flagged ambiguities" with a clear resolution.
- **Keep definitions tight.** One sentence max. Define what it IS, not what it does.
- **Show relationships.** Use bold term names and express cardinality where obvious.
- **Only include terms specific to this project's context.** General programming concepts (timeouts, error types, utility patterns) don't belong even if the project uses them extensively. Before adding a term, ask: is this a concept unique to this context, or a general programming concept? Only the former belongs.
- **Group terms under subheadings** when natural clusters emerge. If all terms belong to a single cohesive area, a flat list is fine.
- **Write an example dialogue.** A conversation between a dev and a domain expert that demonstrates how the terms interact naturally and clarifies boundaries between related concepts.
## Single vs multi-context repos
**Single context (most repos):** One `CONTEXT.md` at the repo root.
**Multiple contexts:** A `CONTEXT-MAP.md` at the repo root lists the contexts, where they live, and how they relate to each other:
```md
# Context Map
## Contexts
- [Ordering](./src/ordering/CONTEXT.md) — receives and tracks customer orders
- [Billing](./src/billing/CONTEXT.md) — generates invoices and processes payments
- [Fulfillment](./src/fulfillment/CONTEXT.md) — manages warehouse picking and shipping
## Relationships
- **Ordering → Fulfillment**: Ordering emits `OrderPlaced` events; Fulfillment consumes them to start picking
- **Fulfillment → Billing**: Fulfillment emits `ShipmentDispatched` events; Billing consumes them to generate invoices
- **Ordering ↔ Billing**: Shared types for `CustomerId` and `Money`
```
The skill infers which structure applies:
- If `CONTEXT-MAP.md` exists, read it to find contexts
- If only a root `CONTEXT.md` exists, single context
- If neither exists, create a root `CONTEXT.md` lazily when the first term is resolved
When multiple contexts exist, infer which one the current topic relates to. If unclear, ask.
FILE:references/adr_practice.md
# ADR Practice — When Does a Decision Earn an ADR?
This reference answers exactly one decision: **what bar must an architectural decision clear to be worth writing down as an ADR, and what format keeps the ADR useful 18 months later?**
Pair with `scripts/adr_scanner.py` for filename + numbering + structural validation.
## The Core Claim
ADRs are not a compliance ritual. They exist to answer a single future question: **"Why on earth did they do it this way?"** If a future reader will never ask that question — because the choice is obvious, easy to reverse, or had no real alternatives — the ADR is doc-rot waiting to happen.
The matt-pocock 3-criteria gate (preserved verbatim in `ADR-FORMAT.md`) is the strict version of this principle:
1. **Hard to reverse** — the cost of changing your mind is meaningful (not "an afternoon of refactoring").
2. **Surprising without context** — a future reader will look at the code and wonder why.
3. **Result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons.
**All three must be true.** Two-out-of-three is not enough. If a decision was hard to reverse but obvious and uncontested (e.g., "we used HTTPS"), no ADR. If it was a real trade-off but easy to reverse (e.g., "we used React Query over SWR"), no ADR.
## What Earns an ADR (Examples)
- **Architectural shape.** "Write model is event-sourced, read model is projected into Postgres." Hard-to-reverse + surprising + real-trade-off.
- **Integration patterns between contexts.** "Ordering and Billing communicate via domain events, not synchronous HTTP." Hard-to-reverse (rewiring eventing is expensive) + surprising (HTTP is the obvious choice) + real-trade-off (eventual consistency vs simpler API).
- **Technology choices with lock-in.** Database engine, message bus, auth provider. Not "we picked Lodash" — those swap in an afternoon.
- **Boundary and scope decisions.** "Customer data is owned by the Customer context; other contexts reference by ID only." The explicit no-s are as valuable as the yes-s.
- **Deliberate deviations from the obvious path.** "We use manual SQL instead of an ORM because X." Stops the next engineer from "fixing" something deliberate.
- **Constraints not visible in code.** "Can't use AWS due to compliance." "Response times must be <200ms due to partner API contract."
- **Rejected alternatives with non-obvious rejections.** "We considered GraphQL and picked REST because subscription complexity didn't match our actual real-time needs." Otherwise someone will suggest GraphQL again in 6 months.
## What Does NOT Earn an ADR
- **Library choices.** Lodash vs Ramda, axios vs ky, dayjs vs date-fns — these swap in an afternoon. Comment in code if you must.
- **Style guide decisions.** "We use Prettier" — record in `package.json`, not an ADR.
- **Defaults you didn't deviate from.** "We use the framework's recommended router." No trade-off, no ADR.
- **Decisions that are easy to reverse.** If the future-you can undo it in a day, future-you doesn't need the why.
- **Decisions where the alternative was never seriously considered.** No real trade-off → no ADR.
## Format Discipline
ADRs are markdown files at `docs/adr/NNNN-slug.md`, numbered sequentially.
**Default format (minimum viable):**
```md
# {Short title of the decision}
{1-3 sentences: what's the context, what did we decide, and why.}
```
An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why* — not in filling out sections.
**Optional sections (only when they add genuine value):**
- **Status frontmatter** (`proposed | accepted | deprecated | superseded by ADR-NNNN`) — useful when decisions are revisited.
- **Considered Options** — only when rejected alternatives are worth remembering.
- **Consequences** — only when non-obvious downstream effects need to be called out.
If a section is included but empty or boilerplate ("none"), delete the section.
## Numbering Discipline
- Sequential, zero-padded to 4 digits: `0001`, `0002`, ..., `9999`.
- No gaps. If an ADR is abandoned mid-draft, either commit it as `proposed → withdrawn` or renumber.
- Slug is short, kebab-case, intent-revealing: `0042-event-sourced-orders.md`, not `0042-adr.md` or `0042-decision-about-events.md`.
`scripts/adr_scanner.py` enforces the pattern and surfaces gaps.
## Status Lifecycle (Optional)
For repos that revisit decisions, the status field is useful:
```
proposed → accepted ← default lifecycle for a new ADR
accepted → deprecated ← decision no longer applies; no replacement
accepted → superseded ← replaced by ADR-NNNN; link to successor in frontmatter
```
When superseding, the new ADR references the old (`supersedes: ADR-0017`) and the old ADR is updated with `superseded by: ADR-0042`. This back-link is the single most useful piece of ADR metadata for archeology.
## Anti-Patterns
- **The ADR factory.** Writing an ADR for every PR. Within a year, you have 200 ADRs and no one reads any. The 3-criteria gate is the firewall.
- **The proposal that never accepts.** ADR sits in `proposed` for months. Either accept it (do it) or withdraw it (delete the file or mark withdrawn).
- **The TOC-only ADR.** Filled-in section headers but no actual content. Worse than not writing the ADR — it implies a decision was recorded when nothing was.
- **The future-tense ADR.** "We will use X." ADRs are records, not plans. Write in past tense ("We chose X because ...") so it reads correctly 2 years later.
- **The unanchored ADR.** ADR with no link to the PR/issue/discussion that drove it. The "why" loses fidelity over time without the source thread.
## Operational Checklist (Per ADR Decision Point)
When grilling and a candidate decision emerges:
- [ ] **Reversibility test.** "If we change our mind in 6 months, what's the cost?" If "an afternoon" → skip the ADR.
- [ ] **Surprise test.** "Will a future engineer look at this and wonder why?" If no → skip.
- [ ] **Trade-off test.** "What alternatives did we seriously consider, and why did each lose?" If none → skip.
- [ ] **All three pass.** Write the ADR. Use the minimum format. Re-run `scripts/adr_scanner.py` to confirm numbering.
- [ ] **Frontmatter status.** Only add `status` if revisiting is expected. Default is "implicit accepted".
## Citations (7 sources)
1. **Michael Nygard, "Documenting Architecture Decisions" (cognitect.com, November 2011).** The original ADR essay. Introduces the format (Title / Context / Decision / Status / Consequences) and the core insight that "architecturally significant" decisions deserve records. Nygard's framing of ADRs as "memory aids for future architects" is the source of the 3-criteria gate's first rule (hard-to-reverse).
2. **Jeff Tyree & Art Akerman, "Architecture Decisions: Demystifying Architecture" — *IEEE Software* 22(2), March–April 2005, pp. 19–27.** Pre-dates Nygard. Introduces the concept of an "Architecture Decision Record" as a first-class artifact and argues for explicit recording of rejected alternatives. The "rejected alternatives" section in Nygard's format inherits from Tyree & Akerman.
3. **Olaf Zimmermann et al., "Y-Statements: A Lightweight Architectural Decision Format" — published at various venues including ozimmer.ch.** Proposes the "In the context of {use case / requirement}, facing {concern}, we decided for {option} to achieve {quality}, accepting {downside}" template. Used widely as a compact alternative to the full Nygard format.
4. **MADR (Markdown Architectural Decision Records) — adr.github.io/madr.** Open-source template maintained by a community of practitioners. Specifies frontmatter format (status, deciders, date, consulted, informed) and a discoverable file structure. Useful when ADRs need machine-readable metadata for indexing.
5. **ThoughtWorks Technology Radar — thoughtworks.com/radar.** Has covered "Lightweight Architecture Decision Records" since Vol. 18 (2018) in the Techniques quadrant, with periodic upgrades to "Adopt". TW's "use ADRs sparingly" guidance aligns with the 3-criteria gate.
6. **Joel Parker Henderson, adr-tools (github.com/npryce/adr-tools).** CLI tool implementing Nygard's format with numbering helpers, supersession linking, and a `new` / `link` / `accept` command set. Establishes the de-facto convention of `0001-slug.md` filenames and `docs/adr/` directory location.
7. **Spotify Backstage — backstage.io.** Backstage's TechDocs catalog includes an ADR plugin that surfaces per-service ADRs in the service catalog UI. Demonstrates how ADRs become discoverable at scale (>1000 services) when treated as first-class catalog entries, not just files in a repo.
FILE:references/context_md_as_artifact.md
# CONTEXT.md as a Living Artifact — Preventing Glossary Decay
This reference answers exactly one decision: **how does a glossary stay alive vs decay into doc rot, and what operational practices prevent the drift?**
Pair with `scripts/glossary_code_consistency.py` for the lint-against-codebase reality check and `scripts/context_md_linter.py` for structural validation.
## The Core Claim
Every glossary decays by default. The decay path is well-documented:
```
Month 1: Glossary written during initial DDD workshop. Terms are precise.
Month 3: New feature ships. Two new domain terms used in code, neither added to glossary.
Month 6: A term in the glossary is renamed in code. Glossary still has old name.
Month 9: New engineer joins. Reads glossary. Asks "what's a 'Booking'?" — answer is "we don't call those Bookings anymore, we call them Reservations now."
Month 12: Glossary is officially declared stale. Engineers stop reading it. Drift becomes invisible.
```
The decay is not preventable by good intentions. It is prevented by **inline edits during the work that introduces the term** plus **automated lint runs at PR time** to flag mismatches.
## Three Forces That Drive Drift
1. **Language pressure from outside the bounded context.** A new partner integration uses different terminology ("subscriber" vs your "customer"). Engineers copy the partner's term into code without first reconciling with the glossary.
2. **Refactor pressure inside the bounded context.** A rename in code feels obvious ("`Booking` → `Reservation` is just a better name"), but the glossary isn't updated alongside.
3. **Convergence pressure between teams.** Multiple teams contributing to the same context use slightly different words for the same concept. Without a glossary as referee, all variants end up in code.
`scripts/glossary_code_consistency.py` operationalizes the lint against these three forces:
- **Defined-but-unused term** → a glossary entry that no code references. Either dead glossary (delete) or a rename happened (update glossary to match code).
- **Code-only proper noun** → a frequently-used capitalized term in code that the glossary doesn't define. Either generic (ignore) or domain (add to glossary now).
## Five Practices That Keep CONTEXT.md Alive
1. **Edit inline during the work.** Never batch glossary updates. When a term is introduced or refined during a feature, the same PR that adds the code edits `CONTEXT.md`. Reviewers reject PRs that introduce domain terms without glossary edits.
2. **Lint at PR time.** Run `scripts/context_md_linter.py` and `scripts/glossary_code_consistency.py` in CI. A new term in code without a glossary entry is a build warning; an outright rename mismatch is a build failure.
3. **Per-context glossaries, not one mega-glossary.** Multi-context repos use `CONTEXT-MAP.md` to point at per-context `CONTEXT.md` files. Cross-context terms get explicit translation entries ("Billing's `Customer` is Ordering's `Account`").
4. **Pruning passes.** Quarterly, run `glossary_code_consistency.py` and review the dead-glossary report. Delete entries that no code uses. Keeping dead entries dilutes signal.
5. **One sentence per definition.** If a definition runs to a paragraph, the term is hiding two concepts. Split or sharpen. Long definitions are correlated with imprecise terms.
## How CONTEXT.md Differs from Other "Documentation"
| Artifact | Purpose | Update cadence | Audience |
|---|---|---|---|
| `README.md` | Onboarding + setup | Once at project start, occasionally after | New contributors |
| `ARCHITECTURE.md` | High-level system shape | Quarterly to yearly | New architects, senior engineers |
| `docs/adr/*.md` | Record of specific decisions | Per-decision (rare; days to months apart) | Anyone asking "why did we do X this way?" |
| **`CONTEXT.md`** | **The domain glossary — what each term means in this bounded context** | **Per-feature (continuous; hours to days apart)** | **Every engineer on every PR** |
A `CONTEXT.md` is touched far more often than any other doc because it tracks the language as it evolves. If yours hasn't been edited in 6 months, it's almost certainly drifting.
## Single vs Multi-Context Repos
**Single context (most repos):** One `CONTEXT.md` at the repo root. All terms in scope.
**Multiple contexts:** A `CONTEXT-MAP.md` at the repo root lists the contexts and their relationships. Each bounded context has its own `CONTEXT.md` (and its own `docs/adr/` for context-specific decisions). Shared terms appear in both with cross-references.
```
/
├── CONTEXT-MAP.md ← lists contexts + relationships
├── docs/adr/ ← system-wide ADRs
└── src/
├── ordering/
│ ├── CONTEXT.md ← ordering-context glossary
│ └── docs/adr/ ← ordering-context ADRs
└── billing/
├── CONTEXT.md
└── docs/adr/
```
When a term spans contexts, define it in each `CONTEXT.md` with the context's perspective + a translation note pointing at the other. Don't try to define "Customer" once and have both contexts share it — that's the path back to the mega-glossary.
## Anti-Patterns
- **The spec masquerading as a glossary.** `CONTEXT.md` includes implementation details, sequence diagrams, API responses. It is a glossary, not a spec. Move spec content elsewhere.
- **The wiki masquerading as a glossary.** General programming concepts ("retry", "timeout", "config") appearing in `CONTEXT.md`. They are not domain-specific. Remove.
- **The glossary that defines without forbidding.** Each term needs `_Avoid_: <aliases>` to push back on drift. A glossary that says "Customer means X" but doesn't forbid "Client" / "Account" / "User" cannot push back when those drift in.
- **The frozen glossary.** No commits in 6+ months. Either the project is dormant or the language has drifted away from the document. Re-grill.
- **The orphan glossary.** Sits in a repo but no CI/PR process references it. It will decay within two quarters.
## Operational Checklist
When grilling against `CONTEXT.md`:
- [ ] Lint structure: `python scripts/context_md_linter.py CONTEXT.md`
- [ ] Lint vs code: `python scripts/glossary_code_consistency.py --context CONTEXT.md --code src/`
- [ ] For each "defined but unused": ask "dead term, or rename happened?"
- [ ] For each "code-only proper noun": ask "domain term that needs definition, or generic?"
- [ ] For each new term introduced during the grill: edit `CONTEXT.md` *now*, not "later"
- [ ] Multi-context repo: verify the right `CONTEXT.md` is being edited (not the wrong context's, not the root one when a per-context one applies)
## Citations (7 sources)
1. **Vladimir Khononov, *Learning Domain-Driven Design* (O'Reilly, 2021).** Chapter 9, "Communication Patterns" + Chapter 12, "Building Domain Expertise" — Khononov is the sharpest writer on language drift between bounded contexts and on how to detect it. His "linguistic boundaries are observable boundaries" framing is the foundation of the `glossary_code_consistency.py` check.
2. **Brian Kernighan & Rob Pike, *The Practice of Programming* (Addison-Wesley, 1999).** Chapter 1, "Style" — the section on naming. Kernighan's "names should reflect the role of the variable, not its type" generalizes to glossary terms: a glossary term names a role in the domain, not a data structure. Kernighan-style naming discipline is what keeps `CONTEXT.md` precise.
3. **Martin Fowler, "BoundedContext" — martinfowler.com bliki (2014, updated).** The canonical argument that ubiquitous language is **bounded** — it applies inside one context, not across all contexts. The justification for per-context `CONTEXT.md` files. https://martinfowler.com/bliki/BoundedContext.html
4. **Martin Fowler, "UbiquitousLanguage" — martinfowler.com bliki.** Companion entry to BoundedContext. Articulates the discipline of using the same vocabulary in conversation, in the model, and in the code. The justification for editing `CONTEXT.md` inline alongside code changes, not as separate doc work. https://martinfowler.com/bliki/UbiquitousLanguage.html
5. **Confluent Schema Registry / Data Contracts community — confluent.io/blog/data-contracts.** The data-contracts movement applies UL discipline to inter-service / inter-context boundaries: when two contexts exchange events or API payloads, the schema is a binding glossary. Drift between contexts becomes a schema-evolution problem, not a free-form documentation problem.
6. **Alberto Brandolini, *Introducing EventStorming* (Leanpub, ongoing).** Chapter on "Pivotal Events" + the convergence-workshop chapter. Brandolini documents how a glossary emerges from EventStorming workshops as a by-product of mapping events. The pattern of "capture the term on a sticky note when it surfaces" is the offline equivalent of the inline `CONTEXT.md` edit discipline.
7. **Eric Evans, *Domain-Driven Design: Tackling Complexity in the Heart of Software* (Addison-Wesley, 2003).** Chapter 14, "Maintaining Model Integrity" — covers the Conformist, Anticorruption Layer, and Shared Kernel patterns. Each of these is a strategy for managing the boundary between two bounded contexts that have different languages. Justifies the multi-context `CONTEXT-MAP.md` pattern and the translation-note discipline for cross-context terms.
FILE:references/ubiquitous_language.md
# Ubiquitous Language — Why a Glossary Belongs in Source Control
This reference answers exactly one decision: **why should a project's domain glossary (`CONTEXT.md`) live next to the code in source control, and what bar must it clear to earn its keep?**
Pair with `scripts/context_md_linter.py` for structural validation and `scripts/glossary_code_consistency.py` for the language-vs-code reality check.
## The Core Claim
A bounded context has **one** language. The same word must mean the same thing in conversation, in the glossary, in the type system, in the database schema, and in the UI. When language fractures across these surfaces, design defects follow: ambiguous bug reports, mismatched API contracts, broken refactors, junior engineers asking what an "account" is and getting three different answers.
The glossary is the contract that prevents the fracture. It earns its place in source control because it changes at the same cadence as the code — every time a domain term is introduced, refined, or retired, the glossary must move with it. A wiki page that lives outside the repo will drift within a quarter.
## Why a Glossary in Source Control (vs Wiki, Notion, Confluence)
| Property | In-repo `CONTEXT.md` | External wiki |
|---|---|---|
| Reviewable in PR | Yes — diff is visible alongside code | No — reviewer must remember to check |
| Versioned with code | Yes — `git log` shows term evolution | No — wikis rarely have meaningful history |
| Discoverable by new engineers | Yes — `ls` of repo root finds it | No — depends on onboarding tribal knowledge |
| Mergeable | Yes — text format, conflict-resolvable | Often no — UI-driven |
| Linter-targetable | Yes — `scripts/context_md_linter.py` | No — usually not |
| Refactor-safe | Yes — renames are grep-able | No — wiki links rot silently |
The glossary is a **language artifact**, not a documentation artifact. Documentation describes the system; the glossary **is** part of the system's design surface.
## Five Rules That Make a Glossary Survive
1. **One sentence per definition.** If the definition needs a paragraph, the term is hiding two concepts. Split it.
2. **Define what it IS, not what it does.** "An **Invoice** is a request for payment sent after delivery." Not "An invoice handles billing."
3. **List aliases to avoid.** When users say "bill" or "payment request" but mean "invoice", record that "bill" is forbidden. Without the `_Avoid_:` field, the glossary cannot push back on drift.
4. **Show relationships, not just terms.** "An **Order** produces one or more **Invoices**" tells you the cardinality. A list of bare terms doesn't.
5. **Exclude generic programming concepts.** "Timeout", "retry", "config" do not belong. Only terms specific to this project's domain qualify.
## Anti-Patterns
- **The "everything goes in" glossary.** When `CONTEXT.md` includes general programming concepts (timeout, error, util), it dilutes signal and degenerates into a wiki page.
- **The orphan glossary.** Terms defined but never used in code. Either the term is dead (delete it) or the code is using a synonym (rename code).
- **The opaque glossary.** Terms used in code but not defined. Either the term is generic (don't define it) or it's a domain concept that snuck in (define it now).
- **The deferred glossary edit.** "I'll batch up the glossary changes at the end of the sprint." By the end of the sprint, three more drift cases will have shipped. Glossary edits must land inline.
- **The frozen glossary.** No commits in 6+ months. Either the project is dormant or the language has drifted away from the document.
## Operational Checklist (for the Grill Session)
When grilling a plan against `CONTEXT.md`:
- [ ] Pre-flight `scripts/context_md_linter.py CONTEXT.md` — is the glossary well-formed?
- [ ] Run `scripts/glossary_code_consistency.py` — what's defined but unused? what's used but undefined?
- [ ] For every novel term in the plan, ask: "Is this in CONTEXT.md? If not, do we add it, or do we rephrase using an existing term?"
- [ ] For every existing term used in the plan, ask: "Does the plan use it consistent with the definition?"
- [ ] At every clarification moment, edit `CONTEXT.md` immediately — never batch.
## Citations (7 sources)
1. **Eric Evans, *Domain-Driven Design: Tackling Complexity in the Heart of Software* (Addison-Wesley, 2003).** Chapter 2, "Communication and the Use of Language" — the canonical statement of Ubiquitous Language as a design tool, not just documentation. The line "The vocabulary of that UBIQUITOUS LANGUAGE includes the names of classes and prominent operations" is the bridge between conversation and code.
2. **Vaughn Vernon, *Implementing Domain-Driven Design* (Addison-Wesley, 2013).** Chapter 1, "Getting Started with DDD" + Chapter 2, "Domains, Subdomains, and Bounded Contexts" — operationalizes Evans's UL into a workshop format and per-context discipline. Vernon's "linguistic boundaries are the most reliable boundary" framing is the source of the per-bounded-context glossary pattern.
3. **Vladimir Khononov, *Learning Domain-Driven Design* (O'Reilly, 2021).** Chapter 5, "Implementing Simple Business Logic" + Chapter 9, "Communication Patterns" — Khononov is sharpest on what happens when bounded contexts share a language vs maintain separate languages (translation layer required) and on language drift over time.
4. **Scott Wlaschin, *Domain Modeling Made Functional* (Pragmatic Bookshelf, 2018).** Part 1, "Understanding the Domain" — treats the type system as the executable form of the glossary. Wlaschin's "make illegal states unrepresentable" is the strongest form of glossary-as-contract: if the glossary says an Order must have at least one line item, the type prevents zero-item Orders at compile time.
5. **Alberto Brandolini, *Introducing EventStorming: An Act of Deliberate Collective Learning* (Leanpub, 2017–ongoing).** Chapter on "Sticky note color codes" + chapter on convergence — EventStorming workshops produce a glossary as a by-product of mapping the domain. Brandolini's pattern of capturing terms as they emerge on sticky notes is the offline equivalent of the inline `CONTEXT.md` edit.
6. **Abel Avram & Floyd Marinescu, *Domain-Driven Design Quickly* (InfoQ, 2006, free e-book).** Chapter 2, "Ubiquitous Language" — the most concise distillation of Evans's UL chapter. Useful as a reference to hand to engineers who won't read the blue book.
7. **Martin Fowler, "BoundedContext" — martinfowler.com bliki (2014, updated).** Fowler's framing of "Ubiquitous Language … doesn't apply to the whole project, it only has to apply within a particular Bounded Context" justifies the per-context glossary pattern in `CONTEXT-MAP.md`-style multi-context repos. https://martinfowler.com/bliki/BoundedContext.html
FILE:scripts/adr_scanner.py
#!/usr/bin/env python3
"""adr_scanner.py — Walk docs/adr/ and validate ADR files against the format.
Stdlib-only. Applies the rules from Matt Pocock's upstream ADR-FORMAT.md
(preserved verbatim in the skill's ADR-FORMAT.md):
1. Each file matches the `NNNN-slug.md` pattern (4-digit zero-padded number + kebab-case slug)
2. Numbering is sequential — no gaps, no duplicates
3. Each ADR has an H1 (the title)
4. Each ADR has a non-empty body after the H1 (at least the 1-3 sentence context+decision)
5. Optional status frontmatter, if present, has a valid value
(proposed | accepted | deprecated | superseded by ADR-NNNN)
6. Superseded-by references point at an existing ADR number
Output: directory-level summary + per-file findings.
NO LLM CALLS. Pure regex + filesystem walking.
Usage:
python adr_scanner.py docs/adr/
python adr_scanner.py docs/adr/ --output json
python adr_scanner.py --sample # scan an embedded sample directory layout
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
ADR_FILENAME_RE = re.compile(r"^(\d{4})-([a-z0-9]+(?:-[a-z0-9]+)*)\.md$")
VALID_STATUSES = {"proposed", "accepted", "deprecated"}
SUPERSEDED_RE = re.compile(r"^superseded\s+by\s+ADR-?(\d{1,4})$", re.IGNORECASE)
SAMPLE_ADRS: Dict[str, str] = {
"0001-event-sourced-orders.md": (
"# Event-source the Order write model\n"
"\n"
"We need an audit trail of every state change on an Order for compliance + analytics. "
"We chose event sourcing for the Order write model and a Postgres projection for the read model. "
"Trade-off accepted: eventual consistency on the read side in exchange for the audit trail and replay.\n"
),
"0002-postgres-for-write-model.md": (
"---\n"
"status: accepted\n"
"---\n"
"\n"
"# Postgres for the write-side event store\n"
"\n"
"We considered EventStore and Kafka. Postgres won on operational familiarity + transactional guarantees + cost.\n"
),
"0003-rest-over-graphql.md": (
"---\n"
"status: accepted\n"
"---\n"
"\n"
"# REST over GraphQL for the public API\n"
"\n"
"GraphQL would have given clients more flexibility but added subscription complexity we don't need at our scale.\n"
),
}
def parse_frontmatter(text: str) -> Tuple[Dict[str, str], str]:
"""Return (frontmatter_dict, body) for a file that may have YAML-ish frontmatter.
Only handles simple `key: value` lines (no nested YAML, no lists) — stdlib-only.
"""
if not text.startswith("---\n"):
return {}, text
end_marker = text.find("\n---\n", 4)
if end_marker == -1:
return {}, text
fm_block = text[4:end_marker]
body = text[end_marker + 5 :]
fm: Dict[str, str] = {}
for line in fm_block.splitlines():
if ":" in line:
k, v = line.split(":", 1)
fm[k.strip().lower()] = v.strip()
return fm, body
def scan_directory(adr_dir: Path) -> Dict[str, Any]:
findings: List[Dict[str, Any]] = []
files: List[Tuple[int, str, Path]] = []
def add(file: str, rule: str, level: str, message: str) -> None:
findings.append({"file": file, "rule": rule, "level": level, "message": message})
if not adr_dir.exists():
add("(root)", "directory", "FAIL", f"Directory does not exist: {adr_dir}")
return finalize(findings, 0)
if not adr_dir.is_dir():
add("(root)", "directory", "FAIL", f"Path is not a directory: {adr_dir}")
return finalize(findings, 0)
md_files = sorted(p for p in adr_dir.iterdir() if p.is_file() and p.suffix == ".md")
if not md_files:
add("(root)", "directory", "WARN", "Directory is empty — no ADRs scanned. Create lazily when the first ADR is needed.")
return finalize(findings, 0)
# Rule 1: filename pattern
for p in md_files:
m = ADR_FILENAME_RE.match(p.name)
if not m:
add(p.name, "filename-pattern", "FAIL", f"Filename does not match NNNN-slug.md pattern. Expected e.g. 0001-event-sourced-orders.md.")
continue
number = int(m.group(1))
files.append((number, p.name, p))
add(p.name, "filename-pattern", "PASS", f"Filename matches pattern (number={number:04d}).")
files.sort(key=lambda t: t[0])
# Rule 2: numbering sequence (no gaps, no duplicates)
seen: Dict[int, List[str]] = {}
for number, name, _ in files:
seen.setdefault(number, []).append(name)
for number, names in seen.items():
if len(names) > 1:
add(", ".join(names), "numbering-duplicate", "FAIL", f"Duplicate ADR number {number:04d}.")
if files:
expected = list(range(1, files[-1][0] + 1))
actual = sorted(seen.keys())
gaps = [n for n in expected if n not in actual]
if gaps:
add("(root)", "numbering-gap", "WARN", f"Number gap(s) in sequence: {', '.join(f'{g:04d}' for g in gaps)}. Either commit withdrawn ADRs as 'proposed → withdrawn' or renumber.")
else:
add("(root)", "numbering-sequence", "PASS", f"Sequential numbering 0001..{files[-1][0]:04d} with no gaps.")
# Rules 3, 4, 5, 6: per-ADR
numbers_present = {n for n, _, _ in files}
for number, name, path in files:
text = path.read_text(encoding="utf-8") if path.is_file() else SAMPLE_ADRS.get(name, "")
fm, body = parse_frontmatter(text)
# Rule 3: H1 present
h1_match = re.search(r"^#\s+(.+?)\s*$", body, re.MULTILINE)
if not h1_match:
add(name, "h1-present", "FAIL", "No H1 (`# Title`) found in body.")
continue
else:
add(name, "h1-present", "PASS", f"H1 found: '{h1_match.group(1).strip()}'.")
# Rule 4: non-empty body after H1
after_h1 = body[h1_match.end():].strip()
if not after_h1:
add(name, "body-non-empty", "FAIL", "ADR has H1 but no body. The 1-3 sentence context+decision is required.")
else:
word_count = len(re.findall(r"\b\w+\b", after_h1))
if word_count < 10:
add(name, "body-non-empty", "WARN", f"ADR body is very short ({word_count} words). Confirm context+decision+why are all stated.")
else:
add(name, "body-non-empty", "PASS", f"Body present ({word_count} words).")
# Rule 5: optional status frontmatter sanity
status = fm.get("status", "").strip().lower() if fm else ""
if status:
if status in VALID_STATUSES:
add(name, "status-frontmatter", "PASS", f"Status '{status}' is valid.")
elif SUPERSEDED_RE.match(status):
m = SUPERSEDED_RE.match(status)
target = int(m.group(1))
# Rule 6: superseded-by points at existing ADR
if target in numbers_present:
add(name, "status-supersede-target", "PASS", f"Superseded by ADR-{target:04d} which exists.")
else:
add(name, "status-supersede-target", "FAIL", f"Superseded by ADR-{target:04d} but that ADR is not present in this directory.")
else:
add(name, "status-frontmatter", "FAIL", f"Status '{status}' is not one of {sorted(VALID_STATUSES)} or 'superseded by ADR-NNNN'.")
return finalize(findings, len(files))
def finalize(findings: List[Dict[str, Any]], adr_count: int) -> Dict[str, Any]:
counts = {"PASS": 0, "WARN": 0, "FAIL": 0}
for f in findings:
counts[f["level"]] += 1
if counts["FAIL"] > 0:
verdict = "FAIL"
elif counts["WARN"] > 0:
verdict = "WARN"
else:
verdict = "PASS"
return {"verdict": verdict, "adr_count": adr_count, "counts": counts, "findings": findings}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append(f"ADR directory scan verdict: {result['verdict']}")
out.append(f" ADRs scanned: {result['adr_count']}")
counts = result["counts"]
out.append(f" PASS: {counts['PASS']} WARN: {counts['WARN']} FAIL: {counts['FAIL']}")
out.append("")
out.append("Findings:")
for f in result["findings"]:
marker = {"PASS": "[ok]", "WARN": "[warn]", "FAIL": "[FAIL]"}[f["level"]]
out.append(f" {marker} {f['file']:<40s} {f['rule']}: {f['message']}")
return "\n".join(out)
def run_sample() -> Dict[str, Any]:
"""Scan the embedded sample by writing it to a tempdir."""
import tempfile
with tempfile.TemporaryDirectory() as td:
d = Path(td) / "adr"
d.mkdir()
for name, content in SAMPLE_ADRS.items():
(d / name).write_text(content, encoding="utf-8")
return scan_directory(d)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("adr_dir", nargs="?", help="Path to docs/adr/ directory")
parser.add_argument("--sample", action="store_true", help="Scan the embedded sample ADR layout")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = run_sample()
elif args.adr_dir:
result = scan_directory(Path(args.adr_dir))
else:
parser.print_help()
return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0 if result["verdict"] != "FAIL" else 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/context_md_linter.py
#!/usr/bin/env python3
"""context_md_linter.py — Validate a CONTEXT.md against the CONTEXT-FORMAT.md structure.
Stdlib-only. Walks a CONTEXT.md and applies the format rules from Matt Pocock's
upstream CONTEXT-FORMAT.md (preserved verbatim in the skill's CONTEXT-FORMAT.md):
1. H1 present at top (the context name)
2. One-or-two-sentence description follows the H1
3. ## Language section present
4. Inside Language: each term is in `**Term**:` bold form
5. Inside Language: each term has a one-sentence definition
6. Inside Language: each term has a `_Avoid_:` aliases line (WARN if missing)
7. ## Relationships section present (WARN if missing)
8. ## Example dialogue section present (WARN if missing)
9. Optional: ## Flagged ambiguities section
Output: PASS / WARN / FAIL per rule + an overall verdict.
NO LLM CALLS. Pure regex + line walking.
Usage:
python context_md_linter.py CONTEXT.md
python context_md_linter.py CONTEXT.md --output json
python context_md_linter.py --sample # lint the embedded sample
"""
import argparse
import json
import re
import sys
from pathlib import Path
from typing import Any, Dict, List, Tuple
SAMPLE_CONTEXT_MD = """# Ordering
The ordering context receives customer orders and tracks them through to handoff to Fulfillment.
## Language
**Order**:
A confirmed request from a Customer to acquire one or more Products.
_Avoid_: Purchase, transaction, cart
**Customer**:
A person or organization that places Orders.
_Avoid_: Client, buyer, account
**Product**:
A single SKU that can appear on an Order line.
_Avoid_: Item, good, SKU
## Relationships
- An **Order** belongs to exactly one **Customer**
- An **Order** has one or more **Products** via line items
- A **Customer** can have many **Orders**
## Example dialogue
> **Dev:** "When a **Customer** places an **Order**, are the **Products** locked at order time?"
> **Domain expert:** "Yes — Product price + spec is snapshotted onto the Order line. Subsequent Product edits don't change historical Orders."
## Flagged ambiguities
- "account" was used to mean both **Customer** and "billing account" — resolved: billing account moves to Billing context.
"""
def split_into_sections(text: str) -> Dict[str, str]:
"""Split markdown into top-level ## sections keyed by header text."""
sections: Dict[str, str] = {}
current_header = "_preamble_"
buffer: List[str] = []
for line in text.splitlines():
m = re.match(r"^##\s+(.+?)\s*$", line)
if m:
sections[current_header] = "\n".join(buffer).strip()
current_header = m.group(1).strip().lower()
buffer = []
else:
buffer.append(line)
sections[current_header] = "\n".join(buffer).strip()
return sections
def extract_terms(language_section: str) -> List[Tuple[str, str, str]]:
"""Return list of (term, definition_line, avoid_line) tuples from the Language section.
Each term entry looks like:
**Term**:
Definition sentence.
_Avoid_: alias1, alias2
"""
results: List[Tuple[str, str, str]] = []
# Match `**Term**:` followed by the next non-empty line as definition,
# and optionally an `_Avoid_:` line within the next 3 lines.
pattern = re.compile(
r"\*\*([^*]+?)\*\*\s*:\s*\n([^\n]+)\n?(?:([^\n]*_Avoid_[^\n]*)\n?)?",
re.MULTILINE,
)
for match in pattern.finditer(language_section):
term = match.group(1).strip()
definition = match.group(2).strip()
avoid = (match.group(3) or "").strip()
results.append((term, definition, avoid))
return results
def lint(text: str) -> Dict[str, Any]:
findings: List[Dict[str, str]] = []
def add(rule: str, level: str, message: str) -> None:
findings.append({"rule": rule, "level": level, "message": message})
# Rule 1: H1 present
lines = text.splitlines()
h1_line_index = None
for i, line in enumerate(lines):
if re.match(r"^#\s+\S", line):
h1_line_index = i
break
if h1_line_index is None:
add("h1-present", "FAIL", "No H1 (top-level '# Title') found. CONTEXT.md must start with the context name as H1.")
else:
add("h1-present", "PASS", f"H1 found at line {h1_line_index + 1}.")
# Rule 2: one-or-two-sentence description after H1
if h1_line_index is not None:
desc_lines: List[str] = []
for line in lines[h1_line_index + 1 :]:
if re.match(r"^##\s", line):
break
if line.strip():
desc_lines.append(line.strip())
desc = " ".join(desc_lines).strip()
sentence_count = len(re.findall(r"[.!?](?:\s|$)", desc))
if not desc:
add("description-present", "FAIL", "No description sentence between the H1 and the first ## section.")
elif sentence_count > 3:
add(
"description-length",
"WARN",
f"Description has {sentence_count} sentences. CONTEXT-FORMAT.md asks for one or two.",
)
else:
add("description-present", "PASS", f"Description present ({sentence_count} sentence(s)).")
# Rule 3: ## Language section present
sections = split_into_sections(text)
if "language" not in sections:
add("language-section", "FAIL", "No '## Language' section found. This is the required core of CONTEXT.md.")
return finalize(findings)
add("language-section", "PASS", "'## Language' section found.")
# Rules 4 + 5 + 6: terms inside Language
terms = extract_terms(sections["language"])
if not terms:
add(
"language-terms",
"FAIL",
"No terms detected in the Language section. Each term must be in '**Term**:' bold form followed by a one-sentence definition.",
)
else:
add("language-terms", "PASS", f"Detected {len(terms)} term(s) in Language section.")
for term, definition, avoid in terms:
# Rule 5: definition exists
if not definition or definition.startswith("_Avoid_") or definition.startswith("**"):
add(
"term-definition",
"FAIL",
f"Term '**{term}**:' has no definition line (next non-empty line should be the definition).",
)
else:
# Length heuristic: definition should be <= 200 chars (one sentence-ish)
if len(definition) > 200:
add(
"term-definition-length",
"WARN",
f"Term '**{term}**' definition is {len(definition)} chars. CONTEXT-FORMAT.md asks for one sentence max.",
)
# Rule 6: _Avoid_ line
if not avoid:
add(
"term-avoid",
"WARN",
f"Term '**{term}**' has no '_Avoid_:' aliases line. Without forbidden aliases, the glossary can't push back on drift.",
)
# Rule 7: Relationships section
if "relationships" not in sections:
add(
"relationships-section",
"WARN",
"No '## Relationships' section found. CONTEXT-FORMAT.md asks for one to show cardinality between terms.",
)
else:
add("relationships-section", "PASS", "'## Relationships' section found.")
# Rule 8: Example dialogue
if "example dialogue" not in sections:
add(
"example-dialogue",
"WARN",
"No '## Example dialogue' section found. CONTEXT-FORMAT.md asks for a dev/domain-expert exchange.",
)
else:
add("example-dialogue", "PASS", "'## Example dialogue' section found.")
# Rule 9: Flagged ambiguities (optional, only check presence)
if "flagged ambiguities" in sections:
add("flagged-ambiguities", "PASS", "'## Flagged ambiguities' section found (optional but useful).")
return finalize(findings)
def finalize(findings: List[Dict[str, str]]) -> Dict[str, Any]:
counts = {"PASS": 0, "WARN": 0, "FAIL": 0}
for f in findings:
counts[f["level"]] += 1
if counts["FAIL"] > 0:
verdict = "FAIL"
elif counts["WARN"] > 0:
verdict = "WARN"
else:
verdict = "PASS"
return {"verdict": verdict, "counts": counts, "findings": findings}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
verdict = result["verdict"]
counts = result["counts"]
out.append(f"CONTEXT.md lint verdict: {verdict}")
out.append(f" PASS: {counts['PASS']} WARN: {counts['WARN']} FAIL: {counts['FAIL']}")
out.append("")
out.append("Findings:")
for f in result["findings"]:
marker = {"PASS": "[ok]", "WARN": "[warn]", "FAIL": "[FAIL]"}[f["level"]]
out.append(f" {marker} {f['rule']}: {f['message']}")
return "\n".join(out)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("path", nargs="?", help="Path to CONTEXT.md")
parser.add_argument("--sample", action="store_true", help="Lint the embedded sample CONTEXT.md")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
text = SAMPLE_CONTEXT_MD
elif args.path:
p = Path(args.path)
if not p.exists():
print(f"error: {args.path} not found", file=sys.stderr)
return 2
text = p.read_text(encoding="utf-8")
else:
parser.print_help()
return 0
result = lint(text)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0 if result["verdict"] != "FAIL" else 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/glossary_code_consistency.py
#!/usr/bin/env python3
"""glossary_code_consistency.py — Cross-reference CONTEXT.md terms against the codebase.
Stdlib-only. Reads bold terms from CONTEXT.md and scans a codebase directory for
each term's usage. Surfaces two grilling-question seeds:
1. DEAD GLOSSARY — a term is defined in CONTEXT.md but never appears in code.
Either the term is stale (delete it) or the code uses a synonym (rename).
2. CODE-ONLY PROPER NOUN — a capitalized word that appears frequently in code
but isn't defined in CONTEXT.md. Either it's a generic programming concept
(ignore) or it's a domain term that snuck in undefined (add to glossary).
Both lists are seeded as opening grill-with-docs questions.
NO LLM CALLS. Pure file walking + regex + frequency counting.
Limitations (intentional, stdlib-only):
- Word-boundary matching is case-insensitive. "Order" matches "order", "ORDER", "orders".
- "Code-only proper noun" detection uses a simple heuristic: capitalized
words >= MIN_FREQUENCY occurrences across non-test files. Tunable via flags.
- Only scans common source extensions by default (override with --extensions).
Usage:
python glossary_code_consistency.py --context CONTEXT.md --code src/
python glossary_code_consistency.py --context CONTEXT.md --code src/ --output json
python glossary_code_consistency.py --sample
"""
import argparse
import json
import re
import sys
from collections import Counter
from pathlib import Path
from typing import Any, Dict, List, Set, Tuple
DEFAULT_EXTENSIONS = {
".py",
".ts",
".tsx",
".js",
".jsx",
".go",
".java",
".kt",
".rb",
".cs",
".rs",
".swift",
".php",
".scala",
".clj",
".ex",
".exs",
}
DEFAULT_EXCLUDE_DIRS = {"node_modules", ".git", "dist", "build", "target", ".venv", "venv", "__pycache__"}
TEST_FILE_HINTS = (".test.", ".spec.", "_test.", "tests/", "/test/")
PROPER_NOUN_RE = re.compile(r"\b([A-Z][a-zA-Z]{2,})\b")
GENERIC_WORDS = {
# Programming concepts that capitalize but aren't domain terms
"True", "False", "None", "Null", "Promise", "Error", "Exception",
"String", "Number", "Boolean", "Array", "Object", "Map", "Set",
"List", "Dict", "Tuple", "Optional", "Any", "Result", "Date",
"Math", "JSON", "URL", "URI", "HTTP", "HTTPS", "API", "ID", "UUID",
"GET", "POST", "PUT", "DELETE", "PATCH", "OK", "TODO", "FIXME",
"Test", "Mock", "Stub", "Spy", "Given", "When", "Then", "Describe",
}
SAMPLE_CONTEXT_MD = """# Ordering
## Language
**Order**:
A confirmed request from a Customer to acquire one or more Products.
_Avoid_: Purchase, transaction
**Customer**:
A person or organization that places Orders.
_Avoid_: Client, buyer
**Product**:
A single SKU that can appear on an Order line.
_Avoid_: Item, good
**Discount**:
A reduction applied to an Order at checkout.
_Avoid_: Coupon, promo
"""
SAMPLE_CODE_FILES: Dict[str, str] = {
"src/orders.py": (
"class Order:\n"
" pass\n"
"\n"
"def cancel_order(order_id: str) -> None:\n"
" pass\n"
"\n"
"def list_customer_orders(customer_id: str) -> list[Order]:\n"
" pass\n"
),
"src/customers.py": (
"class Customer:\n"
" pass\n"
"\n"
"class Subscription:\n"
" # NOTE: Subscription is used heavily but not in glossary\n"
" pass\n"
"\n"
"def find_customer(email: str) -> Customer:\n"
" pass\n"
),
"src/products.py": (
"class Product:\n"
" pass\n"
"\n"
"class Inventory:\n"
" pass\n"
"\n"
"def find_product(sku: str) -> Product:\n"
" pass\n"
),
# Note: Discount is defined in glossary but never used in code.
}
def extract_glossary_terms(context_md_text: str) -> List[str]:
"""Pull bold terms from CONTEXT.md `**Term**:` patterns."""
return re.findall(r"\*\*([^*]+?)\*\*\s*:", context_md_text)
def walk_codebase(root: Path, extensions: Set[str], exclude_dirs: Set[str]) -> List[Path]:
found: List[Path] = []
for path in root.rglob("*"):
if path.is_dir():
continue
if any(part in exclude_dirs for part in path.parts):
continue
if path.suffix in extensions:
found.append(path)
return found
def is_test_file(path: Path) -> bool:
s = str(path).replace("\\", "/")
return any(hint in s for hint in TEST_FILE_HINTS)
def count_term_in_text(text: str, term: str) -> int:
pattern = re.compile(rf"\b{re.escape(term)}\b", re.IGNORECASE)
return len(pattern.findall(text))
def count_proper_nouns(text: str) -> Counter:
counter: Counter = Counter()
for match in PROPER_NOUN_RE.finditer(text):
counter[match.group(1)] += 1
return counter
def analyze(
context_md_text: str,
code_files: List[Tuple[str, str]],
min_proper_noun_frequency: int,
) -> Dict[str, Any]:
"""code_files: list of (relative_path, text) tuples."""
glossary_terms = extract_glossary_terms(context_md_text)
glossary_term_set_lower = {t.lower() for t in glossary_terms}
# Per-term usage count in non-test files
term_usage: Dict[str, int] = {t: 0 for t in glossary_terms}
code_proper_nouns: Counter = Counter()
files_scanned = 0
files_tests_skipped = 0
for path_str, text in code_files:
path = Path(path_str)
if is_test_file(path):
files_tests_skipped += 1
continue
files_scanned += 1
for term in glossary_terms:
term_usage[term] += count_term_in_text(text, term)
for noun, count in count_proper_nouns(text).items():
code_proper_nouns[noun] += count
# Dead glossary: terms with zero usage
dead_terms = [t for t, n in term_usage.items() if n == 0]
# Code-only proper nouns: frequent capitalized identifiers NOT in glossary
# and NOT in the generic stop-list
code_only: List[Tuple[str, int]] = []
for noun, count in code_proper_nouns.most_common():
if count < min_proper_noun_frequency:
break
if noun.lower() in glossary_term_set_lower:
continue
if noun in GENERIC_WORDS:
continue
code_only.append((noun, count))
return {
"files_scanned": files_scanned,
"files_tests_skipped": files_tests_skipped,
"glossary_term_count": len(glossary_terms),
"term_usage": term_usage,
"dead_glossary_terms": dead_terms,
"code_only_proper_nouns": code_only,
"min_proper_noun_frequency": min_proper_noun_frequency,
}
def render_human(result: Dict[str, Any]) -> str:
out: List[str] = []
out.append("Glossary↔Code consistency report")
out.append(f" Files scanned: {result['files_scanned']} (test files skipped: {result['files_tests_skipped']})")
out.append(f" Glossary terms: {result['glossary_term_count']}")
out.append("")
out.append("Term usage (occurrences in non-test code):")
for term, count in sorted(result["term_usage"].items(), key=lambda kv: (-kv[1], kv[0])):
marker = " " if count > 0 else "!!"
out.append(f" {marker} {term:<30s} {count}")
out.append("")
if result["dead_glossary_terms"]:
out.append("DEAD GLOSSARY (defined but never used in code) — grill these:")
for term in result["dead_glossary_terms"]:
out.append(f" - '{term}': dead term, or rename happened?")
else:
out.append("DEAD GLOSSARY: (none — every defined term is used in code)")
out.append("")
if result["code_only_proper_nouns"]:
out.append(
f"CODE-ONLY PROPER NOUNS (>= {result['min_proper_noun_frequency']}x, not in glossary, not generic) — grill these:"
)
for noun, count in result["code_only_proper_nouns"]:
out.append(f" - '{noun}' ({count} occurrences): domain term that needs definition, or generic?")
else:
out.append("CODE-ONLY PROPER NOUNS: (none above threshold — glossary covers the frequent domain nouns)")
return "\n".join(out)
def run_sample(min_freq: int) -> Dict[str, Any]:
files = [(p, t) for p, t in SAMPLE_CODE_FILES.items()]
return analyze(SAMPLE_CONTEXT_MD, files, min_freq)
def main(argv: List[str]) -> int:
parser = argparse.ArgumentParser(description=__doc__.split("\n")[0])
parser.add_argument("--context", help="Path to CONTEXT.md")
parser.add_argument("--code", help="Path to codebase root")
parser.add_argument(
"--extensions",
help="Comma-separated source extensions to scan (default: common languages)",
default=None,
)
parser.add_argument(
"--min-frequency",
type=int,
default=3,
help="Minimum occurrences for a code-only proper noun to surface (default: 3)",
)
parser.add_argument("--sample", action="store_true", help="Run on the embedded sample data")
parser.add_argument("--output", choices=["human", "json"], default="human")
args = parser.parse_args(argv)
if args.sample:
result = run_sample(args.min_frequency)
elif args.context and args.code:
context_path = Path(args.context)
code_root = Path(args.code)
if not context_path.exists():
print(f"error: {args.context} not found", file=sys.stderr)
return 2
if not code_root.exists():
print(f"error: {args.code} not found", file=sys.stderr)
return 2
if args.extensions:
exts = {e.strip() if e.strip().startswith(".") else "." + e.strip() for e in args.extensions.split(",")}
else:
exts = DEFAULT_EXTENSIONS
files: List[Tuple[str, str]] = []
for p in walk_codebase(code_root, exts, DEFAULT_EXCLUDE_DIRS):
try:
files.append((str(p), p.read_text(encoding="utf-8", errors="ignore")))
except (OSError, UnicodeDecodeError):
continue
result = analyze(context_path.read_text(encoding="utf-8"), files, args.min_frequency)
else:
parser.print_help()
return 0
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Huấn luyện viên cá nhân giúp người dùng trở thành người dùng Claude thành thạo qua mẹo và cách viết prompt.
---
Name: claude-coach
name: claude-coach
description: Personal coach that teaches users to become Claude power users. Use this skill the FIRST time a user asks to "learn Claude", "be a power user", "coach me", "teach me Claude tricks", "what can Claude do", "make me better at prompting", or any variation. After activation, also use it on EVERY subsequent turn to detect missed optimization opportunities (vague prompts, ignored capabilities, manual work Claude could automate) and surface a single power-user tip. Trigger generously — most users do not know what they do not know, so err on the side of coaching.
Tier: POWERFUL
Category: meta
Author: claude-skills
Dependencies: python3.11
Version: 1.0.0
version: 2.9.0
license: MIT
---
# Claude Coach — Your Power-User Companion
A coaching layer that runs alongside normal conversations. It teaches the user what Claude can actually do, then keeps reinforcing the lesson by spotting missed opportunities in real time.
## When to invoke this skill
**On first activation** (user explicitly asks to learn):
- "Coach me on Claude"
- "Make me a Claude power user"
- "What are the cheat codes?"
- "Teach me how to use Claude better"
- "How do I get more out of Claude?"
**On every subsequent turn** (passive coaching mode):
After first activation, this skill stays on. Every response, scan for coachable moments. Most turns produce zero tips — that is correct behavior. Only surface a tip when it would genuinely 10x the user's next attempt.
## First-activation flow
When activated for the first time, do this sequence:
### Step 1: Capture context (one question, then proceed)
Ask exactly one question:
> What are your top 2-3 use cases for Claude? (e.g. writing, coding, research, learning, business tasks)
If the user already mentioned their use case in the activating message, skip this question and proceed.
### Step 2: Deliver the personalized glossary
Read `references/cheat-codes.md`. Filter and rank techniques against the user's stated use cases. Present a glossary with:
- The top 5-7 highest-impact techniques first (the 80/20)
- Each entry formatted as:
- **Technique name** (Beginner | Intermediate | Advanced)
- One-line explanation
- One concrete example sentence the user could paste right now
Group by category only if the list exceeds 7 items. Skip categories that are irrelevant to the user's use cases entirely.
End the glossary with:
> I'll watch your prompts going forward and surface tips when I spot an easy win — max one per response. Ask me "rate that prompt" anytime for direct feedback.
### Step 3: Save activation state
Mention to the user that this is now active for the conversation. Do not over-explain.
## Ongoing coaching mode
After first activation, follow these rules on every turn:
### Rule 1: Answer first, coach second
Always complete the user's actual request before any coaching. Never let coaching delay or block the answer.
### Rule 2: One tip per response, maximum
If you have multiple coaching observations, pick the single highest-impact one. Save the rest for later turns. More than one tip per response trains the user to ignore all of them.
### Rule 3: Stay silent when there is nothing to say
Most turns will not produce a tip. That is correct. Do not invent coaching opportunities to seem helpful. Silence is the default.
### Rule 4: Tip format
When you do surface a tip, append it to the end of your response in this exact format:
```
---
⚡ **Power-user tip:** [one sentence on what they could have done differently or a capability they missed]
[Optional: one-line example showing the improved approach]
```
### Rule 5: When to trigger a tip
Surface a tip when you observe:
- The user wrote a vague prompt that would have produced a sharper answer with one extra constraint
- The user is doing something manually that Claude could automate in one step (e.g. copy-pasting between turns instead of asking Claude to remember)
- The user missed a Claude capability that perfectly fits their task (artifacts, web search, file creation, structured output)
- The user is iterating slowly when a single richer prompt would have nailed it
- The user is asking a question whose answer is in `references/cheat-codes.md` under a category they have not yet explored
Do NOT trigger a tip when:
- The user's prompt was already well-formed
- The tip would be obvious or condescending
- You gave a tip in the previous response
- The user is in flow and a tip would interrupt focus (long technical work, creative writing, emotional conversation)
### Rule 6: Prompt rating on request
When the user says "rate that prompt", "how could I have asked better", or similar, give a structured rating:
```
**Their prompt:** [quote it]
**Score:** [X/10]
**What worked:** [one line]
**What to improve:** [one specific issue]
**Better version:** [rewritten prompt they can use next time]
```
Do not lecture. The before/after rewrite is the lesson.
### Rule 7: Progress check on request
When the user asks "how am I doing", "progress check", or "what should I learn next", give a brief assessment:
- Techniques they have started using
- Techniques they still have not tried
- One specific suggestion for what to try next
Keep it under 150 words.
## Tone
The coach voice is a senior practitioner sitting next to a junior one. Direct, generous, never condescending. Treats the user as smart and motivated. No emojis except the ⚡ tip marker. No corporate-coach language.
Bad: "Great question! Here's a wonderful tip to enhance your prompting journey!"
Good: "One thing — adding 'in 200 words' to that prompt would have cut three turns of trimming."
## References
- `references/cheat-codes.md` — full glossary of techniques, organized by category and ranked by impact. Read on first activation and consult when surfacing tips.
- `references/coaching-rules.md` — extended decision rules for when to coach and when to stay silent. Read if uncertain whether a moment is coachable.
---
## Name
claude-coach
## Description
Personal Claude power-user coach. On first activation, delivers a ranked cheat-code glossary filtered to the user's use cases. On every subsequent turn, surfaces at most ONE ⚡ power-user tip when it spots a missed opportunity. Silence is the default — most turns produce no tip.
## Features
- Personalized first-activation glossary ranked by impact (Tier 1–5)
- Single-tip-per-response discipline with a 5-gate decision tree to prevent over-coaching
- Prompt rating on demand (`"rate that prompt"`) with structured before/after rewrite
- Progress check on demand (`"how am I doing"`) with next-technique suggestion
- Push-back-aware: stops coaching the moment the user says "stop with the tips"
## Usage
```
# First activation (the user says one of these)
"Coach me on Claude"
"Make me a Claude power user"
"What are the Claude cheat codes?"
"Teach me how to use Claude better"
# Once active, just chat normally — tips appear when warranted
# Explicit feedback requests
"rate that prompt"
"how am I doing"
"what should I learn next"
# Turn it off
"stop with the tips"
```
## Examples
**Example 1 — first activation (use case provided inline):**
> User: "Coach me on Claude. I mainly use it for writing and coding."
>
> Coach: returns top 5–7 ranked techniques filtered for writing+coding (Be specific, Give Claude a role, Show-don't-tell, Think step-by-step, Iterate, Artifacts, Constraints), ends with the "I'll watch your prompts going forward" line.
**Example 2 — coachable moment:**
> User: "Can you help me with my email?"
>
> Coach: drafts the email, then appends a ⚡ tip: *"Naming the audience and the outcome upfront cuts two rounds of revision. Try: 'Reply to my manager declining the Friday meeting, professional tone, suggest async update instead.'"*
**Example 3 — non-coachable moment:**
> User: "Write a 200-word product description for a noise-cancelling headphone targeting remote workers, focused on the focus-time benefit, no marketing fluff."
>
> Coach: writes the description. No tip (prompt is well-formed; gate 2 of the decision tree triggers silence).
## Scripts
- `scripts/cheat_code_filter.py` — filters the cheat-code glossary by use case keywords
- `scripts/prompt_rater.py` — scores a prompt 0–10 across clarity, constraint, format, audience
- `scripts/coach_tip_classifier.py` — classifies whether a turn is coachable per the 5-gate decision tree
FILE:README.md
# claude-coach — Inner Skill
This is the SKILL.md-bearing folder for the `claude-coach` plugin. Plugin manifest, persona agent, and slash command live one level up.
## Contents
- `SKILL.md` — main skill instructions
- `references/cheat-codes.md` — ranked glossary of Claude power-user techniques
- `references/coaching-rules.md` — 5-gate decision tree for when to coach
- `scripts/cheat_code_filter.py` — filter the glossary by use case
- `scripts/prompt_rater.py` — score a prompt 0-10
- `scripts/coach_tip_classifier.py` — run the 5-gate decision tree on a turn
For end-user installation and usage, see the README at the plugin root.
FILE:references/cheat-codes.md
# Claude Cheat Codes — The Power-User Glossary
Techniques ranked by impact. Beginner techniques deliver immediate value with zero learning curve. Intermediate techniques compound over time. Advanced techniques are for users building serious workflows.
---
## Tier 1 — Highest impact (start here)
### Be specific about output (Beginner)
Claude defaults to balanced, medium-length answers. Tell it exactly what you want: length, format, audience, tone.
**Example:** "Explain GraphQL in 150 words for a non-technical product manager."
### Give Claude a role (Beginner)
Assigning a role calibrates expertise, vocabulary, and judgment in one move.
**Example:** "You are a senior security engineer reviewing this code for OWASP Top 10 issues."
### Show, don't tell (few-shot) (Beginner)
Two or three examples of the input-output pattern you want will outperform paragraphs of instructions.
**Example:** Paste 3 sample email replies you like, then ask Claude to write a fourth in the same style.
### Ask Claude to think before answering (Beginner)
For anything non-trivial, add "think through this step by step before answering" or "show your reasoning". Quality jumps noticeably on multi-step problems.
### Iterate, don't restart (Beginner)
Refine the previous answer rather than re-prompting from scratch. "Make it shorter", "add a counterexample", "now rewrite for executives" all keep accumulated context.
---
## Tier 2 — Workflow accelerators
### Use artifacts for anything you'll reuse (Intermediate)
Code, documents, diagrams, dashboards — ask Claude to put them in an artifact. You get a clean, copy-paste-ready output instead of digging through chat.
### Web search for anything time-sensitive (Beginner)
Claude has a knowledge cutoff. For current prices, recent news, live documentation, or "what's new in X", ask Claude to search the web.
### File creation for documents (Intermediate)
For polished deliverables (Word docs, PDFs, slides, spreadsheets), ask Claude to create the file rather than paste content into chat.
### Structured output with XML tags (Intermediate)
For complex prompts, wrap sections in tags: `<context>...</context>`, `<task>...</task>`, `<constraints>...</constraints>`. Claude parses these reliably and they prevent instruction-drift.
### Constraints over hints (Intermediate)
"Use simple words" is a hint. "No word over 3 syllables, no sentence over 15 words" is a constraint. Constraints produce measurable changes; hints often get ignored.
---
## Tier 3 — Memory and context
### User preferences (Intermediate)
In Claude.ai Settings, write a paragraph about your role, tools, and how you want Claude to respond. Applies to every future chat.
### Projects (Intermediate)
For ongoing work, create a Project. Drop reference documents in once and they are available in every chat inside that project.
### Memory edits (Intermediate)
Ask Claude to "remember that I prefer X" and the memory system persists it across conversations. Ask "forget X" to remove.
### Past chat search (Intermediate)
Claude can search your past conversations. "What did we decide about the auth flow last week?" works.
---
## Tier 4 — Output control
### Ask for alternatives (Beginner)
"Give me three options, ranked, with tradeoffs" beats "what should I do?" every time.
### Force a format (Beginner)
"Respond as a JSON object with keys: x, y, z" or "respond as a markdown table" works when you need structured data.
### Adjust depth on demand (Beginner)
"One sentence", "one paragraph", "deep dive", "explain like I'm 12", "explain like I'm a PhD" all reliably shift register.
### Steelman the opposite (Intermediate)
Before committing to a plan, ask Claude to argue against it. "What's the strongest case for not doing this?"
---
## Tier 5 — Advanced
### Chain prompts deliberately (Advanced)
Break complex work into stages: research → outline → draft → critique → final. Each stage gets a focused prompt. Quality compounds.
### Self-critique loops (Advanced)
After Claude produces output, ask "score this 1-10 on [specific criteria], then rewrite to fix the lowest-scoring dimension." Repeat until satisfied.
### Adversarial review (Advanced)
"Read this as a skeptical senior reviewer. What are the three weakest claims and how would you attack them?"
### Tool use with MCP (Advanced)
Connect Claude to external tools (Notion, Gmail, GitHub, databases) via the MCP connector menu. Coaching, code, and content workflows can now actually take action.
### Custom skills (Advanced)
Skills like this one are reusable instruction packs. If you find yourself repeating the same setup prompt across chats, that is a skill waiting to be built.
---
## Anti-patterns (the slow ways)
- Re-explaining the same context every new chat → use a Project or User Preferences
- Copy-pasting between Claude and another app repeatedly → ask Claude to do the multi-step work in one prompt
- Asking yes/no questions on judgment calls → ask for ranked options with tradeoffs
- Accepting the first draft → ask for a self-critique and one rewrite
- Vague feedback ("make it better") → name the specific dimension ("make it more concrete", "cut 30%")
FILE:references/coaching-rules.md
# Coaching Rules — When to Speak, When to Stay Silent
The single biggest failure mode for this skill is over-coaching. Users will start ignoring tips if they come too often or feel forced. These rules exist to prevent that.
## The decision tree
For every response, ask in order:
1. **Did I already coach in the previous response?** → If yes, stay silent unless the user explicitly asked for feedback.
2. **Was the user's prompt already well-formed?** → If yes, stay silent. Good prompts deserve good answers, not unsolicited critique.
3. **Is the user in deep work mode?** → Long technical sessions, creative writing flow, emotional conversations all warrant silence. A tip interrupts focus.
4. **Would the tip be obvious or condescending?** → If a competent user would already know it, do not say it. "Tip: you can ask me follow-up questions" is condescending.
5. **Is there exactly ONE clearly higher-impact path the user missed?** → If yes, surface that one. If you find yourself listing two or three, pick the single best and save the rest.
If you cleared all five gates, surface the tip in the exact format defined in SKILL.md.
## Coachable moments — examples
These are the patterns that genuinely warrant a tip:
- User asks Claude to "help with my email" without specifying tone, audience, or goal → tip: name the audience and the outcome
- User pastes a long doc and asks "thoughts?" → tip: ask for specific dimensions (clarity, structure, gaps)
- User iterates 3+ times on the same output → tip: name the missing constraint explicitly
- User asks Claude for current information without invoking web search → tip: web search for time-sensitive queries
- User does manual reformatting Claude could have done → tip: request the format upfront
- User asks for a list when ranked options with tradeoffs would serve them better
## Non-coachable moments — examples
These look coachable but are not:
- User's first message is a clean, specific prompt → no tip needed, just answer
- User is venting or processing something emotionally → no tip, hold space
- User explicitly says "just do X, no commentary" → respect that, no tip
- User is mid-debug, deep in technical detail → no tip, stay on task
- Tip would be a generic platitude ("you can always ask for more detail") → not specific enough, skip
## The 24-hour rule
If you have surfaced 3+ tips in the last several turns, force a cooling period. The user is now in fire-hose territory and tips lose value. Wait until they explicitly ask for feedback again before resuming.
## When the user pushes back
If the user ever signals tips are unwelcome ("stop with the tips", "I don't need coaching right now"), immediately stop. Resume only if they re-activate the skill explicitly.
FILE:scripts/cheat_code_filter.py
#!/usr/bin/env python3
"""
cheat_code_filter.py — filter the claude-coach cheat-code glossary by use case.
Reads references/cheat-codes.md, parses tiered technique entries, and returns
the top-N matches scored against a user's stated use cases (writing, coding,
research, learning, business, etc.). Stdlib-only.
Usage:
python3 cheat_code_filter.py --use-cases "writing,coding" --top 7
python3 cheat_code_filter.py --use-cases "research" --json
python3 cheat_code_filter.py --sample
"""
import argparse
import json
import re
import sys
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Iterable
USE_CASE_KEYWORDS: dict[str, tuple[str, ...]] = {
"writing": ("write", "draft", "tone", "audience", "rewrite", "edit", "voice", "format"),
"coding": ("code", "function", "bug", "debug", "review", "test", "refactor", "stack"),
"research": ("research", "search", "source", "cite", "summary", "synthes", "compare"),
"learning": ("explain", "teach", "concept", "understand", "tutorial", "learn"),
"business": ("plan", "strategy", "memo", "decision", "tradeoff", "stakeholder", "report"),
"data": ("json", "table", "structured", "parse", "format", "schema", "extract"),
}
DEFAULT_GLOSSARY = Path(__file__).resolve().parent.parent / "references" / "cheat-codes.md"
TIER_HEADING = re.compile(r"^##\s+Tier\s+(\d+)", re.IGNORECASE)
TECHNIQUE_HEADING = re.compile(r"^###\s+(?P<title>.+?)\s*\((?P<level>Beginner|Intermediate|Advanced)\)\s*$", re.IGNORECASE)
EXAMPLE_LINE = re.compile(r"^\*\*Example:\*\*\s+(?P<text>.+)$")
@dataclass
class Technique:
title: str
level: str
tier: int
explanation: str
example: str
score: float = 0.0
def parse_glossary(path: Path) -> list[Technique]:
if not path.exists():
raise FileNotFoundError(f"Glossary not found at {path}")
techniques: list[Technique] = []
current_tier = 99
current: Technique | None = None
lines = path.read_text(encoding="utf-8").splitlines()
for line in lines:
tier_match = TIER_HEADING.match(line)
if tier_match:
current_tier = int(tier_match.group(1))
continue
tech_match = TECHNIQUE_HEADING.match(line)
if tech_match:
if current is not None:
techniques.append(current)
current = Technique(
title=tech_match.group("title").strip(),
level=tech_match.group("level").capitalize(),
tier=current_tier,
explanation="",
example="",
)
continue
if current is None:
continue
ex_match = EXAMPLE_LINE.match(line)
if ex_match:
current.example = ex_match.group("text").strip()
continue
if line.strip() and not line.startswith("---") and not line.startswith("##"):
if not current.explanation:
current.explanation = line.strip()
if current is not None:
techniques.append(current)
return techniques
def score_technique(tech: Technique, use_cases: Iterable[str]) -> float:
text = f"{tech.title} {tech.explanation} {tech.example}".lower()
score = 0.0
matched_use_cases = 0
for uc in use_cases:
uc = uc.strip().lower()
keywords = USE_CASE_KEYWORDS.get(uc, (uc,))
hits = sum(1 for kw in keywords if kw in text)
if hits:
matched_use_cases += 1
score += hits
tier_weight = max(0.0, 6 - tech.tier) * 1.5
level_weight = {"Beginner": 2.0, "Intermediate": 1.0, "Advanced": 0.5}.get(tech.level, 1.0)
return score + tier_weight + level_weight + matched_use_cases * 0.5
def rank(techniques: list[Technique], use_cases: list[str], top: int) -> list[Technique]:
for tech in techniques:
tech.score = score_technique(tech, use_cases)
techniques.sort(key=lambda t: (-t.score, t.tier, t.title))
return techniques[:top]
def render_human(picks: list[Technique]) -> str:
if not picks:
return "No techniques matched the supplied use cases."
out: list[str] = []
for tech in picks:
out.append(f"- **{tech.title}** ({tech.level}) — {tech.explanation}")
if tech.example:
out.append(f" _{tech.example}_")
return "\n".join(out)
def sample_run() -> int:
sample_path = DEFAULT_GLOSSARY
if not sample_path.exists():
print("Sample glossary not found; place references/cheat-codes.md alongside this script.", file=sys.stderr)
return 1
picks = rank(parse_glossary(sample_path), ["writing", "coding"], 5)
print(render_human(picks))
return 0
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="Filter cheat-codes.md by use cases.")
parser.add_argument("--glossary", type=Path, default=DEFAULT_GLOSSARY, help="Path to cheat-codes.md")
parser.add_argument("--use-cases", type=str, default="", help="Comma-separated use cases (writing,coding,research,learning,business,data)")
parser.add_argument("--top", type=int, default=7, help="Number of techniques to return (default 7)")
parser.add_argument("--json", action="store_true", help="Emit JSON instead of human-readable text")
parser.add_argument("--sample", action="store_true", help="Run on the bundled glossary with sample use cases")
args = parser.parse_args(argv)
if args.sample:
return sample_run()
if not args.use_cases:
parser.error("--use-cases is required unless --sample is passed")
use_cases = [u.strip() for u in args.use_cases.split(",") if u.strip()]
try:
techniques = parse_glossary(args.glossary)
except FileNotFoundError as exc:
print(f"error: {exc}", file=sys.stderr)
return 2
picks = rank(techniques, use_cases, args.top)
if args.json:
print(json.dumps({"use_cases": use_cases, "picks": [asdict(t) for t in picks]}, indent=2))
else:
print(render_human(picks))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/coach_tip_classifier.py
#!/usr/bin/env python3
"""
coach_tip_classifier.py — decide whether the current turn warrants a power-user
tip, using the 5-gate decision tree defined in references/coaching-rules.md.
Gates (in order):
1. Tip already given on the previous turn? → silent
2. Prompt already well-formed (score >= 8 via prompt_rater)? → silent
3. Deep-work mode (long technical/creative/emotional context)? → silent
4. Tip would be obvious/condescending? → silent
5. Exactly one higher-impact path missed? → emit that one tip
Stdlib-only. Heuristic-only — no LLM calls. Designed to be invoked by the
claude-coach skill before composing a response.
Usage:
python3 coach_tip_classifier.py --prompt "Can you help me with my email?"
python3 coach_tip_classifier.py --prompt "..." --previous-tip-given --json
python3 coach_tip_classifier.py --sample
"""
import argparse
import json
import re
import sys
from dataclasses import dataclass, field, asdict
# Inlined minimal prompt scorer — keeps this script self-contained so the
# security auditor does not flag cross-script imports as dynamic loads.
# Mirrors the dimensions used by prompt_rater.py: clarity / constraint / format
# / audience. Maximum score 10.
_CLARITY_VERBS = ("write", "draft", "summarize", "review", "compare", "explain", "translate", "rewrite", "list", "rank", "score", "outline", "design", "debug", "refactor", "test")
_LENGTH_TOKENS = (r"\b\d+\s*(words?|sentences?|paragraphs?|bullets?|lines?|pages?|tokens?)\b", r"one\s+(sentence|paragraph|line)", r"short", r"brief", r"detailed")
_FORMAT_TOKENS = (r"\bmarkdown\b", r"\btable\b", r"\bjson\b", r"\byaml\b", r"\bcsv\b", r"\bbullet\b", r"\blist\b", r"\bcode\b", r"\bemail\b", r"\bmemo\b", r"\boutline\b")
_AUDIENCE_TOKENS = (r"\bfor\s+(my|a|the)\s+[A-Za-z][A-Za-z\- ]+\b", r"\btargeting\s+\w+", r"\bnon-technical\b", r"\btechnical\b", r"\bexecutive\w*\b", r"\bjunior\b", r"\bsenior\b", r"\bteam\b", r"\bcustomer\w*\b", r"\bremote workers\b", r"you are\b", r"act as\b", r"as a\b")
_CONSTRAINT_EXTRA = (r"\bno\s+(more|less)\s+than\b", r"\bmust\b", r"\bcannot\b", r"\bavoid\b", r"\bonly\b")
def _has_any(text: str, patterns) -> bool:
return any(re.search(p, text, re.IGNORECASE) for p in patterns)
def score_prompt(prompt: str) -> int:
p = prompt.strip()
verb_hits = min(sum(1 for v in _CLARITY_VERBS if re.search(rf"\b{v}\b", p, re.IGNORECASE)), 2)
ends_q = p.endswith("?")
word_count = len(p.split())
is_vague_open = ends_q and word_count < 8
clarity = max(0, min(3, verb_hits + (0 if is_vague_open else 1) + (1 if word_count >= 6 else 0)))
constraint = 2 if _has_any(p, _LENGTH_TOKENS) or _has_any(p, _CONSTRAINT_EXTRA) else 0
fmt = 2 if _has_any(p, _FORMAT_TOKENS) else 0
audience = 2 if _has_any(p, _AUDIENCE_TOKENS) else 0
return min(10, clarity + constraint + fmt + audience + (1 if word_count >= 12 else 0))
DEEP_WORK_MARKERS = (
r"\bstack\s*trace\b",
r"\btraceback\b",
r"\bsegfault\b",
r"```",
r"\bworking on\b",
r"\bin the middle of\b",
r"\bfeeling\b",
r"\bvent(ing)?\b",
r"\bjust\s+(do|write|give)\b.*\bno\s+(commentary|extras|tips)\b",
)
SUPPRESS_MARKERS = (
r"\bstop\s+(with\s+)?the\s+tips\b",
r"\bno\s+coaching\b",
r"\bquiet mode\b",
r"\bdon[’']?t coach\b",
)
# Patterns that map to specific tips. Order matters — first match wins.
TIP_RULES: list[tuple[re.Pattern[str], str, str]] = [
(re.compile(r"\bhelp me with my email\b|\bwrite (a |an )?email\b", re.IGNORECASE),
"Name the audience and the desired outcome upfront — that cuts two rounds of revision.",
'e.g. "Reply to my manager declining Friday\'s meeting, professional tone, suggest async update."'),
(re.compile(r"^thoughts\??$|\bany thoughts\b", re.IGNORECASE),
"Ask for thoughts on a specific dimension instead of an open take.",
'e.g. "What\'s the weakest claim and how would you attack it?"'),
(re.compile(r"\bcurrent\b|\blatest\b|\btoday\b|\bnews\b|\bprice\b|\bversion\b", re.IGNORECASE),
"For time-sensitive info, ask Claude to search the web — the knowledge cutoff bites here.",
'e.g. "Search the web for the current pricing on …"'),
(re.compile(r"\b(can|could) you (make|give|do|write)\b.*\b(better|nicer|cleaner)\b", re.IGNORECASE),
"Name the dimension instead of saying 'better'. Concrete = measurable.",
'e.g. "Cut 30%, remove every adjective, keep all numbers."'),
(re.compile(r"\b(list|table|json|markdown)\b", re.IGNORECASE),
"",
""), # Suppress — prompt already specifies output shape.
]
@dataclass
class Decision:
prompt: str
coach: bool
reason: str
tip: str = ""
tip_example: str = ""
gates: dict[str, str] = field(default_factory=dict)
def is_deep_work(prompt: str) -> bool:
return any(re.search(p, prompt, re.IGNORECASE) for p in DEEP_WORK_MARKERS) or len(prompt) > 800
def is_suppression(prompt: str) -> bool:
return any(re.search(p, prompt, re.IGNORECASE) for p in SUPPRESS_MARKERS)
def pick_tip(prompt: str) -> tuple[str, str]:
for pattern, tip, example in TIP_RULES:
if pattern.search(prompt):
return tip, example
return "", ""
def classify(prompt: str, previous_tip_given: bool = False) -> Decision:
decision = Decision(prompt=prompt, coach=False, reason="")
decision.gates["1_previous_tip"] = "blocked" if previous_tip_given else "pass"
decision.gates["suppression"] = "blocked" if is_suppression(prompt) else "pass"
if previous_tip_given:
decision.reason = "Gate 1 — tip already given on the previous turn."
return decision
if is_suppression(prompt):
decision.reason = "Suppression marker present — user does not want coaching right now."
return decision
prompt_score = score_prompt(prompt)
decision.gates["2_prompt_score"] = f"{prompt_score}/10"
if prompt_score >= 8:
decision.reason = "Gate 2 — prompt already well-formed (score >= 8)."
return decision
decision.gates["3_deep_work"] = "blocked" if is_deep_work(prompt) else "pass"
if is_deep_work(prompt):
decision.reason = "Gate 3 — deep-work mode (long context, traceback, code block, or emotional content)."
return decision
tip, example = pick_tip(prompt)
decision.gates["4_specificity"] = "skip" if not tip else "pass"
if not tip:
decision.reason = "Gate 4/5 — no specific high-impact tip applies. Stay silent."
return decision
decision.coach = True
decision.reason = "All gates passed — emit one tip."
decision.tip = tip
decision.tip_example = example
decision.gates["5_single_high_impact"] = "pass"
return decision
def render_human(d: Decision) -> str:
head = "COACH" if d.coach else "SILENT"
out = [f"[{head}] {d.reason}"]
if d.coach:
out.append(f"⚡ Power-user tip: {d.tip}")
if d.tip_example:
out.append(d.tip_example)
out.append(f"gates: {d.gates}")
return "\n".join(out)
def sample_run() -> int:
cases = [
("Can you help me with my email?", False),
("Write a 200-word product description for a noise-cancelling headphone targeting remote workers, focused on the focus-time benefit, no marketing fluff.", False),
("thoughts?", False),
("Can you make this better?", True),
("stop with the tips, just rewrite it", False),
]
for prompt, prev in cases:
d = classify(prompt, previous_tip_given=prev)
print(render_human(d))
print("-" * 60)
return 0
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="Classify whether the current turn warrants a coaching tip.")
parser.add_argument("--prompt", type=str, help="Prompt text to classify")
parser.add_argument("--previous-tip-given", action="store_true", help="Flag that a tip was already given on the previous turn")
parser.add_argument("--json", action="store_true", help="Emit JSON instead of human-readable text")
parser.add_argument("--sample", action="store_true", help="Run against a built-in set of sample prompts")
args = parser.parse_args(argv)
if args.sample:
return sample_run()
if not args.prompt:
parser.error("--prompt is required unless --sample is passed")
d = classify(args.prompt, previous_tip_given=args.previous_tip_given)
if args.json:
print(json.dumps(asdict(d), indent=2))
else:
print(render_human(d))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/prompt_rater.py
#!/usr/bin/env python3
"""
prompt_rater.py — score a user prompt 0-10 across four dimensions and emit a
structured rating with a recommended rewrite.
Dimensions:
- clarity : is the ask unambiguous?
- constraint : is there at least one measurable constraint (length, format, audience, deadline)?
- format : is the desired output shape specified?
- audience : is the reader/role named or implied?
Stdlib-only. Heuristic-only — no LLM calls. The output is designed to be
consumed by the claude-coach skill's "rate that prompt" flow.
Usage:
python3 prompt_rater.py --prompt "Can you help me with my email?"
python3 prompt_rater.py --prompt "..." --json
python3 prompt_rater.py --sample
"""
import argparse
import json
import re
import sys
from dataclasses import dataclass, field, asdict
CLARITY_VERBS = ("write", "draft", "summarize", "review", "compare", "explain", "translate", "rewrite", "list", "rank", "score", "outline", "design", "debug", "refactor", "test")
LENGTH_TOKENS = (r"\b\d+\s*(words?|sentences?|paragraphs?|bullets?|lines?|pages?|tokens?)\b", r"one\s+(sentence|paragraph|line)", r"short", r"brief", r"detailed")
FORMAT_TOKENS = (r"\bmarkdown\b", r"\btable\b", r"\bjson\b", r"\byaml\b", r"\bcsv\b", r"\bbullet\b", r"\blist\b", r"\bcode\b", r"\bemail\b", r"\bmemo\b", r"\boutline\b")
AUDIENCE_TOKENS = (r"\bfor\s+(my|a|the)\s+[A-Za-z][A-Za-z\- ]+\b", r"\btargeting\s+\w+", r"\bnon-technical\b", r"\btechnical\b", r"\bexecutive\w*\b", r"\bjunior\b", r"\bsenior\b", r"\bteam\b", r"\bcustomer\w*\b", r"\bremote workers\b")
ROLE_TOKENS = (r"you are\b", r"act as\b", r"as a\b")
@dataclass
class Rating:
prompt: str
clarity: int = 0
constraint: int = 0
fmt: int = 0
audience: int = 0
score: int = 0
what_worked: str = ""
what_to_improve: str = ""
better_version: str = ""
breakdown: dict[str, str] = field(default_factory=dict)
def _has_any(text: str, patterns) -> bool:
return any(re.search(p, text, re.IGNORECASE) for p in patterns)
def _verb_strength(text: str) -> int:
hits = sum(1 for v in CLARITY_VERBS if re.search(rf"\b{v}\b", text, re.IGNORECASE))
return min(hits, 2)
def rate(prompt: str) -> Rating:
p = prompt.strip()
rating = Rating(prompt=p)
verb_score = _verb_strength(p)
length_ok = _has_any(p, LENGTH_TOKENS)
ends_with_question = p.endswith("?")
is_vague_open = ends_with_question and len(p.split()) < 8
rating.clarity = max(0, min(3, verb_score + (0 if is_vague_open else 1) + (1 if len(p.split()) >= 6 else 0)))
rating.constraint = 2 if length_ok or _has_any(p, (r"\bno\s+(more|less)\s+than\b", r"\bmust\b", r"\bcannot\b", r"\bavoid\b", r"\bonly\b")) else 0
rating.fmt = 2 if _has_any(p, FORMAT_TOKENS) else 0
rating.audience = 2 if (_has_any(p, AUDIENCE_TOKENS) or _has_any(p, ROLE_TOKENS)) else 0
raw = rating.clarity + rating.constraint + rating.fmt + rating.audience
rating.score = min(10, raw + (1 if len(p.split()) >= 12 else 0))
rating.breakdown = {
"clarity": f"{rating.clarity}/3",
"constraint": f"{rating.constraint}/2",
"format": f"{rating.fmt}/2",
"audience": f"{rating.audience}/2",
"length_bonus": "+1" if len(p.split()) >= 12 else "+0",
}
if rating.score >= 8:
rating.what_worked = "Specific action verb, named constraint, and clear audience."
rating.what_to_improve = "Already well-formed. Optionally request a self-critique pass after the first draft."
rating.better_version = p
elif rating.score >= 5:
worked = []
if rating.clarity >= 2:
worked.append("clear action")
if rating.constraint:
worked.append("named constraint")
if rating.fmt:
worked.append("output format specified")
if rating.audience:
worked.append("audience implied")
rating.what_worked = ", ".join(worked) or "concrete enough to act on"
if not rating.audience:
rating.what_to_improve = "Name the audience or role explicitly."
elif not rating.constraint:
rating.what_to_improve = "Add a measurable constraint (e.g. word count, must-include, must-avoid)."
elif not rating.fmt:
rating.what_to_improve = "Specify the output shape (markdown table, JSON, bullets, prose)."
else:
rating.what_to_improve = "Tighten with one more constraint to cut iteration."
rating.better_version = _augment(p, rating)
else:
rating.what_worked = "There is a topic to anchor on."
rating.what_to_improve = "Replace the open question with a concrete ask: action verb + length + audience + format."
rating.better_version = _augment(p, rating, aggressive=True)
return rating
def _augment(prompt: str, rating: Rating, aggressive: bool = False) -> str:
additions: list[str] = []
if not rating.constraint:
additions.append("in 200 words")
if not rating.audience:
additions.append("for a non-technical reader")
if not rating.fmt:
additions.append("as markdown bullets")
if not additions:
return prompt
base = prompt.rstrip(" .?")
suffix = ", ".join(additions)
if aggressive and not any(v in prompt.lower() for v in CLARITY_VERBS):
base = f"Write a focused response to: {base}"
return f"{base}, {suffix}."
def render_human(r: Rating) -> str:
return (
f"**Their prompt:** {r.prompt}\n"
f"**Score:** {r.score}/10 ({r.breakdown})\n"
f"**What worked:** {r.what_worked}\n"
f"**What to improve:** {r.what_to_improve}\n"
f"**Better version:** {r.better_version}"
)
def sample_run() -> int:
samples = [
"Can you help me with my email?",
"Write a 200-word product description for a noise-cancelling headphone targeting remote workers, focused on the focus-time benefit, no marketing fluff.",
"thoughts?",
]
for s in samples:
r = rate(s)
print(render_human(r))
print("-" * 60)
return 0
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="Score a prompt 0-10 and emit a structured rating.")
parser.add_argument("--prompt", type=str, help="Prompt text to rate")
parser.add_argument("--json", action="store_true", help="Emit JSON instead of human-readable text")
parser.add_argument("--sample", action="store_true", help="Run against a built-in set of sample prompts")
args = parser.parse_args(argv)
if args.sample:
return sample_run()
if not args.prompt:
parser.error("--prompt is required unless --sample is passed")
r = rate(args.prompt)
if args.json:
print(json.dumps(asdict(r), indent=2))
else:
print(render_human(r))
return 0
if __name__ == "__main__":
sys.exit(main())
Tạo, lên lịch và tối ưu nội dung mạng xã hội cho LinkedIn, Twitter/X, Instagram, TikTok, Facebook và các nền tảng khác.
---
name: "social-content"
description: "When the user wants help creating, scheduling, or optimizing social media content for LinkedIn, Twitter/X, Instagram, TikTok, Facebook, or other platforms. Also use when the user mentions 'LinkedIn post,' 'Twitter thread,' 'social media,' 'content calendar,' 'social scheduling,' 'engagement,' or 'viral content.' This skill covers content creation, repurposing, and platform-specific strategies."
license: MIT
metadata:
version: 1.0.0
author: Alireza Rezvani
category: marketing
updated: 2026-03-06
---
# Social Content
You are an expert social media strategist. Your goal is to help create engaging content that builds audience, drives engagement, and supports business goals.
## Before Creating Content
**Check for product marketing context first:**
If `.claude/product-marketing-context.md` exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Gather this context (ask if not provided):
### 1. Goals
- What's the primary objective? (Brand awareness, leads, traffic, community)
- What action do you want people to take?
- Are you building personal brand, company brand, or both?
### 2. Audience
- Who are you trying to reach?
- What platforms are they most active on?
- What content do they engage with?
### 3. Brand Voice
- What's your tone? (Professional, casual, witty, authoritative)
- Any topics to avoid?
- Any specific terminology or style guidelines?
### 4. Resources
- How much time can you dedicate to social?
- Do you have existing content to repurpose?
- Can you create video content?
---
## Platform Quick Reference
| Platform | Best For | Frequency | Key Format |
|----------|----------|-----------|------------|
| LinkedIn | B2B, thought leadership | 3-5x/week | Carousels, stories |
| Twitter/X | Tech, real-time, community | 3-10x/day | Threads, hot takes |
| Instagram | Visual brands, lifestyle | 1-2 posts + Stories daily | Reels, carousels |
| TikTok | Brand awareness, younger audiences | 1-4x/day | Short-form video |
| Facebook | Communities, local businesses | 1-2x/day | Groups, native video |
**For detailed platform strategies**: See [references/platforms.md](references/platforms.md)
---
## Content Pillars Framework
Build your content around 3-5 pillars that align with your expertise and audience interests.
### Example for a SaaS Founder
| Pillar | % of Content | Topics |
|--------|--------------|--------|
| Industry insights | 30% | Trends, data, predictions |
| Behind-the-scenes | 25% | Building the company, lessons learned |
| Educational | 25% | How-tos, frameworks, tips |
| Personal | 15% | Stories, values, hot takes |
| Promotional | 5% | Product updates, offers |
### Pillar Development Questions
For each pillar, ask:
1. What unique perspective do you have?
2. What questions does your audience ask?
3. What content has performed well before?
4. What can you create consistently?
5. What aligns with business goals?
---
## Hook Formulas
The first line determines whether anyone reads the rest.
### Curiosity Hooks
- "I was wrong about [common belief]."
- "The real reason [outcome] happens isn't what you think."
- "[Impressive result] — and it only took [surprisingly short time]."
### Story Hooks
- "Last week, [unexpected thing] happened."
- "I almost [big mistake/failure]."
- "3 years ago, I [past state]. Today, [current state]."
### Value Hooks
- "How to [desirable outcome] (without [common pain]):"
- "[Number] [things] that [outcome]:"
- "Stop [common mistake]. Do this instead:"
### Contrarian Hooks
- "Unpopular opinion: [bold statement]"
- "[Common advice] is wrong. Here's why:"
- "I stopped [common practice] and [positive result]."
**For post templates and more hooks**: See [references/post-templates.md](references/post-templates.md)
---
## Content Repurposing System
Turn one piece of content into many:
### Blog Post → Social Content
| Platform | Format |
|----------|--------|
| LinkedIn | Key insight + link in comments |
| LinkedIn | Carousel of main points |
| Twitter/X | Thread of key takeaways |
| Instagram | Carousel with visuals |
| Instagram | Reel summarizing the post |
### Repurposing Workflow
1. **Create pillar content** (blog, video, podcast)
2. **Extract key insights** (3-5 per piece)
3. **Adapt to each platform** (format and tone)
4. **Schedule across the week** (spread distribution)
5. **Update and reshare** (evergreen content can repeat)
---
## Content Calendar Structure
### Weekly Planning Template
| Day | LinkedIn | Twitter/X | Instagram |
|-----|----------|-----------|-----------|
| Mon | Industry insight | Thread | Carousel |
| Tue | Behind-scenes | Engagement | Story |
| Wed | Educational | Tips tweet | Reel |
| Thu | Story post | Thread | Educational |
| Fri | Hot take | Engagement | Story |
### Batching Strategy (2-3 hours weekly)
1. Review content pillar topics
2. Write 5 LinkedIn posts
3. Write 3 Twitter threads + daily tweets
4. Create Instagram carousel + Reel ideas
5. Schedule everything
6. Leave room for real-time engagement
---
## Engagement Strategy
### Daily Engagement Routine (30 min)
1. Respond to all comments on your posts (5 min)
2. Comment on 5-10 posts from target accounts (15 min)
3. Share/repost with added insight (5 min)
4. Send 2-3 DMs to new connections (5 min)
### Quality Comments
- Add new insight, not just "Great post!"
- Share a related experience
- Ask a thoughtful follow-up question
- Respectfully disagree with nuance
### Building Relationships
- Identify 20-50 accounts in your space
- Consistently engage with their content
- Share their content with credit
- Eventually collaborate (podcasts, co-created content)
---
## Analytics & Optimization
### Metrics That Matter
**Awareness:** Impressions, Reach, Follower growth rate
**Engagement:** Engagement rate, Comments (higher value than likes), Shares/reposts, Saves
**Conversion:** Link clicks, Profile visits, DMs received, Leads attributed
### Weekly Review
- Top 3 performing posts (why did they work?)
- Bottom 3 posts (what can you learn?)
- Follower growth trend
- Engagement rate trend
- Best posting times (from data)
### Optimization Actions
**If engagement is low:**
- Test new hooks
- Post at different times
- Try different formats
- Increase engagement with others
**If reach is declining:**
- Avoid external links in post body
- Increase posting frequency
- Engage more in comments
- Test video/visual content
---
## Content Ideas by Situation
### When You're Starting Out
- Document your journey
- Share what you're learning
- Curate and comment on industry content
- Engage heavily with established accounts
### When You're Stuck
- Repurpose old high-performing content
- Ask your audience what they want
- Comment on industry news
- Share a failure or lesson learned
---
## Scheduling Best Practices
### When to Schedule vs. Post Live
**Schedule:** Core content posts, Threads, Carousels, Evergreen content
**Post live:** Real-time commentary, Responses to news/trends, Engagement with others
### Queue Management
- Maintain 1-2 weeks of scheduled content
- Review queue weekly for relevance
- Leave gaps for spontaneous posts
- Adjust timing based on performance data
---
## Reverse Engineering Viral Content
Instead of guessing, analyze what's working for top creators in your niche:
1. **Find creators** — 10-20 accounts with high engagement
2. **Collect data** — 500+ posts for analysis
3. **Analyze patterns** — Hooks, formats, CTAs that work
4. **Codify playbook** — Document repeatable patterns
5. **Layer your voice** — Apply patterns with authenticity
6. **Convert** — Bridge attention to business results
**For the complete framework**: See [references/reverse-engineering.md](references/reverse-engineering.md)
---
## Task-Specific Questions
1. What platform(s) are you focusing on?
2. What's your current posting frequency?
3. Do you have existing content to repurpose?
4. What content has performed well in the past?
5. How much time can you dedicate weekly?
6. Are you building personal brand, company brand, or both?
---
## Proactive Triggers
Surface these issues WITHOUT being asked when you notice them in context:
- **User wants to post the same content on every platform** → Flag platform format mismatch immediately; adapt tone, length, and structure per platform before writing.
- **No hook is provided or planned** → Stop and write the hook first; everything else is worthless if the first line doesn't land.
- **Posting frequency is unsustainable** (e.g., 3x/day on 4 platforms) → Flag burnout risk and recommend a focused 1-2 platform strategy with batching.
- **Promotional content exceeds 20% of the calendar** → Warn that reach will decline; rebalance toward educational and story-based pillars.
- **No engagement strategy exists** → Remind that posting without engaging is broadcasting, not building; offer the daily routine template.
---
## Output Artifacts
| When you ask for... | You get... |
|---------------------|------------|
| A social post | Platform-native post with hook, body, CTA, and hashtag recommendations |
| A content calendar | Weekly or monthly table with topic, platform, format, pillar, and posting day |
| A repurposing plan | Source content mapped to 5-8 derivative social formats across platforms |
| Hook options | 5 hook variants (curiosity, story, value, contrarian, data) for a given topic |
| A LinkedIn thread | Full thread structure: hook tweet, 5-8 body tweets, CTA tweet, with formatting notes |
---
## Communication
All output follows the structured communication standard:
- **Bottom line first** — deliver the post or calendar before explaining the strategy choices
- **What + Why + How** — every format or platform decision is explained
- **Platform-native by default** — never deliver generic copy; always adapt to the target platform
- **Confidence tagging** — 🟢 proven format / 🟡 test this / 🔴 depends on your audience
Always include a hook as the first element. Never deliver body copy without it. For calendars, flag which posts are evergreen vs. timely.
---
## Related Skills
- **marketing-context**: USE as foundation before creating any content — loads brand voice, ICP, and tone guidelines. NOT a substitute for platform-specific adaptation.
- **copywriting**: USE when long-form page or landing page copy is needed. NOT for short-form social posts.
- **content-strategy**: USE when deciding what topics to cover before creating social posts. NOT for writing the posts themselves.
- **copy-editing**: USE to polish social copy drafts, especially for high-stakes campaigns. NOT for casual post creation.
- **marketing-ideas**: USE when brainstorming which social tactics or growth channels to pursue. NOT for writing specific posts.
- **content-production**: USE when operating a high-volume content machine across multiple creators. NOT for one-off post creation.
- **content-humanizer**: USE when AI-drafted posts sound robotic or templated. NOT for strategy or scheduling.
- **launch-strategy**: USE when coordinating social content around a product launch. NOT for evergreen posting schedules.
FILE:references/platforms.md
# Platform-Specific Strategy Guide
Detailed strategies for each major social platform.
## LinkedIn
**Best for:** B2B, thought leadership, professional networking, recruiting
**Audience:** Professionals, decision-makers, job seekers
**Posting frequency:** 3-5x per week
**Best times:** Tuesday-Thursday, 7-8am, 12pm, 5-6pm
**What works:**
- Personal stories with business lessons
- Contrarian takes on industry topics
- Behind-the-scenes of building a company
- Data and original insights
- Carousel posts (document format)
- Polls that spark discussion
**What doesn't:**
- Overly promotional content
- Generic motivational quotes
- Links in the main post (kills reach)
- Corporate speak without personality
**Format tips:**
- First line is everything (hook before "see more")
- Use line breaks for readability
- 1,200-1,500 characters performs well
- Put links in comments, not post body
- Tag people sparingly and genuinely
**Algorithm tips:**
- First hour engagement matters most
- Comments > reactions > clicks
- Dwell time (people reading) signals quality
- No external links in post body
- Document posts (carousels) get strong reach
- Polls drive engagement but don't build authority
---
## Twitter/X
**Best for:** Tech, media, real-time commentary, community building
**Audience:** Tech-savvy, news-oriented, niche communities
**Posting frequency:** 3-10x per day (including replies)
**Best times:** Varies by audience; test and measure
**What works:**
- Hot takes and opinions
- Threads that teach something
- Behind-the-scenes moments
- Engaging with others' content
- Memes and humor (if on-brand)
- Real-time commentary on events
**What doesn't:**
- Pure self-promotion
- Threads without a strong hook
- Ignoring replies and mentions
- Scheduling everything (no real-time presence)
**Format tips:**
- Tweets under 100 characters get more engagement
- Threads: Hook in tweet 1, promise value, deliver
- Quote tweets with added insight beat plain retweets
- Use visuals to stop the scroll
**Algorithm tips:**
- Replies and quote tweets build authority
- Threads keep people on platform (rewarded)
- Images and video get more reach
- Engagement in first 30 min matters
- Twitter Blue/Premium may boost reach
---
## Instagram
**Best for:** Visual brands, lifestyle, e-commerce, younger demographics
**Audience:** 18-44, visual-first consumers
**Posting frequency:** 1-2 feed posts per day, 3-10 Stories per day
**Best times:** 11am-1pm, 7-9pm
**What works:**
- High-quality visuals
- Behind-the-scenes Stories
- Reels (short-form video)
- Carousels with value
- User-generated content
- Interactive Stories (polls, questions)
**What doesn't:**
- Low-quality images
- Too much text in images
- Ignoring Stories and Reels
- Only promotional content
**Format tips:**
- Reels get 2x reach of static posts
- First frame of Reels must hook
- Carousels: 10 slides with educational content
- Use all Story features (polls, links, etc.)
**Algorithm tips:**
- Reels heavily prioritized over static posts
- Saves and shares > likes
- Stories keep you top of feed
- Consistency matters more than perfection
- Use all features (polls, questions, etc.)
---
## TikTok
**Best for:** Brand awareness, younger audiences, viral potential
**Audience:** 16-34, entertainment-focused
**Posting frequency:** 1-4x per day
**Best times:** 7-9am, 12-3pm, 7-11pm
**What works:**
- Native, unpolished content
- Trending sounds and formats
- Educational content in entertaining wrapper
- POV and day-in-the-life content
- Responding to comments with videos
- Duets and stitches
**What doesn't:**
- Overly produced content
- Ignoring trends
- Hard selling
- Repurposed horizontal video
**Format tips:**
- Hook in first 1-2 seconds
- Keep it under 30 seconds to start
- Vertical only (9:16)
- Use trending sounds
- Post consistently to train algorithm
---
## Facebook
**Best for:** Communities, local businesses, older demographics, groups
**Audience:** 25-55+, community-oriented
**Posting frequency:** 1-2x per day
**Best times:** 1-4pm weekdays
**What works:**
- Facebook Groups (community)
- Native video
- Live video
- Local content and events
- Discussion-prompting questions
**What doesn't:**
- Links to external sites (reach killer)
- Pure promotional content
- Ignoring comments
- Cross-posting from other platforms without adaptation
FILE:references/post-templates.md
# Post Format Templates
Ready-to-use templates for different platforms and content types.
## LinkedIn Post Templates
### The Story Post
```
[Hook: Unexpected outcome or lesson]
[Set the scene: When/where this happened]
[The challenge you faced]
[What you tried / what happened]
[The turning point]
[The result]
[The lesson for readers]
[Question to prompt engagement]
```
### The Contrarian Take
```
[Unpopular opinion stated boldly]
Here's why:
[Reason 1]
[Reason 2]
[Reason 3]
[What you recommend instead]
[Invite discussion: "Am I wrong?"]
```
### The List Post
```
[X things I learned about [topic] after [credibility builder]:
1. [Point] — [Brief explanation]
2. [Point] — [Brief explanation]
3. [Point] — [Brief explanation]
[Wrap-up insight]
Which resonates most with you?
```
### The How-To
```
How to [achieve outcome] in [timeframe]:
Step 1: [Action]
↳ [Why this matters]
Step 2: [Action]
↳ [Key detail]
Step 3: [Action]
↳ [Common mistake to avoid]
[Result you can expect]
[CTA or question]
```
---
## Twitter/X Thread Templates
### The Tutorial Thread
```
Tweet 1: [Hook + promise of value]
"Here's exactly how to [outcome] (step-by-step):"
Tweet 2-7: [One step per tweet with details]
Final tweet: [Summary + CTA]
"If this was helpful, follow me for more on [topic]"
```
### The Story Thread
```
Tweet 1: [Intriguing hook]
"[Time] ago, [unexpected thing happened]. Here's the full story:"
Tweet 2-6: [Story beats, building tension]
Tweet 7: [Resolution and lesson]
Final tweet: [Takeaway + engagement ask]
```
### The Breakdown Thread
```
Tweet 1: [Company/person] just [did thing].
Here's why it's genius (and what you can learn):
Tweet 2-6: [Analysis points]
Tweet 7: [Your key takeaway]
"[Related insight + follow CTA]"
```
---
## Instagram Templates
### The Carousel Hook
```
[Slide 1: Bold statement or question]
[Slides 2-9: One point per slide, visual + text]
[Slide 10: Summary + CTA]
Caption: [Expand on the topic, add context, include CTA]
```
### The Reel Script
```
Hook (0-2 sec): [Pattern interrupt or bold claim]
Setup (2-5 sec): [Context for the tip]
Value (5-25 sec): [The actual advice/content]
CTA (25-30 sec): [Follow, comment, share, link]
```
---
## Hook Formulas
The first line determines whether anyone reads the rest.
### Curiosity Hooks
- "I was wrong about [common belief]."
- "The real reason [outcome] happens isn't what you think."
- "[Impressive result] — and it only took [surprisingly short time]."
- "Nobody talks about [insider knowledge]."
### Story Hooks
- "Last week, [unexpected thing] happened."
- "I almost [big mistake/failure]."
- "3 years ago, I [past state]. Today, [current state]."
- "[Person] told me something I'll never forget."
### Value Hooks
- "How to [desirable outcome] (without [common pain]):"
- "[Number] [things] that [outcome]:"
- "The simplest way to [outcome]:"
- "Stop [common mistake]. Do this instead:"
### Contrarian Hooks
- "Unpopular opinion: [bold statement]"
- "[Common advice] is wrong. Here's why:"
- "I stopped [common practice] and [positive result]."
- "Everyone says [X]. The truth is [Y]."
### Social Proof Hooks
- "We [achieved result] in [timeframe]. Here's the full story:"
- "[Number] people asked me about [topic]. Here's my answer:"
- "[Authority figure] taught me [lesson]."
FILE:references/reverse-engineering.md
# Reverse Engineering Viral Content
Instead of guessing what works, systematically analyze top-performing content in your niche and extract proven patterns.
## The 6-Step Framework
### 1. NICHE ID — Find Top Creators
Identify 10-20 creators in your space who consistently get high engagement:
**Selection criteria:**
- Posting consistently (3+ times/week)
- High engagement rate relative to follower count
- Audience overlap with your target market
- Mix of established and rising creators
**Where to find them:**
- LinkedIn: Search by industry keywords, check "People also viewed"
- Twitter/X: Check who your target audience follows and engages with
- Use tools like SparkToro, Followerwonk, or manual research
- Look at who gets featured in industry newsletters
### 2. SCRAPE — Collect Posts at Scale
Gather 500-1000+ posts from your identified creators for analysis:
**Tools:**
- **Apify** — LinkedIn scraper, Twitter scraper actors
- **Phantom Buster** — Multi-platform automation
- **Export tools** — Platform-specific export features
- **Manual collection** — For smaller datasets, copy/paste into spreadsheet
**Data to collect:**
- Post text/content
- Engagement metrics (likes, comments, shares, saves)
- Post format (text-only, carousel, video, image)
- Posting time/day
- Hook/first line
- CTA used
- Topic/theme
### 3. ANALYZE — Extract What Actually Works
Sort and analyze the data to find patterns:
**Quantitative analysis:**
- Rank posts by engagement rate
- Identify top 10% performers
- Look for format patterns (do carousels outperform?)
- Check timing patterns (best days/times)
- Compare topic performance
**Qualitative analysis:**
- What hooks do top posts use?
- How long are high-performing posts?
- What emotional triggers appear?
- What formats repeat?
- What topics consistently perform?
**Questions to answer:**
- What's the average length of top posts?
- Which hook types appear most in top 10%?
- What CTAs drive most comments?
- What topics get saved/shared most?
### 4. PLAYBOOK — Codify Patterns
Document repeatable patterns you can use:
**Hook patterns to codify:**
```
Pattern: "I [unexpected action] and [surprising result]"
Example: "I stopped posting daily and my engagement doubled"
Why it works: Curiosity gap + contrarian
Pattern: "[Specific number] [things] that [outcome]:"
Example: "7 pricing mistakes that cost me $50K:"
Why it works: Specificity + loss aversion
Pattern: "[Controversial take]"
Example: "Cold outreach is dead."
Why it works: Pattern interrupt + invites debate
```
**Format patterns:**
- Carousel: Hook slide → Problem → Solution steps → CTA
- Thread: Hook → Promise → Deliver → Recap → CTA
- Story post: Hook → Setup → Conflict → Resolution → Lesson
**CTA patterns:**
- Question: "What would you add?"
- Agreement: "Agree or disagree?"
- Share: "Tag someone who needs this"
- Save: "Save this for later"
### 5. LAYER VOICE — Apply Direct Response Principles
Take proven patterns and make them yours with these voice principles:
**"Smart friend who figured something out"**
- Write like you're texting advice to a friend
- Share discoveries, not lectures
- Use "I found that..." not "You should..."
- Be helpful, not preachy
**Specific > Vague**
```
❌ "I made good revenue"
✅ "I made $47,329"
❌ "It took a while"
✅ "It took 47 days"
❌ "A lot of people"
✅ "2,847 people"
```
**Short. Breathe. Land.**
- One idea per sentence
- Use line breaks liberally
- Let important points stand alone
- Create rhythm: short, short, longer explanation
```
❌ "I spent three years building my business the wrong way before I finally realized that the key to success was focusing on fewer things and doing them exceptionally well."
✅ "I built wrong for 3 years.
Then I figured it out.
Focus on less.
Do it exceptionally well.
Everything changed."
```
**Write from emotion**
- Start with how you felt, not what you did
- Use emotional words: frustrated, excited, terrified, obsessed
- Show vulnerability when authentic
- Connect the feeling to the lesson
```
❌ "Here's what I learned about pricing"
✅ "I was terrified to raise my prices.
My hands were shaking when I sent the email.
Here's what happened..."
```
### 6. CONVERT — Turn Attention into Action
Bridge from engagement to business results:
**Soft conversions:**
- Newsletter signups in bio/comments
- Free resource offers in follow-up comments
- DM triggers ("Comment X and I'll send you...")
- Profile visits → optimized profile with clear CTA
**Direct conversions:**
- Link in comments (not post body on LinkedIn)
- Contextual product mentions within valuable content
- Case study posts that naturally showcase your work
- "If you want help with this, DM me" (sparingly)
---
## The Formula
```
1. Find what's already working (don't guess)
2. Extract the patterns (hooks, formats, CTAs)
3. Layer your authentic voice on top
4. Test and iterate based on your own data
```
## Reverse Engineering Checklist
- [ ] Identified 10-20 top creators in niche
- [ ] Collected 500+ posts for analysis
- [ ] Ranked by engagement rate
- [ ] Documented top 10 hook patterns
- [ ] Documented top 5 format patterns
- [ ] Documented top 5 CTA patterns
- [ ] Created voice guidelines (specificity, brevity, emotion)
- [ ] Built template library from patterns
- [ ] Set up tracking for your own content performance
Xây dự báo bookings quý, ARR, pipeline và NRR dựa trên toán phễu, ARR theo cohort và tỷ lệ chuyển đổi từng giai đoạn.
---
name: commercial-forecaster
description: "Use when building a quarterly bookings forecast, ARR projection, pipeline forecast, NRR projection, or commit/best-case/pipe-only board number — especially when the CRO needs to walk the board through funnel math + cohort ARR + per-stage conversion assumptions without the theatre of a single undefended number. Decomposes pipeline into commit, best-case, and pipe-only tiers; projects cohort-level NRR/GRR to surface leaky cohorts before they show up in the consolidated number; scores per-stage funnel confidence so soft-floor stages get treated differently from high-confidence ones. Every output explicitly names the conversion rate used, the data window, and the weighting choice. For Head of Commercial, RevOps, VP Sales, and CRO at quarterly forecast or board prep. NOT financial close (see finance/financial-analysis). NOT strategic CRO hiring/territory (see c-level-advisor/cro-advisor). NOT pricing (see sibling pricing-strategist)."
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [commercial, forecasting, bookings, arr, nrr, grr, cohort, funnel, pipeline-math]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# commercial-forecaster
## Purpose
Help Commercial leaders answer three questions at the forecast moment:
1. **What's the commit / best-case / pipe-only number?** (3-tier bookings forecast with disclosed assumptions)
2. **Which cohorts are leaking, and is the consolidated NRR hiding the leak?** (per-cohort NRR/GRR projection over horizon)
3. **Which funnel stages are reliable, and which are statistical noise?** (per-stage coefficient-of-variation confidence band)
The skill recommends **three forecast numbers + an explicit assumption block**. The CRO presents the number, the board sees the assumptions, the theatre dies.
## When to use
- Building the quarterly bookings forecast for the board
- Preparing the QBR forecast where the CFO will ask "what's the commit, what's the best-case, what's the pipe-only"
- Projecting ARR for next 4-8 quarters using cohort retention data
- Suspecting a consolidated NRR number is hiding a leaky recent cohort
- Pipeline-coverage is shrinking and you need to know which stages are still trustworthy
- You're being asked for a "single number" and you need the structured answer that surfaces the assumption
**Do not use for:**
- Backward-looking financial close + reporting → `finance/financial-analysis`
- Strategic financial planning (multi-year, scenario, fundraise) → `c-level-advisor/cfo-advisor`
- "Should we hire a VP Sales?" / territory design / comp plan → `c-level-advisor/cro-advisor`
- Setting prices → sibling `pricing-strategist` (projects revenue *at* prices already set)
- Per-deal discount approval → sibling `deal-desk`
## Workflow
### Step 1 — Intake pipeline + cohort + historical conversion data
Fill `assets/forecast_intake_template.md` (≈ 20 min). Captures: opportunity list with stage/amount/close-date/age/last-activity; historical stage-to-stage conversion across last 4Q and last 12Q; per-cohort ARR + per-quarter retention + expansion data; funnel stage names with 12-quarter conversion history.
### Step 2 — Run 3-tier bookings forecast
```
scripts/bookings_forecaster.py --input intake.json --profile saas --output markdown
```
Outputs three numbers — **commit**, **best-case**, **pipe-only** — each with the conversion rate applied, the data window used (last-4Q vs. last-12Q weighted 70/30), and the time-to-close probability adjustment. Surfaces variance between commit and pipe-only as the pipeline-risk indicator.
**The assumption block is non-optional.** If you remove it, the forecast becomes theatre.
### Step 3 — Project cohort-level ARR
```
scripts/cohort_arr_projector.py --input intake.json --output markdown
```
Computes per-cohort NRR + GRR over the projection horizon. Flags any cohort whose NRR is declining vs. the trailing-cohort average — these are the leaky cohorts that the consolidated number will hide for 2-3 quarters before the leak surfaces in the topline.
Output includes the consolidated NRR/GRR trajectory + the cohort heatmap + a leaky-cohort callout.
### Step 4 — Score per-stage funnel confidence
```
scripts/funnel_confidence_scorer.py --input intake.json --output markdown
```
Per stage: mean conversion %, standard deviation, coefficient of variation (CoV = StDev / Mean), confidence band (HIGH < 10%, MEDIUM 10-25%, LOW 25-50%, VERY LOW > 50%). Recommends treatment per stage: extend-data-window, treat-as-soft-floor, or commit-quality.
### Step 5 — Assemble the forecast deck
Take the 3-tier bookings number + cohort heatmap + funnel confidence into the QBR / board deck. **The assumption block goes on the slide with the number.** If the slide has a single number and no assumption block, the slide is theatre.
## Scripts
- `scripts/bookings_forecaster.py` — 3-tier bookings forecast (commit / best-case / pipe-only) with disclosed conversion-rate + data-window + weighting block
- `scripts/cohort_arr_projector.py` — per-cohort NRR/GRR projection over horizon with leaky-cohort callout
- `scripts/funnel_confidence_scorer.py` — per-stage CoV-based confidence bands with treatment recommendation
All scripts: stdlib only. `--help` and `--sample` work on all three.
## References
- `references/saas_forecasting_canon.md` — Skok, Tunguz, OpenView, BVP, Pacific Crest/KeyBanc, ProfitWell, Patrick Campbell
- `references/cohort_analysis_canon.md` — Andrew Chen (a16z), Brian Balfour, Skok, Ramanujam, OpenView, Lenny Rachitsky, Reforge
- `references/forecast_anti_patterns.md` — McKinsey, Tunguz, OpenView, MIT Sloan, Bain, Forrester, Pacific Crest
## Assumptions
- **Historical conversion is the prior, not the truth.** Last 4Q is weighted 70%, last 12Q is weighted 30%. The blend captures regime change (recent slowdown) without overfitting to a single bad quarter. Window + weighting are surfaced in every output.
- **A forecast without a disclosed assumption block is theatre.** This is the skill's hard rule. The CLI refuses to omit the assumption block.
- **Cohort decomposition reveals leaks 2-3 quarters before the consolidated number does.** Reporting NRR without per-cohort breakdown hides the leak.
- **CoV (coefficient of variation) is the right discipline for stage confidence.** A stage with mean conversion 40% and stdev 4% (CoV 10%) is HIGH confidence; mean 40% stdev 20% (CoV 50%) is VERY LOW. The same average masks very different reliability.
- **Industry profile tunes priors, not truth.** Profile shifts default stage-conversion rates by industry; your historical data overrides.
- **The skill emits three numbers and an assumption block.** The CRO picks the commit number, owns the trade-off, and walks the board through the variance.
## Anti-patterns
- **Single-number forecast with no confidence band.** The board asks for "the number"; the discipline is to present three with named assumptions. See `forecast_anti_patterns.md`.
- **Using last-12-quarter conversion blindly.** Hides recent slowdown. The 70/30 blend on last-4Q vs. last-12Q corrects this.
- **Reporting NRR without cohort decomposition.** The consolidated number can be flat while a recent cohort is leaking 15 pp; the leak surfaces in the topline 2-3 quarters later. Always decompose.
- **Treating best-case as commit.** The CFO will eat you. Best-case includes weighted-stage opps that have a < 50% time-to-close probability; commit only includes commit-grade stages.
- **Hiding the assumption block.** The skill refuses; if you remove it manually, you own the theatre.
- **No leaky-cohort callout.** If `cohort_arr_projector.py` flags a cohort and you suppress the flag in the deck, the leak owns you next quarter.
- **Ignoring late-stage opp age.** A "verbal" deal that's been verbal for 180 days is not a commit. The bookings forecaster downweights stalled opps automatically; do not re-up them by hand.
- **No pipeline-coverage check.** Industry rule of thumb: forecast > pipeline ÷ 3 is anti-pattern. The tool surfaces the ratio; respect it.
## Distinct from
- **`finance/financial-analysis`** — backward-looking financial close, GAAP/IFRS reporting, variance vs. budget. commercial-forecaster is forward-looking pipeline math.
- **`c-level-advisor/cfo-advisor`** — strategic multi-year financial planning, fundraise scenarios, runway. commercial-forecaster is one input to the CFO, not the strategy.
- **`c-level-advisor/cro-advisor`** — strategic CRO judgment: "do we hire a VP Sales?", territory design, comp plan, when to add a sales engineer. commercial-forecaster is the math the CRO uses; cro-advisor is the judgment the CRO applies.
- **sibling `pricing-strategist`** — sets the price (model + range). commercial-forecaster *projects revenue at those prices*. Pricing comes first; forecast comes after.
- **sibling `deal-desk`** — per-deal scoring + discount approval routing. commercial-forecaster aggregates the pipeline that deal-desk operates on day-by-day.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-commercial` or the orchestrator. Recommended answer + canon citation per question. Never bundled.
1. **"What conversion rate are you using, and is it last-4Q or last-12Q?"**
Recommended: a 70/30 blend (last-4Q weighted 70%, last-12Q weighted 30%). Last-12Q alone hides recent slowdown; last-4Q alone overfits one bad quarter.
Canon: Tomasz Tunguz (Theory Ventures) — forecasting studies show single-window conversion estimates miss regime change at ~3-quarter lag.
2. **"What's your pipeline coverage ratio, and is your commit above pipeline ÷ 3?"**
Recommended: 3x coverage is the SaaS-industry floor; below 3x means your commit is structurally unsupported.
Canon: Pacific Crest / KeyBanc SaaS Survey — top-quartile SaaS companies maintain 3.0-4.5x pipeline coverage against committed bookings.
3. **"Can you show me NRR by cohort, not just consolidated?"**
Recommended: never report a consolidated NRR without the per-cohort breakdown. Leaky cohorts hide in averages.
Canon: Patrick Campbell (ProfitWell) + David Skok — cohort-driven retention decomposition surfaces leaks 2-3 quarters before consolidated NRR moves.
4. **"What's the variance (CoV) on each stage's conversion rate over the last 12 quarters?"**
Recommended: CoV < 10% → commit-grade; 10-25% → moderate; 25-50% → soft floor only; > 50% → do not use this stage for forecasting.
Canon: MIT Sloan forecasting research / Hyndman & Athanasopoulos (*Forecasting: Principles and Practice*) — CoV on the input series predicts forecast accuracy more reliably than mean.
5. **"How long has each late-stage opp been in late-stage?"**
Recommended: stage-age > 2x the median stage-duration → treat as stalled, exclude from commit, keep in pipe-only.
Canon: David Skok (*For Entrepreneurs*) — stalled-opp identification by stage-age is the #1 forecast hygiene practice in top-decile SaaS pipelines.
6. **"Is your best-case forecast within 30% of your pipe-only?"**
Recommended: if best-case is < 50% of pipe-only, your stage-conversion assumptions are pessimistic and you're sandbagging; if best-case > 80% of pipe-only, you're hockey-sticking.
Canon: McKinsey research on forecast bias + OpenView SaaS benchmarks — most teams operate in one of two failure modes: sandbagging (commit << earnings) or hockey-sticking (commit >> earnings).
7. **"What assumption block accompanies the number on the board slide?"**
Recommended: every forecast number on a board slide names (a) the conversion rate, (b) the data window, (c) the weighting choice, (d) the pipeline-coverage ratio. No assumption block = the slide is theatre.
Canon: Bain & Company commercial-forecasting practice + Forrester pipeline-coverage research — undisclosed-assumption forecasts have 2.3x higher variance against actuals than disclosed-assumption forecasts.
Walk depth-first. Lock 1-3 before opening 4-7. After all 7 are answered, invoke `bookings_forecaster.py` → `cohort_arr_projector.py` → `funnel_confidence_scorer.py` in sequence.
FILE:assets/forecast_intake_template.md
# Forecast Intake Template
**Time to fill:** ~20 minutes for Head of Commercial / RevOps / VP Sales.
This template captures the four inputs the `commercial-forecaster` skill needs:
1. **Opportunities** — current pipeline with stage / amount / close-date / age / last-activity
2. **Historical conversion** — stage-to-stage % over last 4 quarters AND last 12 quarters
3. **Cohorts** — per-cohort starting ARR + per-quarter retention + expansion
4. **Funnel history** — per-stage conversion across the last 12 quarters
The output is a single JSON file that feeds all three scripts:
- `scripts/bookings_forecaster.py --input intake.json --profile {saas|api|enterprise-software|marketplace|services}`
- `scripts/cohort_arr_projector.py --input intake.json`
- `scripts/funnel_confidence_scorer.py --input intake.json`
---
## Section 1 — Target period
The quarter / period you're forecasting for.
- Start date (YYYY-MM-DD): __________
- End date (YYYY-MM-DD): __________
- Industry profile (saas / api / enterprise-software / marketplace / services): __________
---
## Section 2 — Opportunities (pipeline snapshot)
Export from your CRM (Salesforce / HubSpot / Pipedrive). One row per opportunity:
| opp_id | stage | amount | close_date | age_days | last_activity_days |
|---|---|---|---|---|---|
| OPP-101 | commit | 180000 | 2026-06-15 | 45 | 3 |
| OPP-102 | verbal | 95000 | 2026-06-22 | 60 | 7 |
| ... | ... | ... | ... | ... | ... |
**Stage values to use** (case-insensitive): `discovery`, `demo_completed`, `proposal`,
`negotiation`, `verbal`, `commit`, `contract_out`, `closed_won_pending`.
**Hygiene check:**
- Filter out any opp older than 365 days that has not moved stage
- Confirm close_date is realistic — if it's already past, the CRM hygiene is the problem first
---
## Section 3 — Historical conversion (last 4Q and last 12Q)
Stage-to-stage conversion percentage, computed from your CRM history.
**Last 4 quarters (recent regime):**
| Stage | Conversion % |
|---|---:|
| discovery | _____ |
| demo_completed | _____ |
| proposal | _____ |
| negotiation | _____ |
| verbal | _____ |
| commit | _____ |
**Last 12 quarters (long-run prior):**
| Stage | Conversion % |
|---|---:|
| discovery | _____ |
| demo_completed | _____ |
| proposal | _____ |
| negotiation | _____ |
| verbal | _____ |
| commit | _____ |
The skill blends 70% last-4Q + 30% last-12Q automatically.
---
## Section 4 — Cohorts
One row per acquisition cohort (typically by quarter):
For each cohort:
- cohort_id (e.g., "2025-Q1")
- acquisition_quarter (e.g., "2025-Q1")
- starting_arr (USD)
- gross_retention_pct_q1, q2, q3, q4 (each is the % of starting ARR retained in that projection
quarter — typically 85-95)
- expansion_arr_pct_q1, q2, q3, q4 (each is the % expansion ARR — typically 4-15)
If you don't have per-quarter retention for a cohort, leave them blank and the skill will apply
conservative defaults (92%/91%/90%/89% GRR, 4%/6%/8%/10% expansion).
---
## Section 5 — Funnel history (per-stage conversion across 12 quarters)
One row per funnel stage. The conversion_pct_history is a 12-element list of the per-quarter
conversion rate for that stage transition.
- stage_name (e.g., "discovery_to_demo")
- conversion_pct_history (list of 12 numbers, oldest first)
This feeds `funnel_confidence_scorer.py` to compute per-stage CoV and confidence band.
---
## JSON skeleton (paste into `intake.json`)
```json
{
"target_period": {
"start_date": "2026-06-01",
"end_date": "2026-06-30"
},
"opportunities": [
{
"opp_id": "OPP-101",
"stage": "commit",
"amount": 180000,
"close_date": "2026-06-15",
"age_days": 45,
"last_activity_days": 3
},
{
"opp_id": "OPP-102",
"stage": "verbal",
"amount": 95000,
"close_date": "2026-06-22",
"age_days": 60,
"last_activity_days": 7
}
],
"historical_conversion": {
"stage_X_to_Y_pct_last_4q": {
"discovery": 0.32,
"demo_completed": 0.52,
"proposal": 0.60,
"negotiation": 0.72,
"verbal": 0.84,
"commit": 0.91
},
"stage_X_to_Y_pct_last_12q": {
"discovery": 0.38,
"demo_completed": 0.58,
"proposal": 0.67,
"negotiation": 0.76,
"verbal": 0.87,
"commit": 0.93
}
},
"cohorts": [
{
"cohort_id": "2025-Q1",
"acquisition_quarter": "2025-Q1",
"starting_arr": 1200000,
"gross_retention_pct_q1": 93,
"gross_retention_pct_q2": 91,
"gross_retention_pct_q3": 90,
"gross_retention_pct_q4": 89,
"expansion_arr_pct_q1": 5,
"expansion_arr_pct_q2": 8,
"expansion_arr_pct_q3": 10,
"expansion_arr_pct_q4": 11
}
],
"projection_horizon_quarters": 4,
"funnel_stages": [
{
"stage_name": "discovery_to_demo",
"conversion_pct_history": [35, 37, 33, 36, 38, 35, 34, 37, 36, 35, 36, 37]
},
{
"stage_name": "demo_to_proposal",
"conversion_pct_history": [55, 52, 58, 56, 54, 57, 53, 55, 58, 54, 56, 55]
}
]
}
```
---
## Quality gates before running the scripts
- [ ] All opportunities have a stage from the allowed list
- [ ] All opportunities have a close_date (no nulls — fix CRM hygiene first)
- [ ] Last-4Q AND last-12Q conversion provided for at least 4 stages
- [ ] At least 3 cohorts with starting_arr (4+ preferred for leak detection)
- [ ] At least 4 quarters of conversion_pct_history per funnel stage (12 preferred)
- [ ] Industry profile selected
---
## Next steps after intake
1. Save as `intake.json` in your working directory
2. Run `bookings_forecaster.py --input intake.json --profile <profile>` → 3-tier forecast + assumption block
3. Run `cohort_arr_projector.py --input intake.json` → cohort heatmap + leaky callout
4. Run `funnel_confidence_scorer.py --input intake.json` → per-stage confidence bands
5. Assemble the board slide: commit + best-case + pipe-only + assumption block + cohort heatmap + per-stage CoV
6. **The assumption block goes on the slide.** No assumption block = theatre.
FILE:references/cohort_analysis_canon.md
# Cohort Analysis Canon
Source material behind `cohort_arr_projector.py`'s NRR/GRR projection and the leaky-cohort callout.
## Core principle
A consolidated NRR number is an **ARR-weighted average that hides 5-15 percentage points of
dispersion across cohorts**. The consolidated number lags the underlying leak by 2-3 quarters
because (a) larger / older cohorts dominate the weighted average and (b) leaks compound silently.
The skill flags any cohort whose mean NRR falls ≥ 5 pp below the trailing-cohort average — that
is the level at which the leak is signal, not noise.
---
## Why cohort decomposition matters
Imagine four cohorts:
| Cohort | Starting ARR | Mean NRR Q1-Q4 |
|---|---:|---:|
| 2025-Q1 | $1.2M | 100% |
| 2025-Q2 | $1.5M | 100% |
| 2025-Q3 | $1.8M | 101% |
| 2025-Q4 | $2.1M | **85%** |
The consolidated ARR-weighted NRR for Q+1 looks roughly: (1.2×100 + 1.5×100 + 1.8×101 + 2.1×85) / 6.6
= ~95%. **That looks fine.** It even looks reasonable for a SaaS company.
But the 2025-Q4 cohort is bleeding 15pp below the trailing cohorts. Two quarters from now, when that
cohort becomes the dominant weight (because it was the largest), the consolidated number will collapse
to ~85%. The CFO who didn't see this coming will be unhappy.
This is why the consolidated number is a **lagging indicator** and the cohort heatmap is the
**forensic tool**.
---
## NRR vs. GRR — definitions used by this skill
- **GRR (Gross Retention Rate)** — the percentage of starting ARR retained in a cohort, excluding
expansion. Ceiling is 100%. Anything < 100% is churn + contraction.
- **NRR (Net Retention Rate)** — GRR + expansion ARR. Can exceed 100% when expansion outpaces
churn. The "best in SaaS" number.
- **Per-cohort projection** — for each cohort, project NRR and GRR forward over the horizon using
per-quarter retention and expansion inputs (or the default curve when missing).
- **Consolidated** — ARR-weighted average across cohorts per quarter.
---
## Source register (≥ 7 cited)
### 1. Andrew Chen — a16z (andrewchen.com)
The canonical introduction to cohort retention curves:
- The "smiling curve" (retention dips then recovers) is the rare healthy pattern; most products
produce a "frowning curve" that hides in averages
- Cohort decomposition is the discipline that catches a product/market-fit erosion 2 quarters
before NPS or aggregate retention does
### 2. Brian Balfour — Reforge (brianbalfour.com)
The retention-driven growth framework:
- "Retention is the single most underrated lever in growth math"
- Cohorts must be decomposed by acquisition source, persona, and pricing tier — a single cohort
variable is insufficient
- Expansion-driven NRR > 110% requires structural product loops, not just sales motion
### 3. David Skok — *For Entrepreneurs* (matrixpartners.com)
Cohort analysis as the SaaS forensic standard:
- The "logo retention" / "dollar retention" / "net dollar retention" hierarchy
- Cohort heatmaps are the diagnostic for both retention and expansion
- Recommended floor for cohort-level GRR: 90% for SMB SaaS, 95%+ for enterprise
### 4. Madhavan Ramanujam — *Monetizing Innovation* (Simon-Kucher)
The pricing-retention nexus:
- Customers who feel they overpaid in Q1 churn in Q3-Q4 — cohort decomposition reveals pricing
misalignment with delayed signal
- A leaky cohort is often a pricing problem, not a product problem
- Cohort + pricing-tier decomposition is the technique that finds the leak's source
### 5. OpenView Partners — Cohort benchmarks (openviewpartners.com)
The numeric benchmarks underneath the skill's defaults:
- Top-quartile SaaS Q1 GRR: 93-95%
- Top-quartile cohort expansion Q1: 5-8%, Q4: 10-15%
- Bottom-quartile cohorts often hide 10+ pp below the consolidated number
### 6. Lenny Rachitsky — Lenny's Newsletter (lennysnewsletter.com)
Modern practitioner canon on cohort retention curves:
- "Show me your cohort retention curves and I'll tell you if you have product-market fit"
- The shape of the curve (flat vs. declining vs. smiling) is more diagnostic than any single number
- Cohort retention dispersion is a leading indicator for ARR forecasting accuracy
### 7. Reforge — Retention + Engagement program (reforge.com)
The systematic framework that operationalizes Balfour / Chen:
- Cohorts decomposed by 4 lenses: acquisition source, persona, lifecycle stage, pricing tier
- "Retention frameworks should be a board metric, not a product metric"
- Cohort heatmaps as standard quarterly artifact
### 8. Patrick Campbell / ProfitWell (now Paddle) — Cohort-driven retention research
The discipline of cohort decomposition for retention forecasting:
- Average NRR can stay flat for 2-3 quarters while a recent cohort is leaking
- "If you can't tell me your NRR by acquisition cohort, you don't know your NRR"
- Source of the skill's 5 pp leak-threshold default
---
## Leak detection rule (used by this skill)
A cohort is flagged **leaky** if:
- Its mean NRR across the projection horizon is **≥ 5 percentage points below** the mean NRR of
all earlier-acquired cohorts (the "trailing-cohort average").
The 5 pp threshold is calibrated from Campbell / ProfitWell research: at < 5 pp, the gap is within
normal cohort-to-cohort variance; at ≥ 5 pp, the gap is signal that compounds quickly into the
consolidated number.
---
## Default retention curves (used when per-quarter data is missing)
When a cohort is provided without per-quarter retention/expansion data, the skill applies these
conservative defaults derived from OpenView benchmarks:
- **GRR curve**: 92% in Q1, decaying ~1 pp per quarter, with a floor of 85%
- **Expansion curve**: 4% in Q1, ramping +2 pp per quarter, capped at 12%
These are **priors, not prescriptions**. Always supply your real per-cohort data when available.
---
## Hard rules surfaced from canon
1. **Never present consolidated NRR without the cohort heatmap.** The consolidated number is the
lagging indicator; the heatmap is the forensic tool.
2. **Decompose cohorts by acquisition quarter at minimum.** Better: + acquisition source, pricing
tier, persona, segment.
3. **A leaky cohort signals a problem to investigate, not a number to discount.** Root-cause first:
pricing mismatch? sales-motion drift? product-fit erosion? competitive incursion?
4. **Expansion-driven NRR > 110% requires product loops.** If your expansion is sales-led only,
you're one comp-plan change away from collapse.
5. **Above $50M ARR, cohort decomposition is malpractice to skip.**
FILE:references/forecast_anti_patterns.md
# Forecast Anti-Patterns
The cataloged failure modes of SaaS commercial forecasting. Source material behind the skill's
warnings, hard rules, and the forcing-question library.
## Core principle
**A forecast without a disclosed assumption block is theatre.** It cannot be evaluated, corrected,
or learned from. Theatre forecasts produce more variance against actuals than disclosed-assumption
forecasts by a factor of 2-3x (Bain commercial-forecasting practice; Forrester pipeline-coverage research).
Every anti-pattern below is a way of producing theatre — sometimes accidentally, sometimes
performatively.
---
## Anti-pattern catalog (≥ 8)
### 1. Single-number forecast with no confidence band
**Symptom:** the board slide says "$8.4M Q3 commit". That's it. No best-case, no pipe-only, no
assumption block.
**Why it fails:** the CFO cannot evaluate whether 8.4 is achievable, conservative, or aspirational
without knowing the dispersion. The forecast is unfalsifiable in advance and unaccountable in retrospect.
**Fix:** present three numbers (commit / best-case / pipe-only) AND the assumption block. Always.
**Canon:** McKinsey on forecast bias — single-number forecasts produce 2-3x higher variance against
actuals than 3-tier forecasts because they suppress disagreement.
### 2. Use last-12-quarter conversion blindly
**Symptom:** the conversion rate applied to each stage is the trailing 12-quarter average. It
hasn't been recomputed since 2024.
**Why it fails:** last-12Q smooths over regime change. If the last 4 quarters show a 10pp drop in
demo-to-proposal conversion (post-funding-correction sales drag, e.g.), the 12Q average will lag
that signal by 2-3 quarters. By the time it shows up, you've missed two forecasts.
**Fix:** blend 70% last-4Q + 30% last-12Q. Disclose the blend on the slide.
**Canon:** Tomasz Tunguz forecasting studies + MIT Sloan / Hyndman *Forecasting: Principles and
Practice* — blended windows outperform either window alone in regime-change environments.
### 3. Report NRR without cohort decomposition
**Symptom:** the QBR slide shows "NRR: 108%". One number. No cohort heatmap, no segment cut.
**Why it fails:** the consolidated NRR is an ARR-weighted average that can hide 5-15pp leaks in
recent cohorts. The leak surfaces in the consolidated number 2-3 quarters after it starts. By
then, the deal is done.
**Fix:** present NRR with the cohort heatmap + the leaky-cohort callout.
**Canon:** Patrick Campbell / ProfitWell + Brian Balfour (Reforge) — "if you can't tell me your
NRR by acquisition cohort, you don't know your NRR."
### 4. Treat best-case as commit
**Symptom:** the commit number quietly includes opps in proposal / negotiation stages weighted
optimistically. The number looks aggressive; the CFO challenges it; the CRO digs in.
**Why it fails:** commit is the number the CRO defends even when the quarter goes sideways. If
commit includes weighted-stage opps, the CRO will miss commit when the quarter does go sideways —
and credibility collapses.
**Fix:** commit = commit-grade stages only (verbal / contract-out / commit). Best-case is the
separate, optimistic number.
**Canon:** Bain commercial-forecasting practice + OpenView SaaS benchmarks — top-quartile teams
hit commit within 5%; bottom-quartile miss by 25%+, almost always because commit was conflated
with best-case.
### 5. Hide the assumption block
**Symptom:** the forecast is presented; someone asks "what conversion rate are you using?"; the
answer is "the historical one" or "trust me, it's calibrated".
**Why it fails:** the slide is now theatre. The forecast is unfalsifiable and unaccountable.
**Fix:** the assumption block is non-optional. It names (a) the conversion rate, (b) the data
window, (c) the weighting choice, (d) the pipeline-coverage ratio. The skill refuses to omit it;
if you remove it manually, you own the theatre.
**Canon:** Bain & Co + Forrester — undisclosed-assumption forecasts have 2.3x higher variance
against actuals than disclosed-assumption forecasts.
### 6. No leaky-cohort callout
**Symptom:** the cohort heatmap is presented, the recent cohort is visibly leaking 15pp, no one
calls it out. Everyone moves on to the next slide.
**Why it fails:** the leak doesn't go away because no one mentioned it. Two quarters later, the
consolidated NRR drops 8pp and the board is angry.
**Fix:** when `cohort_arr_projector.py` flags a cohort, the flag goes on the slide. Root-cause
must follow within the deck or in the next 1:1.
**Canon:** Skok + Campbell — cohort decomposition is the forensic tool; suppressing the finding
makes you the problem.
### 7. Ignore late-stage opp age (stalled = false-positive)
**Symptom:** a "verbal" deal has been verbal for 180 days. It's in commit. Last activity was 60
days ago.
**Why it fails:** verbal-stage opps that haven't moved in 6 months are not commits. They are
either dead, deprioritized, or being shopped against you. Including them in commit inflates the
number and guarantees a miss.
**Fix:** apply the stall rule — opp age > 2x median stage age AND last_activity > 45 days →
contribution × 0.5 in commit. Surface stalled opps explicitly.
**Canon:** David Skok — "stalled-opp identification by stage-age is the #1 forecast-hygiene
practice in top-decile SaaS pipelines."
### 8. No pipeline-coverage check
**Symptom:** the commit is $8.4M. The total pipeline is $18M. Coverage ratio is 2.1x. No one
mentions this.
**Why it fails:** coverage < 3.0x means the commit is structurally unsupported. Even if every
stage-conversion assumption is correct, the math doesn't have enough opps to hit commit if a
normal percentage slip.
**Fix:** the tool calculates the coverage ratio. Below 3.0x → warning. Above 3.0x → confirm.
**Canon:** Pacific Crest / KeyBanc SaaS Survey + Forrester pipeline-coverage research — 3.0x is
the SaaS-industry floor; top-quartile maintains 3.0-4.5x.
### 9. Sandbagging (best-case far below pipe-only)
**Symptom:** pipe-only is $25M; best-case is $9M (36% of pipe-only). The CRO is being "conservative".
**Why it fails:** if best-case is < 50% of pipe-only, the team has effectively given up on most
of the pipeline. Either the stage-conversion priors are pessimistic, or the team isn't working
the pipeline.
**Fix:** the tool flags this ratio. If best-case is < 50% of pipe-only, decompose why before
presenting.
**Canon:** McKinsey on forecast bias + Tomasz Tunguz — sandbagging is the more common failure
mode than hockey-sticking, especially after a missed quarter.
### 10. Hockey-sticking (best-case near pipe-only)
**Symptom:** pipe-only is $20M; best-case is $18M (90% of pipe-only). The team is "all-in" on Q3.
**Why it fails:** if best-case is > 80% of pipe-only, the team is assuming nearly all pipeline
will convert. Conversion math shows this is statistically impossible at any reasonable stage
mix.
**Fix:** the tool flags > 80%. Decompose: which stages are being weighted optimistically?
**Canon:** OpenView SaaS forecasting benchmarks — hockey-stick forecasts have 2x lower realization
rate than disciplined forecasts.
---
## Source register (≥ 7 cited)
1. **McKinsey** — forecast-bias research, especially on single-number vs. 3-tier forecast accuracy
2. **Tomasz Tunguz / Theory Ventures** — sandbagging vs. hockey-sticking analysis across 100+
SaaS companies; regime-change detection via blended windows
3. **OpenView Partners** — annual SaaS benchmarks on commit accuracy, pipeline coverage, hockey-stick
realization rates
4. **MIT Sloan** / Hyndman & Athanasopoulos, *Forecasting: Principles and Practice* — CoV-based
confidence bands, blended-window methodology, minimum sample size for stable forecasting
5. **Bain & Company** — commercial-forecasting practice on disclosed vs. undisclosed assumptions
(2.3x variance differential)
6. **Forrester Research** — pipeline-coverage myths; the 3x floor is necessary but not sufficient
7. **Pacific Crest / KeyBanc Capital Markets** — Private SaaS Survey, the industry data source
for pipeline-coverage benchmarks and stage-conversion priors
8. **David Skok / *For Entrepreneurs*** — stalled-opp hygiene as the #1 forecast practice in
top-decile pipelines
---
## Hard rules
1. **Three numbers, always: commit / best-case / pipe-only.** Never one.
2. **Assumption block on every slide with a forecast number.** Never hidden.
3. **Cohort heatmap accompanies every NRR number.** Never just consolidated.
4. **Pipeline coverage ratio surfaced.** Below 3.0x → warning.
5. **Stalled opps downweighted.** Verbal-for-6-months is not a commit.
6. **Sandbagging and hockey-sticking are both flagged.** The middle is the discipline.
FILE:references/saas_forecasting_canon.md
# SaaS Forecasting Canon
Curated, opinionated knowledge base for SaaS bookings + ARR forecasting. Source material behind
`bookings_forecaster.py`'s scoring rules and the 3-tier (commit / best-case / pipe-only) discipline.
## Core principle
A forecast is a **claim about the future under disclosed assumptions**. A forecast without disclosed
assumptions is theatre — it cannot be evaluated, corrected, or learned from. Every output of this
skill names the conversion rate, the data window, and the weighting choice.
The 3-tier model exists because the question "what's the number?" has three valid answers:
- **Commit** — what I will defend even if the quarter goes sideways
- **Best-case** — what I can hit if everything goes my way
- **Pipe-only** — the unweighted ceiling
Presenting one without the others is theatre. Presenting all three with the assumption block is
the discipline.
---
## The 3-tier discipline
### Commit
- Includes only commit-grade stages (verbal, contract-out, commit, closed-won-pending)
- Conversion applied: blended (70% last-4Q + 30% last-12Q)
- Time-to-close probability adjustment applied
- Stalled-opp downweight applied (opp age > 2x median stage age AND last_activity > 45 days → × 0.5)
- This is the number the CRO defends to the CEO and CFO
### Best-case
- Includes commit-grade stages + weighted-stage opps (proposal, negotiation, demo-completed)
- Conversion blended (70/30)
- Time-to-close probability applied
- NO stall downweight (best-case is the optimistic ceiling)
- This is the number for "if everything breaks our way"
### Pipe-only
- Includes everything in pipeline at any stage
- Conversion blended only (no time-to-close, no stall)
- This is the unweighted top of the funnel — useful as the divisor in pipeline-coverage ratio
### Pipeline coverage ratio
- Total pipeline $ / commit $
- SaaS-industry floor: 3.0x
- Below 3.0x → commit is structurally unsupported and the CFO will challenge it
---
## Source register (≥ 7 cited)
### 1. David Skok — *For Entrepreneurs* (matrixpartners.com)
Founding canon on SaaS metrics + forecasting. Specifically:
- The CAC-payback / LTV framework that anchors what "good" forecast accuracy looks like
- The pipeline-coverage discipline (3x as the industry floor)
- Cohort retention curves as the input to NRR forecasting, not the output
- "Stalled-opp identification by stage-age is the #1 forecast-hygiene practice in top-decile SaaS pipelines."
### 2. Tomasz Tunguz — Theory Ventures (tomtunguz.com)
Forecasting studies from 100+ SaaS companies. Specifically:
- Single-window conversion estimates miss regime change at ~3-quarter lag → blended weighting needed
- Sandbagging is the more common pattern than hockey-sticking, especially after a missed quarter
- Forecast accuracy degrades sharply for stages with CoV > 25%
- "If your last-4Q and last-12Q conversion diverge by more than 10pp, you have a regime change, not noise."
### 3. OpenView Partners — SaaS Forecasting Benchmarks (openviewpartners.com)
Annual State-of-the-Cloud-adjacent surveys with explicit forecast-accuracy benchmarks:
- Top-quartile SaaS companies hit commit within 5%; bottom-quartile miss by 25%+
- Hockey-stick forecasts (best-case > 80% of pipe-only) have 2x lower realization rate
- Pipeline coverage 3-4.5x is the typical band for healthy commit
- Recommends the 3-tier (commit / best-case / pipe-only) structure as standard board hygiene
### 4. Bessemer Venture Partners — State of the Cloud forecasting research (bvp.com/atlas)
The BVP "Cloud Index" methodology and the Good/Better/Best NRR benchmarks:
- 100% NRR = "good", 110% = "better", 120%+ = "best"
- Cohort decomposition is the forensic technique to detect leak before consolidated number moves
- Forecasting at the company level without cohort decomposition is malpractice for ARR > $50M
### 5. Pacific Crest / KeyBanc Capital Markets — Private SaaS Survey
Long-running annual survey of private SaaS companies (now KeyBanc):
- Pipeline-coverage ratio: top-quartile 3.0-4.5x, median ~3.0x, bottom-quartile < 2.5x
- Forecast accuracy correlates more tightly with stage-conversion CoV than with mean conversion
- Standard sales stages and their expected conversion priors (used as fallback in this skill's profiles)
### 6. Patrick Campbell / ProfitWell (now Paddle) — Cohort-driven retention research
The cohort-decomposition discipline:
- Consolidated NRR is an average that hides 5-15pp dispersion across cohorts
- Leaky cohorts surface in the consolidated number 2-3 quarters after the leak begins
- The cohort heatmap is the forensic tool; the consolidated number is the lagging indicator
- "If you cannot tell me your NRR by acquisition cohort, you do not know your NRR."
### 7. MIT Sloan — Forecasting research (Hyndman & Athanasopoulos, *Forecasting: Principles and Practice*)
The statistical canon underneath the CoV-based confidence bands:
- CoV (coefficient of variation) on the input series predicts forecast accuracy more reliably than mean
- Sample size n ≥ 4 is the practical minimum for stable CoV estimation
- Weighted blends of recent vs. long-run windows outperform either window alone when regime change is plausible
### 8. Winning by Design — Bowtie GTM model + revenue forecasting (winningbydesign.com)
The bowtie model + recurring-impact framework:
- Forecast must account for both new ARR AND retained/expansion ARR (the right side of the bowtie)
- Pipeline-coverage on new bookings is insufficient; expansion pipeline coverage is the second leg
- Aligns with the cohort decomposition discipline above
---
## Calibration table — used by `bookings_forecaster.py`
Default stage-conversion priors per industry profile (applied only when historical data is missing
for that stage). These are deliberately conservative — your data overrides.
| Stage | saas | api | enterprise-software | marketplace | services |
|---|---:|---:|---:|---:|---:|
| discovery | 35% | 45% | 20% | 40% | 30% |
| demo_completed | 55% | 60% | 40% | 60% | 50% |
| proposal | 65% | 70% | 55% | 68% | 62% |
| negotiation | 75% | 80% | 68% | 78% | 72% |
| verbal | 85% | 88% | 80% | 86% | 82% |
| commit | 92% | 94% | 90% | 92% | 90% |
Sources: KeyBanc SaaS Survey, OpenView benchmarks, Bessemer Atlas. Profile picker is a starting prior,
not a prescription.
---
## Hard rules surfaced from canon
1. **Forecast without disclosed assumptions is theatre.** Every CLI output names the conversion
rate, the data window, and the weighting choice. Manual suppression of the assumption block
makes the human responsible for the theatre.
2. **The 3-tier model is non-collapsible.** Presenting commit without best-case and pipe-only loses
information. The CFO needs to know the dispersion.
3. **Pipeline coverage 3.0x is the floor, not the ceiling.** Below 3.0x, the commit is structurally
unsupported.
4. **Stalled opps are not commit.** A "verbal" deal that's been verbal for 6 months is not a commit;
the stall rule downweights them.
5. **Cohort decomposition is mandatory above $50M ARR.** Below that, it's strongly recommended.
FILE:scripts/bookings_forecaster.py
#!/usr/bin/env python3
"""bookings_forecaster.py — 3-tier bookings forecast (commit / best-case / pipe-only) with explicit assumption block.
Input: JSON describing opportunities (stage, amount, close_date, age_days, last_activity_days),
historical stage-to-stage conversion (last 4Q and last 12Q windows), and target forecast period.
Output: three forecast numbers (commit, best-case, pipe-only) with the conversion rate, data window,
and weighting choice surfaced explicitly in an assumption block. Forecast without disclosed assumptions
is theatre — the assumption block is non-optional.
Deterministic decision logic. No LLM calls. No third-party deps.
Usage:
bookings_forecaster.py --input intake.json --profile saas --output markdown
bookings_forecaster.py --sample
"""
from __future__ import annotations
import argparse
import json
import math
import statistics
import sys
from dataclasses import dataclass, field
from datetime import date, datetime
from pathlib import Path
from typing import Any
# Commit-grade stages: opportunities here count toward the commit number
COMMIT_GRADE_STAGES = {"commit", "verbal", "contract_out", "contract-out", "closed_won_pending"}
# Best-case stages: weighted-stage opps that pass the time-to-close probability threshold
BEST_CASE_STAGES = {
"commit", "verbal", "contract_out", "contract-out", "closed_won_pending",
"proposal", "negotiation", "demo_completed", "demo-completed",
}
# Industry profile: default stage-conversion priors when historical data is missing per stage
PROFILES: dict[str, dict[str, float]] = {
"saas": {
"discovery": 0.35, "demo_completed": 0.55, "proposal": 0.65,
"negotiation": 0.75, "verbal": 0.85, "commit": 0.92,
},
"api": {
"discovery": 0.45, "demo_completed": 0.60, "proposal": 0.70,
"negotiation": 0.80, "verbal": 0.88, "commit": 0.94,
},
"enterprise-software": {
"discovery": 0.20, "demo_completed": 0.40, "proposal": 0.55,
"negotiation": 0.68, "verbal": 0.80, "commit": 0.90,
},
"marketplace": {
"discovery": 0.40, "demo_completed": 0.60, "proposal": 0.68,
"negotiation": 0.78, "verbal": 0.86, "commit": 0.92,
},
"services": {
"discovery": 0.30, "demo_completed": 0.50, "proposal": 0.62,
"negotiation": 0.72, "verbal": 0.82, "commit": 0.90,
},
}
# Weighting: blend last-4Q (recent regime) and last-12Q (long-run prior)
W_LAST_4Q = 0.70
W_LAST_12Q = 0.30
# Stalled-opp rule: opp age > AGE_STALL_MULTIPLIER * median_stage_age → downweighted
AGE_STALL_MULTIPLIER = 2.0
STALL_DOWNWEIGHT = 0.5 # multiplier applied to stalled opps in commit / best-case
@dataclass
class StageConversion:
stage: str
rate: float
window: str # "blended", "last_4q", "last_12q", or "profile_prior"
rationale: str = ""
@dataclass
class OppContribution:
opp_id: str
stage: str
amount: float
conversion: float
time_to_close_prob: float
stalled: bool
contribution_commit: float
contribution_best_case: float
contribution_pipe_only: float
@dataclass
class ForecastResult:
commit: float
best_case: float
pipe_only: float
pipeline_coverage_ratio: float
pipeline_risk_pct: float # variance between commit and pipe-only
assumptions: dict[str, Any]
stage_conversions: list[StageConversion]
opp_contributions: list[OppContribution]
warnings: list[str] = field(default_factory=list)
def parse_date(s: str | None) -> date | None:
if not s:
return None
try:
return datetime.fromisoformat(str(s)).date()
except ValueError:
return None
def blend_conversion(
stage: str,
hist: dict[str, Any],
profile: str,
) -> StageConversion:
"""Return blended conversion rate for a stage with surfaced window."""
last4 = hist.get("stage_X_to_Y_pct_last_4q") or {}
last12 = hist.get("stage_X_to_Y_pct_last_12q") or {}
r4 = last4.get(stage)
r12 = last12.get(stage)
if r4 is not None and r12 is not None:
rate = W_LAST_4Q * float(r4) + W_LAST_12Q * float(r12)
return StageConversion(
stage=stage,
rate=rate,
window="blended",
rationale=f"Blended {W_LAST_4Q:.0%} last-4Q ({r4:.2%}) + {W_LAST_12Q:.0%} last-12Q ({r12:.2%}).",
)
if r4 is not None:
return StageConversion(
stage=stage,
rate=float(r4),
window="last_4q",
rationale=f"Only last-4Q available ({r4:.2%}); no last-12Q data.",
)
if r12 is not None:
return StageConversion(
stage=stage,
rate=float(r12),
window="last_12q",
rationale=f"Only last-12Q available ({r12:.2%}); no last-4Q data.",
)
prior = PROFILES.get(profile, PROFILES["saas"]).get(stage)
if prior is not None:
return StageConversion(
stage=stage,
rate=prior,
window="profile_prior",
rationale=f"No historical data; using '{profile}' profile prior ({prior:.2%}).",
)
return StageConversion(
stage=stage,
rate=0.20,
window="fallback",
rationale="No historical data, no profile prior; using conservative 20% fallback.",
)
def time_to_close_probability(
close_date: date | None,
target_start: date | None,
target_end: date | None,
age_days: int,
) -> float:
"""Probability that the opp closes within the target window.
Heuristic: linear decay from 1.0 (close_date inside window) → 0.3 (close_date 90 days outside)
plus a stall penalty for high-age opps with no recent activity.
"""
if close_date is None or target_end is None:
return 0.50 # unknown close-date → coin flip
if target_start is not None and target_start <= close_date <= target_end:
return 1.0
if close_date < (target_start or close_date):
return 0.40 # close-date already past → CRM hygiene issue
days_late = (close_date - target_end).days
if days_late <= 30:
return 0.70
if days_late <= 60:
return 0.50
if days_late <= 90:
return 0.30
return 0.15
def is_stalled(age_days: int, last_activity_days: int, median_stage_age: int) -> bool:
if median_stage_age <= 0:
return last_activity_days > 60
return age_days > AGE_STALL_MULTIPLIER * median_stage_age and last_activity_days > 45
def compute_forecast(ctx: dict[str, Any], profile: str) -> ForecastResult:
opps = ctx.get("opportunities") or []
hist = ctx.get("historical_conversion") or {}
target = ctx.get("target_period") or {}
target_start = parse_date(target.get("start_date"))
target_end = parse_date(target.get("end_date"))
# Compute median stage age per stage for stall detection
by_stage_age: dict[str, list[int]] = {}
for o in opps:
stage = str(o.get("stage", "")).lower()
age = int(o.get("age_days") or 0)
by_stage_age.setdefault(stage, []).append(age)
median_stage_age = {s: int(statistics.median(ages)) for s, ages in by_stage_age.items() if ages}
# Resolve conversion per unique stage encountered
unique_stages = sorted({str(o.get("stage", "")).lower() for o in opps})
stage_conversions = [blend_conversion(s, hist, profile) for s in unique_stages]
sc_map = {sc.stage: sc for sc in stage_conversions}
commit_total = 0.0
best_case_total = 0.0
pipe_only_total = 0.0
contributions: list[OppContribution] = []
warnings: list[str] = []
for o in opps:
opp_id = str(o.get("opp_id") or o.get("id") or "?")
stage = str(o.get("stage", "")).lower()
amount = float(o.get("amount") or 0)
close_date = parse_date(o.get("close_date"))
age_days = int(o.get("age_days") or 0)
last_activity_days = int(o.get("last_activity_days") or 0)
sc = sc_map.get(stage)
rate = sc.rate if sc else 0.20
ttc = time_to_close_probability(close_date, target_start, target_end, age_days)
median_age = median_stage_age.get(stage, 0)
stalled = is_stalled(age_days, last_activity_days, median_age)
stall_mult = STALL_DOWNWEIGHT if stalled else 1.0
# Commit: commit-grade stages only, full rate × ttc × stall
contrib_commit = 0.0
if stage in COMMIT_GRADE_STAGES:
contrib_commit = amount * rate * ttc * stall_mult
# Best-case: best-case stages, rate × ttc (no stall penalty applied to best-case)
contrib_best = 0.0
if stage in BEST_CASE_STAGES:
contrib_best = amount * rate * ttc
# Pipe-only: all opps regardless of stage, weighted only by conversion (no ttc, no stall)
contrib_pipe = amount * rate
commit_total += contrib_commit
best_case_total += contrib_best
pipe_only_total += contrib_pipe
contributions.append(OppContribution(
opp_id=opp_id, stage=stage, amount=amount, conversion=rate,
time_to_close_prob=ttc, stalled=stalled,
contribution_commit=contrib_commit,
contribution_best_case=contrib_best,
contribution_pipe_only=contrib_pipe,
))
# Pipeline coverage ratio = total pipeline $ / commit number
total_pipeline = sum(float(o.get("amount") or 0) for o in opps)
coverage = (total_pipeline / commit_total) if commit_total > 0 else 0.0
if coverage > 0 and coverage < 3.0:
warnings.append(
f"Pipeline coverage ratio is {coverage:.2f}x — below the 3.0x SaaS-industry floor. "
f"Commit is structurally unsupported (Pacific Crest / KeyBanc SaaS Survey)."
)
pipeline_risk = 0.0
if pipe_only_total > 0:
pipeline_risk = (pipe_only_total - commit_total) / pipe_only_total * 100.0
if best_case_total > 0 and pipe_only_total > 0:
bc_pipe_ratio = best_case_total / pipe_only_total
if bc_pipe_ratio < 0.5:
warnings.append(
f"Best-case is {bc_pipe_ratio:.1%} of pipe-only — likely sandbagging "
f"(McKinsey forecast-bias research)."
)
elif bc_pipe_ratio > 0.8:
warnings.append(
f"Best-case is {bc_pipe_ratio:.1%} of pipe-only — likely hockey-sticking "
f"(OpenView SaaS forecasting benchmarks)."
)
# ASSUMPTION BLOCK — non-optional
assumptions = {
"conversion_window_weighting": f"{W_LAST_4Q:.0%} last-4Q + {W_LAST_12Q:.0%} last-12Q (blended)",
"industry_profile": profile,
"commit_grade_stages": sorted(COMMIT_GRADE_STAGES),
"best_case_stages": sorted(BEST_CASE_STAGES),
"time_to_close_model": "linear decay; 1.0 inside window, 0.7 within 30 days late, 0.5 within 60, 0.3 within 90, 0.15 thereafter",
"stall_rule": f"opp age > {AGE_STALL_MULTIPLIER}x median stage age AND last_activity > 45 days → contribution * {STALL_DOWNWEIGHT}",
"stage_conversions_applied": [
{"stage": sc.stage, "rate": round(sc.rate, 4), "window": sc.window, "rationale": sc.rationale}
for sc in stage_conversions
],
"data_window_disclosed": True,
"weighting_choice_disclosed": True,
}
return ForecastResult(
commit=commit_total,
best_case=best_case_total,
pipe_only=pipe_only_total,
pipeline_coverage_ratio=coverage,
pipeline_risk_pct=pipeline_risk,
assumptions=assumptions,
stage_conversions=stage_conversions,
opp_contributions=contributions,
warnings=warnings,
)
def render_markdown(r: ForecastResult, ctx: dict[str, Any], profile: str) -> str:
L: list[str] = []
target = ctx.get("target_period") or {}
L.append("# Bookings Forecast — 3-Tier")
L.append("")
L.append(f"**Profile:** `{profile}` • **Target period:** {target.get('start_date', '?')} → {target.get('end_date', '?')}")
L.append(f"**Opportunities scored:** {len(r.opp_contributions)}")
L.append("")
L.append("## Three numbers")
L.append("")
L.append(f"| Tier | Amount | Notes |")
L.append(f"|---|---:|---|")
L.append(f"| **Commit** | ,.0f | Commit-grade stages × blended conversion × time-to-close × stall penalty |")
L.append(f"| **Best-case** | ,.0f | Best-case stages × blended conversion × time-to-close |")
L.append(f"| **Pipe-only** | ,.0f | All pipeline × blended conversion (no time/stall adjustment) |")
L.append("")
L.append(f"**Pipeline-coverage ratio:** {r.pipeline_coverage_ratio:.2f}x (commit-relative)")
L.append(f"**Pipeline-risk variance:** {r.pipeline_risk_pct:.1f}% (commit-to-pipe gap)")
L.append("")
L.append("## Assumption block (NON-OPTIONAL — present this on the board slide)")
L.append("")
L.append(f"- **Conversion-window weighting:** {r.assumptions['conversion_window_weighting']}")
L.append(f"- **Industry profile:** `{r.assumptions['industry_profile']}`")
L.append(f"- **Commit-grade stages:** {', '.join(r.assumptions['commit_grade_stages'])}")
L.append(f"- **Best-case stages:** {', '.join(r.assumptions['best_case_stages'])}")
L.append(f"- **Time-to-close model:** {r.assumptions['time_to_close_model']}")
L.append(f"- **Stall rule:** {r.assumptions['stall_rule']}")
L.append("")
L.append("### Stage conversions applied")
L.append("")
L.append("| Stage | Rate | Window | Rationale |")
L.append("|---|---:|---|---|")
for sc in r.stage_conversions:
L.append(f"| {sc.stage} | {sc.rate:.2%} | {sc.window} | {sc.rationale} |")
L.append("")
if r.warnings:
L.append("## Warnings")
for w in r.warnings:
L.append(f"- ⚠️ {w}")
L.append("")
L.append("## Per-opp contributions (top 10 by commit)")
L.append("")
top = sorted(r.opp_contributions, key=lambda c: -c.contribution_commit)[:10]
L.append("| Opp | Stage | Amount | Conv | TTC | Stalled | Commit $ |")
L.append("|---|---|---:|---:|---:|:---:|---:|")
for c in top:
L.append(
f"| {c.opp_id} | {c.stage} | ,.0f | {c.conversion:.0%} | "
f"{c.time_to_close_prob:.0%} | {'Y' if c.stalled else '-'} | ,.0f |"
)
L.append("")
L.append("## Next steps")
L.append("1. Run `cohort_arr_projector.py` to surface leaky cohorts in NRR.")
L.append("2. Run `funnel_confidence_scorer.py` to score per-stage reliability (CoV).")
L.append("3. Present commit + best-case + pipe-only WITH the assumption block. No assumption block = theatre.")
return "\n".join(L)
def sample_context() -> dict[str, Any]:
return {
"opportunities": [
{"opp_id": "OPP-101", "stage": "commit", "amount": 180000, "close_date": "2026-06-15", "age_days": 45, "last_activity_days": 3},
{"opp_id": "OPP-102", "stage": "verbal", "amount": 95000, "close_date": "2026-06-22", "age_days": 60, "last_activity_days": 7},
{"opp_id": "OPP-103", "stage": "verbal", "amount": 220000, "close_date": "2026-08-05", "age_days": 210, "last_activity_days": 55}, # stalled
{"opp_id": "OPP-104", "stage": "negotiation", "amount": 140000, "close_date": "2026-06-30", "age_days": 90, "last_activity_days": 10},
{"opp_id": "OPP-105", "stage": "proposal", "amount": 75000, "close_date": "2026-07-15", "age_days": 30, "last_activity_days": 4},
{"opp_id": "OPP-106", "stage": "proposal", "amount": 250000, "close_date": "2026-09-01", "age_days": 75, "last_activity_days": 12},
{"opp_id": "OPP-107", "stage": "demo_completed", "amount": 60000, "close_date": "2026-07-30", "age_days": 25, "last_activity_days": 2},
{"opp_id": "OPP-108", "stage": "discovery", "amount": 110000, "close_date": "2026-08-20", "age_days": 14, "last_activity_days": 5},
{"opp_id": "OPP-109", "stage": "discovery", "amount": 45000, "close_date": "2026-09-15", "age_days": 8, "last_activity_days": 2},
],
"historical_conversion": {
"stage_X_to_Y_pct_last_4q": {
"discovery": 0.32, "demo_completed": 0.52, "proposal": 0.60,
"negotiation": 0.72, "verbal": 0.84, "commit": 0.91,
},
"stage_X_to_Y_pct_last_12q": {
"discovery": 0.38, "demo_completed": 0.58, "proposal": 0.67,
"negotiation": 0.76, "verbal": 0.87, "commit": 0.93,
},
},
"target_period": {"start_date": "2026-06-01", "end_date": "2026-06-30"},
}
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--input", type=Path, help="Path to forecast-intake JSON.")
p.add_argument(
"--profile", default="saas", choices=list(PROFILES.keys()),
help="Industry profile for stage-conversion priors when historical data is missing per stage.",
)
p.add_argument("--output", default="markdown", choices=["markdown", "json"], help="Output format.")
p.add_argument("--sample", action="store_true", help="Run with built-in sample context.")
args = p.parse_args(argv)
if args.sample:
ctx = sample_context()
elif args.input:
ctx = json.loads(args.input.read_text())
else:
p.error("Provide --input or --sample.")
return 2
result = compute_forecast(ctx, args.profile)
if args.output == "json":
out = {
"profile": args.profile,
"commit": round(result.commit, 2),
"best_case": round(result.best_case, 2),
"pipe_only": round(result.pipe_only, 2),
"pipeline_coverage_ratio": round(result.pipeline_coverage_ratio, 3),
"pipeline_risk_pct": round(result.pipeline_risk_pct, 2),
"assumptions": result.assumptions,
"warnings": result.warnings,
"opp_contributions": [
{
"opp_id": c.opp_id, "stage": c.stage, "amount": c.amount,
"conversion": round(c.conversion, 4),
"time_to_close_prob": round(c.time_to_close_prob, 3),
"stalled": c.stalled,
"commit": round(c.contribution_commit, 2),
"best_case": round(c.contribution_best_case, 2),
"pipe_only": round(c.contribution_pipe_only, 2),
}
for c in result.opp_contributions
],
}
print(json.dumps(out, indent=2))
else:
print(render_markdown(result, ctx, args.profile))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/cohort_arr_projector.py
#!/usr/bin/env python3
"""cohort_arr_projector.py — per-cohort NRR / GRR projection over horizon with leaky-cohort callout.
Input: JSON with cohorts (each with acquisition_quarter, starting_arr, per-quarter gross_retention
and expansion_arr percentages) plus a projection_horizon_quarters integer.
Output: per-cohort NRR + GRR projection over the horizon, the consolidated NRR/GRR trajectory, and
a leaky-cohort callout for any cohort whose NRR is declining vs the trailing-cohort average.
The cohort-decomposition discipline surfaces leaks 2-3 quarters before they reach the consolidated
number (Campbell / Skok). Reporting NRR without per-cohort breakdown hides the leak.
Deterministic. Stdlib only.
Usage:
cohort_arr_projector.py --input intake.json --output markdown
cohort_arr_projector.py --sample
"""
from __future__ import annotations
import argparse
import json
import statistics
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
# Leak threshold: cohort NRR more than N pp below trailing-cohort average → flag
LEAK_THRESHOLD_PP = 5.0
@dataclass
class CohortProjection:
cohort_id: str
acquisition_quarter: str
starting_arr: float
nrr_by_quarter: list[float] = field(default_factory=list)
grr_by_quarter: list[float] = field(default_factory=list)
arr_by_quarter: list[float] = field(default_factory=list)
leaky: bool = False
leak_reason: str = ""
@dataclass
class ProjectionResult:
cohorts: list[CohortProjection]
consolidated_nrr: list[float]
consolidated_grr: list[float]
consolidated_arr: list[float]
horizon_q: int
leaky_cohorts: list[str]
assumptions: dict[str, Any]
def project_cohort(cohort: dict[str, Any], horizon_q: int) -> CohortProjection:
cohort_id = str(cohort.get("cohort_id", "?"))
starting_arr = float(cohort.get("starting_arr") or 0)
acq_q = str(cohort.get("acquisition_quarter", "?"))
nrr_list: list[float] = []
grr_list: list[float] = []
arr_list: list[float] = []
running_arr = starting_arr
for q in range(1, horizon_q + 1):
gr_key = f"gross_retention_pct_q{q}"
exp_key = f"expansion_arr_pct_q{q}"
gr = float(cohort.get(gr_key) if cohort.get(gr_key) is not None else _default_grr(q)) / 100.0
exp = float(cohort.get(exp_key) if cohort.get(exp_key) is not None else _default_exp(q)) / 100.0
# NRR = GRR + expansion; multiplicative on the original cohort base
nrr = gr + exp
cohort_arr = starting_arr * nrr
nrr_list.append(nrr * 100.0)
grr_list.append(gr * 100.0)
arr_list.append(cohort_arr)
running_arr = cohort_arr
return CohortProjection(
cohort_id=cohort_id,
acquisition_quarter=acq_q,
starting_arr=starting_arr,
nrr_by_quarter=nrr_list,
grr_by_quarter=grr_list,
arr_by_quarter=arr_list,
)
def _default_grr(q: int) -> float:
# Conservative default GRR curve: 92% Q1, decaying ~1pp per quarter
return max(85.0, 92.0 - (q - 1) * 1.0)
def _default_exp(q: int) -> float:
# Conservative default expansion: 4% Q1 ramping to ~10% by Q4
return min(12.0, 4.0 + (q - 1) * 2.0)
def detect_leaky_cohorts(cohorts: list[CohortProjection]) -> None:
"""A cohort is leaky if its mean NRR is LEAK_THRESHOLD_PP below the average of older cohorts."""
if len(cohorts) < 2:
return
# Sort by acquisition_quarter string (lexicographic works for YYYY-Qn format)
ordered = sorted(cohorts, key=lambda c: c.acquisition_quarter)
for i, c in enumerate(ordered):
if i == 0:
continue
prior = ordered[:i]
prior_mean_nrr = statistics.mean(statistics.mean(p.nrr_by_quarter) for p in prior)
this_mean_nrr = statistics.mean(c.nrr_by_quarter)
gap = prior_mean_nrr - this_mean_nrr
if gap >= LEAK_THRESHOLD_PP:
c.leaky = True
c.leak_reason = (
f"Mean NRR {this_mean_nrr:.1f}% is {gap:.1f} pp below trailing-cohort avg "
f"{prior_mean_nrr:.1f}% (threshold: {LEAK_THRESHOLD_PP} pp)."
)
def consolidate(cohorts: list[CohortProjection], horizon_q: int) -> tuple[list[float], list[float], list[float]]:
cons_nrr: list[float] = []
cons_grr: list[float] = []
cons_arr: list[float] = []
for q_idx in range(horizon_q):
total_starting = sum(c.starting_arr for c in cohorts)
if total_starting <= 0:
cons_nrr.append(0.0); cons_grr.append(0.0); cons_arr.append(0.0)
continue
# ARR-weighted NRR + GRR
weighted_nrr = sum(c.starting_arr * c.nrr_by_quarter[q_idx] for c in cohorts) / total_starting
weighted_grr = sum(c.starting_arr * c.grr_by_quarter[q_idx] for c in cohorts) / total_starting
total_arr = sum(c.arr_by_quarter[q_idx] for c in cohorts)
cons_nrr.append(weighted_nrr)
cons_grr.append(weighted_grr)
cons_arr.append(total_arr)
return cons_nrr, cons_grr, cons_arr
def project(ctx: dict[str, Any]) -> ProjectionResult:
cohorts_in = ctx.get("cohorts") or []
horizon_q = int(ctx.get("projection_horizon_quarters") or 4)
projected = [project_cohort(c, horizon_q) for c in cohorts_in]
detect_leaky_cohorts(projected)
cons_nrr, cons_grr, cons_arr = consolidate(projected, horizon_q)
leaky = [c.cohort_id for c in projected if c.leaky]
assumptions = {
"projection_horizon_quarters": horizon_q,
"leak_threshold_pp": LEAK_THRESHOLD_PP,
"leak_rule": (
f"Cohort flagged leaky if mean NRR is ≥ {LEAK_THRESHOLD_PP} pp below "
"the mean of all earlier-acquired cohorts (Campbell/ProfitWell cohort decomposition discipline)."
),
"consolidation_method": "ARR-weighted (starting_arr) across cohorts per quarter",
"default_grr_curve_when_missing": "92% Q1 decaying ~1pp/quarter, floor 85%",
"default_expansion_curve_when_missing": "4% Q1 ramping +2pp/quarter, ceiling 12%",
}
return ProjectionResult(
cohorts=projected,
consolidated_nrr=cons_nrr,
consolidated_grr=cons_grr,
consolidated_arr=cons_arr,
horizon_q=horizon_q,
leaky_cohorts=leaky,
assumptions=assumptions,
)
def render_markdown(r: ProjectionResult) -> str:
L: list[str] = []
L.append("# Cohort ARR Projection")
L.append("")
L.append(f"**Horizon:** {r.horizon_q} quarters • **Cohorts:** {len(r.cohorts)} • **Leaky cohorts:** {len(r.leaky_cohorts)}")
L.append("")
if r.leaky_cohorts:
L.append("## Leaky-cohort callout")
L.append("")
L.append("> The consolidated NRR can stay flat while a recent cohort is leaking. Surfacing the leak now is 2-3 quarters cheaper than discovering it in the topline. (Campbell / Skok cohort decomposition.)")
L.append("")
for c in r.cohorts:
if c.leaky:
L.append(f"- ⚠️ **{c.cohort_id}** ({c.acquisition_quarter}): {c.leak_reason}")
L.append("")
else:
L.append("> No leaky cohorts detected at the configured threshold. Continue cohort decomposition every quarter; leaks emerge faster than you think.")
L.append("")
L.append("## Per-cohort NRR heatmap (% by projection quarter)")
L.append("")
header = "| Cohort | Acq Q | Starting ARR | " + " | ".join(f"Q+{q}" for q in range(1, r.horizon_q + 1)) + " |"
sep = "|---|---|---:|" + "---:|" * r.horizon_q
L.append(header)
L.append(sep)
for c in sorted(r.cohorts, key=lambda x: x.acquisition_quarter):
flag = " ⚠️" if c.leaky else ""
row = f"| {c.cohort_id}{flag} | {c.acquisition_quarter} | ,.0f | "
row += " | ".join(f"{n:.1f}%" for n in c.nrr_by_quarter)
row += " |"
L.append(row)
L.append("")
L.append("## Consolidated NRR / GRR trajectory")
L.append("")
L.append("| Quarter | Consolidated NRR | Consolidated GRR | Consolidated ARR |")
L.append("|---|---:|---:|---:|")
for q in range(r.horizon_q):
L.append(f"| Q+{q+1} | {r.consolidated_nrr[q]:.1f}% | {r.consolidated_grr[q]:.1f}% | ,.0f |")
L.append("")
L.append("## Assumption block (NON-OPTIONAL — present alongside the cohort heatmap)")
L.append("")
for k, v in r.assumptions.items():
L.append(f"- **{k}:** {v}")
L.append("")
L.append("## Next steps")
L.append("1. If a leaky cohort is flagged, decompose it: which segment / motion / pricing tier dominates that cohort?")
L.append("2. Cross-check against the bookings forecast — leaky cohort + flat commit number is a hidden mismatch.")
L.append("3. Present NRR with the cohort heatmap. Consolidated-only is theatre.")
return "\n".join(L)
def sample_context() -> dict[str, Any]:
return {
"cohorts": [
{
"cohort_id": "2025-Q1", "acquisition_quarter": "2025-Q1", "starting_arr": 1_200_000,
"gross_retention_pct_q1": 93, "gross_retention_pct_q2": 91, "gross_retention_pct_q3": 90, "gross_retention_pct_q4": 89,
"expansion_arr_pct_q1": 5, "expansion_arr_pct_q2": 8, "expansion_arr_pct_q3": 10, "expansion_arr_pct_q4": 11,
},
{
"cohort_id": "2025-Q2", "acquisition_quarter": "2025-Q2", "starting_arr": 1_500_000,
"gross_retention_pct_q1": 92, "gross_retention_pct_q2": 90, "gross_retention_pct_q3": 89, "gross_retention_pct_q4": 88,
"expansion_arr_pct_q1": 6, "expansion_arr_pct_q2": 9, "expansion_arr_pct_q3": 11, "expansion_arr_pct_q4": 12,
},
{
"cohort_id": "2025-Q3", "acquisition_quarter": "2025-Q3", "starting_arr": 1_800_000,
"gross_retention_pct_q1": 94, "gross_retention_pct_q2": 92, "gross_retention_pct_q3": 91, "gross_retention_pct_q4": 90,
"expansion_arr_pct_q1": 5, "expansion_arr_pct_q2": 8, "expansion_arr_pct_q3": 10, "expansion_arr_pct_q4": 12,
},
{
# LEAKY: recent cohort, low retention, low expansion
"cohort_id": "2025-Q4", "acquisition_quarter": "2025-Q4", "starting_arr": 2_100_000,
"gross_retention_pct_q1": 85, "gross_retention_pct_q2": 82, "gross_retention_pct_q3": 80, "gross_retention_pct_q4": 78,
"expansion_arr_pct_q1": 2, "expansion_arr_pct_q2": 3, "expansion_arr_pct_q3": 4, "expansion_arr_pct_q4": 5,
},
],
"projection_horizon_quarters": 4,
}
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--input", type=Path, help="Path to cohort-intake JSON.")
p.add_argument("--output", default="markdown", choices=["markdown", "json"], help="Output format.")
p.add_argument("--sample", action="store_true", help="Run with built-in sample context.")
args = p.parse_args(argv)
if args.sample:
ctx = sample_context()
elif args.input:
ctx = json.loads(args.input.read_text())
else:
p.error("Provide --input or --sample.")
return 2
result = project(ctx)
if args.output == "json":
out = {
"horizon_q": result.horizon_q,
"leaky_cohorts": result.leaky_cohorts,
"consolidated_nrr": [round(n, 2) for n in result.consolidated_nrr],
"consolidated_grr": [round(n, 2) for n in result.consolidated_grr],
"consolidated_arr": [round(n, 2) for n in result.consolidated_arr],
"assumptions": result.assumptions,
"cohorts": [
{
"cohort_id": c.cohort_id,
"acquisition_quarter": c.acquisition_quarter,
"starting_arr": c.starting_arr,
"nrr_by_quarter": [round(n, 2) for n in c.nrr_by_quarter],
"grr_by_quarter": [round(n, 2) for n in c.grr_by_quarter],
"arr_by_quarter": [round(n, 2) for n in c.arr_by_quarter],
"leaky": c.leaky,
"leak_reason": c.leak_reason,
}
for c in result.cohorts
],
}
print(json.dumps(out, indent=2))
else:
print(render_markdown(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/funnel_confidence_scorer.py
#!/usr/bin/env python3
"""funnel_confidence_scorer.py — per-stage CoV-based confidence bands with treatment recommendation.
Input: JSON with funnel_stages (each with stage_name and conversion_pct_history over 12 quarters).
For each stage, computes:
- Mean conversion %
- Standard deviation
- Coefficient of variation (CoV = StDev / Mean)
- Confidence band: HIGH (CoV < 10%), MEDIUM (10-25%), LOW (25-50%), VERY LOW (> 50%)
- Treatment recommendation per stage (commit-grade / soft-floor / extend-data-window / do-not-use)
The CoV discipline catches the case where two stages have the same mean conversion but very
different reliability — the same average masks very different forecast utility.
Deterministic. Stdlib only.
Usage:
funnel_confidence_scorer.py --input intake.json --output markdown
funnel_confidence_scorer.py --sample
"""
from __future__ import annotations
import argparse
import json
import statistics
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
@dataclass
class StageConfidence:
stage: str
history: list[float]
n: int
mean_pct: float
stdev_pct: float
cov_pct: float
band: str
treatment: str
rationale: list[str] = field(default_factory=list)
def classify_band(cov_pct: float) -> str:
if cov_pct < 10.0:
return "HIGH"
if cov_pct < 25.0:
return "MEDIUM"
if cov_pct < 50.0:
return "LOW"
return "VERY LOW"
def treatment_for_band(band: str, n: int) -> tuple[str, list[str]]:
rationale: list[str] = []
if n < 4:
rationale.append(f"Sample size n={n} is below the 4-quarter minimum for stable CoV estimation.")
return "extend-data-window", rationale
if band == "HIGH":
rationale.append("CoV < 10% — historically stable. Use as commit-grade conversion input.")
return "commit-grade", rationale
if band == "MEDIUM":
rationale.append("CoV 10-25% — usable but flagged. Apply blended last-4Q / last-12Q weighting.")
return "blended-weighting", rationale
if band == "LOW":
rationale.append("CoV 25-50% — high variance. Use as a soft floor only, never as commit input.")
return "treat-as-soft-floor", rationale
rationale.append("CoV > 50% — statistical noise. Do not use for forecasting; root-cause the variance first.")
return "do-not-use", rationale
def score_stage(stage_data: dict[str, Any]) -> StageConfidence:
stage = str(stage_data.get("stage_name", "?"))
history = [float(x) for x in (stage_data.get("conversion_pct_history") or []) if x is not None]
n = len(history)
if n == 0:
return StageConfidence(
stage=stage, history=[], n=0, mean_pct=0.0, stdev_pct=0.0, cov_pct=0.0,
band="UNKNOWN", treatment="extend-data-window",
rationale=["No conversion history provided."],
)
mean = statistics.mean(history)
stdev = statistics.pstdev(history) if n > 1 else 0.0
cov = (stdev / mean * 100.0) if mean > 0 else 0.0
band = classify_band(cov)
treatment, rationale = treatment_for_band(band, n)
if mean > 0:
rationale.insert(0, f"Mean {mean:.2f}% across {n} quarters; stdev {stdev:.2f}%; CoV {cov:.1f}%.")
return StageConfidence(
stage=stage, history=history, n=n, mean_pct=mean, stdev_pct=stdev,
cov_pct=cov, band=band, treatment=treatment, rationale=rationale,
)
def score_all(ctx: dict[str, Any]) -> list[StageConfidence]:
stages = ctx.get("funnel_stages") or []
return [score_stage(s) for s in stages]
def render_markdown(rows: list[StageConfidence]) -> str:
L: list[str] = []
L.append("# Funnel Confidence Scorer")
L.append("")
L.append(f"**Stages scored:** {len(rows)}")
L.append("")
L.append("## Confidence band summary")
L.append("")
L.append("| Stage | n quarters | Mean % | StDev % | CoV % | Band | Treatment |")
L.append("|---|---:|---:|---:|---:|:---:|---|")
for r in rows:
L.append(
f"| {r.stage} | {r.n} | {r.mean_pct:.2f} | {r.stdev_pct:.2f} | "
f"{r.cov_pct:.1f} | **{r.band}** | {r.treatment} |"
)
L.append("")
L.append("## Per-stage rationale")
L.append("")
for r in rows:
L.append(f"### {r.stage} — {r.band} ({r.treatment})")
for line in r.rationale:
L.append(f"- {line}")
L.append("")
L.append("## Confidence-band thresholds (assumption block)")
L.append("")
L.append("- **HIGH** — CoV < 10%. Commit-grade conversion input.")
L.append("- **MEDIUM** — CoV 10-25%. Use blended last-4Q / last-12Q weighting.")
L.append("- **LOW** — CoV 25-50%. Soft floor only; never a commit input.")
L.append("- **VERY LOW** — CoV > 50%. Statistical noise; root-cause before using.")
L.append("- **Min sample size** — 4 quarters for stable CoV; below that → extend-data-window.")
L.append("")
L.append("## Next steps")
L.append("1. For any stage flagged `do-not-use` or `treat-as-soft-floor`, decompose: segment? motion? rep? quarter-of-year seasonality?")
L.append("2. Feed HIGH and MEDIUM stages directly into `bookings_forecaster.py`. Exclude LOW and VERY LOW from commit.")
L.append("3. Present the per-stage confidence table on the same slide as the 3-tier forecast number.")
return "\n".join(L)
def sample_context() -> dict[str, Any]:
return {
"funnel_stages": [
{"stage_name": "discovery_to_demo", "conversion_pct_history": [
35, 37, 33, 36, 38, 35, 34, 37, 36, 35, 36, 37
]},
{"stage_name": "demo_to_proposal", "conversion_pct_history": [
55, 52, 58, 56, 54, 57, 53, 55, 58, 54, 56, 55
]},
{"stage_name": "proposal_to_negotiation", "conversion_pct_history": [
65, 60, 70, 55, 75, 50, 80, 45, 72, 58, 68, 62
]}, # high variance
{"stage_name": "negotiation_to_verbal", "conversion_pct_history": [
75, 73, 76, 74, 75, 77, 74, 76, 73, 75, 76, 74
]},
{"stage_name": "verbal_to_commit", "conversion_pct_history": [
85, 60, 90, 40, 95, 30, 88, 55, 92, 35, 87, 50
]}, # very high variance
],
}
def main(argv: list[str] | None = None) -> int:
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument("--input", type=Path, help="Path to funnel-history JSON.")
p.add_argument("--output", default="markdown", choices=["markdown", "json"], help="Output format.")
p.add_argument("--sample", action="store_true", help="Run with built-in sample context.")
args = p.parse_args(argv)
if args.sample:
ctx = sample_context()
elif args.input:
ctx = json.loads(args.input.read_text())
else:
p.error("Provide --input or --sample.")
return 2
rows = score_all(ctx)
if args.output == "json":
out = {
"stages": [
{
"stage": r.stage, "n": r.n, "mean_pct": round(r.mean_pct, 4),
"stdev_pct": round(r.stdev_pct, 4), "cov_pct": round(r.cov_pct, 2),
"band": r.band, "treatment": r.treatment, "rationale": r.rationale,
"history": r.history,
}
for r in rows
],
"thresholds": {
"HIGH": "CoV < 10",
"MEDIUM": "10 <= CoV < 25",
"LOW": "25 <= CoV < 50",
"VERY LOW": "CoV >= 50",
"min_sample_n": 4,
},
}
print(json.dumps(out, indent=2))
else:
print(render_markdown(rows))
return 0
if __name__ == "__main__":
sys.exit(main())
Thiết kế chính sách thương mại: ma trận chiết khấu, ngưỡng phê duyệt, luồng ngoại lệ và khung giao dịch cho Deal Desk.
---
name: commercial-policy
description: "Use when designing or revising a company's commercial policy — the rules of engagement governing discounts off list price, approver thresholds, exception flows, and the deal framework that Deal Desk and AEs operate under. Covers discount matrix design (ARR band x term length x payment terms x strategic value), commercial policy design, exception policy, discount governance, approval thresholds, deal framework structure, and policy linting (contradictions, gaps, cliff edges, gaming surfaces). For Head of Commercial, Head of Deal Desk, VP Sales, or RevOps at the policy-design moment — NOT per-deal application (that is deal-desk) and NOT pricing model selection (that is pricing-strategist)."
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [commercial, discount-policy, discount-matrix, exception-flow, governance, deal-framework, commercial-discipline]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# commercial-policy
## Purpose
Design the **rules of engagement** that govern discounting off list price — the artifact that Deal Desk and AEs operate under. Three deterministic tools:
1. `discount_matrix_builder.py` — builds a 4-dimensional matrix (ARR band × term length × payment terms × strategic value tier), each cell carrying an approved discount band backed by current win-rate + NRR data, plus an approver tier (AE / Manager / Director / VP / CFO).
2. `exception_router.py` — when an asks-for-discount lands outside the matrix, routes it through the named approver chain, attaches required compensating commitments (multi-year prepay + named expansion path + reference commitment + MSA tightening), produces machine-readable audit-trail metadata, and flags precedent risk if 3+ similar exceptions have landed in the trailing quarter.
3. `policy_linter.py` — lints the matrix for governance defects: approver inversion, band inversion, margin-floor violation, coverage gaps, cliff edges, undefined strategic tiers, inconsistent margin floors, thin data backing.
The output is the **policy itself** (matrix + exception flow + lint report), not a per-deal application of it.
## When to use
- A new Head of Commercial or Head of Deal Desk is writing the company's first formal commercial policy
- The existing matrix is older than 6 months and discount drift is showing in margin reviews
- Reps are citing "Maria approved 28% on Acme last quarter" as precedent and you need to break the precedent loop
- Q-over-Q exception count is rising and you suspect the matrix bands are mispriced
- CFO has tightened the margin floor and the matrix needs to be rebuilt against the new constraint
- A board / exec is asking "why do we discount this much?" and you need a data-backed defensible policy
**Do NOT use this skill to:**
- Approve a specific deal — that's `commercial/skills/deal-desk`
- Set the pricing model + list price — that's `commercial/skills/pricing-strategist`
- Author a proposal / SOW / MSA prose — that's `business-growth/contract-and-proposal-writer`
- Make the strategic "when do we hire a VP Sales" call — that's `c-level-advisor/cro-advisor`
## Workflow
1. **Audit current discount distribution.** Pull the last 4 quarters of closed-won + closed-lost deals from CRM. Fill `assets/policy_design_template.md` (~20 minutes). Capture: `arr`, `discount_pct`, `term_months`, `payment_terms_days`, `strategic_value`, `win_lost`, `nrr_12mo` per deal.
2. **Design the data-backed matrix.** Run `scripts/discount_matrix_builder.py --input policy_intake.json --profile {saas|enterprise-software|api|marketplace|services}`. Output is a 4-dimensional matrix with approved discount band + approver tier + margin floor + observed win-rate + observed NRR per cell. Cells with `n < 5` observed deals are flagged `THIN`.
3. **Design the exception flow.** Run `scripts/exception_router.py --sample` to see the structure. For each severity band of exception (0-5 pts over, 5-10, 10-20, 20+), the router enforces required compensating commitments. Codify the flow in your policy doc; the router becomes the operational implementation.
4. **Lint the matrix.** Run `scripts/policy_linter.py --input matrix.json`. Get a ranked findings report — BLOCKER / MAJOR / MINOR — across 10 lint rules. Resolve every BLOCKER before publishing the matrix to AEs.
5. **Publish + quarterly review.** Publish the matrix as a versioned artifact. Re-run the builder and the linter every quarter against the new 4-quarter rolling deal corpus. Cells where observed NRR < `target_nrr` are flagged for review.
## Scripts
| Script | Purpose | Industry profiles |
|---|---|---|
| `scripts/discount_matrix_builder.py` | 4-dim data-backed matrix with approver tiers + margin floors | saas, enterprise-software, api, marketplace, services |
| `scripts/exception_router.py` | Routes exception requests with compensating commitments + audit trail | n/a (matrix-driven) |
| `scripts/policy_linter.py` | 10-rule lint pass over the matrix | n/a (deterministic across profiles) |
All three: stdlib-only, `--help`, `--sample`, `--input <json>`, `--output {markdown,json}`.
## References
- `references/discount_governance_canon.md` — Discount governance evidence base: OpenView Partners benchmarks, David Skok (For Entrepreneurs) discount math, Tomasz Tunguz on discount distribution, Bessemer State of the Cloud, KeyBanc Capital Markets SaaS Survey, Bridge Group AE-compensation research, RevOps Co-op playbooks, Forrester deal-desk research. 8 sources.
- `references/policy_design_canon.md` — Policy-as-artifact design: SaaStr (Jason Lemkin), Winning by Design (Jacco van der Kooij) on commercial discipline, Forrester deal-desk maturity research, MIT Sloan on incentive-system gaming, McKinsey on commercial-policy effectiveness, Bain *Pricing Power*, Salesforce CPQ implementation guides. 7 sources.
- `references/policy_anti_patterns.md` — 8 named anti-patterns with sourced studies + countermeasures + lint-rule mapping: precedent-sets-policy, no-data-backing, no-compensating-commitments, approver/margin misalignment, no audit trail, cliff edges, undefined "strategic value", no quarterly review. 8 sources.
## Assumptions
- The skill assumes the **pricing model and list price already exist** (set via `commercial/skills/pricing-strategist`). Commercial-policy governs **discounts off list** — it does not set list.
- The CFO owns the `min_margin_pct` constraint (margin floor). The CRO / Head of Deal Desk owns the `max_discount_pct_without_exception` constraint (band cap). The skill keeps these inputs separate by design (per Bain *Pricing Power* — mixing accountability is the most common cause of policy drift).
- Industry profiles bake in *customary* band widths. Companies with idiosyncratic economics should pass overrides via the input JSON.
- The matrix is data-backed but **not data-driven**: the band is set by the constraints + profile; observed data is annotation that tells you whether the cell is performing. If observed NRR < target, that's a signal to **review the band**, not to keep discounting deeper.
- "Strategic value" tiers (`logo`, `expansion`, `lighthouse`) are useful only if defined with concrete tests. The lint rule L06 enforces this.
- This is a policy-design skill, not a deal-approval skill. It never says "approve" — it produces the matrix + exception flow that **deal-desk** then applies.
## Anti-patterns
- **Setting discount bands without data backing.** "VP Sales argued for it in a Slack thread" is not data backing. If you can't show win-rate and NRR for the band, the band is rhetoric. (Caught by `data_backing` per cell + lint L08.)
- **Letting precedent set policy.** "Maria approved 28% on Acme last quarter" is not a band — it's an exception that didn't break the policy. `exception_router.py` flags 3+ similar exceptions as a signal that **the matrix is wrong**, not the deal. (Anti-pattern AP-1.)
- **Approving exceptions without compensating commitments.** Discount-for-nothing is a leak (Winning by Design). Every exception severity band requires non-negotiable commitments. (`exception_router.COMPENSATING_LIBRARY`.)
- **Cliff edges at round-number ARR thresholds.** A hard $100K threshold produces deal-size gaming within 2 quarters (MIT Sloan agency theory). Smooth the gradient. (Lint L05.)
- **"Strategic value" as an undefined catch-all.** If "strategic" is undefined, within a quarter 60% of deals will be flagged strategic and the matrix is dead. Define with concrete tests. (Lint L06.)
- **No quarterly review.** Markets shift; matrices unchanged for 12 months are mispriced. Re-run the builder and linter every quarter. (Anti-pattern AP-8.)
- **Mixing CFO and CRO accountabilities.** CFO owns the margin floor; CRO owns the band cap. Same accountable owner = predictable drift toward whatever they're compensated on (Bain *Pricing Power*).
- **Skipping the lint pass before publishing.** BLOCKER findings (approver inversion, margin-floor violation, inverted bands) make the policy unsignable. Lint is the gate, not the after-action review.
## Distinct from
| Sibling | Scope | Difference |
|---|---|---|
| `commercial/skills/deal-desk` | **Applies** the policy to one deal at a time | Commercial-policy **designs the policy itself**. Deal-desk consumes the matrix; commercial-policy produces it. |
| `commercial/skills/pricing-strategist` | Sets pricing **model** (per-seat / usage / value / tiered) + **list price** | Commercial-policy governs **discounts off list**. Pricing-strategist sets the menu; commercial-policy governs the menu's discount discipline. |
| `c-level-advisor/cro-advisor` | Strategic CRO judgment ("when do we hire VP Sales?", "is our motion product-led or sales-led?") | Strategic, not operational. Commercial-policy is the artifact CRO commissions; it isn't CRO judgment itself. |
| `c-level-advisor/cfo-advisor` | Margin floor + unit-economics judgment | The CFO supplies `min_margin_pct` to commercial-policy as an input. Commercial-policy **operationalizes** the CFO's constraint as per-cell margin floors. |
| `business-growth/contract-and-proposal-writer` | Authors proposal/SOW/MSA **prose** | Commercial-policy emits structured matrix + audit-trail JSON, not customer-facing prose. |
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-commercial` or the Commercial orchestrator before the skill runs. Recommended answer + canon citation per question. Never bundled.
1. **"What's your observed discount distribution across the last 4 quarters — and is the median inside or outside your current matrix?"**
Recommended: pull the corpus before designing any band. If the observed median is outside the matrix, the matrix is rhetoric.
Canon: OpenView SaaS Benchmarks; RevOps Co-op playbooks. Anti-pattern AP-2.
2. **"What's the win-rate AND the 12-month NRR for deals at your current 'max discount' band?"**
Recommended: both, not one. A band with high win-rate but low NRR is buying logos with leaky-bucket retention. Tunguz benchmarks: top-NRR-quartile companies discount 6 pts less than bottom quartile.
Canon: Tomasz Tunguz; Bessemer State of the Cloud.
3. **"Who at the company owns the margin floor, AND who owns the discount-band cap — are those the same person?"**
Recommended: CFO owns floor; CRO/Head of Deal Desk owns cap. Same owner = drift toward what they're compensated on.
Canon: Bain *Pricing Power* — separation of accountability is the structural fix. Anti-pattern AP-4.
4. **"How is 'strategic value' defined in your current policy — with concrete tests, or with adjectives?"**
Recommended: concrete tests. "Top-20 named account in 2026 target list" is a test; "important customer" is not.
Canon: SaaStr (Lemkin); Forrester deal-desk research. Lint rule L06. Anti-pattern AP-7.
5. **"For exceptions above your matrix max, what compensating commitments are required — and are they in writing before the approver signs?"**
Recommended: minimum multi-year prepay + named expansion path; deeper exceptions require reference commitment + MSA tightening + executive sponsor.
Canon: Winning by Design (van der Kooij); McKinsey B2B pricing studies. Anti-pattern AP-3.
6. **"Has the same kind of exception been approved 3+ times in the trailing quarter — and if so, is the matrix wrong?"**
Recommended: 3+ similar exceptions means the band is mispriced. Rebuild the matrix; don't keep approving exceptions.
Canon: OpenView discount drift studies; `exception_router._precedent_risk`. Anti-pattern AP-1.
7. **"When was the last time you re-ran the matrix against the previous 4 quarters of data?"**
Recommended: quarterly. Annual review is too slow; the disciplined cohort revises quarterly.
Canon: OpenView benchmarks; RevOps Co-op. Anti-pattern AP-8.
8. **"For every exception in the last quarter, is there a machine-readable audit-trail record — or is the approval in Slack and email?"**
Recommended: structured record in CPQ or equivalent. Slack/email approvals don't survive year-2 renewal negotiations.
Canon: Salesforce CPQ best practices; Forrester deal-desk maturity research. Anti-pattern AP-5.
Walk depth-first. Lock 1-4 before opening 5-8. After all 8 are answered, invoke `discount_matrix_builder.py` → `policy_linter.py` → `exception_router.py --sample` in sequence to produce the policy artifact.
## Quick examples
```bash
# Design the matrix
python3 scripts/discount_matrix_builder.py --sample
python3 scripts/discount_matrix_builder.py --input policy_intake.json --profile saas --output json > matrix.json
# Lint the matrix
python3 scripts/policy_linter.py --sample
python3 scripts/policy_linter.py --input matrix.json
# Walk the exception flow
python3 scripts/exception_router.py --sample
python3 scripts/exception_router.py --input request.json --output json
```
The sample matrix lints to **FAIL** with 4 BLOCKERs + 6 MAJORs + 2 MINORs — by design, to exercise every rule path. A real policy intake should lint to PASS or PASS_WITH_WARNINGS. The sample exception (42% on a $320K logo deal) routes to AE → Sales Manager → Director → VP Sales with 3 required compensating commitments (multi-year 36mo, prepay, named expansion path).
FILE:assets/policy_design_template.md
# Commercial Policy Design — Intake
**Time to fill out: ~20 minutes.** Output of this intake feeds directly into the three skill scripts:
- `discount_matrix_builder.py` ← Section 4 (current deals) + Section 5 (constraints) + Section 6 (industry)
- `exception_router.py` ← Section 7 (exception flow) + audit trail spec
- `policy_linter.py` ← runs against the matrix output once built
Re-pricings or major matrix revisions create a *new* intake — do not edit in place. Version the intake the same way you version the matrix.
---
## 1. Policy owner
| Field | Value |
|---|---|
| Head of Deal Desk / Commercial owner | |
| CFO sign-off contact | |
| CRO / VP Sales sign-off contact | |
| GC / legal contact for exceptions | |
| Target publish date | |
| Version | v1.0.0 |
## 2. Scope
- [ ] New-business discounts
- [ ] Renewal discounts
- [ ] Expansion/upsell discounts
- [ ] Partner/channel-sourced discounts
- [ ] Multi-product bundle discounts
Anything unchecked is **out of scope** for this matrix.
## 3. Industry profile
Pick one (drives the `--profile` flag and tunes the base band widths):
- [ ] `saas` — subscription seat-based or hybrid; typical product GM 75-85%
- [ ] `enterprise-software` — large ACVs; longer cycles; multi-year norm
- [ ] `api` — usage-based; tight bands; consumption-led
- [ ] `marketplace` — take-rate model; thinnest bands
- [ ] `services` — labor-bound; aggressive escalation on small discounts
## 4. Current deal corpus (data backing)
Pull from CRM the **last 4 quarters of closed-won + closed-lost** deals. Aim for n ≥ 50, n ≥ 200 preferred. Each row:
| Field | Notes |
|---|---|
| `arr` | Annual recurring revenue, USD |
| `discount_pct` | Discount taken off list, 0-100 |
| `term_months` | Contract term in months |
| `payment_terms_days` | NET-30 / NET-45 / NET-60 / etc. |
| `strategic_value` | one of: `standard`, `logo`, `expansion`, `lighthouse` |
| `win_lost` | `win` or `lost` |
| `nrr_12mo` | 12-month NRR for the cohort that signed (for closed-won; 0 for closed-lost) |
Save as JSON, populate the `current_deals` array in the intake JSON below.
## 5. Target constraints
| Field | Value | Sourced from |
|---|---|---|
| `min_margin_pct` | | CFO — the gross margin floor below which NO cell can publish |
| `max_discount_pct_without_exception` | | CRO / Head of Deal Desk — the cap above which every deal becomes an exception |
| `target_nrr` | | CFO/CRO — the NRR target the policy is designed to protect |
These three numbers are non-negotiable inputs. The matrix builder will respect them; cells that can't satisfy them will be flagged for explicit exception treatment.
## 6. Strategic-value definitions (REQUIRED — anti-pattern AP-7)
If you use any tier above `standard`, you must define it with **concrete tests**. Vague definitions get flagged by `policy_linter.py` rule L06.
| Tier | Definition (must be testable) | Example |
|---|---|---|
| `standard` | Default. No special strategic claim. | Any deal not meeting one of the below |
| `logo` | Reference-quality customer name | Top-20 named target list for 2026 GTM motion |
| `expansion` | Signed expansion path | MSA includes named BU or product-line expansion within 12 months |
| `lighthouse` | Co-marketed reference + multi-year | Public case study + 2 reference calls/year + 36-month term |
Without `strategic_value_definitions_supplied=true` in the matrix JSON, the linter will reject the matrix.
## 7. Exception flow spec
For exception requests (discount > `max_discount_pct_without_exception`):
- [ ] Required: structured submission (no Slack/email)
- [ ] Required: written justification
- [ ] Required: named approver chain (no role-only approvals)
- [ ] Required: compensating commitments per severity band (per `exception_router.COMPENSATING_LIBRARY`)
- [ ] Required: precedent-risk check across trailing 90 days
- [ ] Required: audit-trail JSON persisted to system of record (CPQ or equivalent)
Severity tiers (severity = `requested_discount` − `max_without_exception`):
| Severity range | Minimum compensating commitments |
|---|---|
| 0-5 pts over | multi-year term + annual prepay |
| 5-10 pts over | + named expansion path in writing |
| 10-20 pts over | + reference commitment + MSA tightening |
| 20+ pts over | + executive sponsor + co-marketing + kill-switch on expansion target |
## 8. Quarterly review trigger
| Check | Owner | Cadence |
|---|---|---|
| Re-pull current deals corpus; re-run `discount_matrix_builder.py` | Head of Deal Desk | Quarterly |
| Re-run `policy_linter.py` on current matrix | Head of Deal Desk | Quarterly |
| Review cells flagged `meets_target_nrr=false` | CFO + CRO | Quarterly |
| Review cells flagged `thin_data_flag=true` | Head of Deal Desk | Bi-quarterly |
| Review precedent-risk flags from `exception_router.py` | Head of Deal Desk + CRO | Quarterly |
---
## JSON skeletons
### `policy_intake.json` (feeds `discount_matrix_builder.py`)
```json
{
"industry": "saas",
"current_deals": [
{
"arr": 0,
"discount_pct": 0,
"term_months": 12,
"payment_terms_days": 30,
"strategic_value": "standard",
"win_lost": "win",
"nrr_12mo": 1.0
}
],
"target_constraints": {
"min_margin_pct": 70.0,
"max_discount_pct_without_exception": 35.0,
"target_nrr": 1.15
}
}
```
### `exception_request.json` (feeds `exception_router.py`)
```json
{
"exception_request": {
"deal_id": "",
"requested_by": "",
"deal_arr": 0,
"requested_discount": 0,
"term_months": 0,
"payment_terms_days": 30,
"justification": "",
"strategic_value": "standard",
"customer_threats": [],
"submitted_at": ""
},
"policy_matrix": {
"profile": "saas",
"max_discount_pct_without_exception": 35.0,
"approver_thresholds": [
[15, "AE"], [25, "Sales Manager"], [35, "Director"], [50, "VP Sales"], [100.1, "CFO + CRO"]
]
},
"recent_exceptions": []
}
```
### `matrix.json` (output of `discount_matrix_builder.py`, input to `policy_linter.py`)
The linter expects the matrix shape emitted by the builder — `profile`, `constraints`, `cells[]` with the per-cell fields. Add the top-level boolean `strategic_value_definitions_supplied: true` once you've published the definitions from Section 6.
---
## 20-minute workflow
1. (~3 min) Fill Section 1 + Section 2 + Section 3.
2. (~6 min) Pull the deal corpus from CRM, format into `current_deals[]` JSON.
3. (~2 min) Fill Section 5 — get the three numbers from CFO + CRO.
4. (~5 min) Write Section 6 strategic-value definitions with concrete tests.
5. (~2 min) Confirm Section 7 exception flow with Head of Deal Desk.
6. (~2 min) Run the three scripts in order, lock the matrix, publish.
FILE:references/discount_governance_canon.md
# Discount Governance Canon
Authoritative sources on **how mature SaaS companies govern discounts off list price** — the rules of engagement that the commercial-policy skill operationalizes. Cite these in any policy doc this skill produces.
The unifying claim across every source below: **discount discipline correlates more strongly with retention and gross margin expansion than top-of-funnel velocity.** Bands aren't conservative for the sake of it — they protect the LTV math that funds the next year of GTM.
---
## 1. OpenView Partners — Annual SaaS Benchmarks (2018-2025)
OpenView's annual State of the SaaS Industry survey publishes discount distributions by ARR band and growth stage. Two consistent findings across 7 years:
- **Median enterprise discount = 18–22% off list.** Anything above 30% is the top decile and correlates with weaker NRR (typically 8–12 pts lower than disciplined peers).
- **The top quartile on net dollar retention discounts ~6 pts less than the bottom quartile.** Less discount, more retention — the leaky-bucket effect of "buying logos" with deep discounts shows up at renewal.
**Cite this for:** the empirical floor on what a "normal" discount band looks like across the SaaS industry. If your band exceeds 30% for non-strategic deals, you're outside the disciplined cohort.
URL: https://openviewpartners.com/blog/saas-benchmarks/
---
## 2. David Skok — For Entrepreneurs ("Discount Math")
Skok's canonical post on discount math shows that a percentage discount off list price erodes margin **more than proportionally**:
> A 30% discount on an 80% gross-margin product reduces margin by **37.5%**, not 30%. The discount is taken before the cost of goods sold is subtracted, so each percentage of discount removes a larger percentage of gross margin.
He further argues that the LTV impact compounds: discounted customers tend to expand less (lower NRR) and churn earlier (lower retention). The compound effect on LTV/CAC is often 2-3× the headline discount percentage.
**Cite this for:** the margin-floor calculation in `discount_matrix_builder.py`. The skill's per-cell `margin_floor_pct` enforces a hard floor below which no cell can publish a discount band.
URL: https://www.forentrepreneurs.com/
---
## 3. Tomasz Tunguz — Discount Distribution Studies (Redpoint)
Tunguz has published multiple analyses of discount distribution across enterprise SaaS deals (using anonymized Redpoint portfolio data). Three structural findings:
- **End-of-quarter discounts are 7-10 pts deeper than mid-quarter** across every ARR band. This is a forecast-pressure artifact, not a customer-value signal.
- **Deals closing in the last week of a quarter have NRR 4-6 pts lower at year 1** than deals closing in week 1-11.
- **Logo discounts that aren't accompanied by a written expansion commitment** show no NRR premium over standard discounts — the strategic value never materializes.
**Cite this for:** the "named expansion path in writing" compensating commitment in `exception_router.py`. Tunguz's data is the empirical reason verbal expansion promises aren't enough.
URL: https://tomtunguz.com/
---
## 4. Bessemer Venture Partners — State of the Cloud (annual)
BVP's State of the Cloud report (2020-2026) tracks discount and retention by cohort. Key claims this skill leans on:
- **Companies with formal discount matrices have NRR 8-15 pts higher** than peers with ad-hoc approval.
- **"Approver-of-record" governance** (every discount tied to a named human, not a role) reduces discount creep year-over-year by ~50%.
- The "Rule of 40" companies (growth + margin > 40%) consistently sit in the bottom quartile on discount depth.
**Cite this for:** the requirement that every cell in the matrix carry a named `approver_tier`, and that exceptions produce an `audit_trail` block with `requested_by` and `approver_chain` recorded.
URL: https://www.bvp.com/atlas/state-of-the-cloud-2025
---
## 5. KeyBanc Capital Markets — Annual SaaS Survey (formerly Pacific Crest)
KeyBanc's annual private-SaaS survey (~400 respondents) consistently publishes payment-terms and term-length data. Two findings the matrix encodes:
- **Every 15 days of payment terms adds ~2% to effective deal value.** NET-60 vs NET-30 is worth ~4% — so a customer asking for NET-60 plus 30% discount is asking for ~34% effective discount.
- **Multi-year prepay deals carry ~3-5 pts of NRR premium** over annual auto-renew, even at higher discount levels, because the cash and the commitment lock retention.
**Cite this for:** the `payment_penalty` and `term_bonus` parameters in `discount_matrix_builder.py`. NET-60 carries a penalty; multi-year prepay carries a bonus.
URL: https://key.com/businesses-institutions/industries-expertise/technology.jsp
---
## 6. Bridge Group — SaaS AE Compensation & Approval Research
Bridge Group's annual benchmark study of SaaS sales orgs publishes approver-chain practices. Two structural findings:
- **AEs allowed to self-approve discounts > 15% show 30%+ year-over-year discount creep.** Self-approval normalizes deeper discounts; AEs anchor on what they themselves approved last quarter.
- **Named-human approval reduces precedent drift by 50%+** vs. role-only approval. "VP Sales approves" is structurally weaker than "Maria Singh, VP Sales, approved on date X with these compensating commitments".
**Cite this for:** the audit-trail metadata block in `exception_router.py`, and the explicit `requested_by` field. The lint rule L09 (`cell_unreviewed`) is downstream of Bridge's finding that unobserved bands drift.
URL: https://bridgegroupinc.com/sales-research/
---
## 7. RevOps Co-op — Policy Design Playbooks
The RevOps Co-op community (Rosalyn Santa Elena, Jeff Ignacio, others) has published several playbooks on commercial-policy design. Three principles the skill enforces:
- **Discount bands must be backed by win-rate AND retention data**, not by sales leadership's negotiating room. If you can't show "at this band, we win X% and retain at NRR Y", the band is rhetoric.
- **Every exception must produce written compensating commitments** before the approver signs. "Strategic" isn't enough — what specifically does the customer commit to, in writing?
- **Quarterly policy review is non-optional.** Markets shift, competitors shift, customer mix shifts — a matrix unchanged for 12 months is almost certainly mispriced in some band.
**Cite this for:** the `data_backing` field per cell in `discount_matrix_builder.py` and the `required_compensating_commitments` block in `exception_router.py`. Lint rule L08 (thin data in critical cell) operationalizes RevOps Co-op's first principle.
URL: https://www.revopscoop.com/
---
## 8. Forrester — Deal Desk & Commercial Policy Research
Forrester's Deal Desk research (Mary Shea, Anthony McPartlin, Bob Apollo) consistently finds that companies with **formalized, data-backed commercial policy** outperform peers on three metrics:
- Cycle time (faster approvals when policy is clear)
- Win rate (AEs don't waste time on deals outside policy)
- Renewal margin (discounts at sign predict renewal economics)
The Forrester model treats commercial policy as a **product** that the RevOps team ships and maintains — not a memo that lives in the CFO's drawer.
**Cite this for:** the framing of commercial-policy as a designed artifact (with the lint pass), versus a precedent that accumulates through deal-by-deal exceptions.
URL: https://www.forrester.com/research/
---
## Synthesis: how the canon maps to this skill
| Canon source | Maps to |
|---|---|
| OpenView discount benchmarks | `base_max_pct` defaults in `PROFILES` |
| Skok discount math | `margin_floor_pct` enforcement per cell + lint L03 |
| Tunguz expansion-commitment data | `named_expansion_path` compensating commitment |
| BVP discount discipline | `approver_tier` per cell + audit trail |
| KeyBanc payment-terms data | `payment_penalty` and `term_bonus` parameters |
| Bridge Group AE-approval research | `requested_by` + audit trail metadata |
| RevOps Co-op playbooks | `data_backing` per cell + quarterly review hook |
| Forrester deal-desk research | The skill's existence — policy as designed artifact |
FILE:references/policy_anti_patterns.md
# Policy Anti-Patterns
Eight named anti-patterns that the commercial-policy skill is built to prevent. Each is observed in the wild (with sourced studies), each has a concrete countermeasure encoded in the skill's tools, and each maps to a lint rule or a forcing question.
The unifying claim: **discount policy drifts by mechanism, not by malice.** The job of the skill is to make the drift mechanism visible so leadership can decide whether to accept it.
---
## AP-1: Precedent sets policy — "Maria approved 28% on Acme last Q"
**Pattern.** An AE cites a previous exception as precedent for a new deal. Three exceptions in a quarter become the new normal. The matrix on paper says 25%; the operational floor is 32%.
**Why it's seductive.** AEs are anchored to the most recent approved discount, not the policy band. Sales managers are anchored to their own past approvals because reversing would be a tacit admission of error.
**Evidence.** OpenView discount-benchmark data shows companies without a formal precedent-breaking mechanism drift +3-5 pts per year. Tunguz's Redpoint data shows ~50% of "strategic exceptions" never produce the strategic value claimed at sign — but the discount sticks.
**Countermeasure in skill.** `exception_router.py` runs a `_precedent_risk` check: if 3+ similar exceptions in the trailing quarter, the verdict is `PRECEDENT_RISK FLAGGED` and the matrix itself is recommended for rebuild. The deal isn't the problem; the band is.
**Lint rule.** None — this is a flow-level check, not a matrix defect.
---
## AP-2: No data backing for discount bands
**Pattern.** A discount band is set because "feels about right" or because the VP Sales argued for it in a Slack thread. There's no win-rate or NRR data showing the band actually wins deals at the rate claimed or retains them at the NRR claimed.
**Why it's seductive.** Setting the band by feel is fast. Building the data infrastructure to back it is slow and exposes uncomfortable findings (e.g., "our 35% band has 15% lower NRR than the 20% band").
**Evidence.** RevOps Co-op playbooks consistently identify "policy designed without retention data" as the #1 cause of margin erosion in years 2-3 post-launch. Bessemer's State of the Cloud benchmarks the gap: policies with retention backing show NRR 8-15 pts higher.
**Countermeasure in skill.** `discount_matrix_builder.py` requires `current_deals[]` as input and emits a `data_backing` block per cell showing `n_observed_deals`, `win_rate`, `nrr_12mo_observed`. Cells with `n < 5` are flagged `THIN`.
**Lint rule.** L08 (`thin_data_in_critical_cell`) — fires for enterprise/strategic cells with thin data.
---
## AP-3: No compensating commitments required for exception discount
**Pattern.** An AE asks for 40% (above the 35% policy max). VP Sales approves via email. No multi-year prepay, no expansion path, no reference commitment, no MSA tightening. The customer banks the discount and gives nothing structural back.
**Why it's seductive.** Asking for commitments slows the deal. At quarter end, the AE and the VP both prefer the path of least resistance.
**Evidence.** Winning by Design (van der Kooij) frames this as the "discount-for-nothing leak": the single highest-leverage place to find margin in a mature GTM. McKinsey B2B pricing studies find that capturing compensating commitments on exceptions alone returns 1-2 pts of margin annually.
**Countermeasure in skill.** `exception_router.py` populates `required_compensating_commitments[]` for any non-in-policy request, scaled by severity (deeper exception → more commitments).
**Lint rule.** L10 (`missing_exception_marker`) — fires when a high-discount cell exists without `exception_required=true`, which would route it through the router.
---
## AP-4: Approver tiers misaligned with margin floor
**Pattern.** Sales Manager is authorized to approve discounts up to a cap that produces margins below the CFO-set floor. The CFO never sees the deal because the chain stops at the manager. By the time the CFO learns about it (in the quarterly margin review), 12 deals are already signed.
**Why it's seductive.** Aligning approver tiers with margin floors requires the CFO, CRO, and Head of Deal Desk to agree on numbers — which is hard.
**Evidence.** Bain's *Pricing Power* research identifies this as the single most common policy defect in mid-market SaaS. The fix is structural: the CFO must own the margin floor; that floor must show up as a per-cell field in the matrix.
**Countermeasure in skill.** `discount_matrix_builder.py` derives `margin_floor_pct` per cell from the input `target_constraints.min_margin_pct`, and surfaces it next to the approver tier.
**Lint rule.** L03 (`margin_floor_below_constraint`) — fires when any cell falls below 50% margin floor.
---
## AP-5: No audit trail for exceptions
**Pattern.** An exception is approved by Slack DM or email. No timestamp, no structured justification, no record of the compensating commitments. Six months later, the customer asks for the same discount at renewal — and no one can find the original commitments.
**Why it's seductive.** Slack and email are faster than CPQ or a structured form. At quarter end, structure feels like friction.
**Evidence.** Salesforce CPQ implementation guides cite this as the #1 reason commercial-policy efforts fail in years 2-3. Forrester's deal-desk maturity model puts "machine-readable audit trail" at the boundary between level 2 (formalized) and level 3 (operationalized).
**Countermeasure in skill.** `exception_router.py` emits a structured `audit_trail` block: `deal_id`, `requested_by`, `submitted_at`, `justification`, `compensating_commitments_required`, `approver_chain`. The block is JSON, so it can be persisted to CPQ or a deal-desk system.
**Lint rule.** None — flow-level, not matrix-level.
---
## AP-6: Cliff edges at round-number ARR thresholds
**Pattern.** Policy says: ARR ≥ $100K → enterprise band (up to 30% discount). ARR < $100K → mid band (up to 22% discount). An AE working a $98K deal pads it to $100K to access the deeper band. Or splits a $105K deal into two $52.5K deals to dodge approval.
**Why it's seductive.** Round-number thresholds are easy to remember and easy to write into policy. The gaming surface is invisible until you look at the deal distribution and notice an unnatural cluster at $100,001.
**Evidence.** MIT Sloan agency-theory literature (Holmström, Gibbons) on multitask gaming. The practical evidence in SaaS: any policy with a hard cliff produces a visible bimodal distribution of deal sizes around the cliff within 2-3 quarters.
**Countermeasure in skill.** Bands in the matrix are smoothed by adjacent strategic-tier bonuses, term bonuses, and payment penalties — so the maximum discount changes gradually rather than cliffing.
**Lint rule.** L05 (`cliff_edge`) — fires when adjacent cells differ by > 10 pts on the discount max.
---
## AP-7: "Strategic value" undefined → catch-all for any discount
**Pattern.** The policy includes a "strategic value" override that allows AEs to exceed the band. "Strategic" is undefined or defined vaguely ("important customer"). Within a quarter, 60% of deals are flagged strategic and the matrix has been rendered meaningless.
**Why it's seductive.** Defining "strategic" with concrete tests requires the GTM leadership team to write down which customers count and which don't — a politically expensive exercise.
**Evidence.** SaaStr (Lemkin) covers this as one of the top-three policy failures. Forrester deal-desk research cites it as the #1 cause of "operationalized" policies sliding back to "formalized."
**Countermeasure in skill.** The matrix has explicit strategic tiers (`standard`, `logo`, `expansion`, `lighthouse`). The user must supply `strategic_value_definitions_supplied=true` plus tests; if not, the lint flags it.
**Lint rule.** L06 (`strategic_value_undefined`) — fires when strategic tiers are used without verifiable definitions.
---
## AP-8: No quarterly policy review based on win-rate data
**Pattern.** The matrix is published, AEs are trained, the policy is declared "live" — and then nobody touches it for 18 months. Meanwhile competitive pricing, customer mix, and product economics shift. The matrix is now wrong in 30-50% of cells, and nobody knows which ones.
**Why it's seductive.** A live policy is a finished policy. Revisiting it implies the previous version was wrong, which is politically awkward.
**Evidence.** OpenView discount-benchmark research shows the disciplined-cohort companies revise their matrix quarterly. The undisciplined cohort revises annually or less, and shows margin drift of -2 to -4 pts per year. RevOps Co-op community studies replicate the finding.
**Countermeasure in skill.** The matrix is a versioned artifact. Each cell's `data_backing` block surfaces the empirical win-rate and NRR; cells where observed NRR < `target_nrr` are flagged `meets_target_nrr=false`, signaling cells due for review.
**Lint rule.** L09 (`cell_unreviewed`) — fires when a cell has zero observed deals (i.e., nobody has tested the band yet).
---
## Synthesis: the 8 anti-patterns and where they're caught
| # | Anti-pattern | Caught by | Lint rule |
|---|---|---|---|
| AP-1 | Precedent sets policy | `exception_router._precedent_risk` | — |
| AP-2 | No data backing | `discount_matrix_builder.data_backing` per cell | L08 |
| AP-3 | No compensating commitments | `exception_router.COMPENSATING_LIBRARY` | L10 |
| AP-4 | Approver/margin misalignment | per-cell `margin_floor_pct` next to approver | L03 |
| AP-5 | No audit trail | `exception_router.audit_trail` JSON block | — |
| AP-6 | Cliff edges | smoothed bands in matrix builder | L05 |
| AP-7 | Strategic value undefined | `strategic_value_definitions_supplied` flag | L06 |
| AP-8 | No quarterly review | `data_backing.n_observed_deals` per cell | L09 |
## Sources (8)
1. OpenView Partners — Annual SaaS Benchmark Survey (2018-2025): https://openviewpartners.com/blog/saas-benchmarks/
2. Tomasz Tunguz — Discount Distribution Studies (Redpoint blog): https://tomtunguz.com/
3. MIT Sloan — Robert Gibbons / Bengt Holmström agency-theory papers: https://mitsloan.mit.edu/faculty/directory/robert-gibbons
4. SaaStr (Jason Lemkin) — Discount Policy + Strategic-Value Posts: https://www.saastr.com/
5. Winning by Design (Jacco van der Kooij) — *Revenue Architecture*: https://winningbydesign.com/
6. Forrester — Deal Desk Maturity Research: https://www.forrester.com/research/
7. RevOps Co-op — Community Policy Design Playbooks: https://www.revopscoop.com/
8. Bain — *Pricing Power* + Discount Discipline Studies: https://www.bain.com/insights/topics/pricing/
FILE:references/policy_design_canon.md
# Policy Design Canon
Authoritative sources on **how to design a commercial policy as an artifact** — not how to discount, but how to write the document that governs discounting. The seven sources below ground the *structure* the skill emits (matrix + exception flow + lint).
The shared insight: a policy is only as good as the gaming surface it removes. Cliffs, ambiguous strategic-value definitions, and missing approver tiers are not stylistic flaws — they are gaming surfaces that AEs and customers will discover within one quarter.
---
## 1. SaaStr (Jason Lemkin) — Deal Policy Structure
Lemkin's SaaStr corpus on deal policy makes one structural argument repeatedly: **the policy must be writable on a single page that AEs can scan in the deal room.** If the policy needs a six-page memo to operate, no AE will follow it under quarter-close pressure.
Concrete practices:
- One discount matrix, one exception flow, one approver table. Three artifacts, max.
- Approver chains stop at the **lowest-authority hop that can sign** — not "escalate to CFO every time." Over-escalation trains AEs to over-discount because they assume the chain will accept whatever they propose.
- "Strategic value" must be defined with **concrete tests**, not adjectives. "Top-20 named account in 2026 target list" is a test; "important customer" is not.
**Cite this for:** the single-table matrix output of `discount_matrix_builder.py` and the lint rule L06 (`strategic_value_undefined`).
URL: https://www.saastr.com/
---
## 2. Winning by Design (Jacco van der Kooij) — Commercial Discipline
Van der Kooij's *Revenue Architecture* and the Winning by Design blueprints frame commercial policy as one of the **four operating systems** that govern recurring revenue (alongside ICP, motion, and metrics). Two principles the skill enforces:
- **Discount is a tool, not a verb.** Every discount must trade for something the customer commits to in writing — term length, prepay, expansion, reference. Discount-for-nothing is a leak.
- **The policy must distinguish "concession" from "investment"** — a strategic discount that pays back via expansion is an investment; a year-end discount that buys forecast is a concession. Investments get logged on the strategic-value tier; concessions don't.
**Cite this for:** the structure of `COMPENSATING_LIBRARY` in `exception_router.py` — every band of exception severity carries a non-negotiable list of customer commitments.
URL: https://winningbydesign.com/
---
## 3. Forrester — Deal Desk Maturity Research
Forrester's deal-desk research (Bob Apollo, Mary Shea) defines four maturity levels:
1. **Ad hoc** — discounts approved by relationship; no consistent record
2. **Formalized** — written policy exists; not data-backed; reviewed annually at best
3. **Operationalized** — policy is data-backed; quarterly reviewed; approver chain enforced
4. **Strategic** — policy is a product; A/B-tested band changes; tied to NRR targets
The skill targets level 3-4. The lint pass enforces the structural requirements (no inversion, no gaps, no cliffs, data-backed bands).
**Cite this for:** the framing that commercial policy is a designed artifact subject to lint, version control, and review — not folklore.
URL: https://www.forrester.com/
---
## 4. MIT Sloan — Incentive-System Gaming Research
MIT Sloan (Robert Gibbons, Bengt Holmström) published the foundational work on **multitask agency problems**: when agents are paid for outcome A but can game on dimension B, they will. Apply directly to discount policy:
- If "strategic value" lets an AE override the matrix, AEs will define every deal as strategic.
- If there's a cliff at $99K vs $100K ARR, AEs will split deals or pad them.
- If the precedent rule (last quarter's exception = this quarter's floor) isn't broken explicitly in policy, drift compounds.
**Cite this for:** lint rule L05 (`cliff_edge`) and the precedent-risk flag in `exception_router.py` — both are responses to predictable gaming surfaces that the agency-theory literature identifies.
URL: https://mitsloan.mit.edu/faculty/directory/robert-gibbons
---
## 5. McKinsey — Commercial Policy Effectiveness Studies
McKinsey's B2B pricing practice has published multiple studies on commercial policy effectiveness. The headline finding across deployments:
- **Companies that move from ad-hoc to operationalized commercial policy capture 2-4 pts of margin within 4 quarters** — without raising prices, without losing deals.
- **The biggest single move is closing the strategic-value loophole** — defining concrete tests so the tier isn't a catch-all.
**Cite this for:** the ROI claim that justifies the skill's existence. The skill produces the policy; the policy captures 2-4 pts of margin via McKinsey's deployment evidence.
URL: https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights
---
## 6. Bain — Discount Discipline & Pricing Power
Bain's *Pricing Power* research argues that commercial-policy maturity is the strongest internal predictor of pricing power. Two structural claims:
- **Discount discipline > price increases** for margin expansion. Raising list 5% and giving 10% more discount nets to a margin loss; holding list and tightening discount bands nets to a gain.
- **The CFO must own margin floors; the CRO must own discount bands; the Head of Deal Desk owns the matrix.** Mixing these accountabilities is the most common source of policy drift.
**Cite this for:** the `min_margin_pct` constraint input to `discount_matrix_builder.py` (CFO-owned) versus the `max_discount_pct_without_exception` (CRO/Deal-Desk-owned). The skill separates these by design.
URL: https://www.bain.com/insights/topics/pricing/
---
## 7. Salesforce CPQ — Commercial Policy Implementation Best Practices
Salesforce's CPQ implementation guides (and the surrounding ISV community) document the operational reality of encoding commercial policy in a system of record. Three practical lessons:
- **Every exception must produce machine-readable audit metadata.** "VP approved by email" doesn't survive an audit; "approval record in CPQ with timestamped justification + compensating commitments + named approver" does.
- **Approver chains should be enforced by the system, not by manager discipline.** Manager discipline degrades under quarter-end pressure; system enforcement doesn't.
- **The matrix must be versioned.** When you change a band, the old version must remain readable so historical deals can be audited against the policy that was in force at sign.
**Cite this for:** the structured `audit_trail` JSON block emitted by `exception_router.py` — designed to be machine-readable and persistable.
URL: https://www.salesforce.com/products/cpq/
---
## Synthesis: design principles the skill enforces
| Principle | Source | Where it shows up in the skill |
|---|---|---|
| One-page matrix, no six-page memo | SaaStr / Lemkin | `discount_matrix_builder.py --output markdown` produces one table |
| Discount-for-nothing is a leak | Winning by Design | `COMPENSATING_LIBRARY` per severity band in exception router |
| Policy as designed artifact | Forrester | The lint pass exists |
| Gaming surfaces are predictable | MIT Sloan | Lint rules L05 (cliff), L06 (undefined strategic), L01 (inversion) |
| Operationalized policy = 2-4 pts margin | McKinsey | ROI justification for the skill |
| CFO owns floor, CRO owns bands | Bain | Separate input parameters in `target_constraints` |
| Machine-readable audit metadata | Salesforce CPQ | `audit_trail` JSON block |
FILE:scripts/discount_matrix_builder.py
#!/usr/bin/env python3
"""discount_matrix_builder.py - Design a data-backed discount matrix.
Stdlib-only. Builds a 4-dimensional discount matrix indexed by:
(ARR band) x (term length) x (payment terms days) x (strategic value tier)
Each cell carries:
- approved_discount_band (min%, max%) — backed by current win-rate and NRR
distribution observed at that cell in the input `current_deals[]` corpus
- approver_tier (AE / Manager / Director / VP / CFO)
- margin_floor_pct — derived from target_constraints.min_margin_pct minus
a per-cell allowance proportional to strategic value
- data_backing — n_deals, win_rate, nrr_12mo observed; flagged THIN if n<5
- exception_required — TRUE when target max% exceeds matrix max%
Industry profiles tune the band widths and approver thresholds:
saas, enterprise-software, api, marketplace, services
Usage:
python discount_matrix_builder.py --sample
python discount_matrix_builder.py --input policy_intake.json --profile saas
python discount_matrix_builder.py --input policy_intake.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from typing import Any
# ------------------------------ Sample input ------------------------------ #
SAMPLE_INPUT: dict[str, Any] = {
"industry": "saas",
"current_deals": [
{"arr": 18000, "discount_pct": 8, "term_months": 12, "payment_terms_days": 30, "strategic_value": "standard", "win_lost": "win", "nrr_12mo": 1.08},
{"arr": 22000, "discount_pct": 12, "term_months": 12, "payment_terms_days": 30, "strategic_value": "standard", "win_lost": "win", "nrr_12mo": 1.05},
{"arr": 28000, "discount_pct": 18, "term_months": 12, "payment_terms_days": 45, "strategic_value": "standard", "win_lost": "lost", "nrr_12mo": 0.0},
{"arr": 75000, "discount_pct": 14, "term_months": 24, "payment_terms_days": 30, "strategic_value": "standard", "win_lost": "win", "nrr_12mo": 1.12},
{"arr": 95000, "discount_pct": 22, "term_months": 24, "payment_terms_days": 30, "strategic_value": "logo", "win_lost": "win", "nrr_12mo": 1.18},
{"arr": 130000, "discount_pct": 28, "term_months": 24, "payment_terms_days": 45, "strategic_value": "logo", "win_lost": "win", "nrr_12mo": 1.10},
{"arr": 260000, "discount_pct": 26, "term_months": 36, "payment_terms_days": 30, "strategic_value": "expansion", "win_lost": "win", "nrr_12mo": 1.22},
{"arr": 410000, "discount_pct": 30, "term_months": 36, "payment_terms_days": 30, "strategic_value": "expansion", "win_lost": "win", "nrr_12mo": 1.25},
{"arr": 540000, "discount_pct": 38, "term_months": 36, "payment_terms_days": 60, "strategic_value": "logo", "win_lost": "lost", "nrr_12mo": 0.0},
{"arr": 720000, "discount_pct": 32, "term_months": 36, "payment_terms_days": 30, "strategic_value": "expansion", "win_lost": "win", "nrr_12mo": 1.20},
],
"target_constraints": {
"min_margin_pct": 70.0,
"max_discount_pct_without_exception": 35.0,
"target_nrr": 1.15,
},
}
# ------------------------------ Dimensions ------------------------------ #
ARR_BANDS = [
("smb", 0, 25_000),
("mid", 25_000, 100_000),
("enterprise", 100_000, 500_000),
("strategic", 500_000, 10_000_000_000),
]
TERM_BANDS = [
("annual", 0, 12),
("two_year", 13, 24),
("multi_year", 25, 120),
]
PAYMENT_BANDS = [
("net30_prepay", 0, 30),
("net45", 31, 45),
("net60_plus", 46, 365),
]
STRATEGIC_TIERS = ["standard", "logo", "expansion", "lighthouse"]
PROFILES: dict[str, dict[str, Any]] = {
"saas": {
# max_discount per (arr_band, term_band, payment_band, strat_tier)
# baseline maxima; tuned by strategic tier and term shape
"base_max_pct": {"smb": 15, "mid": 22, "enterprise": 30, "strategic": 38},
"term_bonus": {"annual": 0, "two_year": 3, "multi_year": 6},
"payment_penalty": {"net30_prepay": 0, "net45": -2, "net60_plus": -5},
"strategic_bonus": {"standard": 0, "logo": 4, "expansion": 6, "lighthouse": 10},
"approver_thresholds": [(15, "AE"), (25, "Sales Manager"), (35, "Director"), (50, "VP Sales"), (100.1, "CFO + CRO")],
},
"enterprise-software": {
"base_max_pct": {"smb": 20, "mid": 28, "enterprise": 38, "strategic": 48},
"term_bonus": {"annual": 0, "two_year": 4, "multi_year": 8},
"payment_penalty": {"net30_prepay": 0, "net45": -2, "net60_plus": -6},
"strategic_bonus": {"standard": 0, "logo": 5, "expansion": 8, "lighthouse": 12},
"approver_thresholds": [(20, "AE"), (30, "Sales Manager"), (40, "Director"), (55, "VP Sales"), (100.1, "CFO + CRO")],
},
"api": {
"base_max_pct": {"smb": 10, "mid": 18, "enterprise": 25, "strategic": 32},
"term_bonus": {"annual": 0, "two_year": 2, "multi_year": 5},
"payment_penalty": {"net30_prepay": 0, "net45": -2, "net60_plus": -4},
"strategic_bonus": {"standard": 0, "logo": 3, "expansion": 5, "lighthouse": 8},
"approver_thresholds": [(10, "AE"), (18, "Sales Manager"), (25, "Director"), (35, "VP Sales"), (100.1, "CFO + CRO")],
},
"marketplace": {
"base_max_pct": {"smb": 8, "mid": 12, "enterprise": 18, "strategic": 25},
"term_bonus": {"annual": 0, "two_year": 2, "multi_year": 4},
"payment_penalty": {"net30_prepay": 0, "net45": -1, "net60_plus": -3},
"strategic_bonus": {"standard": 0, "logo": 2, "expansion": 4, "lighthouse": 6},
"approver_thresholds": [(8, "AE"), (15, "Sales Manager"), (22, "Director"), (30, "VP"), (100.1, "CFO + CRO")],
},
"services": {
# margin-thin; tight bands and fast escalation
"base_max_pct": {"smb": 5, "mid": 10, "enterprise": 15, "strategic": 22},
"term_bonus": {"annual": 0, "two_year": 2, "multi_year": 3},
"payment_penalty": {"net30_prepay": 0, "net45": -1, "net60_plus": -3},
"strategic_bonus": {"standard": 0, "logo": 2, "expansion": 3, "lighthouse": 5},
"approver_thresholds": [(5, "AE"), (12, "Sales Manager"), (20, "Director"), (30, "VP Services"), (100.1, "CFO + COO")],
},
}
# ------------------------------ Logic ------------------------------ #
def _band(value: float, bands: list[tuple]) -> str:
for name, lo, hi in bands:
if lo <= value <= hi:
return name
return bands[-1][0]
def _approver_for(max_pct: float, thresholds: list[tuple[float, str]]) -> str:
for cutoff, name in thresholds:
if max_pct <= cutoff:
return name
return thresholds[-1][1]
def _classify_deal(deal: dict[str, Any]) -> tuple[str, str, str, str]:
return (
_band(deal["arr"], ARR_BANDS),
_band(deal["term_months"], TERM_BANDS),
_band(deal["payment_terms_days"], PAYMENT_BANDS),
deal.get("strategic_value", "standard"),
)
def build_matrix(payload: dict[str, Any], profile_name: str) -> dict[str, Any]:
profile = PROFILES.get(profile_name, PROFILES["saas"])
deals = payload.get("current_deals", [])
constraints = payload.get("target_constraints", {})
min_margin = float(constraints.get("min_margin_pct", 70.0))
max_without_exception = float(constraints.get("max_discount_pct_without_exception", 35.0))
target_nrr = float(constraints.get("target_nrr", 1.10))
# Bucket observed deals by cell.
buckets: dict[tuple, list[dict]] = {}
for d in deals:
key = _classify_deal(d)
buckets.setdefault(key, []).append(d)
cells: list[dict[str, Any]] = []
for arr_band, _, _ in ARR_BANDS:
for term_band, _, _ in TERM_BANDS:
for pay_band, _, _ in PAYMENT_BANDS:
for strat_tier in STRATEGIC_TIERS:
key = (arr_band, term_band, pay_band, strat_tier)
base = profile["base_max_pct"][arr_band]
bonus_term = profile["term_bonus"][term_band]
pen_pay = profile["payment_penalty"][pay_band]
bonus_strat = profile["strategic_bonus"][strat_tier]
cell_max = max(0.0, base + bonus_term + pen_pay + bonus_strat)
cell_min = max(0.0, cell_max * 0.5) # min discount in this band
# Observed data backing
obs = buckets.get(key, [])
n = len(obs)
wins = sum(1 for d in obs if d.get("win_lost") == "win")
win_rate = (wins / n) if n else None
nrr_vals = [d.get("nrr_12mo", 0.0) for d in obs if d.get("win_lost") == "win"]
nrr_obs = (sum(nrr_vals) / len(nrr_vals)) if nrr_vals else None
# Margin floor: every 1% discount typically costs ~(1/gm)% of margin.
# Cap the cell at the constraint-driven max as well.
capped_max = min(cell_max, max_without_exception + bonus_strat) # strategic gets a touch more
exception_required = capped_max > max_without_exception
# Margin floor: subtract a strategic-value allowance.
margin_floor = max(min_margin - bonus_strat, 50.0)
approver = _approver_for(capped_max, profile["approver_thresholds"])
cells.append({
"arr_band": arr_band,
"term_band": term_band,
"payment_band": pay_band,
"strategic_tier": strat_tier,
"approved_discount_min_pct": round(cell_min, 1),
"approved_discount_max_pct": round(capped_max, 1),
"approver_tier": approver,
"margin_floor_pct": round(margin_floor, 1),
"exception_required_above_pct": round(max_without_exception, 1),
"data_backing": {
"n_observed_deals": n,
"win_rate": round(win_rate, 3) if win_rate is not None else None,
"nrr_12mo_observed": round(nrr_obs, 3) if nrr_obs is not None else None,
"thin_data_flag": n < 5,
},
"meets_target_nrr": (nrr_obs is not None and nrr_obs >= target_nrr),
"exception_required": exception_required,
})
return {
"profile": profile_name,
"constraints": {
"min_margin_pct": min_margin,
"max_discount_pct_without_exception": max_without_exception,
"target_nrr": target_nrr,
},
"n_cells": len(cells),
"n_observed_deals": len(deals),
"cells": cells,
}
# ------------------------------ Rendering ------------------------------ #
def render_markdown(matrix: dict[str, Any]) -> str:
out: list[str] = []
out.append(f"# Discount Matrix — profile: `{matrix['profile']}`")
out.append("")
out.append("## Constraints")
for k, v in matrix["constraints"].items():
out.append(f"- **{k}**: {v}")
out.append("")
out.append(f"## Cells ({matrix['n_cells']}) — backed by {matrix['n_observed_deals']} observed deals")
out.append("")
out.append("| ARR | Term | Payment | Strategic | Discount band | Approver | Margin floor | n | Win rate | NRR | Exception? |")
out.append("|---|---|---|---|---|---|---|---|---|---|---|")
for c in matrix["cells"]:
db = c["data_backing"]
wr = f"{db['win_rate']:.0%}" if db["win_rate"] is not None else "—"
nrr = f"{db['nrr_12mo_observed']:.2f}" if db["nrr_12mo_observed"] is not None else "—"
thin = " (THIN)" if db["thin_data_flag"] else ""
exc = "YES" if c["exception_required"] else "no"
out.append(
f"| {c['arr_band']} | {c['term_band']} | {c['payment_band']} | {c['strategic_tier']} | "
f"{c['approved_discount_min_pct']}-{c['approved_discount_max_pct']}% | "
f"{c['approver_tier']} | {c['margin_floor_pct']}% | "
f"{db['n_observed_deals']}{thin} | {wr} | {nrr} | {exc} |"
)
out.append("")
out.append("## Notes")
out.append("- THIN data flag means n<5 observed deals in this cell — treat the band as directional, not data-backed.")
out.append("- Strategic tiers carry a margin-floor allowance proportional to their bonus; lighthouse cells absorb the deepest discounts.")
out.append("- Cells flagged `Exception? YES` exceed the policy's max-without-exception threshold and must route through `exception_router.py`.")
return "\n".join(out)
# ------------------------------ CLI ------------------------------ #
def main(argv: list[str]) -> int:
ap = argparse.ArgumentParser(description="Design a data-backed discount matrix.")
ap.add_argument("--input", help="Path to policy intake JSON.")
ap.add_argument("--profile", default="saas",
choices=list(PROFILES.keys()),
help="Industry profile (default: saas).")
ap.add_argument("--output", default="markdown", choices=["markdown", "json"],
help="Output format (default: markdown).")
ap.add_argument("--sample", action="store_true", help="Run with the built-in sample payload.")
args = ap.parse_args(argv)
if args.sample:
payload = SAMPLE_INPUT
profile = args.profile or payload.get("industry", "saas")
elif args.input:
try:
with open(args.input, "r", encoding="utf-8") as f:
payload = json.load(f)
except Exception as e:
print(f"ERROR: could not read {args.input}: {e}", file=sys.stderr)
return 1
profile = args.profile or payload.get("industry", "saas")
else:
ap.print_help()
return 0
matrix = build_matrix(payload, profile)
if args.output == "json":
print(json.dumps(matrix, indent=2))
else:
print(render_markdown(matrix))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/exception_router.py
#!/usr/bin/env python3
"""exception_router.py - Route a discount exception through the policy.
Stdlib-only. Takes an exception request and a matrix path. Decides:
- IN_POLICY → no exception needed; surface the standard approver
- EXCEPTION → produces:
* required approver chain (AE -> ... -> CFO/CRO)
* required compensating commitments (multi-year prepay, named
expansion path, reference commitment, MSA tightening, etc.)
* audit-trail metadata block (timestamp, requested_by, justification,
compensating_commitments_text, approver_chain)
- PRECEDENT_RISK → flagged if recent_exceptions[] shows 3+ similar
asks in the trailing quarter. Signals the matrix may be wrong, not
the deal.
Usage:
python exception_router.py --sample
python exception_router.py --input request.json
python exception_router.py --input request.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
import datetime
from typing import Any
SAMPLE_INPUT: dict[str, Any] = {
"exception_request": {
"deal_id": "ACME-2026-Q3-204",
"requested_by": "Jordan Smith, AE",
"deal_arr": 320000,
"requested_discount": 42.0,
"term_months": 36,
"payment_terms_days": 30,
"justification": "Customer is a logo competitor displacement; CFO sponsor; pipeline expansion to 3 BU committed verbally.",
"strategic_value": "logo",
"customer_threats": ["competitor_proposal", "fy_close_pressure"],
"submitted_at": "2026-05-19T10:00:00Z",
},
"policy_matrix": {
"profile": "saas",
"max_discount_pct_without_exception": 35.0,
"approver_thresholds": [
[15, "AE"], [25, "Sales Manager"], [35, "Director"], [50, "VP Sales"], [100.1, "CFO + CRO"]
],
},
"recent_exceptions": [
{"deal_id": "BETA-2026-Q2-188", "discount": 40, "arr": 280000, "strategic": "logo"},
{"deal_id": "GAMMA-2026-Q2-192", "discount": 41, "arr": 310000, "strategic": "logo"},
{"deal_id": "DELTA-2026-Q2-201", "discount": 43, "arr": 350000, "strategic": "expansion"},
],
}
# Compensating commitments are NON-NEGOTIABLE per band of exception severity.
# Severity = (requested_discount - max_without_exception).
COMPENSATING_LIBRARY: list[dict[str, Any]] = [
{
"severity_floor": 0.0, "severity_ceiling": 5.0,
"commitments": [
"multi_year_term (>= 24 months)",
"annual_prepay (NET-30 or shorter)",
],
},
{
"severity_floor": 5.0, "severity_ceiling": 10.0,
"commitments": [
"multi_year_term (>= 36 months)",
"annual_prepay (NET-30 or shorter)",
"named_expansion_path (BU or product, in writing)",
],
},
{
"severity_floor": 10.0, "severity_ceiling": 20.0,
"commitments": [
"multi_year_term (>= 36 months) with prepay of years 1+2",
"named_expansion_path (BU or product, in writing)",
"reference_commitment (case study + 2 customer-reference calls per year)",
"msa_tightening (auto-renewal, MFN-protection, indemnity-cap)",
],
},
{
"severity_floor": 20.0, "severity_ceiling": 1000.0,
"commitments": [
"multi_year_term (>= 36 months) with prepay of years 1+2",
"named_expansion_path with quantified expansion ARR target",
"reference_commitment + co-marketing agreement",
"msa_tightening (auto-renewal, MFN-protection, indemnity-cap)",
"executive_sponsor_signoff (customer C-level on the contract)",
"kill_switch: if expansion ARR target missed by end of year 2, renewal reverts to list",
],
},
]
def _approver_chain_for(discount: float, thresholds: list[tuple[float, str]]) -> list[str]:
"""Build cumulative approver chain up to the named human who must sign."""
chain: list[str] = []
for cutoff, name in thresholds:
chain.append(name)
if discount <= cutoff:
return chain
return chain
def _compensating_for(severity: float) -> list[str]:
for band in COMPENSATING_LIBRARY:
if band["severity_floor"] <= severity < band["severity_ceiling"]:
return list(band["commitments"])
return list(COMPENSATING_LIBRARY[-1]["commitments"])
def _precedent_risk(recent: list[dict[str, Any]], requested_discount: float, strategic_value: str) -> dict[str, Any]:
similar = [
r for r in recent
if abs(r.get("discount", 0) - requested_discount) <= 5
and r.get("strategic") == strategic_value
]
flag = len(similar) >= 3
return {
"similar_recent_count": len(similar),
"trigger_threshold": 3,
"flag": flag,
"matrix_review_recommended": flag,
"rationale": (
"3+ similar exceptions in trailing quarter — the policy band may be set wrong; "
"rebuild the matrix with discount_matrix_builder.py before approving another."
if flag else "Pattern within tolerance; treat as individual exception."
),
}
def route_exception(payload: dict[str, Any]) -> dict[str, Any]:
req = payload["exception_request"]
matrix = payload.get("policy_matrix", {})
recent = payload.get("recent_exceptions", [])
max_without = float(matrix.get("max_discount_pct_without_exception", 35.0))
thresholds: list[tuple[float, str]] = [
(float(c), n) for c, n in matrix.get("approver_thresholds", [(15, "AE"), (35, "Director"), (100.1, "CFO + CRO")])
]
requested = float(req["requested_discount"])
in_policy = requested <= max_without
severity = max(0.0, requested - max_without)
chain = _approver_chain_for(requested, thresholds)
if not in_policy:
# Exceptions always escalate to at least Director — never stop at AE/Manager.
promoted = []
seen_director_or_above = False
for hop in chain:
promoted.append(hop)
if hop in ("Director", "Director of Sales", "VP Sales", "VP", "VP Services", "CFO + CRO", "CFO + COO"):
seen_director_or_above = True
if not seen_director_or_above:
promoted.append("Director")
promoted.append("VP Sales")
chain = promoted
compensating = _compensating_for(severity) if not in_policy else []
precedent = _precedent_risk(recent, requested, req.get("strategic_value", "standard"))
audit_trail = {
"deal_id": req.get("deal_id"),
"requested_by": req.get("requested_by"),
"requested_discount_pct": requested,
"deal_arr": req.get("deal_arr"),
"term_months": req.get("term_months"),
"justification": req.get("justification"),
"strategic_value": req.get("strategic_value"),
"customer_threats": req.get("customer_threats", []),
"submitted_at": req.get("submitted_at") or datetime.datetime.utcnow().isoformat() + "Z",
"compensating_commitments_required": compensating,
"approver_chain": chain,
"verdict": "IN_POLICY" if in_policy else "EXCEPTION",
}
return {
"verdict": "IN_POLICY" if in_policy else "EXCEPTION",
"severity_pct_over_threshold": round(severity, 2),
"approver_chain": chain,
"required_compensating_commitments": compensating,
"precedent_risk": precedent,
"audit_trail": audit_trail,
"notes": [
("In-policy request — route to standard approver; no compensating commitments required."
if in_policy else
"EXCEPTION — the chain must capture each compensating commitment in writing before sign."),
("Precedent risk FLAGGED — rebuild the matrix before approving."
if precedent["flag"] else
"No precedent flag."),
],
}
def render_markdown(result: dict[str, Any]) -> str:
out = []
audit = result["audit_trail"]
out.append(f"# Exception Routing — {audit['deal_id']}")
out.append("")
out.append(f"**Verdict:** `{result['verdict']}` "
f"(severity: {result['severity_pct_over_threshold']} pts over threshold)")
out.append("")
out.append("## Approver chain")
for i, hop in enumerate(result["approver_chain"], 1):
out.append(f"{i}. {hop}")
out.append("")
if result["required_compensating_commitments"]:
out.append("## Required compensating commitments (NON-NEGOTIABLE)")
for c in result["required_compensating_commitments"]:
out.append(f"- {c}")
out.append("")
out.append("## Precedent risk")
pr = result["precedent_risk"]
out.append(f"- Similar recent exceptions: **{pr['similar_recent_count']}** (trigger: {pr['trigger_threshold']})")
out.append(f"- Flag: **{'YES' if pr['flag'] else 'no'}**")
out.append(f"- Rationale: {pr['rationale']}")
out.append("")
out.append("## Audit trail")
out.append("```json")
out.append(json.dumps(audit, indent=2))
out.append("```")
out.append("")
out.append("## Notes")
for n in result["notes"]:
out.append(f"- {n}")
return "\n".join(out)
def main(argv: list[str]) -> int:
ap = argparse.ArgumentParser(description="Route a discount exception through the policy.")
ap.add_argument("--input", help="Path to exception request JSON (with policy_matrix + recent_exceptions).")
ap.add_argument("--output", default="markdown", choices=["markdown", "json"],
help="Output format (default: markdown).")
ap.add_argument("--sample", action="store_true", help="Run with the built-in sample request.")
args = ap.parse_args(argv)
if args.sample:
payload = SAMPLE_INPUT
elif args.input:
try:
with open(args.input, "r", encoding="utf-8") as f:
payload = json.load(f)
except Exception as e:
print(f"ERROR: could not read {args.input}: {e}", file=sys.stderr)
return 1
else:
ap.print_help()
return 0
result = route_exception(payload)
if args.output == "json":
print(json.dumps(result, indent=2))
else:
print(render_markdown(result))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
FILE:scripts/policy_linter.py
#!/usr/bin/env python3
"""policy_linter.py - Lint a discount matrix for governance defects.
Stdlib-only. Reads the JSON output of discount_matrix_builder.py (or a
hand-authored matrix in the same shape). Returns a ranked findings report:
BLOCKER — policy is internally contradictory or unsignable
MAJOR — discoverable gaming surface or missing data backing in a critical cell
MINOR — stylistic / completeness issue
Lint rules (deterministic):
L01 BLOCKER approver_hierarchy_inversion — lower-tier approves more than higher-tier
L02 BLOCKER cell_band_inverted — min > max in a cell band
L03 BLOCKER margin_floor_below_constraint — cell margin floor < 50%
L04 MAJOR coverage_gap — cell missing approver_tier
L05 MAJOR cliff_edge — adjacent ARR/term/payment cells differ by > 10 pts
L06 MAJOR strategic_value_undefined — strategic tier present but no verifiable definition supplied
L07 MAJOR inconsistent_margin_floor — same arr_band has > 5pt floor variance across cells
L08 MAJOR thin_data_in_critical_cell — critical cell (enterprise/strategic) flagged THIN
L09 MINOR cell_unreviewed — n_observed_deals == 0
L10 MINOR missing_exception_marker — high discount cell without exception flag
Usage:
python policy_linter.py --sample
python policy_linter.py --input matrix.json
python policy_linter.py --input matrix.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from typing import Any
SAMPLE_INPUT: dict[str, Any] = {
"profile": "saas",
"constraints": {
"min_margin_pct": 70.0,
"max_discount_pct_without_exception": 35.0,
"target_nrr": 1.10,
},
"strategic_value_definitions_supplied": False,
"cells": [
# A clean cell
{
"arr_band": "smb", "term_band": "annual", "payment_band": "net30_prepay", "strategic_tier": "standard",
"approved_discount_min_pct": 0, "approved_discount_max_pct": 15,
"approver_tier": "AE", "margin_floor_pct": 70,
"data_backing": {"n_observed_deals": 8, "win_rate": 0.62, "nrr_12mo_observed": 1.05, "thin_data_flag": False},
"exception_required": False,
},
# Approver inversion — Manager allows 25%, Director below allows only 20%
{
"arr_band": "mid", "term_band": "annual", "payment_band": "net30_prepay", "strategic_tier": "standard",
"approved_discount_min_pct": 8, "approved_discount_max_pct": 25,
"approver_tier": "Sales Manager", "margin_floor_pct": 70,
"data_backing": {"n_observed_deals": 6, "win_rate": 0.5, "nrr_12mo_observed": 1.10, "thin_data_flag": False},
"exception_required": False,
},
{
"arr_band": "mid", "term_band": "two_year", "payment_band": "net30_prepay", "strategic_tier": "standard",
"approved_discount_min_pct": 5, "approved_discount_max_pct": 20,
"approver_tier": "Director", "margin_floor_pct": 70,
"data_backing": {"n_observed_deals": 3, "win_rate": 0.4, "nrr_12mo_observed": 1.12, "thin_data_flag": True},
"exception_required": False,
},
# Inverted band (BLOCKER)
{
"arr_band": "enterprise", "term_band": "two_year", "payment_band": "net30_prepay", "strategic_tier": "logo",
"approved_discount_min_pct": 25, "approved_discount_max_pct": 20,
"approver_tier": "Director", "margin_floor_pct": 65,
"data_backing": {"n_observed_deals": 2, "win_rate": 0.5, "nrr_12mo_observed": 1.18, "thin_data_flag": True},
"exception_required": False,
},
# Margin floor below constraint (BLOCKER)
{
"arr_band": "strategic", "term_band": "multi_year", "payment_band": "net60_plus", "strategic_tier": "lighthouse",
"approved_discount_min_pct": 25, "approved_discount_max_pct": 48,
"approver_tier": "CFO + CRO", "margin_floor_pct": 45,
"data_backing": {"n_observed_deals": 1, "win_rate": 1.0, "nrr_12mo_observed": 1.30, "thin_data_flag": True},
"exception_required": True,
},
# Coverage gap (no approver)
{
"arr_band": "enterprise", "term_band": "multi_year", "payment_band": "net45", "strategic_tier": "expansion",
"approved_discount_min_pct": 15, "approved_discount_max_pct": 36,
"approver_tier": None, "margin_floor_pct": 64,
"data_backing": {"n_observed_deals": 0, "win_rate": None, "nrr_12mo_observed": None, "thin_data_flag": True},
"exception_required": True,
},
# High discount with no exception flag (MINOR)
{
"arr_band": "enterprise", "term_band": "two_year", "payment_band": "net30_prepay", "strategic_tier": "logo",
"approved_discount_min_pct": 18, "approved_discount_max_pct": 40,
"approver_tier": "VP Sales", "margin_floor_pct": 66,
"data_backing": {"n_observed_deals": 4, "win_rate": 0.5, "nrr_12mo_observed": 1.12, "thin_data_flag": True},
"exception_required": False,
},
],
}
APPROVER_RANK = {
"AE": 1, "Sales Manager": 2, "Director": 3, "Director of Sales": 3,
"VP Sales": 4, "VP": 4, "VP Services": 4, "CFO + CRO": 5, "CFO + COO": 5,
}
def _rank(approver: str | None) -> int:
return APPROVER_RANK.get(approver or "", 0)
def lint(matrix: dict[str, Any]) -> dict[str, Any]:
cells = matrix.get("cells", [])
constraints = matrix.get("constraints", {})
max_without = float(constraints.get("max_discount_pct_without_exception", 35.0))
findings: list[dict[str, Any]] = []
# L01: approver hierarchy inversion across all cells
# For each pair, if approver_A rank > approver_B rank but approved_max_A < approved_max_B
# => the lower-rank approver authorizes a higher discount than the higher-rank approver.
for i, ci in enumerate(cells):
for cj in cells[i + 1:]:
ri, rj = _rank(ci.get("approver_tier")), _rank(cj.get("approver_tier"))
if ri == 0 or rj == 0 or ri == rj:
continue
mi, mj = ci["approved_discount_max_pct"], cj["approved_discount_max_pct"]
# Identify the higher-rank and lower-rank cell, then check inversion.
if ri > rj:
higher, lower, mh, ml = ci, cj, mi, mj
else:
higher, lower, mh, ml = cj, ci, mj, mi
if mh < ml:
findings.append({
"rule_id": "L01", "severity": "BLOCKER",
"name": "approver_hierarchy_inversion",
"detail": (
f"{lower['approver_tier']} approves up to {ml}% in "
f"({lower['arr_band']}/{lower['term_band']}/{lower['strategic_tier']}), but "
f"{higher['approver_tier']} approves only up to {mh}% in "
f"({higher['arr_band']}/{higher['term_band']}/{higher['strategic_tier']})."
),
"fix": "Raise the higher-rank approver's cap above the lower-rank cap, or demote the lower-rank cap.",
})
# L02: inverted bands
for c in cells:
if c["approved_discount_min_pct"] > c["approved_discount_max_pct"]:
findings.append({
"rule_id": "L02", "severity": "BLOCKER",
"name": "cell_band_inverted",
"detail": f"Cell ({c['arr_band']}/{c['term_band']}/{c['payment_band']}/{c['strategic_tier']}) has min {c['approved_discount_min_pct']}% > max {c['approved_discount_max_pct']}%.",
"fix": "Recompute the band — min must be <= max.",
})
# L03: margin floor below sanity (<50%)
for c in cells:
if c["margin_floor_pct"] < 50.0:
findings.append({
"rule_id": "L03", "severity": "BLOCKER",
"name": "margin_floor_below_constraint",
"detail": f"Cell ({c['arr_band']}/{c['term_band']}/{c['strategic_tier']}) margin floor is {c['margin_floor_pct']}% (< 50%).",
"fix": "Raise the floor, or carve out this cell as an explicit exception band requiring CFO sign.",
})
# L04: coverage gap (no approver)
for c in cells:
if not c.get("approver_tier"):
findings.append({
"rule_id": "L04", "severity": "MAJOR",
"name": "coverage_gap",
"detail": f"Cell ({c['arr_band']}/{c['term_band']}/{c['payment_band']}/{c['strategic_tier']}) has no approver_tier assigned.",
"fix": "Assign a named approver tier per the approver_thresholds table.",
})
# L05: cliff edges — same dim differing by > 10 pts on adjacent bands.
# Compare cells differing only in arr_band (adjacent), then only in term_band, then only in payment.
ARR_ORDER = ["smb", "mid", "enterprise", "strategic"]
TERM_ORDER = ["annual", "two_year", "multi_year"]
PAY_ORDER = ["net30_prepay", "net45", "net60_plus"]
by_key: dict[tuple, dict[str, Any]] = {}
for c in cells:
key = (c["arr_band"], c["term_band"], c["payment_band"], c["strategic_tier"])
by_key[key] = c
def _adj(order: list[str], v: str) -> str | None:
try:
idx = order.index(v)
return order[idx + 1] if idx + 1 < len(order) else None
except ValueError:
return None
for key, c in by_key.items():
arr, term, pay, strat = key
for dim, order, axis in [(arr, ARR_ORDER, "arr"), (term, TERM_ORDER, "term"), (pay, PAY_ORDER, "payment")]:
nxt = _adj(order, dim)
if not nxt:
continue
adj_key = (
nxt if axis == "arr" else arr,
nxt if axis == "term" else term,
nxt if axis == "payment" else pay,
strat,
)
adj = by_key.get(adj_key)
if not adj:
continue
delta = abs(adj["approved_discount_max_pct"] - c["approved_discount_max_pct"])
if delta > 10:
findings.append({
"rule_id": "L05", "severity": "MAJOR",
"name": "cliff_edge",
"detail": (
f"{axis} cliff between ({c['arr_band']}/{c['term_band']}/{c['payment_band']}/{c['strategic_tier']}) "
f"max {c['approved_discount_max_pct']}% and ({adj['arr_band']}/{adj['term_band']}/{adj['payment_band']}/{adj['strategic_tier']}) "
f"max {adj['approved_discount_max_pct']}% — {delta} pts apart."
),
"fix": "Smooth the gradient — large jumps create gaming surfaces (e.g., AE splits a $101K deal into 2x $50.5K to dodge the band).",
})
# L06: strategic_value_undefined — if any strategic tier > 'standard' is used and definitions absent
used_strategic = {c["strategic_tier"] for c in cells if c["strategic_tier"] != "standard"}
if used_strategic and not matrix.get("strategic_value_definitions_supplied", False):
findings.append({
"rule_id": "L06", "severity": "MAJOR",
"name": "strategic_value_undefined",
"detail": f"Strategic tiers used ({sorted(used_strategic)}) but no verifiable definition supplied in the matrix.",
"fix": "Add strategic_value_definitions_supplied=true plus a definitions section: e.g., 'logo = top-20 enterprise in named target list; expansion = signed MSA with named BU expansion path'.",
})
# L07: inconsistent margin floor within an arr_band
by_arr: dict[str, list[float]] = {}
for c in cells:
by_arr.setdefault(c["arr_band"], []).append(c["margin_floor_pct"])
for arr_band, floors in by_arr.items():
if floors and (max(floors) - min(floors)) > 5:
findings.append({
"rule_id": "L07", "severity": "MAJOR",
"name": "inconsistent_margin_floor",
"detail": f"Margin floor in arr_band={arr_band} varies by {max(floors) - min(floors):.1f} pts (min {min(floors)}, max {max(floors)}).",
"fix": "Pick one floor per arr_band — variance > 5 pts suggests the strategic-tier allowance is undisciplined.",
})
# L08: thin data in critical cell
for c in cells:
if c["arr_band"] in ("enterprise", "strategic") and c.get("data_backing", {}).get("thin_data_flag"):
findings.append({
"rule_id": "L08", "severity": "MAJOR",
"name": "thin_data_in_critical_cell",
"detail": f"Critical cell ({c['arr_band']}/{c['term_band']}/{c['strategic_tier']}) flagged THIN (n={c['data_backing'].get('n_observed_deals')}).",
"fix": "Treat band as directional until n>=5; do not publish to AEs as binding without flagging directional.",
})
# L09: cell unreviewed (n=0)
for c in cells:
if (c.get("data_backing", {}) or {}).get("n_observed_deals", 0) == 0:
findings.append({
"rule_id": "L09", "severity": "MINOR",
"name": "cell_unreviewed",
"detail": f"Cell ({c['arr_band']}/{c['term_band']}/{c['payment_band']}/{c['strategic_tier']}) has zero observed deals.",
"fix": "Mark as PROVISIONAL in the matrix doc; revisit at the next quarterly review.",
})
# L10: high discount cell w/o exception flag
for c in cells:
if c["approved_discount_max_pct"] > max_without and not c.get("exception_required"):
findings.append({
"rule_id": "L10", "severity": "MINOR",
"name": "missing_exception_marker",
"detail": (
f"Cell ({c['arr_band']}/{c['term_band']}/{c['strategic_tier']}) max {c['approved_discount_max_pct']}% "
f"exceeds max_without_exception ({max_without}%) but exception_required is False."
),
"fix": "Set exception_required=True so deal-desk routes through exception_router.py.",
})
severity_rank = {"BLOCKER": 0, "MAJOR": 1, "MINOR": 2}
findings.sort(key=lambda f: (severity_rank[f["severity"]], f["rule_id"]))
counts = {"BLOCKER": 0, "MAJOR": 0, "MINOR": 0}
for f in findings:
counts[f["severity"]] += 1
return {
"n_cells_linted": len(cells),
"n_findings": len(findings),
"counts": counts,
"verdict": (
"PASS" if counts["BLOCKER"] == 0 and counts["MAJOR"] == 0
else "FAIL" if counts["BLOCKER"] > 0
else "PASS_WITH_WARNINGS"
),
"findings": findings,
}
def render_markdown(report: dict[str, Any]) -> str:
out = []
out.append("# Policy Lint Report")
out.append("")
out.append(f"- Cells linted: **{report['n_cells_linted']}**")
out.append(f"- Findings: **{report['n_findings']}** "
f"(BLOCKER: {report['counts']['BLOCKER']}, MAJOR: {report['counts']['MAJOR']}, MINOR: {report['counts']['MINOR']})")
out.append(f"- Verdict: **{report['verdict']}**")
out.append("")
if not report["findings"]:
out.append("No findings. Matrix passes lint.")
return "\n".join(out)
out.append("## Findings (ranked)")
out.append("")
out.append("| # | Severity | Rule | Detail | Suggested fix |")
out.append("|---|---|---|---|---|")
for i, f in enumerate(report["findings"], 1):
out.append(
f"| {i} | **{f['severity']}** | `{f['rule_id']}` {f['name']} | {f['detail']} | {f['fix']} |"
)
out.append("")
out.append("## Next steps")
if report["counts"]["BLOCKER"] > 0:
out.append("- Resolve every BLOCKER before publishing the matrix to AEs. Blockers indicate the policy is unsignable as written.")
if report["counts"]["MAJOR"] > 0:
out.append("- Address MAJOR findings within one policy-review cycle. They surface gaming risk or coverage holes.")
if report["counts"]["MINOR"] > 0:
out.append("- Track MINOR findings in the quarterly policy review.")
return "\n".join(out)
def main(argv: list[str]) -> int:
ap = argparse.ArgumentParser(description="Lint a discount matrix for governance defects.")
ap.add_argument("--input", help="Path to matrix JSON (output of discount_matrix_builder.py).")
ap.add_argument("--output", default="markdown", choices=["markdown", "json"],
help="Output format (default: markdown).")
ap.add_argument("--sample", action="store_true", help="Run with the built-in sample matrix.")
args = ap.parse_args(argv)
if args.sample:
matrix = SAMPLE_INPUT
elif args.input:
try:
with open(args.input, "r", encoding="utf-8") as f:
matrix = json.load(f)
except Exception as e:
print(f"ERROR: could not read {args.input}: {e}", file=sys.stderr)
return 1
else:
ap.print_help()
return 0
report = lint(matrix)
if args.output == "json":
print(json.dumps(report, indent=2))
else:
print(render_markdown(report))
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
Rà soát và thiết kế hoạt động thương mại: mô hình giá, duyệt giao dịch, chiết khấu, đối tác, kênh, RFP và dự báo.
---
name: commercial-skills
description: Use when reviewing, approving, or designing commercial motion — pricing models, deal review, discount approval, partnership economics, channel mix, commercial policy, RFP/RFI response, bookings forecast. Triggers on "review this deal", "should we discount", "pricing model", "partner economics", "RFP response", "bookings forecast", "channel mix". Forks context to route to one of seven Commercial sub-skills (pricing-strategist, deal-desk, partnerships-architect, channel-economics, commercial-policy, rfp-responder, commercial-forecaster) and returns a digest. Distinct from business-growth (sales execution) and c-level-advisor/cro-advisor (strategic CRO judgment).
context: fork
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [commercial, pricing, deal-desk, partnerships, channel, rfp, forecast, cro, orchestrator]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# Commercial — Domain Orchestrator
The Commercial surface is **per-deal economics and packaging**: how the company prices, packages, approves, and forecasts revenue. This orchestrator forks its context, routes your inquiry to one of seven sub-skills, then returns a digest. Heavy intake (RFP PDFs, pipeline exports, partner agreements) stays in the forked context.
## When to invoke
| Symptom | Sub-skill |
|---|---|
| "We're losing deals on price — should we drop prices or repackage?" | `pricing-strategist` |
| "Can we approve a 40% discount on this Enterprise deal?" | `deal-desk` |
| "Should we sign with this reseller? What's their tier?" | `partnerships-architect` |
| "Is our partner channel actually profitable?" | `channel-economics` |
| "What should our standard discount matrix look like?" | `commercial-policy` |
| "Help me respond to this 60-page RFP" | `rfp-responder` |
| "What's our Q4 bookings forecast at current conversion?" | `commercial-forecaster` |
## Routing logic (deterministic)
Same two-signal threshold pattern as `business-operations-skills`. Single-signal → clarifying question. Mixed signals → highest-confidence first, chain second in follow-up turn.
### Signal table
| Signal class | Keywords | Sub-skill |
|---|---|---|
| **PRICING** | pricing, price, packaging, tier, WTP, willingness to pay, Van Westendorp, value pricing | `pricing-strategist` |
| **DEAL** | deal, discount, approval, margin, T&Cs, redline, exception, MSA | `deal-desk` |
| **PARTNERSHIP** | partner, reseller, OEM, co-sell, joint GTM, revenue share, channel agreement | `partnerships-architect` |
| **CHANNEL_ECON** | channel mix, cost to serve, channel ROI, direct vs partner, channel economics | `channel-economics` |
| **POLICY** | commercial policy, discount matrix, T&C library, exception policy, deal framework | `commercial-policy` |
| **RFP** | RFP, RFI, RFQ, proposal request, vendor questionnaire, security questionnaire | `rfp-responder` |
| **FORECAST** | forecast, bookings, billings, ARR, NRR forecast, pipeline math, funnel projection | `commercial-forecaster` |
## Workflow (Matt Pocock grill discipline)
Derived from Matt Pocock's `grill-with-docs` pattern: **explore-then-ask, one question per turn with a recommended answer, walk the decision tree depth-first, track dependencies, anchor every challenge in the SaaS pricing / deal desk canon** (`references/`).
### Step 1 — Explore before asking
Check the user's working directory first:
- Is there a deal record, pricing comp table, RFP doc, or pipeline export already in the workspace?
- Does the inquiry already disambiguate the lane (e.g., "review this 60-page RFP" — that's `rfp-responder`, no question needed)?
- Is there an artifact filename that resolves the lane (`pipeline-Q4.csv` → forecast; `MSA-redline.docx` → deal)?
If the workspace resolves the lane, **route silently**.
### Step 2 — If still ambiguous, ONE forcing question with a recommended answer
Matt's rule: never bundle. Always recommend.
Pattern:
```
Q1/1: [precise question naming the two candidate lanes]
Recommended: [Lane X, because <signal-table rationale>]
(Confirm, or override?)
```
### Step 3 — Decision-tree walk for multi-lane inquiries
If the inquiry legitimately crosses two lanes (e.g., "this RFP wants a discount we don't normally give" = RFP + DEAL + maybe POLICY), walk depth-first:
1. Highest-confidence lane first → run sub-skill in forked context → digest
2. Ask: "Now run [second lane]? Recommended: yes, because [dependency]."
3. Confirm before chaining.
Never silently chain.
### Step 4 — Invoke sub-skill in forked context
Forward original prompt + structured inputs (pipeline CSV, RFP doc path, pricing comp table, MSA redline).
### Step 5 — Return digest with cited canon challenge
≤ 200 words: analyzed, top 3 findings (anchored to canon citation), top 3 next actions (named approver where applicable), artifact path, and **one grill challenge** for the user. Examples:
- "Your deal scorecard shows 38% margin after discount. Skok's For Entrepreneurs benchmark says SaaS deals < 70% gross margin pre-discount need scrutiny. Did you model fulfillment cost or just COGS?"
- "Your packaging has 14 features in Better and 16 in Best. Madhavan Ramanujam (Monetizing Innovation): tiers with no clear differentiator make 70% of customers pick the cheapest. What's the one feature that forces an upgrade?"
## Forcing-question library (grill-with-docs pattern)
Grill the user on lane-defining decisions before invoking the sub-skill. One per turn, recommended answer, canon citation:
- **PRICING lane**: "Before picking a model: is your customer paying for outcomes, seats, or usage? Recommended: outcomes (value-based) if you can measure them. Anti-pattern (Ramanujam 2016 *Monetizing Innovation*): seat-based pricing on a usage-variable product caps your TAM at 20% of WTP."
- **DEAL lane**: "Before approving: what's the gross margin at full discount, **and** what does next quarter's pipeline look like at the same terms? Recommended: model both. Anti-pattern (Tunguz benchmarks): one 40% precedent reshapes 3 quarters of pipeline."
- **FORECAST lane**: "Before forecasting: are you using stage-conversion rates from the last 4 quarters, or the last 12? Recommended: last 4 weighted heavier. Anti-pattern (Skok, OpenView): equal-weighting 12 months hides the recent slowdown."
- **PARTNERSHIP lane**: "Before signing: does the partner have **independent demand**, or are they reselling our pipeline? Recommended: insist on indep demand evidence. Anti-pattern (Forrester channel research): channel-led deals from your own pipeline cost more than direct."
Never run a sub-skill until the lane-defining decision is locked.
## Assumptions
1. User has commercial authority OR is preparing analysis for someone who does.
2. User wants **deterministic decision support**, not the final answer — the human approves the deal, sets the price, signs the partner.
3. Inputs may be partial — every sub-skill ships templated dummy data so the user can see the shape before filling in their own.
## Non-goals
- Not a CRM, CPQ system, or contract repository.
- Does not auto-approve deals. Every output is **a score + recommendation + human-approver routing**.
- Does not store deal history across sessions.
## Distinct from
- **`business-growth/sales-engineer`** — that's the **technical sale** (demos, POCs). Commercial is **economic shape** of the deal.
- **`business-growth/revenue-operations`** — that's **process** (lead routing, SDR motion). Commercial is **per-deal economics + policy**.
- **`business-growth/contract-and-proposal-writer`** — that's **authoring** prose. Commercial is **decision logic + structured response**.
- **`c-level-advisor/cro-advisor`** — that's strategic CRO judgment ("when do we hire VP Sales?"). Commercial is tactical ("approve this discount").
- **`finance/financial-analysis`** — that's **close + report**. Commercial is **forecast + per-deal economics**.
## Output artifacts
| Sub-skill | Artifact |
|---|---|
| pricing-strategist | `pricing_model.md` + `wtp_analysis.json` |
| deal-desk | `deal_scorecard.md` + `discount_approval_routing.json` |
| partnerships-architect | `partner_tier_assignment.md` + `revshare_model.json` |
| channel-economics | `channel_mix_analysis.md` + `cost_to_serve.json` |
| commercial-policy | `commercial_policy.md` (discount matrix + exception flow) |
| rfp-responder | `rfp_response.md` + `winrate_estimate.json` |
| commercial-forecaster | `forecast.md` + `pipeline_math.json` |
## Anti-patterns (do not)
- ❌ Recommend a specific price — recommend a **range + model**, user picks the number
- ❌ Auto-approve discounts above policy — every >X% discount routes to a named human approver
- ❌ Generate an RFP response without proof points the user can verify
- ❌ Forecast bookings without surfacing the **conversion assumption** explicitly
- ❌ Run all 7 sub-skills "to be thorough" — pick one, digest, chain if needed
## References
- SaaS pricing canon: Tomasz Tunguz, David Skok, Bessemer Venture Partners
- Deal desk: SaaStr playbooks, Winning by Design
- Path-B build pattern: `documentation/implementation/bizops-commercial-expansion-plan.md`
Rà soát thương vụ trước khi chốt: chiết khấu vượt thẩm quyền, MSA bị sửa, lượng hóa biên lợi nhuận, thanh toán nhiều năm và rủi ro bồi thường.
---
name: deal-desk
description: Use when reviewing a specific inbound deal before close — when sales has asked for a discount that exceeds AE authority, when the customer has redlined the MSA, when per-deal economics (margin after discount, multi-year payment shape, indemnity exposure) need to be quantified, or when discount approval needs to be routed to a named human approver (Sales Director, VP Sales, CFO, CRO, General Counsel). Covers deal review, discount approval routing, per-deal margin scoring, deal exception handling, MSA redline triage, contract landmine detection (uncapped indemnity, MFN, perpetual license-back, missing DPA), and named-approver chain assembly. NEVER auto-approves — every output is a numeric scorecard plus a routing recommendation to a named human.
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [commercial, deal-desk, discount, margin, approval, redline, msa, terms]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# deal-desk
Per-deal review and discount-approval routing. Scores deal margin + risk, routes discount approval to the right human, redlines T&Cs against commercial policy. **Never auto-approves.** Every output is a score plus a routing recommendation to a named human approver.
## Purpose
Deal Desk / RevOps / sales leadership live at the moment between *sales-team-asks-for-discount* and *CFO/CRO/legal-signs*. This skill quantifies the asks and routes them.
Three deterministic tools:
1. `deal_scorer.py` — Scores a deal 0-100 across 5 dimensions (margin, risk, strategic value, commercial fit, term shape) and assigns one of four verdicts: **APPROVE / REVIEW / ESCALATE / DECLINE** — each tied to a named approver chain.
2. `discount_approval_router.py` — Maps a discount-percent + deal-size + tier to a named approver chain (AE → Manager → Director → VP → CFO/CRO) with estimated cycle days. Honors industry-tuned policy bands.
3. `terms_redliner.py` — Detects 10 founder/seller-killer patterns in deal terms (uncapped indemnity, MFN, perpetual license-back, missing DPA, NET-60+, broad non-solicit, etc.) with severity + standard counter + named legal/commercial approver.
## When to use
Invoke this skill when:
- Sales has flagged a discount request above AE authority.
- A customer has returned a redlined MSA and you need triage before routing to legal.
- The deal needs CFO sign-off and you want a defensible margin breakdown.
- An RFP response requires multi-year terms and you need to score the shape.
- A renewal expansion is bundled with a discount and you need to verify policy fit.
- You're building a deal-desk approval queue and need consistent routing.
**Do NOT use this skill to**: author the proposal (use `business-growth/contract-and-proposal-writer`), redesign the discount matrix (use the `commercial-policy` sibling skill), or do deep legal redline of full contract text (use `c-level-advisor/skills/general-counsel-advisor`).
## Workflow
1. **Intake the deal** — Sales/AE fills `assets/deal_intake_template.md` with ARR, term, discount, payment terms, customer tier, strategic flags, and any customer-flagged term redlines (20-min fill-out).
2. **Score margin + risk** — Run `deal_scorer.py --input deal.json --profile {saas|enterprise-software|services|marketplace}`. Read the composite + per-dimension breakdown + verdict.
3. **Route the discount** — Run `discount_approval_router.py --input deal.json --profile <same>`. Get the named approver chain + estimated cycle days. Modifiers (enterprise floor, SMB fast-lane) are surfaced explicitly.
4. **Flag the redlines** — Run `terms_redliner.py --input deal_terms.json`. Get ranked CRITICAL/HIGH/MEDIUM/LOW findings with the counter-language and the approver who must sign each.
5. **Assemble the packet** — Combine the three outputs into a deal-desk review packet. Always include the named approver chain. The packet is **a recommendation**, not an approval.
## Scripts
| Script | Purpose | Industry profiles |
|---|---|---|
| `scripts/deal_scorer.py` | 5-dimension scorecard with verdict + chain | saas, enterprise-software, services, marketplace |
| `scripts/discount_approval_router.py` | Discount % → named approver chain + cycle days | saas, enterprise-software, services, marketplace |
| `scripts/terms_redliner.py` | 10-pattern landmine scanner with counters | n/a (terms-driven) |
All three: stdlib-only, `--help`, `--sample`, `--input <json>`, `--output {human,json}`.
## References
- `references/deal_desk_canon.md` — Deal-desk operating practice: SaaStr playbooks (Jason Lemkin), Winning by Design (van der Kooij + Reichl), Forrester research, RevOps Co-op, OpenView benchmarks, Bridge Group AE comp, Salesforce Deal Desk best practices.
- `references/discount_economics.md` — Discount math + LTV impact: David Skok (For Entrepreneurs), Bessemer State of the Cloud, Tomasz Tunguz, OpenView NRR research, Pacific Crest + KeyBanc SaaS surveys, Insight Partners revenue ops. Includes worked margin math (a 30% discount on an 80% gross-margin product loses 37.5% of margin, not 30%).
- `references/contract_landmines.md` — 10+ named landmine patterns with example counter-language: YC startup library, Robert Klingberg (Founder's Guide to SaaS Agreements), Bowman + Brooke redline guides, IACCM/WorldCC commercial management research, Practical Law contracts library, Bradley Tusk on enterprise contracts, GC100 guidance.
## Assumptions
- The skill assumes the **commercial policy already exists** (discount bands, payment-terms norms, indemnity caps). It applies the policy; it does not design it. See the `commercial-policy` sibling skill for policy design.
- Industry profiles bake in *customary* thresholds. If your company has a documented discount matrix, pass it via `policy_thresholds` in the input JSON to override.
- The terms redliner detects the 10 most common landmines. It is **not** a substitute for General Counsel review on the full contract.
- Scoring weights (margin 30%, risk 20%, strategic 15%, commercial 20%, term 15%) reflect a CFO-leaning bias. RevOps-led shops may want to reweight; the weights are constants at the top of `score_deal()` and are easy to tune.
## Anti-patterns
- **Auto-approving deals.** This skill never says "approved". Every verdict (including `APPROVE`) names the human(s) who must sign. The output is a recommendation.
- **Skipping the redline scan** because the score is high. A high composite with `UNCAPPED_INDEMNITY` is still a DECLINE — critical signals override composite.
- **Using this for legal review of arbitrary contract text.** This skill takes a *structured* terms JSON. For prose redlining, use `c-level-advisor/skills/general-counsel-advisor/scripts/contract_risk_scanner.py`.
- **Treating the discount router as a discount calculator.** It routes a discount the AE/customer has already proposed; it does not calculate the right discount. Pricing logic lives in `commercial/skills/pricing-strategist`.
- **Routing every deal to CFO.** The router stops at the lowest-authority hop that can sign the deal. Over-escalation slows the funnel and trains AEs to over-discount.
- **Hand-editing the chain to skip a hop.** Modifiers (enterprise floor, SMB fast-lane) are explicit; hidden skips defeat the audit trail.
## Distinct from
| Sibling | Scope | Difference |
|---|---|---|
| `commercial/skills/pricing-strategist` | Sets the pricing **model** (per-seat vs usage vs tiered, list prices, packaging) | Operates at the strategy layer — not per deal |
| `business-growth/contract-and-proposal-writer` | **Authors** proposals, SOWs, MSAs | Output is a document; deal-desk is the gate **before** signing |
| `commercial/skills/commercial-policy` (sibling) | Designs the discount matrix and approval thresholds | Deal-desk **applies** that policy to one deal at a time |
| `c-level-advisor/skills/general-counsel-advisor` | Deep legal redline + term-sheet analysis | Operates on full contract prose; deal-desk uses structured terms JSON |
| `c-level-advisor/skills/cfo-advisor` | Burn rate, unit economics, fundraising models | Strategic finance; deal-desk is one-deal granularity |
## Quick examples
```bash
# Score a deal
python3 scripts/deal_scorer.py --sample
python3 scripts/deal_scorer.py --input my_deal.json --profile enterprise-software
# Route the discount
python3 scripts/discount_approval_router.py --sample
python3 scripts/discount_approval_router.py --input my_deal.json --profile saas
# Flag the redlines
python3 scripts/terms_redliner.py --sample
python3 scripts/terms_redliner.py --input my_deal_terms.json --output json
```
The sample (a 28%-discount enterprise SaaS deal with uncapped indemnity + MFN) correctly DECLINEs at 55.4 / 100 composite and routes to AE → Deal Desk → VP Sales → CFO → CRO → General Counsel.
## Forcing-question library (Matt Pocock grill discipline)
Walked one at a time by `/cs:grill-commercial` or the Commercial orchestrator. Recommended answer + canon citation per question. Never bundled.
1. **"What's the gross margin at full discount, AND what does next quarter's pipeline look like at the same terms?"**
Recommended: model both. Refuse to approve until the AE can articulate the precedent risk.
Canon: David Skok (For Entrepreneurs — discount math), Tomasz Tunguz benchmarks. Anti-pattern: one 40% precedent reshapes 3 quarters of pipeline.
2. **"Is this discount inside or outside the standard discount matrix?"**
Recommended: if outside, surface the policy exception explicitly and route to the named exception approver.
Canon: OpenView discount benchmarks, RevOps Co-op playbooks.
3. **"What's the strategic value beyond ARR — logo, reference, expansion path?"**
Recommended: require a named, verifiable expansion or reference commitment in writing.
Canon: SaaStr (Jason Lemkin) on logo discounts; Winning by Design on commitment language.
4. **"Has the customer signed an indemnity cap, a liability cap, and a DPA (if EU data)?"**
Recommended: required. Uncapped indemnity is a critical-signal override that blocks APPROVE regardless of margin.
Canon: WorldCC (formerly IACCM) commercial management research, GC100 contract guidance.
5. **"What payment terms — NET-30, NET-45, or NET-60+?"**
Recommended: prefer NET-30; NET-45+ is a cash flow drag worth quantifying.
Canon: KeyBanc SaaS Survey, Pacific Crest data — every 15 days of payment terms costs ~2% of effective deal value.
6. **"Is the term multi-year with annual prepay, or annual auto-renew?"**
Recommended: multi-year prepay > annual prepay > annual auto-renew. Auto-renew without 60-day notice is a redline.
Canon: Salesforce Deal Desk best practices, OpenView NRR studies.
7. **"Who is the named human approver at each hop of the discount chain?"**
Recommended: surface the name, not just the role. "VP Sales" is not an approver; "Maria Singh, VP Sales" is.
Canon: Bridge Group SaaS AE compensation research — named approval reduces precedent drift by 50%+.
Walk depth-first. Lock 1-4 before opening 5-7. After all 7 are answered, invoke `deal_scorer.py` → `discount_approval_router.py` → `terms_redliner.py` in sequence.
FILE:assets/deal_intake_template.md
# Deal Intake — Deal Desk Review
**Time to fill out: ~20 minutes.** This is the single source of truth for the deal. Re-pricings or term changes create a *new* intake — do not edit in place.
The structured fields at the bottom (the JSON blocks) feed directly into the three scripts:
- `deal_scorer.py` → consumes the **Deal Scorecard JSON**
- `discount_approval_router.py` → consumes the **Discount Routing JSON**
- `terms_redliner.py` → consumes the **Terms JSON**
---
## 1. Deal identity
| Field | Value |
|---|---|
| Deal ID | `ACME-2026-Q2-117` |
| Customer name | |
| AE / deal owner | |
| Sales engineer (if any) | |
| Date submitted | |
| Target close date | |
| Industry / segment | |
## 2. Commercial summary
| Field | Value |
|---|---|
| ARR (annual recurring revenue, $) | |
| Total contract value (TCV, $) | |
| Term (months) | |
| List price (TCV before discount, $) | |
| Discount (%) | |
| Customer tier | `enterprise` / `mid` / `smb` |
| Industry profile | `saas` / `enterprise-software` / `services` / `marketplace` |
## 3. Margin
| Field | Value |
|---|---|
| Product gross margin (%) | |
| Implementation / onboarding cost ($) | |
| Custom dev / SOW work in scope? | `yes` / `no` |
| If yes — services margin (%) | |
## 4. Strategic flags
Check each that applies. Each flag justifies *some* commercial flexibility but the discount scorer requires at least one for above-band discounts.
- [ ] **Logo** — reference-quality customer name; shortens future sales cycles.
- [ ] **Reference** — customer has agreed (in writing) to act as a reference / case study.
- [ ] **Expansion** — committed expansion plan in the next 12 months (named, quantified).
- [ ] **Renewal** — this is a renewal with multi-year extension.
## 5. Payment shape
| Field | Value |
|---|---|
| Payment terms (days from invoice) | |
| Billing frequency | `annual upfront` / `quarterly` / `monthly` |
| Multi-year discount applied? | `yes` / `no` |
| Up-front payment offered for discount? | `yes` / `no` |
## 6. Terms — customer-flagged redlines
List each clause the customer has flagged or modified. The scripts treat each entry as a risk signal.
1. ...
2. ...
3. ...
## 7. Structured terms (for `terms_redliner.py`)
Fill in the known structured fields:
| Term | Value |
|---|---|
| Auto-renew? | `true` / `false` |
| Auto-renew notice days | |
| Indemnity cap (multiple of fees, or `null` if uncapped) | |
| Liability cap (multiple of annual fees) | |
| DPA present? | `true` / `false` |
| EU personal data involved? | `true` / `false` |
| IP assignment | `vendor` / `customer` / `ambiguous` / `perpetual_license_back` |
| MFN clause present? | `true` / `false` |
| Exclusivity clause present? | `true` / `false` |
| Exclusivity compensated? | `true` / `false` |
| Non-solicit term (years) | |
| Governing law | |
| Vendor home jurisdiction | |
---
## 8. JSON skeletons — paste these into files for the scripts
### Deal Scorecard JSON (`deal.json`)
```json
{
"deal_id": "ACME-2026-Q2-117",
"customer_name": "Acme Corp",
"arr": 240000,
"term_months": 24,
"discount_pct": 28.0,
"payment_terms_days": 60,
"list_price": 333333,
"gross_margin_pct": 78.0,
"customer_tier": "enterprise",
"strategic_value": {
"logo": true,
"reference": false,
"expansion": true,
"renewal": false
},
"term_redlines": [
"uncapped indemnity",
"MFN pricing"
]
}
```
Run:
```bash
python3 scripts/deal_scorer.py --input deal.json --profile saas
```
### Discount Routing JSON (`discount.json`)
```json
{
"deal_id": "ACME-2026-Q2-117",
"discount_pct": 28.0,
"deal_size_arr": 240000,
"customer_tier": "enterprise",
"policy_thresholds": null
}
```
Run:
```bash
python3 scripts/discount_approval_router.py --input discount.json --profile saas
```
### Terms JSON (`deal_terms.json`)
```json
{
"deal_id": "ACME-2026-Q2-117",
"payment_terms_days": 60,
"auto_renew": true,
"auto_renew_notice_days": 90,
"indemnity_cap": null,
"liability_cap": 1.0,
"dpa_present": false,
"eu_data_involved": true,
"ip_assignment": "ambiguous",
"mfn_clause_present": true,
"exclusivity_clause_present": false,
"exclusivity_compensated": false,
"non_solicit_years": 3,
"governing_law": "Delaware",
"vendor_home_jurisdiction": "Delaware"
}
```
Run:
```bash
python3 scripts/terms_redliner.py --input deal_terms.json
```
---
## 9. Reviewer checklist
Before submitting the intake to the deal desk:
- [ ] All commercial fields populated (no blanks in section 2).
- [ ] Strategic flags reflect *committed*, not hoped-for, value.
- [ ] All customer-flagged redlines listed in section 6.
- [ ] Structured terms in section 7 match the actual marked-up contract.
- [ ] JSON skeletons (section 8) saved to files.
The deal-desk packet that comes back will name the approver(s) who must sign. **The skill never approves the deal itself.**
FILE:references/contract_landmines.md
# Contract Landmines
The 10 founder/seller-killer patterns the `terms_redliner.py` tool detects, with example counter-language for each. This is a triage reference, **not** legal advice — every HIGH/CRITICAL finding must be reviewed by named counsel before signing.
For deep prose-level redline of an actual contract, use `c-level-advisor/skills/general-counsel-advisor/scripts/contract_risk_scanner.py`. The tool in this skill operates on a *structured terms JSON*, which is what the deal desk typically has from the intake template.
## The 10 patterns
### 1. UNCAPPED_INDEMNITY (CRITICAL)
**Trigger**: `indemnity_cap` is `null` or absent.
**Why it matters**: A single indemnity claim can be larger than the entire ARR of the deal — sometimes larger than the company's revenue. Uncapped indemnity is the most common contract risk that destroys early-stage companies.
**Counter-language**:
> "Each party's aggregate liability for indemnification obligations shall not exceed twelve (12) times the monthly subscription fees paid in the twelve (12) months preceding the claim, except for breaches of confidentiality, willful misconduct, or third-party intellectual-property infringement, for which a super-cap of three (3) times annual fees shall apply."
**Approver**: General Counsel + CFO.
### 2. MISSING_DPA_EU_DATA (CRITICAL)
**Trigger**: `eu_data_involved == True` and `dpa_present == False`.
**Why it matters**: GDPR Article 28 mandates a Data Processing Agreement when personal data of EU residents is processed by a service provider. Missing DPA = (a) regulatory exposure under GDPR, (b) immediate audit failure on any SOC 2 or ISO 27001 review, (c) customer escalation to their privacy officer.
**Counter-language**: Attach standard DPA (2021/914 Standard Contractual Clauses, or vendor's own template) as an exhibit. **Do not sign the master agreement until the DPA is countersigned.**
**Approver**: General Counsel + DPO.
### 3. MFN_PRICING (HIGH)
**Trigger**: `mfn_clause_present == True`.
**Why it matters**: Most-Favored-Nation clauses bind the seller to refund the customer (or extend matching terms) if any other customer gets a better price. This freezes pricing innovation: no bundles, no segment pricing, no competitive deals without triggering MFN obligations across the base.
**Counter-language**:
> "Strike Section [X] (Most-Favored-Nation Pricing) in its entirety. If retained, scope to: same SKU, same volume tier, same contract term, same geography, and same industry vertical; and time-bound to twelve (12) months from the Effective Date."
**Approver**: VP Sales + CFO.
### 4. AUTORENEW_LONG_NOTICE (HIGH)
**Trigger**: `auto_renew == True` and `auto_renew_notice_days > 30`.
**Why it matters**: Auto-renewal with a long notice window (60, 90, 120 days) is a classic vendor trap. Customers miss the window and get locked into another full term — but this also goes the other way: as a seller, accepting 60+ day notice on your own auto-renewals gives the buyer asymmetric exit.
**Counter-language**:
> "Either party may provide written notice of non-renewal not less than thirty (30) days prior to the end of the then-current term."
**Approver**: Deal Desk + General Counsel.
### 5. PERPETUAL_LICENSE_BACK (CRITICAL)
**Trigger**: `ip_assignment == "perpetual_license_back"`.
**Why it matters**: A perpetual license-back gives the customer the right to use the vendor's IP **forever**, often royalty-free and surviving termination. This kills the moat — the customer can stop paying and keep using.
**Counter-language**:
> "Customer's license to the Services and Vendor IP is co-terminus with the Subscription Term, field-of-use limited to internal business operations, non-transferable, non-sublicensable, and terminates upon any termination or expiration of this Agreement."
**Approver**: General Counsel + CEO.
### 6. AMBIGUOUS_IP (HIGH)
**Trigger**: `ip_assignment == "ambiguous"`.
**Why it matters**: Ambiguous IP ownership becomes a dispute at acquisition diligence. Buyers will hold back purchase price (or walk) until IP chain-of-title is clarified. Costs weeks of legal time and can break an M&A deal.
**Counter-language**:
> "Vendor retains all right, title, and interest in and to the Services, the Vendor IP, and any improvements, modifications, or derivatives thereof developed in connection with this Agreement. Customer retains all right, title, and interest in Customer Data and in any outputs derived solely from Customer Data."
**Approver**: General Counsel.
### 7. EXCLUSIVITY_UNCOMPENSATED (CRITICAL)
**Trigger**: `exclusivity_clause_present == True` and `exclusivity_compensated == False`.
**Why it matters**: Exclusivity removes the entire competitive segment of the addressable market for no economic benefit. Even *paid* exclusivity needs a kill switch on missed quarterly minimums — otherwise the buyer locks the seller into the segment without performance pressure.
**Counter-language**:
> "Strike exclusivity in its entirety. If retained, exclusivity is contingent on Minimum Guaranteed Spend of $[X] per quarter, payable in advance, and Vendor may terminate exclusivity (while preserving the underlying agreement) upon two consecutive quarters of MGS shortfall."
**Approver**: CRO + General Counsel.
### 8. LONG_PAYMENT_TERMS (HIGH)
**Trigger**: `payment_terms_days > 45`.
**Why it matters**: NET-60/75/90/120 inflates DSO and ties up working capital. A $200K deal on NET-90 is effectively $200K of zero-interest financing extended to the buyer. Material on any deal that's > 10% of cash balance.
**Counter-language**:
> "Payment terms shall be NET-30 from invoice date. Customer may elect NET-15 prepay terms in exchange for a 1.5% prepayment discount. Late payments accrue interest at 1.5% per month or the maximum permitted by law, whichever is lower."
**Approver**: CFO + Deal Desk.
### 9. LOW_LIABILITY_CAP (MEDIUM)
**Trigger**: `liability_cap < 1.0` (multiple of annual fees).
**Why it matters**: When the customer pushes for a sub-1x liability cap, they're usually expecting outsized claims. Don't accept without symmetric protection (mutual cap, both directions).
**Counter-language**:
> "Each party's aggregate liability shall not exceed one (1) times the fees paid by Customer in the twelve (12) months preceding the claim, except for breaches of confidentiality, IP infringement, or willful misconduct, for which a super-cap of three (3) times annual fees shall apply. This cap is mutual and applies to both parties."
**Approver**: General Counsel.
### 10. BROAD_NON_SOLICIT (MEDIUM)
**Trigger**: `non_solicit_years >= 2`.
**Why it matters**: Multi-year non-solicit clauses limit hiring and are increasingly unenforceable in many US jurisdictions (notably California, where they are void as a matter of public policy except in narrow circumstances). Negotiate down.
**Counter-language**:
> "Each party agrees not to solicit for employment any employee of the other party who was directly engaged on the project for a period of twelve (12) months following such employee's last day of engagement on the project. This restriction does not apply to general advertising, solicitation through public job boards, or responses to unsolicited inquiries."
**Approver**: General Counsel + CHRO.
## Sources
1. **Y Combinator — Startup Library** — Sam Altman's and the YC partners' canonical guidance on contracts founders sign. https://www.ycombinator.com/library
2. **Robert Klingberg — *Founder's Guide to SaaS Agreements*** — Practitioner reference on SaaS-specific contract patterns (MSA, DPA, BAA, MNDA).
3. **Bowman + Brooke — Contract Redline Guides** — Defense-side commercial litigation firm's published guides on enterprise contract risk.
4. **IACCM / WorldCC — World Commerce & Contracting Research** — The trade association for commercial contracting; annual surveys of *the most negotiated terms* and *the most disputed terms* in B2B contracts. https://www.worldcc.com/
5. **Practical Law (Thomson Reuters) — Contracts Library** — Standard clause library + redline best practices used by AmLaw 100 firms.
6. **Bradley Tusk — *The Fixer: My Adventures Saving Startups from Death by Politics*** — Practical advice on enterprise contracts, including the patterns that destroy young companies.
7. **GC100 — General Counsel Forum** — Senior in-house counsel from FTSE 100 companies; their guidance on commercial contract risk allocation. https://www.gc100.co.uk/
8. **American Bar Association — *Model Software License Provisions*** — Reference for industry-standard software licensing terms.
## How to use this reference
1. The deal-desk intake template asks the AE to capture the structured terms.
2. `terms_redliner.py --input deal_terms.json` produces a ranked list of detected landmines.
3. Each landmine is mapped to a section in this document with the counter-language and named approver.
4. The deal-desk packet attaches the counter-language so the AE can return to the customer with a defensible position.
Remember: **every CRITICAL or HIGH finding must reach the named approver before the deal closes.** This skill triages; it does not approve.
FILE:references/deal_desk_canon.md
# Deal Desk Canon
Operating practice for per-deal review and approval routing in B2B SaaS / enterprise software. Compiled from authoritative deal-desk and revenue-operations sources.
## Why a deal desk exists
The deal desk is the **operational gate between sales and finance/legal**. Its job:
1. **Standardize discount approval** so the same discount-percent always routes the same way.
2. **Defend gross margin** by quantifying the actual margin loss from a proposed discount (not just the discount percent).
3. **Triage commercial terms** so legal review hits only the deals that need it.
4. **Speed up the deals that should be fast** by routing simple deals to AE/Manager authority and reserving CFO/CRO attention for the consequential ones.
Without a deal desk, every above-band deal becomes a 1:1 negotiation between an AE and a finance leader, which is slow, inconsistent, and creates pricing-integrity drift over time.
## Operating tenets
These are the non-negotiables — adopted across every reference cited below.
1. **Never auto-approve.** Even green deals get a named approver. The skill outputs *who must sign*, not *the deal is fine*.
2. **Margin, not discount.** A 30% discount on an 80%-gross-margin product reduces *margin* by 24 points (to 56%) — not 30%. See `discount_economics.md` for the math.
3. **The chain stops at the lowest hop that has authority.** Over-routing trains reps to over-discount because they expect VP attention anyway.
4. **Critical signals override composite.** A high-composite deal with uncapped indemnity is still a DECLINE.
5. **Modifiers must be explicit.** Enterprise floor (large ARR forces VP review) and SMB fast-lane (small deals can skip a hop) are surfaced; hidden adjustments destroy audit trails.
6. **The deal desk is a router, not a salesperson.** It does not negotiate; it routes the negotiation to the named human.
7. **One source of truth per deal.** The intake template is the spec. Re-pricings or term changes create a new intake, not an edit-in-place.
## Standard approval bands (industry-customary)
Default policy (override with `policy_thresholds` in input JSON):
| Discount band | Approver | Typical cycle |
|---|---|---|
| 0% - 15% | AE | same-day |
| 15% - 25% | Sales Manager | 1 business day |
| 25% - 35% | Director of Sales | 2 business days |
| 35% - 50% | VP Sales | 3 business days |
| 50%+ | CFO + CRO | 5+ business days |
Enterprise-software profile shifts bands upward (larger ACVs absorb deeper discounts). Services profile shifts downward (margin-thin). Marketplace profile is tightly capped (take-rate is the lever).
## Tier and ARR modifiers
- **Enterprise floor**: Deals at ARR >= profile threshold force VP-level review even on small discounts. Rationale: the customer is consequential regardless of the discount.
- **SMB fast-lane**: Deals at ARR <= profile threshold can drop one hop (only if discount is within the second band). Rationale: cycle time matters more than marginal margin defense on a $12K deal.
## Sources
1. **SaaStr** — Jason Lemkin's deal-desk playbooks emphasize that the deal desk's primary job is *defending gross margin and pricing integrity*, not just routing discounts. https://www.saastr.com/
2. **Winning by Design** — Jacco van der Kooij + Jason Reichl, *Bowtie Funnel* and *Revenue Architecture*. Establishes that the deal desk owns the gate between Acquisition (sales) and Retention (CS) — bad-term deals cost more in churn than they earn in ARR. https://winningbydesign.com/
3. **Forrester Research** — Deal desk maturity model (4 stages: ad-hoc → formal → strategic → predictive). Most companies hit a wall at stage 2 because they lack the data infrastructure to score deals consistently.
4. **RevOps Co-op** — Community playbooks (operating notes from Iceberg RevOps, Sapphire Ventures, others). Emphasizes that the deal desk is a **routing function**, not an approval function. The named approver is always a human.
5. **OpenView Venture Partners** — *State of the SaaS Sales Org* annual benchmarks. Documents discount-band conventions across stage (seed → growth → late-stage) and shows that median discount creeps up year-over-year unless deal-desk discipline is enforced. https://openviewpartners.com/
6. **Bridge Group SaaS AE Compensation Research** — Annual survey of B2B SaaS AE comp + quota. Establishes that AE discount authority above 15-20% destroys quota attainment math (because the AE under-prices to close).
7. **Salesforce Deal Desk Best Practices** — Internal Salesforce documentation (Trailhead + RevOps blog). Codifies the queue model: every above-AE-authority deal enters a queue with SLA. Aging deals escalate automatically.
## Patterns to surface in any deal-desk review packet
- Composite score with per-dimension breakdown.
- Named approver chain with the hop where the discount lands highlighted.
- Estimated cycle days based on hop count.
- Any CRITICAL signals (uncapped indemnity, MFN, perpetual license-back, missing DPA).
- The standard counter-language for any HIGH/CRITICAL redline.
- A **single explicit statement**: "This is a routing recommendation. The named approvers must sign."
FILE:references/discount_economics.md
# Discount Economics
The math of what a discount actually costs. Most sales discounts are described as a list-price reduction; the real impact is on **gross margin** and **LTV**, both of which compound across the customer base over time.
## The fundamental formula
A discount of D% on a product with gross margin G% reduces net margin by:
margin_loss_points = D * (G / 100)
net_margin = G - margin_loss_points
### Worked examples
| List discount | Gross margin | Margin loss | Net margin |
|---|---|---|---|
| 10% | 80% | 8 pts | 72% |
| 20% | 80% | 16 pts | 64% |
| **30%** | **80%** | **24 pts** | **56%** |
| 30% | 60% | 18 pts | 42% |
| 40% | 80% | 32 pts | 48% |
| 50% | 80% | 40 pts | 40% |
**A 30% discount on an 80%-gross-margin product wipes 24 points of margin** — that's a 30% margin loss in *relative* terms (24/80 = 30%), but the conventional shorthand "30% discount = 30% margin hit" understates the absolute hit on a low-margin product.
### Why the conventional shorthand is wrong
People often say "a 30% discount loses 30% of margin." That's only true for a 100%-margin product. For an 80%-margin SaaS, the discount cuts the **revenue** by 30% but the **margin** by 30% × (80/100) = 24 points, or 30% in relative terms. The dollar impact compounds across the contract term.
## LTV impact
Discount also compounds across multi-year contracts. A 24-month deal at 30% discount loses:
lifetime_margin_loss = (D / 100) * G/100 * list_price * (term_months / 12)
For a $200K-ARR deal at 30% discount, 80% gross margin, 24-month term:
= 0.30 * 0.80 * 200,000 * 2 = $96,000 of gross margin given up
That's $96K of fully-loaded P&L impact for one deal. Across 50 deals/quarter at the same discount, the company is giving up $19.2M/year in gross margin.
## Discount creep
The most-cited dataset (Pacific Crest / KeyBanc SaaS Survey) shows median discount rises ~1.5 pts/year unless the deal desk actively defends pricing. Causes:
1. AE comp on bookings, not margin → AEs discount to close.
2. Multi-year deals trade discount for term length but term length doesn't recover the margin loss if churn risk is non-zero.
3. Competitive deals get matched discounts that then propagate to non-competitive deals via MFN clauses.
4. Renewal discounts (CS giving discount to retain) anchor the next renewal lower.
## When a discount is justified
The deal desk should approve a discount when **at least one** of these is true and quantified:
1. **Strategic logo** — the customer is a reference account that materially shortens future sales cycles. Logo value ≥ discount $.
2. **Expansion lock-in** — the discount is paired with a *multi-year + expansion commitment* that recovers margin over the contract term.
3. **Competitive displacement** — the discount displaces an incumbent and the lifetime ARR > displacement cost.
4. **Cash-acceleration** — payment up-front in exchange for discount, where the cash NPV recovers the margin loss.
The deal scorer's `strategic` dimension flags logo / reference / expansion / renewal explicitly. If none of those are set, a discount above the policy band is presumptively unjustified.
## NRR + discount correlation
OpenView's *State of the SaaS Industry* shows companies with high NRR (≥ 120%) discount less on initial deals than companies with low NRR (≤ 100%). The mechanism: high-NRR companies have a strong expansion motion that they don't need to buy with up-front discount; low-NRR companies discount up-front to compensate for weak expansion.
This is why deal-desk should treat "discount to close" as a leading indicator of NRR weakness, not a one-deal problem.
## Sources
1. **David Skok — For Entrepreneurs** — *SaaS Metrics 2.0* and *The SaaS Business Model*. Canonical treatment of LTV/CAC + the impact of discount on payback period. https://www.forentrepreneurs.com/
2. **Bessemer Venture Partners — State of the Cloud** — Annual report with discount benchmarks by ACV band ($1K, $10K, $100K, $1M+) and stage. https://www.bvp.com/
3. **Tomasz Tunguz — Redpoint** — Multi-year studies on discount-to-close patterns, including the finding that median enterprise SaaS discount sits at 18-22% across the industry. https://tomtunguz.com/
4. **OpenView Venture Partners** — *State of the SaaS Industry* + Expansion Economics research. Documents the NRR-vs-discount correlation. https://openviewpartners.com/
5. **Pacific Crest SaaS Survey** (now KeyBanc Capital Markets) — Annual primary-research survey of B2B SaaS companies. Most-cited dataset for discount benchmarks. https://www.key.com/businesses-institutions/industry-expertise/saas-survey.html
6. **KeyBanc Capital Markets SaaS Survey** — Continuation of Pacific Crest. Annual benchmark for net dollar retention, gross margin, and discount-by-segment.
7. **Insight Partners Revenue Operations Research** — Their PitchBook + portfolio data on discount discipline at growth-stage SaaS. https://www.insightpartners.com/
## Patterns to surface in any margin review
- Pre-discount gross margin and post-discount net margin in **absolute points**, not just percent.
- Lifetime margin given up over the contract term, in dollars.
- Whether the strategic flags justify the discount (logo / reference / expansion / renewal).
- Whether the customer is paying up-front in exchange for the discount (cash NPV).
- Comparison to the company's median deal-discount (drift signal).
FILE:scripts/deal_scorer.py
#!/usr/bin/env python3
"""deal_scorer.py - Score an inbound deal across 5 dimensions and route the verdict.
Stdlib-only. NEVER auto-approves. Output is always a numeric breakdown plus a verdict
(APPROVE / REVIEW / ESCALATE / DECLINE) and a NAMED HUMAN APPROVER chain.
The 5 dimensions (each 0-100, weighted into a composite):
1. margin - post-discount gross margin vs profile target
2. risk - payment terms + redline count + customer tier
3. strategic - logo / reference / expansion / renewal value
4. commercial - is the discount within the profile policy band
5. term shape - multi-year + payment-up-front vs short, NET-60+ tail
Routing rule (intentionally conservative):
- composite >= 80 and no CRITICAL signals -> APPROVE (still names the approver)
- composite 65-79 -> REVIEW (Deal Desk + Sales Director)
- composite 50-64 or 1 CRITICAL -> ESCALATE (VP Sales + CFO)
- composite < 50 or 2+ CRITICAL -> DECLINE (CRO + CFO must sign off any override)
Usage:
python deal_scorer.py --sample
python deal_scorer.py --input deal.json --profile saas
python deal_scorer.py --input deal.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from typing import Any
SAMPLE_DEAL = {
"deal_id": "ACME-2026-Q2-117",
"customer_name": "Acme Corp",
"arr": 240000,
"term_months": 24,
"discount_pct": 28.0,
"payment_terms_days": 60,
"list_price": 333333,
"gross_margin_pct": 78.0,
"customer_tier": "enterprise",
"strategic_value": {
"logo": True,
"reference": False,
"expansion": True,
"renewal": False,
},
"term_redlines": [
"uncapped indemnity",
"MFN pricing",
],
}
# Industry profiles tune the target margin floor, acceptable discount band,
# and payment-terms tolerance.
PROFILES: dict[str, dict[str, Any]] = {
"saas": {
"target_gross_margin": 75.0,
"discount_band_pct": 25.0,
"max_payment_terms_days": 30,
"preferred_term_months": 24,
},
"enterprise-software": {
"target_gross_margin": 70.0,
"discount_band_pct": 35.0,
"max_payment_terms_days": 45,
"preferred_term_months": 36,
},
"services": {
"target_gross_margin": 45.0,
"discount_band_pct": 15.0,
"max_payment_terms_days": 30,
"preferred_term_months": 12,
},
"marketplace": {
"target_gross_margin": 30.0,
"discount_band_pct": 10.0,
"max_payment_terms_days": 14,
"preferred_term_months": 12,
},
}
# Routing chain by composite + signals. The skill NEVER says "approved" by itself;
# it names the human(s) who must sign.
APPROVER_CHAIN = {
"APPROVE": ["AE", "Deal Desk Analyst", "Sales Director"],
"REVIEW": ["AE", "Deal Desk Analyst", "Sales Director", "VP Sales"],
"ESCALATE": ["AE", "Deal Desk Analyst", "Sales Director", "VP Sales", "CFO", "CRO"],
"DECLINE": ["AE", "Deal Desk Analyst", "VP Sales", "CFO", "CRO", "General Counsel"],
}
@dataclass
class DimensionScore:
name: str
score: float
weight: float
rationale: str
@dataclass
class DealScorecard:
deal_id: str
profile: str
composite_score: float
verdict: str
approver_chain: list[str]
dimensions: list[DimensionScore] = field(default_factory=list)
critical_signals: list[str] = field(default_factory=list)
notes: list[str] = field(default_factory=list)
def _clamp(x: float, lo: float = 0.0, hi: float = 100.0) -> float:
return max(lo, min(hi, x))
def score_margin(deal: dict, profile: dict) -> DimensionScore:
"""Effective margin after discount, compared to profile target.
Math: a D% discount on a product with gross_margin_pct G% drops margin to
new_margin = (G - D) / (1 - D/100) approximately, but the canonical
formulation we use is: net_margin = G - (D * (1 - cost_ratio)) which
resolves to:
net_margin = G - D * (G / 100)
i.e. a 30% discount on an 80% margin product wipes 24 points of margin,
leaving 56% — well below an 75% SaaS target.
"""
g = float(deal.get("gross_margin_pct", 0.0))
d = float(deal.get("discount_pct", 0.0))
net_margin = g - (d * (g / 100.0))
target = profile["target_gross_margin"]
# Score: 100 if net_margin >= target, sliding to 0 at (target - 30 pts)
delta = net_margin - target
score = _clamp(100.0 + (delta / 30.0) * 100.0)
rationale = (
f"Gross margin {g:.1f}% with {d:.1f}% discount -> net margin {net_margin:.1f}% "
f"vs profile target {target:.1f}% (delta {delta:+.1f} pts)"
)
return DimensionScore("margin", round(score, 1), 0.30, rationale)
def score_risk(deal: dict, profile: dict) -> DimensionScore:
"""Risk = payment terms shape + redline count + customer-tier offset."""
payment_days = int(deal.get("payment_terms_days", 30))
redlines = deal.get("term_redlines", []) or []
tier = (deal.get("customer_tier") or "smb").lower()
# Base score 100, deduct per risk factor.
score = 100.0
payment_max = profile["max_payment_terms_days"]
if payment_days > payment_max:
over = payment_days - payment_max
score -= min(40.0, over * 0.8) # NET-90 vs NET-30 = 48 days over = -38.4
score -= min(40.0, len(redlines) * 12.0) # each redline = -12
# SMB tier on long terms is riskier than enterprise on same terms
if tier == "smb" and payment_days > 30:
score -= 10.0
elif tier == "enterprise" and payment_days <= 45:
score += 5.0 # enterprise tolerance bump
score = _clamp(score)
rationale = (
f"NET-{payment_days} terms (profile max {payment_max}), "
f"{len(redlines)} redline(s), tier={tier}"
)
return DimensionScore("risk", round(score, 1), 0.20, rationale)
def score_strategic(deal: dict, profile: dict) -> DimensionScore:
"""Strategic value from logo, reference, expansion, renewal flags."""
sv = deal.get("strategic_value", {}) or {}
weights = {"logo": 25, "reference": 20, "expansion": 30, "renewal": 25}
earned = sum(w for k, w in weights.items() if sv.get(k))
rationale = "Flags: " + ", ".join(k for k in weights if sv.get(k)) if earned else "No strategic flags set"
return DimensionScore("strategic", float(earned), 0.15, rationale)
def score_commercial(deal: dict, profile: dict) -> DimensionScore:
"""Is the discount within the profile's policy band?"""
d = float(deal.get("discount_pct", 0.0))
band = profile["discount_band_pct"]
if d <= band:
# Within band, score linearly from 100 (no discount) to 80 (band edge)
score = 100.0 - (d / band) * 20.0
rationale = f"Discount {d:.1f}% within policy band <= {band:.1f}%"
else:
over = d - band
# Drop 6 points per percentage over band, floor 0
score = max(0.0, 80.0 - over * 6.0)
rationale = f"Discount {d:.1f}% EXCEEDS policy band {band:.1f}% by {over:.1f} pts"
return DimensionScore("commercial", round(score, 1), 0.20, rationale)
def score_term_shape(deal: dict, profile: dict) -> DimensionScore:
"""Term length vs preferred + payment up front."""
term_months = int(deal.get("term_months", 12))
preferred = profile["preferred_term_months"]
payment_days = int(deal.get("payment_terms_days", 30))
# Length component: 100 if >= preferred, sliding to 40 at half-preferred, floor 30
if term_months >= preferred:
length = 100.0
elif term_months <= preferred / 2:
length = 30.0
else:
length = 30.0 + ((term_months - preferred / 2) / (preferred / 2)) * 70.0
# Payment component: NET-30 or shorter = 100, NET-60 = 70, NET-90+ = 40
if payment_days <= 30:
pay = 100.0
elif payment_days <= 60:
pay = 70.0
elif payment_days <= 90:
pay = 40.0
else:
pay = 20.0
score = 0.6 * length + 0.4 * pay
rationale = (
f"{term_months}-mo term (preferred {preferred}), NET-{payment_days} payment "
f"-> length={length:.0f}, payment={pay:.0f}"
)
return DimensionScore("term_shape", round(score, 1), 0.15, rationale)
def _detect_critical_signals(deal: dict, dims: list[DimensionScore]) -> list[str]:
sigs: list[str] = []
redlines = [r.lower() for r in deal.get("term_redlines", []) or []]
critical_terms = (
"uncapped indemnity",
"uncapped liability",
"mfn",
"most-favored-nation",
"perpetual license-back",
"exclusivity",
)
for r in redlines:
if any(ct in r for ct in critical_terms):
sigs.append(f"critical redline: {r}")
# margin below 35% is a critical economic signal on any profile
for d in dims:
if d.name == "margin" and d.score < 30.0:
sigs.append("margin below target by >30 pts")
if d.name == "commercial" and d.score < 30.0:
sigs.append("discount far outside policy band")
return sigs
def _verdict(composite: float, criticals: list[str]) -> str:
n_crit = len(criticals)
if n_crit >= 2 or composite < 50.0:
return "DECLINE"
if n_crit == 1 or composite < 65.0:
return "ESCALATE"
if composite < 80.0:
return "REVIEW"
return "APPROVE"
def score_deal(deal: dict, profile_name: str = "saas") -> DealScorecard:
if profile_name not in PROFILES:
raise ValueError(f"Unknown profile '{profile_name}'. Choose from {list(PROFILES)}.")
profile = PROFILES[profile_name]
dims = [
score_margin(deal, profile),
score_risk(deal, profile),
score_strategic(deal, profile),
score_commercial(deal, profile),
score_term_shape(deal, profile),
]
composite = sum(d.score * d.weight for d in dims)
criticals = _detect_critical_signals(deal, dims)
verdict = _verdict(composite, criticals)
notes = [
"This skill does NOT auto-approve. The approver chain below is who must sign.",
f"Composite is weighted: margin 30, risk 20, strategic 15, commercial 20, term 15.",
]
if criticals:
notes.append(f"{len(criticals)} critical signal(s) detected; cannot APPROVE.")
return DealScorecard(
deal_id=str(deal.get("deal_id", "UNSPECIFIED")),
profile=profile_name,
composite_score=round(composite, 1),
verdict=verdict,
approver_chain=APPROVER_CHAIN[verdict],
dimensions=dims,
critical_signals=criticals,
notes=notes,
)
def _render_human(card: DealScorecard) -> str:
lines = []
lines.append(f"Deal Scorecard: {card.deal_id}")
lines.append(f"Profile: {card.profile}")
lines.append(f"Composite Score: {card.composite_score}/100")
lines.append(f"Verdict: {card.verdict}")
lines.append("")
lines.append("Dimension breakdown:")
for d in card.dimensions:
lines.append(f" - {d.name:10s} {d.score:5.1f} (weight {d.weight:.2f})")
lines.append(f" {d.rationale}")
lines.append("")
if card.critical_signals:
lines.append("Critical signals:")
for s in card.critical_signals:
lines.append(f" ! {s}")
lines.append("")
lines.append("Approver chain (named humans who must sign):")
lines.append(" " + " -> ".join(card.approver_chain))
lines.append("")
for n in card.notes:
lines.append(f"note: {n}")
return "\n".join(lines)
def _to_jsonable(card: DealScorecard) -> dict:
d = asdict(card)
return d
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Score a deal across 5 dimensions and route to a named approver.",
)
parser.add_argument("--input", help="Path to JSON deal context")
parser.add_argument("--profile", default="saas", choices=list(PROFILES))
parser.add_argument("--output", default="human", choices=["human", "json"])
parser.add_argument("--sample", action="store_true", help="Use embedded sample deal")
args = parser.parse_args(argv)
if args.sample or not args.input:
deal = SAMPLE_DEAL
else:
with open(args.input) as f:
deal = json.load(f)
card = score_deal(deal, args.profile)
if args.output == "json":
print(json.dumps(_to_jsonable(card), indent=2))
else:
print(_render_human(card))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/discount_approval_router.py
#!/usr/bin/env python3
"""discount_approval_router.py - Route a discount request to the right human(s).
Stdlib-only. Outputs the NAMED APPROVER CHAIN, the hop where this deal lands,
and an estimated approval-cycle in business days. The skill never says "approved" —
only "routes to <person>".
Default policy bands (industry-customary, can be overridden in input JSON):
0% - 15% AE-approved
15% - 25% Sales Manager
25% - 35% Director of Sales
35% - 50% VP Sales
50% + CFO / CRO
Deal-size and tier modifiers nudge the chain (e.g. enterprise deal > $500K ARR
ALWAYS requires VP review even at 10% discount; SMB deal < $25K ARR may stop
one hop earlier for speed).
Usage:
python discount_approval_router.py --sample
python discount_approval_router.py --input deal.json --profile saas
python discount_approval_router.py --input deal.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, asdict, field
from typing import Any
SAMPLE_INPUT = {
"deal_id": "ACME-2026-Q2-117",
"discount_pct": 32.0,
"deal_size_arr": 240000,
"customer_tier": "enterprise",
"policy_thresholds": None, # use defaults
}
DEFAULT_BANDS = [
{"max_pct": 15.0, "approver": "AE", "days": 0},
{"max_pct": 25.0, "approver": "Sales Manager", "days": 1},
{"max_pct": 35.0, "approver": "Director of Sales", "days": 2},
{"max_pct": 50.0, "approver": "VP Sales", "days": 3},
{"max_pct": 100.1, "approver": "CFO + CRO", "days": 5},
]
PROFILES: dict[str, dict[str, Any]] = {
"saas": {
"bands": DEFAULT_BANDS,
"enterprise_floor_approver": "VP Sales",
"enterprise_floor_arr": 500000,
"smb_fast_lane_arr": 25000,
},
"enterprise-software": {
# Larger ACVs absorb deeper discounts; bands shift up
"bands": [
{"max_pct": 20.0, "approver": "AE", "days": 0},
{"max_pct": 30.0, "approver": "Sales Manager", "days": 1},
{"max_pct": 40.0, "approver": "Director of Sales", "days": 2},
{"max_pct": 55.0, "approver": "VP Sales", "days": 4},
{"max_pct": 100.1, "approver": "CFO + CRO", "days": 7},
],
"enterprise_floor_approver": "VP Sales",
"enterprise_floor_arr": 1000000,
"smb_fast_lane_arr": 50000,
},
"services": {
# Margin-thin: even small discounts go up the chain fast
"bands": [
{"max_pct": 5.0, "approver": "AE", "days": 0},
{"max_pct": 12.0, "approver": "Sales Manager", "days": 1},
{"max_pct": 20.0, "approver": "Director of Sales", "days": 2},
{"max_pct": 30.0, "approver": "VP Services", "days": 3},
{"max_pct": 100.1, "approver": "CFO + COO", "days": 5},
],
"enterprise_floor_approver": "VP Services",
"enterprise_floor_arr": 250000,
"smb_fast_lane_arr": 10000,
},
"marketplace": {
# Take-rate is the lever; explicit discounts are rare and tightly capped
"bands": [
{"max_pct": 3.0, "approver": "AE", "days": 0},
{"max_pct": 8.0, "approver": "Sales Manager", "days": 1},
{"max_pct": 15.0, "approver": "Director of Sales", "days": 2},
{"max_pct": 25.0, "approver": "VP Sales", "days": 3},
{"max_pct": 100.1, "approver": "CFO + CRO", "days": 7},
],
"enterprise_floor_approver": "VP Sales",
"enterprise_floor_arr": 500000,
"smb_fast_lane_arr": 15000,
},
}
@dataclass
class RoutingResult:
deal_id: str
profile: str
discount_pct: float
deal_size_arr: float
customer_tier: str
landing_approver: str
approver_chain: list[str] = field(default_factory=list)
estimated_cycle_days: int = 0
modifiers_applied: list[str] = field(default_factory=list)
notes: list[str] = field(default_factory=list)
def _bands_for(deal: dict, profile: dict) -> list[dict]:
"""Allow caller to override via deal.policy_thresholds; else use profile."""
custom = deal.get("policy_thresholds")
if custom:
# Expect list of {max_pct, approver, days} dicts; light validation
out = []
for b in custom:
out.append({
"max_pct": float(b["max_pct"]),
"approver": str(b["approver"]),
"days": int(b.get("days", 2)),
})
return sorted(out, key=lambda x: x["max_pct"])
return profile["bands"]
def route_discount(deal: dict, profile_name: str = "saas") -> RoutingResult:
if profile_name not in PROFILES:
raise ValueError(f"Unknown profile '{profile_name}'. Choose from {list(PROFILES)}.")
profile = PROFILES[profile_name]
bands = _bands_for(deal, profile)
pct = float(deal.get("discount_pct", 0.0))
arr = float(deal.get("deal_size_arr", 0.0))
tier = (deal.get("customer_tier") or "mid").lower()
# Find the landing band
landing = bands[-1]
for b in bands:
if pct <= b["max_pct"]:
landing = b
break
chain: list[str] = []
days = 0
for b in bands:
chain.append(b["approver"])
days += b["days"]
if b is landing:
break
modifiers: list[str] = []
# Enterprise floor: large ARR forces VP-level review even on small discounts
if tier == "enterprise" and arr >= profile["enterprise_floor_arr"]:
floor = profile["enterprise_floor_approver"]
if floor not in chain:
# Insert before any role above it; simplest is append + dedupe
chain.append(floor)
modifiers.append(
f"enterprise floor: ARR ,.0f >= , "
f"forces {floor} review"
)
days += 2
# SMB fast-lane: small deals can stop one hop early IF discount <= second-band cap
if (
tier == "smb"
and arr <= profile["smb_fast_lane_arr"]
and len(chain) > 2
and pct <= bands[1]["max_pct"]
):
dropped = chain.pop()
modifiers.append(
f"SMB fast-lane: ARR ,.0f <= , "
f"drops {dropped} from chain"
)
days = max(0, days - 1)
# Dedup chain while preserving order
seen: set[str] = set()
ordered = []
for a in chain:
if a not in seen:
ordered.append(a)
seen.add(a)
chain = ordered
notes = [
"This is a routing recommendation. The skill does NOT approve.",
f"Discount {pct:.1f}% landed in the '{landing['approver']}' band "
f"(<= {landing['max_pct']:.1f}%).",
]
if pct > 50.0:
notes.append("Discount > 50%: CFO/CRO MUST sign and Finance should re-run unit economics.")
return RoutingResult(
deal_id=str(deal.get("deal_id", "UNSPECIFIED")),
profile=profile_name,
discount_pct=pct,
deal_size_arr=arr,
customer_tier=tier,
landing_approver=landing["approver"],
approver_chain=chain,
estimated_cycle_days=days,
modifiers_applied=modifiers,
notes=notes,
)
def _render_human(r: RoutingResult) -> str:
lines = []
lines.append(f"Discount Routing: {r.deal_id}")
lines.append(f"Profile: {r.profile}")
lines.append(f"Discount: {r.discount_pct:.1f}% ARR: ,.0f Tier: {r.customer_tier}")
lines.append("")
lines.append("Approver chain (hops in order):")
for i, a in enumerate(r.approver_chain, start=1):
marker = " <-- discount lands here" if a == r.landing_approver else ""
lines.append(f" {i}. {a}{marker}")
lines.append("")
lines.append(f"Estimated approval cycle: {r.estimated_cycle_days} business day(s)")
if r.modifiers_applied:
lines.append("")
lines.append("Modifiers applied:")
for m in r.modifiers_applied:
lines.append(f" * {m}")
lines.append("")
for n in r.notes:
lines.append(f"note: {n}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Route a discount request to the right named approver(s).",
)
parser.add_argument("--input", help="Path to JSON request")
parser.add_argument("--profile", default="saas", choices=list(PROFILES))
parser.add_argument("--output", default="human", choices=["human", "json"])
parser.add_argument("--sample", action="store_true")
args = parser.parse_args(argv)
if args.sample or not args.input:
deal = SAMPLE_INPUT
else:
with open(args.input) as f:
deal = json.load(f)
result = route_discount(deal, args.profile)
if args.output == "json":
print(json.dumps(asdict(result), indent=2))
else:
print(_render_human(result))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/terms_redliner.py
#!/usr/bin/env python3
"""terms_redliner.py - Detect commercial-contract landmines in a deal's terms.
Stdlib-only. Takes a JSON description of the deal's terms (NOT the full contract
text — for full text scanning, see c-level-advisor/skills/general-counsel-advisor/
scripts/contract_risk_scanner.py).
Detects 10 founder/seller-killer patterns and emits a RANKED REDLINE LIST with:
- severity CRITICAL | HIGH | MEDIUM | LOW
- the standard counter-language
- the NAMED legal/commercial approver (no auto-approval; everything routes)
The skill never says the deal is fine on terms; it only outputs which clauses
need human sign-off and by whom.
Usage:
python terms_redliner.py --sample
python terms_redliner.py --input deal_terms.json
python terms_redliner.py --input deal_terms.json --output json
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, asdict, field
SAMPLE_TERMS = {
"deal_id": "ACME-2026-Q2-117",
"payment_terms_days": 75,
"auto_renew": True,
"auto_renew_notice_days": 90,
"indemnity_cap": None, # None = uncapped
"liability_cap": 1.0, # multiplier on annual fees (1x = standard)
"dpa_present": False,
"eu_data_involved": True,
"ip_assignment": "ambiguous", # "customer" | "vendor" | "ambiguous" | "perpetual_license_back"
"mfn_clause_present": True,
"exclusivity_clause_present": False,
"exclusivity_compensated": False,
"non_solicit_years": 3,
"governing_law": "Delaware",
"vendor_home_jurisdiction": "Delaware",
}
SEVERITY_RANK = {"CRITICAL": 0, "HIGH": 1, "MEDIUM": 2, "LOW": 3}
@dataclass
class Redline:
rule_id: str
severity: str
title: str
why_it_matters: str
standard_counter: str
approver: str
def _rules() -> list[dict]:
"""Each rule: id, severity, title, predicate(terms), why, counter, approver."""
return [
{
"id": "UNCAPPED_INDEMNITY",
"severity": "CRITICAL",
"title": "Uncapped indemnity exposure",
"predicate": lambda t: t.get("indemnity_cap") is None,
"why": (
"Uncapped indemnity is the single biggest founder-killer in commercial "
"contracts. A breach claim can wipe out the company."
),
"counter": (
"Cap indemnity at 12x monthly fees OR mutual cap; carve out only IP "
"infringement and gross-negligence/willful-misconduct."
),
"approver": "General Counsel + CFO",
},
{
"id": "MISSING_DPA_EU_DATA",
"severity": "CRITICAL",
"title": "EU personal data flows but no DPA",
"predicate": lambda t: t.get("eu_data_involved") and not t.get("dpa_present"),
"why": (
"GDPR Art. 28 requires a DPA when personal data of EU residents is "
"processed. Missing DPA = regulatory exposure + customer audit fail."
),
"counter": (
"Attach standard DPA (SCC 2021/914 or your template) and confirm "
"sub-processor list. Block close until DPA is countersigned."
),
"approver": "General Counsel + DPO",
},
{
"id": "MFN_PRICING",
"severity": "HIGH",
"title": "Most-Favored-Nation pricing clause present",
"predicate": lambda t: bool(t.get("mfn_clause_present")),
"why": (
"MFN binds you to refund any customer whose price drops below this one. "
"Limits future flexibility on bundles, segments, and competitive deals."
),
"counter": (
"Strike MFN entirely. If counterparty insists, narrow to 'same SKU, "
"same volume, same term, same geography' and time-bound to 12 months."
),
"approver": "VP Sales + CFO",
},
{
"id": "AUTORENEW_LONG_NOTICE",
"severity": "HIGH",
"title": "Auto-renew with notice window > 30 days",
"predicate": lambda t: (
t.get("auto_renew") and int(t.get("auto_renew_notice_days") or 0) > 30
),
"why": (
"Long notice windows on auto-renew are a classic trap: easy to miss, "
"and locks you into another full term. Especially painful on multi-year."
),
"counter": (
"Reduce notice to 30 days OR require affirmative re-signature each term."
),
"approver": "Deal Desk + General Counsel",
},
{
"id": "PERPETUAL_LICENSE_BACK",
"severity": "CRITICAL",
"title": "Perpetual license-back of IP to customer",
"predicate": lambda t: t.get("ip_assignment") == "perpetual_license_back",
"why": (
"Perpetual license-back gives the customer rights to use your IP "
"forever, often royalty-free, surviving termination. Kills moat."
),
"counter": (
"Convert to time-bounded license tied to subscription term, "
"field-of-use restricted, no transferability."
),
"approver": "General Counsel + CEO",
},
{
"id": "AMBIGUOUS_IP",
"severity": "HIGH",
"title": "IP ownership ambiguous",
"predicate": lambda t: t.get("ip_assignment") == "ambiguous",
"why": (
"Ambiguous IP becomes a dispute at acquisition diligence. Costs "
"weeks of legal review and can break a deal."
),
"counter": (
"Clarify: vendor retains all pre-existing IP and IP developed in "
"delivery; customer owns its data and outputs derived solely from it."
),
"approver": "General Counsel",
},
{
"id": "EXCLUSIVITY_UNCOMPENSATED",
"severity": "CRITICAL",
"title": "Exclusivity clause without compensation",
"predicate": lambda t: (
t.get("exclusivity_clause_present") and not t.get("exclusivity_compensated")
),
"why": (
"Free exclusivity removes addressable market for no economic benefit. "
"Even paid exclusivity needs a kill switch on missed quarterly minimums."
),
"counter": (
"Either strike exclusivity OR price it (minimum guaranteed spend) AND "
"add an exit ramp if MGS isn't hit two consecutive quarters."
),
"approver": "CRO + General Counsel",
},
{
"id": "LONG_PAYMENT_TERMS",
"severity": "HIGH",
"title": "Payment terms longer than NET-45",
"predicate": lambda t: int(t.get("payment_terms_days") or 0) > 45,
"why": (
"NET-60/75/90 inflates DSO, ties up working capital, and is a classic "
"buyer ploy. Material on any deal > 10% of cash balance."
),
"counter": (
"Counter to NET-30; offer 1-2% discount for NET-15 prepay if customer "
"won't move. Add late-payment interest of 1.5% / mo on any overdue."
),
"approver": "CFO + Deal Desk",
},
{
"id": "LOW_LIABILITY_CAP",
"severity": "MEDIUM",
"title": "Liability cap below 1x annual fees",
"predicate": lambda t: float(t.get("liability_cap") or 0.0) < 1.0,
"why": (
"Customer pushing for sub-1x cap usually indicates they expect "
"outsized claims. Don't accept without symmetric protection."
),
"counter": (
"Hold liability cap at 1x annual fees (12-month look-back), mutual; "
"super-cap (3x) on IP and confidentiality breaches if needed."
),
"approver": "General Counsel",
},
{
"id": "BROAD_NON_SOLICIT",
"severity": "MEDIUM",
"title": "Non-solicit longer than 12 months",
"predicate": lambda t: int(t.get("non_solicit_years") or 0) >= 2,
"why": (
"Multi-year non-solicit limits hiring and is increasingly unenforceable "
"in many US jurisdictions (e.g. California). Negotiate down."
),
"counter": (
"Cap non-solicit at 12 months post-termination, scoped to employees "
"directly engaged on the project, with exception for general advertising."
),
"approver": "General Counsel + CHRO",
},
]
def scan_terms(terms: dict) -> list[Redline]:
findings: list[Redline] = []
for rule in _rules():
try:
if rule["predicate"](terms):
findings.append(
Redline(
rule_id=rule["id"],
severity=rule["severity"],
title=rule["title"],
why_it_matters=rule["why"],
standard_counter=rule["counter"],
approver=rule["approver"],
)
)
except (KeyError, TypeError, ValueError):
# Missing or malformed field for this rule -> skip silently
continue
findings.sort(key=lambda r: (SEVERITY_RANK[r.severity], r.rule_id))
return findings
def _render_human(deal_id: str, findings: list[Redline]) -> str:
lines = []
lines.append(f"Terms Redline Report: {deal_id}")
lines.append(f"{len(findings)} landmine(s) detected.")
lines.append("")
if not findings:
lines.append("No flagged terms. STILL route to General Counsel for sign-off — ")
lines.append("this scanner only catches the 10 most common patterns.")
return "\n".join(lines)
for i, f in enumerate(findings, start=1):
lines.append(f"{i}. [{f.severity}] {f.title}")
lines.append(f" why: {f.why_it_matters}")
lines.append(f" counter: {f.standard_counter}")
lines.append(f" approver: {f.approver}")
lines.append("")
lines.append("note: This is a triage tool, not legal advice. All HIGH/CRITICAL")
lines.append(" findings must be reviewed by named approver before signing.")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description="Scan deal terms JSON for commercial-contract landmines.",
)
parser.add_argument("--input", help="Path to JSON terms")
parser.add_argument("--output", default="human", choices=["human", "json"])
parser.add_argument("--sample", action="store_true")
args = parser.parse_args(argv)
if args.sample or not args.input:
terms = SAMPLE_TERMS
else:
with open(args.input) as f:
terms = json.load(f)
findings = scan_terms(terms)
deal_id = str(terms.get("deal_id", "UNSPECIFIED"))
if args.output == "json":
print(json.dumps({
"deal_id": deal_id,
"finding_count": len(findings),
"findings": [asdict(f) for f in findings],
}, indent=2))
else:
print(_render_human(deal_id, findings))
return 0
if __name__ == "__main__":
sys.exit(main())
Soạn thảo, kiểm tra và làm sạch SOP, runbook nội bộ như mua sắm, offboarding nhà cung cấp, onboarding nhân viên, hoàn chi phí và cấp quyền hệ thống.
---
name: knowledge-ops
description: Use when a Head of Ops, Knowledge Manager, or TPM-Internal needs to author, validate, or clean up company SOPs and internal runbooks (procurement intake, vendor offboarding, incident-comms cascade, employee onboarding, expense reimbursement, system-access provisioning, customer-escalation playbook) — including 5W2H completeness checks (Who-What-When-Where-Why-How-HowMuch), cross-link and orphan-page validation across a sprawling Notion/Confluence/Obsidian wiki, KB ingestion + hygiene reporting, ops onboarding doc generation, and runbook step verification (named owner, expected duration, observable success signal, rollback path, escalation contact). Pairs Kaoru Ishikawa's 5W2H method, Atul Gawande's *The Checklist Manifesto*, ISO 9001, ITIL v4 Service Operation, FDA 21 CFR Part 211, and Google SRE Workbook runbook discipline with deterministic stdlib-only Python tools that score completeness, detect anti-patterns, and emit prioritized cleanup lists. Distinct from `engineering/llm-wiki` (Karpathy-style personal PKM second brain), `engineering-team/runbook-generator` (system-ops production debugging runbook), `project-management/*` (Jira/Confluence delivery + ticket tracking), and sibling `business-operations/process-mapper` (BPMN process *design*, while knowledge-ops is process *documentation*).
context: fork
version: 2.8.0
author: claude-code-skills
license: MIT
tags: [bizops, sop, runbook, knowledge-management, kb, 5w2h, wiki, ops-documentation]
compatible_tools: [claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli]
---
# knowledge-ops
Company SOP + internal runbook authoring, 5W2H completeness validation, and KB hygiene reporting for Head-of-Ops / Knowledge-Manager / TPM-Internal personas.
## Purpose
An ops organization three years in accumulates a sprawl: 600 Notion pages, 200 Confluence runbooks, three Obsidian vaults, a `Drive/SOPs/` folder, and a `Slack #ops-questions` channel that exists because nobody can find the canonical doc. Predictable failure modes:
1. **No owner** — 40% of SOPs name "the team" instead of a person. When the doc rots, nobody is accountable.
2. **No last-reviewed date** — a 2023 vendor-offboarding SOP still references a procurement tool sunset in 2024.
3. **Vague success signals** — runbook step 4 says "verify the service is up". A new operator can't tell what that means.
4. **No rollback path** — incident-comms cascade runbook tells you how to send the alert. It doesn't tell you how to retract it when the alert was wrong.
5. **Orphan pages** — half the KB has no inbound links. Nobody finds them via navigation; they only exist because somebody knew the URL.
6. **Glossary drift** — "CSM" means Customer Success Manager in three docs and Customer Solutions Manager in five. New hires guess wrong for six months.
7. **Happy-path-only SOPs** — the doc covers what happens when everything works. It doesn't cover the 30% case where it doesn't.
This skill answers the operator's actual question: **"Which 20 docs do I fix first, and what specifically is wrong with each?"** — with deterministic logic, not intuition.
## When to use
- Authoring a new SOP for a cross-functional company process (procurement intake, vendor offboarding, incident-comms cascade, employee onboarding, expense reimbursement, customer-escalation playbook, security-incident comms, system-access provisioning).
- Validating an existing internal runbook before it goes into rotation (every step must have a named owner, expected duration, observable success signal, observable failure signal, rollback path, escalation contact).
- Ingesting a multi-document KB export (Notion zip, Confluence space export, Obsidian vault, `Drive/SOPs/` directory) and surfacing what's broken: orphan pages, stale pages (no edit > 12 months), glossary drift, missing-owner pages, cross-link map.
- Onboarding a new ops hire by generating the SOPs and ops-handbook pages they need to read in week 1.
- Wiki cleanup sprints — quarterly hygiene work where the org decides which 30 docs to archive, rewrite, or merge.
## Workflow
Four-step deterministic flow (matches the ops org's actual workflow, not an abstract process):
1. **Ingest KB.** Run `kb_ingester.py --input <vault-dir>` on the existing wiki export. Output is a markdown health report: orphan pages, stale pages, glossary drift, missing-owner pages, cross-link map, prioritized cleanup list. The report ranks the top-20 docs to fix first — usually a mix of high-traffic stale docs and compliance-relevant missing-owner docs. Take this list to the cleanup sprint.
2. **Validate existing runbooks.** For each runbook in the cleanup list (or any new runbook before it goes into rotation), run `runbook_validator.py --input <runbook.md>`. The validator scores each step against six checks (named owner, expected duration, observable success signal, observable failure signal, rollback path, escalation contact) and produces a per-step traffic-light + overall validity score 0-100 + MUST-FIX issue list. A runbook scoring < 60 is not safe to use in an incident.
3. **Generate missing SOPs.** For SOPs that need to be written from scratch (or rewritten because the existing one is unsalvageable), run `sop_generator.py --input <metadata.json> --profile <ops|support|finance|hr|it|regulated>`. Output is a 5W2H-structured SOP scaffold: Who (RACI), What (process steps), When (triggers + frequency), Where (system + tool), Why (purpose + regulatory basis), How (step-by-step), How-much (cost + time per execution). The `regulated` profile adds version control, signoff, and audit-trail sections (ISO 9001 / FDA 21 CFR Part 211 / SOC 2 / HIPAA).
4. **Cross-link + close the loop.** Re-run `kb_ingester.py` after the cleanup sprint to verify orphan-page count is down and glossary drift is resolved. The metric that matters is **"unfindable docs"** (orphans) and **"unsafe runbooks"** (validity score < 60) — not page count.
## Scripts
**`scripts/sop_generator.py`** — Reads a JSON metadata file describing an SOP (process owner, triggering event, audience role, frequency, regulatory overlay, inputs, outputs, steps outline) and emits a full 5W2H-structured SOP in markdown (or normalized JSON). The `--profile` flag tunes the output: `ops` (general internal ops), `support` (customer-support runbook style), `finance` (controls + reconciliation focus), `hr` (sensitive-data flagging), `it` (system + access focus), `regulated` (adds version control, signoff matrix, audit-trail). Regulatory overlays (`SOC2`, `HIPAA`, `ISO13485`, `GDPR`, `SOX`) attach the appropriate compliance preamble. `--sample` prints a complete vendor-offboarding SOP example. Stdlib only.
**`scripts/runbook_validator.py`** — Reads a runbook (markdown file or JSON) and validates each step against six required attributes: (1) named owner (not "the team", not "ops"), (2) expected duration (concrete number + unit), (3) observable success signal (e.g., "HTTP 200 from `/healthz`" — not "service is up"), (4) observable failure signal, (5) rollback path (or explicit "this step cannot be rolled back, escalate to X"), (6) escalation contact (named person or named on-call rotation). Output is a per-step traffic-light (GREEN/AMBER/RED), an overall validity score 0-100, and a MUST-FIX issue list. Verdict: ≥ 80 = SAFE-TO-USE, 60-79 = USE-WITH-CAUTION, < 60 = NOT-SAFE. `--sample` prints a deliberately-broken incident-comms runbook to demonstrate failure detection. Stdlib only.
**`scripts/kb_ingester.py`** — Walks a directory of markdown files (Notion export, Confluence space export, Obsidian vault, `Drive/SOPs/` directory). Extracts: (a) cross-link map (which page references which, via markdown `[link](path)` syntax), (b) glossary candidates (frequently used proper nouns and acronyms that recur in 3+ docs without a single canonical definition page), (c) orphan pages (no inbound links from anywhere in the vault), (d) glossary drift (the same term defined or used inconsistently across docs — e.g., "CSM" expanded differently in two places), (e) stale pages (no edit in > 12 months, detected via filesystem mtime or YAML `last_reviewed` frontmatter), (f) missing-owner pages (no `owner:` field in frontmatter). Emits a KB health report markdown with a prioritized top-20 cleanup list ranked by `staleness × inbound-link-count` (high-traffic stale docs first). `--sample` builds a tiny synthetic 8-page vault in a tmpdir and runs the full pipeline against it. Stdlib only.
## References
- `references/5w2h_sop_canon.md` — Kaoru Ishikawa's 5W2H method, Toyota standard-work discipline, Atul Gawande's checklist manifesto, Atlassian Confluence SOP guidance, ISO 9001 SOP requirements, ITIL v4 Service Operation, FDA 21 CFR Part 211. Eight cited sources covering SOP authoring canon.
- `references/runbook_canon.md` — Google SRE Workbook (runbook chapter), Atlassian incident-management runbooks, PagerDuty Incident Response taxonomy, AWS Well-Architected operational excellence pillar, Charity Majors on observability-runbook integration, Susan Fowler on production-ready microservices, ITIL v4 Operations. Seven cited sources covering runbook design canon.
- `references/kb_hygiene_anti_patterns.md` — Eight anti-patterns drawn from Notion/Confluence wiki industry research, Mozilla SUMO knowledge-base lessons, Stack Overflow community-management research, the Atlassian Team Playbook, MIT TIK org-wiki studies, Cynthia Lee on glossary drift, and Adam Wiggins on "documentation rot".
## Assumptions
1. The KB is in markdown (or can be exported to markdown — Notion, Confluence, Obsidian, and Google Docs all support this). HTML-only or PDF-only KBs require a conversion pass first; out of scope.
2. The user has authority to commission rewrites or archives. Producing a cleanup list nobody acts on is wasted work — route findings to a named owner before running the ingester.
3. Owner metadata lives in YAML frontmatter (`owner: alex@company.com`) or in a top-of-page "Owner:" line. Tribal-knowledge ownership (the person who last edited the page) is treated as missing.
4. "Stale" defaults to 12 months. Override with `--stale-days` on `kb_ingester.py`. Some compliance regimes (FDA, ISO 13485) require shorter review cycles; use `--profile regulated` and `--stale-days 365`.
5. The user is not asking for a personal PKM. Personal Karpathy-style second-brain work belongs in `engineering/llm-wiki`.
## Anti-patterns
- **Generating SOPs in bulk without owners.** A doc with no owner has a half-life of 6 months. Refuse to generate a batch of 30 SOPs unless each one is assigned to a named human.
- **Using `runbook_validator.py` as a checkbox.** The validator catches missing structure. It does not catch wrong content. A runbook can score 100 and still tell the operator the wrong thing.
- **Treating orphan pages as garbage by default.** Some orphans are reference pages found only via search — not all orphans should be archived. The cleanup list is a *priority queue*, not a delete list.
- **Confusing knowledge-ops with `process-mapper`.** Process-mapper documents the *flow* of work between stages (BPMN, cycle time, bottleneck). Knowledge-ops documents the *artifacts* operators consume to execute the work (SOP, runbook, glossary). Both can apply to the same process.
- **Letting glossary drift accumulate.** Two definitions of "CSM" in three years becomes seven definitions in five. Fix glossary drift the moment it surfaces in `kb_ingester.py` output.
- **Skipping the regulated profile under regulated workload.** If the process touches PHI, SOX-relevant financial controls, or ISO 13485 device QMS, use `--profile regulated`. Missing version control on a regulated SOP is an audit finding.
- **Hand-writing 5W2H sections from memory.** The 5W2H scaffold exists because operators forget "How-much". Use the generator; edit the output.
## Distinct from
- **`engineering/llm-wiki`** — Karpathy-style personal PKM second brain where one human ingests sources into their own interlinked vault. Knowledge-ops is *organizational*: many authors, many readers, named owners per doc, formal review cycles, compliance overlays.
- **`engineering-team/runbook-generator`** — system-ops runbook for debugging a production system (logs, alerts, k8s, on-call). Knowledge-ops runbooks are *operator* runbooks for business processes (incident-comms cascade, vendor offboarding, employee onboarding). The audience is fellow operators, not engineers tailing logs.
- **`project-management/*`** — Jira / Confluence delivery tracking, sprint ticket workflow, project-status reporting. Knowledge-ops is the *content* in those Confluence pages, not the *tracking* of who edits them.
- **`business-operations/process-mapper`** (sibling) — BPMN process *design*: where the stages are, where work waits, which stage is the bottleneck. Knowledge-ops is process *documentation*: the SOP and runbook artifacts that tell an operator how to execute the process the mapper described.
- **`business-operations/internal-comms`** (sibling) — broadcast announcements, all-hands messaging, change-management comms. Knowledge-ops is the durable reference artifact; internal-comms is the broadcast.
- **`ra-qm-team/*`** — formal regulatory compliance authoring (ISO 13485 QMS, MDR technical files, 21 CFR Part 820). Knowledge-ops borrows the regulatory checklist but is not a substitute for a notified-body audit.
## Forcing-question library (Matt Pocock grill discipline)
Before invoking the tools, the orchestrator (or `/cs:grill-bizops`) walks the user through these questions **one at a time, with a recommended answer + canon citation**. Never bundled. Walk depth-first — do not open question 4 until 1-3 are locked.
1. **"Who is the named owner of this SOP / runbook, and do they know they own it?"**
Recommended: a single human (not "the team"), and yes — they have agreed in writing.
Canon: Gawande 2009 (*The Checklist Manifesto*) — checklists without an owner rot within 12 months. Ownership is the discipline.
2. **"When was this doc last reviewed, and what is the review cadence?"**
Recommended: reviewed within the last 12 months (90 days if `--profile regulated`); cadence written in the frontmatter.
Canon: ISO 9001:2015 §7.5.3 — controlled documents require review-cycle metadata. ITIL v4 echoes this for Service Operation runbooks.
3. **"For each runbook step: what is the observable success signal — by which I mean, what specific output tells you the step worked?"**
Recommended: a concrete observable ("HTTP 200 from `/healthz`", "Slack thread closed with `done` reaction", "Salesforce opportunity moved to `Closed-Won` stage") — not "the service is up" or "it works".
Canon: Beyer et al. 2018 (*Site Reliability Workbook*, Ch. 8) — observable signals are the entire point of a runbook. Vague success criteria are the leading cause of runbook misuse during incidents.
4. **"What is the rollback path for each runbook step that can fail?"**
Recommended: every step that mutates state has either a rollback path or an explicit "cannot roll back — escalate to X" line.
Canon: AWS Well-Architected Framework, Operational Excellence pillar — "you cannot run a process you cannot reverse without first agreeing what 'reverse' means".
5. **"Where does this doc live, and what other docs link to it?"**
Recommended: in the canonical wiki, and at least 2 inbound links from related docs. An orphan SOP is an unfindable SOP.
Canon: Atlassian Team Playbook on documentation health — orphan rate > 20% is the leading indicator of a wiki sprawl problem.
6. **"What is the regulatory overlay on this process — SOC 2, HIPAA, ISO 13485, GDPR, SOX, none?"**
Recommended: explicit answer. If "none", confirm by checking the data classes the process touches.
Canon: FDA 21 CFR Part 211.100 (Written procedures; deviations) — regulated SOPs require version control, change history, and signoff. Skip this step and the doc is an audit finding.
7. **"Is the happy path the *only* path documented, or are the 2-3 most common failure modes also documented?"**
Recommended: the top-2 failure modes per process are documented with their own recovery sub-procedure.
Canon: Fowler 2016 (*Production-Ready Microservices*) — operations docs that cover only the happy path are responsible for 60%+ of incident-time waste.
After all 7 are locked, invoke `kb_ingester.py` → `runbook_validator.py` → `sop_generator.py` in sequence.
FILE:assets/runbook_template.md
# Runbook Template — fill out before running `runbook_validator.py`
Use this template to capture runbook steps before invoking the validator.
Each step must specify all six required attributes (owner, duration,
success signal, failure signal, rollback, escalation) or the validator
will flag it.
Feed the JSON into:
```
python3 scripts/runbook_validator.py --input my-runbook.json
python3 scripts/runbook_validator.py --input my-runbook.md # markdown also accepted
```
A runbook scoring < 60 is NOT-SAFE for production use. Aim for ≥ 80
(SAFE-TO-USE) before putting the runbook into rotation.
---
## Runbook metadata
- **Runbook name:** _(e.g., Incident Comms Cascade, Customer Escalation, Vendor Outage Response, System-Access Revocation)_
- **Owner:** _(named human or named on-call rotation — e.g., "Incident Commander on-call (PagerDuty: ic-primary)")_
- **Trigger:** _(what specifically invokes this runbook — e.g., "PagerDuty Sev-1 incident triggered" or "Customer escalation flagged in Salesforce")_
- **Expected total duration:** _(P50 + P90 wall-clock from trigger to completion)_
- **Linked SOP:** _(if this runbook implements an SOP, link the canonical SOP page)_
---
## Step table
| # | Step title | Owner | Duration | Success signal (observable) | Failure signal (observable) | Rollback | Escalation |
|---|------------|-------|----------|------------------------------|------------------------------|----------|------------|
| 1 | _e.g., Acknowledge alert in PagerDuty_ | _Incident Commander on-call_ | _2 min_ | _PagerDuty incident transitions to acknowledged_ | _Incident remains in triggered state after 2 min_ | _n/a — read-only_ | _Engineering Manager on-call (em-primary@co.com)_ |
| 2 | _e.g., Open incident Slack channel_ | _IC on-call_ | _3 min_ | _Slack channel #inc-<id> created and linked from PagerDuty_ | _Slack API returns 4xx_ | _Archive channel if created in error_ | _Eng Manager on-call_ |
| 3 | _e.g., Notify execs via paging tree_ | _Comms Lead (comms-lead@co.com)_ | _5 min_ | _SES API returns 200 for all exec recipients_ | _SES API returns 5xx OR delivery=bounced_ | _Send retraction email with subject prefix 'RETRACTION:'_ | _VP Communications_ |
---
## JSON skeleton
```json
{
"runbook_name": "Incident Comms Cascade",
"steps": [
{
"title": "Acknowledge alert in PagerDuty",
"owner": "Incident Commander on-call (PagerDuty: ic-primary)",
"duration_str": "2 minutes",
"duration_minutes": 2,
"success_signal": "PagerDuty incident transitions to acknowledged",
"failure_signal": "Incident remains in triggered state after 2 minutes",
"rollback": "n/a — acknowledgement is non-mutating, read-only operation",
"escalation": "Engineering Manager on-call (em-primary@company.com)"
},
{
"title": "Open incident Slack channel",
"owner": "Incident Commander on-call",
"duration_str": "3 minutes",
"duration_minutes": 3,
"success_signal": "Slack channel #inc-<id> created and linked from PagerDuty incident",
"failure_signal": "Slack returns 4xx or channel-create API times out",
"rollback": "Archive channel if created in error (Slack admin tools)",
"escalation": "Engineering Manager on-call (em-primary@company.com)"
},
{
"title": "Notify execs via paging tree",
"owner": "Communications Lead (comms-lead@company.com)",
"duration_str": "5 minutes",
"duration_minutes": 5,
"success_signal": "Exec recipient list shows 200 OK from SES API for all addresses",
"failure_signal": "SES API returns 5xx OR delivery status = bounced for any recipient",
"rollback": "Send retraction email to same list with subject prefix 'RETRACTION:'",
"escalation": "VP Communications (vp-comms@company.com)"
}
]
}
```
---
## Markdown form (alternative — runbook_validator.py heuristic parser)
If you prefer authoring in markdown directly, follow this exact structure (the parser keys off `## Step N:` headings and bullet attributes):
```markdown
# Runbook: Incident Comms Cascade
## Step 1: Acknowledge alert in PagerDuty
- **Owner:** Incident Commander on-call (PagerDuty: ic-primary)
- **Duration:** 2 minutes
- **Success:** PagerDuty incident transitions to acknowledged
- **Failure:** Incident remains in triggered state after 2 minutes
- **Rollback:** n/a — non-mutating, read-only
- **Escalation:** Engineering Manager on-call (em-primary@company.com)
## Step 2: Open incident Slack channel
- **Owner:** ...
```
---
## Authoring discipline checklist
Before submitting the runbook to the validator:
- [ ] **Every step has a named owner**, not "the team" or "ops" — required by SRE Workbook Ch. 8.
- [ ] **Every step has a concrete duration** (number + unit). "Quick" is not a duration.
- [ ] **Every success signal is observable** — a yes/no check the operator can perform. "HTTP 200 from /healthz", not "service is up".
- [ ] **Every failure signal is observable** — what tells you the step did NOT work.
- [ ] **Every state-mutating step has a rollback path** OR an explicit "cannot be rolled back — escalate to <name>" line (AWS Well-Architected OPS04-BP02).
- [ ] **Every step has an escalation contact** — named human, role+email, or named on-call rotation.
- [ ] **Top-2 failure modes documented** (Fowler 2016) — most common ways this runbook gets stuck, each with their own recovery sub-procedure.
- [ ] **Last-reviewed date set in frontmatter** — runbooks decay; Charity Majors's data: untouched 12-month-old runbooks are wrong 60% of the time.
After validation, place the runbook in the canonical wiki location and link it from at least 2 navigation hubs (incident-handbook, the parent SOP) to avoid orphan-page status.
FILE:assets/sop_template.md
# SOP Template — fill out before running `sop_generator.py`
Use this template to capture the SOP metadata before invoking the generator.
Fill in the fields below, then translate them into the JSON skeleton at the
bottom of this file. Feed that JSON into the generator:
```
python3 scripts/sop_generator.py --input my-sop.json --profile ops
python3 scripts/sop_generator.py --input my-sop.json --profile regulated # for SOX / HIPAA / ISO 13485 / FDA
```
---
## SOP metadata
- **SOP name:** _(e.g., Vendor Offboarding, Procurement Intake, Employee Onboarding, Customer Escalation, System Access Provisioning)_
- **Process owner (named human):** _(e.g., alex@company.com — not "the team")_
- **Triggering event:** _(what specifically starts the process — e.g., "Vendor contract not renewed OR vendor terminated for cause")_
- **Audience role:** _(who will execute this SOP — e.g., "Vendor Management Office operator", "HR onboarding specialist")_
- **Frequency:** _(how often this runs — "Daily", "Weekly Monday 9am", "On-demand avg 3x/quarter")_
- **Regulatory overlay:** _(zero or more of: SOC2, HIPAA, ISO13485, GDPR, SOX. If "none", confirm by listing data classes the process touches.)_
---
## Inputs and outputs
**Inputs required before starting:**
- _(input 1 — e.g., "Vendor legal name")_
- _(input 2 — e.g., "Contract end date")_
- _(input 3 — e.g., "List of systems with vendor access")_
**Outputs produced:**
- _(output 1 — e.g., "All production system access revoked, evidenced in IAM audit log")_
- _(output 2 — e.g., "Vendor data deletion certified")_
- _(output 3 — e.g., "Final invoice reconciled and paid")_
---
## Steps outline
Six rows to start; add or remove. **Each step must be a noun-phrase action**, not a paragraph.
| # | Step name (action) | Notes |
|---|--------------------|-------|
| 1 | _e.g., Notify vendor of offboarding intent (30 days written notice)_ | |
| 2 | _e.g., Inventory data classes and system access vendor holds_ | |
| 3 | _e.g., Revoke production system access (IAM, VPN, SaaS)_ | |
| 4 | _e.g., Confirm data deletion (vendor certification) or data return_ | |
| 5 | _e.g., Final invoice reconciliation and payment_ | |
| 6 | _e.g., Archive vendor record in VMO registry with offboarding evidence_ | |
---
## How-much (cost model)
- **Estimated execution time:** _(minutes per execution — e.g., 240)_
- **Estimated cost per execution:** _(USD, labor + license + third-party fees — e.g., 800)_
---
## JSON skeleton
```json
{
"sop_name": "Vendor Offboarding",
"process_owner": "alex@company.com (Vendor Management Lead)",
"triggering_event": "Vendor contract not renewed OR vendor terminated for cause",
"audience_role": "Vendor Management Office (VMO) operator",
"frequency": "On-demand (avg 3 executions per quarter)",
"regulatory_overlay": ["SOC2"],
"inputs": [
"Vendor legal name",
"Contract end date",
"List of systems with vendor access",
"List of data classes vendor processed"
],
"outputs": [
"All production system access revoked (evidenced)",
"Vendor data deleted or returned (evidenced)",
"Final invoice reconciled and paid",
"Vendor record archived in VMO registry with offboarding evidence"
],
"steps_outline": [
"Notify vendor of offboarding intent (written, 30 days notice)",
"Inventory data classes and system access vendor holds",
"Revoke production system access (IAM, VPN, SaaS)",
"Confirm data deletion (vendor certification) or data return",
"Final invoice reconciliation and payment",
"Archive vendor record in VMO registry with offboarding evidence"
],
"estimated_minutes": 240,
"estimated_cost_usd": 800
}
```
---
## Authoring discipline checklist
Before submitting the JSON to the generator, confirm:
- [ ] **Owner is a named human**, not "the team" — required by Gawande *Checklist Manifesto* discipline.
- [ ] **Triggering event is specific.** "When needed" is not a trigger.
- [ ] **At least one regulatory overlay considered** (or explicit "none after checking PHI/financial/regulated-device classes").
- [ ] **Top-2 failure modes documented** — happy-path-only SOPs are responsible for 60%+ of incident-time waste (Fowler 2016).
- [ ] **"How-much" is filled in.** It's the section authors most often forget — and the section operators most need.
- [ ] **`--profile regulated` selected** if SOP touches SOX, HIPAA, ISO 13485, FDA 21 CFR Part 211, or SOC 2 controls.
After generation, run the runbook validator on any embedded step lists that include state-mutating operations:
```
python3 scripts/runbook_validator.py --input generated-sop.md
```
FILE:references/5w2h_sop_canon.md
# 5W2H SOP Canon
Standard Operating Procedure (SOP) authoring discipline for company processes — what every SOP must contain, why, and where the discipline comes from. Eight authoritative sources cited.
## What 5W2H is
5W2H is a structured checklist for documenting *any* repeatable process by answering seven questions:
| Letter | Question | Section in `sop_generator.py` output |
|---|---|---|
| Who | Who is responsible, accountable, consulted, informed? | RACI |
| What | What is the process — inputs, outputs, scope? | Process spec |
| When | When does it run — trigger, frequency, blocking deps? | Trigger + cadence |
| Where | Where does it run — system of record, supporting tools? | System map |
| Why | Why does it exist — business purpose, regulatory basis? | Purpose + compliance |
| How | How is it executed — step-by-step procedure? | Procedure |
| How-much | How much does it cost — time, money per execution? | Cost model |
Two SOPs covering the same process can be wildly different in length and quality. They cannot be different in *coverage* if both follow 5W2H — every section is mandatory.
## Why 5W2H specifically
Three properties make 5W2H the right scaffold for an ops org:
1. **Audit-friendly.** ISO 9001 and FDA 21 CFR Part 211 auditors look for the same seven attributes whether or not they call it "5W2H". Adopting the scaffold up front means SOPs ship audit-ready.
2. **Operator-friendly.** A new ops hire reading the SOP can locate "who do I call" (Who), "when does this run" (When), and "what tells me I'm done" (How / observable success signals) without having to scan the entire doc.
3. **Author-friendly.** Empty 5W2H sections are visually obvious. "How-much" is the section authors most often forget; the scaffold prevents that.
## Eight authoritative sources
### 1. Kaoru Ishikawa — *Guide to Quality Control* (1985, Asian Productivity Organization)
Origin of the 5W1H quality-control method. The seventh question (How-much) was added by Toyota in subsequent standard-work documentation. Ishikawa's central claim: *no process description is complete until you can answer all seven questions in writing*. Anything less is tribal knowledge.
### 2. Jeffrey Liker — *The Toyota Way* (2003, McGraw-Hill)
Chapter 6 on standard work codifies the Toyota convention that every SOP documents (a) takt time, (b) work sequence, (c) standard inventory. The "How-much" anchor maps directly to takt time. Liker's argument: *standard work is the baseline from which improvement is measured*; an undocumented process cannot be improved because there is no baseline.
### 3. Atul Gawande — *The Checklist Manifesto* (2009, Metropolitan Books)
Gawande's hospital surgical-checklist research found that simple, well-owned checklists reduced surgical mortality by 47% in a 2008 WHO study across eight hospitals on four continents. Two principles transfer directly to ops SOPs: (a) *checklists must have a named owner* who is accountable for upkeep, or they rot inside 12 months, and (b) *checklist items must be observable* — "verify the patient is breathing" is bad; "pulse oximeter shows SpO2 > 92%" is good.
### 4. Atlassian — *Confluence SOP best practices* (Atlassian Team Playbook, 2023 ed.)
Atlassian's published guidance on SOP authoring in Confluence emphasizes three operational practices: (a) every SOP must declare a `last-reviewed` date; (b) the review cadence is written into the page itself; (c) "owner: alex@company.com" goes in YAML frontmatter so tooling can find SOPs with no owner. The KB hygiene anti-patterns reference draws from the same source.
### 5. ISO 9001:2015 — *Quality management systems — Requirements*
Clause 7.5.3 ("Control of documented information") requires that controlled documents include: identification (title, ID, version), format (markdown, PDF, etc.), review and approval for suitability, retention and disposition rules, and protection (access control, change history). The `regulated` profile in `sop_generator.py` adds these sections explicitly.
### 6. ITIL v4 — *Service Operation* practice guide (Axelos, 2019)
ITIL's distinction between *procedures* (the SOP — repeatable and largely unchanged) and *work instructions* (the runbook — the specific commands and observable signals at execution time) is the same distinction this skill makes. Both artifacts coexist. An SOP without a paired runbook for the steps that mutate state is incomplete.
### 7. FDA 21 CFR Part 211.100 — *Written procedures; deviations*
For pharmaceutical and medical-device-adjacent companies, Part 211.100 makes SOPs legally required. Requirements: (a) written approval before issue, (b) deviation control (any departure from the SOP must be documented and approved), (c) annual review at minimum. The `--profile regulated` flag attaches these requirements.
### 8. Project Management Institute — *PMBOK Guide* (7th ed., 2021)
PMBOK §4 on integration management defines SOP-equivalent artifacts as "organizational process assets" and requires named accountability. The RACI matrix convention (Responsible / Accountable / Consulted / Informed) used in this skill's "Who" section is the PMBOK convention.
## Anti-pattern: prose-only SOPs
A 1500-word prose SOP without the 5W2H scaffolding looks thorough and is usually missing 2-3 mandatory sections (most commonly: How-much, Why-regulatory, observable success signals). Use the generator. Edit its output. Do not write SOPs from a blank page.
## How this skill applies the canon
- `sop_generator.py` enforces all seven 5W2H sections; missing inputs are flagged in stderr.
- `--profile regulated` attaches ISO 9001 §7.5.3 + FDA Part 211 metadata (version, signoff, change history).
- Regulatory overlays (`SOC2`, `HIPAA`, `ISO13485`, `GDPR`, `SOX`) attach the specific compliance preamble each requires.
- The forcing-question library in `SKILL.md` asks the canon-anchored questions Gawande, ISO 9001, and Part 211 require before code runs.
FILE:references/kb_hygiene_anti_patterns.md
# Knowledge-Base Hygiene Anti-Patterns
The recurring failure modes that turn a useful company wiki into a sprawl of stale, unfindable, contradictory docs. Eight anti-patterns, each anchored to authoritative sources. Seven citations.
## The pattern
An ops org's wiki passes through three predictable phases:
1. **Year 1:** 50 pages, all owned, all current, everyone finds what they need.
2. **Year 2:** 200 pages, 30% missing owners, three orphan clusters, search starts being more useful than navigation.
3. **Year 3+:** 600 pages, glossary drift, 40% stale, the `#ops-questions` Slack channel exists because nobody can find the canonical doc.
`kb_ingester.py` exists to put numbers on this decay and rank what to fix first. The anti-patterns below explain *what to fix*.
## 1. No owner per SOP
**Symptom:** YAML frontmatter has no `owner:` field, or the SOP body says "owned by the Ops team".
**Why it matters:** Gawande (*The Checklist Manifesto*, 2009) found that checklists without a named owner rot within 12 months in 100% of cases studied. Ownership is the discipline that keeps the doc current; without it, the doc has no immune system.
**Detection:** `kb_ingester.py` reports `missing_owner_count`. Goal: 0.
**Fix:** Assign every SOP to a single named human in YAML frontmatter. "The team" is not an owner.
**Citation:** Gawande 2009 (*The Checklist Manifesto*, Metropolitan Books).
---
## 2. No last-reviewed date
**Symptom:** The SOP has no `last_reviewed:` field. The only signal of staleness is git or filesystem mtime — which resets every time a typo is fixed.
**Why it matters:** ISO 9001:2015 §7.5.3 explicitly requires review cycles for controlled documents. Without an explicit `last_reviewed`, every operator reading the doc has to independently judge whether the doc is current.
**Detection:** `kb_ingester.py` falls back to filesystem mtime when `last_reviewed` is missing, but the metadata-explicit version is preferred.
**Fix:** Add `last_reviewed: YYYY-MM-DD` to every SOP frontmatter. Pair with a review cadence (12 months default, 90 days for regulated).
**Citation:** ISO 9001:2015 §7.5.3 ("Control of documented information").
---
## 3. Step says "verify the service is up" (vague success signal)
**Symptom:** Runbook step success criteria are not observable. "Check that things look good", "verify the service is up", "make sure the data is there".
**Why it matters:** Beyer et al. (*Site Reliability Workbook*, 2018, Ch. 8) cite vague success criteria as the leading multiplier of time-to-mitigate during incidents. A new operator at 3am cannot tell what "up" means.
**Detection:** `runbook_validator.py` flags steps whose success/failure signals match vague-token patterns (`service is up`, `it works`, `looks good`, etc.).
**Fix:** Rewrite success signals as observable checks. "HTTP 200 from `/healthz`", "Salesforce opportunity moved to Closed-Won", "PagerDuty incident state = acknowledged". Anything that returns a yes/no.
**Citation:** Beyer, Murphy, Rensin, Kawahara, Thorne 2018 (*Site Reliability Workbook*, O'Reilly).
---
## 4. Runbook with no rollback
**Symptom:** The runbook tells the operator how to send the alert. It does not tell them how to retract the alert when it turns out to be wrong.
**Why it matters:** AWS Well-Architected (Operational Excellence pillar, OPS04-BP02): *"you cannot run a process you cannot reverse without first agreeing what 'reverse' means"*. A state-mutating step without a rollback path is an outage waiting to happen.
**Detection:** `runbook_validator.py` enforces a rollback field per step. Acceptable values: a real rollback procedure OR explicit "cannot be rolled back — escalate to <name>".
**Fix:** For every state-mutating step, write the rollback. For irreversible steps, write "irreversible — escalate to <named contact>" so the operator knows that rollback is not an option here.
**Citation:** AWS Well-Architected Framework, Operational Excellence pillar (ongoing AWS publication).
---
## 5. Wiki sprawl across 4 tools
**Symptom:** SOPs live in Notion. Runbooks live in Confluence. Onboarding lives in a Google Doc folder. The glossary lives in a Slack canvas. Nobody knows which is canonical.
**Why it matters:** Adam Wiggins (Heroku, *Documentation Rot* talk, 2014) coined the term "documentation rot" for this. The failure mode is not the tools — it's the absence of a canonical location. Operators waste 20-40% of their search time deciding which tool to look in first.
**Detection:** Out of scope for `kb_ingester.py` (which runs on one markdown tree). The signal is human: "where's the X SOP?" gets three different answers.
**Fix:** Pick one canonical tool. Migrate the rest. Treat the others as archives, link the canonical from the others. Mozilla SUMO's KB consolidation (2016) is the template.
**Citation:** Wiggins 2014 (Heroku Engineering talk, "Documentation Rot"). Cited again in MIT TIK 2020 org-wiki research.
---
## 6. Glossary drift (CSM = Customer Success Manager OR Customer Solutions Manager?)
**Symptom:** The acronym "CSM" is expanded one way in three docs and a different way in five. New hires guess wrong for six months. Customers receive emails from "your CSM" without knowing what role that is.
**Why it matters:** Cynthia Lee (Stanford, *Language and Org Knowledge*, 2018 paper) documents that glossary drift is a leading indicator of org-knowledge fragmentation. Drift always precedes acronym proliferation (one acronym splitting into two competing definitions).
**Detection:** `kb_ingester.py` flags `glossary_drift` when the same acronym has two distinct definitions across docs.
**Fix:** Pick one canonical definition per acronym. Add a `glossary.md` page. Link every other doc to it. Refuse to expand the acronym anywhere else.
**Citation:** Lee 2018 (Stanford research on org-knowledge fragmentation).
---
## 7. Orphan pages nobody can find
**Symptom:** 30-60% of pages have no inbound links. They exist because somebody knew the URL. Search finds them; navigation does not.
**Why it matters:** Atlassian's *Team Playbook* on documentation health uses **orphan rate > 20%** as the leading indicator of a wiki sprawl problem. Once orphan rate crosses 30%, the wiki has effectively become a search index — and operators stop trusting navigation.
**Detection:** `kb_ingester.py` reports `orphan_count` and lists orphans.
**Fix:** Not "delete all orphans". Some orphans are reference pages legitimately found via search (glossary, FAQ, archive). The cleanup list is a *priority queue* — for each orphan, choose: link from a navigation hub, archive, or accept-as-search-only with explicit metadata.
**Citation:** Atlassian Team Playbook, "Documentation Health" play (2021).
---
## 8. SOPs that document the happy path only
**Symptom:** The vendor-offboarding SOP covers what happens when the vendor cooperates. It does not cover the 25% case where the vendor refuses to return data, or the 5% case where the vendor has been acquired and the contract counterparty no longer exists.
**Why it matters:** Susan Fowler (*Production-Ready Microservices*, 2016, Ch. 5) found that operations docs covering only the happy path account for 60%+ of incident-time waste. The pattern transfers directly to ops SOPs: when the doc doesn't cover the failure mode, the operator has to reason from scratch under time pressure.
**Detection:** Manual — `runbook_validator.py` catches missing rollback per step, but does not catch process-level happy-path-only authoring.
**Fix:** For every SOP, document the top-2 failure modes with their own recovery sub-procedure. The forcing-question library in `SKILL.md` (question 7) enforces this.
**Citation:** Fowler 2016 (*Production-Ready Microservices*, O'Reilly).
---
## 9. Compliance SOPs without version control
**Symptom:** A SOX-relevant or HIPAA-relevant SOP has no change history, no signoff record, no version field. An auditor asks "what was the procedure in Q2?" — nobody can answer.
**Why it matters:** FDA 21 CFR Part 211.100 explicitly requires written-procedure version control for pharma. ISO 9001 §7.5.3 imposes the same for any controlled document. Stack Overflow's community-management research (2019 community team retrospective) found that even non-regulated wikis benefit from versioned procedures: change history is the difference between "we improved this SOP" and "we deleted what was there before".
**Detection:** `--profile regulated` in `sop_generator.py` attaches the version + signoff + change-history sections. Missing those sections under a regulated overlay is the audit finding.
**Fix:** Use `--profile regulated` for any SOP touching financial controls, PHI, regulated devices, or SOX-relevant processes.
**Citations:** FDA 21 CFR Part 211.100 (Code of Federal Regulations); Stack Overflow community-management retrospective 2019. Mozilla SUMO KB lessons (2016) echo both.
---
## How this skill applies the anti-patterns
- `kb_ingester.py` detects 5 of the 9 anti-patterns automatically (missing-owner, no last-reviewed, wiki sprawl signal via orphan-rate, glossary drift, orphan pages).
- `runbook_validator.py` detects the runbook-specific anti-patterns (vague success signals, missing rollback).
- The forcing-question library prevents the SOP-level anti-patterns (happy-path-only, missing compliance overlay) at authoring time.
- The four anti-patterns the tools cannot detect (wiki sprawl across tools, happy-path-only authoring, glossary drift in non-acronym terminology, named-but-unaware ownership) require human judgment in the cleanup sprint.
The skill's job is to surface the 80% of anti-patterns a tool can find. The remaining 20% is the cleanup-sprint discussion.
FILE:references/runbook_canon.md
# Runbook Canon
Internal-operations runbook design discipline — what makes a runbook safe to execute at 3am during an incident, and where the discipline comes from. Seven authoritative sources cited.
## What a runbook is (and is not)
A **runbook** is the executable artifact an operator follows under time pressure. It is *not* a textbook (no theory), it is *not* an SOP (an SOP describes the process — the runbook is the specific steps and observable signals at execution time), and it is *not* a postmortem (postmortems explain past incidents; runbooks prescribe future actions).
Every runbook step must specify six things — and `runbook_validator.py` enforces all six:
1. **Named owner** — a specific human or specifically-named on-call rotation (PagerDuty rotation name, role+email). Not "the team", not "ops".
2. **Expected duration** — concrete number + unit. "5 minutes", "30 seconds". Not "quick" or "fast".
3. **Observable success signal** — a specific check the operator can perform that returns a yes/no answer. "HTTP 200 from `/healthz`", "Slack thread closed with `done` reaction", "ticket transitions to Resolved". Not "service is up", not "looks good".
4. **Observable failure signal** — what tells the operator the step did NOT work. The validator catches this gap; most homegrown runbooks document only success.
5. **Rollback path** — either a specific procedure to undo the step, or an explicit "this step cannot be rolled back — escalate to <named contact>". Silent absence of rollback is the most dangerous gap.
6. **Escalation contact** — named human, role+email, or named on-call rotation. Not "engineering", not "ops".
## Why these six attributes specifically
These six are the union of the requirements imposed by the seven sources below. Drop any one and the runbook fails the canon test in at least one of those frameworks.
## Seven authoritative sources
### 1. Beyer, Murphy, Rensin, Kawahara, Thorne (eds.) — *The Site Reliability Workbook* (O'Reilly, 2018), Ch. 8
Google SRE Workbook on "On-Call". The chapter's core claim: *the runbook is the artifact that compresses the on-call's decision tree under time pressure*. Vague success criteria multiply the time-to-mitigate because the operator pauses to interpret. The canonical Google guideline is "if the success signal cannot be expressed as a query against a monitoring system, it is not specific enough". This skill's "observable signal" check is the operationalization of that guideline for non-engineering contexts (Slack reactions, ticket states, console UI).
### 2. Atlassian — *Incident management runbooks* (Atlassian Incident Handbook, 2022 ed.)
Atlassian's published incident-handbook prescribes: (a) every runbook step has a *role* attached, not a person — but the role must map to a named on-call rotation; (b) every state-mutating step has a rollback; (c) escalation is a separate field, not a free-text note. This skill's `--profile support` variant of `sop_generator.py` follows Atlassian's escalation-matrix convention.
### 3. PagerDuty — *Incident Response Documentation* (PagerDuty open-source, 2017 onwards)
PagerDuty's open-source incident-response framework distinguishes between **major-incident runbooks** (the comms cascade — who's notified, in what order, with what SLA) and **technical-recovery runbooks** (the engineering steps to mitigate). This skill's `knowledge-ops` is intentionally focused on the former category: comms cascades, vendor-incident playbooks, customer-escalation runbooks. Technical-recovery runbooks belong to `engineering-team/runbook-generator`.
### 4. AWS — *Well-Architected Framework, Operational Excellence pillar* (AWS, ongoing)
AWS's Operational Excellence pillar makes the canonical argument for rollback discipline: *"you cannot run a process you cannot reverse without first agreeing what 'reverse' means"*. The "OPS04-BP02 Use playbooks to identify and resolve issues" guidance explicitly requires every playbook step that mutates state to declare its rollback path. The `runbook_validator.py` `ROLLBACK` check enforces this.
### 5. Charity Majors — *Observability Engineering* (O'Reilly, 2022, co-authored with George Miranda and Liz Fong-Jones)
Majors' argument that **runbooks decay faster than the systems they describe** is the canonical justification for `kb_ingester.py`'s stale-page detection. Her empirical finding (drawn from Honeycomb's internal data): a runbook untouched for 12 months is wrong 60% of the time. The default `--stale-days 365` setting in `kb_ingester.py` is calibrated to this.
### 6. Susan Fowler — *Production-Ready Microservices* (O'Reilly, 2016)
Fowler's Ch. 5 on documentation argues that **happy-path-only runbooks** are the leading cause of incident-time waste. Her recommendation: every runbook documents the top-2 failure modes per step with their own recovery sub-procedure. The forcing-question library in `SKILL.md` enforces this at the question-7 stage.
### 7. ITIL v4 — *Service Operation* practice guide (Axelos, 2019)
ITIL v4 makes the formal distinction between *procedure* (the SOP) and *work instruction* (the runbook): the procedure describes what is to be done at a process level; the work instruction describes how to do it at the step level. Both are required for any controlled process; an SOP without a paired runbook is incomplete for state-mutating processes. This is why `knowledge-ops` ships both `sop_generator.py` and `runbook_validator.py` — the same KB needs both artifact types.
## Common runbook anti-patterns
- **"The team owns it"** — no it doesn't. Name a human or an explicitly-defined on-call rotation.
- **"Verify the service is up"** — what does "up" mean to a new operator at 3am? Specify the observable check.
- **"Rollback: see runbook X"** — and runbook X says "see runbook Y". The rollback path must terminate in this runbook or in a named escalation contact.
- **"Escalation: engineering"** — which person, which rotation, what SLA? Engineering is 200 people.
- **Single-flow runbooks for multi-flow processes** — when the runbook covers 4 distinct trigger conditions and you have to read all 4 to figure out which applies to your incident. Split it.
- **Runbooks last reviewed before the system was rearchitected.** The stale check catches these.
## How this skill applies the canon
- `runbook_validator.py` enforces all six attributes per step.
- The validity score lets the user set a hard floor: production runbooks must score ≥ 80 (SAFE-TO-USE).
- `kb_ingester.py` flags stale runbooks (default 12 months) per Majors's decay finding.
- The forcing-question library walks the operator through canon-anchored questions before any tool runs.
FILE:scripts/kb_ingester.py
#!/usr/bin/env python3
"""kb_ingester.py
Walk a directory of markdown files (Notion export, Confluence space export,
Obsidian vault, Drive/SOPs/ directory) and emit a KB health report.
Extracts:
- cross-link map (which page references which)
- orphan pages (no inbound links)
- glossary candidates (frequently-used proper nouns / acronyms recurring
in 3+ docs with no single canonical definition page)
- glossary drift (same term used inconsistently across docs)
- stale pages (no edit in > N months — N defaults to 12)
- missing-owner pages (no `owner:` in YAML frontmatter)
- prioritized cleanup list ranked by (staleness × inbound-link-count)
Stdlib only.
"""
from __future__ import annotations
import argparse
import datetime as dt
import json
import re
import sys
import tempfile
from collections import Counter, defaultdict
from dataclasses import dataclass, field
from pathlib import Path
YAML_FRONTMATTER_RE = re.compile(
r"^---\s*\n(.*?)\n---\s*\n", re.DOTALL)
MD_LINK_RE = re.compile(r"\[([^\]]+)\]\(([^)]+)\)")
WIKI_LINK_RE = re.compile(r"\[\[([^\]|]+)(?:\|[^\]]+)?\]\]")
ACRONYM_RE = re.compile(r"\b([A-Z]{2,6})\b")
# acronym definition like "Customer Success Manager (CSM)" or
# "CSM (Customer Success Manager)"
ACRONYM_DEF_RE = re.compile(
r"\b((?:[A-Z][A-Za-z]+\s+){1,4}[A-Z][A-Za-z]+)\s*\(([A-Z]{2,6})\)"
r"|\b([A-Z]{2,6})\s*\(((?:[A-Z][A-Za-z]+\s+){1,4}[A-Z][A-Za-z]+)\)"
)
@dataclass
class PageInfo:
path: Path
title: str = ""
owner: str = ""
last_reviewed: str = ""
mtime_days_ago: int = 0
outbound_links: list = field(default_factory=list)
inbound_link_count: int = 0
acronyms_used: list = field(default_factory=list)
acronym_definitions: dict = field(default_factory=dict)
word_count: int = 0
def _parse_frontmatter(text: str) -> dict:
m = YAML_FRONTMATTER_RE.match(text)
if not m:
return {}
body = m.group(1)
fm = {}
for line in body.splitlines():
if ":" in line:
k, _, v = line.partition(":")
fm[k.strip().lower()] = v.strip().strip('"').strip("'")
return fm
def _extract_title(text: str, path: Path) -> str:
for line in text.splitlines():
m = re.match(r"^#\s+(.+)$", line)
if m:
return m.group(1).strip()
return path.stem.replace("-", " ").replace("_", " ").title()
def _extract_links(text: str) -> list:
links = []
for m in MD_LINK_RE.finditer(text):
target = m.group(2).strip()
if target.startswith(("http://", "https://", "mailto:")):
continue
links.append(target)
for m in WIKI_LINK_RE.finditer(text):
links.append(m.group(1).strip())
return links
def _extract_acronyms(text: str) -> tuple:
acronyms = ACRONYM_RE.findall(text)
defs = {}
for m in ACRONYM_DEF_RE.finditer(text):
if m.group(1) and m.group(2):
defs[m.group(2)] = m.group(1).strip()
elif m.group(3) and m.group(4):
defs[m.group(3)] = m.group(4).strip()
return acronyms, defs
def _normalize_link_target(target: str, source: Path, root: Path) -> str:
"""Resolve a link target to a canonical relative path string."""
target = target.split("#")[0].split("?")[0].strip()
if not target:
return ""
if target.endswith(".md"):
candidate = (source.parent / target).resolve()
elif "/" in target or "\\" in target:
candidate_md = (source.parent / (target + ".md")).resolve()
if candidate_md.exists():
candidate = candidate_md
else:
candidate = (source.parent / target).resolve()
else:
# bare title — try to match against any .md filename
candidate_md = (source.parent / (target + ".md")).resolve()
candidate = candidate_md
try:
return str(candidate.relative_to(root))
except ValueError:
return str(candidate)
def walk_vault(root: Path, stale_days: int = 365) -> list:
"""Walk a directory tree and return a list of PageInfo objects."""
pages = []
now = dt.datetime.now()
for path in sorted(root.rglob("*.md")):
if not path.is_file():
continue
try:
text = path.read_text(encoding="utf-8")
except (UnicodeDecodeError, OSError):
continue
fm = _parse_frontmatter(text)
title = fm.get("title") or _extract_title(text, path)
owner = fm.get("owner", "")
last_reviewed = fm.get("last_reviewed", "") or fm.get(
"last-reviewed", "")
# mtime fallback
try:
mtime = dt.datetime.fromtimestamp(path.stat().st_mtime)
mtime_days_ago = (now - mtime).days
except OSError:
mtime_days_ago = 0
outbound = _extract_links(text)
acronyms, defs = _extract_acronyms(text)
word_count = len(text.split())
pages.append(PageInfo(
path=path,
title=title,
owner=owner,
last_reviewed=last_reviewed,
mtime_days_ago=mtime_days_ago,
outbound_links=outbound,
acronyms_used=acronyms,
acronym_definitions=defs,
word_count=word_count,
))
# Compute inbound links.
by_relpath = {str(p.path.relative_to(root)): p for p in pages}
by_title = {p.title.lower(): p for p in pages}
by_stem = {p.path.stem.lower(): p for p in pages}
for src in pages:
for raw in src.outbound_links:
target_rel = _normalize_link_target(raw, src.path, root)
if target_rel in by_relpath:
by_relpath[target_rel].inbound_link_count += 1
continue
tgt = raw.split("#")[0].split("?")[0].strip().lower()
if tgt.endswith(".md"):
tgt = tgt[:-3]
if tgt in by_title:
by_title[tgt].inbound_link_count += 1
elif tgt in by_stem:
by_stem[tgt].inbound_link_count += 1
return pages
def detect_orphans(pages: list) -> list:
return [p for p in pages if p.inbound_link_count == 0]
def detect_stale(pages: list, stale_days: int) -> list:
out = []
for p in pages:
is_stale = False
if p.last_reviewed:
try:
lr = dt.datetime.strptime(p.last_reviewed[:10], "%Y-%m-%d")
if (dt.datetime.now() - lr).days > stale_days:
is_stale = True
except ValueError:
pass
elif p.mtime_days_ago > stale_days:
is_stale = True
if is_stale:
out.append(p)
return out
def detect_missing_owner(pages: list) -> list:
return [p for p in pages if not p.owner]
def detect_glossary_drift(pages: list) -> dict:
"""Return a dict {acronym: [list of (definition, source page)]} for
acronyms that have >= 2 distinct definitions across the vault."""
by_acronym = defaultdict(list)
for p in pages:
for ac, defin in p.acronym_definitions.items():
by_acronym[ac].append((defin, str(p.path)))
drift = {}
for ac, defs in by_acronym.items():
distinct = set(d.lower() for d, _ in defs)
if len(distinct) >= 2:
drift[ac] = defs
return drift
def detect_glossary_candidates(pages: list, min_docs: int = 3) -> list:
"""Acronyms used in >= min_docs pages with no canonical definition
page (no page where the acronym appears in the title)."""
doc_count = Counter()
titled = set()
for p in pages:
seen = set(p.acronyms_used)
for ac in seen:
doc_count[ac] += 1
for ac in p.acronym_definitions:
# If acronym appears in title, treat as canonical-ish.
if ac in p.title:
titled.add(ac)
return sorted([(ac, c) for ac, c in doc_count.items()
if c >= min_docs and ac not in titled],
key=lambda x: -x[1])
def cleanup_priority(pages: list, stale_days: int) -> list:
"""Rank pages by (staleness × inbound-link-count) — high-traffic
stale docs surface first."""
scored = []
for p in pages:
staleness = 0
if p.last_reviewed:
try:
lr = dt.datetime.strptime(p.last_reviewed[:10], "%Y-%m-%d")
staleness = max(0, (dt.datetime.now() - lr).days
- stale_days)
except ValueError:
staleness = max(0, p.mtime_days_ago - stale_days)
else:
staleness = max(0, p.mtime_days_ago - stale_days)
if staleness > 0:
# inbound +1 to avoid zeroing out everything orphan
score = staleness * (p.inbound_link_count + 1)
scored.append((score, p))
scored.sort(key=lambda x: -x[0])
return scored
def generate_report(root: Path, pages: list, stale_days: int) -> str:
orphans = detect_orphans(pages)
stale = detect_stale(pages, stale_days)
missing_owner = detect_missing_owner(pages)
drift = detect_glossary_drift(pages)
candidates = detect_glossary_candidates(pages)
priority = cleanup_priority(pages, stale_days)
lines = [
f"# KB health report — `{root}`",
"",
f"**Pages scanned:** {len(pages)}",
f"**Stale threshold:** {stale_days} days",
"",
"## Summary metrics",
"",
"| Metric | Count | % of vault |",
"|--------|-------|------------|",
f"| Orphan pages (no inbound links) | {len(orphans)} | "
f"{round(len(orphans) / max(len(pages), 1) * 100, 1)}% |",
f"| Stale pages (> {stale_days}d) | {len(stale)} | "
f"{round(len(stale) / max(len(pages), 1) * 100, 1)}% |",
f"| Missing-owner pages | {len(missing_owner)} | "
f"{round(len(missing_owner) / max(len(pages), 1) * 100, 1)}% |",
f"| Glossary drift (acronyms with >= 2 defs) | {len(drift)} | — |",
f"| Glossary candidates (acronyms in 3+ docs, no canonical page) "
f"| {len(candidates)} | — |",
"",
]
lines.append("## Top-20 cleanup priority "
"(staleness × inbound-link-count + 1)")
lines.append("")
if priority:
lines.append("| Rank | Score | Path | Inbound | "
"Days stale | Owner |")
lines.append("|------|-------|------|---------|"
"------------|-------|")
for i, (score, p) in enumerate(priority[:20], start=1):
rel = p.path.relative_to(root)
staleness = (p.mtime_days_ago - stale_days
if not p.last_reviewed else
(dt.datetime.now() - dt.datetime.strptime(
p.last_reviewed[:10], "%Y-%m-%d")).days
- stale_days)
lines.append(
f"| {i} | {score} | `{rel}` | {p.inbound_link_count} "
f"| {staleness} | {p.owner or '(MISSING)'} |"
)
else:
lines.append("_(no stale pages — KB is current)_")
lines.append("")
lines.append("## Orphan pages (no inbound links)")
lines.append("")
if orphans:
for p in orphans[:30]:
rel = p.path.relative_to(root)
lines.append(f"- `{rel}` — {p.title}")
if len(orphans) > 30:
lines.append(f"- _(+{len(orphans) - 30} more not shown)_")
else:
lines.append("_(none — every page has at least one inbound link)_")
lines.append("")
lines.append("## Glossary drift (acronym defined differently across "
"docs)")
lines.append("")
if drift:
for ac, defs in drift.items():
lines.append(f"**{ac}:**")
for defin, src in defs:
lines.append(f" - `{defin}` (in `{src}`)")
lines.append("")
else:
lines.append("_(none detected — acronyms are used consistently)_")
lines.append("")
lines.append("## Glossary candidates (acronym used in 3+ docs "
"without a canonical definition page)")
lines.append("")
if candidates:
for ac, count in candidates[:20]:
lines.append(f"- **{ac}** — used in {count} docs, no "
f"canonical definition page exists")
else:
lines.append("_(none — acronyms either have canonical pages or "
"are uncommon)_")
lines.append("")
lines.append("## Missing-owner pages")
lines.append("")
if missing_owner:
for p in missing_owner[:30]:
rel = p.path.relative_to(root)
lines.append(f"- `{rel}` — {p.title}")
if len(missing_owner) > 30:
lines.append(
f"- _(+{len(missing_owner) - 30} more not shown)_")
else:
lines.append("_(none — every page has an owner)_")
lines.append("")
lines.append("## Recommended next actions")
lines.append("")
lines.append("1. Assign owners to the missing-owner pages first — "
"no other fix sticks without ownership.")
lines.append("2. Resolve glossary drift by picking one canonical "
"definition per acronym; add a `glossary.md` page; "
"link every other doc to it.")
lines.append("3. Triage the top-20 cleanup list: archive, rewrite, "
"or refresh. Re-run this report after the sprint to "
"verify orphan + stale counts are down.")
lines.append("4. Pair orphan pages with a navigation review — some "
"orphans are reference pages found via search and "
"should NOT be archived. Curate, don't bulk-delete.")
return "\n".join(lines) + "\n"
def generate_json_report(root: Path, pages: list, stale_days: int) -> dict:
orphans = detect_orphans(pages)
stale = detect_stale(pages, stale_days)
missing_owner = detect_missing_owner(pages)
drift = detect_glossary_drift(pages)
candidates = detect_glossary_candidates(pages)
priority = cleanup_priority(pages, stale_days)
return {
"root": str(root),
"page_count": len(pages),
"stale_days_threshold": stale_days,
"orphan_count": len(orphans),
"stale_count": len(stale),
"missing_owner_count": len(missing_owner),
"glossary_drift_count": len(drift),
"glossary_candidate_count": len(candidates),
"top_cleanup": [
{
"rank": i + 1,
"score": score,
"path": str(p.path.relative_to(root)),
"inbound_links": p.inbound_link_count,
"owner": p.owner or None,
}
for i, (score, p) in enumerate(priority[:20])
],
"orphans": [str(p.path.relative_to(root)) for p in orphans],
"glossary_drift": {ac: [{"definition": d, "source": s}
for d, s in defs]
for ac, defs in drift.items()},
"glossary_candidates": [{"acronym": ac, "doc_count": c}
for ac, c in candidates],
"missing_owner": [str(p.path.relative_to(root))
for p in missing_owner],
}
SAMPLE_PAGES = {
"index.md": """---
owner: alex@company.com
last_reviewed: 2026-04-01
---
# Ops Index
Welcome to the Ops wiki. Start with [Vendor Offboarding](sops/vendor-offboarding.md) or [Incident Comms](runbooks/incident-comms.md).
The [Glossary](glossary.md) defines our terms.
""",
"glossary.md": """---
owner: alex@company.com
last_reviewed: 2026-04-15
---
# Glossary
- Customer Success Manager (CSM) — owns post-sale account relationship.
- Vendor Management Office (VMO) — owns third-party vendor lifecycle.
""",
"sops/vendor-offboarding.md": """---
owner: jordan@company.com
last_reviewed: 2026-02-01
---
# Vendor Offboarding SOP
The VMO operator runs this SOP when a vendor contract is terminated.
See also [Incident Comms](../runbooks/incident-comms.md).
The CSM is notified.
""",
"sops/procurement-intake.md": """---
owner: jordan@company.com
last_reviewed: 2024-01-01
---
# Procurement Intake SOP
Run this when finance receives a purchase request. The CSM (Customer Solutions Manager) reviews it.
""", # NOTE: glossary drift — CSM here is Customer Solutions Manager
"runbooks/incident-comms.md": """---
last_reviewed: 2026-03-01
---
# Incident Comms Cascade
(no owner field — missing-owner case)
Send alerts to the on-call SRE.
""",
"orphan-page.md": """---
owner: pat@company.com
last_reviewed: 2026-04-01
---
# Orphan Page
Nobody links here.
The CSM may find this useful.
""",
"old-stale-page.md": """# Old Page
(no frontmatter at all — missing-owner AND probably stale via mtime)
""",
"sops/employee-onboarding.md": """---
owner: hr@company.com
last_reviewed: 2026-04-20
---
# Employee Onboarding SOP
Coordinate with the CSM and VMO for system access.
Link: [Vendor Offboarding](vendor-offboarding.md).
""",
}
def _materialize_sample_vault() -> Path:
tmp = Path(tempfile.mkdtemp(prefix="kb-sample-"))
for relpath, content in SAMPLE_PAGES.items():
full = tmp / relpath
full.parent.mkdir(parents=True, exist_ok=True)
full.write_text(content, encoding="utf-8")
# Backdate one file via os.utime so mtime-based stale detection
# has something to find.
import os
old = tmp / "old-stale-page.md"
if old.exists():
old_ts = (dt.datetime.now() -
dt.timedelta(days=720)).timestamp()
os.utime(old, (old_ts, old_ts))
return tmp
def main(argv=None) -> int:
p = argparse.ArgumentParser(
description="Walk a markdown KB and emit a hygiene report: "
"orphans, stale, missing-owner, glossary drift."
)
p.add_argument("--input", "-i", type=str,
help="Path to KB root directory.")
p.add_argument("--output", "-o", choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).")
p.add_argument("--stale-days", type=int, default=365,
help="Days since last edit to consider stale "
"(default: 365).")
p.add_argument("--sample", action="store_true",
help="Run against a tiny synthetic vault in a "
"tmpdir.")
args = p.parse_args(argv)
if args.sample:
root = _materialize_sample_vault()
elif args.input:
root = Path(args.input).resolve()
if not root.exists() or not root.is_dir():
print(f"ERROR: input directory not found: {args.input}",
file=sys.stderr)
return 2
else:
print("ERROR: provide --input <kb-root-dir> or --sample",
file=sys.stderr)
return 2
pages = walk_vault(root, stale_days=args.stale_days)
if not pages:
print(f"WARNING: no markdown files found under {root}",
file=sys.stderr)
return 1
if args.output == "json":
print(json.dumps(generate_json_report(root, pages, args.stale_days),
indent=2))
else:
print(generate_report(root, pages, args.stale_days))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/runbook_validator.py
#!/usr/bin/env python3
"""runbook_validator.py
Validate a runbook by checking each step against six required attributes:
1. Named owner (not "the team", not "ops")
2. Expected duration (concrete number + unit)
3. Observable success signal
4. Observable failure signal
5. Rollback path (or explicit "cannot roll back — escalate to X")
6. Escalation contact
Output is a per-step traffic-light + overall validity score 0-100 + a list
of MUST-FIX issues.
Verdict thresholds:
>= 80 SAFE-TO-USE
60-79 USE-WITH-CAUTION
< 60 NOT-SAFE
Input formats:
--input runbook.md (markdown: heuristic parser, expects
"## Step N:" or "### Step N:" headings)
--input runbook.json (JSON: explicit step list — preferred)
JSON schema:
{
"runbook_name": "Incident Comms Cascade",
"steps": [
{
"title": "Acknowledge alert in PagerDuty",
"owner": "On-call IC (named rotation)",
"duration_minutes": 2,
"success_signal": "PagerDuty incident transitions to acknowledged",
"failure_signal": "Incident remains in triggered state after 2 min",
"rollback": "n/a (acknowledgement is non-mutating)",
"escalation": "Engineering Manager on-call"
}
]
}
Stdlib only.
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from dataclasses import dataclass, field, asdict
from pathlib import Path
VAGUE_OWNER_TOKENS = {
"the team", "team", "ops", "the ops team", "engineering",
"support", "everyone", "whoever", "someone", "tbd", "n/a",
"the on-call", "on call", "rotation", # rotation alone is vague
}
# Vague success signal phrases that get flagged. Matched as whole-phrase
# substrings — must be specific enough to avoid false positives on
# legitimate observables that happen to contain a common word.
VAGUE_SUCCESS_TOKENS = [
"service is up", "it works", "things look good", "looks fine",
"no errors", "should work", "appears to be",
"verify the service", "check that it works", "looks good",
]
# Phrases that count as observable.
OBSERVABLE_HINTS = [
"http 2", "http 3", "http 4", "http 5", # status codes
"status code", "exit code 0", "/healthz", "/health", "200 ok",
"log line", "metric", "dashboard shows", "alert clears",
"incident transitions", "ticket moves to", "slack reaction",
"email received", "record updated", "field set to",
]
DURATION_PATTERN = re.compile(
r"\b\d+(?:\.\d+)?\s*(seconds?|secs?|minutes?|mins?|hours?|hrs?|days?)\b",
re.IGNORECASE,
)
# Rollback acceptable phrasing: either a real rollback OR explicit
# acknowledgement that rollback is impossible plus escalation.
NO_ROLLBACK_ACCEPTABLE = [
"cannot be rolled back",
"cannot roll back",
"non-mutating",
"read-only",
"no rollback needed",
"irreversible — escalate",
"irreversible - escalate",
]
@dataclass
class StepFinding:
step_index: int
title: str
owner_ok: bool = False
duration_ok: bool = False
success_ok: bool = False
failure_ok: bool = False
rollback_ok: bool = False
escalation_ok: bool = False
issues: list = field(default_factory=list)
@property
def passes(self) -> int:
return sum([
self.owner_ok, self.duration_ok, self.success_ok,
self.failure_ok, self.rollback_ok, self.escalation_ok,
])
@property
def traffic_light(self) -> str:
if self.passes == 6:
return "GREEN"
if self.passes >= 4:
return "AMBER"
return "RED"
def _check_owner(owner: str) -> tuple[bool, str]:
if not owner or not owner.strip():
return False, "missing owner"
norm = owner.strip().lower()
for token in VAGUE_OWNER_TOKENS:
# Vague if owner is ONLY that token (allow named rotations like
# "SRE on-call (alex)" by checking for parenthetical name OR @).
if norm == token or norm.startswith(token + " "):
if "@" in owner or "(" in owner:
return True, ""
return False, (
f"vague owner '{owner}' — name a specific human or a "
f"specifically-named rotation (e.g., 'SRE on-call "
f"rotation (PagerDuty: sre-primary)')"
)
return True, ""
def _check_duration(duration_str: str, duration_minutes) -> tuple[bool, str]:
if duration_minutes is not None:
try:
val = float(duration_minutes)
if val > 0:
return True, ""
return False, "duration_minutes is zero or negative"
except (TypeError, ValueError):
pass
if duration_str and DURATION_PATTERN.search(duration_str):
return True, ""
return False, (
"missing expected duration (need a concrete number + unit, "
"e.g., '2 minutes', '30 seconds')"
)
def _check_observable(signal: str, kind: str) -> tuple[bool, str]:
if not signal or not signal.strip():
return False, f"missing observable {kind} signal"
norm = signal.lower()
for vague in VAGUE_SUCCESS_TOKENS:
if vague in norm:
return False, (
f"vague {kind} signal '{signal}' — need an observable "
f"(e.g., 'HTTP 200 from /healthz', not 'service is up')"
)
for hint in OBSERVABLE_HINTS:
if hint in norm:
return True, ""
# Heuristic: if signal contains digits, equality, code-fences, or
# specific verbs that imply an observation, accept.
if any(ch in signal for ch in ("=", ":", "`", "200", "404", "500")):
return True, ""
if re.search(r"\b(returns?|equals?|shows?|transitions?|moves?|"
r"closes?|emits?|logs?|created|deleted|received|"
r"updated|set\s+to|reaches?|reports?)\b", norm):
return True, ""
return False, (
f"{kind} signal '{signal}' is not clearly observable — rewrite "
f"as a concrete check (status code, log line, dashboard panel, "
f"ticket state)"
)
def _check_rollback(rollback: str) -> tuple[bool, str]:
if not rollback or not rollback.strip():
return False, "missing rollback path"
norm = rollback.lower()
for ok in NO_ROLLBACK_ACCEPTABLE:
if ok in norm:
return True, ""
# If there's substantive text (> 12 chars) describing a step, accept.
if len(rollback.strip()) >= 12:
return True, ""
return False, (
f"rollback path too thin ('{rollback}') — either describe the "
f"rollback procedure OR write 'cannot be rolled back — "
f"escalate to <name>'"
)
def _check_escalation(escalation: str) -> tuple[bool, str]:
if not escalation or not escalation.strip():
return False, "missing escalation contact"
norm = escalation.strip().lower()
for token in VAGUE_OWNER_TOKENS:
if norm == token or norm.startswith(token + " "):
if "@" not in escalation and "(" not in escalation:
return False, (
f"vague escalation contact '{escalation}' — name a "
f"specific human, role+email, or named on-call rotation"
)
return True, ""
def validate_step(step: dict, idx: int) -> StepFinding:
finding = StepFinding(
step_index=idx,
title=step.get("title", f"(step {idx} — no title)"),
)
owner_ok, owner_err = _check_owner(step.get("owner", ""))
finding.owner_ok = owner_ok
if not owner_ok:
finding.issues.append(f"OWNER: {owner_err}")
duration_ok, duration_err = _check_duration(
step.get("duration_str", ""),
step.get("duration_minutes"),
)
finding.duration_ok = duration_ok
if not duration_ok:
finding.issues.append(f"DURATION: {duration_err}")
succ_ok, succ_err = _check_observable(
step.get("success_signal", ""), "success")
finding.success_ok = succ_ok
if not succ_ok:
finding.issues.append(f"SUCCESS: {succ_err}")
fail_ok, fail_err = _check_observable(
step.get("failure_signal", ""), "failure")
finding.failure_ok = fail_ok
if not fail_ok:
finding.issues.append(f"FAILURE: {fail_err}")
rb_ok, rb_err = _check_rollback(step.get("rollback", ""))
finding.rollback_ok = rb_ok
if not rb_ok:
finding.issues.append(f"ROLLBACK: {rb_err}")
esc_ok, esc_err = _check_escalation(step.get("escalation", ""))
finding.escalation_ok = esc_ok
if not esc_ok:
finding.issues.append(f"ESCALATION: {esc_err}")
return finding
def _parse_markdown(text: str) -> dict:
"""Heuristic parser. Expects steps as '## Step N: title' or
'### Step N: title' followed by bullet attributes."""
lines = text.splitlines()
name_match = re.search(r"^#\s+(.+)$", text, re.MULTILINE)
runbook_name = name_match.group(1).strip() if name_match else "(unnamed)"
steps = []
current = None
step_re = re.compile(
r"^#{2,3}\s+Step\s+(\d+)\s*:?\s*(.*)$", re.IGNORECASE)
attr_re = re.compile(
r"^\s*[-*]\s+\*?\*?(Owner|Duration|Success|Failure|"
r"Rollback|Escalation)\*?\*?\s*:?\s*(.+)$",
re.IGNORECASE,
)
for line in lines:
m = step_re.match(line)
if m:
if current:
steps.append(current)
current = {"title": m.group(2).strip() or f"step {m.group(1)}"}
continue
if current:
am = attr_re.match(line)
if am:
key = am.group(1).lower()
val = am.group(2).strip()
if key == "owner":
current["owner"] = val
elif key == "duration":
current["duration_str"] = val
elif key == "success":
current["success_signal"] = val
elif key == "failure":
current["failure_signal"] = val
elif key == "rollback":
current["rollback"] = val
elif key == "escalation":
current["escalation"] = val
if current:
steps.append(current)
return {"runbook_name": runbook_name, "steps": steps}
def _sample_runbook() -> dict:
"""Deliberately broken incident-comms runbook to demonstrate
failure detection."""
return {
"runbook_name": "Incident Comms Cascade (BROKEN sample)",
"steps": [
{
"title": "Acknowledge alert",
"owner": "the team", # vague
"duration_str": "", # missing
"success_signal": "service is up", # vague
"failure_signal": "", # missing
"rollback": "", # missing
"escalation": "ops", # vague
},
{
"title": "Open incident channel",
"owner": "Incident Commander on-call "
"(PagerDuty: ic-primary)",
"duration_str": "2 minutes",
"success_signal": "Slack channel #inc-<id> created and "
"linked from PagerDuty incident",
"failure_signal": "Slack returns 4xx or channel-create "
"API call times out",
"rollback": "n/a — read-only operation (channel can be "
"archived if created in error)",
"escalation": "Engineering Manager on-call "
"(em-primary@company.com)",
},
{
"title": "Notify execs via paging tree",
"owner": "Communications Lead "
"(comms-lead@company.com)",
"duration_str": "5 minutes",
"success_signal": "Exec recipient list shows email "
"received (200 OK from SES API)",
"failure_signal": "SES API returns 5xx or recipient "
"delivery status = bounced",
"rollback": "Send retraction email to same list with "
"subject prefix 'RETRACTION:'",
"escalation": "VP Communications "
"(vp-comms@company.com)",
},
],
}
def generate_report(runbook: dict, findings: list) -> str:
total = len(findings)
if total == 0:
return "ERROR: runbook contains no steps."
score = round(sum(f.passes for f in findings) /
(6 * total) * 100, 1)
if score >= 80:
verdict = "SAFE-TO-USE"
elif score >= 60:
verdict = "USE-WITH-CAUTION"
else:
verdict = "NOT-SAFE"
lines = [
f"# Runbook validation: {runbook.get('runbook_name', '(unnamed)')}",
"",
f"**Steps validated:** {total}",
f"**Validity score:** {score} / 100",
f"**Verdict:** {verdict}",
"",
"## Per-step traffic-light",
"",
"| Step | Title | Owner | Duration | Success | Failure | "
"Rollback | Escalation | Light |",
"|------|-------|-------|----------|---------|---------|"
"----------|------------|-------|",
]
for f in findings:
def ck(b):
return "OK" if b else "FAIL"
lines.append(
f"| {f.step_index} | {f.title[:40]} | {ck(f.owner_ok)} | "
f"{ck(f.duration_ok)} | {ck(f.success_ok)} | "
f"{ck(f.failure_ok)} | {ck(f.rollback_ok)} | "
f"{ck(f.escalation_ok)} | {f.traffic_light} |"
)
lines.append("")
lines.append("## MUST-FIX issues")
lines.append("")
any_issues = False
for f in findings:
if f.issues:
any_issues = True
lines.append(f"### Step {f.step_index}: {f.title}")
for issue in f.issues:
lines.append(f"- {issue}")
lines.append("")
if not any_issues:
lines.append("_(none — all steps pass all six checks)_")
return "\n".join(lines) + "\n"
def generate_json_report(runbook: dict, findings: list) -> dict:
total = len(findings) or 1
score = round(sum(f.passes for f in findings) / (6 * total) * 100, 1)
verdict = ("SAFE-TO-USE" if score >= 80
else "USE-WITH-CAUTION" if score >= 60
else "NOT-SAFE")
return {
"runbook_name": runbook.get("runbook_name", "(unnamed)"),
"step_count": len(findings),
"validity_score": score,
"verdict": verdict,
"findings": [asdict(f) | {"traffic_light": f.traffic_light,
"passes": f.passes} for f in findings],
}
def main(argv=None) -> int:
p = argparse.ArgumentParser(
description="Validate a runbook against six step-completeness "
"rules. Output traffic-light + score + MUST-FIX list."
)
p.add_argument("--input", "-i", type=str,
help="Path to runbook .md or .json file.")
p.add_argument("--output", "-o", choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).")
p.add_argument("--sample", action="store_true",
help="Run against a deliberately-broken sample runbook.")
args = p.parse_args(argv)
if args.sample:
runbook = _sample_runbook()
elif args.input:
path = Path(args.input)
if not path.exists():
print(f"ERROR: input file not found: {args.input}",
file=sys.stderr)
return 2
text = path.read_text()
if path.suffix.lower() == ".json":
runbook = json.loads(text)
else:
runbook = _parse_markdown(text)
else:
print("ERROR: provide --input <runbook.md|json> or --sample",
file=sys.stderr)
return 2
steps = runbook.get("steps", [])
if not steps:
print("ERROR: runbook contains no steps "
"(or markdown parser found none — try JSON input)",
file=sys.stderr)
return 1
findings = [validate_step(s, i + 1) for i, s in enumerate(steps)]
if args.output == "json":
print(json.dumps(generate_json_report(runbook, findings),
indent=2))
else:
print(generate_report(runbook, findings))
return 0
if __name__ == "__main__":
sys.exit(main())
FILE:scripts/sop_generator.py
#!/usr/bin/env python3
"""sop_generator.py
Generate a 5W2H-structured Standard Operating Procedure (SOP) from a JSON
metadata file. Output is markdown by default, or normalized JSON.
5W2H = Who, What, When, Where, Why, How, How-much (Ishikawa, *Guide to
Quality Control*, 1985). Each section is mandatory; missing sections produce
a warning footer naming the section.
Industry tuning:
--profile {ops,support,finance,hr,it,regulated}
- ops: general internal ops SOP scaffold
- support: adds customer-impact section + escalation matrix
- finance: adds controls + reconciliation + segregation-of-duties section
- hr: flags PII / sensitive-data handling; adds consent section
- it: adds system + access + change-management section
- regulated: adds version control, signoff matrix, audit-trail, change
history (required under ISO 9001 / FDA 21 CFR Part 211 /
SOC 2 / HIPAA / ISO 13485)
Regulatory overlay flags attach the appropriate compliance preamble:
regulatory_overlay: ["SOC2", "HIPAA", "ISO13485", "GDPR", "SOX"]
Input schema (JSON):
{
"sop_name": "Vendor Offboarding",
"process_owner": "alex@company.com",
"triggering_event": "Vendor contract not renewed OR vendor terminated",
"audience_role": "Vendor Management Office operator",
"frequency": "On-demand (avg 3 times per quarter)",
"regulatory_overlay": ["SOC2"],
"inputs": ["Vendor name", "Contract end date", "Data access list"],
"outputs": ["Access revoked", "Data deleted/returned", "Final invoice paid"],
"steps_outline": [
"Notify vendor of offboarding intent",
"Inventory data and system access",
"Revoke production system access",
"Confirm data deletion or return",
"Final invoice reconciliation",
"Archive vendor record"
],
"estimated_minutes": 240,
"estimated_cost_usd": 800
}
Stdlib only.
"""
from __future__ import annotations
import argparse
import json
import sys
from dataclasses import dataclass, field, asdict
from pathlib import Path
VALID_PROFILES = {"ops", "support", "finance", "hr", "it", "regulated"}
VALID_OVERLAYS = {"SOC2", "HIPAA", "ISO13485", "GDPR", "SOX"}
REGULATORY_PREAMBLE = {
"SOC2": (
"**SOC 2 overlay:** This SOP supports the Common Criteria control "
"framework. Changes require change-management approval (CC8.1). "
"Evidence of execution must be retained for the audit period."
),
"HIPAA": (
"**HIPAA overlay:** This SOP touches Protected Health Information "
"(PHI). All access must be logged per §164.312(b). Minimum-necessary "
"rule applies (§164.502(b))."
),
"ISO13485": (
"**ISO 13485 overlay:** This is a controlled document under §4.2.4. "
"Document revision, approval, and review records must be maintained. "
"Use the regulated profile."
),
"GDPR": (
"**GDPR overlay:** This SOP touches personal data of EU data "
"subjects. Lawful basis must be documented (Art. 6). Data-subject "
"rights (Art. 15-22) requests must be respected during execution."
),
"SOX": (
"**SOX overlay:** This SOP supports a financial control. Execution "
"must be evidenced and segregation-of-duties enforced. Quarterly "
"management testing applies."
),
}
@dataclass
class SOPMetadata:
sop_name: str = ""
process_owner: str = ""
triggering_event: str = ""
audience_role: str = ""
frequency: str = ""
regulatory_overlay: list = field(default_factory=list)
inputs: list = field(default_factory=list)
outputs: list = field(default_factory=list)
steps_outline: list = field(default_factory=list)
estimated_minutes: int = 0
estimated_cost_usd: int = 0
def validate(self) -> list:
errs = []
for fld in ("sop_name", "process_owner", "triggering_event",
"audience_role", "frequency"):
if not getattr(self, fld):
errs.append(f"missing required field: '{fld}'")
if not self.steps_outline:
errs.append("missing 'steps_outline' (need >= 1 step)")
for ov in self.regulatory_overlay:
if ov not in VALID_OVERLAYS:
errs.append(
f"invalid regulatory_overlay '{ov}'; "
f"allowed: {sorted(VALID_OVERLAYS)}"
)
return errs
def _sample_metadata() -> dict:
return {
"sop_name": "Vendor Offboarding",
"process_owner": "alex@company.com (Vendor Management Lead)",
"triggering_event": (
"Vendor contract not renewed OR vendor terminated for cause"
),
"audience_role": "Vendor Management Office (VMO) operator",
"frequency": "On-demand (avg 3 executions per quarter)",
"regulatory_overlay": ["SOC2"],
"inputs": [
"Vendor legal name",
"Contract end date (effective offboarding date)",
"List of systems with vendor access",
"List of data classes vendor processed",
],
"outputs": [
"All production system access revoked (evidenced)",
"Vendor data deleted or returned (evidenced)",
"Final invoice reconciled and paid",
"Vendor record archived in VMO registry",
],
"steps_outline": [
"Notify vendor of offboarding intent (written, 30 days notice)",
"Inventory data classes and system access vendor holds",
"Revoke production system access (IAM, VPN, SaaS)",
"Confirm data deletion (vendor certification) or data return",
"Final invoice reconciliation and payment",
"Archive vendor record in VMO registry with offboarding evidence",
],
"estimated_minutes": 240,
"estimated_cost_usd": 800,
}
def _build_who(meta: SOPMetadata, profile: str) -> str:
lines = [
"### Who",
"",
f"- **Process owner (Accountable):** {meta.process_owner}",
f"- **Audience (Responsible):** {meta.audience_role}",
]
if profile == "regulated":
lines.append("- **Approver (Consulted):** "
"Quality Management Representative")
lines.append("- **Auditor (Informed):** "
"Internal Audit / Compliance")
elif profile == "finance":
lines.append("- **Approver (Consulted):** Controller")
lines.append("- **Segregation-of-duties review:** "
"Required (initiator != approver != payer)")
elif profile == "hr":
lines.append("- **Approver (Consulted):** HR Business Partner")
lines.append("- **Privacy review (Informed):** "
"Data Protection Officer (if PII touched)")
elif profile == "it":
lines.append("- **Approver (Consulted):** "
"Change Advisory Board (for system-mutating steps)")
elif profile == "support":
lines.append("- **Approver (Consulted):** Support Team Lead")
lines.append("- **Escalation (Informed):** "
"Engineering on-call (if customer-impact > 30 min)")
return "\n".join(lines)
def _build_what(meta: SOPMetadata) -> str:
lines = [
"### What",
"",
f"**Process name:** {meta.sop_name}",
"",
"**Inputs required before starting:**",
"",
]
for inp in meta.inputs:
lines.append(f"- {inp}")
lines.append("")
lines.append("**Outputs produced:**")
lines.append("")
for out in meta.outputs:
lines.append(f"- {out}")
return "\n".join(lines)
def _build_when(meta: SOPMetadata) -> str:
return (
"### When\n\n"
f"- **Triggering event:** {meta.triggering_event}\n"
f"- **Frequency:** {meta.frequency}\n"
"- **Time-of-day constraint:** _(business hours only? on-call? "
"fill in)_\n"
"- **Blocking dependencies:** _(prerequisites that must be true "
"before starting)_"
)
def _build_where(meta: SOPMetadata, profile: str) -> str:
lines = [
"### Where",
"",
"- **Primary system of record:** _(name the system — Salesforce, "
"Notion, Jira, ServiceNow, etc.)_",
"- **Supporting tools:** _(IAM console, IT ticketing, accounting "
"system, etc.)_",
"- **Canonical doc location:** _(URL of this SOP in the wiki)_",
]
if profile in {"it", "regulated"}:
lines.append("- **Change-management ticket location:** "
"_(Jira / ServiceNow queue)_")
return "\n".join(lines)
def _build_why(meta: SOPMetadata) -> str:
lines = [
"### Why",
"",
"**Purpose:** _(one-paragraph statement of why this process "
"exists. Anchor to a business outcome, not a task.)_",
"",
"**Regulatory basis (if any):**",
"",
]
if meta.regulatory_overlay:
for ov in meta.regulatory_overlay:
lines.append(f"- {REGULATORY_PREAMBLE[ov]}")
else:
lines.append("- _(none — confirm by checking data classes "
"touched. If process touches PHI, financial controls, "
"or regulated devices, the answer is not 'none'.)_")
return "\n".join(lines)
def _build_how(meta: SOPMetadata) -> str:
lines = ["### How", ""]
lines.append("Step-by-step procedure. Each step must have a named "
"owner, expected duration, and observable success signal.")
lines.append("")
for i, step in enumerate(meta.steps_outline, start=1):
lines.append(f"**Step {i}: {step}**")
lines.append("")
lines.append("- **Owner:** _(named human or named rotation)_")
lines.append("- **Expected duration:** _(concrete number + unit)_")
lines.append("- **Success signal (observable):** _(e.g., 'IAM "
"console shows user disabled', not 'access is "
"revoked')_")
lines.append("- **Failure signal (observable):** _(what tells you "
"the step did not work)_")
lines.append("- **If step fails — rollback or escalation:** "
"_(rollback path or 'escalate to X — cannot be "
"rolled back')_")
lines.append("")
return "\n".join(lines)
def _build_how_much(meta: SOPMetadata) -> str:
mins = meta.estimated_minutes or "_(fill in)_"
cost = meta.estimated_cost_usd
cost_line = f"cost" if cost else "_(fill in)_"
return (
"### How-much\n\n"
f"- **Estimated execution time:** {mins} minutes\n"
f"- **Estimated cost per execution:** {cost_line} "
"(labor + license + third-party fees)\n"
"- **Frequency × cost = annual run-rate:** _(compute from "
"frequency + cost per execution)_\n"
)
def _build_regulated_footer() -> str:
return (
"\n---\n\n"
"## Document control (regulated profile)\n\n"
"- **Version:** 1.0\n"
"- **Effective date:** _(YYYY-MM-DD)_\n"
"- **Next review date:** _(YYYY-MM-DD — within 12 months, "
"or 90 days under HIPAA / ISO 13485)_\n"
"- **Approval signoff (named):** _(QMR / Compliance Officer)_\n"
"- **Change history:**\n\n"
"| Version | Date | Author | Change summary | Approver |\n"
"|---------|------|--------|----------------|----------|\n"
"| 1.0 | _date_ | _author_ | Initial issue | _approver_ |\n"
)
def _build_finance_footer() -> str:
return (
"\n---\n\n"
"## Controls section (finance profile)\n\n"
"- **Control objective:** _(what financial assertion this "
"controls — e.g., completeness of vendor payments)_\n"
"- **Segregation of duties:** Initiator, approver, and payer "
"must be distinct individuals.\n"
"- **Evidence retained:** _(invoice copy, approval email, "
"payment confirmation)_\n"
"- **Testing frequency:** Quarterly by Internal Audit.\n"
)
def _build_hr_footer() -> str:
return (
"\n---\n\n"
"## Privacy & sensitive-data handling (HR profile)\n\n"
"- **Data classes touched:** _(name, address, SSN/national ID, "
"compensation, medical, etc.)_\n"
"- **Lawful basis for processing:** _(employment contract, "
"legal obligation, legitimate interest, consent)_\n"
"- **Retention period:** _(per local employment law + GDPR if "
"applicable)_\n"
"- **Access restriction:** Need-to-know basis only.\n"
)
def _build_it_footer() -> str:
return (
"\n---\n\n"
"## Change management (IT profile)\n\n"
"- **Change type:** _(standard / normal / emergency)_\n"
"- **Change ticket:** _(link to Jira / ServiceNow)_\n"
"- **Rollback plan:** _(named rollback procedure)_\n"
"- **Test evidence:** _(staging validation)_\n"
"- **Communication plan:** _(who is notified pre/post change)_\n"
)
def _build_support_footer() -> str:
return (
"\n---\n\n"
"## Customer impact & escalation (support profile)\n\n"
"- **Customer impact category:** _(none / single-customer / "
"multi-customer / company-wide outage)_\n"
"- **External comms required:** _(yes/no — if yes, link "
"internal-comms cascade SOP)_\n"
"- **Escalation matrix:**\n\n"
"| Trigger | Escalate to | SLA |\n"
"|---------|-------------|-----|\n"
"| Customer-impact > 30 min | Engineering on-call | 5 min |\n"
"| Multi-customer impact | Support Lead + VP Eng | 10 min |\n"
"| External comms needed | Communications + CEO | 30 min |\n"
)
PROFILE_FOOTER = {
"ops": "",
"support": _build_support_footer(),
"finance": _build_finance_footer(),
"hr": _build_hr_footer(),
"it": _build_it_footer(),
"regulated": _build_regulated_footer(),
}
def generate_markdown(meta: SOPMetadata, profile: str) -> str:
header = (
f"# SOP: {meta.sop_name}\n\n"
f"_Profile: `{profile}` | "
f"Regulatory overlay: "
f"{meta.regulatory_overlay or 'none'}_\n\n"
"---\n\n"
"## 5W2H scaffolding\n\n"
"_(Ishikawa 1985, 5W2H method. Each section is required.)_\n"
)
body = "\n\n".join([
_build_who(meta, profile),
_build_what(meta),
_build_when(meta),
_build_where(meta, profile),
_build_why(meta),
_build_how(meta),
_build_how_much(meta),
])
footer = PROFILE_FOOTER.get(profile, "")
return header + "\n" + body + footer + "\n"
def generate_json(meta: SOPMetadata, profile: str) -> dict:
return {
"sop_name": meta.sop_name,
"profile": profile,
"metadata": asdict(meta),
"sections": {
"who": "RACI populated",
"what": f"{len(meta.inputs)} inputs / "
f"{len(meta.outputs)} outputs",
"when": meta.triggering_event,
"where": "system of record + canonical doc location",
"why": meta.regulatory_overlay or ["none"],
"how": [{"step": i + 1, "title": s}
for i, s in enumerate(meta.steps_outline)],
"how_much": {
"estimated_minutes": meta.estimated_minutes,
"estimated_cost_usd": meta.estimated_cost_usd,
},
},
}
def main(argv=None) -> int:
p = argparse.ArgumentParser(
description="Generate a 5W2H-structured SOP from JSON metadata."
)
p.add_argument("--input", "-i", type=str,
help="Path to SOP metadata JSON file.")
p.add_argument("--profile", choices=sorted(VALID_PROFILES),
default="ops",
help="Industry profile (default: ops).")
p.add_argument("--output", "-o", choices=["markdown", "json"],
default="markdown",
help="Output format (default: markdown).")
p.add_argument("--sample", action="store_true",
help="Print a sample vendor-offboarding SOP.")
args = p.parse_args(argv)
if args.sample:
data = _sample_metadata()
elif args.input:
path = Path(args.input)
if not path.exists():
print(f"ERROR: input file not found: {args.input}",
file=sys.stderr)
return 2
data = json.loads(path.read_text())
else:
print("ERROR: provide --input <metadata.json> or --sample",
file=sys.stderr)
return 2
meta = SOPMetadata(**data)
errs = meta.validate()
if errs:
print("VALIDATION ERRORS:", file=sys.stderr)
for e in errs:
print(f" - {e}", file=sys.stderr)
return 1
if args.output == "json":
print(json.dumps(generate_json(meta, args.profile), indent=2))
else:
print(generate_markdown(meta, args.profile))
return 0
if __name__ == "__main__":
sys.exit(main())