Skills that make your coding agent think like a product data scientist.
A framework for deciding if a question deserves an analysis, picking the method, and checking the result.
Statistics libraries assume you already know which question you are answering. Product management skill packs know which question matters and carry no statistics. With gallop, every question passes routing first before any query runs, and goes to one of three method buckets, on a measurement floor and under a knowledge ceiling.
| Bucket | Asks | Hands back |
|---|---|---|
| Description | What happened? | A hypothesis |
| Causation | Did this change cause that? | An effect size |
| Prediction | What will happen? Who gets what? | A forecast, a ranking, an allocation |
A question enters at the left and leaves as a decision. Measurement is a foundation because what ships changes the data. Knowledge is a ceiling because what you learn has to outlive the test that produced it.
| Skill | What it decides | Reach for it when |
|---|---|---|
routing-questions |
Whether this becomes work at all, and which skill it becomes | a product, analytics, or experimentation request first arrives, when someone asks for a deep dive or a dashboard, or before opening a query editor on any question about impact, lift, or whether something worked |
defining-metrics |
A metric turned into a computation, a source of truth, a registry entry, and a statement of how it will be gamed | defining a north-star or guardrail metric, when two dashboards disagree on the same number, when arbitrating between conflicting metric definitions, or when a report depends on a metric nobody has validated |
forming-hypotheses |
A what-happened question turned into a localised, sized hypothesis, with the floor checked first and the gap never quoted as the prize | a metric moved and someone asks what happened, when asked for a deep dive, a funnel or segment analysis, a root cause, or an opportunity size before a roadmap commitment, or when an observed gap between two groups is about to be quoted as the value of closing it |
designing-experiments |
The four choices that cannot be repaired after launch, with the MDE from the prior store | planning, powering, or pre-registering an experiment, when deciding whether a question is testable at the available traffic, or when a feature is about to ship without a flag |
reading-experiments |
Whether the result is a result: SRM, exposure, the sequential bound, CUPED, shrinkage | analysing or reviewing A/B test results, when a test looks like a winner, when someone reports a lift, or when deciding whether to ship on an experiment report |
choosing-causal-designs |
The method that matches how assignment happened, and the exit that says there is no comparison group | measuring the impact of something already rolled out, a launch, a migration, a pricing change, or a campaign that reached everyone at once |
building-models |
Whether a forecast or a repeated decision belongs to a model, validated out of time, and the holdout that measures its impact | someone asks for a churn, propensity, LTV, scoring, forecasting, uplift, recommendation or allocation model, when a model's offline accuracy is offered as evidence that something worked, or when deciding who gets an offer, a discount or an intervention |
writing-reports |
The decision rule first, the result last; the belief filed where the next question starts | a test finishes, when documenting a shipped or killed decision, when writing up a null result or a rollback, or when a question needs an entry someone can find in a year |
Each skill is a directory of markdown: SKILL.md is the procedure, the reference/ files one level down are the depth. Read them as prose, or copy them into your agent's skills folder:
git clone https://github.com/0trm/gallop
mkdir -p .claude/skills
cp -r gallop/skills/* .claude/skills/How the skills hand off to each other, and the state they share: docs/orchestration.md.
MIT. What it covers and how to contribute: docs/model.md.