Data AnalyticsStatistics for analytical decisions

Build a statistical analysis plan before opening the results

PK
Pankit Kumar
Sr. Data Scientist at Parexel (a Goldman Sachs–backed company) · 20 September 2026 · 3 min read
Technically reviewed by Ishaan Sharma
In this article (8 sections)

Write the analysis plan before inspecting comparative outcomes. Specify the question, population, metric, uncertainty method and stopping rule so that favorable results do not quietly determine those choices afterward.

A useful plan is concrete enough for another analyst to implement and challenge. It is more than a list of statistical tests.

Begin with the estimand

The worked analysis plan compares a hypothetical simplified enquiry flow with an existing flow. Its primary quantity is the difference in seven-day enquiry probability among eligible assigned account IDs.

That wording defines a population, assignment comparison, outcome and time window. It avoids switching later between clicks, form starts and completed enquiries depending on which metric looks best.

The document is an authored teaching example, not a claim of historical preregistration or an approved live NeuraPath experiment.

Define the unit and population

The example uses stable account IDs, one assignment per account and a predeclared exclusion rule for internal/test accounts. Repeated visits do not become new independent units.

The primary analysis includes all eligible assigned units under the defined outcome contract. Filtering afterward to users who clicked a treatment-specific feature would change the population and can destroy the intended comparison.

Microsoft's pre-experiment design discussion emphasizes hypotheses, metrics, randomization units and engineering preparation. The supplied plan applies those general concerns to a specific synthetic workflow.

Make the metric executable

The plan names assignment and outcome fields, event identity, the seven-day half-open window and the rule that multiple valid enquiries count once per assigned account.

It also states that no observed event means nonconversion only when telemetry coverage is complete. Missing logs are not silently converted to zero outcomes.

These definitions should be tested with boundary fixtures before result analysis. Include an event exactly at the upper boundary, duplicate event delivery, a conflicting event version and an account with incomplete coverage.

Specify information size and stopping

The worked design targets 20,000 units with equal allocation and a fixed administrative recruitment cap. It waits for outcome maturation and a data-arrival grace period. It does not stop when the primary p-value first becomes favorable.

If the administrative cap arrives before the target count, report the actual information size and resulting uncertainty. Do not selectively extend the experiment because the current result is almost significant.

Operational incidents can still justify a stop. Record that reason and distinguish the resulting incomplete experiment from a normally completed analysis.

Declare methods and families of claims

The example specifies one confirmatory primary outcome, a pooled two-proportion z test and an unpooled normal interval as teaching approximations. It explicitly notes that these are different approximations and are not exact inverses near a threshold.

Segment analyses are exploratory. Guardrail limits are documented as operational holds, with their own definitions and review needs. Aggregate conversion counts alone cannot prove that latency or technical-error guardrails passed.

For a live study, choose methods appropriate to sparsity, clustering, repeated observations and the actual assignment design.

Preserve deviations rather than rewriting history

Version the plan and record changes after result access, including their reasons. A justified correction to a broken metric may be necessary, but it should remain visible as a deviation.

Keep original and revised definitions, affected results and sensitivity analyses. A clean final document that erases every change can misrepresent how the conclusion was selected.

Evaluate the plan before evaluating the treatment

Exercise: implement the plan's event-window calculation on a tiny synthetic assignment table. Have another person predict outcomes for each account before running the code. Then list which assumptions the aggregate ab_counts.csv fixture can verify and which require event-level or operational data.

NeuraPath's Data Analytics with Generative AI course connects statistical methods with reproducible decision processes. A strong project includes the plan that made its analysis choices explicit before the result could influence them.

Continue learning

This article is part of the Statistics for analytical decisions sequence. Use the neighbouring tasks when you need the prerequisite or the next application.

PK
Pankit Kumar
Lead Instructor, NeuraPath Academy

Pankit Kumar has 10 years in Data Science & AI, building and shipping production systems in regulated pharma and clinical environments. He is a freelance trainer at Boston Institute of Analytics, AnalytixLabs and Scaler, and has taught this material to thousands of working professionals.

This article is part of our Data Analytics with Generative AI programme — 3–4 months. The full analyst stack — Excel, SQL, Power BI and Python pipelines — then a generative-AI layer you can prove is right.

Explore Data Analytics with Generative AI
Counselling is free · no obligation

Not sure which programme fits?

Tell us your background and we will map it to the right entry point — including saying so when a cheaper programme is the better fit. A counsellor replies within one working day.