Interactive statistics

56 free interactive lessons: from mean and median to A/B tests and Bayes. Drag the charts and build intuition. No sign-up.

Module 12: Capstone

Capstone: analyzing a real dataset end to end

Time to put it all together. Before you is one dataset — 2400 users with revenue, device, A/B group and whether they purchased. We'll walk it through the whole analyst's path: from the first glance to an honest decision. The widget is step-by-step: click through stages 1–6 and watch each tool of the course come into play.

median 11mean 19

2400 users, revenue per user. The distribution is right-skewed: the mean (19) is noticeably above the median (11). Takeaway: the typical user is more honestly described by the median than by the mean.

Steps 1–2 in the widget — getting to know the data. The first thing an analyst does: look at the distribution, not compute the mean blindly. Revenue is skewed, the mean is above the median → we describe it with the median. The outliers are real (large purchases) → we don't remove them, but use the mean cautiously. This is the "Descriptive" and "Distributions" modules in action.

What it means

The analyst's main skill is not knowing the formulas but leading the data through the whole cycle: look at the shape, clean it, slice by segment, choose an appropriate method, check the assumptions and state an honest conclusion. Any single method is useless in isolation from this chain.

Take your own data (expenses, game stats, a work dataset) and walk the same route: distribution → outliers → segments → relationship/comparison → conclusion. Understanding sticks when the sliders and steps are on your own questions to real numbers.

Where it shows up

This is how the daily work of a product and data analyst is built: 80% of the time is inspecting, cleaning and exploring the data, and only then the method. Skipping EDA is the most common cause of wrong conclusions.

And this is how strong write-ups are built too: not "I know the t-test", but "here is how I led the data from question to decision, what I checked and where I honestly flagged the uncertainty".

Definitions
The analysis cycle
question → inspect and clean the data → exploration (distributions, segments) → method (test/model) → validation and interpretation → a decision with caveats.
Exploratory analysis (EDA)
the first stage: distributions, outliers, segments — before any tests or models. It determines which methods are even appropriate.
A decision with caveats
a conclusion that honestly states the uncertainty, assumptions and limits of applicability, not a bare number.
Next → Congratulations — that's the whole path! Next → the second part of the site: metric hierarchies
Congrats — you made it to the end! 🎉

This course is built by the DataSlice Telegram channel (in Russian). Subscribe so you don’t lose it.

Open the Telegram channel