Lesson 1
Introduction to Statistics
What statistics is for, how to read a number in context, ball-park estimation, and a first look at data.
Taught
What is Statistics
Statistics gives meaning to data, or better, it is the science of drawing conclusions from data. It earns its place in a great many parts of life:
- Predicting. Using what has already happened to say something about what has not: elections, weather, the price of an asset.
- Contextualising. Placing a number beside the numbers that make it meaningful, which is most of what economics, politics and business reporting consists of.
- Gathering. Designing the collection of new data so that it can support a prediction at all. Experimental design is what medicine, marketing, finance and the social sciences run on, and it is the part that decides whether the other two are worth anything.
Scrutinising a Claim
Consider a study of microplastics in bubble tea sold in the Bay Area, which reported dangerously high levels of BPA and microplastics in the drinks.
Faced with a result like this we already have ideas about where the plastic came from. Perhaps it is the straws. Perhaps it is the containers the drink is served in. Whatever the guess, it has to be scrutinised, because one possibility is always that the effect arose by chance rather than from the cause we have in mind.
Definition 1.1 (Null Hypothesis).
The null hypothesis is the claim that the pattern in the data arose by chance: that the supposed cause has no effect. A study is an attempt to gather evidence against it.
In this case the tap water used to prepare the drinks was itself already carrying a high concentration of microplastics, and tap water is used at every stage, from boiling the tapioca to making the syrup. Once that is known, the cup and the straw stop looking like the main source: the drink was made from contaminated water before it ever met a container.
Remark.
That finding was not in the headline, nor the subheadline. It appears as a passing remark inside the article. The explanation which most changes what the result means is often not the one the piece is written around.
The evidence points at the tap water rather than at the cup or the straw. Describe a measurement which would settle the question, and say what it would show if the cup really were the main source. Then say why the headline is still a fair description of the result and a misleading one.
Gathering New Data
How does one gather new information to answer a hard question? In medicine two designs are standard.
- In an observational study the researchers watch what happens without intervening. Its aim is to find a cause and its effect as they occur naturally.
- In a randomised double-blind trial the researchers assign the treatment themselves, at random, and neither the patient nor the researcher knows who is receiving the drug and who the placebo. Its aim is to remove bias.
The difference in aim matters. An observational study can tell us what happens in the world as it is; a randomised trial can tell us what the treatment does, because randomisation is what breaks the link between receiving the treatment and whatever else the patient happens to be.
A third design, familiar outside medicine, is A/B testing: users are assigned at random to group A or group B, each group is shown a different version of something, and the versions are compared on whatever response we care about. Advertising a new phone one way or another way is decided this way.
Example 1.2 (The Salk polio trial).
The 1954 trial of the Salk polio vaccine had to overcome a difficulty which had nothing to do with the vaccine. The vaccine was of the killed-virus kind, which carries a smaller risk than a live one but also, at that time, a less certain benefit, and a great many parents refused to let their children take part.
The refusal was not evenly spread. Consent to take part was more common among wealthier, more educated families, and polio incidence was higher in wealthier and more hygienic areas, since children in cleaner surroundings met the virus later, when it does more damage. So comparing children who were vaccinated with children who were not would have compared two groups differing in far more than the vaccine, and in a direction that would have made the vaccine look worse than it was.
Consider the polio trial described above.
- What makes designing a trial for this vaccine difficult, beyond the medicine?
- How would you design one? Say which of the two designs above you would use and what your design does about the fact that consent is not evenly spread across the population.
Prediction
The third use is prediction. Given the data available to us, we want the likely outcome: how risky a stock is, what it is likely to be worth, how much of a portfolio to put in it. Everything later in this course is machinery for doing that honestly.
How Statistics Goes Wrong
A statistic can be true and still mislead, if it answers a different question from the one being asked.
Example 1.3 (Two true statements).
At the trial of O. J. Simpson the defence noted that fewer than in women who are abused by their partner are later murdered by that partner. The number is correct, and it is an answer to the question “given that a woman is abused, how likely is she to be murdered by her abuser?”
That was not the question before the court. The victim had already been murdered, and the relevant figure is that about out of abused women who are murdered are murdered by their partner. Conditioning on what is already known changes the answer completely, and we shall have the machinery to say why in a later lesson.
The second failure is one of procedure rather than arithmetic.
-hacking is the practice of searching a body of data for patterns and then reporting whichever pattern is found as though it had been predicted in advance. Because a large enough search finds apparent effects in data with no effect in it, the reported finding may be a false positive produced by the search itself.
The order of operations is what protects us. Form the theory, then collect data to test it. Collecting data first and looking for patterns afterwards risks finding a pattern that is not there.
What is the difference between forming a hypothesis and then collecting data, and collecting data and then forming a hypothesis? Say what goes wrong in the second order, and give an example of a claim you would distrust for this reason.
Numbers in Context
A number on its own is not informative. It becomes informative when we compare it with a baseline, and choosing the baseline is the work.
Is $10 billion a lot of money? As a person’s net worth, yes. As a country’s annual output, no.
To talk about a number at all we put it through three filters.
- What kind of number is it? Is it a rate, a percentage, a nominal amount, an average? How was it calculated, and who is reporting it?
- What can it be compared with? Is there a natural baseline or denominator? Can it be compared with the past? What is missing that the comparison would need?
- Is it plausible? Is it surprising, or is it about what we should have expected?
Example 1.5 (Traders who lose money).
A widely repeated figure is that between and of day traders lose money.
What kind of number is it? A percentage, which raises at once the question: a percentage of what? The figure traces back to a study of Taiwanese day traders, in which a trader counts as having quit after twelve consecutive months without a trade, and profitability is estimated from a model rather than observed directly. Both of those choices affect the number. Who reports it matters too: it usually reaches the reader through an investment article rather than the paper.
What can it be compared with? One comparison is with someone who put the same money into government bonds, another with someone who bought a broad index and left it alone. Neither is a like-for-like comparison, since neither involves trading at all, and both still give a usable picture of what the activity costs.
Is it plausible? For anyone who has traded, only mildly surprising.
Example 1.6 (Student loan debt).
Education Data Initiative reports that the average federal student loan debt balance is $38,375.
What kind of number is it? An average, and immediately: an average of what, over whom? Averaged over borrowers, or over the whole population? Does it include recent graduates, or people who finished twenty years ago? Does it count medical debt, or international students? Is it a poll, a survey, or a figure taken from government and banking records?
Is it representative? The state figures spread widely around it.
| State | Average borrower debt |
|---|---|
| Maryland | $43,692 |
| California | $38,168 |
| North Dakota | $29,279 |
Source: Education Data Initiative
What can it be compared with? What the money buys, what a degree costs elsewhere in the world, the salary it leads to, the salary without it, and the interest accruing on the balance. Those last are what is missing: a debt figure without the earnings it bought is half a comparison.
Is it plausible? Set against a well-known American university charging on the order of $71,000 a year, a lifetime balance of $38,375 is a good deal lower than one might expect, which is itself a hint that the average is being taken over a wider group than one first imagines: not only recent graduates of expensive private universities, but everyone still carrying a federal balance.
Remark.
Both of the questions asked of the average, whether it represents the population and how far the extremes lie from it, are questions about variability. We formalise it later in the course.
In 2022 there were fatalities from car accidents in the United States.
- What kind of number is it?
- What can it be compared with? Give at least two baselines, and say what each one would tell you that the other would not.
- What would you have expected the number to be?
Ball-park Estimates
Most of the time the number we want for the baseline is not to hand. The fix is to estimate it: split the quantity into smaller pieces we can each guess at, guess each piece, and multiply. Working to the nearest power of ten keeps the arithmetic easy and is usually accurate enough to settle the question.
Example 1.7 (Food stamp fraud).
A case study describes an article reporting that fraud in the American food stamp programme is at an all-time high, costing taxpayers $70 million, and using this to argue for ending the programme.
Is $70 million a big number? Not until we know what to compare it with, and the natural comparison is the size of the programme. That figure is not in the article, so we estimate it:
Take million people. Take one in ten receiving, by analogy with the share of the population using food banks in the UK. For the third factor take $150 a month, roughly what a single person receives on universal credit for food, so $1800 a year. Then
about $54 billion, which makes the reported fraud
The actual budget in 2016 was , so our estimate was within of it, and the conclusion does not depend on the difference: the fraud is a fraction of one per cent either way.
Remark.
Estimating a number from limited information this way is called a Fermi problem, after Enrico Fermi, who was fond of such questions.
- Hours a year an Imperial student spends on problem sheets:
- Money spent on university tuition in the United States each year:
- The circumference of the Earth in miles. The metre was originally defined as one ten-millionth of the distance from the pole to the equator, so that distance is km and the circumference is km. At km to the mile that is miles, against a true value of about .
Give a ball-park estimate for each of the following, setting out the factors you multiply and the guess you make for each.
- How many pet dogs are there in North America?
- How many hours a year does the average American spend sitting in traffic?
- What would it cost to pay the tuition, room and board of every Imperial undergraduate for one year? What fraction of the College endowment is that, and what would you need to know to decide whether that fraction is large?
Cost-Benefit Analysis
Ball-park estimates earn their keep when a decision hangs on them. Faced with a choice between and :
- estimate the cost of each;
- estimate the benefit of each;
- choose the more attractive;
- if the stakes are high, ask how robust the analysis is.
The first two steps are Fermi problems. The fourth asks which inputs we are least sure of, and whether the conclusion survives varying them.
Example 1.9 (Cutting a fleet's fuel bill).
Fuel is one of the largest operating costs of a freight company and one of the most volatile, and it drives maintenance spend, vehicle downtime and emissions liabilities with it.
You are the capital allocation lead at a mid-sized trucking firm. You have a one-off capital budget of $10 million to cut fuel costs across a fleet of trucks the company already owns, and you must choose between two options.
- Fleet replacement. Buy new fuel-efficient tractors to replace the oldest vehicles in the fleet. A truck delivers its efficiency gain only once it has been bought outright and put into service; there is no partial version.
- Retrofit and driver programme. Fit the existing trucks with aerodynamic side skirts, and install telematics which score drivers and coach them on fuel-efficient driving. Fuel burn falls appreciably when drivers follow the coaching.
Before reading on: given only what is written above, which would you choose, and why?
The benefits. Both options promise the same kinds of thing: less volatility in the fuel bill, lower operating cost, lower maintenance spend, more vehicle uptime, less exposure to emissions liabilities, and for the new tractors a resale value at the end. Quantifying all of those separately is more than the decision needs. Money saved per year captures most of it and is the one we shall use.
The costs. These also come in kinds: the money spent upfront, the driver time spent in training rather than driving, the installation or contracting work, and the telematics subscription each year. Two of those are upfront and one is annual, so we quantify cost as
Quantifying the benefit. First estimate what one truck burns:
and at $1.20 a litre that is $42,000 of fuel per truck per year.
For the retrofit, side skirts cut burn by about , worth $2,100 a year. Telematics cut it by about , but only about of drivers follow the coaching, so call it and another $2,100. That is $4,200 saved, less the $600 annual subscription, so about $3,600 per truck per year.
For a new truck, burn falls from to L per km, a fifth less, worth $8,400 a year, and maintenance on a new truck is about $6,000 a year lower. That is about $14,400 per truck per year.
Quantifying the cost. Retrofitting costs about $3,500 per truck upfront, side skirts and telematics installation together, and $600 a year after for the subscription. A new tractor costs $150,000, against $30,000 for the old one sold, so $120,000 upfront and nothing thereafter.
The comparison. Payback is upfront cost divided by net yearly saving:
The budget buys either of two things. Retrofitting the whole fleet costs and saves a year. Replacing trucks buys of them for the full million and saves a year. So the retrofit saves more while spending under a fifth as much, and it leaves most of the budget unspent.
Robustness. The inputs we are least sure of are the price of diesel, the share of drivers who follow the coaching, the fleet size and the maintenance saving. Varying any of them within a plausible range does not reverse the conclusion, because the two options differ by a factor rather than a few per cent.
Take the fleet example above over a ten-year horizon, and suppose the budget stays at $10 million.
- Show that no value of the driver compliance rate, however low, makes the new trucks the better choice. Say which feature of the two options makes this so.
- The retrofit cost of $3,500 per truck is the input a supplier could most easily change. How high would it have to go before the new trucks won? Remember that once it rises above $20,000 the budget no longer covers the whole fleet.
Exploratory Data Analysis
Much of what probability is for is saying something about a whole group when we have only seen part of it.
Definition 1.10 (Population and Sample).
The population is the entire collection of units we want to draw a conclusion about. A sample is a subset of the population which we actually observe.
Old Faithful is a geyser which erupts roughly an hour apart. Asking how faithful Old Faithful is means asking about every eruption it will ever make, from the few hundred that have been timed.
Or take an election with voters choosing between A and B. We cannot ask all of them. Coding a vote for A as and for B as , a sample of voters splitting to gives some idea of how the whole electorate will divide.
Remark.
Nothing so far guarantees that what holds in the sample holds in the population. Making that step safely is the point of the rest of the course.
Returns
Take a financial example. The monthly equity returns for Canada were recorded over months, from February 1988 to December 1996, one number per month.
0.07 0.05 0.02 -0.04 0.08 -0.02 -0.05 0.02 0.03
0.00 0.03 0.08 -0.03 0.01 0.03 0.01 0.02 0.08
0.02 -0.02 0.00 0.01 0.02 -0.09 0.00 0.01 -0.07
0.07 0.00 0.02 -0.05 -0.04 -0.03 0.03 0.04 0.00
0.07 0.00 0.01 0.04 -0.02 0.02 0.01 -0.03 0.05
-0.02 0.00 0.01 -0.01 -0.05 -0.01 0.01 0.00 0.02
-0.02 -0.07 0.03 -0.04 0.03 -0.02 0.06 0.03 0.04
0.01 -0.01 -0.01 0.01 -0.05 0.09 -0.02 0.05 0.06
-0.05 -0.04 -0.01 0.01 -0.06 0.05 0.06 0.02 -0.01
-0.06 0.02 -0.05 0.06 0.04 0.02 0.04 0.02 0.02
0.00 0.00 -0.01 0.04 0.01 0.05 -0.01 0.02 0.04
0.02 -0.03 -0.03 0.05 0.04 0.08 0.07 -0.03
A first question worth asking of a column of numbers like this is where the middle is.
Scanning it, the values look as though they tend positive, and that is what we should expect. An asset whose returns averaged zero would go nowhere over the nine years, and one whose returns averaged negative would end the period worth less than it started, which is not what these equities did. So the middle is somewhere above zero. Reading a column one entry at a time is a poor way to locate it, though, and the first thing to do with data like this is draw it.
For an asset held over a period, beginning at value and ending at value , the return over the period is
and is the factor of return, so that .
If an asset ends the period at having returned , then , so .
Remark.
What counts as the ending value depends on the asset. For a share held and sold it is the sale price, but if the share paid dividends over the period those belong in too, since they are part of what holding it returned.
Histograms
To see a column of numbers we plot it. The values are sorted into bins, and the number of values falling in each bin is drawn as a bar.
Reading it: the centre sits somewhere between and , the shape is a single mound with the two tails dying away at roughly the same rate, and nothing sits far out on its own. The most common outcome is a small positive month.
Remark (A histogram is not a bar chart).
Histograms are for quantitative variables and bar charts are for categorical ones. The vertical axis of a bar chart may be anything; the vertical axis of a histogram is always a count or a frequency.
Remark.
The number of bars changes how smooth the picture looks. Too few and the shape disappears into three or four blocks; too many and every bar is a count of one or two, and the shape is lost in the noise.
Not every set of data makes a symmetric mound. A distribution is skewed to the left when its left tail is drawn out, which puts the bulk of the values on the right, and skewed to the right when the right tail is the long one.
The spread of the values about the centre is the variation. In the Canadian returns it could come from almost anything: credit conditions, commodity prices, a change of policy, a single piece of news. Naming the sources is not the same as measuring the spread, and measuring it is what we do later.
Data, Units and Variables
Definition 1.14 (Data, Observational Units, Variables, Distribution).
Data are the values measured, or the characteristics recorded, on the individual entities of interest.
The observational units are those individual entities on which the data are recorded.
The variables are the measured values or recorded characteristics of the observational units. A variable is quantitative if its values are numbers, such as a height or a return, and categorical if its values are labels, such as a country or a vote.
The distribution of a variable is the pattern of its values across the observational units: which are common, which are rare, and how widely they spread.
A histogram is not the only way to draw a distribution. A dot plot puts one dot per observation above its value, stacking them where values repeat, which keeps every individual observation visible in a way binning does not.
For the example above, the observational units are the individual months, the variable is the return , which is quantitative, and Figures 1.1 and 1.3 are two pictures of its distribution.
Remark.
To work out what the observational units are, it helps to ask first what variables are being measured. The units are whatever each measurement is attached to.
For each study, state the observational units and name two relevant variables, saying for each whether it is quantitative or categorical.
- A study of what makes pairs of college room-mates compatible.
- A study of the availability of vegetarian meals in the College dining halls.
A 2003 study titled Do defaults save lives?, published in Science, investigated whether the phrasing of the sign-up question affects how many people register as organ donors. The participants filled out a fake online driver’s licence application, and were asked to sign up in one of three ways: opt-in, told the default was not being a donor; opt-out, told the default was being a donor; and neutral, not told about any default and asked to choose.
- What are the observational units?
- What are the variables, and are they quantitative or categorical?
Comparing Distributions
Put a second country beside the first. Japanese equity returns were recorded over the same months, and here is each country drawn on the axes that suit it.
Set side by side like this the two pictures say almost nothing. Both show a mound near zero of roughly the same width on the page, so the eye reports that the two countries behave alike. They do not: the horizontal axes differ by a factor of two, and the bins differ with them, so equal-looking widths on the page stand for quite different widths in returns. The vertical axes can play the same trick.
Always check the axes before comparing two histograms, and if they differ, redraw. Here are the two on one scale.
Read together, and only now, the two say something a single number would not. The Canadian returns run from to ; the Japanese run from to . Japan’s tails reach much further in both directions, so it offers both the larger gains and the larger losses, while Canada clusters more tightly about its centre.
Remark.
Which is better depends on what is wanted, and the histograms do not answer it. A tighter distribution carries less risk and leaves the larger gains on the table. That trade is the subject of a later part of the course, and no shape of histogram settles it on its own.
Outliers
An outlier is a value lying far from the bulk of the data. Two cohorts of sixty students were each asked how many hours a week they spend on problem sheets; here are the answers, on one pair of axes. The figures are made up for the purpose, so that the shape being discussed is unmistakable.
The lone bar near hours in cohort B is an outlier. Everything else in that cohort lies below , so a single student sits more than twice as far out as anyone around them.
Comparing the two, the general claim is available: cohort A works more, because its centre sits at larger values, near hours against cohort B’s . The outlier does not disturb that claim, and it does move a summary. Cohort B’s mean is hours with the outlier and without it, so one student out of sixty shifts the average of the whole cohort by nearly three quarters of an hour.
An outlier is not automatically an error to be removed. It may be a mistake in recording, in which case it should go; it may be a student who genuinely works that way, in which case it is the most informative observation in the set. Deciding which requires knowing where the number came from, and that is a question about the data collection rather than about the histogram. The same applies to Japan’s two months below in Figure 1.7: they sit clear of the mound, and any summary of the Japanese returns which throws them away will understate how badly a month can go.
Time Series
The returns above were collected over time, and they came in an order. A histogram throws that order away.
Definition 1.15 (Time Series).
A sequence of observations recorded over time, in the order they occurred, is a time series. Plotting the observations against time gives the time series plot.
The vertical axis carries the return and the horizontal axis the time, so each point is one month and the line joins them in the order they occurred.
Look at Figure 1.10. Do you see a pattern? Say what a run of consecutive months above zero would have to look like before you would call it one.
It is tempting to read this plot as returning towards zero after each excursion, or as being about to turn after a run in one direction. Neither is visible here. The series is close to what repeated coin tossing looks like when it is plotted, and one of the things this course does is give us the means to say what “close to” means, rather than deciding by eye.
Not every series is like that.
Source: monthly beer production in Australia, monthly observations running from January 1956 to August 1995. The figure shows the last four and a half years.
Does the series in Figure 1.11 have a pattern? If it does, say what it is, at what points the series peaks and at what points it falls, and what could produce it. The data are Australian.
The pattern repeats once a year, and knowing it changes the question worth asking. The interesting thing about any single month is no longer how large it is, but how large it is beside the same month in other years, since comparing a peak month with a trough month tells us about the calendar rather than about the brewing.
Remark.
The vertical axis is doing as much work here as the data. A plot whose axis starts well above zero makes a small wobble look like a crisis, and one whose axis runs far past the data flattens a real movement into a straight line. Check the axes first, on your own plots as much as on other people’s.
Exercises
Decide whether each of the following is true or false, and support your answer with a ball-park estimate rather than a search.
- More than million gallons of milk are consumed in London each year.
- I could empty an ornamental fountain in one day using only a teaspoon.
- Imperial students collectively buy many thousands of textbooks each year.
Set up a Fermi-style calculation for the benefit and for the cost of distributing insecticide-treated mosquito nets, using the following.
- An unvaccinated child in sub-Saharan Africa has to malaria infections a year in early childhood.
- The infection fatality rate is about death per clinical cases.
- A net costs $2 to manufacture, stays effective for about years, and is shared by people on average.
- People sleeping under a net are half as likely to contract malaria as those without.
Give the benefit as deaths averted per person-year and the cost as dollars per person-year.
Do the same for a vaccination campaign, using the following.
- The vaccine works best if the first dose is followed by a booster years later, and children receiving both doses are to less likely to die of malaria. Assume children receiving only one dose see no change.
- Only of children who get one dose return for the booster, so about doses are given per fully vaccinated child.
- A single dose costs $10, and the vaccine must be kept cold and administered by a trained professional.
Give the cost per death averted for each of the two programmes above. Which is the better use of a fixed budget, and by roughly what factor? Are you surprised? Say which components of your calculation you are least confident in, and whether the conclusion survives varying them.
Take the state figures for average student loan debt given earlier.
- The national average is $38,375 and California’s is $38,168. Explain why these being close does not mean California is typical.
- Name a number not in the table which would change your reading of the national average, and say in which direction.
A newspaper reports that a treatment “doubles your risk” of a disease.
- What kind of number is this, and what is missing?
- What would you need to know before deciding whether to act on it?
- Give a baseline risk for which the doubling would matter, and one for which it would not.
Sketch, by hand, the histogram you would expect for each of the following, and say whether it is roughly symmetric, skewed to the left, or skewed to the right.
- The heights of adults in a large city.
- The annual incomes of households in a large city.
- The marks of a cohort on an examination that most of them found easy.
Figure 1.1 uses one bin per percentage point, and its tallest bar is the months returning .
- Redraw the histogram in your head with bins of percentage points. How many bars are there, and what happens to the shape?
- Redraw it with bins of a tenth of a percentage point. What goes wrong?
- State in one sentence what the bin width is trading off.
Check Yourself
Fresh questions on the whole lesson — none of them is worked out above. Do each on paper first; the box only tells you whether you got there.
Answers are checked in your browser, as often as you like. Nothing is sent anywhere and
nothing is kept but your own work. A formula may be written with the symbols themselves or
with ~ & | -> <-> ^, and \and, \or, \to expand as you type.
The claim that a pattern in data arose by chance, and that the supposed cause has no effect, is called the
Researchers assign a treatment at random, and neither the patient nor the researcher knows who received the drug. This design is a
Searching a body of data for patterns and then reporting whichever is found as though it had been predicted is called
A histogram is the right picture for which kind of variable?
The vertical axis of a histogram is always
A distribution whose right tail is drawn out far past the bulk of its values is
The country of birth is recorded for each student in a cohort. This variable is
In the Canadian returns data, what are the observational units?
An asset begins a period at and ends it at . What is the return?
An asset begins at and returns over the period. What is ?
Canadian monthly returns run from to and Japanese from to over the same months. Which has the greater variation?
A value lying far from the bulk of the data is called
To the nearest power of ten, how many seconds are there in a year?
Reported fraud of $70 million against a programme budget of about $54 billion is roughly
Which plot keeps the order in which the observations were recorded?