BACK Mascot image.

MA0 2

Introduction to Algorithms and Numerical Analysis

Every lesson so far, in one document · 5 chapters

back to the contents

Lesson 1

Vectors and Groups

Taught

Vectors

Algebra deals with variables and equations, geometry with points, lines, planes and solids, and the two are closely linked. Take the equation y=ax+by = ax + b. On its own it is an algebraic object, a relation between two unknowns; but the collection of points whose coordinates satisfy it traces out a straight line, and the two constants in it give two features of that line. The number aa is the slope, the change in yy against the change in xx, and bb is the height at which the line meets the vertical axis.

xyb1a0
Figure 1.1. The line y=ax+by = ax + b, meeting the vertical axis at height bb and rising by aa over every unit travelled to the right.

So the constants in the equation of a line, which are algebraic data, correspond to the slope and the intercept, which are geometric data. This rests on identifying the points of the plane with pairs of numbers, and pairs are objects MA01 has already built.

Coordinates

An ordered pair (a,b)(a, b) is a pair in which one element is marked as coming first, so that

(a,b)=(a′,b′)if and only ifa=a′ and b=b′.(a, b) = (a', b') \quad \text{if and only if} \quad a = a' \text{ and } b = b'.

Given two sets AA and BB, the Cartesian product A×BA \times B is the set of all ordered pairs whose first entry comes from AA and whose second comes from BB,

A×B={(a,b)∣a∈A,  b∈B},A \times B = \{(a, b) \mid a \in A, \; b \in B\},

and when A=BA = B we write A2A^2 for it. More generally, for n∈Nn \in \mathbb{N} the nn-fold product of sets A1,…,AnA_1, \ldots, A_n consists of the nn-tuples with one entry drawn from each,

A1×A2×⋯×An={(a1,a2,…,an)∣a1∈A1,  a2∈A2,  …,  an∈An},A_1 \times A_2 \times \cdots \times A_n = \{(a_1, a_2, \ldots, a_n) \mid a_1 \in A_1, \; a_2 \in A_2, \; \ldots, \; a_n \in A_n\},

written AnA^n as a Cartesian power when all the factors agree. Throughout this course N={1,2,3,…}\mathbb{N} = \{1, 2, 3, \ldots\}.

The example we care about is R2=R×R\mathbb{R}^2 = \mathbb{R} \times \mathbb{R}, whose elements are the pairs of real numbers. Fixing an origin and two perpendicular axes in the plane identifies each point with exactly one such pair.

Remark (Descartes).

The identification of the points of the plane with R2\mathbb{R}^2 is due to René Descartes, and the product carries his name because of it. The points (a,b)(a, b) and (b,a)(b, a) are different unless a=ba = b, which is why ordered pairs are needed rather than two-element sets. Given a point PP of the plane matched with the pair (a,b)(a, b), we call aa the xx-coordinate and bb the yy-coordinate of PP.

Once points are pairs, two different kinds of quantity are in play. A scalar is a quantity settled by a single number, and for us that number will be real. Temperature, mass and altitude are scalars. Force and velocity are not: a force has a direction as well as a size, and both are needed to name it.

Remark (Magnitude and direction).

A vector is often introduced as a quantity carrying both a magnitude and a direction. That description says what vectors are for, but neither magnitude nor direction has been given a meaning yet, so it cannot be calculated with.

Take a pair (a,b)∈R2(a, b) \in \mathbb{R}^2 and the origin O=(0,0)O = (0, 0). Instead of drawing the single point P=(a,b)P = (a, b) we may draw the arrow that runs from OO to PP, reached by travelling aa units along the horizontal axis and bb units along the vertical one.

xyOabP = (a, b)
Figure 1.2. The point P=(a,b)P = (a, b) and the arrow from the origin that represents it.

Because the arrow and the point determine one another, we treat them as the same object and use the words interchangeably.

Definition 1.1 (Vector).

Let n∈Nn \in \mathbb{N}. An nn-dimensional vector, or nn-vector, is an element x\mathbf{x} of Rn\mathbb{R}^n. It may be written as a column vector

x=(x1x2⋮xn)\mathbf{x} = \begin{pmatrix} x_1 \\ x_2 \\ \vdots \\ x_n \end{pmatrix}

or as a row vector x=(x1,x2,…,xn)\mathbf{x} = (x_1, x_2, \ldots, x_n). For 1⩽i⩽n1 \leqslant i \leqslant n the real number xix_i is the iith component of x\mathbf{x}. The set Rn\mathbb{R}^n of all such vectors is called nn-dimensional space.

Since points and vectors are the same objects here, the words component and coordinate are used interchangeably as well. For n=2n = 2 we usually name the components xx and yy, so that

R2={(x,y)∣x,y∈R}\mathbb{R}^2 = \{(x, y) \mid x, y \in \mathbb{R}\}

is the xyxy-plane; for n=3n = 3 we name them xx, yy and zz, and call R3\mathbb{R}^3 the xyzxyz-space.

Remark (Notation for vectors).

We write vectors in bold, u\mathbf{u} and v\mathbf{v}, and write OA→\overrightarrow{OA} for the vector running from OO to AA. Other texts write u⃗\vec{u} or u‾\underline{u} for the same thing, and a vector of length one often gets a hat, as in x^\hat{x} or e^1\hat{e}_1. Any of these will do, provided scalars and vectors are never written the same way.

Example 1.2 (Vectors in R2\mathbb{R}^2 and R3\mathbb{R}^3).

(23)∈R2,(7−101)∈R3,(2.53/104.001)∈R3.\begin{pmatrix} 2 \\ 3 \end{pmatrix} \in \mathbb{R}^2, \qquad \begin{pmatrix} 7 \\ -10 \\ 1 \end{pmatrix} \in \mathbb{R}^3, \qquad \begin{pmatrix} 2.5 \\ 3/10 \\ 4.001 \end{pmatrix} \in \mathbb{R}^3.

Remark (Rows against columns).

A row vector and a column vector with the same entries are the same element of Rn\mathbb{R}^n, and for everything in this chapter the shape is only a matter of how the thing is printed. It stops being a matter of printing as soon as vectors are multiplied by matrices, since the rules for that multiplication treat a row of nn entries and a column of nn entries as objects of different shapes, and the two cannot then be swapped. We will keep to columns whenever the shape could matter.

Definition 1.3 (Zero vector).

The zero vector of Rn\mathbb{R}^n is the vector all of whose components are 00,

0=(0,0,…,0),\mathbf{0} = (0, 0, \ldots, 0),

also called the null vector. It is the vector OO→\overrightarrow{OO} representing the origin.

The vectors of the form (0,…,0,xi,0,…,0)(0, \ldots, 0, x_i, 0, \ldots, 0), with every component but the iith equal to zero, are exactly the points of the xix_i-axis.

Vector Algebra

Definition 1.4 (Addition and subtraction).

Let u=(u1,u2,…,un)\mathbf{u} = (u_1, u_2, \ldots, u_n) and v=(v1,v2,…,vn)\mathbf{v} = (v_1, v_2, \ldots, v_n) be vectors in Rn\mathbb{R}^n. Their sum and difference are formed component by component:

u+v=(u1+v1,  u2+v2,  …,  un+vn),u−v=(u1−v1,  u2−v2,  …,  un−vn).\begin{aligned} \mathbf{u} + \mathbf{v} &= (u_1 + v_1, \; u_2 + v_2, \; \ldots, \; u_n + v_n), \\ \mathbf{u} - \mathbf{v} &= (u_1 - v_1, \; u_2 - v_2, \; \ldots, \; u_n - v_n). \end{aligned}

Two vectors can be added only when they have the same number of components, since otherwise some component of the answer has nothing to be built from. A vector in R3\mathbb{R}^3 and a vector in R2\mathbb{R}^2 have no sum.

Example 1.5 (A sum in R3\mathbb{R}^3).

(365)+(781)=(10146).\begin{pmatrix} 3 \\ 6 \\ 5 \end{pmatrix} + \begin{pmatrix} 7 \\ 8 \\ 1 \end{pmatrix} = \begin{pmatrix} 10 \\ 14 \\ 6 \end{pmatrix}.

Read as arrows, u+v\mathbf{u} + \mathbf{v} is what we reach by travelling along u\mathbf{u} and then travelling along v\mathbf{v} from wherever that leaves us. Travelling along v\mathbf{v} first and u\mathbf{u} second lands in the same place, which is the picture behind u+v=v+u\mathbf{u} + \mathbf{v} = \mathbf{v} + \mathbf{u}: the two routes are the two ways round a parallelogram.

Ouvu + vOuvv − u
Figure 1.3. The sum of two vectors is the diagonal of the parallelogram they span; the difference v−u\mathbf{v} - \mathbf{u} is the arrow carrying the tip of u\mathbf{u} to the tip of v\mathbf{v}.

The right-hand panel shows the difference of two vectors. The vector v−u\mathbf{v} - \mathbf{u} is the one that translates the point with position vector u\mathbf{u} to the point with position vector v\mathbf{v}, since adding it to u\mathbf{u} returns v\mathbf{v}. For two points PP and QQ this is written

PQ→=OQ→−OP→.\overrightarrow{PQ} = \overrightarrow{OQ} - \overrightarrow{OP}.

Definition 1.6 (Scalar multiplication).

Let u=(u1,u2,…,un)∈Rn\mathbf{u} = (u_1, u_2, \ldots, u_n) \in \mathbb{R}^n and let λ∈R\lambda \in \mathbb{R}. The scalar multiple λu\lambda\mathbf{u} is

λu=(λu1,  λu2,  …,  λun).\lambda\mathbf{u} = (\lambda u_1, \; \lambda u_2, \; \ldots, \; \lambda u_n).

We write −u-\mathbf{u} for (−1)u=(−u1,−u2,…,−un)(-1)\mathbf{u} = (-u_1, -u_2, \ldots, -u_n).

When λ\lambda is a positive integer, λu\lambda\mathbf{u} is the translation obtained by translating λ\lambda times by u\mathbf{u}. Letting λ\lambda range over all of R\mathbb{R}, the points λu\lambda\mathbf{u} trace out the straight line through the origin and the point u\mathbf{u}. Translating by −u-\mathbf{u} undoes a translation by u\mathbf{u}.

O−u½uu2u
Figure 1.4. The scalar multiples of a vector u≠0\mathbf{u} \neq \mathbf{0} fill the line through the origin and u\mathbf{u}.

Example 1.7 (A scalar multiple).

4⋅(279)=(82836).4 \cdot \begin{pmatrix} 2 \\ 7 \\ 9 \end{pmatrix} = \begin{pmatrix} 8 \\ 28 \\ 36 \end{pmatrix}.

The product λ⋅u\lambda \cdot \mathbf{u} is written λu\lambda\mathbf{u} from here on, the dot being dropped as it is for the product of two reals.

With both operations available we can build one vector out of several others.

Definition 1.8 (Linear combination).

Let v1,…,vk∈Rn\mathbf{v}_1, \ldots, \mathbf{v}_k \in \mathbb{R}^n and λ1,…,λk∈R\lambda_1, \ldots, \lambda_k \in \mathbb{R}. The vector

λ1v1+λ2v2+⋯+λkvk\lambda_1\mathbf{v}_1 + \lambda_2\mathbf{v}_2 + \cdots + \lambda_k\mathbf{v}_k

is a linear combination of v1,…,vk\mathbf{v}_1, \ldots, \mathbf{v}_k with coefficients λ1,…,λk\lambda_1, \ldots, \lambda_k. A sum of this shape is abbreviated

∑i=1kλivi,\sum_{i=1}^{k} \lambda_i \mathbf{v}_i ,

the symbol ∑\textstyle\sum instructing us to add the terms obtained as the index ii runs through the integers from the value below the symbol to the value above it.

Definition 1.9 (Standard basis).

Let n∈Nn \in \mathbb{N}. For 1⩽i⩽n1 \leqslant i \leqslant n let e^i∈Rn\hat{e}_i \in \mathbb{R}^n be the vector whose iith component is 11 and whose other components are 00, so that

e^1=(1,0,…,0,0),e^2=(0,1,…,0,0),…,e^n=(0,0,…,0,1).\hat{e}_1 = (1, 0, \ldots, 0, 0), \quad \hat{e}_2 = (0, 1, \ldots, 0, 0), \quad \ldots, \quad \hat{e}_n = (0, 0, \ldots, 0, 1).

These nn vectors form the standard basis, or canonical basis, of Rn\mathbb{R}^n.

Proposition 1.10 (Expansion in the standard basis).

Let n∈Nn \in \mathbb{N} and let x∈Rn\mathbf{x} \in \mathbb{R}^n. There is exactly one nn-tuple (λ1,…,λn)(\lambda_1, \ldots, \lambda_n) of real numbers with

x=∑i=1nλie^i,\mathbf{x} = \sum_{i=1}^{n} \lambda_i \hat{e}_i,

namely the tuple of components of x\mathbf{x} itself.

Discussion.

The statement is an existence and uniqueness claim about the tuple of coefficients, so the proof has two halves. Both are settled by looking at one component at a time, because addition and scalar multiplication were defined component by component: the jjth component of ∑iλie^i\sum_i \lambda_i \hat{e}_i is ∑iλi(e^i)j\sum_i \lambda_i (\hat{e}_i)_j, and by the definition of the standard basis every term of that sum vanishes except the one with i=ji = j. So the jjth component of the combination is λj\lambda_j, whatever the coefficients were. Existence then follows by taking λj=xj\lambda_j = x_j, and uniqueness follows because any tuple that works must have λj\lambda_j equal to the jjth component of x\mathbf{x}, which leaves no freedom.

Proof.

Let λ1,…,λn\lambda_1, \ldots, \lambda_n be any real numbers and fix jj with 1⩽j⩽n1 \leqslant j \leqslant n. By the definition of scalar multiplication the jjth component of λie^i\lambda_i \hat{e}_i is λi(e^i)j\lambda_i (\hat{e}_i)_j, which is λi\lambda_i when i=ji = j and 00 otherwise; and by the definition of addition the jjth component of a sum is the sum of the jjth components. Hence

(∑i=1nλie^i)j=∑i=1nλi(e^i)j=λj.\Bigl( \sum_{i=1}^{n} \lambda_i \hat{e}_i \Bigr)_j = \sum_{i=1}^{n} \lambda_i (\hat{e}_i)_j = \lambda_j.

Now write x=(x1,…,xn)\mathbf{x} = (x_1, \ldots, x_n). Taking λi=xi\lambda_i = x_i for each ii, the displayed identity says that ∑ixie^i\sum_i x_i \hat{e}_i has jjth component xjx_j for every jj, so it equals x\mathbf{x}; this proves existence. If (λ1,…,λn)(\lambda_1, \ldots, \lambda_n) is any tuple with x=∑iλie^i\mathbf{x} = \sum_i \lambda_i \hat{e}_i, then comparing jjth components in that equation and using the identity again gives xj=λjx_j = \lambda_j for every jj; this proves uniqueness.

Remark (The basis in two and three dimensions).

For n=2n = 2 the standard basis is (1,0)(1, 0) and (0,1)(0, 1), commonly written i^\hat{i} and j^\hat{j}, and the proposition says that every u=(x,y)\mathbf{u} = (x, y) is xi^+yj^x\hat{i} + y\hat{j} and is so in only one way. For n=3n = 3 the basis is written i^\hat{i}, j^\hat{j}, k^\hat{k}.

Definition 1.11 (Norm of a vector).

Let u=(u1,u2,…,un)∈Rn\mathbf{u} = (u_1, u_2, \ldots, u_n) \in \mathbb{R}^n. The norm, or magnitude, or length, of u\mathbf{u} is the real number

∥u∥=u12+u22+⋯+un2=∑i=1nui2.\lVert \mathbf{u} \rVert = \sqrt{u_1^2 + u_2^2 + \cdots + u_n^2} = \sqrt{\sum_{i=1}^{n} u_i^2}.

This is also called the Euclidean norm. A vector with ∥u∥=1\lVert \mathbf{u} \rVert = 1 is a unit vector.

Remark (Bars and double bars).

Many texts write ∣u∣|\mathbf{u}| for the norm. We keep the single bars for the absolute value of a real number, which appears in the last chapter of this lesson and again throughout the course, and reserve the double bars for vectors. The choice only affects legibility, but a formula such as ∥λu∥=∣λ∣ ∥u∥\lVert \lambda \mathbf{u} \rVert = |\lambda| \, \lVert \mathbf{u} \rVert is a good deal easier to read when the two are told apart.

Example 1.12 (A norm in R3\mathbb{R}^3).

Let A=(3,4,0)A = (3, 4, 0) and let a=OA→\mathbf{a} = \overrightarrow{OA}. Then

∥a∥=32+42+02=25=5.\lVert \mathbf{a} \rVert = \sqrt{3^2 + 4^2 + 0^2} = \sqrt{25} = 5.

Proposition 1.13 (Properties of the norm).

Let x∈Rn\mathbf{x} \in \mathbb{R}^n and λ∈R\lambda \in \mathbb{R}. Then

  1. ∥x∥⩾0\lVert \mathbf{x} \rVert \geqslant 0;
  2. ∥x∥=0\lVert \mathbf{x} \rVert = 0 if and only if x=0\mathbf{x} = \mathbf{0};
  3. ∥λx∥=∣λ∣ ∥x∥\lVert \lambda\mathbf{x} \rVert = |\lambda| \, \lVert \mathbf{x} \rVert.

Discussion.

All three parts are about the number under the square root, which is a sum of squares of real numbers, so each is settled by a fact about real squares rather than by anything about vectors. For the first, every square is non-negative and so is the sum, and the square root of a non-negative number is non-negative by definition. The second is a biconditional; the reverse direction is a computation, and the forward one uses that a sum of non-negative terms vanishes only if every term does, so each xi2x_i^2 is 00 and hence each xix_i is. The third pulls the constant out of the sum and out of the root, where the identity λ2=∣λ∣\sqrt{\lambda^2} = |\lambda| supplies the absolute value; it is the reason the absolute value appears at all, since the root is the non-negative one and λ\lambda need not be.

Proof.

Write x=(x1,…,xn)\mathbf{x} = (x_1, \ldots, x_n) and S=∑i=1nxi2S = \sum_{i=1}^{n} x_i^2.

For the first part, each xi2⩾0x_i^2 \geqslant 0, so S⩾0S \geqslant 0, and ∥x∥=S⩾0\lVert \mathbf{x} \rVert = \sqrt{S} \geqslant 0 since the square root of a non-negative real is taken to be non-negative.

For the second, if x=0\mathbf{x} = \mathbf{0} then every xi=0x_i = 0, so S=0S = 0 and ∥x∥=0\lVert \mathbf{x} \rVert = 0. Conversely, suppose ∥x∥=0\lVert \mathbf{x} \rVert = 0. Squaring gives S=0S = 0; since SS is a sum of terms each of which is at least 00, no term can be strictly positive, so xi2=0x_i^2 = 0 and hence xi=0x_i = 0 for every ii. Thus x=0\mathbf{x} = \mathbf{0}.

For the third, the definition of scalar multiplication makes the components of λx\lambda\mathbf{x} the numbers λxi\lambda x_i, so

∥λx∥=∑i=1nλ2xi2=λ2 S=λ2 S=∣λ∣ ∥x∥.\lVert \lambda\mathbf{x} \rVert = \sqrt{\sum_{i=1}^{n} \lambda^2 x_i^2} = \sqrt{\lambda^2 \, S} = \sqrt{\lambda^2} \, \sqrt{S} = |\lambda| \, \lVert \mathbf{x} \rVert.

Definition 1.14 (Euclidean distance).

Let u,v∈Rn\mathbf{u}, \mathbf{v} \in \mathbb{R}^n. The Euclidean distance between u\mathbf{u} and v\mathbf{v}, taken as position vectors, is ∥v−u∥\lVert \mathbf{v} - \mathbf{u} \rVert; in components,

∥v−u∥=∑i=1n(vi−ui)2.\lVert \mathbf{v} - \mathbf{u} \rVert = \sqrt{\sum_{i=1}^{n} (v_i - u_i)^2}.

The distance is symmetric in its two arguments, since u−v=−(v−u)\mathbf{u} - \mathbf{v} = -(\mathbf{v} - \mathbf{u}) and the third part of the last proposition, with λ=−1\lambda = -1, gives the two vectors the same norm.

Ouv‖u‖‖v‖‖v − u‖
Figure 1.5. The norms of two vectors and the distance between the points they represent.

Problem 1.1.

Let u,v,w∈Rn\mathbf{u}, \mathbf{v}, \mathbf{w} \in \mathbb{R}^n and λ,μ∈R\lambda, \mu \in \mathbb{R}. Prove that u+v=v+u\mathbf{u} + \mathbf{v} = \mathbf{v} + \mathbf{u}, that (u+v)+w=u+(v+w)(\mathbf{u} + \mathbf{v}) + \mathbf{w} = \mathbf{u} + (\mathbf{v} + \mathbf{w}), that u+0=u\mathbf{u} + \mathbf{0} = \mathbf{u} and u+(−u)=0\mathbf{u} + (-\mathbf{u}) = \mathbf{0}, and that

λ(u+v)=λu+λv,(λ+μ)u=λu+μu,λ(μu)=(λμ)u.\lambda(\mathbf{u} + \mathbf{v}) = \lambda\mathbf{u} + \lambda\mathbf{v}, \qquad (\lambda + \mu)\mathbf{u} = \lambda\mathbf{u} + \mu\mathbf{u}, \qquad \lambda(\mu\mathbf{u}) = (\lambda\mu)\mathbf{u}.

Problem 1.2.

Let u∈Rn\mathbf{u} \in \mathbb{R}^n with u≠0\mathbf{u} \neq \mathbf{0}. Show that u/∥u∥\mathbf{u} / \lVert \mathbf{u} \rVert is a unit vector, and that it is the only unit vector of the form λu\lambda\mathbf{u} with λ>0\lambda > 0.

The Scalar Product

Definition 1.15 (Scalar product).

Let u=(u1,…,un)\mathbf{u} = (u_1, \ldots, u_n) and v=(v1,…,vn)\mathbf{v} = (v_1, \ldots, v_n) be vectors in Rn\mathbb{R}^n. Their scalar product, also called the dot product or the Euclidean inner product, is the real number

u⋅v=∑i=1nuivi.\mathbf{u} \cdot \mathbf{v} = \sum_{i=1}^{n} u_i v_i.

It is also written ⟨u,v⟩\langle \mathbf{u}, \mathbf{v} \rangle.

The scalar product of two vectors is a scalar, not a vector. Comparing the definition with that of the norm gives

u⋅u=∑i=1nui2=∥u∥2,so∥u∥=u⋅u.\mathbf{u} \cdot \mathbf{u} = \sum_{i=1}^{n} u_i^2 = \lVert \mathbf{u} \rVert^2, \qquad \text{so} \qquad \lVert \mathbf{u} \rVert = \sqrt{\mathbf{u} \cdot \mathbf{u}}.

Example 1.16 (A scalar product in R3\mathbb{R}^3).

(253)⋅(719)=2⋅7+5⋅1+3⋅9=46.\begin{pmatrix} 2 \\ 5 \\ 3 \end{pmatrix} \cdot \begin{pmatrix} 7 \\ 1 \\ 9 \end{pmatrix} = 2 \cdot 7 + 5 \cdot 1 + 3 \cdot 9 = 46.

Remark (A glance ahead at the transpose).

A matrix has a transpose, obtained by exchanging its rows and its columns. A column vector is a matrix of one column, so its transpose uT\mathbf{u}^{\mathsf{T}} is a row, and once matrices are multiplied the product uTv\mathbf{u}^{\mathsf{T}} \mathbf{v} is a matrix with one row and one column whose single entry is

u1v1+⋯+unvn=u⋅v.u_1 v_1 + \cdots + u_n v_n = \mathbf{u} \cdot \mathbf{v}.

So the dot product is the matrix product uTv\mathbf{u}^{\mathsf{T}} \mathbf{v}, which is why the row and column shapes of the earlier remark are kept apart.

Remark (Inner products in general).

Later modules replace Rn\mathbb{R}^n by a vector space VV over a field FF and keep the product as a map

⟨⋅,⋅⟩:V×V→F,(u,v)↦⟨u,v⟩,\langle \cdot, \cdot \rangle : V \times V \to F, \qquad (\mathbf{u}, \mathbf{v}) \mapsto \langle \mathbf{u}, \mathbf{v} \rangle,

subject to the properties proved in the next proposition, which are there taken as axioms. Fields and vector spaces are defined in the last chapter of this lesson. The proofs that follow use only those properties, so they hold for any inner product.

Proposition 1.17 (The scalar product is a symmetric bilinear form).

Let u,v,w∈Rn\mathbf{u}, \mathbf{v}, \mathbf{w} \in \mathbb{R}^n and λ,μ∈R\lambda, \mu \in \mathbb{R}. Then

  1. u⋅v=v⋅u\mathbf{u} \cdot \mathbf{v} = \mathbf{v} \cdot \mathbf{u};
  2. (λu+μw)⋅v=λ(u⋅v)+μ(w⋅v)(\lambda\mathbf{u} + \mu\mathbf{w}) \cdot \mathbf{v} = \lambda (\mathbf{u} \cdot \mathbf{v}) + \mu (\mathbf{w} \cdot \mathbf{v});
  3. u⋅(λv+μw)=λ(u⋅v)+μ(u⋅w)\mathbf{u} \cdot (\lambda\mathbf{v} + \mu\mathbf{w}) = \lambda (\mathbf{u} \cdot \mathbf{v}) + \mu (\mathbf{u} \cdot \mathbf{w});
  4. u⋅u=∥u∥2⩾0\mathbf{u} \cdot \mathbf{u} = \lVert \mathbf{u} \rVert^2 \geqslant 0, with u⋅u=0\mathbf{u} \cdot \mathbf{u} = 0 if and only if u=0\mathbf{u} = \mathbf{0}.

Discussion.

Parts 1 to 3 are identities between two real numbers, each of which is a sum over the components, so each is proved by writing both sides as such a sum and reconciling them term by term with the arithmetic of R\mathbb{R}. Symmetry needs only that uivi=viuiu_i v_i = v_i u_i. Linearity in the first argument needs the definitions of the sum and the scalar multiple to compute the iith component of λu+μw\lambda\mathbf{u} + \mu\mathbf{w}, and then the distributive law to split the sum in two. Part 3 does not need a separate argument: once symmetry is available, the two arguments may be exchanged, part 2 applied, and the arguments exchanged back. Part 4 is not new either, since u⋅u=∥u∥2\mathbf{u} \cdot \mathbf{u} = \lVert \mathbf{u} \rVert^2 was read off the two definitions above, and the remaining claims are the first two parts of the proposition on the norm.

Proof.

Write u=(u1,…,un)\mathbf{u} = (u_1, \ldots, u_n), v=(v1,…,vn)\mathbf{v} = (v_1, \ldots, v_n) and w=(w1,…,wn)\mathbf{w} = (w_1, \ldots, w_n).

Multiplication of reals is commutative, so

u⋅v=∑i=1nuivi=∑i=1nviui=v⋅u,\mathbf{u} \cdot \mathbf{v} = \sum_{i=1}^{n} u_i v_i = \sum_{i=1}^{n} v_i u_i = \mathbf{v} \cdot \mathbf{u},

which is the first part. For the second, the iith component of λu+μw\lambda\mathbf{u} + \mu\mathbf{w} is λui+μwi\lambda u_i + \mu w_i, so

(λu+μw)⋅v=∑i=1n(λui+μwi)vi=∑i=1n(λuivi+μwivi)=λ∑i=1nuivi+μ∑i=1nwivi=λ(u⋅v)+μ(w⋅v).\begin{aligned} (\lambda\mathbf{u} + \mu\mathbf{w}) \cdot \mathbf{v} &= \sum_{i=1}^{n} (\lambda u_i + \mu w_i) v_i = \sum_{i=1}^{n} \bigl( \lambda u_i v_i + \mu w_i v_i \bigr) \\ &= \lambda \sum_{i=1}^{n} u_i v_i + \mu \sum_{i=1}^{n} w_i v_i = \lambda (\mathbf{u} \cdot \mathbf{v}) + \mu (\mathbf{w} \cdot \mathbf{v}). \end{aligned}

For the third, we use the symmetry just proved twice, with the second part in between:

u⋅(λv+μw)=(λv+μw)⋅u=λ(v⋅u)+μ(w⋅u)=λ(u⋅v)+μ(u⋅w).\mathbf{u} \cdot (\lambda\mathbf{v} + \mu\mathbf{w}) = (\lambda\mathbf{v} + \mu\mathbf{w}) \cdot \mathbf{u} = \lambda (\mathbf{v} \cdot \mathbf{u}) + \mu (\mathbf{w} \cdot \mathbf{u}) = \lambda (\mathbf{u} \cdot \mathbf{v}) + \mu (\mathbf{u} \cdot \mathbf{w}).

For the fourth, the identity u⋅u=∥u∥2\mathbf{u} \cdot \mathbf{u} = \lVert \mathbf{u} \rVert^2 holds because both sides are ∑i=1nui2\sum_{i=1}^{n} u_i^2. That quantity is non-negative and vanishes exactly when u=0\mathbf{u} = \mathbf{0}, by the first two parts of the proposition on the norm.

Proposition 1.18 (Cauchy–Schwarz inequality).

Let u,v∈Rn\mathbf{u}, \mathbf{v} \in \mathbb{R}^n. Then

∣u⋅v∣⩽∥u∥ ∥v∥,|\mathbf{u} \cdot \mathbf{v}| \leqslant \lVert \mathbf{u} \rVert \, \lVert \mathbf{v} \rVert,

with equality if and only if one of u\mathbf{u} and v\mathbf{v} is a scalar multiple of the other. Two vectors standing in that relation are called linearly dependent.

Discussion.

The proof starts from a quantity known to be non-negative, by part 4 of the last proposition: ∥u+tv∥2⩾0\lVert \mathbf{u} + t\mathbf{v} \rVert^2 \geqslant 0 for every real tt. Expanding that norm with bilinearity turns it into

∥u∥2+2t (u⋅v)+t2∥v∥2,\lVert \mathbf{u} \rVert^2 + 2t \, (\mathbf{u} \cdot \mathbf{v}) + t^2 \lVert \mathbf{v} \rVert^2,

a quadratic in tt whose coefficients are the three quantities the statement mentions. A quadratic with positive leading coefficient that is never negative has at most one real root, so its discriminant is at most 00, and that discriminant is exactly 4(u⋅v)2−4∥u∥2∥v∥24(\mathbf{u} \cdot \mathbf{v})^2 - 4\lVert \mathbf{u} \rVert^2 \lVert \mathbf{v} \rVert^2. The case v=0\mathbf{v} = \mathbf{0} has to be handled on its own, since then the leading coefficient vanishes and the expression is not a quadratic; both sides of the inequality are 00 there. For the equality case, note that the discriminant is 00 exactly when the quadratic has a root t0t_0, and by part 4 again a root means u+t0v=0\mathbf{u} + t_0\mathbf{v} = \mathbf{0}.

Problem 1.3.

Prove the Cauchy–Schwarz inequality, including the description of the case of equality.

Remark (Where the inequality lives).

The proof asked for above uses only the four properties of the previous proposition, so the inequality holds for every inner product, not only for the dot product on Rn\mathbb{R}^n.

Proposition 1.19 (Triangle inequality).

Let u,v∈Rn\mathbf{u}, \mathbf{v} \in \mathbb{R}^n. Then

∥u+v∥⩽∥u∥+∥v∥.\lVert \mathbf{u} + \mathbf{v} \rVert \leqslant \lVert \mathbf{u} \rVert + \lVert \mathbf{v} \rVert.

Discussion.

Both sides are non-negative, so it is enough to compare their squares: the left-hand square is (u+v)⋅(u+v)(\mathbf{u} + \mathbf{v}) \cdot (\mathbf{u} + \mathbf{v}), which bilinearity expands into ∥u∥2+2(u⋅v)+∥v∥2\lVert \mathbf{u} \rVert^2 + 2(\mathbf{u} \cdot \mathbf{v}) + \lVert \mathbf{v} \rVert^2, while the right-hand square is ∥u∥2+2∥u∥∥v∥+∥v∥2\lVert \mathbf{u} \rVert^2 + 2\lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert + \lVert \mathbf{v} \rVert^2. The two differ only in the middle term, so the claim reduces to u⋅v⩽∥u∥∥v∥\mathbf{u} \cdot \mathbf{v} \leqslant \lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert, which is Cauchy–Schwarz together with the fact that a real number is at most its own absolute value.

Proof.

By the identity ∥x∥2=x⋅x\lVert \mathbf{x} \rVert^2 = \mathbf{x} \cdot \mathbf{x} and bilinearity,

∥u+v∥2=(u+v)⋅(u+v)=∥u∥2+2(u⋅v)+∥v∥2.\lVert \mathbf{u} + \mathbf{v} \rVert^2 = (\mathbf{u} + \mathbf{v}) \cdot (\mathbf{u} + \mathbf{v}) = \lVert \mathbf{u} \rVert^2 + 2 (\mathbf{u} \cdot \mathbf{v}) + \lVert \mathbf{v} \rVert^2.

Every real number is at most its absolute value, so u⋅v⩽∣u⋅v∣\mathbf{u} \cdot \mathbf{v} \leqslant |\mathbf{u} \cdot \mathbf{v}|, and the Cauchy–Schwarz inequality bounds the latter by ∥u∥∥v∥\lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert. Hence

∥u+v∥2⩽∥u∥2+2∥u∥∥v∥+∥v∥2=(∥u∥+∥v∥)2.\lVert \mathbf{u} + \mathbf{v} \rVert^2 \leqslant \lVert \mathbf{u} \rVert^2 + 2 \lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert + \lVert \mathbf{v} \rVert^2 = \bigl( \lVert \mathbf{u} \rVert + \lVert \mathbf{v} \rVert \bigr)^2.

Both ∥u+v∥\lVert \mathbf{u} + \mathbf{v} \rVert and ∥u∥+∥v∥\lVert \mathbf{u} \rVert + \lVert \mathbf{v} \rVert are non-negative, and for non-negative reals aa and bb the inequality a2⩽b2a^2 \leqslant b^2 gives a⩽ba \leqslant b. Therefore ∥u+v∥⩽∥u∥+∥v∥\lVert \mathbf{u} + \mathbf{v} \rVert \leqslant \lVert \mathbf{u} \rVert + \lVert \mathbf{v} \rVert.

Ouu + v‖u‖‖v‖‖u + v‖
Figure 1.6. The triangle with vertices 0\mathbf{0}, u\mathbf{u} and u+v\mathbf{u} + \mathbf{v}. Its sides have lengths ∥u∥\lVert \mathbf{u} \rVert, ∥v∥\lVert \mathbf{v} \rVert and ∥u+v∥\lVert \mathbf{u} + \mathbf{v} \rVert.

The name comes from that picture. The side from 0\mathbf{0} to u+v\mathbf{u} + \mathbf{v} has length ∥u+v∥\lVert \mathbf{u} + \mathbf{v} \rVert, and the other two sides have lengths ∥u∥\lVert \mathbf{u} \rVert and ∥v∥\lVert \mathbf{v} \rVert, the second because (u+v)−u=v(\mathbf{u} + \mathbf{v}) - \mathbf{u} = \mathbf{v}. Going from 0\mathbf{0} to u+v\mathbf{u} + \mathbf{v} by way of u\mathbf{u} cannot be shorter than going straight there, and it is exactly as long only when u\mathbf{u} lies on the straight segment between the two.

Problem 1.4.

Let u,v∈Rn\mathbf{u}, \mathbf{v} \in \mathbb{R}^n. Prove the parallelogram law

∥u+v∥2+∥u−v∥2=2∥u∥2+2∥v∥2,\lVert \mathbf{u} + \mathbf{v} \rVert^2 + \lVert \mathbf{u} - \mathbf{v} \rVert^2 = 2\lVert \mathbf{u} \rVert^2 + 2\lVert \mathbf{v} \rVert^2,

and interpret it as a statement about the two diagonals of the parallelogram spanned by u\mathbf{u} and v\mathbf{v}.

Problem 1.5.

Let u1,…,uk∈Rn\mathbf{u}_1, \ldots, \mathbf{u}_k \in \mathbb{R}^n. Prove that

∥u1+u2+⋯+uk∥⩽∥u1∥+∥u2∥+⋯+∥uk∥.\lVert \mathbf{u}_1 + \mathbf{u}_2 + \cdots + \mathbf{u}_k \rVert \leqslant \lVert \mathbf{u}_1 \rVert + \lVert \mathbf{u}_2 \rVert + \cdots + \lVert \mathbf{u}_k \rVert .

Problem 1.6.

Write d(u,v)=∥v−u∥d(\mathbf{u}, \mathbf{v}) = \lVert \mathbf{v} - \mathbf{u} \rVert for the Euclidean distance. Show that d(u,v)⩾0d(\mathbf{u}, \mathbf{v}) \geqslant 0 with equality exactly when u=v\mathbf{u} = \mathbf{v}, that d(u,v)=d(v,u)d(\mathbf{u}, \mathbf{v}) = d(\mathbf{v}, \mathbf{u}), and that

d(u,w)⩽d(u,v)+d(v,w)d(\mathbf{u}, \mathbf{w}) \leqslant d(\mathbf{u}, \mathbf{v}) + d(\mathbf{v}, \mathbf{w})

for all u,v,w∈Rn\mathbf{u}, \mathbf{v}, \mathbf{w} \in \mathbb{R}^n.

Symmetries

A symmetry of a geometrical figure is a way of moving the figure so that it ends up occupying exactly the position it started in. Draw an equilateral triangle on a transparent sheet, then pick the sheet up, turn it or flip it over without tearing or stretching it, and put it back down. If the triangle lands exactly on top of where it was, the movement we performed is a symmetry of the triangle.

To watch what a movement does we number the corners, 11 at the top, 22 at the bottom left and 33 at the bottom right. The numbers travel with the paper; the three corner positions stay where they are. Reading off which number sits in which position after the movement tells us which movement it was.

123do nothing312turn left231turn right132flip in the top axis321flip in the left axis213flip in the right axis
Figure 1.7. The six symmetries of an equilateral triangle. Each mirror line runs through one corner and the midpoint of the opposite side.

The first of the six, the movement which does nothing, is the identity symmetry, written ee.

Six is the whole list, because a symmetry has to send corners to corners, and once we know where two of the corners go the third has nowhere left to be.

The six symmetries can be combined. Performing one symmetry and then another leaves the triangle in its original position, so the result is again a symmetry.

123312321turn leftflip in the top axis
Figure 1.8. A turn followed by a flip. The arrangement at the end is the one a single flip in the left axis would have produced from the start.

Combining also shows that the order in which two symmetries are performed matters.

123312321turn leftflip123132213flipturn left
Figure 1.9. The same two symmetries, performed in the two possible orders. The flip is in the top axis in both rows, and the two results differ.

So three things are true of the symmetries of a figure, and one thing is not. There is an identity; every symmetry can be undone, so it has an inverse; combining is associative, since bracketing a list of movements only says where to pause and not what order to perform them in. But combining does not commute.

Maps

Making that precise takes a little vocabulary about functions and three facts about them. The same ground is covered at length in MA01.

Definition 1.20 (Injections, surjections, bijections).

Let f:A→Bf : A \to B be a function. It is injective if f(a)=f(a′)f(a) = f(a') implies a=a′a = a'; surjective if for every b∈Bb \in B there is an a∈Aa \in A with f(a)=bf(a) = b; and bijective, or a bijection, if it is both.

For f:A→Bf : A \to B and g:B→Cg : B \to C the composite g∘f:A→Cg \circ f : A \to C is the function with (g∘f)(a)=g(f(a))(g \circ f)(a) = g(f(a)). The identity map idA:A→A\mathrm{id}_A : A \to A is the function with idA(a)=a\mathrm{id}_A(a) = a.

Proposition 1.21 (Composition and inverses).

Let f:A→Bf : A \to B, g:B→Cg : B \to C and h:C→Dh : C \to D be functions.

  1. (h∘g)∘f=h∘(g∘f)(h \circ g) \circ f = h \circ (g \circ f);
  2. if ff and gg are bijections, then so is g∘fg \circ f;
  3. if ff is a bijection, there is exactly one function f−1:B→Af^{-1} : B \to A with f−1∘f=idAf^{-1} \circ f = \mathrm{id}_A and f∘f−1=idBf \circ f^{-1} = \mathrm{id}_B, and f−1f^{-1} is itself a bijection.

Discussion.

The first part is an equality of two functions with the same domain and codomain, so it is proved by evaluating both at an arbitrary point and unfolding the definition of a composite twice on each side; the two unfoldings meet at h(g(f(a)))h(g(f(a))), and nothing about ff, gg or hh beyond their being functions is used. The second splits along the definition of bijection: injectivity is proved by peeling the two functions off an equation of images in the order they were applied, surjectivity by producing a preimage in two steps, first under gg and then under ff. The third is a construction rather than a deduction: surjectivity of ff says each b∈Bb \in B has at least one preimage and injectivity says it has at most one, so “the” preimage is a well-defined function of bb, and the two composites collapse to the identity by construction. That f−1f^{-1} is again a bijection follows because ff is an inverse for it.

Proof.

For the first part, let a∈Aa \in A. Then

((h∘g)∘f)(a)=(h∘g)(f(a))=h(g(f(a)))=h((g∘f)(a))=(h∘(g∘f))(a),\bigl((h \circ g) \circ f\bigr)(a) = (h \circ g)\bigl(f(a)\bigr) = h\Bigl(g\bigl(f(a)\bigr)\Bigr) = h\bigl((g \circ f)(a)\bigr) = \bigl(h \circ (g \circ f)\bigr)(a),

and as aa was arbitrary the two functions are equal.

For the second, suppose ff and gg are bijections. If g(f(a))=g(f(a′))g(f(a)) = g(f(a')) then f(a)=f(a′)f(a) = f(a') because gg is injective, and then a=a′a = a' because ff is injective; so g∘fg \circ f is injective. Given c∈Cc \in C, surjectivity of gg supplies b∈Bb \in B with g(b)=cg(b) = c, and surjectivity of ff supplies a∈Aa \in A with f(a)=bf(a) = b, whence (g∘f)(a)=c(g \circ f)(a) = c; so g∘fg \circ f is surjective.

For the third, let ff be a bijection and let b∈Bb \in B. Surjectivity gives at least one a∈Aa \in A with f(a)=bf(a) = b, and injectivity gives at most one, so there is exactly one; define f−1(b)f^{-1}(b) to be it. Then f(f−1(b))=bf(f^{-1}(b)) = b for every bb, and for a∈Aa \in A the element f−1(f(a))f^{-1}(f(a)) is the unique preimage of f(a)f(a), which is aa. So the two composites are the identities. If k:B→Ak : B \to A also satisfies them, then k=k∘idB=k∘(f∘f−1)=(k∘f)∘f−1=idA∘f−1=f−1k = k \circ \mathrm{id}_B = k \circ (f \circ f^{-1}) = (k \circ f) \circ f^{-1} = \mathrm{id}_A \circ f^{-1} = f^{-1} by the first part, which is uniqueness. Finally, the same two equations read with the roles of ff and f−1f^{-1} exchanged say that f−1f^{-1} has an inverse, namely ff; so f−1f^{-1} is injective, since f−1(b)=f−1(b′)f^{-1}(b) = f^{-1}(b') gives b=f(f−1(b))=f(f−1(b′))=b′b = f(f^{-1}(b)) = f(f^{-1}(b')) = b', and surjective, since a=f−1(f(a))a = f^{-1}(f(a)) for every a∈Aa \in A.

Symmetries of the Square

The informal account leaves “moving without tearing or stretching” undefined. Such a movement does not change the distance between any two points of the figure, and we take that as the definition.

Definition 1.22 (Symmetry).

Let P⊂R2P \subset \mathbb{R}^2 be a non-empty subset. A symmetry of PP is a bijection f:P→Pf : P \to P that preserves distances, meaning

∥f(x)−f(y)∥=∥x−y∥for all x,y∈P.\lVert f(x) - f(y) \rVert = \lVert x - y \rVert \qquad \text{for all } x, y \in P.

The norm in this condition is the one from the first chapter of this lesson. The definition does not mention turns or flips; those are consequences.

We work with one figure throughout, the unit square centred at the origin with its edges parallel to the axes.

Definition 1.23 (The square SS).

S={(x,y)∈R2  ∣  −12⩽x⩽12,  −12⩽y⩽12}.S = \Bigl\{ (x, y) \in \mathbb{R}^2 \;\Bigm|\; -\tfrac{1}{2} \leqslant x \leqslant \tfrac{1}{2}, \; -\tfrac{1}{2} \leqslant y \leqslant \tfrac{1}{2} \Bigr\}.

Its four corners, numbered anticlockwise from the top right, are

c1=(12,12),c2=(−12,12),c3=(−12,−12),c4=(12,−12),c_1 = \bigl(\tfrac{1}{2}, \tfrac{1}{2}\bigr), \quad c_2 = \bigl(-\tfrac{1}{2}, \tfrac{1}{2}\bigr), \quad c_3 = \bigl(-\tfrac{1}{2}, -\tfrac{1}{2}\bigr), \quad c_4 = \bigl(\tfrac{1}{2}, -\tfrac{1}{2}\bigr),

and we call them the vertices of SS.

xyc1c2c3c4½−½½−½
Figure 1.10. The square SS and its vertices.

Eight symmetries of SS can be written down by inspection: the movement which does nothing; the clockwise rotations through 90∘90^\circ, 180∘180^\circ and 270∘270^\circ; the reflections in the vertical and the horizontal axis; and the reflections in the two diagonals.

Problem 1.7.

For each of the eight symmetries just listed, determine which vertex it sends c1c_1 to.

To show that the list is complete, we first show that a symmetry sends vertices to vertices, using distances alone.

Proposition 1.24 (Symmetries send vertices to vertices).

Let ff be a symmetry of SS. Then f(c)f(c) is a vertex of SS whenever cc is.

Discussion.

A symmetry is only known to preserve distances, so the proof describes the four vertices by distances: they are the points lying furthest apart. We therefore first show that ∥x−y∥⩽2\lVert x - y \rVert \leqslant \sqrt{2} for all x,y∈Sx, y \in S, with equality exactly when xx and yy are diagonally opposite vertices; this is a computation on coordinates, since each coordinate of x−yx - y has absolute value at most 11 and equality in the sum of squares forces equality in each term. That done, a point of SS is a vertex if and only if some point of SS is at distance 2\sqrt{2} from it, a condition stated purely in distances. Applying ff to such a pair preserves the distance, so the image of a vertex again has a partner at distance 2\sqrt{2} and is therefore a vertex.

Proof.

Let x=(x1,x2)x = (x_1, x_2) and y=(y1,y2)y = (y_1, y_2) lie in SS. Each of x1,y1x_1, y_1 lies between −12-\tfrac{1}{2} and 12\tfrac{1}{2}, so ∣x1−y1∣⩽1|x_1 - y_1| \leqslant 1, and likewise ∣x2−y2∣⩽1|x_2 - y_2| \leqslant 1. Hence

∥x−y∥2=(x1−y1)2+(x2−y2)2⩽1+1=2,\lVert x - y \rVert^2 = (x_1 - y_1)^2 + (x_2 - y_2)^2 \leqslant 1 + 1 = 2,

so ∥x−y∥⩽2\lVert x - y \rVert \leqslant \sqrt{2}. Equality forces (x1−y1)2=(x2−y2)2=1(x_1 - y_1)^2 = (x_2 - y_2)^2 = 1, hence ∣xi−yi∣=1|x_i - y_i| = 1 for i=1,2i = 1, 2; and since both coordinates are confined to an interval of length 11, this happens only when one of them is −12-\tfrac{1}{2} and the other 12\tfrac{1}{2}. So equality holds exactly when xx and yy are vertices with both coordinates opposite, that is, when they are diagonally opposite vertices.

Consequently a point x∈Sx \in S is a vertex if and only if there is some x′∈Sx' \in S with ∥x−x′∥=2\lVert x - x' \rVert = \sqrt{2}: if xx is a vertex, take x′x' diagonally opposite; and conversely the equality case just described makes xx a vertex.

Now let ff be a symmetry and cc a vertex, and choose c′c' with ∥c−c′∥=2\lVert c - c' \rVert = \sqrt{2}. Then

∥f(c)−f(c′)∥=∥c−c′∥=2,\lVert f(c) - f(c') \rVert = \lVert c - c' \rVert = \sqrt{2},

and f(c′)f(c') lies in SS, so f(c)f(c) is a vertex by the criterion.

Example 1.25 (A quarter-turn on the vertices).

The anticlockwise rotation through 90∘90^\circ sends (x,y)(x, y) to (−y,x)(-y, x). On the vertices it acts by

c1↦c2,c2↦c3,c3↦c4,c4↦c1,c_1 \mapsto c_2, \qquad c_2 \mapsto c_3, \qquad c_3 \mapsto c_4, \qquad c_4 \mapsto c_1,

which is the numbering running one step anticlockwise, as it should be.

Next, the vertices determine every other point of the square.

Proposition 1.26 (Two adjacent vertices locate a point).

Let dd and d′d' be vertices of SS with ∥d−d′∥=1\lVert d - d' \rVert = 1, and let p,q∈Sp, q \in S satisfy

∥p−d∥=∥q−d∥and∥p−d′∥=∥q−d′∥.\lVert p - d \rVert = \lVert q - d \rVert \qquad \text{and} \qquad \lVert p - d' \rVert = \lVert q - d' \rVert.

Then p=qp = q.

Discussion.

The claim is that two distances determine a point, so what we want is to recover each coordinate of pp from the two given numbers. Adjacent vertices agree in one coordinate and differ by 11 in the other, so subtracting the two squared distances cancels the coordinate they agree in, and what survives is a multiple of the coordinate they differ in. That coordinate is therefore determined. Feeding it back into either squared distance determines the square of the remaining coordinate’s offset from the shared value, and a square leaves two candidates; the ambiguity is removed by the fact that pp lies in SS, which forces the offset to have a known sign. Since every step determines a quantity from the two given distances alone, qq must produce the same values, and the two points agree.

Proof.

Write d=(d1,d2)d = (d_1, d_2) and d′=(d1′,d2′)d' = (d'_1, d'_2). Since dd and d′d' are vertices at distance 11, they agree in one coordinate and differ in the other. Say they agree in coordinate ll and differ in coordinate kk, where {k,l}={1,2}\{k, l\} = \{1, 2\}, and write

dk=ε2,dk′=−ε2,dl=dl′=δ2,d_k = \tfrac{\varepsilon}{2}, \qquad d'_k = -\tfrac{\varepsilon}{2}, \qquad d_l = d'_l = \tfrac{\delta}{2},

with ε,δ∈{1,−1}\varepsilon, \delta \in \{1, -1\}.

Let p=(p1,p2)∈Sp = (p_1, p_2) \in S. Expanding and cancelling the terms in coordinate ll,

∥p−d′∥2−∥p−d∥2=(pk+ε2)2−(pk−ε2)2=2εpk,\lVert p - d' \rVert^2 - \lVert p - d \rVert^2 = \bigl(p_k + \tfrac{\varepsilon}{2}\bigr)^2 - \bigl(p_k - \tfrac{\varepsilon}{2}\bigr)^2 = 2\varepsilon p_k,

so pkp_k is determined by the two distances. Then

(pl−δ2)2=∥p−d∥2−(pk−ε2)2\bigl(p_l - \tfrac{\delta}{2}\bigr)^2 = \lVert p - d \rVert^2 - \bigl(p_k - \tfrac{\varepsilon}{2}\bigr)^2

is determined as well. Since p∈Sp \in S we have ∣pl∣⩽12|p_l| \leqslant \tfrac{1}{2}, so δ(pl−δ2)=δpl−12⩽0\delta\bigl(p_l - \tfrac{\delta}{2}\bigr) = \delta p_l - \tfrac{1}{2} \leqslant 0; that is, pl−δ2p_l - \tfrac{\delta}{2} is −δ-\delta times a non-negative number, and it is therefore the one square root of the displayed quantity carrying that sign. Hence plp_l is determined too.

Every quantity in this computation depends only on ∥p−d∥\lVert p - d \rVert and ∥p−d′∥\lVert p - d' \rVert. The point qq has the same two distances, so the same computation returns the same coordinates, and p=qp = q.

Remark (The same fact drawn).

Geometrically the proposition says that two circles centred at adjacent vertices meet in at most one point of SS. Two distinct circles meet in at most two points, and those two are mirror images of one another in the line through the centres. Here that line is an edge of the square, so one of the two intersections lies on the square’s side of the edge and the other lies outside SS altogether.

dd′pp′
Figure 1.11. Circles centred at two adjacent vertices. Of their two intersections only pp lies in SS; the other, p′p', is its mirror image in the edge and lies outside.

Proposition 1.27 (A symmetry is determined by the vertices).

Let ff and gg be symmetries of SS with f(ci)=g(ci)f(c_i) = g(c_i) for i=1,2,3,4i = 1, 2, 3, 4. Then f=gf = g.

Discussion.

This is a uniqueness claim about functions, so we fix an arbitrary x∈Sx \in S and prove f(x)=g(x)f(x) = g(x). The tool is the proposition above, which needs two things: a pair of adjacent vertices, and the two points f(x)f(x) and g(x)g(x) standing at equal distances from each of them. The pair to use is f(c1)f(c_1) and f(c2)f(c_2), adjacent because ff preserves the distance ∥c1−c2∥=1\lVert c_1 - c_2 \rVert = 1 and by the previous proposition sends both to vertices. The equal distances come from distance preservation applied to each of ff and gg in turn: both ∥f(x)−f(ci)∥\lVert f(x) - f(c_i) \rVert and ∥g(x)−g(ci)∥\lVert g(x) - g(c_i) \rVert equal ∥x−ci∥\lVert x - c_i \rVert, and the hypothesis makes f(ci)f(c_i) and g(ci)g(c_i) the same point, so the two distances are measured from the same place. That proposition then gives f(x)=g(x)f(x) = g(x).

Proof.

The vertices c1c_1 and c2c_2 satisfy ∥c1−c2∥=1\lVert c_1 - c_2 \rVert = 1. By the previous proposition f(c1)f(c_1) and f(c2)f(c_2) are vertices, and

∥f(c1)−f(c2)∥=∥c1−c2∥=1,\lVert f(c_1) - f(c_2) \rVert = \lVert c_1 - c_2 \rVert = 1,

so they are adjacent. Let x∈Sx \in S and let i∈{1,2}i \in \{1, 2\}. Since ff preserves distances,

∥f(x)−f(ci)∥=∥x−ci∥,\lVert f(x) - f(c_i) \rVert = \lVert x - c_i \rVert,

and since gg does too, together with g(ci)=f(ci)g(c_i) = f(c_i),

∥g(x)−f(ci)∥=∥g(x)−g(ci)∥=∥x−ci∥.\lVert g(x) - f(c_i) \rVert = \lVert g(x) - g(c_i) \rVert = \lVert x - c_i \rVert.

So f(x)f(x) and g(x)g(x) are points of SS at equal distances from each of the adjacent vertices f(c1)f(c_1) and f(c2)f(c_2). The proposition above gives f(x)=g(x)f(x) = g(x), and as xx was arbitrary, f=gf = g.

Remark (Not every shuffle of the corners is a symmetry).

The proposition says a symmetry is determined by its effect on the vertices, not that every map of the vertices to themselves comes from one. Consider the assignment

c1↦c2,c2↦c3,c3↦c1,c4↦c4.c_1 \mapsto c_2, \qquad c_2 \mapsto c_3, \qquad c_3 \mapsto c_1, \qquad c_4 \mapsto c_4.

Here ∥c1−c4∥=1\lVert c_1 - c_4 \rVert = 1 while the images satisfy ∥c2−c4∥=2\lVert c_2 - c_4 \rVert = \sqrt{2}, so the assignment changes a distance and cannot be the restriction of a symmetry.

Proposition 1.28 (The square has exactly eight symmetries).

There are exactly eight symmetries of SS, namely the eight listed above.

Discussion.

The eight are known to exist, so it remains to show that there are no others. By the last proposition a symmetry is settled once its values on the four vertices are known, so it is enough to count the possible sets of values, and that count is made one vertex at a time. There are four choices for f(c1)f(c_1). Given it, f(c2)f(c_2) must be a vertex at distance 11 from f(c1)f(c_1), and each vertex has exactly two neighbours, so two choices. The remaining two values are then forced rather than chosen, because c3c_3 is diagonally opposite c1c_1 and c4c_4 diagonally opposite c2c_2, and a symmetry preserves the distance 2\sqrt{2} that says so. Four times two is eight, and since the eight listed symmetries are distinct, every count is attained.

Proof.

Let ff be a symmetry of SS. By the proposition on vertices, ff carries each cic_i to a vertex, so there are at most four possibilities for f(c1)f(c_1).

Suppose f(c1)f(c_1) is fixed. Since ∥c1−c2∥=1\lVert c_1 - c_2 \rVert = 1, the vertex f(c2)f(c_2) satisfies ∥f(c1)−f(c2)∥=1\lVert f(c_1) - f(c_2) \rVert = 1, so it is one of the two vertices adjacent to f(c1)f(c_1): two possibilities.

Now ∥c1−c3∥=2\lVert c_1 - c_3 \rVert = \sqrt{2}, so ∥f(c1)−f(c3)∥=2\lVert f(c_1) - f(c_3) \rVert = \sqrt{2}, and by the equality case established earlier f(c3)f(c_3) is the vertex diagonally opposite f(c1)f(c_1). It is therefore determined by f(c1)f(c_1). The same argument determines f(c4)f(c_4) as the vertex diagonally opposite f(c2)f(c_2).

So the values of ff on the four vertices are settled by at most 4⋅2=84 \cdot 2 = 8 combinations, and by the previous proposition ff is settled by those values. Hence there are at most eight symmetries. The eight movements listed act differently on the vertices, so they are eight distinct symmetries, and the count is exact.

Combining Symmetries

Proposition 1.29 (Symmetries compose and invert).

Let P⊂R2P \subset \mathbb{R}^2 be non-empty and let f,gf, g be symmetries of PP. Then f∘gf \circ g is a symmetry of PP, the identity map on PP is a symmetry of PP, and f−1f^{-1} is a symmetry of PP.

Discussion.

Each of the three claims asks for two things, that a certain map is a bijection P→PP \to P and that it preserves distances, and the proposition on composition and inverses supplies the first half in every case. What is left is distance, and each case is one line. For the composite, apply the hypothesis on gg and then the hypothesis on ff to the pair it produced. For the identity, there is nothing to check. For the inverse, note that an arbitrary pair of points of PP can be written as f(x),f(y)f(x), f(y), because ff is onto, and then the condition on ff read backwards is the condition on f−1f^{-1}.

Proof.

Let x,y∈Px, y \in P. Since gg and then ff preserve distances,

∥f(g(x))−f(g(y))∥=∥g(x)−g(y)∥=∥x−y∥,\lVert f(g(x)) - f(g(y)) \rVert = \lVert g(x) - g(y) \rVert = \lVert x - y \rVert,

and f∘gf \circ g is a bijection P→PP \to P by the second part of the proposition on composition and inverses, so it is a symmetry. The identity map is a bijection and leaves both sides of the condition untouched.

For the inverse, f−1f^{-1} is a bijection P→PP \to P by the third part of that proposition. Let u,v∈Pu, v \in P; since ff is surjective there are x,y∈Px, y \in P with u=f(x)u = f(x) and v=f(y)v = f(y), and then f−1(u)=xf^{-1}(u) = x, f−1(v)=yf^{-1}(v) = y. Hence

∥f−1(u)−f−1(v)∥=∥x−y∥=∥f(x)−f(y)∥=∥u−v∥.\lVert f^{-1}(u) - f^{-1}(v) \rVert = \lVert x - y \rVert = \lVert f(x) - f(y) \rVert = \lVert u - v \rVert.

Definition 1.30 (Product of symmetries).

For symmetries ff and gg of a set PP we write fgfg for the composite f∘gf \circ g, so that

(fg)(x)=f(g(x))for every x∈P.(fg)(x) = f\bigl(g(x)\bigr) \qquad \text{for every } x \in P.

The symmetry gg is performed first and ff second.

Composition of maps is associative, so a product of several symmetries may be written without brackets. From here on rr denotes the anticlockwise quarter-turn of SS and ss the reflection in the horizontal axis, that is

r(x,y)=(−y,x),s(x,y)=(x,−y).r(x, y) = (-y, x), \qquad s(x, y) = (x, -y).

Example 1.31 (Powers of the quarter-turn).

Repeating rr gives the other rotations:

r2=rotation through 180∘,r3=rotation through 270∘,r4=e.r^2 = \text{rotation through } 180^\circ, \qquad r^3 = \text{rotation through } 270^\circ, \qquad r^4 = e.

Since r4=er^4 = e, the inverse of rr is r3r^3.

Example 1.32 (rs≠srrs \neq sr).

Read off the action on c1c_1. We have s(c1)=c4s(c_1) = c_4 and r(c4)=c1r(c_4) = c_1, so

(rs)(c1)=r(s(c1))=r(c4)=c1.(rs)(c_1) = r\bigl(s(c_1)\bigr) = r(c_4) = c_1.

On the other hand r(c1)=c2r(c_1) = c_2 and s(c2)=c3s(c_2) = c_3, so (sr)(c1)=c3(sr)(c_1) = c_3. The two symmetries disagree at c1c_1, hence rs≠srrs \neq sr.

This is the first operation we have met which is not commutative.

Problem 1.8.

Identify rsrs and srsr among the eight symmetries listed earlier, describing each as a rotation or as a reflection in a named axis.

The two generators are not independent. Computing both sides on a general point,

(sr)(x,y)=s(−y,x)=(−y,−x),(r−1s)(x,y)=r−1(x,−y)=(−y,−x),(sr)(x, y) = s(-y, x) = (-y, -x), \qquad (r^{-1}s)(x, y) = r^{-1}(x, -y) = (-y, -x),

so sr=r−1ssr = r^{-1}s; multiplying on the right by ss and using s2=es^2 = e gives the equivalent form srs=r−1srs = r^{-1}. The relation lets any product of rr‘s and ss‘s be rewritten with all the rr‘s on the left, and the resulting shapes are

e,r,r2,r3,s,rs,r2s,r3s.e, \quad r, \quad r^2, \quad r^3, \quad s, \quad rs, \quad r^2 s, \quad r^3 s.

These are eight symmetries and they are distinct, so by the counting proposition they are exactly the eight symmetries of the square. The first four are the rotations and the last four the reflections.

Groups

The arguments about symmetries of the square used four facts: combining two of them gives a third; there is an identity; every one of them can be undone; and combining is associative. The axioms of a group are exactly those four.

Definition 1.33 (Binary operation).

A binary operation on a set GG is a function ⋆:G×G→G\star : G \times G \to G. We write a⋆ba \star b for the value of ⋆\star at the ordered pair (a,b)(a, b).

Definition 1.34 (Group).

A group is a pair (G,⋆)(G, \star) consisting of a set GG and a binary operation ⋆\star on GG satisfying:

  1. (G1) Associativity. (f⋆g)⋆h=f⋆(g⋆h)(f \star g) \star h = f \star (g \star h) for all f,g,h∈Gf, g, h \in G;
  2. (G2) Identity. there is an e∈Ge \in G with e⋆g=g⋆e=ge \star g = g \star e = g for every g∈Gg \in G;
  3. (G3) Inverses. for every g∈Gg \in G there is an h∈Gh \in G with g⋆h=h⋆g=eg \star h = h \star g = e.

The element ee of (G2) is the identity, or neutral element, and the element hh of (G3) is an inverse of gg. Both turn out to be unique, and the inverse of gg is then written g−1g^{-1}.

The first of the four facts about symmetries, that combining two gives a third, does not appear as an axiom because it is already in the definition of a binary operation: the codomain of ⋆\star is GG, so a product of two elements of GG is an element of GG and there is nothing further to require.

Remark (The closure axiom).

Some texts add a fourth axiom, closure: if g,k∈Gg, k \in G then g⋆k∈Gg \star k \in G. Under our definition that is automatic, for the reason just given. It has to be stated when the operation is introduced as something defined on a larger set and then restricted, since one must then check that the restriction lands back inside GG.

Remark (Writing the operation).

The symbol ⋆\star is usually replaced by a dot, and the dot is usually dropped: we write ghgh for g⋆hg \star h and call it the product of gg and hh, exactly as for a product of numbers. This is a convention about notation, not a claim that the operation resembles multiplication, and in particular it does not license writing gh=hggh = hg.

Associativity means that a product of three elements may be written ghkghk with no bracket, since the two possible bracketings agree. The same then holds for longer products: an expression such as ((a1a2)(a3a4))a5((a_1a_2)(a_3a_4))a_5 has the same value as every other bracketing of a1,…,a5a_1, \ldots, a_5 in that order, and we write it a1a2a3a4a5a_1a_2a_3a_4a_5. From here on products of any length are written without brackets.

Problem 1.9.

Let (G,⋆)(G, \star) be a group and let a1,…,an∈Ga_1, \ldots, a_n \in G. Prove that every way of bracketing the product a1⋆a2⋆⋯⋆ana_1 \star a_2 \star \cdots \star a_n, keeping the terms in that order, gives the same element of GG.

Problem 1.10.

Let cnc_n be the number of ways of bracketing a product a1⋆a2⋆⋯⋆ana_1 \star a_2 \star \cdots \star a_n, keeping the terms in that order, so that c2=1c_2 = 1, c3=2c_3 = 2 and c4=5c_4 = 5, and set c1=1c_1 = 1.

  1. Explain why cn=c1cn−1+c2cn−2+⋯+cn−1c1c_n = c_1 c_{n-1} + c_2 c_{n-2} + \cdots + c_{n-1} c_1.
  2. Use the recurrence to compute c5c_5 and c6c_6.

Definition 1.35 (Abelian group).

A group (G,⋆)(G, \star) is commutative, or abelian, if it satisfies the further condition

gh=hgfor all g,h∈G.gh = hg \qquad \text{for all } g, h \in G.

Example 1.36 (Two groups).

The integers under addition, (Z,+)(\mathbb{Z}, +), form an infinite abelian group: addition is associative, 00 is the identity, and the inverse of nn is −n-n. The symmetries of the square under composition form a finite group which is not abelian, since rs≠srrs \neq sr.

Remark (Associativity is not commutativity).

Associativity says (fg)h=f(gh)(fg)h = f(gh) and commutativity says fg=gffg = gf; the two are easy to confuse and are not related. Associativity holds in every group by definition, while commutativity is an extra condition that a group may or may not satisfy.

For symmetries, associativity holds automatically. Both (fg)h(fg)h and f(gh)f(gh) mean: perform hh, then gg, then ff. The brackets only divide the calculation into stages and do not touch the order in which the movements happen. Operations that are not associative do exist, subtraction on R\mathbb{R} being one, since (a−b)−c(a - b) - c and a−(b−c)a - (b - c) usually differ; but they cannot arise from composing functions, which is associative always.

Proposition 1.37 (The identity is unique).

Let (G,⋆)(G, \star) be a group. If ee and e′e' both satisfy the condition (G2), then e=e′e = e'.

Discussion.

The statement is an equality of two elements, and the only material available is that each of them is neutral. We form the product e⋆e′e \star e', in which both hypotheses can be used, and evaluate it twice: treating ee as an identity leaves e′e', and treating e′e' as an identity leaves ee. No contradiction and no case split is needed, and neither associativity nor inverses enter.

Proof.

Since e′e' satisfies (G2), we have e⋆e′=ee \star e' = e. Since ee satisfies (G2), we have e⋆e′=e′e \star e' = e'. Hence e=e⋆e′=e′e = e \star e' = e'.

Proposition 1.38 (Inverses are unique).

Let (G,⋆)(G, \star) be a group with identity ee, and let g,g′,g′′∈Gg, g', g'' \in G satisfy

gg′=e=g′gandgg′′=e=g′′g.g g' = e = g' g \qquad \text{and} \qquad g g'' = e = g'' g.

Then g′=g′′g' = g''.

Discussion.

Again the goal is an equality of two elements and the hypotheses are two equations, so we build one expression that both can act on: the triple product g′gg′′g' g g''. Read with the brackets to the left it uses g′g=eg'g = e and collapses to g′′g''; read with the brackets to the right it uses gg′′=egg'' = e and collapses to g′g'. Associativity says that the two readings are the same element, and that is the proof.

Proof.

Using associativity in the middle step,

g′′=eg′′=(g′g)g′′=g′(gg′′)=g′e=g′.g'' = e g'' = (g' g) g'' = g' (g g'') = g' e = g'.

Because of this proposition the inverse of gg may be named, and we write it g−1g^{-1}.

Problem 1.11.

The following is offered as a proof of the last proposition: “from the hypotheses, g′=g−1g' = g^{-1} and g′′=g−1g'' = g^{-1}, so g′=g′′g' = g''.” Explain why it is not one.

Proposition 1.39 (Solving an equation in a group).

Let (G,⋅)(G, \cdot) be a group and let g,h∈Gg, h \in G.

  1. For x∈Gx \in G we have gx=hgx = h if and only if x=g−1hx = g^{-1}h;
  2. for y∈Gy \in G we have yg=hyg = h if and only if y=hg−1y = hg^{-1}.

Discussion.

Each part is a biconditional between two equations, so each direction is proved by multiplying the given equation by g−1g^{-1} on the appropriate side and simplifying with the axioms. Which side matters: the first part multiplies on the left throughout, because the unknown sits to the right of gg, and the second multiplies on the right for the mirror-image reason. No step exchanges the two factors, so in a group that is not abelian the two parts are different statements.

Proof.

For the first part, suppose gx=hgx = h. Multiplying on the left by g−1g^{-1} gives g−1(gx)=g−1hg^{-1}(gx) = g^{-1}h, and associativity turns the left side into (g−1g)x=ex=x(g^{-1}g)x = ex = x, so x=g−1hx = g^{-1}h. Conversely, suppose x=g−1hx = g^{-1}h. Multiplying on the left by gg gives gx=g(g−1h)=(gg−1)h=eh=hgx = g(g^{-1}h) = (gg^{-1})h = eh = h.

Problem 1.12.

Prove the second part of the last proposition.

Proposition 1.40 (Inverse of a product).

Let (G,⋅)(G, \cdot) be a group and let g,h∈Gg, h \in G. Then (gh)−1=h−1g−1(gh)^{-1} = h^{-1}g^{-1}.

Discussion.

By the uniqueness of inverses it is enough to check that the proposed element does what an inverse of ghgh has to do, so we multiply h−1g−1h^{-1}g^{-1} by ghgh on one side and cancel from the middle outwards, then do the same on the other side. The order in which the two factors are reversed is forced by exactly this cancellation, since it is g−1g^{-1} that has to meet gg first.

Proof.

Using associativity to bracket at will,

(h−1g−1)(gh)=h−1(g−1g)h=h−1eh=h−1h=e,(h^{-1}g^{-1})(gh) = h^{-1}(g^{-1}g)h = h^{-1} e h = h^{-1}h = e,

and symmetrically (gh)(h−1g−1)=g(hh−1)g−1=gg−1=e(gh)(h^{-1}g^{-1}) = g(hh^{-1})g^{-1} = gg^{-1} = e. So h−1g−1h^{-1}g^{-1} is an inverse of ghgh, and by uniqueness it is (gh)−1(gh)^{-1}.

Problem 1.13.

Let (G,⋅)(G, \cdot) be a group and let g1,…,gn∈Gg_1, \ldots, g_n \in G. Prove that

(g1g2⋯gn)−1=gn−1⋯g2−1g1−1.(g_1 g_2 \cdots g_n)^{-1} = g_n^{-1} \cdots g_2^{-1} g_1^{-1}.

Groups of Small Order

Definition 1.41 (Order of a group).

A group (G,⋅)(G, \cdot) is a finite group if the set GG is finite. The order of a finite group is the cardinality #(G)\#(G), the number of its elements. Every group has an identity, so GG is never empty and its order is at least 11.

Remark (Infinite groups).

Not every group is finite. The unit circle in R2\mathbb{R}^2 has infinitely many symmetries, one rotation for each angle, and (Z,+)(\mathbb{Z}, +) is infinite as well.

Powers are defined as in arithmetic.

Definition 1.42 (Powers).

Let (G,⋅)(G, \cdot) be a group and g∈Gg \in G. Define g0=eg^0 = e and gn+1=gngg^{n+1} = g^n g for n⩾0n \geqslant 0, and set g−n=(g−1)ng^{-n} = (g^{-1})^n for n⩾1n \geqslant 1.

Proposition 1.43 (Laws of exponents).

Let (G,⋅)(G, \cdot) be a group, let g∈Gg \in G and let m,nm, n be integers. Then

gm+n=gmgnand(gm)n=gmn.g^{m+n} = g^m g^n \qquad \text{and} \qquad (g^m)^n = g^{mn}.

Discussion.

Both identities are statements about all integers, but the definition of a power is a recursion on the non-negative ones, so we use induction for m,n⩾0m, n \geqslant 0 followed by a reduction of the remaining sign cases to that one. The induction for the first identity runs on nn: the base case is the definition of g0g^0, and the step is the recursion clause together with associativity. The second identity then follows from the first by a second induction on nn. For negative exponents we use that g−1g^{-1} has the same structure, and that gng^n and g−ng^{-n} are inverse to one another, which converts a negative exponent into a positive one on the inverse element.

Proof.

Take first m⩾0m \geqslant 0 and induct on n⩾0n \geqslant 0. For n=0n = 0 both sides are gmg^m, since g0=eg^0 = e. If gm+n=gmgng^{m+n} = g^m g^n, then

gm+(n+1)=g(m+n)+1=gm+ng=(gmgn)g=gm(gng)=gmgn+1,g^{m + (n+1)} = g^{(m+n)+1} = g^{m+n} g = (g^m g^n) g = g^m (g^n g) = g^m g^{n+1},

which is the claim at n+1n + 1. A second induction on n⩾0n \geqslant 0 gives (gm)n=gmn(g^m)^n = g^{mn}: the case n=0n = 0 reads e=g0e = g^0, and the step is (gm)n+1=(gm)ngm=gmngm=gmn+m=gm(n+1)(g^m)^{n+1} = (g^m)^n g^m = g^{mn} g^m = g^{mn + m} = g^{m(n+1)}, using the first identity.

For the sign cases, note that gng−n=eg^n g^{-n} = e for n⩾0n \geqslant 0, by induction on nn using the first identity for the inverse element; so g−n=(gn)−1g^{-n} = (g^n)^{-1}. Any instance of the two identities with negative exponents now follows by rewriting each negative power as the inverse of a positive one and applying the proposition on the inverse of a product.

Definition 1.44 (Order of an element).

Let (G,⋅)(G, \cdot) be a group and g∈Gg \in G. The order of gg is the smallest n∈Nn \in \mathbb{N} with gn=eg^n = e, and is ∞\infty if there is no such nn.

Example 1.45 (Orders in the group of the square).

In the group of symmetries of the square the identity has order 11; the quarter-turns rr and r3r^3 have order 44; and the remaining five elements, the half-turn r2r^2 and the four reflections ss, rsrs, r2sr^2s, r3sr^3s, all have order 22.

Proposition 1.46 (Orders are bounded by the order of the group).

Let GG be a finite group of order nn. Then every element of GG has order at most nn.

Discussion.

We need a power gk=eg^k = e with 1⩽k⩽n1 \leqslant k \leqslant n, and we use only that GG has nn elements. Listing n+1n + 1 powers of gg therefore forces a repetition, and a repetition ga=gbg^a = g^b with a<ba < b is exactly what we want: cancelling gag^a from both sides, which the group allows, leaves gb−a=eg^{b-a} = e with the exponent between 11 and nn. The definition of order then bounds it by that exponent.

Proof.

Let g∈Gg \in G and consider the n+1n + 1 elements g0,g1,…,gng^0, g^1, \ldots, g^n of GG. Since GG has only nn elements they cannot all be distinct, so there are a,ba, b with 0⩽a<b⩽n0 \leqslant a < b \leqslant n and ga=gbg^a = g^b. Multiplying by (ga)−1=g−a(g^a)^{-1} = g^{-a} and using the laws of exponents gives

e=gbg−a=gb−a,e = g^{b} g^{-a} = g^{b-a},

with 1⩽b−a⩽n1 \leqslant b - a \leqslant n. So the set of positive exponents killing gg is non-empty, and the least of them, which is the order of gg, is at most nn.

Remark (Lagrange).

More is true: in a finite group the order of every element divides the order of the group. That is Lagrange’s theorem, and it belongs to a course on abstract algebra rather than here. The finiteness of GG is what the argument above uses, not merely that orders are finite; there are infinite groups in which every element has finite order.

Proposition 1.47 (Rows and columns of the table).

Let (G,⋅)(G, \cdot) be a group and let x∈Gx \in G. Then the maps y↦xyy \mapsto xy and y↦yxy \mapsto yx are bijections from GG to GG. In a finite group, therefore, each element appears exactly once in every row and exactly once in every column of the multiplication table.

Discussion.

Each map is a bijection because an explicit inverse is available: multiplying by x−1x^{-1} on the same side undoes it, which is the content of the proposition on solving gx=hgx = h. The statement about the table is a translation, since the row labelled xx lists the values of y↦xyy \mapsto xy as yy runs over GG; a bijection from a finite set to itself hits each element exactly once, so no entry is missing and none is repeated.

Proof.

Let L(y)=xyL(y) = xy and M(y)=x−1yM(y) = x^{-1}y. Then M(L(y))=x−1(xy)=(x−1x)y=yM(L(y)) = x^{-1}(xy) = (x^{-1}x)y = y and likewise L(M(y))=yL(M(y)) = y, so LL is a bijection with inverse MM. The argument for y↦yxy \mapsto yx is the same with the multiplications on the other side.

The row of the multiplication table labelled xx has the entry xyxy in the column labelled yy, so its entries are the values of LL. As LL is a bijection of the finite set GG onto itself, every element of GG occurs among them exactly once. Columns are handled by the other map.

Definition 1.48 (Isomorphism).

Let (G,⋅)(G, \cdot) and (G′,⋅)(G', \cdot) be groups. An isomorphism from GG to G′G' is a bijection φ:G→G′\varphi : G \to G' with

φ(ab)=φ(a) φ(b)for all a,b∈G.\varphi(ab) = \varphi(a)\,\varphi(b) \qquad \text{for all } a, b \in G.

If one exists, GG and G′G' are isomorphic, written G≅G′G \cong G'.

Two isomorphic groups may be built from different objects, one from movements of a figure and one from numbers, and still have the same multiplication: the isomorphism matches their elements so that products correspond.

Proposition 1.49 (What an isomorphism preserves).

Let φ:G→G′\varphi : G \to G' be an isomorphism, with identities ee and e′e'. Then φ(e)=e′\varphi(e) = e', and for every a∈Ga \in G and every integer nn,

φ(a−1)=φ(a)−1,φ(an)=φ(a)n.\varphi(a^{-1}) = \varphi(a)^{-1}, \qquad \varphi(a^n) = \varphi(a)^n .

Moreover aa and φ(a)\varphi(a) have the same order.

Discussion.

The hypothesis is a single equation, φ(ab)=φ(a)φ(b)\varphi(ab) = \varphi(a)\varphi(b), so each claim is obtained from it by a substitution. Taking a=b=ea = b = e makes the equation say φ(e)=φ(e)φ(e)\varphi(e) = \varphi(e)\varphi(e), which cancels to the first claim. Taking b=a−1b = a^{-1} then says e′=φ(a)φ(a−1)e' = \varphi(a)\varphi(a^{-1}), which identifies the second. The power law follows by induction on n⩾0n \geqslant 0, the negative case by combining the two claims already made. The statement about orders uses bijectivity: injectivity turns φ(an)=φ(e)\varphi(a^n) = \varphi(e) back into an=ea^n = e, so the positive exponents killing aa are exactly those killing φ(a)\varphi(a), and two sets of positive integers that coincide have the same least element.

Proof.

Putting a=b=ea = b = e gives φ(e)=φ(e)φ(e)\varphi(e) = \varphi(e)\varphi(e), and multiplying by φ(e)−1\varphi(e)^{-1} gives e′=φ(e)e' = \varphi(e). Putting b=a−1b = a^{-1} gives φ(a)φ(a−1)=φ(aa−1)=φ(e)=e′\varphi(a)\varphi(a^{-1}) = \varphi(aa^{-1}) = \varphi(e) = e', and symmetrically on the other side, so φ(a−1)=φ(a)−1\varphi(a^{-1}) = \varphi(a)^{-1}.

For n⩾0n \geqslant 0 induct: φ(a0)=φ(e)=e′=φ(a)0\varphi(a^0) = \varphi(e) = e' = \varphi(a)^0, and if φ(an)=φ(a)n\varphi(a^n) = \varphi(a)^n then φ(an+1)=φ(ana)=φ(an)φ(a)=φ(a)n+1\varphi(a^{n+1}) = \varphi(a^n a) = \varphi(a^n)\varphi(a) = \varphi(a)^{n+1}. For n<0n < 0 write an=(a−1)−na^n = (a^{-1})^{-n} and apply the case just proved to a−1a^{-1}, using φ(a−1)=φ(a)−1\varphi(a^{-1}) = \varphi(a)^{-1}.

Finally, for n∈Nn \in \mathbb{N} we have an=ea^n = e if and only if φ(an)=φ(e)\varphi(a^n) = \varphi(e), since φ\varphi is injective, and φ(an)=φ(a)n\varphi(a^n) = \varphi(a)^n while φ(e)=e′\varphi(e) = e'. So an=ea^n = e if and only if φ(a)n=e′\varphi(a)^n = e'. The two elements are killed by the same positive exponents, hence have the same order.

Proposition 1.50 (Cyclic groups).

Let GG be a group of order nn containing an element gg of order nn. Then

G={e,g,g2,…,gn−1},G = \{e, g, g^2, \ldots, g^{n-1}\},

and gagb=gcg^a g^b = g^c, where cc is the remainder of a+ba + b on division by nn. Any two such groups are isomorphic.

Discussion.

Two things need proving: that the listed powers exhaust GG, and that the operation is forced. For the first, the nn listed powers are distinct, since an equality between two of them would produce a positive exponent smaller than nn killing gg and so contradict the order of gg; being nn distinct elements of a set with nn elements they are all of it. For the second, write a+b=qn+ca + b = qn + c by division with remainder and use the laws of exponents, where gn=eg^n = e makes the multiple of nn disappear. The last sentence is then immediate: the map ga↦(g′)ag^a \mapsto (g')^a between two such groups is a bijection by the first part and respects products by the second.

Proof.

Suppose gi=gjg^i = g^j with 0⩽i<j⩽n−10 \leqslant i < j \leqslant n - 1. Then gj−i=eg^{j-i} = e with 1⩽j−i<n1 \leqslant j - i < n, contradicting the order of gg being nn. So the nn elements e,g,…,gn−1e, g, \ldots, g^{n-1} are distinct, and as #(G)=n\#(G) = n they are all of GG.

Let a,b⩾0a, b \geqslant 0 and write a+b=qn+ca + b = qn + c with 0⩽c<n0 \leqslant c < n. By the laws of exponents,

gagb=ga+b=gqn+c=(gn)qgc=eqgc=gc.g^a g^b = g^{a+b} = g^{qn + c} = (g^n)^q g^c = e^q g^c = g^c .

Finally, let G′G' be another group of order nn with an element g′g' of order nn. Both groups are listed by their powers as above, so φ(ga)=(g′)a\varphi(g^a) = (g')^a for 0⩽a<n0 \leqslant a < n is a well-defined bijection G→G′G \to G', and the displayed rule computes the product on both sides by the same remainder, so φ(gagb)=φ(ga)φ(gb)\varphi(g^ag^b) = \varphi(g^a)\varphi(g^b).

Definition 1.51 (Cyclic group).

The group described by the last proposition is the cyclic group of order nn, written CnC_n. It is abelian, since gagbg^ag^b and gbgag^bg^a are both computed from the remainder of a+ba + b.

We can now work through the small orders. Throughout, the multiplication table of a finite group has a row and a column for each element, and the entry in row xx and column yy is x⋆yx \star y.

Order 1. The identity is the only element and the operation is e⋆e=ee \star e = e. This is the trivial group. The letter “P”, read as a subset of R2\mathbb{R}^2, has trivial symmetry group.

Problem 1.14.

Assuming a reasonably symmetrical font, so that “Y” has a reflection symmetry in a vertical axis and “B” one in a horizontal axis, determine which capital letters have trivial symmetry group.

Order 2. Let G={e,f}G = \{e, f\}. The identity axiom fills in every entry but one:

⋆efeefff?\begin{array}{c|cc} \star & e & f \\ \hline e & e & f \\ f & f & ? \end{array}

If f⋆f=ff \star f = f then multiplying by f−1f^{-1} gives f=ef = e, which is false. So f⋆f=ef \star f = e, the table is forced, and ff has order 22. Hence G≅C2G \cong C_2, and C2C_2 is the only group of order 22.

Remark (One group, two pictures).

This group is the symmetry group of the letter “Z”, whose non-trivial element is the half-turn that carries the letter onto itself, and also of the letter “Y”, whose non-trivial element is a reflection. The two figures are different geometrically, but their symmetry groups are isomorphic.

Order 3. Let G={e,f,g}G = \{e, f, g\} with f,g≠ef, g \neq e. Row ff of the table consists of f⋆e=ff \star e = f together with f⋆ff \star f and f⋆gf \star g, and by the proposition on rows and columns those three entries are ee, ff and gg in some order. So {f⋆f,f⋆g}={e,g}\{f \star f, f \star g\} = \{e, g\}. Were f⋆f=ef \star f = e we should have f⋆g=gf \star g = g, and multiplying by g−1g^{-1} would give f=ef = e. Hence f⋆f=gf \star f = g and f⋆g=ef \star g = e, which fills the table:

⋆efgeefgffgeggef\begin{array}{c|ccc} \star & e & f & g \\ \hline e & e & f & g \\ f & f & g & e \\ g & g & e & f \end{array}

Here g=f2g = f^2 and f3=f⋆g=ef^3 = f \star g = e, so ff has order 33 and G≅C3G \cong C_3.

Order 4. Two structures occur, and we separate them by the largest order of an element. By the proposition bounding orders that largest order is 11, 22, 33 or 44, and it is not 11, since then every element would be ee.

If some element has order 44, the cyclic proposition gives G≅C4G \cong C_4.

Suppose some gg has order 33, and put K={e,g,g2}K = \{e, g, g^2\}. Pick h∈G∖Kh \in G \setminus K, which exists since #(G)=4\#(G) = 4. If hgi=gjhg^i = g^j for some i,ji, j then h=gj−i∈Kh = g^{j-i} \in K, contrary to the choice of hh; so the three elements hh, hghg, hg2hg^2 all lie outside KK, and they are distinct because y↦hyy \mapsto hy is injective. That gives at least 3+3=63 + 3 = 6 elements of GG, which is impossible. So no element has order 33.

The remaining case is that every element other than ee has order 22. Write G={e,f,g,h}G = \{e, f, g, h\}. Then fg≠efg \neq e, since fg=efg = e would make g=f−1=fg = f^{-1} = f; and fg≠ffg \neq f and fg≠gfg \neq g, since either would force one of f,gf, g to be ee. Hence fg=hfg = h, and the same argument applied to each pair fills the table:

⋆efgheefghffehggghefhhgfe\begin{array}{c|cccc} \star & e & f & g & h \\ \hline e & e & f & g & h \\ f & f & e & h & g \\ g & g & h & e & f \\ h & h & g & f & e \end{array}

This is a group, because it sits inside one we have already built: taking f=r2f = r^2, g=sg = s and h=r2sh = r^2s inside the symmetries of the square gives a set of four symmetries closed under composition, each equal to its own inverse, with exactly this table. It is the Klein four-group, written V4V_4, and it is the symmetry group of the letter “H”.

So there are exactly two groups of order 44 up to isomorphism, C4C_4 and V4V_4. They are not isomorphic: one has an element of order 44 and the other does not, and an isomorphism preserves the order of an element.

Remark (How the list grows).

All the groups above are abelian, and so is every group of order 55, though we do not prove it. The smallest non-abelian group has order 66. Tabulating the groups of order nn becomes hard quickly, especially when nn is a large power of a small prime: there are 267267 groups of order 6464.

Symmetry Groups of Regular Polygons

The symmetries of the square, with composition, satisfy the group axioms by the proposition on composing and inverting symmetries. That group has eight elements and is called the dihedral group of degree four, written D4D_4.

Remark ($D_4$ or $D_8$).

The subscript here counts the sides of the square, so D4D_4 has eight elements. Some texts write D8D_8 for the same group, counting its elements instead. Check which convention a source is using before comparing statements.

In terms of the quarter-turn rr and the reflection ss, the group is described by the relations

r4=e,s2=e,srs=r−1.r^4 = e, \qquad s^2 = e, \qquad srs = r^{-1}.

The third relation lets every ss be moved past every rr. Rewriting srsr as r−1sr^{-1}s moves the rr‘s to the left, and any product of rr‘s and ss‘s collapses to one of e,r,r2,r3,s,rs,r2s,r3se, r, r^2, r^3, s, rs, r^2s, r^3s. For instance

sr2=(sr)r=(r−1s)r=r−1(sr)=r−1(r−1s)=r−2s=r2s,sr^2 = (sr)r = (r^{-1}s)r = r^{-1}(sr) = r^{-1}(r^{-1}s) = r^{-2}s = r^2 s,

using r4=er^4 = e at the last step, so the half-turn commutes with ss even though the quarter-turn does not. Similarly

(rs)2=rsrs=r(srs)=rr−1=e,(rs)^2 = rsrs = r(srs) = r r^{-1} = e,

so rsrs is one of the reflections: doing it twice returns the square to where it started.

An equilateral triangle has six symmetries, three rotations and three reflections. The rotations are ee, rr and r2r^2, where rr is the rotation through 120∘120^\circ, and they satisfy r3=er^3 = e. Each reflection fixes one vertex and exchanges the other two; if ss is one of them, the other two are rsrs and r2sr^2s. The group they form is D3D_3, with

r3=e,s2=e,srs=r−1,D3={e,r,r2,s,rs,r2s}.r^3 = e, \qquad s^2 = e, \qquad srs = r^{-1}, \qquad D_3 = \{e, r, r^2, s, rs, r^2s\}.

These are the relations of the square with the order of the basic rotation changed from 44 to 33. The group D3D_3 has order 66 and is not abelian, so it is the smallest non-abelian group.

The same account fits every regular polygon. A regular nn-gon with n⩾3n \geqslant 3 has nn rotational and nn reflectional symmetries, so its symmetry group DnD_n has 2n2n elements. Writing rr for the rotation through 360∘/n360^\circ / n and ss for any one of the reflections,

rn=e,s2=e,srs=r−1,r^n = e, \qquad s^2 = e, \qquad srs = r^{-1},

and the elements are e,r,r2,…,rn−1e, r, r^2, \ldots, r^{n-1}, which are the rotations, together with s,rs,r2s,…,rn−1ss, rs, r^2s, \ldots, r^{n-1}s, which are the reflections.

Remark (Small $n$).

For n<3n < 3 there is no regular polygon to act on, but the three relations still make sense and still define a group of order 2n2n. Neither is new: D1D_1 is C2C_2 under another name, and D2D_2 is the Klein four-group.

Symmetries as Permutations

A symmetry of the square is determined by where it sends the four vertices, and a symmetry of the triangle by where it sends the three. So we can study the rearrangements themselves.

Definition 1.52 (Permutation).

Let XX be a non-empty set. A permutation of XX is a bijection α:X→X\alpha : X \to X. The set of all permutations of XX is written Sym⁡(X)\operatorname{Sym}(X), and for X={1,2,…,n}X = \{1, 2, \ldots, n\} we write Sym⁡(n)\operatorname{Sym}(n) or SnS_n.

A permutation of a finite set is recorded by a table of two rows, the points along the top and their images beneath. For X={1,2,3,4}X = \{1, 2, 3, 4\},

α=(12342413)\alpha = \begin{pmatrix} 1 & 2 & 3 & 4 \\ 2 & 4 & 1 & 3 \end{pmatrix}

means α(1)=2\alpha(1) = 2, α(2)=4\alpha(2) = 4, and so on. The bottom row is a rearrangement of the top one, which is the condition that α\alpha is a bijection.

There is a shorter notation. Take p∈S5p \in S_5 sending (1,2,3,4,5)(1, 2, 3, 4, 5) to (3,5,4,1,2)(3, 5, 4, 1, 2). Following one point at a time, 11 goes to 33, which goes to 44, which goes back to 11; that is a cycle of length three. Separately 22 and 55 exchange, a cycle of length two. Writing each cycle in brackets gives the cycle notation

p=(1 3 4)(2 5).p = (1\,3\,4)(2\,5).

Cycle notation is not unique, since (1 3 4)(1\,3\,4) and (3 4 1)(3\,4\,1) name the same cycle, and by convention we leave out the points a permutation fixes: an index that does not appear is understood to stay where it is.

Example 1.53 (Multiplying in cycle notation).

Let p=(1 3 4)(2 5)p = (1\,3\,4)(2\,5) as above and let q=(1 2)(3 4)q = (1\,2)(3\,4) in S5S_5. Recall that pqpq means qq first, then pp. Following each point through both,

1→  q  2→  p  5,5→  q  5→  p  2,2→  q  1→  p  3,3→  q  4→  p  1,1 \xrightarrow{\;q\;} 2 \xrightarrow{\;p\;} 5, \qquad 5 \xrightarrow{\;q\;} 5 \xrightarrow{\;p\;} 2, \qquad 2 \xrightarrow{\;q\;} 1 \xrightarrow{\;p\;} 3, \qquad 3 \xrightarrow{\;q\;} 4 \xrightarrow{\;p\;} 1,

and 4↦3↦44 \mapsto 3 \mapsto 4. Collecting the cycles,

pq=(1 5 2 3).pq = (1\,5\,2\,3).

Example 1.54 (Inverses and conjugates).

Let r=(1 2 3 4 5)r = (1\,2\,3\,4\,5) and let pp be as above. Then

rp=(1 4 2)(3 5),r−1=(1 5 4 3 2),rpr−1=(1 3)(2 4 5).rp = (1\,4\,2)(3\,5), \qquad r^{-1} = (1\,5\,4\,3\,2), \qquad rpr^{-1} = (1\,3)(2\,4\,5).

The inverse of a cycle is the same cycle traversed backwards, which is where the second of these comes from. The third combination, rpr−1rpr^{-1}, is called a conjugate of pp.

Example 1.55 (Composing in two-row notation).

Let

α=(12342413),β=(12342143).\alpha = \begin{pmatrix} 1 & 2 & 3 & 4 \\ 2 & 4 & 1 & 3 \end{pmatrix}, \qquad \beta = \begin{pmatrix} 1 & 2 & 3 & 4 \\ 2 & 1 & 4 & 3 \end{pmatrix}.

Then

α∘β=(12344231),β∘α=(12341324),\alpha \circ \beta = \begin{pmatrix} 1 & 2 & 3 & 4 \\ 4 & 2 & 3 & 1 \end{pmatrix}, \qquad \beta \circ \alpha = \begin{pmatrix} 1 & 2 & 3 & 4 \\ 1 & 3 & 2 & 4 \end{pmatrix},

which disagree at 11, so composition of permutations is not commutative either.

Theorem 1.56 (The symmetric group).

Let XX be a non-empty set. Then Sym⁡(X)\operatorname{Sym}(X) with composition is a group, called the symmetric group on XX.

Discussion.

There are four things to check and the proposition on composition and inverses has done most of them. First that composition is a binary operation on Sym⁡(X)\operatorname{Sym}(X), which is the statement that a composite of bijections X→XX \to X is again one. Then the three axioms: associativity is associativity of composition of maps, which holds for all functions and not merely for bijections; the identity map is a bijection and satisfies the identity axiom by definition; and the inverse required by (G3) is the inverse function, which exists because α\alpha is a bijection and is itself a bijection.

Proof.

If α,β∈Sym⁡(X)\alpha, \beta \in \operatorname{Sym}(X) then α∘β\alpha \circ \beta is a bijection X→XX \to X by the second part of the proposition on composition and inverses, so composition is a binary operation on Sym⁡(X)\operatorname{Sym}(X).

(G1) Composition of maps is associative by the first part of that proposition, so α∘(β∘γ)=(α∘β)∘γ\alpha \circ (\beta \circ \gamma) = (\alpha \circ \beta) \circ \gamma for all α,β,γ∈Sym⁡(X)\alpha, \beta, \gamma \in \operatorname{Sym}(X); both send xx to α(β(γ(x)))\alpha(\beta(\gamma(x))).

(G2) The identity map ι:X→X\iota : X \to X with ι(x)=x\iota(x) = x is a bijection, so lies in Sym⁡(X)\operatorname{Sym}(X), and ι∘α=α∘ι=α\iota \circ \alpha = \alpha \circ \iota = \alpha for every α\alpha.

(G3) If α∈Sym⁡(X)\alpha \in \operatorname{Sym}(X) then α\alpha is a bijection, so by the third part of that proposition it has an inverse function α−1\alpha^{-1}, itself a bijection, and α∘α−1=α−1∘α=ι\alpha \circ \alpha^{-1} = \alpha^{-1} \circ \alpha = \iota.

As with symmetries we write αβ\alpha\beta for α∘β\alpha \circ \beta, so that β\beta acts first. Some texts write the argument on the left, (x)α(x)\alpha rather than α(x)\alpha(x), and then read products in the opposite order; work one product out by hand before trusting a source’s convention.

Theorem 1.57 (The size of SnS_n).

For n∈Nn \in \mathbb{N} we have #(Sn)=n!\#(S_n) = n!.

Discussion.

An element of SnS_n is settled by its bottom row, a list a1,…,ana_1, \ldots, a_n in which each of 1,…,n1, \ldots, n appears once, so the count is a count of such lists. Building one from left to right, each entry may be any value not already used, so the number of choices falls by one at each step: nn for the first, n−1n - 1 for the second, and so on down to 11. Multiplying the numbers of choices gives the total, and that product is the factorial.

Proof.

A permutation α∈Sn\alpha \in S_n is determined by the list a1=α(1),…,an=α(n)a_1 = \alpha(1), \ldots, a_n = \alpha(n), and a list arises from a permutation exactly when the aia_i are 1,…,n1, \ldots, n in some order.

Choose the entries left to right. There are nn possibilities for a1a_1. Once a1a_1 is chosen, injectivity excludes it from the rest, leaving n−1n - 1 possibilities for a2a_2; after a1a_1 and a2a_2, there are n−2n - 2 possibilities for a3a_3; and so on, with one possibility left for ana_n. Each choice is free of the others in the sense that the number available at each step does not depend on which values were taken earlier, so the number of lists is

n(n−1)(n−2)⋯2⋅1=n! .n (n-1)(n-2) \cdots 2 \cdot 1 = n!\,.

For the triangle there is no constraint at all on how the vertices may be moved: every rearrangement of the three of them is produced by exactly one symmetry. There are 3!=63! = 6 rearrangements and six symmetries, so the symmetry group of the equilateral triangle is the group of all permutations of three objects,

D3≅S3.D_3 \cong S_3 .

The two sides describe different things, movements of a figure on one and rearrangements of three labels on the other, and the isomorphism matches them up.

The square is different. It has eight symmetries while S4S_4 has 4!=244! = 24 elements, and the remark above on shuffles of the corners exhibits a rearrangement that no symmetry produces.

Remark (Where $S_4$ does live).

No figure in the plane has symmetry group S4S_4. In three dimensions it does: S4S_4 is the group of symmetries of a regular tetrahedron, whose four vertices may be permuted in any way at all.

Problem 1.15.

Let XX be a finite set and let α:X→X\alpha : X \to X be a function. Show that α\alpha is injective if and only if it is surjective, and give an example of an infinite XX for which this fails.

Problem 1.16.

In S6S_6, let

α=(123456234561),β=(123456165432).\alpha = \begin{pmatrix} 1 & 2 & 3 & 4 & 5 & 6 \\ 2 & 3 & 4 & 5 & 6 & 1 \end{pmatrix}, \qquad \beta = \begin{pmatrix} 1 & 2 & 3 & 4 & 5 & 6 \\ 1 & 6 & 5 & 4 & 3 & 2 \end{pmatrix}.

Compute αβ\alpha\beta, βα\beta\alpha, α−1\alpha^{-1}, β−1\beta^{-1} and αβα−1\alpha\beta\alpha^{-1}, and write each in cycle notation. Then label the vertices of a regular hexagon 11 to 66 clockwise and identify each of these permutations with a symmetry of the hexagon.

Problem 1.17.

Write out the eight symmetries of the square as permutations of {c1,c2,c3,c4}\{c_1, c_2, c_3, c_4\} in cycle notation, and use the result to exhibit an injective map D4→S4D_4 \to S_4 carrying products to products.

Fields

A group carries one binary operation. The number systems we actually compute in carry two, addition and multiplication, and multiplication distributes over addition. A set with two such operations is a field if it satisfies the following.

Definition 1.58 (Field).

Let FF be a set with two binary operations +:F×F→F+ : F \times F \to F and ⋅:F×F→F\cdot : F \times F \to F. The triple (F,+,⋅)(F, +, \cdot) is a field if

  1. (F1) (F,+)(F, +) is an abelian group, with neutral element written 0F0_F;
  2. (F2) (F∖{0F},⋅)(F \setminus \{0_F\}, \cdot) is an abelian group, with neutral element written 1F1_F;
  3. (F3) multiplication distributes over addition: a⋅(b+c)=a⋅b+a⋅ca \cdot (b + c) = a \cdot b + a \cdot c for all a,b,c∈Fa, b, c \in F.

Condition (F2) says two things at once: every element other than 0F0_F has a multiplicative inverse, and F∖{0F}F \setminus \{0_F\} is closed under multiplication, so a product of two non-zero elements is never 0F0_F.

Since 1F1_F lies in F∖{0F}F \setminus \{0_F\} we have 1F≠0F1_F \neq 0_F, so a field has at least two elements. Two is achievable: the set {0,1}\{0, 1\} with 1+1=01 + 1 = 0 and the obvious multiplication is a field.

Example 1.59 (Fields and near misses).

The rationals Q\mathbb{Q} and the reals R\mathbb{R} are fields, and so are the complex numbers.

The naturals N\mathbb{N} are not: (F1) already fails, since 11 has no additive inverse. The integers Z\mathbb{Z} are not either, though they come closer. They satisfy (F1) and (F3), and multiplication on them is associative and commutative with neutral element 11; what fails is (F2), because no integer other than 11 and −1-1 has a multiplicative inverse in Z\mathbb{Z}.

Remark (The complex numbers).

MA01 does not build C\mathbb{C}, so here is what we take it to be. Its elements are the expressions x+yix + yi with x,y∈Rx, y \in \mathbb{R}, where ii is a formal symbol, added and multiplied by

(x+yi)+(z+wi)=(x+z)+(y+w)i,(x+yi)(z+wi)=(xz−yw)+(xw+yz)i,\begin{aligned} (x + yi) + (z + wi) &= (x + z) + (y + w)i, \\ (x + yi)(z + wi) &= (xz - yw) + (xw + yz)i, \end{aligned}

the second rule being what expanding the brackets gives once i2i^2 is replaced by −1-1. The conjugate of z=x+yiz = x + yi is zˉ=x−yi\bar{z} = x - yi and its modulus is ∣z∣=x2+y2|z| = \sqrt{x^2 + y^2}, so that zzˉ=x2+y2=∣z∣2z\bar{z} = x^2 + y^2 = |z|^2. Every z≠0z \neq 0 therefore has the multiplicative inverse zˉ/∣z∣2\bar{z}/|z|^2, and C\mathbb{C} is a field.

Problem 1.18.

Let z,w∈Cz, w \in \mathbb{C}. Prove that

z+w‾=zˉ+wˉ,zw‾=zˉ wˉ,zˉ‾=z,zzˉ=∣z∣2,\overline{z + w} = \bar{z} + \bar{w}, \qquad \overline{zw} = \bar{z}\,\bar{w}, \qquad \overline{\bar{z}} = z, \qquad z\bar{z} = |z|^2,

and deduce that ∣zw∣=∣z∣ ∣w∣|zw| = |z|\,|w| and that z≠0z\neq 0 has inverse zˉ/∣z∣2\bar{z}/|z|^2.

Problem 1.19.

Show that z=zˉz = \bar{z} if and only if z∈Rz \in \mathbb{R}. Then show that conjugation is an isomorphism of (C,+)(\mathbb{C}, +) onto itself, and also of (C∖{0},⋅)(\mathbb{C} \setminus \{0\}, \cdot) onto itself, and that it is its own inverse in both cases.

A subtler near miss drops commutativity of multiplication rather than invertibility.

Example 1.60 (The quaternions).

Let H\mathbb{H} be the set of expressions

a+bi+cj+dk,a,b,c,d∈R,a + bi + cj + dk, \qquad a, b, c, d \in \mathbb{R},

where ii, jj, kk are formal symbols. Addition is componentwise, and multiplication is determined by requiring it to be associative and to distribute over addition, together with the rules

i2=j2=k2=ijk=−1.i^2 = j^2 = k^2 = ijk = -1 .

From these one finds ij=kij = k and ji=−kji = -k, and similar relations among the other pairs. So multiplication on H\mathbb{H} is not commutative, and H\mathbb{H} satisfies every field axiom except that one. A structure of this kind, where the non-zero elements form a group under multiplication which need not be abelian, is a skew field or division ring.

Problem 1.20.

Show that ij=kij = k and ji=−kji = -k in H\mathbb{H}. Suggestion: first check that i−1=−ii^{-1} = -i, j−1=−jj^{-1} = -j and k−1=−kk^{-1} = -k, then read ijk=−1ijk = -1 as an equation to be solved for ijij, and use the rule for the inverse of a product to get at jiji.

Vectors are built over a field of scalars, and the first chapter of this lesson is the case F=RF = \mathbb{R} of the following.

Definition 1.61 (Vector space).

Let (F,+,⋅)(F, +, \cdot) be a field. An FF-vector space is a set VV with an addition V×V→VV \times V \to V and a scalar multiplication F×V→VF \times V \to V such that

  1. (V,+)(V, +) is an abelian group, with neutral element 0\mathbf{0};
  2. λ(μv)=(λμ)v\lambda(\mu v) = (\lambda\mu)v for all λ,μ∈F\lambda, \mu \in F and v∈Vv \in V;
  3. (λ+μ)v=λv+μv(\lambda + \mu)v = \lambda v + \mu v and λ(v+w)=λv+λw\lambda(v + w) = \lambda v + \lambda w for all λ,μ∈F\lambda, \mu \in F and v,w∈Vv, w \in V;
  4. 1F v=v1_F \, v = v for every v∈Vv \in V.

Elements of VV are vectors and elements of FF are scalars.

The last condition cannot be dropped: without it the rule sending every (λ,v)(\lambda, v) to 0\mathbf{0} would satisfy the other three. With the four in place, any calculation in VV reduces to a linear combination

λ1v1+λ2v2+⋯+λnvn,λi∈F,  vi∈V,\lambda_1 v_1 + \lambda_2 v_2 + \cdots + \lambda_n v_n, \qquad \lambda_i \in F, \; v_i \in V,

which is the shape every computation in the first chapter took. That Rn\mathbb{R}^n satisfies these axioms is what the first problem of that chapter asked you to check.

The Real Numbers

We take the reals as known from school: a field carrying an order relation ⩽\leqslant compatible with the two operations, so that a⩽ba \leqslant b implies a+c⩽b+ca + c \leqslant b + c, and a⩽ba \leqslant b with 0⩽c0 \leqslant c implies ac⩽bcac \leqslant bc.

The rationals are a field with an order as well, so the order alone does not separate the two. What separates them is a property of how the order behaves.

For a,b,c∈Ra, b, c \in \mathbb{R} we write a<b<ca < b < c as shorthand for ”a<ba < b and b<cb < c”. Between any two distinct reals there is a third, namely their average.

Definition 1.62 (Intervals).

Let a,b∈Ra, b \in \mathbb{R} with a⩽ba \leqslant b. The closed interval and the open interval are

[a,b]={x∈R∣a⩽x⩽b},(a,b)={x∈R∣a<x<b},[a, b] = \{x \in \mathbb{R} \mid a \leqslant x \leqslant b\}, \qquad (a, b) = \{x \in \mathbb{R} \mid a < x < b\},

the first a single point when a=ba = b and the second empty then. The half-open intervals are

(a,b]={x∈R∣a<x⩽b},[a,b)={x∈R∣a⩽x<b},(a, b] = \{x \in \mathbb{R} \mid a < x \leqslant b\}, \qquad [a, b) = \{x \in \mathbb{R} \mid a \leqslant x < b\},

again empty when a=ba = b. The half-infinite intervals are

(−∞,a)={x∈R∣x<a},(a,∞)={x∈R∣x>a},(-\infty, a) = \{x \in \mathbb{R} \mid x < a\}, \quad (a, \infty) = \{x \in \mathbb{R} \mid x > a\},

together with (−∞,a](-\infty, a] and [a,∞)[a, \infty), defined with ⩽\leqslant and ⩾\geqslant in place of << and >>.

Remark (Reading the notation).

Three things about these symbols.

The ∞\infty and −∞-\infty are notation and nothing else. They do not name elements of R\mathbb{R}, and there is no such set as [a,∞][a, \infty].

The notation (a,b)(a, b) for an open interval collides with the notation for an ordered pair. Some authors avoid this by writing the open ends with reversed square brackets, ]a,b[]a, b[ and [a,∞[[a, \infty[. We tolerate the ambiguity, since context always says whether a subset of R\mathbb{R} or an element of R×R\mathbb{R} \times \mathbb{R} is meant.

If a<ba < b then (a,b)(a, b) contains infinitely many reals, since the averaging remark above produces a new one from any two. The same is true of the other bounded intervals, which contain (a,b)(a, b).

Finally, the notation works over any ordered number system, not only R\mathbb{R}. Where it is not clear from context we write [0,1]R[0, 1]_{\mathbb{R}} or [0,1]Q[0, 1]_{\mathbb{Q}}.

What does it mean for a subset of R\mathbb{R} to have a largest element? We want an m∈Rm \in \mathbb{R} with

  1. m∈Sm \in S, and
  2. x⩽mx \leqslant m for every x∈Sx \in S.

An mm satisfying the second condition alone is called an upper bound for SS, and SS is bounded above if it has one. Lower bounds and bounded below are defined the same way with the inequality reversed, and SS is bounded if it is both.

Proposition 1.63 (A largest element is unique).

Let S⊂RS \subset \mathbb{R}. If mm and m′m' both satisfy conditions 1 and 2, then m=m′m = m'.

Discussion.

The conditions come in a pair, one saying the candidate belongs to SS and one saying it dominates SS, and the proof combines the membership of each with the upper-bound property of the other. This gives the two inequalities m⩽m′m \leqslant m' and m′⩽mm' \leqslant m, and antisymmetry of the order turns the pair into an equality. Nothing about R\mathbb{R} beyond the order is used, so the same argument works in any ordered set.

Proof.

Since m∈Sm \in S and m′m' is an upper bound for SS, we have m⩽m′m \leqslant m'. Since m′∈Sm' \in S and mm is an upper bound for SS, we have m′⩽mm' \leqslant m. Hence m=m′m = m'.

We may therefore speak of the largest element of SS and write it max⁡(S)\max(S). What we may not do is write max⁡(S)\max(S) into an argument before knowing it exists, because many sets have no largest element. This can fail for three separate reasons.

A set may be empty, so that no mm can satisfy the first condition. A set may be unbounded above, so that no mm satisfies the second; R\mathbb{R} itself is one, and so is

S=[0,1)∪(2,3]∪[4,5)∪(6,7]∪⋯ .S = [0, 1) \cup (2, 3] \cup [4, 5) \cup (6, 7] \cup \cdots .

And a set may be non-empty and bounded above and still have no element meeting both conditions at once.

Example 1.64 (A bounded set with no largest element).

Let S=(−∞,1)S = (-\infty, 1) and let m∈Sm \in S, so m<1m < 1. Put m′=m+12m' = \tfrac{m + 1}{2}. Then m<m′<1m < m' < 1, so m′∈Sm' \in S and m′>mm' > m; hence mm is not an upper bound for SS. As mm was arbitrary, SS has no largest element, even though 11 is an upper bound for it.

Everything said about largest elements applies to smallest ones with the inequalities reversed: a smallest element is unique when it exists, is written min⁡(S)\min(S), and need not exist. For a finite non-empty set there is no difficulty at all, which is why max⁡(x,y)\max(x, y), meaning the larger of xx and yy, may be written down freely.

Problem 1.21.

Show that every non-empty subset of N\mathbb{N} has a smallest element. (Harder.)

Remark (Least upper bounds).

In the example above, 11 clearly acts as an upper limit of S=(−∞,1)S = (-\infty, 1), and the reason SS has no largest element is that 11 was left out. The following construction makes this precise. For S⊂RS \subset \mathbb{R} let

UB⁡(S)={m∈R∣x⩽m for every x∈S}\operatorname{UB}(S) = \{m \in \mathbb{R} \mid x \leqslant m \text{ for every } x \in S\}

be the set of upper bounds of SS. If SS has a largest element then UB⁡(S)\operatorname{UB}(S) has a smallest one and min⁡(UB⁡(S))=max⁡(S)\min(\operatorname{UB}(S)) = \max(S). But min⁡(UB⁡(S))\min(\operatorname{UB}(S)) can exist when max⁡(S)\max(S) does not: for S=(−∞,1)S = (-\infty, 1) we get UB⁡(S)=[1,∞)\operatorname{UB}(S) = [1, \infty), whose smallest element is 11. That number is the least upper bound, or supremum, of SS.

The property that separates R\mathbb{R} from Q\mathbb{Q} is that in R\mathbb{R} this never fails: every non-empty subset bounded above has a least upper bound. In Q\mathbb{Q} it does fail, as the set of rationals with square less than 22 shows. Making that statement into a construction of R\mathbb{R} is a course in itself.

Problem 1.22.

Let S⊂RS \subset \mathbb{R} be non-empty. Show that if max⁡(S)\max(S) exists then min⁡(UB⁡(S))\min(\operatorname{UB}(S)) exists and the two are equal.

Absolute Value

Definition 1.65 (Absolute value).

The absolute value is the function ∣⋅∣:R→R|\cdot| : \mathbb{R} \to \mathbb{R} given by

∣x∣={xif x⩾0,−xif x<0,|x| = \begin{cases} x & \text{if } x \geqslant 0, \\ -x & \text{if } x < 0, \end{cases}

equivalently ∣x∣=max⁡(x,−x)|x| = \max(x, -x).

Theorem 1.66 (Properties of the absolute value).

For all x,y∈Rx, y \in \mathbb{R}:

  1. ∣x∣⩾0|x| \geqslant 0, and ∣x∣=0|x| = 0 if and only if x=0x = 0 (positive definiteness);
  2. ∣xy∣=∣x∣ ∣y∣|xy| = |x| \, |y| (homogeneity);
  3. ∣x+y∣⩽∣x∣+∣y∣|x + y| \leqslant |x| + |y| (triangle inequality).

Discussion.

The first part is read straight off the definition, one case at a time. The second could also be done by cases, four of them, but there is a shorter route: both sides are non-negative by the first part, and two non-negative reals with equal squares are equal, so it is enough to check that the squares agree, which they do because ∣t∣2=t2|t|^2 = t^2 for every tt. The third rests on the two inequalities −∣t∣⩽t⩽∣t∣-|t| \leqslant t \leqslant |t|, which hold by inspection of the definition; adding the versions for xx and for yy traps x+yx + y between −(∣x∣+∣y∣)-(|x|+|y|) and ∣x∣+∣y∣|x|+|y|, which is the inequality ∣x+y∣⩽∣x∣+∣y∣|x+y| \leqslant |x|+|y|.

Proof.

For the first part, if x⩾0x \geqslant 0 then ∣x∣=x⩾0|x| = x \geqslant 0, and if x<0x < 0 then ∣x∣=−x>0|x| = -x > 0; so ∣x∣⩾0|x| \geqslant 0 always, and ∣x∣=0|x| = 0 forces the first case with x=0x = 0. Conversely ∣0∣=0|0| = 0.

For the second, note that ∣t∣2=t2|t|^2 = t^2 for every real tt, since ∣t∣|t| is tt or −t-t. Hence

∣xy∣2=(xy)2=x2y2=∣x∣2 ∣y∣2=(∣x∣ ∣y∣)2.|xy|^2 = (xy)^2 = x^2 y^2 = |x|^2 \, |y|^2 = \bigl(|x| \, |y|\bigr)^2 .

Both ∣xy∣|xy| and ∣x∣∣y∣|x||y| are non-negative by the first part, and for non-negative reals equal squares give equal values, so ∣xy∣=∣x∣∣y∣|xy| = |x||y|.

For the third, −∣t∣⩽t⩽∣t∣-|t| \leqslant t \leqslant |t| holds for every tt: if t⩾0t \geqslant 0 the right-hand inequality is an equality and the left is clear, and if t<0t < 0 the two swap roles. Adding the inequalities for xx and for yy,

−(∣x∣+∣y∣)⩽x+y⩽∣x∣+∣y∣.-\bigl(|x| + |y|\bigr) \leqslant x + y \leqslant |x| + |y| .

Now ∣x+y∣|x+y| is either x+yx + y or −(x+y)-(x+y), and both are bounded above by ∣x∣+∣y∣|x| + |y| by the two halves of the display. Hence ∣x+y∣⩽∣x∣+∣y∣|x+y| \leqslant |x| + |y|.

Remark (Distance on the line).

Setting d(x,y)=∣x−y∣d(x, y) = |x - y| turns the absolute value into a measure of distance between points of R\mathbb{R}, and the three properties above become exactly the three properties of the Euclidean distance proved in the first chapter. A function d:X×X→Rd : X \times X \to \mathbb{R} with those properties is called a metric, and ∣⋅∣|\cdot| on R\mathbb{R} is the case n=1n = 1 of the norm on Rn\mathbb{R}^n.

Problem 1.23.

Show that E⊂RE \subset \mathbb{R} is bounded if and only if there is an M>0M > 0 with ∣x∣<M|x| < M for every x∈Ex \in E. Then give examples of subsets of R\mathbb{R} that are bounded above only, bounded below only, and unbounded, with bounds where they exist.

Problem 1.24.

Sketch the set E={x∈R∣∣x−3∣⩽6}E = \{x \in \mathbb{R} \mid |x - 3| \leqslant 6\} and write it as an interval.

Problem 1.25.

Let x,y∈Rx, y \in \mathbb{R} with y≠0y \neq 0. Show that ∣xy∣=∣x∣∣y∣\left| \dfrac{x}{y} \right| = \dfrac{|x|}{|y|}, and that ∣x−y∣+∣x+y∣=2max⁡(∣x∣,∣y∣)|x - y| + |x + y| = 2\max(|x|, |y|).

Useful Inequalities

Theorem 1.67 (Reverse triangle inequality).

For all x,y∈Rx, y \in \mathbb{R},

∣x−y∣⩾∣x∣−∣y∣and∣x+y∣⩾∣x∣−∣y∣.|x - y| \geqslant |x| - |y| \qquad \text{and} \qquad |x + y| \geqslant |x| - |y| .

Discussion.

Each claim bounds a quantity from below, and the triangle inequality bounds things from above, so we write xx as a sum of the pieces the statement mentions. Taking x=(x−y)+yx = (x - y) + y and applying the triangle inequality to that sum gives ∣x∣⩽∣x−y∣+∣y∣|x| \leqslant |x-y| + |y|, which rearranges into the first claim. The second is the same argument with yy replaced by −y-y, using ∣−y∣=∣y∣|-y| = |y|, which is the homogeneity part of the previous theorem with the factor −1-1.

Proof.

Write x=(x−y)+yx = (x - y) + y and apply the triangle inequality:

∣x∣⩽∣x−y∣+∣y∣,|x| \leqslant |x - y| + |y|,

so ∣x−y∣⩾∣x∣−∣y∣|x - y| \geqslant |x| - |y|. Replacing yy by −y-y throughout and using ∣−y∣=∣y∣|-y| = |y| gives ∣x∣⩽∣x+y∣+∣y∣|x| \leqslant |x + y| + |y|, hence ∣x+y∣⩾∣x∣−∣y∣|x + y| \geqslant |x| - |y|.

Theorem 1.68 (Bernoulli's inequality).

Let x∈Rx \in \mathbb{R} with x⩾−1x \geqslant -1. Then (1+x)n⩾1+nx(1 + x)^n \geqslant 1 + nx for every n∈Nn \in \mathbb{N}.

Discussion.

The exponent ranges over N\mathbb{N} and appears on both sides, so the proof is an induction on nn. The base case is an equality. In the step we multiply the inductive hypothesis by 1+x1 + x, which is legitimate precisely because x⩾−1x \geqslant -1 makes that factor non-negative and so preserves the inequality; this is the only place the hypothesis on xx is used, and the inequality is false without it. Expanding the product leaves an extra term nx2nx^2, which is non-negative and can be dropped to reach the claim at n+1n + 1.

Proof.

For n=1n = 1 both sides equal 1+x1 + x. Suppose (1+x)n⩾1+nx(1 + x)^n \geqslant 1 + nx. Since x⩾−1x \geqslant -1 we have 1+x⩾01 + x \geqslant 0, so multiplying the inductive hypothesis by 1+x1 + x preserves the inequality:

(1+x)n+1⩾(1+nx)(1+x)=1+(n+1)x+nx2⩾1+(n+1)x,(1 + x)^{n+1} \geqslant (1 + nx)(1 + x) = 1 + (n+1)x + nx^2 \geqslant 1 + (n+1)x,

the last step because nx2⩾0nx^2 \geqslant 0. This is the claim at n+1n + 1.

Theorem 1.69 (Young's inequality).

Let a,b∈Ra, b \in \mathbb{R} and ε>0\varepsilon > 0. Then

ab⩽12(a2+b2),ab⩽12(a2ε+εb2),∣ab∣⩽12(a2+b2).ab \leqslant \tfrac{1}{2}\bigl(a^2 + b^2\bigr), \qquad ab \leqslant \tfrac{1}{2}\Bigl(\tfrac{a^2}{\varepsilon} + \varepsilon b^2\Bigr), \qquad |ab| \leqslant \tfrac{1}{2}\bigl(a^2 + b^2\bigr).

Discussion.

All three follow from the fact that a square is never negative. Expanding (a−b)2⩾0(a - b)^2 \geqslant 0 produces the first. For the second, the weights ε\varepsilon and 1/ε1/\varepsilon have to appear, so the square to expand is that of a/ε−ε ba/\sqrt{\varepsilon} - \sqrt{\varepsilon}\,b, whose cross term is again −2ab-2ab while the outer terms carry the weights. The third is the first applied twice: expanding (a+b)2⩾0(a + b)^2 \geqslant 0 instead bounds −ab-ab by the same quantity, and a number bounded above together with its negative is exactly a number whose absolute value is bounded.

Proof.

From 0⩽(a−b)2=a2−2ab+b20 \leqslant (a - b)^2 = a^2 - 2ab + b^2 we get 2ab⩽a2+b22ab \leqslant a^2 + b^2, which is the first inequality. Since ε>0\varepsilon > 0,

0⩽(aε−ε b)2=a2ε−2ab+εb2,0 \leqslant \Bigl(\tfrac{a}{\sqrt{\varepsilon}} - \sqrt{\varepsilon}\,b\Bigr)^2 = \tfrac{a^2}{\varepsilon} - 2ab + \varepsilon b^2,

which is the second. From 0⩽(a+b)2=a2+2ab+b20 \leqslant (a + b)^2 = a^2 + 2ab + b^2 we get −ab⩽12(a2+b2)-ab \leqslant \tfrac{1}{2}(a^2 + b^2); combined with the first inequality, both abab and −ab-ab are at most 12(a2+b2)\tfrac{1}{2}(a^2+b^2), and ∣ab∣|ab| is one of them.

Theorem 1.70 (Arithmetic and geometric mean).

Let a,b∈Ra, b \in \mathbb{R} with a,b⩾0a, b \geqslant 0. Then ab⩽12(a+b)\sqrt{ab} \leqslant \tfrac{1}{2}(a + b).

Discussion.

The statement is Young’s first inequality with a\sqrt{a} and b\sqrt{b} in place of aa and bb: since aa and bb are non-negative they have square roots, and substituting a\sqrt{a} and b\sqrt{b} for the two variables there turns the product into ab\sqrt{ab} and the squares into aa and bb. Equivalently one may expand (a−b)2⩾0(\sqrt{a} - \sqrt{b})^2 \geqslant 0 directly, which is the same computation written out.

Proof.

Both aa and bb are non-negative, so a\sqrt{a} and b\sqrt{b} exist. Then

0⩽(a−b)2=a−2ab+b,0 \leqslant \bigl(\sqrt{a} - \sqrt{b}\bigr)^2 = a - 2\sqrt{ab} + b,

which rearranges to ab⩽12(a+b)\sqrt{ab} \leqslant \tfrac{1}{2}(a + b).

Standard Functions

The simplest real-valued functions are built from the field operations alone. A polynomial function is one of the form x↦cnxn+⋯+c1x+c0x \mapsto c_n x^n + \cdots + c_1 x + c_0, such as x3−17x+5x^3 - 17x + 5; later in the course we approximate arbitrary functions by these. A rational function is a quotient of two polynomials, such as

xx2+3andx3+7x+2.\frac{x}{x^2 + 3} \qquad \text{and} \qquad \frac{x^3 + 7}{x + 2}.

The second is not a function on R\mathbb{R}, since division by zero is undefined and the denominator vanishes at x=−2x = -2; it is a function on the subset {x∈R∣x≠−2}\{x \in \mathbb{R} \mid x \neq -2\}.

Definition 1.71 (Exponential function).

The exponential function is the unique differentiable exp⁡:R→R\exp : \mathbb{R} \to \mathbb{R} with

exp⁡(0)=1andexp⁡′(x)=exp⁡(x) for every x∈R.\exp(0) = 1 \qquad \text{and} \qquad \exp'(x) = \exp(x) \text{ for every } x \in \mathbb{R}.

We write e=exp⁡(1)=2.71828…e = \exp(1) = 2.71828\ldots, and exe^x for exp⁡(x)\exp(x).

That such a function exists and is unique is proved in analysis; we take it, and the addition law

exp⁡(a+b)=exp⁡(a)exp⁡(b)for all a,b∈R,\exp(a + b) = \exp(a) \exp(b) \qquad \text{for all } a, b \in \mathbb{R},

as given.

xy11exp x
Figure 1.12. The natural exponential function. It is positive everywhere, strictly increasing, and takes the value 11 at 00.

Proposition 1.72 (Rules for the exponential).

For a,b∈Ra, b \in \mathbb{R} and n∈Nn \in \mathbb{N},

exp⁡(−a)=1exp⁡a,exp⁡(a−b)=exp⁡aexp⁡b,exp⁡(na)=(exp⁡a)n,\exp(-a) = \frac{1}{\exp a}, \qquad \exp(a - b) = \frac{\exp a}{\exp b}, \qquad \exp(na) = (\exp a)^n,

and exp⁡(x)>0\exp(x) > 0 for every xx. The function exp⁡\exp is strictly increasing, and is a bijection from R\mathbb{R} onto (0,∞)(0, \infty).

Discussion.

Every rule here is the addition law used once or repeatedly. Setting b=−ab = -a in it makes the left side exp⁡(0)=1\exp(0) = 1, which identifies exp⁡(−a)\exp(-a) as a reciprocal and in passing shows exp⁡\exp never vanishes; the quotient rule is then the addition law applied to a+(−b)a + (-b). The rule for exp⁡(na)\exp(na) is an induction whose step is one more application. Positivity needs a separate observation: exp⁡(x)\exp(x) is the square of exp⁡(x/2)\exp(x/2), hence non-negative, and it is not zero by what has just been shown. Strict increase then comes from the derivative, which equals the function and is therefore positive; and a strictly increasing function is injective, which leaves only surjectivity onto (0,∞)(0, \infty), and that is the part we take from analysis along with existence.

Proof.

Putting b=−ab = -a in the addition law gives exp⁡(a)exp⁡(−a)=exp⁡(0)=1\exp(a)\exp(-a) = \exp(0) = 1, so neither factor is zero and exp⁡(−a)=1/exp⁡(a)\exp(-a) = 1/\exp(a). Hence

exp⁡(a−b)=exp⁡(a)exp⁡(−b)=exp⁡aexp⁡b.\exp(a - b) = \exp(a)\exp(-b) = \frac{\exp a}{\exp b}.

For exp⁡(na)=(exp⁡a)n\exp(na) = (\exp a)^n, induct on nn: the case n=1n = 1 is trivial, and exp⁡((n+1)a)=exp⁡(na)exp⁡(a)=(exp⁡a)nexp⁡(a)=(exp⁡a)n+1\exp((n+1)a) = \exp(na)\exp(a) = (\exp a)^n \exp(a) = (\exp a)^{n+1}.

For positivity, exp⁡(x)=exp⁡(x2+x2)=exp⁡(x2)2⩾0\exp(x) = \exp\bigl(\tfrac{x}{2} + \tfrac{x}{2}\bigr) = \exp\bigl(\tfrac{x}{2}\bigr)^2 \geqslant 0, and exp⁡(x)≠0\exp(x) \neq 0 by the first paragraph, so exp⁡(x)>0\exp(x) > 0. Consequently exp⁡′=exp⁡>0\exp' = \exp > 0 everywhere, so exp⁡\exp is strictly increasing and therefore injective. Its range is (0,∞)(0, \infty).

Being a bijection onto (0,∞)(0, \infty), the exponential has an inverse.

Definition 1.73 (Natural logarithm).

The natural logarithm is the inverse of exp⁡:R→(0,∞)\exp : \mathbb{R} \to (0, \infty), written

log⁡:(0,∞)→R,\log : (0, \infty) \to \mathbb{R},

so that log⁡(exp⁡x)=x\log(\exp x) = x for every x∈Rx \in \mathbb{R} and exp⁡(log⁡y)=y\exp(\log y) = y for every y>0y > 0.

xy1log x
Figure 1.13. The natural logarithm, the inverse of the exponential. It is defined only for x>0x > 0, is strictly increasing, and vanishes at 11.

Proposition 1.74 (Rules for the logarithm).

For a,b>0a, b > 0 and n∈Nn \in \mathbb{N},

log⁡(ab)=log⁡a+log⁡b,log⁡ ⁣(ab)=log⁡a−log⁡b,log⁡ ⁣(1a)=−log⁡a,\log(ab) = \log a + \log b, \qquad \log\!\left(\frac{a}{b}\right) = \log a - \log b, \qquad \log\!\left(\frac{1}{a}\right) = -\log a,

together with log⁡(an)=nlog⁡a\log(a^n) = n \log a and log⁡1=0\log 1 = 0. Moreover log⁡\log is differentiable with log⁡′(x)=1/x\log'(x) = 1/x.

Discussion.

An identity about log⁡\log becomes an identity about exp⁡\exp once exp⁡\exp is applied to both sides, because exp⁡\exp is injective and undoes log⁡\log. So each rule is proved by applying exp⁡\exp to the proposed right-hand side and using the addition law. The value log⁡1=0\log 1 = 0 is exp⁡(0)=1\exp(0) = 1 read backwards. The derivative is different in kind: differentiating the identity exp⁡(log⁡x)=x\exp(\log x) = x by the chain rule produces exp⁡(log⁡x)log⁡′(x)=1\exp(\log x)\log'(x) = 1, and the left factor is xx, which solves for log⁡′(x)\log'(x).

Proof.

For a,b>0a, b > 0,

exp⁡(log⁡a+log⁡b)=exp⁡(log⁡a)exp⁡(log⁡b)=ab,\exp(\log a + \log b) = \exp(\log a)\exp(\log b) = ab,

and applying log⁡\log to both sides gives log⁡a+log⁡b=log⁡(ab)\log a + \log b = \log(ab). Replacing bb by 1/b1/b and using the same computation gives the quotient rule, and taking a=1a = 1 in it gives log⁡(1/b)=−log⁡b\log(1/b) = -\log b since log⁡1=0\log 1 = 0; and log⁡1=0\log 1 = 0 holds because exp⁡(0)=1\exp(0) = 1. The rule log⁡(an)=nlog⁡a\log(a^n) = n\log a follows by induction from the product rule.

Differentiating exp⁡(log⁡x)=x\exp(\log x) = x by the chain rule gives exp⁡(log⁡x)log⁡′(x)=1\exp(\log x)\log'(x) = 1, and exp⁡(log⁡x)=x\exp(\log x) = x, so log⁡′(x)=1/x\log'(x) = 1/x for x>0x > 0.

Remark (Which logarithm is $\log$?).

In mathematics an unadorned log⁡\log always means the natural logarithm, to base ee. Some software takes the opposite convention, using log⁡\log for the base-1010 logarithm and ln⁡\ln for the natural one; check before trusting a numerical result.

Remark (Why there is no logarithm of a negative number).

We have not defined log⁡x\log x for x⩽0x \leqslant 0, and no definition is possible. The logarithm is defined by the identity exp⁡(log⁡x)=x\exp(\log x) = x, and exp⁡\exp takes only positive values, so no value of log⁡x\log x could satisfy it for x⩽0x \leqslant 0. The same restriction is inherited by the power functions below.

Definition 1.75 (Powers and logarithms to a general base).

For a>0a > 0 define a ⋅:R→(0,∞)a^{\,\cdot} : \mathbb{R} \to (0, \infty) by

ax=exp⁡(xlog⁡a).a^x = \exp(x \log a).

It is a bijection, and its inverse is the logarithm to base aa, written log⁡a:(0,∞)→R\log_a : (0, \infty) \to \mathbb{R}.

This agrees with the elementary meaning of a power whenever that meaning is available: for y∈Zy \in \mathbb{Z} the definition returns xx multiplied by itself yy times when y>0y > 0, returns 11 when y=0y = 0, and returns 1/x−y1/x^{-y} when y<0y < 0. Its advantage is that it makes sense for every real exponent.

xy1.5x2x2.5x11
Figure 1.14. The functions x↦axx \mapsto a^x for a=1.5a = 1.5, 22 and 2.52.5. Each takes the value 11 at 00, and each is exp⁡(xlog⁡a)\exp(x \log a).
xylog10log2log1
Figure 1.15. Logarithms to base 1010, base 22 and base ee. All three vanish at 11, and each is a fixed multiple of the others.

Proposition 1.76 (Changing the base).

For a,b,x>0a, b, x > 0 with a≠1a \neq 1 and b≠1b \neq 1,

log⁡ax=log⁡xlog⁡aandlog⁡ax=log⁡ab⋅log⁡bx.\log_a x = \frac{\log x}{\log a} \qquad \text{and} \qquad \log_a x = \log_a b \cdot \log_b x .

Discussion.

The base-aa logarithm was defined as an inverse, so we compute it by applying the function it inverts. Writing y=log⁡axy = \log_a x means ay=xa^y = x, which by the definition of a power is exp⁡(ylog⁡a)=x\exp(y \log a) = x; taking log⁡\log of both sides turns it into ylog⁡a=log⁡xy \log a = \log x, and dividing gives the first identity. The second is then arithmetic: substituting the first identity three times, once for each of the three logarithms appearing, reduces the claim to cancelling a common factor.

Proof.

Let y=log⁡axy = \log_a x, so ay=xa^y = x, that is exp⁡(ylog⁡a)=x\exp(y \log a) = x. Applying log⁡\log gives ylog⁡a=log⁡xy \log a = \log x, and log⁡a≠0\log a \neq 0 since a≠1a \neq 1, so y=log⁡x/log⁡ay = \log x / \log a.

For the second identity, substituting the first three times,

log⁡ab⋅log⁡bx=log⁡blog⁡a⋅log⁡xlog⁡b=log⁡xlog⁡a=log⁡ax.\log_a b \cdot \log_b x = \frac{\log b}{\log a} \cdot \frac{\log x}{\log b} = \frac{\log x}{\log a} = \log_a x .

Example 1.77 (Halving a fish population).

Suppose a lake holds CC fish and that fishing reduces the population according to f(t)=Ce−λtf(t) = Ce^{-\lambda t} with λ>0\lambda > 0. The time TT at which half the fish are gone satisfies

Ce−λT=C2  ⟺  e−λT=12  ⟺  −λT=log⁡12=−log⁡2  ⟺  T=log⁡2λ,Ce^{-\lambda T} = \frac{C}{2} \iff e^{-\lambda T} = \frac{1}{2} \iff -\lambda T = \log \tfrac{1}{2} = -\log 2 \iff T = \frac{\log 2}{\lambda},

which does not depend on CC.

Trigonometric Functions

Definition 1.78 (Periodic function).

A function f:X→Yf : X \to Y with X⊂RX \subset \mathbb{R} is TT-periodic for T∈RT \in \mathbb{R} if f(x+T)=f(x)f(x + T) = f(x) for every x∈Xx \in X with x+T∈Xx + T \in X.

The trigonometric functions are read off the unit circle, the circle of radius 11 centred at the origin O=(0,0)O = (0, 0). Angles are measured in radians and counted anticlockwise from the positive xx-axis, so that a full circuit is 2π2\pi.

xyαMαO11cos αsin α
Figure 1.16. The unit circle. The point MαM_\alpha has coordinates (cos⁡α,sin⁡α)(\cos\alpha, \sin\alpha), and the outer arrow shows the direction in which angles are counted.
radians2πππ2π3π4π6degrees36018090604530\begin{array}{c|cccccc} \text{radians} & 2\pi & \pi & \tfrac{\pi}{2} & \tfrac{\pi}{3} & \tfrac{\pi}{4} & \tfrac{\pi}{6} \\ \hline \text{degrees} & 360 & 180 & 90 & 60 & 45 & 30 \end{array}

The angle α\alpha may be any real number, and MαM_\alpha denotes the point of the unit circle for which the angle from the positive xx-axis to OMαOM_\alpha is α\alpha. A full circuit is 2π2\pi, so MαM_\alpha, Mα+2πM_{\alpha + 2\pi} and Mα+4πM_{\alpha + 4\pi} are the same point; it is often convenient to restrict α\alpha to (−π,π](-\pi, \pi] or to [0,2π)[0, 2\pi). We say in this situation that α\alpha is defined modulo 2π2\pi.

Definition 1.79 (Congruence modulo a real number).

Let a,b,c∈Ra, b, c \in \mathbb{R}. We say aa is equal to bb modulo cc, written a=b mod ca = b \bmod c, if

a∈{b+ck∣k∈Z},a \in \{b + ck \mid k \in \mathbb{Z}\},

a set also written b+cZb + c\mathbb{Z}.

Example 1.80 (An angle modulo 2π2\pi).

The equation α=π3 mod 2π\alpha = \tfrac{\pi}{3} \bmod 2\pi says that

α∈{π3+2kπ  ∣  k∈Z}={…,−5π3,  π3,  7π3,…}.\alpha \in \Bigl\{\tfrac{\pi}{3} + 2k\pi \;\Bigm|\; k \in \mathbb{Z}\Bigr\} = \Bigl\{\ldots, -\tfrac{5\pi}{3}, \; \tfrac{\pi}{3}, \; \tfrac{7\pi}{3}, \ldots\Bigr\}.

Definition 1.81 (Sine, cosine and tangent).

Let α∈R\alpha \in \mathbb{R} and let MαM_\alpha be as above. The cosine cos⁡α\cos\alpha is the xx-coordinate of MαM_\alpha and the sine sin⁡α\sin\alpha is its yy-coordinate. Where cos⁡α≠0\cos\alpha \neq 0, that is where α≠π2 mod π\alpha \neq \tfrac{\pi}{2} \bmod \pi, the tangent is

tan⁡α=sin⁡αcos⁡α.\tan\alpha = \frac{\sin\alpha}{\cos\alpha}.

Thus cos⁡,sin⁡:R→[−1,1]\cos, \sin : \mathbb{R} \to [-1, 1] and tan⁡:R∖{π2+kπ∣k∈Z}→R\tan : \mathbb{R} \setminus \bigl\{\tfrac{\pi}{2} + k\pi \mid k \in \mathbb{Z}\bigr\} \to \mathbb{R}.

Since a full circuit returns MαM_\alpha to itself, sin⁡\sin and cos⁡\cos are 2π2\pi-periodic; tan⁡\tan turns out to be π\pi-periodic, which the shift formulas below explain.

xy−2π−ππ2π1sincos
Figure 1.17. Sine (solid) and cosine (dashed) over two periods. Both take values in [−1,1][-1, 1] and repeat every 2π2\pi.
xy−2π−ππ2π
Figure 1.18. The tangent. It repeats every π\pi and is undefined at the dashed lines x=π2+kπx = \tfrac{\pi}{2} + k\pi.

We take the addition formulas as known,

sin⁡(x+y)=sin⁡xcos⁡y+cos⁡xsin⁡y,cos⁡(x+y)=cos⁡xcos⁡y−sin⁡xsin⁡y,\begin{aligned} \sin(x + y) &= \sin x \cos y + \cos x \sin y, \\ \cos(x + y) &= \cos x \cos y - \sin x \sin y, \end{aligned}

and derive the rest from them together with the definition.

Proposition 1.82 (Trigonometric identities).

For all real xx and yy at which the expressions are defined:

  1. cos⁡2x+sin⁡2x=1\cos^2 x + \sin^2 x = 1, and 1+tan⁡2x=1cos⁡2x1 + \tan^2 x = \dfrac{1}{\cos^2 x};
  2. sin⁡(−x)=−sin⁡x\sin(-x) = -\sin x, cos⁡(−x)=cos⁡x\cos(-x) = \cos x, tan⁡(−x)=−tan⁡x\tan(-x) = -\tan x;
  3. sin⁡(x−y)=sin⁡xcos⁡y−cos⁡xsin⁡y\sin(x - y) = \sin x \cos y - \cos x \sin y and cos⁡(x−y)=cos⁡xcos⁡y+sin⁡xsin⁡y\cos(x - y) = \cos x \cos y + \sin x \sin y;
  4. sin⁡(x+π)=−sin⁡x\sin(x + \pi) = -\sin x, cos⁡(x+π)=−cos⁡x\cos(x + \pi) = -\cos x, and tan⁡(x+π)=tan⁡x\tan(x + \pi) = \tan x;
  5. sin⁡(π−x)=sin⁡x\sin(\pi - x) = \sin x and cos⁡(π−x)=−cos⁡x\cos(\pi - x) = -\cos x;
  6. sin⁡(x+π2)=cos⁡x\sin\bigl(x + \tfrac{\pi}{2}\bigr) = \cos x and cos⁡(x+π2)=−sin⁡x\cos\bigl(x + \tfrac{\pi}{2}\bigr) = -\sin x;
  7. sin⁡(x−π2)=−cos⁡x\sin\bigl(x - \tfrac{\pi}{2}\bigr) = -\cos x and cos⁡(x−π2)=sin⁡x\cos\bigl(x - \tfrac{\pi}{2}\bigr) = \sin x;
  8. sin⁡2x=2sin⁡xcos⁡x\sin 2x = 2\sin x \cos x and cos⁡2x=cos⁡2x−sin⁡2x\cos 2x = \cos^2 x - \sin^2 x;
  9. tan⁡(x+y)=tan⁡x+tan⁡y1−tan⁡xtan⁡y\tan(x + y) = \dfrac{\tan x + \tan y}{1 - \tan x \tan y}.

Discussion.

Only the first two parts need anything beyond the addition formulas. The Pythagorean identity is the statement that MxM_x lies on the unit circle, since its coordinates are cos⁡x\cos x and sin⁡x\sin x and the circle has radius 11; dividing it by cos⁡2x\cos^2 x, where that is non-zero, gives the companion identity for the tangent. The parity statements are read off the circle as well: reflecting in the xx-axis carries MxM_x to M−xM_{-x}, which negates the second coordinate and fixes the first. Everything after that is substitution. Part 3 is the addition formulas with −y-y in place of yy, using part 2. Parts 4 to 7 are the addition formulas evaluated at the special angles π\pi and π2\tfrac{\pi}{2}, whose sines and cosines are 00 and ±1\pm 1, so most terms vanish; part 5 combines part 4 with part 2, and part 7 combines part 6 with part 2. Part 8 is the addition formulas with y=xy = x, and part 9 is part 8’s method applied to the quotient defining the tangent, dividing numerator and denominator by cos⁡xcos⁡y\cos x \cos y.

Proof.

The point Mx=(cos⁡x,sin⁡x)M_x = (\cos x, \sin x) lies on the circle of radius 11 about the origin, so cos⁡2x+sin⁡2x=1\cos^2 x + \sin^2 x = 1 by the formula for the Euclidean norm. Where cos⁡x≠0\cos x \neq 0, dividing by cos⁡2x\cos^2 x gives 1+tan⁡2x=1/cos⁡2x1 + \tan^2 x = 1/\cos^2 x. Reflecting the circle in the xx-axis sends MxM_x to M−xM_{-x} and negates the second coordinate, so cos⁡(−x)=cos⁡x\cos(-x) = \cos x and sin⁡(−x)=−sin⁡x\sin(-x) = -\sin x; the statement for tan⁡\tan follows by dividing.

Substituting −y-y for yy in the addition formulas and using the parity just proved gives part 3. Taking y=πy = \pi, where cos⁡π=−1\cos\pi = -1 and sin⁡π=0\sin\pi = 0, gives part 4, and the statement for tan⁡\tan follows by dividing the two; part 5 is part 3 with x=πx = \pi together with parity. Taking y=π2y = \tfrac{\pi}{2}, where cos⁡π2=0\cos\tfrac{\pi}{2} = 0 and sin⁡π2=1\sin\tfrac{\pi}{2} = 1, gives part 6, and part 3 with y=π2y = \tfrac{\pi}{2} gives part 7.

Taking y=xy = x in the addition formulas gives part 8. For part 9, divide

tan⁡(x+y)=sin⁡xcos⁡y+cos⁡xsin⁡ycos⁡xcos⁡y−sin⁡xsin⁡y\tan(x+y) = \frac{\sin x \cos y + \cos x \sin y}{\cos x \cos y - \sin x \sin y}

above and below by cos⁡xcos⁡y\cos x \cos y, which is non-zero wherever both tangents are defined.

xyα−απ − αα + πα + π/2O
Figure 1.19. The five angles related to α\alpha by the parity and shift identities, read off the same circle.

With the labelling of the figure below, where α\alpha, β\beta, γ\gamma are the angles at the three vertices and aa, bb, cc are the sides opposite them,

sin⁡αa=sin⁡βb=sin⁡γc(law of sines),\frac{\sin\alpha}{a} = \frac{\sin\beta}{b} = \frac{\sin\gamma}{c} \qquad \text{(law of sines)}, c2=a2+b2−2abcos⁡γ(law of cosines).c^2 = a^2 + b^2 - 2ab\cos\gamma \qquad \text{(law of cosines)}.

Between them they recover the unknown sides and angles of a triangle from the known ones. The law of cosines is proved once the angle between two vectors has been defined.

ABCαβγcab
Figure 1.20. A triangle labelled for the law of sines and the law of cosines: each side is named by the lower-case letter matching the angle opposite it.

Back to Vectors

The angle can now be brought back to Rn\mathbb{R}^n, and with it the case of equality in the triangle inequality.

Proposition 1.83 (Equality in the triangle inequality).

Let u,v∈Rn\mathbf{u}, \mathbf{v} \in \mathbb{R}^n with v≠0\mathbf{v} \neq \mathbf{0}. Then

∥u+v∥=∥u∥+∥v∥\lVert \mathbf{u} + \mathbf{v} \rVert = \lVert \mathbf{u} \rVert + \lVert \mathbf{v} \rVert

if and only if u=λv\mathbf{u} = \lambda\mathbf{v} for some real λ⩾0\lambda \geqslant 0.

Discussion.

The proof of the triangle inequality used two inequalities, first replacing u⋅v\mathbf{u}\cdot\mathbf{v} by ∣u⋅v∣|\mathbf{u}\cdot\mathbf{v}| and then bounding that by ∥u∥∥v∥\lVert\mathbf{u}\rVert\lVert\mathbf{v}\rVert. Equality at the end forces equality at both, so the condition is that u⋅v\mathbf{u}\cdot\mathbf{v} is non-negative and that Cauchy–Schwarz is tight. The Cauchy–Schwarz proposition already says when the second happens, namely when one vector is a multiple of the other, and v≠0\mathbf{v} \neq \mathbf{0} lets us take the multiple in the direction u=λv\mathbf{u} = \lambda\mathbf{v}. The sign condition then decides λ\lambda, because u⋅v=λ∥v∥2\mathbf{u}\cdot\mathbf{v} = \lambda\lVert\mathbf{v}\rVert^2 has the sign of λ\lambda. The converse is a direct computation with λ⩾0\lambda \geqslant 0, where the homogeneity of the norm produces the factor λ\lambda without an absolute value.

Proof.

Suppose first that u=λv\mathbf{u} = \lambda\mathbf{v} with λ⩾0\lambda \geqslant 0. Then

∥u+v∥=∥(λ+1)v∥=(λ+1)∥v∥=λ∥v∥+∥v∥=∥u∥+∥v∥,\lVert \mathbf{u} + \mathbf{v} \rVert = \lVert (\lambda + 1)\mathbf{v} \rVert = (\lambda + 1)\lVert \mathbf{v} \rVert = \lambda\lVert \mathbf{v} \rVert + \lVert \mathbf{v} \rVert = \lVert \mathbf{u} \rVert + \lVert \mathbf{v} \rVert,

using ∣λ+1∣=λ+1|\lambda + 1| = \lambda + 1 and ∣λ∣=λ|\lambda| = \lambda.

Conversely, suppose the norms are equal. Squaring and expanding as in the proof of the triangle inequality,

∥u∥2+2(u⋅v)+∥v∥2=∥u∥2+2∥u∥∥v∥+∥v∥2,\lVert \mathbf{u} \rVert^2 + 2(\mathbf{u} \cdot \mathbf{v}) + \lVert \mathbf{v} \rVert^2 = \lVert \mathbf{u} \rVert^2 + 2\lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert + \lVert \mathbf{v} \rVert^2,

so u⋅v=∥u∥∥v∥\mathbf{u} \cdot \mathbf{v} = \lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert. In particular ∣u⋅v∣=∥u∥∥v∥|\mathbf{u} \cdot \mathbf{v}| = \lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert, which is the case of equality in the Cauchy–Schwarz inequality, so one of u\mathbf{u}, v\mathbf{v} is a multiple of the other; as v≠0\mathbf{v} \neq \mathbf{0} we may write u=λv\mathbf{u} = \lambda\mathbf{v}. Then u⋅v=λ∥v∥2\mathbf{u} \cdot \mathbf{v} = \lambda\lVert \mathbf{v} \rVert^2 is non-negative and ∥v∥2>0\lVert \mathbf{v} \rVert^2 > 0, so λ⩾0\lambda \geqslant 0.

Restricted to [0,π][0, \pi] the cosine is a bijection onto [−1,1][-1, 1], so it has an inverse cos⁡−1:[−1,1]→[0,π]\cos^{-1} : [-1, 1] \to [0, \pi], and the Cauchy–Schwarz inequality says exactly that the quotient below is an admissible input to it.

Definition 1.84 (Angle between vectors).

Let u,v∈Rn\mathbf{u}, \mathbf{v} \in \mathbb{R}^n be non-zero. The angle between them is

θ=cos⁡−1 ⁣(u⋅v∥u∥ ∥v∥)∈[0,π].\theta = \cos^{-1}\!\left( \frac{\mathbf{u} \cdot \mathbf{v}}{\lVert \mathbf{u} \rVert \, \lVert \mathbf{v} \rVert} \right) \in [0, \pi].

The vectors are orthogonal, or perpendicular, if u⋅v=0\mathbf{u} \cdot \mathbf{v} = 0, that is if θ=π2\theta = \tfrac{\pi}{2}.

The definition is legitimate because Cauchy–Schwarz gives ∣u⋅v∣⩽∥u∥∥v∥|\mathbf{u} \cdot \mathbf{v}| \leqslant \lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert, so the quotient lies in [−1,1][-1, 1]. Taking the value of cos⁡−1\cos^{-1} in [0,π][0, \pi] measures the smaller of the two angles between the vectors. For n=2n = 2 this agrees with the angle read off the unit circle.

Corollary 1.85 (Law of cosines).

Let u,v∈Rn\mathbf{u}, \mathbf{v} \in \mathbb{R}^n be non-zero, with angle θ\theta between them. Then

∥u−v∥2=∥u∥2+∥v∥2−2∥u∥∥v∥cos⁡θ.\lVert \mathbf{u} - \mathbf{v} \rVert^2 = \lVert \mathbf{u} \rVert^2 + \lVert \mathbf{v} \rVert^2 - 2 \lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert \cos\theta .

Taking u\mathbf{u} and v\mathbf{v} to be two sides of a triangle issuing from the vertex where the angle is γ\gamma, this is the law of cosines c2=a2+b2−2abcos⁡γc^2 = a^2 + b^2 - 2ab\cos\gamma.

Proof.

By bilinearity of the scalar product,

∥u−v∥2=(u−v)⋅(u−v)=∥u∥2−2(u⋅v)+∥v∥2,\lVert \mathbf{u} - \mathbf{v} \rVert^2 = (\mathbf{u} - \mathbf{v}) \cdot (\mathbf{u} - \mathbf{v}) = \lVert \mathbf{u} \rVert^2 - 2(\mathbf{u} \cdot \mathbf{v}) + \lVert \mathbf{v} \rVert^2,

and the definition of the angle gives u⋅v=∥u∥∥v∥cos⁡θ\mathbf{u} \cdot \mathbf{v} = \lVert \mathbf{u} \rVert \lVert \mathbf{v} \rVert \cos\theta. The third side of the triangle with sides u\mathbf{u} and v\mathbf{v} is u−v\mathbf{u} - \mathbf{v}, whose length is cc, while ∥u∥\lVert \mathbf{u} \rVert and ∥v∥\lVert \mathbf{v} \rVert are bb and aa.

Example 1.86 (An angle in R4\mathbb{R}^4).

Let u=(1,2,3,1)\mathbf{u} = (1, 2, 3, 1) and v=(−1,0,2,3)\mathbf{v} = (-1, 0, 2, 3). Their norms are

∥u∥2=1+4+9+1=15,∥v∥2=1+0+4+9=14,\lVert \mathbf{u} \rVert^2 = 1 + 4 + 9 + 1 = 15, \qquad \lVert \mathbf{v} \rVert^2 = 1 + 0 + 4 + 9 = 14,

so ∥u∥=15\lVert \mathbf{u} \rVert = \sqrt{15} and ∥v∥=14\lVert \mathbf{v} \rVert = \sqrt{14}. Their scalar product is

u⋅v=1⋅(−1)+2⋅0+3⋅2+1⋅3=8,\mathbf{u} \cdot \mathbf{v} = 1 \cdot (-1) + 2 \cdot 0 + 3 \cdot 2 + 1 \cdot 3 = 8,

so the angle between them is

θ=cos⁡−1 ⁣(81514)=cos⁡−1 ⁣(8210)≈0.986 radians.\theta = \cos^{-1}\!\left( \frac{8}{\sqrt{15}\sqrt{14}} \right) = \cos^{-1}\!\left( \frac{8}{\sqrt{210}} \right) \approx 0.986 \text{ radians}.

Proposition 1.87 (Orthogonal projection).

Let u,v∈Rn\mathbf{u}, \mathbf{v} \in \mathbb{R}^n with v≠0\mathbf{v} \neq \mathbf{0}. There is exactly one real λ\lambda for which u−λv\mathbf{u} - \lambda\mathbf{v} is orthogonal to v\mathbf{v}, namely

λ=u⋅v∥v∥2,\lambda = \frac{\mathbf{u} \cdot \mathbf{v}}{\lVert \mathbf{v} \rVert^2},

and the resulting decomposition is

u=λv+(u−λv).\mathbf{u} = \lambda\mathbf{v} + (\mathbf{u} - \lambda\mathbf{v}).

Discussion.

The condition to be met is a single scalar equation, (u−λv)⋅v=0(\mathbf{u} - \lambda\mathbf{v}) \cdot \mathbf{v} = 0, in the single unknown λ\lambda. Bilinearity expands its left side into u⋅v−λ∥v∥2\mathbf{u}\cdot\mathbf{v} - \lambda\lVert\mathbf{v}\rVert^2, which is a linear expression in λ\lambda with coefficient −∥v∥2-\lVert\mathbf{v}\rVert^2; that coefficient is non-zero exactly because v≠0\mathbf{v} \neq \mathbf{0}, so the equation has exactly one solution and existence and uniqueness are settled together. The decomposition is then a matter of adding and subtracting the same vector.

Proof.

For λ∈R\lambda \in \mathbb{R}, bilinearity of the scalar product gives

(u−λv)⋅v=u⋅v−λ(v⋅v)=u⋅v−λ∥v∥2.(\mathbf{u} - \lambda\mathbf{v}) \cdot \mathbf{v} = \mathbf{u} \cdot \mathbf{v} - \lambda (\mathbf{v} \cdot \mathbf{v}) = \mathbf{u} \cdot \mathbf{v} - \lambda \lVert \mathbf{v} \rVert^2 .

Since v≠0\mathbf{v} \neq \mathbf{0} we have ∥v∥2>0\lVert \mathbf{v} \rVert^2 > 0, so this vanishes for exactly one λ\lambda, namely λ=(u⋅v)/∥v∥2\lambda = (\mathbf{u} \cdot \mathbf{v}) / \lVert \mathbf{v} \rVert^2. Writing u=λv+(u−λv)\mathbf{u} = \lambda\mathbf{v} + (\mathbf{u} - \lambda\mathbf{v}) is then an identity.

Definition 1.88 (Vector projection).

With λ\lambda as in the last proposition, the vector

proj⁡vu=λv=u⋅v∥v∥2 v\operatorname{proj}_{\mathbf{v}} \mathbf{u} = \lambda\mathbf{v} = \frac{\mathbf{u} \cdot \mathbf{v}}{\lVert \mathbf{v} \rVert^2}\,\mathbf{v}

is the component of u\mathbf{u} in the direction of v\mathbf{v}, or the vector projection of u\mathbf{u} onto v\mathbf{v}, and u−λv\mathbf{u} - \lambda\mathbf{v} is the component of u\mathbf{u} perpendicular to v\mathbf{v}.

Ovuλvu − λv
Figure 1.21. The components of u\mathbf{u} relative to v\mathbf{v}: the projection λv\lambda\mathbf{v} along v\mathbf{v}, and u−λv\mathbf{u} - \lambda\mathbf{v} at right angles to it.

Problem 1.26.

Let u=(2,−1,2)\mathbf{u} = (2, -1, 2) and v=(1,2,2)\mathbf{v} = (1, 2, 2) in R3\mathbb{R}^3. Compute the angle between them, the projection of u\mathbf{u} onto v\mathbf{v}, and the component of u\mathbf{u} perpendicular to v\mathbf{v}, and verify that the two components are orthogonal.

Problem 1.27.

Let v∈Rn\mathbf{v} \in \mathbb{R}^n be non-zero. Show that proj⁡v\operatorname{proj}_{\mathbf{v}} is unchanged when v\mathbf{v} is replaced by μv\mu\mathbf{v} for any real μ≠0\mu \neq 0.

Problem 1.28.

With λ\lambda as in the last proposition, prove that

∥u∥2=∥λv∥2+∥u−λv∥2,\lVert \mathbf{u} \rVert^2 = \lVert \lambda\mathbf{v} \rVert^2 + \lVert \mathbf{u} - \lambda\mathbf{v} \rVert^2,

and deduce that ∥proj⁡vu∥⩽∥u∥\lVert \operatorname{proj}_{\mathbf{v}} \mathbf{u} \rVert \leqslant \lVert \mathbf{u} \rVert.

One more construction is available in three dimensions only.

Definition 1.89 (Vector product).

For u=(u1,u2,u3)\mathbf{u} = (u_1, u_2, u_3) and v=(v1,v2,v3)\mathbf{v} = (v_1, v_2, v_3) in R3\mathbb{R}^3, the vector product, or cross product, is

u×v=(u2v3−u3v2,  u3v1−u1v3,  u1v2−u2v1)∈R3.\mathbf{u} \times \mathbf{v} = (u_2 v_3 - u_3 v_2, \; u_3 v_1 - u_1 v_3, \; u_1 v_2 - u_2 v_1) \in \mathbb{R}^3 .

Unlike the scalar product it returns a vector rather than a number, and it exists only for n=3n = 3.

Exercises on Vectors

Exercise 1.1.

Let a=(2,1)\mathbf{a} = (2, 1) and b=(8,2)\mathbf{b} = (8, 2). Find a+b\mathbf{a} + \mathbf{b} twice, once by drawing and once by computing, and check that the two agree.

Exercise 1.2.

Let a=(1,2,4)\mathbf{a} = (1, 2, 4), b=(2,1,−1)\mathbf{b} = (2, 1, -1) and c=3c = 3. Compute

a+b,a−b,a⋅b,ca,∥a∥,∥ca∥,\mathbf{a} + \mathbf{b}, \quad \mathbf{a} - \mathbf{b}, \quad \mathbf{a} \cdot \mathbf{b}, \quad c\mathbf{a}, \quad \lVert \mathbf{a} \rVert, \quad \lVert c\mathbf{a} \rVert,

and normalise a\mathbf{a} and b\mathbf{b}.

Exercise 1.3.

Give the coordinates of the eight corners of a cube of edge length 11 positioned so that three of its edges lie along the xx-, yy- and zz-axes.

Exercise 1.4.

Romeo is at (3,4,0)(3, 4, 0) and Juliet is at (2,1,5)(2, 1, 5). How far apart are they?

Exercise 1.5.

Find a vector in R3\mathbb{R}^3 orthogonal to (2,−3,4)(2, -3, 4), and describe all of them.

Exercises on Symmetries and Groups

Exercise 1.6.

Let S={a,b}S = \{a, b\} be a two-element set. Show that there are exactly 1616 binary operations on SS, and determine how many of them make SS a group. Then find a formula for the number of binary operations on a set of nn elements.

Exercise 1.7.

Prove that multiplication of complex numbers is associative.

Exercise 1.8.

Which of the following are groups? Justify each answer.

  1. The complex numbers zz with ∣z∣=1|z| = 1, under multiplication;
  2. {x∈R∣x⩾0}\{x \in \mathbb{R} \mid x \geqslant 0\} under x⋆y=max⁡(x,y)x \star y = \max(x, y);
  3. the rationals with odd denominator, under addition;
  4. {a,b}\{a, b\} with a≠ba \neq b and a⋆a=aa \star a = a, b⋆b=bb \star b = b, a⋆b=ba \star b = b, b⋆a=bb \star a = b;
  5. {a,b}\{a, b\} with a≠ba \neq b and a⋆a=aa \star a = a, b⋆b=ab \star b = a, a⋆b=ba \star b = b, b⋆a=bb \star a = b;
  6. R3\mathbb{R}^3 under the vector product v⋆w=v×w\mathbf{v} \star \mathbf{w} = \mathbf{v} \times \mathbf{w}.

Exercise 1.9.

Let SS be the set of all real numbers except −1-1, and for a,b∈Sa, b \in S define a⋆b=ab+a+ba \star b = ab + a + b. Show that (S,⋆)(S, \star) is a group. Check in particular that ⋆\star really is a binary operation on SS.

Exercise 1.10.

Let GG be a group and a,b,c∈Ga, b, c \in G. Prove that

  1. if ab=acab = ac then b=cb = c;
  2. the equation axb=caxb = c has exactly one solution x∈Gx \in G;
  3. (a−1)−1=a(a^{-1})^{-1} = a.

Exercise 1.11.

Let GG be a group with identity ee in which x⋆x=ex \star x = e for every x∈Gx \in G. Show that GG is abelian. Then produce infinitely many groups with this property.

Exercise 1.12.

Let XX be a non-empty set and let α,β\alpha, \beta be permutations of XX such that every element moved by α\alpha is fixed by β\beta and every element moved by β\beta is fixed by α\alpha. Prove that α∘β=β∘α\alpha \circ \beta = \beta \circ \alpha. (Harder.)

Exercise 1.13.

Let (G,⋅)(G, \cdot) be a group and let g,h∈Gg, h \in G with gh=hggh = hg. Show that (gh)r=grhr(gh)^r = g^r h^r for every r∈Nr \in \mathbb{N}. Then exhibit g,h∈S3g, h \in S_3 with (gh)2≠g2h2(gh)^2 \neq g^2 h^2.

Exercise 1.14.

Prove that SnS_n is abelian only for n<3n < 3.

Exercises on Fields and Real Functions

Exercise 1.15.

Solve sin⁡x=cos⁡x\sin x = \cos x for x∈Rx \in \mathbb{R}.

Exercise 1.16.

Solve each of the following for θ∈[0,2π)\theta \in [0, 2\pi).

  1. sin⁡θ=3cos⁡θ\sin\theta = \sqrt{3}\cos\theta;
  2. sin⁡θ=cos⁡2θ\sin\theta = \cos 2\theta;
  3. sin⁡θcos⁡θ=34\sin\theta\cos\theta = \tfrac{\sqrt{3}}{4};
  4. 2cos⁡2θ−3cos⁡θ+1=02\cos^2\theta - 3\cos\theta + 1 = 0;
  5. 3sin⁡θ−cos⁡2θ+3=03\sin\theta - \cos^2\theta + 3 = 0.

Exercise 1.17.

In a triangle labelled as in the figure of the last chapter, let α=π6\alpha = \tfrac{\pi}{6}, β=π3\beta = \tfrac{\pi}{3} and a=1a = 1. Compute cc.

Exercise 1.18.

Let FF be a field. Prove that 0F⋅a=0F0_F \cdot a = 0_F and (−1F) a=−a(-1_F)\,a = -a for every a∈Fa \in F.

Exercise 1.19.

Show that {0,1}\{0, 1\} with 1+1=01 + 1 = 0 and the usual multiplication is a field, and that no field has exactly three elements in which 1+1=01 + 1 = 0.

Check Yourself

 

Fresh questions on the whole lesson — none of them is worked out above. Do each on paper first; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 1.20.

Let u=(3,−4)\mathbf{u} = (3, -4) in R2\RR^2. What is ∥2u∥\lVert 2\mathbf{u} \rVert?

answer one of these

Exercise 1.21.

Which of these is a unit vector in R3\RR^3?

answer one of these

Exercise 1.22.

Which of these vectors is orthogonal to (1,2,3)(1, 2, 3)?

answer one of these

Exercise 1.23.

What is the projection of u=(4,0)\mathbf{u} = (4, 0) onto v=(1,1)\mathbf{v} = (1, 1)?

answer one of these

Exercise 1.24.

What is the angle between (1,0)(1, 0) and (1,1)(1, 1) in R2\RR^2?

answer one of these

Exercise 1.25.

How many unit vectors in R2\RR^2 are orthogonal to (3,4)(3, 4)?

answer one of these

Exercise 1.26.

How many elements has the symmetry group of a regular hexagon?

answer one of these

Exercise 1.27.

How many reflections are there among the symmetries of a regular pentagon?

answer one of these

Exercise 1.28.

What is the order of the permutation (1 2 3)(4 5)(1\,2\,3)(4\,5) in S5S_5?

answer one of these

Exercise 1.29.

What is #(S5)\#(S_5)?

answer one of these

Exercise 1.30.

In S4S_4, what is the inverse of (1 2 3 4)(1\,2\,3\,4)?

answer one of these

Exercise 1.31.

Exactly one of these is not a group. Which?

answer one of these

Exercise 1.32.

In the two-element field {0,1}\{0, 1\}, where 1+1=01 + 1 = 0, what is 1+1+11 + 1 + 1?

answer one of these

Exercise 1.33.

What is log⁡232\log_2 32?

answer one of these

Exercise 1.34.

What is exp⁡(log⁡3+log⁡4)\exp(\log 3 + \log 4)?

answer one of these

Exercise 1.35.

What is tan⁡(−π4)\tan\left(-\tfrac{\pi}{4}\right)?

answer one of these

Exercise 1.36.

How many solutions has cos⁡θ=0\cos\theta = 0 in [0,2π)[0, 2\pi)?

answer one of these

Exercise 1.37.

What is the least upper bound of { x∈R∣x2<2 }\{\, x \in \RR \mid x^2 < 2 \,\}?

answer one of these

Exercise 1.38.

Which of these sets has no largest element?

answer one of these

Exercise 1.39.

For a,b⩾0a, b \geqslant 0, when is ab=12(a+b)\sqrt{ab} = \tfrac{1}{2}(a + b)?

answer one of these

Lesson 2

Recurrences and Sums

Taught

Recurrences

A recurrence defines a quantity at one value of nn in terms of its values at smaller ones, together with enough starting values. Counting problems often lead to recurrences: removing one piece from a problem of size nn leaves a problem of the same kind and smaller size. Computing the millionth term from a recurrence means computing the first million, and in this chapter we replace such definitions by formulas that can be evaluated directly.

The Tower of Hanoi

Édouard Lucas put the following puzzle on sale in 1883. Three pegs stand in a row, and on the first of them sit nn discs of distinct sizes, stacked largest at the bottom and smallest on top. A move lifts the top disc off one peg and drops it onto another, and it is illegal to place a disc on top of a smaller one. The object is to move the whole stack onto a different peg.

ABC
Figure 2.1. The Tower of Hanoi with four discs. No disc may ever rest on a smaller one.

Write TnT_n for the least number of legal moves that carries a stack of nn discs from one peg to another. With one disc a single move does it, so T1=1T_1 = 1. With two, the small disc goes to the spare peg, the large one to the target, and the small one on top of it: T2=3T_2 = 3. With three the shortest solution takes seven moves,

d1→C,d2→B,d1→B,d3→C,d1→A,d2→C,d1→C,\begin{aligned} &d_1 \to C, \quad d_2 \to B, \quad d_1 \to B, \quad d_3 \to C, \\ &d_1 \to A, \quad d_2 \to C, \quad d_1 \to C, \end{aligned}

and with four it takes fifteen. The counts so far run 1,3,7,15,31,631, 3, 7, 15, 31, 63.

Proposition 2.1 (The Hanoi recurrence).

For every n∈Nn \in \mathbb{N},

Tn=2Tn−1+1,T_n = 2T_{n-1} + 1,

and T0=0T_0 = 0.

Discussion.

An equality between two counts is two inequalities, and each is argued differently. For Tn⩽2Tn−1+1T_n \leqslant 2T_{n-1} + 1 we exhibit a strategy costing that many moves and appeal to TnT_n being the least cost: shift the top n−1n-1 discs to the spare peg, move the largest, shift the n−1n-1 back on top of it. The two shifts are legal because the largest disc is out of the way in the first and sits below everything in the second, so neither is obstructed. For Tn⩾2Tn−1+1T_n \geqslant 2T_{n-1} + 1 we argue about an arbitrary solution rather than a chosen one, and look at the largest disc: it has to move at some point, and at the first moment it moves, the other n−1n-1 discs are on one peg and not on either of the two the largest disc occupies, since none of them may sit on top of it or under it in the destination. Reaching that position costs at least Tn−1T_{n-1}, the move itself costs one, and rebuilding the stack afterwards costs at least Tn−1T_{n-1} again.

Proof.

There are no discs to move when n=0n = 0, so T0=0T_0 = 0.

Let n∈Nn \in \mathbb{N}. For the upper bound, carry out the following. Move the top n−1n-1 discs from the source peg to the spare peg, which is possible in Tn−1T_{n-1} moves and remains legal with the largest disc left in place, since that disc is larger than all of them and lies at the bottom of the source peg. Move the largest disc to the target peg: one move. Move the n−1n-1 discs from the spare peg onto the target peg, again Tn−1T_{n-1} moves, legal because they all land on the largest disc. That is 2Tn−1+12T_{n-1} + 1 moves in all, so Tn⩽2Tn−1+1T_n \leqslant 2T_{n-1} + 1.

For the lower bound, take any legal sequence of moves carrying the stack from the source to the target. The largest disc must move at least once. Consider the first move that lifts it. Just before that move the other n−1n-1 discs lie neither on the peg it leaves nor on the peg it arrives at: not the first, because they are all smaller and would have to be above it; not the second, because it may not land on a smaller disc. So all n−1n-1 of them are stacked on the remaining peg, and getting them there from the source peg took at least Tn−1T_{n-1} moves. After the largest disc has reached the target for the last time, the n−1n-1 discs must be brought from that third peg onto it, which takes at least Tn−1T_{n-1} moves more, and the moves of the largest disc themselves account for at least one. Hence Tn⩾2Tn−1+1T_n \geqslant 2T_{n-1} + 1.

Closed Forms

Computing T100T_{100} from the recurrence means computing T1T_1 through T99T_{99} first. What we want instead is a function ff with f(n)=Tnf(n) = T_n for every nn, evaluable on its own. The sequence 0,1,3,7,15,31,630, 1, 3, 7, 15, 31, 63 sits one below the powers of two, which is enough to make a guess.

Proposition 2.2 (Closed form for the Tower of Hanoi).

For every n∈N0n \in \mathbb{N}_0 we have Tn=2n−1T_n = 2^n - 1.

Discussion.

The recurrence defines TnT_n from Tn−1T_{n-1}, so a claim about all nn is proved by induction, and the induction has exactly the shape of the recurrence: one base case at n=0n = 0, and a step that rewrites TnT_n as 2Tn−1+12T_{n-1} + 1, replaces Tn−1T_{n-1} by the inductive hypothesis, and simplifies. The formula was guessed from the first few terms, and the induction checks it.

Proof.

For n=0n = 0 we have T0=0T_0 = 0 and 20−1=02^0 - 1 = 0.

Suppose Tn−1=2n−1−1T_{n-1} = 2^{n-1} - 1 for some n∈Nn \in \mathbb{N}. By the recurrence,

Tn=2Tn−1+1=2(2n−1−1)+1=2n−2+1=2n−1,T_n = 2T_{n-1} + 1 = 2\bigl(2^{n-1} - 1\bigr) + 1 = 2^n - 2 + 1 = 2^n - 1,

which is the claim at nn.

Induction proves a formula but does not find one, and guessing from a few terms does not always work. Two ways of finding the formula follow.

Unrolling. Apply the recurrence to itself repeatedly:

Tn=2Tn−1+1=2(2Tn−2+1)+1=4Tn−2+2+1=4(2Tn−3+1)+2+1=8Tn−3+4+2+1.\begin{aligned} T_n &= 2T_{n-1} + 1 \\ &= 2\bigl(2T_{n-2} + 1\bigr) + 1 = 4T_{n-2} + 2 + 1 \\ &= 4\bigl(2T_{n-3} + 1\bigr) + 2 + 1 = 8T_{n-3} + 4 + 2 + 1 . \end{aligned}

The pattern is

Tn=2kTn−k+2k−1+2k−2+⋯+2+1=2kTn−k+2k−1,T_n = 2^k T_{n-k} + 2^{k-1} + 2^{k-2} + \cdots + 2 + 1 = 2^k T_{n-k} + 2^k - 1,

and one more application confirms that it reproduces itself:

2kTn−k+2k−1=2k(2Tn−k−1+1)+2k−1=2k+1Tn−k−1+2k+1−1,2^k T_{n-k} + 2^k - 1 = 2^k\bigl(2T_{n-k-1} + 1\bigr) + 2^k - 1 = 2^{k+1}T_{n-k-1} + 2^{k+1} - 1,

which is the same expression with k+1k+1 in place of kk. Since the case k=1k = 1 is the recurrence itself, the displayed identity holds for every kk with 1⩽k⩽n1 \leqslant k \leqslant n by induction on kk. Taking k=nk = n leaves TnT_n in terms of a value we know,

Tn=2nT0+2n−1=2n−1.T_n = 2^n T_0 + 2^n - 1 = 2^n - 1 .

Remark (The geometric sum).

For the step from 2k−1+⋯+2+12^{k-1} + \cdots + 2 + 1 to 2k−12^k - 1, write S=2k−1+⋯+2+1S = 2^{k-1} + \cdots + 2 + 1 and doubling gives 2S=2k+2k−1+⋯+22S = 2^k + 2^{k-1} + \cdots + 2, so 2S−S=2k−12S - S = 2^k - 1, that is S=2k−1S = 2^k - 1. The same trick evaluates 1+q+⋯+qk−11 + q + \cdots + q^{k-1} as (qk−1)/(q−1)(q^k - 1)/(q - 1) for any q≠1q \neq 1.

Substitution. Rather than solve the recurrence, change the unknown so that the recurrence becomes one we can already solve. Adding 11 to both sides of Tn=2Tn−1+1T_n = 2T_{n-1} + 1 gives

Tn+1=2Tn−1+2=2(Tn−1+1),T_n + 1 = 2T_{n-1} + 2 = 2\bigl(T_{n-1} + 1\bigr),

so if we set Un=Tn+1U_n = T_n + 1 then U0=1U_0 = 1 and Un=2Un−1U_n = 2U_{n-1} for n⩾1n \geqslant 1. That recurrence doubles at every step, so Un=2nU_n = 2^n, and therefore Tn=Un−1=2n−1T_n = U_n - 1 = 2^n - 1.

Problem 2.1.

Show that in the strategy of the proof of the Hanoi recurrence every disc moves at least once, and that the smallest disc moves on every second move.

Problem 2.2.

Suppose the three pegs are arranged in a row and a disc may only be moved between adjacent pegs, so that a move from the left peg to the right peg is forbidden. Let SnS_n be the least number of moves needed to transfer a stack of nn discs from the left peg to the right peg. Find a recurrence for SnS_n and solve it.

Lines in the Plane

Here is a second problem of the same shape. What is the largest number LnL_n of regions into which nn straight lines can cut the plane?

With no lines there is one region, so L0=1L_0 = 1. One line cuts the plane in two however it is drawn, so L1=2L_1 = 2. Two lines do best when they are not parallel, giving L2=4L_2 = 4. At this point 1,2,41, 2, 4 invites the guess Ln=2nL_n = 2^n, and the guess fails at once: a third line meets the two existing ones in at most two points, which divide it into at most three pieces, and each piece splits one old region in two. So the third line adds at most three regions and L3=7L_3 = 7, not 88.

1234567
Figure 2.2. Three lines, no two parallel and no three through a common point, cut the plane into seven regions.

Proposition 2.3 (The line recurrence).

For every n∈Nn \in \mathbb{N},

Ln=Ln−1+n,L_n = L_{n-1} + n,

and L0=1L_0 = 1.

Discussion.

Again the equality splits into two inequalities about the nnth line added to n−1n-1 already drawn. For the upper bound, the new line gains exactly as many regions as the number of old regions it passes through, and it passes through one more region than the number of points at which it meets the old lines. Two distinct lines meet in at most one point, so the new line meets n−1n-1 old lines in at most n−1n-1 points and gains at most nn regions. For the lower bound we must place the new line so that both estimates are attained: not parallel to any old line, which forces it to meet each of them, and not through any existing intersection point, which keeps the n−1n-1 meeting points distinct. Only finitely many directions and finitely many points have to be avoided, so such a line exists.

Proof.

With no lines drawn the plane is a single region, so L0=1L_0 = 1.

Let n∈Nn \in \mathbb{N} and suppose n−1n-1 lines have been drawn. Adding a line ℓ\ell increases the number of regions by exactly the number of old regions that ℓ\ell passes through, since each such region is cut into two and no other region is touched. The old lines meet ℓ\ell in some set of points, and those points cut ℓ\ell into one more piece than there are points, with each piece lying in one old region. Two distinct lines meet in at most one point, so ℓ\ell meets the n−1n-1 old lines in at most n−1n-1 points and therefore passes through at most nn old regions. Hence Ln⩽Ln−1+nL_n \leqslant L_{n-1} + n.

For the reverse, take n−1n-1 lines realising Ln−1L_{n-1} regions. The old lines have n−1n-1 directions and finitely many pairwise intersection points, so we may choose ℓ\ell parallel to none of them and passing through none of those points. Then ℓ\ell meets every old line, in n−1n-1 distinct points, so it passes through exactly nn old regions and adds nn new ones. Hence Ln⩾Ln−1+nL_n \geqslant L_{n-1} + n.

Unrolling this recurrence gives

Ln=Ln−1+n=Ln−2+(n−1)+n=Ln−3+(n−2)+(n−1)+n=L0+1+2+⋯+n=1+n(n+1)2,\begin{aligned} L_n &= L_{n-1} + n = L_{n-2} + (n-1) + n = L_{n-3} + (n-2) + (n-1) + n \\ &= L_0 + 1 + 2 + \cdots + n = 1 + \frac{n(n+1)}{2}, \end{aligned}

using the sum of the first nn positive integers.

Problem 2.3.

Prove by induction that Ln=1+12n(n+1)L_n = 1 + \tfrac{1}{2}n(n+1) for every n∈N0n \in \mathbb{N}_0.

Problem 2.4.

Prove that 1+2+⋯+n=12n(n+1)1 + 2 + \cdots + n = \tfrac{1}{2}n(n+1) for every n∈N0n \in \mathbb{N}_0, by pairing the first term with the last, the second with the second-last, and so on.

Problem 2.5.

What is the largest number of regions into which nn lines can cut the plane if all nn lines are required to pass through a common point?

Problem 2.6.

A zig is a bent line, made of two rays issuing from a common point. Find and solve a recurrence for the largest number of regions into which nn zigs can cut the plane.

Linear Recurrences

Guessing and unrolling only work when the recurrence is simple. The linear recurrences below can all be solved by one fixed method.

How many ways are there to climb a staircase of nn steps, if each stride goes up either one step or two? For four steps there are five ways:

1,1,1,12,22,1,11,2,11,1,2\begin{aligned} &1,1,1,1 \qquad 2,2 \qquad 2,1,1 \\ &1,2,1 \qquad 1,1,2 \end{aligned}

There is one way to climb no steps, namely to do nothing, and one way to climb one step. For n⩾2n \geqslant 2 any ascent begins with either a stride of one, leaving an ascent of n−1n-1 steps, or a stride of two, leaving an ascent of n−2n-2; the two cases are exclusive and exhaust the possibilities. Writing f(n)f(n) for the number of ascents of nn steps,

f(0)=1,f(1)=1,f(n)=f(n−1)+f(n−2)(n⩾2).f(0) = 1, \qquad f(1) = 1, \qquad f(n) = f(n-1) + f(n-2) \quad (n \geqslant 2).

This is the Fibonacci recurrence, introduced in 1202 to model rabbit populations and unsolved in closed form for nearly six centuries. We use f(n)f(n) rather than fnf_n from here on, since the argument will shortly be something other than an integer.

Definition 2.4 (Linear recurrence).

A homogeneous linear recurrence of order dd is one of the form

f(n)=a1f(n−1)+a2f(n−2)+⋯+adf(n−d),f(n) = a_1 f(n-1) + a_2 f(n-2) + \cdots + a_d f(n-d),

where a1,…,ada_1, \ldots, a_d are constants with ad≠0a_d \neq 0. Values of ff prescribed at finitely many points are its boundary conditions. Adding a further term g(n)g(n) on the right, where gg is a fixed function, gives an inhomogeneous linear recurrence.

The Fibonacci recurrence has order 22 with a1=a2=1a_1 = a_2 = 1, and the boundary conditions f(0)=f(1)=1f(0) = f(1) = 1.

Theorem 2.5 (Solutions form a linear space).

Let ff and gg both satisfy the homogeneous linear recurrence with coefficients a1,…,ada_1, \ldots, a_d. Then for all s,t∈Rs, t \in \mathbb{R} the function h(n)=sf(n)+tg(n)h(n) = s f(n) + t g(n) satisfies it too.

Discussion.

The recurrence is an identity that must hold at every nn, so the proof evaluates the right-hand side of the recurrence at hh and pushes it back to h(n)h(n). Two facts are used, both about arithmetic rather than about recurrences: the right-hand side is a sum of terms each of which is a constant times a value of the function, so it distributes over the combination sf+tgsf + tg; and the hypotheses on ff and gg let each of the two resulting sums collapse to f(n)f(n) and g(n)g(n). Regrouping the terms is the argument.

Proof.

Let n⩾dn \geqslant d. Using the definition of hh, then the hypotheses on ff and gg,

a1h(n−1)+⋯+adh(n−d)=a1(sf(n−1)+tg(n−1))+⋯+ad(sf(n−d)+tg(n−d))=s(a1f(n−1)+⋯+adf(n−d))+t(a1g(n−1)+⋯+adg(n−d))=sf(n)+tg(n)=h(n).\begin{aligned} a_1 h(n-1) + \cdots + a_d h(n-d) &= a_1\bigl(sf(n-1) + tg(n-1)\bigr) + \cdots + a_d\bigl(sf(n-d) + tg(n-d)\bigr) \\ &= s\bigl(a_1 f(n-1) + \cdots + a_d f(n-d)\bigr) + t\bigl(a_1 g(n-1) + \cdots + a_d g(n-d)\bigr) \\ &= s f(n) + t g(n) = h(n). \end{aligned}

Linear recurrences tend to have exponential solutions, so we look for one of the form f(n)=xnf(n) = x^n and find which xx work. Substituting into f(n)=f(n−1)+f(n−2)f(n) = f(n-1) + f(n-2) gives xn=xn−1+xn−2x^n = x^{n-1} + x^{n-2}, and dividing by xn−2x^{n-2}, which is legitimate since x=0x = 0 does not satisfy the boundary conditions, leaves

x2=x+1.x^2 = x + 1 .

Its roots are

φ=1+52=1.618…,φ^=1−52=−0.618…,\varphi = \frac{1 + \sqrt{5}}{2} = 1.618\ldots, \qquad \hat{\varphi} = \frac{1 - \sqrt{5}}{2} = -0.618\ldots,

so φn\varphi^n and φ^n\hat{\varphi}^n both satisfy the recurrence, and by the theorem above so does sφn+tφ^ns\varphi^n + t\hat{\varphi}^n for every ss and tt. The two boundary conditions determine the two constants.

Definition 2.6 (Characteristic equation).

The characteristic equation of the homogeneous linear recurrence f(n)=a1f(n−1)+⋯+adf(n−d)f(n) = a_1f(n-1) + \cdots + a_df(n-d) is

xd=a1xd−1+a2xd−2+⋯+ad−1x+ad,x^d = a_1 x^{d-1} + a_2 x^{d-2} + \cdots + a_{d-1}x + a_d ,

obtained by substituting f(n)=xnf(n) = x^n and dividing by xn−dx^{n-d}. Its coefficients are read straight off the recurrence.

Theorem 2.7 (Binet's formula).

Let ff satisfy f(0)=f(1)=1f(0) = f(1) = 1 and f(n)=f(n−1)+f(n−2)f(n) = f(n-1) + f(n-2) for n⩾2n \geqslant 2. Then for every n∈N0n \in \mathbb{N}_0,

f(n)=15(φ n+1−φ^ n+1),φ=1+52,φ^=1−52.f(n) = \frac{1}{\sqrt{5}}\left( \varphi^{\,n+1} - \hat{\varphi}^{\,n+1} \right), \qquad \varphi = \frac{1 + \sqrt{5}}{2}, \quad \hat{\varphi} = \frac{1 - \sqrt{5}}{2}.

Discussion.

The two exponentials φn\varphi^n and φ^n\hat\varphi^n satisfy the recurrence because φ\varphi and φ^\hat\varphi solve the characteristic equation, and the previous theorem then makes every combination sφn+tφ^ns\varphi^n + t\hat\varphi^n a solution as well. What remains is to choose ss and tt so that the combination also meets the two boundary conditions, and each condition is one linear equation in the two unknowns. Solving the pair uses φ−φ^=5\varphi - \hat\varphi = \sqrt5 and 1−φ^=φ1 - \hat\varphi = \varphi. Since the recurrence and the two boundary conditions determine ff at every point, the combination found is ff.

Proof.

Both φ\varphi and φ^\hat\varphi satisfy x2=x+1x^2 = x + 1, so multiplying by xn−2x^{n-2} shows that φn\varphi^n and φ^n\hat\varphi^n satisfy the recurrence; by the previous theorem so does h(n)=sφn+tφ^nh(n) = s\varphi^n + t\hat\varphi^n for any reals s,ts, t.

The boundary conditions require h(0)=1h(0) = 1 and h(1)=1h(1) = 1, that is

s+t=1,sφ+tφ^=1.s + t = 1, \qquad s\varphi + t\hat\varphi = 1 .

Substituting t=1−st = 1 - s into the second gives s(φ−φ^)=1−φ^s(\varphi - \hat\varphi) = 1 - \hat\varphi. Now φ−φ^=5\varphi - \hat\varphi = \sqrt5 and 1−φ^=12(1+5)=φ1 - \hat\varphi = \tfrac{1}{2}(1 + \sqrt5) = \varphi, so s=φ/5s = \varphi/\sqrt5; and then t=1−φ/5=(5−φ)/5=−φ^/5t = 1 - \varphi/\sqrt5 = (\sqrt5 - \varphi)/\sqrt5 = -\hat\varphi/\sqrt5, since 5−φ=12(5−1)=−φ^\sqrt5 - \varphi = \tfrac{1}{2}(\sqrt5 - 1) = -\hat\varphi. Therefore

h(n)=φ5φn−φ^5φ^n=15(φ n+1−φ^ n+1).h(n) = \frac{\varphi}{\sqrt5}\varphi^n - \frac{\hat\varphi}{\sqrt5}\hat\varphi^n = \frac{1}{\sqrt5}\left(\varphi^{\,n+1} - \hat\varphi^{\,n+1}\right).

The recurrence together with the values at 00 and 11 determines f(n)f(n) for every nn, and hh satisfies all three, so f=hf = h.

Remark (Integers out of square roots).

Every value of ff is a whole number, yet the formula is built entirely from 5\sqrt5; the irrational parts cancel at every nn. Since ∣φ^∣<1|\hat\varphi| < 1 the second term is smaller than 12\tfrac12 in absolute value at every nn, so f(n)f(n) is the nearest integer to φ n+1/5\varphi^{\,n+1}/\sqrt5 throughout. For instance φ20/5=6765.000029…\varphi^{20}/\sqrt5 = 6765.000029\ldots, and f(19)=6765f(19) = 6765. The same estimate shows that consecutive values have ratio tending to φ\varphi, the golden ratio.

Example 2.8 (A recurrence of order two).

Let f(0)=0f(0) = 0, f(1)=1f(1) = 1 and f(n)=5f(n−1)−6f(n−2)f(n) = 5f(n-1) - 6f(n-2) for n⩾2n \geqslant 2. The characteristic equation is x2=5x−6x^2 = 5x - 6, that is (x−2)(x−3)=0(x-2)(x-3) = 0, with roots 22 and 33. So f(n)=s 2n+t 3nf(n) = s\,2^n + t\,3^n, and the boundary conditions give s+t=0s + t = 0 and 2s+3t=12s + 3t = 1, whence t=1t = 1 and s=−1s = -1. Therefore f(n)=3n−2nf(n) = 3^n - 2^n.

When the characteristic equation has a repeated root, the powers of that root alone do not supply enough independent solutions, and the missing ones carry a factor of nn.

Proposition 2.9 (Repeated roots).

Let r≠0r \neq 0 be a root of multiplicity at least 22 of the characteristic equation of a homogeneous linear recurrence. Then f(n)=nrnf(n) = n r^n satisfies the recurrence.

Discussion.

Saying that rr is a root of multiplicity at least two means that the characteristic polynomial and its derivative both vanish at rr, and differentiating a power xnx^n produces the factor nn in the claimed solution. So the proof multiplies the characteristic polynomial by xn−dx^{n-d} to obtain a polynomial whose vanishing at rr is the statement that rnr^n solves the recurrence, differentiates it, and multiplies by xx; the result, evaluated at rr, is the statement that nrnnr^n solves the recurrence. The hypothesis r≠0r \neq 0 ensures that rr is still a root of multiplicity at least 22 after multiplying by xn−dx^{n-d}.

Proof.

Write p(x)=xd−a1xd−1−⋯−adp(x) = x^d - a_1x^{d-1} - \cdots - a_d for the characteristic polynomial and fix n⩾dn \geqslant d. Put

q(x)=xn−dp(x)=xn−a1xn−1−⋯−adxn−d.q(x) = x^{n-d}p(x) = x^n - a_1x^{n-1} - \cdots - a_d x^{n-d} .

Since r≠0r \neq 0 is a root of pp of multiplicity at least 22, it is a root of qq of multiplicity at least 22, so q(r)=0q(r) = 0 and q′(r)=0q'(r) = 0. Differentiating,

q′(x)=nxn−1−a1(n−1)xn−2−⋯−ad(n−d)xn−d−1,q'(x) = n x^{n-1} - a_1(n-1)x^{n-2} - \cdots - a_d(n-d)x^{n-d-1},

and multiplying by xx,

x q′(x)=nxn−a1(n−1)xn−1−⋯−ad(n−d)xn−d.x\,q'(x) = n x^{n} - a_1(n-1)x^{n-1} - \cdots - a_d(n-d)x^{n-d} .

Evaluating at x=rx = r and using q′(r)=0q'(r) = 0 gives

nrn=a1(n−1)rn−1+⋯+ad(n−d)rn−d,n r^n = a_1 (n-1) r^{n-1} + \cdots + a_d (n-d) r^{n-d},

which says exactly that f(n)=nrnf(n) = nr^n satisfies the recurrence.

More generally, a root rr of multiplicity kk contributes the kk solutions rnr^n, nrnnr^n, n2rnn^2r^n, up to nk−1rnn^{k-1}r^n, and a recurrence of order dd is solved by taking a linear combination of the dd solutions collected in this way from all the roots. If the characteristic equation of an order-four recurrence has roots ss, tt and uu twice, the general solution is

f(n)=a sn+b tn+c un+d n un,f(n) = a\,s^n + b\,t^n + c\,u^n + d\,n\,u^n,

and four boundary conditions give four linear equations in a,b,c,da, b, c, d.

Inhomogeneous Recurrences

The Hanoi recurrence f(n)=2f(n−1)+1f(n) = 2f(n-1) + 1 is not homogeneous: the extra 11 is a term g(n)g(n) on the right. Such recurrences are solved in five steps.

  1. Delete g(n)g(n) and find the roots of the characteristic equation of what is left.
  2. Write down the general solution of that homogeneous recurrence, leaving its constants undetermined. This is the homogeneous solution.
  3. Restore g(n)g(n) and find any one function satisfying the full recurrence, ignoring the boundary conditions. This is a particular solution.
  4. Add the two. This is the general solution.
  5. Use the boundary conditions to fix the constants.

Example 2.10 (Hanoi with heavy discs).

Suppose moving a disc costs its size in seconds, so that moving the nnth disc takes nn seconds rather than one. The total time obeys

f(1)=1,f(n)=2f(n−1)+n(n⩾2).f(1) = 1, \qquad f(n) = 2f(n-1) + n \quad (n \geqslant 2).

Deleting the nn leaves f(n)=2f(n−1)f(n) = 2f(n-1), whose characteristic equation is x=2x = 2, so the homogeneous solution is c 2nc\,2^n.

For a particular solution, g(n)=ng(n) = n is a polynomial of degree one, so try f(n)=an+bf(n) = an + b. Substituting,

an+b=2(a(n−1)+b)+n⟺0=(a+1)n+(b−2a),an + b = 2\bigl(a(n-1) + b\bigr) + n \quad\Longleftrightarrow\quad 0 = (a+1)n + (b - 2a),

which holds for every nn exactly when a=−1a = -1 and b=−2b = -2. So −n−2-n-2 is a particular solution and the general solution is f(n)=c 2n−n−2f(n) = c\,2^n - n - 2.

The boundary condition f(1)=1f(1) = 1 gives 2c−3=12c - 3 = 1, so c=2c = 2 and

f(n)=2 n+1−n−2.f(n) = 2^{\,n+1} - n - 2 .

Against 2n−12^n - 1 for the original puzzle, the heavy discs cost roughly twice as long.

Remark (Guessing a particular solution).

Finding a particular solution is the one step that involves a guess, and the guess is usually shaped like g(n)g(n) itself.

If g(n)g(n) is constant, try f(n)=cf(n) = c; if that fails, try bn+cbn + c, then an2+bn+can^2 + bn + c, and so on. If g(n)g(n) is a polynomial, try a polynomial of the same degree first and then of higher degree. If g(n)g(n) is an exponential such as 3n3^n, try f(n)=c 3nf(n) = c\,3^n, then bn3n+c3nbn3^n + c3^n, and so on. A guess fails when substituting it leaves an equation with no constant solution, and raising the degree by one is then the next move.

Problem 2.7.

Solve f(0)=0f(0) = 0, f(1)=1f(1) = 1, f(n)=4f(n−1)−4f(n−2)f(n) = 4f(n-1) - 4f(n-2) for n⩾2n \geqslant 2. (The characteristic equation has a repeated root.)

Problem 2.8.

Solve f(0)=1f(0) = 1 and f(n)=3f(n−1)+2nf(n) = 3f(n-1) + 2^n for n⩾1n \geqslant 1.

Problem 2.9.

Let rr be a root of multiplicity at least 33 of the characteristic equation of a homogeneous linear recurrence, with r≠0r \neq 0. Prove that n2rnn^2 r^n satisfies the recurrence.

Problem 2.10.

How many ways are there to climb nn stairs if a stride may go up one, two or three steps? Write down the recurrence, its characteristic equation, and compute the number of ways for n⩽6n \leqslant 6.

Problem 2.11.

Let ff satisfy the Fibonacci recurrence with f(0)=f(1)=1f(0) = f(1) = 1. Prove that f(n)f(n) is the nearest integer to φ n+1/5\varphi^{\,n+1}/\sqrt{5} for every n⩾1n \geqslant 1.

The Josephus Problem

Flavius Josephus, a first-century historian, was said to have been trapped in a cave with forty-one other rebels who preferred death to capture and agreed to stand in a circle and kill every third man. Josephus worked out where to stand.

The version we take is this. Number nn people 11 to nn around a circle and go round eliminating every second person until one is left. Which number survives? Call it J(n)J(n).

12346789105
Figure 2.3. Ten people in a circle. Eliminating every second one leaves number 55, so J(10)=5J(10) = 5.

Proposition 2.11 (The Josephus recurrence).

For every ℓ∈N\ell \in \mathbb{N},

J(2ℓ)=2J(ℓ)−1,J(2ℓ+1)=2J(ℓ)+1,J(2\ell) = 2J(\ell) - 1, \qquad J(2\ell + 1) = 2J(\ell) + 1,

and J(1)=1J(1) = 1.

Discussion.

With one person, that person survives. Otherwise, after one lap of the circle exactly the even-numbered people have gone, and what is left is a circle of half the size with the same rule about to be applied to it. So the problem of size nn reduces to the problem of size ⌊n/2⌋\lfloor n/2 \rfloor, and it remains to translate the numbering of the survivors back into their original numbers. The two cases differ only in where the second lap starts. If nn is even the lap ends by killing person nn and the next to die is the second of the survivors; if nn is odd the lap ends by killing person n−1n-1, then person 11 dies immediately, and the survivors start from person 33. Reading off the kkth survivor’s original number in each case, 2k−12k-1 and 2k+12k+1, converts J(ℓ)J(\ell) into J(n)J(n).

Proof.

With one person there is nobody to eliminate, so J(1)=1J(1) = 1.

Suppose n=2ℓn = 2\ell with ℓ∈N\ell \in \mathbb{N}. Going round once eliminates 2,4,…,2ℓ2, 4, \ldots, 2\ell and leaves the ℓ\ell odd-numbered people 1,3,…,2ℓ−11, 3, \ldots, 2\ell - 1, with the next elimination falling on the second of them. So what remains is the same problem for ℓ\ell people, counted from 11, and the kkth of those people carries the original number 2k−12k - 1. The survivor is the J(ℓ)J(\ell)th of them, so its original number is 2J(ℓ)−12J(\ell) - 1.

Suppose instead n=2ℓ+1n = 2\ell + 1 with ℓ∈N\ell \in \mathbb{N}. Going round once eliminates 2,4,…,2ℓ2, 4, \ldots, 2\ell; the count then passes from 2ℓ+12\ell+1 to 11, which is eliminated next. That leaves the ℓ\ell people 3,5,…,2ℓ+13, 5, \ldots, 2\ell + 1, with the next elimination falling on the second of them. Again this is the problem for ℓ\ell people, and the kkth of them carries the original number 2k+12k + 1. The survivor is therefore numbered 2J(ℓ)+12J(\ell) + 1.

Tabulating the first sixteen values, grouped by powers of two, makes the pattern plain:

n12345678910111213141516J(n)1131357135791113151\begin{array}{c|cc|cccc|cccccccc|c} n & 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 & 10 & 11 & 12 & 13 & 14 & 15 & 16 \\ \hline J(n) & 1 & 1 & 3 & 1 & 3 & 5 & 7 & 1 & 3 & 5 & 7 & 9 & 11 & 13 & 15 & 1 \end{array}

Each block begins at 11 when nn reaches a power of two and climbs through the odd numbers. That suggests measuring nn from the power of two below it.

Theorem 2.12 (Closed form for the Josephus problem).

Let n∈Nn \in \mathbb{N} and write n=2m+rn = 2^m + r where 2m2^m is the largest power of two with 2m⩽n2^m \leqslant n, so that m⩾0m \geqslant 0 and 0⩽r<2m0 \leqslant r < 2^m. Then

J(n)=2r+1.J(n) = 2r + 1 .

Discussion.

The recurrence takes nn to ⌊n/2⌋\lfloor n/2 \rfloor, so the induction hypothesis is needed at a smaller value than n−1n - 1, and we use strong induction on nn. The proof follows how the decomposition n=2m+rn = 2^m + r changes under halving. If nn is even then rr is even too, since r=n−2mr = n - 2^m and m⩾1m \geqslant 1, and halving nn gives 2m−1+r/22^{m-1} + r/2 with the remainder still in range. If nn is odd then rr is odd, and ⌊n/2⌋=2m−1+(r−1)/2\lfloor n/2 \rfloor = 2^{m-1} + (r-1)/2, again in range. In both cases the hypothesis supplies the survivor for the halved circle, and the two clauses of the recurrence turn it into 2r+12r+1; both cases give the same answer.

Proof.

We induct on nn, assuming the claim for all smaller values.

If n=1n = 1 then m=0m = 0 and r=0r = 0, and J(1)=1=2⋅0+1J(1) = 1 = 2 \cdot 0 + 1.

Let n⩾2n \geqslant 2 and write n=2m+rn = 2^m + r with 0⩽r<2m0 \leqslant r < 2^m; since n⩾2n \geqslant 2 we have m⩾1m \geqslant 1.

Suppose nn is even, say n=2ℓn = 2\ell. Then r=n−2mr = n - 2^m is even, and

ℓ=n2=2m−1+r2,0⩽r2<2m−1,\ell = \frac{n}{2} = 2^{m-1} + \frac{r}{2}, \qquad 0 \leqslant \frac{r}{2} < 2^{m-1},

so the decomposition of ℓ\ell has exponent m−1m-1 and remainder r/2r/2. As ℓ<n\ell < n, the inductive hypothesis gives J(ℓ)=2(r/2)+1=r+1J(\ell) = 2(r/2) + 1 = r + 1, and the recurrence gives

J(n)=2J(ℓ)−1=2(r+1)−1=2r+1.J(n) = 2J(\ell) - 1 = 2(r+1) - 1 = 2r + 1 .

Suppose instead nn is odd, say n=2ℓ+1n = 2\ell + 1 with ℓ⩾1\ell \geqslant 1. Then r=n−2mr = n - 2^m is odd, and

ℓ=n−12=2m−1+r−12,0⩽r−12<2m−1,\ell = \frac{n-1}{2} = 2^{m-1} + \frac{r-1}{2}, \qquad 0 \leqslant \frac{r-1}{2} < 2^{m-1},

the upper bound because r<2mr < 2^m forces r−1<2m−1r - 1 < 2^m - 1 and hence (r−1)/2<2m−1(r-1)/2 < 2^{m-1}. The inductive hypothesis gives J(ℓ)=2⋅r−12+1=rJ(\ell) = 2\cdot\frac{r-1}{2} + 1 = r, and the recurrence gives

J(n)=2J(ℓ)+1=2r+1.J(n) = 2J(\ell) + 1 = 2r + 1 .

Reading the Answer in Binary

The decomposition n=2m+rn = 2^m + r can be read off from the binary expansion of nn. Writing

n=(1 bm−1 bm−2⋯b1 b0)2n = (1\,b_{m-1}\,b_{m-2}\cdots b_1\,b_0)_2

with leading digit 11, the remainder is r=(0 bm−1⋯b1 b0)2r = (0\,b_{m-1}\cdots b_1\,b_0)_2, so 2r=(bm−1⋯b1 b0 0)22r = (b_{m-1}\cdots b_1\,b_0\,0)_2 and

J(n)=2r+1=(bm−1 bm−2⋯b1 b0 1)2.J(n) = 2r + 1 = (b_{m-1}\,b_{m-2}\cdots b_1\,b_0\,1)_2 .

Since the leading digit that was dropped is a 11, the answer is obtained from nn by moving its leading digit to the end: a one-bit cyclic shift to the left.

Example 2.13 (A cyclic shift).

Take n=100=(1100100)2n = 100 = (1100100)_2. Shifting the leading digit round to the end gives (1001001)2=73(1001001)_2 = 73, so J(100)=73J(100) = 73. Checking against the closed form, 100=64+36100 = 64 + 36, so r=36r = 36 and 2r+1=732r + 1 = 73.

Remark (Iterating the shift).

Applying JJ repeatedly does not cycle back to nn after m+1m+1 shifts, because J(n)⩽nJ(n) \leqslant n always, and once a value drops it can never climb again. What happens instead is that whenever the leading digit is a 00 it is dropped, so each application deletes a zero. After enough applications only the ones remain, and the value settles at

(11⋯1⏟ν(n))2=2ν(n)−1,(\underbrace{11\cdots1}_{\nu(n)})_2 = 2^{\nu(n)} - 1,

where ν(n)\nu(n) is the number of ones in the binary representation of nn. For instance J(11)=J((1011)2)=(111)2=7J(11) = J((1011)_2) = (111)_2 = 7, and J(7)=7J(7) = 7.

Proposition 2.14 (When the survivor is halfway round).

Let n=2m+rn = 2^m + r with 0⩽r<2m0 \leqslant r < 2^m. Then J(n)=n/2J(n) = n/2 if and only if mm is odd and r=(2m−2)/3r = (2^m - 2)/3.

Discussion.

Substituting the closed form turns the condition J(n)=n/2J(n) = n/2 into a linear equation in rr and 2m2^m, and solving it gives r=(2m−2)/3r = (2^m-2)/3; it remains to find the mm for which that number is an integer lying below 2m2^m. The bound is immediate. Integrality is a statement about 2m2^m modulo 33, and the powers of two alternate between 11 and 22 there, since doubling exchanges the two; so 2m−22^m - 2 is divisible by 33 exactly for odd mm.

Proof.

By the closed form, J(n)=n/2J(n) = n/2 says 2r+1=(2m+r)/22r + 1 = (2^m + r)/2, that is 4r+2=2m+r4r + 2 = 2^m + r, that is

3r=2m−2,sor=2m−23.3r = 2^m - 2, \qquad\text{so}\qquad r = \frac{2^m - 2}{3}.

This rr satisfies r<2mr < 2^m automatically. It is an integer exactly when 2m≡2(mod3)2^m \equiv 2 \pmod 3. Now 20=12^0 = 1 leaves remainder 11, and doubling a number that leaves remainder 11 gives one that leaves remainder 22, while doubling a number that leaves remainder 22 gives 44, which leaves remainder 11. So the remainders alternate 1,2,1,2,…1, 2, 1, 2, \ldots as m=0,1,2,3,…m = 0, 1, 2, 3, \ldots, and 2m≡22^m \equiv 2 exactly when mm is odd.

The first few such nn are

mrn=2m+rJ(n)=n/2n in binary10211032105101051042211010107421708510101010\begin{array}{c|c|c|c|c} m & r & n = 2^m + r & J(n) = n/2 & n \text{ in binary} \\ \hline 1 & 0 & 2 & 1 & 10 \\ 3 & 2 & 10 & 5 & 1010 \\ 5 & 10 & 42 & 21 & 101010 \\ 7 & 42 & 170 & 85 & 10101010 \end{array}

These are the nn for which shifting the leading bit to the end has the same effect as deleting the last bit.

The Repertoire Method

Guessing worked for the Josephus recurrence because its answers were small numbers with a visible pattern. When no guess is available, the following method finds the formula. Consider the recurrence with three undetermined constants,

f(n)={α,n=1,2f(ℓ)+β,n=2ℓ,  ℓ∈N,2f(ℓ)+γ,n=2ℓ+1,  ℓ∈N,f(n) = \begin{cases} \alpha, & n = 1, \\ 2f(\ell) + \beta, & n = 2\ell, \; \ell \in \mathbb{N}, \\ 2f(\ell) + \gamma, & n = 2\ell + 1, \; \ell \in \mathbb{N}, \end{cases}

of which the Josephus recurrence is the case α=1\alpha = 1, β=−1\beta = -1, γ=1\gamma = 1. Its first few values are

nf(n)1α22α+β32α+γ44α+3β54α+2β+γ64α+β+2γ74α+3γ88α+7β98α+6β+γ\begin{array}{c|l} n & f(n) \\ \hline 1 & \alpha \\ 2 & 2\alpha + \beta \\ 3 & 2\alpha + \gamma \\ 4 & 4\alpha + 3\beta \\ 5 & 4\alpha + 2\beta + \gamma \\ 6 & 4\alpha + \beta + 2\gamma \\ 7 & 4\alpha + 3\gamma \\ 8 & 8\alpha + 7\beta \\ 9 & 8\alpha + 6\beta + \gamma \end{array}

Every entry is a combination of α\alpha, β\beta and γ\gamma with coefficients depending only on nn, so we may write

f(n)=A(n) α+B(n) β+C(n) γ,f(n) = A(n)\,\alpha + B(n)\,\beta + C(n)\,\gamma ,

and the task becomes finding the three functions AA, BB, CC. The method is to substitute values of α,β,γ\alpha, \beta, \gamma, or functions ff, for which the recurrence can be solved by inspection; each substitution yields one equation relating AA, BB and CC, and three independent equations determine them.

Proposition 2.15 (The coefficient of α\alpha).

Let n=2m+rn = 2^m + r with 0⩽r<2m0 \leqslant r < 2^m. Then A(n)=2mA(n) = 2^m.

Discussion.

Setting α=1\alpha = 1 and β=γ=0\beta = \gamma = 0 makes f(n)f(n) equal to A(n)A(n) and collapses the recurrence to A(1)=1A(1) = 1 with A(n)=2A(⌊n/2⌋)A(n) = 2A(\lfloor n/2 \rfloor) in both the even and the odd case. That recurrence doubles once per halving and never distinguishes the two cases, so it depends on nn only through how often nn can be halved, which is mm. The induction is the same strong induction as before, using that ⌊n/2⌋\lfloor n/2 \rfloor has exponent m−1m - 1 in its own decomposition.

Proof.

Putting α=1\alpha = 1 and β=γ=0\beta = \gamma = 0 gives f=Af = A and

A(1)=1,A(2ℓ)=2A(ℓ),A(2ℓ+1)=2A(ℓ).A(1) = 1, \qquad A(2\ell) = 2A(\ell), \qquad A(2\ell+1) = 2A(\ell) .

We induct on nn. For n=1n = 1 we have m=0m = 0 and A(1)=1=20A(1) = 1 = 2^0. For n⩾2n \geqslant 2 write n=2m+rn = 2^m + r with m⩾1m \geqslant 1, and put ℓ=⌊n/2⌋\ell = \lfloor n/2 \rfloor. As in the proof of the closed form for JJ, the decomposition of ℓ\ell has exponent m−1m-1, whichever parity nn has. Both clauses of the recurrence give A(n)=2A(ℓ)A(n) = 2A(\ell), and the inductive hypothesis gives A(ℓ)=2m−1A(\ell) = 2^{m-1}, so A(n)=2mA(n) = 2^m.

Two more substitutions finish the job, and this time we choose the function rather than the constants.

Take f(n)=1f(n) = 1 for every nn. The three clauses become 1=α1 = \alpha, 1=2+β1 = 2 + \beta and 1=2+γ1 = 2 + \gamma, so this constant function is the solution for α=1\alpha = 1, β=γ=−1\beta = \gamma = -1. Substituting those values into f(n)=A(n)α+B(n)β+C(n)γf(n) = A(n)\alpha + B(n)\beta + C(n)\gamma gives

1=A(n)−B(n)−C(n).1 = A(n) - B(n) - C(n) .

Take f(n)=nf(n) = n. The clauses become 1=α1 = \alpha, 2ℓ=2ℓ+β2\ell = 2\ell + \beta and 2ℓ+1=2ℓ+γ2\ell + 1 = 2\ell + \gamma, so this is the solution for α=1\alpha = 1, β=0\beta = 0, γ=1\gamma = 1, and

n=A(n)+C(n).n = A(n) + C(n) .

Now solve. From A(n)=2mA(n) = 2^m and A(n)+C(n)=n=2m+rA(n) + C(n) = n = 2^m + r we get C(n)=rC(n) = r; and then B(n)=A(n)−C(n)−1=2m−1−rB(n) = A(n) - C(n) - 1 = 2^m - 1 - r.

Theorem 2.16 (Solution of the generalised Josephus recurrence).

Let α,β,γ\alpha, \beta, \gamma be constants and let ff satisfy the recurrence above. Then for n=2m+rn = 2^m + r with 0⩽r<2m0 \leqslant r < 2^m,

f(n)=2mα+(2m−1−r)β+rγ.f(n) = 2^m \alpha + \bigl(2^m - 1 - r\bigr)\beta + r\gamma .

Discussion.

The repertoire method produced this formula but did not prove it, since it assumed at the outset that f(n)f(n) is a combination of α\alpha, β\beta and γ\gamma with coefficients independent of them. That assumption is easy to justify after the fact: the right-hand side is a specific function of nn, so it is enough to check that it satisfies the three clauses of the recurrence, and a solution of the recurrence is unique because the clauses determine f(n)f(n) from f(⌊n/2⌋)f(\lfloor n/2 \rfloor) and f(1)f(1) is given. The verification is the same case split on the parity of nn used twice already, with the decomposition of ⌊n/2⌋\lfloor n/2 \rfloor read off as before.

Proof.

Write F(2m+r)=2mα+(2m−1−r)β+rγF(2^m + r) = 2^m\alpha + (2^m - 1 - r)\beta + r\gamma for 0⩽r<2m0 \leqslant r < 2^m. Since the clauses determine f(n)f(n) from f(⌊n/2⌋)f(\lfloor n/2\rfloor) for n⩾2n \geqslant 2 and fix f(1)f(1), at most one function satisfies them; so it suffices to check that FF does.

At n=1n = 1 we have m=r=0m = r = 0 and F(1)=αF(1) = \alpha.

Let n=2ℓ⩾2n = 2\ell \geqslant 2, so m⩾1m \geqslant 1 and rr is even, and ℓ=2m−1+r/2\ell = 2^{m-1} + r/2 with 0⩽r/2<2m−10 \leqslant r/2 < 2^{m-1}. Then

2F(ℓ)+β=2[2m−1α+(2m−1−1−r2)β+r2γ]+β=2mα+(2m−2−r+1)β+rγ,2F(\ell) + \beta = 2\left[2^{m-1}\alpha + \left(2^{m-1} - 1 - \tfrac{r}{2}\right)\beta + \tfrac{r}{2}\gamma\right] + \beta = 2^m\alpha + \bigl(2^m - 2 - r + 1\bigr)\beta + r\gamma,

which is F(n)F(n).

Let n=2ℓ+1⩾3n = 2\ell + 1 \geqslant 3, so m⩾1m \geqslant 1 and rr is odd, and ℓ=2m−1+(r−1)/2\ell = 2^{m-1} + (r-1)/2 with 0⩽(r−1)/2<2m−10 \leqslant (r-1)/2 < 2^{m-1}. Then

2F(ℓ)+γ=2[2m−1α+(2m−1−1−r−12)β+r−12γ]+γ=2mα+(2m−1−r)β+rγ,2F(\ell) + \gamma = 2\left[2^{m-1}\alpha + \left(2^{m-1} - 1 - \tfrac{r-1}{2}\right)\beta + \tfrac{r-1}{2}\gamma\right] + \gamma = 2^m\alpha + \bigl(2^m - 1 - r\bigr)\beta + r\gamma,

which is F(n)F(n) again.

Setting α=1\alpha = 1, β=−1\beta = -1, γ=1\gamma = 1 recovers 2m−(2m−1−r)+r=2r+12^m - (2^m - 1 - r) + r = 2r + 1, the Josephus answer.

Radix Notation

The generalised recurrence has a shorter description if the two constants added at each step are indexed by the bit that decides between them. Write β0=β\beta_0 = \beta and β1=γ\beta_1 = \gamma; then the three clauses become

f(1)=α,f(2ℓ+j)=2f(ℓ)+βj(j∈{0,1}),f(1) = \alpha, \qquad f(2\ell + j) = 2f(\ell) + \beta_j \quad (j \in \{0, 1\}),

and jj is the last binary digit of nn. Unrolling the second clause strips one digit at a time:

f((1 bm−1⋯b1b0)2)=2f((1 bm−1⋯b1)2)+βb0=4f((1 bm−1⋯b2)2)+2βb1+βb0    ⋮=2mα+2m−1βbm−1+⋯+2βb1+βb0.\begin{aligned} f\bigl((1\,b_{m-1}\cdots b_1 b_0)_2\bigr) &= 2 f\bigl((1\,b_{m-1}\cdots b_1)_2\bigr) + \beta_{b_0} \\ &= 4 f\bigl((1\,b_{m-1}\cdots b_2)_2\bigr) + 2\beta_{b_1} + \beta_{b_0} \\ &\;\;\vdots \\ &= 2^m \alpha + 2^{m-1}\beta_{b_{m-1}} + \cdots + 2\beta_{b_1} + \beta_{b_0} . \end{aligned}

The right-hand side has the shape of a binary expansion whose digits are the constants rather than 00 and 11, so we write it

f(n)=(α  βbm−1  βbm−2⋯βb1  βb0)2,f(n) = \bigl(\alpha\;\beta_{b_{m-1}}\;\beta_{b_{m-2}}\cdots\beta_{b_1}\;\beta_{b_0}\bigr)_2 ,

meaning that each entry is multiplied by the appropriate power of two and the results added. Recovering the earlier table is a matter of reading off digits: 6=(110)26 = (110)_2 gives 22α+2γ+β2^2\alpha + 2\gamma + \beta, and 9=(1001)29 = (1001)_2 gives 23α+22β+2β+γ2^3\alpha + 2^2\beta + 2\beta + \gamma.

Example 2.17 (The Josephus survivor for n=100n = 100).

With α=1\alpha = 1, β0=−1\beta_0 = -1 and β1=1\beta_1 = 1, and 100=(1100100)2100 = (1100100)_2,

J(100)=26(1)+25(1)+24(−1)+23(−1)+22(1)+2(−1)+(−1)=64+32−16−8+4−2−1=73,\begin{aligned} J(100) &= 2^6(1) + 2^5(1) + 2^4(-1) + 2^3(-1) + 2^2(1) + 2(-1) + (-1) \\ &= 64 + 32 - 16 - 8 + 4 - 2 - 1 = 73 , \end{aligned}

agreeing with the cyclic shift (1100100)2↦(1001001)2(1100100)_2 \mapsto (1001001)_2.

Nothing in the argument used the base 22 or the multiplier 22 separately, and separating them gives the general statement.

Theorem 2.18 (Recurrences that strip a digit).

Let d⩾2d \geqslant 2 and cc be constants, let α1,…,αd−1\alpha_1, \ldots, \alpha_{d-1} and β0,…,βd−1\beta_0, \ldots, \beta_{d-1} be constants, and let ff satisfy

f(i)=αi(1⩽i⩽d−1),f(dn+j)=c f(n)+βj(n⩾1,  0⩽j<d).f(i) = \alpha_i \quad (1 \leqslant i \leqslant d-1), \qquad f(dn + j) = c\,f(n) + \beta_j \quad (n \geqslant 1, \; 0 \leqslant j < d).

Then for n=(bm bm−1⋯b1 b0)dn = (b_m\,b_{m-1}\cdots b_1\,b_0)_d with bm≠0b_m \neq 0,

f(n)=(αbm  βbm−1⋯βb1  βb0)c=cmαbm+cm−1βbm−1+⋯+c βb1+βb0.f(n) = \bigl(\alpha_{b_m}\;\beta_{b_{m-1}}\cdots\beta_{b_1}\;\beta_{b_0}\bigr)_c = c^m \alpha_{b_m} + c^{m-1}\beta_{b_{m-1}} + \cdots + c\,\beta_{b_1} + \beta_{b_0}.

Discussion.

The second clause removes exactly one base-dd digit of its argument, because writing n=dn′+jn = dn' + j with 0⩽j<d0 \leqslant j < d is the same as splitting off the last digit b0=jb_0 = j and leaving n′=(bm⋯b1)dn' = (b_m \cdots b_1)_d. So each application of the clause removes a digit and multiplies what has accumulated by cc, and after mm applications the argument has been reduced to its leading digit bmb_m, which lies between 11 and d−1d-1 and is therefore covered by the first clause. The proof is an induction on the number of digits, and the point to check is that the split of nn into dn′+jdn' + j corresponds to the split of the digit string, which is where the uniqueness of base-dd representation is used.

Proof.

We induct on mm, the number of digits after the leading one.

If m=0m = 0 then n=b0n = b_0 with 1⩽b0⩽d−11 \leqslant b_0 \leqslant d-1, and the first clause gives f(n)=αb0f(n) = \alpha_{b_0}, which is the claim.

Let m⩾1m \geqslant 1 and write n=(bm⋯b1b0)dn = (b_m \cdots b_1 b_0)_d, that is

n=bmdm+bm−1dm−1+⋯+b1d+b0=d(bmdm−1+⋯+b1)+b0.n = b_m d^m + b_{m-1}d^{m-1} + \cdots + b_1 d + b_0 = d\bigl(b_m d^{m-1} + \cdots + b_1\bigr) + b_0 .

Setting n′=(bm⋯b1)dn' = (b_m \cdots b_1)_d and j=b0j = b_0, we have n=dn′+jn = dn' + j with 0⩽j<d0 \leqslant j < d and n′⩾1n' \geqslant 1. The second clause and the inductive hypothesis applied to n′n', which has m−1m-1 digits after its leading one, give

f(n)=c f(n′)+βb0=c(cm−1αbm+cm−2βbm−1+⋯+βb1)+βb0,f(n) = c\,f(n') + \beta_{b_0} = c\left(c^{m-1}\alpha_{b_m} + c^{m-2}\beta_{b_{m-1}} + \cdots + \beta_{b_1}\right) + \beta_{b_0},

and expanding the bracket is the claim at mm.

Representation in a Base

The theorem above took for granted that nn has a base-dd digit string and that the string is determined by nn. Both facts need proof, and the tools are the floor function and the remainder.

Definition 2.19 (Floor, ceiling and remainder).

For x∈Rx \in \mathbb{R},

⌊x⌋=max⁡{k∈Z∣k⩽x},⌈x⌉=min⁡{k∈Z∣k⩾x}.\lfloor x \rfloor = \max\{k \in \mathbb{Z} \mid k \leqslant x\}, \qquad \lceil x \rceil = \min\{k \in \mathbb{Z} \mid k \geqslant x\} .

For x∈Zx \in \mathbb{Z} and y∈Ny \in \mathbb{N} the remainder of xx on division by yy is

x mod y=x−y⌊xy⌋.x \bmod y = x - y\left\lfloor \frac{x}{y} \right\rfloor .

Rearranging the last definition gives the identity we shall use repeatedly: for z∈N0z \in \mathbb{N}_0 and b∈Nb \in \mathbb{N},

z=b⌊zb⌋+(z mod b),0⩽z mod b<b.z = b\left\lfloor \frac{z}{b} \right\rfloor + (z \bmod b), \qquad 0 \leqslant z \bmod b < b .

Example 2.20 (Floors and remainders).

⌊13.2⌋=13\lfloor 13.2 \rfloor = 13 and ⌈13.2⌉=14\lceil 13.2 \rceil = 14. For the remainder,

23 mod 6=23−6⌊236⌋=23−6⋅3=5,23=3⋅6+5.23 \bmod 6 = 23 - 6\left\lfloor \tfrac{23}{6} \right\rfloor = 23 - 6 \cdot 3 = 5, \qquad 23 = 3 \cdot 6 + 5 .

Theorem 2.21 (Representation in a base).

Let b,n∈Nb, n \in \mathbb{N} with b>1b > 1. Every z∈N0z \in \mathbb{N}_0 with 0⩽z⩽bn−10 \leqslant z \leqslant b^n - 1 can be written as

z=∑i=0n−1zibi,zi∈{0,1,…,b−1},z = \sum_{i=0}^{n-1} z_i b^i, \qquad z_i \in \{0, 1, \ldots, b-1\},

and the digits z0,…,zn−1z_0, \ldots, z_{n-1} are uniquely determined by zz.

Discussion.

There are two claims, and they are proved by different means.

Existence is an induction on the number of digits nn, and the step is division by bb. Given zz below bn+1b^{n+1}, split it as z=bz^+z0z = b\hat{z} + z_0 with z0=z mod bz_0 = z \bmod b the last digit and z^=⌊z/b⌋\hat{z} = \lfloor z/b \rfloor what remains. The point to check is that z^\hat{z} falls in the range the inductive hypothesis covers, namely below bnb^n, and it does because dividing by bb shrinks the bound by a factor of bb. The hypothesis then supplies nn digits for z^\hat{z}, and multiplying the whole expansion by bb shifts every digit up one place, leaving room at the bottom for z0z_0.

Uniqueness is an argument by contradiction that needs no induction. Suppose two different digit strings give the same value, and look at the highest place mm where they disagree. Above mm the terms are equal and cancel, so the difference of the two sums is a single term (zm−z^m)bm(z_m - \hat{z}_m)b^m, which is at least bmb^m, plus lower terms, each of which is at least −(b−1)bi-(b-1)b^i. Summing those lower bounds telescopes to 1−bm1 - b^m, so the total is at least bm+1−bm=1b^m + 1 - b^m = 1; but the total is 00.

Proof.

Existence. We induct on nn.

For n=0n = 0 the only zz in range is z=0z = 0, represented by the empty sum.

Suppose every z^\hat{z} with 0⩽z^⩽bn−10 \leqslant \hat{z} \leqslant b^n - 1 has a representation with nn digits, and let 0⩽z⩽bn+1−10 \leqslant z \leqslant b^{n+1} - 1. Put

z0=z mod b,z^=⌊zb⌋,soz=bz^+z0,0⩽z0<b.z_0 = z \bmod b, \qquad \hat{z} = \left\lfloor \frac{z}{b} \right\rfloor, \qquad\text{so}\qquad z = b\hat{z} + z_0, \quad 0 \leqslant z_0 < b .

We check that z^\hat{z} lies in the range covered by the hypothesis. From z⩽bn+1−1<bn+1z \leqslant b^{n+1} - 1 < b^{n+1} we get z/b<bnz/b < b^n, and z^⩽z/b\hat z \leqslant z/b, so z^<bn\hat{z} < b^n; being an integer, z^⩽bn−1\hat{z} \leqslant b^n - 1. Also z^⩾0\hat{z} \geqslant 0. So the hypothesis applies and gives digits z^0,…,z^n−1\hat{z}_0, \ldots, \hat{z}_{n-1} in {0,…,b−1}\{0, \ldots, b-1\} with z^=∑i=0n−1z^ibi\hat{z} = \sum_{i=0}^{n-1}\hat{z}_i b^i. Then

z=b∑i=0n−1z^ibi+z0=∑i=0n−1z^ibi+1+z0=∑i=1nz^i−1bi+z0=∑i=0nzibi,z = b\sum_{i=0}^{n-1}\hat{z}_i b^i + z_0 = \sum_{i=0}^{n-1}\hat{z}_i b^{i+1} + z_0 = \sum_{i=1}^{n}\hat{z}_{i-1} b^{i} + z_0 = \sum_{i=0}^{n} z_i b^i,

where the third step reindexed the sum by i↦i+1i \mapsto i+1 and the last set zi=z^i−1z_i = \hat{z}_{i-1} for 1⩽i⩽n1 \leqslant i \leqslant n, with z0z_0 as chosen. Every digit lies in {0,…,b−1}\{0, \ldots, b-1\}, and there are n+1n+1 of them, which is the claim for n+1n+1.

Uniqueness. Suppose

∑i=0n−1zibi=∑i=0n−1z^ibi\sum_{i=0}^{n-1} z_i b^i = \sum_{i=0}^{n-1} \hat{z}_i b^i

with all digits in {0,…,b−1}\{0, \ldots, b-1\}, and suppose the two strings are not identical. Let

m=max⁡{ i∣0⩽i⩽n−1,  zi≠z^i },m = \max\{\, i \mid 0 \leqslant i \leqslant n-1, \; z_i \neq \hat{z}_i \,\},

and assume without loss of generality that zm>z^mz_m > \hat{z}_m. All terms with i>mi > m agree and cancel, so subtracting one sum from the other leaves

0=∑i=0m(zi−z^i) bi=(zm−z^m) bm+∑i=0m−1(zi−z^i) bi.0 = \sum_{i=0}^{m} (z_i - \hat{z}_i)\,b^i = (z_m - \hat{z}_m)\,b^m + \sum_{i=0}^{m-1}(z_i - \hat{z}_i)\,b^i .

Now bound the two pieces from below. The digits are integers with zm>z^mz_m > \hat z_m, so zm−z^m⩾1z_m - \hat{z}_m \geqslant 1 and the first term is at least bmb^m. For i<mi < m we have zi⩾0z_i \geqslant 0 and z^i⩽b−1\hat{z}_i \leqslant b-1, so zi−z^i⩾1−bz_i - \hat{z}_i \geqslant 1 - b and

∑i=0m−1(zi−z^i) bi  ⩾  (1−b)∑i=0m−1bi=∑i=0m−1bi−∑i=0m−1bi+1=∑i=0m−1bi−∑i=1mbi=b0−bm,\sum_{i=0}^{m-1}(z_i - \hat{z}_i)\,b^i \;\geqslant\; (1-b)\sum_{i=0}^{m-1} b^i = \sum_{i=0}^{m-1}b^i - \sum_{i=0}^{m-1}b^{i+1} = \sum_{i=0}^{m-1}b^i - \sum_{i=1}^{m}b^{i} = b^0 - b^m,

the last step because the two sums share every term with 1⩽i⩽m−11 \leqslant i \leqslant m-1. Adding the two bounds,

0  ⩾  bm+(1−bm)=1,0 \;\geqslant\; b^m + \bigl(1 - b^m\bigr) = 1,

which is false. So no two distinct digit strings represent the same zz.

Corollary 2.22 (Radix notation).

Let b>1b > 1. For a digit string zn−1,…,z0z_{n-1}, \ldots, z_0 with entries in {0,…,b−1}\{0, \ldots, b-1\} write

(zn−1 zn−2⋯z1 z0)b=∑i=0n−1zibi.(z_{n-1}\,z_{n-2}\cdots z_1\,z_0)_b = \sum_{i=0}^{n-1} z_i b^i .

Then every z∈Nz \in \mathbb{N} is (zn−1⋯z0)b(z_{n-1}\cdots z_0)_b for exactly one digit string with leading digit zn−1≠0z_{n-1} \neq 0, and its length is n=⌊log⁡bz⌋+1n = \lfloor \log_b z \rfloor + 1.

Proof.

Let z∈Nz \in \mathbb{N} and put n=⌊log⁡bz⌋+1n = \lfloor \log_b z \rfloor + 1, so that n−1⩽log⁡bz<nn - 1 \leqslant \log_b z < n and hence bn−1⩽z<bnb^{n-1} \leqslant z < b^n. The theorem applies with this nn and gives digits z0,…,zn−1z_0, \ldots, z_{n-1}, unique for this length. The leading digit is not zero: if zn−1=0z_{n-1} = 0 then z=∑i=0n−2zibi⩽(b−1)∑i=0n−2bi=bn−1−1z = \sum_{i=0}^{n-2}z_ib^i \leqslant (b-1)\sum_{i=0}^{n-2}b^i = b^{n-1} - 1, contradicting z⩾bn−1z \geqslant b^{n-1}.

Conversely, suppose z=(zn′−1⋯z0)bz = (z_{n'-1}\cdots z_0)_b for some length n′n' with zn′−1≠0z_{n'-1} \neq 0. Then z⩾bn′−1z \geqslant b^{n'-1}, and also z⩽(b−1)∑i=0n′−1bi=bn′−1<bn′z \leqslant (b-1)\sum_{i=0}^{n'-1}b^i = b^{n'} - 1 < b^{n'}. So bn′−1⩽z<bn′b^{n'-1} \leqslant z < b^{n'}, which forces n′=nn' = n, and the theorem then makes the digits the ones already found.

Remark (Horner's scheme).

Evaluating (zn−1⋯z0)b(z_{n-1}\cdots z_0)_b by computing each power bib^i and multiplying costs about 2n2n multiplications. Nesting the expression,

∑i=0n−1zibi=z0+b(z1+b(z2+⋯+b(zn−2+b zn−1)⋯ )),\sum_{i=0}^{n-1} z_i b^i = z_0 + b\Bigl(z_1 + b\bigl(z_2 + \cdots + b(z_{n-2} + b\,z_{n-1})\cdots\bigr)\Bigr),

evaluates it from the inside out with n−1n-1 multiplications and n−1n-1 additions, and never forms a power of bb explicitly. This is Horner’s scheme, and the same nesting evaluates any polynomial at any point.

Problem 2.12.

Compute J(1000)J(1000) two ways, from the closed form and from the cyclic shift, and check that they agree.

Problem 2.13.

Show that J(n)=nJ(n) = n if and only if n=2k−1n = 2^k - 1 for some k∈N0k \in \mathbb{N}_0.

Problem 2.14.

For which nn does person 11 survive? Show that person 22 never survives, whatever nn may be.

Problem 2.15.

Use the repertoire method to solve

f(1)=α,f(2ℓ)=3f(ℓ)+β,f(2ℓ+1)=3f(ℓ)+γ.f(1) = \alpha, \qquad f(2\ell) = 3f(\ell) + \beta, \qquad f(2\ell + 1) = 3f(\ell) + \gamma .

Problem 2.16.

Write ν(n)\nu(n) for the number of ones in the binary representation of nn. Prove that repeated application of JJ to any n∈Nn \in \mathbb{N} eventually reaches 2ν(n)−12^{\nu(n)} - 1 and stays there.

Problem 2.17.

Let b>1b > 1 and k∈Nk \in \mathbb{N}. Show that the last kk digits of z∈Nz \in \mathbb{N} in base bb are the digits of z mod bkz \bmod b^{k}, padded with leading zeros if need be.

Sums

A recurrence adds one term to what came before, so every recurrence of the form Sn=Sn−1+anS_n = S_{n-1} + a_n is a sum, and every sum is such a recurrence. We set up the notation for sums and the laws for rearranging them, and use the correspondence with recurrences to evaluate sums.

Sequences

Definition 2.23 (Sequence).

A sequence of elements of a set AA is a function f:N0→Af : \mathbb{N}_0 \to A. The value f(n)f(n) is the nnth term and is written ana_n, and the sequence itself is written

{an}n∈N0,or {an} when the index set is understood.\{a_n\}_{n \in \mathbb{N}_0}, \qquad \text{or } \{a_n\} \text{ when the index set is understood.}

A sequence can be given either as a function or as the list of its values. Defining f:N0→Rf : \mathbb{N}_0 \to \mathbb{R} by f(n)=n+nf(n) = n + \sqrt{n} is the same as writing an=n+na_n = n + \sqrt{n}, and the same again as writing out

0,2,2+2,3+3,…,n+n,…0, \quad 2, \quad 2 + \sqrt{2}, \quad 3 + \sqrt{3}, \quad \ldots, \quad n + \sqrt{n}, \quad \ldots

A sequence in this sense is infinite, since its domain is, and it is countably infinite: the terms can be laid out in a list indexed by 0,1,2,…0, 1, 2, \ldots with no term left out.

Definition 2.24 (Countably infinite).

A set TT is countably infinite if there is a bijection N0→T\mathbb{N}_0 \to T, and we then write #T=#N0\#T = \#\mathbb{N}_0.

Removing finitely many elements from N0\mathbb{N}_0 leaves a countably infinite set, so the index set of a sequence need not be N0\mathbb{N}_0 itself; any countably infinite set of indices will do, and the terms can then be listed in the order that a bijection with N0\mathbb{N}_0 supplies.

Example 2.25 (Indexing by another set).

The even numbers E={0,2,4,…}E = \{0, 2, 4, \ldots\} are countably infinite, since n↦2nn \mapsto 2n is a bijection N0→E\mathbb{N}_0 \to E. So {ak}k∈E\{a_k\}_{k \in E} is a sequence, with terms a0,a2,a4,…a_0, a_2, a_4, \ldots listed in that order.

A formula may be undefined at some indices. For

an=n(n−2)(n−5)a_n = \frac{n}{(n-2)(n-5)}

the denominator vanishes at n=2n = 2 and n=5n = 5, so the sequence is the function f:N∖{2,5}→Rf : \mathbb{N} \setminus \{2, 5\} \to \mathbb{R}, and its index set is again countably infinite.

Definition 2.26 (Finite sequence).

A finite sequence of elements of AA is a function f:K→Af : K \to A with KK finite. When KK is a non-empty set of natural numbers we take K={1,2,…,n}K = \{1, 2, \ldots, n\} and call nn the length, writing f={ak}k=1,…,nf = \{a_k\}_{k=1,\ldots,n}. For K=∅K = \emptyset the function is empty and ff is the empty sequence, written ε\varepsilon.

Sigma Notation

Definition 2.27 (Finite sum).

Let {ak}\{a_k\} be a sequence of real numbers. For n∈Nn \in \mathbb{N} we write

∑k=1nak=a1+a2+⋯+an.\sum_{k=1}^{n} a_k = a_1 + a_2 + \cdots + a_n .

Each aka_k appearing is a term, the expression aka_k after the ∑\textstyle\sum is the summand, and kk is the index of summation, which is bound to the ∑\textstyle\sum and has no meaning outside it.

The notation says: include exactly those terms aka_k whose index is an integer between the lower and upper limits, inclusive. That reading has a delimited form and a general form, and the two are interchangeable:

∑k=1nak  =  ∑1⩽k⩽nak  =  ∑k∈{1,…,n}ak  =  ∑Kakfor K={1,…,n}.\sum_{k=1}^{n} a_k \;=\; \sum_{1 \leqslant k \leqslant n} a_k \;=\; \sum_{k \in \{1, \ldots, n\}} a_k \;=\; \sum_{K} a_k \quad\text{for } K = \{1, \ldots, n\}.

More generally, one or more conditions written under the ∑\textstyle\sum specify which indices take part. The sum of the squares of the odd positive integers below 100100 is

∑1⩽k<100k oddk2,\sum_{\substack{1 \leqslant k < 100 \\ k \text{ odd}}} k^2 ,

whose delimited form ∑k=049(2k+1)2\sum_{k=0}^{49}(2k+1)^2 is harder to read; and the sum of the reciprocals of the primes up to NN is

∑p⩽Np prime1p,\sum_{\substack{p \leqslant N \\ p \text{ prime}}} \frac{1}{p} ,

whose delimited form needs a function counting the primes before it can even be written down.

The general form also makes a change of index easier. Replacing kk by k+1k+1 turns

∑1⩽k⩽nakinto∑1⩽k+1⩽nak+1,\sum_{1 \leqslant k \leqslant n} a_k \qquad\text{into}\qquad \sum_{1 \leqslant k+1 \leqslant n} a_{k+1} ,

where the substitution is made directly, while in delimited form the same change reads

∑k=1nak=∑k=0n−1ak+1,\sum_{k=1}^{n} a_k = \sum_{k=0}^{n-1} a_{k+1} ,

and the limits have to be recomputed by hand.

Definition 2.28 (Summation over a property).

Let P(k)P(k) be a statement about integers which is true or false for each kk, and suppose only finitely many kk with P(k)P(k) true have ak≠0a_k \neq 0. Then

∑P(k)ak=∑k∈Kak=∑Kak,K={ k∈Z∣P(k) }.\sum_{P(k)} a_k = \sum_{k \in K} a_k = \sum_{K} a_k, \qquad K = \{\, k \in \mathbb{Z} \mid P(k) \,\} .

If P(k)P(k) is false for every kk the sum is empty, and its value is 00.

Example 2.29 (Reading a condition).

Let P(n)P(n) be the property ”1⩽n<1001 \leqslant n < 100 and nn is odd”. Then

K={ n∈N∣P(n) }={1,3,5,…,99}={ 2j+1∣0⩽j⩽49 },K = \{\, n \in \mathbb{N} \mid P(n) \,\} = \{1, 3, 5, \ldots, 99\} = \{\, 2j+1 \mid 0 \leqslant j \leqslant 49 \,\},

so

∑P(n)an=∑Kak=∑j=049a2j+1=a1+a3+⋯+a99.\sum_{P(n)} a_n = \sum_{K} a_k = \sum_{j=0}^{49} a_{2j+1} = a_1 + a_3 + \cdots + a_{99} .

Dropping the parity condition gives K={1,2,…,99}K = \{1, 2, \ldots, 99\} and ∑P(n)an=∑k=199ak\sum_{P(n)}a_n = \sum_{k=1}^{99}a_k. Taking the summand to be an=(2n+1)2a_n = (2n+1)^2 instead gives

∑1⩽n<100(2n+1)2=32+52+⋯+1992.\sum_{1 \leqslant n < 100} (2n+1)^2 = 3^2 + 5^2 + \cdots + 199^2 .

Remark (Keeping limits simple).

Terms equal to zero do no harm, and it is usually simpler to leave them in. In

∑k=0nk(k−1)(n−k)\sum_{k=0}^{n} k(k-1)(n-k)

the terms at k=0k = 0, k=1k = 1 and k=nk = n all vanish, and one is tempted to write ∑k=2n−1\sum_{k=2}^{n-1} instead. That form is worse: it is harder to manipulate, and its meaning is unclear when n=0n = 0 or n=1n = 1.

Remark (Two conventions for empty sums).

An empty sum is 00, which fixes the value of ∑k=abak\sum_{k=a}^{b}a_k when b<ab < a:

∑k=abak=0whenever b<a.\sum_{k=a}^{b} a_k = 0 \qquad \text{whenever } b < a .

In particular ∑k=10ak=0\sum_{k=1}^{0}a_k = 0 and ∑k=0−1ak=0\sum_{k=0}^{-1}a_k = 0. With this convention the splitting rules hold without exception: for every n∈N0n \in \mathbb{N}_0,

∑k=1n+1ak=∑k=1nak+an+1,∑k=0nak=∑k=0n−1ak+an,\sum_{k=1}^{n+1} a_k = \sum_{k=1}^{n} a_k + a_{n+1}, \qquad \sum_{k=0}^{n} a_k = \sum_{k=0}^{n-1} a_k + a_n ,

the second reading correctly at n=0n = 0 as a0=0+a0a_0 = 0 + a_0.

Iverson Brackets

Kenneth Iverson introduced a device that removes conditions from beneath the ∑\textstyle\sum altogether.

Definition 2.30 (Iverson bracket).

For a statement PP that is either true or false, write

[P]={1if P is true,0if P is false.[P] = \begin{cases} 1 & \text{if } P \text{ is true}, \\ 0 & \text{if } P \text{ is false}. \end{cases}

A term multiplied by a bracket that evaluates to 00 is taken to be 00 even when the other factor is undefined.

With brackets, a sum over a condition becomes a sum over all integers,

∑P(k)ak=∑kak [P(k)],\sum_{P(k)} a_k = \sum_{k} a_k\,[P(k)] ,

since the terms failing PP contribute nothing. The index may then be changed freely, without keeping track of the limits. The reciprocals of the primes up to NN become

∑p[ p prime ] [ p⩽N ]p,\sum_{p} \frac{[\,p \text{ prime}\,]\,[\,p \leqslant N\,]}{p} ,

and the term at p=0p = 0 is 00 rather than a division by zero, by the convention in the definition.

Proposition 2.31 (Laws of summation).

Let KK be a finite set of integers, let {ak}\{a_k\} and {bk}\{b_k\} be sequences of reals, and let c∈Rc \in \mathbb{R}. Then

  1. ∑k∈Kc ak=c∑k∈Kak\displaystyle \sum_{k \in K} c\,a_k = c \sum_{k \in K} a_k;
  2. ∑k∈K(ak+bk)=∑k∈Kak+∑k∈Kbk\displaystyle \sum_{k \in K} (a_k + b_k) = \sum_{k \in K} a_k + \sum_{k \in K} b_k;
  3. ∑k∈Kak=∑k∈Kaσ(k)\displaystyle \sum_{k \in K} a_k = \sum_{k \in K} a_{\sigma(k)} for every bijection σ:K→K\sigma : K \to K.

Discussion.

Each law is a property of addition of reals lifted to a finite list of terms, so each is proved by induction on the size of KK, removing one element at a time. The first is distributivity, the second is associativity and commutativity used together to interleave two lists, and the third is commutativity alone: reordering the terms of a finite sum does not change it, and a bijection of the index set with itself is exactly a reordering. The third is the one that needs the index set to be finite, since rearranging infinitely many terms can change a sum.

Proof.

We induct on #K\#K. If K=∅K = \emptyset all three sums are empty and every claim reads 0=00 = 0.

Let KK be non-empty, pick j∈Kj \in K and write K′=K∖{j}K' = K \setminus \{j\}, so that ∑k∈Kxk=xj+∑k∈K′xk\sum_{k \in K}x_k = x_j + \sum_{k \in K'}x_k for any sequence xx.

For the first law, ∑k∈Kcak=caj+∑k∈K′cak=caj+c∑k∈K′ak=c(aj+∑k∈K′ak)\sum_{k\in K}ca_k = ca_j + \sum_{k \in K'}ca_k = ca_j + c\sum_{k\in K'}a_k = c\bigl(a_j + \sum_{k\in K'}a_k\bigr), using the inductive hypothesis and then distributivity in R\mathbb{R}.

For the second, ∑k∈K(ak+bk)=(aj+bj)+∑k∈K′ak+∑k∈K′bk\sum_{k\in K}(a_k+b_k) = (a_j+b_j) + \sum_{k\in K'}a_k + \sum_{k\in K'}b_k by the hypothesis, and regrouping the four terms by commutativity and associativity gives (aj+∑k∈K′ak)+(bj+∑k∈K′bk)\bigl(a_j + \sum_{k\in K'}a_k\bigr) + \bigl(b_j + \sum_{k\in K'}b_k\bigr).

For the third, let σ:K→K\sigma : K \to K be a bijection and let j∈Kj \in K be arbitrary. Then σ\sigma restricts to a bijection K∖{σ−1(j)}→K′K \setminus \{\sigma^{-1}(j)\} \to K', and splitting both sums so that the term aja_j is taken first reduces the claim to the same statement on a set with one fewer element.

The laws above all concern a single index set. Two different index sets combine as well.

Proposition 2.32 (Combining index sets).

Let KK and K′K' be finite sets of integers. Then

∑k∈Kak+∑k∈K′ak=∑k∈K∩K′ak+∑k∈K∪K′ak.\sum_{k \in K} a_k + \sum_{k \in K'} a_k = \sum_{k \in K \cap K'} a_k + \sum_{k \in K \cup K'} a_k .

Discussion.

Counting elements suggests the shape of the answer, since #K+#K′=#(K∪K′)+#(K∩K′)\#K + \#K' = \#(K \cup K') + \#(K \cap K'): an element of both sets is counted twice on each side. Iverson brackets turn membership into numbers that can be added, and the proof adds them. The identity to establish is then [k∈K]+[k∈K′]=[k∈K∩K′]+[k∈K∪K′][k \in K] + [k \in K'] = [k \in K \cap K'] + [k \in K \cup K'] for each single kk, which is checked by looking at the four possible cases; summing it over all kk and using the law for sums of sequences gives the claim.

Proof.

Fix an integer kk and compare the two sides of

[k∈K]+[k∈K′]=[k∈K∩K′]+[k∈K∪K′].[k \in K] + [k \in K'] = [k \in K \cap K'] + [k \in K \cup K'] .

If kk lies in both sets, both sides are 22. If it lies in exactly one, both sides are 11, the right-hand side because the intersection bracket is 00 and the union bracket is 11. If it lies in neither, both sides are 00.

Multiplying by aka_k and summing over all integers kk, the second law of summation gives

∑kak[k∈K]+∑kak[k∈K′]=∑kak[k∈K∩K′]+∑kak[k∈K∪K′],\sum_k a_k[k \in K] + \sum_k a_k[k \in K'] = \sum_k a_k[k \in K \cap K'] + \sum_k a_k[k \in K \cup K'],

and each of the four sums is the corresponding sum over its index set.

Problem 2.18.

Show that [P] [Q]=[ P and Q ][P]\,[Q] = [\,P \text{ and } Q\,] and [P]+[Q]−[P] [Q]=[ P or Q ][P] + [Q] - [P]\,[Q] = [\,P \text{ or } Q\,] for all statements PP and QQ, and use the first of these to write

∑1⩽k⩽nk oddak\sum_{\substack{1 \leqslant k \leqslant n \\ k \text{ odd}}} a_k

as a sum over all integers kk with no condition beneath the sign.

Geometric Sums

Definition 2.33 (Geometric sequence).

A sequence {an}\{a_n\} of reals is geometric with ratio qq if an+1=q ana_{n+1} = q\,a_n for every n∈N0n \in \mathbb{N}_0, equivalently if an=a0qna_n = a_0 q^n for every nn.

Proposition 2.34 (Geometric sum).

Let q≠1q \neq 1 and n∈N0n \in \mathbb{N}_0. Then

∑k=0na0qk=a0 1−q n+11−q.\sum_{k=0}^{n} a_0 q^k = a_0\,\frac{1 - q^{\,n+1}}{1 - q} .

Discussion.

Multiplying the sum by qq shifts every term one place along, so the product and the original share all their terms but two: the original has a0a_0 where the product has nothing, and the product has a0qn+1a_0q^{n+1} where the original has nothing. Subtracting therefore cancels everything in the middle and leaves those two terms, which is a linear equation for the sum. The hypothesis q≠1q \neq 1 is needed to divide at the end: at q=1q = 1 every term is a0a_0 and the sum is (n+1)a0(n+1)a_0.

Proof.

Write S=∑k=0na0qk=a0+a0q+⋯+a0qnS = \sum_{k=0}^{n}a_0q^k = a_0 + a_0q + \cdots + a_0q^n. By the first law of summation,

qS=a0q+a0q2+⋯+a0qn+a0q n+1.qS = a_0q + a_0q^2 + \cdots + a_0q^{n} + a_0q^{\,n+1} .

Every term of qSqS with exponent between 11 and nn appears in SS as well, so subtracting leaves only the extremes:

S−qS=a0−a0q n+1,that isS(1−q)=a0(1−q n+1).S - qS = a_0 - a_0 q^{\,n+1}, \qquad\text{that is}\qquad S(1 - q) = a_0\bigl(1 - q^{\,n+1}\bigr).

Since q≠1q \neq 1 we may divide by 1−q1 - q.

Example 2.35 (Halving).

With a0=1a_0 = 1 and q=12q = \tfrac12,

∑k=0n(12)k=1−(1/2)n+11−1/2=2−(12)n.\sum_{k=0}^{n} \left(\tfrac{1}{2}\right)^{k} = \frac{1 - (1/2)^{n+1}}{1 - 1/2} = 2 - \left(\tfrac{1}{2}\right)^{n}.

Starting the sum at k=1k = 1 removes the term 11, so

∑k=1n(12)k=1−(12)n,\sum_{k=1}^{n} \left(\tfrac{1}{2}\right)^{k} = 1 - \left(\tfrac{1}{2}\right)^{n},

that is, repeatedly halving what remains of 11 leaves a gap of 2−n2^{-n}.

Sums and Recurrences

Writing Sn=∑k=0nakS_n = \sum_{k=0}^{n} a_k and splitting off the last term gives Sn=Sn−1+anS_n = S_{n-1} + a_n, so a sum is a recurrence and the methods of the previous chapter apply to it. The correspondence runs both ways.

Proposition 2.36 (A sum is a recurrence).

Let {ak}\{a_k\} be a sequence of reals and put Sn=∑k=0nakS_n = \sum_{k=0}^{n}a_k. Then SS is the unique function N0→R\mathbb{N}_0 \to \mathbb{R} with

S0=a0,Sn=Sn−1+an(n⩾1),S_0 = a_0, \qquad S_n = S_{n-1} + a_n \quad (n \geqslant 1),

and conversely any function satisfying those two conditions is given by that sum.

Discussion.

Both halves come from the splitting rule for sums, and the convention that an empty sum is 00 makes S0S_0 come out as a0a_0 with no separate argument. For the forward direction, split the last term off the sum defining SnS_n. For the converse, the two conditions determine the value at every nn from the value below it, so at most one function satisfies them, and the sum has just been shown to be one.

Proof.

At n=0n = 0 the sum has the single term a0a_0, so S0=a0S_0 = a_0. For n⩾1n \geqslant 1, splitting off the term at k=nk = n gives

Sn=∑k=0nak=∑k=0n−1ak+an=Sn−1+an.S_n = \sum_{k=0}^{n} a_k = \sum_{k=0}^{n-1}a_k + a_n = S_{n-1} + a_n .

Conversely, suppose R:N0→RR : \mathbb{N}_0 \to \mathbb{R} satisfies R0=a0R_0 = a_0 and Rn=Rn−1+anR_n = R_{n-1} + a_n for n⩾1n \geqslant 1. Then R0=S0R_0 = S_0, and if Rn−1=Sn−1R_{n-1} = S_{n-1} then Rn=Rn−1+an=Sn−1+an=SnR_n = R_{n-1} + a_n = S_{n-1} + a_n = S_n. By induction R=SR = S.

Used from left to right, this turns a sum into a recurrence to be solved by the techniques already available. The repertoire method of the previous chapter applies without change, and the recurrence to use is the general first-order one with a linear term.

Theorem 2.37 (A linear sum-recurrence).

Let α,β,γ∈R\alpha, \beta, \gamma \in \mathbb{R} and let RR satisfy

R0=α,Rn=Rn−1+β+γn(n⩾1).R_0 = \alpha, \qquad R_n = R_{n-1} + \beta + \gamma n \quad (n \geqslant 1).

Then

Rn=α+βn+γ n2+n2.R_n = \alpha + \beta n + \gamma\,\frac{n^2 + n}{2} .

Discussion.

Computing R1=α+β+γR_1 = \alpha + \beta + \gamma, R2=α+2β+3γR_2 = \alpha + 2\beta + 3\gamma and R3=α+3β+6γR_3 = \alpha + 3\beta + 6\gamma shows every value to be a combination A(n)α+B(n)β+C(n)γA(n)\alpha + B(n)\beta + C(n)\gamma whose coefficients do not depend on the three constants, so the repertoire method applies: choose functions RR that satisfy the recurrence for some values of α,β,γ\alpha, \beta, \gamma, and read off one equation in AA, BB, CC from each. The constant function 11 forces β=γ=0\beta = \gamma = 0 and gives AA; the function nn forces β=1\beta = 1, γ=0\gamma = 0 and gives BB; and the function n2n^2 forces β=−1\beta = -1, γ=2\gamma = 2, which involves BB and CC together and so needs the previous two results to finish. Three substitutions give three equations, and the system is triangular.

Proof.

Write Rn=A(n)α+B(n)β+C(n)γR_n = A(n)\alpha + B(n)\beta + C(n)\gamma, which the first few values show to be the right shape, and determine the coefficients by substitution. Each substitution takes a function RR, asks which α,β,γ\alpha, \beta, \gamma make it satisfy the recurrence, and then reads the displayed identity at those values.

Take Rn=1R_n = 1. Then R0=1R_0 = 1 forces α=1\alpha = 1, and the recurrence reads 1=1+β+γn1 = 1 + \beta + \gamma n, that is β+γn=0\beta + \gamma n = 0 for every nn, which forces β=γ=0\beta = \gamma = 0. Substituting (α,β,γ)=(1,0,0)(\alpha, \beta, \gamma) = (1, 0, 0) gives

1=A(n).1 = A(n) .

Take Rn=nR_n = n. Then R0=0R_0 = 0 forces α=0\alpha = 0, and the recurrence reads n=(n−1)+β+γnn = (n-1) + \beta + \gamma n, that is 1=β+γn1 = \beta + \gamma n for every nn, which forces β=1\beta = 1 and γ=0\gamma = 0. Substituting (0,1,0)(0, 1, 0) gives

n=B(n).n = B(n) .

Take Rn=n2R_n = n^2. Then α=0\alpha = 0, and the recurrence reads n2=(n−1)2+β+γnn^2 = (n-1)^2 + \beta + \gamma n, that is

0=−2n+1+β+γn=(1+β)+(γ−2)n0 = -2n + 1 + \beta + \gamma n = (1 + \beta) + (\gamma - 2)n

for every nn, which forces β=−1\beta = -1 and γ=2\gamma = 2. Substituting (0,−1,2)(0, -1, 2) gives

n2=−B(n)+2C(n)=−n+2C(n),soC(n)=n2+n2.n^2 = -B(n) + 2C(n) = -n + 2C(n), \qquad\text{so}\qquad C(n) = \frac{n^2+n}{2}.

Assembling, Rn=α+βn+γ(n2+n)/2R_n = \alpha + \beta n + \gamma(n^2+n)/2. Finally this function does satisfy the recurrence, as substituting it into Rn−1+β+γnR_{n-1} + \beta + \gamma n confirms, and the recurrence with its boundary condition has only one solution.

Example 2.38 (Summing an arithmetic progression).

Let a,b∈Ra, b \in \mathbb{R} and Sn=∑k=0n(a+bk)S_n = \sum_{k=0}^{n}(a + bk). By the correspondence above, S0=aS_0 = a and Sn=Sn−1+(a+bn)S_n = S_{n-1} + (a + bn), which is the recurrence of the theorem with α=a\alpha = a, β=a\beta = a and γ=b\gamma = b. Hence

∑k=0n(a+bk)=a+na+n2+n2b=(n+1)a+n(n+1)2b.\sum_{k=0}^{n}(a + bk) = a + na + \frac{n^2+n}{2}b = (n+1)a + \frac{n(n+1)}{2}b .

The same answer comes out of the summation laws directly: splitting the summand and pulling out the constants,

∑k=0n(a+bk)=∑k=0na+b∑k=0nk=(n+1)a+b n(n+1)2,\sum_{k=0}^{n}(a + bk) = \sum_{k=0}^{n}a + b\sum_{k=0}^{n}k = (n+1)a + b\,\frac{n(n+1)}{2},

since a sum of n+1n+1 copies of aa is (n+1)a(n+1)a. The repertoire method is needed for recurrences that are not sums of anything recognisable.

The Summation Factor

Used from right to left, the correspondence turns a recurrence into a sum. A recurrence of the form Sn=Sn−1+cnS_n = S_{n-1} + c_n is already a sum, so the problem is to bring a given recurrence into that shape, which is done by multiplying through by a suitable factor.

Theorem 2.39 (Summation factor).

Let {an}\{a_n\}, {bn}\{b_n\}, {cn}\{c_n\} be sequences of reals with an≠0a_n \neq 0 and bn≠0b_n \neq 0 for every n⩾1n \geqslant 1, and let TT satisfy

anTn=bnTn−1+cn(n⩾1),a_n T_n = b_n T_{n-1} + c_n \qquad (n \geqslant 1),

with T0T_0 given. Put s1=1s_1 = 1 and

sn=a1a2⋯an−1b2b3⋯bn(n⩾2).s_n = \frac{a_1 a_2 \cdots a_{n-1}}{b_2 b_3 \cdots b_n} \qquad (n \geqslant 2).

Then snbn=sn−1an−1s_n b_n = s_{n-1}a_{n-1} for every n⩾2n \geqslant 2, and

Tn=1snan(s1b1T0+∑k=1nskck)(n⩾1).T_n = \frac{1}{s_n a_n}\left( s_1 b_1 T_0 + \sum_{k=1}^{n} s_k c_k \right) \qquad (n \geqslant 1).

Discussion.

The obstruction to reading the recurrence as a sum is that TnT_n and Tn−1T_{n-1} carry different coefficients. Multiplying the whole equation by a factor sns_n removes the obstruction provided the new coefficient of Tn−1T_{n-1}, namely snbns_nb_n, equals the coefficient that Tn−1T_{n-1} had at the previous step, namely sn−1an−1s_{n-1}a_{n-1}; that condition is a recurrence for sns_n itself, and unrolling it gives the stated product. With this factor, Sn=snanTnS_n = s_na_nT_n satisfies Sn=Sn−1+sncnS_n = S_{n-1} + s_nc_n, which the previous proposition evaluates as a sum. Dividing by snans_na_n at the end is legitimate because both factors are non-zero, which the hypotheses on ana_n and bnb_n guarantee.

Proof.

For n⩾2n \geqslant 2, the definition of sns_n gives

snbn=a1⋯an−1b2⋯bn bn=a1⋯an−1b2⋯bn−1=a1⋯an−2b2⋯bn−1 an−1=sn−1an−1,s_n b_n = \frac{a_1 \cdots a_{n-1}}{b_2 \cdots b_n}\,b_n = \frac{a_1 \cdots a_{n-1}}{b_2 \cdots b_{n-1}} = \frac{a_1 \cdots a_{n-2}}{b_2\cdots b_{n-1}}\,a_{n-1} = s_{n-1}a_{n-1},

with the case n=2n = 2 reading s2b2=a1=s1a1s_2b_2 = a_1 = s_1a_1.

Multiply the recurrence by sns_n and set Sn=snanTnS_n = s_na_nT_n. For n⩾2n \geqslant 2,

Sn=snanTn=snbnTn−1+sncn=sn−1an−1Tn−1+sncn=Sn−1+sncn,S_n = s_na_nT_n = s_nb_nT_{n-1} + s_nc_n = s_{n-1}a_{n-1}T_{n-1} + s_nc_n = S_{n-1} + s_nc_n,

and at n=1n = 1 the recurrence gives S1=s1a1T1=s1(b1T0+c1)=s1b1T0+s1c1S_1 = s_1a_1T_1 = s_1(b_1T_0 + c_1) = s_1b_1T_0 + s_1c_1. So SS satisfies a sum-recurrence started at s1b1T0s_1b_1T_0, and unrolling it gives

Sn=s1b1T0+∑k=1nskck.S_n = s_1b_1T_0 + \sum_{k=1}^{n}s_kc_k .

Since sn≠0s_n \neq 0 and an≠0a_n \neq 0, dividing Sn=snanTnS_n = s_na_nT_n by snans_na_n gives the stated formula.

Example 2.40 (The Tower of Hanoi again).

The recurrence Tn=2Tn−1+1T_n = 2T_{n-1} + 1 has an=1a_n = 1, bn=2b_n = 2 and cn=1c_n = 1, so a summation factor is sn=2−ns_n = 2^{-n}, which satisfies snbn=2−(n−1)=sn−1an−1s_nb_n = 2^{-(n-1)} = s_{n-1}a_{n-1}. Multiplying through,

Tn2n=Tn−12 n−1+12n,\frac{T_n}{2^n} = \frac{T_{n-1}}{2^{\,n-1}} + \frac{1}{2^n},

so Sn=Tn/2nS_n = T_n/2^n obeys S0=0S_0 = 0 and Sn=Sn−1+2−nS_n = S_{n-1} + 2^{-n}. That is a geometric sum,

Sn=∑k=1n12k=1−12n,S_n = \sum_{k=1}^{n}\frac{1}{2^k} = 1 - \frac{1}{2^n},

and multiplying back by 2n2^n gives Tn=2n−1T_n = 2^n - 1.

Remark (Where the factor comes from).

The condition snbn=sn−1an−1s_nb_n = s_{n-1}a_{n-1} is itself a recurrence, sn=sn−1an−1/bns_n = s_{n-1}a_{n-1}/b_n, and unrolling it from s1=1s_1 = 1 produces the product in the theorem. Any non-zero constant multiple of that product serves equally well, since multiplying every sns_n by the same constant leaves the condition and the final formula unchanged; in the Hanoi example the product gives 2−(n−1)2^{-(n-1)} and we used 2−n2^{-n}. The method needs every ana_n and every bnb_n to be non-zero, and fails otherwise.

Definition 2.41 (Harmonic numbers).

For n∈N0n \in \mathbb{N}_0 the nnth harmonic number is

Hn=∑k=1n1k=1+12+13+⋯+1n,H_n = \sum_{k=1}^{n}\frac{1}{k} = 1 + \frac{1}{2} + \frac{1}{3} + \cdots + \frac{1}{n},

so that H0=0H_0 = 0. The name comes from music: the kkth harmonic of a vibrating string is the tone produced by a string 1/k1/k times as long.

Harmonic numbers have no closed form in the sense of this chapter, and they appear when a summation factor is used on a recurrence whose coefficients vary with nn.

Example 2.42 (A recurrence with variable coefficients).

The Hanoi recurrence had constant ana_n and bnb_n, so the factor was a constant power. Consider instead

T0=0,n Tn=(n+1) Tn−1+2n(n⩾1),T_0 = 0, \qquad n\,T_n = (n+1)\,T_{n-1} + 2n \qquad (n \geqslant 1),

with an=na_n = n, bn=n+1b_n = n+1 and cn=2nc_n = 2n. The summation factor is

sn=a1a2⋯an−1b2b3⋯bn=(n−1)!(n+1)!/2=2n(n+1),s_n = \frac{a_1 a_2 \cdots a_{n-1}}{b_2 b_3 \cdots b_n} = \frac{(n-1)!}{(n+1)!/2} = \frac{2}{n(n+1)} ,

which agrees with s1=1s_1 = 1. Then snanTn=2Tn/(n+1)s_n a_n T_n = 2T_n/(n+1) and skck=4/(k+1)s_k c_k = 4/(k+1), while s1b1T0=0s_1 b_1 T_0 = 0, so the theorem gives

2Tnn+1=∑k=1n4k+1,that isTn=2(n+1)∑k=1n1k+1.\frac{2T_n}{n+1} = \sum_{k=1}^{n}\frac{4}{k+1}, \qquad\text{that is}\qquad T_n = 2(n+1)\sum_{k=1}^{n}\frac{1}{k+1} .

The remaining sum is a harmonic number with its index shifted. Shifting the index by one,

∑k=1n1k+1=∑2⩽k⩽n+11k=∑1⩽k⩽n1k−1+1n+1=Hn−nn+1,\sum_{k=1}^{n}\frac{1}{k+1} = \sum_{2 \leqslant k \leqslant n+1}\frac{1}{k} = \sum_{1 \leqslant k \leqslant n}\frac{1}{k} - 1 + \frac{1}{n+1} = H_n - \frac{n}{n+1},

where the middle step dropped the term at k=1k = 1 and added the one at k=n+1k = n+1. Therefore

Tn=2(n+1)(Hn−nn+1)=2(n+1)Hn−2n,T_n = 2(n+1)\left(H_n - \frac{n}{n+1}\right) = 2(n+1)H_n - 2n ,

and checking at n=1n = 1 gives T1=4H1−2=2T_1 = 4H_1 - 2 = 2, which the recurrence confirms.

Problem 2.19.

Use a summation factor to solve T0=1T_0 = 1 and nTn=(n+1)Tn−1+n(n+1)nT_n = (n+1)T_{n-1} + n(n+1) for n⩾1n \geqslant 1.

Problem 2.20.

Prove that H2n−Hn⩾12H_{2n} - H_n \geqslant \tfrac{1}{2} for every n∈Nn \in \mathbb{N}, and deduce that HnH_n can be made as large as we please by taking nn large enough.

Exercises on Recurrences

Exercise 2.1.

Compute TnT_n for the Tower of Hanoi for n=1,…,8n = 1, \ldots, 8 from the recurrence, and check each against 2n−12^n - 1.

Exercise 2.2.

Suppose a fourth peg is added to the Tower of Hanoi. Give a strategy for nn discs using the extra peg, count its moves, and compare the count with 2n−12^n - 1 for n=4,8,16n = 4, 8, 16.

Exercise 2.3.

Solve each recurrence by unrolling, then confirm the answer by induction.

  1. f(0)=2f(0) = 2 and f(n)=3f(n−1)f(n) = 3f(n-1) for n⩾1n \geqslant 1;
  2. f(0)=0f(0) = 0 and f(n)=f(n−1)+2n−1f(n) = f(n-1) + 2n - 1 for n⩾1n \geqslant 1;
  3. f(1)=1f(1) = 1 and f(n)=f(n−1)+1/nf(n) = f(n-1) + 1/n for n⩾2n \geqslant 2, as far as a sum in closed form allows.

Exercise 2.4.

Solve f(0)=1f(0) = 1, f(n)=3f(n−1)+4f(n) = 3f(n-1) + 4 for n⩾1n \geqslant 1 by the substitution Un=f(n)+cU_n = f(n) + c, choosing cc so that the recurrence for UnU_n has no constant term.

Exercise 2.5.

What is the largest number of regions into which nn circles can cut the plane? Find the recurrence and solve it.

Exercise 2.6.

What is the largest number of pieces into which nn planes can cut three-dimensional space? Argue that the answer satisfies Pn=Pn−1+Ln−1P_n = P_{n-1} + L_{n-1}, where LnL_n is the count for lines in the plane, and solve.

Exercise 2.7.

For each recurrence, write down the characteristic equation, find its roots, and give the solution meeting the stated boundary conditions.

  1. f(0)=0f(0) = 0, f(1)=1f(1) = 1, f(n)=6f(n−1)−8f(n−2)f(n) = 6f(n-1) - 8f(n-2);
  2. f(0)=1f(0) = 1, f(1)=2f(1) = 2, f(n)=2f(n−1)−f(n−2)f(n) = 2f(n-1) - f(n-2);
  3. f(0)=1f(0) = 1, f(1)=0f(1) = 0, f(2)=1f(2) = 1, f(n)=3f(n−2)−2f(n−3)f(n) = 3f(n-2) - 2f(n-3).

Exercise 2.8.

Solve f(0)=0f(0) = 0 and f(n)=2f(n−1)+n2f(n) = 2f(n-1) + n^2 for n⩾1n \geqslant 1.

Exercise 2.9.

Let ff satisfy the Fibonacci recurrence with f(0)=f(1)=1f(0) = f(1) = 1. Prove that f(0)+f(1)+⋯+f(n)=f(n+2)−1f(0) + f(1) + \cdots + f(n) = f(n+2) - 1.

Exercise 2.10.

Compute J(n)J(n) for n=17,…,32n = 17, \ldots, 32 and check that the values run through the odd numbers below 3232 and then restart at 11.

Exercise 2.11.

Suppose the circle is counted in the other direction, so that the first person eliminated is the one before the leader rather than after. Work out the survivor for n=1,…,8n = 1, \ldots, 8 and find a recurrence.

Exercise 2.12.

Use the repertoire method on f(1)=αf(1) = \alpha, f(2ℓ)=2f(ℓ)+βℓf(2\ell) = 2f(\ell) + \beta\ell, f(2ℓ+1)=2f(ℓ)+γℓf(2\ell+1) = 2f(\ell) + \gamma\ell, taking f(n)=1f(n) = 1 and f(n)=nf(n) = n as two of the substitutions.

Exercise 2.13.

Convert 20262026 to base 22, base 33 and base 1616, and check each answer by Horner’s scheme.

Exercise 2.14.

Let b>1b > 1 and z∈Nz \in \mathbb{N}. Show that zz and z+1z+1 have the same number of base-bb digits unless z+1z + 1 is a power of bb, and use this to count how many numbers have exactly nn digits in base bb.

Exercise 2.15.

Show that ⌊x⌋+⌈x⌉=2⌊x⌋\lfloor x \rfloor + \lceil x \rceil = 2\lfloor x \rfloor when x∈Zx \in \mathbb{Z} and 2⌊x⌋+12\lfloor x \rfloor + 1 otherwise, and that ⌊x+k⌋=⌊x⌋+k\lfloor x + k \rfloor = \lfloor x \rfloor + k for every k∈Zk \in \mathbb{Z}.

Exercises on Sums

Exercise 2.16.

Some of the regions cut out by nn lines in the plane are bounded and the rest are not. What is the largest possible number of bounded regions?

Exercise 2.17.

Let H(n)=J(n+1)−J(n)H(n) = J(n+1) - J(n), with JJ the Josephus survivor. The recurrence for JJ gives H(2n)=2H(2n) = 2 and

H(2n+1)=J(2n+2)−J(2n+1)=(2J(n+1)−1)−(2J(n)+1)=2H(n)−2H(2n+1) = J(2n+2) - J(2n+1) = \bigl(2J(n+1) - 1\bigr) - \bigl(2J(n) + 1\bigr) = 2H(n) - 2

for every n⩾1n \geqslant 1. It therefore looks possible to prove H(n)=2H(n) = 2 for all nn by induction. Compute H(1)H(1), H(2)H(2) and H(3)H(3), and say exactly what is wrong with the argument.

Exercise 2.18.

Let α,β∈R\alpha, \beta \in \mathbb{R} and define Q0=αQ_0 = \alpha, Q1=βQ_1 = \beta and

Qn=1+Qn−1Qn−2(n⩾2),Q_n = \frac{1 + Q_{n-1}}{Q_{n-2}} \qquad (n \geqslant 2),

assuming α\alpha and β\beta are such that no denominator ever vanishes. Compute Q2,…,Q6Q_2, \ldots, Q_6 and prove that Qn+5=QnQ_{n+5} = Q_n for every n∈N0n \in \mathbb{N}_0.

Exercise 2.19.

Evaluate

∑k [ 1⩽j⩽k⩽n ]\sum_{k} \,[\,1 \leqslant j \leqslant k \leqslant n\,]

as a function of jj and nn, taking care over the values of jj for which the sum is empty.

Exercise 2.20.

Prove the rule for summation by parts: for every n∈N0n \in \mathbb{N}_0,

∑0⩽k<n(ak+1−ak) bk=anbn−a0b0−∑0⩽k<nak+1 (bk+1−bk).\sum_{0 \leqslant k < n} (a_{k+1} - a_k)\,b_k = a_n b_n - a_0 b_0 - \sum_{0 \leqslant k < n} a_{k+1}\,(b_{k+1} - b_k) .

Use only the distributive, associative and commutative laws of summation together with a shift of the index.

Exercise 2.21.

Find a closed form for

∑k=0n(−1)kk2.\sum_{k=0}^{n} (-1)^k k^2 .

Exercise 2.22.

Use a summation factor to solve

T0=5,2Tn=n Tn−1+3⋅n!(n⩾1),T_0 = 5, \qquad 2T_n = n\,T_{n-1} + 3 \cdot n! \quad (n \geqslant 1),

and check your answer at n=1n = 1 and n=2n = 2.

Exercise 2.23.

Evaluate ∑k=0nk 2k\sum_{k=0}^{n} k\,2^k by writing the recurrence it satisfies and solving it, and check the answer for n⩽4n \leqslant 4.

Exercise 2.24.

Show that ∑k=1n1k(k+1)=1−1n+1\sum_{k=1}^{n} \dfrac{1}{k(k+1)} = 1 - \dfrac{1}{n+1}, first by unrolling the corresponding recurrence and then by writing 1k(k+1)=1k−1k+1\dfrac{1}{k(k+1)} = \dfrac{1}{k} - \dfrac{1}{k+1} and cancelling.

Exercise 2.25.

Let KK and K′K' be finite sets of integers with K∩K′=∅K \cap K' = \emptyset. Deduce from the law for combining index sets that ∑k∈K∪K′ak=∑k∈Kak+∑k∈K′ak\sum_{k \in K \cup K'}a_k = \sum_{k \in K}a_k + \sum_{k \in K'}a_k, and give an example showing the hypothesis is needed.

Exercise 2.26.

Write each of the following as a sum with the index running from 00, using Iverson brackets where a condition is needed.

  1. ∑1⩽k⩽100k divisible by 3k\displaystyle\sum_{\substack{1 \leqslant k \leqslant 100 \\ k \text{ divisible by } 3}} k;
  2. ∑k=5201k−4\displaystyle\sum_{k=5}^{20} \frac{1}{k-4};
  3. the sum of aka_k over those kk between 11 and nn that are perfect squares.

Check Yourself

 

Fresh questions on the whole chapter — none of them is worked out above. Do each on paper first; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 2.27.

How many moves does the Tower of Hanoi need for five discs?

answer one of these

Exercise 2.28.

What is the largest number of regions four lines can cut the plane into?

answer one of these

Exercise 2.29.

In how many ways can a staircase of five steps be climbed, one or two steps at a time?

answer one of these

Exercise 2.30.

What are the roots of the characteristic equation of f(n)=3f(n−1)−2f(n−2)f(n) = 3f(n-1) - 2f(n-2)?

answer one of these

Exercise 2.31.

A root r≠0r \neq 0 of multiplicity 22 of the characteristic equation contributes which solutions?

answer one of these

Exercise 2.32.

To find a particular solution of f(n)=2f(n−1)+3nf(n) = 2f(n-1) + 3^n, what should be tried first?

answer one of these

Exercise 2.33.

What is J(20)J(20)?

answer one of these

Exercise 2.34.

What is J(64)J(64)?

answer one of these

Exercise 2.35.

For which nn does the Josephus problem leave the last person standing, that is J(n)=nJ(n) = n?

answer one of these

Exercise 2.36.

What is (1011)2(1011)_2 in base ten?

answer one of these

Exercise 2.37.

What is 101 mod 7101 \bmod 7?

answer one of these

Exercise 2.38.

What is ⌊−3.2⌋\lfloor -3.2 \rfloor?

answer one of these

Exercise 2.39.

How many binary digits does 10001000 have?

answer one of these

Exercise 2.40.

Unrolling Ln=Ln−1+nL_n = L_{n-1} + n from L0=1L_0 = 1, what is L10L_{10}?

answer one of these

Exercise 2.41.

What is ∑k=15k2\sum_{k=1}^{5} k^2?

answer one of these

Exercise 2.42.

What is the value of ∑k=31ak\sum_{k=3}^{1} a_k?

answer one of these

Exercise 2.43.

What is [ 3 is even ]+[ 3 is odd ][\,3 \text{ is even}\,] + [\,3 \text{ is odd}\,]?

answer one of these

Exercise 2.44.

What is ∑k=043⋅2k\sum_{k=0}^{4} 3 \cdot 2^k?

answer one of these

Exercise 2.45.

What is the harmonic number H4H_4?

answer one of these

Exercise 2.46.

Which summation factor turns Tn=3Tn−1+1T_n = 3T_{n-1} + 1 into a sum?

answer one of these

Lesson 3

Introduction to Python

Taught

Imperative Knowledge and Computation

There are two kinds of knowledge a mathematical text can record. Declarative knowledge is a statement of fact: the square root of a positive real xx is the positive yy with y2=xy^2 = x. Imperative knowledge is a recipe: a sequence of instructions which, followed to the letter, produces it. A machine can carry out imperative knowledge, but not declarative knowledge.

At its lowest level a computer does two things: it performs arithmetic, and it stores the results. It does billions of operations a second and can store a great deal.

Definition 3.1 (Algorithm).

An algorithm is a finite sequence of unambiguous instructions which, given an initial state and a set of inputs, passes through well-defined successive states and after finitely many steps produces an output and stops.

Finite rules out an infinite list of instructions. Unambiguous rules out a step whose meaning depends on who is reading it. After finitely many steps rules out a recipe that never returns an answer; of the three conditions it is the hardest to check.

Writing an algorithm down needs a notation for assignment, which has no counterpart in an equation. We write

g=defeg \defeq e

for the instruction “replace the current value of gg by the value of the expression ee”. It is an instruction, not a statement about gg. The instruction g=defg+1g \defeq g + 1 makes sense, while the equation g=g+1g = g + 1 has no solutions.

Here is Heron of Alexandria’s method for the square root of a real x>0x > 0.

  1. Choose any guess g>0g > 0.
  2. If g2g^2 is close enough to xx, stop and return gg.
  3. Otherwise perform g=def12(g+x/g)g \defeq \tfrac{1}{2}\bigl(g + x/g\bigr).
  4. Go back to step 2.
startg ≝ 1|g² − x| < εreturn gstopg ≝ (g + x/g)/2yesno
Figure 3.1. Heron’s method as a flowchart. The test in the diamond decides which arrow is followed, and the arrow from the update box returns to the test, so the accented box may run any number of times.

The procedure has the two ingredients of every algorithm: a list of operations, and control flow, the rule deciding which operation comes next. Control flow depends on tests that answer yes or no, and it lets a fixed piece of text describe a computation whose length is not fixed.

Problem 3.1.

Let x=25x = 25 and g=1g = 1, and read “close enough” as ∣g2−x∣<0.01|g^2 - x| < 0.01. Compute the first three values of gg produced by Heron’s method. Does the method stop within those three steps?

Computability

Early machines were fixed-program computers, wired to solve one problem. The stored-program computer keeps the instructions and the data they act on in the same memory, so that a single interpreter can execute any legal instruction sequence handed to it. A program counter walks the interpreter through the instructions in order, deviating only where control flow says to jump. Since the output of a computation may itself be a sequence of instructions, a machine of this kind can write its own programs.

Remark (The Church–Turing thesis).

Turing’s 1936 model, the Universal Turing Machine, is the standard formalisation of what a stored-program machine can do. The Church–Turing thesis asserts that a function is computable by any effective procedure exactly when some Turing machine computes it. It is not a theorem: “effective procedure” is an informal notion, and the thesis says that the formal model captures it.

A language is Turing complete if it can simulate a Universal Turing Machine. Python is, and so is every language in ordinary use, which is why any algorithm expressible in one of them is expressible in all of them.

There is, however, no algorithm which, given an arbitrary program and its input, decides in finitely many steps whether that program eventually stops or runs forever. This is the halting problem. Because it is unsolvable, the clause “stops after finitely many steps” in the definition of an algorithm has to be proved for each algorithm separately; it cannot be checked mechanically.

Syntax and Semantics

A program is a piece of text, and like a mathematical formula it has a meaning only if it obeys three kinds of rule.

  1. Primitives. The atoms of the language: numeric literals such as 3.2, strings of text, and operators such as + and *.
  2. Syntax. The rules saying which arrangements of primitives are well formed. 3.2 + 3.2 is well formed; 3.2 3.2 is not. Violations are caught before a single instruction runs.
  3. Static semantics. The rules saying which well-formed arrangements have a meaning. Adding a number to a piece of text has the right shape, operand-operator-operand, and still means nothing.

A program that obeys all three has a semantics: exactly one meaning, fixed by the language and not by the reader. An English sentence, by contrast, can have several.

So when a program misbehaves the machine has not misunderstood it: the error is in what was written. It shows up in one of three ways: the program stops with an error, it runs forever, or it finishes and returns the wrong answer. The third gives no sign that anything is wrong, which is why a program needs a proof of correctness and not only a run that looked right.

Problem 3.2.

Classify each of the following as a syntax error, a static semantic error, or a program that runs to completion and returns the wrong answer. Justify each answer in a sentence.

  1. x = 5 + * 3
  2. x = "hello" + 7
  3. An implementation of Heron’s method that stops as soon as g2>xg^2 > x rather than when ∣g2−x∣|g^2 - x| is small.

Problems, Words and Languages

To compare algorithms for a problem, the problem has to be stated in a fixed form. A machine reads and writes finite strings of characters, so problems are stated in terms of strings.

Definition 3.2 (Alphabet, word and language).

Let AA be a non-empty finite set, called an alphabet. For k∈N0k \in \mathbb{N}_0 let AkA^k denote the set of functions {1,…,k}→A\{1, \ldots, k\} \to A. Such a function ff is written as the sequence

f(1) f(2)⋯f(k)f(1)\,f(2)\cdots f(k)

and is called a word, or string, of length kk over AA. The set A0A^0 has a single element, the empty word, of length 00. Writing

A∗=⋃k∈N0AkA^{*} = \bigcup_{k \in \mathbb{N}_0} A^{k}

for the set of all words over AA, a language over AA is a subset of A∗A^{*}.

A set SS is finite if there is an injection S→{1,…,n}S \to \{1, \ldots, n\} for some n∈Nn \in \mathbb{N}, and infinite otherwise; the number of elements of a finite set is written ∣S∣|S|. A set admitting an injection into N\mathbb{N} is countable, and the infinite ones among them are the countably infinite sets of the last lesson.

Definition 3.3 (Computational problem).

A computational problem is a relation P⊆D×EP \subseteq D \times E such that every d∈Dd \in D has at least one e∈Ee \in E with (d,e)∈P(d, e) \in P. The elements of DD are the instances of PP, and ee is a correct output for the instance dd whenever (d,e)∈P(d, e) \in P.

The problem is unique if PP is a function, so that every instance has exactly one correct output. It is discrete if DD and EE are languages over a finite alphabet, and numerical if D⊆RmD \subseteq \mathbb{R}^m and E⊆RnE \subseteq \mathbb{R}^n for some m,n∈Nm, n \in \mathbb{N}. A unique problem P:D→EP : D \to E with ∣E∣=2|E| = 2 is a decision problem.

A decision problem is one whose answer is yes or no, and we state such a problem by naming its instances and asking the question.

Example 3.4 (Primality as a decision problem).

Taking N\mathbb{N} as a language over the alphabet {0,1,…,9}\{0, 1, \ldots, 9\}, the relation

{ (n,e)∈N×{0,1}  ∣  e=1 exactly when n is prime }\bigl\{\, (n, e) \in \mathbb{N} \times \{0, 1\} \;\bigm|\; e = 1 \text{ exactly when } n \text{ is prime} \,\bigr\}

is a decision problem, which we write as

Primality
Input:    n ∈ N.
Question: is n prime?

A machine works directly only on discrete problems. A numerical problem must have its instances and its answers written as finite strings, and no finite string names an arbitrary real; rounding error is the difference between a real number and the string used for it.

Pseudocode

An algorithm does not depend on a programming language, and is usually written in pseudocode: the control flow and the assignments written out, without the details a particular language would require. An algorithm in pseudocode names its input and its output, and uses =def\defeq for assignment and output for the result.

Example 3.5 (The square of an integer).

Squaring a natural number takes a single instruction, and the algorithm written out in full reads

The Square of an Integer
Input:  x ∈ N.
Output: the square of x.

    result ≝ x · x
    output result

The variables here are the input x and the working variable result; the assignment puts the value of x⋅xx \cdot x into result, and output yields what result holds.

We use pseudocode to describe an algorithm and Python to run it.

Objects and Types

A Python program is a sequence of definitions and commands, evaluated in order by an interpreter. A command is called a statement, and it instructs the interpreter to do something. A call to the built-in print writes to the screen: it takes any number of arguments, separates them by spaces, and follows them by a newline unless told otherwise.

print("Algorithm")
print("terminated.")

# several arguments, and a different ending
print("Algorithm", "terminated", end=".\n")

Writing an f before the opening quote of a string makes it an f-string, in which any expression placed in braces is evaluated and its value substituted:

n = 7
print(f"{n} squared is {n * n}")   # 7 squared is 49

In mathematics the legality of an operation depends on what it is applied to: addition is defined for numbers, composition for maps. In Python each value is an object, and every object has a type, which fixes the operations allowed on it. The interpreter uses types to enforce static semantics.

Types divide into scalar, meaning indivisible, and non-scalar, meaning having internal structure. Python has four primitive scalar types.

  1. int, the integers, written as usual: 5, -12.
  2. float, an approximation to the reals, written with a decimal point (3.0, -28.72) or in scientific notation (1.6e-19). Memory is finite and the reals are not, so a float holds one of finitely many values and float arithmetic is not the arithmetic of R\mathbb{R}: the expression 0.1 + 0.2 == 0.3 evaluates to False.
  3. bool, inhabited by exactly two values, True and False.
  4. None, inhabited by one value, used for the absence of a result.

Expressions and Relational Operators

Objects combined with operators according to the syntax form an expression, and every expression evaluates to a single object of some type. The built-in type reports which type an object has:

type(5)     # int
type(5.0)   # float

Control flow needs tests that come out True or False, and these are built with the relational operators <, <=, >, >=, == and !=. Every one of them evaluates to a bool.

Python’s notation differs from the pseudocode here. Mathematical assignment g=defeg \defeq e is written in Python with a single equals sign, g = e. Testing whether two expressions have the same value is ==, and testing that they do not is !=.

3.0 + 2.0   # the float 5.0
3 != 2      # the bool True

Problem 3.3.

Give the value of type(4 == 4) and of type(4.0). Then say what happens if the test in step 2 of Heron’s method is written with = in place of ==, and which of the three kinds of rule of the previous section it breaks.

Arithmetic Operators

Addition, subtraction and multiplication are +, - and *, and exponentiation is **. Division splits into three operators, two of which are the floor and the remainder of the last lesson.

  1. / always evaluates to a float: 5 / 3 gives 1.6666666666666667.
  2. // is floor division, the largest integer not exceeding the quotient. So 5 // 3 is 1 and -4 // 3 is -2.
  3. % is the remainder: 5 % 3 is 2. A modern processor carries out % and // in a few clock cycles, so we count each of them as a single elementary operation.

Remark (Floors and remainders in Python).

For integers a and b > 0, a // b is ⌊a/b⌋\lfloor a/b \rfloor and a % b is a mod ba \bmod b as defined in Definition 2.19. The identity z=b⌊z/b⌋+(z mod b)z = b\lfloor z/b\rfloor + (z \bmod b) therefore reads

a == b * (a // b) + a % b

and holds for every integer a. It holds for negative a because // rounds down, not towards zero.

Problem 3.4.

Verify that a % b and a - (a // b) * b are equal, first for b=3b = 3 and a=7,6,0,−1,−4a = 7, 6, 0, -1, -4, and then in general.

Suppose integer division were defined by truncation towards zero instead, so that −4-4 divided by 33 gave −1-1. What must the remainder then be for the identity to survive, and what happens to the guarantee 0⩽a mod b<b0 \leqslant a \bmod b < b?

Compound expressions are read by precedence: ** binds most tightly, then *, /, // and %, then + and -.

print(2 + 3 * 4)   # 14, not 20
print(5 + 4 % 3)   # 6, not 0   (% binds as tightly as * and /)
print(2 ** 3 * 4)  # 32, not 4096

Operators of equal precedence are read by associativity. Most associate to the left; exponentiation associates to the right.

print(5 - 4 - 3)    # -2, not 4
print(4 ** 3 ** 2)  # 4 ** 9 = 262144, not 64 ** 2 = 4096
print((4 ** 3) ** 2)  # 4096

Logical Operators

The three logical connectives, negation, conjunction and disjunction, are implemented as the keywords not, and and or. They act on bool values and are fixed by these tables, in which T and F abbreviate True and False.

anot a
TF
FT
aba and ba or b
TTTT
TFFT
FTFT
FFFF

So a and b holds exactly when both do, a or b when at least one does, and not a reverses the value. All three are reserved words: they are part of the syntax and cannot be reassigned or used as names.

Both and and or stop as soon as the answer is settled: in a and b the expression b is never evaluated when a is False, and in a or b it is never evaluated when a is True.

(3 != 2) and (5 % 3 == 0)   # False, since 5 % 3 is 2

The Operators in One Place

CategoryOperators
Arithmetic+, -, *, /, //, **, %, unary -, unary +
Relational<, <=, >, >=, ==, !=
Assignment=, +=, -=, *=, /=, //=, **=, %=, <<=, >>=
Logicaland, or, not

The compound assignments abbreviate an update in place: total += x means total = total + x.

The bitwise operators (<<, >>, &, |, ^, ~, &=, |=, ^=), and with them <<= and >>=, are not used in this lesson.

Variables and State

An equation such as y=x2y = x^2 ties yy to xx permanently: change xx and yy changes with it. Assignment does not. It creates a binding, at one moment in time, between a name and one object in memory, and the binding stands until it is replaced.

pi = 3.14159
radius = 11.0
area = pi * (radius ** 2)
radius = 14.0

The third statement evaluates pi * (radius ** 2) to the single float 380.13239 and binds area to it. The fourth rebinds radius, and area is untouched: it is bound to a number, not to a formula. This lets Heron’s method overwrite its guess gg at every step while the input xx stays fixed.

Names and Comments

A program has to be checked by a reader, and a correct program with opaque names is hard to check. The # symbol begins a comment, which the interpreter ignores.

a = 3.14159
b = 11.2
c = a * (b ** 2)

pi = 3.14159
diameter = 11.2
area = pi * ((diameter / 2.0) ** 2)

The two blocks compute different numbers, and only the second makes it visible which one is the area of a circle: the first squares a diameter where it should square a radius, and its names hide the fact.

Multiple Assignment

Several names may be bound at once. Every expression on the right of the = is evaluated in full before any name on the left is rebound, so the two sides do not interfere.

x, y = 2, 3
x, y = y, x

print("x is", x)   # x is 3
print("y is", y)   # y is 2

Exchanging two values therefore needs no third name to hold one of them.

Problem 3.5.

Start from a, b = 1, 1 and perform a, b = b, a + b three times. Give the values of a and b after each of the three steps, and name the sequence they run through. Then say what a, b = b, a + b would produce if the right-hand side were evaluated one name at a time, left to right.

Numbers in Other Bases

The string 28122812 is a representation of a number rather than the number itself. Each digit occupies a position whose value is a power of ten:

2812=2⋅103+8⋅102+1⋅101+2⋅100.2812 = 2 \cdot 10^3 + 8 \cdot 10^2 + 1 \cdot 10^1 + 2 \cdot 10^0 .

The rightmost digit sits in the units place, 100=110^0 = 1, and each position to its left is worth ten times the one before, so the leftmost 22 of 28122812 stands for two thousand while the rightmost stands for two. This is a positional numeral system: what a digit means depends on where it sits.

Base ten comes from counting on fingers, not from mathematics. Ten is divisible by only four numbers, 1,2,51, 2, 5 and 1010, which limits how easily fractions can be handled; twelve, divisible by 1,2,3,4,61, 2, 3, 4, 6 and 1212, would serve better. A digital circuit holds two voltage levels, high and low, so machines work in base two, where arithmetic is carried out directly by sequences of logic gates.

That every natural number has exactly one representation of this shape in any base b>1b > 1 is Theorem 2.21 of the last lesson, and the notation is fixed by its corollary. Read as a statement about digit strings, it says that the base-bb expansion of n∈N0n \in \mathbb{N}_0 is the string drdr−1⋯d0d_r d_{r-1} \cdots d_0 satisfying three conditions:

  1. n=drbr+dr−1br−1+⋯+d0b0n = d_r b^r + d_{r-1}b^{r-1} + \cdots + d_0 b^0;
  2. every digit satisfies 0⩽di<b0 \leqslant d_i < b;
  3. if n>0n > 0 then the leading digit drd_r is not 00, and the expansion of 00 is the string 00 in every base.

The first condition says that the digits record how many copies of each power of bb are wanted; in decimal the places from the right are units, tens, hundreds, and in binary they are 1,2,4,8,16,…1, 2, 4, 8, 16, \ldots. The second restricts the available digits to 0,1,…,b−10, 1, \ldots, b-1, and without it uniqueness fails: if X\text{X} were a decimal digit standing for ten, then X2\text{X}2 and 102102 would both represent one hundred and two. The third bans leading zeros, without which 0142301423 and 14231423 would be different strings for one number.

Following the corollary we write

(drdr−1⋯d0)b=drbr+dr−1br−1+⋯+d0b0(d_r d_{r-1} \cdots d_0)_b = d_r b^r + d_{r-1} b^{r-1} + \cdots + d_0 b^0

for the number with this expansion, and a string carrying no subscript is decimal.

Remark (Names of the small bases).

The base-22, 33, 88, 1010 and 1616 expansions are called binary, ternary, octal, decimal and hexadecimal. Bases past ten need more than the ten digits, and the convention is to carry on with letters: A=10\text{A} = 10, B=11\text{B} = 11, and so on up to Z=35\text{Z} = 35.

Example 3.6 (One number in six bases).

The decimal expansion of 10231023 is 10231023 itself. Since 1023=210−11023 = 2^{10} - 1, every power of two from 202^0 to 292^9 occurs exactly once and the binary expansion is ten ones. For base thirty-six, 1023=28⋅36+151023 = 28 \cdot 36 + 15, and 2828 is S\text{S} while 1515 is F\text{F}. Altogether

1023=(1111111111)2=(1101220)3=(1777)8=(1023)10=(3FF)16=(SF)36.1023 = (1111111111)_2 = (1101220)_3 = (1777)_8 = (1023)_{10} = (3\text{FF})_{16} = (\text{SF})_{36} .

Python writes the three bases a machine uses with bin, oct and hex, each returning a string carrying a prefix that names the base:

n = 1023

bin(n)   # '0b1111111111'
oct(n)   # '0o1777'
hex(n)   # '0x3ff'

The same prefixes may be typed directly as literals, so the conversions can be checked against each other:

a = 0b1111111111
b = 0o1777
c = 0x3FF

a == b == c == 1023   # True

For any other base, int takes the base as a second argument and reads a string written in it:

int('1101220', 3)   # 1023
int('sf', 36)       # 1023

Problem 3.6.

Find the binary, ternary, octal, hexadecimal and base-3636 expansions of 17291729 by hand, using A\text{A} to F\text{F} as the extra hexadecimal digits and A\text{A} to Z\text{Z} for base thirty-six. Check each against int, and against bin, oct and hex.

Problem 3.7.

A number has octal expansion (2745)8(2745)_8. Give its decimal value and its hexadecimal expansion.

Problem 3.8.

Let b>1b > 1. Write down, in terms of bb, the smallest and the largest number whose base-bb expansion has exactly three digits.

Branching

Everything written so far is a straight-line program: the statements run in the order they appear, each exactly once. Such a program is easy to reason about, and limited. If one atomic operation takes one unit of time, a straight-line program of NN lines takes at most NN units however large its input, because no line ever runs twice.

The first way to let a program do more is branching: a test is evaluated, and the block of statements that runs next depends on the answer.

Conditional Statements

The basic branching construct is if–else. It consists of a test evaluating to True or False, a block run when the test is True, and an optional else block run when it is False.

Python marks the extent of a block by indentation. Where other languages use braces or an end keyword, Python uses the whitespace itself, so the layout of a program shows its structure.

x = 14

if x % 2 == 0:
    print("The integer is even.")
else:
    print("The integer is odd.")

print("Branching complete.")

The test uses x % 2, the remainder on division by two, which is the last digit of the binary expansion of x. It is 0 for even x and 1 for odd. The final print is not indented, so it belongs to neither block and runs either way; indenting it by four spaces would make it part of the else and suppress it for even x.

Because indentation carries meaning, a long expression cannot be broken across lines anywhere. A backslash at the end of a line continues it explicitly, and a line inside unclosed brackets continues implicitly.

# explicit continuation
alpha = 1.61803398875 + \
        2.71828182845 + \
        3.14159265359

# implicit continuation, inside parentheses
beta = (1.61803398875 +
        2.71828182845 +
        3.14159265359)

Nested Conditionals

A block inside a branch may itself branch, which builds a decision tree.

Where the cases are mutually exclusive, a chain of else blocks each containing one if grows an extra level of indentation per case. The keyword elif collapses the chain: its test is evaluated exactly when every test above it has come out False.

n = 15

if n % 2 == 0:
    if n % 3 == 0:
        print("n is a multiple of 6.")
    else:
        print("n is even, but not divisible by 3.")
elif n % 3 == 0:
    print("n is odd and divisible by 3.")
else:
    print("n is not divisible by 2 or 3.")

Compound Tests and Updating State

Tests may be combined with and, or and not. How they are combined changes how much work the program does.

Take the problem of returning the largest positive number among aa, bb and cc, or, if none of them is positive, the smallest of the three. Each of the three is positive or not, so there are 23=82^3 = 8 cases, and a program with one branch per case tests the same three comparisons over and over.

Binding a provisional answer and updating it only when a later value beats it replaces the eight cases by three independent tests:

a, b, c = -5, 12, 7

# if none is positive the answer is the smallest, so start there
target = min(a, b, c)

if a > 0:
    target = a
if b > 0 and b > target:
    target = b
if c > 0 and c > target:
    target = c

print(target)   # 12

The built-in min returns the smallest of its arguments, and max the largest. Each of the three variables has its sign examined once.

Conditional Expressions

For a choice between two values rather than between two blocks of work, Python supplies the conditional expression, which packs an if–else into a single expression:

expression_if_true if test else expression_if_false

A variable may therefore be bound according to a test without a multi-line block. The absolute value of Definition 1.65 is given by cases, so there are several ways to compute it. The most direct is a standard if block:

x = -10

if x < 0:
    x = -x

For a body consisting of a single short statement, Python permits writing the block on the same line as the if:

if x < 0: x = -x   # only for very short statements

Branching can also be avoided entirely by using the fact that True and False behave as 11 and 00 in arithmetic:

y = (x < 0) * (-x) + (x >= 0) * x   # works, but hard to read

This works, but it is hard to read, it evaluates both branches every time, and it hides the case split inside a multiplication. The conditional expression is clearer:

y = x if x >= 0 else -x

Python also provides the built-in abs for this purpose, so abs(-10) returns 10.

Problem 3.9.

Using conditional expressions and no if statements, write single-line definitions of

  1. the sign of xx, which is −1-1, 00 or 11 according as xx is negative, zero or positive;
  2. the larger of xx and yy, without using max;
  3. the distance ∣x−y∣|x - y| between two reals, without using abs.

Constant Time

Branching does not remove the limit on straight-line programs. Each block is entered at most once on a run, so if one atomic operation takes one unit of time, a branching program of NN lines can never exceed NN units, and its maximum running time is hard-bounded by a constant k⩽Nk \leqslant N.

Definition 3.7 (Constant time).

An algorithm runs in constant time if there is a constant kk, depending on the algorithm alone, such that the algorithm performs at most kk atomic operations on every input, whatever the size or magnitude of that input.

Every straight-line program runs in constant time, and so does every branching program, with kk the number of lines.

Remark (Beyond constant time).

Constant time is a strong restriction. Consider computing the factorial of an integer nn. It requires n−1n - 1 multiplications, and nn is not bounded, so no fixed number of multiplication statements serves for every nn; a program with one hard-coded branch per value of nn would need infinitely many branches and would violate the finiteness in the definition of an algorithm. The same holds for reading the base-bb digits of a number, whose count grows like log⁡bn\log_b n. Computations whose length grows with the input need control flow that can return to an earlier instruction.

Problem 3.10.

Let a, b and c be the coefficients of ax2+bx+cax^2 + bx + c. Write a program that computes the discriminant Δ=b2−4ac\Delta = b^2 - 4ac and reports whether the polynomial has two distinct real roots, one repeated real root, or none. Treat a=0a = 0 separately, where the expression is linear rather than quadratic, and say what your program should report when a=b=0a = b = 0.

Problem 3.11.

Consider the rule which replaces an integer nn by 3n+13n + 1 when nn is odd and by n/2n/2 when nn is even. Using conditional expressions, write a program that applies the rule three times in succession to a given positive integer. Check that starting from n=7n = 7 it produces 3434, and find a starting value below 1010 from which three applications return the starting value itself.

Iteration

Branching programs are bound by constant time: each instruction is executed at most once, so the whole computation is capped by the static length of the source. The base expansions of the first chapter show the limit. Every step applied the same pair of operations, // and %, to the current quotient, and with no way of repeating those operations automatically we performed each step by hand. Computing the factorial of nn, testing whether a number is prime, extracting the digits of an arbitrarily large base expansion: these tasks take longer on larger inputs, and no branching program can perform them.

Iteration sends execution back to an instruction already passed.

Definition 3.8 (Iteration).

An iteration, or loop, is a control structure that sends execution back to an instruction it has already passed, so that a block of statements runs repeatedly. How many times it runs is decided by the state of the computation while it runs, not by the length of the program.

testnext statementloop bodytruefalse
Figure 3.2. The shape of a loop. The test is evaluated, the body runs when it holds, and control returns to the test; the loop is left by the other exit.

The While Loop

Python’s while statement has the syntax of if: a test, a colon, and an indented body. The test is evaluated; if it is True the body runs in full and control returns to the test; if it is False the body is skipped and execution continues after the loop.

The following program computes x2x^2 by adding xx to a running total xx times.

x = 3
ans = 0
num_iterations = 0

while num_iterations != x:
    ans = ans + x
    num_iterations = num_iterations + 1

print(f"{x} squared is {ans}")

Recording the values of the variables each time the test is reached, as though one were the interpreter, is called hand simulation, and it is how we check what a loop does.

Test evaluationxansnum_iterationsTest
1st300True
2nd331True
3rd362True
4th393False

On the fourth evaluation num_iterations has reached x, and the program prints 3 squared is 9.

Whether it stops at all depends on x, and there are three cases. If x=0x = 0, the test fails at once, the body never runs, and 0 squared is 0 is printed. If x>0x > 0, the counter starts below x and rises by exactly one per pass, so after xx passes it equals x and the loop ends with the right answer. If x<0x < 0, the counter runs through 0,1,2,…0, 1, 2, \ldots and never equals a negative number: the test is never False, and the program runs forever. This is an infinite loop.

Weakening the test to num_iterations < abs(x) stops the loop after abs(x) passes, but each pass still adds the negative number x, so the program would announce that (−3)2=−9(-3)^2 = -9. The body has to be corrected too:

x = -3
ans = 0
num_iterations = 0

while num_iterations < abs(x):
    ans = ans + abs(x)
    num_iterations = num_iterations + 1

print(f"{x} squared is {ans}")   # -3 squared is 9

Example 3.9 (The leading digit).

Floor division by ten discards the last decimal digit, so applying it until one digit is left leaves the first digit. How many times it must be applied is not known before the loop starts, which is what a while loop is for.

n = 72658489290098
n = abs(n)

while n >= 10:
    n = n // 10

print(n)   # 7

Two further pieces of Python are needed for the problems below. Writing + between two strings joins them end to end, so 'X' + 'X' is 'XX'. And input(prompt) prints its prompt, waits for the user to type a line, and returns what was typed as a string; int converts a string of digits to the integer it denotes, so int(input('n? ')) reads a number.

Problem 3.12.

The program below should print the letter X a given number of times. Replace the comment by a while loop that appends 'X' to to_print exactly num_x times, and say what your loop does when the user enters 00 or a negative number.

num_x = int(input('How many times should I print the letter X? '))
to_print = ''
# append X to to_print num_x times
print(to_print)

Leaving a Loop Early

A break statement ends the loop containing it at once and passes control to the first statement after it, without returning to the test.

# the smallest positive integer divisible by both 11 and 12
x = 1

while True:
    if x % 11 == 0 and x % 12 == 0:
        break
    x = x + 1

print(x, 'is divisible by 11 and 12')   # 132 is divisible by 11 and 12

The test while True never fails, so the break is the only way out, and the proof that the loop stops is about the break. This arrangement suits a loop whose exit condition is natural to check partway through the body rather than at the top.

Where one loop sits inside another, a break ends only the loop that immediately contains it; the outer loop carries on.

Problem 3.13.

Write a program that reads ten integers, one at a time, and then prints the largest odd number among them, or a message saying that none was odd. Do not use max.

The For Loop

The loops above share a pattern: a counter is set up, tested at the top, and advanced at the bottom of the body. The for statement does that bookkeeping itself. Its form is

for variable in sequence:
    body

The variable is bound to the first entry of the sequence and the body runs; then to the second, and the body runs again; and so on until the sequence is exhausted or a break intervenes.

total = 0

for num in (77, 11, 3):
    total = total + num

print(total)   # 91

The object (77, 11, 3) is a tuple, an ordered finite sequence written in parentheses.

The Range Function

The sequence is most often produced by range, which generates a progression of integers and takes one, two or three arguments.

With one argument, range(stop) gives 0,1,…,stop−10, 1, \ldots, \text{stop} - 1.

for i in range(4):
    print(i)     # 0, then 1, then 2, then 3

With two, range(start, stop) gives start,start+1,…,stop−1\text{start}, \text{start}+1, \ldots, \text{stop}-1. The lower end is included and the upper end is not.

total = 0

for x in range(5, 11):
    total = total + x

print(total == 5 + 6 + 7 + 8 + 9 + 10)   # True

With three, range(start, stop, step) gives start,start+step,start+2 step,…\text{start}, \text{start} + \text{step}, \text{start} + 2\,\text{step}, \ldots, and stops before reaching stop. For positive step the last entry is the largest start+i step\text{start} + i\,\text{step} below stop; a negative step descends instead, so range(40, 5, -10) gives 40,30,20,1040, 30, 20, 10.

total = 0

for x in range(10, 3, -1):
    if x % 2 == 1:
        total = total + x

print(total)   # 9 + 7 + 5 = 21

Choosing the step to land on the right numbers is often more trouble than testing them inside the body. The following sums the odd numbers between m and n whatever the parity of m:

m, n = 4, 10
total = 0

for x in range(m, n + 1):
    if x % 2 == 1:
        total = total + x

print(total)   # 5 + 7 + 9 = 21

The two-argument form covers the one-argument form: range(0, 3) produces the same sequence as range(3). The entries are produced one at a time as the loop asks for them rather than stored all at once, so range(1000000) costs no more memory than range(3).

Remark (Which loop to use).

A for loop is the right choice when the number of passes is settled before the loop begins, since range then states that number and no counter can be mismanaged. A while loop is the right choice when it is not: the leading-digit program above stops when the number falls below ten, and how many divisions that takes is a fact about the input rather than about the program.

Example 3.10 (Squaring by repeated addition again).

The squaring program loses its counter entirely when written with for:

x = -3
ans = 0

for num_iterations in range(abs(x)):
    ans = ans + abs(x)

print(f"{x} squared is {ans}")   # -3 squared is 9

There is no explicit test and no explicit increment; range supplies both.

Reassigning the Loop Variable

Assigning to the loop variable inside the body does not disturb the loop.

for i in range(2):
    print(i)
    i = 0
    print(i)

This prints 0, 0, 1, 0 and stops. The sequence is fixed when the for statement is first reached, and at the start of each pass the variable is rebound to the next entry of it, whatever happened to the variable in between. The loop is equivalent to

index = 0
last_index = 1

while index <= last_index:
    i = index
    print(i)
    i = 0
    print(i)
    index = index + 1

For the same reason, changing a variable that was used in the range call has no effect, because the call is evaluated once:

x = 1

for i in range(x):
    print(i)
    x = 4

# prints 0, and nothing else

An inner for is a different matter: its range call is reached afresh on every pass of the outer loop and so is evaluated again.

x = 3

for j in range(x):
    print('Outer')
    for i in range(x):
        print('  Inner')
        x = 2

The outer range(x) is evaluated once, with x=3x = 3, so the outer loop makes three passes. The inner range(x) sees x=3x = 3 on the first pass and x=2x = 2 afterwards, giving 3+2+2=73 + 2 + 2 = 7 inner passes in total.

Iterating Over a String

A string is a sequence of characters. Its positions are numbered from 00, so a string of length kk occupies positions 0,1,…,k−10, 1, \ldots, k-1: len(s) gives the length, s[i] gives the character at position i, and the slice s[i:j] gives the characters at positions i up to but not including j. That upper end is excluded for the same reason it is excluded from range, and the two conventions agree: s[0:len(s)] is the whole of s, and the positions it covers are exactly those produced by range(len(s)).

Combined with in, the for statement walks a string directly, binding the variable to one character at a time and dispensing with the positions altogether.

total = 0

for c in '12345678':
    total = total + int(c)

print(total)   # 36

Problem 3.14.

Write a program that computes 5+6+⋯+1005 + 6 + \cdots + 100 with a for loop, and check the result against the closed form for an arithmetic progression obtained in the last lesson. Then do the same for 5+7+9+⋯+995 + 7 + 9 + \cdots + 99, once by choosing the step of the range and once by testing inside the body.

Nested Loops

A loop inside a loop runs the inner loop to completion on every pass of the outer one.

Example 3.11 (Patterns).

An n×nn \times n block of asterisks needs one loop for the rows and one for the columns:

n = 5

for row in range(n):
    for col in range(n):
        print('*', end='')
    print()
*****
*****
*****
*****
*****

The bare print() after the inner loop ends the line. Letting the inner range depend on the outer variable changes the shape:

n = 5

for row in range(n):
    print(row, end=' ')
    for col in range(row):
        print('*', end=' ')
    print()
0
1 *
2 * *
3 * * *
4 * * * *

Row kk carries kk asterisks, so the block becomes a triangle.

Continue and Pass

Two keywords sit alongside break. A continue abandons the rest of the current pass and goes straight to the next one: back to the test for a while, on to the next entry for a for. A pass does nothing, and exists because Python’s syntax requires a statement in places where no action is wanted.

for n in range(200):
    if n % 3 == 0:
        continue
    elif n == 8:
        break
    else:
        pass
    print(n, end=' ')

# 1 2 4 5 7

The multiples of three never reach the print, the loop ends when n reaches 88, and the remaining values pass through the else branch and are printed.

Divisibility and Primality

Definition 3.12 (Divisibility).

Let d,n∈Zd, n \in \mathbb{Z}. We say dd divides nn, written d∣nd \mid n, if n=dqn = dq for some q∈Zq \in \mathbb{Z}; equivalently, for d≠0d \neq 0, if n mod d=0n \bmod d = 0. In Python this is the test n % d == 0.

Definition 3.13 (Prime and composite).

An integer n⩾2n \geqslant 2 is prime if its only positive divisors are 11 and nn, and composite otherwise. The integers 00 and 11 and the negative integers are neither.

Checking that 9191 is composite takes a single operation once the divisor 77 is known, since 91 % 7 == 0 settles it; finding the divisor is where the work lies. Read as an instruction, the definition says to test every candidate between 22 and n−1n - 1 and to stop at the first one that divides.

n = 91
is_prime = n >= 2

for factor in range(2, n):
    if n % factor == 0:
        is_prime = False
        break

if is_prime:
    print(f'{n} is prime')
else:
    print(f'{n} is not prime')

# 91 is not prime

One divisor settles the question, so the loop stops at the first one it finds. Wrapping the whole thing in a second loop lists the primes below a bound:

for n in range(2, 100):
    is_prime = True
    for factor in range(2, n):
        if n % factor == 0:
            is_prime = False
            break
    if is_prime:
        print(n, end=' ')

# 2 3 5 7 11 13 17 19 23 29 31 37 41 43 47 53 59 61 67 71 73 79 83 89 97

Trial Division to the Square Root

Testing every candidate up to n−1n - 1 is more than necessary.

Proposition 3.14 (A composite number has a small divisor).

Let n⩾2n \geqslant 2. Then nn is composite if and only if some integer dd with 2⩽d⩽n2 \leqslant d \leqslant \sqrt{n} divides nn.

Discussion.

Divisors come in pairs: if d∣nd \mid n then dd and n/dn/d are both divisors and their product is nn. A product of two numbers both exceeding n\sqrt{n} exceeds nn, so the two members of a pair cannot both lie above n\sqrt{n}, and one of them is at or below it. The proof takes the least divisor above 11; its partner is then the larger of the two, and the inequality d⩽n/dd \leqslant n/d can be squared. The converse direction is immediate: a divisor in that range is neither 11 nor nn, because n<n\sqrt{n} < n once n⩾2n \geqslant 2.

Proof.

Suppose nn is composite. The set D={ d∈Z:d⩾2 and d∣n }D = \{\, d \in \mathbb{Z} : d \geqslant 2 \text{ and } d \mid n \,\} contains nn, so it is non-empty; let dd be its least element and write n=dqn = dq with q∈Zq \in \mathbb{Z}, so that q⩾1q \geqslant 1 and q∣nq \mid n.

We first rule out q=1q = 1. If q=1q = 1 then d=nd = n. But nn is composite, so it has a divisor ee with 1<e<n1 < e < n; that ee lies in DD and is smaller than n=dn = d, contradicting the minimality of dd. Hence q⩾2q \geqslant 2, so q∈Dq \in D and therefore q⩾dq \geqslant d. Consequently

d2⩽dq=n,d^2 \leqslant dq = n ,

so d⩽nd \leqslant \sqrt{n}.

Conversely, suppose 2⩽d⩽n2 \leqslant d \leqslant \sqrt{n} and d∣nd \mid n. From n⩾2n \geqslant 2 we get n<n\sqrt{n} < n, so d≠nd \neq n, and d≠1d \neq 1 by hypothesis. Thus nn has a positive divisor other than 11 and nn, so it is composite.

Only the integers up to n\sqrt{n} need be tested, and 22 can be dealt with separately so that the loop skips the even candidates:

n = 1010809
is_prime = n >= 2

if n > 2 and n % 2 == 0:
    is_prime = False
else:
    max_factor = round(n ** 0.5)
    for factor in range(3, max_factor + 1, 2):
        if n % factor == 0:
            is_prime = False
            break

print(f'{n} is prime: {is_prime}')   # 1010809 is prime: True

round(n ** 0.5) returns either ⌊n⌋\lfloor \sqrt{n} \rfloor or one more, never less, so the loop may test one candidate too many and never one too few.

The first program tests up to n−2n - 2 candidates and the second about n/2\sqrt{n}/2: for n=1010809n = 1010809, about a million tests against about five hundred. Python’s time module measures it. The statement import time makes the module’s tools available, and time.time() returns the current time in seconds. A colon inside the braces of an f-string says how to format the value: :.2f prints it with two digits after the decimal point.

import time

n = 1010809

time0 = time.time()
is_prime_basic = True
for factor in range(2, n):
    if n % factor == 0:
        is_prime_basic = False
        break
time1 = time.time()
print(f'Basic: {(time1 - time0) * 1000:.2f} ms')

time0 = time.time()
is_prime_fast = n >= 2
if n > 2 and n % 2 == 0:
    is_prime_fast = False
else:
    max_factor = round(n ** 0.5)
    for factor in range(3, max_factor + 1, 2):
        if n % factor == 0:
            is_prime_fast = False
            break
time1 = time.time()
print(f'Optimised: {(time1 - time0) * 1000:.2f} ms')

print(is_prime_basic == is_prime_fast)   # True

The two agree, and on an ordinary machine the first takes tens of milliseconds where the second takes a small fraction of one.

Searching for the nnth Object

Finding the nnth number with a given property is a different kind of problem: the answer is not known in advance, so there is no range to run over. A counter of successes and a candidate that advances by one at a time turn it into a while loop.

target = 5
found = 0
guess = 0

while found <= target:
    guess = guess + 1
    if guess % 4 == 0 or guess % 7 == 0:
        found = found + 1

print(f'Counting from zero, entry {target} of the multiples of 4 or 7 is {guess}')
# entry 5 is 16

The multiples of four or seven begin 4,7,8,12,14,164, 7, 8, 12, 14, 16, and counting from zero the entry numbered 55 is 1616. Substituting the primality test for the divisibility test finds the nnth prime:

target = 10
found = 0
guess = 1

while found <= target:
    guess = guess + 1
    is_prime = guess >= 2
    if guess > 2 and guess % 2 == 0:
        is_prime = False
    else:
        max_factor = round(guess ** 0.5)
        for factor in range(3, max_factor + 1, 2):
            if guess % factor == 0:
                is_prime = False
                break
    if is_prime:
        found = found + 1

print(f'Counting from zero, prime {target} is {guess}')   # prime 10 is 31

The primes are 2,3,5,7,11,13,17,19,23,29,312, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, and counting from zero the one numbered 1010 is 3131.

Problem 3.15.

Write a program that prints the sum of the primes strictly between 22 and 10001000, using a primality test nested inside a loop over the odd integers from 33 to 999999. Then say how many candidate divisors your program tests in total, and how many the version without the square-root bound would test.

Digits in an Arbitrary Base

In the first chapter we read off base-bb expansions by hand, dividing by bb, recording the remainder as the next digit from the right, and continuing with the quotient until it reached zero. Automating that needs a way of repeating instructions until a condition is met, which the while loop provides. The number of repetitions is the number of digits, and it is not known before the divisions are done.

Proposition 3.15 (Digit extraction).

Let b>1b > 1 and n∈Nn \in \mathbb{N}. Define n0=nn_0 = n and nk+1=⌊nk/b⌋n_{k+1} = \lfloor n_k / b \rfloor. Then nk=0n_k = 0 for some kk, the least such kk is m=⌊log⁡bn⌋+1m = \lfloor \log_b n \rfloor + 1, and for 0⩽i<m0 \leqslant i < m

ni mod b=zi,n_i \bmod b = z_i ,

where zm−1⋯z1z0z_{m-1} \cdots z_1 z_0 is the base-bb expansion of nn.

Discussion.

One pass of the loop performs the split z=b⌊z/b⌋+(z mod b)z = b\lfloor z/b \rfloor + (z \bmod b), which removes the last digit. The proof identifies nkn_k explicitly, as the number whose digits are the top m−km - k digits of nn, and proves this by induction on kk. The step is the split applied to nkn_k, and the remainder is the digit zkz_k by the uniqueness of quotient and remainder, as in the uniqueness half of Theorem 2.21. The number of steps then follows: the sum defining nkn_k is empty exactly when k=mk = m, and is at least 11 before that, because its leading digit is not zero.

Proof.

By Corollary 2.22 the number nn has a unique expansion

n=∑i=0m−1zibi,zi∈{0,…,b−1},zm−1≠0,m=⌊log⁡bn⌋+1.n = \sum_{i=0}^{m-1} z_i b^i, \qquad z_i \in \{0, \ldots, b-1\}, \quad z_{m-1} \neq 0, \quad m = \lfloor \log_b n \rfloor + 1 .

We claim that

nk=∑i=km−1zib i−k(0⩽k⩽m).n_k = \sum_{i=k}^{m-1} z_i b^{\,i-k} \qquad (0 \leqslant k \leqslant m).

For k=0k = 0 this is the expansion itself. Assume it for some k<mk < m and split off the term i=ki = k:

nk=zk+b∑i=k+1m−1zib i−k−1.n_k = z_k + b\sum_{i=k+1}^{m-1} z_i b^{\,i-k-1} .

The sum on the right is an integer and 0⩽zk<b0 \leqslant z_k < b, so this is the division of nkn_k by bb with quotient and remainder, and by the uniqueness of that division

nk+1=⌊nkb⌋=∑i=k+1m−1zib i−k−1,nk mod b=zk,n_{k+1} = \left\lfloor \frac{n_k}{b} \right\rfloor = \sum_{i=k+1}^{m-1} z_i b^{\,i-k-1}, \qquad n_k \bmod b = z_k ,

which is the claim at k+1k+1 together with the stated identity for the digits.

At k=mk = m the sum is empty, so nm=0n_m = 0. For k<mk < m every term is non-negative and the term i=m−1i = m-1 is zm−1b m−1−k⩾1z_{m-1}b^{\,m-1-k} \geqslant 1, so nk⩾1n_k \geqslant 1. Hence mm is the least index at which the sequence vanishes.

The loop below carries out the proposition. Digits are produced from the last to the first, so each new one is put in front of what has been built so far, and digits above nine are turned into letters: str(d) turns the number d into the string of its decimal digits, ord(c) gives the code number of the character c, chr(k) gives the character with code k, and the codes of A to Z run consecutively.

n = 12345
b = 17
digits = ''

while n > 0:
    remainder = n % b
    if remainder < 10:
        digits = str(remainder) + digits
    else:
        digits = chr(ord('A') + remainder - 10) + digits
    n = n // b

print(digits)   # 28C3

Example 3.16 (Checking the expansion).

Reading (28C3)17(28\text{C}3)_{17} back gives

2⋅173+8⋅172+12⋅17+3=9826+2312+204+3=12345,2 \cdot 17^3 + 8 \cdot 17^2 + 12 \cdot 17 + 3 = 9826 + 2312 + 204 + 3 = 12345 ,

as the proposition says.

Problem 3.16.

The loop above prints nothing when n=0n = 0. Say why, in terms of the hypotheses of the proposition, and repair it so that it prints '0'. Then use it to find the base-77 and hexadecimal expansions of 99999999, and check them against the by-hand method of the first chapter.

Problem 3.17.

Write a program that repeatedly asks the user for a string and prints it back, stopping when the user enters 'done'. It should then print 'Bye!' followed by the number of strings entered, not counting 'done'.

Running Time

Timing the two primality tests measured one machine at one time. The number of instructions a program performs depends on the compiler and on the hardware, and we do not in any case know it exactly. We want a statement about the algorithm instead: how the amount of work grows with the input.

Definition 3.17 (Elementary operation and running time).

An elementary operation is one the machine performs in a bounded number of clock cycles independently of the values involved: an arithmetic operation or a comparison on numbers of bounded size, a read or a write of one stored value. Both a // b and a % b are elementary in this sense.

The running time of an algorithm is the number of elementary operations it performs, counted as a function of its input.

Two things remain to be fixed: how the size of an input is measured, and how precisely we count.

Landau Notation

Definition 3.18 (Landau's OO).

Let g:N→R⩾0g : \mathbb{N} \to \mathbb{R}_{\geqslant 0}. Then O(g)O(g) is the set of functions f:N→R⩾0f : \mathbb{N} \to \mathbb{R}_{\geqslant 0} for which there exist α∈R>0\alpha \in \mathbb{R}_{>0} and n0∈Nn_0 \in \mathbb{N} with

f(n)⩽α⋅g(n)for all n⩾n0.f(n) \leqslant \alpha \cdot g(n) \qquad \text{for all } n \geqslant n_0 .

The constant α\alpha must not depend on nn; were it allowed to, every function would lie in O(1)O(1) and the definition would say nothing. Nor does the inequality have to hold everywhere: it may fail for finitely many nn, so changing ff at the first million values leaves the statement untouched.

Instead of f∈O(g)f \in O(g) one usually writes

f=O(g),orf(x)=O(g(x)) as x→∞,f = O(g), \qquad\text{or}\qquad f(x) = O\bigl(g(x)\bigr) \text{ as } x \to \infty ,

read ”ff is big-Oh of gg”. The equals sign here is not the symmetric one, since O(g)O(g) is a set and ff is a member of it; the notation is standard nonetheless.

Example 3.19 (A cubic is O(x4)O(x^4)).

Take f(x)=4x3+7xf(x) = 4x^3 + 7x and g(x)=x4g(x) = x^4. For x⩾1x \geqslant 1 we have x3⩽x4x^3 \leqslant x^4 and x⩽x4x \leqslant x^4, so

4x3+7x⩽4x4+7x4=11x4,4x^3 + 7x \leqslant 4x^4 + 7x^4 = 11x^4 ,

and α=11\alpha = 11 with n0=1n_0 = 1 meets the definition. Hence 4x3+7x=O(x4)4x^3 + 7x = O(x^4).

Example 3.20 (A tight bound).

Take f(n)=3n2+5n+7f(n) = 3n^2 + 5n + 7 and g(n)=n2g(n) = n^2. For n⩾1n \geqslant 1 we have n⩽n2n \leqslant n^2 and 1⩽n21 \leqslant n^2, so

3n2+5n+7⩽3n2+5n2+7n2=15n2,3n^2 + 5n + 7 \leqslant 3n^2 + 5n^2 + 7n^2 = 15n^2 ,

and α=15\alpha = 15 with n0=1n_0 = 1 serves. Hence 3n2+5n+7=O(n2)3n^2 + 5n + 7 = O(n^2).

Here the bound grows at the same rate as ff itself, where the previous example bounded a cubic by a quartic. Both statements are true, but the cubic-by-quartic one loses information, since x4x^4 grows strictly faster than 4x3+7x4x^3 + 7x; the same argument with α=11\alpha = 11 gives the sharper 4x3+7x=O(x3)4x^3 + 7x = O(x^3). A function lies in O(g)O(g) for many different gg, and the useful statement uses the slowest-growing gg available.

Example 3.21 (When the definition fails).

Not every pair of functions is related this way: n2n^2 is not O(n)O(n). Suppose it were, so that n2⩽αnn^2 \leqslant \alpha n for some α>0\alpha > 0 and all n⩾n0n \geqslant n_0. Dividing by nn, which is positive, gives n⩽αn \leqslant \alpha for all n⩾n0n \geqslant n_0. But

n=max⁡(n0,⌈α⌉)+1n = \max\bigl(n_0, \lceil \alpha \rceil\bigr) + 1

satisfies n⩾n0n \geqslant n_0 and n>αn > \alpha, which contradicts it. So no constant serves, and the direction of an OO statement is not reversible: n=O(n2)n = O(n^2) holds while n2=O(n)n^2 = O(n) does not.

The argument of the first two examples works for any polynomial.

Proposition 3.22 (Polynomials).

Let f(n)=adnd+ad−1nd−1+⋯+a1n+a0f(n) = a_d n^d + a_{d-1}n^{d-1} + \cdots + a_1 n + a_0 with every ai⩾0a_i \geqslant 0 and d∈N0d \in \mathbb{N}_0. Then f=O(nd)f = O(n^d).

Discussion.

The proof uses one inequality: for n⩾1n \geqslant 1 and i⩽di \leqslant d we have ni⩽ndn^i \leqslant n^d, because raising a number at least 11 to a larger exponent cannot decrease it. Applying it to each term replaces every power by the top power, and what is left is the sum of the coefficients, which serves as the constant α=ad+⋯+a0\alpha = a_d + \cdots + a_0 of the definition. Non-negative coefficients let each term be bounded separately; with negative coefficients the same bound holds after taking absolute values of the coefficients.

Proof.

Let n⩾1n \geqslant 1. For each ii with 0⩽i⩽d0 \leqslant i \leqslant d we have ni⩽ndn^i \leqslant n^d, so aini⩽ainda_i n^i \leqslant a_i n^d since ai⩾0a_i \geqslant 0. Adding these d+1d+1 inequalities,

f(n)=∑i=0daini  ⩽  (∑i=0dai)nd.f(n) = \sum_{i=0}^{d} a_i n^i \;\leqslant\; \left(\sum_{i=0}^{d} a_i\right) n^d .

Taking α=∑i=0dai\alpha = \sum_{i=0}^{d} a_i, which is a positive constant unless every aia_i is 00, and n0=1n_0 = 1, the definition is satisfied.

Remark (The limit form).

For a reader who has met limits there is a shorter route to most OO statements. Suppose g(n)≠0g(n) \neq 0 from some point on and the ratio f(n)/g(n)f(n)/g(n) tends to a finite limit,

lim⁡n→∞f(n)g(n)=L<∞.\lim_{n \to \infty} \frac{f(n)}{g(n)} = L < \infty .

Then f=O(g)f = O(g): beyond some n0n_0 the ratio stays below L+1L + 1, so α=L+1\alpha = L + 1 meets the definition. The cubic example is settled in one line this way, since (4x3+7x)/x4→0(4x^3 + 7x)/x^4 \to 0, and so is the tight bound, since (3n2+5n+7)/n2→3(3n^2 + 5n + 7)/n^2 \to 3.

The converse fails only because the ratio need not converge at all, and replacing the limit by the limit superior repairs it: f=O(g)f = O(g) holds exactly when

lim sup⁡n→∞f(n)g(n)<∞.\limsup_{n \to \infty} \frac{f(n)}{g(n)} < \infty .

A limit of 00 says more than OO does. It says that ff is negligible against gg rather than merely bounded by a multiple of it, and that stronger relation is written f=o(g)f = o(g); so 4x3+7x=o(x4)4x^3 + 7x = o(x^4), while 3n2+5n+7=O(n2)3n^2 + 5n + 7 = O(n^2) is not o(n2)o(n^2).

Remark (What $O$ does and does not see).

The notation is insensitive to scaling: if f=O(g)f = O(g) then Cf=O(Dg)Cf = O(Dg) for any non-zero constants CC and DD. It is equally insensitive to any finite initial stretch of the two functions. What it describes is therefore the asymptotic behaviour of ff and gg, their behaviour for large inputs rather than at any one input. Constant factors depend on the machine and the compiler, which we do not model, so this is the level of precision we want; for the same reason an OO bound alone does not decide which of two programs to run on a given machine.

The Cost of Schoolbook Arithmetic

We all learned at school how to add and multiply natural numbers written in decimal: once m+nm + n and mnmn are known for the single digits 0⩽m,n⩽90 \leqslant m, n \leqslant 9, sums and products of arbitrary numbers follow by working column by column and carrying. These schoolbook methods are algorithms in the sense of the first chapter, and we can count their cost.

Remark (The base does not matter).

There is nothing special about ten here. The same algorithms work, with minor changes, in any base, binary included. Binary is convenient for multiplication, because the table of single-digit products that has to be known in advance is very much smaller.

Throughout this section the size of the input is the number of digits, which for xx written in base bb is ⌊log⁡bx⌋+1\lfloor \log_b x \rfloor + 1 by Corollary 2.22.

Proposition 3.23 (Cost of schoolbook addition).

Fix a base b⩾2b \geqslant 2. Computing x+yx + y by the schoolbook algorithm takes O(n)O(n) elementary operations, where nn is the larger of the numbers of base-bb digits of xx and of yy.

Discussion.

The algorithm treats one column at a time, and each column costs at most a fixed amount, however large the numbers are. A single column adds three quantities: the digit of xx there, the digit of yy there, and the carry coming in from the column to its right. The first two lie in {0,…,b−1}\{0, \ldots, b-1\} and the third is 00 or 11, so there are at most 2b22b^2 possible columns to deal with, a number fixed once bb is fixed and independent of nn. So each column costs at most some constant C=C(b)C = C(b), there are nn columns and at most one extra step to write a final carry, and the count is Cn+CCn + C. Padding the shorter number with leading zeros makes both numbers nn digits long.

Proof.

Write xx and yy in base bb, padding the shorter expansion with leading zeros so that both have nn digits.

The algorithm works through the columns from right to left. At each column it adds three quantities: the digit of xx in that column, the digit of yy in that column, and the carry from the preceding column. The two digits lie in {0,…,b−1}\{0, \ldots, b-1\} and the carry is 00 or 11, so there are finitely many possible single-column computations, their number depending on bb alone. Since bb is fixed, there is a constant CC such that every single-column computation takes at most CC elementary operations.

The algorithm performs one such computation for each of the nn columns and at most one further step to write a final carry, so its running time is at most Cn+CCn + C. By the proposition on polynomials this is O(n)O(n).

No algorithm does better than O(n)O(n) here, since the output has about nn digits and writing it down takes that long. Schoolbook multiplication costs more.

Proposition 3.24 (Cost of schoolbook multiplication).

Fix a base b⩾2b \geqslant 2. Computing xyxy by the schoolbook algorithm takes O(n2)O(n^2) elementary operations, with nn as above.

Discussion.

The algorithm has two stages, bounded separately. In the first it forms one partial product for each digit of yy: multiplying the whole of xx by a single digit is a column-by-column pass of the kind already costed, so it is O(n)O(n), and there are nn digits of yy, giving O(n2)O(n^2). Shifting a partial product left only decides where its digits are written. In the second stage the nn partial products are added, and each addition involves numbers of at most 2n2n digits, so by the previous proposition each costs O(n)O(n) and the nn of them cost O(n2)O(n^2). Two stages of O(n2)O(n^2) make O(n2)O(n^2).

Proof.

Write xx and yy in base bb, padded to nn digits each.

The algorithm forms nn partial products, one for each digit yjy_j of yy. Forming the one belonging to yjy_j means multiplying yjy_j by each of the nn digits of xx, one column at a time and carrying where needed, and then shifting the result jj places to the left to obtain x⋅yjb jx \cdot y_j b^{\,j}. Each single-digit product with a carry is one of finitely many computations depending on bb alone, so each partial product costs O(n)O(n) and all nn of them cost O(n2)O(n^2).

It then adds the nn partial products. Each is at most 2n2n digits long, so by the previous proposition each addition costs O(n)O(n), and nn of them cost O(n2)O(n^2).

Both stages are O(n2)O(n^2), so there are constants bounding each by a multiple of n2n^2 beyond some point; adding the two bounds gives a constant bounding the total, and the running time is O(n2)O(n^2).

Remark (Faster multiplication).

Multiplication algorithms asymptotically better than the schoolbook one do exist, though they are a great deal more complicated. The first was found by Karatsuba in 1960 and multiplies two nn-digit numbers in O(nα)O(n^{\alpha}) operations with α=log⁡3/log⁡2≈1.58\alpha = \log 3 / \log 2 \approx 1.58. More recently Harvey and van der Hoeven gave an O(nlog⁡n)O(n \log n) algorithm, which is believed to be the best possible.

This last result also shows a limitation of the notation. The running time is bounded by Cnlog⁡nCn\log n for some constant CC, but CC is very large: the algorithm beats its rivals only on enormous inputs, and at the sizes that arise in practice methods with worse asymptotic behaviour and smaller constants are faster.

The Cost of Trial Division

Remark (The cost of the obvious method).

The definition of primality suggests testing, for a given n⩾2n \geqslant 2, every pair a,ba, b with 2⩽a,b⩽n/22 \leqslant a, b \leqslant n/2 to see whether n=abn = ab. That is up to (n/2−1)2(n/2 - 1)^2 multiplications. Today’s machines perform several billion operations a second, and even so an eight-digit input on this method would occupy one of them for several hours.

Trial division does much better.

Proposition 3.25 (Cost of trial division).

Deciding whether n⩾2n \geqslant 2 is prime by testing every candidate divisor from 22 to ⌊n⌋\lfloor \sqrt{n} \rfloor takes O(n)O(\sqrt{n}) elementary operations.

Discussion.

Each pass of the loop performs a fixed amount of work, one remainder and one comparison, both elementary, so the running time is a constant multiple of the number of passes. The loop runs over the integers from 22 to ⌊n⌋\lfloor \sqrt n \rfloor and may stop early at a divisor, so the count is at most ⌊n⌋−1\lfloor \sqrt n \rfloor - 1; the floor is at most n\sqrt n, and the constants are absorbed by the OO. That stopping at n\sqrt n rather than at n−1n - 1 is correct is the proposition on small divisors.

Proof.

By the proposition on small divisors, nn is composite exactly when some dd with 2⩽d⩽n2 \leqslant d \leqslant \sqrt{n} divides it, so the loop over d=2,…,⌊n⌋d = 2, \ldots, \lfloor \sqrt{n} \rfloor decides the question. It makes at most ⌊n⌋−1⩽n\lfloor \sqrt{n}\rfloor - 1 \leqslant \sqrt{n} passes, and each pass computes one remainder and one comparison, which is a bounded number CC of elementary operations. The total is at most CnC\sqrt{n}, which is O(n)O(\sqrt{n}).

Remark (Avoiding the square root).

It is not obvious that ⌊n⌋\lfloor \sqrt{n} \rfloor can itself be computed with an elementary operation, and the algorithm need not compute it. Increasing the candidate ii by 11 and stopping as soon as i⋅i>ni \cdot i > n tests exactly the same candidates and uses only multiplication and comparison. Every step of the algorithm is then well defined for every admissible input, the algorithm stops after finitely many steps, and it returns the right answer, so it computes the function f:N→{yes,no}f : \mathbb{N} \to \{\text{yes}, \text{no}\} taking the value yes exactly at the primes. We write {true,false}\{\text{true}, \text{false}\} or {0,1}\{0, 1\} for the two values as well.

Remark (Fast in $n$, slow in the input).

An O(n)O(\sqrt{n}) bound looks good, but nn is not the size of the input. The instance handed to the algorithm is the digit string of nn, whose length is m=⌊log⁡bn⌋+1m = \lfloor \log_b n \rfloor + 1, so nn is about bmb^{m} and n\sqrt{n} is about bm/2b^{m/2}. Measured against the length of its input, trial division takes exponentially many steps, so a three-hundred-digit number is out of its reach at any realistic speed.

Problem 3.18.

Give the running time of each of the following in OO notation, as a function of nn, and justify each answer.

  1. Summing the integers from 11 to nn with a loop.
  2. Summing the integers from 11 to nn with the closed form.
  3. Printing every pair (i,j)(i, j) with 1⩽i<j⩽n1 \leqslant i < j \leqslant n.
  4. Extracting the base-bb digits of nn by repeated division.

Problem 3.19.

A product xyxy may be computed by adding xx to a running total yy times, using no multiplication at all. Let xx and yy have nn digits in base bb.

  1. Give the running time of this method as a function of nn and bb.
  2. Evaluate that count and the n2n^2 of the schoolbook algorithm at b=10b = 10 for n=1n = 1, n=3n = 3 and n=20n = 20. On a machine performing 10910^9 operations a second, say for which of the three the repeated-addition method is still usable.
  3. Both methods perform additions of nn-digit numbers, and only one of them is out of reach for a twenty-digit input. Say which quantity in your answer to the first part is responsible, and why measuring the input by nn rather than by yy is what makes the difference visible.

Problem 3.20.

Two programs settle the same problem on inputs of size nn. The first performs 100n100n elementary operations, the second n2/100n^2/100.

  1. Find every nn at which the second is the faster, and the size at which the first overtakes it.
  2. Give the running time of each in OO notation, and say what those two statements do and do not tell you about which program to run.
  3. A machine performs 10910^9 operations a second and no input ever exceeds n=5000n = 5000. Which program should be run, and how long does it take?

Search and Approximation

The loops of the previous chapters mostly checked something: whether any candidate divides nn, whether every character of a string is a digit. The loops of this chapter search for a value: they try candidates until one works.

Exhaustive Enumeration

The simplest search strategy is the one the primality test already used: try every candidate in some collection until one works. When the collection is finite, or can be made finite by bounding the range, this is exhaustive enumeration, and it finds a solution whenever one exists.

Cube Roots

Suppose we want the integer cube root of a perfect cube. Given an integer nn, we seek an integer xx with x3=nx^3 = n, or a report that no such integer exists.

The strategy is direct: test x=0,1,2,…x = 0, 1, 2, \ldots in order until either x3=∣n∣x^3 = |n|, a success, or x3>∣n∣x^3 > |n|, a failure, the latter being conclusive because a<ba < b implies a3<b3a^3 < b^3 for non-negative integers. For negative nn the cube root is the negative of the cube root of ∣n∣|n|.

n = 27
x = 0

while x ** 3 < abs(n):
    x = x + 1

if x ** 3 != abs(n):
    print(f'{n} is not a perfect cube')
else:
    if n < 0:
        x = -x
    print(f'Cube root of {n} is {x}')

# Cube root of 27 is 3

Hand simulating for n=27n = 27:

Test evaluationxx ** 3x ** 3 < 27
1st00True
2nd11True
3rd28True
4th327False

The loop stops with x=3x = 3, and since 33=27=∣n∣3^3 = 27 = |n| the program reports a cube.

Problem 3.21.

Trace the program above for n=8n = 8, n=−8n = -8 and n=9n = 9. In each case build the hand-simulation table, state the value of x when the loop stops, and say whether the program reports a perfect cube.

Example 3.26 (The speed of exhaustive search).

Take n=1,957,816,251n = 1{,}957{,}816{,}251. Its cube root is 1,2511{,}251, so the loop makes 1,2511{,}251 passes and finishes at once. Take instead n=7,406,961,012,236,344,616n = 7{,}406{,}961{,}012{,}236{,}344{,}616, whose cube root is 1,949,3061{,}949{,}306: the loop makes nearly two million passes and still finishes in well under a second. At billions of instructions a second, a million passes take a fraction of a second.

Termination and Decrementing Functions

Every loop in a correct program must stop, unless it is deliberately endless. To prove that one does, we exhibit a quantity that strictly decreases at every pass and cannot go below zero.

Definition 3.27 (Decrementing function).

A decrementing function for a loop is an integer-valued expression DD in the program’s variables such that

  1. D⩾0D \geqslant 0 whenever the loop test holds, and
  2. DD strictly decreases at every pass of the body.

A loop admitting a decrementing function stops after at most D0D_0 passes, where D0D_0 is the initial value of DD.

With it, “the loop eventually stops” can be proved. For the cube-root search a suitable choice is D=⌈∣n∣1/3⌉−xD = \lceil |n|^{1/3}\rceil - x, written with the ceiling of the last lesson, which may also be had from the floor as ⌈x⌉=−⌊−x⌋\lceil x\rceil = -\lfloor -x\rfloor.

Theorem 3.28 (Termination of the cube-root search).

The exhaustive cube-root search stops for every integer nn.

Discussion.

The variable xx increases rather than decreases, so it is not itself a decrementing function; what decreases is the distance from xx to the value at which the loop stops, and the loop stops once x3x^3 reaches ∣n∣|n|, that is, once xx reaches ∣n∣1/3|n|^{1/3}. Since xx is an integer this value is rounded up, which is where the ceiling enters. The two conditions of the definition then have to be checked separately: non-negativity comes from the loop test, since the test holding means xx has not yet reached the target, and the strict decrease comes from the body, since the body adds exactly 11 to xx and does not change the target.

Proof.

Put D=⌈∣n∣1/3⌉−xD = \lceil |n|^{1/3} \rceil - x. Initially x=0x = 0, so D=⌈∣n∣1/3⌉⩾0D = \lceil |n|^{1/3}\rceil \geqslant 0.

Suppose the loop test x3<∣n∣x^3 < |n| holds. Then x<∣n∣1/3⩽⌈∣n∣1/3⌉x < |n|^{1/3} \leqslant \lceil |n|^{1/3}\rceil, and both sides being integers gives D⩾1D \geqslant 1, so in particular D⩾0D \geqslant 0.

Each pass replaces xx by x+1x + 1 and changes nothing else, so DD falls by exactly 11. Being a non-negative integer that falls by 11 each pass, DD can do so at most ⌈∣n∣1/3⌉\lceil |n|^{1/3}\rceil times, and the loop stops after at most that many passes.

Note (Decrementing functions as a diagnostic).

When a program appears to run forever, the definition suggests where to look. Identify the quantity that ought to be decreasing; if there is none, the loop has no decrementing function, which suggests that the loop may never end. Otherwise print the candidate at every pass, and if the printed values fail to fall, the fault is in the body. Deleting x = x + 1 from the cube-root search, for instance, leaves D=⌈∣n∣1/3⌉D = \lceil |n|^{1/3}\rceil at every pass, constant instead of decreasing.

Remark (The cost of exhaustive enumeration).

Exhaustive enumeration is correct and simple, but it can be slow. The cube-root search tests ⌈∣n∣1/3⌉\lceil |n|^{1/3}\rceil candidates, so its running time is O(n1/3)O(n^{1/3}). For n=1012n = 10^{12} that is 10410^4 passes, which is quick; for n=1030n = 10^{30} it is 101010^{10} passes, minutes or hours. Bisection search needs about a hundred steps for the same nn.

Approximate Solutions

The cube-root search demands an exact integer answer, so it works only on perfect cubes. Many numerical problems have no exact answer in the integers, or even in the rationals: 2\sqrt{2} is irrational, and no finite program returns it. We ask instead for an approximate answer, within a stated tolerance.

Definition 3.29 (ε\varepsilon-approximation).

Let ff be a function, yy a target value and ε>0\varepsilon > 0 a prescribed tolerance. An ε\varepsilon-approximation to a solution of f(x)=yf(x) = y is a value x^\hat{x} with

∣f(x^)−y∣<ε.|f(\hat{x}) - y| < \varepsilon .

The tolerance is chosen by whoever writes the program. A smaller ε\varepsilon demands a more accurate answer, and can take much more computation.

Exhaustive Search for Square Roots

Exhaustive enumeration adapts to the approximate setting. Rather than the integers 0,1,2,…0, 1, 2, \ldots we test the values 0,δ,2δ,3δ,…0, \delta, 2\delta, 3\delta, \ldots for a small step δ>0\delta > 0, and accept the first x^\hat{x} with ∣x^2−n∣<ε|\hat{x}^2 - n| < \varepsilon.

n = 25
epsilon = 0.01
step = 0.0001
guess = 0.0
num_guesses = 0

while abs(guess ** 2 - n) >= epsilon and guess <= n:
    guess = guess + step
    num_guesses = num_guesses + 1

if abs(guess ** 2 - n) >= epsilon:
    print(f'Failed to find sqrt({n})')
else:
    print(f'{guess} is close to sqrt({n})')
    print(f'Number of guesses: {num_guesses}')

# 4.999000000001688 is close to sqrt(25)
# Number of guesses: 49990

Roughly n/δn/\delta candidates are tested before the neighbourhood of n\sqrt{n} is reached, here 49,99049{,}990 of them. The answer is not 55: 4.9992=24.9900014.999^2 = 24.990001 is within ε=0.01\varepsilon = 0.01 of 2525, which is all that was asked.

Example 3.30 (When the search space misses the answer).

Run the same program with n=0.25n = 0.25:

n = 0.25
epsilon = 0.01
step = 0.0001
guess = 0.0

while abs(guess ** 2 - n) >= epsilon and guess <= n:
    guess = guess + step

if abs(guess ** 2 - n) >= epsilon:
    print(f'Failed to find sqrt({n})')

# Failed to find sqrt(0.25)

The search fails because 0.25=0.5\sqrt{0.25} = 0.5, while the guard guess <= n stops it at 0.250.25. The guard was written for n⩾1n \geqslant 1, where n⩽n\sqrt{n} \leqslant n; for 0<n<10 < n < 1 we have n>n\sqrt{n} > n and the upper bound has to be raised.

Example 3.31 (When the step is too large).

Now take n=123,456n = 123{,}456 with the same δ=0.0001\delta = 0.0001. The program runs a long time and then reports failure: the step carries it over every value within ε\varepsilon of 123456≈351.363\sqrt{123456} \approx 351.363 without ever landing on one. Shrinking δ\delta to 10−610^{-6} repairs that and obliges the program to test some 351,000,000351{,}000{,}000 candidates. Starting nearer the answer would help, and presumes we already know roughly where the answer is.

The step δ\delta controls both the accuracy and the running time, in opposite directions: a smaller step is more accurate and slower.

Exhaustive enumeration does not use whether x^2\hat{x}^2 was too small or too large. Bisection uses that comparison to discard half the remaining candidates at every step.

Looking up a word in a dictionary works the same way: open it near the middle, and if the word comes later than the page shown, discard the first half and open the second half near its middle.

We state the algorithm on an interval, in the notation of the first lesson.

Definition 3.32 (Bisection search).

Let ff preserve order on [ℓ,h][\ell, h], so that a<ba < b implies f(a)<f(b)f(a) < f(b) throughout. Bisection search solves f(x)=yf(x) = y by maintaining the invariant that a solution lies in [ℓ,h][\ell, h] and halving the interval:

  1. compute the midpoint m=(ℓ+h)/2m = (\ell + h)/2;
  2. if f(m)f(m) is too large, replace hh by mm; if f(m)f(m) is too small, replace ℓ\ell by mm;
  3. repeat until ∣f(m)−y∣<ε|f(m) - y| < \varepsilon.

After kk steps the interval has width (h0−ℓ0)/2k(h_0 - \ell_0)/2^{k}, where h0−ℓ0h_0 - \ell_0 is its initial width.

0123450510152025midpoint of the surviving interval
Figure 3.3. The first six steps of a bisection search for 25\sqrt{25} on [0,25][0, 25]. Each bar is the interval still under consideration, the dot on it is the midpoint tested, and the dashed line marks the root at x=5x = 5.

Implementation

n = 25
epsilon = 0.01
low = 0.0
high = max(1.0, n)
guess = (low + high) / 2.0
num_guesses = 0

while abs(guess ** 2 - n) >= epsilon:
    num_guesses = num_guesses + 1
    if guess ** 2 < n:
        low = guess
    else:
        high = guess
    guess = (low + high) / 2.0

print(f'{guess} is close to sqrt({n})')
print(f'Number of guesses: {num_guesses}')

# 5.00030517578125 is close to sqrt(25)
# Number of guesses: 13

Exhaustive enumeration needed about fifty thousand guesses; bisection needs thirteen. The upper end is max(1.0, n) rather than n, so that n\sqrt{n} lies in [0,h][0, h] whether n⩾1n \geqslant 1 or 0<n<10 < n < 1, which repairs the failure on n=0.25n = 0.25.

Example 3.33 (Hand simulating the bisection).

The first four steps for 25\sqrt{25} on [0,25][0, 25]:

Steplowhighguessguess ** 2Action
00.025.012.5156.25too high, high = 12.5
10.012.56.2539.0625too high, high = 6.25
20.06.253.1259.765625too low, low = 3.125
33.1256.254.687521.972656too low, low = 4.6875

Four steps have taken the interval from width 2525 to width 6.25−4.6875=1.56256.25 - 4.6875 = 1.5625, and eight more bring it below 0.010.01.

Example 3.34 (Bisection on a larger input).

Exhaustive approximation failed on n=123,456n = 123{,}456 because no step served: too large and it skipped the root, too small and it needed hundreds of millions of guesses. Bisection starts from [0,123456][0, 123456] and, by the theorem below, needs at most ⌈log⁡2(123456/0.01)⌉=⌈log⁡212,345,600⌉=24\lceil \log_2(123456/0.01)\rceil = \lceil \log_2 12{,}345{,}600 \rceil = 24 guesses to narrow the interval that far. Run, it stops after 3030, the excess coming from the test being on ∣x^2−n∣|\hat{x}^2 - n| rather than on the width of the interval.

Convergence

Bisection produces guesses m0,m1,m2,…m_0, m_1, m_2, \ldots, and each step halves the interval holding the root, so the distance from the guess to the root falls by a factor of two every time. The bound this gives is stated with the logarithm to base 22, written log⁡2N\log_2 N: the power to which 22 must be raised to give NN, so that log⁡28=3\log_2 8 = 3 and log⁡21024=10\log_2 1024 = 10, and log⁡22500≈11.29\log_2 2500 \approx 11.29 for an NN that is not a power of two.

Theorem 3.35 (Convergence of bisection search).

Let [a,b][a, b] be the initial interval and ε>0\varepsilon > 0 the tolerance. Bisection search narrows the interval below ε\varepsilon after at most ⌈log⁡2((b−a)/ε)⌉\lceil \log_2((b-a)/\varepsilon)\rceil steps.

Discussion.

The proof follows the width of the interval. One step replaces the interval by one of its two halves, so the width is halved whichever half survives, and after kk steps it is the initial width divided by 2k2^k. It remains to find the least kk for which this is below ε\varepsilon, by taking log⁡2\log_2 of both sides; the ceiling appears because kk counts steps and must be an integer. The argument resembles the one for decrementing functions, with a width halved at each step in place of a counter decreased by one, and so the number of steps is logarithmic rather than linear.

Proof.

Each step replaces [ℓ,h][\ell, h] by either [ℓ,m][\ell, m] or [m,h][m, h] with mm the midpoint, so the new width is half the old one. After kk steps the width is therefore (b−a)/2k(b-a)/2^{k}, and since the midpoint of an interval of width ww is within w/2w/2 of every point of it, the guess is within (b−a)/2k+1(b-a)/2^{k+1} of the root.

We need (b−a)/2k<ε(b-a)/2^{k} < \varepsilon, that is 2k>(b−a)/ε2^{k} > (b-a)/\varepsilon, that is

k>log⁡2 ⁣(b−aε).k > \log_2\!\left(\frac{b-a}{\varepsilon}\right) .

The least integer meeting this is ⌈log⁡2((b−a)/ε)⌉\lceil \log_2((b-a)/\varepsilon)\rceil.

Example 3.36 (Checking the bound).

For 25\sqrt{25} on [0,25][0, 25] with ε=0.01\varepsilon = 0.01 the theorem gives

⌈log⁡2 ⁣(250.01)⌉=⌈log⁡22500⌉=⌈11.29⌉=12,\left\lceil \log_2\!\left(\frac{25}{0.01}\right) \right\rceil = \lceil \log_2 2500 \rceil = \lceil 11.29 \rceil = 12 ,

and the program used 1313: the test is on ∣x^2−n∣|\hat{x}^2 - n| rather than on the width, and floating-point rounding can add a step.

Remark (Logarithmic against linear).

Exhaustive enumeration with step δ\delta tests about n/δn/\delta candidates; bisection tests about log⁡2(n/ε)\log_2(n/\varepsilon). For n=109n = 10^9 and ε=0.01\varepsilon = 0.01 that is of the order of 101110^{11} guesses against roughly 3737. Doubling the range doubles the work of the first and adds a single step to the second.

Problem 3.22.

Adapt the bisection program to approximate 273\sqrt[3]{27} to within ε=0.001\varepsilon = 0.001. How many guesses does it need? Compare that with the number exhaustive enumeration with step δ=0.001\delta = 0.001 would need, and with the bound of the theorem.

Machine Arithmetic

The methods above assume that arithmetic on reals is exact: that 0.1+0.1+0.10.1 + 0.1 + 0.1 is 0.30.3, that halving an interval halves it, and that the only error is the tolerance we chose. On a machine none of these holds exactly.

Binary Fractions

A computer stores numbers in binary. Just as the decimal system writes fractions with negative powers of ten, so that 0.375=3/10+7/100+5/10000.375 = 3/10 + 7/100 + 5/1000, the binary system uses negative powers of two:

(0.b1b2b3…)2=b12+b24+b38+⋯ .(0.b_1 b_2 b_3 \ldots)_2 = \frac{b_1}{2} + \frac{b_2}{4} + \frac{b_3}{8} + \cdots .

Example 3.37 (Exact binary fractions).

The decimal 0.3750.375 has an exact binary form:

0.375=14+18=0⋅12+1⋅14+1⋅18=(0.011)2.0.375 = \frac{1}{4} + \frac{1}{8} = 0 \cdot \tfrac{1}{2} + 1 \cdot \tfrac{1}{4} + 1 \cdot \tfrac{1}{8} = (0.011)_2 .

Likewise 0.5=(0.1)20.5 = (0.1)_2 and 0.625=1/2+1/8=(0.101)20.625 = 1/2 + 1/8 = (0.101)_2. These are exact because their denominators are powers of two: 0.375=3/80.375 = 3/8, 0.5=1/20.5 = 1/2 and 0.625=5/80.625 = 5/8.

Not every decimal fraction has a finite binary form.

Theorem 3.38 (One tenth is not a finite binary fraction).

The number 1/101/10 has no finite binary representation.

Discussion.

A finite binary fraction has a power of two as its denominator, and the proof compares that with the factor 55 in 1010. Suppose the representation existed with kk bits. Multiplying it through by 2k2^{k} clears every denominator at once and leaves an integer on the right, so the supposed identity becomes 2k=10M2^{k} = 10M for an integer MM. The right-hand side is divisible by 55 and the left is a power of two, and no power of two has 55 among its factors. The same argument gives the general case: a fraction in lowest terms is a finite binary fraction exactly when its denominator is a power of two.

Proof.

Suppose 1/10=(0.b1b2…bk)21/10 = (0.b_1 b_2 \ldots b_k)_2 for some bits b1,…,bkb_1, \ldots, b_k. Then

110=b12+b24+⋯+bk2k=M2k,M=b12k−1+b22k−2+⋯+bk∈N0.\frac{1}{10} = \frac{b_1}{2} + \frac{b_2}{4} + \cdots + \frac{b_k}{2^{k}} = \frac{M}{2^{k}}, \qquad M = b_1 2^{k-1} + b_2 2^{k-2} + \cdots + b_k \in \mathbb{N}_0 .

Cross-multiplying gives 2k=10M2^{k} = 10M, so 55 divides 2k2^{k}. But 2k2^{k} is a product of kk factors of 22, and 55 is a prime different from 22, so 55 divides no power of 22. The supposition is therefore false.

The same holds in general: a rational p/qp/q in lowest terms has a finite binary expansion exactly when qq is a power of two, and since 10=2⋅510 = 2 \cdot 5, the fraction 1/101/10 needs an infinite repeating binary expansion, just as 1/3=0.333…1/3 = 0.333\ldots repeats for ever in decimal.

What This Costs in Practice

Python’s float uses the IEEE 754 double-precision format: 6464 bits carrying a sign, an 1111-bit exponent and a 5252-bit significand with one further bit implied. That is about 1515 to 1717 significant decimal digits.

The consequence is that a decimal constant which looks exact in the source is silently rounded to the nearest representable binary fraction. Putting the format specifier :.20f inside an f-string prints twenty digits after the point and exposes it:

print(0.1)              # 0.1                     (the display is rounded)
print(f'{0.1:.20f}')    # 0.10000000000000000555

print(0.1 + 0.1 + 0.1 == 0.3)     # False
print(f'{0.1 + 0.1 + 0.1:.20f}')  # 0.30000000000000004441
print(f'{0.3:.20f}')              # 0.29999999999999998890

The sum overshoots by about 5.5×10−175.5 \times 10^{-17}, while the literal 0.3 is itself rounded downwards from three tenths. The two errors go opposite ways, so == returns False.

Example 3.39 (A loop that never lands).

Counting from 00 to 11 in steps of a tenth:

x = 0.0
count = 0

while x != 1.0:
    x = x + 0.1
    count = count + 1
    if count > 20:
        print('Gave up')
        break

The loop never stops of its own accord. After ten additions x is 0.9999999999999999, not 1.0; the eleventh takes it to 1.0999999999999999, and it has stepped over 1.0 without touching it. Replacing != by < 1.0, or by abs(x - 1.0) >= epsilon, repairs it.

Remark (Never compare floats with $==$).

Do not test floating-point numbers for exact equality. The expression x == 0.3 is almost certainly wrong even when x was computed by a formula mathematically equal to three tenths. Test instead that the difference is small:

x = 0.1 + 0.1 + 0.1
print(abs(x - 0.3) < 1e-9)   # True

This is an ε\varepsilon-approximation applied to equality; the tolerance should match the precision of the computation, not of the display.

Example 3.40 (Rounding error accumulates).

Adding a tenth a thousand times:

total = 0.0

for i in range(1000):
    total = total + 0.1

print(f'{total:.20f}')     # 99.99999999999859312538
print(total == 100.0)      # False
print(abs(total - 100.0))  # about 1.4e-12

Each addition contributes a rounding error of the order of 10−1710^{-17}, and over a thousand additions they accumulate to about 10−1210^{-12}: small, but enough to make an exact test fail.

Remark (When it matters).

For the bisection search this error is harmless: an approximate answer was wanted, and ε\varepsilon is much larger than the rounding. The problem arises in code that assumes exact arithmetic, by testing x == 0.0 where it should test abs(x) < epsilon, or by expecting a counter incremented by 0.1 to arrive exactly at 1.0 after ten steps.

Problem 3.23.

The theorem above rules out 1/101/10, and the paragraph following it makes the general claim: a rational p/qp/q in lowest terms has a finite binary expansion exactly when qq is a power of two. Prove both directions, following the argument of the theorem. Then say how many bits the expansion of p/2kp/2^{k} needs when pp is odd, and give the expansions of 7/167/16 and 5/65/6 as far as each can be written.

Problem 3.24.

The loop of the example above adds 0.10.1 a thousand times and lands about 1.4×10−121.4 \times 10^{-12} short of 100100.

  1. Write a second program that adds the integer 11 a thousand times and divides by ten at the end, and compare its result with 100.0 using ==.
  2. Say why the second is exact where the first is not, in terms of which numbers have a finite binary expansion.
  3. A sum of money is to be accumulated over many transactions, each an exact number of pounds and pence. Say which of the two arrangements should be used, and what the other would cost after a million transactions.

Exercises on Types and Expressions

Exercise 3.1.

Give the type and the value of each expression.

  1. 7 / 2;
  2. 7 // 2;
  3. 7 % 2;
  4. 7.0 // 2;
  5. 2 ** 0.5;
  6. 1 == 1.0.

Exercise 3.2.

State the value of 2 ** 3 ** 2, (2 ** 3) ** 2, -2 ** 2 and (-2) ** 2, and say which rule of precedence or associativity settles each one.

Exercise 3.3.

Predict the value of 0.1 + 0.2 == 0.3 and of 0.5 + 0.25 == 0.75, then check both. Explain why the two come out differently, and give another pair of float values whose sum can be tested exactly.

Exercise 3.4.

Let n be a positive integer.

  1. Write an expression for its last two decimal digits.
  2. Write an expression for the digit in its hundreds place.
  3. Write an expression that is True exactly when n is a multiple of 33 but not of 99.

Exercise 3.5.

Give the value of int('101', b) for b=2,3,8b = 2, 3, 8 and 1616, and find the base bb for which it equals 122122.

Exercise 3.6.

Explain why x != 0 and 100 % x == 0 may be evaluated for any integer x, while 100 % x == 0 and x != 0 may not.

Exercise 3.7.

Explain what a, b = b, a does, and why performing a = b and then b = a does not do the same. Say what the second pair leaves in a and b.

Exercises on Branching

Exercise 3.8.

Write a program that reads three integers and prints them in increasing order, using conditional statements only and no built-in sorting.

Exercise 3.9.

A year is a leap year when it is divisible by 44, except that centuries are not, except that those divisible by 400400 are. Write a program that reads a year and reports whether it is a leap year, and check it on 19001900, 20002000, 20232023 and 20242024.

Exercise 3.10.

Write a program that reads a real number x and prints which of the intervals (−∞,−1)(-\infty, -1), [−1,0)[-1, 0), [0,1][0, 1] and (1,∞)(1, \infty) contains it.

Exercise 3.11.

Write a program that reads three numbers a, b and c and prints how many of them are strictly positive. Use branching only, with no loop and no arithmetic on the results of the tests.

Exercise 3.12.

The expressions 1/x if x != 0 else 0 and (x != 0) * (1/x) agree for every non-zero x. Say what each does when x is 0, and explain the difference in terms of which arms of a case split get evaluated.

Exercise 3.13.

Explain why a branching program of NN statements performs at most NN atomic operations on any input, and give a three-statement program whose count of operations depends on its input.

Exercise 3.14.

Write a program that reads three positive reals and reports whether they can be the side lengths of a triangle, that is, whether each of them is smaller than the sum of the other two. Use conditional statements only.

Exercises on Iteration

Exercise 3.15.

Write a program that computes n!n! for a given n⩾0n \geqslant 0 with a for loop, and a second that does the same with a while loop. State how many multiplications each performs, and say what each returns for n=0n = 0.

Exercise 3.16.

Write a program that counts the digits of nn in base bb with a loop, for n∈Nn \in \mathbb{N} and b>1b > 1. Check its count against Corollary 2.22 for several nn and bb.

Exercise 3.17.

Write a program that sums the decimal digits of a positive integer using // and % only, with no strings. Extend it to repeat the process on the result until a single digit is left, and compare that digit with n mod 9n \bmod 9.

Exercise 3.18.

Write a program that prints the first nn Fibonacci numbers, using a single multiple assignment inside the loop to advance the pair.

Exercise 3.19.

Write a program that takes positive integers aa and bb and repeatedly replaces the pair by the smaller number and the remainder of the larger on division by it, stopping when the remainder is 00. Hand simulate the loop on (48,18)(48, 18), and identify the surviving number as the largest integer dividing both aa and bb.

Exercise 3.20.

Write a program that prints every pair (a,b)(a, b) with 1⩽a<b⩽N1 \leqslant a < b \leqslant N and a∣ba \mid b, for a given NN. Say how many divisibility tests it performs as a function of NN.

Exercise 3.21.

Modify the trial-division program so that, when nn is composite, it also prints the least divisor of nn above 11. Run it on three numbers near one million of your own choosing, and say which of them are prime.

Exercise 3.22.

Write a program that converts a string of base-bb digits to an integer with a loop, for a given b>1b > 1, using Horner’s scheme rather than forming any power of bb. Count the multiplications, and check the result against int.

Exercise 3.23.

Write a program that runs the loop replacing nn by n/2n/2 when nn is even and by 3n+13n + 1 when nn is odd, stopping when nn reaches 11, and counts its passes. Run it on each starting value from 11 to 3030, and report which takes the most.

Exercises on Running Time

Exercise 3.24.

Show directly from the definition that ∑k=1nk=O(n2)\sum_{k=1}^{n} k = O(n^2), giving an explicit α\alpha and n0n_0, and that n2=O(∑k=1nk)n^2 = O\bigl(\sum_{k=1}^{n} k\bigr) as well, so that each of the two is OO of the other. Then do both again with the limit form.

Exercise 3.25.

Decide which of the following hold, with a proof or a counterexample in each case.

  1. 2n+1=O(2n)2^{n+1} = O(2^n);
  2. 22n=O(2n)2^{2n} = O(2^n);
  3. log⁡2n=O(log⁡10n)\log_2 n = O(\log_{10} n);
  4. nlog⁡2n=O(n2)n\log_2 n = O(n^2).

Exercise 3.26.

Suppose f1=O(g)f_1 = O(g) and f2=O(g)f_2 = O(g). Prove that f1+f2=O(g)f_1 + f_2 = O(g) and that c f1=O(g)c\,f_1 = O(g) for every constant c>0c > 0. Then give functions with f1=O(g)f_1 = O(g) and f2=O(g)f_2 = O(g) for which f1f2=O(g)f_1 f_2 = O(g) fails, so that OO is closed under sums and constant multiples but not under products.

Exercise 3.27.

An input of nn decimal digits denotes a number of size about 10n10^n. Express the running time of trial division as a function of the number of digits of its input rather than of the number itself, and say how many digits a number may have before an algorithm performing 10910^9 operations a second needs more than an hour.

Exercise 3.28.

Count the elementary operations performed by the schoolbook algorithms on two nn-digit numbers exactly, rather than up to OO: give the number of single-column additions performed by the addition, and the number of single-digit multiplications performed by the multiplication.

Exercises on Search and Approximation

Exercise 3.29.

Give a decrementing function for each of the following loops and state the bound on the number of passes it yields.

  1. while n >= 10: n = n // 10, for n∈Nn \in \mathbb{N};
  2. while a != b: a, b = (a - b, b) if a > b else (a, b - a), for positive integers aa and bb;
  3. the exhaustive square-root search of this chapter, whose body is guess = guess + step.

Exercise 3.30.

The positions of a string are numbered from 00. Write a program that prints the characters of a string my_str at the even positions, so that 'abcdefg' produces aceg. Write it twice, once with range and indexing, and once with a for loop over the characters and a counter.

Exercise 3.31.

Let NN be an integer with 0⩽N⩽10000 \leqslant N \leqslant 1000. Write a program that finds NN by bisection search, printing the number of guesses it needed and the value found. When the midpoint of the current interval falls between two integers, take the smaller.

Exercise 3.32.

A positive integer nn is a perfect power if n=rpn = r^{p} for integers r⩾1r \geqslant 1 and p⩾2p \geqslant 2. Write a program that reads nn and prints integers root and pwr with 1<pwr<61 < \text{pwr} < 6 and rootpwr=n\text{root}^{\text{pwr}} = n, or reports that no such pair exists. Test it on 6464, which admits 828^2, 434^3 and 262^6, on 2727, and on 7272.

Exercise 3.33.

Write a program that approximates log⁡2x\log_2 x for a positive real xx to within ε=0.001\varepsilon = 0.001 by bisection. It must first find an interval [L,H][L, H] containing log⁡2x\log_2 x: note that 20=12^0 = 1, that 2k2^{k} grows without bound, and that 2−k2^{-k} becomes arbitrarily small, so LL may have to be negative. Test it on x=1x = 1, x=32x = 32 and x=0.1x = 0.1.

Exercise 3.34.

Run the exhaustive square-root search and the bisection search on the same nn for ε=10−2,10−4\varepsilon = 10^{-2}, 10^{-4} and 10−610^{-6}, recording the number of guesses each makes. Say which of the two counts grows with 1/ε1/\varepsilon and which with log⁡(1/ε)\log(1/\varepsilon), and check the readings against the two bounds.

Exercise 3.35.

Find two float values a and b for which (a + b) + c and a + (b + c) differ for some c, and explain the difference in terms of rounding. What does this say about summing a list of numbers in different orders?

Exercise 3.36.

Which of the decimals 0.50.5, 0.20.2, 0.250.25, 0.30.3 and 0.1250.125 are exact as float values? Predict the answer from the theorem on binary fractions before checking each with :.20f.

An Applied Exercise

Exercise 3.37.

A band puts N=5000N = 5000 tickets on sale for one night and prices them dynamically. The first ticket costs p0=£40.00p_0 = \pounds 40.00. The tickets are sold in blocks of k=250k = 250, and after each block the price is raised by r=8%r = 8\% of the current price. The band will not charge more than a cap of C=£120.00C = \pounds 120.00: once a rise would take the price above the cap, the price is set to the cap and stays there for every remaining block.

  1. Write a program that sells all NN tickets under this rule and prints the price of each block together with the total revenue. Report the revenue.
  2. Give, in terms of p0p_0, rr and CC, the number of rises after which the uncapped price would first exceed the cap, and hence the number of the first block sold at the cap. Check your formula against the output of your program.
  3. Consider the sequence of price increases between consecutive blocks. Show that before the cap binds these form a geometric sequence and give its ratio. Exactly one increase belongs to neither the geometric stretch nor the capped stretch: identify it, give its value, and say what it would have been without the cap.
  4. Show that once the cap binds the total revenue is a linear function of the number of tickets sold, and give its slope. Say what the revenue would look like as a function of NN if the cap were removed.
  5. Give the running time of your program in OO notation as a function of NN and kk, and say which of the two the running time really depends on.
  6. Prices in pounds are not exact float values. Rewrite the simulation in integer pence, rounding each new price down to the nearest penny, and compare the two totals. The two disagree for two separate reasons; say which reason accounts for most of the gap, and which of the two totals the band should quote.

Check Yourself

 

Fresh questions on the whole lesson — none of them is worked out above. Work each one out on paper or in your head before opening Python; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 3.38.

What is the type of 5 % 2 == 1?

answer one of these

Exercise 3.39.

What is the value of 7 // -2?

answer one of these

Exercise 3.40.

What is the value of -7 % 3?

answer one of these

Exercise 3.41.

What is the value of 9 // 2 ** 2?

answer one of these

Exercise 3.42.

After total = 0 and then total += 5 twice, what is total?

answer one of these

Exercise 3.43.

What is the value of 2 + 3 * 4 ** 2?

answer one of these

Exercise 3.44.

After x, y = 1, 2 and then x, y = y, x + y, what is y?

answer one of these

Exercise 3.45.

After x = 5, y = x and x = 7, what is y?

answer one of these

Exercise 3.46.

How many values does range(5, 20, 4) produce?

answer one of these

Exercise 3.47.

What is the last value produced by range(10, 3, -2)?

answer one of these

Exercise 3.48.

In a for i in range(3) loop whose body is a for j in range(i) loop, how many times does the inner body run in total?

answer one of these

Exercise 3.49.

Starting from n = 4096, how many times does the body of while n >= 10: n = n // 10 run?

answer one of these

Exercise 3.50.

Testing n=200n = 200 for primality by trial division, what is the largest candidate divisor that has to be tried?

answer one of these

Exercise 3.51.

What is the value of int('ff', 16)?

answer one of these

Exercise 3.52.

What are the binary digits of 2020?

answer one of these

Exercise 3.53.

Testing whether nn is prime by trial division up to ⌊n⌋\lfloor\sqrt{n}\rfloor has which running time?

answer one of these

Exercise 3.54.

Multiplying two nn-digit numbers by the schoolbook algorithm has which running time?

answer one of these

Exercise 3.55.

How many bisection steps are needed to narrow [0,64][0, 64] to a width below 11?

answer one of these

Exercise 3.56.

Starting from x = 0, how many times does the body of while x ** 3 < 64: x = x + 1 run?

answer one of these

Exercise 3.57.

What is the value of (3 > 2) or (1 / 0 == 0)?

answer one of these

Lesson 4

Data Structures and Graphs

Taught

Data Structures

Every program of the last lesson held a bounded number of values at once: a guess, a counter, a running total, the string being built up. A program that has to keep many values, for example to sort a list of numbers or to remember which candidates have been ruled out, needs somewhere to store them and a way to reach any one of them. A data structure is such a place together with the operations for using it.

We compare algorithms by explicit models of cost, and refine an implementation only once the broad algorithmic choice has been made. For every algorithm there are two things to prove: that it terminates, and that it is correct, meaning that the output it returns is a correct output for its input in the sense of a computational problem.

Functions

From here on an algorithm is usually written as a function: a named piece of code that takes inputs, called its arguments, and returns an output.

def square(x):
    return x * x

print(square(7))       # 49
print(square(-3) + 1)  # 10

The line def square(x): names the function and its parameter x; the indented body is its code. A call square(7) binds the parameter x to the argument 7, runs the body, and evaluates to the value named by return, which ends the call at once. A function whose body finishes without reaching a return returns None. Names bound inside the body are local: they exist only while that call runs and do not disturb names of the same spelling elsewhere.

def smallest(a, b, c):
    target = a
    if b < target:
        target = b
    if c < target:
        target = c
    return target

target = 100
print(smallest(4, -2, 9))   # -2
print(target)              # 100, untouched by the call

A function may return several values at once as a tuple, and the caller may unpack them by multiple assignment: return p, q and then p, q = f(p, q).

In pseudocode we give a function a name and write its arguments in brackets, as in Fact(n), and return has its Python meaning.

Recursion

A function may call itself. The call f(n) then waits while f(n - 1) runs, and that call has its own local names, separate from those of the call that made it. The part of memory that keeps track of this is organised as a stack: each call places a frame holding its local names on top, and returning removes the top frame, so the most recent call is always the first to finish. The number of frames present at once is the depth of the recursion.

Recursion needs a case that does not call the function again, reached after finitely many calls; otherwise the stack grows until Python stops the program with a RecursionError.

def count_down(n):
    if n == 0:
        print('done')
    else:
        print(n)
        count_down(n - 1)

count_down(3)   # 3, 2, 1, done

Problem 4.1.

Write a function digit_sum(n) that returns the sum of the decimal digits of a positive integer nn, once with a while loop and once recursively, using // and % only. State the depth of the recursion in terms of nn.

Induction and Loop Invariants

For a statement P(n)P(n) about every nonnegative integer, proof by induction has two parts: prove P(0)P(0), then assume P(n)P(n) for an arbitrary nn and use that assumption to prove P(n+1)P(n+1). If the statement begins at n=1n = 1, start with P(1)P(1). The base case gives P(0)P(0), and the step then gives P(1)P(1), P(2)P(2) and so on. The last lesson used this pattern for the closed forms of recurrences; we now use it for recursive algorithms and for loops.

Definition 4.1 (Loop invariant).

A loop invariant is a statement about the program state that

  1. holds before the first iteration (initialisation),
  2. is preserved by every iteration (maintenance), and
  3. gives the desired conclusion when the loop stops (use at termination).

An invariant plays the part of the induction hypothesis: initialisation is the base case, and maintenance is the step. It says nothing about whether the loop stops. That is proved separately by a decrementing function, often called a variant in this context.

Theorem 4.2 (Chocolate-bar strategy).

A rectangular chocolate bar has a marked corner square. A move removes a nonempty strip by one horizontal or vertical cut, retaining the piece containing the marked square. The player who receives the 1×11 \times 1 bar loses. The first player has a winning strategy exactly when the starting rectangle is not a square.

Discussion.

Write (p,q)(p, q) for the current numbers of rows and columns. The strategy is to hand the opponent a square every time: from a nonsquare, cut the longer side down to the shorter. The invariant is the pair of facts that the strategy user always receives a nonsquare and always hands over a square; initialisation is the starting position, and maintenance holds because every legal move from a square produces a nonsquare. The product pqpq falls with every move, so it serves as the variant and play reaches (1,1)(1, 1). Since (1,1)(1, 1) is a square, the strategy user never receives it. The two cases of the statement say which player can use the strategy.

Proof.

Write (p,q)(p, q) for the current positive numbers of rows and columns. If p≠qp \neq q, the player to move can cut the larger coordinate down to the smaller one and hand the other player a square. Every legal move from a square changes exactly one coordinate, so it makes the coordinates unequal; the next player can therefore make a square again. The invariant is that the player using this strategy always receives a nonsquare rectangle and always hands over a square.

Every move strictly reduces the positive integer pqpq, so play eventually reaches (1,1)(1, 1). The strategy user cannot receive (1,1)(1, 1), since that state is square. If the initial rectangle is nonsquare, Player One uses the strategy; if it is square, Player Two uses it after Player One’s first cut.

Example 4.3 (A game from (5,3)(5, 3)).

Starting from (5,3)(5, 3), Player One cuts to (3,3)(3, 3). If the opponent cuts to (3,1)(3, 1), Player One cuts to (1,1)(1, 1) and the opponent loses.

One full round can be implemented as a loop: check whether Player One has received (1,1)(1, 1), let Player One move, check whether the opponent has received (1,1)(1, 1), then let the opponent move. Player One’s move is the square-producing cut:

Player One's Move
Input:  the numbers p, q of rows and columns of the retained rectangle.
Output: the rectangle after the move.

    if p > q then p ≝ q
    else if q > p then q ≝ p
    else p ≝ p − 1
    return p, q
def player(p, q):
    if p > q:
        p = q
    elif q > p:
        q = p
    else:
        p = p - 1   # reached only from a square
    return p, q

For a nonsquare starting position, the last branch is never reached under the winning strategy. The opponent may choose any legal cut; the proof above covers all such choices.

Recursive and Iterative Factorials

Recursive and iterative factorial functions show the connection between induction and invariants. Define 0!=10! = 1 and n!=n⋅(n−1)!n! = n \cdot (n-1)! for n>0n > 0.

Recursive Factorial                    Iterative Factorial
Input:  n ∈ N₀.                        Input:  n ∈ N₀.
Output: n!.                            Output: n!.

    Fact(n):                               r ≝ 1
        if n = 0 then return 1             for j ≝ 1 to n do r ≝ r · j
        else return n · Fact(n − 1)        return r

The recursive version returns 11 at n=0n = 0 and otherwise returns n⋅Fact(n−1)n \cdot \mathrm{Fact}(n-1); induction on nn proves it correct, the base case being the first branch and the step the second. The iterative version begins with r=1r = 1 and multiplies rr by jj for j=1,…,nj = 1, \ldots, n. Before iteration jj the invariant is r=(j−1)!r = (j-1)!; after the multiplication it becomes r=j!r = j!, so at exit r=n!r = n!. The variant n−j+1n - j + 1 decreases to zero.

def fact_recursive(n):
    if n == 0:
        return 1
    return n * fact_recursive(n - 1)

def fact_iterative(n):
    r = 1
    for j in range(1, n + 1):
        r = r * j
    return r

print(fact_recursive(10), fact_iterative(10))   # 3628800 3628800

The recursive version has depth n+1n + 1: the call for nn waits on the call for n−1n - 1, down to 00. The iterative version uses a single frame. Python integers have no fixed size, so both return n!n! exactly however large it is. A product of large integers is then not an elementary operation, and costs what the schoolbook bound says.

Problem 4.2.

The loop below is meant to compute ana^n for an integer aa and n∈N0n \in \mathbb{N}_0 with about log⁡2n\log_2 n multiplications.

def power(a, n):
    result, base, e = 1, a, n
    while e > 0:
        if e % 2 == 1:
            result = result * base
        base = base * base
        e = e // 2
    return result

Show that result⋅base e=an\text{result} \cdot \text{base}^{\,e} = a^n is a loop invariant, give a variant, and deduce that power is correct. How many multiplications does it perform, in terms of the binary expansion of nn?

Problem 4.3.

In the chocolate-bar game a move may instead remove a strip of width one or two only. Decide, for each starting rectangle (p,q)(p, q) with 1⩽p,q⩽41 \leqslant p, q \leqslant 4, which player has a winning strategy, and state and prove a rule covering every (p,q)(p, q).

Measuring Cost

A program can spend most of its running time in a small part of its code. Choosing a better algorithm for that part usually does more than rewriting individual instructions, and much of the low-level rewriting is done automatically in any case by the software that translates a program into machine instructions. Costs other than time and memory can matter too, energy and money among them. For example, a “Penny Sort” measure asks how many externally stored records can be sorted for a US cent, after the purchase cost of a fixed machine is spread over an assumed working life. That measure combines hardware cost and algorithm throughput.

Example 4.4 (Which method is faster depends on nn).

Suppose method A uses exactly 100n100n operations on an input of size nn, while method B uses exactly n2n^2. At n=10n = 10, A uses 10001000 operations and B uses 100100; at n=1000n = 1000, A uses 100,000100{,}000 and B uses 1,000,0001{,}000{,}000. The method with fewer operations depends on the input size, and the two cross at n=100n = 100.

The running time of the last lesson counted elementary operations. We also measure memory.

Definition 4.5 (Time cost and auxiliary space).

The time cost of an algorithm counts operations in a stated model. Its auxiliary space, or memory footprint, is the maximum amount of extra storage present at any instant of a run, excluding the input.

Storage may be released and reused later, which is why the footprint is a maximum and not a cumulative total. In a conventional model allocating or touching a memory cell costs at least one operation, so an algorithm’s footprint cannot exceed a constant multiple of its running time. This claim depends on the model, and it does not say that every algorithm uses as much space as time.

The unit-cost assumption has to be stated with care. A comparison of two strings is not necessarily a single operation: it may inspect many characters. A comparison of two integers of bounded size is commonly counted as one operation, but for arbitrarily large integers or variable-length records the cost must be stated.

Lower and Two-Sided Bounds

Landau’s OO from the last lesson is an upper bound. It has a lower counterpart, and the two together give a two-sided bound.

Definition 4.6 (Ω\Omega and Θ\Theta).

Let f,g:N→R⩾0f, g : \mathbb{N} \to \mathbb{R}_{\geqslant 0}. We write

f=Ω(g)if there are c>0 and n0 with c g(n)⩽f(n) for all n⩾n0,f=Θ(g)if f=O(g) and f=Ω(g).\begin{aligned} f &= \Omega(g) &&\text{if there are } c > 0 \text{ and } n_0 \text{ with } c\,g(n) \leqslant f(n) \text{ for all } n \geqslant n_0, \\ f &= \Theta(g) &&\text{if } f = O(g) \text{ and } f = \Omega(g). \end{aligned}

As with OO, the equals sign is shorthand for membership of a set of functions, not equality of functions. The last lesson wrote f=o(g)f = o(g), read ”ff grows strictly more slowly than gg”, when f(n)/g(n)f(n)/g(n) tends to zero. Written out without limits: g(n)>0g(n) > 0 from some point on, and for every ε>0\varepsilon > 0 there is an n0n_0 with f(n)⩽ε g(n)f(n) \leqslant \varepsilon\, g(n) for every n⩾n0n \geqslant n_0. This is stronger than f=O(g)f = O(g), since the constant multiplier can be made as small as we wish by taking nn large enough.

The growth rates met most often are, from slowest to fastest,

log⁡n,n,nlog⁡n,n3/2,n2,n3,2n.\log n, \quad n, \quad n \log n, \quad n^{3/2}, \quad n^2, \quad n^3, \quad 2^n .

Logarithm bases differ only by a constant factor, since log⁡an=log⁡bn/log⁡ba\log_a n = \log_b n / \log_b a by the change of base; so Θ(log⁡2n)\Theta(\log_2 n) and Θ(log⁡10n)\Theta(\log_{10} n) are the same class and the base is usually left off.

Problem 4.4.

Show that 12n2−3n=Θ(n2)\tfrac12 n^2 - 3n = \Theta(n^2) by exhibiting the constants, and that nlog⁡2n=o(n2)n \log_2 n = o(n^2).

Lower Bounds

The comparison model of computation acts on a set of comparable objects. The objects are treated as black boxes supporting only binary tests called comparisons, namely <<, ⩽\leqslant, >>, ⩾\geqslant, == and ≠\neq: each takes two objects and returns True or False according to their relative order. Nothing else about the objects may be inspected, so an algorithm in this model learns about its input only through the answers to comparisons, and its cost is counted as the number of comparisons it makes.

Definition 4.7 (Problem lower bound and optimality).

A problem lower bound is a bound that applies to every algorithm solving the problem in a specified model. An algorithm is asymptotically optimal for a cost measure when its upper bound matches the problem’s lower bound up to constants.

An upper bound is about one algorithm, and a lower bound is about every algorithm for the problem. A lower bound therefore has to fix a model, which lists the operations an algorithm may use.

Proposition 4.8 (Finding a maximum).

Finding the position of the maximum among nn pairwise distinct, otherwise unordered comparable elements requires at least n−1n - 1 comparisons in the comparison model, and n−1n - 1 comparisons suffice.

Discussion.

For the lower bound we count losses: an element can be ruled out as the maximum only once it has been seen to be smaller than something, and one comparison produces exactly one loser. All n−1n - 1 non-maximal elements must be ruled out, so at least n−1n - 1 comparisons are needed. For the upper bound a single left-to-right scan with a running maximum makes one comparison per element after the first, and its invariant says that the stored index is maximal in the prefix scanned so far.

Proof.

Before an element can be ruled out as the maximum, it must lose a comparison against a larger element. One comparison can give a first loss to at most one candidate. All but the true maximum, namely n−1n - 1 candidates, must be ruled out, so at least n−1n - 1 comparisons are necessary.

A scan keeps the index of the largest element seen and compares each of the remaining n−1n - 1 elements with it once; its invariant is that the stored index is maximal in the scanned prefix. Thus it meets the lower bound.

A second proof uses an adversary, an imagined opponent who may change the input as long as every answer already given stays true.

Proof.

Fix an input of distinct values T[0],…,T[n−1]T[0], \ldots, T[n-1] and a nonmaximum element T[j]T[j]. Suppose T[j]T[j] is never compared with an element larger than itself. Change only its value, to one larger than every original value. Every comparison involving T[j]T[j] previously had a smaller other operand, so its outcome stays the same; all other comparisons are unchanged. A deterministic comparison algorithm therefore follows the same sequence of branches and returns the same index. But jj is now the true maximum, a contradiction. Thus each nonmaximum element must lose a comparison, which gives n−1n - 1 losses.

The scan in pseudocode and in Python:

Position of the Maximum
Input:  a sequence T[0], …, T[n − 1] of comparable elements, n ⩾ 1.
Output: an index k with T[k] maximal.

    k ≝ 0
    for i ≝ 1 to n − 1 do
        if T[i] > T[k] then k ≝ i
    return k
def arg_max(T):
    k = 0
    for i in range(1, len(T)):
        if T[i] > T[k]:
            k = i
    return k

print(arg_max((3, 9, 2, 9, 4)))   # 1

Python’s len(T) gives the length of a tuple, as it does for a string, and T[i] its entry at position i.

Problem 4.5.

Give an algorithm that finds both the maximum and the minimum of nn distinct elements with at most ⌈3n/2⌉−2\lceil 3n/2 \rceil - 2 comparisons. Then use an adversary to show that at least ⌈3n/2⌉−2\lceil 3n/2 \rceil - 2 comparisons are necessary.

The Word-RAM Model

To calculate the resources an algorithm uses we need to say how long a computer takes to perform basic operations. Fixing such a set of operations gives a model of computation, on which the analysis is then based. We use the ww-bit Word-RAM model, which treats a computer as a random-access array of machine words called memory, together with a processor that performs operations on that memory.

Definition 4.9 (Word-RAM).

A machine word is a sequence of ww bits, read as an integer in {0,…,2w−1}\{0, \ldots, 2^w - 1\}. A Word-RAM processor performs each of the following in constant time:

  1. addition, subtraction, multiplication, integer division, remainder, bitwise operations and comparisons of two machine words;
  2. given a word aa, reading or writing the word stored in memory at address aa.

A machine word of ww bits can name at most 2w2^w addresses, so the processor can read and write at most 2w2^w locations of memory. When a problem’s input occupies nn machine words we therefore always assume a word size of w>log⁡2nw > \log_2 n bits, or the machine could not reach all of its input. For comparison, a Word-RAM model of a byte-addressable 6464-bit machine allows inputs of up to about 101010^{10} gigabytes.

Arrays

An array is a fixed number of storage slots in a row, numbered from 00, any of which may be read or written in a single elementary operation. The iith slot of an array pp is written p[i]p[i]. In the Word-RAM an array of nn words is a block of nn consecutive addresses, and reaching p[i]p[i] means reading the address of p[0]p[0] plus ii: one addition and one read, whatever ii is and whatever the length of pp.

Python Lists

Python’s counterpart of an array is a list, written in square brackets. A list is a sequence like a tuple, but its entries may be changed.

p = [3, 1, 4, 1, 5]
print(len(p), p[0], p[4])   # 5 3 5
p[1] = 9
print(p)                    # [3, 9, 4, 1, 5]

q = [True] * 4              # [True, True, True, True]
r = [None] * 3              # [None, None, None]

Positions are numbered from 00, as they are for strings, and p[i] = x writes to position i. The expression [x] * n builds a list of n copies of x. Several further operations will be used below.

  1. p.append(x) adds x at the end, and p.pop() removes the last entry and returns it; p.pop(i) removes and returns the entry at position i.
  2. The slice p[i:j] is a new list holding the entries at positions i up to but not including j, with the same convention as for strings; p[i:] runs to the end and p[:j] starts at the beginning. Building a slice copies its j−ij - i entries.
  3. p + q is a new list holding the entries of p followed by those of q.
  4. The comprehension [f(a) for a in X] builds the list of values f(a) as a runs through X.
  5. x is None tests whether x is the object None.

Unlike a tuple, a list can grow and shrink; how Python does this is described under dynamic arrays below. Used with a fixed length it behaves as an array.

Listing the Primes

Deciding whether one number is prime was a decision problem; listing the primes up to a bound is a general discrete computational problem. It is the first problem here whose best algorithm uses an array, and it has a much faster algorithm than testing each number for primality.

List of Prime Numbers
Input: n ∈ N.
Task:  compute all prime numbers p with p ⩽ n.

Running the trial-division test on each of 2,…,n2, \ldots, n in turn settles it in O(nn)O(n\sqrt{n}) operations. The sieve of Eratosthenes does better by never testing divisibility at all: it writes down every candidate, then crosses out the multiples of each survivor in turn. It uses an array pp indexed by 0,…,n0, \ldots, n.

Example 4.10 (The sieve of Eratosthenes).

The algorithm marks every index as a candidate and then strikes out the multiples of each index that is still marked.

The Sieve of Eratosthenes
Input:  n ∈ N.
Output: all prime numbers less than or equal to n.

    for i ≝ 2 to n do p[i] ≝ "yes"
    for i ≝ 2 to n do
        if p[i] = "yes" then
            output i
            for j ≝ i to ⌊n / i⌋ do p[i · j] ≝ "no"

The inner loop starts at j=ij = i rather than at j=2j = 2: the multiples i⋅2,…,i⋅(i−1)i \cdot 2, \ldots, i \cdot (i-1) have a factor smaller than ii and were struck out already.

The running time depends on the total work of the inner loops, which is a sum of harmonic numbers, which have no closed form. What we need instead is a bound on them. The lower bound is not needed for the sieve; it is used for quicksort in the next lesson.

Proposition 4.11 (Bounds on the harmonic numbers).

For every n∈Nn \in \mathbb{N},

12⌊log⁡2n⌋  ⩽  Hn  ⩽  1+log⁡2n.\tfrac12 \lfloor \log_2 n \rfloor \;\leqslant\; H_n \;\leqslant\; 1 + \log_2 n .

Discussion.

We group the terms into blocks between consecutive powers of two. The block running from k=2tk = 2^t to k=2t+1−1k = 2^{t+1}-1 has 2t2^t terms, each at most 1/2t1/2^t and more than 1/2t+11/2^{t+1}, so the block contributes between 12\tfrac12 and 11 whatever tt is. The number of blocks needed to cover 1,…,n1, \ldots, n is about log⁡2n\log_2 n, and both bounds follow. At the top the last block may be incomplete: for the upper bound we enlarge the sum to the end of that block, and for the lower bound we drop it.

Proof.

Let m=⌊log⁡2n⌋m = \lfloor \log_2 n \rfloor, so that 2m⩽n<2m+12^m \leqslant n < 2^{m+1}. For t⩾0t \geqslant 0 the block of indices 2t⩽k⩽2t+1−12^t \leqslant k \leqslant 2^{t+1} - 1 has 2t+1−2t=2t2^{t+1} - 2^t = 2^t terms, and each has 2t⩽k<2t+12^t \leqslant k < 2^{t+1}, so

12=2t⋅12t+1  ⩽  ∑k=2t2t+1−11k  ⩽  2t⋅12t=1.\frac12 = 2^t \cdot \frac{1}{2^{t+1}} \;\leqslant\; \sum_{k=2^{t}}^{2^{t+1}-1} \frac{1}{k} \;\leqslant\; 2^t \cdot \frac{1}{2^t} = 1 .

Every term is positive, so enlarging the range of summation increases the sum and shrinking it decreases the sum. The blocks t=0,…,mt = 0, \ldots, m cover 1,…,2m+1−1⩾n1, \ldots, 2^{m+1} - 1 \geqslant n, and the blocks t=0,…,m−1t = 0, \ldots, m - 1 cover 1,…,2m−1⩽n1, \ldots, 2^m - 1 \leqslant n. Hence

Hn  ⩽  H2m+1−1=∑t=0m  ∑k=2t2t+1−11k  ⩽  m+1  ⩽  log⁡2n+1,Hn  ⩾  H2m−1=∑t=0m−1  ∑k=2t2t+1−11k  ⩾  m2.H_n \;\leqslant\; H_{2^{m+1}-1} = \sum_{t=0}^{m} \; \sum_{k=2^{t}}^{2^{t+1}-1} \frac{1}{k} \;\leqslant\; m + 1 \;\leqslant\; \log_2 n + 1, \qquad H_n \;\geqslant\; H_{2^{m}-1} = \sum_{t=0}^{m-1} \; \sum_{k=2^{t}}^{2^{t+1}-1} \frac{1}{k} \;\geqslant\; \frac{m}{2} .

Theorem 4.12 (The sieve is correct and runs in O(nlog⁡n)O(n \log n)).

The sieve of Eratosthenes outputs exactly the primes less than or equal to nn, and performs O(nlog⁡n)O(n \log n) elementary operations.

Discussion.

Correctness and running time are proved separately.

Correctness has two directions. No prime is ever struck out, because an entry is written to only as p[i⋅j]p[i \cdot j] with both ii and jj at least 22, and such an index is composite by definition. In the other direction every composite kk must be struck before the outer loop reaches it, and the index that strikes it is its least divisor above 11: that divisor is itself prime, so it still carries “yes” when the outer loop arrives at it, and its partner kk divided by it is large enough to fall in the range the inner loop covers. That the partner is at least the divisor is Proposition 3.14, which is why the inner loop may start at j=ij = i.

For the running time, the outer loop costs O(n)O(n) by itself, and the inner loop belonging to ii runs at most n/in/i times. Summing n/in/i over ii gives nn times a harmonic sum, and the upper bound just proved turns that into nlog⁡2nn\log_2 n.

Proof.

Correctness. An entry of pp is set to “no” only in the inner loop, where the index written to is i⋅ji \cdot j with i⩾2i \geqslant 2 and j⩾i⩾2j \geqslant i \geqslant 2. Such an index is a product of two integers greater than 11 and so is composite; hence no prime is ever struck out, and every prime ⩽n\leqslant n still carries “yes” when the outer loop reaches it and is output.

Conversely let k⩽nk \leqslant n be composite, and let ii be its least divisor with i⩾2i \geqslant 2. Then ii is prime: a divisor dd of ii with 1<d<i1 < d < i would divide kk as well and contradict minimality. By Proposition 3.14, i⩽ki \leqslant \sqrt{k}, so writing j=k/ij = k / i we have

j=ki  ⩾  kk=k  ⩾  i,j=ki⩽ni,j = \frac{k}{i} \;\geqslant\; \frac{k}{\sqrt{k}} = \sqrt{k} \;\geqslant\; i, \qquad j = \frac{k}{i} \leqslant \frac{n}{i},

and jj is an integer, so i⩽j⩽⌊n/i⌋i \leqslant j \leqslant \lfloor n/i \rfloor. Since ii is prime it is not struck out, so when the outer loop reaches ii the test succeeds and the inner loop runs, setting p[i⋅j]=p[k]p[i \cdot j] = p[k] to “no”. Finally i⩽k<ki \leqslant \sqrt{k} < k, so this happens before the outer loop reaches kk, and kk is not output. The algorithm therefore outputs the primes and nothing else.

Running time. The first loop performs n−1n - 1 assignments. In the second loop, each of the n−1n-1 values of ii costs a bounded amount for the test and the output, contributing O(n)O(n) in total. The inner loop belonging to ii runs only when p[i]p[i] is “yes”, and then makes at most ⌊n/i⌋−i+1⩽n/i\lfloor n/i \rfloor - i + 1 \leqslant n/i passes, each of bounded cost. Summing over ii,

∑i=2nni=n(Hn−1)⩽nlog⁡2n\sum_{i=2}^{n} \frac{n}{i} = n\left(H_n - 1\right) \leqslant n \log_2 n

by the previous proposition. Adding the three contributions, the running time is O(n)+O(nlog⁡2n)=O(nlog⁡n)O(n) + O(n\log_2 n) = O(n \log n).

In Python the array pp is a list of n + 1 entries, True for “yes” and False for “no”; positions 00 and 11 are never read.

def sieve(n):
    p = [True] * (n + 1)
    primes = []
    for i in range(2, n + 1):
        if p[i]:
            primes.append(i)
            for j in range(i, n // i + 1):
                p[i * j] = False
    return primes

print(sieve(50))
# [2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37, 41, 43, 47]

Problem 4.6.

Count the assignments p[i * j] = False performed by the sieve for n=50n = 50, and compare the count with the bound nlog⁡2nn\log_2 n of the theorem. Then count how many of them write “no” to an entry that was already “no”, and say which composite numbers are struck more than once.

Problem 4.7.

Modify the sieve so that the inner loop starts at j=2j = 2 rather than j=ij = i. Show that the output is unchanged, and say how the count of assignments changes for n=50n = 50.

Problem 4.8.

Show that the sieve may stop its outer loop at i=⌊n⌋i = \lfloor \sqrt{n} \rfloor provided the surviving indices above that point are output afterwards. Which part of the proof of the theorem does this rely on?

Data Structures and Interfaces

A data structure is a way to store a non-constant amount of data, supporting a set of operations to interact with that data. The set of operations a data structure supports is its interface. Many data structures may support the same interface and differ in the cost of each operation, and many problems become easy once the data are stored in a suitable structure.

The most primitive data structure native to the Word-RAM is the static array: a contiguous sequence of words reserved in memory, supporting the static sequence interface.

  1. StaticArray(n): allocate a new static array of size nn, every entry initialised to 00, in Θ(n)\Theta(n) time.
  2. get_at(i): return the word stored at index ii, in Θ(1)\Theta(1) time.
  3. set_at(i, x): write the word xx to index ii, in Θ(1)\Theta(1) time.

The operations get_at(i) and set_at(i, x) run in constant time because every item of the array has the same size, one machine word. To store a larger object at an index, the machine word there is read as the memory address of a larger piece of memory holding the object. A Python tuple is like a static array without set_at(i, x).

Classes

Python writes a data structure as a class: a bundle of stored values, called attributes, together with the functions that act on them, called methods.

class Counter:
    def __init__(self, start):
        self.value = start

    def increment(self):
        self.value = self.value + 1

    def __len__(self):
        return self.value

c = Counter(5)
c.increment()
print(c.value, len(c))   # 6 6

Calling Counter(5) creates a new object of the class and runs __init__ on it, with self bound to the new object and start to 5; the assignment self.value = start creates the attribute value. A call c.increment() runs the method with self bound to c. A few method names are special: len(c) calls c.__len__(), and a for loop over c calls c.__iter__().

Inside a method, yield x hands x to the for loop that is running over the object, then carries on from the same point when the loop asks for the next value; yield from X hands on every value of X in turn. This is how range produces its entries one at a time. The statement raise IndexError stops the program with an error naming a bad index, and assert test stops it with an error if test is False.

A class may be built on another: class B(A): gives B every method of A, and any method B defines with the same name replaces the one from A. Inside B, super().__init__() runs the __init__ of A.

A Static Array in Python

Python has no static array, so we imitate one with a list whose length we never change.

class StaticArray:
    def __init__(self, n):
        self.data = [None] * n

    def get_at(self, i):
        if not (0 <= i < len(self.data)): raise IndexError
        return self.data[i]

    def set_at(self, i, x):
        if not (0 <= i < len(self.data)): raise IndexError
        self.data[i] = x

Birthday Matching

Given the students of a class, each with a name and a birthday, we want two students who share a birthday, or a report that there are none. The algorithm keeps a record of the students seen so far, and checks each new student against it before adding them.

Birthday Match
Input:  n students, each a pair (name, birthday).
Output: the names of two students with the same birthday, or None.

    record ≝ a static array of length n
    for k ≝ 0 to n − 1 do
        (name₁, bday₁) ≝ student k
        for i ≝ 0 to k − 1 do
            (name₂, bday₂) ≝ record[i]
            if bday₁ = bday₂ then return (name₁, name₂)
        record[k] ≝ (name₁, bday₁)
    return None
def birthday_match(students):
    # students: tuple of (name, bday) tuples
    n = len(students)                          # O(1)
    record = StaticArray(n)                    # O(n)
    for k in range(n):                         # n passes
        (name1, bday1) = students[k]           # O(1)
        for i in range(k):                     # k passes: check the record
            (name2, bday2) = record.get_at(i)  # O(1)
            if bday1 == bday2:                 # O(1)
                return (name1, name2)          # O(1)
        record.set_at(k, (name1, bday1))       # O(1)
    return None                                # O(1)

print(birthday_match((('Ada', 'Dec 10'), ('Alan', 'Jun 23'), ('Emmy', 'Mar 23'), ('Kurt', 'Jun 23'))))
# ('Kurt', 'Alan')

We assume that each name and each birthday fits into a constant number of machine words, so that one student’s information can be read and compared in constant time. This allows names and birthdays of O(w)O(w) characters from a fixed alphabet, and since w>log⁡2nw > \log_2 n it still allows every student’s information to be distinct.

Every line then takes constant time except three. Building record takes Θ(n)\Theta(n) time; the outer loop makes at most nn passes; and the inner loop on pass kk runs through the kk entries already in the record. The running time is therefore at most

O(n)+∑k=0n−1(O(1)+k⋅O(1))=O(n)+O(n)+O ⁣(n(n−1)2)=O(n2),O(n) + \sum_{k=0}^{n-1} \bigl( O(1) + k \cdot O(1) \bigr) = O(n) + O(n) + O\!\left(\frac{n(n-1)}{2}\right) = O(n^2),

using the sum of an arithmetic progression. This is quadratic in nn. A different data structure for the record does better, and the hashing section at the end of this chapter gives one.

Problem 4.9.

Suppose birthdays are given as integers 0,…,3650, \ldots, 365 (with 365365 for 29 February). Rewrite birthday_match so that it runs in O(n)O(n) time, using a static array of length 366366 indexed by birthday. Where does your running-time argument use the fact that the number of possible birthdays does not grow with nn?

Sequences and Sets

We use two interfaces, which differ in what decides the order of the stored items.

Sequences maintain a collection of items in an extrinsic order: each stored item has a rank in the sequence, including a first item and a last item. Extrinsic means that the first item is first not because of what the item is, but because some external party put it there. An iterable below is anything a for loop can run over, such as a tuple, a string, a list or a range.

OperationMeaning
Containerbuild(X)given an iterable X, build a sequence from the items of X
len()return the number of stored items
Staticiter_seq()return the stored items one by one in sequence order
get_at(i)return the item of rank i
set_at(i, x)replace the item of rank i with x
Dynamicinsert_at(i, x)add x as the item of rank i
delete_at(i)remove and return the item of rank i
insert_first(x)add x as the first item
delete_first()remove and return the first item
insert_last(x)add x as the last item
delete_last()remove and return the last item

The insert and delete operations change the rank of every item after the one inserted or deleted. Two restricted forms of a sequence have names of their own.

Definition 4.13 (Stack and queue).

A stack is a sequence used only through insert_last and delete_last: the item removed is always the one most recently added, “last in, first out”. A queue is a sequence used only through insert_last and delete_first: the item removed is always the one added longest ago, “first in, first out”.

The call stack of the functions section is a stack in this sense. With a Python list, append and pop() are insert_last and delete_last.

Sets, by contrast, maintain a collection of items based on an intrinsic property of what the items are, usually a unique key x.key attached to each item x. Sets generalise dictionaries and other databases queried by content.

OperationMeaning
Containerbuild(X)given an iterable X, build a set from the items of X
len()return the number of stored items
Staticfind(k)return the stored item with key k
Dynamicinsert(x)add x to the set, replacing the item with key x.key if there is one
delete(k)remove and return the stored item with key k
Orderiter_ord()return the stored items one by one in key order
find_min()return the stored item with smallest key
find_max()return the stored item with largest key
find_next(k)return the stored item with smallest key larger than k
find_prev(k)return the stored item with largest key smaller than k

The find operations return None if no qualifying item exists. In Python an item with a key can be an object of a small class:

class Item:
    def __init__(self, key, value):
        self.key = key
        self.value = value

We now give three data structures for the sequence interface. None of them supports insertion or deletion at an arbitrary rank in less than linear time.

Array Sequences

Computer memory is a finite resource. On a modern computer many running programs share the same main memory, so the operating system assigns a fixed range of memory addresses to each of them. When a program asks to store a variable it must say how much memory, how many bits, the variable needs; the operating system finds that much free memory in the program’s assigned range and reserves it, or allocates it, until it is no longer needed. Python hides memory management from the programmer, but whenever Python is asked to store something it makes such a request, for a fixed amount of memory, behind the scenes.

Now suppose a program wants to store two arrays, each of ten 6464-bit words. It makes two requests, for 640640 bits each, and the operating system might reserve the first ten words of the program’s range for the first array AA and the next ten for the second array BB. Later an eleventh word ww has to be added to AA, and there is no room next to AA: the start of the range is to its left, and BB is to its right. One could shift BB right to make room, but much other data may already be reserved beyond BB and would have to move too. It is better to request eleven new words, copy AA into the start of the new allocation, store ww at the end, and release the old ten words for later requests.

Memory itself is one large fixed-length array, from which the operating system allocates. Implementing a sequence with an array, so that index ii of the array holds the item of rank ii, makes get_at and set_at take O(1)O(1) time by random access. Inserting or deleting, however, means moving items and resizing the array, and these operations take linear time in the worst case.

class Array_Seq:
    def __init__(self):                     # O(1)
        self.A = []
        self.size = 0

    def __len__(self): return self.size     # O(1)
    def __iter__(self): yield from self.A   # O(n) iter_seq

    def build(self, X):                     # O(n)
        self.A = [a for a in X]             # stands in for a static array
        self.size = len(self.A)

    def get_at(self, i): return self.A[i]   # O(1)
    def set_at(self, i, x): self.A[i] = x   # O(1)

    def _copy_forward(self, i, n, A, j):    # O(n)
        for k in range(n):
            A[j + k] = self.A[i + k]

    def _copy_backward(self, i, n, A, j):   # O(n)
        for k in range(n - 1, -1, -1):
            A[j + k] = self.A[i + k]

    def insert_at(self, i, x):              # O(n)
        n = len(self)
        A = [None] * (n + 1)
        self._copy_forward(0, i, A, 0)
        A[i] = x
        self._copy_forward(i, n - i, A, i + 1)
        self.build(A)

    def delete_at(self, i):                 # O(n)
        n = len(self)
        A = [None] * (n - 1)
        self._copy_forward(0, i, A, 0)
        x = self.A[i]
        self._copy_forward(i + 1, n - i - 1, A, i)
        self.build(A)
        return x

    def insert_first(self, x): self.insert_at(0, x)             # O(n)
    def delete_first(self): return self.delete_at(0)            # O(n)
    def insert_last(self, x): self.insert_at(len(self), x)      # O(n)
    def delete_last(self): return self.delete_at(len(self) - 1) # O(n)

_copy_forward(i, n, A, j) copies the nn items starting at index ii into the array A starting at index jj, from left to right; _copy_backward copies the same items from right to left, which matters when source and destination overlap. A method name starting with an underscore is a convention for “used only inside the class”.

In a sorted array, whose items are arranged in the order of their keys, an item can be found far faster than by scanning, by the bisection of the last lesson; binary search in the next lesson makes this precise. Inserting an object while preserving the order is then harder, and costs O(n)O(n) time. Deleting an element, even when its index is known, also costs O(n)O(n) time if the array is to have no gaps and keep the order of the remaining elements.

Problem 4.10.

Trace Array_Seq on build((5, 7, 9)), then insert_at(1, 6), then delete_at(0), giving the list self.A after each operation. Count the item copies each of the two dynamic operations makes, and give that count for insert_at(i, x) on a sequence of length nn.

Linked Lists

In a linked list, inserting or deleting an item does not move the others. Its items are kept in a certain order, as in an array, but they can be stored anywhere in memory, in places independent of one another. With each item we store a reference to the place of the next item, its successor; for the last item this reference is None, which marks the end of the list. One further reference to the first item, the head, is needed to reach the list at all.

Linked lists can be singly or doubly linked. In a doubly linked list each item also stores a reference to its predecessor, and the list keeps a reference to its last item, the tail. This allows an item whose place is known to be deleted with a number of steps bounded by a constant independent of the length of the list: the running time is O(1)O(1).

prevdatanextitem 1prevdatanextitem 2prevdatanextitem 3firstlast
Figure 4.1. A doubly linked list with three items. Every item holds a data entry, a reference to the previous item and a reference to the next; a small open circle is None. The names first and last refer to the two ends. Deleting the accented parts leaves a singly linked list.

The Python below is a singly linked list. A node holds an item and a reference next; later_node(i) walks ii steps along the list.

class Linked_List_Node:
    def __init__(self, x):                  # O(1)
        self.item = x
        self.next = None

    def later_node(self, i):                # O(i)
        if i == 0: return self
        assert self.next
        return self.next.later_node(i - 1)

class Linked_List_Seq:
    def __init__(self):                     # O(1)
        self.head = None
        self.size = 0

    def __len__(self): return self.size     # O(1)

    def __iter__(self):                     # O(n) iter_seq
        node = self.head
        while node:
            yield node.item
            node = node.next

    def build(self, X):                     # O(n)
        for a in reversed(X):
            self.insert_first(a)

    def get_at(self, i):                    # O(i)
        node = self.head.later_node(i)
        return node.item

    def set_at(self, i, x):                 # O(i)
        node = self.head.later_node(i)
        node.item = x

    def insert_first(self, x):              # O(1)
        new_node = Linked_List_Node(x)
        new_node.next = self.head
        self.head = new_node
        self.size += 1

    def delete_first(self):                 # O(1)
        x = self.head.item
        self.head = self.head.next
        self.size -= 1
        return x

    def insert_at(self, i, x):              # O(i)
        if i == 0:
            self.insert_first(x)
            return
        new_node = Linked_List_Node(x)
        node = self.head.later_node(i - 1)
        new_node.next = node.next
        node.next = new_node
        self.size += 1

    def delete_at(self, i):                 # O(i)
        if i == 0:
            return self.delete_first()
        node = self.head.later_node(i - 1)
        x = node.next.item
        node.next = node.next.next
        self.size -= 1
        return x

    def insert_last(self, x): self.insert_at(len(self), x)       # O(n)
    def delete_last(self): return self.delete_at(len(self) - 1)  # O(n)

reversed(X) runs through X from its last entry to its first, so that inserting each at the front leaves them in their original order; a test while node: holds as long as node is not None.

Linked lists have a disadvantage: bisection cannot be applied to them, since reaching the middle item means walking to it. Scanning a linked list is also slower than scanning an array, by a constant factor only, because today’s computers reach consecutive storage places substantially faster than places far apart.

Problem 4.11.

Add a method reverse() to Linked_List_Seq that reverses the order of the items in O(n)O(n) time and O(1)O(1) auxiliary space, by changing the next references rather than the items. State the invariant your loop maintains.

Dynamic Arrays

The array sequence’s dynamic operations take time linear in the length of the array. One way to add items without paying a linear transfer cost every time is to over-allocate: request more space than the array currently needs, so that inserting an item means writing it into the next empty slot. This trades a little extra space for constant-time insertion. Any extra allocation is bounded, however; repeated insertions eventually fill it, and the array must be reallocated and copied again. Extra space reserved also means less space for the rest of the program.

Python does not append to the end of a list in worst-case O(1)O(1) time. Sometimes appending to a Python list requires O(n)O(n) time to transfer the array to a larger allocation, so sometimes appending takes linear time. Allocating extra space in the right way guarantees that any sequence of nn insertions takes O(n)O(n) time in total, because the linear-time transfers happen rarely, so insertion takes O(1)O(1) time per insertion on average over the sequence.

Definition 4.14 (Amortized cost).

An operation has amortized cost T(n)T(n) if every sequence of kk operations, starting from an empty data structure, takes at most k⋅T(n)k \cdot T(n) time in total, where nn is the largest size the structure reaches.

The cost of an expensive operation is amortized, that is spread, across the many cheap ones. To achieve amortized constant-time insertion at the end of an array, the strategy is to allocate extra space in proportion to the size of the array stored. Allocating Θ(n)\Theta(n) extra space ensures that a linear number of insertions must occur before an insertion overflows the allocation. A typical implementation allocates double the space needed for the current array, which is called table doubling; any constant fraction of extra space achieves the same bound. The list implementation of CPython, the standard Python interpreter, has used the rule

new_allocated = newsize // 8 + (3 if newsize < 9 else 6)   # extra slots

when a list must grow to newsize items, translated here from C; recent versions use a rule of the same shape. The extra allocation is modest, about one eighth of the size of the array, but it is still linear in that size, so on average n/8n/8 insertions are performed for every linear-time reallocation: amortized constant time.

Now consider removing items from the end. Popping the last item can be done in constant time by decrementing a stored length, which Python does. But if many items are removed from a large list, the unused allocation can hold a large amount of memory that is not available for other purposes. Once the array is small enough we transfer its contents to a smaller allocation and release the larger one. The new allocation cannot be exactly the size of the array, since an immediate insertion would then trigger another reallocation. For constant amortized time over any sequence of appends and pops, there must remain a linear fraction of unused space whenever we rebuild into a smaller array, which guarantees that Ω(n)\Omega(n) further operations must occur before the next reallocation.

The implementation below does both with table-doubling proportions. When an append would pass the end of the allocation, the contents move to an allocation twice as large. When removals bring the array down to a quarter of its allocation, the contents move to an allocation half as large. Python lists already work this way; the code shows how amortized constant-time append and pop can be implemented.

class Dynamic_Array_Seq(Array_Seq):
    def __init__(self, r = 2):              # O(1)
        super().__init__()
        self.size = 0
        self.r = r
        self._compute_bounds()
        self._resize(0)

    def __len__(self): return self.size     # O(1)

    def __iter__(self):                     # O(n)
        for i in range(len(self)): yield self.A[i]

    def build(self, X):                     # O(n)
        for a in X: self.insert_last(a)

    def _compute_bounds(self):              # O(1)
        self.upper = len(self.A)
        self.lower = len(self.A) // (self.r * self.r)

    def _resize(self, n):                   # O(1) or O(n)
        if (self.lower < n < self.upper): return
        m = max(n, 1) * self.r
        A = [None] * m
        self._copy_forward(0, self.size, A, 0)
        self.A = A
        self._compute_bounds()

    def insert_last(self, x):               # O(1) amortized
        self._resize(self.size + 1)
        self.A[self.size] = x
        self.size += 1

    def delete_last(self):                  # O(1) amortized
        self.A[self.size - 1] = None
        self.size -= 1
        self._resize(self.size)

    def insert_at(self, i, x):              # O(n)
        self.insert_last(None)
        self._copy_backward(i, self.size - (i + 1), self.A, i + 1)
        self.A[i] = x

    def delete_at(self, i):                 # O(n)
        x = self.A[i]
        self._copy_forward(i + 1, self.size - (i + 1), self.A, i)
        self.delete_last()
        return x

    def insert_first(self, x): self.insert_at(0, x)       # O(n)
    def delete_first(self): return self.delete_at(0)      # O(n)

def __init__(self, r = 2) gives the parameter r a default value: Dynamic_Array_Seq() uses r = 2, and Dynamic_Array_Seq(3) uses r = 3. The class inherits get_at, set_at and the two copying methods from Array_Seq.

Proposition 4.15 (Appending is amortized constant time).

Starting from an empty Dynamic_Array_Seq with r=2r = 2, any sequence of nn calls of insert_last takes O(n)O(n) time in total.

Discussion.

Each call does a constant amount of work apart from reallocations, so we bound the total cost of the reallocations. A reallocation at size ss allocates about 2s2s slots and copies ss items, so it costs O(s)O(s). The sizes at which reallocations happen are the points where the allocation fills up, and because the allocation doubles each time these sizes grow geometrically. Their sum is therefore dominated by the last one, which is less than nn, and a geometric sum bounds the total by a constant times nn.

Proof.

Without removals lower never exceeds a quarter of the allocation, so _resize(s + 1) reallocates exactly when s+1s + 1 reaches the current allocation upper. The initial call _resize(0) allocates 22 slots. After a reallocation triggered at size s+1=ms + 1 = m the allocation becomes 2m2m. So the allocations are 2,4,8,…2, 4, 8, \ldots, and the insertions that reallocate are those that bring the size to 2,4,8,…2, 4, 8, \ldots; the one bringing the size to 2t2^t copies 2t−12^t - 1 items and allocates 2t+12^{t+1} slots, at cost at most C 2t+1C\,2^{t+1} for a constant CC.

Over nn insertions these are the tt with 2t⩽n2^t \leqslant n, that is 1⩽t⩽m1 \leqslant t \leqslant m with m=⌊log⁡2n⌋m = \lfloor \log_2 n \rfloor. Their total cost is at most

∑t=1mC 2t+1=4C (2m−1)<4C n\sum_{t=1}^{m} C\, 2^{t+1} = 4C\,(2^m - 1) < 4C\,n

by the geometric sum. Every other part of every call costs O(1)O(1), contributing O(n)O(n). The total is O(n)O(n).

The worst-case costs of the three sequence structures are collected below, with (a) marking an amortized bound.

Data structurebuild(X)get_at(i), set_at(i, x)insert_first(x), delete_first()insert_last(x), delete_last()insert_at(i, x), delete_at(i)
Arraynn11nnnnnn
Linked listnnnn11nnnn
Dynamic arraynn11nn11 (a)nn

Each entry is the O(⋅)O(\cdot) bound as a function of the number nn of stored items.

Problem 4.12.

Take r=2r = 2 and start from an empty Dynamic_Array_Seq. Show that any sequence of nn operations, each an insert_last or a delete_last on a nonempty array, takes O(n)O(n) time in total. Then show that if the array were halved as soon as it fell to half full, rather than a quarter, some sequence of nn operations would take Θ(n2)\Theta(n^2) time.

Problem 4.13.

Extend Linked_List_Seq with a reference to its last node so that insert_last takes O(1)O(1) time, and extend Dynamic_Array_Seq so that insert_first and delete_first take O(1)O(1) amortized time. Which operation cannot be made O(1)O(1) in a singly linked list with a tail reference, and why?

Hashing

The set interface asks for find(k). Stored in an array in no particular order, a set answers find(k) by scanning, in O(n)O(n) time. With comparisons alone we cannot do much better: the lower bound for searching in the next lesson shows that Ω(log⁡n)\Omega(\log n) comparisons are needed for a search among nn items. Reading memory at an address computed from the key is not a comparison, and it reaches any memory cell in one step. Hashing uses this.

Direct Access Arrays

A direct access array is a static array with a meaning attached to each index: an item xx with key kk is stored at index kk. This makes sense only when keys are integers. Anything stored in a computer can be associated with an integer, for example its sequence of bits read as a binary number, or its address in memory, so from now on keys are integers.

Suppose we want to store a set of nn items whose unique integer keys lie in the range 0,…,u−10, \ldots, u - 1. We store them in a direct access array of length uu, whose slot ii holds the item with key ii if there is one. To find the item with key ii, look in slot ii: worst-case constant time. The order operations are slow: the first, last or next item could be in any slot, so they may take uu time.

class DirectAccessArray:
    def __init__(self, u): self.A = [None] * u   # O(u)
    def find(self, k): return self.A[k]          # O(1)
    def insert(self, x): self.A[x.key] = x       # O(1)
    def delete(self, k): self.A[k] = None        # O(1)

    def find_next(self, k):                      # O(u)
        for i in range(k + 1, len(self.A)):
            if self.A[i] is not None:
                return self.A[i]

    def find_max(self):                          # O(u)
        for i in range(len(self.A) - 1, -1, -1):
            if self.A[i] is not None:
                return self.A[i]

    def delete_max(self):                        # O(u)
        for i in range(len(self.A) - 1, -1, -1):
            x = self.A[i]
            if x is not None:
                self.A[i] = None
                return x

A direct access array needs a slot for every possible key in the range. When uu is very large compared with the number of items stored, the array is wasteful, or impossible to store at all. Suppose we wanted find(k) on ten-letter names with a direct access array. There are u=2610≈1.4×1014u = 26^{10} \approx 1.4 \times 10^{14} possible names, and even an array of one bit per name would need 17.617.6 terabytes.

The following counting principle is used below to show that collisions cannot be avoided.

Proposition 4.16 (The pigeonhole principle).

If NN objects are placed in rr boxes, some box contains at least ⌈N/r⌉\lceil N/r \rceil objects.

Discussion.

The argument is by contradiction on the total. If every box held fewer than ⌈N/r⌉\lceil N/r \rceil objects, each would hold at most ⌈N/r⌉−1\lceil N/r \rceil - 1, and the rr boxes together would hold fewer than NN. It remains to check the arithmetic step that ⌈N/r⌉−1<N/r\lceil N/r \rceil - 1 < N/r, which is the defining property of the ceiling.

Proof.

Suppose every box contains at most ⌈N/r⌉−1\lceil N/r \rceil - 1 objects. By the definition of the ceiling, ⌈N/r⌉−1<N/r\lceil N/r \rceil - 1 < N/r, so the total number of objects is at most r(⌈N/r⌉−1)<r⋅N/r=Nr\bigl(\lceil N/r \rceil - 1\bigr) < r \cdot N/r = N, a contradiction.

Hash Functions

To keep fast search while using only O(n)O(n) space when nn is much smaller than uu, we store the items in a smaller direct access array of m=O(n)m = O(n) slots, growing and shrinking it like a dynamic array according to the number of items stored. This needs a way to send each key to one of the mm slots.

Definition 4.17 (Hash function and hash table).

A hash function is a function

h:{0,…,u−1}→{0,…,m−1},h : \{0, \ldots, u - 1\} \to \{0, \ldots, m - 1\} ,

and h(k)h(k) is the hash of the key kk. The smaller direct access array of mm slots in which an item with key kk is stored at slot h(k)h(k) is a hash table. Two keys k1≠k2k_1 \neq k_2 collide if h(k1)=h(k2)h(k_1) = h(k_2).

If hh happens to be injective on the nn keys being stored, so that no two of them collide, the hash table acts as a direct access array over the smaller range {0,…,m−1}\{0, \ldots, m-1\} and supports worst-case constant-time search. When m<um < u, however, the pigeonhole principle puts at least two of the uu possible keys in some slot, and when the keys to be stored are not known in advance it is very unlikely that a chosen hash function avoids collisions among them. (If all the keys are known in advance, a scheme called perfect hashing can be designed to avoid collisions between them.)

A slot can hold one item, so colliding items must be stored somewhere. Either they are stored elsewhere in the same array, which is called open addressing and is how most hash tables are implemented in practice, though it is harder to analyse; or they are stored in a separate structure, which is called chaining and is the strategy we adopt.

Chaining

In chaining each slot of the hash table holds a reference to a chain, a separate data structure supporting the dynamic set operations find(k), insert(x) and delete(k). A chain is usually a linked list or a dynamic array, and any implementation will do provided each operation takes at most linear time in the length of the chain. To insert an item xx, insert it into the chain at slot h(x.key)h(x.\mathrm{key}); to find or delete a key kk, find or delete it in the chain at slot h(k)h(k).

01212473823634table
Figure 4.2. Chaining with m=5m = 5 slots and h(k)=k mod 5h(k) = k \bmod 5. The keys 1212 and 4747 collide in slot 22, and 88, 2323 and 6363 in slot 33; each slot refers to a chain holding the items that hash to it.

A chain in Python can be a list of items searched from the front:

class Chain:
    def __init__(self): self.items = []          # O(1)
    def __iter__(self): yield from self.items    # O(length)

    def find(self, k):                           # O(length)
        for x in self.items:
            if x.key == k: return x
        return None

    def insert(self, x):                         # O(length)
        for i in range(len(self.items)):
            if self.items[i].key == x.key:
                self.items[i] = x                # replace, nothing added
                return False
        self.items.append(x)
        return True

    def delete(self, k):                         # O(length)
        for i in range(len(self.items)):
            if self.items[i].key == k:
                return self.items.pop(i)
        return None

We want chains to be short: if every chain holds a constant number of items, the dynamic set operations run in constant time. If instead the hash function sends every stored key to the same slot, one chain has linear length and the operations can take linear time. A good hash function keeps collisions rare, so that no chain grows long.

Choosing a Hash Function

The simplest map from keys in {0,…,u−1}\{0, \ldots, u-1\} to {0,…,m−1}\{0, \ldots, m-1\} is the division method: h(k)=k mod mh(k) = k \bmod m, in Python k % m. If the keys stored are spread evenly over the range, it spreads them roughly evenly among the slots and the chains stay short. But if all of them happen to leave the same remainder on division by mm, every one lands in one chain. We want performance that does not depend on which keys are stored, and no single hash function gives that.

Remark (Every fixed hash function has bad inputs).

If u>nmu > nm, then every hash function hh from {0,…,u−1}\{0, \ldots, u-1\} to {0,…,m−1}\{0, \ldots, m-1\} sends some nn keys to the same slot. By the pigeonhole principle some slot receives at least ⌈u/m⌉\lceil u/m \rceil keys, and u/m>nu/m > n.

Instead we choose the hash function at random, from a large family, after the keys are fixed. Then no set of keys is bad for most of the family, and we can bound the cost on average over the choice of function. The averages we need are over a finite set of equally likely choices.

Definition 4.18 (Uniform choice from a finite family).

Let H\mathcal{H} be a finite nonempty set, and let hh be chosen from H\mathcal{H} with every member equally likely. For a property PP of members of H\mathcal{H}, and a function X:H→RX : \mathcal{H} \to \mathbb{R}, the probability of PP and the expectation of XX are

Pr⁡h∈H[P(h)]=#{h∈H:P(h)}#H,Eh∈H[X(h)]=1#H∑h∈HX(h).\Pr_{h \in \mathcal{H}}\bigl[P(h)\bigr] = \frac{\#\{h \in \mathcal{H} : P(h)\}}{\#\mathcal{H}}, \qquad \mathbb{E}_{h \in \mathcal{H}}\bigl[X(h)\bigr] = \frac{1}{\#\mathcal{H}} \sum_{h \in \mathcal{H}} X(h) .

Two facts follow from the laws of summation. Expectation is linear: E[X+Y]=E[X]+E[Y]\mathbb{E}[X + Y] = \mathbb{E}[X] + \mathbb{E}[Y] and E[cX]=c E[X]\mathbb{E}[cX] = c\,\mathbb{E}[X], because the sum defining the left side splits into the sums defining the right. And the expectation of an Iverson bracket is a probability: E[[P(h)]]=Pr⁡[P(h)]\mathbb{E}\bigl[[P(h)]\bigr] = \Pr[P(h)], since the bracket contributes 11 for each hh with P(h)P(h) and 00 otherwise. The expectation here is over the choice of hash function, which is made independently of the input. It is not an average over possible input keys.

Definition 4.19 (Universal family).

A finite family H\mathcal{H} of hash functions from {0,…,u−1}\{0, \ldots, u-1\} to {0,…,m−1}\{0, \ldots, m-1\} is universal if for any two keys ki≠kjk_i \neq k_j in {0,…,u−1}\{0, \ldots, u-1\},

Pr⁡h∈H[h(ki)=h(kj)]⩽1m.\Pr_{h \in \mathcal{H}}\bigl[h(k_i) = h(k_j)\bigr] \leqslant \frac{1}{m} .

A family that performs well is

H(m,p)={ hab(k)=((ak+b) mod p) mod m  ∣  a,b∈{0,…,p−1}, a≠0 },\mathcal{H}(m, p) = \bigl\{\, h_{ab}(k) = \bigl((ak + b) \bmod p\bigr) \bmod m \;\bigm|\; a, b \in \{0, \ldots, p-1\},\ a \neq 0 \,\bigr\},

where pp is a prime larger than uu. A single function of the family is specified by choosing concrete values of aa and bb. This family is universal. The proof uses arithmetic modulo a prime, which these notes have not developed, and we take it as given.

Proposition 4.20 (Expected chain length).

Let H\mathcal{H} be a universal family, and let nn distinct keys k0,…,kn−1k_0, \ldots, k_{n-1} be stored in a hash table of mm slots with chaining, using hh chosen uniformly from H\mathcal{H}. For each ii, the expected number of stored keys in the chain at slot h(ki)h(k_i) is at most 1+(n−1)/m1 + (n-1)/m.

Discussion.

The chain holding kik_i contains exactly the stored keys that collide with kik_i, together with kik_i itself. So its length is a sum of Iverson brackets, one for each stored key, and linearity turns the expectation of the sum into a sum of expectations. Each bracket’s expectation is a collision probability: the one for j=ij = i is 11, and each of the n−1n - 1 others is at most 1/m1/m by universality.

Proof.

For each jj let Xij(h)=[ h(ki)=h(kj) ]X_{ij}(h) = [\,h(k_i) = h(k_j)\,], which is 11 if kik_i and kjk_j collide under hh and 00 otherwise. The number of stored keys in the chain at slot h(ki)h(k_i) is Xi=∑jXijX_i = \sum_{j} X_{ij}, and Xii=1X_{ii} = 1 for every hh. By linearity and universality,

Eh∈H[Xi]=∑jE[Xij]=1+∑j≠iPr⁡h∈H[h(ki)=h(kj)]⩽1+∑j≠i1m=1+n−1m.\mathbb{E}_{h \in \mathcal{H}}[X_i] = \sum_{j} \mathbb{E}[X_{ij}] = 1 + \sum_{j \neq i} \Pr_{h \in \mathcal{H}}\bigl[h(k_i) = h(k_j)\bigr] \leqslant 1 + \sum_{j \neq i} \frac{1}{m} = 1 + \frac{n-1}{m} .

If the table is at least linear in the number of items stored, m=Ω(n)m = \Omega(n), the expected length of any chain is 1+(n−1)/Ω(n)=O(1)1 + (n-1)/\Omega(n) = O(1). A hash table with chaining and a hash function chosen at random from a universal family therefore performs the dynamic set operations in expected constant time, the expectation being over the choice of hash function and not over the input keys. To keep m=Θ(n)m = \Theta(n), insertions and deletions may have to rebuild the table at a different size and reinsert every item, as a dynamic array does; this makes the bounds for the dynamic operations amortized as well.

A Hash Table in Python

The statement from random import randint makes the function randint available: randint(a, b) returns an integer chosen uniformly from a,a+1,…,ba, a+1, \ldots, b. The keys are assumed to be integers below the prime p=231−1p = 2^{31} - 1. In [Chain() for _ in range(m)] the name _ is the usual name for a loop variable whose value is not used.

from random import randint

class Hash_Table_Set:
    def __init__(self):                          # O(1)
        self.A = []
        self.size = 0
        self.p = 2**31 - 1                       # a prime larger than every key
        self.a = randint(1, self.p - 1)
        self.b = randint(0, self.p - 1)
        self._compute_bounds()
        self._resize(0)

    def __len__(self): return self.size          # O(1)

    def __iter__(self):                          # O(n)
        for chain in self.A:
            yield from chain

    def build(self, X):                          # O(n) expected
        for x in X: self.insert(x)

    def _hash(self, k, m):                       # O(1)
        return ((self.a * k + self.b) % self.p) % m

    def _compute_bounds(self):                   # O(1)
        self.upper = len(self.A)
        self.lower = len(self.A) // 4

    def _resize(self, n):                        # O(n)
        if self.lower < n < self.upper: return
        m = max(n, 1) * 2
        A = [Chain() for _ in range(m)]
        for x in self:
            A[self._hash(x.key, m)].insert(x)
        self.A = A
        self._compute_bounds()

    def find(self, k):                           # O(1) expected
        h = self._hash(k, len(self.A))
        return self.A[h].find(k)

    def insert(self, x):                         # O(1) amortized expected
        self._resize(self.size + 1)
        h = self._hash(x.key, len(self.A))
        added = self.A[h].insert(x)
        if added: self.size += 1
        return added

    def delete(self, k):                         # O(1) amortized expected
        assert len(self) > 0
        h = self._hash(k, len(self.A))
        x = self.A[h].delete(k)
        if x is not None:
            self.size -= 1
            self._resize(self.size)
        return x

    def find_min(self):                          # O(n)
        out = None
        for x in self:
            if (out is None) or (x.key < out.key):
                out = x
        return out

    def find_max(self):                          # O(n)
        out = None
        for x in self:
            if (out is None) or (x.key > out.key):
                out = x
        return out

    def find_next(self, k):                      # O(n)
        out = None
        for x in self:
            if x.key > k:
                if (out is None) or (x.key < out.key):
                    out = x
        return out

    def find_prev(self, k):                      # O(n)
        out = None
        for x in self:
            if x.key < k:
                if (out is None) or (x.key > out.key):
                    out = x
        return out

    def iter_ord(self):                          # O(n^2)
        x = self.find_min()
        while x:
            yield x
            x = self.find_next(x.key)

The number of items stays strictly between a quarter of the number of slots and the number of slots; when it leaves that range the table is rebuilt with twice as many slots as items, which is the rule of Dynamic_Array_Seq with r=2r = 2. So m=Θ(n)m = \Theta(n) throughout, and by the proposition every chain has expected constant length.

Problem 4.14.

Insert the keys 3,13,23,33,8,183, 13, 23, 33, 8, 18 in that order into a hash table of m=10m = 10 slots with chaining and the division method h(k)=k mod 10h(k) = k \bmod 10, and draw the table. Then do the same with m=7m = 7. Describe every set of keys that makes the division method with m=10m = 10 put all keys into one chain.

Problem 4.15.

Rewrite birthday_match using a Hash_Table_Set keyed by birthday, with birthdays given as integers, so that it runs in expected O(n)O(n) time. Explain why the O(n)O(n) bound is an expectation and what it is taken over.

Graphs

A graph records which pairs of objects are related: towns joined by roads, people who know each other, states of a computation that lead to one another. The objects are the vertices and the related pairs are the edges.

Undirected Graphs

Definition 4.21 (Undirected graph).

An undirected graph is a pair G=(S,A)G = (S, A), where SS is a finite set of vertices and AA is a set of unordered pairs {si,sj}\{s_i, s_j\} of vertices. Such a pair is an edge. Unless stated otherwise, the two vertices of an edge are distinct.

The vertices si,sjs_i, s_j joined by an edge are adjacent, and we write

Adj⁡(si)={ sj∈S:{si,sj}∈A }\operatorname{Adj}(s_i) = \bigl\{\, s_j \in S : \{s_i, s_j\} \in A \,\bigr\}

for the set of vertices adjacent to sis_i. We draw an edge as a line between its endpoints.

156243
Figure 4.3. A graph with S={1,…,6}S = \{1, \ldots, 6\} and four edges. Vertex 44 lies on no edge.

Here S={1,2,3,4,5,6}S = \{1, 2, 3, 4, 5, 6\} and A={{1,2},{1,5},{2,5},{3,6}}A = \bigl\{\{1, 2\}, \{1, 5\}, \{2, 5\}, \{3, 6\}\bigr\}. Vertex 44 is isolated: it belongs to SS but to no edge. For example, Adj⁡(2)={1,5}\operatorname{Adj}(2) = \{1, 5\} and Adj⁡(4)=∅\operatorname{Adj}(4) = \varnothing.

Definition 4.22 (Loops, simple graphs and multigraphs).

A loop joins a vertex to itself. A graph is simple if it has no loops and at most one edge between any two vertices. If repeated edges are allowed, the edge collection is a multiset and the graph is a multigraph. We normally work with simple graphs.

Definition 4.23 (Order and size).

The order of a graph is its number ∣S∣|S| of vertices. Its size is its number ∣A∣|A| of edges, counted with multiplicity for a multigraph.

The graph of Figure 4.3 has order 66 and size 44, even though one vertex is isolated. The pair {1,5}\{1, 5\} is the same unordered edge as {5,1}\{5, 1\}; listing both would not add an edge to a simple graph.

Running times of graph algorithms are functions of the graph, not of a single number, and the notation of the last chapter is extended to cover them. For a graph GG we write S(G)S(G) and A(G)A(G) for its vertex set and edge set.

Definition 4.24 (Landau notation on graphs).

Let G\mathcal{G} be the set of all graphs, and let f,g:G→R⩾0f, g : \mathcal{G} \to \mathbb{R}_{\geqslant 0}. We say that f=O(g)f = O(g) if there exist α>0\alpha > 0 and n0∈Nn_0 \in \mathbb{N} such that

f(G)⩽α⋅g(G)for all G∈G with ∣S(G)∣+∣A(G)∣⩾n0.f(G) \leqslant \alpha \cdot g(G) \qquad \text{for all } G \in \mathcal{G} \text{ with } |S(G)| + |A(G)| \geqslant n_0 .

The notations Ω\Omega and Θ\Theta are extended in the same way.

In other words, if ff is greater than gg, then by at most a constant factor, with exceptions allowed only among graphs with fewer than n0n_0 vertices and edges together. The same definition can be made on any countable set of inputs, with a measure of size in place of ∣S(G)∣+∣A(G)∣|S(G)| + |A(G)|. The function ff usually describes the running time of an algorithm or some memory requirement, and gg often depends only on the numbers of vertices and edges, which we write n=∣S(G)∣n = |S(G)| and m=∣A(G)∣m = |A(G)| throughout. Thus O(n+m)O(n + m) means at most a constant times the number of vertices plus edges.

Directed Graphs

Definition 4.25 (Directed graph).

A directed graph is a pair G=(S,A)G = (S, A) with finite vertex set SS and arc set A⊆S×SA \subseteq S \times S. An arc (si,sj)(s_i, s_j) starts at sis_i and ends at sjs_j, and is drawn si→sjs_i \to s_j. The vertex sjs_j is a successor of sis_i, and sis_i a predecessor of sjs_j.

We write

Succ⁡(si)={ sj:(si,sj)∈A },Pred⁡(si)={ sj:(sj,si)∈A }.\operatorname{Succ}(s_i) = \{\, s_j : (s_i, s_j) \in A \,\}, \qquad \operatorname{Pred}(s_i) = \{\, s_j : (s_j, s_i) \in A \,\}.
124563
Figure 4.4. A directed graph. The two curved arcs between 44 and 55 point in opposite directions, and 44 carries a loop.

In Figure 4.4,

A={(1,2),(2,4),(2,5),(4,1),(4,4),(4,5),(5,4),(6,3)},A = \{(1, 2), (2, 4), (2, 5), (4, 1), (4, 4), (4, 5), (5, 4), (6, 3)\},

so for example Succ⁡(4)={1,4,5}\operatorname{Succ}(4) = \{1, 4, 5\} and Pred⁡(4)={2,4,5}\operatorname{Pred}(4) = \{2, 4, 5\}.

Definition 4.26 (Directed loops and pp-graphs).

A directed loop is an arc (v,v)(v, v). A directed graph without loops is sometimes called elementary. A directed pp-graph permits at most pp parallel arcs with any given ordered endpoints; in a 11-graph the arcs form a set, not a multiset.

The pair (4,5)(4, 5) and the pair (5,4)(5, 4) are different arcs, and removing one changes the graph. Reversing every arc of a drawing exchanges each successor set with the corresponding predecessor set.

Problem 4.16.

Let S={1,…,12}S = \{1, \ldots, 12\} and put an arc (a,b)(a, b) whenever a≠ba \neq b and aa divides bb. List Succ⁡(2)\operatorname{Succ}(2), Pred⁡(12)\operatorname{Pred}(12) and every vertex with no successors, and count the arcs.

Degree

Definition 4.27 (Degree).

In an undirected graph, the degree d(v)d(v) of a vertex vv is the number of edge-ends at vv, a loop counting twice. In a directed graph, the out-degree d+(v)d^+(v) counts the arcs beginning at vv, the in-degree d−(v)d^-(v) counts the arcs ending at vv, and d(v)=d+(v)+d−(v)d(v) = d^+(v) + d^-(v). A directed loop contributes one to each of d+(v)d^+(v) and d−(v)d^-(v).

We write δ(v)\delta(v) for the set of edges with an end at vv in an undirected graph, and δ+(v)\delta^+(v) and δ−(v)\delta^-(v) for the sets of arcs beginning and ending at vv in a directed graph.

In a graph without loops d(v)=∣δ(v)∣d(v) = |\delta(v)|, and in a directed graph d+(v)=∣δ+(v)∣d^+(v) = |\delta^+(v)| and d−(v)=∣δ−(v)∣d^-(v) = |\delta^-(v)|, counting parallel arcs separately. In a simple graph d(v)=∣Adj⁡(v)∣d(v) = |\operatorname{Adj}(v)|, and in a directed 11-graph d+(v)=∣Succ⁡(v)∣d^+(v) = |\operatorname{Succ}(v)| and d−(v)=∣Pred⁡(v)∣d^-(v) = |\operatorname{Pred}(v)|. For multigraphs, arcs are counted with multiplicity rather than by taking the sizes of these sets.

In Figure 4.3 the degrees of the vertices 1,2,3,4,5,61, 2, 3, 4, 5, 6 are 2,2,1,0,2,12, 2, 1, 0, 2, 1. Their sum is 8=2∣A∣8 = 2|A|. In Figure 4.4, d+(4)=3d^+(4) = 3 and d−(4)=3d^-(4) = 3: the loop counts once in each.

Proposition 4.28 (Counting edge-ends).

For an undirected graph and for a directed graph respectively,

∑v∈Sd(v)=2∣A∣,∑v∈Sd+(v)=∑v∈Sd−(v)=∣A∣.\sum_{v \in S} d(v) = 2|A|, \qquad \sum_{v \in S} d^+(v) = \sum_{v \in S} d^-(v) = |A| .

Discussion.

Both sides of each identity count the same thing in two ways. The left side of the first sums, vertex by vertex, the edge-ends at that vertex; the right side counts the same edge-ends edge by edge, and every edge has exactly two ends, a loop included. For a directed graph every arc has one tail and one head, so summing out-degrees counts each arc once by its tail, and summing in-degrees counts it once by its head.

Proof.

Count the pairs (v,e)(v, e) in which vv is an end of the edge ee, a loop at vv giving two such pairs. Grouping by vv gives ∑vd(v)\sum_{v} d(v), by the definition of degree. Grouping by ee gives 22 for every edge, so 2∣A∣2|A|. For a directed graph, count the pairs (v,a)(v, a) in which the arc aa begins at vv: grouping by vv gives ∑vd+(v)\sum_v d^+(v), and grouping by aa gives 11 for every arc, so ∣A∣|A|. Counting the pairs in which aa ends at vv gives ∑vd−(v)=∣A∣\sum_v d^-(v) = |A| in the same way.

Example 4.29 (Odd degrees and repeated degrees).

The degree sum of an undirected graph is 2m2m, an even number. A sum of integers is even exactly when it contains an even number of odd terms: each even degree contributes zero modulo 22, and each odd degree contributes one. Thus the number of vertices of odd degree is even.

Also, among n⩾2n \geqslant 2 vertices of a simple graph, two have the same degree. Every degree lies in {0,1,…,n−1}\{0, 1, \ldots, n-1\}. However, degree n−1n - 1 means a vertex is adjacent to every other vertex, so no vertex can have degree 00 at the same time. At most n−1n - 1 of the nn listed values can occur, and putting nn vertices into at most n−1n - 1 degree classes forces two into one class by the pigeonhole principle, Proposition 4.16 .

Kinds of Graphs

We may attach numbers to edges or arcs, to express distance, cost or capacity.

Definition 4.30 (Weighted graph).

A weighted graph G=(S,A,ν)G = (S, A, \nu) is a graph together with a function ν:A→R\nu : A \to \mathbb{R}, the weight. An undirected weight belongs to the unordered edge {u,v}\{u, v\}.

Definition 4.31 (Complete graph).

A simple undirected graph is complete if every pair of distinct vertices is joined; it is written KnK_n when it has nn vertices. A loop-free directed 11-graph is complete when both (u,v)(u, v) and (v,u)(v, u) are arcs for every u≠vu \neq v.

K3123K512345
Figure 4.5. The complete graphs K3K_3 and K5K_5, with 33 and 1010 edges.

To count the edges of KnK_n, choose two distinct vertices without regard to order. There are nn choices for the first vertex and n−1n - 1 for the second. This counts each unordered pair twice, once in each order, so KnK_n has n(n−1)/2n(n-1)/2 edges. We write

(n2)=n(n−1)2,\binom{n}{2} = \frac{n(n-1)}{2},

read ”nn choose 22”, for this number of unordered pairs among nn objects. In the complete directed graph both directions are arcs, so it has n(n−1)n(n-1) arcs.

Definition 4.32 (Subgraphs).

Let G=(S,A)G = (S, A) be a directed or undirected graph. For S′⊆SS' \subseteq S, the induced subgraph G[S′]G[S'] has vertex set S′S' and contains all edges or arcs of GG whose endpoints lie in S′S'. A subgraph (S′,A′)(S', A') may keep only some of these: A′⊆A∩(S′×S′)A' \subseteq A \cap (S' \times S') in the directed case, with the analogous rule for unordered edges. A spanning subgraph, also called a partial subgraph, keeps every vertex but possibly deletes edges.

For a simple undirected graph, the possible edges of a subgraph on S′S' are the unordered pairs of distinct vertices of S′S' that were already edges of GG. Figure 4.6 shows both operations on a four-vertex directed graph.

G1234G′1234G[{1, 2, 4}]124
Figure 4.6. A directed graph GG (left), the spanning subgraph G′G' obtained by deleting (2,2)(2, 2), (3,2)(3, 2) and (3,4)(3, 4) (middle), and the induced subgraph G[{1,2,4}]G[\{1, 2, 4\}] (right).

Here G′G' is the spanning subgraph obtained by deleting (2,2)(2, 2), (3,2)(3, 2) and (3,4)(3, 4). In contrast, the induced subgraph G[{1,2,4}]G[\{1, 2, 4\}] retains the arcs (2,1)(2, 1), (4,1)(4, 1), (4,2)(4, 2) and (2,2)(2, 2): vertex 33 disappears, while every arc of GG joining the remaining vertices stays. An arc between retained vertices cannot be omitted from an induced subgraph.

Definition 4.33 (Clique and independent set).

In a simple undirected graph, a clique is a vertex set CC for which G[C]G[C] is complete, and an independent set, also called a stable set, is a vertex set II for which G[I]G[I] has no edges. For a loop-free directed 11-graph, a directed clique contains both arcs between each pair of its vertices, while a directed stable set has no arcs at all.

72851634
Figure 4.7. The accented triangle {1,5,7}\{1, 5, 7\} is a clique of maximum size; {2,4,5}\{2, 4, 5\} is an independent set of maximum size.

In Figure 4.7, {1,5,7}\{1, 5, 7\} is a clique of maximum size and {2,4,5}\{2, 4, 5\} is an independent set of maximum size. To test whether {1,5,7}\{1, 5, 7\} is a clique, check the three pairs {1,5}\{1, 5\}, {1,7}\{1, 7\}, {5,7}\{5, 7\}; to test whether {2,4,5}\{2, 4, 5\} is independent, check that none of its three pairs is an edge. Checking a particular set requires only pairwise tests. Proving that it is of maximum size also requires ruling out every larger set, and finding maximum cliques in arbitrary graphs is much harder than checking whether a proposed set is a clique.

Definition 4.34 (Bipartite and regular graphs).

An undirected graph is bipartite if its vertices can be divided into disjoint sets S1,S2S_1, S_2 so that every edge has one endpoint in each. It is complete bipartite, written Kn1,n2K_{n_1, n_2}, if ∣Si∣=ni|S_i| = n_i and every pair with one vertex in each set is an edge. A graph is kk-regular if every vertex has degree kk.

Example 4.35 (K2,3K_{2,3} and the four-cycle).

In K2,3K_{2,3}, each of the two vertices of S1S_1 is joined to each of the three vertices of S2S_2, giving 2⋅3=62 \cdot 3 = 6 edges. Its degrees are 33 on S1S_1 and 22 on S2S_2, so it is not regular. The graph with vertices 1,2,3,41, 2, 3, 4 and edges {1,2},{2,3},{3,4},{4,1}\{1,2\}, \{2,3\}, \{3,4\}, \{4,1\}, the cycle on four vertices, is 22-regular and bipartite, with classes {1,3}\{1, 3\} and {2,4}\{2, 4\}.

In general Kn1,n2K_{n_1, n_2} has n1n2n_1 n_2 edges. Fix a vertex of S1S_1; it is joined to each of the n2n_2 vertices of S2S_2. There are n1n_1 such vertices, and every edge has exactly one endpoint in S1S_1, so this counts each edge exactly once.

Example 4.36 (Counting the same thing twice).

Let a bipartite graph with parts S1,S2S_1, S_2 be kk-regular, where k>0k > 0, and let it have mm edges. Sum the degrees of the vertices in S1S_1. Every edge has exactly one endpoint there, so this sum counts every edge once and equals mm. Each of the ∣S1∣|S_1| vertices has degree kk, so the same sum is k∣S1∣k|S_1|. Therefore m=k∣S1∣m = k|S_1|. Repeating the count over S2S_2 gives m=k∣S2∣m = k|S_2|. The two expressions for mm are equal, and dividing by the positive number kk gives ∣S1∣=∣S2∣|S_1| = |S_2|.

Example 4.37 (A monochromatic triangle).

Colour each edge of K6K_6 red or blue. Then there is a triangle whose three edges have the same colour. Pick a vertex xx. Five edges leave xx, so by the pigeonhole principle at least ⌈5/2⌉=3\lceil 5/2 \rceil = 3 of them, say xy1,xy2,xy3xy_1, xy_2, xy_3, share a colour; suppose it is red. If any edge among y1,y2,y3y_1, y_2, y_3 is red, that edge and the two red edges to xx make a red triangle. If none is red, all three edges among y1,y2,y3y_1, y_2, y_3 are blue, making a blue triangle. These cases cover every colouring, so the argument works for all of them.

Problem 4.17.

Show that K5K_5 has a red and blue colouring of its edges with no triangle all of one colour, so that 66 in the last example cannot be replaced by 55.

Problem 4.18.

Show that a graph is bipartite if it has no cycle of odd length, in the sense of the next section, by colouring each vertex according to the parity of its distance from a fixed vertex in its component. Show conversely that a bipartite graph has no cycle of odd length.

Walks, Paths, Cycles and Distance

Definition 4.38 (Walks and paths in a directed graph).

In a directed graph, a walk from uu to vv is a sequence ⟨s0,s1,…,sk⟩\langle s_0, s_1, \ldots, s_k \rangle of vertices with s0=us_0 = u, sk=vs_k = v and (si−1,si)∈A(s_{i-1}, s_i) \in A for every 1⩽i⩽k1 \leqslant i \leqslant k. Its length is kk. The vertex vv is reachable from uu if such a walk exists. A simple path is a walk with no repeated vertex. A closed walk has s0=sks_0 = s_k and k⩾1k \geqslant 1; a simple directed cycle is a closed walk with no further repeated vertices. A loop is a cycle of length one.

A walk may repeat vertices; a simple path may not. We allow the walk ⟨u⟩\langle u \rangle of length zero from uu to itself, so that every vertex is reachable from itself. This convention is used for matrix powers below. When a positive number of arcs is required we say positive-length reachability explicitly.

123456
Figure 4.8. The accented arcs form the simple directed cycle ⟨1,2,5,4,1⟩\langle 1, 2, 5, 4, 1 \rangle.

In Figure 4.8, ⟨1,4,2,5⟩\langle 1, 4, 2, 5 \rangle is a simple path, ⟨3,6,6,6⟩\langle 3, 6, 6, 6 \rangle is a walk with the repeated vertex 66, ⟨1,2,5,4,1⟩\langle 1, 2, 5, 4, 1 \rangle is a simple directed cycle, and ⟨1,2,5,4,2,5,4,1⟩\langle 1, 2, 5, 4, 2, 5, 4, 1 \rangle is a closed walk that is not a simple cycle.

Definition 4.39 (Walks, trails and cycles in an undirected graph).

In an undirected graph, consecutive vertices of a walk are joined by an edge, and walks, length, reachability and simple paths are as in the directed case. A trail is a walk that uses no edge twice. A cycle is a closed trail of positive length; a simple cycle additionally has no repeated vertices except its first and last. A graph with no cycles is acyclic. In a simple undirected graph a simple cycle has length at least three.

Proposition 4.40 (A walk contains a simple path).

If there is a walk from uu to vv, there is a simple path from uu to vv.

Discussion.

If a walk visits some vertex twice, the part between the two visits can be cut out, leaving a shorter walk between the same endpoints. So a walk with the fewest edges can have no repeated vertex. The proof takes such a shortest walk, which exists because walk lengths are nonnegative integers and at least one walk exists.

Proof.

Among all walks from uu to vv, choose one, ⟨s0,…,sk⟩\langle s_0, \ldots, s_k \rangle, with the fewest edges. If it repeated a vertex, say si=sjs_i = s_j with i<ji < j, then ⟨s0,…,si,sj+1,…,sk⟩\langle s_0, \ldots, s_i, s_{j+1}, \ldots, s_k \rangle would be a walk from uu to vv with k−(j−i)<kk - (j - i) < k edges, since (si,sj+1)=(sj,sj+1)(s_i, s_{j+1}) = (s_j, s_{j+1}) is an arc, or edge, of the graph. Hence no vertex repeats.

Definition 4.41 (Distance and diameter).

The distance d(u,v)d(u, v) is the length of a shortest walk from uu to vv, or +∞+\infty if vv is not reachable from uu. The diameter of a nonempty graph is max⁡u,v∈Sd(u,v)\max_{u, v \in S} d(u, v), which may be +∞+\infty. In directed graphs, d(u,v)d(u, v) and d(v,u)d(v, u) may differ.

The two-argument d(u,v)d(u, v) is a distance and the one-argument d(v)d(v) a degree; the number of arguments tells them apart. A shortest walk to a different vertex is a simple path, by the proof of the last proposition, so it has at most n−1n - 1 arcs.

In Figure 4.8, d(1,5)=2d(1, 5) = 2 via 1→2→51 \to 2 \to 5, and d(5,1)=2d(5, 1) = 2 via 5→4→15 \to 4 \to 1. The arc 3→63 \to 6 gives d(3,6)=1d(3, 6) = 1, but there is no directed route from 66 to 33, so d(6,3)=∞d(6, 3) = \infty.

Problem 4.19.

Give the distance d(u,v)d(u, v) for every ordered pair of vertices of Figure 4.8, as a 6×66 \times 6 table, and find the diameter. Which vertices can reach every other vertex?

Representing a Graph

To run an algorithm on a graph we must store it. Let G=(S,A)G = (S, A) have its vertices numbered 1,…,n1, \ldots, n.

Adjacency Lists

An adjacency list T[i]T[i] lists the vertices jj for which (i,j)(i, j) is an arc, or {i,j}\{i, j\} is an edge. For an undirected edge {i,j}\{i, j\} with i≠ji \neq j, jj appears in T[i]T[i] and ii appears in T[j]T[j]. The order within a list is arbitrary.

More generally, an adjacency list representation keeps for every vertex vv a list of all the edges incident with it, the set δ(v)\delta(v); for simple graphs a list of the adjacent vertices, as above, is often enough. For a directed graph one keeps two lists for each vertex, one of the outgoing arcs δ+(v)\delta^+(v) and one of the incoming arcs δ−(v)\delta^-(v).

Example 4.42 (Adjacency lists of a directed graph).

The directed graph GG of Figure 4.6 has

T[1]=(),T[2]=(2,1),T[3]=(1,4,3,2),T[4]=(1,2,3).T[1] = (), \qquad T[2] = (2, 1), \qquad T[3] = (1, 4, 3, 2), \qquad T[4] = (1, 2, 3).

Its arc set is exactly the nine ordered pairs described by these lists. The loop at 22 contributes the entry 22 to T[2]T[2], and the loop at 33 contributes the entry 33 to T[3]T[3]. There are 0+2+4+3=90 + 2 + 4 + 3 = 9 entries, hence nine arcs; counting list entries is a direct way to check the arc count.

Storage. There are nn list headers. The total number of entries is mm for a directed graph with mm arcs and 2m2m for an undirected graph with mm edges, a loop also taking two entries if loops are allowed. Thus storage is O(n+m)O(n + m).

Operations. Reading T[i]T[i] gives the successors and the out-degree of ii quickly. Testing for one particular arc (i,j)(i, j) may require scanning T[i]T[i]. Finding all predecessors of jj requires scanning all the lists, unless reverse lists are stored as well. Visiting all arcs takes O(n+m)O(n + m) time, not merely O(m)O(m), when isolated vertices must also be visited.

Adjacency Lists in Two Arrays

We can store the adjacency lists one after another in a single array Succ[1..m]\mathrm{Succ}[1..m]. A second array Head[1..n]\mathrm{Head}[1..n] gives the final position of each list, and Head[0]=0\mathrm{Head}[0] = 0 gives the starting boundary of the first. The successors of vv then occupy

Succ[Head[v−1]+1],…,Succ[Head[v]],\mathrm{Succ}\bigl[\mathrm{Head}[v-1] + 1\bigr], \ldots, \mathrm{Succ}\bigl[\mathrm{Head}[v]\bigr],

and an empty list has equal consecutive boundaries. For the graph of the last example,

(Head[1],…,Head[4])=(0,2,6,9),Succ=(2,1,1,4,3,2,1,2,3).\bigl(\mathrm{Head}[1], \ldots, \mathrm{Head}[4]\bigr) = (0, 2, 6, 9), \qquad \mathrm{Succ} = (2, 1, 1, 4, 3, 2, 1, 2, 3).

The empty first list is encoded by Head[1]=0\mathrm{Head}[1] = 0; the second list ends at position 22, and so on. These boundaries are cumulative counts: if d+(v)d^+(v) is the length of list vv, then Head[v]=d+(1)+⋯+d+(v)\mathrm{Head}[v] = d^+(1) + \cdots + d^+(v).

To find the predecessors of 22, scan each list and record its owner vv whenever an entry equals 22; the result is 2,3,42, 3, 4. This costs O(n+m)O(n + m) even though only three names are recorded. The same cumulative-count idea produces all the predecessor lists in O(n+m)O(n + m): count how many times each vertex occurs in Succ\mathrm{Succ}, which gives the in-degrees; take cumulative sums to allocate a contiguous block to each vertex; then scan the arcs once more to fill the blocks.

Predecessor Lists
Input:  n, and the arrays Head[0..n] and Succ[1..m] of a directed graph.
Output: arrays PHead[0..n] and Pred[1..m] storing the predecessor lists
        in the same way.

    for v ≝ 1 to n do c[v] ≝ 0
    for i ≝ 1 to m do c[Succ[i]] ≝ c[Succ[i]] + 1
    PHead[0] ≝ 0
    for v ≝ 1 to n do
        PHead[v] ≝ PHead[v − 1] + c[v]
        free[v] ≝ PHead[v − 1] + 1
    for u ≝ 1 to n do
        for i ≝ Head[u − 1] + 1 to Head[u] do
            w ≝ Succ[i]
            Pred[free[w]] ≝ u
            free[w] ≝ free[w] + 1

Python numbers list positions from 00, so the successors of vv are the slice Succ[Head[v - 1]:Head[v]] and the positions of the blocks shift down by one.

def predecessor_lists(n, Head, Succ):
    count = [0] * (n + 1)
    for w in Succ:
        count[w] += 1
    PHead = [0] * (n + 1)
    free = [0] * (n + 1)
    for v in range(1, n + 1):
        PHead[v] = PHead[v - 1] + count[v]
        free[v] = PHead[v - 1]
    Pred = [None] * len(Succ)
    for u in range(1, n + 1):
        for w in Succ[Head[u - 1]:Head[u]]:
            Pred[free[w]] = u
            free[w] += 1
    return PHead, Pred

PHead, Pred = predecessor_lists(4, [0, 0, 2, 6, 9], [2, 1, 1, 4, 3, 2, 1, 2, 3])
print(PHead, Pred)           # [0, 3, 6, 8, 9] [2, 3, 4, 2, 3, 4, 3, 4, 3]
print(Pred[PHead[1]:PHead[2]])   # [2, 3, 4], the predecessors of 2

Each of the loops runs over the vertices or over the arcs once, so the running time is O(n+m)O(n + m).

Matrices

Definition 4.43 (Matrix).

An r×sr \times s matrix is a table of numbers with rr rows and ss columns. Its entry in row ii and column jj is written MijM_{ij}; rows run across and columns down. An n×nn \times n matrix is square, and the identity matrix InI_n is the square matrix with 11 on its diagonal and 00 elsewhere.

Two matrices of the same shape are added entry by entry. If BB is r×sr \times s and CC is s×ts \times t, their product BCBC is the r×tr \times t matrix with

(BC)ij=∑ℓ=1sBiℓ Cℓj,(BC)_{ij} = \sum_{\ell=1}^{s} B_{i\ell}\, C_{\ell j},

and the powers of a square matrix are M0=InM^0 = I_n and Mk+1=MkMM^{k+1} = M^k M. This is a definition, not ordinary entrywise multiplication. For example,

(1101)(1101)=(1201),\begin{pmatrix} 1 & 1 \\ 0 & 1 \end{pmatrix} \begin{pmatrix} 1 & 1 \\ 0 & 1 \end{pmatrix} = \begin{pmatrix} 1 & 2 \\ 0 & 1 \end{pmatrix},

because the top-right entry is 1⋅1+1⋅1=21 \cdot 1 + 1 \cdot 1 = 2. We use only this rule and ordinary arithmetic below. Computing one entry of the product of two n×nn \times n matrices takes nn multiplications, so the whole product takes O(n3)O(n^3) arithmetic operations.

Adjacency Matrices

For a directed 11-graph with vertices 1,…,n1, \ldots, n, the adjacency matrix MM is the n×nn \times n matrix with

Mij={1,(i,j)∈A,0,(i,j)∉A.M_{ij} = \begin{cases} 1, & (i, j) \in A, \\ 0, & (i, j) \notin A . \end{cases}

For a simple undirected graph we put Mij=1M_{ij} = 1 when {i,j}\{i, j\} is an edge. Then Mij=MjiM_{ij} = M_{ji}, so MM is symmetric, and one can store only the entries on and above the diagonal and recover the rest by reflection; this uses about half as many entries, but still Θ(n2)\Theta(n^2) space. For a directed multigraph one may instead put the number of parallel arcs into MijM_{ij}.

Example 4.44 (An adjacency matrix).

The graph GG of Figure 4.6 has

M=(0000110011111110).M = \begin{pmatrix} 0 & 0 & 0 & 0 \\ 1 & 1 & 0 & 0 \\ 1 & 1 & 1 & 1 \\ 1 & 1 & 1 & 0 \end{pmatrix}.

The rows match the adjacency lists: the third row has four ones, one for each successor of 33.

A matrix stores n2n^2 entries. Testing whether an arc exists takes one entry lookup; scanning the successors of a vertex takes a row scan of nn entries; and scanning all arcs takes O(n2)O(n^2) time, however few arcs there are.

A second matrix records which vertices lie on which edges.

Definition 4.45 (Incidence matrix).

Let GG be a graph without loops, with vertices 1,…,n1, \ldots, n and edges e1,…,eme_1, \ldots, e_m. Its incidence matrix is the n×mn \times m matrix NN with, for an undirected graph,

Nij={1,i is an end of ej,0,otherwise,N_{ij} = \begin{cases} 1, & i \text{ is an end of } e_j, \\ 0, & \text{otherwise}, \end{cases}

and for a directed graph Nij=−1N_{ij} = -1 if the arc eje_j begins at ii, Nij=1N_{ij} = 1 if it ends at ii, and Nij=0N_{ij} = 0 otherwise.

Example 4.46 (An incidence matrix).

Number the edges of the graph of Figure 4.3 as e1={1,2}e_1 = \{1, 2\}, e2={1,5}e_2 = \{1, 5\}, e3={2,5}e_3 = \{2, 5\}, e4={3,6}e_4 = \{3, 6\}. Its incidence matrix, with rows for the vertices 1,…,61, \ldots, 6, is

N=(110010100001000001100001).N = \begin{pmatrix} 1 & 1 & 0 & 0 \\ 1 & 0 & 1 & 0 \\ 0 & 0 & 0 & 1 \\ 0 & 0 & 0 & 0 \\ 0 & 1 & 1 & 0 \\ 0 & 0 & 0 & 1 \end{pmatrix}.

Every column has exactly two ones, the two ends of its edge, and the sum of the entries of row ii is the degree d(i)d(i). Adding all the entries both ways is the count of edge-ends.

The memory requirements of the adjacency matrix and the incidence matrix are therefore Θ(n2)\Theta(n^2) and Θ(nm)\Theta(nm). For graphs with Θ(n)\Theta(n) edges, which is often the case, this is far more than required: adjacency lists need memory proportional to n+mn + m.

For most purposes an adjacency list is the preferred data structure. It allows δ(v)\delta(v), or δ+(v)\delta^+(v) and δ−(v)\delta^-(v), to be scanned for every vertex vv in time linear in their size, and its memory requirement is proportional to the number of vertices and edges, if we assume, as usual, that references, vertex numbers and edge numbers each need only a constant amount of memory. In the Word-RAM this is the assumption that one machine word holds any of them, which w>log⁡2nw > \log_2 n allows for vertex numbers. Graph algorithms are therefore stated for this representation, and moving to the next edge in a list of edges, or reading an endpoint of an edge, is counted as an elementary operation. An adjacency matrix remains the better choice for a graph with very many edges, or when many single-arc tests are needed.

Proposition 4.47 (Powers of the adjacency matrix count walks).

Let MM be the adjacency matrix of a directed 11-graph, and k∈N0k \in \mathbb{N}_0. Then (Mk)ij(M^k)_{ij} is the number of walks of length kk from ii to jj.

Discussion.

The definition of the power is a recursion on kk, so the proof is an induction on kk. For the step, a walk of length k+1k+1 from ii to jj is a walk of length kk from ii to some vertex ℓ\ell followed by one arc from ℓ\ell to jj, and ℓ\ell, the second-to-last vertex, is determined by the walk. So the walks can be sorted by ℓ\ell, and the number through ℓ\ell is the number of kk-walks from ii to ℓ\ell times MℓjM_{\ell j}, which is 11 or 00 according as the last arc exists. Summing over ℓ\ell is exactly the formula for the entry of MkMM^k M. Walks, not paths, are counted: vertices may repeat.

Proof.

For k=0k = 0, InI_n has 11 in position (i,i)(i, i) and 00 elsewhere, and there is exactly one walk of length 00 from each vertex to itself and none to any other vertex.

Suppose the claim holds for kk. Every walk ⟨i=v0,…,vk,vk+1=j⟩\langle i = v_0, \ldots, v_k, v_{k+1} = j \rangle of length k+1k + 1 has a unique second-to-last vertex ℓ=vk\ell = v_k, and consists of a walk of length kk from ii to ℓ\ell followed by the arc (ℓ,j)(\ell, j). For fixed ℓ\ell there are (Mk)iℓ(M^k)_{i\ell} choices of the first part, by the hypothesis, and MℓjM_{\ell j} choices of the last arc, so (Mk)iℓMℓj(M^k)_{i\ell} M_{\ell j} walks pass through ℓ\ell last. Summing over ℓ\ell,

#{walks of length k+1 from i to j}=∑ℓ=1n(Mk)iℓ Mℓj=(MkM)ij=(Mk+1)ij.\#\{\text{walks of length } k+1 \text{ from } i \text{ to } j\} = \sum_{\ell=1}^{n} (M^k)_{i\ell}\, M_{\ell j} = (M^k M)_{ij} = (M^{k+1})_{ij} .

The directed graph 1→2→11 \to 2 \to 1 has adjacency matrix (0110)\begin{pmatrix} 0 & 1 \\ 1 & 0 \end{pmatrix}. Its square is I2I_2: there is one walk of length two from each vertex back to itself, and none to the other vertex.

Reachability from Matrix Powers

To turn walk counts into yes-or-no answers, let sgn⁡(x)=0\operatorname{sgn}(x) = 0 when x=0x = 0 and 11 when x>0x > 0, applied to every entry of a matrix.

Definition 4.48 (Transitive closure).

For a directed graph on nn vertices with adjacency matrix MM, the reflexive transitive closure is

R=sgn⁡(I+M+M2+⋯+Mn−1),R = \operatorname{sgn}\bigl(I + M + M^2 + \cdots + M^{n-1}\bigr),

and the transitive closure is

T=sgn⁡(M+M2+⋯+Mn).T = \operatorname{sgn}\bigl(M + M^2 + \cdots + M^{n}\bigr).

By the proposition, Rij=1R_{ij} = 1 exactly when jj is reachable from ii, the walk of length zero included: a shortest walk to a different vertex is a simple path and uses at most n−1n - 1 arcs. Likewise Tij=1T_{ij} = 1 exactly when there is a walk of positive length from ii to jj, and T=sgn⁡(MR)T = \operatorname{sgn}(MR). The power MnM^n matters on the diagonal, since a shortest closed walk of positive length may be a cycle of length nn. Thus TiiT_{ii} need not be 11, while Rii=1R_{ii} = 1 always.

In the two-vertex graph 1→21 \to 2 with no return arc,

M=(0100),R=(1101),T=(0100),M = \begin{pmatrix} 0 & 1 \\ 0 & 0 \end{pmatrix}, \qquad R = \begin{pmatrix} 1 & 1 \\ 0 & 1 \end{pmatrix}, \qquad T = \begin{pmatrix} 0 & 1 \\ 0 & 0 \end{pmatrix},

and the difference on the diagonal is exactly the zero-length walk.

Computing RR from its definition takes n−2n - 2 matrix products; repeated squaring needs far fewer.

Proposition 4.49 (Closure by repeated squaring).

For every q∈Nq \in \mathbb{N},

∏j=0q−1(I+M2j)=I+M+M2+⋯+M2q−1.\prod_{j=0}^{q-1} \bigl(I + M^{2^j}\bigr) = I + M + M^2 + \cdots + M^{2^q - 1} .

Discussion.

Multiplying out the product, each term chooses from every factor either II or M2jM^{2^j}, and the chosen powers multiply to MM raised to a sum of distinct powers of two. The claim is that every exponent 0,…,2q−10, \ldots, 2^q - 1 arises exactly once, which is the existence and uniqueness of the binary expansion of a number below 2q2^q. The proof is an induction on qq: multiplying the sum up to M2q−1M^{2^q - 1} by I+M2qI + M^{2^q} keeps the sum and adds a copy shifted up by 2q2^q, which fills in the exponents 2q,…,2q+1−12^q, \ldots, 2^{q+1} - 1.

Proof.

For q=1q = 1 both sides are I+MI + M. Suppose the identity holds for qq. Since powers of MM commute with one another,

∏j=0q(I+M2j)=(∑e=02q−1Me)(I+M2q)=∑e=02q−1Me+∑e=02q−1Me+2q=∑e=02q+1−1Me.\prod_{j=0}^{q} \bigl(I + M^{2^j}\bigr) = \Bigl(\sum_{e=0}^{2^q - 1} M^{e}\Bigr)\bigl(I + M^{2^q}\bigr) = \sum_{e=0}^{2^q - 1} M^{e} + \sum_{e=0}^{2^q - 1} M^{e + 2^q} = \sum_{e=0}^{2^{q+1} - 1} M^{e} .

Let q=⌈log⁡2n⌉q = \lceil \log_2 n \rceil, so that 2q−1⩾n−12^q - 1 \geqslant n - 1. Powers beyond Mn−1M^{n-1} reveal no new reachable vertex, so R=sgn⁡(∏j<q(I+M2j))R = \operatorname{sgn}\bigl(\prod_{j<q}(I + M^{2^j})\bigr), and we may replace every positive entry by 11 after each product. The same result follows by working throughout with Boolean matrix addition (OR) and multiplication (AND followed by OR). For n⩾2n \geqslant 2 there are q−1q - 1 squarings to form M2,M4,…,M2q−1M^2, M^4, \ldots, M^{2^{q-1}} and q−1q - 1 products to combine the qq factors: at most 2q−22q - 2 matrix products in total. One more multiplication by MM gives TT. For n=1n = 1, R=IR = I and no product is needed. For comparison, MKM^K takes K−1K - 1 products by repeated multiplication when K⩾1K \geqslant 1, and at most ⌊log⁡2K⌋+popcount⁡(K)−1\lfloor \log_2 K \rfloor + \operatorname{popcount}(K) - 1 by squaring, where popcount⁡(K)\operatorname{popcount}(K) is the number of digits 11 in the binary expansion of KK. These counts are of matrix products; each product of two n×nn \times n matrices takes O(n3)O(n^3) arithmetic operations.

Reflexive Transitive Closure
Input:  the adjacency matrix M of a directed graph on n vertices.
Output: the matrix R with R[i][j] = 1 exactly when j is reachable from i.

    R ≝ sgn(I + M);  P ≝ M;  k ≝ 2
    while k < n do
        P ≝ sgn(P · P)
        R ≝ sgn(R · (I + P))
        k ≝ 2k
    return R

At each test of the loop, P=sgn⁡(Mk/2)P = \operatorname{sgn}(M^{k/2}) and R=sgn⁡(I+M+⋯+Mk−1)R = \operatorname{sgn}(I + M + \cdots + M^{k-1}), by the proposition; the loop stops once k⩾nk \geqslant n, when RR covers every power up to Mn−1M^{n-1}. In Python a matrix is a list of rows, each row a list, and M[i][j] is the entry in row i and column j, numbered from 00.

def product(X, Y):              # Boolean product, O(n^3)
    n = len(X)
    Z = [[0] * n for _ in range(n)]
    for i in range(n):
        for j in range(n):
            for l in range(n):
                if X[i][l] == 1 and Y[l][j] == 1:
                    Z[i][j] = 1
    return Z

def plus_identity(X):           # sgn(I + X)
    n = len(X)
    return [[1 if i == j else X[i][j] for j in range(n)] for i in range(n)]

def closure(M):
    n = len(M)
    R = plus_identity(M)
    P = M
    k = 2
    while k < n:
        P = product(P, P)
        R = product(R, plus_identity(P))
        k = 2 * k
    return R

M = [[0, 1, 0], [0, 0, 1], [0, 0, 0]]   # 1 -> 2 -> 3, numbered 0, 1, 2
print(closure(M))                       # [[1, 1, 1], [0, 1, 1], [0, 0, 1]]

The inner [0] * n must be built afresh for each row, which is what the comprehension does: [[0] * n] * n would make nn references to one and the same row.

Weight Matrices

For a weighted directed graph G=(S,A,ν)G = (S, A, \nu) with vertices s1,…,sns_1, \ldots, s_n, the weight matrix is

Wij={ν(si,sj),(si,sj)∈A,+∞,(si,sj)∉A.W_{ij} = \begin{cases} \nu(s_i, s_j), & (s_i, s_j) \in A, \\ +\infty, & (s_i, s_j) \notin A . \end{cases}

The value +∞+\infty means that there is no direct arc; it is a convention of computation, not an arc of the graph, and WW has entries in R∪{+∞}\mathbb{R} \cup \{+\infty\}. A zero-weight arc is present and must not be confused with an absent arc.

AGBCEF−262−34−1289
Figure 4.9. A weighted directed graph. Each arc carries its weight.

The graph of Figure 4.9 has, with rows and columns in the order A,B,C,E,F,GA, B, C, E, F, G,

W=ABCEFGA∞6∞∞∞−2B∞∞∞∞∞2C∞4∞∞−1∞E∞28∞9∞F∞∞∞∞∞∞G∞∞−3∞∞∞W = \begin{array}{c|rrrrrr} & A & B & C & E & F & G \\ \hline A & \infty & 6 & \infty & \infty & \infty & -2 \\ B & \infty & \infty & \infty & \infty & \infty & 2 \\ C & \infty & 4 & \infty & \infty & -1 & \infty \\ E & \infty & 2 & 8 & \infty & 9 & \infty \\ F & \infty & \infty & \infty & \infty & \infty & \infty \\ G & \infty & \infty & -3 & \infty & \infty & \infty \end{array}

Reading row EE: the arcs E→BE \to B, E→CE \to C and E→FE \to F have weights 22, 88 and 99. The entry WEA=∞W_{EA} = \infty says that there is no arc E→AE \to A; it says nothing about longer routes. For an undirected weighted graph the weight matrix is symmetric. A diagonal entry is ∞\infty unless a loop is present; a distance matrix, which puts 00 on the diagonal, is a different object.

Problem 4.20.

For the graph of Figure 4.8, write down the adjacency lists, the arrays Head\mathrm{Head} and Succ\mathrm{Succ}, and the adjacency matrix MM. Compute M2M^2 and M3M^3, and check the entries (M2)15(M^2)_{15} and (M3)11(M^3)_{11} against walks you can list.

Problem 4.21.

In a directed graph a universal sink is a vertex of in-degree n−1n - 1 and out-degree 00. Give an algorithm that decides whether a graph given by its adjacency matrix has a universal sink, using O(n)O(n) matrix lookups.

Exercises on Data Structures

Exercise 4.1.

An increasing subarray of an array of integers is a run of consecutive entries whose values strictly increase. Write a Python function count_long_subarrays(A) which takes a tuple A=(a0,a1,…,an−1)A = (a_0, a_1, \ldots, a_{n-1}) of n>0n > 0 positive integers and returns the number of longest increasing subarrays of AA, that is, the number of increasing subarrays whose length is at least that of every other. For A=(1,3,4,2,7,5,6,9,8)A = (1, 3, 4, 2, 7, 5, 6, 9, 8) it should return 22, since the longest increasing subarrays have length three and there are two of them, (1,3,4)(1, 3, 4) and (5,6,9)(5, 6, 9). Your function should run in O(n)O(n) time.

Exercise 4.2.

Order the following functions so that if faf_a appears before fbf_b then fa=O(fb)f_a = O(f_b), and indicate which pairs satisfy both fa=O(fb)f_a = O(f_b) and fb=O(fa)f_b = O(f_a). Here log⁡\log means log⁡2\log_2.

f1=log⁡(nn),f2=(log⁡n)n,f3=log⁡(n6006),f4=(log⁡n)6006,f5=log⁡log⁡(6006n).f_1 = \log(n^n), \qquad f_2 = (\log n)^n, \qquad f_3 = \log(n^{6006}), \qquad f_4 = (\log n)^{6006}, \qquad f_5 = \log\log(6006n).

Exercise 4.3.

Let ff and gg be functions N→R⩾0\mathbb{N} \to \mathbb{R}_{\geqslant 0}. Using the definition of Θ\Theta, prove that max⁡(f(n),g(n))=Θ(f(n)+g(n))\max\bigl(f(n), g(n)\bigr) = \Theta\bigl(f(n) + g(n)\bigr).

Exercise 4.4.

A data structure DD supports the sequence operations D.build(X) in O(n)O(n) time, and D.insert_at(i, x) and D.delete_at(i) each in O(log⁡n)O(\log n) time, where nn is the number of items stored at the time of the operation. Using only these operations, describe algorithms for the following, each running in O(klog⁡n)O(k \log n) time. Recall that delete_at returns the deleted item.

  1. reverse(D, i, k): reverse the order of the kk items of DD starting at index ii, that is, those at indices ii to i+k−1i + k - 1.
  2. move(D, i, k, j): move the kk items of DD starting at index ii, in order, to be in front of the item at index jj, where i⩽j<i+ki \leqslant j < i + k is false.

Exercise 4.5.

Each node x of a doubly linked list keeps a reference x.prev to the node before it as well as x.next to the node after it, and the list L keeps L.head and L.tail, its first and last nodes. The list does not store its length.

  1. Describe algorithms for insert_first(x), insert_last(x), delete_first() and delete_last(), each in O(1)O(1) time.
  2. Given two nodes x1 and x2 of a list L, with x1 before x2, describe a constant-time algorithm that removes all nodes from x1 to x2 inclusive from L and returns them as a new doubly linked list.
  3. Given a node x of a list L1 and a second list L2, describe a constant-time algorithm that splices L2 into L1 after x, leaving L2 empty.
  4. Implement these operations in Python, in a class built like Linked_List_Seq.

Exercise 4.6.

A student keeps nn pages of notes in a binder, the first at index 00 and the last at index n−1n - 1, and has two bookmarks AA and BB. Describe a data structure supporting the following operations, where nn is the number of pages at the time of the operation. Assume both bookmarks are placed before any shift or move, and that AA is always at a lower index than BB. For each operation, say whether your running time is worst-case or amortized.

OperationEffectTime
build(X)initialise with the pages of the iterable XO(∣X∣)O(\lvert X \rvert)
place_mark(i, m)place bookmark m∈{A,B}m \in \{A, B\} between the pages at indices ii and i+1i+1O(n)O(n)
read_page(i)return the page at index iiO(1)O(1)
shift_mark(m, d)move bookmark mm, in front of the page at index ii, to be in front of the page at index i+di + d, for d∈{−1,1}d \in \{-1, 1\}O(1)O(1)
move_page(m)move the page in front of bookmark mm to be in front of the other bookmarkO(1)O(1)

Exercise 4.7.

  1. Insert the integer keys 47,61,36,52,56,33,9247, 61, 36, 52, 56, 33, 92 in this order into a hash table of size 77 using the hash function h(k)=(10k+4) mod 7h(k) = (10k + 4) \bmod 7. Each slot stores a linked list of the keys hashing to it, later insertions being appended at the end. Draw the table after all keys have been inserted.
  2. Suppose instead h(k)=((10k+4) mod c) mod 7h(k) = \bigl((10k + 4) \bmod c\bigr) \bmod 7 for a positive integer cc. Find the smallest cc for which no collisions occur when inserting these keys.

Exercise 4.8.

A university assigns 2n2n new students to nn rooms, numbered 00 to n−1n - 1, by hashing their IDs. Each ID is a positive integer less than uu, with uu much larger than 2n2n; no two students have the same ID, and students choose their own IDs. The university publishes a family H\mathcal{H} of hash functions before IDs are chosen, and afterwards chooses the rooming function uniformly from H\mathcal{H}. Two students want to be roommates. For each family below, either show that they can choose IDs k1,k2k_1, k_2 that guarantee it, or prove that no choice guarantees it and find the highest probability of being roommates they can achieve.

  1. H={ hab(k)=(ak+b) mod n  :  a,b∈{0,…,n−1}, a≠0 }\mathcal{H} = \bigl\{\, h_{ab}(k) = (ak + b) \bmod n \;:\; a, b \in \{0, \ldots, n-1\},\ a \neq 0 \,\bigr\}.
  2. H={ ha(k)=(⌊kn/u⌋+a) mod n  :  a∈{0,…,u−1} }\mathcal{H} = \bigl\{\, h_{a}(k) = \bigl(\lfloor kn/u \rfloor + a\bigr) \bmod n \;:\; a \in \{0, \ldots, u-1\} \,\bigr\}.

Exercise 4.9.

A wall is lined with nn boxes of paper, box ii standing ii feet from the left end and containing bib_i reams, where the bib_i are distinct positive integers. A pair of boxes (bi,bj)(b_i, b_j) is close if ∣i−j∣<n/10\lvert i - j \rvert < n/10, and it fulfils an order of rr reams if bi+bj=rb_i + b_j = r. Given B=(b0,…,bn−1)B = (b_0, \ldots, b_{n-1}) and rr, describe an algorithm running in expected O(n)O(n) time that decides whether BB contains a close pair fulfilling the order.

Exercise 4.10.

Imagine inserting the keys 0,1,2,…,n0, 1, 2, \ldots, n, in that order, into a hash table of size 77 that resolves collisions by chaining, with h(k)=k mod 7h(k) = k \bmod 7. Draw the table after the insertion of the keys up to n=9n = 9. Explain how the table evolves for arbitrary nn, and derive the worst-case time for the whole operation of inserting the n+1n + 1 keys.

Exercises on Graphs

Exercise 4.11.

Construct a 33-regular graph on 88 vertices. Is there a 33-regular graph on 99 vertices?

Exercise 4.12.

Show that every graph whose average degree is dd has a subgraph in which every vertex has degree at least d/2d/2.

Exercise 4.13.

Let GG be a graph with no loops in which every vertex has the same odd degree kk. Show that the number of edges is a multiple of kk and that the number of vertices is even.

Exercise 4.14.

  1. Can 77 line segments be drawn in the plane so that each intersects exactly 55 others? Prove your answer.
  2. Suppose a simple graph has exactly two vertices of odd degree. Prove that there is a path between them.

Check Yourself

 

Fresh questions on the whole lesson — none of them is worked out above. Work each one out on paper before opening Python; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 4.15.

Which class contains 3n2+10 nlog⁡2n3n^2 + 10\, n \log_2 n?

answer one of these

Exercise 4.16.

What is the least number of comparisons that finds the maximum of 1010 distinct numbers in the worst case?

answer one of these

Exercise 4.17.

How many numbers does the sieve of Eratosthenes output for n=30n = 30?

answer one of these

Exercise 4.18.

Starting from an empty Dynamic_Array_Seq with r=2r = 2, how many slots are allocated after 1717 calls of insert_last?

answer one of these

Exercise 4.19.

What is the worst-case cost of get_at(i) in Linked_List_Seq with nn items?

answer one of these

Exercise 4.20.

Twenty-five objects are placed in seven boxes. What is the largest number that is certain to be in some single box?

answer one of these

Exercise 4.21.

Ten distinct keys are stored with chaining in a table of 55 slots, the hash function drawn from a universal family. What bound does the proposition on chain length give for the expected length of the chain holding a given stored key?

answer one of these

Exercise 4.22.

A simple graph has 1212 edges. What is the sum of its vertex degrees?

answer one of these

Exercise 4.23.

How many edges does K7K_7 have?

answer one of these

Exercise 4.24.

How many edges does K3,4K_{3,4} have?

answer one of these

Exercise 4.25.

For the directed graph with arcs 1→21 \to 2 and 2→12 \to 1 and adjacency matrix MM, what is the entry (M3)12(M^3)_{12}?

answer one of these

Lesson 5

Trees, Searching and Sorting

Taught

Connectivity and Trees

Connectivity

Definition 5.1 (Connected graph).

An undirected graph is connected if every vertex is reachable from every other vertex. By the proposition on walks and simple paths, equivalently, every pair of distinct vertices is joined by a simple path.

abcdefg
Figure 5.1. A graph with two components, {a,b,c,d}\{a, b, c, d\} and {e,f,g}\{e, f, g\}. The accented edge {e,f}\{e, f\} is a bridge.

The graph of Figure 5.1 is disconnected: no path leads from aa to ee. Its vertices split into {a,b,c,d}\{a, b, c, d\} and {e,f,g}\{e, f, g\}.

Definition 5.2 (Connected component).

A connected component is a maximal connected induced subgraph: no further vertex of the graph can be included while keeping it connected. A graph is connected exactly when it has one component, and an isolated vertex forms a component by itself.

Theorem 5.3 (Edges of a connected graph).

A connected simple undirected graph with n⩾1n \geqslant 1 vertices and mm edges has m⩾n−1m \geqslant n - 1.

Discussion.

The first proof grows the graph from one vertex. At each stage connectivity supplies an edge leaving the part reached so far, and that edge brings in exactly one new vertex, so reaching all nn vertices uses n−1n - 1 distinct edges of the graph.

The second proof is an induction on nn that removes a vertex of smallest degree kk. If k⩾2k \geqslant 2, counting edge-ends gives m⩾nm \geqslant n directly and no induction is needed. If k=1k = 1, removing a vertex of degree one loses exactly one edge, and what is left is still connected, because a vertex of degree one cannot lie in the middle of a path. The hypothesis then applies to the smaller graph. The hypothesis has to be stated for every connected graph on n−1n - 1 vertices, since it is applied to a graph that the proof constructs.

Proof.

Choose one vertex and mark it reached. If fewer than nn vertices are reached, connectivity guarantees an edge from some reached vertex to an unreached one; otherwise no path could lead out of the reached set. Mark that new vertex and keep the connecting edge. Each step reaches exactly one new vertex and keeps an edge not kept before. After n−1n - 1 steps all nn vertices are reached, and we have found n−1n - 1 distinct edges of GG. Therefore m⩾n−1m \geqslant n - 1.

A second proof, by induction, removes a vertex at each step.

Proof.

For each integer n⩾1n \geqslant 1, let P(n)P(n) be the statement: every connected simple undirected graph with nn vertices has at least n−1n - 1 edges.

Base case, P(1)P(1). A simple graph with one vertex has no pair of distinct vertices to join, so it has m=0m = 0 edges. It is connected, and 0=1−10 = 1 - 1.

Step. Fix n⩾2n \geqslant 2 and assume P(n−1)P(n-1). Take an arbitrary connected simple graph G=(S,A)G = (S, A) with ∣S∣=n|S| = n and ∣A∣=m|A| = m, and let k=min⁡v∈Sd(v)k = \min_{v \in S} d(v) be its smallest degree. Because n⩾2n \geqslant 2, a vertex of degree zero would have no path to any other vertex, contradicting connectivity, so k⩾1k \geqslant 1. Either k=1k = 1 or k⩾2k \geqslant 2.

Case k⩾2k \geqslant 2. Every one of the nn vertices has degree at least 22, so ∑v∈Sd(v)⩾2n\sum_{v \in S} d(v) \geqslant 2n. Counting edge-ends gives ∑v∈Sd(v)=2m\sum_{v \in S} d(v) = 2m. Therefore 2m⩾2n2m \geqslant 2n, so m⩾n⩾n−1m \geqslant n \geqslant n - 1. This case needs neither the removal of a vertex nor P(n−1)P(n-1).

Case k=1k = 1. Some vertex vv has exactly one edge, say {v,w}\{v, w\}. Remove vv and that edge, and call the result G′G'. No other edge is removed, because no other edge touches vv, so G′G' has n−1n - 1 vertices and m−1m - 1 edges.

Before applying P(n−1)P(n-1) we check that G′G' is connected. Take distinct vertices a,ba, b of G′G'. Since GG is connected, a simple path joins aa to bb in GG. Neither endpoint is vv, because aa and bb remain. Nor can vv be an internal vertex of the path: the path would have to enter vv along one edge and leave along a second, distinct edge, and vv has only the edge {v,w}\{v, w\}. So the path uses neither vv nor its edge, and it is a path in G′G'. This holds for every pair a,ba, b, so G′G' is connected; when n=2n = 2, G′G' has a single vertex and is connected by definition.

Now G′G' satisfies the hypotheses of P(n−1)P(n-1): it is simple, connected, and has n−1n - 1 vertices. Hence m−1⩾(n−1)−1=n−2m - 1 \geqslant (n - 1) - 1 = n - 2, and adding 11 to both sides gives m⩾n−1m \geqslant n - 1. Both cases prove P(n)P(n) from P(n−1)P(n-1), and with P(1)P(1) induction proves the theorem for every n⩾1n \geqslant 1.

A connected graph with five vertices needs at least four edges, and a path through all five vertices attains the bound. Adding one edge to that path gives five edges and creates a cycle; the theorem does not claim that every connected graph has exactly n−1n - 1 edges.

Definition 5.4 (Articulation vertex, bridge and vertex cut).

An articulation vertex is a vertex whose deletion, together with its edges, increases the number of connected components. A bridge, also called an isthmus, is an edge whose deletion increases that number. A vertex cut of a connected graph is a set of vertices whose deletion disconnects it, provided at least two vertices remain afterwards.

In Figure 5.1, the edge {e,f}\{e, f\} is a bridge: deleting it separates ee from ff and gg. The vertex ff is an articulation vertex: deleting it leaves ee and gg in separate components. All three edges of the left component are bridges too, since that component is a path through four vertices.

Strong Connectivity

For a directed graph, weak connectivity means connectivity once the directions of the arcs are ignored. Requiring directed routes both ways is a stronger condition.

Definition 5.5 (Strong connectivity and the reduced graph).

A directed graph is strongly connected if for every u,vu, v there is a directed walk from uu to vv and one from vv to uu. A strongly connected component is a maximal strongly connected induced subgraph. The reduced graph, or condensation, has one vertex for each strongly connected component, and an arc C→DC \to D whenever C≠DC \neq D and some arc of the original graph goes from a vertex of CC to a vertex of DD.

abcdabcd
Figure 5.2. The left graph is strongly connected. In the right graph no arc returns towards dd or leaves cc, and each vertex is a strongly connected component by itself.

In Figure 5.2, the left graph is strongly connected. In the right one every arc points away from dd and towards cc, so each of its four vertices is a strongly connected component by itself.

abcdefg⟶C1C2
Figure 5.3. A directed graph whose strongly connected components are {a,b,c,d}\{a, b, c, d\} and {e,f,g}\{e, f, g\}, and its reduced graph. The accented arcs are the ones crossing between components.

In Figure 5.3, {a,b,c,d}\{a, b, c, d\} and {e,f,g}\{e, f, g\} are each strongly connected, and the only arcs between them point from the first set to the second. The reduced graph has two vertices C1,C2C_1, C_2 and one arc C1→C2C_1 \to C_2. There cannot be an arc back: that would give reachability both ways and merge the two into one component.

Proposition 5.6 (The reduced graph has no directed cycle).

The reduced graph of any directed graph contains no simple directed cycle.

Discussion.

A cycle through several components would let every vertex in any one of them reach every vertex in the others and come back, so all of them would lie in a single strongly connected component. That contradicts maximality. The proof replaces each arc of the reduced graph by a walk in the original graph.

Proof.

The reduced graph has no loops, since its arcs join distinct components. Suppose C1→C2→⋯→Ck→C1C_1 \to C_2 \to \cdots \to C_k \to C_1 were a simple directed cycle of it with k⩾2k \geqslant 2. Each arc Ci→Ci+1C_i \to C_{i+1} comes from an arc of the original graph from some vertex of CiC_i to some vertex of Ci+1C_{i+1}, and within each component any two vertices are joined by walks both ways. Concatenating, any vertex of C1C_1 can reach any vertex of C2C_2, and any vertex of C2C_2 can reach any vertex of C1C_1 around the rest of the cycle. So C1∪C2C_1 \cup C_2 induces a strongly connected subgraph larger than C1C_1, contradicting the maximality of C1C_1.

Problem 5.1.

Show that a simple undirected graph on nn vertices with more than (n−12)\binom{n-1}{2} edges is connected, and give a disconnected graph with exactly (n−12)\binom{n-1}{2} edges.

Problem 5.2.

Find the strongly connected components and the reduced graph of the graph in Figure 4.8 of Lesson 4.

Trees

Definition 5.7 (Tree and forest).

A tree is a connected acyclic undirected graph. A forest is an undirected graph whose components are trees.

abcdefgabcdefg
Figure 5.4. A tree on seven vertices (left), and the out-arborescence obtained by orienting its edges away from aa (right).

The tree on the left of Figure 5.4 has seven vertices and six edges.

Proposition 5.8 (Counting the edges of a forest).

An acyclic undirected graph with nn vertices, mm edges and cc components satisfies n=m+cn = m + c.

Discussion.

The proof is an induction on the number of edges, and the step removes one. In an acyclic graph every edge is a bridge: if its endpoints were still joined after deleting it, that route together with the edge would be a cycle. So deleting an edge lowers mm by one and raises cc by one, and m+cm + c is unchanged. The base case, a graph with no edges, has every vertex as its own component.

Proof.

For m=0m = 0 every vertex is its own component, so c=nc = n and n=m+cn = m + c. Assume the formula for all acyclic graphs with m−1m - 1 edges, and take one with m>0m > 0 edges. Delete one edge. Its endpoints cannot remain joined by another path, or that path together with the deleted edge would be a cycle. Hence the component containing the edge splits in two and the number of components rises from cc to c+1c + 1. The remaining graph is acyclic with m−1m - 1 edges, so the hypothesis gives n=(m−1)+(c+1)=m+cn = (m - 1) + (c + 1) = m + c.

In particular a forest with nn vertices and cc components has n−cn - c edges.

Theorem 5.9 (Six characterisations of a tree).

Let GG be a finite simple undirected graph with n⩾1n \geqslant 1 vertices. The following are equivalent.

  1. GG is connected and acyclic.
  2. GG is acyclic and has n−1n - 1 edges.
  3. GG is connected and has n−1n - 1 edges.
  4. GG is acyclic, and adding any missing edge creates exactly one simple cycle.
  5. GG is connected, and deleting any edge disconnects it.
  6. Every pair of distinct vertices is joined by exactly one simple path.

In particular each condition characterises a tree. For K1K_1 the clauses about adding or deleting an edge are vacuous.

Discussion.

Six conditions are proved equivalent by a cycle of implications, 1⇒6⇒5⇒4⇒2⇒3⇒11 \Rightarrow 6 \Rightarrow 5 \Rightarrow 4 \Rightarrow 2 \Rightarrow 3 \Rightarrow 1, so that each follows from each. The proof uses two facts. One is the relation between a cycle and two different paths between the same vertices: a cycle gives two routes around it, and two different paths give a cycle where they separate and rejoin. The other is the edge count n=m+cn = m + c for acyclic graphs, which converts between “has n−1n - 1 edges” and “has one component”. The last implication, 3⇒13 \Rightarrow 1, uses the growing argument from the edge bound for connected graphs: it produces n−1n - 1 edges forming a connected acyclic spanning subgraph, and a graph with only n−1n - 1 edges has no others.

Proof.

1⇒61 \Rightarrow 6. Connectivity supplies a simple path between every pair. If two different simple paths joined the same pair, follow them from the common start until they first diverge, and then along the first until it next meets the second; the two segments between those meeting points form a cycle. Thus the path is unique.

6⇒56 \Rightarrow 5. GG is connected by 6. The unique simple path between the endpoints of an edge is that edge itself, so deleting it leaves the endpoints with no path between them.

5⇒45 \Rightarrow 4. If GG had a cycle, deleting one of its edges would leave the endpoints of that edge connected along the rest of the cycle, so it would not disconnect GG. So GG is acyclic. By connectivity a simple path joins the endpoints of any missing edge, and adding the edge closes the path into a simple cycle. Now GG is connected and acyclic, so by 1⇒61 \Rightarrow 6 that path is unique, and every simple cycle through the new edge consists of the new edge and a simple path of GG between its endpoints; so the new simple cycle is unique.

4⇒24 \Rightarrow 2. If GG had two components, adding an edge between them would create no cycle, contrary to 4. Thus c=1c = 1, and since GG is acyclic the edge count n=m+cn = m + c gives m=n−1m = n - 1.

2⇒32 \Rightarrow 3. Since GG is acyclic, the edge count gives c=n−m=n−(n−1)=1c = n - m = n - (n - 1) = 1. One component means that GG is connected.

3⇒13 \Rightarrow 1. Because GG is connected, start at one vertex and repeatedly add an edge to a new vertex until every vertex is reached, as in the first proof of the edge bound. The chosen edges form a spanning subgraph with exactly n−1n - 1 edges, one for each vertex other than the first. It is connected, since every vertex was joined to one reached earlier, and acyclic, since each edge was added to a new vertex and so cannot close a cycle among those already reached. Since GG itself has exactly n−1n - 1 edges, it has no others, so GG is this subgraph and is acyclic.

In the tree of Figure 5.4 the unique path from dd to ff is d,b,c,g,fd, b, c, g, f, of length 44. Removing {b,c}\{b, c\} separates it into two components of sizes 33 and 44. Adding the missing edge {d,g}\{d, g\} creates the unique cycle d,b,c,g,dd, b, c, g, d.

Rooted Trees

Definition 5.10 (Out-arborescence).

An out-arborescence rooted at rr is a directed graph in which there is exactly one directed simple path from rr to each vertex, and whose underlying undirected graph is a tree. Equivalently, rr has in-degree zero, every other vertex has in-degree one, and every vertex is reachable from rr. It has n−1n - 1 arcs.

Orienting the tree of Figure 5.4 away from aa, as a→ba \to b, b→cb \to c, b→db \to d, c→ec \to e, c→gc \to g and g→fg \to f, gives the out-arborescence with root aa on the right of the figure.

Definition 5.11 (Rooted and binary trees).

A rooted tree is a tree with one vertex rr chosen as its root. For a vertex v≠rv \neq r, the vertex before vv on the unique path from rr to vv is the parent of vv, and vv is a child of its parent. A vertex with no children is a leaf, and the others are internal. The depth of vv is d(r,v)d(r, v), the length of the path from rr to vv, and the height of the tree is the largest depth of a vertex. The subtree at vv consists of vv and every vertex whose path from rr passes through vv, rooted at vv.

A binary tree is a rooted tree in which every vertex has at most two children, each labelled as a left child or a right child, with at most one of each. It is full if every internal vertex has exactly two children.

Orienting each edge of a rooted tree from parent to child gives an out-arborescence rooted at rr, since the unique path from rr to vv in the tree becomes the unique directed path. The subtrees at the children of the root of a binary tree are again binary trees, of height one less at most.

Proposition 5.12 (How large a binary tree of given height can be).

A binary tree of height at most hh has at most 2h2^h leaves and at most 2h+1−12^{h+1} - 1 vertices.

Discussion.

A binary tree consists of its root and at most two smaller binary trees below it, so both counts are proved by induction on the height. Each subtree at a child of the root has height at most h−1h - 1, so by induction each has at most 2h−12^{h-1} leaves and 2h−12^h - 1 vertices, and there are at most two of them. The leaves of the whole tree are the leaves of the subtrees, unless the root has no children, and the vertices are those of the subtrees plus the root. The recurrence T(h)=2T(h−1)+1T(h) = 2T(h-1) + 1 with T(0)=1T(0) = 1 is the Tower of Hanoi recurrence shifted by one.

Proof.

We use induction on hh. A binary tree of height 00 is a single vertex, which is a leaf: 1=201 = 2^0 leaves and 1=21−11 = 2^1 - 1 vertices.

Let h⩾1h \geqslant 1 and suppose the claim holds for height at most h−1h - 1. If the root has no children the tree has one leaf and one vertex. Otherwise the root is not a leaf, and every other vertex lies in the subtree at exactly one of the at most two children of the root; each such subtree is a binary tree of height at most h−1h - 1. So there are at most 2⋅2h−1=2h2 \cdot 2^{h-1} = 2^h leaves, and at most 1+2 (2h−1)=2h+1−11 + 2\,(2^h - 1) = 2^{h+1} - 1 vertices.

Corollary 5.13 (Height of a binary tree).

A binary tree with NN vertices has height at least ⌈log⁡2(N+1)⌉−1\lceil \log_2 (N+1) \rceil - 1, and a binary tree with KK leaves has height at least ⌈log⁡2K⌉\lceil \log_2 K \rceil.

Proof.

If the height is hh, the proposition gives N⩽2h+1−1N \leqslant 2^{h+1} - 1, so h+1⩾log⁡2(N+1)h + 1 \geqslant \log_2(N + 1), and since h+1h + 1 is an integer, h+1⩾⌈log⁡2(N+1)⌉h + 1 \geqslant \lceil \log_2(N+1) \rceil. Likewise K⩽2hK \leqslant 2^h gives h⩾log⁡2Kh \geqslant \log_2 K, and so h⩾⌈log⁡2K⌉h \geqslant \lceil \log_2 K \rceil.

Problem 5.3.

Show that every tree with at least two vertices has at least two leaves, that is, vertices of degree one, and that a full binary tree with KK leaves has exactly K−1K - 1 internal vertices.

Problem 5.4.

Count the binary trees with 11, 22, 33 and 44 vertices, where two binary trees are the same only if they have the same shape including the left and right labels. Find a recurrence for the number with NN vertices by considering the sizes of the two subtrees at the root.

Searching

To search an array T[0],…,T[n−1]T[0], \ldots, T[n-1] for a key xx is to return an index ii with key⁡(T[i])=x\operatorname{key}(T[i]) = x, or to report that there is none. This is the find(k) operation of the set interface on an array.

Linear search inspects the entries in order and returns the first index whose key equals xx. The entries need not be in any order.

Linear Search
Input:  an array T[0], …, T[n − 1] and a key x.
Output: an index i with key(T[i]) = x, or −1 if there is none.

    for i ≝ 0 to n − 1 do
        if key(T[i]) = x then return i
    return −1
def linear_search(T, x):
    for i in range(len(T)):
        if T[i] == x:
            return i
    return -1

T = [7, 3, 11, 2, 5]
print(linear_search(T, 2), linear_search(T, 4))   # 3 -1

The invariant before the pass for ii is that xx does not occur among T[0],…,T[i−1]T[0], \ldots, T[i-1]. It holds trivially at i=0i = 0, a pass that does not return extends it by one position, and when the loop ends without returning it says that xx does not occur at all. The variant is n−in - i.

Proposition 5.14 (Cost of linear search).

Linear search on an array of nn entries makes at most nn comparisons with xx, and exactly nn when xx does not occur. If xx occurs exactly once, at each of the nn positions with equal likelihood, the mean number of comparisons is (n+1)/2(n+1)/2. So linear search takes Θ(n)\Theta(n) time in the worst case and on average.

Discussion.

Each pass makes one comparison, and the pass for ii is reached only if xx was not among the first ii entries, so the count is the position at which the search stops, plus one. The worst case is an absent key, which runs through all nn passes. For the mean, a key at position ii costs i+1i + 1 comparisons, and averaging 1,2,…,n1, 2, \ldots, n is the sum of an arithmetic progression divided by nn.

Proof.

The loop makes one comparison per pass and at most nn passes, and it makes all nn when no entry has key xx. If xx occurs exactly once, at position ii, the search stops after the (i+1)(i+1)st comparison. Averaging over the nn equally likely positions and using the sum of an arithmetic progression,

1n∑i=0n−1(i+1)=1n⋅n(n+1)2=n+12.\frac{1}{n}\sum_{i=0}^{n-1} (i + 1) = \frac{1}{n} \cdot \frac{n(n+1)}{2} = \frac{n+1}{2} .

Each comparison and each pass cost a bounded number of elementary operations, so the running time is Θ(n)\Theta(n) in both the worst case and the mean.

On an unsorted array no comparison algorithm does better. We prove this with an adversary, as in the second proof of the proposition on finding a maximum.

Proposition 5.15 (Searching an unsorted array).

Every deterministic algorithm that decides whether xx occurs in an array of nn entries, and learns about the entries only by comparing them with xx, compares xx with all nn entries on some input.

Discussion.

An adversary answers every comparison as though the entry compared were larger than xx, and fixes the entries only as far as those answers require. If the algorithm stops before comparing xx with some entry, that entry is still free. The adversary then sets it equal to xx if the algorithm answered “absent”, and to anything else if it answered “present at that entry”, and in either case the answers given stay true while the output becomes wrong.

Proof.

Answer every comparison of xx with an entry T[j]T[j] as x<key⁡(T[j])x < \operatorname{key}(T[j]). Suppose that on these answers the algorithm stops having compared xx with only some of the entries, and let T[k]T[k] be one it never compared. Every compared entry can be given a key larger than xx, consistently with all the answers. If the algorithm reports that xx does not occur, give T[k]T[k] the key xx: the answers are unchanged, so a deterministic algorithm gives the same wrong report. If it reports an index ii, then either T[i]T[i] was compared and has key larger than xx, or i=ki = k and we give T[k]T[k] a key different from xx; in both cases the report is wrong. So on some input every entry is compared with xx.

Problem 5.5.

Suppose that xx is absent with probability 12\tfrac12, and otherwise lies at each of the nn positions with probability 12n\tfrac{1}{2n}. Find the mean number of comparisons made by linear search.

Problem 5.6.

A sentinel search appends xx at the end of the array before searching, so that the loop needs no test of the index against nn. Write it in Python using append and pop, show that it is correct, and count its comparisons of keys and of indices, compared with the linear search above.

When the array is sorted we can search it by bisection, as in Lesson 3, halving a range of positions instead of an interval of reals.

Let T[0],…,T[n−1]T[0], \ldots, T[n-1] be sorted in nondecreasing key order, meaning key⁡(T[0])⩽key⁡(T[1])⩽⋯⩽key⁡(T[n−1])\operatorname{key}(T[0]) \leqslant \operatorname{key}(T[1]) \leqslant \cdots \leqslant \operatorname{key}(T[n-1]). Binary search for a target key xx keeps a candidate interval of positions [L,R)[L, R), that is L⩽i<RL \leqslant i < R. It compares xx with the key at the middle position m=L+⌊(R−L)/2⌋m = L + \lfloor (R - L)/2 \rfloor. If they are equal it returns mm; if xx is smaller it sets R=mR = m; otherwise it sets L=m+1L = m + 1. It returns −1-1 when the interval is empty, L=RL = R.

Binary Search
Input:  an array T[0], …, T[n − 1] sorted by key, and a key x.
Output: an index m with key(T[m]) = x, or −1 if there is none.

    L ≝ 0;  R ≝ n
    while L < R do
        m ≝ L + ⌊(R − L) / 2⌋
        if x = key(T[m]) then return m
        if x < key(T[m]) then R ≝ m
        else L ≝ m + 1
    return −1
def binary_search(T, x):
    L, R = 0, len(T)
    while L < R:
        m = L + (R - L) // 2
        if x == T[m]:
            return m
        if x < T[m]:
            R = m
        else:
            L = m + 1
    return -1

T = [2, 3, 5, 7, 11, 13, 17]
print(binary_search(T, 11), binary_search(T, 4))   # 4 -1

The invariant is: if xx occurs in TT, then at least one matching index lies in [L,R)[L, R). It holds initially, when the interval is the whole array. Sortedness maintains it: if x<key⁡(T[m])x < \operatorname{key}(T[m]) then every position from mm on has key at least key⁡(T[m])>x\operatorname{key}(T[m]) > x and can be discarded, and symmetrically when x>key⁡(T[m])x > \operatorname{key}(T[m]). At termination either a match has been returned or [L,R)[L, R) is empty, and then by the invariant xx does not occur. The length R−LR - L is a variant: it strictly decreases at every pass. We count one comparison per pass, a three-way comparison of xx with key⁡(T[m])\operatorname{key}(T[m]) whose outcomes are smaller, equal and larger.

Theorem 5.16 (Binary search takes O(log⁡n)O(\log n) comparisons).

For n⩾1n \geqslant 1, binary search on a sorted array of length nn makes at most ⌊log⁡2n⌋+1\lfloor \log_2 n \rfloor + 1 passes of its loop.

Discussion.

The proof is the one for the convergence of bisection, with the length of the interval in place of its width. Because the middle position is removed from the interval, a pass leaves at most ⌊ℓ/2⌋\lfloor \ell/2 \rfloor of the ℓ\ell positions rather than exactly half. After tt passes the length is therefore at most ⌊n/2t⌋\lfloor n/2^t \rfloor, which is zero once 2t>n2^t > n, and the least such tt is ⌊log⁡2n⌋+1\lfloor \log_2 n \rfloor + 1.

Proof.

Let ℓ=R−L⩾1\ell = R - L \geqslant 1 at the start of a pass, so that m−L=⌊ℓ/2⌋m - L = \lfloor \ell/2 \rfloor. If the pass does not return, the new interval is [L,m)[L, m), of length ⌊ℓ/2⌋\lfloor \ell/2 \rfloor, or [m+1,R)[m+1, R), of length ℓ−⌊ℓ/2⌋−1=⌈ℓ/2⌉−1⩽⌊ℓ/2⌋\ell - \lfloor \ell/2 \rfloor - 1 = \lceil \ell/2 \rceil - 1 \leqslant \lfloor \ell/2 \rfloor. So after each pass the length is at most ⌊ℓ/2⌋\lfloor \ell / 2 \rfloor, and since ⌊⌊a⌋/2⌋=⌊a/2⌋\bigl\lfloor \lfloor a \rfloor / 2 \bigr\rfloor = \lfloor a/2 \rfloor for a⩾0a \geqslant 0, after tt passes it is at most ⌊n/2t⌋\lfloor n / 2^t \rfloor. For t=⌊log⁡2n⌋+1t = \lfloor \log_2 n \rfloor + 1 we have 2t>n2^t > n, so the length is 00 and the loop has stopped.

The bound is attained, for example by an unsuccessful search for a key smaller than every key in the array: it moves left at every pass, and the length goes n,⌊n/2⌋,⌊n/4⌋,…n, \lfloor n/2 \rfloor, \lfloor n/4 \rfloor, \ldots down to 00.

3150246<>
Figure 5.5. The indices binary search tests on an array of length 77. A search starts at the root and moves left when the target is smaller than the tested key and right when it is larger; it may stop at any vertex on finding the target.

For n=7=23−1n = 7 = 2^3 - 1 the positions tested form the binary tree of Figure 5.5: the root is the first position tested, and the left and right children of a position are the ones tested next when the target is smaller or larger. A search that stops at depth tt has made t+1t + 1 comparisons.

Proposition 5.17 (Mean cost of a successful search).

Let n=2k−1n = 2^k - 1 and suppose the target is one of the nn keys, each equally likely, with all keys distinct. The mean number of comparisons made by binary search is

12k−1∑i=1ki 2i−1=k−1+k2k−1.\frac{1}{2^k - 1} \sum_{i=1}^{k} i\, 2^{i-1} = k - 1 + \frac{k}{2^k - 1} .

Discussion.

For n=2k−1n = 2^k - 1 every interval that arises has odd length and splits into two equal halves, so the tree of tested positions is full, with 2i−12^{i-1} positions at depth i−1i - 1 for i=1,…,ki = 1, \ldots, k, and a target at depth i−1i - 1 is found after ii comparisons. That gives the sum. To evaluate Sk=∑i=1ki 2i−1S_k = \sum_{i=1}^{k} i\,2^{i-1} we proceed as for the geometric sum and compare SkS_k with 2Sk−12S_{k-1}, whose terms are those of SkS_k shifted by one place with the multiplier lowered by one, so that the difference is a geometric sum.

Proof.

An interval of length 2j−12^j - 1 with j⩾2j \geqslant 2 has middle position with 2j−1−12^{j-1} - 1 positions on each side, so by induction on jj the positions tested form a full binary tree with 2i−12^{i-1} positions at depth i−1i - 1, for 1⩽i⩽k1 \leqslant i \leqslant k. The target at a position of depth i−1i - 1 is found at the iith comparison, so the total over the nn equally likely targets is Sk=∑i=1ki 2i−1S_k = \sum_{i=1}^{k} i\,2^{i-1}.

For k⩾2k \geqslant 2,

Sk−2Sk−1=∑i=1ki 2i−1−∑i=2k(i−1) 2i−1=1+∑i=2k2i−1=2k−1S_k - 2S_{k-1} = \sum_{i=1}^{k} i\,2^{i-1} - \sum_{i=2}^{k} (i-1)\,2^{i-1} = 1 + \sum_{i=2}^{k} 2^{i-1} = 2^k - 1

by the geometric sum. So Sk=2Sk−1+2k−1S_k = 2S_{k-1} + 2^k - 1 with S1=1S_1 = 1, and induction gives Sk=(k−1)2k+1S_k = (k-1)2^k + 1: at k=1k = 1 both sides are 11, and 2((k−2)2k−1+1)+2k−1=(k−1)2k+12\bigl((k-2)2^{k-1} + 1\bigr) + 2^k - 1 = (k-1)2^k + 1. Finally (k−1)2k+1=(k−1)(2k−1)+k(k-1)2^k + 1 = (k-1)(2^k - 1) + k, and dividing by 2k−12^k - 1 gives the stated mean.

The mean is Θ(log⁡n)\Theta(\log n). For arbitrary nn interval halving gives an O(log⁡n)O(\log n) worst case, by the theorem. For a lower bound on the mean, the positions found within tt comparisons are those of depth less than tt in the tree of tested positions, and by the proposition on binary trees there are at most 2t−12^t - 1 of them. Take t=⌊log⁡2n⌋−1t = \lfloor \log_2 n \rfloor - 1 for n⩾4n \geqslant 4; then 2t−1<n/22^t - 1 < n/2, so fewer than half the nn positions are found within tt comparisons. At least half need more than tt, and the mean over equally likely targets is Ω(log⁡n)\Omega(\log n).

A Lower Bound for Searching

A deterministic comparison search algorithm can be pictured as a fixed binary decision tree of all its possible executions, in which each internal vertex is a comparison the algorithm makes. The algorithm walks down the tree from the root: the first comparison it makes is at the root, and according to the outcome it continues at one of the two children. It stops on reaching a leaf, which records its output, so there must be a leaf for every possible output. The number of comparisons on a given input is the depth of the leaf reached, and the worst-case number is the height of the tree.

Theorem 5.18 (Searching needs Ω(log⁡n)\Omega(\log n) comparisons).

Every deterministic algorithm that searches a set of nn items with distinct keys for a given key, using only comparisons of keys, makes at least ⌈log⁡2(n+1)⌉\lceil \log_2(n+1) \rceil comparisons on some input.

Discussion.

The proof counts outputs. A search can end in n+1n + 1 different ways, one for each stored item and one for “not present”, and each must appear at some leaf of the decision tree. A binary tree with that many leaves has height at least log⁡2(n+1)\log_2(n+1) by the corollary on heights, and height is the worst-case number of comparisons.

Proof.

The algorithm has n+1n + 1 possible outputs, the nn stored items and the report that no item has the key, and each occurs for some input, so its decision tree has at least n+1n + 1 leaves. By the corollary on heights of binary trees its height is at least ⌈log⁡2(n+1)⌉\lceil \log_2(n+1) \rceil, and some input follows a path of that length.

So binary search makes the least possible number of comparisons, up to a constant factor. The same argument with any fixed number of outcomes per comparison in place of two still gives Ω(log⁡n)\Omega(\log n). Hashing does better because it is not a comparison algorithm: reading a direct access array at an index computed from the key can go to any one of its slots in a single step.

Problem 5.7.

Modify binary_search so that it returns the least index ii with key⁡(T[i])⩾x\operatorname{key}(T[i]) \geqslant x, or nn if there is none, still in O(log⁡n)O(\log n) time. State the invariant, and use your function to count the entries of a sorted array lying in an interval [a,b)[a, b).

Problem 5.8.

A programmer writes L = m in place of L = m + 1. Give an array and a key for which the modified loop never terminates, and say which part of the termination argument fails.

Sorting

The Sorting Problem

We sort records, each carrying a key that can be compared with other keys. A comparison sort learns about the keys only by comparing them, as in the comparison model.

Definition 5.19 (Sorting, in place and stable).

A sort returns the same records as its input, rearranged in nondecreasing key order. A sort is in place if it uses O(1)O(1) auxiliary cells besides the input array, apart from stack space for recursion if that is being counted separately. A sort is stable if records with equal keys keep their input order.

Checking only that the output keys are in order is not enough to certify a sort: the output must also be a rearrangement of the input records, with none lost or duplicated. Stability can be forced in any comparison sort by comparing the pairs (key,original index)(\text{key}, \text{original index}) lexicographically, first by key and then by index, though storing the original indices may change the space requirements.

Example 5.20 (Stable and unstable).

Sorting (2,a),(1,b),(2,c)(2, a), (1, b), (2, c) stably by the numeric key gives (1,b),(2,a),(2,c)(1, b), (2, a), (2, c): the two records with key 22 remain in their original relative order. A routine that returns (1,b),(2,c),(2,a)(1, b), (2, c), (2, a) has sorted the records, but not stably.

A rearrangement of nn records is a permutation, written as the list π1,…,πn\pi_1, \ldots, \pi_n of the input positions in output order, and there are n!n! of them. A brute-force method tests permutations one by one until it finds an arrangement in increasing order of pairwise distinct keys. A candidate may need n−1n - 1 comparisons of adjacent keys to certify that it is sorted, and there are n!n! candidates.

Definition 5.21 (Inversion).

An inversion of an array T[0],…,T[n−1]T[0], \ldots, T[n-1] is a pair of positions i<ji < j with key⁡(T[i])>key⁡(T[j])\operatorname{key}(T[i]) > \operatorname{key}(T[j]).

A sorted array has no inversions, and an array of distinct keys in decreasing order has all (n2)\binom{n}{2} possible ones, one for each unordered pair of positions. An exchange of two adjacent records T[j],T[j+1]T[j], T[j+1] with key⁡(T[j])>key⁡(T[j+1])\operatorname{key}(T[j]) > \operatorname{key}(T[j+1]) removes exactly one inversion, that pair itself, since the relative order of every other pair is unchanged.

Selection Sort

Selection sort repeatedly removes the largest remaining key and puts it at the next final position from right to left. Having already sorted the largest items into the subarray A[i+1:], it scans A[:i+1] for the largest item not yet placed and swaps it with A[i].

Selection Sort
Input:  an array A[0], …, A[n − 1].
Output: the same array, sorted.

    for i ≝ n − 1 down to 1 do
        m ≝ i
        for j ≝ 0 to i − 1 do
            if key(A[m]) < key(A[j]) then m ≝ j
        swap A[m] and A[i]
def selection_sort(A):
    for i in range(len(A) - 1, 0, -1):   # O(n) passes
        m = i                            # O(1) index of the largest so far
        for j in range(i):               # O(i) search A[:i] for a larger item
            if A[m] < A[j]:              # O(1)
                m = j                    # O(1) new largest found
        A[m], A[i] = A[i], A[m]          # O(1) swap

A = [5, 2, 9, 1, 5, 6]
selection_sort(A)
print(A)   # [1, 2, 5, 5, 6, 9]

The invariant before the pass for ii is that A[i+1:] holds the n−1−in - 1 - i largest records in sorted order, and that every key in A[:i+1] is at most every key in A[i+1:]. The pass moves the largest key of A[:i+1] to position ii, which extends the sorted suffix by one. When the loop ends, at i=0i = 0, the suffix A[1:] is sorted and A[0] holds the smallest key.

Finding the maximum among kk remaining keys takes k−1k - 1 comparisons, which is the least possible by the proposition on finding a maximum, so the method uses

∑k=1n(k−1)=n(n−1)2=Θ(n2)\sum_{k=1}^{n} (k - 1) = \frac{n(n-1)}{2} = \Theta(n^2)

comparisons, even on an input that is already sorted. It performs at most n−1n - 1 swaps. It is in place: apart from the array it uses the names i, j and m.

Example 5.22 (Selection on (1,2,4)(1, 2, 4)).

On (1,2,4)(1, 2, 4), selection chooses 44, then 22, then 11. Each chosen maximum occupies its final, rightmost open position, and the comparison counts are 2+1+0=32 + 1 + 0 = 3.

Selection sort, swapping as above, is not stable. On the records (2,a),(1,b),(1,c)(2, a), (1, b), (1, c), compared by key alone, the first pass finds (2,a)(2, a) at position 00 as the maximum and swaps it with (1,c)(1, c) at position 22, giving (1,c),(1,b),(2,a)(1, c), (1, b), (2, a); the second pass makes no change, and (1,c)(1, c) ends before (1,b)(1, b).

Bubble Sort

A left-to-right pass of bubble sort compares T[j]T[j] with T[j+1]T[j+1] and exchanges them if they are out of order. The largest key in the scanned part moves to its right end, so a full pass places the maximum of the active prefix at its final position. Repeating on successively shorter prefixes sorts the array. If a pass makes no exchange, the array is already sorted and the algorithm can stop early.

Bubble Sort
Input:  an array T[0], …, T[n − 1].
Output: the same array, sorted.

    for last ≝ n − 1 down to 1 do
        changed ≝ false
        for j ≝ 0 to last − 1 do
            if key(T[j]) > key(T[j + 1]) then
                swap T[j] and T[j + 1];  changed ≝ true
        if not changed then stop
def bubble_sort(T):
    for last in range(len(T) - 1, 0, -1):
        changed = False
        for j in range(last):
            if T[j] > T[j + 1]:
                T[j], T[j + 1] = T[j + 1], T[j]
                changed = True
        if not changed:
            return

The invariant before the pass with a given last is that T[last+1:] holds the largest records in sorted order, each at least every key in T[:last+1].

For n⩾1n \geqslant 1, the worst-case number of comparisons is n(n−1)/2n(n-1)/2, and the best case, on a sorted input, is n−1n - 1 with the early stop. Over the n!n! orders of nn distinct keys, counted equally, the mean number of comparisons is still Θ(n2)\Theta(n^2). For the lower bound, the smallest key begins in the last half of the array, at one of the ⌈n/2⌉\lceil n/2 \rceil positions from ⌊n/2⌋\lfloor n/2 \rfloor on, in at least half of the n!n! orders. A left-to-right pass moves the smallest key left by at most one position, so on those orders at least ⌊n/2⌋\lfloor n/2 \rfloor passes are needed, and each of those passes makes at least n/2n/2 comparisons. Thus at least half the orders cost at least ⌊n/2⌋⋅n/2\lfloor n/2 \rfloor \cdot n/2 comparisons, and the mean is Ω(n2)\Omega(n^2). The upper bound O(n2)O(n^2) holds on every input.

Proposition 5.23 (Exchanges in bubble sort).

The number of exchanges bubble sort makes on an array equals its number of inversions. Over the n!n! orders of nn distinct keys, counted equally, the mean number of inversions is n(n−1)/4n(n-1)/4.

Discussion.

For the first statement, every exchange removes exactly one inversion, and the algorithm stops with a sorted array, which has none; so the number of exchanges is the number of inversions at the start. For the mean, pair each order with its reverse. A pair of positions is inverted in exactly one of the two, so the inversion counts of an order and of its reverse add up to (n2)\binom{n}{2}. The reversal pairs the n!n! orders off, so the average over all of them is half of (n2)\binom{n}{2}.

Proof.

Bubble sort exchanges T[j]T[j] and T[j+1]T[j+1] only when key⁡(T[j])>key⁡(T[j+1])\operatorname{key}(T[j]) > \operatorname{key}(T[j+1]), and such an exchange of adjacent records removes exactly one inversion. It stops with a sorted array, which has no inversions. So if the input has II inversions, it makes exactly II exchanges.

For an order π\pi of distinct keys let inv⁡(π)\operatorname{inv}(\pi) be its number of inversions, and let πˉ\bar\pi be the reverse order. For each of the (n2)\binom{n}{2} pairs of keys, exactly one of π\pi and πˉ\bar\pi lists the larger key first, so inv⁡(π)+inv⁡(πˉ)=(n2)\operatorname{inv}(\pi) + \operatorname{inv}(\bar\pi) = \binom{n}{2}. Reversal is a bijection from the set of orders to itself, so summing over all orders,

2∑πinv⁡(π)=∑π(inv⁡(π)+inv⁡(πˉ))=n!(n2),2\sum_{\pi} \operatorname{inv}(\pi) = \sum_{\pi} \bigl(\operatorname{inv}(\pi) + \operatorname{inv}(\bar\pi)\bigr) = n!\binom{n}{2},

and the mean is 12(n2)=n(n−1)/4\tfrac12\binom{n}{2} = n(n-1)/4.

Exchanges happen only on strict inversions, so records with equal keys are never exchanged with each other and keep their original order: bubble sort is stable. Small keys near the right end move slowly, one position per pass, and are sometimes called turtles; large keys near the left move quickly to the right and are called hares. Bidirectional bubble sort alternates the direction of its passes to move turtles faster, and gnome sort steps back to recheck the previous adjacent pair after each exchange.

Insertion Sort

Insertion sort treats T[0..i-1] as sorted and inserts T[i] into that prefix by shifting the larger records one position right.

Insertion Sort
Input:  an array T[0], …, T[n − 1].
Output: the same array, sorted.

    for i ≝ 1 to n − 1 do
        x ≝ T[i];  j ≝ i − 1
        while j ⩾ 0 and key(x) < key(T[j]) do
            T[j + 1] ≝ T[j];  j ≝ j − 1
        T[j + 1] ≝ x
def insertion_sort(T):
    for i in range(1, len(T)):           # O(n) passes
        x = T[i]                         # the record to insert
        j = i - 1
        while j >= 0 and x < T[j]:       # O(i) shift larger records right
            T[j + 1] = T[j]
            j = j - 1
        T[j + 1] = x                     # fill the gap

The invariant before the pass for ii is that T[0..i-1] holds the first ii input records in sorted order. The final assignment T[j+1] = x places the saved record in the gap left by the shifts, and it is essential: writing xx anywhere else would duplicate one record and lose another. The while test relies on and stopping as soon as j >= 0 is False, so that T[-1] is never read.

Example 5.24 (Insertion on (4,2,3)(4, 2, 3)).

On (4,2,3)(4, 2, 3), save 22, shift 44 right and write 22 at index 00, giving (2,4,3)(2, 4, 3). Next save 33, shift 44 right and write 33 at index 11, giving (2,3,4)(2, 3, 4).

Insertion sort is Θ(n2)\Theta(n^2) in the worst case, on an input in decreasing order, where the pass for ii shifts all ii records of the prefix, and Θ(n)\Theta(n) on an already sorted input, where every pass makes one comparison and no shift.

In Place and Stable

Selection sort, bubble sort and insertion sort are all in place: each uses a constant amount of space besides the array, and acts on the array only by comparisons and by exchanging or moving records. Bubble sort and insertion sort are stable, insertion sort because a record is never shifted past one with an equal key. Selection sort, as implemented above, is not.

Problem 5.9.

Show that insertion sort performs exactly as many shifts T[j + 1] = T[j] as the input has inversions, and deduce the mean number of shifts over the n!n! orders of nn distinct keys. Give an input of length nn on which selection sort makes n−1n - 1 swaps and insertion sort makes no shift.

Problem 5.10.

Run bubble sort and insertion sort on (3,1,4,1,5,9,2,6)(3, 1, 4, 1, 5, 9, 2, 6), keeping the two 11s apart as 1a1_a and 1b1_b. Record the array after each pass of each algorithm, count the comparisons and exchanges or shifts, and confirm that both outputs keep 1a1_a before 1b1_b.

Merge Sort

Merge sort splits the array in half, sorts each half recursively, and merges the two sorted halves into one.

Proposition 5.25 (Merging two sorted arrays).

Two sorted arrays of lengths a,b>0a, b > 0 can be merged into one sorted array with at most a+b−1a + b - 1 key comparisons. If ties take the record from the left array first, the merge is stable.

Discussion.

The smallest remaining record of the whole is always at the front of one of the two arrays, because each array is sorted, so one comparison of the two front records decides which comes next. Each comparison outputs one record, and once one array is exhausted the rest of the other is copied without comparing. At least one record, the last, is output without a comparison, which gives a+b−1a + b - 1. For stability, when the two front keys are equal, taking the left one keeps records from the left array ahead of equal records from the right.

Proof.

Keep a cursor at the first unmerged element of each array. While both arrays have elements left, compare their current heads and copy the smaller one to the output. Its key is no larger than any remaining key, because each input array is already sorted. So the output stays sorted, and no record is lost or copied twice. Each comparison advances one cursor, so after at most a+b−1a + b - 1 comparisons one array must be empty; the other array’s sorted tail is then appended without further comparisons. On equal keys, choosing the left head first preserves the original order between the two arrays, while the order within each array is unchanged.

Merge
Input:  sorted arrays X and Y.
Output: a sorted array Z holding the records of X and Y.

    i ≝ 0;  j ≝ 0;  Z ≝ empty
    while i < length(X) and j < length(Y) do
        if key(X[i]) ⩽ key(Y[j]) then append X[i] to Z;  i ≝ i + 1
        else append Y[j] to Z;  j ≝ j + 1
    append the remaining part of X, then of Y, to Z
    return Z

Merge Sort
Input:  an array A.
Output: a sorted array holding the records of A.

    if length(A) ⩽ 1 then return A
    h ≝ ⌊length(A) / 2⌋
    return Merge(Merge Sort(A[0..h − 1]), Merge Sort(A[h..length(A) − 1]))
def merge(X, Y):
    i, j, Z = 0, 0, []
    while i < len(X) and j < len(Y):
        if X[i] <= Y[j]:
            Z.append(X[i])
            i = i + 1
        else:
            Z.append(Y[j])
            j = j + 1
    return Z + X[i:] + Y[j:]

def merge_sort(A):
    if len(A) <= 1:
        return A
    h = len(A) // 2
    return merge(merge_sort(A[:h]), merge_sort(A[h:]))

print(merge_sort([15, 5, 64, 8, 12, 6, 4, 35]))   # [4, 5, 6, 8, 12, 15, 35, 64]

Theorem 5.26 (Merge sort).

On an array of nn records with comparable keys, merge sort returns a sorted rearrangement of the input, and choosing the left head on equal keys makes it stable. It uses O(nlog⁡n)O(n \log n) time and O(n)O(n) auxiliary storage. For n=2kn = 2^k, its worst-case number of key comparisons is exactly nlog⁡2n−n+1n \log_2 n - n + 1.

Discussion.

Correctness is an induction on nn whose step is the merge proposition: the recursive calls are on shorter arrays, so by the hypothesis they return sorted rearrangements of the halves, and merging those gives a sorted rearrangement of the whole. For the time, we count level by level: the arrays at one level are disjoint pieces of the input, so merging all of them costs O(n)O(n), and halving gives about log⁡2n\log_2 n levels. For the exact count, the merge proposition allows a merge of two halves of size 2k−12^{k-1} up to 2k−12^k - 1 comparisons, which gives the recurrence Ck=2Ck−1+2k−1C_k = 2C_{k-1} + 2^k - 1. It remains to find an input that attains the bound at every merge, and to solve the recurrence.

Proof.

We prove correctness by induction on nn. For n⩽1n \leqslant 1 the array is already sorted. If n>1n > 1, both halves have fewer than nn elements, so the recursive calls terminate and, by induction, return sorted rearrangements of their halves. The merge proposition combines those into a sorted rearrangement of the whole input. Its tie rule preserves the order of equal-key records across the halves, so induction also proves stability. The recursion terminates because every call on more than one record is made on strictly shorter arrays.

At any fixed level of the recursion the subarrays are disjoint parts of the original array, so their lengths add up to at most nn. Merging each costs time proportional to its length, so the total work per level is O(n)O(n). Halving gives at most ⌈log⁡2n⌉\lceil \log_2 n \rceil levels of merging, and therefore O(nlog⁡n)O(n \log n) time. At any instant the arrays alive along the current chain of calls have lengths at most n,n/2,n/4,…n, n/2, n/4, \ldots together with their merge outputs, so at most O(n)O(n) auxiliary storage is present at once; the recursion stack has O(log⁡n)O(\log n) frames.

When n=2kn = 2^k, a merge of two sorted halves of equal size can need all 2k−12^k - 1 comparisons the merge proposition allows: give the left half the records of odd rank and the right half those of even rank, so that their sorted values alternate until only one remains. Apply the same odd–even assignment recursively within each half. Arranging the input by these assignments makes every merge attain its worst case. Writing CkC_k for the worst case at size 2k2^k, this gives C0=0C_0 = 0 and

Ck=2Ck−1+2k−1.C_k = 2C_{k-1} + 2^k - 1 .

We claim Ck=k2k−2k+1C_k = k2^k - 2^k + 1. At k=0k = 0 both sides are 00. Assuming the formula at k−1k - 1,

Ck=2((k−1)2k−1−2k−1+1)+2k−1=k2k−2k+1.C_k = 2\bigl((k-1)2^{k-1} - 2^{k-1} + 1\bigr) + 2^k - 1 = k2^k - 2^k + 1 .

With n=2kn = 2^k this is nlog⁡2n−n+1n \log_2 n - n + 1. It is a count of key comparisons, and an array of one record makes none.

Problem 5.11.

Write an iterative merge sort that merges adjacent runs of length 11, then 22, then 44, and so on, with no recursion. Prove it correct with a loop invariant on the run length, and show that it makes at most n⌈log⁡2n⌉n \lceil \log_2 n \rceil comparisons.

Quicksort

Quicksort chooses a pivot, partitions the array so that smaller keys precede the pivot and larger or equal keys follow it, and then recursively sorts the two sides. For the partition in place, with the pivot initially at position LL, the invariant before inspecting position kk is: positions L+1,…,pL+1, \ldots, p hold keys smaller than the pivot, positions p+1,…,k−1p+1, \ldots, k-1 hold keys larger than or equal to it, and positions k,…,R−1k, \ldots, R-1 are unexamined. A newly found smaller record is swapped into position p+1p + 1 and pp is increased. At the end the pivot is swapped with the record at position pp. Partitioning ss records compares each of the other s−1s - 1 records with the pivot once.

The pseudocode below sorts the half-open slice [L,R)[L, R) of the original array in place; no slice is copied. The partition moves only records strictly smaller than the pivot to its left.

Quicksort
Input:  an array A and positions L ⩽ R.
Output: A with A[L..R − 1] sorted.

    Sort(A, L, R):
        if R − L ⩽ 1 then return
        choose q uniformly from L, …, R − 1;  swap A[L] and A[q]
        pivot ≝ A[L];  p ≝ L
        for k ≝ L + 1 to R − 1 do
            if key(A[k]) < key(pivot) then
                p ≝ p + 1;  swap A[p] and A[k]
        swap A[L] and A[p]
        Sort(A, L, p);  Sort(A, p + 1, R)

The statement from random import randrange makes randrange available: randrange(L, R) returns an integer chosen uniformly from L,…,R−1L, \ldots, R - 1, the same range as range(L, R).

from random import randrange

def sort_slice(A, L, R):
    if R - L <= 1:
        return
    q = randrange(L, R)
    A[L], A[q] = A[q], A[L]
    pivot = A[L]
    p = L
    for k in range(L + 1, R):
        if A[k] < pivot:
            p = p + 1
            A[p], A[k] = A[k], A[p]
    A[L], A[p] = A[p], A[L]
    sort_slice(A, L, p)
    sort_slice(A, p + 1, R)

def quicksort(A):
    sort_slice(A, 0, len(A))

If the pivot has rank rr among nn distinct keys, meaning r−1r - 1 keys are smaller, the number of comparisons satisfies

C(n)=n−1+C(r−1)+C(n−r),C(0)=C(1)=0.C(n) = n - 1 + C(r - 1) + C(n - r), \qquad C(0) = C(1) = 0 .

Example 5.27 (One partition).

For the input (15,5,64,8,12,6,4,35)(15, 5, 64, 8, 12, 6, 4, 35), take the first value 1515 as pivot, with no random swap. The partition puts the five smaller values 5,8,12,6,45, 8, 12, 6, 4 before it and 64,3564, 35 after it, and the in-place result is (4,5,8,12,6,15,64,35)(4, 5, 8, 12, 6, 15, 64, 35). Recursively sorting the two sides gives (4,5,6,8,12,15,35,64)(4, 5, 6, 8, 12, 15, 35, 64). The order within each side before the recursion depends on the partition routine.

Proposition 5.28 (Worst and best case of quicksort).

For nn distinct keys, any run of quicksort makes at most (n2)=n(n−1)/2\binom{n}{2} = n(n-1)/2 key comparisons. This bound is attained when every pivot is the smallest or largest key in its current subarray, so the worst-case number of comparisons is Θ(n2)\Theta(n^2). If every pivot splits its subarray as evenly as possible, the number of comparisons is Θ(nlog⁡n)\Theta(n \log n).

Discussion.

For the upper bound we count pairs of keys. Two keys are compared only when one of them is the pivot, and a pivot is left out of all later calls, so no pair is compared twice; there are (n2)\binom{n}{2} pairs. When every pivot is extreme, one side of each partition is empty, and the recurrence becomes Wn=(n−1)+Wn−1W_n = (n-1) + W_{n-1}, which sums to the same number. For even splits, the depth of the recursion is logarithmic and each level costs at most nn comparisons, which gives O(nlog⁡n)O(n \log n); the matching lower bound needs a count of how many levels still have large subarrays, and how many comparisons each such level makes.

Proof.

Partitioning a subarray of size ss compares its pivot with each of the other s−1s - 1 keys exactly once. A given pair of keys can be compared only when one of them is the pivot, and that pivot is then excluded from all recursive subarrays. Thus no pair is compared twice, and there are at most (n2)\binom{n}{2} comparisons. If every pivot is extreme, one recursive side is empty and the other has size n−1n - 1. The count obeys Wn=(n−1)+Wn−1W_n = (n-1) + W_{n-1} with W0=W1=0W_0 = W_1 = 0, and hence

Wn=(n−1)+(n−2)+⋯+1=n(n−1)2.W_n = (n-1) + (n-2) + \cdots + 1 = \frac{n(n-1)}{2} .

For balanced splits, the recursion has O(log⁡n)O(\log n) levels. At any level the active subarrays are disjoint, so their sizes add up to at most nn and their partitions make at most nn comparisons in total. This proves O(nlog⁡n)O(n \log n).

For the reverse bound, number the levels from ℓ=0\ell = 0. A balanced split of ss records leaves at least ⌊(s−1)/2⌋\lfloor (s-1)/2 \rfloor records on either side. Applying this at each level shows that every subarray at level ℓ\ell has at least ⌊(n+1)/2ℓ⌋−1\lfloor (n+1)/2^\ell \rfloor - 1 records. For 0⩽ℓ⩽⌊log⁡2n⌋−20 \leqslant \ell \leqslant \lfloor \log_2 n \rfloor - 2 this is at least 33, so no branch has yet ended at a call on one record. Before such a level at most 2ℓ−12^\ell - 1 pivots have been removed, and the remaining records all belong to at most 2ℓ2^\ell active subarrays. A subarray of size ss uses s−1s - 1 comparisons, so level ℓ\ell uses at least n−(2ℓ−1)−2ℓ=n−2ℓ+1+1⩾n/2n - (2^\ell - 1) - 2^\ell = n - 2^{\ell+1} + 1 \geqslant n/2 comparisons. There are Θ(log⁡n)\Theta(\log n) such levels for large nn, which proves Ω(nlog⁡n)\Omega(n \log n).

The in-place partition is not stable in general. The recursive calls use O(log⁡n)O(\log n) stack space for balanced splits and O(n)O(n) when the splits are as uneven as possible.

With a random pivot the cost is an average over the random choices, in the sense of the expectation defined for hashing in the last lesson.

Remark (Expected cost of a randomized algorithm).

The expectation used for hashing is an average over one uniform choice. Randomized quicksort makes a uniform choice at every call, and its expected cost is defined in the same way one call at a time: if a call on ss keys chooses the pivot rank rr with probability 1/s1/s, and the rest of the run then has expected cost crc_r, the expected cost of the call is its own comparisons plus 1s∑rcr\tfrac{1}{s}\sum_{r} c_r. Equivalently, it is the average over all complete runs, each run weighted by the product of the probabilities 1/s1/s of the choices it makes.

Theorem 5.29 (Expected cost of randomized quicksort).

Fix any input order of nn distinct keys. At every recursive call choose the pivot uniformly from the current subarray, and compare it once with each other key there. Let EnE_n be the expected total number of key comparisons. Then

E0=0,En=2(n+1)Hn−4n(n⩾1).E_0 = 0, \qquad E_n = 2(n+1)H_n - 4n \quad (n \geqslant 1) .

In particular En=Θ(nlog⁡n)E_n = \Theta(n \log n). The expectation is over the pivot choices; the input order is fixed and need not be random.

Discussion.

The proof has four steps. Conditioning on the rank of the first pivot turns the expectation into a recurrence in which EnE_n depends on all of E0,…,En−1E_0, \ldots, E_{n-1} through their sum. Writing the recurrence at nn and at n−1n - 1 and subtracting removes the sum, leaving a first-order recurrence with variable coefficients, of the kind the summation factor of the last lesson solves; the example there is the same recurrence with 2n2n in place of 2(n−1)2(n-1). Here the answer is checked by induction instead. Finally the bounds on HnH_n give the growth rate.

Proof.

Step 1: condition on the first pivot. A subarray of size at most 11 needs no comparisons, so E0=E1=0E_0 = E_1 = 0. For n⩾2n \geqslant 2 the first partition costs exactly n−1n - 1 comparisons. The pivot is equally likely to have any rank r∈{1,…,n}r \in \{1, \ldots, n\} among the keys. Rank rr leaves r−1r - 1 smaller keys on the left and n−rn - r larger keys on the right. Because every later pivot is chosen uniformly within its own subarray, the expected costs of those two calls are Er−1E_{r-1} and En−rE_{n-r}, whatever their internal order. Averaging over the nn possible ranks gives

En=1n∑r=1n((n−1)+Er−1+En−r)=n−1+2n∑j=0n−1Ej,E_n = \frac{1}{n}\sum_{r=1}^{n} \bigl((n-1) + E_{r-1} + E_{n-r}\bigr) = n - 1 + \frac{2}{n}\sum_{j=0}^{n-1} E_j ,

since both r−1r - 1 and n−rn - r run once through 0,…,n−10, \ldots, n-1 as rr runs through 1,…,n1, \ldots, n.

Step 2: remove the sum. Multiply this equation by nn, and write the same equation for n−1n - 1 multiplied by n−1n - 1:

nEn=n(n−1)+2∑j=0n−1Ej,(n−1)En−1=(n−1)(n−2)+2∑j=0n−2Ej.\begin{aligned} nE_n &= n(n-1) + 2\sum_{j=0}^{n-1} E_j, \\ (n-1)E_{n-1} &= (n-1)(n-2) + 2\sum_{j=0}^{n-2} E_j . \end{aligned}

Subtract the second from the first. The two sums differ only by En−1E_{n-1}, while n(n−1)−(n−1)(n−2)=2(n−1)n(n-1) - (n-1)(n-2) = 2(n-1). Therefore nEn−(n−1)En−1=2(n−1)+2En−1nE_n - (n-1)E_{n-1} = 2(n-1) + 2E_{n-1}, that is,

nEn=(n+1)En−1+2(n−1).nE_n = (n+1)E_{n-1} + 2(n-1) .

Step 3: solve by induction. At n=1n = 1 the claimed expression is 2⋅2⋅H1−4=0=E12 \cdot 2 \cdot H_1 - 4 = 0 = E_1. Assume En−1=2nHn−1−4(n−1)E_{n-1} = 2nH_{n-1} - 4(n-1) for some n⩾2n \geqslant 2. Substituting into the last recurrence and using Hn−1=Hn−1/nH_{n-1} = H_n - 1/n,

nEn=(n+1)(2nHn−1−4(n−1))+2(n−1)=2n(n+1)Hn−1−4n2+2n+2=2n(n+1)(Hn−1n)−4n2+2n+2=2n(n+1)Hn−4n2.\begin{aligned} nE_n &= (n+1)\bigl(2nH_{n-1} - 4(n-1)\bigr) + 2(n-1) \\ &= 2n(n+1)H_{n-1} - 4n^2 + 2n + 2 \\ &= 2n(n+1)\left(H_n - \frac{1}{n}\right) - 4n^2 + 2n + 2 \\ &= 2n(n+1)H_n - 4n^2 . \end{aligned}

Dividing by nn gives En=2(n+1)Hn−4nE_n = 2(n+1)H_n - 4n.

Step 4: the growth rate. By the bounds on the harmonic numbers, 12⌊log⁡2n⌋⩽Hn⩽1+log⁡2n\tfrac12\lfloor \log_2 n \rfloor \leqslant H_n \leqslant 1 + \log_2 n. The upper bound gives En⩽2(n+1)(1+log⁡2n)=O(nlog⁡n)E_n \leqslant 2(n+1)(1 + \log_2 n) = O(n \log n). The lower bound gives En⩾(n+1)(log⁡2n−1)−4nE_n \geqslant (n+1)(\log_2 n - 1) - 4n, and for all sufficiently large nn the term (n+1)log⁡2n(n+1)\log_2 n exceeds twice the rest, so En=Ω(nlog⁡n)E_n = \Omega(n \log n). Thus En=Θ(nlog⁡n)E_n = \Theta(n \log n).

Example 5.30 (Three keys).

For three distinct keys, a smallest or largest first pivot costs two comparisons and leaves a call on two keys costing one more, for a total of 33. A middle first pivot costs two comparisons and leaves only calls on one key, for a total of 22. The three ranks are equally likely, so E3=(3+2+3)/3=8/3E_3 = (3 + 2 + 3)/3 = 8/3. The formula agrees: 2⋅4⋅(1+12+13)−12=8/32 \cdot 4 \cdot \bigl(1 + \tfrac12 + \tfrac13\bigr) - 12 = 8/3.

Problem 5.12.

Show that on the input (1,2,…,n)(1, 2, \ldots, n) quicksort with the first record always taken as pivot makes n(n−1)/2n(n-1)/2 comparisons, and that with a uniformly random pivot the probability of making all n(n−1)/2n(n-1)/2 is 2n−1/n!2^{n-1}/n! for n⩾2n \geqslant 2.

Problem 5.13.

Two keys of ranks i<ji < j are compared by randomized quicksort exactly when the first pivot chosen from the keys of ranks i,…,ji, \ldots, j is one of the two. Deduce that they are compared with probability 2/(j−i+1)2/(j - i + 1), and use linearity of expectation to give a second proof that En=2(n+1)Hn−4nE_n = 2(n+1)H_n - 4n.

Lower Bounds for Comparison Sorting

A deterministic comparison sort is described, like a comparison search, by a binary decision tree: each internal vertex is a comparison, the two branches below it are its two outcomes, and each leaf records the rearrangement the algorithm outputs. An algorithm may ask a comparison whose result is already forced by earlier answers, and then one branch below it is reached by no input.

a₀ < a₁?a₀ < a₂?a₁ < a₂?a₁ < a₂?a₁ < a₀?a₀ < a₂?a₁ < a₀?012021×201102120210×yesno
Figure 5.6. A comparison tree for three records a0,a1,a2a_0, a_1, a_2. Left branches answer yes and right branches no; a leaf such as 021021 means a0<a2<a1a_0 \lt a_2 \lt a_1. The two accented leaves cannot be reached, because the comparison above them was already settled by an earlier answer.

The comparison tree of Figure 5.6 names the compared records a0,a1,a2a_0, a_1, a_2 rather than their changing positions in the array. Its six reachable leaves are the six possible strict orders; the crossed leaves are impossible by transitivity.

Theorem 5.31 (Worst-case lower bound for sorting).

Every deterministic sorting algorithm that learns about nn distinct, otherwise arbitrary keys only by comparing pairs makes at least ⌈log⁡2(n!)⌉\lceil \log_2(n!) \rceil comparisons on some input. Hence its worst-case number of comparisons is Ω(nlog⁡n)\Omega(n \log n).

Discussion.

As for searching, the proof counts outputs, and there are now n!n! of them. Two inputs whose keys are in different relative orders need different rearrangements, so they must reach different leaves, and the decision tree has at least n!n! reachable leaves. The corollary on heights of binary trees then bounds the height below by log⁡2(n!)\log_2(n!). For the growth rate, half of the factors of n!n! are at least n/2n/2, so log⁡2(n!)\log_2(n!) is at least about n2log⁡2n2\tfrac{n}{2}\log_2 \tfrac{n}{2}.

Proof.

There are n!n! possible relative orders of the distinct input keys: nn choices for the rank of the first record, n−1n - 1 for the second, and so on. Represent the algorithm by its binary decision tree. Each input follows one path from the root to a leaf, and the length of that path is its number of comparisons. Two different relative orders cannot end at the same leaf: that leaf prescribes one output rearrangement for both, and for at least one of them it is wrong. Therefore the tree has at least n!n! reachable leaves, and after removing the unreachable branches it is a binary tree with at least n!n! leaves. By the corollary on heights its height hh satisfies h⩾⌈log⁡2(n!)⌉h \geqslant \lceil \log_2(n!) \rceil.

The last ⌊n/2⌋\lfloor n/2 \rfloor factors of n!=n(n−1)⋯1n! = n(n-1)\cdots 1 are each at least n/2n/2, so n!⩾(n/2)⌊n/2⌋n! \geqslant (n/2)^{\lfloor n/2 \rfloor} and

log⁡2(n!)⩾⌊n/2⌋log⁡2(n/2)=Ω(nlog⁡n).\log_2(n!) \geqslant \lfloor n/2 \rfloor \log_2(n/2) = \Omega(n \log n) .

Removing the branches that no input can follow, and contracting each vertex left with one child, turns the decision tree into a full binary tree with one leaf per possible outcome.

before pruningafter pruninga₁a₂a₃a₄impossibleimpossiblea₁a₂a₃a₄
Figure 5.7. Pruning. Branches that no input can follow are drawn dashed on the left; removing them and contracting each vertex left with a single child gives the full binary tree on the right, with one leaf for each of the four possible outcomes.

Merge sort makes at most n⌈log⁡2n⌉n\lceil \log_2 n\rceil comparisons, so it is asymptotically optimal among comparison sorts in the worst case.

The mean number of comparisons over the n!n! equally weighted orders of distinct keys is the mean depth of the leaves of the decision tree, one leaf per order. A full binary tree with KK leaves is balanced if its leaf depths differ by at most one.

Proposition 5.32 (Balanced trees minimise the total depth).

Among full binary trees with KK leaves, a balanced one has the least sum of leaf depths, and in a balanced full binary tree with KK leaves every leaf has depth at least ⌊log⁡2K⌋\lfloor \log_2 K \rfloor.

Discussion.

The first claim is proved by an exchange that lowers the depth sum of an unbalanced tree. If some leaf cc is at least two levels above the deepest leaves, move a deepest pair of sibling leaves a,ba, b together with their parent into the place of cc, and put cc where their parent was. Three depths change, and the sum falls. The depth sum is a nonnegative integer, so the exchange can be repeated only finitely often, and it stops at a balanced tree. For the second claim, look at the shallowest leaf: every level above it is full, and all leaves lie on its level or the next, which bounds KK by twice the size of its level.

Proof.

Take a full binary tree that is not balanced. Choose deepest sibling leaves a,ba, b, at depth pp, with parent nn, and a leaf cc at depth p′⩽p−2p' \leqslant p - 2. Exchange the subtree consisting of nn, aa and bb with the leaf cc. The tree remains a full binary tree with the same leaves, aa and bb move to depth p′+1p' + 1, and cc moves to depth p−1p - 1. The sum of these three depths changes from 2p+p′2p + p' to 2(p′+1)+(p−1)2(p' + 1) + (p - 1), a decrease of p−p′−1>0p - p' - 1 > 0, and no other depth changes. Repeating the operation reaches a balanced tree, because the nonnegative integer depth sum strictly decreases each time. Every tree with KK leaves can be transformed in this way into a balanced one with no larger depth sum, and all balanced full binary trees with KK leaves have the same multiset of leaf depths, as the count below shows; so a balanced tree has the least sum.

Let the shallowest leaf of a balanced full binary tree have depth dd. There is no leaf above depth dd, so every vertex at depth less than dd is internal with two children, and depth dd holds 2d2^d vertices. The leaves are at depth dd or d+1d + 1, and those at depth d+1d + 1 are the children of the internal vertices at depth dd, so K=2d+#{internal vertices at depth d}K = 2^d + \#\{\text{internal vertices at depth } d\}. At least one vertex at depth dd is a leaf, so 2d⩽K<2d+12^d \leqslant K < 2^{d+1}. This determines d=⌊log⁡2K⌋d = \lfloor \log_2 K \rfloor, and with it the number of leaves at each depth, and every leaf has depth at least ⌊log⁡2K⌋\lfloor \log_2 K \rfloor.

beforeafterc⋯abndepth p′depth p⋯abcnp′ + 1p − 1
Figure 5.8. The exchange. The two deepest sibling leaves aa, bb and their parent nn swap places with a leaf cc at least two levels higher; aa and bb rise to depth p′+1p' + 1 and cc falls to depth p−1p - 1.

For example, six leaves can occur at depths 22 and 33. The proposition gives the asymptotic bound: the mean depth of any full binary tree with KK leaves is at least ⌊log⁡2K⌋\lfloor \log_2 K \rfloor, which is Ω(log⁡K)\Omega(\log K). The exact bound log⁡2K\log_2 K is proved differently.

Theorem 5.33 (Average lower bound for sorting).

Suppose the n!n! relative orders of nn distinct keys are equally likely. Every deterministic comparison sort has mean number of comparisons at least log⁡2(n!)\log_2(n!), and therefore Ω(nlog⁡n)\Omega(n \log n).

Discussion.

The proof turns the leaf depths into lengths of intervals. Each leaf is reached by a word of left and right choices, and a word of length dd picks out a subinterval of [0,1)[0, 1) of length 2−d2^{-d} by halving repeatedly. Since no leaf’s word begins another leaf’s word, the intervals do not overlap, so their lengths add up to at most 11. The mean depth is then bounded below by comparing the arithmetic mean of the numbers 2−di2^{-d_i} with their geometric mean. The inequality of the means was proved for two numbers in Lesson 1; the proof extends it to KK numbers by doubling up to a power of two and padding.

Proof.

Keep one reachable leaf of the decision tree for each of the K=n!K = n! relative orders, and let their depths be d1,…,dKd_1, \ldots, d_K; the mean number of comparisons is K−1∑idiK^{-1}\sum_i d_i. Each path from the root to a leaf is a word of left and right choices. No leaf’s word can be a prefix of another leaf’s word, because a computation stops when it reaches a leaf. To a word of length did_i associate the subinterval of [0,1)[0, 1) of length 2−di2^{-d_i} obtained by choosing the left or right half at each successive letter. The intervals for distinct leaf words do not overlap, so their total length is at most 11:

∑i=1K2−di⩽1.\sum_{i=1}^{K} 2^{-d_i} \leqslant 1 .

Write xi=2−di>0x_i = 2^{-d_i} > 0 and A=K−1∑ixiA = K^{-1}\sum_i x_i, so that A⩽1/KA \leqslant 1/K. We need the inequality of arithmetic and geometric means for KK positive numbers, (∏ixi)1/K⩽A\bigl(\prod_i x_i\bigr)^{1/K} \leqslant A. For two numbers u,vu, v it is uv⩽(u+v)/2\sqrt{uv} \leqslant (u + v)/2. For a list of 2r2^r numbers, split it into two equal halves with geometric means G1,G2G_1, G_2 and arithmetic means A1,A2A_1, A_2; by induction on rr, G1⩽A1G_1 \leqslant A_1 and G2⩽A2G_2 \leqslant A_2, so the whole list has geometric mean G1G2⩽A1A2⩽(A1+A2)/2\sqrt{G_1 G_2} \leqslant \sqrt{A_1 A_2} \leqslant (A_1 + A_2)/2, its arithmetic mean. For general KK, choose a power of two N⩾KN \geqslant K and append N−KN - K copies of AA to the list x1,…,xKx_1, \ldots, x_K. The enlarged list still has arithmetic mean AA, so the power-of-two case gives (∏ixi)AN−K⩽AN\bigl(\prod_i x_i\bigr) A^{N-K} \leqslant A^N, and cancelling the positive factor AN−KA^{N-K} gives ∏ixi⩽AK\prod_i x_i \leqslant A^K.

Substituting xi=2−dix_i = 2^{-d_i},

2−1K∑idi⩽A⩽1K.2^{-\frac{1}{K}\sum_i d_i} \leqslant A \leqslant \frac{1}{K} .

Taking base-two logarithms and multiplying by −1-1, which reverses the inequality, gives K−1∑idi⩾log⁡2K=log⁡2(n!)K^{-1}\sum_i d_i \geqslant \log_2 K = \log_2(n!). Finally log⁡2(n!)⩾⌊n/2⌋log⁡2(n/2)=Ω(nlog⁡n)\log_2(n!) \geqslant \lfloor n/2 \rfloor \log_2(n/2) = \Omega(n \log n), as in the worst-case bound.

Example 5.34 (Six outcomes).

With six equally likely outcomes, a full binary tree may have two leaves at depth 22 and four at depth 33. Its mean depth is (2⋅2+4⋅3)/6=8/3(2 \cdot 2 + 4 \cdot 3)/6 = 8/3, which is at least log⁡26≈2.585\log_2 6 \approx 2.585. The same leaf-depth argument applies to the n!n! possible input orders.

Sorting algorithms that are faster than nlog⁡nn \log n exist, such as counting sort and radix sort, but only because they use information about the keys beyond pairwise comparison, for instance that they are small integers usable as array indices.

Problem 5.14.

Draw a decision tree for insertion sort on three distinct keys a0,a1,a2a_0, a_1, a_2, and give its height and the mean depth of its leaves. Compare both with ⌈log⁡23!⌉\lceil \log_2 3! \rceil and log⁡23!\log_2 3!.

Problem 5.15.

Show that five distinct keys can be sorted with 77 comparisons in the worst case, and that no comparison sort does it with 66.

Sorted Arrays as Sets

A sorted array implements the set interface. Building it is sorting, O(nlog⁡n)O(n \log n) by merge sort. find(k) is binary search, O(log⁡n)O(\log n). The smallest and largest keys are at the two ends, and find_next(k) and find_prev(k) are binary searches for the position where kk would go. Inserting or deleting while keeping the order shifts up to nn items.

Data structurebuild(X)find(k)insert(x), delete(k)find_min(), find_max()find_prev(k), find_next(k)
Arraynnnnnnnnnn
Sorted arraynlog⁡nn \log nlog⁡n\log nnn11log⁡n\log n
Direct access arrayuu1111uuuu
Hash tablenn (e)11 (e)11 (a)(e)nnnn

Each entry is an O(⋅)O(\cdot) bound; (e) marks an expected bound and (a) an amortized one. For the sorted array, find is optimal among comparison algorithms by the lower bound for searching, and the bound does not apply to the hash table, which is not a comparison algorithm.

Problem 5.16.

Implement the set interface on a sorted array as a class Sorted_Array_Set, using merge_sort in build and binary search in find, find_next and find_prev, with the running times of the table.

Exercises on Connectivity and Trees

Exercise 5.1.

How many spanning subgraphs of K4K_4 are trees?

Exercise 5.2.

Prove that every connected graph with at least two vertices has a vertex that is not an articulation vertex.

Exercise 5.3.

Let n⩾2n \geqslant 2, and let d1⩽d2⩽⋯⩽dnd_1 \leqslant d_2 \leqslant \cdots \leqslant d_n be integers. Show that there is a tree with vertex degrees d1,…,dnd_1, \ldots, d_n if and only if d1⩾1d_1 \geqslant 1 and ∑idi=2n−2\sum_i d_i = 2n - 2.

Exercise 5.4.

Let T1,…,TkT_1, \ldots, T_k be subtrees of a tree TT, meaning subgraphs that are trees, any two of which have at least one vertex in common. Prove that some vertex lies in every TiT_i.

Exercise 5.5.

The radius of a connected graph is min⁡vmax⁡wd(v,w)\min_{v} \max_{w} d(v, w), the least over all vertices vv of the greatest distance from vv to another vertex. Let CnC_n be the cycle on nn vertices, with vertices 1,…,n1, \ldots, n and edges {i,i+1}\{i, i+1\} and {n,1}\{n, 1\}. Find the radius and the diameter of K6K_6, of C9C_9 and of K5,7K_{5,7}.

Exercise 5.6.

A directed graph GG has 10001000 vertices. Its underlying undirected graph has one connected component of size 700700, and GG has one strongly connected component of size 500500; its other strongly connected components are smaller.

  1. Suppose 650650 vertices are reachable from a vertex uu. Is uu in the component of size 700700? Explain.
  2. Is uu in the strongly connected component of size 500500? Explain.
  3. Suppose 600600 vertices are reachable from a vertex vv, and 600600 vertices are reachable from vv once every arc is reversed. Is vv in the strongly connected component of size 500500? Explain.
  4. Prove that for every vertex ww of the component of size 700700, there is a directed walk from vv to ww or one from ww to vv.

Exercise 5.7.

The divisibility graph D(n)D(n) has vertices 1,2,…,n1, 2, \ldots, n and an edge between i≠ji \neq j whenever one of them divides the other.

  1. Draw D(12)D(12). How many connected components does it have, and what is the size of its largest clique?
  2. Prove that for every d⩾2d \geqslant 2, a graph with no loops in which every vertex has degree at least dd contains a simple cycle through at least d+1d + 1 vertices.

Exercises on Searching

Exercise 5.8.

A narrow island runs north–south for nn kilometres, and a searcher must locate a friend to the nearest kilometre. A tracking device tells the searcher whether the friend is north or south of the current position, but not how far, and a teleporter jumps to any given kilometre in constant time. If the friend is kk kilometres from the nearer end of the island, at kilometre kk or n−kn - k, describe an algorithm that finds the friend after visiting O(log⁡k)O(\log k) locations.

Exercise 5.9.

A ridge is given as an array of nn distinct altitudes. A point is a good collection point if it is lower than all its neighbours in the array. Design an algorithm running in O(log⁡n)O(\log n) time that finds a good collection point, prove it correct, and prove its running time.

Exercises on Sorting

Exercise 5.10.

Implementations of insertion sort and merge sort run on the same machine. On inputs of size nn, insertion sort takes 8n28n^2 steps and merge sort takes 64 nlog⁡2n64\, n \log_2 n steps. For which values of nn does insertion sort beat merge sort?

Exercise 5.11.

Person ii, for i∈{1,…,n}i \in \{1, \ldots, n\}, enters a room at time aia_i and leaves at time bi>aib_i > a_i, all the ai,bia_i, b_i being distinct. The lights are off at the start of the day; the first person to enter switches them on, and a person who leaves an empty room switches them off. Given (a1,b1),…,(an,bn)(a_1, b_1), \ldots, (a_n, b_n), we want the number of times the lights are switched on. Design, and prove correct and costed,

  1. a Θ(n2)\Theta(n^2) algorithm, and
  2. an O(nlog⁡n)O(n \log n) algorithm.

Exercise 5.12.

For each scenario choose selection sort, insertion sort or merge sort, and justify the choice by asymptotic running time.

  1. A data structure DD maintains an extrinsic order on nn items, with D.get_at(i) in worst-case Θ(1)\Theta(1) time and D.set_at(i, x) in worst-case Θ(nlog⁡n)\Theta(n \log n) time. Sort the items of DD in place.
  2. A static array holds references to nn comparable objects, any two of which take Θ(log⁡n)\Theta(\log n) time to compare. Sort the references so that the objects appear in nondecreasing order.
  3. A sorted array of nn integers, each fitting in a machine word, has had log⁡log⁡n\log \log n exchanges made between pairs of adjacent items. Re-sort it.

Exercise 5.13.

  1. Describe the principle of merge sort, and show the steps it takes to sort the array (9,3,6,2,4,1,5)(9, 3, 6, 2, 4, 1, 5).
  2. Insertion sort can be seen as a merge sort in which each step splits an array of size nn into one of size 11, the element to be inserted, and one of size n−1n - 1. By solving the appropriate recurrence, show that this recursive insertion sort takes O(n2)O(n^2) time, assuming that merging two arrays takes O(n)O(n) time.
  3. Show that merge sort on a linked list takes O(nlog⁡n)O(n \log n) time, and that it can be done with O(1)O(1) auxiliary space apart from the recursion, showing how. A programmer who can merge arrays only with O(n)O(n) extra space proposes to convert arrays to linked lists before sorting them, to save space. Comment on this strategy.

Exercise 5.14.

  1. Suppose that quicksort always partitions into two parts of relative sizes α\alpha and 1−α1 - \alpha, for a constant 0<α<120 < \alpha < \tfrac12. Ignoring rounding, find the least depth of a leaf in the recursion tree as a function of nn and α\alpha.
  2. How long does the quicksort of these notes take if all the keys are equal? Explain.
  3. What are the advantages and disadvantages of choosing the pivot at random? How does it affect the worst-case and the average-case running time?

Exercise 5.15.

How would you construct an input that makes randomized quicksort take quadratic time, without access to the state of the random number generator?

Exercise 5.16.

Merge sort can be implemented as the usual two-way merge sort, or as a three-way merge sort that splits its input into three and sorts each part recursively.

  1. Find the worst-case number of comparisons needed to merge two sorted arrays of length n/2n/2.
  2. Find the worst-case number of comparisons needed to merge three sorted arrays of length n/3n/3, both by merging all three at once and by merging in pairs.
  3. Using these, and solving suitable recurrences, find the total number of comparisons made by two-way and by three-way merge sort.
  4. If comparisons dominate the cost, which would you expect to be faster on an arbitrary array?

Exercise 5.17.

Describe an algorithm, running in time strictly better than O(n2)O(n^2), that takes a positive integer ss and a set AA of nn positive integers and decides whether two distinct elements of AA add up to exactly ss. Give its running time.

Exercise 5.18.

Each of nn computers has run the same computation, and we want to know whether strictly more than n/2n/2 of them arrived at the same result. The only available query takes two computers and reports whether they produced the same result. Design an algorithm that decides this with O(nlog⁡n)O(n \log n) queries, prove it correct, and prove the bound on the number of queries.

Exercise 5.19.

Find an asymptotically tight upper bound for the recurrence T(1)=1T(1) = 1, T(n)=T(n−1)+log⁡2nT(n) = T(n-1) + \log_2 n, and explain your answer.

Check Yourself

 

Fresh questions on the whole lesson — none of them is worked out above. Work each one out on paper before opening Python; the box only tells you whether you got there.

Answers are checked in your browser, as often as you like. Nothing is sent anywhere and nothing is kept but your own work. A formula may be written with the symbols themselves or with ~ & | -> <-> ^, and \and, \or, \to expand as you type.

Exercise 5.20.

A forest has 2020 vertices and 44 components. How many edges does it have?

answer one of these

Exercise 5.21.

What is the largest number of leaves a binary tree of height 55 can have?

answer one of these

Exercise 5.22.

What is the least height of a binary tree with 2020 vertices?

answer one of these

Exercise 5.23.

At most how many passes does binary search make on a sorted array of length 100100?

answer one of these

Exercise 5.24.

How many comparisons does linear search make on an array of 1212 entries that does not contain xx?

answer one of these

Exercise 5.25.

How many inversions does the array (4,1,3,2)(4, 1, 3, 2) have?

answer one of these

Exercise 5.26.

How many comparisons does selection sort make on an array of 88 records?

answer one of these

Exercise 5.27.

What is the worst-case number of key comparisons made by merge sort on 1616 records?

answer one of these

Exercise 5.28.

What is the expected number E4E_4 of comparisons made by randomized quicksort on four distinct keys?

answer one of these

Exercise 5.29.

What is the smallest integer hh with 2h⩾4!2^h \geqslant 4!?

answer one of these

Exercise 5.30.

In how many relative orders can 55 distinct keys arrive?

answer one of these

Exercise 5.31.

Which of the three quadratic sorts of these notes is not stable as implemented?

answer one of these

Exercise 5.32.

At most how many comparisons does quicksort make on 66 distinct keys?

answer one of these

Exercise 5.33.

At most how many comparisons does merging sorted arrays of lengths 33 and 55 take?

answer one of these

Exercise 5.34.

How many comparisons does bubble sort with the early stop make on an already sorted array of 1010 records?

answer one of these

Exercise 5.35.

How many shifts does insertion sort make on (5,4,3,2,1)(5, 4, 3, 2, 1)?

answer one of these